Addresses the medium-severity finding in @abhigyanpatwari's review of #1175:
the four `pair`-with-arrow patterns in `query.ts` anchored
`@declaration.function` on the outer `pair` node instead of the inner
`arrow_function` / `function_expression`. For multi-action object literals
like Zustand's
persist((set) => ({
addItem: (item) => doA(item),
removeItem: (item) => doB(item),
fetchData: () => doC(),
}))
`pass2AttachDeclarations.atPosition(pair.startLine, pair.startCol)`
resolved to the *parent* `(set) => ({...})` callback's scope (because the
pair node starts at the property-key token, before the inner arrow's
`@scope.function` range). All three pair-function defs landed in the
same parent's `ownedDefs`, and `resolveCallerGraphId.ownedDefs.find(...)`
returned the FIRST one — `addItem` — for every walk-up. Calls inside
`removeItem` and `fetchData` mis-attributed to `addItem`; those two
functions had zero outgoing CALLS edges in the registry-primary path.
Single-pair fixtures (`bump` in `store.ts`, `queryFn` in `query-hook.ts`)
masked the defect because there is no ambiguity when only one
Function-like def lives in the parent's `ownedDefs` — `find()` is
deterministic over a single-element set.
Fix: move the `@declaration.function` anchor from the outer `pair` to
the inner `arrow_function` / `function_expression`, mirroring the
`lexical_declaration` patterns above (`const fn = () => {}`). The def
then lands in the arrow's own scope's `ownedDefs`, the
`rangesEqual(anchor.range, innermost.range)` auto-hoist promotes the
binding to the parent scope (so importers + lookups still find the name
in the surrounding scope), and each pair-arrow becomes an independent
caller anchor in the walk.
Tests:
* Updated `useFeature → fetchData` expectation to `queryFn → fetchData`
in `typescript-hof-callbacks.test.ts`. The new attribution is
structurally correct: `fetchData()` is called from inside the named
pair-arrow `queryFn: () => fetchData()`. The pre-fix expectation
only worked because the pair-pattern bug rerouted the walk past the
syntactic owner.
* Added `multi-action-store.ts` fixture with three pair-arrows
(`addItem` / `removeItem` / `fetchData`) plus three top-level call
targets (`doA` / `doB` / `doC`). Four new tests pin per-action
attribution: positive (each action calls its own target), negative
(no sibling leakage), exact-set (the full pair set is what we
expect), and the regression fingerprint (`addItem → doB` MUST be
empty).
Validation:
* `REGISTRY_PRIMARY_TYPESCRIPT=1 vitest run` on
typescript-hof-callbacks (12 tests, +4 new), typescript-jsx-as-call
(7), typescript (236), typescript-finalize, typescript-cross-file-imports,
call-attribution-issue-1166 (18), all scope-resolution unit suites:
886/886 pass on registry-primary AND legacy DAG paths.
* Legacy DAG attribution was already correct via @abhigyanpatwari's
`tsExtractFunctionName` pair-parent handling (#1179, merged into this
PR earlier); this fix brings the registry-primary path to the same
behavior, restoring parity for multi-action objects.
* `npx prettier --check .`, `tsc --noEmit`, and `eslint` clean on the
three modified/added files.
Made-with: Cursor
Single line-length fix in `gitnexus/src/core/ingestion/languages/typescript.ts`
flagged by `quality / format` CI on commit ef96603f. The unformatted block came
from the merge of upstream PR #1179 (`fix/issue-1166-calls-edges`) where the
`pair`-with-arrow / `pair`-with-string-key handling was added; prettier wanted
the `.find` callback inlined onto a single line.
No behavior change. Pre-commit hook would have caught this locally if the
husky postinstall step had been able to write `.git/config` on this dev machine.
Made-with: Cursor
Introduced new scripts in package.json for GitNexus analysis:
- `gitnexus:refresh`: analyzes with embeddings and skills.
- `gitnexus:full`: forces analysis with embeddings and skills.
No production behavior changes. This enhances the development workflow for GitNexus users.
Addresses the automated review findings on PR #1175:
- prettier --write the 3 files flagged by `quality / format` CI check
(query.ts, typescript-hof-callbacks.test.ts, typescript-jsx-as-call.test.ts).
- [medium] typescript-jsx-as-call.test.ts: tighten the combined HOF+JSX
assertion from `toBeGreaterThan(0)` to `toHaveLength(1)`. A single
`<Foo />` is one logical invocation; the bounds-only assertion would
have masked a duplicate-CALLS-edge regression (e.g. if both
`jsx_self_closing_element` and a generic call pattern matched the
same site).
- [medium] typescript-hof-callbacks.test.ts: replace the vacuously-true
`for (c of calls) expect(...)` Zustand assertion with a structural
one. Old form passed unconditionally when `calls` was empty (any
change that silenced ALL CALLS edges from store.ts would have
slipped through). New form asserts both: (a) at least one File-rooted
edge exists (proving the `isCallerAnchorLabel` fallback fires), and
(b) no edge sources from anything else (proving the fallback fires
exclusively).
- [low] finalize-algorithm.ts (`findExportByName`): rephrase the
comment to make the language-agnostic nature of the tie-break rule
explicit. The implementation was already correct for all migrated
languages; only the comment overplayed the TypeScript specificity.
- [low] captures.ts (arity synthesis): add a comment explaining why
JSX call anchors (`jsx_self_closing_element` / `jsx_opening_element`)
intentionally don't synthesize `@reference.arity`. Name-only
resolution is correct for React (components aren't overloaded in the
current graph model); a JSX-aware synthesizer counting jsx_attribute
children would be needed if that ever changes.
No production behavior change. All 8/8 HOF + 7/7 JSX + 236/236
typescript + 11/11 api-deep-flow integration tests still pass.
gitnexus and gitnexus-shared typechecks clean.
Made-with: Cursor
Two roots in `findEnclosingFunctionId` (parse-worker) and the parallel
`findEnclosingFunction` (call-processor):
A. `genericFuncName` scanned `arrow_function` / `function_expression`
children for the first identifier and returned it. For unparenthesized
arrows like `file => processFile(file)` the first identifier is the
parameter `file`, so calls inside got attributed to a phantom
`Function file` ID and emitted dangling CALLS edges that never showed
up in `(:Function)-[:CALLS]->()` queries.
B. `tsExtractFunctionName` only named arrows whose parent was
`variable_declarator`. Object-property arrows like
`addItem: (item) => set(...)` (Zustand stores, TanStack queryFn,
React Context providers, config objects) live under a `pair`, so they
were treated as anonymous. With no named ancestor up to the file,
every call inside fell back to the File and became invisible to
`context()` / `impact()`.
Fix:
- `genericFuncName` returns null for anonymous JS/TS function-likes —
the language hook is authoritative.
- `tsExtractFunctionName` resolves names from `pair` parents
(property_identifier / string keys; computed keys stay anonymous).
- Mirror the new shape in `TYPESCRIPT_QUERIES` / `JAVASCRIPT_QUERIES` /
the scope-resolution query so pair-with-arrow becomes a Function
declaration node — call sourceIds resolve to a real graph node.
Adds 18 unit tests pinning attribution and definition behaviour for
plain helpers, `arr.map(x => fn(x))`, Promise constructor callbacks,
Zustand-style nested HOFs, TanStack query factories, string-keyed
pairs, and computed-key anonymity.
Fixes#1166
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two distinct gaps in the TypeScript scope-resolution path were silently
dropping call edges in real-world React + TanStack + Zustand codebases.
On the bug reporter's repo (Sourcerer-fe, 1185 src/ functions), 504
missing Function->Function CALLS edges are now captured (+61.6%) and
the no-outgoing-CALLS orphan rate drops from 73.2% to 60.3%.
HOF / arrow-callback caller-attribution (3 cooperating fixes):
- typescript/query.ts: @declaration.function anchor moved from the
wrapping lexical_declaration to the inner arrow_function /
function_expression, so anchor.range aligns with @scope.function and
pass2AttachDeclarations lands the def on the arrow's own scope.
- finalize-algorithm.ts: findExportByName prefers callable / class-
like defs over Variable when localDefs contains both for the same
name (TS emits two defs per `const fn = () => {}`).
- graph-bridge/ids.ts: resolveCallerGraphId's walk-up class-fallback
now uses isCallerAnchorLabel restricted to Function / Method /
Constructor / Class / Interface / Struct / Enum, so module-level
calls fall through to the File node instead of mis-attributing to
sibling Variable defs (the Zustand `create()(devtools(...))`
phantom-self-loop regression).
JSX as a CALLS edge (2 cooperating fixes):
- typescript/query.ts: new TSX_JSX_QUERY_SUFFIX (TSX-grammar only)
captures jsx_self_closing_element / jsx_opening_element as
@reference.call.free / @reference.call.member. PascalCase predicate
filters native HTML elements (<div>, <span>) so they don't emit
edges to nonexistent targets.
- typescript/captures.ts: shouldEmitReadMember extended with
jsx_self_closing_element / jsx_opening_element parent cases to
suppress phantom ACCESSES edges on member-form JSX names.
Tests: 8 HOF assertions + 7 JSX assertions across two new integration
test files plus 13 minimal fixtures. typescript.test.ts (236),
api-deep-flow.test.ts (11), and scope-resolution / scope-extractor unit
tests (613) pass with no regressions.
Made-with: Cursor
The "Web UI (browser-based)" section described an old client-side
architecture. Today gitnexus.vercel.app is a thin frontend that
auto-connects to a local `gitnexus serve` backend — there is no
ZIP drag-and-drop and no fully self-contained mode.
- Drop "No server, no install" claim
- Replace "drag & drop a ZIP" tagline with the actual onboarding step
- Add the missing `gitnexus serve` step to the local-dev block
Closes#1110
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: add platform-aware semantic fallback
Make VECTOR an optional capability so Windows analysis remains stable while semantic embeddings can fall back to exact scan when native vector indexing is unavailable.
Made-with: Cursor
* fix: remove stale vector pool import
Keep the merge with main lint-clean after VECTOR loading moved out of the read pool.
* fix(swift): use official prebuilt parser runtime
Vendor the official tree-sitter-swift 0.7.1 runtime package so Swift parsing works without source-building, while keeping the repo on the current tree-sitter runtime until the broader upgrade is ready. Also preserves Swift resolver correctness for overloaded owned functions and extension-backed type duplicates now that Swift is available by default.
Made-with: Cursor
* fix(swift): move duplicate type ordering into provider
Keep Swift extension candidate ordering behind the LanguageProvider contract and cover the Swift 0.7 init scanner path so parser runtime changes do not leak language-specific logic into shared resolution.
Made-with: Cursor
* fix(swift): address parser runtime review
Add explicit Swift prebuild checks and vendor guidance so parser runtime packaging remains observable and maintainable.
* fix(hooks): ignore global registry during staleness checks
* test(hooks): cover indexed repos under global registry
---------
Co-authored-by: laplace young <yangqk12@whu.edu.cn>
* fix(group): add configurable cross-link path exclusions to reduce false positives
Add matching.exclude_links_paths and matching.exclude_links_param_only_paths
to group.yaml config. These filter out noisy HTTP contracts (health checks,
param-only catch-all routes) from cross-link matching while preserving them
in the contract registry for documentation purposes.
Defaults are empty/false for backward compatibility — no behavior change
unless the operator explicitly configures exclusions.
* fix(group): address review findings — filter unmatched, normalize trailing slash, add tests
- Excluded contracts no longer inflate SyncResult.unmatched (isNoisy guard)
- pathPart in buildNoisyContractFilter strips trailing slashes before comparison
- 8 new unit tests for buildNoisyContractFilter covering all code paths
- Config-parser test asserts defaults for new matching fields
* fix(group): normalize configured exclusion paths and add root-path test
- Strip trailing slashes from configured exclude_links_paths at Set-build
time so root path '/' (which normalizes to '') matches correctly
- Add test: exclude_links_paths: ['/'] suppresses http::GET::/ contracts
- Add new matching fields as commented examples in fixture group.yaml (DoD §2.4)
* docs(group): document exclude_links_paths and exclude_links_param_only_paths config fields
Add JSDoc to MatchingConfig interface, update the microservices guide
YAML example and field notes, and scaffold the new fields (commented out)
in the group create template.
* fix(lbug): bound DuckDB extension install via ExtensionManager (closes#1128)
`gitnexus analyze` could hang indefinitely (60% / 85% on Windows) when
DuckDB's `INSTALL fts` or `INSTALL VECTOR` was unable to reach
`extensions.duckdb.org`. The DuckDB driver's INSTALL is a synchronous
network call, so any blocked egress would block the Node event loop
forever.
Replace the ad-hoc, in-process INSTALL/LOAD scattered across
`lbug-adapter.ts` and `pool-adapter.ts` with a single
`ExtensionManager` that owns the lifecycle of optional DuckDB
extensions:
* `LOAD` is always tried first — per-connection, idempotent, no network.
* If `LOAD` fails and policy permits, INSTALL runs in a short-lived
child Node process bounded by `GITNEXUS_LBUG_EXTENSION_INSTALL_TIMEOUT_MS`
(default 15s). The parent loop keeps spinning; on timeout the child is
killed with SIGKILL and the capability is flagged unavailable.
* Capabilities and install attempts are cached per process, so a single
bounded install per extension covers every subsequent call.
Install policy is now an explicit, per-context decision:
* `auto` (default for analyze) — try LOAD, fall back to bounded INSTALL.
* `load-only` — used by `pool-adapter` (serve / MCP read paths) so user
queries never block on a network install.
* `never` — operator escape hatch for offline / airgapped environments.
`createFTSIndex` and `createVectorIndex` now check the boolean return
value before issuing the index DDL, so missing extensions degrade BM25
and semantic search gracefully without ever throwing during analyze.
Tests:
- New unit suite for `ExtensionManager` covering LOAD-first behavior,
all three policies, install caching, observability, and warn dedup.
- Existing vector-extension integration tests pass against the new
boolean return type.
- Existing embedding-pipeline mocks updated to return `true`.
Docs: `gitnexus/README.md` documents `GITNEXUS_LBUG_EXTENSION_INSTALL`
and `GITNEXUS_LBUG_EXTENSION_INSTALL_TIMEOUT_MS` with examples for
offline and slow-network environments.
Made-with: Cursor
* fix(lbug): move DuckDB extension install child into script
Keep the bounded out-of-process INSTALL behavior, but replace the inline child code with a stable packaged ESM script. This makes the child process directly runnable and gives debuggable stack traces without source-vs-dist branching or a runtime transpiler.
Made-with: Cursor
Avoid remote git/SSH downloads for the Dart grammar during Docker and npm installs by resolving tree-sitter-dart from vendored source and building it during postinstall.
Made-with: Cursor
installOpenCodeSkills() was writing to ~/.config/opencode/skill/gitnexus/
but OpenCode only discovers skills from ~/.config/opencode/skills/*/SKILL.md.
Skills installed by `gitnexus setup` were silently ignored by OpenCode.
- Line 590: path.join(opencodeDir, 'skill') → 'skills'
- Line 587: updated JSDoc comment to match
* fix(serve): serve web UI at root path instead of 404
gitnexus serve returned Cannot GET / because no route handler existed
for the root path. Now serves the built gitnexus-web dist at / with
SPA fallback for client-side routing. Falls back to a helpful landing
page with API links when the web UI hasn't been built yet.
Also updates the build script to build and copy gitnexus-web into
gitnexus/web/ for the published npm package.
* fix(serve): address Copilot review feedback
- Use regex SPA fallback that excludes /api paths (avoids serving
index.html for unknown API routes)
- Add rel="noopener noreferrer" to external link (reverse-tabnabbing)
- Move build "done" log after web UI step
* fix(build): use npm run build for web UI, add npm install guard
The build script ran `npx tsc -b && npx vite build` in gitnexus-web/,
but CI only installs node_modules for gitnexus/ — not gitnexus-web/.
npx then resolved the wrong `tsc` package (a trojan on npm), causing
all CI jobs to fail.
Fix: add an npm install guard when node_modules is missing, and use
`npm run build` (which runs the local typescript) instead of npx.
* feat(serve): styled fallback page, asset 404s, build script safety
- Add landingPageHtml() with gitnexus-web design tokens (void bg,
surface cards, accent color, terminal-style build command block).
- Add resolveWebDistDir() helper with non-ENOENT error logging.
- Register express.static with Cache-Control headers (no-cache HTML,
immutable assets) and SPA fallback route.
- Replace wildcard SPA fallback with regex that excludes /api/* AND
asset-like file extensions (.js, .css, .ico, .woff2, .map, etc.).
- Add ordering comment warning about SPA fallback route placement.
scripts/build.js:
- Change npm install to npm ci.
- Add timeout: 120_000 to all execSync calls.
Test coverage:
- 26 new unit tests for design tokens, terminal block, external links,
SPA regex acceptance/exclusion, cache headers, and fs.access edge
cases.
Closes#1048 (review feedback)
* fix: format, lint, and add GITNEXUS_WEB_DIST env var
- Remove unused fsType import from web-ui-serving.test.ts (lint error)
- Run prettier on fallback-page-screenshot.html and test file
- Add GITNEXUS_WEB_DIST env var as primary override in resolveWebDistDir
- Add tests for env var: prefer when set, fallback when dir missing
* fix: use cross-platform path matching in env var tests
Path.includes('/env/dist') fails on Windows where path.join
produces backslashed paths. Normalize via path.sep replacement
before matching.
* fix(serve): address PR #1048 review findings
- Add uncaughtException/unhandledRejection crash guards to HTTP serve path
- Export SPA_FALLBACK_REGEX so tests use the production constant (no drift)
- Export staticCacheControlSetHeaders so tests verify the real production function
- Add real Express dispatch tests for API 404 and asset 404 isolation
- Delete committed debug artifact fallback-page-screenshot.html
* fix(scope-resolution): allow same-range Module-as-parent for top-level scopes (closes#1086)
When a C# file consists of a single top-level `namespace_declaration` that
ends exactly at EOF (no trailing newline, no leading content outside the
namespace's `{}` body), tree-sitter-c-sharp 0.23.1 reports identical byte
ranges for `compilation_unit` and `namespace_declaration`. Pre-fix the
scope-extractor parent-finder relied on strict containment, so the Module
was popped off the stack and the Namespace ended up with `parent === null`
→ `ScopeTreeInvariantError: non-module-requires-parent` →
`extractParsedFile` swallowed the throw and the whole file was dropped
from the registry-primary path. Cross-file IMPORTS / CALLS edges
originating in or terminating at that file vanished.
Hit on three real-world `*.Designer.cs` files in PersistentWindows
(`HotKeyWindow.Designer.cs`, `LaunchProcess.Designer.cs`,
`DbKeySelect.Designer.cs`) — all have the byte signature
`<BOM><CRLF>namespace ... { ... }<EOF>` (last hex = `... 7D 0D 0A 7D`).
The fix is a single carve-out in the parent-validity contract: a `Module`
may parent a same-range non-`Module` child. The relationship stays
acyclic because the carve-out is direction-asymmetric — only Module-as-
outer parents a same-range non-Module, never the reverse.
Two coordinated changes:
* `gitnexus/src/core/ingestion/scope-extractor.ts` — `pass1BuildScopes`
now consults a new `canParentScope` helper instead of
`rangeStrictlyContains` directly. Sort tie-breaker added so a same-
range Module always sorts before a non-Module candidate, ensuring the
Module lands on the parent-stack first regardless of tree-sitter
capture iteration order.
* `gitnexus-shared/src/scope-resolution/scope-tree.ts` — `buildScopeTree`'s
`parent-must-contain-child` check now uses the same `canParentScope`
carve-out so the validator agrees with the extractor on what a
well-formed parent edge looks like. Error message updated to spell
out the new contract.
`rangeStrictlyContains` keeps its strict semantics in both files —
position-index lookups, hook-side range comparisons, and other call
sites are unchanged.
* `gitnexus/test/fixtures/lang-resolution/csharp-namespace-as-root-no-trailing-newline/`
— minimal regression fixture mirroring the PersistentWindows shape:
both `Models/User.cs` and `App/Program.cs` end exactly on the closing
`}` of their namespace with no trailing newline. The trigger is shape-
driven, not size-driven, so the fixture stays small (~250 bytes total).
* New `csharp.test.ts` describe block: scope extraction completes for
both files, and the cross-file `IMPORTS` edge resolves through the
scope-resolution path with `reason: 'csharp-scope: using'`.
* `scope-tree.test.ts`: replaced the prior "rejects child ranges
identical to the parent" case with three new ones — non-Module parent
still rejected at equal range; Module-as-parent of a same-range non-
Module accepted (the #1086 carve-out); Module-as-parent of another
Module still rejected (the asymmetry guard).
* `npx vitest run test/unit/scope-resolution test/integration/resolvers`
→ 2514 passed / 77 skipped / 0 failed (52 test files).
* `npx tsc --noEmit` clean in both `gitnexus/` and `gitnexus-shared/`.
* End-to-end on PersistentWindows (after rebuilding the Docker image
with this branch): 3 prior `scope extraction failed for *.Designer.cs`
warnings → 0. Pre-fix index numbers will be re-checked here once the
branch is built and indexed; the existing post-#1082 baseline is
1113 nodes / 2987 edges / 39 clusters / 97 flows.
`canParentScope` is language-agnostic. Other languages whose query emits
`(compilation_unit) @scope.module` plus a single same-range top-level
scope can naturally hit the same byte shape on minimal files; this fix
applies to all of them uniformly.
Refs: #1086 (issue with full root-cause analysis + 4-case empirical
repro through `extractParsedFile`).
* refactor(scope-resolution): export canParentScope from gitnexus-shared
Addresses #1087 review (medium): the helper was previously duplicated
byte-for-byte in `scope-extractor.ts` and `scope-tree.ts`. Per DoD
"single source of truth in shared", the contract piece belongs in
gitnexus-shared (Ring 2 SHARED #912) and the consuming layer should
import it. Eliminates the silent-drift surface where a future edit
to one copy would produce extractor/validator disagreement on what
a well-formed parent edge looks like.
Changes:
- gitnexus-shared/src/scope-resolution/scope-tree.ts: add `export`
to `canParentScope`.
- gitnexus-shared/src/index.ts: re-export `canParentScope`.
- gitnexus/src/core/ingestion/scope-extractor.ts: remove the local
`canParentScope` definition (and its now-unused local copy of
`rangeStrictlyContains`), import from `gitnexus-shared`. The local
`rangesEqual` stays — it's still used in capture-anchor logic at
two unrelated sites.
Validation (per DoD §4.4 — both CLI and web consumers verified):
- npx tsc --noEmit clean in gitnexus/ and gitnexus-shared/
- cd gitnexus-web && npx tsc -b --noEmit clean
- gitnexus-shared `npm run build` clean
- Targeted: vitest run test/unit/scope-resolution test/integration/resolvers
→ 2522 passed / 0 failed / 77 skipped (54 files)
- Full suite: vitest run → 7238 passed / 1 failed / 97 skipped.
The single failure is `test/unit/ignore-service.test.ts > warns
on EACCES but does not throw`, which cannot run when uid=0 (root
bypasses POSIX permission checks). Pre-existing on this branch
before the refactor; unrelated to scope-resolution.
The `loadIgnoreRules — error handling > warns on EACCES but does not
throw` test relies on `chmod 000` denying read access to a temporary
.gitignore file. On Linux, root bypasses POSIX read-permission checks,
so chmod 000 does NOT trigger EACCES under uid=0 — fs.readFile reads
the file anyway and loadIgnoreRules returns parsed rules instead of
the `null` the test expects.
Symptom under root: assertion fails with `Ignore { _rules: [...] }
to be null`, surfaced as a single test failure in any privileged
test environment (rootful Docker container, CI runners configured to
run tests as root, etc.).
Fix: extend the existing `skipIf(process.platform === 'win32')` guard
with `process.getuid?.() === 0`. The non-root code path still
exercises the real EACCES branch — root just can't reproduce the
failure mode the test asserts on, so skipping there is the correct
posture (matches the win32 skip's reasoning: the OS-level mechanism
the test depends on isn't available there).
Optional chaining (`getuid?.()`) keeps Windows compatibility — Node
on Windows doesn't expose `process.getuid` at all.
The csharp-large-cache-miss-resolution fixture added in #1082 reproduces
the freeze contract failure via tree-sitter cache-miss reparse on >32 KB
files. This adds a complementary trigger for the same root cause that
does not depend on file size: a small-file pair where the importer
locally declares a class with the same simple name as a sibling reached
through `using`.
Pre-#1082 path: scope-extractor pre-populates (and freezes) `User` in
the importer's Module bindings, then populateCsharpNamespaceSiblings'
namespace-import loop calls push() on the frozen array and throws
"Cannot add property N, object is not extensible", aborting the whole
scopeResolution phase.
Post-#1082 the augmentation channel keeps both bindings visible; the
local `Collision.App.User` shadows the namespace-imported one per
origin precedence, so `Program.Run -> new User()` resolves to the
local class.
Three assertions:
- scopeResolution completes (no throw on the colliding bucket).
- both `User` declarations are detected across the two namespaces.
- `Program.Run -> User` constructor edge points at App/Program.cs
(not Models/User.cs), verifying origin:local shadows origin:namespace.
Verified: full csharp.test.ts suite green (207/207). tsc --noEmit clean.
Refs: #1066, #1082, #1083 (closed as superseded).
* fix(csharp): adaptive tree-sitter buffer + frozen-bucket clone for cross-namespace siblings (#1066)
Two coupled regressions surfaced when analyzing real-world C# repos with
large source files (issue #1066):
1. Tree-sitter `parser.parse()` is hard-coded to a 32 KB buffer by
default. Any file exceeding that threshold throws `Invalid argument`
on the worker re-parse path of `populateCsharpNamespaceSiblings`
(and the analogous Python / TypeScript captures fallbacks).
2. After the buffer fix unblocks the AST walk, the hook tries to
`push()` onto the inner `BindingRef[]` array fetched from
`indexes.bindings` — but `materializeBindings` froze that array via
`Object.freeze(refs.slice())`. Result: `Cannot add property N,
object is not extensible`.
Fixes:
- `csharp/captures.ts`, `python/captures.ts`, `typescript/captures.ts`:
pass `bufferSize: getTreeSitterBufferSize(sourceText.length)` to
`parser.parse()` on the cache-miss path so multi-MB files parse.
- `csharp/namespace-siblings.ts`: introduce `cloneBindingBucket` to
copy the frozen array before mutating, then `set()` the new array
back. This is a working but architecturally compromised workaround
(#1050 follow-up will replace it with an explicit augmentation
channel — see docs/plans/2026-04-26-001 plan).
Tests:
- New `csharp-large-cache-miss-resolution` fixture (Models/Services/
Other layout, ~77 KB padded UserService.cs) drives the buffer-size
failure end-to-end through worker mode.
- `csharp.test.ts`: 4 new regression assertions covering both the
parse-time buffer-size failure and the freeze workaround.
- Per-language captures unit tests gain "large cache-miss file uses
adaptive buffer" coverage (TS, Python, C#).
- `csharp-hooks.test.ts`: in-memory freeze regression test that
reproduces the `Cannot add property` crash without invoking the C#
parser at all.
Made-with: Cursor
* refactor(scope-resolution): add bindingAugmentations channel to indexes
Step 1 of the binding-augmentation-channel refactor (issue #1066
follow-up). Pure shape change — no consumers yet.
Adds a new `readonly bindingAugmentations` field to
`ScopeResolutionIndexes` initialized as an empty `Map` by
`finalizeScopeModel`. The new channel is the dedicated post-finalize
write target for hooks like `populateCsharpNamespaceSiblings`, so
`indexes.bindings` can stay frozen and finalize-owned.
Behavior unchanged: nothing reads or writes the new field yet. tsc and
the full unit suite remain green.
Plan: docs/plans/2026-04-26-001-binding-augmentation-channel.md (local
only — `docs/plans/` is gitignored).
Made-with: Cursor
* feat(scope-resolution): add lookupBindingsAt dual-source helper
Step 2 of the binding-augmentation-channel refactor. Introduces a
single primitive every walker uses to read both the finalize-owned
`indexes.bindings` channel and the post-finalize
`indexes.bindingAugmentations` channel.
Contract:
- Finalized refs come first (preserves existing precedence).
- Augmented refs append, deduped by `def.nodeId`.
- Empty input on both channels returns a shared frozen empty array.
- Single-channel hits return the bucket by reference (no allocation).
No consumers are wired yet — Step 3 routes the existing walker
primitives through this helper. Augmentations remain empty for every
language; behavior of the full suite is unchanged.
8 unit tests pin precedence, dedup, identity for single-channel hits,
and the shared-empty-frozen-array sentinel.
Made-with: Cursor
* refactor(scope-resolution): route binding lookups through lookupBindingsAt
Step 3 of the binding-augmentation-channel refactor. Every direct
`indexes.bindings.get(...)` consumer in the post-finalize phase is
now routed through `lookupBindingsAt` (per-name) or `namesAtScope`
+ `lookupBindingsAt` (bulk iteration).
Routed sites:
- `findClassBindingInScope` (walkers.ts) — class-receiver lookups.
- `findCallableBindingInScope` (walkers.ts) — free-call lookups.
- `findExportedDefByName` (walkers.ts) — module-scope-fallback
callable lookups.
- `propagateImportedReturnTypes` (passes/imported-return-types.ts)
— bulk iteration over an importer's binding entries; switched to
`namesAtScope` + per-name `lookupBindingsAt` so post-finalize
augmentations are visible to import-derived typeBinding mirrors.
Behavior unchanged: augmentations are empty across the suite (Step 4
populates them for C# `populateNamespaceSiblings`). 587
scope-resolution unit tests + 50 integration resolver suites green
(4 pre-existing Swift method-implements failures unrelated to this
work).
Adds `namesAtScope` companion helper for the bulk-iteration callers.
Made-with: Cursor
* refactor(csharp): write namespace siblings to bindingAugmentations channel
Step 4 of the binding-augmentation-channel refactor. The C#
`populateNamespaceSiblings` hook is the only consumer that needed
to inject cross-file bindings post-finalize, and prior to this
change it cloned the (frozen) finalized `BindingRef[]` arrays
through a `cloneBindingBucket` helper, then `set()`-back the new
array — a workaround for the `Object.freeze` applied by
`finalize-algorithm.ts` (issue #1066 root cause).
Architecturally that violated `ScopeResolver` Invariant I8 (which
permits post-finalize modifications but not in-place mutation of
finalized buckets). It also forced read-side consumers to be aware
of the workaround.
This change:
* Switches the three C# write sites to append into
`indexes.bindingAugmentations` via `getAugmentationBucket`. The
augmentation channel was added in Step 1 and is mutable by
contract: inner `BindingRef[]` arrays here are NEVER frozen.
* Deletes `cloneBindingBucket` and `getMutableScopeBindings`
(workaround helpers no longer needed).
* `lookupBindingsAt` (Step 2) merges the two channels transparently
for every walker (Step 3), so behavior is unchanged for callers.
* Updates the unit test to assert against both channels: finalized
bucket stays frozen and untouched, cross-file siblings show up in
augmentations only. Renamed the test accordingly.
Validation:
* `npx tsc --noEmit` clean.
* csharp hooks unit + walkers-augmentations unit + csharp integration
resolver suite all green (236/236).
* Wider `test/unit/scope-resolution test/integration/resolvers`
suite: 2507 pass, only 4 pre-existing Swift METHOD_IMPLEMENTS
failures remain (unrelated to this work, present on baseline).
Refs: issue #1066, ADR-pending binding-augmentation-channel.
Made-with: Cursor
* feat(scope-resolution): tighten I8 + add validateBindingsImmutability dev guard
Step 5 of the binding-augmentation-channel refactor. Captures the
new two-channel binding lifecycle in the contract docs and adds a
dev-mode runtime validator so a future hook cannot silently drift
back into mutating `indexes.bindings`.
Contract changes:
* `contract/scope-resolver.ts` — rewrote Invariant I8 to describe
the two channels (`indexes.bindings` is finalize-output and
immutable post-finalize; `indexes.bindingAugmentations` is the
append-only post-finalize channel populated by hooks like
`populateNamespaceSiblings`). Documented `lookupBindingsAt` as
the read-side merger and pointed at the new validator as the
enforcement mechanism.
* `gitnexus-shared/src/scope-resolution/types.ts` — extended the
module-header lifecycle contract to call out
`bindingAugmentations` alongside `ReferenceIndex` as the two
structures populated after the freeze.
Validator:
* New `pipeline/validate-bindings-immutability.ts` mirrors the
shape of `validateOwnershipParity` (#909): runs only when
`NODE_ENV !== 'production' && VALIDATE_SEMANTIC_MODEL !== '0'`,
emits via `onWarn`, never throws. Asserts (a) every inner
`BindingRef[]` in `indexes.bindings` is `Object.isFrozen`, and
(b) every inner array in `indexes.bindingAugmentations` is NOT
frozen.
* Wired into `pipeline/run.ts` after both
`populateNamespaceSiblings` and `propagateImportedReturnTypes`,
before `resolveReferenceSites`. One sweep covers the full
post-finalize surface.
Tests:
* `validate-bindings-immutability.test.ts` — 6 cases pinning happy
path, both drift directions, multi-violation accumulation, and
both production no-op gates.
All scope-resolution + csharp resolver tests green (242/242 in the
focused run; matches the wider Step 4 baseline).
Made-with: Cursor
* fix(ingestion): size tree-sitter buffers from UTF-8 bytes
Tree-sitter buffer sizing is byte-based, so computing adaptive buffers from JavaScript string length under-sized UTF-8-heavy files. Make getTreeSitterBufferSize accept source text directly and compute Buffer.byteLength internally, then update all parse call sites and max-buffer skip checks to use byte length.
Add multibyte cache-miss and cap regressions for C#, Python, TypeScript, and the C# namespace-sibling fallback parse path.
Made-with: Cursor
* test(scope-resolution): pin augmentation read paths
Add focused unit coverage for augmented-only binding reads across the routed walker helpers and imported-return-type propagation path. Clarify I8 wording around lexical Scope.bindings versus post-finalize index channels, and document the intentional local-only behavior of findExportedDef.
Also switch the immutability validator tests to Vitest env stubs, document one intentional validator blind spot, and split C# namespace-sibling tests so UTF-8 parsing and augmentation-channel behavior are asserted independently.
Made-with: Cursor
* test(scope-resolution): avoid slow parser stress fixtures
Replace high-cardinality large-file capture fixtures with large padding plus a trailing declaration. This still proves adaptive tree-sitter buffers parse beyond large ASCII and UTF-8-heavy input, without making query matching process thousands of declarations and risking timeouts.
Made-with: Cursor
* test(scope-resolution): add python and typescript cache-miss resolver regressions
Add worker-mode resolver integration coverage mirroring the C# #1066 scenario for Python and TypeScript. Each test builds a temp fixture with large ASCII and UTF-8-heavy source padding, then asserts trailing declarations and call edges still resolve after scope-resolution cache-miss reparsing.
Made-with: Cursor
* refactor(scope-resolution): gate I8 validator and fast-path namesAtScope
Addresses SPARC reviewer feedback on the binding-augmentation channel:
- Validator gate is now opt-in outside development. Extract
isSemanticModelValidatorEnabled() in utils/env.ts as the single
predicate; both validateBindingsImmutability and phase.ts's warn
handler share it. Default CLI runs no longer pay the O(binding-buckets)
scan, and explicit VALIDATE_SEMANTIC_MODEL=1 now emits warnings even
when NODE_ENV is unset.
- namesAtScope returns Iterable<string> and zero-allocates when at most
one channel is populated (returns Map.keys() directly), only
materializing a Set when both channels carry names. The caller-side
branching and EMPTY_NAMES escape hatch in propagateImportedReturnTypes
are gone -- both helpers handle the empty-augmentation case internally.
- C# namespace-siblings header/JSDoc, model JSDoc, I8 contract prose, and
the #1066 integration-test header rewritten to say post-finalize fanout
appends only to bindingAugmentations; finalized refs come first and win
duplicate def.nodeId metadata; local lexical Scope.bindings remains the
first-tier shadowing channel.
Validator unit-test setup deduplicated via beforeEach and extended with
default-CLI no-op + explicit-opt-in cases.
Made-with: Cursor
* feat(ingestion): TypeScript registry-primary scope resolution (Ring 3)
- Add TypeScript ScopeResolver stack (query/captures/interpret, import decomposition, hooks, arity, merge, receiver binding) and register in SCOPE_RESOLVERS.
- Harden shared compound receiver and receiver-bound CALLS pass for map for-of tuple bindings, dotted typeRef shapes, and callable-alias fallbacks.
- Flip TypeScript into MIGRATED_LANGUAGES; refresh AGENTS.md and type-resolution-system.md.
- Shared finalize-algorithm updates for cross-file scope parity.
- Tests: TS scope-resolution unit suite; legacy call-processor suite forces REGISTRY_PRIMARY_TYPESCRIPT=0; registry-primary flag test opts out TS in override scenario.
Made-with: Cursor
* fix(ingestion): SCC-ordered cross-file return-type propagation + multi-hop re-export resolution
Fix CI failures on PR #1050 (TypeScript registry-primary migration) by
making `propagateImportedReturnTypes` deterministic via reverse-
topological SCC ordering and updating the multi-hop re-export contract
to match `followReexportChain` behavior.
Why: the legacy pass mirrored an intermediate ref instead of the
terminal type when an importer was processed before its source module
had its own typeBindings chain-followed (4-file alias chain regression
in `ts-simple` fixture: `models.User -> service.user -> app.user`
collapsed to `getUser` instead of `User`). Reverse-topological walk of
`indexes.sccs` (leaves first) lets every importer see the source's
already-followed terminal type in a single pass.
Changes:
- `imported-return-types.ts`: rewrite to walk SCCs leaves-first, chain-
follow the source module's typeBindings BEFORE mirroring, and chain-
follow the importer's typeBindings AFTER mirroring. Cyclic SCCs
reach a partial fixpoint (no convergence guarantee, ts-circular only
asserts no-throw).
- `finalize-algorithm.ts`: docstring update on `FinalizeFile.localDefs`
to reflect that `followReexportChain` resolves multi-hop re-exports
through barrels even when intermediates do not surface the name -
surfacing is now a static optimization, not a correctness requirement.
- `contract/scope-resolver.ts` Invariant I3: explicitly document the
SCC ordering requirement.
- `pipeline/run.ts`: split PROF timer into `finalize` and `propagate`
so the pass's cost is observable independently.
- `ARCHITECTURE.md` Performance notes: describe SCC-ordered propagation.
- `imported-return-types.ts`: expand chain-depth comment (2x effective
depth from pre/post follow), add multi-ref break rationale, add
`ts-simple` motivating-fixture pointer.
Tests:
- `finalize-algorithm.test.ts`: add 4 cases (3-hop chain, cyclic
re-export visited-set guard, wildcard re-export fall-through,
multi-source first-match-wins); fix misleading shared nodeId in the
thick variant; rename and update the multi-hop test for the new
contract (transitiveVia assertion on the thin variant).
- `imported-return-types.test.ts` (NEW): unit tests for the SCC pass
pinning topological collapse, local-annotation guard, missing-source
skip, and cyclic-SCC no-throw.
- `cross-file-binding.test.ts` + `ts-deep-alias-chain` fixture (NEW):
5-file integration regression guard for SCC-ordered propagation
through 4 module boundaries.
Validation: 865 scope-resolution + cross-file tests pass on Windows;
typecheck clean across both packages; only pre-existing Swift overload
failures remain (verified on PR base commit, environmental).
Made-with: Cursor
* fix(ingestion): address PR #1050 review findings — side-effect imports, resolve-cache perf, adapter signature
Three independent fixes surfaced by the production-readiness review of
the TypeScript registry-primary scope-resolution migration (RFC #909
Ring 3). All three pass under both REGISTRY_PRIMARY_TYPESCRIPT=0 and =1.
1. Side-effect imports were silently dropped (correctness regression).
The legacy DAG emitted IMPORTS edges for `import './polyfill'` because
its tree-sitter query matches `(import_statement source: (string))`
regardless of clause. The new registry-primary path returned `[]`
from `splitImportStatement()` for clause-less imports, so no
ParsedImport / ImportEdge was ever produced — silent file-level edge
loss. Add a generic 'side-effect' variant to `ParsedImport` and
`ImportEdge['kind']` in `gitnexus-shared`; finalize resolves the
target file and pre-finalizes the edge (no `targetDefId`, no
`BindingRef`) so the SCC fixpoint loop skips it. The TypeScript
provider now emits + interprets the new kind end-to-end. The
variant is intentionally generic so other languages (Rust
`use foo as _`, Python module-init) can adopt it.
2. Per-import re-derivation in `resolveImportTarget` (perf regression).
The TS adapter built `new Set(allFilePaths)` on every call and let
`resolveTsImportTarget` re-derive `allFileList` /
`normalizedFileList` and discard the `resolveCache`. For a workspace
with N files and M imports that's O(N × M) work per pass. Wrap the
adapter in a closure that memoizes all five derived values keyed on
the orchestrator's `ReadonlySet` identity; reset only when the set
reference changes (start of new pass). New cost: O(N + M).
3. Misleading fake `ParsedImport` in the adapter (architecture).
The adapter constructed `{ kind: 'named', localName: '_',
importedName: '_', targetRaw }` to call `resolveTsImportTarget`,
even though only `targetRaw` and the structural-typed context are
read. Extract `resolveTsTarget(targetRaw, ctx)` so the adapter has
an honest signature; `resolveTsImportTarget` still works for other
callers. Also extract `narrowTsContext` for the type narrowing.
Tests: - New 4-file fixture `typescript-side-effect-imports` with two
side-effect imports + one named import.
- New "TypeScript side-effect imports" describe in
`test/integration/resolvers/typescript.test.ts` (parity-gated by
`ci-scope-parity.yml` — runs under both flag states).
- Updated 2 unit tests to expect 1 side-effect ParsedImport and 4
`@import.statement` matches (was 0 / 3).
- 785 / 785 TS scope-resolution tests pass under both
REGISTRY_PRIMARY_TYPESCRIPT=0 and =1.
Made-with: Cursor
* fix(scope): address Codex adversarial review findings on PR #1050
Four findings from the Codex adversarial review broke registry-primary
TypeScript resolution for common patterns. All four now have unit and
integration regression coverage that pass under both
`REGISTRY_PRIMARY_TYPESCRIPT=0` (legacy DAG) and the default
registry-primary path.
[high] tsconfig path aliases dropped:
Threaded `tsconfigPaths` through ScopeResolver via a new opaque
`resolutionConfig` parameter and a `loadResolutionConfig(repoPath)`
hook. The orchestrator (`scopeResolutionPhase` + `runScopeResolution`)
loads it once per workspace pass and forwards into every
`resolveImportTarget` call. TypeScript resolver now resolves
`@/services/user` style imports through the standard resolver's alias
branch.
[high] TSX parsed with the wrong grammar:
`emitTsScopeCaptures` now picks the parser/query by `filePath`
(`.tsx` -> TSX grammar) and validates cached trees against the
expected grammar via the new exported `tsCachedTreeMatchesGrammar`
helper. Stale TS-grammar trees for `.tsx` files no longer leak through
the scope query.
[medium] Literal dynamic imports never linked:
Added `kind: 'dynamic-resolved'` to `ParsedImport` and `ImportEdge`.
The decomposer emits a synthetic `@import.literal` capture for
string-literal dynamic imports; the interpreter maps that to
`dynamic-resolved`; finalize pre-finalizes it as a file-level terminal
(same shape as `side-effect`). `import('./feature')` now produces a
real IMPORTS edge under the registry-primary path. Legacy DAG keeps
its existing behavior — the new integration assertion is gated behind
the flag.
[medium] Namespace re-exports invisible from barrels:
The decomposer now emits TWO captures for `export * as ns from './m'`
— the existing `reexport-namespace` import draft AND a synthetic
`@declaration.namespace` capture (via `buildNamespaceDeclarationMatch`).
The latter creates a Namespace `SymbolDefinition` in the barrel's
`localDefs`, so downstream `import { ns } from './barrel'` resolves
through `findExportByName`.
Regression fixtures under `gitnexus/test/fixtures/lang-resolution/`:
- typescript-tsconfig-aliases (`@/` alias)
- typescript-tsx-jsx (Button.tsx + App.tsx with JSX)
- typescript-dynamic-import (`await import('./feature')`)
- typescript-reexport-namespace (`export * as Models from './base'`)
Validation:
- gitnexus-shared builds clean
- gitnexus typecheck clean
- 385/385 TS scope-resolution tests pass under both
`REGISTRY_PRIMARY_TYPESCRIPT=0` and default
Made-with: Cursor
* perf(scope): O(1) defById lookup + bounded re-export depth (PR #1050 round 3)
Addresses the round-3 PR #1050 reviews (Claude adversarial + xkonjin):
both flagged the existing O(N²) `findDefById` linear scan in
`materializeBindings` and the unbounded recursion in
`followReexportChain` as production-readiness blockers for TypeScript
monorepos. Both fixes land alongside their regression tests under
both `REGISTRY_PRIMARY_TYPESCRIPT=0` and the default registry-primary
path.
[high] materializeBindings O(N_files × N_defs × N_edges) → O(N_defs + N_edges):
Build a `nodeId → SymbolDefinition` index map once at the top of
`materializeBindings` (one O(N_defs) pass), then replace the per-edge
`findDefById(files, edge.targetDefId)` linear scan with an O(1)
`defById.get(edge.targetDefId)` lookup. Also drop the now-unused
`findDefById` helper. At realistic TypeScript monorepo scale (~5k
files × ~50 defs/file × ~100k linked import edges) this is the
difference between ~25 s and a few ms inside finalize. Regression
test in `finalize-algorithm.test.ts` builds 200 leaf files +
1 consumer importing one symbol from each, asserts every binding
materializes correctly.
[medium] followReexportChain unbounded recursion:
The existing `visited` set caps depth at `O(N_files)` but allows
recursion proportional to barrel-chain depth, mismatching the
explicit "Iterative DFS to avoid stack overflow" policy in
`tarjanSccs`. Added a `MAX_REEXPORT_DEPTH = 100` constant and a
`depth` parameter to `followReexportChain` (defaults to 0); each
recursive call passes `depth + 1` and the function returns `null`
when the cap is exceeded. 100 is comfortably above any realistic
hand-authored barrel chain (typical depth 1-5; auto-generated
barrels rarely exceed 20) while staying well below JS engine call
stack limits. Regression test wires a 200-link reexport chain and
verifies the crawl terminates cleanly with `linkStatus: 'unresolved'`
(no terminal def reachable within the budget).
[low] synthesizeInstanceofNarrowings bare-identifier-only limitation:
xkonjin's review #4 noted that the LHS narrowing only handles bare
identifiers (`if (x instanceof Foo)`), not member expressions
(`if (user.address instanceof Address)`). Added a JSDoc note
explaining the constraint and pointing readers at field-type
resolution as the workaround for member-chain receivers.
Validation:
- gitnexus-shared builds clean
- gitnexus typecheck clean
- 413/413 tests pass under both flag states for finalize-algorithm +
TS unit + TS integration suites
- 972/972 tests pass across full scope-resolution + Python +
C# integration smoke (no cross-language regression)
Made-with: Cursor
* refactor(finalize): replace recursive followReexportChain with SCC-condensed iterative closure
The legacy `followReexportChain` walked re-export drafts via mutual
recursion guarded by a per-call visited set + a `MAX_REEXPORT_DEPTH`
ceiling. Recursion is fragile (call-stack ceiling, no bound on depth
that's actually meaningful), so this replaces it with a structurally
better algorithm: a precomputed per-file re-export closure built by
running Tarjan SCC over the re-export sub-graph and propagating names
in reverse-topological order with a bounded intra-SCC fixpoint.
Algorithm (`buildReexportClosures` in finalize-algorithm.ts):
1. Sub-graph: build the directed graph of `reexport` + `wildcard`
drafts only (regular/namespace/dynamic imports do not contribute).
2. SCC condensation: run the same iterative `tarjanSccs` already
used for the file-level import graph; output is in reverse-topo
order so out-of-SCC neighbors are always already-finalized.
3. Per-SCC propagation:
- Acyclic singleton: one pass populates from neighbors' closures.
- Cyclic SCC: bounded fixpoint capped at |SCC|+1 iterations.
With first-wins precedence the closure map is monotone, so
each name needs at most |SCC| hops to traverse the cycle.
Precedence (preserved from the recursive crawl):
- Named re-exports take precedence over wildcards.
- Within each kind, declaration order wins.
Lookup at finalize time becomes O(1) (`lookupReexportedName`), down
from O(chain_depth × drafts) per consult and recursive at that.
Properties vs the legacy implementation:
- Stack-safe by construction; no `MAX_REEXPORT_DEPTH` guard needed.
- 1000-hop barrel chains now resolve in full (legacy capped at 100
and surfaced anything deeper as `unresolved`).
- Cycles handled structurally via SCC, not via per-call visited set.
- Same observable semantics: every existing test passes unchanged.
Tests:
- Replace the obsolete `MAX_REEXPORT_DEPTH (200-hop chain stops
cleanly without stack overflow)` test (which asserted the OLD
bug — that deep chains failed to resolve) with a positive
1000-hop test that asserts full resolution + accurate
`transitiveVia`. Proves both the recursion is gone AND the
closure correctly inherits the leaf def across all hops.
- Update commentary on adjacent re-export tests to reference the
closure mechanism.
- Update `FinalizeFile.localDefs` JSDoc + import-decomposer.ts
inline doc to point at `buildReexportClosures` instead of the
removed function name.
Validation: - gitnexus-shared builds cleanly.
- gitnexus typechecks cleanly.
- 28/28 finalize-algorithm.test.ts tests pass (incl. new 1000-hop).
- 801/801 TypeScript scope-resolution tests pass under default
(registry-primary) AND `REGISTRY_PRIMARY_TYPESCRIPT=0` (legacy DAG).
- 404/404 Python + C# integration tests pass — no regression in
cross-language consumers of the shared `finalize`.
Made-with: Cursor
* fix(scope): remove non-null assertions from scope resolution
Made-with: Cursor
* fix(scope): address TypeScript review follow-ups
Made-with: Cursor
* fix(scope): address TypeScript import review follow-ups
Add regression coverage for non-binding import edges and circular TypeScript bindings so PR #1050 review concerns stay visible without changing runtime semantics.
Made-with: Cursor
On Windows, HOME env is often unset, causing cache to be written to
'./undefined/'. Using os.homedir() ensures cross-platform compatibility
while preserving HF_HOME priority.
Fixes#1068
The early Validate step ran on both workflow_call and push events, but
push events never populate inputs.tag (the tag comes from github.ref).
This regressed every real tag-push release — v1.6.3's Docker Build &
Push failed at that gate. The downstream Verify step already falls back
to GITHUB_REF, so the upfront guard only needs to cover workflow_call.
* chore(deps)(deps): bump lucide-react in /gitnexus-web
Bumps [lucide-react](https://github.com/lucide-icons/lucide/tree/HEAD/packages/lucide-react) from 0.562.0 to 1.11.0.
- [Release notes](https://github.com/lucide-icons/lucide/releases)
- [Commits](https://github.com/lucide-icons/lucide/commits/1.11.0/packages/lucide-react)
---
updated-dependencies:
- dependency-name: lucide-react
dependency-version: 1.8.0
dependency-type: direct:production
update-type: version-update:semver-major
...
Signed-off-by: dependabot[bot] <support@github.com>
* chore(deps)(deps): provide local Github SVG for lucide-react v1
lucide-react 1.0 removed all brand icons (Github, Gitlab, Facebook,
Slack, etc) per https://lucide.dev/guide/react/migration. Our
centralized icon module re-exported `Github` from lucide-react,
which now fails typecheck.
Replace the re-export with a local forwardRef component that mirrors
the lucide v0 GitHub mark and the LucideProps API. All consumers keep
importing `Github` from `@/lib/lucide-icons` unchanged.
Made-with: Cursor
* refactor(web): use Primer Octicons mark for local Github icon
Swap the local lucide v0 outline mark for a verbatim copy of Primer
Octicons `mark-github-{16,24}` — the icon set GitHub itself ships on
github.com (MIT, Copyright (c) GitHub Inc.).
Why this source over the alternatives is documented at the top of
`gitnexus-web/src/lib/lucide-icons.tsx`, including:
* the lucide v1 brand-icon removal context and migration link,
* the trademark vs. license distinction (MIT covers our right to
copy the SVG; trademark rules govern *use*, and we only use the
mark in permitted ways per GitHub's brand toolkit),
* why we didn't add `@primer/octicons-react`, `react-icons`, or
`simple-icons` (zero-dep policy for one icon),
* source URLs for both SVG variants.
The component still implements `LucideProps` and is drop-in compatible
with the existing import sites in Header, RepoAnalyzer and
AnalyzeOnboarding. The mark is now filled (matching github.com) rather
than stroke-outlined; lucide-only stroke props are accepted for type
parity but ignored. Both 16 and 24 variants are shipped so the mark
stays crisp at small sizes when consumers pass an explicit `size`.
Made-with: Cursor
---------
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
Final step of the iterative vite 5 -> 8 migration. This is the
substantive hop: Rolldown replaces Rollup, Oxc replaces esbuild,
Lightning CSS replaces esbuild for CSS, and vitest jumps to v4 (vitest
3 only peers with vite ^5||^6||^7).
Dep changes (gitnexus-web/package.json):
- vite ^7.3.2 -> ^8.0.10
- vitest ^3.2.4 -> ^4.1.5
- @vitest/coverage-v8 ^3.2.4 -> ^4.1.5
- @tailwindcss/vite ^4.1.18 -> ^4.2.4 (vite ^8 peer support starts at 4.2.2)
- tailwindcss ^4.2.2 -> ^4.2.4 (match the vite plugin minor)
- @vitejs/plugin-react already at 5.2.0 from iter 2 (vite ^8 peer included)
Test fix (heartbeat.test.ts):
- vitest 4 enforces [[Construct]] on mock implementations used with `new`.
The arrow function passed to .mockImplementation() in the EventSource
stub is now rejected with "() => { ... } is not a constructor". Switched
to a regular function declaration, which restores constructor semantics
without changing test behaviour. All 7 heartbeat tests pass again.
Coverage threshold tune (vitest.config.ts):
- vitest 4 ships AST-aware coverage remapping by default, which measures
reachable code more accurately than the legacy istanbul-style mapping.
Same 220 tests now report 9.44%/4.47%/7.24%/9.58% instead of just over
10% on each axis. Lowered thresholds to 9/4/7/9 to keep them as soft
regression floors rather than coverage targets. No tests removed.
What we deliberately did NOT change:
- vite.config.ts: the five resolve.alias entries (mermaid, anthropic deep
import, gitnexus-shared, @, @shared) all keep working under Rolldown.
server.fs.allow: ['..'] is unchanged in v8. The mermaid alias is
arguably MORE important now because vite 8.0.10 explicitly removed
format-sniffing module resolution from the JS resolver.
- engines.node: vite 8 has the same Node floor as vite 7
(^20.19.0 || >=22.12.0), already set in iter 2.
- CI setup-node pin: already at 20.19.0 from iter 2.
Verified locally (Node v22.14.0):
- npm install: clean (+11 / -55 / 27 changed; size shrinks because vite 8
bundles deps internally), no ERESOLVE on @tailwindcss/vite
- npx tsc -b --noEmit: clean
- npm test: 220/220 pass, 1.80s (~21x faster than vite 7's 3.05s)
- npm run test:coverage: passes new thresholds
- npm run build: clean, **539ms** with Rolldown (vs 11.41s on vite 7,
~21x speedup), bundle ~1% smaller than vite 7
Closes the iterative vite 5 -> 8 series (#1061 vite 6, #1062 vite 7,
this PR vite 8). Supersedes Dependabot #1040.
Made-with: Cursor
Step 2 of the iterative vite 5 -> 8 migration. Tightens engines.node
to satisfy vite 7's require(esm) floor; no vite.config.ts edits.
Changes:
- vite ^6.4.2 -> ^7.3.2
- @vitejs/plugin-react ^5.1.0 -> ^5.1.4 (npm picked 5.2.0 within ^5.1.4,
which already lists vite ^8 as a peer -> iter 3 won't need to re-bump)
- gitnexus-web engines.node: >=20.0.0 -> ^20.19.0 || >=22.12.0 (vite 7
requirement; gitnexus CLI engines untouched since CLI doesn't use vite)
- .github/actions/setup-gitnexus-web: pin node-version to '20.19.0' so we
don't depend on the floating "20" alias resolving to a high enough patch.
CLI-side actions stay on '20'.
Why no other config changes: vite 7's removed surfaces (sass legacy API,
splitVendorChunkPlugin, transformIndexHtml.transform, optimizeDeps.entries
glob semantics, CORS middleware order) are not used here. The five
resolve.alias entries (@, @shared, gitnexus-shared, anthropic deep import,
mermaid ESM) keep working - alias plugin precedence is unchanged.
Verified locally (Node v22.14.0, well above the new floor):
- npm install: clean, no peer warnings
- npx tsc -b --noEmit: clean
- npm test: 220/220 pass
- npm run build: clean (11.41s, dist tree shape identical, hashes
shifted as expected because vite 7 changed default build.target from
'modules' to 'baseline-widely-available' - bundle is 1-4% smaller)
Iter 3 (vite 8) will follow once this bakes on main.
Made-with: Cursor
Step 1 of the iterative vite 5 -> 8 migration for gitnexus-web. This
PR does the lowest-risk hop: vite 5 -> 6 only. No config or engine
changes are required because:
- @tailwindcss/vite@4.1.18 already lists vite ^6 in its peer range
- @vitejs/plugin-react@5.1.x supports vite ^6
- vitest@3.2.4 supports vite ^6 (peer ^5 || ^6 || ^7)
- vite 6 still supports Node 18/20/22, so engines.node >=20.0.0 stays
- None of vite 6's breaking changes (sass legacy API, postcss-load-config v6,
json.stringify default, environment API, fs.allow auto-detect) touch
this app's vite.config.ts / vitest.config.ts surface
Verified locally:
- npm install: clean, no peer warnings
- npx tsc -b --noEmit: clean
- npm test: 220/220 pass
- npm run build: clean, dist tree shape matches main
Subsequent PRs will land vite 6 -> 7 (engines + setup-node pin) and
vite 7 -> 8 (plugin-react/tailwindcss-vite/vitest co-bumps). This
supersedes Dependabot #1040, which jumped 5 -> 8 in one shot and broke
on @tailwindcss/vite peer resolution.
Made-with: Cursor
Wraps docker/build-push-action with a local composite action that retries
once on failure (upstream keeps retry out of the action per
docker/build-push-action#1422). Adds ignore-error=true on cache-to so GHA
cache export flakes don't fail an otherwise successful push.
- Emit `::notice::` in the resolve step when attempt 2 recovers from a
first-attempt failure, so silent retries are grep-able in run logs and
trending registry/cache flakes stay visible.
- Bind `retry-wait-seconds` via `env:` in the backoff step to match the
env-binding convention used elsewhere in docker.yml (TAG_INPUT, DIGEST,
TAGS) — no direct expression interpolation inside shell bodies.
Preserves existing contract end-to-end: SHA pin, provenance=max, sbom=true,
dual-registry push, `steps.build.outputs.digest` wiring to Cosign and the
build-provenance attestations.
* test(lbug): stabilize rel-csv-split Windows CI with expect.poll
Fixed sleeps assumed readline had already created the first mock stream
within 20ms; windows-latest can lag, causing streams.length===0 and
ENOTEMPTY tempdir cleanup. Poll up to 10s instead (Vitest 4).
Refs #1051
Made-with: Cursor
* test(lbug): use exact toBe assertions in rel-csv-split (DoD §2.7)
- Poll for streams.length === 2 after unblock (two pair keys only)
- disk-full test: streams.length === 1 for single Function|Class row
Made-with: Cursor
* test(lbug): replace rel-csv-split setTimeout waits with expect.poll
Shared pollOpts; drain-listener and disk-full tests now wait on streams.length
instead of fixed 50ms sleeps (DoD §2.7 deterministic tests).
Made-with: Cursor
* ci(docker): mirror signed images to Docker Hub alongside GHCR
docker.yml now publishes to docker.io/abhigyanpatwari/gitnexus{,-web} in
the same build step as the existing GHCR push, so both registries receive
the same digest, the same Cosign keyless signature, and the same SBOM /
build-provenance attestations. The Docker Hub login uses new repo secrets
DOCKERHUB_USERNAME / DOCKERHUB_TOKEN (scoped PAT, not account password).
Supply-chain guarantees carry over unchanged: the signing loop iterates
metadata-action's full tag set, so Docker Hub tags get signed at the
identical digest under the same docker.yml@refs/tags/v* identity. The
ClusterImagePolicy is extended with docker.io / index.docker.io / bare-
namespace globs so admission cannot be sidestepped by registry-prefix
choice. README and .env.example document both registries; RC section in
CONTRIBUTING.md notes the Docker Hub mirror tag.
Closes#1027
* ci(docker): publish to akonlabs Docker Hub namespace; add PR dry-run CI
- Hardcode `akonlabs` as the Docker Hub namespace in metadata-action and
both attestation subject-names (Docker Hub org differs from GitHub org
`abhigyanpatwari`, so `github.repository_owner` would produce the wrong ref)
- Update docs (.env.example, README, CONTRIBUTING) and the Kubernetes
ClusterImagePolicy globs to reference `akonlabs/gitnexus{,-web}`
- Add `pull_request` trigger so the image build runs as CI on every PR
(build only — no push, sign, or attestation)
- Add `workflow_dispatch` with `dry_run: boolean` (default true) for
manual build-only runs; all publish steps gated on
`github.event_name != 'pull_request' && !inputs.dry_run`
tree-sitter-c-sharp ships with `"type": "module"` + `"main": "bindings/node"`
(no file extension) and no `"exports"` field. On Node 22, the bare-package
ESM import hits the deprecated main-field extension resolution and emits
repeated `[DEP0151] DeprecationWarning` lines for every analyze run.
Switch both callsites (parser-loader.ts, parse-worker.ts) to the explicit
subpath `tree-sitter-c-sharp/bindings/node/index.js`. The explicit path
bypasses the deprecated resolution step entirely; types continue to resolve
from the colocated `bindings/node/index.d.ts`. Other tree-sitter grammars
are CommonJS (no `type: "module"`) so they don't trigger DEP0151 and are
left untouched to keep the diff narrow.
Reporter pre-tested the same fix locally (#1013).
* feat(ingestion): GITNEXUS_INDEX_TEST_DIRS opt-in for __tests__ / __mocks__ (#771)
The DEFAULT_IGNORE_LIST hardcodes __tests__ and __mocks__ as
auto-filtered directory names. The comment at ignore-service.ts:273
explicitly documents this as intentional — .gitnexusignore negation
cannot override hardcoded entries. That default is right for the
majority of users, but for Quality Engineering workflows where test
files are the primary index target (tracing coverage via CALLS
edges), there was no escape hatch short of patching the installed
package.
Add an opt-in env var mirroring the GITNEXUS_NO_GITIGNORE /
GITNEXUS_MAX_FILE_SIZE precedent:
`GITNEXUS_INDEX_TEST_DIRS=1` removes __tests__ / __mocks__ from the
effective ignore set. Scope is deliberately limited to these two
names — the issue asked for these specifically, and other
test-adjacent entries (__snapshots__, snapshots, fixtures, .jest)
remain auto-filtered unchanged. .gitnexusignore negation semantics
are not touched; the env var is the orthogonal escape hatch.
Implementation: new `isEffectivelyIgnoredDirectory` helper in
ignore-service.ts wraps the `DEFAULT_IGNORE_LIST.has(name)` check
with the env-var opt-out. Two call-sites swap: shouldIgnorePath
(affects filesystem walker and wiki generator) and
createIgnoreFilter.childrenIgnored (affects directory pruning
during traversal). `isHardcodedIgnoredDirectory` export unchanged —
its contract is "is in the raw list", which remains true for
__tests__ / __mocks__ regardless of env var state (locked in by a
test).
Default behaviour is byte-identical for users who don't set the env
var. 9 new unit tests cover default-unset, opt-in-set, scoped scope
(other hardcoded entries unaffected), and the scope-discipline
guard (future expansion beyond the two named dirs fails loudly).
Env state restored by afterEach to prevent leakage.
Closes#771.
* feat(ingestion): .gitnexusignore negation overrides hardcoded DEFAULT_IGNORE_LIST (#771)
Per @magyargergo's review feedback: rather than add a special-case
GITNEXUS_INDEX_TEST_DIRS env var to unlock __tests__ / __mocks__,
let .gitnexusignore use !pattern negation to override the hardcoded
DEFAULT_IGNORE_LIST — mirroring the .gitignore mental model users
already know.
Implementation:
- New private hasExplicitUnignore(ig, rel) helper that walks ancestor
segments and uses ignore.test(path)'s `unignored` flag to detect
explicit negation. Ancestor-walk is required because .gitignore
negation propagates — !__tests__/ implicitly unignores every
descendant, but ignore.test() only reports unignored: true on the
directly-matched path.
- createIgnoreFilter.ignored() and .childrenIgnored() now check
hasExplicitUnignore BEFORE applying the hardcoded DEFAULT_IGNORE_LIST.
If any ancestor (or the path itself) was explicitly unignored in
.gitnexusignore, the hardcoded block is bypassed.
- shouldIgnorePath stays pure hardcoded-list — the wiki generator and
other callers without per-repo config context keep deterministic
behavior. The #771 override lives only inside createIgnoreFilter,
which IS called with config.
Dropped:
- GITNEXUS_INDEX_TEST_DIRS env var (superseded by the more general
negation mechanism)
- isEffectivelyIgnoredDirectory helper
- Associated env-var help text and unit tests
Added:
- Tip in analyze --help pointing users at .gitnexusignore with
!__tests__/ as the example
- 8 new unit tests covering default behaviour, directory-level
negation, selective overrides, generalisation (!node_modules/),
non-leakage across hardcoded entries, standard non-negation rules
still layering on top, and preservation of shouldIgnorePath /
isHardcodedIgnoredDirectory contracts
Default behaviour (no .gitnexusignore or no negation pattern) is
byte-identical to pre-#771. Users who want to index an auto-filtered
directory add a single !pattern line — no env var, no flag, no
re-install.
Closes#771.
* fix(ingestion): honour re-ignore rules after .gitnexusignore negation (#771)
When .gitnexusignore contains both `!__tests__/` and
`__tests__/generated/`, the parent negation previously short-circuited
and allowed the re-ignored child through. Consult `ig.ignores(rel)`
after `hasExplicitUnignore` so a more-specific rule in the same file
correctly re-ignores a subset — matching .gitignore's last-match-wins
semantics. Adds a compound-pattern test locking this in.
* feat(csharp-scope): unit 1 — scope query + captures orchestrator
First slice of the C# scope-resolution migration (issue #934, RFC #909
Ring 3). Closes `Unit 1` of
docs/plans/2026-04-21-004-feat-csharp-scope-resolution-plan.md.
Adds:
- src/core/ingestion/languages/csharp/query.ts — tree-sitter scope
query covering compilation_unit, namespace (block + file-scoped),
class-like (class/interface/struct/record/enum), method-like
(method/constructor/destructor/local_function/operator), property
and field declarations, using directives, type bindings (parameter
annotations, local variable annotations, constructor inference,
invocation alias), and references (free call, member call including
null-conditional, constructor call, member write).
- src/core/ingestion/languages/csharp/captures.ts — pass-through
orchestrator mirroring python/captures.ts. Import decomposition
(Unit 2), receiver-type-binding synthesis (Unit 3), and arity
metadata synthesis (Unit 5) stub out for future units.
- src/core/ingestion/languages/csharp/cache-stats.ts — PROF
instrumentation mirror of python/cache-stats.ts.
Design notes:
- Return-type / field-type / property-type captures deferred.
tree-sitter-c-sharp does not expose these under a clean named field
that pattern-matches. When Unit 7 parity gate surfaces a gap, add
positional patterns or a post-hoc extractor lookup.
- object_creation_expression with qualified_name type — the qualified
name itself is the reference text; captured as a whole via a
dedicated tag so interpretation in later units can split namespace
+ name.
- Null-conditional calls use positional descendant patterns because
tree-sitter-c-sharp's member_binding_expression and
conditional_access_expression don't expose named fields.
Coverage:
- 23/23 new unit tests in
test/unit/scope-resolution/csharp/csharp-captures.test.ts cover
every capture tag. Confirmed against tree-sitter-c-sharp via the
probe-script loop during development; grammar drift would surface
as a capture-shape assertion failure.
- tsc --noEmit clean.
No changes to shared infrastructure. Resolver wiring + registration
land in Unit 6.
* fix(csharp-scope): capture null-conditional receiver + operator decls
Adversarial review surfaced two Unit 1 bugs that would silently
corrupt the graph once C# is flipped on the scope-resolution path:
- `obj?.Save()` only emitted @reference.name, so receiver-bound
resolution downgraded to the free-call fallback and could mis-link
to an imported `Save`. Capture the conditional_access_expression
receiver under @reference.receiver.
- `operator_declaration` had @scope.function but no @declaration.method
owner, so calls inside operator bodies were attributed to the
enclosing class and the operator itself disappeared from method
lookup. Capture the operator token as @declaration.name (downstream
csharpMethodConfig normalizes to op_Addition etc.).
- `conversion_operator_declaration` was missing from both scope and
declaration sets. Added with the target type as the name anchor.
Arity metadata for overload resolution remains deferred to Unit 5 and
gated behind Unit 7's parity flip, as documented in captures.ts.
* chore(scope-resolution): drop unused python/scopes.scm sibling
The file was documentation-only — the authoritative scope query is
the embedded `PYTHON_SCOPE_QUERY` constant in `python/query.ts`.
Nothing loaded the `.scm` at runtime, so it drifted from the code.
Remove it and update the four doc comments that pointed at it:
- language-provider.ts: "scopes.scm query" → "scope query (embedded
in each language's query.ts)".
- languages/python.ts: capture-vocabulary pointer → query.ts.
- python/query.ts header: drop the "edit both together" note.
- python/receiver-binding.ts: "keeps the .scm declarative" → "keeps
the embedded scope query declarative".
- scope/walkers.ts: "Python's scopes.scm" → "Python's scope query".
Historical plan docs under docs/plans/ still reference scopes.scm but
are frozen artifacts, not living documentation. C# never had a .scm
sibling, so no action needed there.
* feat(csharp-scope): Unit 2 — import interpret + target resolver
Adds the three files Unit 2 of the C# scope-resolution plan calls for:
- `import-decomposer.ts` — inspects each `using_directive` node and
synthesizes `@import.kind/source/name/alias` markers. Kinds:
`namespace` — `using X;` / `using X.Y.Z;`
`alias` — `using Alias = X.Y.Z;` (generics stripped)
`static` — `using static X.Y;`
`global using` maps to namespace (plan's deferred decision); the
`global::` qualifier is stripped before emitting.
- `interpret.ts` — reads the markers and builds `ParsedImport`. Static
using maps to `kind: 'wildcard'` since it brings members into
unqualified scope; Unit 4's merge-bindings tiers wildcards lowest.
Also provides `interpretCsharpTypeBinding` with nullable/single-arg
generic/qualifier stripping so receiver-typed resolution sees the
concrete class name.
- `import-target.ts` — suffix-match adapter returning a single primary
file. Cross-file partial-class aggregation runs later at graph-bridge
time (Unit 6). The csproj-based `resolveCSharpImportInternal` stays
on the legacy path until Unit 7's parity gate surfaces a gap.
- `captures.ts` routes `@import.statement` matches through the
decomposer so the interpreter sees the markers it needs.
Tests cover every using flavor + resolution edge cases. 38/38 scope-
resolution C# unit tests pass; tsc clean.
* feat(csharp-scope): Unit 3 — simple hooks (binding/import/receiver)
Adds simple-hooks.ts mirroring Python's pattern:
- `csharpBindingScopeFor` — delegates to innermost (block scope is
already captured by @scope.block in the query).
- `csharpImportOwningScope` — binds `using` inside a namespace to that
namespace's scope so imports don't leak into sibling namespaces.
File-level using delegates to module. Function-body using (not legal
C# but possible from malformed input) attaches to the function.
- `csharpReceiverBinding` — looks up `this` / `base` in the function
scope's type bindings; returns null for statics, free functions, and
non-Function scopes. `this` / `base` synthesis itself is deferred to
a follow-up (matches Python's receiver-binding.ts pattern).
9 new tests pin delegation semantics. 47/47 C# scope-resolution unit
tests pass; tsc clean.
* feat(csharp-scope): Unit 4 — mergeBindings (using precedence)
Three-tier shadowing, same shape as Python's LEGB merge:
0: local — class members, locals, parameters
1: using — namespace / named / reexport (equal tier; compiler
requires explicit qualifier if two using collide)
2: wildcard — `using static X.Y;` static-member imports
Within the surviving tier, de-dup by DefId (last-write-wins) so a
re-declared `using` cleanly replaces its earlier binding. Explicit
interface implementations bind under their qualified name in the
extractor layer, so they don't collide with plain simple names here.
7 new tests pin precedence + dedup semantics. 54/54 C# scope-resolution
unit tests pass.
* feat(csharp-scope): Unit 5 — arity metadata synthesis + compatibility
Adversarial review flagged overload narrowing as a blocker for the Unit
7 flip. This lands the declaration-side metadata; callsite-side arity
synthesis is a separate gap we'll address if the parity gate surfaces
overload misresolution.
- `arity-metadata.ts` — reads `csharpMethodConfig.extractParameters`
and produces `{ parameterCount, requiredParameterCount,
parameterTypes }`. `params` variadic collapses parameterCount to
undefined (matches Python's `*args` treatment) and appends a literal
`'params'` marker to parameterTypes so the compatibility hook can
detect it without re-reading the AST. Default-valued parameters
contribute to optionalCount → requiredParameterCount = total − optional.
- `arity.ts` — `csharpArityCompatibility(def, callsite)` returns
compatible / incompatible / unknown. Mirrors Python's three-verdict
shape so the central registry's arity filter works without adapter
logic per-verdict.
- `captures.ts` — on every @declaration.method / @declaration.constructor
/ @declaration.function match, synthesize
@declaration.parameter-count, @declaration.required-parameter-count,
and @declaration.parameter-types captures. Covers method_declaration,
constructor_declaration, destructor_declaration, operator_declaration,
conversion_operator_declaration, and local_function_statement.
12 new tests: 5 on captures-side synthesis (method + params + types +
variadic + constructor + local function), 7 on the compatibility hook.
66/66 C# scope-resolution unit tests pass; tsc clean.
* feat(csharp-scope): Unit 6 — wire csharpScopeResolver + register
Creates the public barrel (index.ts) and ScopeResolver (scope-resolver.ts)
and plumbs them into the provider + registry:
- `languages/csharp/index.ts` — re-exports the hook entry points and
documents the 8 known limitations of the registry-primary path
(csproj-driven namespace resolution, multi-file namespace expansion,
type-based overload resolution, nested generics, dynamic, preprocessor
branches, cross-file global using, expression-bodied members).
- `languages/csharp/scope-resolver.ts` — ScopeResolver shape mirroring
Python's. `isSuperReceiver` matches the literal `base` keyword.
`fieldFallbackOnMethodLookup: false` since C# is statically typed
— the type-binding layer already produces precise owner types;
`propagatesReturnTypesAcrossImports: true` since signatures are
authoritative.
- `languages/csharp.ts` — adds the 9 hook entry points to the provider
(emitScopeCaptures, interpretImport, interpretTypeBinding, four
simple hooks, mergeBindings, arityCompatibility, resolveImportTarget).
- `scope-resolution/pipeline/registry.ts` — registers csharpScopeResolver
alongside the Python entry.
MIGRATED_LANGUAGES stays at {Python} — the resolver sits idle until
Unit 7's parity gate confirms ≥99% fixture parity. 368/368
scope-resolution unit tests pass; tsc clean.
* feat(csharp-scope): parity Unit 1 — this/base receiver-binding synthesis
Closes 3 parity failures (51 → 48). Target bucket: Category C from the
parity plan.
Changes:
- `languages/csharp/receiver-binding.ts` (new): walks up from a
function node to the enclosing class/struct/record/interface,
synthesizes `@type-binding.self` captures with boundName `'this'`
(and `'base'` when the enclosing type is a class/record with an
explicit base_list entry). Skips static methods and interface /
struct `base` cases. Anchors to the method's `body` block so the
scope-extractor's positionIndex places the binding inside the
function scope (not the enclosing class scope).
- `languages/csharp/captures.ts`: route `@scope.function` matches
through the synth, emitting the receiver captures as separate
matches.
- `languages/csharp/interpret.ts`: map `@type-binding.self` to
`source: 'self'` (parity with Python).
- `languages/csharp/query.ts`: explicit patterns for `this.X()`,
`base.X()`, and `this.X = ...` / `base.X = ...` assignment writes.
`this` and `base` are anonymous tokens in tree-sitter-c-sharp so
the existing `expression: (_)` pattern (named-only) didn't match.
Tests:
- 8 new unit tests for receiver-binding synthesis edge cases
(class/struct/record/interface, static, nested, constructor,
local function inside method).
- Parity: 48 failed | 127 passed (175) under REGISTRY_PRIMARY_CSHARP=1;
legacy path 175/175 green.
* feat(csharp-scope): parity Unit 2a — foreach + pattern + field captures
Closes 11 parity failures (48 → 37). Partial Unit 2 progress.
Adds type-binding captures for every shape the parity suite exercises
whose resolution path is in-file:
- Typed foreach `foreach (User u in xs)` — @type-binding.annotation
with bindingName `u` and type `User`.
- Var foreach `foreach (var u in xs)` — @type-binding.alias so the
generic-stripper unwraps `List<User>` / `Dictionary<K,V>.Values` to
the element type at chain-follow time. Matches Python's for-loop
alias pattern.
- `is` pattern `if (obj is User u)` — @type-binding.annotation with
scope narrowing simplified to function scope (matches Python's
match-case treatment since we don't emit @scope.block).
- `switch_section > declaration_pattern` (`case User u:`) — no
case_pattern_switch_label wrapper in tree-sitter-c-sharp.
- `recursive_pattern` (`is User { Age: 1 } u` / `case User { ... } u:`)
— named binding via type+name fields on the pattern node.
- Field declaration `private City _city;` — @type-binding.annotation
attached to the class scope for `this._city.X` resolution.
- Property declaration `public User Owner { get; set; }` — same.
- Assignment rebind `alias = Factory()` / `alias = new User()` —
@type-binding.alias / @type-binding.constructor so reassignment
propagates type info to later receiver-typed resolution.
Closed tests: foreach (3), var foreach Tier 1c (2), is-pattern (1),
switch pattern (2), recursive_pattern (3). Remaining 37 include
tests that need cross-file same-namespace visibility (field chains,
assignment chain, cross-file return-type propagation) — deferred to
Unit 5 where the IMPORTS/cross-file work lives.
74/74 scope-resolution unit tests pass; legacy path 175/175 green.
* feat(csharp-scope): parity Unit 2b — same-namespace cross-file visibility
Closes 3 parity failures (37 → 34). Adds the C#-specific implicit
import that has no syntactic counterpart: every type declared in
`namespace X` is visible to every other file also declaring
`namespace X`, without any `using` directive.
Changes:
- `scope-resolution/contract/scope-resolver.ts` — new optional hook
`populateNamespaceSiblings(parsedFiles, indexes, { fileContents })`.
Most languages leave it undefined; Python / TypeScript / Java need
explicit imports so there's no analogous pass.
- `scope-resolution/pipeline/run.ts` — invoke the hook after
`buildWorkspaceResolutionIndex` and before
`propagateImportedReturnTypes` so the return-type pass sees
cross-file sibling class bindings.
- `languages/csharp/namespace-siblings.ts` (new) — groups top-level
class-like defs by namespace name (extracted from source via regex
since `file_scoped_namespace_declaration` scope range covers only
the declaration line, not the rest of the file). Injects sibling
classes into each file's Module AND Namespace scope bindings with
origin='namespace'. Local declarations shadow cross-file siblings
via mergeBindings tier precedence.
- `languages/csharp/scope-resolver.ts` — wire the hook.
74/74 scope-resolution unit tests pass; legacy path 175/175 green;
34 parity failures remain (was 37) under REGISTRY_PRIMARY_CSHARP=1.
* feat(csharp-scope): parity Unit 2c — alias/await/return-type captures
Closes 7 parity failures (34 → 27). Adds the remaining type-binding
shapes the parity suite exercises:
- `var alias = u;` / `alias = u;` — identifier-to-identifier alias.
The resolver's chain-follow walks alias → u → u's declared type.
- `var u = svc.GetUser();` — chained method call alias. Anchors on
the method_access_expression's `name` field; chain-follow picks up
GetUser's return type.
- `var u = await Factory();` / `await svc.Get();` — await propagation.
Strips the `await_expression` wrapper; interpret layer's
`stripGeneric` handles `Task<T>` / `ValueTask<T>` unwrapping.
- `public User GetUser() { ... }` — method return-type annotation
via `@type-binding.return`. Required for `propagateImportedReturnTypes`
to see the return type in later cross-file passes. Covers identifier,
generic_name, qualified_name, and nullable_type return shapes.
74/74 scope-resolution unit tests pass; legacy path 175/175 green;
27 parity failures remain under REGISTRY_PRIMARY_CSHARP=1.
* feat(csharp-scope): parity Unit 3a — cross-namespace `using` binding
Closes 2 parity failures (27 → 25). Extends the namespace-siblings
pass to resolve `using X;` directives against known namespace
buckets: for each `using` that targets a namespace declared
somewhere in the workspace, inject that namespace's classes into
the importer's module scope with origin='namespace'.
This is the scope-resolution analog of legacy's csproj-driven
directory↔namespace mapping. Without it, `new User()` in
`Services/UserService.cs` (namespace MyApp.Services) can't see the
User class in `Models/User.cs` (namespace MyApp.Models) even with
`using MyApp.Models;` — the scope-resolver layer doesn't have
csproj metadata to translate the dotted namespace path into a
directory lookup.
Legacy 175/175 green; 25 parity failures remain.
* feat(csharp-scope): parity Unit 3b — constructor CALLS emission
Closes 3 parity failures (25 → 22). Adds constructor-form CALLS
edge emission + C# 12 primary constructor synthesis.
Changes:
- `scope-resolution/passes/free-call-fallback.ts`: when a site's
callForm === 'constructor', look up the class def (not a callable)
and pick its explicit Constructor def via workspaceIndex's
memberByOwner — or fall back to the Class def itself for implicit
constructors. Matches legacy behavior (targetLabel === 'Constructor'
when explicit, 'Class' when implicit).
- `scope-resolution/pipeline/run.ts`: pass workspaceIndex to the
free-call fallback.
- `languages/csharp/captures.ts`: synthesize @declaration.constructor
for C# 12 primary constructors — `class User(string name, int age)`
/ `record Person(string First, string Last)`. The parameter_list is
a named child of the class_declaration / record_declaration (not a
separate constructor_declaration node). Skip the synthesis when
the type already has an explicit constructor to avoid duplicates.
Emits @declaration.parameter-count + required-parameter-count
alongside.
Legacy 175/175 green; 376/376 scope-resolution unit tests pass;
22 parity failures remain.
* feat(csharp-scope): parity Unit 3c — static call + default-namespace
Closes 2 parity failures (22 → 21).
- `receiver-bound-calls.ts`: add Case 5 for class-as-receiver. When
`Animal.Classify()` has an identifier receiver that resolves to a
Class binding (rather than a variable with a typeBinding), look up
the member on the class's MRO chain. Covers C#-style static calls
and any type-qualified member access. Python doesn't hit this
because `ClassName.method()` is syntactically identical to a free
call there.
- `namespace-siblings.ts`: treat files with no `namespace X;`
declaration as living in the default (empty-name) bucket, so
types declared in no-namespace files share cross-file visibility.
Required for fixtures without explicit namespaces (e.g. the
method-enrichment fixture's Animal/App/Dog classes).
Legacy 175/175 green; 21 parity failures remain.
* feat(csharp-scope): parity Unit 4 — callsite arity synthesis (infra)
Synthesize @reference.arity on every invocation_expression and
object_creation_expression by counting `argument` named children of
the backing `argument_list`. Wires the capture-to-Callsite pipeline
shared extractor already consumes (`scope-extractor.ts:878`).
No parity-count movement: the remaining arity-adjacent failures
(overload disambiguation, optional-parameter dedup, variadic
resolution) need type-based argument inference or member-call dedup,
both explicitly deferred in the plan's Known Limitations section.
This commit is infrastructure — future work lands on top of it.
Legacy 175/175 green; 21 parity failures remain.
* feat(csharp-scope): parity Unit 5a — IMPORTS edge + static-using mapping
Closes 1 parity failure (21 → 20). Fixes cross-file IMPORTS edge
emission for C#:
- `languages/csharp/interpret.ts`: map `using static X.Y;` to
`kind: 'namespace'` rather than `'wildcard'`. The File→File
IMPORTS edge needs a non-wildcard kind to survive finalize's
Phase 4 (wildcard-expanded edges drop to empty when the provider
doesn't implement `expandsWildcardTo`). Unqualified static-member
access is a deferred limitation — covered by the namespace-siblings
cross-namespace pass for type lookups, and documented under the
module's Known Limitations.
- `languages/csharp/import-target.ts`: progressive prefix stripping.
`using CrossFile.Models;` in a repo laid out `Models/User.cs` (no
`CrossFile/` directory) works because the legacy resolver consults
csproj; the scope-resolver tries each suffix of the dotted path
against `.cs` files. Also handles `using static NS.Type;` by
stripping leading segments until a direct match lands.
- `test/unit/scope-resolution/csharp/csharp-imports.test.ts`: update
the `using static` test to the new namespace-kind shape.
376/376 scope-resolution unit tests pass; legacy 175/175 green;
20 parity failures remain.
* feat(csharp-scope): parity Unit 5b — return-type module hoist + chain fallback
Closes 1 parity failure (20 → 19) and lays groundwork for Unit 6.
Based on investigation-agent findings, addresses cluster of 7
cross-file + chain tests whose return-type bindings were stuck at
Class scope and invisible to the chain-follow and propagation passes.
Changes:
- `languages/csharp/simple-hooks.ts::csharpBindingScopeFor`: when the
declaration is a `@type-binding.return`, hoist the binding all the
way to the Module scope. The central extractor's auto-hoist only
promotes one level (Function → Class); for C# methods the parent
is always a Class, so without this override the return binding
never reaches Module where chain-follow and cross-file
`propagateImportedReturnTypes` read from.
- `scope-resolution/passes/compound-receiver.ts`: when the
class-scope typeBindings lookup at `objClass.typeBindings.get(
methodName)` misses, walk up from the class scope through the
parent chain (→ Module) for a return-type binding. Preserves the
existing class-scope fast-path while restoring owner-chain lookup
for languages that hoist to Module.
Python parity suite stays 204/204 green on both flag paths;
legacy C# 175/175 green; 19 C# parity failures remain.
* feat(csharp-scope): parity Unit 5c — switch-expr + reasons + ACCESSES 1.0
Closes 4 parity failures (19 → 15).
- `languages/csharp/query.ts`: add captures for `switch_expression_arm`
with `declaration_pattern` and `recursive_pattern`. C# expression-
switch (`obj switch { User u => ..., Repo { Name: "x" } r => ... }`)
uses a different AST node from classic `switch_statement`'s
`switch_section` — needed separate query patterns.
- `scope-resolution/passes/receiver-bound-calls.ts`: replace the
self-describing `'scope-resolution: *-receiver'` reason strings
(which fail legacy-parity consumer filters) with the legacy
convention: `'import-resolved'` when the resolved member lives in
a different file, `'global'` otherwise. Mirrors
`free-call-fallback.ts`'s existing reason logic.
- `scope-resolution/passes/receiver-bound-calls.ts`: pass
`confidence: 1.0` to `tryEmitEdge` for write/read ACCESSES edges,
matching legacy DAG behavior (default 0.85 was legacy-CALLS).
Python parity 204/204 on both flag paths; legacy C# 175/175;
15 C# parity failures remain.
* feat(csharp-scope): parity Unit 5d — cross-file typeBinding mirror
Closes 3 parity failures (15 → 12).
`languages/csharp/namespace-siblings.ts`: extend the pass to mirror
method return-type bindings from accessible sibling files' Module
scopes into the importer's Module scope. "Accessible" =
same-namespace siblings + `using namespace X;` targets.
Without this mirror, `var u = svc.GetUser()` in App.cs couldn't
chain-follow to User even after Unit 5b's module-scope hoist:
`GetUser → User` lived on User.cs's Module scope, which isn't on
the ancestor chain of App.cs's function scope, and
`propagateImportedReturnTypes` only mirrors across explicit
ImportEdge targets (not same-namespace implicit visibility).
Closes: var-invocation return type, async/await u.Save (ambient
namespace), cross-file return-type propagation (via u.Save /
u.GetName in Program.cs).
Python parity 204/204 on both flag paths; legacy C# 175/175;
12 C# parity failures remain.
* feat(csharp-scope): parity Unit 5e — namespace-prefix bucket matching
Closes 2 parity failures (12 → 10).
`languages/csharp/namespace-siblings.ts`: when matching accessible
namespaces against class buckets, also probe every dotted prefix.
`using static CrossFile.Models.UserFactory;` parses into the
importer's accessible-namespace set as the full type path, but the
matching bucket is keyed on the containing namespace
(`CrossFile.Models`). Walking back through the dotted segments
ensures the static-using importer sees the containing namespace's
sibling files' return-type bindings.
Legacy 175/175 green; 10 C# parity failures remain.
* feat(csharp-scope): parity Unit 6a — class-like owner extension
Closes 1 parity failure (10 → 9). Extends `populateClassOwnedMembers`
to recognize Interface / Struct / Record / Enum / Trait as class-like
owners, not just Class.
The C# scope query collapses interface_declaration / struct_declaration
/ record_declaration / enum_declaration to @scope.class (they share
body-scope semantics), but the declaration-side tags produce defs of
type Interface / Struct / Record / Enum. `populateClassOwnedMembers`
previously only looked for Class-typed defs in class scopes, so
interface members (including C# 8+ default methods) never got
ownerIds — making them invisible to `findOwnedMember` via
`memberByOwner`.
With this fix, `user.Validate()` on a variable typed as `IValidator`
resolves correctly: receiver-bound-calls Case 4 finds IValidator via
findClassBindingInScope (which already accepted Interface), walks the
chain, and findOwnedMember locates Validate now that the interface
default has a proper ownerId.
Legacy C# 175/175 green; Python parity 204/204 on both flag paths;
9 C# parity failures remain.
* feat(csharp-scope): parity Unit 6b — member-call dedup + handled-site fix
Closes 1 parity failure (9 → 8). Adds the missing legacy-parity
behavior: collapse multiple member-call sites from the same caller
to the same target into one CALLS edge.
Changes:
- `scope-resolution/contract/scope-resolver.ts`: new optional
`collapseMemberCallsByCallerTarget` flag. Default false (preserves
the per-site invariant); C# sets it true.
- `scope-resolution/graph-bridge/edges.ts`: dedup key drops
`line:col` when `collapseByCallerTarget` is on AND edgeType is
`CALLS` (ACCESSES writes keep per-site granularity).
- `scope-resolution/passes/receiver-bound-calls.ts`: plumbs
`collapse` through every `tryEmitEdge` call, and crucially marks
`handledSites.add(siteKey)` whenever a resolved def was found —
not only when the edge was freshly emitted. Otherwise the site
leaked through to `emitReferencesViaLookup` which re-emitted a
per-site edge, defeating the collapse.
- `languages/csharp/scope-resolver.ts`: opt in to the collapse.
Python parity 204/204 on both flag paths; legacy C# 175/175 green;
8 C# parity failures remain.
* feat(csharp-scope): parity Unit 6c — Dictionary.Values / .Keys unwrap
Closes 2 parity failures (8 → 6).
Dictionary<K,V>.Values in a foreach binds the element to V; .Keys
binds to K. Without this, `foreach (var user in data.Values)` where
`data: Dictionary<string, User>` couldn't propagate user's type to
User, and `user.Save()` stayed unresolved.
Changes:
- `languages/csharp/interpret.ts`: don't strip the qualifier when
the final dotted segment is a known collection accessor
(`Values` / `Keys`). Preserves the dotted form so downstream
resolvers can unwrap the receiver's generic type based on the
suffix.
- `scope-resolution/passes/compound-receiver.ts`: new
`extractDictionaryArgs` helper splits `Dictionary<K, V>` at the
top-level comma. In the dotted-access walk, detect trailing
`.Values` / `.Keys` and return V/K via findClassBindingInScope
instead of the normal class-walk (Dictionary itself isn't a
local class def).
- Handles nested cases: `this.data.Values` walks `this.data`
recursively (resolving `data` as a field on `this`'s class)
before applying the unwrap.
- `scope-resolution/passes/receiver-bound-calls.ts` Case 3b: when
the typeRef's trailing segment is an accessor, pass the raw
dotted path to `resolveCompoundReceiverClass` without appending
`()` — the extra parens would misroute to the call-expression
branch.
Python parity 204/204 on both flag paths; legacy C# 175/175 green;
6 C# parity failures remain.
* feat(csharp-scope): parity Unit 6d — using-static member injection
Closes 2 parity failures (6 → 4). `using static X.Y.Z;` now injects
every public static method of class Z into the importer's module
scope, so `Record("hi")` (without `Logger.` qualifier) resolves to
`Logger.Record` as a free call.
`languages/csharp/namespace-siblings.ts`: regex-scan each file's
source for `using static X.Y.Z;` directives. For each, look up the
class Z in the `X.Y` namespace bucket, walk its owning file's
localDefs for method/function members with `ownerId === Z.nodeId`,
and inject them as `origin: 'import'` bindings in the importer's
module-scope finalized bindings map. `findCallableBindingInScope`
then picks them up via its imported-bindings check.
Closes: variadic `Record(params string[])` + heritage arity
narrowing `WriteAudit`.
Python parity 204/204 on both flag paths; legacy C# 175/175 green;
4 C# parity failures remain (interface-dispatch pass + type-based
overload disambiguation).
* feat(csharp-scope): parity Unit 6e — overload disambig + interface dispatch + FLAG FLIP
Closes the final 4 parity failures (4 → 0). C# now runs the
registry-primary scope-resolution path by default — added to
MIGRATED_LANGUAGES.
Changes:
- `scope-resolution/scope/walkers.ts`: was already extended in
Unit 6a to recognize Interface/Struct/Record/Enum as class-like
owners (interface default methods get ownerIds).
- `scope-resolution/passes/receiver-bound-calls.ts`: build
IMPLEMENTS edge index → emit secondary `interface-dispatch`
CALLS edges to every implementor's same-named member when the
primary receiver-typed edge targets an Interface method (closes
heritage CreateUser CALLS-count test).
- `scope-resolution/passes/receiver-bound-calls.ts`: new
`pickOverload` helper narrows multi-valued
`membersByOwner.get(owner).get(name)` candidates by arity then
argument types. Replaces the first-seen `findOwnedMember` lookup
in Case 4 so receiver-typed overloaded calls pick the right def.
- `scope-resolution/passes/free-call-fallback.ts`: new
`pickImplicitThisOverload` walks up to the enclosing class scope
and applies the same arity + argument-type narrowing for free
calls inside a class body (`Lookup("alice")` → `Lookup(string)`).
- `scope-resolution/workspace-index.ts`: new `membersByOwner`
multi-valued index (`Map<owner, Map<name, Def[]>>`) preserves
every overload alongside the existing first-seen `memberByOwner`.
- `scope-resolution/graph-bridge/node-lookup.ts` +
`scope-resolution/graph-bridge/ids.ts`: include parameter-types
suffix in the qualified lookup key for Method nodes. Legacy
parse-phase encodes the type tag into the node id (`Method:f.cs:
UserService.Lookup#1~int`); without this two same-arity overloads
collapsed to one lookup entry and routed to the wrong graph node.
- `scope-resolution/contract/scope-resolver.ts`: new
`collapseMemberCallsByCallerTarget` opt-in flag (was added in
Unit 6b for member-call dedup; documented here).
- `gitnexus-shared/src/scope-resolution/reference-site.ts`: new
`argumentTypes` field carrying inferred per-arg types.
- `scope-extractor.ts`: read @reference.parameter-types capture into
`site.argumentTypes` and add it + the declaration-arity tags to
KNOWN_SUB_TAGS so the anchor-detection picks the right anchor.
- `languages/csharp/captures.ts`: synthesize @reference.parameter-types
by inferring arg types from literal AST nodes (integer_literal →
'int', string_literal → 'string', constructor_expression →
type-name, etc).
- `languages/csharp/scope-resolver.ts`: opt in to
`collapseMemberCallsByCallerTarget`.
- `registry-primary-flag.ts`: **add CSharp to MIGRATED_LANGUAGES**.
Final state:
- C# parity: 175/175 green on flag-on AND flag-off.
- Python parity: 204/204 green on both flag paths (no regression).
- TypeScript clean.
51 → 0 failures across 18 commits on `feat/csharp-scope-resolution`.
* refactor(scope-resolution): extract language-specific accessor unwrap to provider hook
Optimizer pass: move C# Dictionary-family `.Values`/`.Keys` handling
out of the shared `compound-receiver.ts` (where it had hardcoded
regex + accessor names) into a provider-level
`unwrapCollectionAccessor` hook. The shared pass now takes an
arbitrary language-specific unwrap function; C# supplies its
Dictionary implementation in `languages/csharp/accessor-unwrap.ts`.
Related cleanup in `receiver-bound-calls.ts` Case 3b: replace the
hardcoded `tail === 'Values' || tail === 'Keys'` accessor check with
a try-dotted-walk-first / fall-back-to-call-form strategy. This
removes the last C#-specific branch in the shared pass and makes the
logic generalize cleanly to other languages that use property-style
accessors for collection views (Kotlin `.size`, future languages).
Changes:
- `scope-resolution/contract/scope-resolver.ts`: new optional
`unwrapCollectionAccessor(receiverType, accessor) => string | undefined`
hook. Documented as language-specific with examples.
- `scope-resolution/passes/compound-receiver.ts`: delete
`extractDictionaryArgs`, accept `unwrapCollectionAccessor` via
options, call it for trailing accessor segments.
- `scope-resolution/passes/receiver-bound-calls.ts`: plumb the hook
through to `resolveCompoundReceiverClass`, remove the
C#-hardcoded Case 3b accessor check.
- `languages/csharp/accessor-unwrap.ts` (new): C# Dictionary-family
regex + element-type extraction.
- `languages/csharp/scope-resolver.ts`: opt in.
Audit outcome: everything else added across the 19 C# migration
commits is either correctly scoped to `languages/csharp/` (query,
captures, namespace-siblings, receiver-binding, interpret, imports)
or correctly generic in shared paths (argumentTypes field,
collapseMemberCallsByCallerTarget flag, overload narrowing via
parameterTypes, interface-dispatch via IMPLEMENTS edges, class-like
owner extension for Interface/Struct/Record/Enum, type-tagged node
IDs, module-scope return-type lookup fallback).
175/175 C# green on both flag paths; 204/204 Python green on both
flag paths; TypeScript clean.
* refactor(scope-resolution): gate module-scope typeBinding walk-up on hook
Add optional `hoistTypeBindingsToModule` to the ScopeResolver contract
and gate the Module-scope walk-up in `resolveCompoundReceiverClass` on
it. Only providers that hoist method return-type bindings to Module
scope (C#) opt in; Python and other providers no longer traverse that
fallback path.
Closes the architectural leak flagged in the production-readiness
review: the walk-up was unconditional and therefore widened Python's
code path despite existing only for C#.
No behavior change for C# (hook=true restores the prior lookup). No
behavior change for Python (hook undefined = walk-up skipped, matching
pre-PR behavior).
Verified:
- npx tsc --noEmit clean
- C# unit suite 74/74 passing
- C# + Python integration 388/388 passing
* refactor(csharp-scope): remove as-unknown-as double casts in scope-resolver
Tighten three type boundaries that were previously papered over with
`as unknown as` casts:
* `CsharpResolveContext.allFilePaths`: `Set<string>` → `ReadonlySet<string>`.
The orchestrator only hands out a read-only view; drop the widening
cast at the resolver-adapter site.
* `resolveCsharpImportTarget`: call passes the narrow context directly.
`WorkspaceIndex` is `unknown` in the shared contract, so the
`as unknown as WorkspaceIndex` cast was gratuitous — structural
assignability covers it.
* `csharpMergeBindings`: drop unused `_scope: Scope` parameter. The
implementation never read it; the cast chain in `scope-resolver.ts`
existed only to satisfy an unused slot. LanguageProvider.mergeBindings
now wraps with a tiny arrow adapter; ScopeResolver.mergeBindings
passes through directly.
No runtime behavior change. `grep 'as unknown as' csharp/scope-resolver.ts`
returns zero matches.
Verified:
- npx tsc --noEmit clean
- C# unit + integration 462/462 passing (incl. Python integration)
* test(csharp-scope): integration fixtures for Units 6c/6d/6e runtime behavior
Close the integration-coverage gap flagged in the production-readiness
review. Units 6c (collection-accessor unwrap), 6d (using-static member
injection), and 6e (overload disambig + interface dispatch) previously
had only hook-level unit tests; the end-to-end wiring was exercised
only by the parity harness.
Three minimal fixtures + four new it() blocks:
* csharp-collection-accessor — RenderAll iterates
Dictionary<string, Widget>.Values and calls .Render(); asserts the
CALLS edge lands on Widget.Render.
* csharp-using-static — `using static Helpers.MathUtils;` makes
Square(int) a free-callable in the consumer; asserts the CALLS
edge lands on MathUtils.Square.
* csharp-overload-interface — three assertions:
1. Run → Log binds to the 2-arg overload only (arity narrowing);
verified via target Method node's parameterTypes.length === 2.
2. Run → Greet emits one primary edge to IGreeter.Greet plus two
reason='interface-dispatch' siblings to En/FrGreeter.Greet.
3. Interface-dispatch fan-out excludes the primary target.
Verified:
- csharp integration 189/189 passing
* docs(scope-resolution): de-c#-ify optional-hook doc-comments on contract
Rewrite the doc-comments on four optional hooks so they describe the
behavior and when a provider would enable it, rather than naming C#
as the sole consumer. Hook names were already generic — only the
comments had baked in one-language framing, which risked discouraging
future reuse.
Affected hooks:
* unwrapCollectionAccessor
* collapseMemberCallsByCallerTarget
* populateNamespaceSiblings
* hoistTypeBindingsToModule
Language-specific rationale stays where it belongs — next to the hook
assignment in `languages/csharp/scope-resolver.ts`. Zero-match grep for
`C#|csharp|CSharp` in the contract file confirms the separation.
No code change.
* docs(csharp-scope): justify regex-based namespace-sibling detection
Record why `namespace-siblings.ts` uses regex over AST walks and
enumerate the known misses so the next reader has ground to stand on:
* `global using static X.Y;` — no plain `using static` token.
* Aliased `using static X = Y.Z;` — `=` breaks the pattern.
* Attributed namespace declarations between `]` and `{`.
* Multi-namespace files — first-wins attribution.
* Preprocessor-gated namespace declarations — textual branch only.
Rationale: the pass is file-path-driven and the tree-sitter tree isn't
available at its call site (the orchestrator feeds raw fileContents);
re-parsing to count namespaces would cost more than the regex walk.
Refactor to AST-driven detection is deferred to a separate PR.
Mirrored the known-miss list into `csharp/index.ts`'s limitations
ledger so the operator-visible surface and the in-code justification
stay in sync.
No code change.
* refactor(csharp-scope): AST-driven namespace detection with treeCache reuse
Replace regex-over-source-content with tree-sitter AST walks in
namespace-siblings.ts; thread the orchestrator's treeCache through
the populateNamespaceSiblings hook so the pass reuses the same parse
trees `extractParsedFile` already consumed (single-source-of-truth
for the AST — no double-parse).
Behavior gains (no longer "known misses"):
* `global using static X.Y;` is now detected.
* Aliased `using static X = Y.Z;` is now detected.
* Attributed namespace declarations (`[attr] namespace X`) parse
correctly because tree-sitter sees them as one node.
* Preprocessor-gated namespace declarations parse via the grammar.
Contract change (additive, optional):
* `populateNamespaceSiblings` ctx now carries an optional
`treeCache?: { get(filePath): unknown }`. Existing providers that
don't set it on `RunScopeResolutionInput` see undefined, and the
hook falls back to a fresh parse (current behavior preserved on
cache miss).
Limitation ledger updated in csharp/index.ts: the AST-based detection
removes 4 of the 5 prior known misses; only "first-wins multi-namespace
file attribution" remains.
Verified:
- npx tsc --noEmit clean
- C# + Python integration 393/393 passing
* refactor(python-scope): remove as-unknown-as casts in scope-resolver (mirrors Unit 2)
Replay the C# scope-resolver cleanup on the Python side so both
providers share a single clean pattern:
* Drop `ws as unknown as WorkspaceIndex` — `WorkspaceIndex` is
`unknown` in the shared contract, so the narrow context assigns
structurally without a cast.
* Drop `{ id: scopeId } as unknown as Scope` — `pythonMergeBindings`
never read the scope (the parameter was `_scope`), so the stub
was a type-only ghost. Signature is now `(bindings)` and the
LanguageProvider slot wraps with an arrow adapter.
* Drop `allFilePaths as Set<string>` — the orchestrator hands a
`ReadonlySet<string>`; we copy it into a `Set` at the resolver
adapter so the legacy downstream `resolvePythonImportInternal`
chain (typed for mutable `Set<string>`) keeps working. The copy
is O(N) once per import, trivial cost.
Left intact on purpose: the `(callsite, def) → (def, callsite)`
arrow wrapper on `arityCompatibility`. That's a documented shape
difference between `LanguageProvider.arityCompatibility(def, callsite)`
and `ScopeResolver.arityCompatibility(callsite, def)`; both providers
(Python + C#) carry the same wrapper. Reconciling is a separate
refactor across both contracts.
No runtime behavior change.
Verified:
- npx tsc --noEmit clean
- Python + C# unit + integration suites 529/529 passing
* docs(scope-resolution): document I1-I8 invariants, source-of-truth, and same-graph guarantee
Promote contract knowledge that was implicit in code into the canonical docs
so future migrations and the next reviewer don't have to reverse-engineer it.
contract/scope-resolver.ts:
* Migration cookbook lists every optional hook (was: only the two
booleans), with one-line guidance per hook including when to enable
`hoistTypeBindingsToModule`.
* Contract Invariants I1-I7 are now spelled out in full (was: only
I1/I3/I5 summarized with a pointer to a plan file). Added new I8
"post-finalize hooks may mutate Scope.typeBindings and indexes.bindings;
consumers must not freeze or snapshot before all post-finalize hooks
have run".
* New "Semantic-model source of truth" section: ParsedFile is the
single semantic model; passes that need AST-level facts must reuse
the orchestrator's treeCache rather than re-parse.
* New "Same-graph guarantee" section: legacy DAG and scope-resolution
emit indistinguishable edges (node identity, edge vocabulary,
confidence). CI parity workflow enforces this.
gitnexus-shared/src/scope-resolution/parsed-file.ts:
* Added "Source-of-truth invariant" pointer paragraph.
ARCHITECTURE.md (Coexistence section):
* Updated migrated-language list (Python + C#).
* Added "Same-graph guarantee" subsection.
* Added "Semantic-model source of truth" subsection.
* Filled in the ScopeResolver hook table with the five optional hooks
that landed in this branch (unwrapCollectionAccessor,
collapseMemberCallsByCallerTarget, populateNamespaceSiblings,
hoistTypeBindingsToModule, fieldFallbackOnMethodLookup).
* Added C# rows to the code-references table.
Verified:
- npx tsc --noEmit clean
- C# + Python integration 393/393 passing
* refactor(scope-resolution): consume SemanticModel as single authoritative store
Unify scope-resolution and legacy parse into one symbol index per the
industry pattern (Roslyn / tsc / rust-analyzer). Scope-resolution
passes now consume `SemanticModel.methods` / `SemanticModel.fields` /
`SemanticModel.symbols` for all symbol-keyed lookups. The legacy DAG
already read from these; the drift — two parallel owner-keyed indexes
populated by two writers with divergent ownerId semantics — is closed.
Changes:
* `MethodRegistry.lookupAllByOwner(owner, name)`: new API returning
every overload without arity narrowing. Powers `findOwnedMember` /
`pickOverload`.
* `pipeline/run.ts` reconciliation pass: after
`provider.populateOwners(parsed)`, iterate `parsed.localDefs[i]`
and register methods/fields into the SemanticModel under the
corrected ownerId. Idempotent — skips defs already present under
`(ownerId, simple)` by nodeId, so unmigrated languages whose
legacy extractor already set ownerId (C#) don't double-register.
Closes the Python gap where class-body methods were invisible to
`MethodRegistry` because the legacy Python method extractor
couldn't resolve `enclosingClassId` at parse time.
* `WorkspaceResolutionIndex` slimmed to Scope-valued maps only
(`classScopeByDefId`, `moduleScopeByFile`). Dropped `memberByOwner`,
`membersByOwner`, `defsByFileAndName`, `callablesBySimpleName` —
all symbol-keyed duplicates of SemanticModel indexes.
* Walker helpers now consume SemanticModel:
- `findOwnedMember(owner, name, model)` → methods then fields
fallback (ACCESSES writes target Property/Variable defs too).
- `findExportedDefByName` fallback walks every Module scope's
`origin === 'local'` bindings via `index.moduleScopeByFile`
(preserves the module-export-visibility filter that
SymbolTable.fileIndex can't cheaply encode).
- `findExportedDef` reads `moduleScope.bindings` directly.
* `pickOverload` in receiver-bound-calls.ts falls back to
`model.fields.lookupFieldByOwner` when method lookup returns empty,
fixing ACCESSES write edges that receive a Property target.
* `phase.ts` threads `resolutionContext.model` into
`RunScopeResolutionInput`.
Boundary rule, enforced by file placement:
- symbol-indexed lookups (key = nodeId / name / filePath) →
`SemanticModel`
- Scope-valued lookups (value = `Scope`) →
`WorkspaceResolutionIndex`
Research synthesized from web-researcher + Explore + best-practices +
system-architect agents; canonical references: Roslyn Overview,
rust-analyzer architecture, stack-graphs paper.
Verified:
- npx tsc --noEmit clean
- C# + Python integration 393/393 passing
* docs(scope-resolution): refresh comments after dropping duplicated indexes
Replace references to the now-deleted `memberByOwner` /
`callablesBySimpleName` index fields with comments that describe the
actual lookup path (`SemanticModel` registries + scope-tied module
bindings). Pure doc cleanup; no behavior change.
* feat(scope-resolution): extract reconciliation pass + add parity validator
Extract the SemanticModel reconciliation pass (previously inline in
`pipeline/run.ts`) into a dedicated module with:
* `reconcileOwnership(parsedFiles, model)` — pure function returning
stats (methodsRegistered / fieldsRegistered / skippedAlreadyPresent).
Idempotent; safe to re-run.
* `validateOwnershipParity(parsedFiles, model, onWarn)` — dev-mode
runtime validator for Contract Invariant I9. Walks every def with
an `ownerId` and asserts it is reachable via
`model.methods.lookupAllByOwner` or `model.fields.lookupFieldByOwner`.
Soft-fails via `onWarn`; never throws.
Validator is gated on both `NODE_ENV !== 'production'` and
`VALIDATE_SEMANTIC_MODEL !== '0'` so production incurs zero cost but
development surfaces any drift between `parsed.localDefs` ownership and
the registries.
12 new unit tests cover:
* happy path: method, property, Variable registration
* edge case: defs without ownerId are skipped
* idempotency: second call is a no-op
* coexistence: defs the legacy extractor already registered (via
`model.symbols.add`) are skipped on reconcile
* overloads: multiple methods under the same (owner, name)
* validator: no warnings after reconciliation
* validator: warns on drift
* validator: no-op under NODE_ENV=production
* validator: no-op when VALIDATE_SEMANTIC_MODEL=0
* validator: warns on missing Property same as missing Method
Verified:
- npx tsc --noEmit clean
- reconcile-ownership unit tests 12/12 passing
- C# + Python integration 393/393 passing
* refactor(scope-resolution): narrow handles + tighten required params
Two small hygiene fixes that fell out of the unified-model work:
* Introduce `readonlyModel: SemanticModel` in `runScopeResolution`
immediately after reconciliation so the write/read phase boundary
is explicit at the code level. Downstream passes (receiver-bound,
free-call) receive the narrowed `SemanticModel` rather than the
`MutableSemanticModel` that only the reconciliation pass needs.
The type system now rejects accidental writes in the read phase.
* Make `emitFreeCallFallback`'s `workspaceIndex` parameter required.
It's now always passed (every caller threads it through), and the
`workspaceIndex?` guard was dead code. Also drops the `| undefined`
branch from `pickConstructorOrClass` which no caller can hit.
No behavior change.
* docs(semantic-model): document unified single-source-of-truth invariant (I9)
Add Contract Invariant I9 to the ScopeResolver contract and write the
single-source-of-truth + write/read phase contract into both the
SemanticModel file-head and ARCHITECTURE.md.
Three landing points so the rule is reachable from every entry:
* contract/scope-resolver.ts — new I9 entry in the Contract
Invariants list: scope-resolution passes consult SemanticModel
exclusively for symbol-keyed lookups; WorkspaceResolutionIndex is
reserved for Scope-valued maps. Documents the two-phase write
(legacy parse + reconcileOwnership) and the narrowed-handle read
posture. Calls out the reconciliation shim as transitional.
* model/semantic-model.ts — new "Single-source-of-truth invariant"
and "Write / read phase contract" sections in the file-head.
Three ordered write phases (parse → reconcile → attachScopeIndexes),
then frozen for readers.
* ARCHITECTURE.md § "Semantic-model source of truth" — expanded
subsection covering both invariants (ParsedFile = AST truth,
SemanticModel = symbol truth), the write/read phase diagram, and
the reconciliation-shim rationale.
No code change.
* test(scope-resolution): rewrite workspace-index test for slimmed index
The test file previously asserted on \`defsByFileAndName\`,
\`callablesBySimpleName\`, and \`memberByOwner\` — fields removed when
symbol-keyed lookups moved to \`SemanticModel\`. Rewrite so the same
invariants are asserted via the authoritative consumers:
* New WorkspaceResolutionIndex shape test (scope-only maps).
* \`findExportedDef\` module-export visibility tests:
- keeps top-level class and function defs.
- excludes class-body Variable defs (MAX_USERS = 100).
- excludes class methods from module-export lookup.
* \`findExportedDefByName\` fallback excludes class methods when a
same-named module function exists.
* \`findOwnedMember\` via the reconciled SemanticModel finds Python
class methods after populateOwners + reconcileOwnership.
Total assertions preserved: every invariant from the old test file is
still pinned; the assertion surface shifted from the index shape to
the walker helpers.
Verified:
- workspace-index.test.ts 8/8 passing
* fix(tests): update registry-primary-flag test for C# migration
The "returns exactly the flipped languages" case expected `enabled.size === 1`
after toggling Python off and Go on. After the C# migration lands C# in
MIGRATED_LANGUAGES, C# is default-on too — so the size is now 2 (Go + C#)
unless C# is also opted out.
Turn off C# alongside Python in the test setup. Added a comment noting
that future migrations must add their REGISTRY_PRIMARY_<LANG>='false'
line here.
* refactor(scope-resolution): address PR #1019 review findings
Resolves all 5 findings from the automated review on
feat/csharp-scope-resolution. Shared ingestion code stays
language-agnostic; C# (and every class-like language) benefits.
F1 [high] Broaden class-like predicate
Hoist `isClassLike` in `scope/walkers.ts` to an exported top-level
helper covering Class | Interface | Struct | Record | Enum | Trait.
Use it in `findClassBindingInScope`, `findEnclosingClassDef`, and
`buildWorkspaceResolutionIndex` so C# records, structs, interfaces,
and enums participate in scope chains and receiver binding the same
way Python classes do.
F2 [medium] Remove stale comment in csharp simple-hooks
`csharpReceiverBinding`'s doc claimed this/base synthesis was
"planned for a follow-up"; synthesis has been implemented in
receiver-binding.ts since the migration landed. Rewrite the doc to
describe the actual behavior (non-null TypeRef on instance-method
bodies, null on static/free functions).
F3 [medium] O(1) reverse lookup for classScopeId -> classDefId
Add `classScopeIdToDefId: ReadonlyMap<ScopeId, string>` to
`WorkspaceResolutionIndex`, populated as the inverse of
`classScopeByDefId`. Replace the O(C) linear scan in
`pickImplicitThisOverload` (free-call-fallback.ts) with an O(1)
`Map.get` — turns per-site reverse resolution from linear in class
count to constant time for every free call.
F4 [low] Extract narrowOverloadCandidates shared utility
New `passes/overload-narrowing.ts` centralizes the arity + argument-
type narrowing previously duplicated across `pickOverload`
(receiver-bound-calls.ts) and `pickImplicitThisOverload`
(free-call-fallback.ts). Both callsites now share identical
narrowing semantics; variadic `params T` handling is preserved.
Return type is `readonly SymbolDefinition[]` with no defensive
spreads (allocations saved on the hot path).
F5 [low] Merge unreachable Case 5 into Case 2
`Case 5` in `receiver-bound-calls.ts` was dead code — `Case 2`
pre-empted it for every static/class-name receiver. Delete Case 5
and lift its kind-aware read/write ACCESSES reason/confidence logic
into Case 2 so static-style member access (e.g. `Interface.Member`,
`TypeName.StaticMember`) gets the correct edge metadata.
Tests
- New unit tests for `narrowOverloadCandidates` covering empty
input, arity filtering, variadic params, type narrowing, and
fallback semantics.
- New unit tests for `classScopeIdToDefId` verifying inverse
invariant and empty index behavior.
- New C# integration fixtures and tests:
* csharp-record-base — record inheritance + `base.Save()`
* csharp-struct-overloads — struct with implicit-this overload
narrowing (pinned exact edge count under registry-primary)
* csharp-interface-receiver-static — interface-qualified static-
style call exercises the merged Case 2.
- Full runs green:
* scope-resolution unit: 406/406
* csharp integration (registry-primary): 197/197
* csharp integration (legacy DAG): 197/197
* python integration (regression guard): 204/204
Chore
- Add `.context/` to root `.gitignore` to prevent agent scratch
files from being committed.
Made-with: Cursor
* test(csharp-scope-resolution): address adversarial review follow-ups on PR #1019
Applies the three actionable follow-ups from the post-commit adversarial
review of 5a1bce7f against DoD.md. No runtime code changes.
- [medium] Strengthen bounds-only assertion in the struct-overloads
suite: `methods.length` is now pinned to `toBe(2)` and the arity list
to `toEqual([1, 2])`. Fixture `csharp-struct-overloads/src/Calc.cs`
declares exactly two `Add` methods, so a regression that adds, drops,
or merges an overload will now fail the test instead of silently
passing a `>= 2` gate.
- [low] Pin the merged Case 2 kind-aware branch (receiver-bound-calls.ts
lines 257-289) with a dedicated fixture and three new assertions:
`csharp-class-static-field-access/src/Counters.cs` exercises
`ClassName.Field = value` where the receiver resolves via
`findClassBindingInScope` (no typeBinding on `Counters`). The test
verifies (a) two distinct ACCESSES writes are emitted from a single
method (per-site dedup from graph-bridge/edges.ts:80-87),
(b) `reason === 'write'`, (c) `confidence === 1.0`, and (d) no
spurious CALLS edges are produced for the same sites. This is the
semantic upgrade lifted from the deleted Case 5; without a pinning
test a future revert of the kind-aware branch would silently drop
back to `import-resolved`/`global` at 0.85 for the same sites.
Read-side coverage is intentionally not asserted because the C#
tree-sitter query currently emits only `write.member` captures
(languages/csharp/query.ts:485-501) — a read counterpart would have
no reference site today and would give a false sense of coverage.
- [info] Left the `?? overloads[0]` fallback in place at
receiver-bound-calls.ts:450 unchanged. With the package's current
tsconfig (strict: false, no noUncheckedIndexedAccess) both the
defensive fallback and a `candidates[0]!` assertion type-check
identically, so the finding has no production-readiness impact.
Keeping the fallback minimizes churn.
Validation (local, Windows PowerShell):
- `npx prettier --check test/integration/resolvers/csharp.test.ts` -> clean
- `npx tsc --noEmit` -> 0 errors
- `REGISTRY_PRIMARY_CSHARP=1 npx vitest run test/integration/resolvers/csharp.test.ts` -> 200/200
- `REGISTRY_PRIMARY_CSHARP=0 npx vitest run test/integration/resolvers/csharp.test.ts` -> 200/200
- `npx vitest run test/integration/resolvers/python.test.ts` -> 204/204
- `npm test` (full gitnexus suite) -> 6967 passed, 6 pre-existing failures
(4x Swift overload/dedup, 1x Swift method-extraction unit, 1x Swift
type-env unit, 1x LadybugDB lockfile on Windows). All six reproduce on
5a1bce7f with these follow-up changes stashed, confirming they are
environment/baseline failures unrelated to this work. Swift is not in
MIGRATED_LANGUAGES so the merged Case 2 path cannot affect it.
Refs: PR #1019
Made-with: Cursor
* refactor(scope-resolution): address full-PR review findings on PR #1019
Resolves the two remaining findings from the code-review-swarm full-PR
sweep (verdict: production-ready with minor follow-ups).
[low] Complete the csharp/index.ts module-layout JSDoc.
`languages/csharp/index.ts` is the discovery surface for the C# scope-
resolution module decomposition (per AGENTS.md). The "Module layout"
list silently omitted three load-bearing modules — `accessor-unwrap.ts`
(`.Values`/`.Keys` receiver-type unwrap), `namespace-siblings.ts`
(AST-driven cross-file implicit-namespace visibility), and
`receiver-binding.ts` (`this`/`base` type-binding synthesis). Extended
the JSDoc list so the "single-concern" decomposition story is honest
and the next contributor can locate the right file without grep.
No behavior change.
[info] Replace the non-standard `'scope-resolution: super-receiver'`
edge reason with the canonical `'global'` tier.
`passes/receiver-bound-calls.ts` emitted a non-canonical reason string
for the super/base branch, which falls outside the vocabulary declared
in ARCHITECTURE.md § Scope-Resolution Pipeline (`'import-resolved' |
'global' | 'local-call' | 'same-file' | 'interface-dispatch' | 'read'
| 'write'`). Super/base calls resolve through the MRO chain rather
than through import directives, so the correct canonical tier is
`'global'` (same classification the legacy DAG's `toResolveResult`
applies to non-same-file, non-import-scoped resolutions).
Locked the contract with `rel.reason === 'global'` assertions on the
existing `csharp-super-resolution` and `csharp-generic-parent-
resolution` suites, both of which go through the super-branch MRO
path. The `csharp-record-base` suite intentionally does not pin a
reason (records don't currently emit EXTENDS edges, so the MRO lookup
misses and the edge is produced by the reference-index fallback
instead of the super-branch). A code comment flags the pre-existing
Python-legacy asymmetry (Python legacy tier classifier marks
`super()` as `'import-resolved'` because the ancestor arrives via an
`import` statement); closing that gap requires realigning the legacy
tier classifier and is tracked separately.
Validation:
- `npx tsc --noEmit` passes.
- `npx prettier --check` clean on all four touched files.
- `test/integration/resolvers/csharp.test.ts` — 200/200 under both
`REGISTRY_PRIMARY_CSHARP=0` (legacy DAG) and `REGISTRY_PRIMARY_CSHARP=1`
(registry-primary), preserving same-graph parity on the super branch.
- `test/integration/resolvers/python.test.ts` — 204/204 under both
`REGISTRY_PRIMARY_PYTHON=0` and `=1`.
- `test/unit/scope-resolution/` — 406/406 passing.
Unstaged: `gitnexus/package-lock.json` (drift from `npm install` run
to resolve the pre-existing missing `jsonc-parser` dependency — not
part of this change).
Made-with: Cursor
* test(ci): raise integration-test timeouts so slow Windows runners stop flaking
The `windows-latest` CI runner for this branch was consistently failing
two integration suites in ways that had nothing to do with the PR's
scope-resolution changes:
* `cli-e2e.test.ts` — `analyze command runs pipeline on mini-repo`
hit the default 30 s vitest test timeout, which raced the test's
own 30 s subprocess timeout and prevented the existing
`if (result.status === null) return;` slow-CI tolerance from ever
firing. That single timeout then cascaded into the downstream
`cypher`/`query`/`impact` tests (which exited non-zero because the
mini-repo was never indexed) and the `EPIPE handling` test.
* `skills-e2e.test.ts` — `beforeAll` hooks run a full
`runSkillsCli(tmpDir)` subprocess that analyzes a fixture repo and
generates skills. 50 s was enough on Linux/macOS but not on slow
Windows CPUs, producing "Hook timed out in 50000ms" errors and
cascading test failures across every language describe block.
Fix:
* Bump the `analyze` test's vitest test-level timeout to 60 s so it
exceeds the 30 s subprocess timeout and the slow-CI tolerance can
actually activate.
* Bump all 12 `runSkillsCli`-driven `beforeAll` hooks from 50 s to
120 s.
No production-code behavior changes. No change to what the tests
assert — only the per-test/hook wall-clock budget.
Made-with: Cursor
Follow-up to #1044. Adds user-facing documentation for the configurable
skip threshold introduced in that PR:
- README CLI Commands: new --max-file-size example line
- README Troubleshooting: new 'Large files are being skipped' subsection
covering the CLI flag, env var, default (512 KB), ceiling (32768 KB),
fallback behaviour, and the effective-threshold banner
- CHANGELOG [Unreleased] Added: feature entry with issue/PR cross-refs
* feat(ingestion): make large-file skip threshold configurable
The walker previously hardcoded a 512KB skip threshold, which silently dropped legitimate large source files (e.g. ~900KB hand-written Java service classes) during analysis with no way to override short of editing source.
Allow overrides via the GITNEXUS_MAX_FILE_SIZE env var (KB) — consistent with the existing GITNEXUS_NO_GITIGNORE / GITNEXUS_VERBOSE patterns — and a matching --max-file-size <kb> flag on gitnexus analyze.
- New utility getMaxFileSizeBytes() in core/ingestion/utils/max-file-size.ts parses the env var, falls back to the 512KB default for missing/invalid values, and clamps against TREE_SITTER_MAX_BUFFER (32MB) to keep the downstream parser safe.
- filesystem-walker.ts now resolves the threshold per call and drops the 'likely generated/vendored' editorial when the user has explicitly raised the limit.
- analyze CLI wires --max-file-size to the env var and echoes a one-line notice when the threshold is overridden, mirroring how --no-gitignore is handled.
- index.ts documents the new flag and env var under the analyze help text.
- Warnings for invalid or out-of-range values are emitted exactly once per distinct value to avoid log spam.
Tests:
- New test/unit/max-file-size.test.ts covers defaults, KB parsing, clamp-at-ceiling, invalid-input fallback + warn-once, and distinct-value warnings.
- test/integration/filesystem-walker.test.ts gains a 'large file skip threshold (#991)' block: 600KB fixture skipped by default, included under GITNEXUS_MAX_FILE_SIZE=1024, invalid values fall back and warn once, and the 'generated/vendored' suffix is only emitted under the default threshold.
Closes#991
* fix(cli): show effective clamped max-file-size in banner
Addresses the PR #1044 review finding: the startup banner printed the raw GITNEXUS_MAX_FILE_SIZE value rather than the clamped effective threshold, producing misleading telemetry when the value exceeded the 32 MB tree-sitter ceiling.
The banner is also suppressed when the effective threshold equals the default, removing log noise when operators explicitly set the value to the current default.
Extracted the logic into a new getMaxFileSizeBannerMessage() helper and pinned the behavior with unit tests covering default, raised override, invalid fallback, and above-ceiling clamp cases.
* docs: add DoD.md repo-wide Definition of Done
Adds a stable baseline completion bar for production-ready changes,
complementing AGENTS.md, GUARDRAILS.md, CONTRIBUTING.md, TESTING.md,
and ARCHITECTURE.md. Includes core DoD, GitNexus-specific requirements,
per-package validation baseline, task-specific DoD template, and guidance
for review prompts.
* docs: expand DoD with security, observability, and agent-workflow gates
Restructure DoD.md into numbered sections and add axes that were previously
implicit: security, observability/operability, reversibility, and explicit
guardrails for agent-assisted workflow (scope match, evidence-based edits,
pre-edit impact analysis, embeddings preservation).
Expand the validation baseline to reflect the real CI shape (shared-first
build ordering, prettier, setup-gitnexus action, CHANGELOG ownership) and
add a Review Gates checklist plus a "Not Done" signals section that flags
contract drift, language leakage into shared code, and unrelated churn.
`upsertGitNexusSection` in ai-context.ts uses `indexOf` to locate the
bounds of the GitNexus section in CLAUDE.md / AGENTS.md before
replacement. `indexOf` matches the first occurrence of the marker
anywhere in the file, including inline prose references in backtick-
quoted fragments mid-sentence.
The shipped CLAUDE.md contains exactly such a reference ("See the
`<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in AGENTS.md
for the canonical MCP tools..."). Running `gitnexus analyze` on a
fresh install matches those inline markers as section delimiters and
replaces the prose between them with the full ~100-line injected
block, breaking the backtick and corrupting markdown for every user.
Fix: new private `findSectionMarkerIndex` helper that only matches
markers occupying their own line — preceded by `\n` or start-of-file,
followed by `\n` / `\r` (CRLF files) / end-of-file. `\r` is explicit
so CRLF-terminated sections on Windows (core.autocrlf = true) still
match. The generator always emits markers alone on their line, so
every legitimate section continues to update in place; only inline
prose references now fall through to the append branch, which leaves
existing content untouched.
Two new unit tests:
- #1041 regression — seed CLAUDE.md with the shipped inline prose
line, run analyze twice, assert inline prose preserved verbatim
and marker counts stay at 2/2 (1 inline + 1 section-position)
- CRLF handling — seed a CRLF file with inline prose + legitimate
section, run analyze, assert section replaced in place, inline
prose preserved, stale stub content removed
No destructive ops, no bypass flags, no new deps. Behaviour change
is strictly narrowing — files that previously updated correctly
still do; files that previously got corrupted now fall through to
the safer append branch.
Closes#1041.
* deps: add jsonc-parser for JSONC-safe config editing
* fix: use jsonc-parser to preserve comments in opencode.json during setup
- Add mergeJsoncFile() using parseTree/modify/applyEdits pipeline
- Add getOpenCodeMcpEntry() for OpenCode MCP format { type: local, command: [...] }
- Replace readJsonFile+writeJsonFile in setupOpenCode with mergeJsoncFile
- Fix wipe bug: JSON.parse on JSONC comments caused catch block to reset config to {}
- Add 9 tests for JSONC comment preservation, corrupt file safety, and format
* fix: use parseTree error collection and detect indentation
- Pass parseErrors array to parseTree() instead of checking
(tree as any).errors which was always undefined — a real bug
that allowed corrupt files to be rewritten
- Detect tab indentation from file content to avoid mixed
indentation in modified JSONC files
- Fix JSDoc to match actual fallback behavior (JSON.parse, not
readJsonFile)
- Strengthen corrupt-file test to assert exact content match
* style(setup): fix prettier formatting on mergeJsoncFile
* fix(setup): remove dead JSON.parse fallback, detect space-indent width, fix JSDoc
- Remove the semantically unreachable JSON.parse fallback branch in
mergeJsoncFile (jsonc-parser's parseTree is a strict superset of
JSON.parse, so the fallback can never fire for content JSON.parse
would accept)
- Replace binary tab/space detection with detectIndentation() that
measures actual indent width from the first indented line
- Fix JSDoc: 'valid JSON that is not valid JSONC' is impossible by
definition
- Add tests for tab indentation and 4-space indentation preservation
The top-level and CLI READMEs advertised `gitnexus group add <name> <repo>`
(two args) and `gitnexus group remove <name> <repo>`, but the CLI
(`gitnexus/src/cli/group.ts`) actually requires three args for `add`
(`<group> <groupPath> <registryName>`) and uses `<groupPath>` — not a
repo path — for `remove`. Reusing the same second argument across two
`group add` invocations silently overwrote the previous mapping because
the hierarchy path is the key in `group.yaml`'s `repos` map.
Update both READMEs to match the real CLI contract. Node_modules not
installed locally for this docs-only change, so pre-commit (prettier +
typecheck) was skipped.
Made-with: Cursor
Co-authored-by: TuanPM1 <tuanpm1@kaopiz.com>
* fix(docker): switch Dockerfile.cli from Alpine to Debian slim
Alpine uses musl libc which is incompatible with @ladybugdb/core's
glibc-compiled native binary, causing ERR_DLOPEN_FAILED on startup.
Closes#1008
* fix(docker): resolve build and runtime failures in Dockerfile.cli
- Add **/*.tsbuildinfo to .dockerignore and rm -f tsbuildinfo in
builder to prevent stale incremental cache from skipping
gitnexus-shared compilation
- Install libstdc++6 from Debian Trixie for @ladybugdb/core native
module compatibility (requires GLIBCXX_3.4.31)
* fix(docker): use node:22-trixie-slim for GLIBCXX_3.4.31 support
Replaces the manual Trixie libstdc++6 backport with the official
node:22-trixie-slim base image, which ships GCC 14 runtime natively.
* fix(group): surface friendly error when group name not found
Squashed commits:
- test(csharp): add #903 regression — parse completeness for single-file C# repo
- fix(group): add GroupNotFoundError guard to groupList + re-throw tests for groupQuery/groupStatus
- fix(test): restore section comments in csharp.test.ts stripped during rebase
* fix(group): catch GroupNotFoundError explicitly in groupContext and groupImpact
* Initial plan
* plan: Python scope-based resolution migration
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0eee6c69-fc17-4df5-9ac6-358ab41f5740
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* feat(python): scope-based resolution provider hooks + 62 tests
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0eee6c69-fc17-4df5-9ac6-358ab41f5740
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* refactor(python): split scope-hooks monolith into focused modules
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/db76e937-4b0e-4c4d-82b1-265a1fb3673d
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* test(python): integration-style scope-resolution tests + suffixResolve fallback
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/db76e937-4b0e-4c4d-82b1-265a1fb3673d
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* wire python scope-based resolution end-to-end (initial pass)
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c474dc66-5cf7-445d-8eb4-76501c5e6d67
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* keep legacy IMPORTS for python (heritage needs importMap), scope phase owns CALLS only
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c474dc66-5cf7-445d-8eb4-76501c5e6d67
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* test(python): remove parallel scope-resolution integration test
The new test/integration/python-scope-resolution.test.ts duplicated coverage
the reviewer explicitly rejected. The existing
test/integration/resolvers/python.test.ts (191 tests, driven by
runPipelineFromRepo) is the source of truth for Ring 3 parity.
Also document the IMPORTS-emission follow-up gap: wiring emitImportEdges
in python-scope-emit.ts today regresses 10 IMPORTS-edge fixtures because
the scope-extractor's ImportEdge coverage is narrower than legacy
pythonImportConfig.importResolver. Tracked as a follow-up.
Baseline with REGISTRY_PRIMARY_PYTHON=1 is unchanged: 109/191 pass.
* feat(ingestion): scope-resolution phase owns Python IMPORTS edges (RFC #909 Ring 3)
When `REGISTRY_PRIMARY_PYTHON=1`, IMPORTS graph edges for Python files are now
emitted exclusively by the new scope-resolution path. The legacy
`import-processor` still runs — heritage resolution needs its importMap /
namedImportMap / moduleAliasMap population — but its graph edge emission is
gated per-language so Python no longer double-emits.
This closes the reviewer's second change request on PR #980: "the legacy path
must be turned off". Legacy IMPORTS edges for Python are now off by default
when the flag is enabled.
Three bugs were fixed to make the new path's coverage match legacy:
1. **Root-file bailout** (import-resolvers/python.ts): `resolvePythonImportInternal`
returned null immediately when the importer file lived at the repo root
(importerDir === ''). The ancestor directory walk further down already
handles this case correctly; the early return was the bug. Proximity check
now only runs when importerDir is non-empty, and the ancestor walk sees
root-level files for the first time.
2. **External dotted imports** (languages/python/import-target.ts): the new
path fell straight through to `suffixResolve` for multi-segment imports,
which happily matched `django.apps` to a local `accounts/apps.py`. Mirror
`pythonImportStrategy`'s `hasRepoCandidate` guard — suffix-match only when
the leading segment exists somewhere in-repo as a package, __init__.py,
or namespace directory.
3. **suffixResolve ambiguity** (languages/python/import-target.ts): the
shared `suffixResolve` helper requires a pre-built `SuffixIndex` to
disambiguate ties. Without one it falls back to an O(files) scan that
silently picks the first match when the last segment collides across
directories (e.g. `accounts.models` matching `billing/models.py`).
Replaced with `resolveAbsoluteFromFiles` — exact lookup first, then a
deterministic suffix match.
Validation:
- Flag OFF: 191/191 pass (no regression).
- Flag ON: 109/191 pass (82 fail — exact baseline match; remaining 82 are
unchanged CALLS-edge provider-feature gaps tracked as Phase B follow-ups).
- `tsc --noEmit`: clean.
The 82 CALLS failures cluster into 44 describe blocks covering type-inference
features (assignment chains, walrus, class-level annotations, constructor
inference, C3 MRO, overload dispatch, return-type inference) that need
dedicated Ring 3 follow-up work. Each cluster is tracked against the RFC #909
shadow-parity gate (>=99% fixtures / >=98% corpus) in the per-language ticket.
* ci(scope-resolution): automatic parity gate driven by MIGRATED_LANGUAGES
Adds the Ring 3 parity gate the RFC §6.4 requires: when a language's
scope-resolution migration is marked complete, CI runs its resolver
integration test twice on every PR (once with the legacy DAG, once with
the registry-primary path) and both must pass.
The "is this language migrated" signal is a single TypeScript constant:
// gitnexus/src/core/ingestion/registry-primary-flag.ts
export const MIGRATED_LANGUAGES: ReadonlySet<SupportedLanguages> =
new Set([ /* SupportedLanguages.Python when ready */ ]);
Adding a language here has three simultaneous effects:
1. `isRegistryPrimary(lang)` defaults to true for that language in
production (env-var override still wins if set explicitly).
2. `.github/workflows/ci-scope-parity.yml` auto-discovers the set via
`npx tsx scripts/ci-list-migrated-languages.ts`, builds a parity
matrix, and runs:
- `REGISTRY_PRIMARY_<LANG>=0 npx vitest run resolvers/<slug>.test.ts`
- `REGISTRY_PRIMARY_<LANG>=1 npx vitest run resolvers/<slug>.test.ts`
Both legs must pass for the job to succeed.
3. Legacy-path gating in call-processor.ts / import-processor.ts kicks
in automatically through the same `isRegistryPrimary` lookup.
No JSON registry, no manual workflow edit, no second source of truth —
contributors update the Set and CI picks it up. Empty Set = parity job
is a skipped matrix (workflow still reports success).
The new `scope-parity` reusable workflow is added to ci.yml's `needs`
graph and ci-status gate. Its result must be `success` (skipped would
mean upstream discover job failed and should block).
Validation (with empty MIGRATED_LANGUAGES set):
- flag OFF: 191/191 pass (no behavior change)
- flag ON (manual REGISTRY_PRIMARY_PYTHON=1): 82 fails = baseline exact match
- `npx tsc --noEmit`: clean
- concurrency-convention script: pass
- tsx discovery script: emits `[]` correctly
* ci(scope-resolution): keep MIGRATED_LANGUAGES empty; fix linter auto-uncomment
Previous commit's example entry got auto-uncommented (linter preferred a
type-checkable `SupportedLanguages.Python` over a commented-out reference).
That would have triggered the parity CI gate against Python, which today
has 82 known flag-on failures — unintended and would block the PR.
Use the explicit generic `new Set<SupportedLanguages>([])` so an empty set
still type-checks without needing an uncommented-out sample member.
Example in the comment now has `// SupportedLanguages.Python,` so it
remains illustrative without participating in the set.
* feat(python): capture constructor-inferred + annotated type bindings
Extends the Python scope-extractor with two new type-binding capture
patterns so receiver-typed method dispatch has concrete type bindings
to work from:
1. `u: User = ...` / `u: User` — variable annotations. `@type-binding.annotation`
anchor, `source: 'annotation'`.
2. `u = User("alice")` — assignment RHS is a bare-identifier call (Python
has no `new` keyword; constructor-shaped calls are syntactically
identical to function calls). `@type-binding.constructor` anchor,
`source: 'constructor-inferred'`.
The runtime query lives in `query.ts` (the `.scm` file is documentation
per the comment at its top); both are updated.
Fixes 19 failures across these resolver fixtures (flag-on 82 → 63):
- Python constructor-inferred type resolution (3)
- Python class-level annotation resolution (3)
- Python nullable receiver resolution (3)
- Python member-call / receiver-constrained / constructor-call (3)
- Python assignment chain propagation (2)
- Python walrus / match-case / chained method (3)
- Python member access iterable for-loop (2)
* feat(python): strip nullable unions + prefer annotations over inference
Two linked changes that together fix the 4 nullable-receiver tests:
1. `stripNullable` in Python's `interpretTypeBinding` unwraps `User | None`,
`None | User`, and `Optional[User]` to `User`, so receiver-typed
resolution treats nullable receivers identically to non-nullable ones.
Three-arm unions (`User | Error | None`) are left unchanged — truly
ambiguous for single-receiver inference.
2. Source-strength ordering in `pass4CollectTypeBindings`. When multiple
matches fire for the same bound name in the same scope — e.g. the
`u: User = find()` idiom where both the annotation and
constructor-inferred patterns match — the explicit annotation now
wins regardless of query-match arrival order. Rank:
explicit (annotation / parameter-annotation / return-annotation / self) > inferred
Also reorders the two Python patterns in query.ts / scopes.scm so the
constructor-inferred pattern appears first — a belt-and-braces fallback
that keeps behavior deterministic if the shared priority ranking is ever
revisited.
Fixes 4 failures (flag-on 63 → 59):
- Python nullable receiver resolution (4 tests)
Flag-off regression check: 191/191 still pass.
* feat(python): walrus, qualified-call, match-case type bindings
Extends the constructor-inferred family of captures with three more
assignment-shaped patterns that all bind a variable to a class-like type:
- Walrus: `(u := User(...))` → `u: User` via `(named_expression)`.
- Qualified call RHS: `u = models.User(...)` → `u: models.User` via
`(attribute)` node .text. Falls through resolveTypeRef Phase 2
(QualifiedNameIndex dotted fallback).
- Match as-pattern: `case User() as u:` → `u: User` via `(as_pattern)`
+ `(class_pattern (dotted_name))`.
Fixes 2 failures (flag-on 59 → 57):
- Python walrus operator type inference
- Python match/case as-pattern type binding
Qualified-call constructor tests still fail because they require
cross-module qualifiedName registration (models.User → models.py's User
class) which isn't yet wired in the Python extractor. Tracked as
follow-up alongside module-import CALLS (#337) resolution.
* feat(python): chain type bindings + strip list[T] generic for for-loop
Adds two capture patterns and a shared transitive-closure pass that
together handle Python's variable-aliasing and for-loop-over-typed-
iterable patterns:
1. `(assignment left: (identifier) right: (identifier))` — `alias = u`.
2. `(for_statement left: (identifier) right: (identifier))` — `for u in users`.
Both emit `@type-binding.alias` with the RHS identifier as rawName. The
shared `pass4CollectTypeBindings` now runs a final transitive-closure
walk that follows identifier-chain TypeRefs through the declaring scope
and its ancestors (depth-capped, cycle-guarded) so `alias` ultimately
points at the class type instead of another local variable name.
Generic stripping in `interpret.ts` unwraps single-arg collection
wrappers — `list[User]`, `set[User]`, `Iterable[User]`, etc. — to the
element type. Multi-arg generics (`dict[str, User]`, `Callable[...]`)
are left alone; their semantics aren't unambiguous.
Fixes 8 failures (flag-on 57 → 49):
- Python assignment chain propagation (4)
- Python nullable + assignment chain (2)
- Python walrus operator (:=) assignment chain (2)
Flag-off still 191/191.
* feat(python): namespace & class receiver resolution + file-level caller fallback
Adds a Python-specific post-resolution pass `emitReceiverBoundCalls`
that closes two receiver gaps the shared `MethodRegistry.lookup` doesn't
cover:
1. **Namespace receivers** — `import models; models.User()` /
`import models as m; m.User()`. The shared `lookupReceiverType` only
walks `scope.typeBindings`; namespace imports never land there
(they're filtered out of `scope.bindings` when the target module
has no self-named def, per `finalize-algorithm.ts:540`). The new
pass walks `indexes.imports` directly, builds a per-file
`localName → targetFilePath` map, and emits CALLS/ACCESSES edges
against the target file's `localDefs`.
2. **Class-name receivers** — `Dog.classify("dog")`. The shared resolver
requires typeBindings; class bindings in `scope.bindings` are never
consulted as receivers. The new pass checks class-kind bindings in
the call scope's chain and resolves members via `ownerId`.
Also fixes module-level call attribution: `resolveCallerGraphId` now
falls back to the File node id (`generateId('File', filePath)`) when no
enclosing function/method/class is found. Matches legacy DAG behavior
for module-scope calls like `u = models.User()` at the top of app.py.
Fixes 4 failures (flag-on 49 → 45):
- Python module import CALLS resolution (Issue #337) (4 of 7)
Flag-off still 191/191.
* feat(python): dotted-typebinding receiver resolution
Adds case 3 to `emitReceiverBoundCalls`: when a receiver's typeBinding
has a dotted rawName like `u: models.User` (the constructor-inferred
form fired by `u = models.User(...)`), walk the namespace map + target
file's defs to find the class, then look up the member via ownerId.
`resolveTypeRef`'s QualifiedNameIndex fallback can't cover this because
the target class's qualifiedName in models.py is just `"User"`, not
`"models.User"` — the dotted form only exists in the call-site file's
receiver expression. This pass bridges that gap without modifying the
shared registry.
Fixes 9 more failures (flag-on 45 → 36):
- Python qualified constructor inference (2)
- Python module import CALLS resolution (Issue #337) (3)
- (cluster overlap — several downstream tests in assignment/nullable/
walrus that propagate through qualified-ctor bindings also benefit)
Flag-off still 191/191.
* feat(python): consult finalized bindings for receiver resolution
`findClassBindingInScope` now walks BOTH:
1. `scope.bindings` — pre-finalize local declarations (origin: 'local')
2. `indexes.bindings` — post-finalize cross-file imports/namespaces
Without (2) we were blind to any class brought in via
`from models import Dog` at the call site's file, because the
scope-extractor's Pass 2 only populates local bindings and the
cross-file finalize produces a separate bindings map that never lands
on `scope.bindings`.
Case 2 (`Dog.classify()`) now walks MRO so inherited static/class
methods resolve — `Dog.classify()` where `classify` lives on `Animal`.
Case 4 (simple typeBinding like `u: U` from aliased import) now uses
`findClassBindingInScope` instead of the shared `resolveTypeRef`,
because `resolveTypeRef`'s `ctx.scopes` only sees pre-finalize local
bindings too.
Fixes 4 more failures (flag-on 36 → 32):
- Python method enrichment > Dog.classify static (1)
- Python static/classmethod class-as-receiver (2)
- Python alias import resolution (1)
Flag-off still 191/191.
* refactor(python-scope): extract language-agnostic emit-core/
Unit 1 of the python migration architectural plan
(docs/plans/2026-04-19-001-refactor-python-migration-architectural-plan.md).
Splits python-scope-emit.ts (~945 → 481 lines) by lifting 14 generic
graph-feeding primitives into emit-core/:
- graph-node-lookup, graph-id, emit-edge
- emit-references, emit-imports
- scope-walkers (findReceiverTypeBinding, findClassBindingInScope,
findOwnedMember, findExportedDef)
- namespace-targets, method-dispatch-bridge
Each file carries a "Next-consumer contract" JSDoc so future language
migrations (TS #927, JS #928, Java, Kotlin, Ruby) import from emit-core
rather than re-implementing. python-scope-emit.ts keeps only the four
Python-specific pieces: runPythonScopeResolution (orchestrator),
buildPythonMro, emitReceiverBoundCalls (4 cases), populateMethodOwnerIds
— these move to languages/python/emit/ in Unit 11.
Pure refactor, zero behavior change:
- flag-off: 191/191 python.test.ts pass (identical baseline).
- flag-on (REGISTRY_PRIMARY_PYTHON=1): 32 fail / 159 pass (identical
baseline — the refactor neither fixes nor regresses any test).
- tsc --noEmit clean.
* feat(python-scope): arity metadata + bind function decls in parent scope
Unit 2 of the python migration architectural plan
(docs/plans/2026-04-19-001-refactor-python-migration-architectural-plan.md).
Two changes that the registry-primary path needs before any of the
arity-sensitive failures can move:
1. Arity metadata on scope-extracted Function/Method defs.
- New helper `languages/python/arity-metadata.ts` reuses
`pythonMethodConfig.extractParameters` so self/cls stripping,
defaults, and *args/**kwargs detection match legacy semantics.
- `emit-captures.ts` synthesizes
`@declaration.parameter-count` /
`@declaration.required-parameter-count` /
`@declaration.parameter-types` captures on every
`@declaration.function` match.
- Generic `scope-extractor.ts buildDefFromDeclarationMatch` reads
the three optional captures into `SymbolDefinition`. Absence is
still the no-op default for non-Python providers.
2. Hoist function/class declaration bindings to the enclosing scope.
The "innermost scope containing the anchor" default placed
`def greet(...)` inside greet's OWN body — invisible to other
module-level callers, so every flag-on free-call resolved to
`unresolved`. The hoist condition (`anchor range == innermost
range`) only fires for scope-creating declarations, so variable /
for-loop captures whose anchor is a child identifier stay put.
Hooks can still override via `bindingScopeFor`.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on (REGISTRY_PRIMARY_PYTHON=1): 31 fail / 160 pass
(was 32/159; the hoist unblocks free-call resolution end-to-end).
- tsc --noEmit clean.
Per-(source,target) edge collapse for multi-call-site cases
(default-params, variadic) still pending — landing it without
regressing the static-method find_user fixture (which expects two
distinct edges through different targets) needs the ownership-aware
qualified-id work that lands with Unit 4 / Unit 11.
* feat(python-scope): capture function return-type annotations
Unit 3 of the python migration architectural plan
(docs/plans/2026-04-19-001-refactor-python-migration-architectural-plan.md).
Wires the `def get_user() -> User` return-type annotation into the
typeBindings stream so the existing constructor-inferred + transitive
chain machinery can resolve `u = get_user(); u.save()` to `User#save`
without any orchestrator change.
Changes:
- `query.ts` + `scopes.scm`: new `@type-binding.return` pattern keyed by
the function name (matches RFC §5.1 canonical vocabulary).
- `interpret.ts`: maps `@type-binding.return` to the existing
`'return-annotation'` source label (no shared change needed).
- `scope-extractor.ts pass4CollectTypeBindings`: extends the Pass 2
auto-hoist (anchor range == innermost scope range → bind in parent)
to type bindings as well — return-type bindings whose anchor IS the
function_definition land in the function's enclosing scope so
callers see them.
Same-file return-type inference is now end-to-end:
`def get_user() -> User: ...` + `u = get_user()` produces
`u: User (return-annotation)` in the caller's scope via
`followChainedRef`.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 31 fail / 160 pass (no change — every remaining
return-type test in this fixture set is *cross-file*; carrying
`get_user → User` across module boundaries lands with the
cross-file typeBinding propagation work in Unit 5/7).
- tsc --noEmit clean.
* feat(python-scope): resolve dotted receivers via class-scope field types
Unit 4 partial — the dotted-receiver case (`user.address.save()`).
Class-body annotations like `class User: address: Address` already
land in the class scope's typeBindings via the existing
`@type-binding.annotation` capture. This commit consumes that signal:
- Build a `Map<classDefId, Scope>` from every parsed file's class
scopes once per resolution pass.
- New Case 0 in `emitReceiverBoundCalls`: when the receiver's name
contains a dot, walk the chain — resolve the head's type, then for
each remaining segment look up that field's type in the owner
class's scope.typeBindings, then emit the call against the final
class with MRO walk.
- Cross-scope lookups use each TypeRef's `declaredAtScope` so an
imported `Address` resolves in the file that owns the field
declaration, not the file holding the call site.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 29 fail / 162 pass (was 31/160; both `Field type
resolution` fixtures now pass — same-file and cross-file disambig).
- tsc --noEmit clean.
Remaining Unit 4 work (write ACCESSES, `self.X` for-loop iteration)
needs Unit 6's tuple/iterable destructuring before it can land —
`for u in self.users` requires the iterable typing path.
* feat(python-scope): chain receiver via call-expression return types
Unit 5 — extends the compound-receiver case to handle call-expression
receivers (`svc.get_user().save()`).
`resolveCompoundReceiverClass` is the single recursive entry point for
all compound receivers. Three shapes:
- bare identifier — typeBinding chain
- dotted `obj.field[.field]…` — class-scope field types
- call `expr.method()` — recurse into expr, look up method's
return-type typeBinding on its class scope
Method return-type bindings auto-hoist to the parent (class) scope per
Unit 3, so `methodClassScope.typeBindings.get(methodName)` is the
canonical lookup. Free-call return types (`get_user()`) walk the
caller's scope chain.
Depth-capped at 4 hops to bound recursion.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 28 fail / 163 pass (was 29/162; `Python chained method
call resolution` now passes).
- tsc --noEmit clean.
Two related tests (`city.save() via method chain`, `c.greet().save()
depth-2 MRO`) still fail because the captures yield typeBindings
shaped like `city → user.get_city` (no trailing parens — the capture
grabs the attribute text). Resolving those needs a follow step that
detects the call-shape rawName and feeds it through the compound
recurser. Lands with the chain-typeBinding work in a follow-up.
* feat(python-scope): free-call fallback consults finalized bindings
Unit 7 — closes the cross-file free-call gap.
The shared `MethodRegistry.lookup` walks `scope.bindings` (pre-finalize
local-only) for free-call resolution. Cross-file imports land in
`indexes.bindings` (post-finalize). Without the dual-source lookup,
`from x import f; f()` resolves to "unresolved" and no CALLS edge is
emitted.
Two changes:
- `emit-core/scope-walkers.ts`: new `findCallableBindingInScope` —
same dual-source pattern as `findClassBindingInScope`, but accepts
Function/Method/Constructor. Promoted to emit-core because every
language with cross-file imports needs the same lookup.
- `python-scope-emit.ts emitFreeCallFallback`: post-pass that walks
every free-call reference site, looks up the callee with the new
helper, and emits via `tryEmitEdge`. Pre-seeds `seen` from the
shared resolver's emissions so we never double-count.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 22 fail / 169 pass (was 28/163; +6 tests including
the Python overload dispatch fixtures, ancestor-directory imports,
and same-name module-alias collision).
- tsc --noEmit clean.
* feat(python-scope): super() receiver dispatches up the MRO
Unit 8 — `super().method()` inside a class method walks the enclosing
class's MRO chain (skipping self) and resolves to the first ancestor
that owns the method.
New receiver branch in `emitReceiverBoundCalls` recognizes
`super(...)` syntactically (regex-cheap), finds the enclosing class
via a new `findEnclosingClassDef` scope-walk helper, then re-uses
`scopes.methodDispatch.mroFor` + `findOwnedMember` from the existing
class-receiver path. Handled before the compound-receiver case so
`super()` doesn't fall into the bare-identifier branch where `super`
isn't a binding.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 21 fail / 170 pass (was 22/169; `super().save() inside
User to BaseModel.save` now passes).
- tsc --noEmit clean.
* feat(python-scope): suppress shared resolver on member-call sites
Unit 9 — `app_metrics.get_metrics()` (namespace import alias) was
emitting two CALLS edges: a wrong self-call from the shared
resolver's free-call fallback, plus the correct namespace-receiver
edge from the Python post-pass.
Mechanism:
- `emit-core/emit-references.ts`: new optional `skipSites` parameter
(`Set<string>` of `${filePath}:${line}:${col}` keys). When supplied,
references at those positions are skipped — the provider has
already emitted (or chosen not to emit) for that site.
- `python-scope-emit.ts`: reorders Phase 4 — receiver-bound + free-
call fallback run FIRST, populating `handledSites`. The shared
`emitReferencesViaLookup` then runs with that set so the resolver's
fallback can't fight a precise per-receiver emission. Site keys are
added only on successful tryEmitEdge (not for sites the post-pass
saw but couldn't resolve — those still get a chance from the shared
path).
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 20 fail / 171 pass (was 21/170; same-name module-alias
collision now resolves correctly).
- tsc --noEmit clean.
* feat(python-scope): propagate return-type bindings across imports
Closes the cross-file return-type propagation gap that left tests
like `u = get_user(); u.save()` (where get_user lives in another
file) with `u` typed as the function name instead of its return type.
The shared finalize pass copies callable bindings (`from x import f`
puts `f` in the importer's bindings) but typeBindings stay file-local
because they live on `Scope.typeBindings`, not on the index. Mutate
post-finalize:
- For each module-scope import binding (`origin: 'import'` or
`'reexport'`), look up the source file's module-scope typeBinding
for the def's simple name. If present (return-annotation source),
mirror it under the importer's local alias. Skip when the importer
already has its own typeBinding for the name (explicit local always
wins).
- After propagation, re-run a chain-follow on every scope's
typeBindings — pass-4 ran before propagation and missed any chain
whose terminal lived in a foreign file. Same algorithm as
`followChainedRef` in scope-extractor, but operates on the
finalized scopes so propagated entries are visible.
Mutating `Scope.typeBindings` is safe — `draftToScope` constructs a
plain `new Map(...)`, not a frozen one.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 16 fail / 175 pass (was 20/171; +4 — both cross-file
return-type tests, plus two related propagation cases).
- tsc --noEmit clean.
* feat(python-scope): for-loop call-iterable typeBinding
Adds `(for_statement left: (identifier) right: (call function:
(identifier)))` to the typeBinding capture set. Combined with Unit 3's
return-type capture and the cross-file return-type propagation pass,
this makes `for u in get_users(): u.save()` resolve to `User.save`
even when `get_users` is imported from another module.
Captured as `@type-binding.alias` (rawName = function identifier,
without parens) so the existing chain-follow walks the alias to the
function's return-type binding without any new code path.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 12 fail / 179 pass (was 16/175; +4 for-loop call-iterable
tests across get_users / get_repos fixtures).
- tsc --noEmit clean.
* feat(python-scope): collapse free-call edges per (caller, target)
Free calls (no explicit receiver) now emit a single CALLS edge per
(caller, target) pair regardless of how many call sites the caller
contains. Mirrors the legacy DAG's per-pair dedup contract — what
the `default-params`, `variadic`, and `overload` fixtures expect.
Member calls keep position-based dedup so distinct resolved targets
(e.g. UserService.find_user vs AdminService.find_user from the same
caller) still produce distinct edges.
Implementation: bypass `tryEmitEdge` (which dedupes positionally) and
hand-roll the relationship with a position-independent rel.id
(`rel:CALLS:<caller>-><target>`). Site handling is now unconditional —
even when the dedup-collapse skips the actual emit, we mark the site
handled so the shared `emit-references` doesn't fight us with its
fallback.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 10 fail / 181 pass (was 12/179; +2 — both `default
parameter arity` tests now pass).
- tsc --noEmit clean.
* fix(python-scope): match legacy CALLS reason for import-resolved free calls
The arity-narrowing test asserts \`rel.reason === 'import-resolved'\`
for cross-file free-call edges. Switch the free-call fallback's
reason to mirror legacy DAG semantics:
- target-file !== source-file → 'import-resolved'
- same file → 'local-call'
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 9 fail / 182 pass (was 10/181; +1 arity-narrowing test).
- tsc --noEmit clean.
* fix(python-scope): drop dead pre-seeding from receiver-bound pass
The pre-seeding loop at the top of \`emitReceiverBoundCalls\` populated
\`seen\` with every reference the shared resolver had already resolved.
That was useful when emit-references ran FIRST. After Unit 9 reversed
the order (emit-references runs after the Python passes and uses
\`handledSites\` to skip what we processed), the pre-seed only causes
harm: when an MRO walk in Case 0 (compound receiver) and Case 4
(simple typeBinding) both touch the same site at the same position
but resolve to different targets, the pre-seed suppresses the second
emission because the shared resolver had already entered the wrong
target into \`seen\`.
Concrete case: \`c.greet().save()\` — Case 0 emits the outer save edge
to Greeting.save; Case 4 then resolves the inner \`c.greet()\` to
A.greet via MRO walk. With pre-seed both edges should emit (different
targets, different rel.ids); without removing the pre-seed the inner
emission was being deduped against an already-seeded entry and the
A.greet edge was lost.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 8 fail / 183 pass (was 9/182; +1 — \`c.greet() to A#greet
via MRO walk\` now passes).
- tsc --noEmit clean.
* feat(python-scope): enumerate(X) for-loop tuple destructuring
Adds two new typeBinding capture patterns for the canonical enumerate
pattern:
for (i, u) in enumerate(users): ... ; tuple_pattern
for i, u in enumerate(users): ... ; pattern_list
Both bind the second tuple element (u) to the iterable identifier
(users). The chain-follow then unwraps users → its element type via
the existing generic-strip in interpret.ts (List[User] → User).
The #eq? predicate scopes the pattern to enumerate specifically;
generic tuple destructuring of arbitrary callables is left to a
future iteration once we have a richer signal for "what does this
call yield".
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 7 fail / 184 pass (was 8/183; +1 — `parenthesized tuple:
for (i, u) in enumerate(users)` now passes).
- tsc --noEmit clean.
* feat(python-scope): dict.items() value-type unwrapping
Two changes that together resolve `for k, v in data.items(): v.save()`:
- `interpret.ts stripGeneric`: extends to `dict[K, V]` /
`Dict[K, V]` / `Mapping[K, V]` etc., stripping to the value type V.
Previously only single-arg generics (list[User] → User) were
stripped; multi-arg ones returned the raw text.
- `query.ts` + `scopes.scm`: new typeBinding patterns for
`for k, v in X.items()` (both pattern_list and tuple_pattern). The
second tuple element binds to X; the chain-follow then unwraps X's
dict annotation to V via the new stripGeneric branch.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 6 fail / 185 pass (was 7/184; +1 — `dict.items() loop`
test now passes).
- tsc --noEmit clean.
* feat(python-scope): nested tuple destructuring for enumerate(d.items())
Two more for-loop typeBinding patterns:
- `for i, (k, v) in enumerate(d.items())` — nested tuple destructuring
where v is the value of the dict's items() yield.
- `for v in d.values()` — explicit values() form (companion to items).
Both bind the loop var to the dict identifier; the chain-follow
unwraps via the dict-aware stripGeneric to the value type.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 5 fail / 186 pass (was 6/185; +1 nested tuple test).
- tsc --noEmit clean.
* feat(python-scope): 3-var flat destructuring for enumerate(d.items())
Adds the \`for i, k, v in enumerate(d.items())\` shape — flat
3-variable destructuring of the (i, (k, v)) tuple yielded by
\`enumerate\` over \`items()\`. Binds v (the last identifier in the
pattern_list) to the dict identifier; the existing dict-aware
stripGeneric unwraps to the value type.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 4 fail / 187 pass (was 5/186; +1).
- tsc --noEmit clean.
* feat(python-scope): write ACCESSES edges for attribute assignments
Three changes that together produce ACCESSES (write) edges for
\`obj.field = value\` assignments:
- New \`@reference.write.member\` capture in query.ts and scopes.scm
matching \`(assignment left: (attribute object: ... attribute: ...))\`.
Reuses the existing receiver/name capture shape so the
receiver-bound emit pass can resolve obj's class and look up the
field.
- \`populateMethodOwnerIds\` now sets ownerId on class-body fields too,
not only on methods. Previously it only walked Function scopes
whose parent was Class; class-body annotations like \`name: str\`
live directly in the Class scope's ownedDefs and were missed, so
\`findOwnedMember(User, "name")\` returned undefined.
- \`emit-core isLinkableLabel\` extends to Variable and Property so
field nodes appear in the graph-node lookup (the legacy parser
emits both kinds for class-body annotations).
- Case 4 in receiver-bound pass now uses the kind word as the edge
reason for read/write sites — matches the legacy DAG convention
the test asserts on.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 3 fail / 188 pass (was 4/187; +1 — write-ACCESSES test).
- tsc --noEmit clean.
* feat(python-scope): chain-typebinding + field-fallback method lookup
Reaches the architectural-plan target of >= 189/191 flag-on passing.
Two intertwined changes:
- Field-fallback in resolveCompoundReceiverClass: when method lookup
on the receiver's class (and its MRO) fails, walk the class's
fields and try the same lookup on each field's type. Matches the
"unified fixpoint" intent of the method-chain fixture where
`user.get_city()` reaches `Address.get_city` through User's
`address: Address` field.
- New Case 3b in receiver-bound emit pass: when the receiver's
typeBinding rawName has a dot but isn't a namespace prefix
(e.g. `city -> user.get_city` from the constructor-inferred capture
for `city = user.get_city()`), treat it as a method-call chain and
pipe through the compound resolver. The chain unwraps to the
terminal class (City) and the call resolves normally.
Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 2 fail / 189 pass (was 3/188; +1 city.save method chain).
- tsc --noEmit clean.
Remaining 2 failures are fixture-driven (self.users / self.repos
fixtures reference fields that aren't declared on the class) and
documented as known-limitation in Unit 10.
* feat(python-scope): flip Python to registry-primary (191/191 parity)
Adds the \`for u in self.X\` heuristic typeBinding capture (binds u to
the attribute name X so the chain-follow can resolve via the enclosing
method's parameter typeBinding) — closes the last two failing
fixtures whose classes reference \`self.X\` for fields that are
actually method parameters.
With 191/191 passing on BOTH legacy and registry-primary paths,
flips \`MIGRATED_LANGUAGES\` to include \`SupportedLanguages.Python\`.
Effects:
- Production default for Python files: registry-primary path.
- CI parity gate auto-discovers Python via the script + workflow
(\`scripts/ci-list-migrated-languages.ts\` /
\`.github/workflows/ci-scope-parity.yml\`) and runs the resolver
integration test BOTH ways on every PR.
- Operators retain the \`REGISTRY_PRIMARY_PYTHON=0\` escape hatch.
Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- Default (unset, post-flip): 191/191 (uses registry).
- tsc --noEmit clean.
This concludes RFC #909 Ring 3 — Python migration.
* refactor(emit-core): EmitProvider interface + promote 5 generic helpers
G-Units 1-2 of the emit-pipeline generalization plan.
Adds:
- emit-core/emit-provider.ts — typed EmitProvider contract (6 required +
2 optional fields). Will be consumed by the generic orchestrator in
G-Unit 6. Documents the LanguageProvider vs EmitProvider boundary.
- emit-core/emit-free-call.ts — emitFreeCallFallback promoted as-is
(drops the unused referenceIndex pre-seed parameter; underscore-prefixed
to keep the signature compatible).
- emit-core/propagate-return-types.ts — propagateImportedReturnTypes +
followChainPostFinalize. Documents the mutation contract (Invariant
I3 + I6 from the plan): runs after finalize, before resolve, mutates
the non-frozen Scope.typeBindings map.
- emit-core/scope-walkers.ts: + findEnclosingClassDef +
findExportedDefByName. Both were already generic in the Python
source.
python-scope-emit.ts shrinks 1055 → 799 lines (–256). Imports the
promoted helpers from emit-core. No behavior change.
Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- tsc --noEmit clean.
* refactor(emit-core): promote receiver-bound dispatcher + compound resolver
G-Unit 3 of the emit-pipeline generalization plan.
- emit-core/emit-compound-receiver.ts — resolveCompoundReceiverClass
+ matchingOpenParen + COMPOUND_RECEIVER_MAX_DEPTH. Field-fallback
is now an option (default true) so strictly-typed languages can
opt out via EmitProvider.fieldFallbackOnMethodLookup.
- emit-core/emit-receiver-bound.ts — the 7-case dispatcher (super,
Cases 0/1/2/3/3b/4). Accepts a ReceiverBoundProviderSubset
(isSuperReceiver + fieldFallbackOnMethodLookup) so partial wiring
works during the rest of the migration. Documents Contract
Invariants I4 (case order) and I5 (no pre-seeding).
python-scope-emit.ts shrinks 799 → 384 lines. The orchestrator now
calls the generic emitReceiverBoundCalls with an inline minimal
provider (pythonEmitProviderInline) — full provider lands in G-Unit 6
when the orchestrator itself moves to languages/python/emit/.
Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- tsc --noEmit clean.
* refactor(emit-core): promote MRO walk + populateClassOwnedMembers
G-Units 4-5 of the emit-pipeline generalization plan.
- emit-core/build-mro.ts — generic buildMro takes a LinearizeStrategy
hook receiving (classDefId, directParents, parentsByDefId). Three
shared steps (collect EXTENDS, build defId-by-graphId, walk per
class) + parametric linearization. Default strategy is BFS-with-
visited (Python's depth-first first-seen, also correct for
single-inheritance languages).
- emit-core/scope-walkers.ts: + populateClassOwnedMembers — generic
OO ownership rule (methods + class-body fields). Both rules ship
together because every OO language migrated so far (Python; planned
TS/JS/Java/Kotlin) wants both. Languages that need different rules
can compose with this as a base step.
python-scope-emit.ts shrinks 384 → 255 lines.
Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- tsc --noEmit clean.
* refactor(scope-resolution): generic orchestrator + language-agnostic phase
G-Units 6-7 of the emit-pipeline generalization plan, plus the
pipeline-phase generalization (the user's observation that the phase
itself is generic once the orchestrator is).
Changes:
- emit-core/orchestrator.ts — runScopeResolution(input, provider).
The 180 lines of pipeline glue moved here, parametrized by
EmitProvider. Provider supplies LanguageProvider, importEdgeReason,
and the 6 emit-side hooks.
- emit-core/emit-provider.ts — EmitProvider gains languageProvider
and importEdgeReason fields so the orchestrator needs nothing else.
resolveImportTarget now takes (targetRaw, fromFile, allFilePaths).
- languages/python/emit/index.ts — pythonEmitProvider + thin
runPythonScopeResolution wrapper. The first reference impl every
next-language migration copies.
- emit-providers-registry.ts (NEW) — registry of per-language
EmitProviders keyed by SupportedLanguages. Adding a language is
one line here + the provider file.
- pipeline-phases/scope-resolution.ts (NEW) — language-agnostic phase
iterating EMIT_PROVIDERS ∩ MIGRATED_LANGUAGES. Replaces
pipeline-phases/python-scope.ts (deleted).
- python-scope-emit.ts deleted.
- pipeline.ts swaps pythonScopePhase → scopeResolutionPhase.
The next language migration is now: implement EmitProvider, register
it, add to MIGRATED_LANGUAGES. No new pipeline phase, no orchestrator
copy-paste. The Python migration's 700+ lines of glue collapse to
~80 lines per future language.
Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- Default (post MIGRATED_LANGUAGES flip): 191/191.
- tsc --noEmit clean.
* docs(emit-provider): migration cookbook for next-language porters
* refactor(scope-resolution): rename emit-core/ → scope-resolution/, EmitProvider → ScopeResolver
Reorganizes the registry-primary resolution layer for clarity and
contributor onboarding. Driven by feedback that "emit" was triple-
overloaded (graph-edge emission + tree-sitter capture extraction +
the provider name itself), and the flat 16-file emit-core/ folder
mixed five concerns.
External research (rust-analyzer hir-def/nameres, Pyright analyzer/,
TypeScript binder/checker, Roslyn Binder, IntelliJ Resolver, swc
semantic/, biome semantic/, semgrep naming/, JDT Binding, clangd
Sema) consistently uses **the phase name** for this layer, never an
output verb. "Scope resolution" matches our pipeline-phase name, the
plan, and the RFC.
## Folder rename
emit-core/ → scope-resolution/
├── (16 flat files) → ├── contract/scope-resolver.ts
├── pipeline/{run,registry,phase}.ts
├── passes/{receiver-bound-calls,
│ free-call-fallback,
│ compound-receiver,
│ imported-return-types,
│ mro}.ts
├── graph-bridge/{node-lookup,ids,
│ edges,references-to-edges,
│ imports-to-edges,
│ method-dispatch}.ts
└── scope/{walkers,namespace-targets}.ts
Each subfolder maps to one concern a new contributor needs to find:
*the contract I implement / the runner that calls me / the helpers I
reuse / the graph layer I shouldn't touch / the scope walkers*.
## Symbol renames
EmitProvider → ScopeResolver
pythonEmitProvider → pythonScopeResolver
runPythonScopeResolution → resolvePythonScope
EMIT_PROVIDERS → SCOPE_RESOLVERS
getEmitProvider → getScopeResolver
RunPythonScopeResolution{Input,Stats} → ResolvePythonScope{Input,Stats}
## File renames (per-language)
languages/python/emit/index.ts → languages/python/scope-resolver.ts
languages/python/emit-captures.ts → languages/python/captures.ts
(kills the parse-side "emit" collision)
## Mechanics
- Used `git mv` for all files so blame history is preserved.
- Updated ~30 import lines across 18 files plus the pipeline-phases
barrel and pipeline.ts.
- Updated JSDoc cross-references throughout to match the new vocabulary.
Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- Default (post MIGRATED_LANGUAGES flip): 191/191.
- tsc --noEmit clean.
Migration cookbook in `scope-resolution/contract/scope-resolver.ts`
JSDoc points the next-language porter at all the new names and
folder locations.
* docs(scope-resolution): finalize phase JSDoc + drop python emoji from generic log line
* perf(scope-resolution): O(1) workspace lookup index
Introduces `WorkspaceResolutionIndex` — a precomputed bundle of
lookup tables built ONCE per resolution run, after `populateOwners`
and after finalize, before any pass that needs to find members,
exported defs, or class scopes by id.
What it replaces (all are pre-existing O(N×D) linear scans of
parsedFiles, called inside the receiver-bound MRO chain):
- `findOwnedMember(ownerId, name, parsedFiles)` → `Map.get` via
`index.memberByOwner.get(ownerId)?.get(name)`. Was the worst
offender — receiver-bound dispatcher calls this O(sites × MRO
depth) times.
- `findExportedDef(filePath, name, parsedFiles)` → `Map.get` via
`index.defsByFileAndName`. Hot for namespace-receiver case.
- `findExportedDefByName` workspace-wide fallback scan → `Map.get`
via `index.callablesBySimpleName`.
- `classScopeByDefId` (rebuilt inside `emitReceiverBoundCalls` on
every invocation) — moved to one-shot build during finalize, read
from `index.classScopeByDefId` everywhere.
- `moduleScopeByFile` (rebuilt inside `propagateImportedReturnTypes`
on every invocation) — read from `index.moduleScopeByFile`.
Findings from a synthetic 100-file Python workload (60 model files
each defining 5 classes × 3 methods + 40 user files calling them
heavily):
scope-resolution wall time: 764ms → 710ms (median, 5 iters)
That's a ~7% in-layer win. The smaller-than-expected gain was
informative: profiling the synthetic workload shows scope-resolution
breakdown is `extract=62% resolve=30% emit=4%`; the index touched
the 4% slice (emit + walker calls inside it). Larger O(D) per owner
classes will benefit more.
Profiling the FULL pipeline (49 fixtures × 3 iters) shows
scope-resolution accounts for ~1% of pipeline wall time — the
remaining 99% is parse (tree-sitter), heritage, ORM, MRO, processes,
and DB writes. So further optimization of this specific layer has
marginal pipeline impact; the next-biggest wins live in those
phases. Documented as the "double-parse" finding in the audit
(captures.ts re-parses each Python file even though the parse phase
already produced a tree-sitter Tree) — that's a separate plumbing
project across phase boundaries.
Bonus: opt-in PROF_SCOPE_RESOLUTION=1 env var prints a per-phase
ms breakdown to stderr, so future perf work can measure without
extra code changes.
Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- tsc --noEmit clean.
* perf(parse/heritage/mro): typed graph iterator + cross-phase tree cache
Two structural perf wins targeting the parse / heritage / MRO
layers, identified by the post-WorkspaceResolutionIndex profiling
(scope-resolution = ~1% of pipeline; the bulk lives upstream).
## 1. KnowledgeGraph.iterRelationshipsByType (PHM-Units 1-2)
- Adds a per-type `Map<RelationshipType, Map<id, Relationship>>`
index inside `createKnowledgeGraph`, maintained on add / remove /
removeNode / removeNodesByFile.
- New `iterRelationshipsByType(type)` returns a typed iterator that
yields only the requested type. Backwards-compatible: existing
`iterRelationships()` / `forEachRelationship()` callers untouched.
- Migrated two MRO call sites:
- `mro-processor.ts buildAdjacency`: split the single
`forEachRelationship` (which scanned every edge in the graph and
type-filtered per-iteration) into three typed iterations
(EXTENDS, IMPLEMENTS, HAS_METHOD).
- `scope-resolution/passes/mro.ts buildMro`: replaced
`for (const rel of graph.iterRelationships()) if (rel.type !== 'EXTENDS') continue`
with `for (const rel of graph.iterRelationshipsByType('EXTENDS'))`.
- Heritage-processor (PHM-Unit 3) was a no-op: it only WRITES
EXTENDS/IMPLEMENTS edges, never re-reads. Index is still useful
for the seven other graph-iter consumers (community-processor,
csv-generator, wildcard-synthesis, process-processor, etc.) — those
follow-ups can switch to the typed iterator without touching the
graph layer.
- Adds 5 unit tests for the new method (add/remove/dedupe semantics,
empty-type fresh iterator, removeNode index sync).
## 2. Cross-phase tree cache (PHM-Units 4-5)
The audit's #2 finding: Python files are parsed by tree-sitter once
in the parse phase, then re-parsed inside scope-resolution's
`captures.ts`. Eliminate the second parse by sharing the Tree across
phases.
- `parse-impl.ts` now maintains TWO ASTCaches with distinct lifetimes:
- `astCache` (chunk-local, cleared between chunks) — unchanged;
used by call/heritage/import processors during parse.
- `scopeTreeCache` (total-parseable-sized, never cleared) — new,
exposed via `ParseOutput.astCache` for cross-phase consumption.
- `parsing-processor.ts` writes every sequentially-parsed Tree to
BOTH caches. Worker-mode parses skip the persistent cache too
(Trees can't cross MessageChannels).
- `LanguageProvider.emitScopeCaptures` gains an optional `cachedTree`
parameter (typed `unknown` to keep the tree-sitter dep out of the
contract).
- `captures.ts` short-circuits its own `parser.parse(sourceText)`
when a cached Tree is supplied. Cache miss falls back to a fresh
parse — same correctness path as before.
- `runScopeResolution` accepts an optional `treeCache` and forwards
per-file `cachedTree` to `extractParsedFile`.
- `scope-resolution/pipeline/phase.ts` reads
`getPhaseOutput<{astCache}>(deps, 'parse')` and passes through.
Verified end-to-end: a small fixture run with PROF_SCOPE_RESOLUTION=1
shows 6/6 cache hits (100% hit rate) on the python-grandparent fixture
that exercises the full pipeline below the worker-pool threshold.
## Verification
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- New graph.test.ts: 25/25 (was 20).
- tsc --noEmit clean.
## Where the win lands
Wall-clock on the 49-fixture integration suite: 14050ms → 14080ms
(within noise). Fixtures are 1-3 files each, dominated by per-fixture
pipeline overhead (worker-pool init, DB writes, fixture startup).
The cache + typed-iterator wins are constant-factor improvements
that scale linearly with workload size and visible only on larger
repos. The dev-mode `PROF_SCOPE_RESOLUTION` instrumentation +
`getPythonCaptureCacheStats()` are kept for future perf work.
## Plan
docs/plans/2026-04-20-002-perf-parse-heritage-mro-plan.md.
PHM-Unit 3 (heritage-processor migration) intentionally collapsed
to a no-op — heritage only writes, never re-reads.
* perf(scope-resolution): bound tree-cache lifetime + gate population
Address P1 residuals from ce:review of 8c6f5cee:
- Dispose scopeTreeCache at end of scopeResolutionPhase via
astCache.clear(). Trees were previously retained for the full
pipeline (10-100x memory regression on large repos). Downstream
phases (mro, community, csv-generator) never read them.
- Gate scopeTreeCache.set on provider.emitScopeCaptures !== undefined.
Polyglot repos no longer retain Trees for languages with no
scope-resolution consumer.
- PROF_SCOPE_RESOLUTION=1 now warns when workers engage, since
Trees can't cross MessageChannels so the cache will be empty for
worker-parsed files — prevents a silent perf cliff once a repo
crosses the worker-pool threshold.
Tests: 26/26 graph unit, 299/299 scope-resolution unit, 191/191
python integration both flag paths.
* refactor(scope-resolution): clean up P2/P3 review residuals
P2:
- WASM dual-ownership invariant documented on ASTCache dispose:
a Tree must live in AT MOST ONE disposing ASTCache. Native
tree-sitter today is unaffected; WASM adoption would require
tree.copy() or a non-disposing secondary cache.
- mro-processor C3 ordering test: pins EXTENDS-before-IMPLEMENTS
parent grouping for classes with interleaved edge additions.
Asserts exact MRO ['Base', 'Iface'] — a revert to single-loop
insertion-order iteration would produce ['Iface', 'Base'] and
fail loudly.
- cached-tree parity test: emitPythonScopeCaptures(src, path, T)
returns identical CaptureMatch[] to emitPythonScopeCaptures(src,
path). Pins the cache-hit path's correctness so a regression
that silently returns stale captures would break the test.
P3:
- Dev-mode cache counters moved from captures.ts to cache-stats.ts.
Production hot-path module no longer carries the module-global
export surface; PROF gating behavior preserved.
- ParseOutput field rename astCache → scopeTreeCache. Clarifies
that the surfaced cache is the persistent cross-phase one, not
the chunk-local astCache parse-impl clears between chunks.
Single consumer (scopeResolutionPhase) updated; no other readers.
- ASTCacheReader interface extracted. scopeResolutionPhase now
reads the phase dep via a shared type instead of a hand-rolled
inline structural shape that could drift from ASTCache's contract.
- graph.ts dual-index invariant enforced through writeRel/deleteRel
private helpers instead of duplicated add/delete at 3 mutation
sites. Adding a new mutation method only needs to call the
helpers — forgetting to update one index becomes structurally
impossible.
Tests: 382/382 unit (incl. 2 new), 191/191 python integration both
flag paths. tsc clean.
* fix(ci): prettier formatting + Python-migration test adjustments
CI run 24666612657 failed on three jobs. Fixes:
quality/format:
- Prettier --check flagged 3 files after the accumulated branch work.
Ran prettier --write from repo root (CI's invocation cwd) to apply:
simple-hooks.ts, resolve-references.ts, python-hooks.test.ts.
tests/{ubuntu,macos,windows} — 9 assertion failures, all traceable to
Python landing in MIGRATED_LANGUAGES (default-on registry-primary):
- registry-primary-flag.test.ts (3 tests): the 'returns false by
default' / 'primaryLanguages empty' / 'Python mid-process
mutation' assertions were written in Ring 2 when MIGRATED_LANGUAGES
was empty. Rewrote to assert MIGRATED_LANGUAGES membership is the
default, use Java (unmigrated) for the no-stale-cache test, and
verify env overrides work in both directions (migrated-off,
unmigrated-on).
- call-processor.test.ts (6 tests in SM-10 + D2-widen blocks):
these exercise the LEGACY call-resolution DAG on .py fixtures.
processCalls now gates Python out (isRegistryPrimary === true by
default), returning 0 edges. Added REGISTRY_PRIMARY_PYTHON=false
override in the relevant beforeEach + restore in afterEach, so
the legacy DAG runs for these test-local fixtures without
affecting the production-default behavior.
Local verification: 4126/4126 unit tests pass, prettier clean.
* docs(python): known-limitation block on scope-resolution public API
Unit 10 — document what the Python registry-primary path intentionally
does not resolve, so reviewers and future maintainers can distinguish
conscious trade-offs from latent bugs:
- Dynamic attribute access (getattr / setattr)
- Dynamic imports (importlib, __import__)
- Metaclass-driven dispatch
- Union / Optional branch-picking behavior
- Arbitrary signature-rewriting decorators
- typing.TYPE_CHECKING-guarded imports
- *args / **kwargs type flow-through
- super() outside a directly-bound method
Each item names the file that owns the relevant hook so a future
follow-up knows where to start. Shadow-harness corpus parity + the
CI parity gate remain the authoritative signal for which of these
matter at fleet scale.
* docs: record scope-resolution pipeline alongside legacy call DAG
Capture what shipped in #980 so future readers don't have to reverse-
engineer the coexistence of the legacy call-resolution DAG and the new
scope-resolution pipeline:
- ARCHITECTURE.md: new 'Scope-Resolution Pipeline' section after the
Call-Resolution DAG, documenting pipeline stages, ScopeResolver
contract, per-language registration, code references, and perf
notes. Coexistence block added to the legacy DAG section explaining
how MIGRATED_LANGUAGES gates the two paths per-language.
- AGENTS.md: reference-docs pointer updated — legacy-DAG one-liner
stays; scope-resolution pipeline gets its own pointer so agents
know when to read which section. Changelog bumped.
- type-resolution-system.md: callout at the 'call-processor.ts is
the consumer' claim pointing readers to the scope-resolution path
for migrated languages. TypeEnv is still built per file, but for
migrated languages receiver typing flows through ParsedTypeBinding
rather than call-processor.ts.
CHANGELOG.md intentionally not touched — owned by the release process.
* chore: remove obsolete scheduled_tasks.lock file
* fix(scope-resolution): qualified-name keys for same-file method collisions
Review feedback from PR #980 reviewer flagged a BLOCKING correctness
bug: when two classes in the same file define a method with the same
simple name (e.g. class User: def save + class Document: def save),
every d.save() CALLS edge silently resolved to User.save because the
graph node lookup keyed only by (filePath, simpleName) and first-wins
took User's method.
Three-layer fix:
1. populateClassOwnedMembers now promotes a nested def's
qualifiedName from `save` to `ClassName.save` when the def sits
inside a class scope. Python's scopes.scm doesn't emit
@declaration.qualified_name for methods, so without this the
finalized SymbolDefinition carried only the simple name.
2. buildGraphNodeLookup adds a second key per node:
(filePath, qualifiedName). For Method/Function nodes the qualifier
is parsed deterministically out of the node id
(`Method:file.py:User.save#N` → `User.save`), which is robust to
Windows-style filePath colons. Simple-name key retained as a
fallback for callers that don't know the qualifier.
3. resolveDefGraphId now tries the qualified key first, then falls
back to the simple-name lookup.
Also addresses the non-blocking review items:
- scopeResolutionPhase.deps now includes `crossFile` so the Kahn's
runner can't schedule scope-resolution before crossFile finishes
writing heritage edges that buildMro consumes.
- run.ts no longer mutates the finalized ScopeResolutionIndexes via
`as` cast — spreads into a fresh object with the populated
methodDispatch field instead.
- Doc nits: scope-resolver.ts registry path + phase.ts Ring number.
Test coverage:
- New fixture test/fixtures/lang-resolution/python-same-file-method-collision
with User.save + Document.save in one file and app.py calling both
through typed receivers.
- Three new integration assertions pin that u.save() and d.save()
target the correct qualified node id. Fail before the fix, pass
after. Confirmed by running once without populateClassOwnedMembers
qualifier promotion — reproduces the original User.save-for-both bug.
Verification: 194/194 test/integration/resolvers/python.test.ts pass
both REGISTRY_PRIMARY_PYTHON=0 and =1. 523/523 related unit tests.
tsc --noEmit clean.
* fix(scope-resolution): filter export index to module-level defs + label-prefixed qualified key
Codex adversarial review on PR #980 flagged that
buildWorkspaceResolutionIndex feeds defsByFileAndName and
callablesBySimpleName from parsed.localDefs — the flat set of every
def in the file including methods, fields, and nested functions.
findExportedDef / findExportedDefByName treat those maps as
file-level exports, so `mod.save()` could silently bind to User.save
whenever a method's simple name appeared first in parse order.
Plan: docs/plans/2026-04-21-001-fix-workspace-index-module-scope-only-plan.md
Fix layers:
1. workspace-index.ts: split the single parsed.localDefs loop into
two passes:
- Module-export pass: iterate moduleScope.ownedDefs PLUS ownedDefs
of every child scope whose parent is the module scope. Top-level
class and function declarations each live in their own scope
with parent=module, not in moduleScope.ownedDefs directly, so
the "parent === moduleScope.id" walk is required to reach them.
Methods (scope.parent === Class scope) and nested functions
(scope.parent === another Function scope) are excluded.
- Member-by-owner pass: keeps iterating parsed.localDefs since
that map is keyed on ownerId and correctly saw class-owned defs
before this change.
2. graph-bridge/node-lookup.ts: qualified keys now live in a separate
keyspace (`<q>:filePath::<label>::<qualifiedName>`) and include
the node label. Without the label prefix, a top-level `def save`
(Function, qualifier `save`) would collide with a class method
`User.save` (Method, simple name `save`) in the same simple-key
slot because the Function's qualifier happens to equal the
Method's simple name. The label differentiates them.
3. graph-bridge/ids.ts: resolveDefGraphId uses the new
type-prefixed qualified key when def.type is set. Simple-name
fallback retained for languages that don't yet synthesize
qualifiers on their defs.
Test fixture: python-module-export-vs-method-collision places
`class User: def save` BEFORE top-level `def save` — parse order
that exposes the bug (class method enters the index first). Three
new integration assertions:
- `mod.save(x)` resolves to the module-level Function, not User.save
- `u.save()` resolves to User.save Method
- Exactly two CALLS edges to `save` exist, one per intended target
Fixture confirmed failing before the workspace-index fix (bug
reproduced), passing after.
Verification: 197/197 test/integration/resolvers/python.test.ts pass
both REGISTRY_PRIMARY_PYTHON=0 and =1. 523/523 related unit tests.
tsc --noEmit clean.
* fix(scope-resolution): drive module export index from moduleScope.bindings
Codex round-2 adversarial review flagged that the workspace-index
module-export pass iterated every def in every direct-child scope of
the module, including class-body Variable defs like
`class User: MAX_USERS = 100`. `defsByFileAndName[file][MAX_USERS]`
silently aliased to the class attribute. Latent today because Python
doesn't emit ACCESSES edges for `mod.NAME` member access, but the
index-layer leak would surface the moment reference capture widens.
Plan: docs/plans/2026-04-21-002-fix-codex-round2-scope-resolution-plan.md
Drive the module-export index from the extractor invariant instead of
a scope-kind → allowed-label switch:
moduleScope.bindings already contains exactly the names visible at
module level — top-level class/function declarations, module-level
variable assignments, imports. Class methods, class-body attributes,
and nested-function defs bind to their containing (Class or Function)
scope, not the module, so they're naturally excluded.
Filter to `BindingRef.origin === 'local'` so imports and wildcard
re-exports stay out of the index (matches the pre-fix invariant when
the source was `parsed.localDefs`).
No per-kind predicates, no scope-kind / def-kind enumeration, no
two-pass merge between moduleScope.ownedDefs and direct-child scope
walks — one loop, language-agnostic.
Codex also flagged `propagateImportedReturnTypes` as potentially
broken for function-local imports, but scope-dump probing showed the
finalize algorithm puts `from svc import get_user` into the MODULE
scope's finalized bindings even when declared inside a function, so
the existing module-scope propagation already handles the case. The
new python-function-local-import-chain integration test pins that
working behavior as a regression guard; no code change required.
Coverage:
- test/unit/scope-resolution/workspace-index.test.ts (new, 5 tests) —
directly asserts the index shape. The "excludes class-body Variable
defs" test fails without this fix and passes after (confirmed via
stash-pop probe).
- test/integration/resolvers/python.test.ts — 4 new integration
assertions across two describe blocks (python-class-attr-export-leak,
python-function-local-import-chain) pin end-to-end invariants.
- Two new fixtures under test/fixtures/lang-resolution/.
Verification: 201/201 test/integration/resolvers/python.test.ts both
REGISTRY_PRIMARY_PYTHON=0 and =1. 528/528 related unit tests (was
523). tsc clean.
* test(scope-resolution): pin local-namespace-import behavior + document empirical finalize hoisting
Codex round-3 adversarial review raised three concerns about
scope-resolution passes assuming module-scope semantics that would
contradict `pythonImportOwningScope`'s documented per-scope contract.
Empirical verification via scope-dump probes resolved each:
Plan: docs/plans/2026-04-21-003-fix-codex-round3-scope-aware-resolution-plan.md
1. Function- and class-local namespace imports: VERIFIED WORKING.
`def outer(): import svc as s; s.call()` and `class A: import mod;
def use(self): mod.helper()` both emit CALLS edges with reason
"scope-resolution: namespace-receiver". finalize-algorithm hoists
the ImportEdges onto `indexes.imports[moduleScope]` regardless of
where the `import` statement appears, so collectNamespaceTargets'
module-scope read finds them.
2. Imported return-type propagation module-scope-only: VERIFIED
WORKING (already pinned in round 2). `from svc import get_user`
inside a function body lands in indexes.bindings[moduleScope], so
propagateImportedReturnTypes' module-scope read still finds it.
3. Nested method-local defs stamped as class members: VERIFIED FALSE.
The scope extractor creates nested Function scopes for inner
`def`s; `def helper` inside `def save` inside `class User` lives
in helper's own Function scope whose parent is save's Function
scope (NOT the Class scope). populateClassOwnedMembers'
`parentScope.kind === 'Class'` branch correctly skips it;
helper.ownerId stays undefined.
Instead of implementing speculative scope-aware refactors that the
tests would pass regardless, this commit:
- Adds regression fixtures and integration assertions that pin each
working behavior. If finalize routing ever changes to honor the
hook's per-scope contract, these assertions flip red and signal the
need for the scope-chain-aware refactor.
- Adds defensive JSDoc to the three flagged call sites
(collectNamespaceTargets, propagateImportedReturnTypes,
populateClassOwnedMembers) documenting the empirical invariant so
future reviewers don't re-derive Codex's theoretical concern
without the benefit of the probe.
Files:
- Two new fixtures under test/fixtures/lang-resolution/ covering the
function-local and class-body namespace-import patterns.
- Two new describe blocks in test/integration/resolvers/python.test.ts
(3 assertions, positive-pin intent).
- Defensive comments in namespace-targets.ts, imported-return-types.ts,
and scope-resolution/scope/walkers.ts.
Verification: 204/204 test/integration/resolvers/python.test.ts both
REGISTRY_PRIMARY_PYTHON=0 and =1. tsc clean.
* perf(graph): reverse-adjacency + file indexes drop removeNode/removeNodesByFile from O(N)
PR #980 in-line review flagged that `removeNode` iterated the full
relationshipMap to find edges touching a node (O(E)), and
`removeNodesByFile` called removeNode for every matching node after
a full nodeMap scan (O(N × E)). Pre-existing, but worth fixing
properly since the writeRel/deleteRel helpers we just added make the
index-maintenance story coherent.
Two new indexes maintained on every mutation path:
- `edgeIdsByNode: Map<nodeId, Set<relId>>` — reverse adjacency. Every
edge records both endpoints, so removeNode iterates
edgeIdsByNode.get(id) instead of every relationship. Self-edges
skip the duplicate-endpoint write to keep the Set dedup explicit.
- `nodeIdsByFile: Map<filePath, Set<nodeId>>` — file index.
removeNodesByFile reaches its file's nodes directly.
Complexity:
- removeNode: O(edges-touching-node), was O(total-edges).
- removeNodesByFile: O(file-nodes × avg-edges-per-node + scan of the
file bucket), was O(total-nodes + file-nodes × total-edges).
Index maintenance is centralized in writeRel/deleteRel + new
addToBucket/removeFromBucket helpers. Empty buckets are pruned to
keep the indexes compact. Existing dual-invariant (relationshipMap ↔
relationshipsByType) preserved.
Nodes without a `filePath` property (e.g. Community/Cluster nodes)
are intentionally NOT indexed in nodeIdsByFile — they can't belong
to any file, so removeNodesByFile correctly leaves them alone.
Coverage: 7 new unit tests (33/33 total, was 26). Added cases:
- removes only edges touching the removed node
- handles self-edges
- removes orphan node with no edges
- removeNodesByFile removes only matching nodes
- returns 0 when no match
- also removes edges whose endpoints lived on the removed file
- does not index nodes without a filePath property
Verification: 204/204 test/integration/resolvers/python.test.ts both
REGISTRY_PRIMARY_PYTHON=0 and =1. 4235/4235 unit tests. tsc clean.
* refactor(ingestion): merge python/ast-utils into utils/ast-helpers; iterative findNodeAtRange
python/ast-utils.ts held three language-agnostic helpers
(nodeToCapture, syntheticCapture, findNodeAtRange) plus two
duplicates of the shared utils version (findChildOfType ==
findChild; findIdentifierChild was unused). Consolidating into
utils/ast-helpers.ts so the next language migrating to the
scope-resolution pipeline imports from one place.
findNodeAtRange rewritten iteratively using an explicit stack.
Previous implementation was recursive — fine for shallow Python
trees today, but a landmine for languages with deeper nesting
(Kotlin sealed-hierarchy decomposition, Rust macro expansion,
etc.) and the task hooks explicitly call out "no recursion".
Children are pushed reverse-index so LIFO pop visits them
left-to-right; row-bound pruning preserves the prior early-skip
optimization (the `break` shortcut is replaced with `continue`
since a stack can't leverage ordered sibling termination).
findChildOfType consumers migrated to the existing findChild
helper. findIdentifierChild deleted — no callers remained.
Coverage: 204/204 test/integration/resolvers/python.test.ts both
REGISTRY_PRIMARY_PYTHON=0 and =1. 339/339 scope-resolution +
graph unit tests. tsc clean.
* refactor(scope-resolution): remove unused shouldShadow / shouldCreateScope hooks
Both LanguageProvider hooks were dead weight:
- `shouldShadow` had zero call sites — the interface declared it,
Python implemented a trivial always-true no-op, but no consumer
ever read it. The shadowing decision lives in pythonMergeBindings
and the central merge algorithm, not in a per-scope predicate.
- `shouldCreateScope` had one call site in pass1BuildScopes but the
only language implementing it (Python) always returned true. No
producer ever emits a `@scope.block` for Python, so the hook's
"declines to create" branch was unreachable. Other languages
didn't implement it at all.
Removing both:
- Drops the interface declarations in language-provider.ts.
- Drops `shouldCreateScope` from ScopeExtractorHooks Pick and from
the pass1BuildScopes conditional — the stack-based parent-resolve
loop becomes unconditional.
- Drops pythonShouldShadow / pythonShouldCreateScope from simple-hooks,
the Python index barrel, and the python.ts provider wiring.
- Drops the tests that exercised the removed hooks: one block-
suppression scenario in scope-extractor.test.ts, one shouldCreateScope
test in parse-worker-scope-integration.test.ts, and the
pythonShouldShadow / pythonShouldCreateScope always-true assertions
in python-hooks.test.ts. pythonBindingScopeFor's delegate-to-default
test is preserved in its own describe block.
Shadowing itself is unchanged: pythonMergeBindings still runs, LEGB
ordering still applies, wildcard transparency is still handled via
the merge precedence rules. The hook API just no longer has a
vestigial per-scope toggle we decided not to use.
Verification: 204/204 test/integration/resolvers/python.test.ts both
REGISTRY_PRIMARY_PYTHON=0 and =1. 335/335 scope-resolution + graph
unit tests (was 339, net -4 after removing the hook-specific
assertions). tsc clean.
* refactor(scope-resolution): drop dead exports surfaced by knip
Knip flagged 44+ dead exports in the PR surface. Cleanup:
Barrel deletion:
- Remove src/core/ingestion/scope-resolution/index.ts entirely.
It re-exported 30+ symbols but only one file
(languages/python/scope-resolver.ts) imported from it, and only
7 symbols. Matches the project's "no barrel re-exports" preference
and removes a drift surface. scope-resolver.ts now imports from
concrete files (passes/mro.ts, scope/walkers.ts, contract/...).
Dead functions/interfaces removed:
- resolvePythonScope + ResolvePythonScopeInput + ResolvePythonScopeStats
in languages/python/scope-resolver.ts — never called. pipelinePhase
reaches pythonScopeResolver via SCOPE_RESOLVERS, not via a
per-language entry point.
- getScopeResolver in scope-resolution/pipeline/registry.ts — had zero
callers. Consumers read SCOPE_RESOLVERS directly.
Exports demoted to module-internal (used only within their own file):
- PYTHON_SCOPE_QUERY (query.ts) + its re-export from python/index.ts
- PROF (cache-stats.ts)
- PythonArityMetadata (arity-metadata.ts)
- ReferenceSiteSkipSet (graph-bridge/references-to-edges.ts)
- ReceiverBoundProviderSubset (passes/receiver-bound-calls.ts)
- ResolveCompoundReceiverOptions interface (passes/compound-receiver.ts)
- matchingOpenParen function (passes/compound-receiver.ts)
- followChainPostFinalize function (passes/imported-return-types.ts)
- RunScopeResolutionInput + RunScopeResolutionStats (pipeline/run.ts)
Also removed:
- Redundant `export type { Scope }` re-export from contract/scope-resolver.ts
(consumers import Scope directly from gitnexus-shared).
Verification: knip reports zero dead exports in PR-touched files.
204/204 test/integration/resolvers/python.test.ts both flag paths.
335/335 scope-resolution + graph unit tests. tsc clean.
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* feat(cli): gitnexus remove <target> to unindex a registered repo by name or path (#664)
Add a `remove` CLI command that deletes the `.gitnexus/` index AND
unregisters a repo from the global registry (~/.gitnexus/registry.json),
addressing the lifecycle gap flagged in #664: previously users had to
cd into the repo to run `clean`, and there was no path-based or
alias-based remove for an already-deleted working tree.
- New command `gitnexus remove <target> [-f|--force]`. `<target>` is
alias / basename-derived name / remote-inferred name / absolute path.
- New helper `resolveRegistryEntry(entries, target)` in repo-manager.ts
with path > name precedence; throws RegistryNotFoundError or
RegistryAmbiguousTargetError (typed, `kind`-discriminated).
- Atomicity mirrors `clean`: fs.rm first, then unregisterRepo; partial
failures self-heal on next `listRegisteredRepos({ validate: true })`.
- Idempotent on unknown targets (exit 0 with warning) per the #664
spec: "behave atomically and idempotently so retries are safe".
- `--force` uses `clean`-style confirmation-skip semantics — distinct
from `analyze --force` (pipeline re-index); here there is no pipeline
so no conflation.
- 7 new unit tests cover resolver precedence, case sensitivity,
ambiguity, and not-found hints; 2 integration tests cover the real
CLI -> registry -> filesystem chain including the --allow-duplicate-name
(#829) ambiguity case.
* fix(cli): canonicalize repo paths so remove/register match across platforms (#1003 review)
Address review feedback from @evander-wang and @magyargergo on PR #1003
plus the Windows + macOS CI failure (same root cause).
Problem:
- macOS: /var is a symlink to /private/var. `path.resolve` does NOT
follow symlinks, so a child running analyze in /var/folders/X stores
/private/var/folders/X (realpath from OS cwd) but an outer caller
passing the symlink form misses.
- Windows: GitHub runners surface tmpdirs in 8.3 short-name form
(RUNNERA~1) while process.cwd() returns the long form (runneradmin).
Same divergence.
Fix: new `canonicalizePath(p)` helper wraps `path.resolve` plus
`fs.realpathSync.native`, falling back to `path.resolve` when the path
doesn't exist (preserves idempotent-on-missing semantics needed by
`remove <unknown>`). Applied at 3 call-sites — registerRepo,
unregisterRepo, resolveRegistryEntry — canonicalising BOTH the input
and each stored `entry.path` at compare time. That last bit is the
backward-compat story: registries written by older versions
(pre-canonicalisation) still match correctly, so we don't need a
migration script.
Test side: the ambiguous-target integration test now reads the path
from the registry snapshot rather than passing the outer `repoA`
variable directly, so it exercises the registry contract regardless of
which path form the platform stores. 4 new unit tests cover the helper
(idempotent, fallback-on-missing, absolute-for-relative) plus the
backward-compat resolver path.
* fix(cli): store resolved (non-canonical) path, compare via canonicalizePath (#1003 CI)
Follow-up to c5eceba0. The previous commit canonicalised the repo path
at BOTH write-time AND compare-time in registerRepo — that expanded
Windows 8.3 short names (RUNNER~1) to long names (runneradmin) when
storing `entry.path`. Pre-existing #829 unit tests that assert
`path.resolve(err.existingPath) === path.resolve(tmpPath)` then broke
because `tmpPath` is still short-form (path.resolve doesn't expand
8.3) while `entry.path` was long-form (canonicalizePath does).
Fix: split storage from comparison.
- entry.path stores `path.resolve(repoPath)` — whatever form the
caller passed. `list` output and error messages show the path the
user typed.
- All compare points (existing-entry lookup in registerRepo, the
collision guard, unregisterRepo, resolveRegistryEntry path tier)
canonicalise BOTH sides via `canonicalizePath`. That is where the
/var ↔ /private/var and RUNNER~1 ↔ runneradmin divergence actually
matters.
Net effect: storage is tolerant (preserves user input), matching is
strict (canonical-vs-canonical). Pre-existing #829 tests stay green
because `err.existingPath` is unchanged from what `path.resolve` gives
back; the cross-platform CI failure from #1003 stays fixed because
every comparison path goes through `canonicalizePath`.
* fix(cli): refuse destructive fs.rm when registry storagePath isn't <repo>/.gitnexus (#1003 review)
Address @magyargergo's inline review finding on remove.ts:89 and the
sibling vulnerability in clean.ts --all (caught during a pre-commit
safety audit). ~/.gitnexus/registry.json is a user-writable plain-text
file, so a corrupted or hand-edited entry could point storagePath at
the repo root (catastrophic: rm the working tree), an empty string
(→ cwd), a parent dir, or anywhere else. fs.rm(recursive: true,
force: true) on any of those is a runtime disaster.
- New UnsafeStoragePathError + exported assertSafeStoragePath() in
repo-manager.ts. Pure lexical string check (Windows-case-
insensitive) asserting entry.storagePath === path.join(entry.path,
'.gitnexus').
- Guard wired into BOTH destructive registry-trusting sites:
- remove.ts: exit 1 with actionable hint
- clean.ts --all: skip the poisoned entry with a warning and
continue (preserves existing per-repo error tolerance — one bad
entry doesn't halt the batch)
- clean.ts default path and server/api.ts are safe-by-construction
(they recompute storagePath from findRepo / getStoragePath rather
than trusting the registry field).
- 8 unit tests cover the guard (valid, repo-root, parent, empty,
unrelated, sibling, error payload, Windows case).
- 2 integration tests prove the full CLI path: remove-poisoned exits
1 without touching the working tree; clean --all with a poisoned
sibling entry cleans the good entry, skips the bad one, and leaves
the poisoned repo intact.
* test(cli): assert full remove dry-run + success output shape (#1003 NIT)
Address the one NIT from the senior-reviewer pass on PR #1003: the
integration test was only checking for the "Run with --force" hint in
dry-run output, not verifying that the three actual console.log lines
(alias, repo path, storage path) appear. Same weak check on the
success-branch "Removed" output.
Tighten both assertions to toContain(alias), toContain(entry.path),
toContain(storagePath). Catches silent format regressions — e.g. a
future refactor that drops a console.log line or swaps
entry.name/entry.path in the output.
No code change; +20 test lines. All assertions in the happy-path
integration test now fire for a meaningful reason.
When the Phase 1 local-impact leg returned a structured { error: ... }
payload (missing symbol, graph-load failure, or an exception wrapped by
safeLocalImpact), runGroupImpact previously buried it inside a zero-hit
GroupImpactResult with empty cross / outOfScope arrays and risk 'UNKNOWN'.
Callers branch on top-level `error` (CLI, MCP wrapper), so the failure
path surfaced as a silent "no impact across the group" — a false
negative on a safety-critical blast-radius tool.
Fail closed: bubble the error as a top-level { error } prefixed with the
repoPath, matching how runGroupImpact already handles resolveGroupRepo,
config-load, and bridgePrep failures. Chose option 1 (bubble the error)
over option 2 (partial-result discriminant) because runGroupImpact only
runs local impact for a single member repo at this point — cross-repo
fan-out happens later via the bridge, so there is no partial success
data to preserve on the local-phase failure path.
Added two regression tests covering both the port-returned { error }
case and the thrown-exception case (wrapped by safeLocalImpact).
Made-with: Cursor
* docs(group): add gRPC microservices group guide (#906)
Adds `docs/guides/microservices-grpc.md`, a walkthrough for using
GitNexus across multiple repositories whose services communicate over
gRPC. Covers the group mental model, per-repo `gitnexus analyze`, the
`group.yaml` schema, `group sync`, inspecting `contracts.json`,
running cross-repo `impact` with `@<group>` routing, the gRPC
extractor's provider/consumer signals per language, the
`config.links` manifest escape hatch, and a short troubleshooting
list. Wires the new page from the group-mode note in AGENTS.md.
Closes#906.
Made-with: Cursor
* docs(grpc-guide): drop hard line wraps, rely on editor soft wrap
Made-with: Cursor
vite.config.ts reads engines.node from ../gitnexus/package.json,
but Dockerfile.web only copied gitnexus-shared and gitnexus-web,
causing the build to fail with "Cannot find module" during
`npm run build --prefix gitnexus-web`.
Co-authored-by: wangjichao <wangjichao@inke.cn>
In a reusable workflow, github.event_name inherits the caller's
event (e.g. "push"), not "workflow_call". This caused the
type=raw tag to be disabled when docker.yml was called from
release-candidate.yml, producing no Docker tags at all and
failing the build.
Fix: check `inputs.tag != ''` instead, since inputs.tag is only
populated for workflow_call invocations.
Co-authored-by: wangjichao <wangjichao@inke.cn>
Extend the PHP tree-sitter plugin to emit consumer HttpDetections for
three common PHP HTTP call shapes, matching Node plugin parity:
- Laravel HTTP client: Http::get/post/put/delete/patch($url)
- Guzzle / generic: $client->get/post/...($url)
- file_get_contents($url) when the URL is absolute http(s)://
String-literal URLs only. Paths built via binary concatenation
(`$base . '/path'`), sprintf, or config lookups are intentionally
deferred — they need constant-folding of the enclosing scope to be
useful and are tracked as follow-up work.
Refs #992
Co-authored-by: Jonas Vanderhaegen <jonasvanderh+claude.ai@gmail.com>
* fix(bm25): return FTS-matched symbols instead of arbitrary LIMIT 3 nodes
Previously, bm25Search fetched up to 3 arbitrary symbols from the matched
file using MATCH (n) WHERE n.filePath = $filePath LIMIT 3 (no ORDER BY).
This meant the specific function or class that actually scored highest in
the BM25 index could be completely absent from the results.
Fix: propagate nodeId from each FTS hit through searchFTSFromLbug, then
use those nodeIds in bm25Search to look up the exact matched nodes via
WHERE n.id IN $nodeIds. Falls back to the old filePath-based lookup when
nodeIds are unavailable.
Also switches the per-file score aggregation from naive sum-of-all to
sum-of-top-3, which prevents files with many mediocre matches (e.g. test
files) from outranking files with a single highly-relevant symbol.
* test(bm25): add unit tests for top-3 aggregation and nodeIds propagation
Covers the new logic paths added in the previous commit:
- top-3 score aggregation (file with 5+ matches → only top-3 contribute)
- nodeIds propagation through BM25SearchResult
- empty nodeId filtering
- cross-table merge for the same file
- result ranking by aggregated score
Also fixes in-place entries.sort() mutation (bm25-index.ts:125) to use
[...entries].sort() so the Map value is not silently modified.
* style: apply prettier formatting
* fix(test): use importOriginal to avoid missing export errors in vi.mock
* fix(bm25): align queryFTSViaExecutor nodeId extraction to match lbug-adapter
Use node.nodeId || node.id || '' in queryFTSViaExecutor to match the
fallback logic in lbug-adapter.ts:1040. Without this, the MCP pool path
could silently return empty nodeIds if LadybugDB surfaces the node id
under node.nodeId rather than node.id.
---------
Co-authored-by: jisue0224 <>
findFunctionNode and findDeclarationNode had no depth limit, causing
stack overflow on deeply nested or auto-generated ASTs, especially
when --stack-size is not applied (e.g. heap already large enough to
skip ensureHeap re-exec).
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(embeddings): structural chunking with data-driven CHUNKING_RULES dispatch
Replace hardcoded label comparisons with a CHUNKING_RULES lookup table
that drives chunking strategy and text generation. Key changes:
- Data-driven dispatch: CHUNKING_RULES table maps labels to chunking
mode (ast-function / ast-declaration), prefix/suffix, field grouping,
and structural text mode
- Struct support: add Struct to AST declaration chunking with field
grouping (same as Class)
- Multi-chunk context: preceding chunk tail (prevTail) injected into
embedding text for cross-chunk coherence
- Version-gated hashes: EMBEDDING_TEXT_VERSION prefix in content hashes
invalidates stale vectors when text template changes
- Compact container context: first declaration line preserved in every
structural chunk for identity
* fix(embeddings): address PR review findings for CHUNKING_RULES refactor
- Remove LABEL_ENUM from STRUCTURAL_LABELS to avoid wasted AST parses
- Add maintenance note about extractStructuralNames and EMBEDDING_TEXT_VERSION
- Clarify CHUNK_MODE_CHARACTER is a no-op in CHUNKING_RULES
- Strengthen EMBEDDING_TEXT_VERSION test assertion to exact value
---------
Co-authored-by: wangjichao <wangjichao@inke.cn>
* feat(ingestion): shadow-mode parity harness + static dashboard (#923, RFC #909 Ring 2 PKG)
Side-car observability for the RFC #909 registry rollout. Callers that
dual-run legacy-DAG + `Registry.lookup` feed their result pairs into
the harness; the harness diffs each pair via shared `diffResolutions`
(#918), aggregates via `aggregateDiffs`, and persists a per-language
parity report that the static dashboard can render offline.
## Shipped
### `gitnexus/src/core/ingestion/shadow-harness.ts` (new)
```ts
createShadowHarness(): ShadowHarness
```
API:
- `enabled` — `true` iff `GITNEXUS_SHADOW_MODE` is truthy at
construction. Captured once; later env-var mutations don't flip it.
- `record({ language, callsite, legacy, newResult, primary })` —
accumulator. No-op when `enabled === false` (near-zero overhead).
- `size()` — diagnostic counter.
- `snapshot(now?)` — deterministic `ShadowParityReport` from the
accumulated diffs.
- `persist(outputDir, now?)` — writes BOTH a timestamped
`<runId>.json` and a `latest.json` pointer. Creates outputDir if
absent. Returns the per-run file path.
- `clear()` — resets the accumulator; preserves `enabled`.
Activation: `GITNEXUS_SHADOW_MODE` accepts `'true'` / `'1'` / `'yes'`
(case-insensitive, trimmed); same truthy convention as
`REGISTRY_PRIMARY_<LANG>` from #924. Typos → disabled (fail-safe).
Persisted payload (`PersistedShadowReport`) is schema-versioned (`v1`):
```jsonc
{
"schemaVersion": 1,
"runId": "YYYYMMDD-HHMMSS-xxxxxxxx",
"generatedAt": "ISO 8601",
"primaryByLanguage": { "python": "legacy", ... },
"report": { /* ShadowParityReport from #918 aggregateDiffs */ }
}
```
`runId` prefix is the timestamp so files sort chronologically; the
entropy suffix prevents collisions within a clock-second.
### `gitnexus/shadow-parity-dashboard/index.html` (new)
Minimal static dashboard — one HTML file, zero build step, zero runtime
deps. Fetches `./latest.json` and renders:
- Overall summary cards (total calls, both agree, disagree, overall parity %)
- Per-language table: language tag ("primary: legacy" / "primary:
registry" pill) + total / agree / only-legacy / only-new / disagree
/ both-empty / parity%
- Parity cells colored by threshold: ≥95% green, ≥80% amber, <80% red
- Light / dark via `prefers-color-scheme`
- Empty-state message when no records yet
File-serving is static: `cp .gitnexus/shadow-parity/latest.json
gitnexus/shadow-parity-dashboard/` + open in a browser.
## Tests (14, all passing)
- **Flag detection** (5): default off · truthy variants case-insensitive ·
falsy / typo → off · record() is no-op when disabled · env flip
AFTER construction doesn't enable (constructed-once semantics)
- **Record + snapshot** (4): multi-language accumulation ·
per-language rows with correct outcomes · snapshot determinism ·
`clear()` resets accumulator + `primaryByLanguage`
- **Persistence** (5): mkdir-p on missing outputDir · per-run +
latest.json match byte-for-byte · schema v1 payload shape ·
runId timestamp prefix sorts chronologically · empty report
persists gracefully
Tests use a per-test tmpdir (`fs.mkdtemp`), cleaned in `afterEach`,
so parallel vitest runs don't collide. `GITNEXUS_SHADOW_MODE` is
saved + restored per-test.
## What's deliberately NOT in this PR (call-out in harness docstring)
- **Dual-run dispatch.** The harness is a side-car — it does NOT
invoke either resolution path. Call-processor integration that
actually runs both legacy + registry paths lands as a follow-up.
Without that integration, `record()` is never called in production
today. The harness is tested in isolation with synthetic inputs.
- **CI artifact publishing.** Config work to upload
`latest.json` + the dashboard HTML per CI run. Tracked separately;
the harness + dashboard are ready when the CI job wires in.
- **Fixture-level drill-down.** The issue mentions per-fixture AST
snippet + evidence trace drill-down. MVP dashboard shows per-language
rows only; drill-down extends the static JSON format + the dashboard
JS in a focused follow-up.
## Verification
- `tsc --noEmit` clean (both `gitnexus-shared` and `gitnexus`)
- 14/14 new tests pass
- Full scope-resolution / shadow / model / flag suite: **335/335 pass**
## Part of
- Parent: #909
- Depends on (code): #917 (registries), #918 (diff + aggregate)
- Unblocks Ring 3 language flips: the parity dashboard becomes the
checkpoint before flipping `REGISTRY_PRIMARY_<LANG>=true` for a
language — once per-language parity stabilizes, the flip ships.
* chore: prettier format on shadow-parity-dashboard index.html
Bridges the CLI's existing per-language `ImportResolverFn`s (16 languages
already implemented) to the shared `FinalizeHooks.resolveImportTarget`
contract consumed by `finalize()` (#915) and
`finalizeScopeModel` (#921).
No resolver logic is reimplemented — the adapter wraps
`provider.importResolver` from each `LanguageProvider` verbatim.
## Shipped
### `import-target-adapter.ts` (new)
```ts
buildImportTargetWorkspace(providers, resolveCtx): ImportTargetWorkspace
resolveImportTargetAcrossLanguages(targetRaw, fromFile, workspaceIndex): string | null
```
- `ImportTargetWorkspace` is the opaque `workspaceIndex` shape the
adapter recognizes: `{ perLanguage: Map<SupportedLanguages,
{ resolver, ctx }> }`. Callers build it once per ingestion run from
the active language providers.
- `resolveImportTargetAcrossLanguages` is the `FinalizeHook`
implementation. It:
1. Reads `getLanguageFromFilename(fromFile)`.
2. Looks up the per-language entry.
3. Calls the existing `ImportResolverFn` — same signature, same
code path the legacy DAG uses today.
4. Picks `result.files[0]` (covers both `'files'` and `'package'`
result kinds; the legacy pipeline's richer multi-file + dirSuffix
semantics stay accessible through `importResolver` directly).
5. Returns `null` on any null result, empty files[], unknown
extension, missing workspace, or resolver exception.
- Exceptions from resolvers are swallowed — the finalize algorithm
treats `null` as `linkStatus: 'unresolved'`, which is the right
fallback for malformed inputs.
### What's deliberately NOT here
- **Re-implementation of any per-language resolver.** Wraps the
existing `importResolver` field on each provider.
- **Dynamic-import handling.** The shared finalize algorithm short-
circuits `ParsedImport { kind: 'dynamic-unresolved' }` before
calling `resolveImportTarget`, so the adapter never sees them.
- **`importPathPreprocessor`.** Preprocessing belongs inside the
provider's `interpretImport` hook that produces
`ParsedImport.targetRaw`; the adapter forwards that verbatim.
## Tests (12, all passing)
- **`buildImportTargetWorkspace`** (3): registers providers with
importResolver · skips providers without · threads shared ctx
into every entry
- **`resolveImportTargetAcrossLanguages`** (9): forwards targetRaw +
fromFile · dispatches by extension · null resolver result →
null · `package`-kind takes first file · empty files[] → null ·
no registered resolver → null · unknown extension → null ·
undefined/malformed workspace → null · resolver throw → null
Real per-language resolver correctness is covered by the existing
per-language resolver test suites — the adapter is the bridge layer.
## Verification
- `tsc --noEmit` clean (both `gitnexus-shared` and `gitnexus`)
- `gitnexus-shared` build clean
- 12/12 new tests pass
- Full scope-resolution / shadow / model / flag suite: **333/333 pass**
## Integration flow
```ts
const workspace = buildImportTargetWorkspace(providers, resolveCtx);
const indexes = finalizeScopeModel(parsedFiles, {
hooks: { resolveImportTarget: resolveImportTargetAcrossLanguages },
workspaceIndex: workspace,
});
model.attachScopeIndexes(indexes);
```
## Closes part of #909. Unblocks
- Ring 3 language migrations (#926+): a language flipping to
`REGISTRY_PRIMARY_<LANG>=true` now has correct import-target
resolution out of the box via its existing `importResolver`.
- #923 shadow harness — can run the dual-path comparison knowing
both sides use the same per-language resolution semantics.
Ties the Ring 2 pipeline together. Takes the `ParsedFile[]` produced by
#920's parse-worker integration, feeds them to shared `finalize()`
(#915), and bundles every workspace-wide index for attachment onto
`MutableSemanticModel`. Thin integration glue per issue #884's boundary
— all algorithm lives in `gitnexus-shared`.
## Shipped
### `model/scope-resolution-indexes.ts` (new)
```ts
interface ScopeResolutionIndexes {
readonly scopeTree: ScopeTree;
readonly defs: DefIndex;
readonly qualifiedNames: QualifiedNameIndex;
readonly moduleScopes: ModuleScopeIndex;
readonly methodDispatch: MethodDispatchIndex;
readonly imports: ReadonlyMap<ScopeId, readonly ImportEdge[]>;
readonly bindings: ReadonlyMap<ScopeId, ReadonlyMap<string, readonly BindingRef[]>>;
readonly referenceSites: readonly ReferenceSite[];
readonly sccs: readonly FinalizedScc[];
readonly stats: FinalizeStats;
}
```
The bundle produced by the orchestrator, consumed by the resolution
phase. `ReferenceIndex` is deliberately NOT here — it's populated in
the next phase (#925).
### `model/semantic-model.ts` — extended
- `SemanticModel.scopes?: ScopeResolutionIndexes` — undefined until
attached; once attached, frozen.
- `MutableSemanticModel.attachScopeIndexes(indexes)` — one-shot write.
Throws on second call; `Object.freeze`s the bundle on write. `clear()`
resets the slot back to `undefined` so re-ingestion can re-attach.
### `finalize-orchestrator.ts` (new)
```ts
finalizeScopeModel(parsedFiles, options?): ScopeResolutionIndexes
```
Orchestration steps:
1. Map `ParsedFile[]` → `FinalizeInput` (`FinalizeFile` is a structural
subset, so no shape-shifting).
2. Call shared `finalize()` with provider hooks (defaults provided for
the zero-provider case today).
3. Build the four workspace indexes (`DefIndex`, `QualifiedNameIndex`,
`ModuleScopeIndex`, `ScopeTree`) from per-file unions.
4. Build an empty `MethodDispatchIndex` as a placeholder (owners=[],
both callbacks return []). Real MRO wiring lands with the
per-language adapters in #922.
5. Bundle + return.
**Empty-input safety.** Zero parsedFiles → valid but empty bundle with
all zero-sized indexes and `stats.totalFiles === 0`. Downstream code
can consult `model.scopes` without branching on presence — only on
`stats`.
**Hook defaults** (`withDefaultHooks`) for missing provider hooks:
- `resolveImportTarget: () => null` — every import goes `unresolved`
- `expandsWildcardTo: () => []` — wildcards don't materialize
- `mergeBindings: (a, b) => [...a, ...b]` — append without precedence
Providers override these in #922 (per-language import adapters).
## Tests (10, all passing)
- **Empty input** (1): zero parsedFiles → valid empty bundle
- **Single file** (2): all per-file indexes populated · referenceSites
aggregated
- **Cross-file imports** (3): resolveImportTarget threads through +
links · default-null resolver → unresolved · stats reflect graph
- **MutableSemanticModel integration** (4): undefined initially · attach
once · Object.freeze applied · throws on re-attach · clear() resets
## Verification
- `tsc --noEmit` clean in both packages
- `gitnexus-shared` build clean
- 10/10 new tests pass
- Full scope-resolution / shadow / model / flag suite: **321/321 pass**
## What's deferred (not this PR, per RFC #909 scope)
- **Per-language hook adapters** (#922): `resolveImportTarget` +
`expandsWildcardTo` + `mergeBindings` wired per language.
- **MethodDispatchIndex wiring via HeritageMap**: populate MRO + implements
via the existing CLI-package HeritageMap strategies. Likely companion
to #922 or a focused follow-up.
- **Pipeline invocation**: actually calling `finalizeScopeModel` from
the real ingestion pipeline. The orchestrator is callable today; the
ingestion entry point wiring lands with the shadow harness (#923).
- **`ReferenceIndex` population**: RFC §3.2 Phase 4 / #925.
## Closes part of #909. Unblocks
- #923 shadow harness — now has a fully materialized `model.scopes` to
query against the legacy DAG for parity measurement
- #925 ReferenceIndex → LadybugDB emission — consumes `model.scopes`
- Ring 3 language migrations (#926+) — a language flipping to
`REGISTRY_PRIMARY_<LANG>=true` can now expect `model.scopes` to be
populated when the pipeline wires the orchestrator in
Plumbs the ScopeExtractor (#919) into the real parsing pipeline.
`ParsedFile` artifacts now flow from workers to the parsing-processor
without changing any legacy-DAG behavior.
## Shipped
### `gitnexus/src/core/ingestion/scope-extractor-bridge.ts` (new)
- `extractParsedFile(provider, sourceText, filePath, onWarn?)`
- Short-circuits (returns `undefined`) when the provider has not
implemented `emitScopeCaptures`. True for every language today —
this is the default no-op path.
- Invokes the hook + `ScopeExtractor.extract`, returns a `ParsedFile`.
- **Swallows exceptions on both sides.** Failures route through the
optional `onWarn` callback (or `console.warn`) and return
`undefined`. Scope-extraction errors NEVER break legacy parsing on
the same file.
- Standalone module (not nested in `parse-worker.ts`) so tests can
import it directly without triggering the worker's top-level
`parentPort!.on(...)`.
### `gitnexus/src/core/ingestion/workers/parse-worker.ts`
- `ParseWorkerResult.parsedFiles: ParsedFile[]` added.
- `processFileGroup` calls `extractParsedFile` AFTER tree parse,
BEFORE legacy extraction. Worker provides an `onWarn` callback that
routes bridge warnings through `parentPort.postMessage({ type:
'warning', message })`.
- `mergeResult` includes `parsedFiles` in the sub-batch merge.
- Initial + reset accumulator templates include `parsedFiles: []`.
### `gitnexus/src/core/ingestion/parsing-processor.ts`
- `WorkerExtractedData.parsedFiles: ParsedFile[]` added.
- Empty-result branch and the across-chunk aggregation both include
`parsedFiles`. Aggregation is tolerant of workers that don't emit
the field (older builds / partial rollouts).
### Ring 1 tweak: `emitScopeCaptures` sync return
`readonly CaptureMatch[]` (was `Promise<readonly CaptureMatch[]>`).
Tree-sitter and COBOL's regex tagger are both synchronous; no
foreseeable need for async work inside this hook. Sync lets the
already-sync worker pipeline invoke it inline without cascading
`async` up through the batch driver + IPC handler.
## Tests (9 new; full suite 311/311)
`gitnexus/test/unit/scope-resolution/parse-worker-scope-integration.test.ts`:
- Not-migrated (2): undefined-returning hook · never-invokes-extractor
- Migrated (3): happy path · argument threading · honors
`shouldCreateScope` override
- Error resilience (4): hook throws · extractor throws (no Module) ·
extractor throws (sibling overlap) · `onWarn` gets routed
message with filePath + error body
## Verification
- `tsc --noEmit` clean in both packages
- `gitnexus-shared` build clean
- 311/311 combined scope-resolution / shadow / model / flag suite
- 9/9 new bridge tests
## What's NOT in this PR (still deferred to #921)
- Actually using the `parsedFiles` — that's the finalize orchestrator.
- `ModuleScopeIndex.byFilePath` materialization — belongs alongside
the rest of the SemanticModel indexes in #921.
## Closes part of #909. Unblocks
- #921 finalize-orchestrator — consumes `WorkerExtractedData.parsedFiles`
Adds the per-language feature flag primitive that gates the Ring 3
registry-primary rollout. Single source of truth for whether a given
language uses `Registry.lookup` (new) or the legacy DAG (current).
## Shipped
### `gitnexus/src/core/ingestion/registry-primary-flag.ts`
- `isRegistryPrimary(lang): boolean` — reads
`REGISTRY_PRIMARY_<UPPER(enum-value)>` from `process.env`.
- `envVarNameFor(lang): string` — exposed for CI tooling that
cross-references flag flips (and for test assertions).
- `primaryLanguages(): ReadonlySet<SupportedLanguages>` — all
currently-on languages; useful for startup logging + the #923
shadow dashboard which distinguishes "primary: legacy" vs
"primary: registry" rows.
### Contract
- Default: `false` for every language. A language must explicitly
opt in by setting its env var.
- Truthy: `'true'`, `'1'`, `'yes'` (case-insensitive, whitespace-
trimmed). Anything else — typos, empty string, `'off'` — is
`false`. Fail-safe posture: a misspelled flag doesn't accidentally
flip a language.
- No per-process caching. `process.env` is read per call; overhead
is negligible (one lookup per file at resolution time), and
test isolation is lexical (no cache-reset coordination).
### Env-var mapping
Uses the enum VALUE, not the TS key, for the env-var suffix:
- `SupportedLanguages.Python` → `REGISTRY_PRIMARY_PYTHON`
- `SupportedLanguages.CPlusPlus` → `REGISTRY_PRIMARY_CPP` (value `'cpp'`)
- `SupportedLanguages.CSharp` → `REGISTRY_PRIMARY_CSHARP`
Users flip languages by their canonical name, not the TS symbol.
## Tests (16, all passing)
- `envVarNameFor` (3): upper-casing · enum-VALUE-not-KEY mapping ·
all-languages uniqueness smoke-test
- `isRegistryPrimary` (9): default false · `'true'` / `'1'` / `'yes'`
truthy · mixed-case + whitespace-padded · falsy-looking values ·
unrecognized tokens (typo-safe) · per-language isolation · no
stale cache on mid-process mutation · CPlusPlus mapping
- `primaryLanguages` (3): empty · exact membership · Set instanceof
Tests scrub every `REGISTRY_PRIMARY_*` env var in `beforeEach` +
`afterEach` so parallel vitest runs on the same process don't bleed state.
## What's NOT in this PR (deferred by design)
The actual integration in `call-processor.ts` belongs in #921
(finalize-orchestrator). Reason: the "new path" requires a populated
`SemanticModel` to call `Registry.lookup` against, and the model
becomes accessible only after #921 orchestrates finalize. Wiring a
dead branch now would just get rewritten then.
This PR ships the flag primitive in isolation so #921 has a clean,
tested utility to consult — and so `#923` (shadow harness) has a
stable boolean to read for its "which row is primary?" rendering.
## Closes part of #909. Unblocks
- #921 finalize-orchestrator — can now consult `isRegistryPrimary`
at resolution time
- #923 shadow harness — can distinguish primary-flipped rows
* feat(ingestion): ScopeExtractor driver — 5-pass CaptureMatch → ParsedFile (#919, RFC #909 Ring 2 PKG)
Kicks off Ring 2 PKG. Implements RFC §5.3 + §3.2 Phase 1: the central,
source-agnostic driver that turns a language provider's `CaptureMatch[]`
into a `ParsedFile` — the per-file artifact the finalize orchestrator
(#921) feeds into the shared `finalize()` algorithm (#915).
## Files
### New shared contracts
- `gitnexus-shared/src/scope-resolution/parsed-file.ts`
Per-file extraction artifact: scopes, parsedImports, localDefs,
referenceSites. Structural superset of `FinalizeFile` so the
finalize orchestrator threads `ParsedFile` through unchanged.
- `gitnexus-shared/src/scope-resolution/reference-site.ts`
Pre-resolution usage fact: name, atRange, inScope, kind, optional
callForm/explicitReceiver/arity. Converted to `Reference` records
by the resolution phase (populates `ReferenceIndex`).
### Ring 1 collateral tweak
- `language-provider.ts: emitScopeCaptures` now returns
`Promise<readonly CaptureMatch[]>` (was `readonly Capture[]`).
Pre-grouping per tree-sitter match is the provider's job — the
extractor expects coherent matches, not flat captures. No
consumers yet (all languages still on legacy DAG), so no breakage.
Docstring updated.
### New CLI module
- `gitnexus/src/core/ingestion/scope-extractor.ts`
Single entry point: `extract(matches, filePath, provider): ParsedFile`.
Five-pass pipeline:
Pass 1 — Build scope tree. `@scope.*` → `ScopeDraft[]` via
range-containment parent derivation. Honors
`provider.shouldCreateScope` (skip-but-reparent-children) and
`provider.resolveScopeKind`. Throws `ScopeTreeInvariantError`
via `buildScopeTree` on malformed input.
Pass 2 — Attach declarations + local bindings. `@declaration.*`
→ `SymbolDefinition` + `BindingRef { origin: 'local' }`.
Default attachment: innermost containing scope. Hoisting via
`provider.bindingScopeFor`.
Pass 3 — Collect raw imports. `@import.*` → `ParsedImport` via
`provider.interpretImport`. Attached to ParsedFile
(finalize resolves owning scope in Phase 2).
Pass 4 — Collect type bindings. `@type-binding.*` →
`TypeRef` via `provider.interpretTypeBinding` →
`scope.typeBindings`. Hoistable via `bindingScopeFor`.
Pass 5 — Collect reference sites. `@reference.*` →
`ReferenceSite[]`. Call form from declarative sub-tag
(`@reference.call.member`) or `provider.classifyCallForm`.
### Tests
- `gitnexus/test/unit/scope-resolution/scope-extractor.test.ts`
23 tests organized by pass + one end-to-end fixture exercising
all 5 passes together. MockProvider emits synthetic
`CaptureMatch[]` with no AST — extractor is pure given those.
## Design notes
- **Source-agnostic.** No `Tree` / `SyntaxNode` types leak into the
driver. Works for tree-sitter providers and COBOL's regex tagger.
- **One AST walk per language.** Providers do the walk inside
`emitScopeCaptures`; this driver does zero traversal.
- **Invariants delegated.** `ScopeTree.buildScopeTree` enforces
structural rules (non-Module has parent, parent contains child,
siblings don't overlap). The extractor doesn't try to repair
malformed captures.
- **Sub-tag whitelist.** `@reference.receiver`, `@declaration.name`,
`@import.source`, etc. are known sub-tags — excluded from anchor
selection so the broadest-range heuristic doesn't mis-identify them
as anchors for their topic. Bug surfaced in the end-to-end fixture
test (member call with a large-range receiver) and was fixed before
commit.
## Verification
- `tsc --noEmit` clean (both `gitnexus-shared` and `gitnexus`)
- `gitnexus-shared` build clean
- 23/23 new tests pass
- Full scope-resolution / model / shadow suite: **285/285 pass**
## Closes part of #909. Unblocks
- #920 parse-worker integration (emit ParsedFile from the worker)
- #921 finalize orchestrator (consume ParsedFile[] workspace-wide)
- #922 per-language import adapters
* chore(ingestion): address #919 review findings on the extractor
Addresses all 5 items from the PR #965 review in-PR.
## Structural changes
- **Extract `ScopeExtractorHooks` as the narrow dependency surface.**
The extractor now declares its dependency on a `Pick`-narrowed subset
of `LanguageProvider` (just the 6 scope-resolution hooks it actually
reads). Test mocks implement exactly that interface — no more
`as unknown as LanguageProvider` cast hiding missing-field bugs.
Adding a new hook read becomes a compile error, not a silent test
pass. (Finding 3.2)
- **Remove dead `ownerDefIdFor` stub + `isOwnerKind` helper.** The
function always returned `undefined` with `void innermost; void
drafts;` suppressors — an incomplete-implementation signal. The code
path was also misleading: creating a clone of the def with
`ownerId: undefined` is structurally identical to keeping the
original. Pass 2 now keeps the def as-is. Contract is documented in
a code comment: providers that need `ownerId` set it from their
declaration hook; `finalize` (via #914 `MethodDispatchIndex`) fills
in method/field `ownerId` in a post-extraction pass that has full
def visibility. (Finding 2.1)
- **Standardize `filePath` threading across passes 4 and 5.** Pass 4
was reading `drafts[0]!.filePath`; pass 5 was reading
`anyFilePathFromScopeTree(scopeTree)`. Both equivalent but
inconsistent. Both now take `filePath` as a parameter from the
top-level `extract()` call. The `anyFilePathFromScopeTree` helper is
removed. (Finding 2.2)
## Documentation
- **Snapshot-semantics comment on `scopeTree` + `positionIndex`.** The
hooks called during Passes 2-5 receive a `scopeTree` built BEFORE any
bindings/ownedDefs/typeBindings were written. Hooks MUST NOT rely on
`scope.bindings` etc. being populated — they're for parent/range/kind
queries only. Added a doc block at the `scopeTree`/`positionIndex`
construction site so future Ring 3 implementers don't write a
`classifyCallForm` that reads bindings. (Finding 2.3)
## Tests
- **Regression for the anchor-vs-receiver bug** (Finding 3.1): a
member-call match where `@reference.receiver` spans columns 0-10
(wider) and the call name spans 11-15 (narrower). Without the
`KNOWN_SUB_TAGS` exclusion, the broadest-range heuristic would have
picked the receiver; the test pins that the call name is the one
that ends up in `referenceSites[0].name`.
- **Mock provider now types exactly `ScopeExtractorHooks`**, no more
double-cast. Any future hook added to `extract()` that isn't in
`ScopeExtractorHooks` is a compile error.
## Verification
- `tsc --noEmit` clean in both `gitnexus-shared` and `gitnexus`
- `gitnexus-shared` build clean
- 24/24 scope-extractor tests pass (+1 regression)
- Full scope-resolution / model / shadow suite: **286/286 pass**
* chore(shared): apply Ring 2 SHARED review follow-ups in one diff
Aggregates all actionable follow-ups from the 9 Ring 2 SHARED PRs
(#949–#963) before proceeding to Ring 2 PKG. No behavior changes;
docstring edits, test refinements, and one structural cleanup.
## #913 (DefIndex / ModuleScopeIndex / QualifiedNameIndex)
- Rename `freezeIndex` → `wrapIndex` across all three index builders.
The old name implied `Object.freeze` on the wrapper, which we never
applied; `wrapIndex` more accurately describes the lightweight
readonly-interface wrap. Safety surface (frozen bucket arrays,
frozen miss-empty array, readonly Maps) is unchanged.
- Document in `buildModuleScopeIndex` JSDoc that callers must
pre-normalize `filePath` keys (no path-separator canonicalization
happens here). Prevents silent cross-platform misses.
- Add an explicit hit-path freeze assertion in
`qualified-name-index.test.ts` (the existing test covered only the
miss-path `EMPTY` array).
## #914 (MethodDispatchIndex)
- Differentiate the C3 and BFS test cases: both tests now use
distinct MRO orderings so they prove the materializer stores
whatever order the `computeMro` callback produces (not that C3 and
BFS yield identical output).
- Add `implementsOfCalls` counter in the first-write-wins test, and
document the call-count contract in `MethodDispatchInput.implementsOf`
JSDoc: `implementsOf` fires **per occurrence** in `input.owners`
(not per unique owner); `computeMro` fires at most once per unique
owner. Callers with expensive `implementsOf` implementations should
pre-dedupe `owners`.
## #916 (resolveTypeRef)
- Document the deliberate exclusion of `'Type'` from `TYPE_KINDS`
(verified no extractor in `gitnexus/src/core/ingestion/` emits
`type: 'Type'` for annotation-relevant symbols).
- Rename the namespace-origin test from `'resolves ...'` to
`'returns null for a namespace-origin binding whose def is not a
type kind'`, matching the failure-case intent.
## #918 (shadow diff + aggregate)
- Remove the partial re-export `export type { ShadowAgreement, ShadowDiff };`
from `aggregate.ts` — it omitted `ShadowCallsite` and diverged
from the top-level barrel. Consumers import all three from the
`gitnexus-shared` entry point.
- Fix the invalid `'wildcard'` evidence kind in `diff.test.ts` fixture
(that kind is not a valid `ResolutionEvidence.kind`). Replaced with
`'global-name'`, a real kind the test treats identically.
## #912 (ScopeTree / PositionIndex / makeScopeId)
- Document the touching-boundary semantics on `PositionIndex.atPosition`:
when siblings share a boundary point, the right (later-start) sibling
wins per the existing innermost-wins sort contract.
- Resolve the layer-inversion flagged by review: move `ScopeLookup`
from `resolve-type-ref.ts` to `types.ts` (its natural home in the
data-model layer). `scope-tree.ts` now imports `ScopeLookup` from
`types.js` directly; the old re-export from `resolve-type-ref.ts`
is removed per repo convention (`feedback_no_reexport`). Barrel
export moved alongside.
## #917 (ClassRegistry / MethodRegistry / FieldRegistry)
- Replace the dangling "try a name-match among class-like defs"
comment in `lookupReceiverType` with explicit prose that callers
must pre-resolve via `resolveTypeRef` if they want richer semantics.
No behavior change — the function already returned `undefined` on
ambiguous/missing qnames.
- Fix `tieBreakKey.origin` default for pure Step-2 candidates.
Type-binding-only hits no longer falsely inherit `'local'` from
`ensureCandidate`'s neutral default; they now demote to `'import'`
on their first type-binding hit, and only a later Step-1 lexical
hit can upgrade them back to `'local'`. Keeps the Appendix B
cascade faithful to the true origin.
- Document `'global-name'` in `evidence.ts`: currently reserved for
Ring 3's byName global index; `lookupCore` never emits it today.
The weight stays live so `composeEvidence` remains exhaustive over
the origin union.
- Rename the mislabeled Step-7 test from `'confidence DESC is the
primary key'` (which actually tested hard-shadow baseline) to
`'inner scope shadows outer, yielding single result'`, and add a
separate test that actually exercises multi-candidate confidence
ordering (local vs wildcard at the same scope).
## Verification
- `tsc --noEmit` clean (both `gitnexus-shared` and `gitnexus`)
- `gitnexus-shared` build clean
- Combined scope-resolution / model / shadow suite: **260/260 pass**
(+1 from the new multi-candidate ordering test in #917)
## Not addressed (non-actionable)
- #949 CI "failure with zero failing tests": pre-existing Swift Node 22
grammar flake unrelated to #910 scope.
- #950: the two non-blocking findings were already addressed in
follow-up commit `cbac32ba` (ParsedImport discriminated union +
`ScopeId | null` on the two hooks).
- #915: the five in-scope findings were already addressed in
follow-up commit `54515a7e` (dead code, unused params, multi-hop
docs, cap-hit test, stats granularity).
- #915 LanguageProvider.resolveImportTarget signature divergence +
`findDefById` O(F×D) perf: tracked separately as follow-up issues
for the Ring 3 migration window.
* chore(shared): address ce:review findings on the follow-up diff
ce:review (interactive) on PR #964 surfaced two P2s and several P3s. This
commit applies all `safe_auto` fixes + both manual tests in-line so the
PR ships with a cleaner review trail.
## P2 fixes
- **Complete `freezeIndex` → `wrapIndex` rename.** The prior commit renamed
3 of 5 sibling index files; `method-dispatch-index.ts` and
`position-index.ts` still carried the old name. Now all 5 helpers use
the consistent `wrapIndex` naming.
(maintainability + project-standards reviewers both flagged this.)
- **Add regression tests for the `recordTypeBindingHit` origin demotion.**
The prior commit introduced the `tieBreakKey.origin = 'import'`
demotion for Step-2-only candidates without a direct test. Added:
- `registries.test.ts`: two Step-2-only siblings under the same
interface, asserting deterministic DefId.localeCompare tie-break
AND the stronger invariant that composeEvidence never emits a
where-found signal for Step-2-only candidates (no `signals.origin`).
- `position-index.test.ts`: touching-boundary test proving the
right-sibling-wins rule documented in the new JSDoc.
(testing + kieran-typescript + api-contract reviewers all flagged these gaps.)
## P3 fixes
- Fix wrong comment in `recordTypeBindingHit` that claimed Step 1 could
later upgrade a demoted origin. Step 1 runs BEFORE Step 2 — the actual
upgrade path is Step 3 (`seedFromOwnerScopedContributor`). Comment now
describes execution order correctly.
- Fix inaccurate "re-exported there" comment in `index.ts`. `types.ts`
*defines* ScopeLookup natively; it's not a re-export. Phrasing now
says "defined in types.ts and exported from the type-export block
above — not from this module."
- Update stale `scope-tree.ts` file-header prose that still referenced
`ScopeLookup` as living in #916/resolve-type-ref.ts. Now points to
`./types.js` with a cross-ref to both #916 and #917 consumers.
- Expand `atPosition` touching-boundary JSDoc to name the mechanism
(backward scan through start-sorted array) so readers can trace the
binary-search code to the claim.
- Add breadcrumb to `aggregate.ts` module header pointing future readers
to `./diff.ts` / the top-level barrel for `ShadowAgreement`,
`ShadowCallsite`, and `ShadowDiff`.
- Remove unnecessary non-null assertion in `recordTypeBindingHit`. Local
`const existingMroDepth = ...` lets TS narrow to `number` in the
else-branch, eliminating the `!` without behavior change.
## Verification
- `tsc --noEmit` clean (both `gitnexus-shared` and `gitnexus`)
- `gitnexus-shared` build clean
- Combined scope-resolution / model / shadow suite: **262/262 pass** (+2
from the new origin-demotion + touching-boundary regression tests)
Capstone of Ring 2 SHARED. Implements RFC §4 — the shared, scope-aware
resolution surface the rest of the semantic model feeds into.
## Modules (`gitnexus-shared/src/scope-resolution/registries/`)
- `context.ts` — `RegistryContext` bundling ScopeTree / DefIndex
/ QualifiedNameIndex / ModuleScopeIndex /
MethodDispatchIndex + provider hooks.
Narrows Ring 1's opaque `RegistryContributor`
to concrete `OwnerScopedContributor`.
- `tie-breaks.ts` — `compareByConfidenceWithTiebreaks`, the RFC
Appendix B cascade: confidence DESC → scope
depth ASC → MRO depth ASC → ORIGIN_PRIORITY
ASC → DefId.localeCompare.
- `evidence.ts` — `composeEvidence(signals)` / `confidenceFromEvidence`.
Translates raw walk signals into the typed
`ResolutionEvidence[]` using authoritative
`EvidenceWeights`. No magic numbers.
- `lookup-qualified.ts`— RFC §4.5. Qualified-name fast path consumed
by `resolveTypeRef` dotted fallback and by
Step 6 of lookup-core.
- `lookup-core.ts` — The 7-step canonical algorithm. Pure. Param-
eterized by `CoreLookupParams`.
- `{class,method,field}-registry.ts`
— Thin wrappers over `lookupCore` that fix
`acceptedKinds` + `useReceiverTypeBinding` per
kind. `buildClassRegistry` / `buildMethodRegistry`
/ `buildFieldRegistry` factory functions.
## RFC §4.2 algorithm contract (honored verbatim)
1. Lexical scope-chain walk. Hard shadow on any `scope.bindings.has(name)`
regardless of kind survivorship.
2. Type-binding resolution (methods/fields only, opt-in via
`useReceiverTypeBinding`). MRO walk via `MethodDispatchIndex.mroFor`.
MRO-depth-decayed weight via `typeBindingWeightAtDepth`.
3. Owner-scoped contributor — when the caller knows the receiver owner,
its direct members merge in as `origin: 'local'`.
4. Kind filter — `acceptedKinds` per registry; `kind-match` evidence
at weight 0 is always emitted for debuggability.
5. Arity filter — `provider.arityCompatibility` per candidate. When at
least one compatible candidate exists, incompatibles are dropped;
otherwise the −0.15 penalty alone disambiguates (they stay in the
result, just ranked lower).
6. Global fallback — fires only when Steps 1-3 produced NO candidates
AND the name is dotted. Delegates to `lookupQualified`.
7. Rank + tie-break — evidence list sorted by the Appendix B cascade.
## §4.7 invariants asserted in tests
- No tier vocabulary in the return type (`Resolution`, not `TierXResult`).
- Confidence is per-candidate (not per-tier).
- Shadowing is a HARD filter; globals are consulted ONLY when lexically
empty.
- Caller can read `[0]` for one-shot answers.
- `Resolution.confidence` is capped at 1.0.
- `kind-match` is always emitted (weight 0).
## Unresolved-import + dynamic-unresolved evidence shape
- `BindingRef.via.linkStatus === 'unresolved'` applies the
`unlinkedImportMultiplier` (0.5×) to the where-found signal only.
Corroborators (`arity-match`, `owner-match`, `type-binding`) remain
unaffected — the RFC §4v2 capped-signal rule applies per-signal, not
per-candidate.
- `BindingRef.via.kind === 'dynamic-unresolved'` adds a degraded
`dynamic-import-unresolved` evidence signal at weight 0.02.
## Tests (28 in registries.test.ts, 259/259 combined)
Organized per RFC §4.2 step so a regression localizes to the step it broke:
- Step 1: local + walk-to-parent + hard-shadow + origin=import
- Step 2: explicit receiver type-binding + MRO depth decay on ancestor
- Step 3: owner-scoped contributor + owner-match
- Step 5: drop-incompatible-when-compatible-exists + soft-penalty-when-all-
incompatible + unknown-when-no-provider
- Step 6: global-qualified fires only when lexically empty + never for
non-dotted names + not consulted when lexical hit exists
- Step 7: tie-break cascade (inner shadows outer; defId.localeCompare
final)
- Corroborators: unresolved-import 0.5× cap per-signal + dynamic-
unresolved 0.02 degraded signal
- §4.5: lookupQualified kind filter + empty on miss + deterministic defId
order for partial classes
- §4.7: invariants — confidence per-candidate, capped at 1.0, kind-match
always present, [0]-for-one-shot
## Known follow-up optimizations
`collectOwnedMembers` in `lookup-core.ts` iterates `defs.byId.values()`
for each MRO hop — O(D) per call. Acceptable for Ring 2 fixtures; a
by-owner index should land before Ring 3 migrates large-workspace
languages. Tracked alongside the existing `findDefById` follow-up from
#915 review.
## Module placement
All under `gitnexus-shared/src/scope-resolution/registries/` — consistent
with the Ring 2 SHARED folder layout (#912/#913/#914/#915/#916/#918).
Slight deviation from the issue's `gitnexus-shared/src/registries/`
suggestion for consistency with siblings.
## Part of
- Parent: #909
- Depends on (code): #910, #911, #912, #913, #914, #915, #916, #918.
- Closes the Ring 2 SHARED delivery band. Unblocks Ring 2 PKG (#919–#925
bridges to the gitnexus/ CLI package) and Ring 3 language migrations.
* feat(shared): SCC-aware finalize algorithm with bounded fixpoint (#915, RFC #909 Ring 2 SHARED)
Implements RFC §3.2 Phase 2 as pure logic in `gitnexus-shared`. Takes
per-file parse output and returns linked `ImportEdge[]` + materialized
module-scope bindings, fully language-agnostic (target resolution,
wildcard expansion, and binding precedence all go through caller hooks).
Three-phase algorithm:
1. Tarjan SCC over the file-level import graph (iterative, deterministic
node order, O(V+E)). Returns SCCs in reverse-topological order so
leaves finalize before dependents — and so disjoint SCCs are
explicitly surfaced for parallel-processing callers.
2. Per-SCC bounded fixpoint. For each SCC in topo order, iterate up to
`N = |intra-SCC edges|`; each pass tries to resolve every still-
unlinked edge by looking up the imported name in the target file's
local defs. Stops early when no progress. Edges still unlinked after
the cap get `linkStatus: 'unresolved'` — keeps malformed inputs
bounded and preserves the RFC §4v2 capped-signal contract for
unresolved markers.
3. Wildcard expansion + module-scope binding materialization. For each
`wildcard` ParsedImport that linked to a module, expand via
`expandsWildcardTo` into one `wildcard-expanded` ImportEdge per
exported name. Bindings per module scope are the merge of local defs
(`origin: 'local'`), named / alias / reexport imports
(`origin: 'import' | 'reexport'`), namespace imports (`origin:
'namespace'`), and wildcard expansions (`origin: 'wildcard'`), with
precedence delegated to `provider.mergeBindings`.
Dynamic imports rule: `kind: 'dynamic-unresolved'` passes through as an
ImportEdge with `targetFile: null` and no BindingRef.
Re-export flattening: reexport edges land with `transitiveVia: [targetFile]`.
Multi-hop chains settle iteratively across the fixpoint.
Types:
- Adds `'wildcard'` variant to ParsedImport (parse-time signal for
`import * from M`). The finalize-only `'wildcard-expanded'` ImportEdge
kind is unchanged and remains finalize output only, as documented.
- Exports `finalize` + `FinalizeFile` / `FinalizeInput` / `FinalizeHooks`
/ `FinalizeOutput` / `FinalizedScc` / `FinalizeStats`.
Simple-name derivation: `deriveSimpleName` uses `def.qualifiedName` as the
authoritative source (tail after the last `.`). Defs without a
qualifiedName are not name-resolvable by this algorithm — an explicit
design choice that trades strictness for predictability (no heuristic
nodeId parsing).
Tests (20, all passing):
- Trivial: empty workspace · acyclic resolution · unresolvable target
(file + name) · dynamic-unresolved passthrough.
- Cycles: A↔B two-file cycle linked · cycles packed into SCC with
isCycle=true · disjoint cycles produce disjoint SCCs · mixed
linked/unresolved edges reported correctly in stats.
- Wildcards: one ImportEdge per exported name · unresolved wildcards
survive as single edges · expanded bindings carry origin='wildcard'.
- Reexports: transitiveVia carries the intermediate file path.
- Aliased + namespace: alias preserves targetExportedName under its
local name · namespace links to module scope even without a module-def.
- Bindings: locals land as origin='local' · imports layer on via
mergeBindings · mergeBindings can drop existing (last-write-wins
precedence honored).
- SCC-DAG: reverse-topological ordering verified (leaf first).
Combined scope-resolution / model / shadow suite: 229/229 pass.
`tsc --noEmit` clean in both `gitnexus-shared` and `gitnexus`.
Closes part of #909. Unblocks #917 (Registry.lookup's import-chain fast
path consumes finalized ImportEdges); unblocks Ring 3 language migrations
(per-language providers supply FinalizeHooks implementations).
* chore(shared): address #915 review findings — dead code, docs, tests
Review thread on PR #962.
Code changes:
- Remove dead `resolvedTargets` map + `keyFor` + `ParsedImportKey`
type alias. The map was populated but never read; originally intended
to cache / dedup resolutions for later phases but that path was never
wired (finding 1.1).
- Drop unused params (`_edgeIndex`, `_hooks`, `_workspace`) from
`tryFinalize`. No planned fixpoint-state consultation; no reason to
keep them reserved (finding 2.1).
Documentation:
- `FinalizeFile.localDefs` now documents the multi-hop re-export
contract explicitly: `finalize` looks names up in the target's
static `localDefs`; if B only re-exports from C and doesn't surface
the name in its own localDefs, A's import of that name from B will
hit the cap and be marked unresolved. Parsers that want multi-hop
chains to settle end-to-end must include re-exported names in the
intermediate file's localDefs (finding 1.2).
- `FinalizeStats` now documents its counting granularity: all edge
counters are per-`ParsedImport`, not per-materialized-`ImportEdge`.
A wildcard expanding to N exports counts as one linked edge;
dynamic-unresolved pass-throughs count as linked. The bindings map
is the authoritative "has a BindingRef" source (finding 3.2).
Tests (2 added, 22 total in finalize-algorithm.test.ts, 231/231 combined):
- Explicit cap-hit → `linkStatus: 'unresolved'` assertion for a cycle
where the name-level lookup never succeeds (distinct from
`targetFile: null`; cap exhaustion path) (finding 3.1).
- Multi-hop re-export contract test: demonstrates both variants —
intermediate B WITHOUT X in localDefs → unresolved; B WITH X in
localDefs → resolved to the original source DefId (finding 1.2).
Not addressed (filed as follow-up issues):
- LanguageProvider.resolveImportTarget vs FinalizeHooks signature
divergence (finding 1.3) — pre-Ring-3 concern.
- findDefById O(F×D) scan in Phase 5 (finding 4.1) — acceptable for
Ring 2; optimize before large-workspace Ring 3 migrations.
Implements the scope-tree spine and position-indexed lookup as pure logic
in `gitnexus-shared`. Generalizes the `enclosingFunctions` pattern from
closed PR #902 to arbitrary `ScopeKind`s.
Three modules under `gitnexus-shared/src/scope-resolution/`:
1. `scope-id.ts` — `makeScopeId({filePath, range, kind})` builds the
canonical RFC §2.2 shape
`scope:{filePath}#{startLine}:{startCol}-{endLine}:{endCol}:{kind}`
and interns the result through a process-local pool so repeated calls
with structurally identical inputs return the same string reference.
`clearScopeIdInternPool()` exported for test isolation.
2. `scope-tree.ts` — `buildScopeTree(scopes)` validates invariants and
returns an immutable `ScopeTree`:
- `getScope(id)` / `getParent(id)` / `getChildren(id)` / `getAncestors(id)`
- Implements the `ScopeLookup` contract from #916, so `resolveTypeRef`
can consume a `ScopeTree` directly (test included).
Invariants enforced (throw `ScopeTreeInvariantError` on violation):
- Non-Module scopes must have a parent.
- Parent must exist in the supplied set.
- Parent range STRICTLY contains child range (equal ranges rejected).
- Sibling ranges under the same parent do not overlap. Ranges that
merely touch at the boundary (`a.end == b.start`) are accepted.
- Parent and child live in the same filePath.
- Duplicate scope ids are rejected.
3. `position-index.ts` — `buildPositionIndex(scopes)` produces a
`PositionIndex` with `atPosition(filePath, line, col)`. Per-file sorted
array; binary-search the upper bound of `start ≤ query`, scan backward
through the prefix, return the first containing hit.
Complexity: `O(log N_file + D)` typical (D = lexical depth ≤ ~10);
degrades to `O(N_file)` only under pathological inputs (many scopes
starting at the same position). "Innermost wins" falls out of the sort
+ backward-scan contract because `ScopeTree`'s invariants guarantee
that scopes containing a point form an ancestor chain.
Types:
- `ScopeTree` now exported from `scope-tree.ts`. The Ring 1 opaque
placeholder in `types.ts` has been removed; LanguageProvider hooks
that previously took `ScopeTree = unknown` now receive the concrete
interface (CLI `tsc --noEmit` passes — no existing callers rely on
the opaque shape).
Tests (39, all passing):
- scope-id: canonical shape · all six ScopeKinds encoded · identity
equality (same inputs → same reference) · distinguished by
filePath / range / kind · purity under repeated calls · intern-pool
clear preserves canonical shape.
- scope-tree: empty tree · single module · nested Module→Class→Function
· multiple siblings input-order preserved · ScopeLookup integration
with resolveTypeRef · frozen children and ancestor arrays · all six
invariant violations (non-Module orphan, parent-not-found, parent
doesn't contain, parent == child, siblings overlap, cross-file parent,
duplicate id) · boundary-touching siblings accepted.
- position-index: empty · unindexed filePath · before/after-file
queries · start/end inclusivity · innermost-wins for nested / co-
starting / co-ending / same-line scopes · sibling dispatch · multi-
file isolation · size · id-dedup.
Combined scope-resolution / model / shadow suite: 190/190 pass.
`tsc --noEmit` clean in both `gitnexus-shared` and `gitnexus`.
Closes part of #909. Unblocks #917 (`Registry.lookup` needs the scope
spine); makes `ScopeLookup` in #916 concrete without API churn.
* feat(search): per-phase timing instrumentation for the query pipeline
The eval harness already measures search-pipeline latency per phase,
but the *product* query() tool has no timing visibility. That leaves
production latency opaque:
- Is BM25 the tail, or vector search?
- How much Promise.all overlap do concurrent searches actually save?
- Does symbol_lookup dominate when per-symbol Cypher round-trips pile up?
None of this is answerable from the outside, which blocks the
latency-quality Pareto work tracked in #546 / #553.
Changes:
* New PhaseTimer class at src/core/search/phase-timer.ts.
Supports three APIs:
- start(phase) / stop() for sequential phases (per issue spec)
- mark(phase, durationMs) for pre-measured durations
- time(phase, promise) to wrap a promise inside Promise.all
The issue's original spec was sequential-only, which doesn't work
for BM25 + vector inside Promise.all — the second start() would
auto-stop the first and only one phase would get timed. The mark()
and time() variants resolve that without changing the sequential
API for the other phases.
* local-backend.ts query() instrumented across seven phase markers:
bm25, vector (concurrent via timer.time inside Promise.all)
merge (RRF reciprocal-rank-fusion)
symbol_lookup (per-symbol process + cohesion + content Cypher)
ranking (in-memory priority sort)
formatting (response object construction + dedup)
wall (end-to-end; separate mark so callers can compare
sum(phases) vs wall and see Promise.all savings)
* logQueryTiming() helper next to logQueryError(), same console-based
pattern (repo has no structured logger). Emits
GitNexus [query:timing] query="..." totalMs=N phases={...}
to stdout — greppable prefix, JSON-parseable payload, no new deps.
* timing: Record<string, number> added as a top-level field on the
query() response. Strict superset of the previous shape — existing
tests only assert field presence, so no regression. Other MCP tools
use the same top-level-metadata convention (status, row_count,
warning) rather than a nested _meta wrapper.
Tests:
- 6 new unit tests for PhaseTimer covering start/stop, implicit
stop-on-start, additive mark(), Promise.all-safe time(),
negative/NaN rejection, and totalMs auto-stop.
- 3 new assertions on the existing query integration test verifying
timing.wall is a non-negative number and at least one of
bm25/vector fired.
Verification:
npx vitest run test/unit/phase-timer.test.ts -> 6 pass
npx vitest run test/unit/calltool-dispatch.test.ts -> 65 pass
npx vitest run test/integration/local-backend-calltool.test.ts -> 18 pass
npm run test:unit -> 3777 pass
(4 pre-existing env failures unchanged: skip-git-cli needs
built dist/, git-utils tmpdir on Windows worktree)
npx tsc --noEmit -> clean
Scope declined for v1:
- In-process histogram aggregation — the log line is enough for
external tooling
- Pareto curve generation — issue asks to enable it, not generate it
- Sub-phases of symbol_lookup (process vs cohesion vs content) —
issue lists them under one bucket; can split later if demand surfaces
Closes#553
* fix(search): route query:timing log to stderr to preserve stdio MCP contract
CI (#953) failed the `query: JSON appears on stdout, not stderr`
e2e test in test/integration/cli-e2e.test.ts with:
SyntaxError: Unexpected token 'G', "GitNexus [..." is not valid JSON
Root cause: my initial logQueryTiming() in 63fbdc4 used console.log,
which writes to stdout. The MCP stdio transport uses stdout
exclusively for JSON-RPC responses (#324), and the CLI e2e test
guards that contract by asserting stdout parses as JSON on every
tool invocation. The "GitNexus [query:timing] ..." line was
interleaving with the response JSON and breaking the parse.
Fix: route logQueryTiming through console.error instead. stderr is
the correct channel for human-readable diagnostics and it is what
the sibling logQueryError already uses for the same reason. The log
line format is otherwise unchanged -- still greppable, still
JSON-parseable payload.
Verification (local, with dist built):
npx vitest run test/integration/cli-e2e.test.ts -t "query: JSON"
-> now passes (was failing across ubuntu/windows/macos in CI)
npx tsc --noEmit -> clean
Two unrelated pre-existing failures on non-git
directory handling remain (same on upstream/main).
Closes the CI regression introduced in 63fbdc4.
Implements RFC §3.1 `MethodDispatchIndex`: a two-way materialized view
keyed by `DefId` for O(1) method-dispatch resolution:
- `mroByOwnerDefId` — owner class → full MRO ancestor chain
(excludes self, per-language strategy order)
- `implsByInterfaceDefId` — interface/trait → classes that implement it
**Not an MRO implementation.** `buildMethodDispatchIndex` is a pure
aggregator that calls back into caller-provided `computeMro` and
`implementsOf` functions. The five existing strategies (Python C3, Ruby
kind-aware, Java/Kotlin linear, Rust qualified-syntax, COBOL none) stay
where they are today (`model/resolve.ts`, `languages/ruby.ts`); this index
does not reimplement them.
Why callbacks rather than a shared registry: the strategies depend on the
CLI's `HeritageMap` + `SemanticModel`. Migrating both to `gitnexus-shared`
is out of scope for #914; callbacks let the shared build stay pure.
Module placement: `gitnexus-shared/src/scope-resolution/method-dispatch-index.ts`
for consistency with the other RFC §3.1 indexes (#913 DefIndex /
ModuleScopeIndex / QualifiedNameIndex; #916 resolveTypeRef).
Safety surface mirrors sibling indexes:
- First-write-wins on duplicate owners.
- Repeated (interface, owner) pairs deduplicated.
- Stored arrays are `Object.freeze`d; caller mutation of the source
array does not leak into the index.
- Miss returns a shared frozen empty array.
Tests (19, all passing): empty input, single-inheritance chain, Python
C3 diamond, Java BFS, Ruby kind-aware mixin, Rust qualified-syntax empty,
interface inversion (single, multiple, ordered), dedup within and across
callback calls, frozen miss + bucket arrays, callback-array isolation,
readonly Map iteration.
Closes part of #909.
Implements RFC §4.6: a strict, pure resolver for `TypeRef`s used by
`Registry.lookup` Step 2 (type-binding propagation) and by any caller that
wants the single best type-target for an annotation without paying for the
full evidence pipeline.
Algorithm (strict):
1. Walk the scope chain from `ref.declaredAtScope`:
- Return the first binding for `rawName` whose origin is in
`{'local','import','namespace','reexport'}` AND whose `def.type` is a
type-kind (class-like, interface-like, enum-like, alias-like).
- If bindings exist but none qualify (non-type shadow, wildcard-only
origin), return null immediately — do NOT fall through to the global
qualified-name index.
2. If `rawName` is dotted and the scope walk produced no match, consult
`QualifiedNameIndex.byQualifiedName`. Only accept a UNIQUE type-kind
hit; ambiguous or non-type results return null.
`'wildcard'` is deliberately excluded from strict origins — a
wildcard-expanded name is too loose to anchor type resolution.
Module placement: `gitnexus-shared/src/scope-resolution/resolve-type-ref.ts`
(alongside sibling indexes) rather than the issue's suggested
`gitnexus-shared/src/resolve-type-ref.ts`, for consistency with the rest of
the RFC §2/§3 surface.
A minimal `ScopeLookup` interface is declared inline so #916 ships
standalone; #912's `ScopeTree` will satisfy this contract without change.
Closes part of #909.
Three flat O(1) indexes + pure build functions over per-file artifacts.
Contract-only; no runtime behavior change yet — consumers (#917 Registry
lookups, #915 SCC finalize, #919 ScopeExtractor) wire in later.
Each index follows the same shape:
- build function: flat input list → frozen immutable index
- public interface: readonly Map + get/has/size accessors
- first-write-wins on id/filePath collisions (upstream bug signal)
- pure, side-effect-free, safe to call repeatedly
DefIndex — the global "what is this id?" lookup
gitnexus-shared/src/scope-resolution/def-index.ts
buildDefIndex(defs: readonly SymbolDefinition[]): DefIndex
byId: ReadonlyMap<DefId, SymbolDefinition>
Consumed by Registry.lookup (#917) to materialize DefId[] hits back to
full SymbolDefinition records.
ModuleScopeIndex — `filePath → moduleScopeId` for cross-file hops
gitnexus-shared/src/scope-resolution/module-scope-index.ts
buildModuleScopeIndex(entries): ModuleScopeIndex
byFilePath: ReadonlyMap<string, ScopeId>
Consumed by the SCC finalize link pass (#915) to resolve
ImportEdge.targetFile to a concrete module scope in constant time.
QualifiedNameIndex — cross-kind qualified-name fast path
gitnexus-shared/src/scope-resolution/qualified-name-index.ts
buildQualifiedNameIndex(defs: readonly SymbolDefinition[]): QualifiedNameIndex
byQualifiedName: ReadonlyMap<string, readonly DefId[]>
Returns DefId[] (not a single DefId) because partial classes, method
overloads, and cross-kind collisions can legitimately share a
qualifiedName. Callers filter by acceptedKinds at the lookup site.
Consumed by Registry.lookup qualified fast path + resolveTypeRef
dotted fallback (#916, #917).
Barrel re-exports added to gitnexus-shared/src/index.ts so consumers
import from 'gitnexus-shared' rather than deep paths.
Tests (gitnexus/test/unit/scope-resolution/, 23 total):
def-index.test.ts (6):
empty, single def, multiple distinct, first-write-wins collision,
missing id returns undefined, byId direct iteration
module-scope-index.test.ts (6):
empty, single entry, multiple files, first-write-wins on duplicate
filePath, missing returns undefined, byFilePath direct iteration
qualified-name-index.test.ts (11):
empty, single qnamed def, partial classes accumulate, input-order
preservation, qname separation, skip undefined/empty qname, pair
dedup, cross-kind indexing, frozen-empty-array on miss, direct
iteration
Verification:
- gitnexus-shared + gitnexus build clean (tsc + scripts/build.js)
- test/unit/scope-resolution: 23/23 pass
- model + shadow + scope-resolution combined: 129/129 pass
- No runtime consumer wiring yet — indexes are standalone library
functions that #915, #917, #919 will import when ready
Depends on #910 (SymbolDefinition, DefId, ScopeId types — already on main).
Unblocks #915 (finalize algorithm), #917 (Registry.lookup), #919
(ScopeExtractor materialization).
Replaces the scaffold stubs with working pure-logic implementations plus
unit-test coverage for both functions. Unblocks Ring 2 PKG #923 (shadow
harness) to consume a concrete library instead of throwing scaffolds.
gitnexus-shared/src/scope-resolution/shadow/diff.ts
`diffResolutions(callsite, legacy, newResult): ShadowDiff`
- [0] on each side is the top match
- both empty → 'both-empty', delta []
- legacy empty only → 'only-new', delta = new top evidence
- new empty only → 'only-legacy', delta = legacy top evidence
- same top nodeId → 'both-agree', delta []
- different nodeIds → 'both-disagree',
delta = symmetric difference of evidence kinds
(legacy-only first in input order, then new-only)
Evidence identity is `ResolutionEvidence.kind` — weight/note differences
for the same kind do NOT produce delta entries. Rationale: the aggregator
wants to know which *signals* explain a disagreement, not fluctuations
in calibration values.
gitnexus-shared/src/scope-resolution/shadow/aggregate.ts
`aggregateDiffs(diffs, now?): ShadowParityReport`
- buckets by `SupportedLanguages`
- tallies agreements, evidence-breakdown (divergences only — agree and
empty rows do not contribute)
- parity = bothAgree / (totalCalls - bothEmpty), yields 0 (not NaN)
when the denominator is 0
- perLanguage sorted alphabetically by enum value for stable output
- evidenceBreakdown internally sorted by kind for stable output
- overall = column-wise sum across languages
- `now` parameter makes generatedAt deterministic in tests
gitnexus-shared/src/index.ts
Re-exports the full shadow API: diffResolutions, aggregateDiffs, and all
their types (ShadowAgreement, ShadowCallsite, ShadowDiff,
LanguageParityRow, ShadowParityReport).
gitnexus/test/unit/shadow/diff.test.ts (13 tests)
- 5 agreement outcomes
- symmetric-by-kind evidence delta (disjoint, overlapping, fully-overlapping)
- weight-only differences produce no delta
- top-match only (ignores indices beyond [0])
- callsite passthrough
- delta ordering (legacy-only first, input order preserved)
gitnexus/test/unit/shadow/aggregate.test.ts (9 tests)
- empty input
- single language, all agree / mixed / all empty
- multi-language bucketing + overall sum
- alphabetical language sort
- evidence breakdown scope
- determinism via injected `now` + JSON round-trip identity
Verification:
- gitnexus-shared + gitnexus build clean (tsc + scripts/build.js)
- test/unit/shadow: 22/22 pass
- test/unit/model + test/unit/shadow combined: 106/106 pass
- No runtime behavior changes (shadow is invoked by #923, not yet wired)
Stacked on main (af1d278a). Depends on types from #910 (merged).
Unblocks: #923 (Ring 2 PKG — shadow harness wiring) — concrete library
to consume instead of scaffold stubs.
Plan: docs/plans/2026-04-18-001-refactor-911-senior-hooks-redesign-plan.md
is about #911; #918's scope is the scaffold+fill-in described in the PR
description of #951.
Adds the 14 optional scope-resolution hooks from RFC #909 §5.2 to
`LanguageProviderConfig` plus the supporting input/output types in
`gitnexus-shared`. Contract-only; no runtime behavior changes.
Review-driven refinements (addresses two non-blocking review comments on #950):
1. `ParsedImport` is now a 5-variant discriminated union, not a flat
record. Each variant carries only its legal fields so invalid shapes
are compile errors:
- 'named', 'alias', 'namespace', 'reexport', 'dynamic-unresolved'
'wildcard-expanded' is deliberately excluded — finalize materializes
that kind; a provider must never emit it at parse time.
'reexport' is a first-class parse-phase variant so syntactically-
detectable re-exports (TS `export { X } from './y'`, Rust
`pub use foo::bar`) keep their parse-time signal through to finalize
rather than being re-derived by the SCC pass.
`namespace` gains an `importedName` field so `import numpy as np`
can carry both `localName: 'np'` and `importedName: 'numpy'`.
`dynamic-unresolved.targetRaw` is `string | null` (was mandatory
null) so providers can emit the unresolvable expression text for
diagnostics when available.
2. `bindingScopeFor` and `importOwningScope` return type changed from
`ScopeId` to `ScopeId | null`, aligning with the X | null convention
used by the 12 sibling optional hooks (receiverBinding,
resolveScopeKind, interpretTypeBinding, …). `null` = delegate to the
central default. Enables partial overrides — a JS provider can
return a hoisted scope for `var` and `null` for `let`/`const`
without re-implementing the default lookup.
Both hooks also gain a purity JSDoc contract: same inputs yield the
same ScopeId (or null) across invocations; no closure over mutable
state. Required to keep scope-tree construction deterministic.
A richer callable-defaults pattern (typed BindingScopeDefaults /
ImportOwningDefaults helper interfaces on a `defaults` parameter)
was considered and deferred to Ring 2 PKG #919, where the concrete
ScopeExtractor will exist to inform the helper shape. Designing that
pattern before the first consumer would set cross-hook precedent
based on a single motivating example.
Supporting types added to gitnexus-shared/src/scope-resolution/types.ts:
- CaptureMatch, ParsedImport, ParsedTypeBinding
- WorkspaceIndex, ScopeTree (opaque placeholders until Ring 2)
- Callsite
14 hooks added to LanguageProviderConfig (all optional):
Parse phase: emitScopeCaptures, interpretImport, receiverBinding,
interpretTypeBinding, resolveScopeKind, shouldCreateScope,
bindingScopeFor
Finalize phase: resolveImportTarget, expandsWildcardTo,
importOwningScope, mergeBindings
Reference-extraction phase: classifyCallForm
Resolution phase: shouldShadow, arityCompatibility
Verification:
- gitnexus-shared builds clean (tsc)
- gitnexus builds clean (scripts/build.js)
- test/unit/model: 84/84 pass — no regressions
- No provider needs updating (all hooks optional)
- No BindingScopeDefaults/ImportOwningDefaults/defaults parameter
introduced (deferred to #919)
Stacked on #910 (merged as afc0a8b6); rebased on main.
Tracking: #909 (meta). Unblocks Ring 2 PKG (#919 ScopeExtractor,
#922 import adapters) and all Ring 3 per-language migrations.
Plan: docs/plans/2026-04-18-001-refactor-911-senior-hooks-redesign-plan.md
Lands the authoritative data model and constants for the pure scope-based
resolution RFC (#909) as Ring 1, part 1. No runtime behavior changes —
types + constants only.
New in gitnexus-shared/src/scope-resolution/:
- types.ts — Scope, ScopeKind, ScopeId, DefId, Range, Capture,
BindingRef, ImportEdge, TypeRef, Resolution, ResolutionEvidence,
Reference, ReferenceIndex, LookupParams, RegistryContributor
- evidence-weights.ts — EvidenceWeights constant map + typeBindingWeightAtDepth
(RFC Appendix A)
- origin-priority.ts — ORIGIN_PRIORITY constant map for deterministic
tie-breaks (RFC Appendix B)
- language-classification.ts — LanguageClassification type +
LanguageClassifications map (production × 14, experimental × 2
for vue/cobol; governs Ring 4 DAG-retirement gate)
- symbol-definition.ts — SymbolDefinition moved from
gitnexus/src/core/ingestion/model/symbol-table.ts so scope-resolution
types can reference it from the shared package
Consumer updates:
- symbol-table.ts: removes local SymbolDefinition declaration; imports
from gitnexus-shared
- model/index.ts: drops SymbolDefinition from barrel re-export per
"direct imports from gitnexus-shared" convention (see
gitnexus-shared feedback in project memory)
- 9 source files + 5 test files: import SymbolDefinition directly
from 'gitnexus-shared'
Verification:
- gitnexus-shared builds clean (tsc)
- gitnexus builds clean (scripts/build.js)
- 131/132 unit test files pass; 3767 tests green
- Zero behavior changes; SymbolDefinition shape unchanged
Blocks: #911 (LanguageProvider hook interface extensions) and all of
Ring 2 (#912-#925). Closes part of #909.
Deterministic fix for the Windows-flaky pipeline-graph-golden test.
Root cause
cli-e2e.test.ts wrote into the SHARED fixture directory
(test/fixtures/mini-repo/) — git init, analyze run that creates
AGENTS.md, CLAUDE.md, .claude/, .gitnexus/. When pipeline-graph-golden
ran in parallel, its `cpSync` of the source directory could capture
the mid-flight pollution before cli-e2e's afterAll cleanup fired.
macOS/Ubuntu won the race often enough that the flake presented as
Windows-only.
Fix
cli-e2e now copies mini-repo into a fresh `mkdtemp`'d parent whose
basename is `mini-repo` (preserving `--repo mini-repo` CLI lookup by
basename), runs git-init there, and rm's the whole tmpdir in afterAll.
The shared fixture source is never touched.
Fallout from the cwd change: bare `--import tsx` specifiers (2
spawnSync + 1 spawn) can't resolve `tsx` from an os.tmpdir cwd where
there is no node_modules. Switched them to the already-existing
`tsxImportUrl` (absolute file:// URL to the tsx loader), matching
the `runCliOutsideProject` pattern that was already set up for this
exact case.
Updated the "MINI_REPO is inside the project tree" comment in the
`status on non-indexed repo` test — MINI_REPO is now in os.tmpdir,
so the rationale for using a separate throwaway tmp git repo is
different (but still valid: previous tests in the suite create
MINI_REPO/.gitnexus, which findRepo() would pick up).
Also updated pipeline-graph-golden's comment explaining WHY it
copies to tmp — it's now defense-in-depth rather than a necessity,
so a future test that adds files to the source can't silently
regress the golden.
Verification
- 5x consecutive `cli-e2e + pipeline-graph-golden` runs: 20/20 pass
(deterministic)
- 3x full suite including pipeline.test: 27/27 pass
- test/fixtures/mini-repo/ post-run contents: only `src/` —
zero pollution from any test
- macOS/Ubuntu behavior unchanged (they were passing; tmpdir
isolation is purely additive)
* feat(mcp): rank context/impact disambiguation candidates and expose kind/file_path hints
The `context` MCP tool already returned `{ status: 'ambiguous', candidates }`
when a name hit multiple symbols, but the candidates were returned in
arbitrary DB order and the only hint it accepted was file_path. The
`impact` tool was worse: when its name resolver found multiple viable
matches it silently picked the first one from a priority UNION, with no
signal back to the caller that a different symbol might have been
intended.
Both failure modes were flagged in issue #470 and reconfirmed in the
comments by a second user who described impact as returning "incorrect
parsing results and meaningless tool calls" in the multi-match case.
Changes:
* Add `resolveSymbolCandidates(repo, query, hints)` private helper on
LocalBackend. Single place that:
- Short-circuits on direct uid (zero-ambiguity)
- Runs the same name-or-qualified-id match as before, with LIMIT 20
(was 10) so the ranker has headroom instead of arbitrary truncation
- Preserves the #480 Class/Constructor preference -- when the only
ambiguity is a Class and its own Constructor, the Class wins
silently
- Scores each candidate (pure TS, no extra DB round-trip): base 0.50,
+0.40 for file_path match, +0.20 for kind match, plus a small
kind-priority tiebreaker (Class > Interface > Function > Method >
Constructor) when no explicit kind hint is given
- Sorts desc by score with stable tiebreakers (shorter filePath,
then lex uid)
- Promotes to a single confident resolve when the top score is
>= 0.95 AND beats the runner-up by >= 0.10 -- lets a strong hint
cut through without forcing the caller through a disambiguation
round-trip
* Rewire `context()` to use the shared helper. Response shape is a
strict superset of today's: candidates gain a `score` field, the
existing `{ uid, name, kind, filePath, line }` keys are preserved so
every downstream consumer (rename, eval-server formatter, etc.) keeps
working. New `kind` input hint accepted.
* Rewire `impact()` to use the shared helper. Now emits the same
`{ status: 'ambiguous', candidates, impactedCount: 0, risk: 'UNKNOWN' }`
shape instead of silent first-pick. New inputs accepted:
`target_uid`, `file_path`, `kind`.
* Update tool schemas in mcp/tools.ts to advertise the new inputs and
describe ranked disambiguation.
Backward compatibility:
The #480 Class/Constructor collapse is preserved and covered by the
existing java-class-impact integration test (still green). The
ambiguous response shape is a strict superset -- `eval-formatters`
unit test that parses the old shape is unchanged and still passes.
`impact` going from silent-first-pick to structured ambiguous is a
semantic improvement that is the entire point of the issue; callers
relying on silent first-pick now get an actionable response.
Scope declined for v1:
module/community hint -- the issue lists it as one of several hints,
but kind + file_path cover the vast majority of disambiguation needs
in practice, and a community-label filter requires an extra graph
query per candidate. Natural v2 follow-up.
Tests: calltool-dispatch.test.ts gains 5 new cases covering file_path
boost, kind hint boost, impact ambiguous shape, impact target_uid
short-circuit, and score field presence on the existing ambiguous
test. Plus the extended assertions on the existing
`context tool returns disambiguation for multiple matches`.
Verification:
npx vitest run test/unit/calltool-dispatch.test.ts -> 64 pass
npx vitest run test/integration/java-class-impact.test.ts -> pass
npm run test:unit -> 3642 pass
(4 pre-existing env failures unchanged: skip-git-cli needs built
dist/, git-utils tmpdir on Windows worktree -- same on main)
npx tsc --noEmit -> clean
Closes#470
* fix(mcp): enrich labels from UNION when labels(n)[0] is empty; address review findings
CI on PR #888 caught 13 integration-test failures I did not cover locally:
my resolver refactor collected candidates via `labels(n)[0] AS type`, but
LadybugDB returns an empty string for that projection on certain node
types (most importantly Class). With an empty `type`, impact's downstream
`_runImpactBFS` no longer recognised `symType === 'Class' | 'Interface'`
and stopped seeding Constructor + File nodes into the frontier, so the
"impact(upstream) surfaces the file importer" assertion broke across 11
language fixtures plus 2 OVERRIDES filter tests.
The original impact resolver worked around this by running a prioritised
UNION across Class/Interface/Function/Method/Constructor and picking the
first hit. My refactor dropped that. Fix: keep the simple candidate MATCH
but enrich types afterward via a single scoped UNION query, so every
candidate carries an accurate label for both scoring and downstream
BFS seeding. The UID direct-lookup path is patched the same way.
Also addresses the findings from the senior reviewer on PR #888:
* MIGRATION.md: document the `impact` behavioural change (silent first-
pick → structured `{ status: 'ambiguous', candidates }`) so downstream
callers know to branch on `result.status` before reading byDepth/
summary. `context` is unchanged shape-wise (strict superset).
* New test: `context tool promotes top candidate via scoring when
multiple rows survive DB pre-filter`. The review flagged that the
existing file_path test works only because the mock ignores WHERE
parameters -- the scored-promotion path (top ≥ 0.95 AND gap > 0.09)
wasn't directly exercised. The new test uses two candidates both in
App.tsx-containing paths plus a kind hint so promotion is decided by
scoring, not DB pre-filtering. Also tightened the comment on the
earlier file_path test to describe the mock vs production divergence
honestly.
* NIT: added a paragraph explaining why `scored.length >= 2` is kept as
a defensive guard even though the `normalized.length === 1` early
return already covers the single-candidate path.
* Integration: two tests in `local-backend-calltool.test.ts` targeted
`'authenticate'`, which now correctly resolves as ambiguous (two
Method nodes: AuthService.authenticate and BaseService.authenticate).
Updated both to pass `file_path: 'src/auth.ts'` so they exercise the
new disambiguation API and still assert the METHOD_OVERRIDES filtering
they were originally about.
Edge case fix in the promotion gap check: IEEE754 makes 0.50 + 0.40 +
0.20 - 0.90 = 0.09999999999999998 instead of exactly 0.10, which would
otherwise break the "winner clearly dominates" intent for legitimate
1.00 vs 0.90 cases. Changed `>= 0.10` to `> 0.09`; same user-facing
intent, no floating-point sensitivity.
Verification (all from gitnexus/):
npx vitest run test/integration/class-impact-all-languages.test.ts
-> 52 pass (was 11 FAIL on CI before this fix)
npx vitest run test/integration/local-backend-calltool.test.ts
-> 18 pass (was 2 FAIL on CI before this fix)
npx vitest run test/integration/java-class-impact.test.ts
-> 10 pass (regression guard for #480 preserved)
npx vitest run test/unit/calltool-dispatch.test.ts
-> 65 pass (1 new test + 4 from original #470 PR)
npm run test:unit
-> 3626 pass, 4 pre-existing env failures unchanged
npx tsc --noEmit
-> clean
* fix: keep worker warnings non-terminal
Treat parse-worker warning messages as informational so a warning can be surfaced without short-circuiting the worker result protocol.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
* style: apply prettier formatting
---------
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
* feat: add docker support
* feat: move docker files to root
* feat: add docker build and push workflow
* fix: pin docker action SHAs to verified commits
Made-with: Cursor
* fix: remove redundant --platform=$TARGETPLATFORM from runtime stage
Made-with: Cursor
* fix: upgrade docker actions to Node.js 24-compatible versions
Made-with: Cursor
* docs: updated readme
* fix: update docker references
* fix(docker-server): reject null bytes in resolvePath
Defensively harden the path traversal guard by returning null
early when the URL contains a null byte, before normalization runs.
Made-with: Cursor
* fix(docker-server): handle createReadStream errors
Attach an error listener before piping so mid-flight read errors
(truncated file, permission change) cleanly destroy the response
instead of being silently swallowed.
Made-with: Cursor
* fix(docker-server): replace existsSync with async stat
Eliminates the TOCTOU race between the initial stat call and the
subsequent existsSync check. Reuses the async stat pattern already
in place and removes the now-unused existsSync import.
Made-with: Cursor
* test(docker-server): add integration tests; fix %00 null-byte bypass
Decode the URL before the null-byte check so percent-encoded null
bytes (%00) are also rejected with 400 instead of falling through
to the SPA fallback. Adds 5 node:test integration tests covering
valid assets, SPA fallback, path traversal, null bytes, and 404.
Made-with: Cursor
* style: fix prettier formatting in docker-server files
Made-with: Cursor
* fix(docker): wire tests into CI, fix resolvePath separator, correct image namespace
- Add `node --test docker-server.test.mjs` step to ci-tests.yml so the
path-traversal guard tests run in every CI pass instead of being silently skipped.
- Fix resolvePath containment check: `startsWith(root)` would allow sibling
directories like `/app/dist-evil/`; now guards with `root + sep` or exact match.
- Update docker-compose.yaml default image from `abhigyanpatwari` namespace to
`brainifii` to match what docker.yml publishes to GHCR.
* fix(docker): update apt-get commands and set user permissions
- Modify Dockerfile and Dockerfile.test to include options for apt-get to bypass validity checks during updates.
- Set ownership of the /app directory to the 'node' user in the runtime stage for improved security and proper permission handling.
* fix(docker): switch to Alpine base images for smaller footprint
- Update Dockerfile to use Alpine-based Node.js images for both builder and runtime stages, reducing image size and improving performance.
- Replace apt-get commands with apk for package installation in the runtime stage.
* fix(docker): update Node.js version in Dockerfile
- Change base image from node:20-alpine to node:22-alpine
* fix(docker): update Node.js version in Dockerfile to 22-alpine for runtime
---------
Co-authored-by: kritik.b <kritik.b@media.net>
* feat(extractors): detect jQuery $.ajax/$.get/$.post and axios object-form as HTTP consumers
The JS/TS HTTP consumer extractor currently recognises fetch() and
axios.<verb>() but misses three patterns extremely common in Laravel
and legacy frontends:
- jQuery shorthand: $.get(url), $.post(url, data)
- jQuery ajax form: $.ajax({ url, method }) / $.ajax({ url, type })
- axios object form: axios({ method, url })
Missing them means the frontend->backend cross-link disappears from
`group sync`, breaking impact analysis for whole classes of repos.
Implementation (node.ts):
- 3 new PatternSpecs alongside the existing FETCH_/AXIOS_ specs
- NodePatternBundle extended with jqueryShorthand / jqueryAjax /
axiosObject slots, compiled for JS / TS / TSX grammars
- readStringProp() helper walks object-literal `pair` children and
resolves `url` / `method` / `type` keys independent of order,
sidestepping the positional S-expression constraint on the
query form proposed in the issue
- 3 new scan loops in scanBundle() emit HttpDetection with
framework 'jquery' (new) or 'axios' (existing), confidence 0.7
to match the existing source-scan consumers, defaulting method
to GET when absent (matches both jQuery and axios runtime)
Tests (http-route-extractor.test.ts): 4 new cases -- 3 positive
(shorthand, ajax with method:/type: and default GET, object-form
with swapped key order and default GET) plus 1 negative control
that asserts unrelated \$.fn.extend / \$.each / non-axios helper
calls with {url, method} literals produce zero consumer contracts.
Closes#828
* test(extractors): cover jQuery $.ajax with template-literal URL
Extend the existing $.ajax fixture with `url: \`/api/orders/\${id}\``
and assert the consumer is emitted as http::GET::/api/orders/{param}.
This makes jQuery + template-URL explicit rather than implicit via the
axios object test (readStringProp already accepts template_string for
both; this is coverage, not new behaviour).
Addresses the single non-blocking finding on PR #887.
* Initial plan
* feat(ingestion): add variable extraction types, factory, configs, and wire into language providers
- Create variable-types.ts with VariableInfo, VariableExtractionConfig, VariableExtractor interfaces
- Create variable-extractors/generic.ts with createVariableExtractor() factory
- Add variableExtractor field to LanguageProvider interface
- Create per-language variable extraction configs for all 16 languages
- Wire variableExtractor into all language providers
- Add variable metadata enrichment to parse-worker for Const/Static/Variable labels
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3cb85c68-1792-473e-9a46-ea2588da0e5e
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* feat(ingestion): add variable extraction tests and fix Python/TS config issues
- Create test/unit/variable-extraction.test.ts with 29 tests covering
TypeScript, JavaScript, Python, Go, Rust, C, C++, Ruby, and factory behavior
- Fix isConst in generic factory to use config.isConst over node-type membership
(TS let/const both use lexical_declaration)
- Fix Python type extraction for annotated assignments at module scope
- Fix Python dunder name visibility (e.g., __name__ is public, not protected)
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3cb85c68-1792-473e-9a46-ea2588da0e5e
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix: address code review feedback — move imports, clarify scope comment, use shared test context
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3cb85c68-1792-473e-9a46-ea2588da0e5e
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix: address review comments, fix prettier formatting and lint errors
- Fix prettier formatting in 5 files (c-cpp, jvm, swift configs, test file)
- Remove unused SyntaxNode imports in php.ts and ruby.ts (lint errors)
- Remove unused constNodeSet/variableNodeSet variables in generic.ts (warnings)
- Remove semantically wrong `methodProps.isReadonly = varInfo.isConst` (review)
- Remove dead `nodeLabel === 'Variable'` guard in parse-worker (review)
- Fix test guard: replace `if (declNode)` with `expect(declNode).toBeDefined()` (review)
- Add comment about Python expression_statement broadness (review)
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/040edbbf-65b5-40e1-80c8-e98f7c4bb54a
* feat(ingestion): add block-scoped variable extraction via tree-sitter queries
Add @definition.const and @definition.variable tree-sitter query patterns
for TypeScript, JavaScript, Python, Go, Java, C, C++, C#, PHP, Ruby, and
Dart. Add parse-worker dedup logic to avoid duplicate nodes when variable
captures overlap with existing function/property captures. Add 'Variable'
label support in getLabelFromCaptures and DEFINITION_CAPTURE_KEYS.
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9fa828c1-87b7-4482-8f26-d2079fb4c58a
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* test: add block-scoped variable extraction tests and query capture tests
Add 6 tests for block-scoped variable extraction (TypeScript, Go, Rust, C,
Python). Add 14 tests verifying @definition.const/@definition.variable
query patterns exist in all language query strings. Import RUBY_QUERIES
in test file.
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9fa828c1-87b7-4482-8f26-d2079fb4c58a
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* test: add Python non-assignment expression statement rejection test
Addresses code review feedback: verify that the Python variable extractor
returns null for expression_statement nodes that contain function calls
rather than assignments (e.g. `print("hello")`).
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9fa828c1-87b7-4482-8f26-d2079fb4c58a
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix: Dart query node type, add Variable schema, update schema counts
- Change `top_level_variable_declaration` → `declaration` in DART_QUERIES
(the former doesn't exist in tree-sitter-dart grammar, causing all
Dart integration tests to fail with TSQueryErrorNodeType)
- Add VARIABLE_SCHEMA to schema.ts and register in initLbug() so that
Variable-labeled nodes are persisted to LadybugDB (not silently dropped)
- Add 'Variable' to MULTI_LANG_TYPES in csv-generator.ts
- Update Dart variable config to remove invalid node type
- Update schema test counts (30→31 node schemas, 32→33 total)
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/f79931d1-207f-4fbb-91da-259d44f7fd88
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix: address code review comment improvements
- Clarify processedDefinitionNodes tracks start indices, not nodes
- Improve Python variableNodeTypes comment wording
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/f79931d1-207f-4fbb-91da-259d44f7fd88
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix: add Variable to NODE_TABLES, RELATION_SCHEMA, update golden snapshot
- Add 'Variable' to NODE_TABLES in gitnexus-shared so validTables.has('Variable')
returns true and Variable graph edges are not silently dropped
- Add FROM File TO Variable, FROM Variable TO Community, FROM Variable TO Process
to RELATION_SCHEMA so KuzuDB can represent edges connecting Variable nodes
- Update schema.test.ts: add Variable to multiLang list, fix count 30→31
- Regenerate pipeline-graph-golden snapshot for mini-repo fixture
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e3aad558-e7bb-40d1-b53f-0a2c0132ca96
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix: isolate golden test from cli-e2e fixture pollution
The pipeline-graph-golden test was non-deterministic because cli-e2e.test.ts
creates AGENTS.md, CLAUDE.md, .claude/skills/, and .gitignore in the shared
mini-repo fixture during analyze. These leftover files caused the golden test
to find 9 files instead of 7 when tests ran in parallel.
Fixes:
- Golden test now copies the fixture to a temp dir before running, making it
immune to concurrent test pollution
- cli-e2e afterAll cleanup now removes ALL generated files (AGENTS.md,
CLAUDE.md, .claude/, .gitignore) not just .git/ and .gitnexus/
- Golden snapshot regenerated from clean 7-file fixture
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/bd378e73-6f37-49c6-aed6-7fabf4dc6183
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* chore(deps): add tree-sitter aware Dependabot config and drift monitoring
Two things Dependabot cannot see on its own:
1. ABI consistency. The tree-sitter runtime supports a known range of
grammar ABIs. When a grammar bumps past that range, require() silently
fails and fallback paths mask the regression in test coverage.
2. Vendored upstream drift. vendor/tree-sitter-proto is a snapshot of
coder3101/tree-sitter-proto regenerated against a pinned cli version.
Upstream keeps moving. Nothing notices until a maintainer remembers to
look.
Dependabot configuration
- Added npm ecosystems for gitnexus, gitnexus-web, gitnexus-shared.
- Grouped all tree-sitter-* grammar bumps into one PR (ecosystem moves in
lockstep, one PR per grammar is noise).
- Pinned the tree-sitter runtime itself. Bumping 0.21 to 0.22+ changes
which grammar ABIs load and requires coordinated updates to the
vendored proto grammar. That stays a deliberate human decision.
- Pinned tree-sitter-cli for the same reason (it controls which ABI
vendor/tree-sitter-proto/src/parser.c emits when regenerated).
Drift check (.github/scripts/check-tree-sitter-drift.py)
- Reads the tree-sitter runtime version from gitnexus/package.json.
- Walks every installed tree-sitter-* grammar plus the vendored proto
and reports its LANGUAGE_VERSION against the runtime's supported ABI
range (table maintained in the script; extend when bumping runtime).
- Fetches coder3101/tree-sitter-proto main parser.c and compares byte
for byte to the vendored copy. Reports the upstream HEAD short SHA
and the upstream ABI so a maintainer can act.
- Prints a Markdown report; exits 0 when everything is in range and
matches upstream, 1 otherwise.
- Stdlib only, no external deps.
Drift workflow (.github/workflows/tree-sitter-drift-check.yml)
- Runs weekly (Mondays 09:00 UTC) to match Dependabot's cadence.
- Also runs on PRs that touch the script or workflow itself, where it
fails the PR check on drift so the drift gate cannot land broken.
- On scheduled runs with drift, opens or updates a single tracking
issue labeled tree-sitter-drift. On scheduled runs that come back
clean, closes the open tracking issue (if any) with a comment.
* refactor(deps): rewrite drift check as tree-sitter 0.25 upgrade readiness monitor
Replace the ABI drift pass/fail gate with a daily upgrade readiness
dashboard that tracks peer-dep compatibility of all 14 grammars with
tree-sitter@0.25.0 and reports which are ready, unreleased, or blocking.
Key changes:
- Rename drift-check → upgrade-readiness (script, workflow, job id)
- Fix P0: pass report via env var, not ${{ }} template interpolation
- Fix P1: npm fetch failure now adds a blocker instead of false-green
- Fix P1: pass GITHUB_TOKEN for authenticated GitHub API calls
- Switch Dependabot to daily for tree-sitter grammars
- Use dict for blockers (no prefix collision), derive TARGET_RUNTIME
constant, reuse GRAMMARS parser_path, normalize CRLF in comparisons
- Reduce per-call HTTP timeout from 15s to 8s for workflow budget
- PR runs warn on blockers instead of hard-failing
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* chore(ci): remove global-upgrade smoke test workflow
The ci-global-upgrade.yml workflow tested npm global install upgrades
over a specific release candidate (1.6.2-rc.8). That RC has shipped
and the workflow is no longer needed. Remove it and all references
from ci.yml (needs, env vars, gate check).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(ci): add changelog comments to upgrade readiness tracking issue
Each daily run now posts a comment summarizing what changed before
updating the issue body. Comments include the ready/blocker counts
and a diff of grammar status changes (e.g. tree-sitter-cpp:
Unreleased -> Ready). Gives a timeline of how the upgrade unblocks.
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
release-drafter v7 (merged in #852) removed the `disable-releaser`
input, causing the autolabel job to attempt creating a release and
fail with "Resource not accessible by integration". Replace with
`dry-run: true` which achieves the same label-only behavior.
Also update stale version comments for release-drafter and
action-semantic-pull-request to match the actual pinned versions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: devendor tree-sitter-proto install lifecycle to fix ENOTEMPTY on global upgrade
PR #843's preinstall cleanup hook cannot address the reported bug because
it runs on the NEW package's staging tree, not the OLD install being
removed. Issue #836 still reproduces on 1.6.2-rc.8.
Root cause: vendor/tree-sitter-proto was declared as `file:` dep with its
own `dependencies` and `install` script, so npm created
`vendor/tree-sitter-proto/node_modules/node-addon-api/` at install time,
which blocked npm's rmdir on global upgrade.
Changes:
- Strip `dependencies` and `install` script from the vendored sub-package's
package.json so npm no longer creates a nested node_modules or runs a
lifecycle script under vendor/.
- Hoist `node-addon-api` and `node-gyp-build` into gitnexus
optionalDependencies; npm resolves them at the consumer's top level.
- Add scripts/build-tree-sitter-proto.cjs modeled on patch-tree-sitter-swift.cjs.
Runs at gitnexus postinstall, best-effort: skips cleanly on missing
toolchain or --ignore-scripts so non-proto functionality keeps working.
- Remove scripts/preinstall-cleanup.cjs — dead code; cannot run against
the old install being removed.
- Keep .npmignore entries from PR #843 (tarball hygiene, still correct).
- Add explicit .gitignore rules for gitnexus/vendor/**/build and
gitnexus/vendor/**/node_modules (closes the repo-side hygiene gap).
- Add .github/workflows/ci-global-upgrade.yml: matrix smoke test that
installs the previously-published rc globally, upgrades to the packed
current branch, and verifies no vendor install-time artifacts survive.
Runs on macOS (reporter's platform), Linux, and Windows. Also includes
an --ignore-scripts degraded-mode lane. Wired into ci.yml gate.
Plan: docs/plans/2026-04-15-002-fix-tree-sitter-proto-vendor-deps-plan.md
Phase 1 (this commit) addresses the reported `node_modules/node-addon-api`
hazard. Phase 2 (follow-up) will migrate to prebuildify + prebuilt .node
binaries in the tarball — the 2026 canonical shape for tree-sitter
grammars, which eliminates the postinstall compile path entirely.
Refs #836
* fix(ci): ci-global-upgrade should be reusable-only and use setup-gitnexus
Three issues caught by CI on PR #846:
1. Concurrency linter rejected the `CIGU-` prefix (allowlist is
`${{ github.workflow }}` or substring `CI-`). The literal-prefix
guidance in ci.yml is specifically about disambiguating when
reusable workflows run in nested contexts, and ci-global-upgrade
doesn't need its own concurrency block at all — the caller
(ci.yml) already governs concurrency for nested invocations.
2. `npm install` in gitnexus/ runs `prepare: node scripts/build.js`,
which depends on gitnexus-shared/dist being built first. Other CI
jobs handle this via the setup-gitnexus composite action. Use it
here too (with build: 'false' — we only need the dep graph, then
npm pack runs prepack which builds gitnexus itself).
3. Removed `pull_request` and `workflow_dispatch` triggers. The
workflow is now pure `workflow_call` — invoked once from ci.yml
via `uses:`. This avoids the duplicate-run problem where both the
top-level pull_request trigger AND the nested workflow_call would
fire on every PR.
* fix(ci): relax vendor build/ guard and use bash shell on Windows
Two fixes for ci-global-upgrade failures on PR #846:
1. The guard after the upgrade step was rejecting vendor/tree-sitter-proto/build/
in the global install. That was too strict. The original #836 bug was
about vendor/tree-sitter-proto/node_modules/ specifically, not build/.
The build/ directory appears because node-gyp-build compiles through the
symlink npm creates at node_modules/gitnexus/node_modules/tree-sitter-proto,
and its contents are plain .node, .obj, .lib files that rmdir handles
without trouble. We know this empirically because the test got past the
upgrade step in the run where the old vendor/node_modules was present.
The guard now only flags nested node_modules, which is what the fix
actually removes.
2. The Windows --ignore-scripts lane failed with ENOENT when npm tried to
open the tarball. The path was computed in a bash step using $(pwd),
which on Windows returns /d/a/... form, but npm install ran in the
default cmd shell and received a mangled Windows path. Adding
shell: bash to the install steps keeps path handling consistent.
When @huggingface/transformers is installed globally (e.g. via npm install -g),
it defaults its cache directory to ./node_modules/.cache inside its own install
dir, which is unwritable by non-root users.
This causes EACCES errors on first use when the model is downloaded:
EACCES: permission denied, mkdir '/usr/lib/node_modules/gitnexus/node_modules/@huggingface/transformers/.cache'
Set env.cacheDir before pipeline() is called in both embedders (CLI and MCP).
Respects HF_HOME env var if set, falls back to ~/.cache/huggingface.
Co-authored-by: Sisyphus <sisyphus@gitnexus.dev>
* ci: standardize workflow concurrency and automate release-note labeling
Concurrency — prevent racing CI jobs
- Every top-level workflow now declares an explicit concurrency block.
- PR runs cancel-in-progress on supersede; main/push/workflow_call/publish
runs queue instead of cancelling so every commit and every release is
validated end-to-end.
- ci.yml uses a literal `CI-` prefix (not `${{ github.workflow }}`) and a
per-run nested group for workflow_call invocations, avoiding a potential
deadlock with publish.yml and release-candidate.yml callers whose own
concurrency groups could otherwise collide with the called workflow.
- ci-report.yml falls back to `<head-repo>/<head-branch>` for fork PRs
(stable across reruns) instead of the per-run-unique workflow_run.id
which did not actually serialize anything.
- ci-quality.yml enforces the convention: fails CI if any non-reusable
workflow lacks a concurrency block or a reusable workflow declares one.
Release-note automation
- New pr-labeler.yml: amannn/action-semantic-pull-request enforces
conventional-commit PR titles on pull_request (fork-safe, read-only);
release-drafter/release-drafter with disable-releaser: true applies the
matching label under pull_request_target (write-scoped). sync-labels in
.github/release-drafter.yml removes managed autolabels that no longer
match (e.g. when `!` or `BREAKING CHANGE:` is dropped from a PR).
- .github/release.yml (unchanged) continues to map labels to categorized
release-notes sections.
- dependabot.yml added for the github-actions ecosystem so pinned SHAs
auto-refresh on a weekly cadence.
Docs
- CONTRIBUTING.md documents the concurrency convention, the
conventional-commit PR-title rules, and the reusable-workflow exception.
Follow-up to verify before relying on the labeler in anger
- gh api repos/amannn/action-semantic-pull-request/git/refs/tags/v5.5.3
- gh api repos/release-drafter/release-drafter/git/refs/tags/v6.0.0
- Confirm release-drafter reads its config from the base ref (not fork
head) when invoked via pull_request_target.
* ci: address PR review feedback on concurrency and labeler workflows
Two blocking fixes
- pr-labeler.yml: separate concurrency slots for pull_request and
pull_request_target. Previously both triggers shared a single group
with cancel-in-progress: true, so the privileged autolabel run could
cancel the title-validation check mid-run and leave a required status
in a permanent cancelled state.
- pr-labeler.yml autolabel job: add contents: read. release-drafter's
context.config() reads .github/release-drafter.yml from the default
branch via the repo-contents API and 403s without the scope. Job-level
permissions nullify all unlisted scopes so an explicit grant is needed.
Two non-blocking improvements
- Replace the hardcoded reusable-workflow allowlist in ci-quality.yml
with dynamic on:-block parsing. New workflow_call-only workflows no
longer produce false-positive convention failures.
- Implement actual group-key validation. The check now also asserts that
every concurrency.group expression references either ${{ github.workflow }}
or the literal CI- prefix (the documented ci.yml exception).
- Script extracted to .github/scripts/check-workflow-concurrency.py so
it is runnable locally and independently testable.
* Initial plan
* fix: stale vectors preserved on content edits and vector index missing after zero-node run
Issue 1: Add contentHash to EMBEDDING_SCHEMA and embedding pipeline.
- contentHash column persisted per CodeEmbedding row
- POST /api/embed queries nodeId+contentHash, compares per-node hash
- Stale rows (hash mismatch) are DELETE'd before re-embedding
- Legacy DBs without contentHash treated as stale (full re-embed)
- loadCachedEmbeddings and run-analyze cache restore include contentHash
Issue 2: createVectorIndex called unconditionally before zero-node early return.
Regression tests:
- contentHashForNode determinism and content-change detection
- EMBEDDING_SCHEMA includes contentHash STRING column
- Pipeline exports verified
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/1581c0c0-f359-4376-b47e-62d24a28fd2d
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix: use parameterized query for stale embedding DELETE, revert package-lock.json
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/1581c0c0-f359-4376-b47e-62d24a28fd2d
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix: address review feedback — config consistency, narrow catches, extract DB logic
Bug #1: Use finalConfig consistently in contentHashForNode (line 224 was
using raw `config` while line 307 used `finalConfig`). Cache precomputed
hashes in filter phase to avoid double computation (Perf #5).
Bug #2: Narrow catch in loadCachedEmbeddings to only fall back on
column/table-missing errors. Rethrow transient/connection errors.
Bug #3: Log non-trivial DELETE failures instead of silently swallowing.
Arch Violation #3: Extract fetchExistingEmbeddingHashes from api.ts into
lbug-adapter.ts. Server layer now calls a single adapter function instead
of re-implementing the DB query logic with nested try-catch.
Tests: Add config consistency test, note that fetchExistingEmbeddingHashes
tests require native module (run in CI).
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b8c4f6b0-4095-4507-a15d-d8469793efac
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix: narrow Column error match to 'contentHash' in lbug-adapter fallback checks
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b8c4f6b0-4095-4507-a15d-d8469793efac
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix: address production-readiness review — eliminate competing state, use schema constants, hard-fail on stale DELETE, add incremental filter tests
Gap A / Arch Violation 1: Remove duplicate vectorExtensionLoaded flag from
embedding-pipeline.ts — delegate to lbug-adapter's loadVectorExtension()
which owns the VECTOR extension lifecycle and resets on DB reconnect.
Arch Violation 2: Replace all hardcoded 'CodeEmbedding' and
'code_embedding_idx' strings in embedding-pipeline.ts and run-analyze.ts
with EMBEDDING_TABLE_NAME, EMBEDDING_INDEX_NAME, and CREATE_VECTOR_INDEX_QUERY
imported from schema.ts. Add EMBEDDING_INDEX_NAME export to schema.ts.
Gap B: Make DELETE failure for stale vectors a hard throw (not just a
warning). Continuing after failed DELETE risks Kuzu vector-index corruption
since the constraint requires DELETE-before-INSERT for vector-indexed
properties. "not found" / "does not exist" errors are still safe to ignore.
STALE_HASH_SENTINEL: Define a named constant in embedding types.ts for the
empty-string sentinel convention. Used consistently in lbug-adapter.ts and
run-analyze.ts so the invariant is self-documenting.
Tests: Add comprehensive unit tests for the incremental filter logic with
mocked embedder:
- New node → embedded
- Unchanged node (hash matches) → skipped
- Stale node (hash mismatch) → DELETE + re-embed
- STALE_HASH_SENTINEL → treated as stale
- Zero nodes after filter → createVectorIndex still called
- DELETE failure with non-trivial error → throws
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b21edee7-c9c5-4742-947b-d0def4fb26aa
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix: tighten error classification — extract isMissingColumnOrTableError helper, remove broad pattern matching
- Extract isMissingColumnOrTableError() helper in lbug-adapter for
consistent schema-error detection (replaces duplicate inline checks)
- Tighten 'contentHash' match: now requires 'property' AND 'contentHash'
(Kuzu-specific pattern) instead of broad 'contentHash' substring
- Tighten DELETE error check: only ignore 'does not exist' (Kuzu's actual
message), not broad 'not found' which could mask connection errors
- Fix test node ID/name/filePath consistency
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b21edee7-c9c5-4742-947b-d0def4fb26aa
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix: CI failures and final review — move STALE_HASH_SENTINEL to schema, tighten error matching, fix test mocking, format
- Move STALE_HASH_SENTINEL from embeddings/types.ts to lbug/schema.ts
(fixes inverted layer dependency: lbug should not import from embeddings)
- Tighten isMissingColumnOrTableError: replace broad msg.includes('not found')
with /(table|column|property).*not found/i regex to avoid matching transient errors
- Add vi.resetModules() in test beforeEach for explicit module isolation
(fixes vi.doMock not intercepting loadVectorExtension in CI)
- Skip precomputedHashes.set() on unchanged (return false) path
- Run prettier on all 5 files flagged by CI format check
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e20311fd-4361-47b4-a137-9adc3e533b35
* fix: address remaining review nits — rename precomputedHashes, generalize error matcher, revert package-lock
- Rename precomputedHashes → computedStaleHashes (hashes are computed
on-demand during filter, only cached for stale nodes being re-embedded)
- Remove contentHash-specific clause from isMissingColumnOrTableError —
the regex /(table|column|property).*not found/i already covers it
- Revert package-lock.json ssh→https protocol change
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e20311fd-4361-47b4-a137-9adc3e533b35
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix(group/sync): wire ManifestExtractor into syncGroup pipeline
ManifestExtractor was fully implemented in extractors/manifest-extractor.ts
but never imported or called in sync.ts. As a result, any links declared in
group.yaml were parsed and validated by config-parser.ts but silently dropped
— config.links was always an empty dead-end as far as syncGroup was concerned.
Changes:
- Import ManifestExtractor in sync.ts
- Call extractFromManifest(config.links, dbExecutors) inside the outer try
block, after all repos are processed but before the finally closes the DB
pools (symbol resolution via resolveSymbol requires open executors)
- Collect the resulting contracts into autoContracts and the cross-links into
a separate manifestCrossLinks array
- Merge manifestCrossLinks into the final crossLinks alongside runExactMatch
results
Without this fix, users who declare explicit service dependencies in
group.yaml links (the documented workaround for HTTP clients that use absolute
URLs and are invisible to the auto-extractors) get 0 cross-links regardless of
what they configure.
* test(group/sync): cover manifest links producing cross-links
Add a unit test that asserts config.links entries produce contract pairs
and a manifest cross-link (matchType: 'manifest') via syncGroup.
Also refactors the manifest extraction call to sit outside the else/try
block so it runs regardless of extractorOverride arity — makes the code
testable without mocked DB pools and ensures links work when callers supply
a zero-arity override (e.g. in tests or programmatic usage).
* style: prettier format sync.ts and sync.test.ts
Also removes the stray empty line in the finally block (noted in review).
* fix(group/sync): dedupe cross-links and warn on dangling manifest repos
Addresses review feedback on PR #827:
1. Dedupe cross-links. Manifest contracts participate in runExactMatch, so a
manifest-declared link also emitted a duplicate matchType:'exact' CrossLink
for the same endpoint pair. Dedupe by (from, to, type, contractId) and
prefer manifest (operator-declared intent).
2. Warn on dangling repos. When a manifest link references a repo not in
config.repos, log a warning. Synthetic UIDs keep the cross-link
deterministic, but the operator probably meant something else.
3. Tests:
- Assert no duplicate 'exact' CrossLink is emitted alongside the manifest one.
- Assert synthetic UID format when no DB executors are available.
- New test: dangling manifest repo still produces a cross-link + logs a warning.
* perf(group/manifest): parallelize and memoize symbol resolution
Previous implementation ran 2N sequential Cypher round-trips per
manifest (one for provider side, one for consumer, awaited in-order
per link). For manifests with tens of links this dominated syncGroup
latency in groups with many declared cross-repo contracts.
Changes:
- Resolve provider + consumer in parallel per link (Promise.all).
- Resolve all links in parallel (outer Promise.all over links.map).
Each repo's executor pool is independent, so cross-repo fan-out
scales with the number of distinct repos in the manifest.
- Memoize by (repo, type, contract). Manifests frequently declare
the same contract from both directions or across sibling groups,
so duplicate triples now hit the DB once instead of 2× per link.
Correctness:
- resolveSymbol is a pure LIMIT 1 read, so caching + concurrent
invocation is safe.
- Iteration order over links is preserved in the final
contracts / crossLinks arrays — result shape is identical.
Test:
- New test asserts that two links sharing (repo, type, contract)
produce exactly one DB call per distinct repo-tuple.
---------
Co-authored-by: jonasvanderhaegen-xve <>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* fix(lbug): wait for read stream close in splitRelCsvByLabelPair (Windows ENOTEMPTY)
The windows-latest CI job intermittently failed:
FAIL test/unit/rel-csv-split.test.ts > splitRelCsvByLabelPair > handles empty CSV (header only) without errors
Error: ENOTEMPTY: directory not empty, rmdir 'C:\Users\RUNNER~1\AppData\Local\Temp\rel-csv-test-XW5KOu'
Cause: splitRelCsvByLabelPair resolved its Promise on readline's 'close'
event, but the underlying fs.ReadStream's file descriptor is released
asynchronously after that — especially on Windows. For the empty-CSV
test the function returns so quickly that afterEach fires rmSync while
the relations.csv fd is still held, so Windows reports ENOTEMPTY on
the directory.
Fixes:
- Production: after readline 'close', wait for inputStream 'close' (or
resolve immediately if already closed/destroyed). Call inputStream
.destroy() defensively so we never hang if the fd never emits 'close'.
- Test: afterEach now retries rmSync up to 5 times on ENOTEMPTY/EBUSY/
EPERM with a brief back-off — defense-in-depth so the test doesn't
flake on slow CI runners independent of the production change.
The production fix benefits every caller, not just the test: any code
that deletes the CSV's parent directory right after the Promise
resolves previously hit the same race on Windows.
* refactor(lbug): replace custom stream state machines with stdlib primitives
Full audit of splitRelCsvByLabelPair's stream usage after the original
ENOTEMPTY fix. Replaced three hand-rolled mechanisms with their
standard-library equivalents — 147 -> 71 lines in the function, and
the caller's WriteStream closure dropped from 13 lines to 5.
- readline: 'on(line)' + pause/resume/waitingForDrain state machine
-> 'for await (const line of rl)'. Async-iterator delivery naturally
serializes line processing with our awaits, so at most one ws is in
backpressure at a time. We just 'await once(ws, "drain")' when
'write()' returns false — the custom Set, the settled flag and the
'only resume when all streams have drained' logic all go away.
- Multi-stream error coordination: hand-rolled cleanup() that had to
be entered exactly once and had to destroy the inputStream and every
pair ws -> single AbortController shared across every 'once(ws,
'drain', { signal })'. Any stream error aborts every pending wait.
- 'stream/promises.finished(inputStream)' in the 'finally' block
replaces the manual 'rl.on('close', () => inputStream.once('close',
...))' dance, and covers both the success and error paths with the
same primitive. This closes the Windows ENOTEMPTY race root cause —
we never return while the fd might still be in flight.
- Caller closure: 'new Promise((res, rej) => ws.end(cb) + remove
listener on error)' -> 'ws.end(); await finished(ws)'.
- Test 'afterEach': custom retry loop -> 'fs.rmSync(..., { maxRetries:
5, retryDelay: 50 })' (Node added these options specifically for
cross-platform tmpdir cleanup).
- Test 'destroys all streams when one errors': old code leaked
backpressure and created multiple pair streams before the first
blocked; new strict serial backpressure doesn't, so the test now
unblocks the first stream once to advance the loop and create the
second stream before triggering the error.
* fix(csv-generator): deduplicate all node types, not just File nodes
The pipeline can produce duplicate node IDs across all symbol types
(Class, Method, Function, etc.). Only File nodes were guarded by a
seenFileIds Set, leaving every other type unprotected. When the CSV
was COPY'd into LadybugDB, duplicate PKs caused mass "Batch execution
error: Found duplicated primary key value" warnings on gitnexus serve.
Replace the per-type seenFileIds with a single seenNodeIds Set checked
at the top of the iteration loop, before the switch, so every label is
covered by the same O(1) deduplication guard.
Fixes: #822
* fix(embeddings): use MERGE instead of CREATE for CodeEmbedding inserts
CREATE fails with duplicate PK when a CodeEmbedding node already exists,
which happens when:
- A PostToolUse hook triggers a concurrent gitnexus analyze during an
active analyze run (git commits fire the hook)
- A partial prior run left some embeddings in the DB before a crash
Switching to MERGE makes the insert idempotent: existing embeddings are
updated in place, new ones are created, no PK violations.
Fixes: #822
* fix(server): skip already-embedded nodes in POST /api/embed to avoid vector-index SET error
Kuzu/LadybugDB forbids SET on a property that is part of a vector index.
The /api/embed endpoint was calling runEmbeddingPipeline without skipNodeIds,
causing it to attempt MERGE+SET on every node including those already embedded.
Fix: query existing CodeEmbedding nodeIds before running the pipeline and pass
them as skipNodeIds so only new (unembedded) nodes are processed.
* fix(server): narrow catch to table-not-exist errors only in POST /api/embed
Bare catch{} would silently swallow connection errors and proceed to
re-embed all nodes, hiding infrastructure issues. Now only swallows
errors where the CodeEmbedding table does not yet exist.
* style: prettier format gitnexus/src/server/api.ts
* fix(server): log skip-embedding count and table-not-found swallow path
Addresses review feedback on PR #823:
- Log count of already-embedded nodes when skipNodeIds is populated
(aids debugging if Kuzu driver row shape changes).
- Log when the 'table does not exist' swallow path fires so ops can
catch it if Kuzu ever changes error wording.
- Document the {} config positional argument with an inline comment
referencing the runEmbeddingPipeline signature.
---------
Co-authored-by: jonasvanderhaegen-xve <>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* feat(ci): add release-candidate publish pipeline
Auto-publishes gitnexus@rc on every merge to main. Version scheme is
canonical semver X.Y.Z-rc.N where the base is the current npm 'latest'
bumped by the 'bump' input (default patch) and N auto-increments by
querying existing rc versions on the registry. First rc for a new base
is rc.1; the counter resets naturally when the base advances after a
stable release.
- Reuses ci.yml via workflow_call so tests must pass before publish
- SHA-pinned actions, per-job permission scoping, provenance enabled
- Guard job dedupes duplicate dispatches against HEAD via v*-rc.* tags
- Docs-only pushes skipped via paths-ignore
- workflow_dispatch inputs: bump (patch/minor/major), force (override guard)
- Publishes under the 'rc' dist-tag so 'latest' is never moved
- Tags commits as v<rc-version> and creates GitHub prereleases
* fix(ci): address release-candidate review feedback
- Sort rc tags by creatordate (handles out-of-order pushes correctly)
- Fail fast on npm registry errors; only fall back to package.json on E404
- Drop unused pull-requests: write permission on the reused CI job
- Add secrets: inherit so any future CI secrets are available to sub-jobs
- Remove unused reltag step output
* fix(ci): address Copilot review comments
- Correct concurrency comment (runs serialize on same ref, not overlap)
- Apply E404-only fallback to 'npm view versions' query, matching the
pattern used for the 'npm view version' query
- README: clarify that docs-only merges don't trigger rc publish
- CONTRIBUTING: drop 'from main' claim for publish.yml; the tag-push
trigger does not enforce branch reachability
* fix(ci): address adversarial review — idempotency, cycle continuity, tag integrity
Codex adversarial review flagged three release-safety issues in the rc
pipeline. Fixes:
1. Cycle continuity (H). Non-patch rc trains no longer collapse back to
patch on the next push. 'bump' input accepts a new 'auto' value
(default) that infers the active rc base from the registry: if any
X.Y.Z-rc.* exists with X.Y.Z > latest, continue that base; otherwise
patch-bump. Explicit patch/minor/major still forces a cycle reset and
now also bypasses the dedup guard so an explicit dispatch on a
tagged HEAD is honored.
2. Idempotency across post-publish failures (H). The guard marker
('rc/<HEAD_SHA>' lightweight tag) and the release tag ('v<RC>'
annotated) are now pushed atomically *before* 'npm publish'. A
publish failure leaves the marker in place and the guard refuses to
re-publish. Added a defensive 'npm view <pkg>@<rc> version' check
before publish to catch registry-level races. Recovery path
documented in CONTRIBUTING.md.
3. Tag ↔ package integrity (M). 'v<RC>' now points at a detached
release commit whose tree contains the rewritten package.json, so
the tag's source archive matches the npm tarball exactly. 'main'
stays pristine; the release commit is reachable only via the tag.
* fix(ci): surface registry errors on defensive version check; drop actions: read
- npm view <pkg>@<rc> version now distinguishes E404 (safe) from network
failures (abort) via the same mktemp+grep pattern used for the other
two npm view calls
- Dropped actions: read on the ci workflow_call — no sub-workflow uses
the Actions API
* fix: add setMaxListeners(50) to relationship pair WriteStreams
Dynamically-created per-pair WriteStreams for relationship CSV splitting
default to Node.js's maxListeners limit of 10. On large repositories with
many relationship types, readline backpressure causes repeated
ws.once('drain', ...) calls that exceed this limit, flooding stderr with
MaxListenersExceededWarning messages.
This matches the existing pattern in csv-generator.ts where
BufferedCSVWriter already calls this.ws.setMaxListeners(50).
* fix: address all 3 stream bugs in relationship CSV splitting
Addresses review feedback from @magyargergo and Claude CI analysis:
Bug 1 (High): Add error handlers to per-pair WriteStreams.
Previously, if a WriteStream errored (disk full, EMFILE) while rl was
paused waiting for drain, the drain callback never fired, rl.resume()
was never called, and the outer Promise hung forever — leaking all
open file descriptors until process kill.
Now each WriteStream gets an error handler that destroys all streams,
closes the readline interface + its input ReadStream, and rejects the
Promise.
Bug 2 (Medium): Add waitingForDrain Set to prevent drain listener
accumulation. rl.pause() is not synchronous — buffered line events
continue firing after pause(), and multiple lines targeting the same
pairKey each added another ws.once('drain', ...) listener. This was the
root cause of MaxListenersExceededWarning.
Now a Set<string> tracks which streams are already waiting for drain.
Only the first backpressure event registers the listener; subsequent
lines for the same stream are silently skipped (they're already written
to the stream buffer). This eliminates listener accumulation entirely
and makes setMaxListeners(50) a safety net rather than a band-aid.
Bug 3 (Low): Close readline and destroy input ReadStream in error
handler. Previously only the WriteStreams were destroyed on error,
leaving the ReadStream FD to linger until GC.
* fix: address review feedback — remove setMaxListeners, harden cleanup
- Remove setMaxListeners(50) entirely. The waitingForDrain guard
guarantees at most 1 drain listener per stream at any time. Tested
with 200 pairs x 500 lines (100k total) — max listeners was always 1,
zero warnings. No hard-coded limit needed.
- Wrap destroy() calls in cleanup() with try/catch so already-destroyed
streams don't throw synchronously (addresses @xkonjin review point 1).
- Add ws.once('error', reject) to the ws.end() phase so flush errors
during stream close properly reject instead of hanging Promise.all
(addresses Claude CI Bug 3b finding).
* test: add 8 regression tests for relationship CSV stream fixes
Covers all bugs fixed in this PR:
- Bug 1: WriteStream error rejects Promise and destroys all streams
- Bug 2: waitingForDrain guard keeps drain listeners at max 1 per stream
- Bug 3: cleanup() handles already-destroyed streams safely
Tests use a MockWriteStream with controllable backpressure and error
injection to verify the exact patterns in loadGraphToLbug() without
needing a real LadybugDB instance.
* style: run prettier on changed files
* fix(test): use backpressure to keep promise pending during error tests
The error tests were racing — readline finished reading the tiny CSV
and resolved the Promise before setTimeout fired the error. Now the
mock streams use blocked=true to trigger backpressure, keeping the
Promise pending so the error fires while the split is still in progress.
* fix: use named error handler in ws.end() to prevent listener leak
ws.once() wraps the callback, so removeListener with the original
function reference won't match. Switch to ws.on() with a named
onError function so removeListener correctly detaches it after
successful close.
* refactor: extract splitRelCsvByLabelPair, fix multi-stream drain
1. Extract splitRelCsvByLabelPair as an exported function with optional
wsFactory parameter for dependency injection. loadGraphToLbug now
delegates to it. Tests import and call the real function instead of
a local reimplementation.
2. Fix multi-stream drain coordination: rl.resume() is now guarded by
waitingForDrain.size === 0, so readline only resumes when ALL
backpressured streams have drained. Previously, any single stream
draining would resume readline while other streams were still full,
allowing unbounded buffer growth.
3. Export WriteStreamFactory type and RelCsvSplitResult interface for
test consumption.
* fix: resolve C/C++ cross-file calls through transitive #include chains
In C/C++, #include is transitive: if a.c includes b.h and b.h includes
c.h, then a.c can call any function declared in c.h. The wildcard import
synthesis only walked direct imports (1 hop), missing symbols reachable
through transitive header chains.
This is the dominant pattern in large C codebases — Redis's db.c includes
server.h which includes dict.h, so db.c should resolve calls to dictFind()
declared in dict.h and defined in dict.c. Before this fix, those cross-file
call edges were missing entirely.
The fix expands the import closure transitively for C/C++ files before
synthesizing wildcard bindings. A BFS walks ctx.importMap and graphImports
to collect all transitively reachable headers, then passes the full closure
to synthesizeForFile.
Tested on Redis (github.com/redis/redis):
- Before: dictFetchValue had 0 cross-file callers, processCommand had 0
- After: dictFetchValue has 9 callers, processCommand has 1, +1946 edges total
Fixes#813
* refactor(ingestion): dispatch wildcard synthesis by import-semantics strategy
Generalize PR #816's C/C++ transitive #include fix into a language-agnostic
strategy pattern. The `wildcard-synthesis.ts` pipeline phase no longer
references `SupportedLanguages.C` / `SupportedLanguages.CPlusPlus` — it
dispatches on `provider.importSemantics` via an exhaustive `switch`.
Also fixes a correctness bug the original BFS introduced: `queue.pop()`
(LIFO/DFS) reversed the iteration order of `#include` directives, which —
combined with first-seen-wins dedup in `synthesizeForFile` — silently
bound overloaded symbols to the wrong header. For the `cpp-calls`
fixture, `write_audit("hello")` was being resolved to `zero.h`'s arity-0
overload instead of `one.h`'s arity-1 overload, breaking arity
narrowing. Switched to FIFO (`queue.shift()`) with direct imports seeded
in declaration order.
Taxonomy (researched across 20+ languages + stack-graphs / SCIP prior art):
| Tag | Traversal | Languages |
|---------------------|-----------------|------------------------------------|
| named | none | TS, JS, Java, C#, Rust, PHP, Kotlin|
| wildcard-transitive | BFS closure | C, C++ |
| wildcard-leaf | single hop | Go, Ruby, Swift, Dart |
| namespace | none at import | Python |
| explicit-reexport | topological DAG | (scaffold; TS `export *` future) |
Changes:
- Widen `ImportSemantics` union from 3 to 5 tags with full taxonomy JSDoc
- Retag 5 providers: c-cpp (x2) → wildcard-transitive; dart, go, ruby,
swift → wildcard-leaf
- Move BFS closure into `wildcard-synthesis.ts` as `expandTransitiveIncludeClosure`
(pipeline-owned; providers stay pure declarations)
- Replace `if (lang === C || CPP)` with `dispatchSynthesis` helper called
by both Loop 1 (ctx.importMap) and Loop 2 (graphImports) so a future
transitive language whose edges arrive via graphImports gets closure
expansion consistently
- `never`-assertion default arm forces compile-time exhaustiveness
- `explicit-reexport` arm falls through to leaf behavior (scaffold;
TODO: implement re-export DAG walk for TS `export *` / Rust `pub use`)
- New unit tests covering circular includes, deep chains, diamond dedup,
graphImports-only paths, and order-preservation (the regression fix)
Verification:
- All existing C/C++ transitive tests pass unchanged
- Previously failing `cpp.test.ts > resolves run → write_audit to one.h
via arity narrowing` now passes
- `tsc --noEmit` clean
- 225/225 tests pass across wildcard-synthesis, cross-file-binding,
cpp resolver, and new closure unit tests
* fix(ingestion): bound closure size, O(1) dequeue, track Strategy 4 (#816 review)
Address @xkonjin's review feedback on the import-resolution strategy refactor:
1. **DoS guard**: cap transitive closures at 5,000 files via
`MAX_TRANSITIVE_CLOSURE_SIZE`. Pathological codebases (boost-style headers,
monoheader kernels) could previously produce closures with tens of thousands
of entries per translation unit. BFS now stops early and returns a partial
closure rather than risking OOM. The closest-headers-first BFS ordering
means the partial closure still contains the files overload resolution
cares about.
2. **Perf**: replace `Array.prototype.shift()` (O(n)) with a head-index queue
(O(1) dequeue). Deep chains previously had quadratic BFS behavior; now
linear in closure size.
3. **Strategy 4 tracking**: change TODO in `dispatchSynthesis` to
`TODO(#821)` referencing the filed issue for TS `export *` / Rust
`pub use` DAG-walk implementation, and clarify that today's leaf
fallthrough preserves correctness for direct imports — only the extra
re-export traversal is missing.
4. **Test**: new unit test exercising the 5,000-file cap on a 10k-file
synthetic chain, verifying partial-closure invariants (starts from
importer side, bounded, deep nodes excluded).
Not addressed in this commit (followups):
- Review point 3 (graphImports-only deep-chain *integration* fixture):
unit tests already exercise the `graphImports` traversal path directly
in isolation and combined with `importMap`. A fixture that stresses
graphImports-only transitive resolution is valuable but requires
understanding when the pipeline populates graphImports distinctly from
ctx.importMap — tracking as a followup rather than blocking this PR.
---------
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* fix(extractors): resolve 3 silent contract mis-resolution bugs (#793)
Addresses Codex adversarial review findings for extractor contract
resolution on the new group extractor surface.
F1 (manifest-extractor): resolveSymbol passed the full "METHOD::path"
contract string through normalizeRoutePath, producing "/GET::/api/orders"
which never matches Route.name. Adds parseHttpContract() helper that
strips the METHOD:: prefix before path normalization. Contract ID
construction (buildContractId) is unchanged.
F2 (http-route-extractor): graph-assisted backfill used path-only
detections.find(), so multi-verb same-URL files attached the wrong
verb/handler to provider rows and inferred the wrong verb on FETCHES
consumer edges. Now requires path+method match when method is known,
and skips backfill when method is unknown and multiple detections tie
on path.
F3 (grpc-extractor): resolveProtoConflict seeded bestScore=-1 and only
replaced on strict >, so all-zero-score ties silently selected
candidates[0]. Now computes all scores, counts ties at the top score,
and returns null on ambiguity (caller skips contract emission and
warns with service name + candidate paths).
All three fixes are test-first; 73 tests pass across the three suites.
No schema changes, no new dependencies, contract ID wire format
(http::METHOD::path, grpc::pkg.Service/Method, http::*::path) preserved.
* fix(extractors): address PR #817 review — ambiguous symbol pick + contract id casing
Copilot + Claude review on PR #817 flagged two follow-up bugs on top of
the F1/F2/F3 fixes:
1. http-route-extractor: ambiguous multi-verb case left handlerName null
but still ran the CONTAINS DB query. pickSymbolUid(syms, null) then
silently picked pool[0] — reintroducing handler mis-attribution via
a different route than the .find() bug F2 fixed. Now gates symbol
enrichment on an ambiguousCandidates flag so the file-basename
fallback wins instead.
2. manifest-extractor: buildContractId passed raw user casing through
for the explicit-method form, so get::/api/orders and
GET::/api/orders produced different contract ids even though
parseHttpContract upper-cases during lookup. Now reuses
parseHttpContract + normalizeRoutePath to canonicalize both method
and path, so logically equivalent manifest inputs share a contract
id (and share a manifestSymbolUid fallback).
Adds one regression test per bug: lowercase vs uppercase manifest
contract ids must match, and ambiguous multi-verb with CONTAINS rows
must not silently attach a real handler or call the CONTAINS query
at all. 75 tests pass across the three extractor suites.
* chore: prettier formatting
* Initial plan
* refactor: move language-specific container node logic into LanguageProvider
- Add resolveEnclosingOwner hook to LanguageProviderConfig
- Add staticOwnerTypes to MethodExtractionConfig
- Implement Ruby resolveEnclosingOwner (singleton_class → class/module)
- Replace hardcoded STATIC_OWNER_TYPES with config.staticOwnerTypes
- Move Ruby static types to rubyMethodConfig
- Move Kotlin static types to kotlinMethodConfig
- Remove Ruby singleton_class branch from findEnclosingClassInfo
- Collapse seqFindEnclosingClassNode/seqFindRawEnclosingContainerNode
into single provider-aware seqFindEnclosingOwnerNode
- Update worker path to pass provider.resolveEnclosingOwner
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/bc9f9d4d-f749-4872-9ff2-17fc86e08787
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* test: add regression tests for config-driven staticOwnerTypes and resolveEnclosingOwner hook
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/bc9f9d4d-f749-4872-9ff2-17fc86e08787
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* refactor: implement DAG-based pipeline architecture with phase extraction
Restructure the ingestion pipeline from a ~1800-line monolithic orchestrator
into a DAG (Directed Acyclic Graph) of named phases with explicit dependencies.
New files under pipeline-phases/:
- types.ts: PipelinePhase, PipelineContext, PhaseResult contracts
- runner.ts: DAG runner with topological sort validation
- scan.ts, structure.ts, markdown.ts, cobol.ts: early phases
- parse.ts + parse-impl.ts: chunked parse + resolve (the core)
- routes.ts, tools.ts, orm.ts: post-parse enrichment phases
- cross-file.ts + cross-file-impl.ts: cross-file binding propagation
- mro.ts, communities.ts, processes.ts: graph analysis phases
- index.ts: barrel export
pipeline.ts reduced from ~1960 lines to ~184 lines:
- DAG phase array declaration
- runPipelineFromRepo as thin orchestrator
- topologicalLevelSort retained for backward compat
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/136bf9c3-2f4f-449b-9fff-001332c8371c
* test: add DAG runner unit tests, update ARCHITECTURE.md with phase DAG docs
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/136bf9c3-2f4f-449b-9fff-001332c8371c
* fix: address code review - pass resolutionContext through parse output, fix worker URL path
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/136bf9c3-2f4f-449b-9fff-001332c8371c
* fix: declare transitive parse dependency explicitly in mro/communities/processes phases
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/136bf9c3-2f4f-449b-9fff-001332c8371c
* refactor: improve pipeline-phases clean code and folder structure
- Extract synthesizeWildcardImportBindings to wildcard-synthesis.ts
- Extract extractORMQueriesInline to orm-extraction.ts
- Create shared constants.ts for AST_CACHE_CAP
- Fix inline type import in orm.ts (use proper top-level import)
- Add comprehensive JSDoc to getPhaseOutput explaining type safety
- Move isDev to module level in cross-file.ts (consistency)
- Improve module-level documentation across files
- Organize barrel exports in index.ts with section comments
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2bd6d4aa-6271-4009-8dd2-332ea8ec73ab
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* address review feedback: fix circular dep, allFetchCalls mutation, progress bugs, remove DAG naming, extract isDev, fix _item naming, fix O(n²) line calc
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6cf53c9b-d55d-4c6f-bf3d-7bfb82d512b6
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* improve JSDoc on lineNumberAtOffset binary search
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6cf53c9b-d55d-4c6f-bf3d-7bfb82d512b6
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* address review: filter deps in runner, move totalFiles to ctx, fix cycle JSDoc, centralize isDev, remove DAG naming
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b388424f-b939-4a94-97de-3855f9465564
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix doc consistency in graph-sort.ts module-level and function-level JSDoc
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b388424f-b939-4a94-97de-3855f9465564
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix(pipeline): wrap phase errors with phase name and emit terminal error progress event
Restores phase diagnostics at CLI/MCP boundary. runPipeline now wraps
phase.execute() in try/catch and rethrows with 'Phase <name> failed: ...'
preserving the original via { cause }. Also emits a terminal
{ phase: 'error' } progress event so subscribers see the failure before
the rejection propagates. Handler errors during error reporting are
swallowed to keep the original cause authoritative.
Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U1)
* fix(pipeline): move bindingAccumulator dispose into crossFile try/finally; make single-use
crossFile.execute() now wraps its body in try/finally so the accumulator
is released on both the happy path and when runCrossFileBindingPropagation
throws. Dev-mode telemetry stays inside the try block before dispose (all
three counters return 0 after dispose clears internal maps).
BindingAccumulator becomes single-use: appendFile after dispose now throws
'BindingAccumulator: use after dispose' instead of silently re-animating
via the old _disposed auto-clear. Docs updated; the only production
construction site (parse-impl) always creates a fresh instance per run,
so no caller relied on the re-use contract.
Residual risk documented in crossFile module JSDoc: a future phase
inserted between parse and crossFile that throws would still leak the
accumulator. Any such phase must manage accumulator lifetime explicitly.
Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U2)
* docs(pipeline): explain why importCtx teardown is safe before crossFile
Investigation (plan U3) confirms: `importCtx` (ImportResolutionContext)
is a scratch workspace with no downstream consumer after parse.
`resolutionContext` (returned to crossFile) is a distinct object that
owns importMap / namedImportMap / packageMap / moduleAliasMap / model,
and never closes over importCtx. cross-file-impl consumes only that
ctx via processCalls. The two confusingly-similar "context" names
were the root of the adversarial reviewer's concern — comment locks
in the invariant so the next reader sees it.
No behavioral change.
Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U3)
* refactor(pipeline): remove ctx.totalFiles side-channel; promote to ParseOutput
totalFiles was a hidden mutable field on PipelineContext written by
parse and read by mro/communities/processes — five reviewers flagged
this as a violation of the immutable-context invariant. Removed from
PipelineContext, which is now fully readonly, and made the implicit
temporal dep explicit: mro/communities/processes now declare 'parse'
as a dep and read totalFiles via getPhaseOutput<ParseOutput>(...).
No behavior change. Topo-sort unchanged because parse was already a
transitive dep through crossFile.
Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U4)
* feat(method-extractor): runtime staticOwnerTypes guard at factory chokepoint
createMethodExtractor now rejects MethodExtractionConfigs that list
companion_object / singleton_class / object_declaration in
typeDeclarationNodes but omit the matching entry from staticOwnerTypes.
Fails loudly at provider construction time instead of producing
silent isStatic=false on the 50000th file analyzed.
Opt-out convention preserved: an explicit `new Set()` (empty Set)
signals intentional exclusion and passes the guard (memory obs #30588).
All 13 existing language configs pass the guard; the new negative test
fails without it. Test-first.
Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U5)
* fix(pipeline): wrap sequential-fallback in try/finally so cleanup survives throws
The sequential-fallback block in runChunkedParseAndResolve now runs
inside a try/finally that guarantees astCache.clear(), accumulator
finalize, and enrichExportedTypeMap execute even if readFileContents
or processCalls throws mid-fallback. Cleanup failures are caught
inside the finally so they can't mask the original error.
Accumulator disposal ownership remains with crossFile (U2) — U6 only
adds astCache cleanup and preserves finalize ordering on the error
path.
Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U6)
* test(pipeline): direct unit coverage for wildcard-synthesis and cross-file-impl
Both modules previously had zero direct unit coverage — branches were
exercised only through integration tests' happy paths.
wildcard-synthesis.test.ts covers: Go graph-IMPORTS fallback, Python
moduleAliasMap build, MAX_SYNTHETIC_BINDINGS_PER_FILE cap, dedup
against existing namedImportMap entries, and empty-exportedSymbols
early return.
cross-file-impl.test.ts covers: gapRatio below threshold no-op,
MAX_CROSS_FILE_REPROCESS cap, graph-only exportedTypeMap fallback,
and empty namedImportMap short-circuit.
Tests assert current behavior — any future regression flips them.
Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U7)
* test(pipeline): golden-file graph-parity regression guard on mini-repo fixture
Pins the current post-P1/P2 graph output (57 symbols, 92 relationships,
4 processes, deterministic edge digest) so future silent refactors
cannot drift behavior unnoticed. If any count changes or any edge
rewires, the test fails with a readable diff listing what changed
and a copy-pasteable UPDATE_GOLDEN=1 regen command.
Edge digest keyed by symbolic (label, name, filePath) triples rather
than raw generateId output — stays meaningful across id-encoding
refactors while still catching real semantic rewiring.
Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U8)
* fix(pipeline): minimal cycle reporting + resolveEnclosingOwner loop safeguards
U9: runner cycle detection now reports only the SCC members via DFS
back-edge trace ('Cycle detected: A -> B -> C -> A') rather than
everything with inDegree > 0 (which mixed cycle members with blocked
dependents). Also emits the 'error' progress event for graph-
validation failures, symmetric with U1's runtime-error path.
U16: findEnclosingClassInfo now defends against language-provider
hooks that return non-container nodes — visitedContainers Set breaks
repeat-visit loops, MAX_ENCLOSING_WALK_ITERATIONS is belt-and-braces.
Documented the hook contract invariant so future provider authors
know the walk-continues-upward expectation.
Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U9, U16)
* refactor(pipeline): type hygiene, dead code cleanup, shared allPathSet, graph-sort naming
Bundles plan units U10, U11, U12, U14, U15:
U10 — Type hygiene: readonly ParseOutput arrays (allExtractedRoutes,
allDecoratorRoutes, allToolDefs, allORMQueries, allPaths); removed
redundant 'as string[] | undefined' cast in routes.ts and 'as URL' in
parse-impl.ts; WorkerPool is now 'import type'. Readonly contract
propagated into processORMQueries (only iterates).
U11 — Dead code & shims: deleted constants.ts shim (AST_CACHE_CAP
inlined into its sole real consumer cross-file-impl.ts; isDev
consumers now import directly from ../utils/env.js). Removed internal
utility re-exports from pipeline-phases/index.ts (no external
consumers). Removed topologicalLevelSort re-export from pipeline.ts;
updated topological-sort.test.ts to import from the canonical
utils/graph-sort.js. Stripped 'Phase 3+4:' stale JSDoc from
parse-impl.ts.
U12 — Perf: StructureOutput now carries allPathSet (ReadonlySet<string>)
built once; cobol, markdown, and cross-file-impl consume the shared
set instead of allocating their own. Parse forwards it via
ParseOutput.allPathSet; processCobol/processMarkdown widened to
ReadonlySet<string>.
U14 — graph-sort.ts: renamed local 'inDegree' to
'pendingImportsPerFile' with expanded JSDoc explaining the reverse-
graph Kahn's formulation and warning future maintainers not to
'correct' it to standard in-degree semantics. Added self-edge test.
U15 — Unconditional worker-fallback logging: removed isDev guard on
the worker-pool-creation-failure console.warn so operators can
diagnose perf degradations in production.
No behavior change. U8 golden-file test confirms pipeline output is
byte-identical.
Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U10, U11, U12, U14, U15)
* docs: fix ARCHITECTURE.md table integrity; bump AGENTS.md/CLAUDE.md to 1.3.0
U13 — documentation fixes:
ARCHITECTURE.md: the prior insertion of the 'Pipeline Phase DAG'
section orphaned 7 rows from the 'Where to change what' header.
Moved those 7 rows back up under their header so the table reads
contiguously; DAG section now follows the completed table.
AGENTS.md + CLAUDE.md: bumped version 1.2.0 -> 1.3.0, updated Last
reviewed to 2026-04-13, added matching Changelog row documenting
the GitNexus index stats refresh after the DAG refactor. Stat
bumps (symbols/relationships/execution flows) that were sitting
uncommitted in the working tree are now landed under a proper
changelog entry per each file's own documented schema.
Plan: docs/plans/2026-04-13-001-fix-pipeline-dag-refactor-review-findings-plan.md (U13)
* refactor(pipeline): drop spurious parse deps, true-readonly ParseOutput.exportedTypeMap, skip redundant wildcard synth
- mro/communities/processes: switch redundant `parse` dep to `structure` —
totalFiles originates in structure, so depending on parse for it was a
spurious data dep that obscured the real DAG.
- ParseOutput.exportedTypeMap: typed as truly ReadonlyMap<...,ReadonlyMap>>;
graph→exports enrichment moved into parse-impl so the snapshot is
fully populated at parse return. crossFile builds its own local mutable
working copy for per-file re-resolution writes — no cast at the boundary.
- parse-impl: hasSynthesized flag guards the unconditional final
synthesizeWildcardImportBindings call when per-chunk/fallback synthesis
already ran (graph-global + idempotent across chunks).
- cross-file-impl: documented the intentional `phase: 'parsing'` progress
label so telemetry bucketing stays consistent with the parse phase.
- cross-file-impl test: replaced the now-moved fallback-enrichment
assertion with a stronger one — crossFile must not mutate the
parse-supplied map.
Addresses PR #809 review pass 5 carry-overs.
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* feat(group): extractor expansion + manifest extractor
Part 2 of 4 in the split of #606 (ticket: #792). Follows #795
(bridge.lbug storage foundation, already merged), but this PR has no
code-level dependency on #795 — it only imports types and the
ContractExtractor interface that existed on upstream main before
either PR. It could have been reviewed in parallel with #795.
## What changed
Expands the 3 existing contract extractors with substantially more
language/framework coverage, and adds a new `manifest-extractor`
that resolves `group.yaml`-declared cross-links against the per-repo
graph via exact-name lookups.
### New file (228 LOC)
- `gitnexus/src/core/group/extractors/manifest-extractor.ts` —
exact graph lookup for `group.yaml`-declared cross-links. HTTP
paths are canonicalized before Route.name matching; gRPC is
resolved by service/method name (NO `.proto`-filename fallback);
topic and lib use exact-name match. Falls back to a synthetic
`manifest::<repo>::<contractId>` uid when the graph has no
matching symbol, so cross-impact traversal still has a stable
anchor for the contract.
### Modified extractors (+958 LOC prod)
- `extractors/grpc-extractor.ts` (+522) — `.proto` parser with
comment and string-literal sanitization (braces inside strings no
longer truncate service bodies); package/service/method canonical
IDs; server/client detection across Go (`grpc.NewServer`,
`RegisterXxxServer`, `XxxGrpc.XxxImplBase`), Java (`@GrpcService`,
`BlockingStub`), Python (`servicer_to_server`, `XxxStub`), and
TypeScript/Node (`@GrpcMethod`, `ClientGrpc`, `loadPackageDefinition`).
- `extractors/http-route-extractor.ts` (+174) — Go gin/echo/stdlib
`HandleFunc`, NestJS `@Controller`+`@Get`/etc, Python FastAPI
decorators, Java Spring `@RequestMapping`/`@GetMapping`,
restTemplate / WebClient / OkHttp consumers.
- `extractors/topic-extractor.ts` (+98) — sarama `ProducerMessage{}`
struct literal detection (replaces a constructor-anchored regex
that missed topics inside producer loops), kafka-go Writer/Reader,
Python NATS (`await nc.subscribe`/`await nc.publish`), JetStream
helpers.
### Modified and new tests (+1264 LOC)
- `grpc-extractor.test.ts` (+539) — full coverage of the new proto
parser (strings-with-braces regression, comments-with-braces
regression), per-language server/client detection
- `http-route-extractor.test.ts` (+240) — per-framework route
extraction + normalization edge cases
- `topic-extractor.test.ts` (+177) — the sarama in-loop regression,
JetStream, Python NATS, kafka-go Writer/Reader
- `manifest-extractor.test.ts` (+308 NEW) — HTTP path normalization,
gRPC exact lookup with proto-fallback regression, lib and topic
exact matching, synthetic-uid fallback behavior
### Self-review fixes folded in
Carried forward from the #606 self-review (commit `d15b8cb`):
- **HIGH #1** — `manifest-extractor.resolveSymbol` was too fuzzy.
Previously used `CONTAINS` on route/name fields plus an
unconditional `filePath ENDS WITH '.proto'` fallback for gRPC.
Consequences: `/orders` matched `/suborders`, and any repo with
any `.proto` file returned a random proto symbol for a gRPC
manifest entry. Replaced with exact equality + deterministic
`ORDER BY` + synthetic-uid fallback for unresolved manifests.
Regression tests included.
- **MED #3** — gRPC proto parser brace-depth counting now sanitizes
strings and comments first (`stripProtoCommentsAndStrings`). A
valid proto with `option deprecated_reason = "use NewService {
instead"` used to have its service body closed early by the `"{"`
inside the literal, silently dropping methods after the offending
string. Regression tests for both string-with-brace and
comment-with-brace cases.
- **MED #4** — sarama Kafka regex changed from
`sarama.NewSyncProducer[\s\S]{0,300}?Topic:` (anchored on
constructor, caught only first topic in a loop) to
`sarama.ProducerMessage{...Topic:}` (matches every struct literal
directly). Regression test with a for-loop that constructs
multiple `ProducerMessage`s.
- **MED #7** — `manifest-extractor.resolveSymbol` no longer has a
silent `catch { /* fall through */ }`. Errors from the graph
executor are logged via `console.warn` with link type, contract
name, repo key, and error message before falling through to the
synthetic-uid path.
## Why
Reviewer focus here is pure regex / parser correctness — no
storage, no Cypher queries, no algorithmic changes to the cross-link
algorithm. Separating this from the bridge foundation PR (#795)
meant reviewers could stay in a single mental mode (parsing logic)
instead of context-switching between DDL, Cypher, and regex.
## How to verify
- `cd gitnexus && npx tsc --noEmit`
- `cd gitnexus && npx vitest run test/unit/group/grpc-extractor.test.ts --pool=forks`
- `cd gitnexus && npx vitest run test/unit/group/http-route-extractor.test.ts --pool=forks`
- `cd gitnexus && npx vitest run test/unit/group/topic-extractor.test.ts --pool=forks`
- `cd gitnexus && npx vitest run test/unit/group/manifest-extractor.test.ts --pool=forks`
Local pre-push: typecheck clean, all 99 extractor unit tests pass
(grpc 43, http 18, topic 30, manifest 8).
## Risk / rollback
**Low.** Extractors have no user-facing surface in this PR — they
produce `ExtractedContract[]` that is consumed by `sync.ts` in the
next split (#793). No existing behavior changes for users who don't
run a `group sync`. Rollback = `git revert` of the merge commit;
the modifications to `grpc-extractor.ts` / `http-route-extractor.ts`
/ `topic-extractor.ts` revert to the pre-PR versions that still
work (they're subsets of the new functionality).
## Scope discipline (per GUARDRAILS.md)
- Only the 8 files above are touched; no drive-by refactors
- No CI/release/security config changes
- No secrets or machine-specific paths
- Content lifted from #606 (CI 11/11 green on `d15b8cb`)
## Dependencies
- **Base:** `main` (upstream already includes #795 as `1ff324c`)
- **Blocks:** sync pipeline (#793) and the cross-impact feature (#794)
- **Tracker issue:** #792
- **Parent PR:** #606
Co-authored-by: Claude <noreply@anthropic.com>
* refactor(group): migrate topic-extractor from regex to tree-sitter queries
Addresses @magyargergo's feedback on #796 that regex-based lookups
should use tree-sitter nodes instead, and that the top-level
extractors must NOT carry language dependencies. This is phase 1 of
a multi-step migration — topic-extractor first because its patterns
are the most uniform (16 "call/annotation with first-arg string
literal" variants), which makes it a clean proof of the approach
before grpc-extractor and http-route-extractor get the same treatment.
## Architecture: language-agnostic orchestrator + per-language plugins
The top-level extractor is a thin orchestrator that never imports a
tree-sitter grammar or a query string. Per-language knowledge lives
in a new `topic-patterns/` folder with one file per language plus a
registry that maps file extensions to compiled plugins:
```
src/core/group/extractors/
├── tree-sitter-scanner.ts # shared, language-agnostic scanning utilities
├── topic-extractor.ts # thin orchestrator (no grammar imports)
└── topic-patterns/
├── types.ts # TopicMeta, Broker
├── index.ts # registry: extension → compiled provider
├── java.ts # tree-sitter-java + JAVA_TOPIC_PROVIDER
├── go.ts # tree-sitter-go + GO_TOPIC_PROVIDER
├── python.ts # tree-sitter-python + PYTHON_TOPIC_PROVIDER
└── node.ts # tree-sitter-javascript + tree-sitter-typescript
# → JAVASCRIPT_/TYPESCRIPT_/TSX_TOPIC_PROVIDER
```
**Shared scanner (`tree-sitter-scanner.ts`)** — defines
`PatternSpec<TMeta>`, `LanguagePatterns<TMeta>`, `CompiledPatterns<TMeta>`
and the `scanFile(parser, plugin, content)` helper. Plugins compile their
queries eagerly at module load via `compilePatterns()`, so a broken
pattern fails loudly at import time instead of silently at scan time.
`unquoteLiteral()` handles single/double/template quotes, Python
triple-quoted strings, and Go raw backtick strings.
**Per-language plugins** own:
- the tree-sitter grammar import (this is the ONLY place in
`src/core/group/` where tree-sitter grammars are imported),
- the query S-expressions,
- the `TopicMeta` payload (role, broker, confidence, symbolName) that
the orchestrator receives back on every match.
Each plugin uses a `@value` capture name to bind the topic literal node.
The JavaScript and TypeScript grammars share AST node names for every
construct we query, so `node.ts` defines the pattern sources once and
compiles them against `JavaScript`, `TypeScript.typescript`, and
`TypeScript.tsx` — exporting three providers because `Parser.Query`
objects are NOT portable across grammar instances.
**Registry (`topic-patterns/index.ts`)** — maps `.java` → Java provider,
`.go` → Go, `.py` → Python, `.js`/`.jsx` → JS, `.ts` → TS, `.tsx` → TSX.
Also exports `TOPIC_SCAN_GLOB` so adding a new language is a single
file-level edit (drop `topic-patterns/<lang>.ts`, import + register it
here — zero edits required in `topic-extractor.ts`).
**Orchestrator (`topic-extractor.ts`)** — ~110 lines, no grammar or
query imports. Per file: `getProviderForFile(rel)` → `scanFile(parser,
provider, content)` → `unquoteLiteral(valueText)` → `makeContract(...)`.
Reuses one `Parser` instance across files; the scanner calls
`setLanguage` per plugin.
## Why this is better than regex
1. **Comments and strings are respected for free.** The old regex
would match `// kafkaTemplate.send("fake.topic")` as a real
producer; tree-sitter never visits comments or string literals as
code nodes, so false positives from commented-out code are
eliminated.
2. **Struct/object literal patterns are structural, not textual.**
`sarama.ProducerMessage{Topic: "..."}` no longer needs a 300-char
lookahead (which was a known cross-match bug partly mitigated by a
loop regression test in the self-review). The new query matches a
specific `composite_literal` with a specific `qualified_type` and
`keyed_element` — exactly one struct literal per match.
3. **No order-of-operations fragility.** Regex for
`channel.publish` vs `channel.consume` was independent and
file-wide; the AST scopes matches to the specific `call_expression`.
4. **Language-agnostic extension.** Adding Ruby, Rust, or C# topic
detection later means dropping one file in `topic-patterns/` — no
changes to shared scanner or orchestrator, and no tree-sitter
imports leak into top-level code.
## Per-file fault tolerance
- Malformed files that tree-sitter can't parse are silently skipped
(`parser.parse` is wrapped by `scanFile`). The ingestion pipeline
already logs unparseable files at index time.
- A syntactically invalid query is caught at `compilePatterns` time,
not scan time — broken plugins fail loudly at import.
- Per-pattern `matches()` failures are swallowed so one broken query
in a plugin doesn't block the rest.
## Tests
All 30 existing `topic-extractor.test.ts` tests pass **without any
changes to the test file** — they were written as input/output contract
tests (given this source file, expect these `ExtractedContract` objects)
and that contract is unchanged. Regression coverage includes:
- Kafka: Java `@KafkaListener` + `kafkaTemplate.send`; Node
`producer.send` + `consumer.subscribe`; Go sarama producer/consumer
(sync and async); kafka-go Writer/Reader; Python `KafkaConsumer` +
`producer.send/produce`
- RabbitMQ: Java `@RabbitListener` + `rabbitTemplate.convertAndSend`;
Node `channel.consume/publish/sendToQueue`; Python `basic_consume/
basic_publish` with keyword args
- NATS: Go and Node `nc.Subscribe/Publish`; Go and Node JetStream
`js.Subscribe/Publish`; Python `await nc.subscribe/publish`
Including the regression test for the sarama `ProducerMessage`
in-loop case — the AST-based query captures every literal in the
file independently, not just the first one after `NewSyncProducer`.
## Neighbor regression check
- `topic-extractor.test.ts` — 30/30 pass (rewritten extractor)
- `http-route-extractor.test.ts` — 18/18 pass (untouched)
- `grpc-extractor.test.ts` — 43/43 pass (untouched)
- `manifest-extractor.test.ts` — 8/8 pass (untouched)
- Full `npx tsc --noEmit` clean
## Scope discipline (per GUARDRAILS.md)
- Only files under `src/core/group/extractors/` are touched; no
changes to other extractors, tests, MCP surface, or pipeline.ts.
- No CI/release/security config changes, no secrets.
- New tree-sitter imports all reference grammars that are already
installed as dependencies (`tree-sitter`, `tree-sitter-javascript`,
`tree-sitter-typescript`, `tree-sitter-python`, `tree-sitter-java`,
`tree-sitter-go` — all in `package.json` for the existing pipeline).
## Phase 2 / phase 3 plan
- **Phase 2 (next commit):** rewrite `http-route-extractor.ts`
Strategy B (regex fallback) on the same plugin pattern. Graph-assisted
Strategy A stays as-is (already uses pipeline-built tree-sitter data
via `HANDLES_ROUTE` Cypher queries).
- **Phase 3 (commit after):** rewrite `grpc-extractor.ts` for Java /
Go / Python / TypeScript detection. `.proto` files are the one
outstanding question — there is no `tree-sitter-proto` grammar
installed; the in-tree string-sanitizing parser stays as a pragmatic
exception with a comment, alternative being to add
`tree-sitter-proto` as a dep (open for the maintainer).
Co-authored-by: Claude <noreply@anthropic.com>
* refactor(group): migrate http-route-extractor Strategy B to tree-sitter plugins
Phase 2 of the extractor refactor requested by @magyargergo on #796.
Same architecture as the phase 1 topic-extractor rewrite: a thin,
language-agnostic orchestrator plus per-language plugins that own
tree-sitter grammars and query sources. The top-level extractor file
no longer imports any tree-sitter grammar or query string.
## Architecture
```
src/core/group/extractors/
├── tree-sitter-scanner.ts # shared, language-agnostic primitives
├── http-route-extractor.ts # thin orchestrator (no grammar imports)
└── http-patterns/
├── types.ts # HttpDetection, HttpLanguagePlugin, HttpRole
├── index.ts # registry: ext → plugin + HTTP_SCAN_GLOB
├── java.ts # tree-sitter-java: Spring + RestTemplate/WebClient/OkHttp
├── go.ts # tree-sitter-go: gin/echo/HandleFunc + http/resty consumers
├── python.ts # tree-sitter-python: FastAPI + requests
├── php.ts # tree-sitter-php: Laravel Route::get/...
└── node.ts # tree-sitter-javascript + tree-sitter-typescript:
# NestJS controllers, Express, fetch, axios
```
**Shared scanner (`tree-sitter-scanner.ts`)** — generalised from phase 1:
- `ScanMatch<TMeta>.captures` is now a full `CaptureMap` (every named
capture the query binds, not just a single `@value`). Topic extractor
updated to read `match.captures.value` accordingly.
- New `runCompiledPatterns(plugin, tree)` helper lets plugins run
multiple query bundles against the same pre-parsed tree. This is
needed for HTTP plugins that combine a class-prefix query with a
method-route query (Spring, NestJS).
- `scanFile` becomes a thin wrapper over `parser.parse + runCompiledPatterns`.
**HTTP plugin shape** — unlike topic plugins, HTTP plugins expose a
`scan(tree)` function rather than a flat pattern list. This reflects
HTTP's more complex extraction: each detection needs method + path +
handler name, and framework patterns like Spring `@RequestMapping` /
NestJS `@Controller` require cross-referencing a class-level prefix
with method-level annotations. Plugins internally use
`compilePatterns` + `runCompiledPatterns` and walk the AST to resolve
the class/method relationships.
**Per-framework coverage:**
- **Java (`java.ts`)**
- Spring: `@RequestMapping("/api/v2")` class prefix + `@(Get|Post|Put|
Delete|Patch)Mapping("/sub")` method routes, joined via the
enclosing `class_declaration` node id.
- `RestTemplate.getForObject/postForEntity/put/delete/patchForObject` →
method derived from API name.
- `WebClient.method(HttpMethod.X, "/path")` → method from
`HttpMethod.X` capture.
- `new Request.Builder().url("/path")` → OkHttp consumer.
- **Go (`go.ts`)**
- gin / echo / chi frameworks: `\w+.GET("/path", handler)` captures
upper-case verb + handler identifier.
- `net/http.HandleFunc("/path", handler)` → provider (default GET).
- `http.Get/Post/Head` consumer, `http.NewRequest("METHOD", ...)`,
resty `client.R().Get/Post/...`.
- **Python (`python.ts`)**
- `@app.get("/path")` FastAPI decorators.
- `requests.get/post/...` and `requests.request("METHOD", "url")`.
- **PHP (`php.ts`)**
- Laravel `Route::get/post/.../patch('/path', ...)` via
`scoped_call_expression`. Uses `PHP.php_only` to match the
existing ingestion pipeline's grammar selection.
- **Node (`node.ts`) — JS + TS + TSX**
- Pattern sources defined once, compiled against three grammar
variants (`JavaScript`, `TypeScript.typescript`, `TypeScript.tsx`)
because `Parser.Query` objects are not portable across grammars.
Exports three plugins sharing the same `scan` logic.
- NestJS: `@Controller('prefix')` decorators are siblings of the
class in `export_statement` / `program`; `@Get(':id')` decorators
are siblings of the method in `class_body`. The plugin walks
decorator → next named sibling to find the decorated class /
method, then combines the class prefix with the method path.
Only emits NestJS detections when the enclosing class has a real
`@Controller` decorator — prevents false positives from generic
classes that happen to use `@Get` from another library.
- Express: `(router|app).<verb>('/path', ...)`.
- `fetch(url)` (default GET) + `fetch(url, { method: 'X' })`
(uses two queries + a SyntaxNode-id dedupe set so URL literals
aren't double-emitted by the options variant).
- `axios.get/post/...`.
## Orchestrator changes
`http-route-extractor.ts` drops every `scanXxxProviders` / `scanXxxConsumers`
regex method and replaces them with a single source-scan loop that
delegates to `getPluginForFile(rel).scan(tree)`. The orchestrator
still owns:
- **Path normalization** (`normalizeHttpPath`, `normalizeConsumerPath`)
— language-agnostic string processing shared by both strategies.
- **Graph-assisted Strategy A** (`HANDLES_ROUTE` / `FETCHES` / `CONTAINS`
Cypher queries) — unchanged in spirit. The only regex helpers it
used (`inferMethodFromFileScan`, `pickJavaHandlerName`) are now
replaced by a lookup against the plugin's detections for the same
file: for each route row, find the detection whose normalized path
matches, and pull the HTTP method + handler name from it.
- **Per-file parse cache** — the orchestrator parses each relevant
file at most once per `extract()` call. Both the graph-assisted
enrichment loop and the source-scan fallback share the same
`cachedDetections` map, so we never run the plugin twice for the
same file.
## Why this is better than the regex version
1. **Comments and strings for free.** The old regex would match
`// router.get('/fake')` as a real Express route; tree-sitter
never visits string/comment nodes.
2. **Structural controller-prefix.** Spring and NestJS class-prefix
joining is now scoped to the enclosing class via `class_declaration`
node ids, eliminating file-wide state that broke when a file had
multiple controllers.
3. **Precise NestJS disambiguation.** The plugin only emits a NestJS
detection when the enclosing class has a real `@Controller`
decorator — the old regex would fire on any `@Get(...)` in the
file regardless of surrounding context.
4. **Language-agnostic extension.** Adding Ruby / Rust / Kotlin HTTP
detection later means dropping one file in `http-patterns/` — no
changes to the shared scanner, the orchestrator, or the Strategy A
Cypher queries.
## Tests
- `http-route-extractor.test.ts` — **18/18 pass** (tests unchanged;
they're contract-style input/output tests and the contract shape is
unchanged). Covers Spring class prefix, Express, gin/echo, stdlib
HandleFunc, NestJS, Laravel, FastAPI for providers and
fetch/axios/python-requests/rest-template/webClient/okhttp/go-stdlib/
resty for consumers, plus graph-first Strategy A for both.
- `topic-extractor.test.ts` — **30/30 pass** after the `captures.value`
API migration.
- `grpc-extractor.test.ts` — 43/43 pass (untouched; phase 3).
- `manifest-extractor.test.ts` — 8/8 pass (untouched).
- `service.test.ts`, `sync.test.ts`, `storage.test.ts` — 41/41 pass.
- `npx tsc -p tsconfig.json --noEmit` clean.
## Scope discipline (per GUARDRAILS.md)
- Only files under `src/core/group/extractors/` are touched.
- No changes to pipeline.ts, MCP surface, ingestion, or tests.
- No CI / release / security / secrets changes.
- Tree-sitter grammars imported by plugins (`tree-sitter-java`,
`tree-sitter-go`, `tree-sitter-python`, `tree-sitter-php`,
`tree-sitter-javascript`, `tree-sitter-typescript`) are all already
in `package.json` for the existing ingestion pipeline.
## Phase 3 plan
- **grpc-extractor** gets the same treatment: plugin-per-language under
`grpc-patterns/` for Java / Go / Python / TS detection. `.proto`
files remain an open question — no `tree-sitter-proto` grammar is
installed, so the in-tree string-sanitizing parser from PR #796's
self-review stays as a pragmatic exception unless the maintainer
wants us to add `tree-sitter-proto` as a new dep.
Co-authored-by: Claude <noreply@anthropic.com>
* refactor(group): migrate grpc-extractor source scans to tree-sitter plugins
Phase 3 (final) of the extractor refactor requested by @magyargergo on
#796. Same architecture as phase 1 (topic) and phase 2 (http): thin
language-agnostic orchestrator + per-language plugins that own
tree-sitter grammars and query sources. With this commit the top-level
extractors under `src/core/group/extractors/` import ZERO tree-sitter
grammars and ZERO query strings — every grammar import lives in a
`*-patterns/<lang>.ts` plugin file, and the orchestrators go through
the registry indirection.
## Architecture
```
src/core/group/extractors/
├── tree-sitter-scanner.ts # shared primitives (unchanged)
├── grpc-extractor.ts # orchestrator (only `.proto` parser left)
└── grpc-patterns/
├── types.ts # GrpcDetection, GrpcLanguagePlugin, GrpcRole
├── index.ts # registry: ext → plugin + GRPC_SCAN_GLOB
├── go.ts # tree-sitter-go: RegisterXxxServer, Unimplemented, NewXxxClient
├── java.ts # tree-sitter-java: @GrpcService + XxxImplBase + newBlockingStub
├── python.ts # tree-sitter-python: add_XxxServicer_to_server + XxxStub
└── node.ts # tree-sitter-javascript + tree-sitter-typescript:
# @GrpcMethod, @GrpcClient field type,
# .getService<X>('Svc'), new XxxServiceClient,
# loadPackageDefinition dynamic constructors
```
## Per-language coverage
**Go (`go.ts`)**
- Provider: `\w+.RegisterXxxServer(...)` via `call_expression →
selector_expression → field_identifier` + JS regex filter
`^Register(\w+)Server$`.
- Provider: `pb.UnimplementedXxxServer` embedded in a struct via
`struct_type → field_declaration_list → field_declaration →
qualified_type → type_identifier` + JS filter.
- Consumer: `\w+.NewXxxClient(...)` via the same call_expression
query + JS filter `^New(\w+)Client$`.
**Java (`java.ts`)**
- Provider: `class X extends YyyGrpc.YyyImplBase` — two queries
handle the scoped and plain forms. `scoped_type_identifier`'s
children are positional (no `scope:`/`name:` fields), so the
query matches the two `type_identifier` children by position.
- `#match? @inner "ImplBase$"` restricts matches at query time.
- Whether the class has `@GrpcService` or not controls only the
`source` metadata label — the plugin walks the class_declaration's
`modifiers` child in JS to detect the marker_annotation.
- Consumer: `YyyGrpc.newStub(ch)` / `newBlockingStub(ch)` via a
`method_invocation` query with `#match? @method
"^new(Blocking)?Stub$"`, service name extracted via
`^(\w+)Grpc$` on the object identifier.
**Python (`python.ts`)**
- Single call-expression query covers both bare identifier and
`obj.method` attribute forms:
`(call function: [(identifier) @fn (attribute attribute: (identifier) @fn)])`.
- Plugin filters `@fn.text` against two JS regexes:
`^add_(\w+)Servicer_to_server$` (provider) and `^(\w+)Stub$`
(consumer), with a reserved-names ignore list for the Stub case
(Mock / Test / Fake / Stub).
**Node — JavaScript + TypeScript + TSX (`node.ts`)**
- Pattern sources defined once, compiled three times (one per grammar)
because `Parser.Query` objects are not portable across grammars.
Exports three `GrpcLanguagePlugin`s sharing the same `scan`.
- `@GrpcMethod('Service', 'Method')`: decorator query captures the
two string literals. Confidence is hard-coded 0.8 regardless of
proto map resolution (matches the original regex version's
behaviour).
- `@GrpcClient(...) field: XxxServiceClient`: decorator query
captures the decorator node, plugin walks up to find the enclosing
`public_field_definition` (decorators on fields are CHILDREN of
the field definition in tree-sitter-typescript, not siblings) and
reads its first `type_annotation → type_identifier`, then runs the
`^(\w+Service)Client$` JS filter.
- `client.getService<X>('AuthService')`: call-expression query on
`member_expression.property = "getService"` + string literal arg.
- `new XxxServiceClient(...)`: `new_expression` with a bare
identifier constructor, filtered by `^(\w+Service)Client$` so
generic `new AuthClient(...)` (missing the `Service` infix) does
NOT falsely register as a consumer. Preserves the regression test
`test_extract_ts_non_service_client_constructor_is_ignored`.
- `loadPackageDefinition` dynamic loader: gated on
`tree.rootNode.text.includes('loadPackageDefinition')`. When set,
`new foo.bar.Xxx(...)` qualified constructors with a capitalised
property name register as consumers.
## Orchestrator changes
`grpc-extractor.ts` loses every `scanGoProviders` / `scanJavaProviders`
/ ... helper and replaces them with a single source-scan loop that:
1. Parses each file with the plugin's grammar (one shared `Parser`
instance across all files, `setLanguage` called per plugin).
2. Calls `plugin.scan(tree)` to get `GrpcDetection[]`.
3. Converts each detection to an `ExtractedContract` via the private
`detectionToContract` helper, which:
- Looks the short service name up in the proto map (filled by
the `.proto` parser).
- Picks confidence = `confidenceWithProto` if resolved, else
`confidenceWithoutProto`.
- Builds a method-level contract id (`grpc::pkg.Svc/Method`) when
the detection carries a `methodName` (TS `@GrpcMethod` only),
otherwise a service-level id (`grpc::pkg.Svc/*`).
Everything else — the `.proto` parser, `buildProtoContext`,
`buildProtoMap`, `resolveProtoConflict`, `serviceContractId`,
`stripProtoCommentsAndStrings`, `extractServiceBlocks`, the dedupe
function — stays exactly as before. The `.proto` parser is kept as a
pragmatic exception to the "no regex in extractors" rule because no
`tree-sitter-proto` grammar is installed in the repo; a comment at the
top of the file explains this and flags the maintainer option of
adding `tree-sitter-proto` as a dependency.
## Why this is better than the regex version
1. **Comments and strings are respected for free.** Matched node types
are only code constructs, never text inside comments or string
literals.
2. **No false positives on partial names.** The old `(\w+?)Grpc`-style
regexes would cross-match unrelated identifiers; structural queries
restrict matches to the exact AST shape (`scoped_type_identifier →
type_identifier` pairs, `method_invocation → identifier` etc.).
3. **NestJS `@GrpcClient` is structural, not regex-based.** The old
regex required a specific textual layout
(`@GrpcClient(...) private readonly foo!: XxxServiceClient`); the
plugin now walks the AST, so modifier order / optional modifiers /
multi-line formatting don't break it.
4. **Language-agnostic extension.** Adding Kotlin / Rust / C# gRPC
detection later is a one-file edit in `grpc-patterns/index.ts` —
no touches to the shared scanner, the orchestrator, or the proto
parser.
## Tests
- `grpc-extractor.test.ts` — **43/43 pass** (tests unchanged; the
contract shape is identical). Covers .proto parsing (including the
brace-inside-string regression), Go provider/consumer,
Java @GrpcService / plain ImplBase provider + newBlockingStub
consumer, Python servicer + stub, TS @GrpcMethod + @GrpcClient +
.getService + new XxxServiceClient + loadPackageDefinition + the
`AuthClient` vs `AuthServiceClient` discrimination, dedupe across
multiple patterns in one file, proto-aware confidence, and the
inherited-package resolution for split proto definitions.
- `topic-extractor.test.ts` — 30/30 pass.
- `http-route-extractor.test.ts` — 18/18 pass.
- `manifest-extractor.test.ts` — 8/8 pass.
- `service.test.ts`, `sync.test.ts`, `storage.test.ts` — 41/41 pass.
- `npx tsc -p tsconfig.json --noEmit` clean.
## Scope discipline (per GUARDRAILS.md)
- Only files under `src/core/group/extractors/` are touched.
- No pipeline.ts, MCP surface, ingestion, CI / release / security, or
test changes.
- New tree-sitter grammar imports (`tree-sitter-go`, `tree-sitter-java`,
`tree-sitter-python`, `tree-sitter-javascript`, `tree-sitter-typescript`)
are all already installed for the ingestion pipeline.
## End of phase series
This commit completes the three-phase extractor refactor:
- **Phase 1** (`ea06d11`): topic-extractor → `topic-patterns/`
- **Phase 2** (`b6015f6`): http-route-extractor → `http-patterns/`
- **Phase 3** (this commit): grpc-extractor → `grpc-patterns/`
Every remaining regex-based extractor helper under the `src/core/group/
extractors/` directory is either (a) language-agnostic string
processing (path normalization, dedupe keys) or (b) the `.proto`
parser, which is documented as an explicit exception.
Co-authored-by: Claude <noreply@anthropic.com>
* feat(group): add tree-sitter-proto for .proto file parsing
Addresses @magyargergo's suggestion on #796 to replace the manual
string-sanitizing .proto parser with a tree-sitter grammar.
- **Vendored `tree-sitter-proto`** in `vendor/tree-sitter-proto/`.
Grammar source from [coder3101/tree-sitter-proto](https://github.com/coder3101/tree-sitter-proto)
(latest `grammar.js`), parser.c regenerated with `tree-sitter-cli
0.24` to produce ABI version 14 — compatible with the project's
`tree-sitter 0.25` runtime (which supports ABI ≤ 14). Added as
`optionalDependency` with `file:./vendor/tree-sitter-proto`.
- **New `grpc-patterns/proto.ts` plugin** — uses the same
`compilePatterns` + `runCompiledPatterns` infrastructure as every
other plugin. Two queries:
- `(package (full_ident) @pkg)` — package declaration
- `(service (service_name) @service_name (rpc (rpc_name) @rpc_name))`
— one match per (service, rpc) pair
- **Graceful fallback** — `tree-sitter-proto` is an optional
dependency. If it fails to install (platform incompatibility) or
fails the runtime smoke-test (`setLanguage` + `parse` on a trivial
proto), `PROTO_GRPC_PLUGIN` stays `null` and the orchestrator
uses the existing manual parser. The smoke-test catches the
`SyntaxNode` TDZ error that occurs in vitest's fork-based test
runner.
- **Orchestrator updated** — when `hasProtoPlugin` is true, `.proto`
files are handled by the plugin loop (they're included in
`GRPC_SCAN_GLOB`), and the manual `parseProtoFile` loop is
skipped. `buildProtoContext` still runs to build the proto map
for cross-referencing source-file detections.
1. **No manual comment/string stripping.** The old parser needed
`stripProtoCommentsAndStrings` (110 lines) to avoid counting
braces inside comments and string literals. tree-sitter handles
this natively.
2. **No brace-depth tracking.** `extractServiceBlocks` used a manual
depth counter to find service boundaries. tree-sitter's AST gives
us `service` → `service_name` + `rpc` → `rpc_name` directly.
3. **Performance.** tree-sitter's C-based parser is faster than
character-by-character JS scanning + regex on large proto files.
- `grpc-extractor.test.ts` — **43/43 pass** (unchanged)
- All other extractor tests — 99/99 pass
- `npx tsc -p tsconfig.json --noEmit` clean
Co-authored-by: Claude <noreply@anthropic.com>
* chore: add .gitignore for vendored tree-sitter-proto build artifacts
https://claude.ai/code/session_01SFUCxgKMMQ8EgRHYw91xPU
* fix: correct .gitignore paths for vendored tree-sitter-proto
Patterns should be relative to the .gitignore file's directory.
https://claude.ai/code/session_01SFUCxgKMMQ8EgRHYw91xPU
* refactor(group): address Copilot review feedback on #796
Six fixes suggested by the Copilot AI review:
1. **`normalizeHttpPath` root-path edge case** — stripping trailing
slashes on the input `/` produced an empty string, yielding
malformed contract ids like `http::GET::`. Now preserves `/` for
the root handler/fetch case.
2. **Dedupe `scanFiles` call** — `extract()` was globbing the
source-scan file list twice (once for the provider fallback, once
for the consumer fallback). Moved to a single lazy call that
memoizes the result for the rest of the method.
3. **HTTP `scanFiles` now ignores `**/vendor/**`** — every other
extractor's glob already ignored vendored sources; the HTTP one
didn't. Fixed for consistency.
4. **`loadPackageDefinition` check is now structural** — was calling
`tree.rootNode.text.includes('loadPackageDefinition')` which forces
materialization of the entire file text from the parse tree
(expensive on large files). Replaced with a dedicated compiled
query on `(call_expression function: [(identifier) | (member_expression)])`
so the check stays in the AST domain.
5. **`grpc-extractor.ts` header docstring updated** — still claimed
".proto parsing is not tree-sitter-based because no grammar is
installed". Now describes the actual behaviour: tree-sitter when
`tree-sitter-proto` is available (optionalDependency), manual
fallback otherwise.
6. **Eliminated the double proto file parse on the fallback path** —
`buildProtoContext` already globs + parses every `.proto` file to
build `servicesByName`. On the `!hasProtoPlugin` branch the
extractor was globbing + parsing again via the now-removed
`parseProtoFile` helper. The fallback branch now iterates the map
that `buildProtoContext` already produced to emit provider
contracts directly — single pass per proto file.
## Tests
- `topic-extractor.test.ts` — 30/30 pass
- `http-route-extractor.test.ts` — 18/18 pass
- `grpc-extractor.test.ts` — 43/43 pass
- `manifest-extractor.test.ts` — 8/8 pass
- `npx tsc -p tsconfig.json --noEmit` clean
Co-authored-by: Claude <noreply@anthropic.com>
* refactor(group): address Claude review feedback (bugs + dedup + hygiene) on #796
Follows up `2f28bfc` with the remaining items from the Claude AI review:
## Bugs
**Bug 2 — Label-unaware Cypher queries in `resolveSymbol`.**
The manifest-extractor's lookup queries were `MATCH (n) WHERE n.name = $x`
with no label filter, so a topic/service/package name could silently match
any node type (File, Variable, Import, Folder, …). Added label filters:
- `topic` → `(n:Function|Method|Class|Interface)` (topics are best-effort
symbol-name matches against listener/publisher symbols)
- `grpc` method → `(n:Function|Method)`
- `grpc` service → `(n:Class|Interface)`
- `lib` → `(n:Package|Module)`
All 8 manifest-extractor tests still pass (mock executor is
label-agnostic, but the production LadybugDB graph now gets correctly
scoped queries).
**Bug 8 — Tautological `!handlerName` condition.**
`http-route-extractor.ts:extractProvidersGraph` had
`let handlerName = null; if (!method || !handlerName) { ... }` — the
`!handlerName` clause was always true since there was no intervening
assignment. Simplified to always run the plugin-scan lookup (we need
the handler name even when `methodFromRouteReason` already resolved
the method).
## Clean code / dedup
**Design 7 — `readSafe` was copy-pasted in all three orchestrators.**
Extracted to `extractors/fs-utils.ts` as the single source of truth
for the path-traversal guard. Dropped the three local copies and the
now-unused `fs`/`path` imports from topic-extractor.
**Style 10 — Language-specific `_test.go` skip in the topic orchestrator.**
Was `if (rel.endsWith('_test.go')) continue;` inside the language-
agnostic extraction loop. Pushed into the glob's ignore list
(`'**/*_test.go'`) alongside the existing `node_modules`, `vendor`,
`dist`, `build` entries, with a comment explaining that other
languages' test file conventions either live in separate directories
(Python `tests/`, Java `src/test/`) or are already covered by the
existing ignores.
## Already addressed in `2f28bfc` (mentioned again in Claude review)
- Bug 3: `normalizeHttpPath('/')` returns `''` — fixed
- Bug 4: double glob + double parse of `.proto` — fixed
- Bug 5: `scanFiles` called twice in HTTP — fixed
- Bug 6: missing `**/vendor/**` in HTTP glob — fixed
- Design 9 partially: `tree.rootNode.text.includes('loadPackageDefinition')`
replaced with a dedicated structural query
## Deferred
- Bug 1 (`http::*::path` vs `http::GET::path` matching) — out of scope;
sync.ts matching logic lands in #793, manifest extractor already
emits correct synthetic uids for unresolved HTTP contracts.
- Design 9 full (change plugin `scan(tree)` → `scan(tree, source)`) —
the only real use case (`loadPackageDefinition` gate) is already
fixed via a structural query, so the interface change would be
cosmetic churn without a concrete consumer.
## Tests
- `topic-extractor.test.ts` — 30/30 pass
- `http-route-extractor.test.ts` — 18/18 pass
- `grpc-extractor.test.ts` — 43/43 pass
- `manifest-extractor.test.ts` — 8/8 pass
- `npx tsc -p tsconfig.json --noEmit` clean
Co-authored-by: Claude <noreply@anthropic.com>
* docs+fix(group): address remaining Claude review items + add pipeline flow chart
## Fixes
**Remaining 🔴 — HTTP contract id wildcard format.** Documented the
`http::*::<path>` format as an intentional wildcard for manifest links
that omit the HTTP method, alongside the explicit-method form
(`GET::/path` → `http::GET::/path`). The docblock on `buildContractId`
now states both forms, notes that wildcard-aware matching is the
responsibility of the sync / cross-impact layer (#793), and
recommends the explicit-method form whenever the author knows the
method (it round-trips through exact equality without needing
wildcard logic downstream). Tests unchanged — the wildcard format is
what they've always asserted.
**Minor 1 — stale comment at `manifest-extractor.ts:124-126`.** The
comment claimed "creates a contract with an empty symbolUid/ref" but
the code switched to `manifestSymbolUid(repo, contractId)` a few
commits back. Updated to describe the actual synthetic-uid fallback
semantics and the cross-impact path that relies on both sides of the
join deriving the same uid.
**Minor 2 — exhaustiveness guard on `buildContractId`.** The
`switch(type)` covered all five current `ContractType` variants but
silently returned `undefined` if a new variant was added. Added a
`default: const _exhaustive: never = type; throw new Error(...)`
clause so the build fails loudly on an unhandled variant.
**Minor 3 — `tree.rootNode.text` in `grpc-patterns/node.ts`.** Already
fixed in `2f28bfc` via a dedicated structural query
(`LOAD_PACKAGE_DEFINITION_SPEC`). No action needed.
## New: pipeline flow chart (per @magyargergo's request)
Added `src/core/group/PIPELINE.md` with four Mermaid diagrams:
1. **High-level overview** — `group.yaml` → extractors + manifest →
contract matching → `bridge.lbug` → `runGroupImpact`.
2. **Per-repo extractor two-strategy shape** — graph-assisted
Strategy A vs. source-scan Strategy B.
3. **Plugin architecture** — orchestrator → registry →
per-language `*-patterns/<lang>.ts` → `tree-sitter-scanner.ts` →
`ExtractedContract`.
4. **Manifest extraction** — label-scoped `resolveSymbol` with the
synthetic-uid fallback.
5. **Cross-impact query (#606)** — local impact → bridge join →
cross-repo fan-out.
Each diagram is annotated with which PRs own which stage (this PR:
extractors + manifest; #795: bridge storage; #606: cross-impact
runtime) and points at the concrete files/functions involved.
## Tests
- 99/99 extractor tests pass
- `npx tsc -p tsconfig.json --noEmit` clean
Co-authored-by: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* Initial plan
* feat(SM-20): extract registries into model/ module with SemanticModel interface
- Create model/type-registry.ts — TypeRegistry interface + factory
- Create model/method-registry.ts — MethodRegistry interface + factory
- Create model/field-registry.ts — FieldRegistry interface + factory
- Create model/semantic-model.ts — SemanticModel interface + factory
- Create model/heritage-map.ts — re-export HeritageMap types
- Create model/binding-accumulator.ts — re-export BindingAccumulator types
- Create model/resolve.ts — move lookupMethodByOwnerWithMRO from call-processor
- Update symbol-table.ts — delegate to SemanticModel for registry ops
- Update call-processor.ts — re-export lookupMethodByOwnerWithMRO from model/resolve
No circular dependencies: model/resolve.ts does NOT import resolution-context.ts.
All 775 related unit tests pass with no regressions.
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/27ad2975-1a31-4f50-815b-178ee8a95277
* fix: clarify re-export comment per code review feedback
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/27ad2975-1a31-4f50-815b-178ee8a95277
* refactor(SM-20): wire up SemanticModel as first-class resolution input
PR #786 extracted TypeRegistry/MethodRegistry/FieldRegistry into model/
behind SemanticModel, but consumers still routed through SymbolTable
delegates. This change completes Phase 6 of the fuzzy-lookup elimination
roadmap by making call-processor, resolution-context, type-env, and
heritage-map query the model directly via `table.model.{types,methods,fields}`.
Also absorbs the open PR #786 review findings so the branch lands clean:
- Removed duplicate JSDoc block on lookupMethodByOwner (symbol-table.ts)
- Added model/index.ts barrel for the public model/ surface
- Fixed O(n) buildParentMapFromHeritage BFS via head-pointer queue
- Clarified re-export facade framing on binding-accumulator.ts and
heritage-map.ts inside model/
- Refined @internal JSDoc on lookupMethodByOwnerWithMRO
Changes:
- symbol-table.ts: expose `readonly model: SemanticModel` on the
SymbolTable interface. SymbolTable delegate wrappers (lookupClassByName
etc.) stay as thin pass-throughs for backward compat; deletion is a
follow-up once all internal callers are migrated.
- model/resolve.ts: lookupMethodByOwnerWithMRO now takes SemanticModel
instead of SymbolTable, removing the last SymbolTable import from the
model/ module. Preserves circular-dependency firewall.
- call-processor.ts: 6 call sites in D0 member resolution, field
resolution, ctor override, and ctor disambiguation migrated to
model.types/methods/fields.
- resolution-context.ts: tier 3 class+impl lookup migrated.
- type-env.ts: 5 sites across lookupClassDefsByName, resolveFieldType,
and resolveMethodReturnType migrated.
- heritage-map.ts: parent/child class-name resolution migrated.
Tests:
- symbol-table.test.ts: +10 parity and feeding-audit tests covering
every model.{types,methods,fields} path (Class, Method, Property,
Impl, Function-with-ownerId, Property-without-ownerId skip, arity
filtering, clear cascade).
- call-processor.test.ts: classLookupSpy now targets
ctx.symbols.model.types since the wrapper is bypassed.
- type-env.test.ts: createMockSymbolTable and the destructured-call
makeSymbolTable helpers gained a model shim that forwards to the
(possibly overridden) top-level lookup stubs.
Validation: full suite 5603 passed / 159 skipped, resolver integration
suite (19 files, 1766 tests) clean, tsc --noEmit clean.
* refactor(SM-21): invert ownership — SemanticModel contains SymbolTable
Follow-up to SM-20. Previously SymbolTable owned a `model` subfield;
this commit turns the ownership direction around so the SemanticModel
is the top-level container and SymbolTable is nested as `.symbols`:
SemanticModel (top-level, passed everywhere)
├── types (TypeRegistry)
├── methods (MethodRegistry)
├── fields (FieldRegistry)
└── symbols (SymbolTable — file-indexed + callable-name index)
The owner-scoped registries live directly on the model; file and
callable-name lookups go through `.symbols`. Consumers receive a
`SemanticModel` and reach into the appropriate field — no more
`table.model.types.X` double-hop.
Core changes:
- symbol-table.ts: createSymbolTable now takes injected
TypeRegistry/MethodRegistry/FieldRegistry via a SymbolTableDeps
argument. When omitted (test fallback), it creates standalone
registries locally and clears them in clear() — production callers
always inject. The five registry convenience delegates
(lookupClassByName, lookupMethodByOwner, lookupFieldByOwner,
lookupClassByQualifiedName, lookupImplByName) remain as thin
forwards to the injected registries so standalone SymbolTable use
(chiefly tests) stays ergonomic.
- model/semantic-model.ts: createSemanticModel() now creates the
three registries AND a SymbolTable wired to them, exposing the
SymbolTable as `.symbols`. clear() cascades through all four.
- resolution-context.ts: `readonly symbols: SymbolTable` field is
replaced with `readonly model: SemanticModel`. Internal factory
builds a SemanticModel and keeps a local `symbols` alias for
backward-compatible inner body.
Consumer migrations (src/):
- call-processor.ts: ctx.symbols.add/.lookupExactAll/
.lookupCallableByName → ctx.model.symbols.*; ctx.symbols.model.X →
ctx.model.X. buildTypeEnv option key renamed symbolTable → model.
- type-env.ts: symbolTable parameter renamed model (type
SemanticModel), all internal call sites rewritten to use
model.types.*, model.methods.*, model.fields.*,
model.symbols.lookupExactAll / .lookupCallableByName.
- heritage-map.ts: 2 class-lookup sites migrated.
- pipeline.ts: ctx.symbols → ctx.model.symbols throughout.
Test migrations:
- symbol-table.test.ts: parity tests (which validated the old
table.model.X hop) replaced with direct SemanticModel coverage via
createSemanticModel(). New tests exercise types/methods/fields/
symbols feeding end-to-end.
- type-env.test.ts: createMockSymbolTable rebuilt as a
SemanticModel-shaped mock that still accepts the legacy flat
override bag for backward compat; inline `makeSymbolTable` helpers
for destructured-call and importedReturnTypes suites rewritten to
match the new shape; buildTypeEnv options `symbolTable: X` and
`{ symbolTable }` shorthand renamed to `model:`; one real
createSymbolTable-based test rewritten to use createSemanticModel.
- call-processor.test.ts, heritage-map.test.ts,
heritage-processor.test.ts, symbol-resolver.test.ts: bulk sed
`ctx.symbols.` → `ctx.model.symbols.`. call-processor.test.ts spy
updated to target `ctx.model.types.lookupClassByName`.
Validation: full test suite 5589 passed / 169 skipped / 0 failed;
tsc --noEmit clean; pre-commit eslint + prettier + typecheck all
green. CLAUDE.md / AGENTS.md stats bumped from an earlier `npx
gitnexus analyze` refresh (3965 symbols / 10012 edges / 243 flows).
* refactor(SM-22/SM-23): dispatch table + DAG rearchitecture
SM-22: Extract registration dispatch table into model/registration-table.ts.
Replaces the if/else ladder inside SymbolTable.add() with an O(1)
Map<NodeLabel, RoutingDecision> fan-out. SemanticModel wires the table
per-instance so hooks close over the correct registries.
SM-23: DAG rearchitecture. symbol-table.ts is now a pure 2-index leaf
(fileIndex + callableByName) with zero imports from model/. All
type/method/field routing lives in the model/ layer. Tests migrated to
createSemanticModel() + model.symbols access pattern.
Tests: 5632 passed, 0 failures.
* refactor: delete dead code (skipCallableIndex + model/ facades)
Removes the unused skipCallableIndex flag from the registration dispatch
table and deletes two facade files that had zero consumers.
skipCallableIndex was declared on RoutingDecision and populated for all
10 entries but never read at runtime — semantic-model.ts explicitly
documented that the flag was NOT consulted. The callable-index gate
lives inside SymbolTable.add() via CALLABLE_TYPES.has(type), which is
the single source of truth. Deleting the flag keeps SymbolTable as the
sole decision point and removes documentation-as-data.
model/binding-accumulator.ts and model/heritage-map.ts were facade
pass-throughs of their parent-directory counterparts. Grep confirms no
consumer imports either from the model/ path — all usage goes through
../binding-accumulator.js and ../heritage-map.js directly. model/index.ts
was the only "user" and re-exported them with a note about unifying the
import boundary, but that boundary has no actual consumers today.
Resolves review findings M-01 and M-03 from
.context/compound-engineering/ce-review/20260411-144641-59605d93/maintainability.json
Tests: 5631 passed, 0 failures (1 less than pre-Unit-1: the
skipCallableIndex-specific assertion was removed).
* refactor: remove lookupMethodByOwnerWithMRO backward-compat shim
call-processor.ts re-exported lookupMethodByOwnerWithMRO from
./model/resolve.js as a backward-compat shim for symbol-table.test.ts.
The function already lives in model/resolve.ts and is re-exported
properly from model/index.ts (the barrel) — the call-processor shim
was a duplicate export path with no durable reason to exist.
Migrated the test import from call-processor.js to model/index.js
(the canonical barrel). Deleted the re-export statement and the stale
"re-exported for backward compatibility" comment block. Hoisted the
remaining import to the top of the file with the other imports; the
bottom-of-file position was a relic of the shim pattern.
Resolves review finding M-02 from
.context/compound-engineering/ce-review/20260411-144641-59605d93/maintainability.json
Tests: 5631 passed, 0 failures.
* refactor: harden registration dispatch runtime safety
Two hardening changes in semantic-model.ts, both closing silent-failure
paths in the SM-series dispatcher-bypass failure mode.
1. model.symbols.clear() now cascades to the owner-scoped registries.
Previously, the SymbolTable facade exposed rawSymbols.clear directly,
which only emptied fileIndex + callableByName — the types/methods/
fields registries stayed populated. Any caller holding a SymbolTable
reference that invoked .clear() left the model in a split state where
subsequent .add() calls double-registered in the registries. No
current caller exercises this path, but it was a latent phantom-
resolution risk that didn't belong in a public API. Extracted the
cascade into a single cascadeClear closure wired into both
model.clear() and the facade's clear field.
2. runExhaustivenessGuard now throws instead of console.warn on drift.
The production short-circuit via NODE_ENV === 'production' is
preserved, so real users never see the throw — but CI and dev runs
now fail loudly if a NodeLabel is added to gitnexus-shared without
being placed in one of the three registration-table allowlists. The
previous warn-only behavior was silent in test output volume; SM-19
already documented dispatcher-bypass as the dominant silent-failure
mode in this codebase.
Test-first: added test/unit/model/semantic-model.test.ts covering
model.symbols.clear() cascade (4 registries × clear = 4 tests), the
existing model.clear() cascade (regression guard), and a happy-path
construction test that verifies the current allowlists have zero drift.
Resolves correctness P2 finding (symbols.clear() partial clear),
correctness P3 (exhaustiveness warn-only), and kieran-typescript KT-03
(same exhaustiveness finding, agreement boost).
Tests: 5638 passed (+7 new), 0 failures.
* docs: fix stale JSDoc references in resolveStaticCall
call-processor.ts:2215-2216 referenced SymbolTable.lookupClassByName
and SymbolTable.lookupMethodByOwner via {@link}. Both methods were
removed from SymbolTable during SM-20 — they now live on TypeRegistry
and MethodRegistry respectively, accessible via model.types and
model.methods.
Other SymbolTable.* references in the codebase (lookupExactFull, add,
lookupCallableByName in call-processor.ts:593, symbol-table.ts:86,
type-extractors/types.ts:57) target methods that are still on
SymbolTable and remain valid.
Resolves correctness P3 and kieran-typescript KT-02 (same finding,
agreement boost).
* refactor: deduplicate ALL_NODE_LABELS constant
ALL_NODE_LABELS was private in semantic-model.ts and duplicated
verbatim in registration-table.test.ts. Two hardcoded lists meant a
new NodeLabel added to gitnexus-shared could land in one copy but not
the other, silently drifting the exhaustiveness invariant.
Exported ALL_NODE_LABELS from semantic-model.ts, re-exported through
model/index.ts for barrel consistency, and switched the test to import
it instead of redeclaring. The explanatory comment now describes the
single-source-of-truth contract.
Resolves maintainability M-04.
Tests: 5638 passed, 0 failures.
* refactor: add compile-time NodeLabel exhaustiveness check
The runtime exhaustiveness guard in semantic-model.ts caught drift at
test time. Added a type-level check in registration-table.ts that
catches drift at BUILD time — if a new NodeLabel is added to
gitnexus-shared without being classified into one of the three
allowlists, TypeScript fails the _exhaustiveCheck assignment and
names the missing label.
The runtime guard stays as belt-and-suspenders: if a future contributor
bypasses the type check with @ts-ignore, the runtime guard still fires
in dev/test.
Implementation: converted the three allowlist Set<NodeLabel> initializers
to use `as const` tuples, then derived a union type from the tuples and
asserted `Exclude<NodeLabel, union> extends never`. Zero runtime impact
— the exported Sets are unchanged, Map.get hot-path performance is
unchanged, the test API is unchanged.
Resolves kieran-typescript KT-04.
Tests: 21/21 registration-table tests pass with zero modifications.
* refactor(test): restore type safety to createMockSymbolTable
createMockSymbolTable was widened to (overrides: any = {}): any with an
eslint-disable-next-line, and every buildTypeEnv call site passed the
mock as `model: mockSymbolTable as any`. The widening masked silent
false-green tests: buildTypeEnv accesses model.types/methods/fields,
and a flat any-typed override could silently return undefined from a
path that TypeScript should have caught at compile time.
Defined LegacyMockOverrides interface with typed stubs for each method
the mock can override (SymbolTable reads + TypeRegistry/MethodRegistry/
FieldRegistry lookups). Return type is now SemanticModel, so the mock
object is compile-checked against the real interface — a missing
registry method is a type error, not a silent runtime undefined.
Removed the eslint-disable and all 9 `as any` casts at call sites
(lines 1287, 1300, 1307, 2124, 2138, 5823, 5835, 5850, 5870). The
mock's return value now flows through buildTypeEnv's typed `model`
option without coercion.
Resolves kieran-typescript KT-01 and testing gap TG-02. This was the
highest-value cleanup in the plan — the only finding representing real
hidden test weakness.
Tests: 360 passed | 7 skipped (type-env.test.ts), typecheck clean.
* test: close coverage gaps in model/ registries
Added direct unit tests for the three owner-scoped registries that
previously had only transitive coverage via symbol-table.test.ts and
registration-table.test.ts. These new tests pin behaviors that were
flagged by the testing reviewer as untested or undertested.
method-registry.test.ts (14 tests):
- T-01: arity-fallback branch — when argCount matches no overload,
fall back to the full pool so fuzzy resolution still has candidates.
Previously untested and would have returned undefined instead of
a valid candidate if the branch regressed.
- T-02: requiredParameterCount range filtering — methods with default
parameters accept any argCount in [requiredParameterCount,
parameterCount]. Previously untested at the registry level.
- Variadic fallback (parameterCount=undefined is retained during arity
narrowing, bypassing range check).
- Return-type dedup paths: shared returnType → first wins, differing
returnTypes → undefined, firstReturnType=undefined → undefined,
single-overload skips dedup entirely.
type-registry.test.ts (9 tests):
- classByName homonym accumulation (two User classes in different
packages both returned).
- classByQualifiedName disambiguation — same simple name, different
FQNs resolve independently.
- Partial classes with identical simple + qualified name accumulate
in both indexes.
- registerImpl stores Rust impl blocks separately from classes.
- Multiple impl blocks per type accumulate.
field-registry.test.ts (6 tests):
- register/lookup round-trip, owner-scope isolation, last-wins on
duplicate key (flat map, not overload list).
- clear + re-register round-trip.
Extended symbol-table.test.ts cascade test (renamed from "both
registries" to "all three registries and the nested symbol table") to
also assert model.methods and model.fields are cleared — the test
name previously implied full coverage but only asserted types + symbols.
Resolves testing findings T-01, T-02, T-03, T-05.
Tests: 5667 passed (+29 new), 0 failures.
* refactor(test): replace brittle reference-equality tests + add intent comments
Two cleanups flagged as low-severity P3 by the testing reviewer:
1. registration-table.test.ts: Replaced three reference-equality tests
(hook identity via toBe) with behavioral tests that survive a future
refactor to per-label closures. The new "class-like behavior group"
describe iterates Class/Struct/Interface/Enum/Record/Trait and
verifies each one writes to types.registerClass. Same pattern for
Method/Constructor. A separate "behavior group isolation" describe
verifies class-like hooks don't leak into methods/fields and Impl
never pollutes registerClass. Strictly more coverage than the
reference-equality tests provided and implementation-independent.
2. symbol-resolver.test.ts: Added a comment above the lookupExactFull
and SM-16: getFiles() describes explaining why they intentionally
use createSymbolTable() directly instead of createSemanticModel().
The DAG leaf-only behaviors they test do not involve registries, so
testing the bare SymbolTable keeps the unit isolated. Prevents a
future reader from "fixing" the inconsistency.
3. qualified-class-lookups.test.ts: Added a comment above
`const symbolTable = model.symbols` explaining that processParsing
writes still reach the owner-scoped registries via SemanticModel's
fan-out — the alias is convenience, not a leaf in isolation.
Resolves testing T-04, kieran-typescript KT-05, kieran-typescript KT-06.
Tests: affected files all green (112 passed in registration-table +
symbol-resolver + qualified-class-lookups).
* refactor(model): collapse RoutingDecision wrapper and trim barrel surface
Two cleanups against the advanced-review findings on post-Unit-9 state:
S2 (cross-reviewer agreement — architecture-strategist + code-simplicity):
Delete the RoutingDecision single-field wrapper interface. Post-Unit-1
it held exactly one field (hook: RegistrationHook) and added pure
ceremony at every call site — `dispatchTable.get(key)!.hook(name, def)`
vs the now-direct `dispatchTable.get(key)!(name, def)`. Change the Map
type from Map<NodeLabel, RoutingDecision> to Map<NodeLabel,
RegistrationHook>, drop the interface, and update 17 test call sites.
A3 (architecture-strategist): Trim model/index.ts barrel surface.
createRegistrationTable, RegistrationHook, and RegistrationTableDeps
were re-exported from the barrel despite having zero legitimate
consumers outside model/ itself. The only callers (semantic-model.ts
and registration-table.test.ts) import directly from
./registration-table.js. Barrel exposure invited external callers to
construct orphan dispatch tables with independent registries,
weakening the SM-21 ownership inversion where SemanticModel is the
composition root. Kept CALLABLE_ONLY_LABELS, INERT_LABELS,
DISPATCH_LABELS exported since those remain useful for downstream
resolution logic and have no construction risk.
Resolves review findings:
- S2 (code-simplicity P3, 0.85) + architecture-strategist residual
- A3 (architecture-strategist P3, 0.82)
Tests: 5674 passed, 0 failures. Typecheck clean.
* refactor(model): replace runtime exhaustiveness guard with compile-time bijection
Replace the three-layer drift protection (hardcoded ALL_NODE_LABELS
array + 3 tuple consts + _ExhaustiveLabelCheck type + runExhaustivenessGuard
runtime + CI taxonomy test) with a single Record<NodeLabel, LabelBehavior>
map that structurally proves every invariant at compile time.
## Before
- ALL_NODE_LABELS hardcoded in semantic-model.ts (36 entries, could drift)
- DISPATCH_LABELS_TUPLE / CALLABLE_ONLY_LABELS_TUPLE / INERT_LABELS_TUPLE
private tuples (36 more entries total, could overlap or miss)
- _ClassifiedLabel / _UncoveredLabel type-level check (caught missing
labels but NOT duplicates across tuples)
- runExhaustivenessGuard runtime throw (only defense against duplicates)
- NodeLabel taxonomy coverage test in CI (same check as runtime guard)
Four defenses for invariants that the type system can express directly.
## After
```ts
type LabelBehavior = 'dispatch' | 'callable-only' | 'inert';
const LABEL_BEHAVIOR = {
Class: 'dispatch',
// ...36 entries...
Tool: 'inert',
} as const satisfies Record<NodeLabel, LabelBehavior>;
```
The `as const satisfies Record<NodeLabel, LabelBehavior>` combo enforces:
1. **Every NodeLabel must be a key** — Record requires all K keys.
Adding a NodeLabel to gitnexus-shared without classifying it here
fails with "Property 'X' is missing in type ..." naming the drifted label.
2. **No non-NodeLabel keys allowed** — `satisfies` with object literals
triggers excess-property checking. A typo'd key fails to compile.
3. **No duplicate classification** — impossible by construction; object
keys are unique at the source level.
4. **Valid category** — LabelBehavior is a narrow union, typos caught.
`ALL_NODE_LABELS`, `DISPATCH_LABELS`, `CALLABLE_ONLY_LABELS`, and
`INERT_LABELS` are now derived via `Object.keys(LABEL_BEHAVIOR)` and
`filter(l => LABEL_BEHAVIOR[l] === ...)` — single source of truth,
structurally impossible to drift.
## Deleted
- runExhaustivenessGuard() function in semantic-model.ts (~18 lines)
- ALL_NODE_LABELS hardcoded array in semantic-model.ts (~38 lines)
- DISPATCH_LABELS_TUPLE / CALLABLE_ONLY_LABELS_TUPLE / INERT_LABELS_TUPLE
private consts in registration-table.ts (~30 lines)
- _ClassifiedLabel / _UncoveredLabel / _exhaustiveCheck type machinery
(~20 lines)
## Kept named proofs: none
The `as const satisfies` on the object literal already catches all four
drift modes. Named type-level proofs (_MissingFromMap / _ExtraKeysInMap)
are pure duplication and were removed per review.
## Also in this commit
- S6: trim wrappedAdd narration comments in semantic-model.ts
(Step 1/2/3 block comments removed; kept the Function+ownerId WHY note)
- A3: tighten model/index.ts barrel — createRegistrationTable,
RegistrationHook, RegistrationTableDeps remain direct-imports only;
ALL_NODE_LABELS and LabelBehavior re-exported from the new home in
registration-table.ts
## Resolves
- Advanced-review S4 (runtime guard per-call cost) — guard no longer exists
- Advanced-review S1 (tuple three-defenses indirection) — single Record replaces all tuples
- Correctness P3 (exhaustiveness warns-only) — structurally impossible to drift
- Unit 6 type-level check — subsumed by the Record type
- Unit 3 runtime throw — no longer needed
Tests: 5674 passed, 0 failures. Typecheck clean.
* test(model): delete duplicate closure-isolation spy tests
S5 (code-simplicity P3): The 'closure isolation — each hook can only
write to its registry' describe block duplicated the 'behavior group
isolation' block's coverage via a different mechanism.
Behavioral tests (lines 151-174, kept):
table.get('Class')!('User', def);
expect(deps.methods.lookupMethodByOwner('unrelated', 'User')).toBeUndefined();
expect(deps.fields.lookupFieldByOwner('unrelated', 'User')).toBeUndefined();
Spy tests (deleted, ~55 lines):
vi.spyOn(deps.methods, 'register')
table.get('Class')!('User', def);
expect(methodsSpy).not.toHaveBeenCalled();
Both assert the same invariant — classHook does not touch the methods or
fields registries. The behavioral form observes the END STATE of the
registry (lookup returns undefined), which is the actual contract.
The spy form asserts the IMPLEMENTATION (a specific method was not
called), which couples to internal wiring — a refactor to a different
register function name would break the spy test while the behavioral
test would still pass.
Also dropped the now-unused `vi` import from vitest.
Tests: 24/24 registration-table.test.ts pass (-4 from spy deletion).
* refactor(model): compile-time cross-invariant between CLASS_TYPES and dispatch classHook
A1 (architecture-strategist P2, 0.90): CLASS_TYPES in symbol-table.ts
and the class-like entries of the dispatch table were two independent
hardcoded sets. Adding a new class-like label (e.g. Swift 'Extension')
to one but not the other would silently degrade qualifiedName
population — the symptom is subtle (partial qualified-name lookups)
and no test asserted the co-extensive invariant.
Fixed with a single source of truth and a two-layer compile-time
enforcement:
## symbol-table.ts
- Add `CLASS_TYPES_TUPLE` as `readonly [...] as const satisfies
readonly NodeLabel[]`. The `satisfies` forces every tuple entry to
be a valid NodeLabel at compile time.
- Export derived type `ClassLikeLabel = typeof CLASS_TYPES_TUPLE[number]`.
- Derive `CLASS_TYPES` Set from the tuple — same runtime shape as
before, now typed `ReadonlySet<NodeLabel>`.
## registration-table.ts
- Import `CLASS_TYPES_TUPLE` and `ClassLikeLabel` from symbol-table.ts.
- Narrow the `satisfies` on `LABEL_BEHAVIOR` via intersection:
Record<NodeLabel, LabelBehavior> & Record<ClassLikeLabel, 'dispatch'>
This forces every class-like label to have value 'dispatch' at
compile time. Adding a label to CLASS_TYPES_TUPLE without
classifying it as dispatch in LABEL_BEHAVIOR fails to compile with
a type error naming the drifted label.
- Build the class-like entries of the dispatch Map by iterating
`CLASS_TYPES_TUPLE` at factory time. Adding a label to the tuple
automatically wires it to classHook — no second place to update.
## What the design prevents
1. Drift scenario A (A1 original): 'Extension' added to CLASS_TYPES_TUPLE
but not to LABEL_BEHAVIOR → compile error on LABEL_BEHAVIOR's
satisfies.
2. Drift scenario B: 'Extension' added to CLASS_TYPES_TUPLE but not
wired to classHook → impossible because the Map is derived from the
tuple.
3. Drift scenario C: class-like label classified as something other
than 'dispatch' in LABEL_BEHAVIOR → compile error on the narrowed
intersection.
Runtime behavior unchanged: same 6 labels in CLASS_TYPES, same 6
class-like entries in the dispatch Map. Tests pin the behavior via
the existing behavior-group tests in registration-table.test.ts.
DAG unchanged: registration-table.ts already imported from symbol-table.ts
(the allowed upward direction). symbol-table.ts still imports nothing
from model/.
Tests: 5670 passed, 0 failures. Typecheck clean.
* test(field-extraction): use SemanticModel facade instead of raw SymbolTable
A6 (architecture-strategist P3, 0.85): field-extraction.test.ts created
its FieldExtractorContext fixture with `symbolTable: createSymbolTable()` —
a raw SymbolTable leaf, not the facade. In production, the context's
symbolTable field is always `model.symbols` (the SemanticModel-wrapped
facade where .add() dispatches through the owner-scoped registries).
The current field extractors don't call symbolTable.add() at all, so
this change is behavior-neutral today. The value is architectural
consistency — matching the test fixture to the production shape
prevents silent drift if a future field extractor starts registering
dynamically-discovered properties via the context. Without the fix,
such writes would hit the raw leaf and skip the fan-out, and tests
would pass even though the symptom (empty FieldRegistry) would
manifest in production.
Tests: 50/50 field-extraction.test.ts pass. Production tsc --noEmit
clean. Test-tsconfig error count unchanged (634 pre-existing errors
in unrelated test files, out of scope).
* refactor(A5): decouple model/resolve.ts from language registry
Move the MroStrategy type into gitnexus-shared and replace the
language: SupportedLanguages parameter on lookupMethodByOwnerWithMRO
with a direct mroStrategy: MroStrategy literal. Callers derive the
strategy from their language provider before invoking the resolver.
model/resolve.ts no longer imports from ../languages/index.js, so the
model/ layer is free of cross-layer coupling with the language
registry — this closes finding A5 from the SM-20/21/22/23 advanced
review (plan 006).
* feat(A4): add MethodRegistry.lookupMethodByName flat-by-name index
Add a secondary `methodsByName: Map<string, SymbolDefinition[]>` index
on MethodRegistry that returns every method with a given unqualified
name, accumulated across owners and overloads. The new index shares
SymbolDefinition references with methodByOwner — no duplication.
This is step 1 of the A4 double-index removal (plan 006). Tier 3
global resolution will switch to this index in Unit 3 so Method and
Constructor can be removed from CALLABLE_TYPES in Unit 4.
* refactor(A4): extend Tier 3 + memberCallByFile to consult method registry
Add model.methods.lookupMethodByName to Tier 3 global resolution in
resolution-context.ts and to the callable-pool build in
call-processor.ts (resolveMemberCallByFile + D2 widen path).
Intentionally behavior-preserving: Method and Constructor are still
in CALLABLE_TYPES so the new lookup returns identical candidates that
already reach Tier 3 through callableByName. Both paths dedup by
nodeId during this intermediate state — Unit 4 shrinks CALLABLE_TYPES
and the dedup is removed.
Part of plan 006 A4 step 2.
* refactor(A4): shrink CALLABLE_TYPES to free callables only
CALLABLE_TYPES = {Function, Macro, Delegate}. Method and Constructor
are no longer double-indexed in callableByName — they reach resolvers
through model.methods.lookupMethodByName instead.
Companion changes:
- Introduce CALL_TARGET_TYPES = CALLABLE_TYPES ∪ {Method, Constructor}
for the resolver's kind filter (filterCallableCandidates,
countCallableCandidates). Separates registration semantics (narrow)
from the resolver's acceptable-target set (wide).
- type-env.ts for-loop return-type inference consults both indexes,
treating the union as the authoritative call pool.
- resolveMemberCallByFile + D2 widen path keep the nodeId dedup in
place: Python/Rust/Kotlin class methods emitted as Function+ownerId
still land in both indexes until Unit 5 unblocks the normalization.
- Tier 3 global resolution (resolution-context.ts) keeps the same
dedup for the same reason.
Test updates reflect the new contract: Method/Constructor live in
methodsByName, not callableByName. Orphan Method-without-ownerId now
lives only in the file index (no registry coverage).
Part of plan 006 — closes A4 for strictly-labeled methods. Python/
Rust/Kotlin Function+ownerId normalization is tracked as Unit 5
(blocked).
* refactor: rename CALLABLE_TYPES → FREE_CALLABLE_TYPES
Pure rename. The constant's meaning changed in Unit 4 (free callables
only — no methods, no constructors) so the name now reflects that
scope: "callables that have no owner scope". Updates the constant
declaration and every consumer in src/ and test/.
Closes plan 006 Unit 6.
* refactor(A2): strict SymbolTableReader (pure reads) + SymbolTableWriter (+add)
Split the SymbolTable interface into three strictly layered surfaces:
- SymbolTableReader: lookups + iteration. NO add, NO clear. Holders
cannot mutate the table in any way.
- SymbolTableWriter extends Reader: + add. NO clear. Holders can
register new symbols but cannot trigger a leaf-index reset.
- InternalSymbolTable (private, not exported): + clear. The cascading
reset capability is reachable only through createSymbolTable's
return type, held exclusively by SemanticModel.rawSymbols.
SemanticModel.symbols is now typed as SymbolTableWriter — external
consumers (workers, processors, pipelines) can register symbols and
query them, but cannot reach .clear(). The A2 LSP fix holds: callers
holding any public reference cannot desync the leaf indexes from the
owner-scoped registries.
Delete the transitional `type SymbolTable = SymbolTableReader` alias
and migrate every consumer (src + test) to the explicit names:
- Field and parameter annotations use SymbolTableReader by default;
only code that calls .add() uses SymbolTableWriter.
- parsing-processor (workers + sequential paths) takes
SymbolTableWriter so it can register extracted symbols.
- field-types, call-processor, named-binding-processor,
workers/parse-worker: use SymbolTableReader (query-only).
- Tests: drop the stale `clear` fields from mock factories and
migrate the semantic-model cascade tests from the removed
model.symbols.clear() path to model.clear().
Closes plan 006 Unit 7. Industry sources: TypeScript compiler API
builder pattern, Salsa ParallelDatabase, .NET IReadOnlyList. See the
a2-lsp-clear-contract-research artifact for full citations.
* feat(A2): add SemanticModel.resetFileIndex() partial-reset entry point
Add a named method that clears only the leaf file and callable
indexes without cascading to the three owner-scoped registries
(types, methods, fields). Replaces the rare partial-reset use case
that was previously reachable via the now-removed symbols.clear()
path from A2 (plan 006 Unit 7).
JSDoc makes the semantic difference with model.clear() explicit so
future readers don't have to guess which method to call for a given
reingestion scenario.
Test-first: three scenarios cover the partial-vs-full semantics,
re-add after reset, and idempotency.
Closes plan 006 Unit 8.
* docs(S7): trim registration-table module JSDoc
Remove the ~24 lines of design-provenance citations from the module
JSDoc. The rust-analyzer, TypeScript-compiler, and Fowler references
are preserved in git history via the original SM-22 commits and in
plan 006 Unit 9.
Keep the ownership diagram, behavior-group table, and the
'How to add a new NodeLabel' checklist — those are load-bearing for
future contributors.
Closes plan 006 Unit 9 (S7 advanced-review finding).
* test(S3): migrate type-env.test.ts off LegacyMockOverrides
Replace the createMockSymbolTable bridge and LegacyMockOverrides
interface with real createSemanticModel() + add() calls across all
14 call sites. Where a test needs a specific registry lookup that
can't be pre-populated cleanly, use vi.spyOn on the real registry
instead.
Pattern breakdown:
- Pattern A (pre-populate via model.symbols.add): 13 sites
- Pattern B (vi.spyOn on registry lookup): 1 site
Deletes LegacyMockOverrides + createMockSymbolTable entirely. The
real MethodRegistry arity/returnType semantics match the hand-rolled
mock behavior in every migrated case, and no 'as any' casts remain
in the file.
Closes plan 006 Unit 10 (S3 advanced-review finding).
* refactor: remove unused MroStrategy type exports from language-provider and resolve modules
* refactor: relocate symbol-table, heritage-map, resolution-context into model/
Use git mv so blame and history follow each file:
- gitnexus/src/core/ingestion/symbol-table.ts → model/symbol-table.ts
- gitnexus/src/core/ingestion/heritage-map.ts → model/heritage-map.ts
- gitnexus/src/core/ingestion/resolution-context.ts → model/resolution-context.ts
These three files are part of the SemanticModel layer (file/callable
indexes, heritage parent map, tiered resolver) and now sit alongside
the registries they collaborate with. Updates every consumer import
path across src/ and test/ to the new locations.
* refactor(model): enforce pure-leaf DAG + delete legacy re-exports
model/ is now a pure leaf: zero upward imports and zero compat
shims in its parent processors. Completes the DAG cleanup started
in the previous commit.
1. walkBindingChain — moved into model/resolution-context.ts;
named-binding-processor.ts deleted.
2. NamedImportMap + NamedImportBinding + isFileInPackageDir —
moved into model/resolution-context.ts. Every consumer now
imports from the canonical location directly. Legacy re-exports
in import-processor.ts deleted.
3. c3Linearize + gatherAncestors — moved into model/resolve.ts.
mro-processor.ts imports them back for computeMRO. Legacy
c3Linearize re-export from mro-processor.ts deleted.
4. ExtractedHeritage type — moved into model/heritage-map.ts.
call-processor.ts, parsing-processor.ts, pipeline.ts,
heritage-processor.ts, and the test files now import it from
the canonical location. Legacy re-exports in parse-worker.ts
and heritage-processor.ts deleted.
5. resolveExtendsType — rewritten in model/heritage-map.ts to
take an explicit HeritageResolutionStrategy (A5-style DI).
buildHeritageMap accepts an optional getHeritageStrategy
callback; production uses getHeritageStrategyForLanguage from
heritage-processor.ts. Legacy resolveExtendsType re-export
from heritage-processor.ts deleted.
Verified:
- grep 'from "..' gitnexus/src/core/ingestion/model → empty
- grep 'Re-export for legacy' gitnexus/src/core/ingestion → empty
- npx tsc --noEmit → clean
- npx vitest run → 5686 passing
* docs(model): strip phase/plan references from module comments
Remove SM-20/21/22/23, A2/A4/A5, plan 006, Unit N labels and historical
phrasing ("previously", "legacy", "model-leaf DAG cleanup") from all 10
files in src/core/ingestion/model/. Preserve domain vocabulary (Tier
1/2/3), invariants, and caveats — only the plan archaeology is gone.
* refactor(model): tighten interface segregation + compile-time invariants
Apply four gated findings from branch-wide code review:
- SemanticModel.symbols now typed as SymbolTableReader; MutableSemanticModel
widens it back to SymbolTableWriter. ResolutionContext.model is typed as
MutableSemanticModel since it owns the lifecycle. Resolvers that only
query symbols can annotate their own fields as SemanticModel to drop
write access at the type level.
- Lookup methods (lookupExactAll, lookupCallableByName, lookupClassByName,
lookupClassByQualifiedName, lookupImplByName) now return
readonly SymbolDefinition[]. The returned arrays are live views into
the internal indexes; the readonly marker prevents accidental caller
mutation. walkBindingChain return type narrowed to match.
- FREE_CALLABLE_TUPLE + FreeCallableLabel exported from symbol-table.ts
as the single source of truth for free-callable labels. LABEL_BEHAVIOR
now satisfies Record<FreeCallableLabel, 'callable-only'> as a second
cross-invariant alongside Record<ClassLikeLabel, 'dispatch'>. Adding a
label to the tuple without classifying it as 'callable-only' fails at
build time. CALLABLE_ONLY_LABELS is now a re-export alias of
FREE_CALLABLE_TYPES so the two sets cannot drift.
- walkBindingChain fast-exits before allocating its cycle-detection Set
when the caller's file has no named bindings. Skips ~200k transient
Set allocations per large-repo resolution pass.
Also fixes five stale comments flagged by the review: duplicate JSDoc
block on RegistrationHook merged; resolve.ts "delegates to mro-processor"
direction corrected; RegistrationTableDeps JSDoc names
createRegistrationTable (not createSymbolTable); mro-processor.ts
"re-exported at top" stale comment removed; gatherAncestors export
comment matches reality.
tsc --noEmit clean, full test suite green (5786 tests).
* refactor(model): resolve four deferred P2 review findings
Address the four gated items from the branch-wide review that needed
design decisions before applying:
F#3 — Method/Constructor without ownerId fallback to callable index.
The dispatch hook silently skips owner-scoped labels that lack an owner
(an extractor contract violation — AST-degraded parse, or a buggy
language extractor). Pre-dispatch-table code let such defs fall through
to callableByName and stay reachable at Tier 3 global resolution. This
restores that fallback in SymbolTable.add so orphaned Methods and
Constructors don't silently vanish. Property deliberately does NOT
participate in the fallback to avoid polluting common names like
id / name / type.
F#4 — Delete MutableSemanticModel.resetFileIndex. The method had zero
production callers (only three tests), documented a "rare partial-
reingestion flow" that was never implemented, and contained the
adversarial-reviewer's double-populate trap: calling resetFileIndex
followed by re-adding the same class symbol would push a duplicate
SymbolDefinition into TypeRegistry.classByName without ever clearing
the first one. If incremental reingestion is ever needed, it can be
designed properly with per-file TypeRegistry invalidation. For now,
deleting the footgun is safer than documenting it.
F#5 — Compile-time dispatch-table completeness check. `LABEL_BEHAVIOR`
already enforces "every NodeLabel is classified" via
`Record<NodeLabel, LabelBehavior>`, but the dispatch-table factory
populated its Map with manual `table.set(...)` calls that TypeScript
could not correlate back to the `'dispatch'` classification. Add a
type-level `DispatchLabel` extracted from `LABEL_BEHAVIOR` via a
conditional mapped type, and build the table from an object literal
that satisfies `Record<DispatchLabel, RegistrationHook>`. Adding a new
dispatch-classified label without wiring it to a hook now fails the
build with a named-key error — no more silent no-op hooks.
F#7 — Tier 3 dedup fast-path via MethodRegistry.hasFunctionMethods.
The Set-based dedup between callableDefs and methodDefs is only needed
when a Python/Rust/Kotlin class method (emitted as Function+ownerId by
the worker) lands in both indexes. For TS/Java/C#/C++/Ruby-only repos
— where the two indexes are disjoint by construction — the dedup was
pure overhead on every global-tier hit. MethodRegistry now tracks
whether any Function-typed def was ever registered, and resolution-
context branches Tier 3 into a concat-only fast path when that flag
is false. Slow path with dedup survives unchanged for mixed-language
repos.
New tests pin the invariants: hasFunctionMethods flag transitions,
Method/Constructor orphan fallback, Property non-fallback, and the
MethodRegistry clear() reset. Full test suite green (5756 tests).
* refactor(model): close remaining P3 review findings + coverage gaps
Address the remaining review items in one batch.
Production refactors:
- Rename classHook → classLikeHook (M05). The hook handles Class /
Struct / Interface / Enum / Record / Trait; the vocabulary used in
surrounding docs and the behavior-group table is "class-like". The
rename makes the code match the taxonomy without forcing readers
through a mental glossary.
- Extract MAX_BINDING_CHAIN_DEPTH constant in resolution-context.ts
and document it as a known silent false-negative source (ADV-003).
Five hops cover the common TypeScript monorepo pattern; raising the
cap is a one-line change if a real repo exceeds it. walkBindingChain
consumes the constant so the 5 magic number no longer floats free.
- Replace defs.filter() allocation in MethodRegistry.lookupMethodByOwner
with a two-pass streaming count + conditional materialization
(PERF-04). Pure-match and pure-reject arity paths now skip the
filtered-array allocation entirely; only the discriminating case
(at least one match AND at least one rejection) pays it.
- Rewrite NOOP_SYMBOL_TABLE in parse-worker.ts and NOOP_SYMBOL_TABLE_SEQ
in parsing-processor.ts to implement all six SymbolTableReader
methods (ADV-005). The `as unknown as SymbolTableReader` cast is
removed in favor of a direct SymbolTableReader annotation, so future
additions to the interface surface as compile errors on the stubs
instead of silently falling through.
- type-env.ts getCallableUnionCount and getFirstCallable now take
`model: SemanticModel` as an explicit argument instead of reaching
into the enclosing `model!` non-null assertion (KT-003). Callers
enter via an `if (model)` guard and pass the narrowed reference, so
the non-null precondition is visible at the type level and the
closures cannot be accidentally extracted into a context without
the guard.
- Tier 3 dedup in resolution-context.ts now covers all four index reads
(classDefs, implDefs, callableDefs, methodDefs) via a pushUnique
helper (C-03). Previously classDefs and implDefs were spread directly
without dedup; any theoretical nodeId collision would have produced
duplicates in globalDefs.
Test infrastructure:
- Extract makeDef / makeMethod factory helpers into
test/unit/model/helpers.ts (T-07). The four registry/table test
files now import the shared helper and specialize with overrides,
removing ~25 lines of duplicated boilerplate and creating a single
point of maintenance.
New test coverage:
- T-01: c3 BFS fallback — cyclic Python hierarchy that fails c3
linearization and must fall back to heritageMap.getAncestors() BFS
order. Added to the lookupMethodByOwnerWithMRO describe block.
- T-02: Tier 2a-named precedence — verifies the binding chain walker
fires before Tier 2a import-scoped when an aliased import
`import { User as U } from B` competes with a raw same-name Tier 2a
hit. Also pins Tier 1 same-file precedence over Tier 2a-named.
- T-03: Tier 3 Function+ownerId dedup — end-to-end test that a Python
class method emitted as `Function + ownerId` yields exactly ONE Tier
3 candidate (not two). Companion test pins the fast-path branch for
hasFunctionMethods === false repos.
- T-06: walkBindingChain guards — circular re-export detection,
depth-cap exceeded drop, and boundary case at exactly
MAX_BINDING_CHAIN_DEPTH hops resolving successfully.
All tests added to a new test/unit/model/resolution-context.test.ts
dedicated to ResolutionContext.resolve() tier-precedence invariants.
Full suite: 5708 passing (minus the known Windows LBUG lock flake
that passes in isolation).
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* feat(group): bridge.lbug storage + contract matching expansion
Part 1 of 4 in the split of #606 (ticket: #791, closes#790 with a
revised plan per @magyargergo's request).
## What changed
Adds the LadybugDB-backed bridge storage infrastructure and extends
the contract matching algorithm with wildcard support. All changes are
additive: storage.ts, sync.ts, service.ts, cli/group.ts, mcp/tools.ts
are left on their upstream main versions and will migrate to the new
bridge in follow-up PRs (#792, #793, #794).
### Files
**New (844 LOC prod):**
- `gitnexus/src/core/group/bridge-db.ts` — atomic write-to-temp with
`retryRename` for Windows EBUSY/EPERM, per-item write tolerance via
`WriteBridgeReport`, `findContractNode` with three-tier symbol
lookup (uid → filePath+name → filePath)
- `gitnexus/src/core/group/bridge-schema.ts` — schema DDL
- `gitnexus/src/core/group/normalization.ts` — contract ID
canonicalization + `dedupeContracts` / `dedupeCrossLinks` helpers
used by both matching and bridge write
**Modified (+137 LOC prod):**
- `gitnexus/src/core/group/matching.ts` — adds `runWildcardMatch` for
`grpc::Service/*` wildcard consumers, `buildProviderIndex` helper,
and canonical gRPC ID handling in `normalizeContractId`
- `gitnexus/src/core/group/types.ts` — `MatchType` gains `'wildcard'`;
new `BridgeHandle` and `BridgeMeta` interfaces
**New tests (658 LOC):**
- `gitnexus/test/unit/group/bridge-db.test.ts` — core write/read round
trip, `WriteBridgeReport` shape, dropped-links counter, retryRename
behavior on EBUSY/ENOENT/EPERM/EACCES
- `gitnexus/test/unit/group/bridge-db-edge.test.ts` — edge cases
(malformed meta, missing contract nodes, concurrent access)
**Modified tests (+225 LOC):**
- `gitnexus/test/unit/group/matching.test.ts` — wildcard consumer
matching, gRPC canonical ID handling, same-service guard
### Self-review fixes folded in
Carried forward from the original #606 self-review:
- `writeBridge` try/finally handle lifecycle + `handleClosed` sentinel
- `openBridgeDbReadOnly` partial-handle cleanup
- `writeBridgeMeta` uses `retryRename` for Windows consistency
- `retryRename` unit tests (was zero coverage)
- Per-item try/catch around every CREATE loop so one malformed contract
doesn't abort the whole write
- Dropped cross-link counter (`linksDroppedMissingNode`)
### Why now
magyargergo asked for the #606 PR to be split so we can iterate with
confidence (https://github.com/abhigyanpatwari/GitNexus/pull/606#issuecomment-4229612271).
This is the foundational layer — pure infra, no user-facing surface,
no callers of the new APIs in this PR. Later PRs wire it in.
### How to verify
- `cd gitnexus && npx tsc --noEmit`
- `cd gitnexus && npx vitest run test/unit/group/bridge-db.test.ts --pool=forks`
- `cd gitnexus && npx vitest run test/unit/group/bridge-db-edge.test.ts --pool=forks`
- `cd gitnexus && npx vitest run test/unit/group/matching.test.ts --pool=forks`
- Pre-commit hook runs clean
### Risk / rollback
**Low.** All new code sits under `src/core/group/` in new files plus a
minimal `+16/-1` diff to `types.ts` and a `+136/-0` diff to `matching.ts`
(both purely additive). No existing callers reference the new APIs
(bridge-db, openBridgeOrFallback, runWildcardMatch) — the PRs that wire
them in come later in the split chain. Rollback = `git revert` of the
merge commit; no state introduced, no schema migration triggered.
### Scope discipline (per GUARDRAILS.md)
- Only the 8 files listed above are touched; no drive-by refactors
- No CI/release/security config changes
- No secrets, tokens, or machine-specific paths
- Content is lifted from the #606 branch which already passed CI 11/11
green on `d15b8cb` (before the split)
### Dependencies
- **Base:** `main` (no dependencies on other split PRs)
- **Blocks:** extractor expansion (#792), sync pipeline (#793),
cross-impact feature (#794)
- **Related ticket:** #791
Co-authored-by: Claude <noreply@anthropic.com>
* fix(group): address @claude review on #795
Addresses the findings from the automated review on PR #795
(https://github.com/abhigyanpatwari/GitNexus/pull/795#issuecomment-4229770000
— posted by @magyargergo / claude-code Action run).
### Medium severity (reviewer flagged as blockers)
- **bridge-db.ts `openBridgeDbReadOnly` bak recovery** — the `.bak`
recovery path used bare `fsp.rename(bakPath, dbPath)`, which is
exactly the scenario most likely to hit Windows EBUSY/EPERM (an
interrupted writer still holding the handle for a few ms). Switched
to `retryRename` for consistency with the rest of the file's
Windows-safe rename path.
- **bridge-db.ts `ensureBridgeSchema` error detection** — the inline
`msg.includes('already exists')` substring match has been lifted
into a named constant `LBUG_ALREADY_EXISTS_MSG` with a comment
documenting the coupling to LadybugDB's error message wording and
why we can't use `IF NOT EXISTS` (LadybugDB DDL doesn't support it)
or typed errors (LadybugDB's JS driver doesn't expose error codes).
Also tightened the `catch (err: any)` to `catch (err: unknown)`.
- **bridge-db.ts `findContractNode` — extracted out of writeBridge**
— the 35-line async closure living inside `writeBridge` has been
lifted to three module-level functions: `createContractLookupIndex`,
`indexContract`, and `findContractNode`. `findContractNode` is now
a pure synchronous function taking a prebuilt index instead of
doing its own DB queries. The `writeBridge` cross-link loop is now
~25 lines instead of ~100.
- **bridge-db.ts `findContractNode` — N+1 query elimination** — the
old inner-closure version issued up to 6 DB round-trips per
cross-link (2 endpoints × up to 3 tiers of fallback queries). For a
group with 1000 cross-links, that's up to 6000 DB queries just to
resolve endpoints. The new version consults an in-memory
`ContractLookupIndex` built incrementally as contracts are inserted
(`indexContract` called AFTER each successful insert so failed
inserts don't poison the index). Cross-link resolution is now
O(1) per link instead of O(3) DB queries per link, with zero DB
round-trips during the cross-link loop.
### Minor severity
- **bridge-db.ts `queryBridge` empty-array guard** — if LadybugDB
ever returns an empty `QueryResult[]` at the top level (shouldn't
happen with single-statement calls, but driver contract isn't
explicit), the old code would call `.getAll()` on `undefined` and
crash with a confusing stack. Added an `unwrapQueryResult` helper
that throws an explicit `'empty QueryResult array'` error instead,
making a potential driver regression visible immediately.
- **normalization.ts `contractRichness` weights** — added a
block-level comment documenting the weight ordering (+3 for
symbolUid, +2 for each symbol-identifying field, +1 for service
tag or non-manifest origin) and explicitly noting that the
absolute numbers don't matter, only the relative ordering. Matches
the "comment for contributors" suggestion in the review.
- **bridge-schema.ts `BRIDGE_SCHEMA_VERSION` migration comment** —
added a 4-point contract explaining what bumping the constant
means ("discard and re-sync" strategy for V1, no in-place
migration yet, new migration logic should live in a separate
`bridge-migrations.ts` module when it becomes necessary).
- **test/unit/group/fixtures.ts** — extracted the `makeContract`
helper previously copy-pasted between `bridge-db.test.ts` and
`bridge-db-edge.test.ts` into a shared fixtures module. Both test
files now import from `./fixtures.js`. Kept the scope minimal:
fixtures is NOT a general-purpose factory module, just the shared
baseline contract builder.
### New tests
Added 9 pure-function unit tests for the now-extracted
`findContractNode` in `bridge-db.test.ts`:
- returns null on empty index
- tier 1 (symbolUid) match, including repo-scope and role-scope
isolation
- tier 2 (filePath + symbolName) fallback when symbolUid is empty
or mismatches
- tier 3 (filePath only) when exactly one contract lives in the
file, and refusal when multiple do
- priority ordering when multiple tiers could resolve
These are fully isolated — no DB, no temp directories, no native
LadybugDB binding — so they run in <10ms total and are
immediately trustworthy as a regression safety net.
### Deliberately deferred (reviewer marked as "fine for now")
- `BridgeHandle._db` / `._conn` typing to `unknown` with casts in
`bridge-db.ts` — reviewer's note: "The typing is fine for now."
- Batch inserts via `UNWIND` — needs LadybugDB support confirmation,
tracked as a follow-up; the per-item pattern remains.
- `queryBridge` prepared-statement lifecycle — the current pattern
(prepare → execute → GC) relies on LadybugDB's internals, worth
verifying against their docs in a separate audit.
### Scope discipline (per `GUARDRAILS.md`)
- Only files touched by this PR (`bridge-db.ts`, `bridge-schema.ts`,
`normalization.ts`, both bridge test files, new `fixtures.ts`) —
no drive-by refactors
- No CI/release/security config changes
- No secrets
### Test + typecheck status
- `npx tsc --noEmit` clean
- `bridge-db.test.ts`: added 9 `findContractNode` tests, all pass in
isolation. The full-file run still hits the pre-existing native
LadybugDB cleanup segfault that flakes the reported count — same
as every prior commit on this branch, not a regression.
- `bridge-db-edge.test.ts`: 4/4 pass
- `matching.test.ts`: 28/28 pass
- `types.test.ts`: 5/5 pass
- `retryRename` tests (4/4) and `findContractNode` tests (9/9)
verified in isolation via `-t` filter
Co-authored-by: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix: restore tree-sitter-swift postinstall patch for macOS ARM64
PR #516 (77dcb06) deleted `scripts/patch-tree-sitter-swift.cjs` and
the `postinstall` script entry when bumping to `tree-sitter-swift@0.7.1`,
since 0.7.1 ships prebuilt darwin-arm64 binaries and no longer needs the
patch. PR #538 (01ddc3e) then had to revert `tree-sitter-swift` back to
`^0.6.0` (and `tree-sitter` back to `^0.21.1`) because `npm overrides`
doesn't apply when gitnexus is installed via `npx -y` (gitnexus isn't the
root project, so overrides are silently ignored, producing ERESOLVE errors).
PR #538 reverted the grammar package changes but did not restore the patch
script, leaving `tree-sitter-swift@0.6.0` unable to build its native
binding on macOS ARM64. The symptom is `gitnexus analyze` printing
"Skipping swift" or "swift parser not available".
`Dockerfile.test` still references `node scripts/patch-tree-sitter-swift.cjs`
(added in the same PR #516), confirming the regression — the test image
build is also broken.
This commit restores the patch script from commit `0c8ec95` (the last
revision before it was deleted) and re-adds the `postinstall` entry to
`package.json`. No logic changes — it is an exact restoration.
The TODO comment in the script ("Remove this script when tree-sitter is
upgraded to ^0.22.x") still applies.
* style: run prettier on patch-tree-sitter-swift.cjs
The C# tree-sitter query set only matched `base_list` on
`class_declaration`, so interfaces extending other interfaces
(`interface IFoo : IBar`) were never captured as heritage edges.
This broke transitive interface implementation chains. For example,
given:
interface IBase { }
interface IFoo : IBase { }
class MyClass : IFoo { }
only `MyClass -> IFoo` was emitted, and the `IFoo -> IBase` edge was
silently dropped. Any analysis that relies on walking the full
interface inheritance chain (e.g. "which classes implement IBase?")
therefore returned incomplete results.
This patch adds two new query patterns mirroring the existing
class_declaration heritage patterns, but targeting
`interface_declaration`:
(interface_declaration name: (identifier) @heritage.class
(base_list (identifier) @heritage.extends)) @heritage
(interface_declaration name: (identifier) @heritage.class
(base_list (generic_name (identifier) @heritage.extends))) @heritage
The existing heritage-processor pipeline already handles these
captures correctly once the query emits them, so no changes are
needed outside of tree-sitter-queries.ts.
Testing:
- New fixture `csharp-interface-heritage/` covering:
* interface : interface (single base)
* interface : interface, interface (multiple bases)
* class : interface (where that interface derives from others)
- 6 new test cases in test/integration/resolvers/csharp.test.ts
asserting exactly 4 IMPLEMENTS edges and 0 EXTENDS edges for the
fixture.
- Full C# resolver suite: 175/175 passing, no regressions.
Co-authored-by: Prota100 <Prota100@users.noreply.github.com>
* Fix: replace recursive AST traversal with iterative stack to prevent stack overflow on large files
Fixes#752
Large PHP files (2000+ lines) with deeply nested AST structures (closures,
array literals, chained method calls) cause "Maximum call stack size exceeded"
during analysis. This converts three recursive tree traversal functions to
iterative loops using explicit stacks:
1. `walk()` in type-env.ts — the main AST walker that processes every node.
On a 2,462-line PHP controller, this recurses through 5,000-10,000+ nodes.
2. `findRelationCall()` in languages/php.ts — recursive search for Eloquent
relationship calls within method bodies.
3. `findDescendant()` in utils/ast-helpers.ts — generic recursive utility
used by PHP property extraction and other parsers.
All three now use a while loop with an array-based stack instead of function
call recursion, eliminating V8's ~10K frame call stack limit as a constraint.
Tested against a production Laravel codebase with 373 PHP files (87,723 lines
total, largest file 2,462 lines) — indexes successfully in 17.4s with zero
errors, where the recursive version would crash with stack overflow.
* Fix: reverse child push order in findRelationCall iterative traversal
The iterative stack-based traversal pushed children in forward order,
causing the last child to be processed first (LIFO). This reversed the
original recursive left-to-right DFS order. Push children in reverse
so the first child ends up on top of the stack.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Fix: move stack declaration before processNode, rename walk to processNode
Move the stack initialization above the function that pushes onto it,
making the data-flow order match the code order. Rename walk to
processNode since it now processes a single node rather than
recursively traversing the tree.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: load VECTOR extension during DB init for semantic search
The VECTOR extension was only loaded inside the embedding generation
pipeline (createVectorIndex). On a fresh gitnexus serve session,
semantic and hybrid search failed because QUERY_VECTOR_INDEX was
unknown.
Now loads the VECTOR extension alongside FTS during database
initialization in both the single-connection and pool-based paths.
Fixes#766
* fix: reset vectorExtensionLoaded on DB close and retry paths
The vectorExtensionLoaded flag was not being reset in closeLbug() or
the busy-retry cleanup path in withLbugDb(). This caused the VECTOR
extension to not be re-loaded after a close+re-init cycle, breaking
semantic search on reconnection.
Also resets shared.ftsLoaded and shared.vectorLoaded in the pool
adapter closeOne() for external DB entries, preventing stale extension
state when the pool is re-opened.
Adds integration tests covering vector extension loading, idempotency,
and state reset on both close and busy-retry paths.
* fix: set ftsLoaded flag in initLbugWithDb to avoid redundant extension reloads
* fix: set shared.vectorLoaded flag in initLbugWithDb to avoid redundant reloads
* fix: map diff hunks to symbol line ranges in detect_changes
The detect_changes tool previously used `git diff --name-only` and
picked the first 20 arbitrary symbols from each changed file. This
produced false positives (unchanged symbols reported as modified) and
false negatives (actually changed symbols dropped by the LIMIT).
Now uses `git diff -U0` to get unified diff with hunk headers, parses
the @@ line ranges, and queries for symbols whose [startLine, endLine]
range overlaps the diff hunks. Only truly touched symbols are reported.
Also fixed the CONTAINS path match to ENDS WITH to prevent cross-file
false positives from substring matching.
Fixes#758
* fix: address review feedback - variable shadowing, batch queries, tests
- Rename `params` to `queryParams` in detectChanges hunk-mapping loop
to avoid shadowing the outer method parameter
- Replace N+1 per-symbol process lookup with a single batched query
using WHERE n.id IN $ids (same pattern as impact BFS traversal)
- Add unit tests for parseDiffHunks covering single/multi file,
single/multi hunk, omitted count, pure-deletion, and empty input
* style: fix prettier formatting in parse-diff-hunks test
* Initial plan
* SM-19: Replace resolveCallTarget with thin dispatcher
Delete the monolithic resolveCallTarget function (~200 lines) and replace it
with a 15-line thin dispatcher that routes to resolveMemberCall,
resolveStaticCall, or resolveFreeCall. Extract module-alias resolution and
file-based member-call fallback into dedicated helper functions.
- resolveCallTarget body reduced from ~200 lines to ~15 lines
- Extract resolveModuleAliasedCall helper (Python/Ruby module imports)
- Extract resolveMemberCallByFile helper (trait dispatch, overload disambiguation)
- Extract singleCandidate helper (constructor alias fallback, name-based fallback)
- Update unit tests for new dispatcher semantics
- Update doc comments referencing deleted D0-D4 paths
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/469eac38-b0c0-4a26-a2ff-3eb06299730b
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* SM-19: Add singleCandidate tail fallback for member calls with unresolvable receiver type
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/469eac38-b0c0-4a26-a2ff-3eb06299730b
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix(SM-19): address all PR #770 review findings + fix CI
Fixes all 5 test failures (2 unit + 3 integration) and addresses 10
review findings from comment 4225312416.
Critical fix — singleCandidate null-route guard
The SM-19 dispatcher chained singleCandidate as an unconditional tail
fallback for member calls with receiverTypeName. This bypassed the
SM-10 R3 null-route contract: when the receiver type IS in the index
but file/owner filtering produced zero matches, the old code returned
null (genuine miss), but the new code fell through to singleCandidate
(false-positive CALLS edge).
Root cause: resolveMemberCallByFile returns null for two semantically
different reasons — (1) type not found in the index at all, and
(2) type found but no candidate matched after narrowing. The dispatcher
treated both as "try the next fallback." The old resolveCallTarget
exited the entire function on case 2.
Fix: after the scoped resolvers both return null, check whether the
receiver type resolves in the index. If it does (case 2), null-route
— the scoped resolvers made the right decision. If it doesn't (case 1,
e.g. PHP 'mixed', dynamic types), singleCandidate is the correct last
resort. ctx.resolve is cached so the check is free.
This fixes:
- Unit: no heritageMap null-route test (was getting 1 edge, expects 0)
- Integration: Rust c.trait_only() negative test
- Integration: 3 PHP heritage + alias tests (singleCandidate correctly
fires when the receiver type is not in the index)
Performance (findings #1, #2, #3)
- Thread pre-computed tiered result into resolveModuleAliasedCall via
new tieredOverride parameter — eliminates the duplicate ctx.resolve
call on every module-alias path.
- Add countCallableCandidates helper that short-circuits at threshold
without allocating an intermediate array — replaces the
filterCallableCandidates(...).length > 1 allocation in skipMember.
- resolveMemberCallByFile lookupCallableByName caching deferred to a
follow-up (finding #2) — the fix requires threading widenCache
through the file-scoped resolver which is a larger change.
Code quality (findings #4, #5)
- Remove dead code: redundant conditional in resolveMemberCallByFile
where both branches returned null.
- Move WidenCache type declaration from mid-file (between JSDoc blocks)
to adjacent to CONSTRUCTOR_TARGET_TYPES with other type declarations.
Formatting
- Applied prettier to call-processor.ts (CI format check was failing).
Verification
- tsc --noEmit clean
- 3188 unit tests pass (0 skipped real tests)
- 1766 resolver integration tests pass
- Zero regressions — all PHP, Rust, and no-heritageMap tests green
Review: https://github.com/abhigyanpatwari/GitNexus/pull/770#issuecomment-4225312416
* fix(SM-19): restore module-alias narrowing and constructor disambiguation
Codex adversarial review on PR #770 surfaced two silent regressions in the
SM-19 thin dispatcher:
Finding 1 [high] — Typed member calls bypassed module-alias narrowing.
When two homonym receiver types are both imported by the caller, the
import-scoped tier no longer narrows and the owner/file resolvers see
genuine ambiguity. The dispatcher null-routed silently, dropping valid
CALLS edges. Fix: consult `resolveModuleAliasedCall` at the top of the
typed-member branch so an active alias on `call.receiverName` picks the
aliased file before the generic resolvers run.
Finding 2 [medium] — Constructor dispatch lost overload disambiguation.
When `resolveStaticCall` bails (ambiguous or ownerless Constructor pool)
and the caller supplied `overloadHints` / `preComputedArgTypes`, the
branch fell straight through to `singleCandidate` — which also bails on
multiple same-arity survivors. Fix: between `resolveStaticCall` and
`singleCandidate`, run constructor-filtered overload disambiguation on
the tiered pool. Only engages when a narrowing signal is present;
preserves SM-10 R3 null-route for genuinely ambiguous cases.
Tests:
- call-processor.test.ts: 3 new dispatcher-level regression tests
covering real-homonym alias narrowing, constructor overload
disambiguation with `argTypes`, and null-route control
- symbol-table.test.ts: update `module alias homonyms` test which
previously codified the Finding 1 regression; now asserts resolution
to the aliased file's method
Verification: 3191 unit + 2398 integration tests pass; tsc --noEmit
clean; prettier clean.
* refactor(SM-19): address code review findings with clean-code pass
Code review on commit f424685e surfaced one P1 correctness regression and
two P2 maintainability concerns. This commit closes all ten findings:
P1 — Alias helper placement regression
- resolveModuleAliasedCall now runs as a FALLBACK in the typed-member
branch, after resolveMemberCall/resolveMemberCallByFile return null.
Previously it short-circuited BEFORE scoped resolvers, leaking unrelated
homonyms from the aliased file when a local var coincidentally matched
a module alias.
- Added type-file verification guard: alias narrowing only fires when the
alias target file is among the receiver type's defining files. Prevents
cross-type false positives and hardens SM-10 R3.
P2 — Thin-dispatcher drift (roadmap Phase 3)
- Extracted disambiguateByOverloadOrArgTypes shared helper. Centralizes
the overloadHints → preComputedArgTypes precedence rule used by both
member and constructor resolvers.
- Folded constructor overload disambiguation into resolveStaticCall as
step 4.5 (between the ambiguous-pool bail and the instantiable-class
fallback). resolveStaticCall now accepts optional overloadHints /
preComputedArgTypes symmetric with resolveMemberCallByFile.
- Dispatcher's constructor branch returns to a 2-line delegation.
- resolveMemberCallByFile now calls the shared helper instead of inlining
the ternary.
P2 — Missing test coverage
- owner-scoped wins over alias narrowing (alias with unrelated target
class must not override unique owner-scoped answer)
- alias narrowing rejects unrelated target type (type-file guard)
- alias fallthrough: receiverName not in alias map
- alias fallthrough: alias target file has no matching method
(overloadHints-for-constructor variant transitively covered via the
extracted helper's member-path tests; direct dispatcher test deferred
as it requires real OverloadHints fixture parsing)
P3 — Clarity and durability
- Stripped "Codex SM-19 Finding N" prefixes from comments. Replaced with
durable explanations of WHY each guarded branch exists.
- Added cross-reference comment at the tail-branch resolveModuleAliasedCall
call site pointing to the typed-member branch usage.
Verification: 3195 unit + 1766 resolver integration + 2398 full integration
tests pass. tsc --noEmit clean. prettier clean.
Plan: docs/plans/2026-04-11-002-fix-sm19-code-review-findings-plan.md
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* fix: resolve false 404s and stale repo context during multi-repo switching on Windows
* test(e2e): add repo-switching tests — hold-queue 503, ?project= URL, Windows path normalization
* test(e2e): fix repo-switching specs — use live backend with ?server= param
* Initial plan
* Update test files for SymbolTable interface changes
Remove lookupFuzzy, lookupFuzzyCallable, globalIndex, and callableIndex
references from all test files. Replace lookupFuzzyCallable with
lookupCallableByName. Update getStats assertions to only expect
{ fileCount }. Remove tests that exclusively tested removed methods.
Files updated:
- symbol-table.test.ts: Remove lookupFuzzy describe block and all
globalIndex/callableIndex tests, update callable method references
- symbol-resolver.test.ts: Remove SM-16 lookupFuzzy test block,
update Tier 3 describe title
- type-env.test.ts: Update all mock SymbolTable objects and spy
variable names
- call-form.test.ts: Update ownerId propagation test
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(SM-18): Remove lookupFuzzy, lookupFuzzyCallable, globalIndex, callableIndex
Remove from SymbolTable interface and implementation:
- lookupFuzzy method
- lookupFuzzyCallable method
- globalIndex Map
- callableIndex Map (renamed to callableByName, backing lookupCallableByName)
Add lookupCallableByName as the targeted replacement for fuzzy callable
lookups. Migrate all production callers:
- resolution-context.ts: lookupFuzzyCallable → lookupCallableByName
- type-env.ts: lookupFuzzyCallable → lookupCallableByName
- call-processor.ts: lookupFuzzy → lookupCallableByName (D2 widen paths)
Remove fuzzyCallCount/fuzzyCallableCallCount stats and globalSymbolCount
from getStats(). Update pipeline.ts logging accordingly.
Memory savings: globalIndex stored every non-Property symbol (typically
the largest index by entry count). Removing it eliminates one Map plus
all its per-name arrays — net savings proportional to unique symbol
count in the project.
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/4a658c69-41a9-4d57-8527-50ca544ca967
* fix(SM-18): address all PR #769 review findings
1. type-env.test.ts mock: add missing lookupImplByName + getFiles methods.
2. Macro/Delegate tests: 2 new tests confirm C/C++ Macro and C# Delegate
are indexed in callableByName.
3. D2 widen path test: module-alias scenario verifying lookupCallableByName
resolves methods in aliased files that shadow same-file definitions.
4. CALLABLE_TYPES unified: exported from symbol-table.ts (single source of
truth), imported in call-processor.ts. Removed duplicate
CALLABLE_SYMBOL_TYPES constant.
5. getStats() observability restored: tier hit counters (tierSameFile,
tierImportScoped, tierGlobal, tierMiss) replace the removed
fuzzyCallCount diagnostic.
* chore(SM-18): remove unnecessary `as any` casts on valid NodeLabel types
Macro, Delegate, TypeAlias, Const, and Variable are all valid NodeLabel
values in gitnexus-shared. The casts suppressed type checking without
purpose and signaled false uncertainty.
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* Initial plan
* chore: initial plan for SM-16 resolveUncached refactor
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0f505332-25be-46a7-b78e-fde58c1fc6fd
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* feat(SM-16): restructure resolveUncached — replace lookupFuzzy with targeted index lookups
- Remove single lookupFuzzy call that fed all Tier 2a/2b/3 in resolveUncached
- Tier 2a: iterate importedFiles with lookupExactAll per file (O(imports) × O(1))
- Tier 2b: iterate symbols.getFiles() filtered by isFileInPackageDir + lookupExactAll
(O(files) × O(1), avoids global name scan)
- Tier 3: replace with lookupClassByName + lookupImplByName + lookupFuzzyCallable
(three O(1) index lookups covering class-like, Rust impl blocks, and callables)
- Add getFiles() to SymbolTable interface (exposes fileIndex.keys() for Tier 2b)
- Add lookupImplByName() to SymbolTable — dedicated Rust Impl index kept separate
from classByName to preserve correct heritage-map resolution
- Remove allDefs parameter from walkBindingChain; always use lookupExactAll directly
- Add 29 new unit tests covering SM-16 changes and per-language fixtures
- fuzzyCallCount in getStats() is now 0 for all resolve() calls (acceptance criterion)
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0f505332-25be-46a7-b78e-fde58c1fc6fd
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix(SM-16): clean up — readable Tier 3 if-else, correct doc comment, remove unused import
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0f505332-25be-46a7-b78e-fde58c1fc6fd
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix(SM-16): address all PR #764 review findings
1. Eager callableIndex — maintained on add() like classByName/implByName,
removing the O(globalIndex) lazy rebuild on the Tier 3 hot path.
2. Tier 2b inverted index — packageDirSuffix→Set<filePath> built lazily on
first Tier 2b hit. Changes O(allFiles×packages) per resolution to
O(packages×filesInPackage).
3. Tier 3 type exclusion documented — TypeAlias, Const, Variable are
intentionally not reachable at Tier 3. 4 negative/positive tests added.
4. Tier 3 allocation guard simplified — single spread replaces 4-way if-else.
5. getFiles() live iterator documented with safety contract.
6. Remaining lookupFuzzy callers in call-processor.ts documented in the
Tier 3 comment block.
7. Tier 2b language fixtures — added Rust, Kotlin, PHP tests (3 new).
Also merges origin/main (SM-15 accumulator fixes).
* fix(SM-16): address Codex adversarial review — Tier 2b cache lifecycle + Macro/Delegate at Tier 3
1. Tier 2b packageDirIndex now invalidated in clearCache() and clear(),
preventing stale snapshots when symbols/packages are added between
chunk processing phases.
2. Macro (C/C++) and Delegate (C#) added to CALLABLE_TYPES in the eager
callableIndex, restoring Tier 3 reachability for these call targets
that the old lookupFuzzy returned.
* fix(SM-16): address ce:review findings — Tier 2b cache lifecycle + Tier 3 perf + test gaps
1. packageDirIndex no longer invalidated in clearCache() — the index
persists across file boundaries since packageMap and symbols are
append-only during the calls phase. Only clear() (pipeline reset)
invalidates. Prevents O(files×dirs) rebuild per-file.
2. Tier 3 short-circuit: return null before spread when all three
indexes are empty, avoiding allocation on the common miss path.
3. Add Macro (C/C++) and Delegate (C#) Tier 3 regression tests —
the only newly-added CALLABLE_TYPES were completely untested.
4. Add packageDirIndex invalidation regression test — verifies clear()
resets the index and newly-added symbols are visible.
* fix(SM-16): address final review — deduplicate NamedImportMap + doc fixes
1. NamedImportMap: removed duplicate definition from resolution-context.ts,
now imported directly from import-processor.ts (no re-export needed —
no consumers imported it from resolution-context).
2. packageDirIndex build cost documented accurately in comment.
3. fuzzyCallCount scope documented in test comment.
4. Tier 2a test suite: added comment about Go/Kotlin/PHP coverage.
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* Initial plan
* Initial setup - Phase 9 BindingAccumulator cross-file return type wiring
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7cee6490-090d-4714-8cb5-a704168ff47a
* feat(SM-15): wire BindingAccumulator into processCallsFromExtracted for Phase 9 cross-file return type propagation
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7cee6490-090d-4714-8cb5-a704168ff47a
* fix(SM-15): address all PR #763 review findings
Performance (R1)
- Changed _fileScopeByFile from Map<string, [string,string][]> to
Map<string, Map<string,string>>. fileScopeGet(filePath, name) is
now O(1) — replaces the O(n) linear scan + defensive-copy alloc
that ran once per ConstructorBinding entry. fileScopeEntries()
reconstructs tuples from Map.entries() for backward compat.
- Updated finalize() dev-mode invariant to compare deduplicated Map
size rather than raw array length (Map.set deduplicates same-name).
Lifecycle (R2)
- Documented that Phase 9 intentionally reads pre-finalize because
finalize() cannot move before both the worker consumer (line 984)
AND the sequential-path writer (line 1061). Pre-finalize reads are
safe because finalize() is write-lock-only with no side effects.
Replaced the ambiguous "populated but not yet finalized" comment
with the full lifecycle ordering explanation.
Sequential-path parity (R3)
- Wired bindingAccumulator into processCalls at line 797 (sequential
path) so verifyConstructorBindings gets the Phase 9 fallback.
- Added bindingAccumulator parameter to processAssignmentsFromExtracted
signature and wired it at the pipeline.ts call site (line 1026).
- Both paths now produce identical Phase 9 behavior for the same code.
Tracking comments (R4)
- Added "Overlapping mechanism (N of 3)" cross-references at:
1. buildImportedReturnTypes (~line 109)
2. collectExportedBindings (~line 168)
3. Phase 9 fallback in verifyConstructorBindings (~line 563)
Each links to the other two and notes future unification.
Language coverage (R5)
- Added 5 new Phase 9 integration test suites in cross-file-binding.test.ts:
JavaScript, C++, C#, PHP, Ruby. Each uses the existing fixture
directories and asserts getUser() → User → user.save() resolves.
Total cross-file binding tests: 52 (was 37).
Quality asymmetry (R6)
- Added inline comment at the Phase 9 fallback noting worker-path
entries are Tier 0/1 only and that binding accuracy is structurally
lower for large repos where the worker path dominates.
Tests (+21 new)
- 6 fileScopeGet unit tests (happy path, unknown file/name, mixed
scopes, post-dispose, duplicate varName last-write-wins)
- 15 integration tests across 5 new language suites
Verification
- tsc --noEmit clean
- 3147 unit tests pass (+6 new)
- 52 cross-file binding integration tests pass (+15 new)
- 1766 resolver integration tests pass
- Zero regressions
Plan: docs/plans/2026-04-10-001-fix-sm15-review-findings-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/763#issuecomment-4220354242
* fix(SM-15): gate accumulator fallback on resolution tier and fix sequential file-order dependency
Two Codex adversarial reviews identified medium-severity bugs in the Phase 9
BindingAccumulator fallback:
1. Local-first violation: the fallback fired regardless of whether ctx.resolve()
found same-file candidates, letting an imported callee shadow a local one
and produce false CALLS edges. Fixed by gating on tiered.tier !== 'same-file'
and callableDefs.length <= 1.
2. Sequential file-order dependency: processCalls flushed and verified per-file,
so consumer files processed before their providers missed accumulator bindings.
Fixed by splitting into a flush pre-pass (all files) then a resolution loop,
mirroring the worker path's "all appends before any reads" pattern.
Also adds 11 consumer-before-provider integration test fixtures (one per
supported language) and 4 unit tests for tier gating edge cases.
* refactor(SM-15): eliminate duplicated prepare logic in processCalls two-pass split
Replace the duplicated pre-pass + legacy-path code (parse → query → heritage
→ TypeEnv → exports) with a single preparation loop followed by a resolution
loop. Both paths now share the same preparation code — the only conditional
is the accumulator flush.
Side benefit: globalParentMap is now fully populated before any resolution
runs, improving cross-file isSubclassOf accuracy regardless of file order.
Net -118 lines (226 removed, 108 added).
* fix(SM-15): address PR #763 third-pass review findings
1. Update stale dispose() JSDoc — remove forward-reference to Phase 9
wiring that is now complete; document actual consumers.
2. Add processAssignmentsFromExtracted Phase 9 unit test — verifies the
accumulator fallback produces ACCESSES write edges when the SymbolTable
has no returnType for the callee.
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* Initial plan
* feat(SM-13): extract resolveFreeCall from resolveCallTarget
Extract the free-function call resolution path into a dedicated
`resolveFreeCall(calledName, filePath, ctx)` function that uses
`lookupExact` + import-scoped resolution via `ctx.resolve()`.
- Free function calls (foo()) now route through `resolveFreeCall`
- Swift/Kotlin implicit constructors (User()) delegate to
`resolveStaticCall` within `resolveFreeCall`
- `resolveCallTarget` dispatches `callForm === 'free'` early,
removing the inline freeFormHasClassTarget logic
- S0 block simplified to only handle `callForm === 'constructor'`
- Global (Tier 3) fallthrough preserved via ctx.resolve() until Phase 5
- 9 new unit tests for resolveFreeCall
- All 163 unit tests pass, all 1199 integration resolver tests pass
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c5f2e73a-259a-438c-b5c8-286b82e3c215
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* chore: revert unrelated package-lock.json change
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c5f2e73a-259a-438c-b5c8-286b82e3c215
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix(SM-13): address PR #756 review findings on resolveFreeCall
Addresses all 7 findings from the PR #756 review comment.
Code (R1, finding #1)
- Replace the literal `'Class' | 'Struct' | 'Record'` check in
`hasClassTarget` with `INSTANTIABLE_CLASS_TYPES.has(c.type)`. Converts
an invariant that was previously comment-enforced ("keep this list
aligned with INSTANTIABLE_CLASS_TYPES") into one enforced structurally.
Any future extension of the set propagates here automatically. The
narrower Swift extension dedup block below still uses literal
`'Class' | 'Struct'` by design — Swift extensions only produce Class
duplicates in practice, Record is deliberately excluded there, and
the inline comment now documents that asymmetry.
Tests (+12 regression scenarios)
Finding #2 — language coverage
- Go free function (doStuff())
- Python free function (def helper(): ... helper())
- Rust free function outside any impl block
- Java statically-imported function
- JavaScript module-level function
Each exercises `_resolveCallTargetForTesting` with `callForm='free'`
and the language-specific file extension. `resolveFreeCall` has no
file-extension branching, so these guard the dispatch chain per
language without assuming extractor-specific symbol shapes.
Finding #3 — argCount threading
- 2-arg overload selected when argCount=2
- 0-arg overload selected when argCount=0
Finding #5 — Tier 3 (global) resolution
- Function globally visible but not imported. Asserts exact
`TIER_CONFIDENCE.global === 0.5` and `reason === 'global'` to catch
silent drift if the tier table is ever refactored.
Finding #6 — preComputedArgTypes worker path
- String overload matched via preComputedArgTypes=['String']
- Int overload matched via preComputedArgTypes=['int'] (lowercase,
mirroring the parse-worker's inferred-literal shape; stored 'Int' is
normalized via normalizeJvmTypeName at comparison time)
Finding #7 — Enum null-route documentation
- Enum-only free call asserts `toBeNull()` with an explanatory comment
linking to the INSTANTIABLE_CLASS_TYPES rationale. NOT marked skipped
— current behavior is intentional, not broken.
Finding #4 — Swift extension dedup guard
- Two same-name Class entries at different path lengths; exercises the
full dispatch chain:
1. filterCallableCandidates with 'free' strips Class → length 0
2. hasClassTarget triggers resolveStaticCall
3. Homonym ambiguity null-routes per SM-12 round-1 contract
4. Constructor-form retry repopulates with both Classes
5. Dedup block sorts by filePath.length → shortest path wins
Verification
- `tsc --noEmit` clean
- 3064 unit tests pass (+12)
- 1766 integration tests pass
- Zero regressions
Plan: docs/plans/2026-04-09-003-fix-sm13-resolve-free-call-review-findings-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/756#issuecomment-4213879002
* refactor(SM-13): extract dedupSwiftExtensionCandidates shared helper
Follow-up to the PR #756 review fix. SM-13 duplicated the Swift
extension same-name collision dedup block between `resolveCallTarget`
and `resolveFreeCall` — two copies of identical 15-line logic with the
same heuristic (`filePath.length` sort, Class/Struct-only, `length > 1`
guard). Extract a single shared helper so the two sites cannot drift.
Changes
- New `dedupSwiftExtensionCandidates(candidates, tier)` helper defined
alongside `tryOverloadDisambiguation`, with JSDoc documenting:
- The Swift extension scenario it addresses
- Why it is intentionally narrower than INSTANTIABLE_CLASS_TYPES
(Class/Struct only, not Record — C#/Kotlin records don't exhibit
the multi-file definition pattern, widening risks accidental
dedup of legitimately distinct record types)
- The return-null-on-no-match contract so callers can fall through
- `resolveCallTarget` tail dedup (was lines 1593-1610): replaced with
a single `dedupSwiftExtensionCandidates` call
- `resolveFreeCall` tail dedup (was lines 1994-2012): same replacement
- Net line count: -32 insertions, -9 deletions in the consumer sites,
+36 for the shared helper + JSDoc
Verification
- `tsc --noEmit` clean
- 3064 unit tests pass (including the R7 Swift dedup guard test added
in the previous commit that exercises the full free-form retry
chain through this helper)
- 1766 integration tests pass
- Zero regressions
Follows-up on: https://github.com/abhigyanpatwari/GitNexus/pull/756
* docs(SM-13): address PR #756 final review — comment cleanup only
Three documentation-only findings from the approval review. No
behavior change, no new tests, no code path modifications.
Finding #1 — stale line-number comment
- The comment inside `resolveFreeCall` at the `hasClassTarget` site
referenced "lines ~1994-2008" for the Swift extension dedup block.
Those lines were the inlined pre-SM-13 version; the block has since
been extracted to `dedupSwiftExtensionCandidates`. Replaced the line
reference with the helper name so future readers don't chase dead
line numbers.
Finding #2 — fuzzy-widening asymmetry undocumented
- `resolveFreeCall` intentionally has no `widenCache` parameter and no
D2 fuzzy-widening pass (unlike `resolveCallTarget`'s member-call
path). Added an explicit "Asymmetry vs `resolveCallTarget`" paragraph
to the JSDoc so a caller comparing the two signatures knows the
skipped pass is deliberate and tied to Phase 5.
Finding #3 — constructor-form retry reasons undocumented
- `resolveStaticCall` can return null for three distinct reasons
(empty instantiable pool, homonym ambiguity, ownerless Constructor
nodes). The retry below it unconditionally re-filters with
`'constructor'` form, which is correct for all three but not
obvious. Added a structured three-case comment enumerating each
reason and linking (a) to the SM-12 null-route contract, (b) to
the R7 dedup test, and (c) to the currently-uncovered ownerless-
Constructor path (noted as a future test candidate).
Verification
- `tsc --noEmit` clean
- 175 `resolveFreeCall` + `resolveStaticCall` + sibling tests pass
(sanity check — no behavior change expected)
- No regressions
Follows-up on: https://github.com/abhigyanpatwari/GitNexus/pull/756#issuecomment-4215739052
* test(SM-13): cover ownerless-Constructor retry + PHP free function
Two low-severity test gaps from PR #756 review comment 4215739052 —
previously addressed doc-only, now have concrete test coverage.
Finding #3 low — ownerless-Constructor retry path (previously comment-only)
- The retry after resolveStaticCall returns null handles three distinct
null-return reasons. Cases (a) and (b) were already tested (Interface/
Trait null-route from SM-12, Swift shadowing dedup from R7). Case (c) —
resolveStaticCall step-4 bailout when the tiered pool contains
ownerless Constructor nodes — was only covered by a comment.
- New test: Class + ownerless Constructor in tiered pool, callForm='free'.
Exercises the full chain:
1. resolveStaticCall step 3 walks classCandidates via
lookupMethodByOwner — ownerless Constructor not in methodByOwner,
nothing found.
2. Step 4 detects Constructor in tiered pool, bails with null.
3. resolveFreeCall retry re-runs filterCallableCandidates with
'constructor' form, which prefers Constructor over Class per
CONSTRUCTOR_TARGET_TYPES ordering.
4. Single survivor returned.
- Asserts the Constructor node (not the Class) is the resolved target.
Low — PHP free function coverage gap
- The language coverage table in the same review flagged PHP free
functions (top-level `function helper()` outside any class) as
uncovered. Added a test mirroring the existing Go/Python/Rust/Java/
JS language tests — exercises the `.php` dispatch path for free
calls. Ruby and C/C++ remain uncovered; deferred to a future round
since those languages also have other gaps in the broader test file.
Verification
- `tsc --noEmit` clean
- 3066 unit tests pass (+2 new regression tests)
- 1766 integration tests pass
- Zero regressions
Follows-up on: https://github.com/abhigyanpatwari/GitNexus/pull/756#issuecomment-4215739052
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* Initial plan
* feat(SM-12): extract resolveStaticCall from resolveCallTarget
- Add resolveStaticCall(className, methodName, currentFile, ctx, argCount?) using
lookupClassByName + lookupMethodByOwner for O(1) constructor/static resolution
- Add S0 fast path in resolveCallTarget for constructor/free-form class calls
- Export resolveStaticCall from call-processor.ts
- Add 11 unit tests covering constructor resolution, confidence tiers,
arity disambiguation, and resolveCallTarget delegation
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c9471ca9-57ff-4dae-956e-e7ffdc326bc4
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* chore: revert unrelated package-lock.json change
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c9471ca9-57ff-4dae-956e-e7ffdc326bc4
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* refactor: shorten verbose test name per code review feedback
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c9471ca9-57ff-4dae-956e-e7ffdc326bc4
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix(SM-12): address PR #754 review findings
Addresses Claude's review comments on PR #754:
Performance
- Pass pre-computed `tiered` result into `resolveStaticCall` as optional
`tieredOverride` parameter, eliminating the duplicate `ctx.resolve(className,
currentFile)` on every constructor call path.
- Cache `freeFormHasClassTarget` in `resolveCallTarget` so the S0 fast path
and the free-form constructor retry share a single `.some()` scan.
Architecture
- Reconcile `CLASS_LIKE_TYPES` (call-processor) with `CLASS_TYPES`
(symbol-table): `CLASS_LIKE_TYPES = [...CLASS_TYPES, 'Impl']`. This makes
the relationship explicit — the call resolver's set is a strict superset
of the heritage-index set, guaranteeing anything reachable via
`lookupClassByName` also passes the resolver filter. Trait is now included
(harmless: traits have no Constructor nodes, so step-3 returns undefined
and step-5 still returns the class-like node when unique). Documented
the Interface inclusion rationale (static methods + MRO walker).
- Collapse `resolveStaticCall`'s `methodName` parameter into `className` —
all call sites passed identical values. Named constructors (Dart
`User.fromJson()`) arrive as member calls and go through
`resolveMemberCall`. Documented the reserved path for when a language
surfaces a static-method-shaped call with a distinct member name.
- Document the known gap: `callForm === 'member'` constructor patterns
(e.g. Python `models.User()`) are handled by the tail fallback, not S0.
Tests
- Add tiered-override test asserting `ctx.resolve` is not re-invoked when
a pre-computed result is passed in.
- Add language-specific `_resolveCallTargetForTesting` integration tests
for Java (`new User()`), Python (`User()`), and Kotlin (`User()`).
Verification: 3031 unit + 1766 integration tests pass, zero regressions.
* fix(SM-12): restrict resolveStaticCall fallback to instantiable kinds
Addresses the high-severity finding from the Codex adversarial review of
PR #754: `resolveStaticCall`'s step-5 "return the class itself when no
Constructor node is found" fallback reused `CLASS_LIKE_TYPES`, which —
after SM-11 and PR #754's reconciliation — now includes `Interface`,
`Trait`, and `Impl`. That is the method-dispatch set, not the
instantiable set, so constructor-shaped calls could resolve to
non-instantiable nodes and emit false `CALLS` edges.
Concrete failure: Rust same-file `impl User { ... }` alongside
`struct User { ... }` — both land at same-file tier, the Impl is not
filtered out, and the step-5 fallback produces a `CALLS` edge to the
`Impl` block instead of the `Struct`. The same widening exposed
Interface / Trait targets in Java / C# / PHP / Scala.
Fix
- Introduce `INSTANTIABLE_CLASS_TYPES = {'Class', 'Struct', 'Record'}`
as a sibling to `CLASS_LIKE_TYPES`, documenting the contract
explicitly and cross-referencing `CONSTRUCTOR_TARGET_TYPES`.
- Update `CLASS_LIKE_TYPES` JSDoc to clarify it is the method-dispatch
set and add an anti-pattern warning against reusing it for
constructor-fallback filtering.
- Tighten `resolveStaticCall` step 5: filter `classCandidates` through
`INSTANTIABLE_CLASS_TYPES` before the `length === 1` check. This
strips `Impl` from the Rust shadowing scenario (leaving `Struct` as
the sole instantiable target) and null-routes Interface / Trait /
`Impl`-alone scenarios, matching the SM-10 R3 null-route precedent.
- Step 3 (explicit Constructor lookup via `lookupMethodByOwner`) is
intentionally unchanged — its `def.type === 'Constructor'` check is
the correct contract, and legitimate Constructor nodes attached to
`Impl` owners still resolve correctly.
Tests (+10 regression scenarios)
- Positive guards: Struct, Record fallback paths.
- Null-route: Interface (Java/C#/TS), PHP Trait, Rust Trait.
- Rust same-file shadowing: Struct wins over Impl.
- Rust Impl-alone: null-routes (no Struct present).
- Step-3 preservation: Constructor owned by Impl still resolves to the
Constructor node, proving step-5 tightening doesn't leak into step 3.
- Full cascade via `_resolveCallTargetForTesting` for Interface and
Trait — confirms no downstream path silently re-introduces the edge.
Verification
- `tsc --noEmit` clean
- 3041 unit tests pass (+10)
- 1766 integration tests pass
- Zero regressions
Plan: docs/plans/2026-04-09-002-fix-sm12-constructor-fallback-instantiable-only-plan.md
Codex review job: review-mnrao7fr-nv9y0e
* fix(SM-12): address PR #754 second review round
Addresses the 9 findings from the follow-up review on PR #754.
Performance
- Align `freeFormHasClassTarget` with `INSTANTIABLE_CLASS_TYPES`: drop
`Enum` (S0 would always return null for it — wasted lookup work) and
add `Record` (C# records and Kotlin data classes were bypassing S0
entirely). The trigger set and the fallback filter set now agree by
construction, documented inline.
Documentation
- Remove stale single-line JSDoc on `CLASS_LIKE_TYPES` (line 57) that
duplicated the full multi-line block immediately below it — tooling
picks up the first block so the old one-liner was shadowing the
current explanation.
- Rewrite the `resolveStaticCall` JSDoc step list to match the actual
step boundaries in the implementation (steps 3, 4, 5 were blurred in
the old description).
- Add inline comment on step 3 documenting the same-name lookup
assumption (`${candidate.nodeId}\0${className}`) and the symmetric
miss case for Python `__init__`-style constructors.
- Add inline comment on step 4 documenting that it also catches the
ambiguous-step-3 case, and warning against removing the check
without handling that path explicitly.
- Add inline comment on step 5 enumerating the three length outcomes
(0 / 1 / >1) so future readers see the dominant null-route case.
- Document Ruby `User.new` as a known gap alongside Python
`models.User()` in the S0 header comment.
Tests (+2 scenarios)
- Record free-form constructor call via `_resolveCallTargetForTesting`
exercises the aligned `freeFormHasClassTarget` trigger end-to-end,
closing the gap where the direct `resolveStaticCall` test passed
but the integration path was silently bypassing S0.
- Arity threading via `_resolveCallTargetForTesting` asserts that
`call.argCount` flows through resolveCallTarget → S0 →
resolveStaticCall → lookupMethodByOwner, catching any future
regression where the argCount is dropped at the S0 call site.
Verification
- `tsc --noEmit` clean
- 3043 unit tests pass (+2)
- 1766 integration tests pass
- Zero regressions
Plan: docs/plans/2026-04-09-002-fix-sm12-constructor-fallback-instantiable-only-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/754#issuecomment-4213536094
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* Initial plan
* feat(SM-11): extract resolveMemberCall from resolveCallTarget
- Create resolveMemberCall(ownerType, methodName, currentFile, ctx, heritageMap?)
that uses owner-scoped + MRO resolution only (no fuzzy lookup)
- resolveCallTarget delegates member calls (D0 path) to resolveMemberCall
- walkMixedChain uses resolveMemberCall for owner-scoped member-call resolution
- Add 7 unit tests for resolveMemberCall covering direct, inherited, MRO,
null cases, and confidence tier assertions
- Export resolveMemberCall for external use
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3b7889a9-5f2f-4572-8904-45084210f10d
* fix(SM-11): address PR #744 review
Blocking fixes:
- B1: Revert unrelated package-lock.json gitnexus-shared addition
- B2: Document confidence-tier semantic change on resolveMemberCall
Performance / coupling fixes:
- S1: walkMixedChain now calls resolveMethodByOwner directly (hot path) to avoid throwaway ResolveResult allocation per chain step
- S2: Thread tier from resolveMethodByOwner via { def, tier } tuple; eliminates double ctx.resolve
Alignment with semantic-model plan (Phase 3 target):
- resolveMethodByOwner now iterates ALL class-like candidates from ctx.resolve, deduplicating matches by nodeId. Absorbs D4's ownerId-filtering into the owner-scoped path.
- Handles homonym classes (two Users in different files) without falling through to D1-D4 fuzzy widening
- Shared-ancestor MRO walks automatically dedup (both homonyms walk to same base method)
- Unified direct-vs-MRO lookup under a single canWalkMRO check
Tests added:
- T1: Three D0 skip-condition tests via new _resolveCallTargetForTesting internal export (overloadHints, preComputedArgTypes, hasActiveModuleAlias)
- T2: Rust qualified-syntax null test (trait-inherited method) + direct impl control
- T3: C++ leftmost-base diamond inheritance test
- B2 lock-in: cross-file class tier assertion
- Homonym disambiguation: only-one-owns-method, both-own-method ambiguity, shared-ancestor MRO convergence
Verification:
- tsc --noEmit: clean
- vitest run test/unit/: 3014 passed
- vitest run test/integration/resolvers/: 1746 passed
* test(SM-11): address second PR #744 review round + per-language integration tests
Review fixes (https://github.com/abhigyanpatwari/GitNexus/pull/744#issuecomment-4211877593):
P1 (Performance): Replace Map allocation in resolveMethodByOwner with a firstDef+ambiguous flag pattern. Zero allocation for the common single-candidate case on the hot path — the previous Map approach allocated on every member call regardless of whether deduplication was needed.
P2 (Test gap): Strengthen the module-alias D0 skip test with a homonym fixture (two Users in different files). Previously the test passed whether or not D0 was actually bypassed; the new version proves D0 must be skipped by showing that resolveMemberCall directly returns null (ambiguous) but D1-D4 with alias narrowing picks the right one. Also fixes the underlying D2-vs-alias widening interaction: when filteredCandidates was narrowed by module-alias disambiguation, D2 no longer widens back to the full fuzzy pool (introduces aliasNarrowed boolean flag).
L1 (Language coverage): Add C# and Kotlin implements-split tests at the resolveMemberCall layer.
L2 (Maintainability): Export OverloadHints as @internal so the test can use a direct cast instead of fragile Parameters<...> type inference.
Per-language integration tests:
- rust-child-extends-parent: Direct impl method resolution via D0 (with honest documentation of the trait-method-as-Function gap that is Phase 5 / SM-16 scope)
- java-interface-default-method: User implements Validator with default method resolved via implements-split MRO
- csharp-interface-default-method: Same pattern for C# 8.0+ default interface methods
- kotlin-interface-default-method: Same pattern for Kotlin interfaces with default implementations
- python-multi-level-mro: 3-level C3 linearization (Grandparent ← Parent ← Child)
- cpp-diamond-inheritance: Classic diamond (Base ← A, B ← Derived) via leftmost-base MRO
Verification:
- tsc --noEmit: clean
- vitest run test/unit/: 3015 passed
- vitest run test/integration/resolvers/: 1763 passed (+17 new per-language tests)
* fix(SM-11): Codex adversarial review corrections + deeper D0 fixes
Addresses the three high-severity findings from the Codex adversarial review of PR #744 (https://github.com/abhigyanpatwari/GitNexus/pull/744#issuecomment-4212075120), plus four deeper fixes discovered during regression triage. All discovered issues are now addressed end-to-end rather than papered over with tail-return fallbacks.
Codex review findings:
R1 (C++ diamond): The cpp-diamond-inheritance fixture used non-virtual inheritance, which is genuinely ambiguous in real C++ (two Base subobjects). Changed A and B to use 'virtual public Base' so there's a single shared Base subobject and d.method() is an unambiguous call that the leftmost-base MRO walk correctly resolves.
R2 (C# default-interface): The csharp-interface-default-method fixture called user.Validate() via a User-typed variable, but C# does not inherit default interface methods as callable class members — the call is only valid through an interface-typed variable. Changed App.cs to 'IValidator user = new User(...)' which is the idiomatic dispatch pattern.
R3 (resolveCallTarget tail-return): When D1-D4 receiver filtering produced zero file-matched and zero owner-matched candidates for a member call, the function fell through to the permissive single-candidate tail return — silently emitting CALLS edges for methods that don't belong to the receiver. Added an explicit null-route inside the D1-D4 block that fires only when both filters yielded 0.
R4 (Rust negative assertion): Added the c.trait_only() negative integration test in rust.test.ts demonstrating that direct member calls on Rust structs do not walk trait ancestry. The test now passes because of R3 (previously fell through to the tail return).
Regression triage discoveries:
1. D0 was dead code on the sequential pipeline. The sequential path sets overloadHints for every call regardless of whether the method is overloaded, and the original D0 skip condition '!overloadHints && !preComputedArgTypes' was therefore always false. The Java/C#/C++ SM-9/SM-10 inheritance tests were passing ONLY via the tail-return fallback. Fix: narrow the skip to 'overloadHints && filteredCandidates.length > 1' — skip D0 only when there are actually multiple candidates that need overload disambiguation.
2. lookupMethodByOwner couldn't disambiguate arity-differing overloads (e.g. C++ greet() vs greet(string)). With D0 now firing on the sequential path, same-name/different-arity overloads would collapse to an arbitrary first pick. Fix: added an optional argCount parameter to lookupMethodByOwner + lookupMethodByOwnerWithMRO that filters the overload set by parameterCount/requiredParameterCount before the returnType dedup.
3. Python and Rust class methods are captured as Function nodes (not Method) with ownerId set to the class. The methodByOwner index only accepted 'Method' and 'Constructor' types, so Python class methods and Rust trait methods were invisible to D0. Fix: extended the methodByOwner indexing condition to include 'Function' when ownerId is set. This also unlocks the Rust trait-method negative assertion by ensuring the qualified-syntax MRO strategy has something to return null for.
4. D0 was being skipped when a local variable shadowed an imported module name (Python 'from models.c import C; c = C()' creates both a module alias 'c → models/c.py' AND a typed local 'c'). Fix: the D0 skip now gates on 'aliasNarrowed' (a new boolean tracking whether the alias block actually narrowed filteredCandidates) instead of 'hasActiveModuleAlias'. If the method isn't in the aliased module, the receiver is a typed local variable and D0 should run.
5. PHP trait walk missed the HasTimestamps trait because lookupClassByName did not include 'Trait' type. buildHeritageMap uses lookupClassByName to resolve parent names, so 'BaseModel use HasTimestamps' was failing to register an ancestor edge for BaseModel → HasTimestamps. Fix: added 'Trait' to CLASS_TYPES. The trait is now a valid class-like type for heritage resolution (PHP use, Rust impl Trait for Struct, Scala traits).
Test updates:
- Updated the 'no heritageMap' unit test in call-processor.test.ts to assert the correct null-route behavior instead of the old tail-return fallback.
- Added a new unit test asserting Trait inclusion in the class set.
- Updated the 'does NOT include other type-like labels' test to remove Trait from its rejection set.
Verification:
- tsc --noEmit: clean
- vitest run test/unit/: 3016 passed (+1 new Trait inclusion test)
- vitest run test/integration/resolvers/: 1764 passed (+1 new Rust negative assertion)
- Zero regressions
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergo Magyar <magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* Initial plan
* Add MRO fast path before D2 fuzzy widening in resolveCallTarget
When receiverTypeName is known, try resolveMethodByOwner (owner-scoped
+ MRO lookup) before falling back to the expensive lookupFuzzy in D2.
This short-circuits cross-file member call resolution for the common
non-overloaded case.
The fast path is skipped when overload disambiguation hints are
available (overloadHints or preComputedArgTypes) to avoid picking the
wrong overload for same-return-type overloaded methods.
Passes heritageMap to resolveCallTarget from all 4 call sites:
- Language seed path (processCalls)
- Sequential path (processCalls)
- walkMixedChain fallback
- Worker path (processCallsFromExtracted)
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9e49521f-2472-47bc-96e9-be4a46b073f0
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix(SM-10): address PR #741 review
Correctness:
- Module-alias guard for D0. When call.receiverName matches an active
entry in ctx.moduleAliasMap for the current file, D0 is now skipped
and resolution falls through to D1-D4 which respects the
alias-narrowed candidate pool. Prevents a homonymous class in a
different file from being picked by ctx.resolve(receiverTypeName)
inside resolveMethodByOwner. New unit test pins the contract.
Unit tests (call-processor.test.ts — 3 new):
- D0 hit: child.parentMethod() resolves via MRO walk when
heritageMap is provided.
- D0 skipped: same scenario still resolves via D1-D4 when heritageMap
is undefined (backward-compat guard).
- Module-alias guard: two files both define class User with a save()
method; 'import auth_mod as auth' in app.py must resolve
auth.user.save() to auth_mod.py, not user_mod.py.
Integration language coverage (+3 fixtures/tests):
- swift-child-extends-parent — first-wins, gated on swiftAvailable.
- ruby-child-extends-parent — first-wins.
- php-child-extends-parent — first-wins (uses ParentClass since
'Parent' is a PHP reserved word).
* test(SM-10): address second PR #741 review round
Unit tests (call-processor.test.ts, +2 new):
- overloadHints guard: Java source with two same-return-type overloads
method(int) and method(String), int added first so lookupMethodByOwner
would return it. processCalls auto-generates overloadHints for Java,
forcing D0 to be skipped. o.method("hello") must resolve to
method(String) via literal-inferred disambiguation.
- preComputedArgTypes guard: worker-path equivalent via
processCallsFromExtracted with ExtractedCall.argTypes=['String'].
Same two overloads, same correctness guarantee.
Integration tests (+2 fixtures + test blocks):
- go-child-extends-parent — struct embedding, first-wins
(Go structs are labeled 'Struct' not 'Class' in GitNexus).
- dart-child-extends-parent — extends, first-wins, gated on
dartAvailable like other Dart tests.
Documentation:
- Expanded the fallthrough comment in resolveMethodByOwner to clarify
that unknown-extension paths land on plain lookupMethodByOwner
without an ancestor walk, and that D1-D4 still runs on D0 miss.
* test(SM-10): D0 miss with heritageMap present falls through to D1-D4
Closes the last remaining gap from PR #741 review round 3. The existing
'D0 skipped' test only covered the heritageMap=undefined case, leaving
the miss-with-heritageMap path implicitly covered by integration tests
only. This adds a focused unit test where:
- Class Obj has a method doWork findable via tiered resolution
(import-scoped) but intentionally NOT registered in methodByOwner
(no ownerId), so lookupMethodByOwner misses.
- heritageMap is provided but built from an empty heritage array, so
getAncestors(class:Obj) returns []. The MRO walk yields no parents.
- lookupMethodByOwnerWithMRO therefore returns undefined → D0 miss.
- D1 resolves the receiver type; D2 widens via lookupFuzzy;
D3 file-filter picks the single matching candidate.
- A CALLS edge must still be emitted — D0 miss must not swallow
the call.
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* Initial plan
* feat(SM-9): add lookupMethodByOwnerWithMRO with HeritageMap parent chain walking
- Export c3Linearize from mro-processor.ts for reuse
- Add lookupMethodByOwnerWithMRO in call-processor.ts with MRO strategy support
- Update resolveMethodByOwner to fall back to MRO walk when HeritageMap available
- Thread heritageMap through walkMixedChain for chain resolution
- Add 10 unit tests covering all acceptance criteria
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc58249b-42f1-45a9-89fb-e3917e4d0171
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* feat(SM-9): add Java integration test with class Child extends Parent fixture
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc58249b-42f1-45a9-89fb-e3917e4d0171
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* docs: address code review comments on MRO strategy documentation
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc58249b-42f1-45a9-89fb-e3917e4d0171
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* perf(SM-9): address PR #740 review comments
- Eliminate double direct lookup in resolveMethodByOwner: delegate
straight to lookupMethodByOwnerWithMRO when a HeritageMap is
available (the MRO helper already does the direct lookup before
walking ancestors). Fallback path handles the no-HeritageMap case.
- Memoize C3 linearization per HeritageMap via a WeakMap keyed cache.
HeritageMap is immutable after build, so C3 results are stable for
its lifetime; WeakMap lets the cache auto-drain when the HeritageMap
is GC'd. Null sentinel caches linearization failures so cyclic
hierarchies are not reprocessed. Eliminates per-call buildParentMap +
c3Linearize on Python codebases.
- ancestors variable typed as readonly to accept the cached result
without copying.
- Add four missing MRO unit tests: Kotlin implements-split, C#
implements-split, JavaScript first-wins (separate provider from TS),
and C++ leftmost-base diamond (first diamond test for C++).
* fix(SM-9): CI prettier + address PR #740 follow-up review
- Fix CI prettier failure in test/integration/resolvers/java.test.ts
(auto-formatted — was introduced in 37563a31 before my first fix
commit but had not been caught locally).
- Pin caller on the SM-9 Java integration test (parentMethodCall.source
=== 'run') so a regression that misattributes the CALLS edge fails.
- Add two implements-split unit tests:
* Ambiguous default from two interfaces → BFS first-wins. Pins the
contract that lookupMethodByOwnerWithMRO returns a defined result
(full ambiguity detection is deferred to computeMRO graph pass).
* Class method precedence over interface default: Child extends Base
implements IFoo where both define handle() — documents that BFS
visits the extends edge first, matching Java's class-wins rule.
- Add @internal JSDoc on lookupMethodByOwnerWithMRO clarifying it is
exported only for testing; resolveMethodByOwner is the proper entry
point for callers.
* test(SM-9): per-language integration fixtures and tests for inherited method resolution
Extends the SM-9 integration coverage beyond Java with six new
child-extends-parent fixtures, one per MRO strategy:
- python-child-extends-parent → C3 strategy
- typescript-child-extends-parent → first-wins
- javascript-child-extends-parent → first-wins (separate provider)
- kotlin-child-extends-parent → implements-split
- csharp-child-extends-parent → implements-split
- cpp-child-extends-parent → leftmost-base
Each fixture follows the java-child-extends-parent pattern:
- Parent class with a single method
- Child class extending Parent, no override
- App class/function that instantiates Child and calls
the parent method — exercises the full ingestion pipeline,
HeritageMap construction, and lookupMethodByOwnerWithMRO walk.
For every fixture the matching integration test asserts:
- Parent and Child classes are detected
- Child → Parent EXTENDS edge is emitted
- The parent-method call resolves to the correct target file
- The caller is pinned (source === 'run' / 'Run') to catch
edge misattribution regressions
Rust is intentionally omitted — its qualified-syntax strategy
returns undefined from lookupMethodByOwnerWithMRO by design, so
there is no inherited-method resolution to assert against.
All 1739 integration resolver tests pass (+18 new SM-9 tests
across 6 languages).
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* Initial plan
* feat(SM-8): add HeritageMap with MRO-aware parent/ancestor lookup
- New heritage-map.ts: HeritageMap interface with getParents() and getAncestors()
- buildHeritageMap() consumes ExtractedHeritage[], resolves names via lookupClassByName
- Cycle protection and bounded depth (MAX_ANCESTOR_DEPTH=32) in getAncestors
- Worker path: HeritageMap built from deferredWorkerHeritage, threaded into processCallsFromExtracted
- Sequential path: Heritage accumulated across chunks, HeritageMap built after all chunks, passed to processCalls
- 18 unit tests covering parent lookup, multi-level, diamond, cycles, missing parent, bounded depth
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c413e0a3-5d63-4ddb-8ece-02fe6ed99efd
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* test: rename cycle test for clarity per code review
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c413e0a3-5d63-4ddb-8ece-02fe6ed99efd
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* refactor(SM-8): merge implementor map into heritage map
- Add `getImplementorFiles(interfaceName)` to HeritageMap interface
- Build implementor index (interface name → file paths) alongside parent
lookup in `buildHeritageMap`, using same `resolveExtendsType` logic
- Remove `ImplementorMap` type, `buildImplementorMap`, `mergeImplementorMaps`
from call-processor.ts
- Update `findInterfaceDispatchTargets`, `processCalls`, and
`processCallsFromExtracted` to use HeritageMap for both parent
lookup and implementor dispatch
- Pipeline: single `buildHeritageMap` call replaces separate
buildImplementorMap + buildHeritageMap for both worker and
sequential paths
- Migrate implementor tests from call-processor.test.ts to
heritage-map.test.ts (4 new getImplementorFiles tests)
- Update interface dispatch test to use buildHeritageMap instead
of hand-constructed ImplementorMap
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/085dffb4-b31e-4aa5-9aa3-4314bc0010e7
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* test: rename implementor test for clarity per code review
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/085dffb4-b31e-4aa5-9aa3-4314bc0010e7
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
* fix(SM-8): address PR #739 review comments
- pipeline.ts: cache chunk file contents from Pass 1 to eliminate
double-read of sequential chunks in Pass 2. Peak memory drains
incrementally as Pass 2 processes each chunk.
- heritage-map.ts: document Rust trait-impl omission from implementor
index and the interface-name collision limitation.
- heritage-map.test.ts: add six tests covering the extends->IMPLEMENTS
path across C# (interfaceNamePattern), Swift (heritageDefaultEdge),
Java (symbol-table Interface lookup), Kotlin, PHP, and the Rust
trait-impl omission.
- pipeline.ts: comment why the heritage accumulation uses a manual
push loop instead of spread (ref #650).
* test(SM-8): address second PR #739 review pass
- Add TypeScript implements test to getImplementorFiles (closes
the .ts coverage gap flagged by the bot reviewer).
- Tighten deep-chain boundary assertion from toBeLessThanOrEqual(32)
to toBe(32) so a future regression returning fewer ancestors
fails loudly. Added an ancestors[31] === 'class:Level32' check
to pin the upper boundary.
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* refactor(call-processor): use class lookup index in phase p
* test(call-processor): cover class lookup fallback
---------
Co-authored-by: 许恩宁 <xuenning@qiyi.com>
* refactor(type-env): use class lookup index for type resolution
* test(type-env): add lookupClassByName regression coverage
* test(type-env): expand class lookup regression coverage
---------
Co-authored-by: 许恩宁 <xuenning@qiyi.com>
Add eagerly-populated methodByOwner index to SymbolTable, keyed by
ownerNodeId\0methodName. Used by walkMixedChain as a fast path for
resolving intermediate method calls in cross-class chains like
user.getAddress().getCity().getZipCode(), avoiding expensive fuzzy
lookups when the owner type is already known.
Handles overloaded methods: returns the first match when all overloads
share the same returnType, undefined when return types differ (ambiguous).
- Add lookupMethodByOwner to SymbolTable interface + implementation
- Add resolveMethodByOwner helper in call-processor.ts
- Add fast path in walkMixedChain before resolveCallTarget fallback
- Add Java cross-class chain fixture + 6 integration tests
- Add 148 unit tests for methodByOwner index behavior
* fix(ignore): respect negation patterns in .gitnexusignore childrenIgnored
childrenIgnored checked `ig.ignores(rel) || ig.ignores(rel + '/')` which
short-circuited on the bare path — directory-only negation patterns like
`!iOS/` were missed because `ig.ignores('iOS')` treats the path as a file.
Now only checks with trailing slash since childrenIgnored is only called
for directories. Bare-name patterns (e.g. `local`) still match per gitignore spec.
Fixes#596
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* test(ignore): add edge-case for bare `!dir` negation pattern
Verifies that `!iOS` (without trailing slash) also un-ignores the iOS/
directory — confirms the `ignore` package normalizes both `!dir` and
`!dir/` forms consistently when tested with a trailing-slash path.
Addresses non-blocking review suggestion on #654.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* docs(ignore): link ignore package docs for bare-name normalization
Adds references to the `ignore` package documentation in both the
childrenIgnored comment and the bare-negation test, explaining why
`!iOS` (without trailing slash) also re-includes the iOS/ directory.
Addresses non-blocking review suggestion on #654.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: replace Array.push(...spread) with loop to prevent stack overflow
On large codebases (78K+ C files), deferred arrays in
runChunkedParseAndResolve accumulate 100K+ entries. The spread
operator in push(...array) puts every element on the call stack
as a function argument, exceeding the maximum call stack size.
Replace all 11 occurrences of `arr.push(...other)` with
`for (const _item of other) arr.push(_item)` which uses
constant stack space regardless of array size.
Fixes#649
* fix: also replace push(...spread) in parsing-processor.ts
* style: format pipeline.ts to match prettier config
---------
Co-authored-by: tian kian tan <tan@example.com>
* feat: same-arity overload disambiguation via type-hash suffix (#651)
Add ~type1,type2 suffix to Method/Constructor node IDs when same-arity
overloads with different parameter types exist in the same class. Also add
$const suffix for C++ const-qualified method overloads via new isConst field.
Key changes:
- typeTagForId() detects same-arity collisions and appends ~typeTag
- constTagForId() detects const/non-const collisions and appends $const
- TS/JS excluded from type-hashing (overload signatures collapse to impl body)
- Sequential findEnclosingFunction fixed: falls through on ambiguous same-class
candidates instead of picking first; fallback path includes typeTag + constTag
- Per-call-site integration tests across Java, C#, Kotlin, C++, TypeScript
- Cross-file + chain resolution tests for all 5 languages
- C++ isConst extraction via tree-sitter type_qualifier in function_declarator
1710 integration + 18 unit tests pass.
* fix: preserve generic/template args in type-hash, perf + type safety fixes
- Add rawType field to ParameterInfo preserving full type text (vector<int>)
while type stays simplified (vector). typeTagForId uses rawType for tags.
- Populate rawType in all 11 language method extractors
- Add buildCollisionGroups() to pre-group methods by name#arity (O(N) once
per class instead of O(N) per method call)
- Cache method extraction in call-processor findEnclosingFunction fallback
- Fix null guards on getLanguageFromFilename in all findEnclosing paths
- Tighten SKIP_TYPE_HASH_LANGUAGES to ReadonlySet<SupportedLanguages>
- Document ID stability invariant on first overload introduction
- C++ integration tests: template overloads (vector<int> vs vector<string>),
cross-file template + chain resolution, out-of-class method definitions
1718 integration + 20 unit tests pass.
* fix: add rawType to method-extraction unit test assertions
All 26 parameter .toEqual() assertions in method-extraction.test.ts
needed the new rawType field added to match ParameterInfo schema change.
* perf: cache tempMap/groups per class, consolidate extractFromNode
- Cache derived method map + collision groups per classNode.id in
parsing-processor (avoids rebuild per method in same class)
- Replace per-call extractFromNode with cached class extraction +
funcName:line lookup in call-processor fallback (avoids AST walk
per call site)
- Remove dead clearEnclosingFunctionCache export, fix JSDoc
* test: add sequential-path integration test for same-arity overloads
Add skipWorkers option to PipelineOptions to force sequential parsing.
New test suite verifies type-hash disambiguation produces identical
results through the sequential path (parsing-processor + call-processor
findEnclosingFunction) as the worker path.
* feat: MethodExtractor configs for Python, PHP, Swift, Dart, Rust, Ruby with exhaustive integration tests
Add per-language MethodExtractionConfig for all remaining tree-sitter languages
(RFC #568 PR 2). Each config follows the established createMethodExtractor()
factory pattern — no new types, no parse-worker changes.
Configs:
- Python: @abstractmethod, @staticmethod/@classmethod, *args/**kwargs, type hints, _/__ visibility
- PHP: abstract/final/static keywords, PHP 8 #[] attributes, __construct/__destruct
- Swift: 5-level visibility, protocol-as-abstract, static/class methods, @ attributes
- Dart: _ convention visibility, abstract (no body), method_signature unwrapping
- Rust: pub visibility, &self receiver, trait_item + impl_item, #[] attributes
- Ruby: positional visibility via sibling-walk, singleton_method as static
Integration fixtures (18 directories) covering 3 resolution patterns:
- Method enrichment: parameterTypes, isAbstract, isFinal, annotations on graph nodes
- Overload dispatch: arity-based CALLS resolution via parameterTypes
- Abstract dispatch: abstract/concrete method distinction (Python, PHP, Rust, Swift)
Go deferred — requires factory changes for receiver-based method extraction.
Closes#571
* fix: address code review findings across 6 MethodExtractor configs
Fix all actionable items from the PR #624 deep-dive review:
Dart (critical — fixes 6 CI failures):
- isDartStatic: check children first, siblings as fallback
- isDartAbstract: handle declaration nodes for abstract methods
- extractSingleParam: detect required keyword as sibling token
- Add declaration to methodNodeTypes, mixin_declaration to typeDeclarationNodes
- Add member call query for variable assignments in tree-sitter-queries
Python:
- hasDecorator now matches dotted paths (e.g. @abc.abstractmethod)
- Fix version comment from ^0.23.6 to 0.23.4
PHP:
- Add enum_declaration to typeDeclarationNodes (PHP 8.1+)
- Add version comment for 0.23.12
Swift:
- Add isOverride using hasKeyword/hasModifier pattern
Rust:
- Fix version comment from ^0.23.2 to 0.23.1
Also: identifier fallback in generic.ts for mixin owner names,
Dart integration test label fix (Method vs Function), version
comment for tree-sitter-dart 1.0.0.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: Dart extension_declaration and Ruby module_function support
Dart:
- Add extension_declaration to typeDeclarationNodes and extension_body
to bodyNodeTypes — extension methods are now extracted into the graph
- Add extension_declaration and mixin_declaration to CLASS_CONTAINER_TYPES
for HAS_METHOD edge resolution
Ruby:
- module_function now maps to visibility 'private' in extractRubyVisibility
- module_function methods marked isStatic via backward-walk in isStatic
- Override semantics: private/public after module_function resets isStatic
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(go): Go MethodExtractor config with receiver-based extraction
Add Go as the 13th language with a per-language MethodExtractor config.
Go methods are top-level (not nested in struct bodies), so this adds
extractFromNode() to the MethodExtractor interface for direct method
node extraction without an enclosing class.
Config extracts:
- Name from field_identifier (methods) / identifier (functions)
- Return type including multi-return (first type from parameter_list)
- Parameters with variadic support
- Visibility via uppercase/lowercase convention
- Receiver type with pointer unwrapping (*User → User)
- isStatic for functions (no receiver)
Infrastructure:
- extractOwnerName optional hook on MethodExtractionConfig
- extractFromNode on MethodExtractor (factory auto-implements)
- Parse-worker uses extractFromNode when no enclosing class found
- method_declaration added to CLASS_CONTAINER_TYPES
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* test: method enrichment integration tests for 7 languages + TS abstract class fix
Add method-enrichment integration test fixtures and test blocks for
Go, C++, Java, Kotlin, TypeScript, JavaScript, and C#. Each fixture
tests: class detection, HAS_METHOD edges, EXTENDS edges, isAbstract,
isStatic, annotations, parameterTypes, and CALLS edge resolution.
Fixes found during testing:
- Remove method_declaration from CLASS_CONTAINER_TYPES (added for Go
but broke Java/C# HAS_METHOD edge resolution — method_declaration
is also Java's method node type)
- Add abstract_class_declaration query to TypeScript tree-sitter
queries (was missing, so abstract classes were invisible to pipeline)
1699 integration tests pass across 20 test files, 0 regressions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* style: format typeDeclarationNodes array for better readability in PHP config
* fix: Go interface methods + Rust impl-for-Struct owner resolution
Go:
- Add method_elem to methodNodeTypes so interface method signatures
are extractable as abstract methods
- Integration test: Animal interface detected, Speak isAbstract,
CALLS edges from app.go
Rust:
- Add extractOwnerName to resolve impl Trait for Struct to the
concrete Struct (not the Trait) — fixes method misattribution
- Fix findEnclosingClassId to generate Struct: label (not Impl:)
for impl blocks so HAS_METHOD edges resolve to struct nodes
- Tighten abstract-dispatch test: assert SqlRepo owns find/save
generic.ts:
- Fix extractOwnerName fallback: when hook returns a value, skip
both name-field and type_identifier scan (was overwriting result)
1703 integration tests pass, 0 regressions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: code review response — Rust impl label, Swift params, Dart async, sequential methodExtractor
Address code review findings from PR #624:
- ast-helpers: Rust `impl Trait for Struct` uses Struct label (matches existing
graph node), plain `impl Struct` uses Impl label (matches definition.impl)
- swift: fix parameter type extraction (user_type not type_annotation), detect
default values as function_declaration siblings, add version comment
- dart: isDartAsync now detects async*/sync* generators, add clarifying comment
for declaration nodes in extension bodies
- python: correct isFinal comment (PEP 591 @typing.final exists, just not modeled)
- parsing-processor: port methodExtractor enrichment to sequential path so
isAbstract/isStatic/visibility/annotations/isFinal populate on <15-file repos
- tests: remove silent `if (prop !== undefined)` guards, assert properties
directly, fix label queries (Dart Method vs Function, Swift Method for protocol
methods), add Rust HAS_METHOD sourceLabel tests, Swift parameterTypes tests,
and Dart async/sync* integration tests with fixture
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: Rust grammar gap + qualified method IDs to resolve same-file collisions
Phase 1 — Rust grammar:
- Add function_signature_item query to RUST_QUERIES so abstract trait methods
(fn speak(&self) -> String;) become graph nodes with isAbstract=true
Phase 2 — Qualified method IDs:
- findEnclosingClassInfo returns {classId, className} for AST-based class lookup
- Both parsing paths (sequential + worker) qualify method/property IDs with
enclosing class: Method:file:ClassName.method instead of Method:file:method
- extractFuncNameFromSourceId handles ClassName.method format
- Fixes silent data loss when same-name methods in different classes shared a
file (e.g., Animal.speak and Dog.speak both now exist as distinct graph nodes)
Test updates:
- Rust: abstract+concrete trait methods both verified, function count adjusted
- Python: static method disambiguation now emits 2 CALLS edges (correct — no
more ID collision masking the second call)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: owner-aware resolution for qualified method IDs
Address Codex adversarial review findings after qualified ID change:
- findEnclosingFunction: disambiguate candidates by ownerId when multiple
same-name methods exist in file; qualify fallback-generated IDs
- findEnclosingFunctionId (worker): qualify sourceIds with enclosing class
name so CALLS source attribution matches definition-phase node IDs
- buildExportedTypeMapFromGraph: use lookupExactAll + nodeId match instead
of lookupExactFull which returns first definition for bare name
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: methodExtractor variadic arity, return type preservation, PHP abstract dispatch
Three bugs in the methodExtractor enrichment path broke 17 integration tests:
1. Variadic parameterCount: buildMethodProps and parse-worker set
parameterCount = info.parameters.length even for variadic functions,
causing arity filtering to reject valid calls. Now checks isVariadic
and sets parameterCount = undefined (matching extractMethodSignature).
2. C++ bare `...` token: extractCppParameters only iterated named
children, missing the unnamed `...` token in C-style variadics like
log_entry(const char* fmt, ...). Added fallback scan of all children.
3. Return type stripping: All 11 language extractReturnType functions
used extractSimpleTypeName() which strips generic parameters
(List<User> → "List", Task<User> → "Task"). Changed to .text?.trim()
to preserve full generic types needed for for-loop iterable resolution,
async-await binding, and return-type inference.
Also fixes PHP abstract dispatch test that matched SqlRepository instead
of the interface due to ambiguous filePath.includes('Repository') filter,
and adds parent-walk fallback in PHP isAbstract for extractFromNode path.
* chore: remove plan and review artifacts from PR
* fix: address Round 4 review findings + infrastructure improvements
- Ruby: add singleton_class support for class << self methods (4 new tests)
- PHP: add enum_declaration to CLASS_CONTAINER_TYPES
- Dart: add mixin/extension labels to CONTAINER_TYPE_TO_LABEL
- Swift: add TODO for unverifiable struct/enum node types on Node 22
- C#: add grammar version comment (0.23.1)
- Ruby: fix version comment range to pin (0.23.1)
- Rust/ast-helpers: add cross-reference comments for impl_item duplication
- ast-helpers: document CLASS_CONTAINER_TYPES ↔ typeDeclarationNodes invariant
- generic.ts: replace Array.includes with Set for O(1) dedup in addNestedBodies
- Go/Python/Ruby: align isAbstract signature with 2-param interface contract
- CLAUDE.md: fix malformed backtick around gitnexus:start HTML comment
- parsing-processor: add per-class method extraction cache (eliminates O(N*M))
- ast-helpers: add scoped_type_identifier to impl_item resolution
- call-processor: add dev-mode warnings at silent candidates[0] fallbacks
- MCP context(): surface methodMetadata for Method/Function/Constructor nodes
- resources.ts: update schema to list all stored Method properties
* fix: singleton_class HAS_METHOD edge regression in findEnclosingClassInfo
singleton_class (class << self) was added to CLASS_CONTAINER_TYPES but
has no name field — its receiver `self` has node type 'self', not
'identifier'. findEnclosingClassInfo now walks up to the enclosing
class/module to inherit its name, matching ruby.ts:extractOwnerName.
Also fixes findEnclosingClassNode in parse-worker.ts to skip
singleton_class and return the actual class/module node.
Adds integration test assertions for from_habitat (class << self method):
HAS_METHOD edge from Animal, isStatic=true, parameterCount=1.
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(vue): add Vue SFC (.vue) support for indexing
Vue Single File Components are now fully supported in the indexing pipeline.
The implementation extracts <script> / <script setup> blocks from .vue files
and parses them using the existing TypeScript tree-sitter grammar — no new
npm dependencies required.
Key changes:
- SFC script extractor: regex-based extraction of <script setup lang="ts">
blocks with correct line offset mapping back to the .vue file
- Vue language provider: reuses TypeScript queries, type config, field
extractors, and named binding extraction
- Import resolution: .vue added to EXTENSIONS so `import Foo from './Foo'`
resolves to Foo.vue; Vue import resolver delegates to TS resolver for
tsconfig path alias support
- Export detection: <script setup> top-level bindings are implicitly exported
- Template component detection: PascalCase tags in <template> emit CALLS edges
- Line offsets applied to all emitted positions (startLine, endLine, route
lineNumbers, decorator positions) in both worker and sequential paths
Validated on a 3,553-file Vue project:
Before: 24,693 nodes | 73,614 edges | 0 symbols from .vue
After: 30,495 nodes | 112,324 edges | 5,213 symbols from .vue
18,682 imports from .vue | 5,826 vue-to-vue imports
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(typescript): track destructured call results in TypeEnv
Extend `extractPendingAssignment` to handle object destructuring from
function calls and await expressions:
const { isMaker } = useUserRole()
const { data } = await fetchData()
const { name } = repo.getProfile()
Previously, only `const { x } = someVariable` (identifier RHS) produced
TypeEnv bindings. Call-expression RHS was silently skipped, leaving
destructured properties untracked.
The fix emits a synthetic `callResult` item plus N `fieldAccess` items
per destructured property, which the existing fixpoint resolver processes
in 2 iterations. No changes needed to type-env.ts, PendingAssignment
types, or call-processor — the existing infrastructure handles it.
Also extracts a `collectDestructuredFields` helper to share the
object_pattern property iteration logic between the identifier and
call-expression branches.
Note: Full property-type resolution requires the callee to have a
declared returnType in the SymbolTable. Arrow-function composables
without type annotations (common in Vue/React) won't resolve property
types until return-type inference is added in a future change.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(vue): address PR review issues for Vue SFC support
- Extract duplicated isVueSetupTopLevel to vue-sfc-extractor.ts shared
utility, removing identical copies from parse-worker.ts and
parsing-processor.ts
- Fix VUE_BUILT_INS to be a superset of TS BUILT_INS by importing and
spreading the TypeScript set, preventing spurious unresolved calls for
standard built-ins (Symbol, BigInt, WeakMap, array methods, etc.)
- Add Vue template component CALLS edge resolution in both sequential
and worker paths (call-processor.ts), matching PascalCase template
tags against imported .vue file basenames via the import map
- Add integration test for template PascalCase CALLS edges
(App.vue → Button.vue)
- Add integration test for isExported: false on non-setup <script>
blocks (OldStyle.vue options API)
- Add comment explaining TEMPLATE_RE greedy regex behavior for nested
template tags
- Fix stale language count comment (14 → 15) and remove dead code
branch in test
Made-with: Cursor
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Spec covers 4 HIGH-priority issues from review: path traversal via
group name, gRPC proto regex nested braces, service boundary detector
directory exclusions, double-close of LadybugDB pools.
Plan: 6 tasks with TDD, ordered by complexity.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Wire extractors into the sync pipeline with service boundary detection.
GroupService provides high-level API for all group operations.
- Sync pipeline: orchestrates extraction (HTTP, gRPC, topics) with
service boundary assignment and exact matching
- GroupService: groupList, groupSync, groupContracts, groupQuery,
groupStatus (groupImpact deferred to cross-repo follow-up PR)
- CLI: group create/add/remove/list/sync/contracts/query/status
- MCP tools: group_list, group_sync, group_contracts, group_query,
group_status
- Monorepo fixture: 3 services (auth/orders/gateway) connected via
gRPC + Kafka + HTTP — all intra-repo cross-links discovered
- Documentation: CLI commands and MCP tools added to both READMEs
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Core foundation for repository group analysis:
- Type system: ContractType, ExtractedContract, StoredContract, CrossLink
with optional `service` field for intra-repo matching
- Config parser for group.yaml (repos, detection flags, matching thresholds)
- Contract registry storage with atomic writes
- Exact matching engine with per-type normalization (HTTP, gRPC, topic)
and intra-repo support (different services within same repo can match)
- Extract LadybugDB pool-adapter from MCP backend for reuse by sync pipeline
- Git staleness checker for group status reporting
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(cpp): C/C++ MethodExtractor config with pure virtual detection (#572)
- Pure virtual (= 0) detected as isAbstract via token scanning
- virtual/final/override via hasKeyword and virtual_specifier children
- Access specifier visibility via backward sibling walk (public:/private:/protected:)
- Pointer/reference parameter types extracted correctly
- Constructor and destructor support via declaration node type
- Static detection via storage_class_specifier
- 16 new tests covering all acceptance criteria
* fix(cpp): isVirtual infers from override/final + out-of-class resolution
- isVirtual returns true for override/final methods (C++ mandates these
are virtual)
- Add findClassNodeByQualifiedName to parse-worker: resolves Foo::bar()
back to the Foo class declaration for method extractor enrichment
- Handles pointer/ref return types, constructors, destructors
- Integration test for virtual/static/constructor inline methods
- 233 unit+integration tests pass, 97 C++ resolver tests pass
* fix(cpp): address review — deep pointers, templates, unions, trailing returns
- Fix extractParamName: recursive unwrap for int** ptr → "ptr" (not "**ptr")
- Fix findFunctionDeclarator: recursive unwrap for multi-level pointer chains
- Template methods: generic extractor unwraps template_declaration to inner node
- union_specifier: added to typeDeclarationNodes, visibility defaults to public
- Trailing return type: auto foo() -> T now extracts T instead of "auto"
- Fix version comment: ^0.22.4 → ^0.23.4 to match package.json
- 4 new tests: double pointer params, template methods, union methods, trailing returns
* fix(cpp): template method visibility + union isTypeDeclaration test
extractCppVisibility now walks from the template_declaration parent
when the node is wrapped by a template, restoring correct access-
specifier resolution for templated class methods.
Also adds missing isTypeDeclaration assertion for union_specifier and
expands the template method test with explicit visibility checks.
* fix(cpp): address deep gap analysis review findings
- findClassNodeByQualifiedName: recursive pointer/reference
declarator unwrap, fixing out-of-class linking for deep pointer
return types (e.g. int** Foo::bar())
- findClassNodeByQualifiedName: recurse into namespace_definition
blocks so namespace-wrapped classes resolve correctly
- Suppress = delete / = default special members from extraction
via delete_method_clause / default_method_clause node detection
- Update known-gaps: namespace-wrapped classes, const-overload collapse
- Add tree-sitter-c version comment for consistency
- toBeFalsy() → toBe(undefined) for precise isVirtual assertion
- Tests: = delete, = default, = 0 non-regression, operator overloads,
deep pointer return types, default visibility (class vs struct),
multiple access specifier sections
* fix(wiki): Azure OpenAI compat and HTML viewer script injection
- Use max_completion_tokens instead of deprecated max_tokens for all models
- Skip sending temperature for Azure provider (some models reject non-default values)
- Simplify Azure interactive setup: endpoint + deployment + key (3 prompts instead of 7)
- Escape </script> in embedded JSON to prevent premature script tag closure
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(test): align wiki-llm-client test with max_completion_tokens change
The test expected max_tokens for non-reasoning models, but the source
now uses max_completion_tokens for all models since max_tokens is
deprecated by newer OpenAI models.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Abhigyan Patwari <abhigyan@Abhigyans-MacBook-Air.local>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(ts,js): MethodExtractor config for TypeScript and JavaScript (#570)
Add per-language method extraction config following the established
JVM and C# patterns. Shared config base mirrors the field extractor's
typescript-javascript.ts pattern — TS-only node types are harmless
no-ops for JS.
Key features:
- isAbstract for abstract class methods and interface methods
- Parameter extraction with isOptional (?:, defaults) and isVariadic (...)
- Decorator extraction from preceding body-level siblings
- isAsync and isOverride detection
- Visibility via accessibility_modifier two-pass pattern
- Return type extraction unwrapping type_annotation
* test(ts,js): add override, getter/setter, destructured param tests
Address code review findings:
- Add override method detection test
- Add getter/setter extraction test
- Add destructured parameter with type annotation test
- Tighten constructor and private method assertions
* refactor(ts,js): address code review findings
- Replace O(M*N) decorator index scan with previousNamedSibling walk
- Remove dead findVisibility 'modifiers' fallback (TS uses
accessibility_modifier, not a modifiers wrapper)
- Document call_signature/construct_signature as known gaps
- Document that TS constructors are method_definition nodes
- Remove unused findVisibility import
* fix(ts,js): type guard before cast, add generator/computed/overload tests
- Use type guard pattern (Set.has check before as-cast) in visibility
extraction to ensure string is validated before narrowing
- Add generator method test (*items()) — confirms extraction works
- Add computed property name test ([Symbol.iterator]) — documents
bracket-in-name behavior as intentional
- Add class-level method overload test — verifies overload signatures
+ implementation are all extracted
* fix(ts,js): detect #private methods as visibility 'private'
ES2022 private class methods (#name) use private_property_identifier
as their name node type. Detect this and return 'private' visibility
instead of the default 'public'.
* fix(ts,js): address review findings + close ingestion gaps
- hasKeyword/findVisibility: skip name field child to prevent false
positives on soft-keyword method names (e.g. `abstract()`, `static()`)
- extractTsJsParameters: filter TS `this` parameter (compile-time only)
- extractMethodSignature: mirror `this`-param skip in fallback path
- tree-sitter queries: capture abstract_method_signature,
method_signature, and private_property_identifier for TS; add
private_property_identifier for JS
- Remove dead childForFieldName('name') fallbacks and typeFromAnnotation
fallback
- Add 10+ unit tests, 4 integration tests through query pipeline
* test(ts): update HAS_METHOD count for interface method_signature capture
The new method_signature query now captures ILogger.log() as a Method
node with a HAS_METHOD edge, increasing the expected count from 4 to 5.
* fix(ts,js): address second review — async generator test, declare module gap
- Add async generator method test (async *values() → isAsync: true)
- Document declare module/global augmentation as known gap
npm runs `prepare` after `prepack` during publish, so the previous
`prepare: tsc` overwrote the rewritten imports before packing.
Both `prepare` and `prepack` now run the full build script so the
tarball always contains rewritten relative imports.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): build gitnexus-shared before publish, use CHANGELOG for release notes
The publish workflow was missing the gitnexus-shared build step that
the setup-gitnexus composite action provides. Since PR #536 unified
the ingestion pipeline, gitnexus imports types from gitnexus-shared,
so it must be built first.
Also replaces generate_release_notes with CHANGELOG.md extraction so
GitHub Releases use the reviewed changelog entry instead of a flat
PR title list.
Made-with: Cursor
* fix: bundle gitnexus-shared into CLI dist to fix module resolution
gitnexus-shared was declared as a file: dependency but never published
to npm, causing ERR_MODULE_NOT_FOUND for users installing gitnexus
globally. The build script now copies gitnexus-shared/dist into
dist/_shared/ and rewrites bare specifiers to relative paths.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: move gitnexus-shared to devDependencies, use tsc for prepare
gitnexus-shared must remain available for tsc to resolve imports during
development/CI, but is not needed at runtime since it's bundled into
dist/_shared/. Moving it to devDependencies keeps it out of production
installs while allowing compilation. The prepare script now runs plain
tsc (no shared bundling needed for local dev).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Abhigyan Patwari <abhigyan@Abhigyans-MacBook-Air.local>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The publish job failed because npm ci triggers the prepare script (tsc)
before gitnexus-shared types are available. Add the same build step
that the setup-gitnexus composite action uses in CI.
* feat(web): add repo landing screen with selectable repo cards
Instead of auto-loading the first indexed repo when the backend is
detected, show a landing screen that lets users choose which repo
to explore or analyze a new one. This addresses the UX gap where
users with multiple indexed repos had no way to pick—they were
always sent to the first one found.
- New RepoLanding component with clickable repo cards (name, stats,
indexed date) and an embedded RepoAnalyzer for new repos
- DropZone gains a 'landing' phase between server detection and
graph loading
- Shared connectToRepo handler replaces the old handleAnalyzeComplete
for both repo selection and post-analysis connection
Made-with: Cursor
* fix(web): update e2e flow for repo landing screen
The new landing screen intentionally stops auto-loading the first indexed
repo, so the existing Playwright tests were still waiting for the explorer
to appear automatically. Update the specs to select a repo from the landing
screen before asserting on the graph, and add a stable test id for repo cards.
Also format DropZone to satisfy the Prettier CI check.
Made-with: Cursor
* fix(e2e): use waitFor instead of instant isVisible for landing card
locator.isVisible() is a non-retrying instant check — the landing card
hadn't rendered yet when it was called, causing the click to be silently
skipped. Switch to waitFor which properly polls until the element appears.
Made-with: Cursor
---------
Co-authored-by: Abhigyan Patwari <abhigyan@Abhigyans-MacBook-Air.local>
* feat(csharp): add C# MethodExtractor config (#573)
Add C# method extraction config mirroring the JVM pattern from PR #576.
Wire csharpMethodConfig into the C# language provider and add 18 tests
covering classes, interfaces, abstract classes, structs, records,
constructors, params/out/ref/optional parameters, sealed methods,
attributes, and visibility modifiers.
* fix(csharp): add destructor, operator, conversion operator, and in-param support
- Add destructor_declaration, operator_declaration, and
conversion_operator_declaration to methodNodeTypes
- Custom extractName for operators (e.g., "operator +", "implicit operator double")
- Fix extractReturnType for operator declarations (use type field, not returns)
- Add in modifier to parameter extraction (alongside out/ref)
- Add 4 new tests: destructor, operator+, implicit conversion, in parameter
* fix(csharp): add ref param test and document compound visibility limitation
- Add test for ref parameter modifier (was only testing out)
- Document that protected internal / private protected resolve to first modifier
* feat(csharp): support compound visibilities (protected internal, private protected)
- Add 'protected internal' and 'private protected' to FieldVisibility union
- Detect compound modifiers in both C# method and field extractors via
collectModifierTexts helper scanning adjacent modifier nodes
- Add 2 tests for compound visibility detection
* feat(csharp): primary constructors, virtual/override/async, primary fields
Address all known limitations from review:
- Primary constructor support (C# 12): add extractPrimaryConstructor to
MethodExtractionConfig and extractPrimaryFields to FieldExtractionConfig.
Record params become public readonly properties; class params become
private captured fields.
- Add isVirtual, isOverride, isAsync optional fields to MethodInfo,
MethodExtractionConfig, NodeProperties, and parse-worker propagation.
- Detect virtual/override/async modifiers in C# method config.
- Move collectModifierTexts to shared helpers.ts (deduplicate).
- Fix destructor name to ~ClassName (disambiguates from constructor).
- Add expression-bodied method test.
- 118 tests total across method + field extraction suites, all passing.
* fix(csharp): review round 2 — annotations, record_struct, grammar pin
- Fix primary constructor annotations: use [] instead of extracting
class-level attributes (C# has no syntax for ctor-specific attributes)
- Add record_struct_declaration to typeDeclarationNodes in both method
and field extractors, CLASS_CONTAINER_TYPES, and isRecord visibility check
- Pin tree-sitter-c-sharp version (^0.23.1) in params comment
* fix(csharp): complete record_struct query + label mapping, sealed override test
- Add record_struct_declaration capture patterns to tree-sitter-queries.ts
(type definition + primary constructor)
- Add record_struct_declaration → 'Struct' in CONTAINER_TYPE_TO_LABEL
- Assert isOverride: true alongside isFinal in sealed override test
* fix(csharp): record_struct label mismatch, add record struct + documented limitation tests
- Fix record_struct_declaration query tag: @definition.struct (not @definition.record)
to match CONTAINER_TYPE_TO_LABEL and prevent broken HAS_METHOD edges
- Add 3 record struct tests: isTypeDeclaration, method extraction, primary constructor
- Add documented limitation tests: partial method (isAbstract: false), generic type
parameter stripping (name excludes <T>)
* fix(csharp): remove record_struct_declaration — not a real tree-sitter node type
tree-sitter-c-sharp 0.23.1 parses 'record struct' as record_declaration
(absorbs the 'struct' keyword as an unnamed child token). The non-existent
record_struct_declaration in queries caused TSQueryErrorNodeType, breaking
ALL C# file processing.
Remove from: tree-sitter-queries.ts, typeDeclarationNodes in both
extractors, CLASS_CONTAINER_TYPES, and CONTAINER_TYPE_TO_LABEL.
Record struct types are already handled via record_declaration.
* feat(csharp): add isPartial support, filter targeted attributes, static ctor test
- Add isPartial optional field to MethodInfo, MethodExtractionConfig,
NodeProperties, and parse-worker propagation pipeline
- Detect partial modifier in C# config — marks both declaration-only
and implemented partial methods
- Filter targeted attribute lists (e.g. [return: MarshalAs(...)]) in
extractCSharpAnnotations — only untargeted attributes collected
- Add static constructor test (isStatic: true, same name as class)
- Add 3 partial method tests: declaration-only, with body, coexisting pair
- Document record_struct/record_class as defensive dead code in
export-detection.ts (grammar absorbs keywords into record_declaration)
* fix(csharp): this param for extension methods, dedup visibility, test fixes
- Handle this modifier on extension method parameters (type prefixed
as 'this string', consistent with out/ref/in handling)
- Deduplicate visibility logic in extractPrimaryConstructor — reuse
csharpMethodConfig.extractVisibility instead of inline compound check
- Fix record struct test title to reflect actual grammar behavior
- Add conversion operator returnType assertion
- Add extension method this parameter test
* fix(csharp): primary constructor line points to param list, empty name guard
- Use paramList.startPosition instead of ownerNode.startPosition for
primary constructor line number (avoids methodInfoCache key collision)
- Guard against empty param names from tree-sitter error recovery nodes
gitnexus-web imports from gitnexus-shared, which requires npm run build
to generate dist/. Without this step, npm run dev fails with module
resolution errors.
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* feat(wiki): extend LLMConfig/CLIConfig with Azure and reasoning model fields
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(wiki): restore cursor model resolution, fix LLMProvider type, clean up regex
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(wiki): remove stale LLMProvider type alias from repo-manager
* fix(wiki): fix Azure auth header, api-version param, reasoning model params, content_filter error
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(wiki): tighten Azure detection, reasoning model regex, content_filter gating
- isReasoningModel: new regex matches only o1/o3 bare + any oN-mini/oN-preview; bare o4/o5/etc now return false
- isAzureProvider: use URL hostname matching to block spoofed subdomain URLs
- callLLM: warn on Azure legacy /deployments/ URL without api-version
- callLLM: gate content_filter error to azure===true; also catch ResponsibleAIPolicyViolation
- tests: add afterEach stub cleanup, spoofed-URL, bare-o4, non-Azure content_filter, and URL-only Azure auto-detect tests
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(wiki): detect content_filter finish_reason in SSE stream and throw clear error
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(wiki): skip delta accumulation after content_filter, use provider-neutral error message
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(wiki): add Azure OpenAI option to interactive setup wizard
Inserts Azure as option [3] in the provider menu (shifting Custom to [4]
and Cursor to [5]), adds guided Azure setup flow with resource/deployment
prompts, v1/legacy URL format selection, reasoning-model flag, and
content_filter error handling in the catch block.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(wiki): store explicit false for non-reasoning Azure deployments, trim resource name inputs
- isReasoningModelDeployment now stores false (not undefined) when user says no
- Always include isReasoningModel in saved azureConfig (no conditional guard needed)
- Trim resourceName and deploymentName prompt inputs to avoid whitespace issues
- Improve reasoning model note to mention Azure requirement
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(wiki): add --api-version and --reasoning-model CLI flags for Azure
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(wiki): include apiVersion and reasoningModel in hasCLIOverrides guard
* fix(wiki): default provider to 'openai' in resolveLLMConfig when not configured
* style: apply prettier formatting
* fix(wiki): address PR review — remove unrelated files, harden inputs
- Remove evidence/, fix-adapter.js, and planning doc accidentally included
- URL-encode apiVersion in buildRequestUrl to prevent query string injection
- Add --no-reasoning-model flag to allow CLI override of saved config
- Simplify verbose ternary in Azure wizard prompt
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(wiki): use execFileSync for EDITOR to prevent shell injection
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor: replace NodeProperties index signature any with unknown
Change [key: string]: any to [key: string]: unknown in NodeProperties.
Remove 19 redundant (node.properties as any) casts in csv-generator.ts
— all accessed properties are already declared on the type.
* refactor: replace any with SyntaxNode across ingestion layer
Mechanical substitution — all tree-sitter AST node parameters and
variables typed as any are now properly typed as SyntaxNode.
- ast-helpers.ts: 13 any → SyntaxNode
- parsing-processor.ts: 8 any → SyntaxNode
- parse-worker.ts: 40 any → SyntaxNode/TreeSitterLanguage/Parser.Query
- php.ts: 11 any → SyntaxNode
Also adds TreeSitterLanguage type alias for optional grammar loading.
* refactor: eliminate remaining any in ingestion layer
- call-processor, call-routing, c-cpp: SyntaxNode substitutions
- parse-worker: typed WorkerIncomingMessage discriminated union
- worker-pool: typed WorkerOutgoingMessage + Error handler
- ast-cache, import-processor: targeted cast for Tree.delete()
- community-processor: graphology AbstractGraph types, LeidenModule
interface for vendored leiden code
Ingestion layer: 130 → 7 any warnings remaining.
* feat(java): method references + worker overload disambiguation (TypeEnv + argTypes)
Fix two Java gaps: (1) method references (obj::method) via
tree-sitter @call + parseJavaMethodReference wired through extractLanguageCallSiteSeed for parse-worker and call-processor; (2) overloaded calls with typed
non-literal args by extending OverloadHints with TypeEnv for identifiers and adding ExtractedCall.argTypes from extractCallArgTypes on the worker path with
matchCandidatesByArgTypes (inferJvmLiteralType remains for literals).
* test(csharp): expect interface-dispatch edge for IRepository.Save in heritage fixture
* refactor(ingestion): move parseJavaMethodReference to call-sites/java.ts
* refactor(ingestion): defer worker call resolution until implementor map is complete
* style: prettier + remove unused import for CI quality checks
Made-with: Cursor
* fix(ingestion): implementor map for C# base_list + sequential pipeline path
- buildImplementorMap: treat extends rows as implements when resolveExtendsType
says IMPLEMENTS (worker heritage mirrors parse-worker, all base_list as extends)
- Worker path: pass ctx into buildImplementorMap(deferredWorkerHeritage, ctx)
- Sequential path: extract heritage before processCalls and pass implementor map
so small repos get interface-dispatch CALLS (fixes csharp-proj integration test)
Made-with: Cursor
* perf(pipeline): accumulate sequential implementor map without O(E) per chunk
- Merge buildImplementorMap(chunk heritage) into one map each sequential chunk so
work is O(heritage) per chunk and interface dispatch sees prior chunks (worker parity)
- Drop unused globalImplementorMap + redundant merge after worker pass
Made-with: Cursor
* feat: configure prettier with pre-commit hook integration
Add prettier, lint-staged, and prettier-plugin-tailwindcss at the repo
root with husky pre-commit hook integration. Moves husky from
gitnexus/ to root package.json for reliable hook installation.
- Root package.json with prepare/format/format:check scripts
- .prettierrc with endOfLine:lf and tailwindStylesheet for TW v4
- .prettierignore excluding fixtures, vendor, generated, *.d.ts, *.md
- .gitattributes enforcing LF line endings for Windows consistency
- Pre-commit hook uses direct node_modules/.bin/ paths (no npx)
* style: apply prettier formatting to entire codebase
One-time bulk format. No logic changes.
Use .git-blame-ignore-revs to skip this commit in git blame.
* chore: add .git-blame-ignore-revs for prettier format commit
* perf: pre-commit hook runs only tests related to staged files
Use vitest --related to scope test execution to tests that import
the changed files, instead of running the full suite on every commit.
* perf: remove vitest from pre-commit hook, keep in CI only
Pre-commit now runs lint-staged + tsc only. Tests run in CI
(ci-tests.yml) where they belong — keeps commits fast.
* ci: add prettier format check to quality workflow
PRs will now fail if code isn't formatted with prettier.
* feat: add server-side ingestion API (POST /api/analyze, SSE progress)
Extract core analysis orchestration from CLI into shared run-analyze.ts
module. Add server-side analyze endpoints so the web app can trigger
ingestion via HTTP instead of running the full pipeline in-browser.
New files:
- src/core/run-analyze.ts — shared runFullAnalysis() orchestrator
- src/server/analyze-job.ts — job manager (single-slot, dedup, SSE events)
- src/server/analyze-worker.ts — forked child process (8GB heap, IPC)
- src/server/git-clone.ts — shallow clone/pull with SSRF protection
API endpoints:
- POST /api/analyze — start analysis (returns 202 + jobId)
- GET /api/analyze/:jobId — poll job status
- GET /api/analyze/:jobId/progress — SSE progress stream
Security: URL validation blocks private IPs and non-HTTP schemes.
Path validation requires absolute paths. Git stderr not leaked to API.
* feat(web): add server-side analyze UI (Phase 2)
Add "Analyze on Server" flow to the web app's Server tab so users
can trigger server-side ingestion from the browser. On completion,
the graph is automatically loaded via the existing connectToServer flow.
New files:
- AnalyzeProgress.tsx — progress bar with phase label, elapsed time, cancel
Modified files:
- backend.ts — startAnalyze(), streamAnalyzeProgress() SSE client
- DropZone.tsx — analyze URL input + button below Connect section
- App.tsx — onServerAnalyze handler wires analyze -> connect flow
* feat: add job cancellation, timeout, and child process tracking (Phase 3)
- DELETE /api/analyze/:jobId — cancel running analysis (SIGTERM to worker)
- 30-minute timeout kills long-running workers automatically
- Child process refs tracked in JobManager for cleanup on shutdown
- dispose() kills all active children on SIGINT/SIGTERM
- Web cancel button now calls server DELETE endpoint
- cancelAnalyze() added to web backend client
* refactor(web): remove browser ingestion pipeline (Phase 4)
Delete 16 duplicated ingestion files, 2 unused service files
(git-clone, zip), and tree-sitter parser-loader from gitnexus-web.
All ingestion now runs server-side via POST /api/analyze.
Deleted (18 files, ~5,000 lines):
- core/ingestion/*.ts (16 pipeline processors)
- core/tree-sitter/parser-loader.ts (WASM tree-sitter loader)
- services/git-clone.ts (isomorphic-git client-side clone)
- services/zip.ts (JSZip extraction)
Simplified:
- DropZone.tsx — server-only (removed ZIP/GitHub tabs)
- ingestion.worker.ts — removed runPipeline/runPipelineFromFiles
- useAppState.tsx — removed pipeline callbacks
- App.tsx — removed handleFileSelect/handleGitClone
- main.tsx — removed Buffer polyfill for isomorphic-git
- types/pipeline.ts — removed PipelineResult/serialize helpers
Kept: cluster-enricher.ts (LLM enrichment, still used by worker)
Dependencies now removable: web-tree-sitter, isomorphic-git,
@isomorphic-git/lightning-fs, jszip (estimated 3-4MB bundle savings)
* refactor(web): sync graph schema from CLI + delete WASM grammars
Sync graph/types.ts and lbug/schema.ts from the CLI (source of truth)
to the web module so the browser LadybugDB can handle all node and
relationship types the server pipeline produces.
Synced types: Route, Tool, Section node labels; HANDLES_ROUTE, FETCHES,
HANDLES_TOOL, ENTRY_POINT_OF, WRAPS, QUERIES relationship types;
description fields on Function/Class/Interface/Method/CodeElement.
Deleted: public/wasm/ directory (14 tree-sitter WASM grammars + core).
Removed deps: web-tree-sitter, isomorphic-git, @isomorphic-git/lightning-fs,
jszip, buffer, @types/jszip (~3-4MB bundle savings).
* feat: create gitnexus-shared package for unified type definitions
Create a new gitnexus-shared package that is the single source of truth
for types shared between the CLI and web modules:
- SupportedLanguages enum (15 languages)
- Graph types: NodeLabel, NodeProperties, RelationshipType, GraphNode, GraphRelationship
- Schema constants: NODE_TABLES, REL_TYPES, REL_TABLE_NAME, EMBEDDING_TABLE_NAME
- Pipeline types: PipelinePhase, PipelineProgress
Both gitnexus (CLI) and gitnexus-web import from gitnexus-shared via
file: dependency. Each package re-exports and extends with platform-specific
additions (CLI: KnowledgeGraph with mutation methods; Web: simpler KnowledgeGraph).
This ensures types can never drift between packages — adding a new
language, node type, or relationship type in gitnexus-shared automatically
propagates to both consumers.
* refactor: import shared types directly from gitnexus-shared at call sites
Replace all re-export patterns with direct imports from gitnexus-shared.
72 files updated across CLI and web:
- SupportedLanguages: 49 CLI files now import from 'gitnexus-shared'
instead of '../config/supported-languages.js'
- GraphNode, GraphRelationship, NodeLabel: 22 CLI + 10 web files now
import from 'gitnexus-shared' instead of local re-export wrappers
- NODE_TABLES: api.ts imports from 'gitnexus-shared'
- PipelineProgress: useAppState.tsx imports from 'gitnexus-shared'
Local types.ts files now only define platform-specific KnowledgeGraph
(CLI has mutation methods, web has add-only). No more re-exports.
* fix: update lock files for gitnexus-shared, remove stale vite polyfills
Add gitnexus-shared@1.0.0 to lock files so npm ci succeeds in CI.
Remove buffer polyfill and global define from vite.config.ts (isomorphic-git was removed).
* fix(security): add write guard to HTTP /api/query, fix CORS proxy bypass
- Add isWriteQuery() check to POST /api/query handler — blocks CREATE,
DELETE, SET, MERGE, DROP, etc. via HTTP API (guard was only in MCP
pool adapter and browser-side, not the HTTP server path)
- Extend CYPHER_WRITE_RE with CALL, INSTALL, LOAD keywords
- Fix CORS proxy subdomain bypass: endsWith('github.com') allowed
'evil-github.com'. Now requires exact match or '.github.com' suffix
* feat(server): enhance /api/search with enrichment, add /api/grep, strip graph content
- POST /api/search: add mode param (hybrid|semantic|bm25), server-side
enrichment returns connections/cluster/processes per result in one call
(collapses 31 sequential HTTP calls to 1 for the agent search tool)
- GET /api/grep: regex search across indexed file contents, eliminates
need to transfer all file contents to browser
- GET /api/graph: strip content field by default (80-95% payload
reduction). Use ?includeContent=true for backward compat
- Add LRU cache invalidation hook point for future caching
* feat(server): add /api/embed endpoint for server-side embedding generation
- POST /api/embed: triggers embedding pipeline via onnxruntime-node
with JobManager for single-slot concurrency, timeout, and dedup
- GET /api/embed/:jobId: poll job status
- GET /api/embed/:jobId/progress: SSE stream with heartbeat, event IDs,
and X-Accel-Buffering:no header for proxy compatibility
- DELETE /api/embed/:jobId: cancel running embedding job
- Maps embedding pipeline phases (ready→complete, error→failed) to
JobManager status conventions
* feat(web): create consolidated BackendClient module
Single HTTP client replacing backend.ts, server-connection.ts, and
worker HTTP helpers. Includes:
- Typed methods: runQuery, search (enriched), grep, readFile, connect
- Generic streamSSE<T> utility extracted from analyze progress pattern
- BackendError with discriminated code field (network/server/client/timeout)
- Embed API: startEmbeddings, streamEmbeddingProgress, cancelEmbeddings
- Search with mode param (hybrid|semantic|bm25) and enrichment
* refactor(web): rewrite Graph RAG tools for backend-only HTTP queries
- Search tool: uses enriched /api/search (1 call replaces 31 sequential queries)
- Cypher tool: removes browser-side embedding; {{QUERY_VECTOR}} routes to
/api/search with mode:'semantic' instead of local transformers.js
- Grep tool: uses /api/grep instead of in-memory fileContents map
- Read tool: uses /api/file instead of fileContents map lookup
- Impact tool: getCallSiteSnippet now async via /api/file
- createGraphRAGTools now accepts GraphRAGBackend interface instead of
7 separate function params + fileContents map
- createGraphRAGAgent simplified to (config, backend, context?)
- Removed imports: embedder, lbug/schema (replaced with gitnexus-shared)
- Net: -205 lines
* refactor(web): delete WASM infrastructure, remove 7 packages (-5242 lines)
Delete browser-side LadybugDB, embeddings, search, and worker:
- gitnexus-web/src/core/lbug/ (adapter, csv-generator, schema, query-result)
- gitnexus-web/src/core/embeddings/ (embedder, pipeline, text-gen, types)
- gitnexus-web/src/core/search/ (bm25-index, hybrid-search)
- gitnexus-web/src/workers/ingestion.worker.ts (828 lines)
- gitnexus-web/src/services/server-connection.ts (merged into backend-client)
- gitnexus-web/src/types/lbug-wasm.d.ts
Remove packages: @ladybugdb/wasm-core, @huggingface/transformers,
comlink, minisearch, vite-plugin-wasm, vite-plugin-top-level-await,
vite-plugin-static-copy
Update vite.config.ts: remove WASM plugins, COOP/COEP headers,
worker config, optimizeDeps exclude
Update imports: App.tsx, DropZone, Header, AnalyzeProgress,
BackendRepoSelector, useBackend → backend-client
* refactor(web): replace Worker/Comlink with direct BackendClient calls
- useAppState: remove Worker instantiation, Comlink.wrap, apiRef.
All queries now go through BackendClient HTTP functions directly.
- Agent runs on main thread (I/O-bound LLM streaming, not CPU-bound)
- initializeAgent: creates GraphRAGAgent with GraphRAGBackend interface
bound to BackendClient methods (runQuery, search, grep, readFile)
- startEmbeddings: calls POST /api/embed + SSE progress instead of
running browser-side transformers.js pipeline
- switchRepo: no longer loads graph into WASM DB or extracts fileContents
- App.tsx: handleServerConnect simplified (no fileContents, no loadServerGraph)
- Delete old backend.ts (replaced by backend-client.ts)
- Net: -396 lines
* fix(web): fix await-in-map build error in agent streaming
Move dynamic import of AIMessage outside .map() callback to avoid
"await can only be used inside an async function" build error.
* fix(web): remove stale apiRef references that broke chat functionality
sendChatMessage referenced apiRef.current (deleted Worker ref) which
would throw TypeError. Replaced with agentRef.current guard since agent
now runs on main thread.
* fix(server): dispose embedJobManager on shutdown, fix job mutation
- Add embedJobManager.dispose() to shutdown handler (was missing,
causing cleanup timer to keep Node process alive)
- Replace direct job.repoName/status mutation with updateJob() to
ensure SSE event emission for initial status change
* fix(server): parameterize Cypher, harden grep, unify SSE endpoints
- Search enrichment: replace string interpolation with executePrepared()
using $nid parameter binding to prevent Cypher injection
- Add executePrepared() to core lbug-adapter (prepare/execute pattern)
- /api/grep: add 200-char pattern length limit (ReDoS protection),
search files on disk instead of loading entire corpus into memory
(constant memory usage regardless of repo size)
- Extract mountSSEProgress() shared helper for SSE streaming — both
analyze and embed endpoints now have consistent heartbeat (30s),
event IDs (reconnection support), and X-Accel-Buffering header
* refactor(web): remove dead code from Worker-era architecture
- Remove loadServerGraph no-op function, interface member, and all consumers
- Remove testArrayParams stub and interface member
- Remove fileContents state from GraphStateProvider (never populated in
server-side architecture)
- Remove forceDevice parameter from startEmbeddings (server-side, no device choice)
- Replace phantom EmbeddingProgress type with inline { phase, percent }
- Replace resolvePathFromContents (needed fileContents Map) with graph-based
file path resolution using filePathIndex built from graph nodes
- Fix: AI citation grounding ([[file.ts:10]]) now works via graph node lookup
instead of broken fileContents-based resolution
* fix(web): use streamAgentResponse for full tool_call/reasoning streaming
Replace naive agent.stream() loop that only handled content chunks with
streamAgentResponse() generator from agent.ts. This properly routes:
- reasoning tokens (before/between tool calls)
- tool_call events (name, args, status)
- tool_result events (completed tool output)
- content tokens (final answer after all tools done)
Previously the onChunk handler for tool_call/tool_result/reasoning was
dead code since the streaming loop only emitted content events.
* fix(web): resolve CI type errors from dead code removal
- Import GraphNode/GraphRelationship from gitnexus-shared in graph.ts
(not re-exported from local types.ts)
- Add Route, Tool entries to NODE_COLORS and NODE_SIZES constants
- Add PipelineResult type to web types/pipeline.ts
- Remove fileContents from CodeReferencesPanel and RightPanel
- Remove testArrayParams and forceDevice from EmbeddingStatus
- Remove forceDevice args from startEmbeddings() calls in App.tsx
- Fix embeddingProgress property accesses for simplified type
* fix(ci): add setup-gitnexus-web action, build shared once per job
- Remove prepare script from gitnexus-shared (tsc not available during
npm ci of consuming packages)
- Create .github/actions/setup-gitnexus-web composite action: builds
gitnexus-shared then runs npm ci for gitnexus-web
- setup-gitnexus action: already builds gitnexus-shared for CLI jobs
- ci-quality typecheck-web: uses setup-gitnexus-web (DRY)
- ci-e2e: uses setup-gitnexus-web (DRY)
- ci-tests: gitnexus-shared already built by setup-gitnexus, just
install web deps without rebuilding
* fix(ci): use prepare script so gitnexus-shared builds during npm ci
Move typescript from devDependencies to dependencies in gitnexus-shared
so the prepare script (tsc) works when npm resolves file: deps during
npm ci. No GHA modifications needed — npm handles the build lifecycle
automatically.
Remove manual gitnexus-shared build steps from setup-gitnexus and
setup-gitnexus-web actions.
* fix(ci): build gitnexus-shared explicitly in setup actions
The file: dependency protocol doesn't reliably run prepare scripts
because devDependencies aren't installed first. Instead of fragile
lifecycle hacks, build gitnexus-shared explicitly in both setup actions:
- setup-gitnexus: npm install && npm run build in gitnexus-shared/
- setup-gitnexus-web: same, before npm ci in gitnexus-web/
- ci-tests: shared already built by setup-gitnexus, web just npm ci
No prepare script, no dist in git, no typescript as a prod dependency.
* fix: remove CALL from CYPHER_WRITE_RE — breaks FTS and vector search
CALL is used by read-only procedures: CALL QUERY_FTS_INDEX(...) and
CALL QUERY_VECTOR_INDEX(...). Adding it to the write guard blocked all
FTS search, causing 3 test failures. The database is opened in read-only
mode as defense-in-depth against write procedures via CALL.
Keep INSTALL and LOAD in the blocklist (genuinely dangerous).
* fix(web): update vercel.json for gitnexus-shared, remove COOP/COEP
- Add installCommand that builds gitnexus-shared before installing
web deps (Vercel doesn't know about the monorepo file: dependency)
- Remove Cross-Origin-Opener-Policy and Cross-Origin-Embedder-Policy
headers (no longer needed — WASM LadybugDB removed)
* fix(web): update tests for deleted modules
- Delete csv-generator.test.ts (tests deleted WASM-only csv-generator)
- Update security-guards.test.ts: import NODE_TABLES/REL_TYPES from
gitnexus-shared instead of deleted src/core/lbug/schema
- Update server-connection.test.ts: import normalizeServerUrl from
backend-client, remove extractFileContents tests (function deleted)
* fix(e2e): remove Server tab click — UI is now server-only
The DropZone no longer has ZIP/GitHub/Server tabs (browser ingestion
was removed). The server URL input is directly visible on the landing
page. Update e2e test to skip the tab click and go straight to input.
All 5 e2e tests pass locally.
* refactor: use gitnexus-shared for PipelinePhase/PipelineProgress types
CLI was duplicating PipelinePhase and PipelineProgress locally instead
of importing from gitnexus-shared. Updated all consumers to import
directly. Also removed dead code: SerializablePipelineResult,
serializePipelineResult(), deserializePipelineResult().
* fix(server): address PR #536 review — security, race conditions, dead code
- Fix path traversal in POST /api/analyze: split into isAbsolute + normalize check
- Add shared repo lock (activeRepoPaths) preventing concurrent analyze+embed on same repo
- Fix 202 response returning actual job.status instead of hardcoded 'queued'
- Add 30-minute timeout for embedding jobs (was missing unlike analyze jobs)
- Fix DropZone calling startAnalyze without setting backend URL first
- Add SSE reconnect with exponential backoff (3 retries) and Last-Event-ID
- Fix normalizeServerUrl to return base URL (no /api suffix) — clear contract
- Delete dead code: proxy.ts, server-graph-hydration.ts, pipeline.ts re-export barrel
- Update LoadingOverlay to import PipelineProgress directly from gitnexus-shared
* fix(server): fix repo lock key mismatch and embed cancel race
- Use getStoragePath(targetPath) as lock key in analyze handler to match
embed handler's entry.storagePath — keys now always align
- Guard embed completion: don't overwrite 'failed' with 'complete' when
job was cancelled while pipeline was still running
- Remove unused jobType parameter from acquireRepoLock
- Log backend.init() errors instead of silently swallowing
* fix: add gitnexus-shared as a local dependency in package-lock.json
* refactor: move language detection to gitnexus-shared, add syntax highlighting for all 15 languages
Move getLanguageFromFilename() from CLI to gitnexus-shared with COBOL
support added. Add getSyntaxLanguageFromFilename() for Prism-compatible
syntax highlighting covering all 15 code languages plus auxiliary
formats (json, yaml, markdown, html, css, bash, sql, xml).
Refactor CodeReferencesPanel to use shared function instead of a local
30-line switch. Delete dead gitnexus-web/src/config/supported-languages.ts
(web already imports SupportedLanguages from gitnexus-shared).
* feat(web): add first-time user onboarding with auto server detection
Replace the manual "Connect to Server" panel with an automatic onboarding
flow that guides first-time users through starting the GitNexus server.
Server detection:
- useBackend hook polls via setTimeout chain (3s, no overlap)
- Page Visibility API pauses polling when tab is hidden
- SSE heartbeat (/api/heartbeat) for instant disconnect detection
Onboarding UI (OnboardingGuide.tsx):
- Step-by-step flow: copy command → run → auto-connect
- Smart command: shows `gitnexus serve` in dev, `npx gitnexus@latest serve` in prod
- Node.js version auto-detected from package.json via Vite define
- Faux terminal windows with copy-to-clipboard, platform tabs, polling indicator
Transitions (DropZone.tsx):
- Crossfade wrapper with snapshot pattern for smooth phase transitions
- Three phases: onboarding → success (1.2s hold) → loading → graph
- Auto-recovery: falls back to onboarding if server dies or connect fails
Server changes:
- GET /api/heartbeat: SSE endpoint for liveness detection
- GET /api/info: version, launch context, Node.js version
- npm run serve script for local development
- app.disable('x-powered-by') hardening
* feat(web): add repo analysis UI, SSE heartbeat, and review fixes
Repo analysis:
- AnalyzeOnboarding: empty-state card when server has zero repos
- RepoAnalyzer: GitHub URL + Local Folder tabs with browse button
- Header repo dropdown: click project badge to switch repos or analyze new
- DropZone 'analyze' phase integrated into Crossfade transitions
Reliability fixes from 5-agent review:
- Polling: stop scheduling timers when tab hidden, restart on visibility return
- Heartbeat: exponential backoff (1s/2s/4s, 3 retries) prevents graph loss on blip
- RepoAnalyzer: completion timer tracked in ref, cleaned up on unmount
- DropZone: standardized card padding (p-7), heading sizes (text-lg)
Accessibility:
- prefers-reduced-motion global CSS rule (WCAG 2.3.3)
- focus-visible rings on CopyButton
- cursor-pointer on all Header buttons
- Consistent rounded-xl on all dropdowns
Cleanup:
- Deleted dead AnalyzeSheet.tsx (219 LOC) and BackendRepoSelector.tsx (89 LOC)
- Fixed AnalyzeProgress lucide import (lucide-react → @/lib/lucide-icons)
* fix(server): resolve analyze worker fork crash in dev mode
The forked analyze worker was crashing immediately with exit code 1
when running via `npm run serve` (tsx). Two issues:
1. Worker path resolved to `analyze-worker.js` but only `.ts` exists
in the source directory — the `.js` file is only in `dist/`.
2. On Windows, bare `--import tsx` in execArgv fails because Node's
ESM resolver for --import uses the child's CWD, not the parent's
node_modules. Windows also rejects raw paths as `d:` is not a
valid URL scheme.
Fix: detect dev vs prod via `import.meta.url` extension. In dev mode,
resolve `tsx/esm` to an absolute `file://` URL via `pathToFileURL()`
anchored to the parent's `createRequire` context. This works on all
platforms and doesn't depend on the child's CWD or PATH.
Also captures child stderr for better crash diagnostics.
Verified: `POST /api/analyze` with GitHub URL completes successfully
in dev mode (tsx) — status goes from cloning → analyzing → complete.
* fix(server): add worker auto-retry, error handling, and crash diagnostics
Worker resilience:
- Auto-retry up to 2 times with exponential backoff (1s, 2s) on crash
- SSE progress shows "Retrying after crash (1/2)..." during retry
- Captures child stderr for crash diagnostics in failure message
- AnalyzeJob tracks retryCount per job
Server error handling:
- app.listen wrapped in Promise so EADDRINUSE/EACCES propagate cleanly
- serve.ts catches startup errors with friendly messages and exit code 1
- EADDRINUSE gets actionable guidance (stop other process or --port flag)
- Global uncaughtException/unhandledRejection handlers prevent silent exits
- DEBUG=1 env var shows full stack traces
* feat: add e2e tests for onboarding flows, worker retry, and error handling
E2E tests (onboarding.spec.ts — 11 tests):
- Flow 1: OnboardingGuide shown when server unreachable (6 tests)
- Flow 2: Auto-connect with success card, analyze phase for zero repos
- Flow 3: Analyze form — GitHub URL validation, Local Folder tab, tab switching
- Flow 4: Repo dropdown in exploring view (skipped without live server)
Updated server-connect.spec.ts:
- Replaced manual Connect button flow with auto-connect waitForGraphLoaded
Server resilience:
- Worker auto-retry (2 attempts with exponential backoff) on crash
- Friendly error messages for serve startup failures (EADDRINUSE etc.)
- Global uncaughtException/unhandledRejection handlers prevent silent exits
- app.listen wrapped in Promise for proper error propagation
* refactor(shared): enforce exhaustive language coverage via Record types
Replace the if/else chain in getLanguageFromFilename with two exhaustive
Record<SupportedLanguages, ...> maps:
- EXTENSION_MAP: every language → its file extensions
- SYNTAX_MAP: every language → its Prism syntax identifier
Adding a new member to the SupportedLanguages enum without adding it to
both maps now produces a TypeScript compile error:
Property '[SupportedLanguages.NewLang]' is missing in type...
This matches the existing pattern in languages/index.ts (providers table)
which already uses `satisfies Record<SupportedLanguages, LanguageProvider>`.
Three compile-time enforcement points now exist:
1. EXTENSION_MAP in language-detection.ts (file extensions)
2. SYNTAX_MAP in language-detection.ts (Prism syntax identifiers)
3. providers in languages/index.ts (LanguageProvider instances)
* feat(web): load source code from server and scroll to selected line
CodeReferencesPanel now fetches file content via GET /api/file when a
node is selected, instead of showing "Code not available in memory".
- Fetches via readFile() from backend-client when selectedFilePath changes
- Shows loading spinner while fetching
- After content loads, auto-scrolls to the selected node's startLine
- Highlights the selected line range with a cyan left border
- Cancels in-flight fetch if selection changes before it completes
Also: refactored language-detection.ts to use exhaustive Record types
(EXTENSION_MAP and SYNTAX_MAP) so adding a new SupportedLanguages enum
member without implementing extensions/syntax is a compile error.
* feat: buffered file reading for Code Inspector
Server: GET /api/file now supports ?startLine=N&endLine=M for reading
a line range instead of the entire file. Returns { content, startLine,
endLine, totalLines }.
Client: readFile() returns ReadFileResult with metadata. When selecting
a symbol (function, class, method), fetches only ±50 lines around the
symbol's startLine/endLine instead of the full file. File nodes still
fetch the entire file.
SyntaxHighlighter startingLineNumber set from the buffer offset so line
numbers are correct even for partial reads.
* fix: adapt readFile callers to new ReadFileResult return type
tools.ts: readFile comes from GraphRAGBackend interface which returns
Promise<string> (the adapter in useAppState extracts .content), so
revert the { content } destructuring back to plain string assignment.
useAppState.tsx: wrap backendReadFile with { repo } options object
and extract .content to satisfy the GraphRAGBackend interface.
* fix(web): ensure new repos appear in list immediately after analysis
Two fixes:
1. DropZone: handleAnalyzeComplete now passes the repoName through to
connectToServer so the specific newly-analyzed repo loads — not the
server's default first repo.
2. App.tsx: fetchRepos() is now awaited BEFORE handleServerConnect in
both the DropZone and Header flows. This ensures the repo list is
populated before the exploring view renders, so the new repo appears
in the header dropdown immediately without a page reload.
* feat: delete repos, re-analyze with force, select after analysis
Server — DELETE /api/repo:
- Acquires repo lock first (409 if analyze/embed in flight)
- Closes LadybugDB, deletes index + clone dir, unregisters, re-inits
- Lock released in finally block
Server — analyze complete:
- backend.init() must succeed before SSE complete fires
- If backend.init() fails, job is marked failed (not complete)
Web — Header repo dropdown:
- Re-analyze: calls POST /api/analyze with force=true, shows spinning
icon + inline progress bar via SSE
- Delete: acquires lock, aborts any running re-analysis SSE for same
repo, refreshes list, switches to next repo
- After analysis completes: refreshes repo list, connects to the
specific repo by name, loads graph, shows in explorer
- Retry with 1.5s backoff on 404 (server may still be reinitializing)
Type safety:
- err: any → err: unknown + instanceof BackendError in retry loop
- Added missing BackendRepo + BackendError imports in App.tsx
Downgrade tree-sitter from ^0.25.0 to ^0.21.1 and align all parser versions to eliminate ERESOLVE peer dependency conflicts that break MCP server install via npx. Also corrects hallucinated tree-sitter-dart SHA. Fixes#537
* feat(phase8): add field type data structures and extractor interface
* feat(phase8): implement TypeScript field extractor
* feat-phase9-add-call-result-binding
* test-phase8-add-field-extraction-unit-tests
* docs: update documentation for Phase 8 and Phase 9
* feat(swift): Phase 8/9 integration tests for field-type and call-result binding
Add Swift field-type resolution and call-result binding integration tests
with fixtures, plus merge-conflict fixes for the FieldExtractor code.
**Swift integration tests:**
- `swift-field-types/` fixture (Models.swift + App.swift) — tests
HAS_PROPERTY edges, field-chain CALLS resolution (user.address.save()
→ Address#save), and ACCESSES edges for field reads.
- `swift-call-result-binding/` fixture — tests call-result binding
(let user = getUser(); user.save() → User#save).
- 2 new describe blocks in swift.test.ts with skipIf(!swiftAvailable).
**Swift arity fix:**
- extractMethodSignature fallback counts direct `parameter` children
when no wrapper list node exists (Swift's tree-sitter grammar places
parameters as direct children of function_declaration). Without this,
all Swift functions had parameterCount: 0 and the arity filter rejected
valid call targets.
**FieldExtractor merge-conflict fixes:**
- field-extractor.ts: update import from removed ./utils.js to
./utils/ast-helpers.js; use typeEnv.fileScope() instead of .get('').
- field-extractors/typescript.ts: same import fix.
- field-types.ts: alias TypeEnvironment as TypeEnv (renamed on main).
- field-extraction.test.ts: mock TypeEnvironment interface properly.
* feat(field-extractors): generic table-driven field extractors for all 14 languages, wired into pipeline
Implements field extractors for all supported languages and integrates
them into the ingestion pipeline as the single source of truth for
Property node metadata.
**Generic field extractor factory** — `field-extractors/generic.ts`
defines a `createFieldExtractor(config)` factory that generates
FieldExtractor instances from a per-language `FieldExtractionConfig`.
Each config specifies AST node types, name/type/visibility extraction
functions, and static/readonly detection — typically 20-40 lines per
language vs 300+ for a hand-written extractor.
**Per-language configs** — `field-extractors/configs/` has 11 config
files covering 13 languages (TS/JS share, Java/Kotlin share).
TypeScript keeps its hand-written extractor for richer handling.
**LanguageProvider integration** — New optional `fieldExtractor` property
on LanguageProviderConfig, set via defineLanguage() in each language
file. Follows the same strategy pattern as typeConfig, exportChecker,
and labelOverride. Removed the separate FieldExtractorRegistry class
and field-extractors/index.ts — extractors are accessed via
getProvider(lang).fieldExtractor.
**Pipeline wiring** — Both parse-worker.ts (worker pool) and
parsing-processor.ts (sequential fallback) now call the FieldExtractor
during Property node creation. Results are cached per class node.
Property nodes are enriched with: declaredType, visibility, isStatic,
isReadonly.
**extractPropertyDeclaredType removed** — The 100-line multi-strategy
function in type-extractors/shared.ts is replaced by the FieldExtractor.
All 14 languages register an extractor, eliminating the need for a
generic fallback. The Python config's extractType was fixed to handle
annotation-without-value patterns (address: Address).
**Integration tests** — Each language's resolver test file gains
pipeline-based assertions verifying visibility/isStatic/isReadonly on
Property nodes via getNodesByLabelFull. Tests run through
runPipelineFromRepo with real fixtures — no direct extractor calls.
* fix(type-env): thread enclosingFunctionFinder through scope resolution, unskip Dart ACCESSES test
The type-env's findEnclosingScopeKey had the same Dart sibling problem
as findEnclosingFunction — it walked parents but never found
function_signature because the call lives inside function_body (a
sibling). Instead of hardcoding a function_body check, thread the
provider's enclosingFunctionFinder hook through BuildTypeEnvOptions →
lookupInEnv → findEnclosingScopeKey. All three buildTypeEnv call sites
(call-processor, parsing-processor, parse-worker) now pass the hook.
This enables the type-env to resolve scoped parameter bindings for Dart
(e.g., `user: User` in processUser), which lets the chain-resolution
tier (Step 1c) walk `user.address` and emit ACCESSES edges.
Dart integration test unskipped — 10/10 passing including ACCESSES.
Reverted CHANGELOG.md to origin/main.
* fix: resolve all PR #494 review findings (10 items)
CRITICAL:
- parse-worker.ts: classNode: any → SyntaxNode on getFieldInfo
and findEnclosingClassNode; removed redundant as number casts
- parsing-processor.ts: classNode: any → SyntaxNode on seqGetFieldInfo
HIGH:
- ruby.ts: attr_accessor now extracts ALL symbol arguments via
extractNames hook in generic factory (was firstNamedChild only)
- typescript.ts: added JSDoc explaining why hand-written extractor
coexists with config-based typescript-javascript.ts
MEDIUM:
- field-types.ts: FieldVisibility union type replaces string
('public'|'private'|'protected'|'internal'|'package'|'fileprivate'|'open')
Propagated through field-extractor.ts, generic.ts, all 7 config files
- typescript.ts: extractFullType collapsed from 12 branches to 3 lines
- generic.ts: added extractNames? optional hook + buildField refactor
LOW:
- ruby.ts: extractVisibility(node) → extractVisibility(_node)
- python.ts: fixed misleading isStatic comment
TypeScript compiles cleanly.
* test: add 24 field extraction tests for generic factory + 5 languages
Generic factory (4 tests):
- createFieldExtractor with TypeScript config validates factory itself
- Body discovery for interfaces, static/readonly modifiers
- Non-type node rejection
Python (4 tests):
- Annotated class field extraction
- Underscore-based visibility: _name=protected, __name=private
Go (5 tests):
- isTypeDeclaration on type_declaration nodes
- Config functions: uppercase=public, lowercase=package visibility
- extractType, isStatic, isReadonly
C++ (5 tests):
- public/private/protected access specifier backward-sibling walk
- Default visibility: class=private, struct=public
- static/const modifier detection
Ruby (6 tests):
- attr_accessor multi-symbol: :name, :email, :age → 3 fields
- attr_reader=readonly, attr_writer=non-readonly
- Multiple attr_* calls in one class
Total: 46 tests passing
* chore: remove plan doc from PR
---------
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* feat: add COBOL language support with regex extraction pipeline
Standalone COBOL processor following the markdown-processor.ts pattern:
- No LanguageProvider modification — COBOL uses regex, not tree-sitter
- No SupportedLanguages enum change — standalone processor pattern
New files:
- cobol-processor.ts — orchestrator (processCobol, isCobolFile, isJclFile)
- cobol/cobol-preprocessor.ts — regex state machine extraction (~888 LOC)
- cobol/cobol-copy-expander.ts — COPY statement expansion with circular detection
- cobol/jcl-parser.ts — JCL job/step/DD extraction
- cobol/jcl-processor.ts — JCL graph node creation
Extraction produces:
- Module nodes (PROGRAM-ID)
- Function nodes (paragraphs)
- Namespace nodes (sections)
- Property nodes (data items)
- CALLS edges (PERFORM intra-file, CALL cross-program)
- IMPORTS edges (COPY statements)
- CONTAINS edges (section → paragraph hierarchy)
Pipeline integration: single processCobol() call in Phase 2.6
54 new tests (33 COBOL + 21 JCL), all 3889 tests pass.
* docs: document custom processor pattern in pipeline.ts
Add comment block at the custom processor integration point
documenting the pattern for future non-tree-sitter language additions.
* feat(cobol): enrich graph with EXEC SQL/CICS, ENTRY points, MOVE data flow, PERFORM THRU
Maps the remaining 60% of CobolRegexResults to the graph:
- EXEC SQL blocks → CodeElement nodes + ACCESSES edges to DB tables
- EXEC CICS LINK/XCTL → CodeElement nodes + cross-program CALLS edges
- ENTRY points → Constructor nodes (registered for cross-program resolution)
- MOVE statements → ACCESSES edges (read/write data flow tracking)
- PERFORM THRU → expanded CALLS edges for range targets
- File declarations → Record nodes with assignment metadata
- Cross-program CALL 2nd pass: resolves unresolved targets after all programs processed
* test(cobol): add 26 integration tests with exact assertions + fix CICS resolution bug
Integration tests (test/integration/resolvers/cobol.test.ts):
- 26 tests covering full COBOL system extraction
- ALL assertions use exact toBe(N) — zero fuzzy assertions
- Fixtures: CUSTUPDT.cbl, AUDITLOG.cbl, CUSTDAT.cpy, RPTGEN.cbl, RUNJOBS.jcl
Bug fix (cobol-processor.ts):
- CICS LINK/XCTL cross-program resolution was broken — edges were
created with "resolved" reason but pointing to <unresolved> targets
- Fix: use cics-link-unresolved / cics-xctl-unresolved suffix pattern
matching the existing cobol-call-unresolved pattern
- Second-pass resolver now patches both CALL and CICS unresolved edges
All 3915 tests pass, 0 failures.
* test(cobol): exhaustive 57-test suite with strict exact assertions
Complete rewrite of COBOL integration tests using ground-truth approach:
dump the full graph, then assert EVERY node and EVERY edge.
57 tests across 9 sections:
- Node completeness: Module(3), Function(13), Namespace(2), Property(21),
Record(1), CodeElement(8), Constructor(1) — exact sorted arrays
- Edge completeness: 22 tests covering every type+reason combination
with exact source→target pairs
- Cross-program resolution: 6 tests verifying CALL, CICS LINK/XCTL, JCL
- COPY expansion: copybook data items in RPTGEN
- Section hierarchy: exact paragraph membership per section
- Data item ownership: exact per-module breakdown
- MOVE data flow: exact read/write pairs
- JCL integration: job/step/dataset containment
- Grand totals: CALLS(22), CONTAINS(48), IMPORTS(1), ACCESSES(7)
Fixture enhancements:
- CUSTUPDT.cbl: added INIT-SECTION + PROCESSING-SECTION, PERFORM THRU
- AUDITLOG.cbl: added ENTRY "AUDITLOG-BATCH"
- RPTGEN.cbl: added EXEC CICS XCTL
Zero fuzzy assertions — every expect uses toBe(N) or toEqual([...sorted]).
* fix(cobol): add removeRelationship API + single-quote CALL/COPY/ENTRY, PERFORM keyword skip
Phase 0A: Add removeRelationship(id) to KnowledgeGraph interface and
implementation (trivial Map.delete wrapper). Required for orphan edge
cleanup in next commit.
Phase 1A (from PR #500 review, modified):
- RE_CALL and RE_COPY_QUOTED now match both "double" and 'single' quotes
- parseSingleCopyStatement in copy-expander updated for single quotes
- PERFORM_KEYWORD_SKIP set prevents UNTIL/VARYING/WITH/TEST/FOREVER
from being stored as false-positive perform targets
- Sequence number stripping uses /[^0-9 ]/ (preserves numeric seq numbers
unlike PR #500's /\S/ which stripped them)
- Normalized || to ?? for regex group extraction in copy-expander
5 new graph unit tests, all 57 COBOL integration tests pass.
* fix(cobol): RE_ENTRY single-quote + remove orphan unresolved CALLS edges
Phase 1B: RE_ENTRY regex now supports both "double" and 'single' quoted
ENTRY targets. Uses named intermediates (entryName, usingClause) with ??
operator. USING capture group shifted from [2] to [3].
Phase 1C: Second-pass resolution now collects resolved orphan edge IDs
during iteration and removes them after the loop completes, using the new
graph.removeRelationship() API. Graph no longer contains phantom
<unresolved>: edges alongside their resolved replacements. CALLS count
drops from 22 to 18 (4 orphan edges removed).
* fix(cobol): Property ID collisions + O(1) Map lookup for MOVE edges
Phase 1D+3C (atomic): Property node IDs now use composite key
filePath:section:level:name instead of filePath:name. This prevents
duplicate data item names in different sections (e.g., STATUS in both
WORKING-STORAGE and LINKAGE) from silently colliding.
New generatePropertyId() helper ensures both node creation and MOVE
edge lookup use the identical key formula. buildDataItemMap() replaces
the O(n) findDataItemNode linear scan with O(1) Map lookup, built once
per file before MOVE processing.
* feat(cobol): MOVE multi-target extraction with OF/IN qualifier filtering
MOVE X TO A B C now produces write edges for all targets, not just the
first. extractMoveTargets() helper handles OF/IN qualified names
(WS-NAME OF WS-RECORD -> target is WS-NAME), subscript stripping
(WS-TABLE(I) -> WS-TABLE), and MOVE_SKIP filtering on targets.
Data model: CobolRegexResults.moves.to:string -> targets:string[]
MOVE CORRESPONDING stays single-target per COBOL standard.
Processor MOVE loop now iterates move.targets.
* feat(cobol): COPY IN/OF library, pseudotext REPLACING, dynamic CALL, PERFORM TIMES, CICS MAP unquoted
Phase 2B: COPY ... IN/OF library-name now captured as metadata in
CopyResolution (IN and OF are synonyms per COBOL-85 standard).
Phase 2C: COPY REPLACING ==pseudotext== support. Tokenizer handles
==...== delimiters alongside "quoted" strings. Pseudotext forces EXACT
type. Two-pass applyReplacing: first pass handles space-containing/
non-identifier pseudotext via global string replace; second pass handles
identifier-level LEADING/TRAILING/EXACT. New test file
cobol-copy-expander.test.ts with 10 tests.
Phase 2E: PERFORM WS-COUNT TIMES no longer produces a false-positive
perform target (checks for TIMES keyword after captured identifier).
Phase 2F: Dynamic CALL via data item (CALL WS-PROG-NAME without quotes)
now emits a CodeElement annotation node with description 'dynamic-call'
instead of silently ignoring. Adds isQuoted:boolean to call results.
Phase 3A: CICS MAP(WS-MAP-NAME) unquoted identifiers now captured.
Phase 3B: Normalized || to ?? in copy-expander (done in Phase 1A).
* feat(cobol): nested program support — capture multiple PROGRAM-IDs per file
Phase 2D: The state machine now captures all PROGRAM-IDs, not just the
first. The primary program name stays in programName; additional nested
programs go into nestedPrograms[]. The processor creates separate Module
nodes for each nested program, contained by the outer module, and
registers them in moduleNodeIds for cross-program CALL resolution.
Paragraphs/data items are not yet scoped per-program (attributed to the
outer module) — full per-program scoping is a future enhancement that
requires END PROGRAM boundary tracking in the state machine.
* test(cobol): expand integration tests for all new language features
New fixtures:
- NESTED.cbl — two PROGRAM-IDs (OUTER-PROG, INNER-PROG) for nested
program support testing
- COPYLIB.cpy — copybook for pseudotext REPLACING test target
Modified fixtures:
- CUSTUPDT.cbl — single-quoted ENTRY 'ALTENTRY', multi-target MOVE
(WS-AMT TO FIELD-A FIELD-B), dynamic CALL WS-PROG-NAME, COPY COPYLIB
with pseudotext REPLACING, LINKAGE SECTION with LS-PARAM
- RPTGEN.cbl — PERFORM WS-COUNT TIMES (false-positive guard), unquoted
MAP(WS-MAP-NAME), additional data items WS-COUNT WS-MAP-NAME
Integration test rewritten with 62 exact assertions covering:
- 5 Module, 17 Function, 33 Property, 9 CodeElement, 2 Constructor nodes
- Nested program containment (OUTER-PROG -> INNER-PROG)
- Dynamic CALL annotation (CodeElement with cobol-dynamic-call)
- Multi-target MOVE (UPDATE-BALANCE: 2 reads, 3 writes)
- Single-quoted ENTRY (ALTENTRY under CUSTUPDT)
- PERFORM TIMES guard (WS-COUNT not in CALLS)
- Orphan unresolved edge removal (zero -unresolved edges)
- Grand totals: 21 CALLS, 68 CONTAINS, 2 IMPORTS, 10 ACCESSES
* fix(cobol): pseudotext REPLACING now applies correctly via isPseudotext flag
Root cause: ==PREFIX-== matched /^[A-Z][A-Z0-9-]*$/i (trailing hyphens
allowed), routing it to the second-pass EXACT identifier match where
PREFIX-RECORD !== PREFIX- failed silently.
Fix: Propagate isPseudotext from parseReplacingClause to CopyReplacing
interface, then use it in applyReplacing first-pass condition to force
global string replacement for all pseudotext entries regardless of
whether the content looks like an identifier.
Result: COPY COPYLIB REPLACING ==PREFIX-== BY ==WS-==. now correctly
transforms PREFIX-RECORD → WS-RECORD, PREFIX-CODE → WS-CODE, etc.
* refactor(cobol): per-program scoping via boundary tracking + line-range grouping
State machine changes (minimal, ~30 lines):
- Add RE_END_PROGRAM regex for END PROGRAM program-name. detection
- Replace nestedPrograms[] with programs[] containing startLine/endLine/
nestingDepth metadata for each PROGRAM-ID in the file
- Reset division/section/paragraph state on new PROGRAM-ID boundary
- EOF finalization flushes remaining stack entries (single-program files)
- Programs sorted by startLine (outer before inner)
Processor changes:
- Uses programs[] with line-range containment to find enclosing parent
Module for nested programs (replaces hardcoded nestedParent logic)
- programModuleIds Map tracks Module node IDs per program name
Fixture: NESTED.cbl now includes END PROGRAM lines for both programs.
Integration test: PREFIX-* Property nodes now correctly appear as WS-*
after the pseudotext REPLACING fix from the previous commit.
* feat(cobol): free-format COBOL support (>>source free)
Auto-detects >>SOURCE FREE directive in the first 500 chars and switches
to free-format line processing:
- No column-position rules (cols 1-6 are program text, not sequence area)
- Comments use *> prefix instead of col 7 indicator
- No continuation line indicator
- Strip inline *> comments
- Skip >>SOURCE directive lines
preprocessCobolSource() skips col-1-6 stripping for free-format files.
Paragraph/section regexes relaxed from fixed 7-space prefix to flexible
whitespace with case-insensitivity (/^\s*([A-Z][A-Z0-9-]+)\.\s*$/i).
EXCLUDED_PARA_NAMES expanded with COBOL verbs (GOBACK, END-READ, etc.)
to prevent false-positive paragraph detection in free-format.
Also fixes: entry-point-scoring.ts crash when language is 'cobol'
(MERGED_ENTRY_POINT_PATTERNS[language] was undefined → optional chaining).
Benchmark on ACAS 3.01 (268 GnuCOBOL free-format programs, 10MB):
- Before: 407 nodes, 393 edges (near-empty, only file nodes)
- After: 4,297 nodes, 3,612 edges, 542 clusters, 11 flows
* fix(cobol): relax data item regexes for free-format (^\s+ to ^\s*)
RE_FD, RE_DATA_ITEM, RE_ANONYMOUS_REDEFINES, and RE_88_LEVEL all used
^\s+ which requires at least 1 leading space. In free-format mode, lines
are trimmed before processing, so data items like "01 WS-FIELD PIC X."
have no leading whitespace after trimming.
Changed to ^\s* (zero or more spaces) which works for both fixed-format
(indented lines still have spaces) and free-format (trimmed lines).
ACAS benchmark (268 GnuCOBOL programs):
- Before: 4,297 nodes, 3,612 edges (paragraphs only)
- After: 13,832 nodes, 8,615 edges (+ data items, FDs, 88-levels)
* feat(cobol): 100% structural feature coverage — GO TO, SCREEN, SD/RD, SORT, SEARCH, CANCEL, Level 66
New extractions: GO TO (CALLS edges), SCREEN SECTION data items,
SD/RD alongside FD (Record nodes), SORT/MERGE USING/GIVING (ACCESSES),
SEARCH (ACCESSES), CANCEL (CALLS), Level 66 RENAMES (Property),
IS EXTERNAL/IS GLOBAL (Property description enrichment).
ACAS: 13,951 nodes | 13,193 edges | 685 clusters | 150 flows
(+53% edges from new GO TO/SORT/SEARCH/CANCEL extractions)
* feat(cobol): enriched CICS extraction — file I/O, dynamic PROGRAM, queues, HANDLE ABEND
EXEC CICS blocks now extract:
- FILE/DATASET clause: captures VSAM file name (literal or data item ref)
for READ/WRITE/REWRITE/DELETE/STARTBR/READNEXT/READPREV → ACCESSES edges
- PROGRAM clause: now handles unquoted variable references (dynamic CICS
program transfer) → CodeElement annotation with cics-dynamic-program reason
- QUEUE clause: captures TS/TD queue names from WRITEQ/READQ → ACCESSES edges
- LABEL clause: captures HANDLE ABEND error handler targets → CALLS edges
- TRANSID: now handles unquoted variable references
CodeElement descriptions enriched with all captured fields (map, program,
transid, file, queue, label).
CardDemo benchmark: +49 nodes, +33 edges from enriched CICS extraction.
* feat(cobol): complete CICS command extraction — all 7 expert recommendations
From COBOL expert agent analysis:
1. ENDBR added to isRead file command list
2. LOAD added to PROGRAM edge commands (alongside LINK/XCTL)
3. Two-word commands expanded: WRITEQ/READQ/DELETEQ TS/TD, HANDLE
ABEND/AID/CONDITION, START TRANSID
4. Queue reason differentiated: cics-queue-read/-write/-delete
5. RETURN/START TRANSID → CALLS edges to synthetic <transid> target
6. MAP → ACCESSES edges for screen traceability
7. INTO/FROM data fields extracted → ACCESSES edges to data items
Also: dataItemMap built before CICS block processing (was declared after),
CodeElement descriptions enriched with all captured CICS fields.
* test(cobol): strict exhaustive integration tests with exact edgeSet assertions
Every edge reason has exact sorted pair assertions via edgeSet(), not
just counts. Any change to extraction that adds, removes, or reorders
edges will produce a precise, descriptive failure.
Updated RPTGEN.cbl fixture with:
- GO TO EXIT-PARAGRAPH, SORT USING/GIVING, SEARCH table
- EXEC CICS READ FILE INTO, WRITEQ TS QUEUE FROM, SEND MAP FROM
- EXEC CICS HANDLE ABEND LABEL, RETURN TRANSID, XCTL PROGRAM(variable)
- ABEND-HANDLER and EXIT-PARAGRAPH paragraphs
46 tests covering 24 CALLS + 79 CONTAINS + 18 ACCESSES + 2 IMPORTS edges
across 15 distinct edge reason codes, all with exact sorted pair lists.
* fix(cobol): address 5 findings from second Claude review (compiler front-end perspective)
Finding #2: Numeric sequence numbers now stripped (changed /[^0-9 ]/ to
/\S/ in preprocessCobolSource). Lines like "000100 MAIN-PARAGRAPH." now
have cols 1-6 blanked so paragraph regex matches correctly.
Finding #11: JCL in-stream PROC ordering fixed — pre-register all PROCs
into moduleNames before step processing. Steps that EXEC a PROC defined
later in the same file now get CALLS edges.
Finding #A: PROCEDURE DIVISION USING no longer captures calling-convention
keywords (BY, VALUE, REFERENCE, CONTENT, ADDRESS, OF) as parameter names.
Finding #C: SORT/MERGE USING/GIVING now captures ALL file references
(multi-file), not just the first. Changed from single-match to section
extraction with split.
Finding #D: Section headers no longer set currentParagraph, preventing
PERFORM caller misattribution to Namespace instead of Function nodes.
* fix(cobol): address code review findings — ReDoS fix, perf, cleanup
P1 CRITICAL — ReDoS in SORT USING/GIVING:
Replaced nested-quantifier regex with safe indexOf+substring+split
approach. No backtracking possible on crafted input.
P2 — readCopy O(M) linear scan:
Added copybookByPath reverse Map for O(1) path-to-content lookup.
P3 — Dead code removal:
Deleted unused RE_SORT_USING and RE_SORT_GIVING constants.
P3 — EXCLUDED_PARA_NAMES simplification:
Replaced 20 END-* entries with startsWith('END-') prefix check.
Auto-covers future END-* verbs.
P3 — Misplaced JSDoc on removeRelationship:
Fixed comment that described removeNodesByFile instead.
Added missing JSDoc to removeNodesByFile.
Review agents: architecture-strategist, performance-oracle,
security-sentinel, code-simplicity-reviewer.
* refactor: add Cobol to SupportedLanguages with parseStrategy: standalone
New languages/cobol.ts — standalone regex processor provider with no-op
tree-sitter fields. Declares parseStrategy: 'standalone' to distinguish
from tree-sitter-based languages.
Added parseStrategy: 'tree-sitter' | 'standalone' to LanguageProviderConfig
for languages that use their own processor instead of tree-sitter.
Removed all 11 'cobol' as any casts — now uses SupportedLanguages.Cobol.
Added empty Cobol entries to entry-point-scoring and framework-detection.
* fix(cobol): 5 fixes from third Claude review + 3 regression tests
Fixes:
- Line numbers now 1-indexed in fixed-format (was 0-indexed, off-by-one
in jump-to-definition links)
- Copybook content preprocessed before COPY expansion (sequence numbers
and patch markers in copybooks no longer survive into expanded source)
- ENTRY USING filters calling-convention keywords (BY, VALUE, REFERENCE,
CONTENT, ADDRESS, OF) — same fix as PROCEDURE DIVISION USING
- SORT/MERGE trailing period stripped from USING/GIVING file tokens
- Paragraph exclusion uses exact match for SECTION/DIVISION (was substring
match that excluded valid names like CROSS-SECTION-ANALYSIS)
USING_KEYWORDS moved to module scope for reuse by both PROCEDURE DIVISION
USING and ENTRY USING handlers.
New unit tests:
- ENTRY USING BY VALUE filtering
- Paragraph names containing SECTION not excluded
- Numeric sequence numbers stripped enabling paragraph detection
* fix(cobol): address 6 findings from fourth Claude review + tests
Fourth review findings fixed:
- New #IV: PERFORM TIMES guard uses perfMatch.index instead of
line.indexOf (prevents wrong match when target appears earlier in line)
- New #V: 88-level condition values now handle single-quoted literals
('Y' no longer stored with embedded quotes)
- New #I: CANCEL edges use two-pass resolution like CALL (no longer
silently dropped when target indexed after source)
- New #3: Multi-line SORT/MERGE accumulation — sortAccum state variable
accumulates lines until period, then extracts USING/GIVING from full
statement (95% of production SORT statements span multiple lines)
- New #II: PROCEDURE DIVISION USING on split lines — pendingProcUsing
flag defers parameter capture to next line if USING not on same line
- New #6 (prior): EXCLUDED_PARA_NAMES exact match for SECTION/DIVISION
Updated fixture: RPTGEN.cbl SORT now uses multi-line format with GIVING
on separate line (period-terminated). New sort-giving integration test.
ACCESSES total: 18 → 19 (new sort-giving edge from multi-line capture).
* fix(cobol): address 4 findings from fifth Claude review
Finding #B (5 reviews old): Section/paragraph node IDs now include
enclosing program name to prevent collision when nested programs share
section/paragraph names. New findOwningProgramName() helper uses
programs[] line ranges to find the innermost enclosing program.
Finding #α: pendingProcUsing now reset in the if(procUsingMatch) branch
(was only set in else branch, could leak across nested programs).
Finding #β: RE_CALL_DYNAMIC uses negative lookbehind (?<![A-Z0-9-]) to
prevent false-positive on compound identifiers like WS-CALL OCCURS.
Finding #γ: sortAccum flushed at EOF (parallel to flushSelect and
pendingFdName EOF cleanup). Prevents silent loss of SORT USING/GIVING
relationships in truncated files.
* fix(cobol): address findings from reviews 5+6 with full test coverage
Review 5 fixes:
- #α: pendingProcUsing reset in if(procUsingMatch) branch
- #β: RE_CALL_DYNAMIC negative lookbehind prevents WS-CALL false positive
- #γ: sortAccum flushed at EOF for truncated files
- #B: Section/paragraph IDs include owning program name
Review 6 fixes:
- #P: sectionNodeIds/paraNodeIds maps use program-scoped keys
(PROGNAME:NAME). New scopedParaLookup/scopedCallerLookup helpers.
findContainingSection updated with programs parameter.
- #Q: RETURNING added to USING_KEYWORDS for COBOL 2002+
- #R: RE_PERFORM matches both THRU and THROUGH via alternation
New unit tests (6):
- PERFORM THROUGH captures thruTarget
- PROCEDURE DIVISION USING RETURNING filters keyword
- RE_CALL_DYNAMIC no false-match on WS-CALL compound identifier
- Multi-line SORT captures USING/GIVING from continuation lines
- PROCEDURE DIVISION USING on split line via pendingProcUsing
- Copybook preprocessing strips sequence numbers
* fix(cobol): address findings from seventh Claude review + 3 tests
Review 7 fixes:
- #i: findContainingSection only updates best when lookup succeeds
(prevents undefined overwriting valid parent section)
- #ii: RE_PROC_SECTION handles segment numbers (SECTION 30.)
- #III: procedureUsing now stored per-program on boundary stack
entries, propagated to programs[] output. Inner programs no longer
overwrite outer program's parameters.
- #δ: Dynamic CANCEL (CANCEL variable) now creates CodeElement
annotation node, matching dynamic CALL behavior. RE_CANCEL_DYNAMIC
with negative lookbehind. cancels[] gains isQuoted field.
- #Q: RETURNING added to USING_KEYWORDS (already in prev commit)
- #R: PERFORM THROUGH already fixed (THRU|THROUGH alternation)
New unit tests:
- Nested programs carry per-program procedureUsing
- SECTION with segment number detected
- Dynamic CANCEL via data item captured with isQuoted=false
* feat(cobol): link PROCEDURE DIVISION USING to LINKAGE data items + close 4 findings
Finding #10 FIXED: procedureUsing parameters now create ACCESSES edges
with reason 'cobol-procedure-using' from Module to matching LINKAGE
SECTION Property nodes. This exposes the program's parameter contract
in the graph (e.g., AUDITLOG → LS-CUST-ID, AUDITLOG → LS-AMOUNT).
Findings closed by expert agent consensus:
- #6 COPY IN library: WONTFIX — captured metadata, no universal
library-to-directory mapping exists. Field costs nothing and is useful
for library queries.
- #14 SQL DELETE: WONTFIX — DB2 requires FROM; existing FROM pattern
handles it. Bare DELETE would risk false positives.
- #E OCCURS DEPENDING ON: WONTFIX — runtime sizing concern, not
structural. The static occurs count is sufficient for indexing.
All 39 findings from 7 Claude reviews now resolved or closed.
* fix(cobol): resolve 48 review findings across 9 review cycles
Ninth deep review resolved all remaining COBOL parser gaps identified
by 5 specialist agents (COBOL expert, architecture strategist,
TypeScript reviewer, security sentinel, code simplicity reviewer).
Fixes (P1 — critical):
- SELECT OPTIONAL now correctly skips OPTIONAL keyword (C1)
- RETURNING params excluded from PROCEDURE DIVISION USING list (C7)
- SORT GIVING no longer captures clause keywords as file names (C5)
- Extract flushSort() helper eliminating 40-line duplication (S2)
- Flush unclosed EXEC blocks at EOF matching SORT/SELECT pattern (S3)
- Guard undefined map key in jcl-processor moduleNames (S1)
- Add MAX_TOTAL_EXPANSIONS=500 to prevent exponential COPY breadth (S4)
Fixes (P2 — important):
- Quote-aware stripInlineComment for | and *> in string literals (C2+C3)
- Fixed-format literal continuation now handles quoted strings (C6)
- PROGRAM-ID detected regardless of division state for siblings (C9)
Fixes (P3 — cleanup):
- EXEC SQL INTO restricted to INSERT INTO to avoid FETCH false-pos (C8)
- Copy expander line numbers fixed from 0-based to 1-based (C11)
- Remove dead code: inInStreamProc, fileIsLiteral, expansionDepth (S7-S10)
Also fixes 8th-review findings: nested program CONTAINS attribution,
multi-PERFORM on same line, INPUT/OUTPUT PROCEDURE IS in SORT,
GO TO DEPENDING ON multi-target, MOVE CORR abbreviation, per-program
procedureUsing ACCESSES edges.
Tests: 145 COBOL tests passing (59 integration + 86 unit)
Benchmarks: CardDemo 12,323 nodes/8,893 edges (7.4s)
ACAS 14,016 nodes/15,452 edges (9.3s, -9% faster)
* docs(cobol): update documentation for ninth review cycle fixes
Update all 4 COBOL documentation files to reflect the 16 fixes
from the ninth review cycle:
- regex-extraction.md: quote-aware comment stripping, SELECT OPTIONAL,
RETURNING exclusion, SORT_CLAUSE_NOISE filter, flushSort() helper,
GO TO multi-target, PROGRAM-ID division-independent detection
- copy-expansion.md: MAX_TOTAL_EXPANSIONS=500 breadth guard, 1-based
line numbers, removed expansionDepth/warnedCircular param
- deep-indexing.md: GO TO DEPENDING ON, INPUT/OUTPUT PROCEDURE IS,
MOVE CORR edge reasons, INSERT INTO restriction, literal continuation
- performance.md: updated benchmarks (CardDemo 12,323n/8,893e/7.4s,
ACAS 14,016n/15,452e/9.3s), COPY breadth guard
* fix(cobol): resolve 10th review findings — nested program edge attribution
Fix 6 findings from the 10th review (PR #498 comment #4132201110):
#A+#F: All CALL/CANCEL/CICS/ENTRY/SQL/SEARCH/file-declaration edges
now use owningModuleId() for nested program attribution instead of
the outer program's parentId. Added helper function owningModuleId()
to centralize the pattern.
#B: Added USING and GIVING to SORT_CLAUSE_NOISE set to prevent MERGE
USING + OUTPUT PROCEDURE from capturing clause keywords as file names.
#C: INPUT/OUTPUT PROCEDURE regex now captures optional THRU/THROUGH
range end paragraph, mirroring RE_PERFORM's THRU support.
#D: scopedCallerLookup fallback now uses programModuleIds.get(pgm)
instead of parentId, so PERFORM/MOVE/GOTO in nested programs with
unresolvable paragraphs fall back to the correct inner module.
#E: pendingProcUsing only set when PROCEDURE DIVISION line is NOT
period-terminated, preventing false USING expectation.
Tests: 145 passing | TypeScript clean
* fix(cobol): resolve 10th review findings — nested program edge attribution
Fix 6 findings from the 10th review (PR #498 comment #4132201110):
#A+#F: All CALL/CANCEL/CICS/ENTRY/SQL/SEARCH/file-declaration edges
now use owningModuleId() for nested program attribution instead of
the outer program's parentId. Added helper function owningModuleId()
to centralize the pattern.
#B: Added USING and GIVING to SORT_CLAUSE_NOISE set to prevent MERGE
USING + OUTPUT PROCEDURE from capturing clause keywords as file names.
#C: INPUT/OUTPUT PROCEDURE regex now captures optional THRU/THROUGH
range end paragraph, mirroring RE_PERFORM's THRU support.
#D: scopedCallerLookup fallback now uses programModuleIds.get(pgm)
instead of parentId, so PERFORM/MOVE/GOTO in nested programs with
unresolvable paragraphs fall back to the correct inner module.
#E: pendingProcUsing only set when PROCEDURE DIVISION line is NOT
period-terminated, preventing false USING expectation.
Tests: 145 passing | TypeScript clean
* fix(cobol): resolve 11th review findings — final nested program + multi-CALL gaps
#1: scopedCallerLookup(null) now uses owningModuleId(lineNum) instead
of parentId, fixing PERFORM/MOVE/GOTO before first paragraph in nested
programs.
#2+#3: CALL and CANCEL extraction now uses matchAll (global flag) to
capture multiple occurrences on the same line. Dynamic CALL/CANCEL
checked independently instead of in else branch.
#4: SORT/MERGE ACCESSES edge IDs now use owningModuleId(sort.line)
instead of parentId for nested program correctness.
#5: preprocessCobolSource free-format detection now uses first 10 lines
(consistent with extractCobolSymbolsWithRegex threshold).
#6: EXCLUDED_PARA_NAMES expanded with DISPLAY, ACCEPT, WRITE, READ,
REWRITE, DELETE, OPEN, CLOSE, RETURN, RELEASE, SORT, MERGE to prevent
false-positive paragraph detection on isolated verbs.
Also removed unused GraphNode import from cobol-processor.ts.
Tests: 145 passing | TypeScript clean
* docs(cobol): deepened full language coverage plan with research findings
3 research agents analyzed Phase 1-2 features and graph value ranking.
Key findings: cobol-call-using is #1 edge type (9.2/10); multi-line
accumulation is dominant challenge; DECLARATIVES is lowest-risk Phase 2
item; SET TO TRUE covers 80-90% of SET usage.
* feat(cobol): implement Phase 1 — high-value data flow edges
4 new extraction features that create new ACCESSES and IMPORTS edges:
1.1: EXEC SQL INCLUDE -> IMPORTS edges with reason 'sql-include'
Handles unquoted (SQLCA), quoted ('DBRMLIB.MEMBER'), and
underscored (CUST_TBL_DCL) member names.
1.2: CALL USING parameter extraction -> ACCESSES edges
Extracts parameters from CALL USING clause, filtering BY/REFERENCE/
CONTENT/VALUE/ADDRESS/OF/LENGTH/OMITTED keywords. Creates
'cobol-call-using' ACCESSES edges (graph value: 9.2/10).
1.4: OCCURS DEPENDING ON -> ACCESSES edges with reason 'cobol-depends-on'
Extended OCCURS regex captures DEPENDING ON field with subscript
stripping. Creates dependency edge from table to controlling field.
1.5: VALUE clause for standard data items
Extracts VALUE from data item clauses: quoted strings with type
prefix (X/N/G/B), ALL literals, numerics (incl negative/decimal),
and figurative constants. Populates Property node values.
Tests: 145 passing (+2 ACCESSES from CALL USING) | TypeScript clean
* feat(cobol): implement Phase 2 — DECLARATIVES, SET, INSPECT, EXEC DLI
4 new extraction features for error handling, data flow, and IMS/DB:
2.1: EXEC DLI (IMS/DB) -> CodeElement + ACCESSES edges
Accumulates EXEC DLI blocks like EXEC SQL. Parses DLI verbs
(GU, GN, ISRT, REPL, DLET, CHKP, SCHD, TERM). Extracts
SEGMENT, PCB, INTO/FROM, PSB. Creates dli-{verb} ACCESSES
edges to <ims>:segment Record nodes.
2.2: DECLARATIVES / USE AFTER EXCEPTION -> ACCESSES edges
Tracks inDeclaratives state. Detects USE AFTER STANDARD
EXCEPTION ON file-name. Creates cobol-error-handler ACCESSES
edge from handler section to file Record.
2.3: SET statement -> ACCESSES edges
Detects SET TO TRUE (80-90% of SET usage) and SET index
TO/UP BY/DOWN BY. Creates cobol-set-condition / cobol-set-index
write edges + cobol-set-read for identifier values.
2.4: INSPECT -> ACCESSES edges with multi-line accumulator
Accumulates INSPECT until period (like SORT). Extracts inspected
field + tally counters. Creates cobol-inspect-read/write/tally
edges. Form detection: tallying/replacing/converting/combined.
Preprocessor: 1398 -> 1597 LOC (+199). Tests: 145 passing.
* feat(cobol): implement Phase 3 — completeness fixes
6 partial features fixed to first-class support:
3.1: CALL RETURNING -> ACCESSES write edge (cobol-call-returning)
3.2: SELECT OPTIONAL flag preserved in FileDeclaration + Record node
3.3: ALTERNATE RECORD KEY extraction (matchAll for multiple keys)
3.4: COMMON attribute on nested programs (RE_PROGRAM_ID extended)
3.5: IS EXTERNAL / IS GLOBAL as first-class boolean properties
(removed usage string hack)
3.6: AUTHOR / DATE-WRITTEN mapped to Module node description
Tests: 145 passing | TypeScript clean
* feat(cobol): implement Phase 4 — INITIALIZE + metadata completeness
4.1: INITIALIZE statement -> ACCESSES write edge (cobol-initialize)
4.2: DATE-COMPILED and INSTALLATION paragraphs extracted and mapped
to Module node description alongside existing AUTHOR/DATE-WRITTEN
All 4 plan phases complete. Coverage: ~95% (up from 71.9%).
Tests: 145 passing | TypeScript clean
* test(cobol): add 24 unit tests for Phase 1-4 features
Coverage for all new extraction features:
Phase 1 (8 tests):
- EXEC SQL INCLUDE (unquoted, quoted, underscored)
- CALL USING (simple, mixed modes, ADDRESS OF, OMITTED)
- CALL RETURNING
- OCCURS DEPENDING ON
- VALUE clause (string, numeric, figurative constant)
Phase 2 (10 tests):
- EXEC DLI GU/ISRT/SCHD (verb, segment, PCB, INTO, FROM, PSB)
- DECLARATIVES USE AFTER EXCEPTION (single + multiple sections)
- SET TO TRUE, SET index UP BY
- INSPECT TALLYING, INSPECT REPLACING
Phase 3-4 (6 tests):
- SELECT OPTIONAL flag
- ALTERNATE RECORD KEY
- PROGRAM-ID IS COMMON
- IS EXTERNAL / IS GLOBAL booleans
- INITIALIZE extraction
- Full programMetadata (AUTHOR, DATE-WRITTEN, DATE-COMPILED, INSTALLATION)
Total: 168 tests passing (145 + 24 - 1 removed duplicate)
* fix(cobol): use /\r?\n/ split for Windows CRLF compatibility
All 4 COBOL source files now split on /\r?\n/ instead of '\n' to
handle CRLF line endings on Windows. Previously, trailing \r in
lines caused RE_GOTO's $ anchor to fail on multi-line GO TO
DEPENDING ON statements, producing only 1 goto edge instead of 4.
Files fixed: cobol-preprocessor.ts (2 sites), cobol-processor.ts,
jcl-parser.ts, cobol-copy-expander.ts
Tests: 168 passing | TypeScript clean
* fix(cobol): resolve 12th review — dynamic CALL/CANCEL dedup + trailing anchors
#1+#2: Removed incorrect hasQuotedCall/hasQuotedCancel deduplication
guards. RE_CALL_DYNAMIC and RE_CANCEL_DYNAMIC require [A-Z] after
CALL/CANCEL, so they CANNOT match quoted targets — the guards were
both unnecessary and actively harmful, suppressing dynamic CALL/CANCEL
in ON EXCEPTION patterns.
#3+#5: Changed RE_CALL_DYNAMIC and RE_CANCEL_DYNAMIC trailing anchor
from (?:\s|\.) to (?=\s|\.|$) (lookahead). The consuming anchor
failed when the identifier was the last token on a physical line.
Tests: 168 passing | TypeScript clean
* feat(cobol): add CALL accumulator + fix SORT double-statement (#4, #6)
Finding #4: Multi-line CALL USING accumulator
Added callAccum state variable that accumulates CALL statements
spanning multiple physical lines until period or END-CALL is found.
Uses flushCallAccum() to re-extract CALL target + USING parameters
from the full accumulated statement. This fixes the silent loss of
ACCESSES parameter edges when USING appears on lines after CALL.
Finding #6: SORT double-statement on same line
After flushSort(), the code now falls through to re-check the
current line for a new SORT/MERGE start (was previously blocked
by the sortAccum === null check evaluating before flushSort ran).
Also fixed: used non-global regex for CALL detection test to avoid
the classic global regex .test() lastIndex bug.
Tests: 168 passing (+1 ACCESSES from multi-line CALL USING)
* fix(cobol): resolve 13th review — CICS LOAD, USING extraction, file scoping
#1: CICS LOAD unresolved edge no longer silently deleted in second pass.
Changed narrow cics-link/cics-xctl check to catch-all pattern:
rel.reason?.startsWith('cics-') && rel.reason.endsWith('-unresolved')
#2: flushCallAccum USING extraction now stops before COBOL statement
verbs (INSPECT, SEARCH, SORT, MERGE, DISPLAY, ACCEPT, MOVE, PERFORM,
GO TO, CALL, IF, EVALUATE). Prevents absorbing adjacent statements
as false USING parameters in legacy pre-COBOL-85 code without END-CALL.
#3: CICS FILE Record nodes now globally-scoped (<cics-file>:FILENAME)
instead of per-file-scoped. Enables cross-program CICS file access
analysis, consistent with SQL table scoping (<db>:TABLE).
#4: callAccum pre-check regex now has (?<![A-Z0-9-]) lookbehind to
prevent false activation on compound identifiers like WS-CALL-FLAG.
Tests: 168 passing | TypeScript clean
* fix(cobol): resolve 14th review — callAccum false paragraph + Area A guard
#1: callAccum continuation lines now check for COBOL statement verb
starts (GO TO, PERFORM, MOVE, etc.) and paragraph/section headers.
If detected, the CALL is flushed as-is and the line processed
normally — prevents false paragraph detection and currentParagraph
corruption from lines like "WS-ADDR." being treated as paragraphs.
#4: callAccum pre-check now guarded by currentDivision === 'procedure'
to prevent unnecessary activations in DATA DIVISION.
#5: Fixed-format paragraph detection now rejects lines with >7 leading
spaces (Area B indentation) as paragraph candidates. Paragraph
names in fixed-format must start in Area A (col 8-11, max 7 spaces).
Free-format mode is unaffected.
Tests: 168 passing | TypeScript clean
* fix(cobol): resolve 15th review — callAccum Area A + verb boundary fixes
#A: Column-position-aware paragraph detection in callAccum flush.
#B: inspectAccum early-flush on paragraph/section/verb headers.
#C: Verb boundary \b → (?:\s|$) prevents MOVE-COUNT false flush.
* test(cobol): add 17 edge-case regression tests + fix USING verb boundary
17 new tests covering all recurring review patterns:
Multi-line CALL USING (7 tests):
- Parameters on separate continuation lines (IBM mainframe style)
- No absorption of INSPECT/GO TO/paragraphs following CALL
- END-CALL scope terminator
- Hyphenated identifiers (MOVE-COUNT) not triggering false flush
- Dual quoted+dynamic CALL on same line (ON EXCEPTION)
Nested program attribution (2 tests):
- CALL in inner program within inner line range
- PERFORM before first paragraph has null caller
CRLF compatibility (1 test):
- GO TO DEPENDING ON with \r\n line endings
Area A paragraph detection (2 tests):
- Area B (>7 spaces) rejected; Area A (7 spaces) accepted
SORT/MERGE (1 test): COLLATING SEQUENCE keywords not captured
PROCEDURE USING (2 tests): RETURNING excluded, period-terminated
Comment stripping (1 test): pipe in quoted string preserved
SELECT OPTIONAL (1 test): correct file name, not OPTIONAL keyword
Bug fix: USING extraction regex verb terminators changed from
\bVERB\b to \bVERB(?=\s|$) in flushCallAccum — prevents truncation
on hyphenated identifiers like MOVE-COUNT, PERFORM-LIMIT.
Total: 185 tests passing
* test(cobol): add 32 comprehensive edge-case regression tests
13 new describe blocks covering all extraction features:
- EXEC DLI: no-SEGMENT, multi-line accumulation (2 tests)
- SET: multiple targets, DOWN BY, TO numeric (3 tests)
- INSPECT: CONVERTING, multiple counters, tallying-replacing,
paragraph flush during accumulation (4 tests)
- DECLARATIVES: no-STANDARD keyword, I-O mode, post-END paragraphs (3)
- COPY REPLACING: pseudotext deletion ==OLD== BY ==== (1 test)
- VALUE: hex literal, negative numeric, ALL literal (3 tests)
- OCCURS: TO range, fixed-size without DEPENDING ON (2 tests)
- Dynamic CALL/CANCEL: end-of-line, multiple CANCELs (3 tests)
- EXEC SQL: INCLUDE skips tables, SELECT INTO host vars, host
variable extraction (3 tests)
- INITIALIZE: target and caller context (1 test)
- Nested programs: sibling scoping, PROGRAM-ID without ID DIV (2)
- EXEC EOF flush: unclosed EXEC SQL flushed (1 test)
- Multi-PERFORM: IF/ELSE dual PERFORM on single line (1 test)
- IS EXTERNAL: USAGE not polluted by external flag (1 test)
Total: 215 tests passing
* fix(cobol): resolve 16th review — CANCEL in CALL block + USING boundary
#1: flushCallAccum now extracts CANCEL statements from within CALL
ON EXCEPTION blocks. Adds RE_CANCEL + RE_CANCEL_DYNAMIC matchAll
passes alongside existing CALL extraction.
#2: Added \bCANCEL(?=\s|$) to USING lookahead regex to prevent CANCEL
keyword being captured as false USING parameter.
#3: Multi-line CALL start now returns immediately to prevent the CALL
start line from simultaneously feeding sortAccum/inspectAccum.
#6: Division transitions now flush all active accumulators (callAccum,
sortAccum, inspectAccum) to prevent state leakage across programs.
Also added CANCEL to callAccum flush trigger verb list.
Tests: 215 passing | TypeScript clean
* refactor(cobol): extract shared verb constants + resolve 17th review
Extract COBOL_STATEMENT_VERBS, RE_STATEMENT_VERB_START, and
RE_USING_PARAMS as shared constants — eliminates 4 duplicated
25-verb regex patterns.
17th review: #1 flushCallAccum before EXEC entry, #2 inspectAccum
verb parity via shared constant.
Tests: 215 passing | TypeScript clean
* test(cobol): replace all fuzzy assertions with exact toBe checks
Replaced 7 toBeGreaterThan/toBeLessThan/toBeGreaterThanOrEqual
assertions with exact toBe values:
- dataItems.length: >= 3 → toBe(3)
- calls.length: >= 1 → toBe(1)
- calls[0].line: range check → toBe(10)
- programs[].startLine/endLine: comparison → exact values
- innerA.endLine/innerB.startLine: comparison → exact values
Also added 11 new edge-case tests (accumulator flush on EXEC/division
transitions, free-format, CANCEL in CALL block, SORT THRU, verb
flush, integration).
226 tests passing — zero fuzzy assertions remain.
* fix(cobol): resolve 19th review + 15 accumulator flush tests
Fixes:
#1: END PROGRAM flushes callAccum/sortAccum/inspectAccum
#2: PROGRAM-ID sibling path flushes all accumulators
#3: Added COMPUTE/ADD/SUBTRACT/MULTIPLY/DIVIDE/STRING/UNSTRING
to COBOL_STATEMENT_VERBS (now 32 verbs)
Tests (15 new):
- END PROGRAM flush: single + nested programs (2)
- PROGRAM-ID sibling flush (1)
- Arithmetic verb flush: COMPUTE/ADD/SUBTRACT/MULTIPLY/DIVIDE (5)
- String verb flush: STRING/UNSTRING (2)
- Arithmetic not captured as false USING params (1)
- SORT flushed at END PROGRAM (1)
- INSPECT flushed at END PROGRAM (1)
- All with exact toBe assertions (2)
Total: 239 tests passing | Zero fuzzy assertions
* fix(cobol): resolve 20th review — INITIALIZE multi-target + 2 tests
Finding 1: INITIALIZE now captures multiple targets with REPLACING
clause keyword filtering. Regex changed to lazy match stopping at
REPLACING/WITH/period boundary. Targets split on whitespace and
filtered against INITIALIZE_CLAUSE_KEYWORDS set.
Tests (2 new):
- INITIALIZE multi-target: WS-CUSTOMER WS-ORDER WS-LINE-ITEM → 3
- INITIALIZE with REPLACING: only WS-RECORD captured, not keywords
Total: 241 tests passing | TypeScript clean
* fix: close remaining Dart language support gaps
Four issues that were not addressed in PR #204:
1. extractFunctionName: add function_signature/method_signature handlers
and add both to FUNCTION_NODE_TYPES. Without this, findEnclosingFunctionId
cannot resolve Dart function scopes — all calls inside Dart functions
have no sourceId, breaking CALLS edge attribution.
2. formal_parameter_list: add to paramListTypes in extractMethodSignature.
Dart's tree-sitter grammar uses this node type (not formal_parameters),
so parameter counting returns 0 for all Dart functions.
3. Write-access queries: add @assignment patterns for obj.field = value
and this.field = value. Without these, no ACCESSES write edges are
emitted for Dart code.
4. initialized_identifier guard in extractDartDeclaration: comma-separated
declarations (String a, b, c) produce initialized_identifier nodes
which are in DART_DECLARATION_NODE_TYPES but were unhandled — the type
lives on the parent node.
Also adds Dart column to the feature matrix in type-resolution-system.md.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(dart): field-type resolution, call attribution, import resolution, and integration tests
Fixes five Dart language support gaps with integration tests and
architectural alignment:
**Tree-sitter queries** — Add field declaration patterns for typed and
nullable class fields (`String name = ''`, `String? name`). Without
these, Dart class fields were invisible to the pipeline (zero Property
nodes, zero HAS_PROPERTY edges).
**Import resolution** — Dart relative imports (`import 'models.dart'`)
don't use a leading `./`. The standard resolver only recognises paths
starting with `.` as relative; bare paths fell through to a Java-style
dot-to-slash conversion that mangled `models.dart` into `models/dart`.
Fix: prepend `./` before calling resolveStandard.
**Call attribution** — Dart's tree-sitter grammar places `function_body`
as a sibling of `function_signature`, not as a child wrapping both. The
`findEnclosingFunction` parent-walk never found the function because the
call lives inside `function_body` which is a sibling of the signature.
Fix: add `enclosingFunctionFinder` hook to LanguageProvider interface
(following the same strategy pattern as `labelOverride`), with the
Dart-specific logic in `languages/dart.ts`. Both `parse-worker.ts` and
`call-processor.ts` consume the hook generically — no Dart-specific code
in the generic processors.
**Receiver chain extraction** — Add `unconditional_assignable_selector`
to `MEMBER_ACCESS_NODE_TYPES` so `inferCallForm` returns `'member'` for
Dart method calls. Add Dart-specific receiver extraction blocks in
`extractReceiverName`, `extractReceiverNode`, and a `selector` handler
in `extractMixedChain` for Dart's flat sibling-selector model (vs the
nested member-expression model used by all other languages).
**Integration tests** — New `dart.test.ts` with field-type resolution
and call-result-binding describe blocks. Fixtures: `dart-field-types/`
(models.dart + app.dart) and `dart-call-result-binding/` (models.dart +
app.dart). 9 passing tests, 1 skipped (ACCESSES edges for field reads
depend on type-env parameter binding propagation — tracked for follow-up).
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
Dart was added as the 14th supported language in PR #204 but the README
was not updated. Adds Dart row to the supported languages table and
updates the language count from 13 to 14.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(web): enable Cypher queries when connected to backend server
Route queries through HTTP API in backend mode instead of checking local WASM database.
Made-with: Cursor
* chore: add Maven/Gradle wrapper files to default ignore list
Add build wrapper scripts and directories to hardcoded ignore lists:
- Directories: .mvn, .gradle, gradle
- Files: mvnw, mvnw.cmd, gradlew, gradlew.bat
These are build infrastructure files, not source code.
Made-with: Cursor
* ci: re-trigger CI (Windows flaky timeout)
Made-with: Cursor
* feat: add more node types in filter panel
* feat: add more node types in filter panel
* revert additional changes
* test(web): add unit tests for filter panel node types
- FILTERABLE_LABELS: verify new types (Enum, Type, Decorator, Variable)
have colors, sizes, and no duplicates
- Filter panel icons: verify every filterable label has an icon mapped
and all icons are exported from lucide-icons
- Color legend: verify new types are included, ordered correctly, and
are a subset of FILTERABLE_LABELS
Made-with: Cursor
The shape-check-regression test uses withTestLbugDB but was running in
the default vitest project with parallel forks, causing LadybugDB
file-lock conflicts on Windows CI. Move it to the lbug-db project
(sequential execution) and exclude from default.
Follows up on #501.
* feat: add PHP response shape extraction for json_encode patterns
Adds extractPHPResponseShapes() to detect response keys from PHP
json_encode() calls with associative array literals. Supports:
- Short array syntax: json_encode(['key' => value])
- Long array syntax: json_encode(array('key' => value))
- Error classification via http_response_code() and header() status
- exit;/die; boundary detection to prevent cross-block status leaking
- Nested array filtering (only top-level keys extracted)
Pipeline integration dispatches PHP files to the new extractor.
Verified on collector project: 10 PHP routes now show responseKeys.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address review — exit boundary, die; offset, CGI Status header
- Replace lastIndexOf('exit;')/lastIndexOf('die;') with regex that
matches exit(N), exit(0), die('msg'), die($var) as boundaries
- Fixes die; off-by-one (was slicing at +5 for a 4-char keyword)
- Add header('Status: NNN') CGI/FastCGI format detection
- Add 3 regression tests for the fixed bugs
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* refactor: extract shared helpers, remove duplicate test block
- Extract lastMatchGroup() and buildShapeResult() to eliminate repeated
patterns in both JS/TS and PHP extractors
- Simplify detectPHPStatusCode to use ?? chaining with lastMatchGroup
- Remove duplicate 9-test PHP describe block (kept the 12-test version)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* test: add PHP response shape integration tests
Adds a PHP fixture (api/items.php, api/submit.php) with multiple
json_encode patterns and a pipeline integration test verifying:
- Route nodes created for PHP endpoints
- responseKeys/errorKeys correctly extracted and separated
- exit(N)/die() boundaries respected
- HANDLES_ROUTE edges point to correct PHP handler files
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* ci: E2E workflow, web typecheck job, pre-commit hook, test suite
CI:
- ci.yml consolidated to reference ci-tests.yml
- ci-quality.yml: add typecheck-web job for gitnexus-web/
- ci-e2e.yml: E2E workflow with dorny/paths-filter (web changes only)
- ci-report.yml: remove dead integration-reports references
- CI gate allows skipped E2E status
- .gitignore: playwright artifacts, eval test artifacts
Pre-commit hook:
- .githooks/pre-commit: typecheck + unit tests for both packages
- Activated via git config core.hooksPath in prepare script
Test infrastructure:
- Vitest + React Testing Library: 58 unit tests
(graph, server-connection, mermaid, settings, constants, utils, paths)
- Playwright E2E: 5 tests + manual recording harness
- vitest.config from vitest/config, engines.node >= 20
- Playwright artifacts retain-on-failure
- wait-on in devDependencies
- vitest/coverage-v8 aligned with vitest 4.x
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* chore: update gitnexus-web package-lock.json
Reflects devDependency additions (vitest, playwright, wait-on,
@testing-library, etc.) from package.json changes in this PR.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(e2e): add missing process-list-loaded testid, increase CI timeouts
- Add data-testid="process-list-loaded" to ProcessesPanel (E2E tests
were waiting for an element that didn't exist)
- Increase server connect timeouts from 5s to 10s for slower CI
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): run gitnexus-web unit tests in CI, remove unused variable
- Add gitnexus-web npm ci + vitest run to ci-tests.yml so web unit
tests are gated by the CI status check (were only running locally)
- Remove unused IS_PLAYWRIGHT_AUTOMATION variable from E2E spec
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(e2e): add process-row testid, wait for networkidle on page load
- Add data-testid="process-row" to ProcessItem component (E2E tests
referenced it but it didn't exist in the source)
- Use waitUntil: 'networkidle' on page.goto to ensure Vite dev server
is fully ready before interacting (fixes first-test timeout in CI)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(e2e): add process-view-button and process-highlight-button testids
E2E tests referenced these data-testid attributes but they didn't
exist in ProcessItem. All 6 E2E testids now have matching source
elements: status-ready, process-list-loaded, process-row,
process-view-button, process-highlight-button, server-url-input.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(e2e): remove networkidle — Vite HMR WebSocket prevents it from resolving
networkidle waits for zero network activity for 500ms, but Vite's HMR
WebSocket stays open permanently, causing page.goto to timeout at 60s
on all tests after the first. The explicit toBeVisible waits on UI
elements are sufficient and deterministic.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(e2e): wait for Server button visibility, add CI retry, all 5 tests pass locally
Root cause: test 1 clicked the Server button before React hydrated,
so the tab content never rendered and the input wasn't found.
Fixes:
- Wait for Server button toBeVisible before clicking
- Increase input wait to 15s
- Remove networkidle (Vite HMR WebSocket prevents it from resolving)
- Add retries: 1 in CI for transient cold-start flakiness
Verified locally: all 5 E2E tests pass, 198 unit tests pass, typecheck clean.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): tolerate LadybugDB native crash during analyze step
gitnexus analyze can crash with "double free or corruption" (known
issue #273) during the LadybugDB native addon shutdown. The index is
usually written successfully before the crash. The workflow now:
1. Allows analyze to exit non-zero with a warning
2. Verifies .gitnexus index was actually created
3. Only fails if no index exists (real failure)
All tests verified locally: 198 unit, 5 E2E pass, typecheck clean.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): fix shell quoting in analyze step, simplify to || true
The previous echo string had special characters that broke bash
quoting in GitHub Actions. Simplified to: analyze || true, then
check if .gitnexus exists.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* docs: add agent development framework, GitHub templates, eval refactor
Agent framework (layered docs for AI-assisted contributions):
- AGENTS.md: canonical instructions, impact analysis, MCP tools
- CLAUDE.md: Claude Code-specific deltas and hooks
- GUARDRAILS.md: safety boundaries, non-negotiables, escalation
- ARCHITECTURE.md: monorepo layout, data flow map
- TESTING.md: test structure, commands, categories
- RUNBOOK.md: copy-paste operations for dev/CI/MCP
- llms.txt: minimal LLM context pointer
Editor integration:
- .cursor/index.mdc + rules/100-monorepo.mdc
GitHub templates:
- PR template with areas-touched checkboxes
- Bug report + feature request issue forms
Eval harness:
- Refactored mcp_bridge, tool_registry, constants
- Error sanitization utilities
- Property-based tests via Hypothesis
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(eval): use format_exception instead of format_exc in sanitize_exception
format_exc() returns the currently handled exception traceback, which
may be unrelated if called outside an active except block. Using
format_exception(type(exc), exc, exc.__traceback__) reliably captures
the passed exception's traceback.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* docs: update CONTRIBUTING.md and TESTING.md for current CI/hook setup
- CONTRIBUTING.md: add gitnexus-web typecheck command, pre-commit hook
checklist item
- TESTING.md: add gitnexus-web typecheck command, pre-commit hook
section (husky), update CI integration to list actual workflow files
(ci-quality, ci-tests, ci-e2e)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* docs: update testing docs to reflect CI/E2E changes from PR #486
- AGENTS.md: update test counts (CLI ~2000 unit, ~1850 integration),
add gitnexus-web testing section (198 unit, 5 E2E with commands)
- RUNBOOK.md: fix Node requirement to >=20, fix E2E local repro command
- TESTING.md: E2E uses data-testid selectors + real servers, not mocks
- .cursor/rules/100-monorepo.mdc: add web test/E2E commands
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* docs: address context engineering review — deduplicate tokens, expand Cursor rules
- Remove ~100-line gitnexus:start block from CLAUDE.md (was duplicated from AGENTS.md)
- Fix gitnexus:start block inlined inside AGENTS.md Reference Docs bullet (doubled)
- Replace CLAUDE.md scope table with pointer to AGENTS.md (single source of truth)
- Expand .cursor/index.mdc with 5 non-negotiable safety rules for always-on context
- Add .cursor/rules/200-eval.mdc with Python/eval commands (glob-scoped to eval/**)
- Improve llms.txt with priority annotations and descriptions
- Bump version headers to 1.2.0, last-reviewed to 2026-03-24
Saves ~1,400 tokens/session with zero information loss.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
- Add process.stdin.isTTY guard before --review prompt to prevent CI hangs
- Rename misleading --verbose e2e test to reflect it checks help output
- Replace DEBUG with GITNEXUS_VERBOSE for error stack traces
Made-with: Cursor
Spawn actual CLI process to verify:
- wiki --help surfaces all new flags
- wiki on non-git directory exits with code 1
- wiki on non-indexed repo fails with "No GitNexus index"
- --provider cursor skips API key prompt in non-TTY mode
- --verbose is accepted as valid flag
Made-with: Cursor
- Cache detectCursorCLI() result to avoid spawning `agent --version`
on every LLM call
- Fix stale JSDoc in cursor-client.ts (no longer uses stdin or
stream-json)
- Remove unused WikiOptions fields (model, baseUrl, apiKey) that were
passed but never read by WikiGenerator
- Fix inconsistent progress callback phase tracking in --review
continuation path
- Add wiki CLI help test covering --provider, --review, --verbose flags
Made-with: Cursor
Add Cursor headless CLI as a 4th provider option for wiki generation,
allowing users to leverage their Cursor subscription for wiki pages.
- Add --provider cursor flag and cursor-client.ts
- Add --review flag for interactive module tree editing
- Add --verbose flag for debugging
- Improve module tree generation (flatten single-child, unique slugs)
- Prevent DB timeout during long LLM calls
Usage: gitnexus wiki --provider cursor --model claude-4.5-opus-high
Made-with: Cursor
* refactor: SICP-informed LanguageProvider architecture for ingestion pipeline
Consolidate 16 scattered dispatch surfaces into a single LanguageProvider
Strategy interface per language. Processors are now fully language-agnostic —
zero SupportedLanguages.X enum access, zero dispatch table imports.
Architecture (5-layer DAG, zero circular dependencies):
L0: Capability modules (dispatch tables, single source of truth)
L1: LanguageProvider interface + createLanguageProvider factory
L2: 13 per-language provider files (Strategy objects)
L3: Registry with satisfies Record<SL, LP> + pre-built lookup maps
L4: Processors (language-agnostic, all behavior via provider.*)
Key changes:
- Add LanguageProvider interface with 15 properties (6 required, 9 optional)
- Create 13 provider files in languages/ + php-helpers.ts
- Migrate all processors to getProvider(language) — cached once per scope
- Replace heritage if-checks with provider.interfaceNamePattern/heritageDefaultEdge
- Replace MRO switch(language) with switch(provider.mroStrategy)
- Replace isNodeExported with provider.exportChecker
- Move PHP description extraction behind provider.descriptionExtractor
- Move Swift implicit imports behind provider.implicitImportWirer
- Move PHP route detection behind provider.isRouteFile
- Move Kotlin wildcard append behind provider.importPathPreprocessor
- Remove deprecated TypeEnvironment.env, add fileScope()/allScopes()
- De-export TypeEnv type (module-private)
- Pre-build extensionMap, WILDCARD_LANGUAGES, SYNTHESIS_LANGUAGES at load
- Remove dead entryPointPatterns/frameworkPatterns from interface
- Derive createLanguageProvider config type via Pick/Partial/Omit
- Tighten callback types from any to SyntaxNode
- Migrate 270+ test call sites from .env to TypeEnvironment API
Adding a new language: 3 files (enum + provider + registry line).
No processor file touched. Ever.
* refactor: clean architecture for LanguageProvider with O(1) AST cache
Address all PR #488 review comments and achieve pristine SICP layer separation:
Interface redesign:
- Split LanguageProvider into Config (input) + Provider (runtime with defaults)
- Rename createLanguageProvider → defineLanguage with explicit DEFAULTS constant
- Add MroStrategy, ImportSemantics named type aliases for better IDE tooltips
- Tighten labelOverride signature: string|null → NodeLabel|null (compile-time safety)
- Tighten descriptionExtractor nodeLabel: string → NodeLabel
- Un-export LanguageProviderConfig (internal to defineLanguage)
CI fixes (all 4 failures resolved):
- isNodeExported: add null guard for unknown languages
- preprocessImportPath tests: pass getProvider() instead of raw enum
- MRO tests: update expected strings to match language-agnostic prefixes
Code deduplication:
- Extract findDescendant/extractStringContent to ast-helpers.ts (single source of truth)
- Unify Kotlin method detection: remove duplicate from extractFunctionName,
use provider.labelOverride as single source of truth via findEnclosingFunctionId
- extractFunctionName return type: string → NodeLabel
Performance (O(1) AST node access):
- Add per-file Map-based memoization in parse-worker for parent-chain walks
- Cache enclosingClassId, enclosingFunctionId, exportStatus per SyntaxNode
- Clear caches before each file parse (not after — handles parse failures)
Architecture (pristine languages/ folder):
- Move php-helpers.ts → helpers/php.ts (L0 capability, not L2 config)
- Create helpers/swift.ts from extracted Swift provider logic
- Extract cppLabelOverride AST walk → isCppInsideClassOrStruct in ast-helpers.ts
- Extract isPhpRouteFile → helpers/php.ts
- All 13 provider files are now pure configuration — zero implementation logic
- Ruby: remove no-op namedBindingExtractor assignment (undefined from dispatch table)
* refactor: eliminate LANGUAGE_QUERIES, typeConfigs, namedBindingExtractors dispatch tables
Phase 1 of L0 dispatch table elimination. Providers now import capabilities
directly instead of indexing into redundant Record<SL, T> dispatch tables:
- LANGUAGE_QUERIES: providers import named query constants directly
(TYPESCRIPT_QUERIES, PYTHON_QUERIES, etc.). Table kept in tree-sitter-queries.ts
for call-processor.ts dynamic lookup + test consumers.
- typeConfigs: providers import from individual type-extractor files
(typescriptConfig from typescript.ts, javaTypeConfig from jvm.ts, etc.).
Dispatch table fully removed from type-extractors/index.ts.
- namedBindingExtractors: providers import extractors directly from
named-binding-extraction.ts (extractTsNamedBindings, etc.).
Dispatch table fully removed from import-resolution.ts.
Net: -48 LOC of dispatch table indirection. L3 satisfies Record<SL, LP>
remains the single exhaustiveness check.
* refactor: eliminate exportCheckers, callRouters, importResolvers dispatch tables
Phase 2 of L0 dispatch table elimination. All 6 dispatch tables are now gone:
- exportCheckers: individual checkers exported directly (tsExportChecker,
pythonExportChecker, etc.). isNodeExported uses a local checkersByLanguage
map to avoid circular dependency with languages/index.ts.
- callRouters: table removed. Providers import noRouting or routeRubyCall
directly. noRouting now exported. Dead import removed from call-processor.ts.
- importResolvers: resolver functions exported with clean names
(resolveTypescriptImport, resolveJavaImport, etc.). Inline lambdas
extracted to named exports. Dispatch functions renamed from *Dispatch
suffix to clean resolve*Import pattern.
Combined with Phase 1, all 6 L0 dispatch tables have been eliminated.
L3 satisfies Record<SL, LanguageProvider> is the single exhaustiveness check.
Providers are now fully self-contained — each imports its capabilities directly.
* perf+refactor: type-env caching, sequential fallback caching, utils.ts split
Phase 3 — performance optimizations and barrel cleanup:
Type-env parent-walk caching:
- Memoize findEnclosingClassName and findEnclosingParentClassName with
per-file Map<SyntaxNode, string|undefined> caches
- Eliminates O(n*m) repeated child scanning in extractParentClassFromNode
- Caches cleared in buildTypeEnv before each file's walk phase
Sequential fallback caching:
- Add classIdCache + exportCache Maps to parsing-processor.ts
- Mirrors the O(1) memoization pattern from parse-worker.ts
- Both paths now have identical caching for parent-chain walks
Split utils.ts barrel into focused modules:
- noise-filter.ts: BUILT_IN_NAMES + isBuiltInOrNoise (167 LOC)
- language-detection.ts: getLanguageFromFilename (58 LOC)
- utils.ts slimmed to re-exports + yieldToEventLoop + isVerboseIngestionEnabled
- Backward compatible — existing imports from utils.ts still work
* refactor: rename resolvers/ → import-resolvers/, restructure tests per-concern
Directory renames (git mv — history preserved):
- src/core/ingestion/resolvers/ → import-resolvers/ (10 files)
- test/unit/call-routing.test.ts → call-routing/ruby.test.ts
- test/unit/named-binding-extraction.test.ts → named-bindings/csharp.test.ts
- test/unit/import-resolution.test.ts → import-resolution/preprocessing.test.ts
All 11 import paths updated to reference new import-resolvers/ location.
Test imports updated for new subdirectory depth.
Note: test/integration/resolvers/ NOT renamed — those tests cover the full
ingestion pipeline per-language, not just import resolution.
* refactor: eliminate utils.ts barrel — all 33 consumers now import directly
Migrated 65 import sites across 33 files to import from the focused source
module instead of the utils.ts barrel:
- ast-helpers.js: SyntaxNode, extractFunctionName, findEnclosingClassId, etc.
- call-analysis.js: inferCallForm, extractReceiverName, countCallArguments, etc.
- noise-filter.js: BUILT_IN_NAMES, isBuiltInOrNoise
- language-detection.js: getLanguageFromFilename
utils.ts reduced to 2 original functions only:
- yieldToEventLoop
- isVerboseIngestionEnabled
Zero re-exports remain. Every import is now direct to its source module.
* refactor: create utils/ folder, move all shared utilities, delete utils.ts barrel
Final phase of module structure migration:
- git mv ast-helpers.ts, call-analysis.ts, noise-filter.ts,
language-detection.ts → utils/ subdirectory (history preserved)
- Extract yieldToEventLoop → utils/event-loop.ts
- Extract isVerboseIngestionEnabled → utils/verbose.ts
- Delete utils.ts (zero re-exports, zero functions remain)
- Update 38 import paths across source and test files
The ingestion/ root is now clean — only processors, capability modules,
and the pipeline orchestrator live at the top level. All shared utilities
are in utils/, all language-specific helpers in helpers/, all import
resolvers in import-resolvers/.
* refactor: move findChild from import-resolvers/utils.ts to utils/ast-helpers.ts
findChild is a generic AST helper (find first named child by type) — it
belongs with the other AST traversal utilities, not in the import resolver
module. 4 consumers updated to import from utils/ast-helpers.js.
* refactor: split named-binding-extraction.ts into per-language files
Rename named-binding-extraction.ts → named-binding-processor.ts (git mv,
history preserved), keeping only walkBindingChain for re-export chain resolution.
7 per-language extractor functions moved to named-bindings/ subdirectory:
- named-bindings/typescript.ts (extractTsNamedBindings — TS + JS)
- named-bindings/python.ts (extractPythonNamedBindings)
- named-bindings/kotlin.ts (extractKotlinNamedBindings)
- named-bindings/rust.ts (extractRustNamedBindings + collectRustBindings)
- named-bindings/php.ts (extractPhpNamedBindings)
- named-bindings/csharp.ts (extractCsharpNamedBindings)
- named-bindings/java.ts (extractJavaNamedBindings)
Each provider now imports its binding extractor from the per-language file.
* refactor: eliminate import-resolution.ts — distribute to natural homes
Split per-language resolvers into import-resolvers/ per-language files and
eliminate the import-resolution.ts catch-all module entirely:
Per-language resolvers moved to import-resolvers/:
- standard.ts: resolveStandard, resolveJavascriptImport, resolveTypescriptImport,
resolveCImport, resolveCppImport
- jvm.ts: resolveJavaImport, resolveKotlinImport
- go.ts: resolveGoImport
- csharp.ts: resolveCSharpImport (helper renamed to Internal)
- php.ts, python.ts, ruby.ts, rust.ts: same pattern
- swift.ts: new file for resolveSwiftImport
Types distributed to their concern directories:
- import-resolvers/types.ts: ImportResult, ImportConfigs, ResolveCtx, ImportResolverFn
- named-bindings/types.ts: NamedBinding, NamedBindingExtractorFn
preprocessImportPath moved to import-processor.ts (its primary consumer).
import-resolution.ts deleted — zero catch-all modules remain.
* refactor: tighten SPR — eliminate re-exports, dead code, type holes, and redundant patterns
12 review findings resolved across the ingestion layer:
Type safety:
- CallRouter callNode: any → SyntaxNode (closes type hole)
- CaptureMap type alias replaces Record<string, any>
- providersWithImplicitWiring filter now type-narrowed (removes ! assertions)
- Ruby exportChecker: unnecessary as-cast removed, named export created
Architecture:
- Circular type dependency eliminated (ImportResolutionContext moved to types.ts)
- LANGUAGE_QUERIES residual dispatch replaced with provider.treeSitterQueries
- noRouting sentinel deleted — callRouter now properly optional on 12 providers
- All 6 re-exports from import-processor/pipeline/languages eliminated
Pattern cleanup:
- Dead checkersByLanguage table + isNodeExported removed from export-detection
- 4 duplicated config interfaces consolidated to language-config.ts
- extractCsharpNamedBindings → extractCSharpNamedBindings (casing consistency)
Simplification:
- import-resolvers/index.ts barrel deleted (dead re-exports)
- helpers/ inlined into languages/ (php.ts, swift.ts) — 1 directory removed
Verified: tsc --noEmit clean, 3837 tests pass, 0 failures.
* refactor: address review — remove LANGUAGE_QUERIES table, type-extractors barrel, fix Windows timeout
Review comment fixes (github.com/abhigyanpatwari/GitNexus/pull/488#issuecomment-4117817648):
1. LANGUAGE_QUERIES dispatch table removed from tree-sitter-queries.ts
— 5 test files migrated to getProvider(lang).treeSitterQueries
— eliminates last parallel dispatch surface
2. type-extractors/index.ts barrel deleted
— type-env.ts now imports TYPED_PARAMETER_TYPES from shared.js directly
3. Windows CI timeout fix: afterAll cleanup hook in test-indexed-db.ts
now passes explicit 120s timeout to prevent KuzuDB C++ destructor
hang from hitting vitest's default 30s testTimeout on Windows
Verified: tsc --noEmit clean, 3835 tests pass, 0 failures.
* refactor: eliminate chained getProvider property access — assign to variable first
All getProvider(lang).property calls now follow the pattern:
const provider = getProvider(language);
const x = provider.property;
5 source files + 4 test files updated (~35 occurrences).
This ensures consistent provider variable usage and avoids
repeated lookups in hot paths.
* refactor: remove last 4 re-exports from import-resolvers, fix stale CaptureMap comment
- Remove `export type { TsconfigPaths }` from standard.ts
- Remove `export type { GoModuleConfig }` from go.ts
- Remove `export type { ComposerConfig }` from php.ts
- Remove `export type { CSharpProjectConfig }` from csharp.ts
All 4 types are canonically defined in language-config.ts;
zero consumers imported via the resolver re-exports.
- Fix stale CaptureMap JSDoc: said "Uses any" but type is SyntaxNode | undefined
Add GLM support using OpenAI-compatible API via ChatOpenAI from LangChain.
Defaults to the Z.AI coding endpoint (https://api.z.ai/api/coding/paas/v4)
with configurable base URL. Supported models: GLM-5, GLM-5-Turbo, GLM-4.7, GLM-4.5.
- Add ADD_TAGS: ['foreignObject'] to all DOMPurify.sanitize calls —
Mermaid uses foreignObject for HTML text labels inside flowchart
nodes. The SVG profile was stripping them, causing empty boxes.
- Remove leftover sub-batch loop lines from prepared statement hoist
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- MarkdownRenderer: wrap handleLinkClick in useCallback, add to
markdownComponents useMemo deps (fixes stale closure)
- GraphCanvas: remove sigmaRef from useEffect deps (ref identity
never changes), extract handleToggleAIHighlights to useCallback
- CodeReferencesPanel: add nodeById Map for O(1) focus-in-graph
lookup (was O(N) graph.nodes.find on every click)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Deep imports from lucide-react/dist/esm/icons/*.js are internal paths
that broke the Vercel production build. Replaced with standard named
re-exports from lucide-react — keeps the centralized module pattern
without relying on fragile internal paths.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Hoists conn.prepare(cypher) outside the sub-batch loop so it's called
once per (fromLabel, toLabel) pair instead of ceil(N/4) times. The
statement is reused for all rows in the group, then closed in finally.
Yields to event loop every 500 relations instead of every sub-batch.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- BackendRepoSelector, EmbeddingStatus, Header, MermaidDiagram,
QueryFAB, RightPanel, ToolCallCard, WebGPUFallbackDialog: switch
from lucide-react barrel imports to @/lib/lucide-icons deep imports
- MermaidDiagram: lazy-load ProcessFlowModal via React.lazy
- Extract ProviderConfigCard from SettingsPanel for cleaner separation
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Performance:
- nodeById Map in GraphCanvas for O(1) click/hover lookups
- fileNodeByPath Map in useAppState for O(1) file path lookups
- Set.has() for HIGHLIGHT_NODES/IMPACT matching (was O(N²))
- useMemo for primaryLanguage in StatusBar
- useCallback on toggleLabelVisibility/toggleEdgeVisibility
- Recursive FileTreePanel search (full subtree, not 1 level)
React fixes:
- Remove stale queryResult dep from clearAICodeReferences
- Cancel RAF chains in CodeReferencesPanel on cleanup
- Clean up timeouts in SettingsPanel and MarkdownRenderer on unmount
- try/catch on localStorage in DropZone (private browsing)
- Await handleServerConnect in App.tsx auto-connect
- pendingToolCalls counter replaces allToolsDone boolean in agent.ts
- JSON.parse try/catch in agent streaming
- sessionStorage JSDoc fix in settings-service
Bundle:
- Centralized lucide icon deep imports
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Mermaid uses foreignObject for HTML text labels inside flowchart
nodes. The SVG profile strips them by default, causing empty boxes.
ADD_TAGS: ['foreignObject'] preserves text while still sanitizing
against XSS.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add NODE_COLORS and NODE_SIZES entries for all 15 multi-language labels
(Struct, Trait, Impl, TypeAlias, Const, Static, Namespace, Union,
Typedef, Macro, Property, Record, Delegate, Annotation, Constructor,
Template) — fixes Record<NodeLabel, ...> completeness
- Escape table names in count/lookup queries with escapeTableName() to
prevent silent failures for backtick-required tables
- Add HAS_PROPERTY and ACCESSES to RelationshipType union
- Update initPromise after db recreation in loadGraphToLbug so subsequent
initLbug() calls return fresh db/conn refs
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- ProcessFlowModal: revert import from non-existent @/lib/lucide-icons back to lucide-react
- embedding-pipeline: remove extra argument in executeQuery call (signature only accepts 1 arg)
Both issues were introduced in PR #475 (web-security-hardening).
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- empty graph produces header-only CSVs (no data rows)
- empty graph relCSV has only header
- double quotes in node names are RFC 4180 escaped
- file node without fileContents gets empty content (no crash)
- community with empty keywords array produces valid CSV
- unknown node labels are silently skipped
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add 3 new test fixtures: swift-if-let-guard-let, swift-await-try, swift-for-loop-inference
- Add integration tests for if let/guard let binding resolution (4 assertions)
- Add integration tests for await/try expression unwrapping (3 assertions)
- Add for-loop-inference fixture (documented as known gap — type-env infrastructure
is in place but call-processor re-parse path doesn't propagate the binding yet)
- Fix cross-chunk Swift implicit imports: standard processImports path now passes
allFileList instead of chunk-only files to addSwiftImplicitImports, matching
the fast-path behavior
- Add Swift type_annotation fallback in type-env declarationTypeNodes population
(handles [User] array sugar where childForFieldName('type') returns null)
- Handle Swift 'pattern' node in extractVarName fallback (pattern wraps simple_identifier)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The readOnly guard was matching keywords inside string literals,
blocking legitimate queries like WHERE n.name CONTAINS "delete".
Now strips single/double-quoted strings before checking, so only
actual Cypher write keywords outside strings are blocked.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The readOnly=true default on executeQuery blocks queries containing
CREATE, which includes CALL CREATE_VECTOR_INDEX. The embedding pipeline
needs write access for this setup step.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Node IDs are generated as Label:filePath (e.g., Function:src/foo.ts:bar),
so forward slashes are expected in legitimate IDs. The over-tightened
regex was dropping all path-based IDs from Cypher queries.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
`import models as m` aliases were stored in namedImportMap (a symbol-binding
map) then cross-referenced in pipeline.ts — semantic misuse and inefficient.
Refactored: NamedBinding gains `isModuleAlias` flag. applyImportResult routes
tagged bindings directly to moduleAliasMap at import time. Removes the
pipeline.ts post-processing loop entirely.
Added test fixture and 5 integration tests for `import X as Y` with
multi-module disambiguation (both models.py and auth.py export User).
.husky/pre-commit is committed to the repo — developers get the hook
by cloning, not by running npm install. prepare only needs to build
TypeScript for npm publish/pack.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
`import numpy as np` and `from models import User as U` previously
generated no CALLS edges because:
1. `import_statement` with an `aliased_import` child was not captured
by the tree-sitter query for Python imports.
2. `extractPythonNamedBindings()` only handled `import_from_statement`,
ignoring plain `import X as Y` forms.
Changes:
- `tree-sitter-queries.ts`: add query pattern for
`(import_statement name: (aliased_import name: (dotted_name)))` so
the import path is captured before named-binding extraction runs.
- `named-binding-extraction.ts`: extend `extractPythonNamedBindings()`
to handle `import_statement` nodes carrying `aliased_import` children.
Records `{ local: "np", exported: "numpy" }` so call-sites using the
alias resolve to the real module.
- `test/fixtures/lang-resolution/python-alias-imports/`: update fixtures
used by `python.test.ts` to exercise `from models import User as U`.
Existing tests in `test/integration/resolvers/python.test.ts`
(suite "Python alias import resolution") cover this path.
16 tests covering the data structures and logic underlying each fix:
createKnowledgeGraph (loadServerGraph data flow):
+ nodes stored correctly via addNode
+ relationships stored correctly via addRelationship
+ deduplication by ID
+ nodeCount reflects unique count
- empty graph has zero counts
- relationships with non-existent nodes still stored
loadServerGraph data flow:
+ server data reconstructs into valid KnowledgeGraph
+ fileContents Map built from server entries
- empty server data produces empty graph
- fileContents replaces (not accumulates) on reload
BM25 index argument type:
+ Map<string, string> has entries() for BM25
- KnowledgeGraph does NOT have entries() (the original bug)
Highlight clearing:
+ clearing Set produces empty set
+ independent highlight sources cleared separately
- clearing highlights doesn't affect node selection
- toggling AI ON doesn't clear process highlights
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Clear progress after handleServerConnect in both auto-connect and
DropZone paths (fixes frozen progress bar on StatusBar)
- Replace shell-based prepare script with Node scripts/prepare.cjs
for Windows cmd.exe compatibility
- Gate LadybugDB load warning behind import.meta.env.DEV (consistent
with finalizePipeline's silent catch)
- Remove misleading "parallel" comment (fetch is sequential after connect)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Per maintainer request. Husky is activated via `cd .. && husky` in the
prepare script. The pre-commit hook mirrors CI: typecheck + unit tests
for both packages when relevant files are staged.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Previously, Python was added to WILDCARD_IMPORT_LANGUAGES which expanded
all exported symbols into namedImportMap using first-seen wins. This caused
`auth.User()` to incorrectly resolve to `models.py:User` when both modules
exported a class named User.
Root cause: Python `import models` is a namespace import, not wildcard
symbol expansion. Expanding all symbols produces ambiguous bindings that
cannot be disambiguated later.
Fix:
- Remove Python from WILDCARD_IMPORT_LANGUAGES
- Add ModuleAliasMap (callerFile → alias → sourceFile) to ResolutionContext
- In synthesizeWildcardImportBindings, build moduleAliasMap for Python
using the filename stem as the module alias
- In resolveCallTarget, add module-alias disambiguation step: when multiple
candidates survive filtering and the receiver name matches a module alias,
narrow candidates to the aliased file
Result: `models.User()` → models.py:User, `auth.User()` → auth.py:User
even when both modules export a class named User.
Adds regression test for the ambiguity case.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Replace result.getAll() with result.getAllRows() across lbug-adapter
(getAll doesn't exist in LadybugDB WASM v0.15.1)
- Add loadServerGraph worker method that pipes server data through
initLbug/loadGraphToLbug for in-browser querying
- Extract finalizePipeline helper (shared by runPipeline, runPipelineFromFiles)
- Fix buildBM25Index called with graph object instead of fileContents Map
- Fix 'Turn off all highlights' to clear sigma selection, AI tool/citation
highlights, and blast radius
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Fix timeout detection: AbortSignal.timeout() throws TimeoutError, not
AbortError. Timeouts are no longer retried (30s fail, not 93s).
- Validate embedding dimensions in both httpEmbed and httpEmbedQuery
against config.dimensions or the 384d schema default. When DIMS is
unset, the error says 'Set GITNEXUS_EMBEDDING_DIMS=N' to guide users.
- Centralize test env var cleanup in afterEach via savedEnv snapshot.
- Test mocks use 384d vectors matching schema default.
- 4 new tests: timeout not retried, network retry success, query path
dim mismatch, unset-dims hint. 23 total, all pass.
* remove friction in onboarding by correcting typo
* Revert "remove friction in onboarding by correcting typo"
This reverts commit 07dec38c2c.
* feat(ui): add HelpPanel component with tabbed reference, node legend, AI query guide, and dual Mac/Windows keyboard shortcuts
* feat(ui): add HelpPanel component with tabbed reference, node legend, AI query guide, and dual Mac/Windows keyboard shortcuts
* made changes based on the suggestions
* minor fix
npm ci was failing with "Missing: hono@4.12.8" and
"Missing: graphology-types@0.24.8" because the lock file was
out of sync after rebase. Regenerated from clean state.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Python repos were producing 0 CALLS edges for module-qualified constructor
calls like `models.User()` where `import models` is a bare module import.
Root causes:
1. `SupportedLanguages.Python` was absent from `WILDCARD_IMPORT_LANGUAGES`,
so `synthesizeWildcardImportBindings` never ran for Python files — bare
module imports never received per-symbol namedImportMap bindings.
2. Synthesis only ran in the Phase 14 pre-pass, after all chunks had already
been call-resolved. When `models.User()` was processed in Phase 3+4,
`namedImportMap` was empty for Python → Tier 2a-named fell through to
Tier 2a which found both `models.py:User` and `auth.py:User` (ambiguous).
3. `filterCallableCandidates` with `callForm='member'` excluded `Class` nodes
(only `CALLABLE_SYMBOL_TYPES` = Function/Method/Constructor/…). With 2
ambiguous Class candidates both were dropped, producing 0 CALLS edges.
Fixes:
- Add `SupportedLanguages.Python` to `WILDCARD_IMPORT_LANGUAGES` so that
`import models` expands to per-symbol namedImportMap entries (first-seen
semantics: `User→models.py:User`, `Admin→auth.py:Admin`).
- Call `synthesizeWildcardImportBindings` inline in the chunk loop, after
`processImportsFromExtracted` but BEFORE `processCallsFromExtracted`. This
ensures Tier 2a-named can disambiguate `module.ClassName()` at initial
call-resolution time. The Phase 14 pre-pass remains as a final safety net.
- Add a fallback in `resolveCallTarget`: if `callForm='member'` yields 0
filtered candidates, retry with `callForm='constructor'`. This handles the
case where a module-qualified class instantiation (e.g. `models.User()`)
is syntactically an attribute-access call but semantically a constructor
call. The fallback only triggers for 0-candidate member calls, so it
cannot over-eagerly promote normal member calls.
Tests: add `python-module-import` fixture (models.py/auth.py/app.py) with
4 regression tests covering IMPORTS edges, name-collision disambiguation
for `models.User()`, and `auth.Admin()`.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Pin exact versions (no ^) to prevent surprise upgrades:
- tree-sitter: "0.22.4" (was "^0.22.4")
- tree-sitter-swift: "0.7.1" (was "^0.7.1")
Add npm overrides to suppress peer dependency warnings from grammar
packages that declare ^0.21.x but work fine with 0.22.4.
Note: tree-sitter-swift 0.6.0 fails to build on current Node (needs
node-gyp + Swift toolchain). 0.7.1 with prebuilt binaries is required
for Swift support to work at all.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
1. Export detection: exclude private(set)/fileprivate(set) from
unexported check. Only the setter is restricted — the symbol
itself is still readable cross-file.
2. For-loop binding: use extractVarName() instead of raw .text
to avoid polluting scopeEnv with non-identifier keys from
tuple destructuring patterns (e.g. `for (a, b) in ...`).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Covers all high and medium impact gaps from the Swift feature coverage
analysis:
1. if let / guard let bindings: add if_statement and guard_statement
to DECLARATION_NODE_TYPES, extract varName and value for Tier 2
return-type propagation (callResult, copy, fieldAccess, methodCallResult)
2. await / try expression unwrapping: add unwrapSwiftExpression() that
strips await_expression and try_expression wrappers before checking
for call_expression. Applied in extractPendingAssignment,
extractInitializer, and scanConstructorBinding.
3. for item in collection: add extractForLoopBinding for Swift with
extractSwiftElementTypeFromTypeNode that handles [User] array sugar
and Array<User> generic types. Registered in typeConfig.
4. Multiple inheritance specifiers: already working — tree-sitter
queries match all inheritance_specifier occurrences automatically.
Verified, no code changes needed.
5. Enum case extraction: add (enum_entry (simple_identifier) @name)
@definition.property query to SWIFT_QUERIES.
6. self/super resolution: unskipped both describe.skip test suites
(tree-sitter-swift 0.7.1 ships prebuilds, Node 22 build issue
resolved). Both pass — 5 previously-skipped tests now running.
7. Optional chaining obj?.method(): already working — tree-sitter-swift
parses the ? transparently. Verified, no code changes needed.
Tests: 3,603 → 3,608 (5 unskipped self/super tests)
Swift tests: 23 → 28 passing
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Addresses reviewer feedback: the new Swift behaviors (implicit imports,
constructor fallback, extension dedup, export detection) had no dedicated
integration tests. Adds 4 fixture directories and 11 new test assertions:
1. swift-implicit-imports: two files, no explicit import, cross-file
constructor + member call resolves via addSwiftImplicitImports
2. swift-extension-dedup: extension creates duplicate Class node,
constructor still resolves to primary definition
3. swift-constructor-fallback: ClassName() without `new` resolves as
constructor via free→constructor retry
4. swift-export-visibility: internal symbols visible cross-file,
public/open visible, private/fileprivate noted as Tier 3 limitation
All 3,603 tests pass (11 new, 0 regressions).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Swift was missing the extractPendingAssignment extractor, which meant
return-type-based variable bindings like `let user = getUser()` couldn't
propagate the return type of `getUser()` to `user`. This broke member
call resolution: `user.save()` couldn't resolve to `User.save()` when
there were competing methods (both User and Repo have save()).
Handles four Swift patterns:
- let user = getUser() → callResult (Tier 2 propagation)
- let result = user.save() → methodCallResult
- let name = user.name → fieldAccess
- let copy = user → copy
All 3,592 tests pass — including the 2 previously-failing Swift
return-type inference tests.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Fix assignment query: tree-sitter-swift 0.7.1 uses named fields
(target:/result:/suffix:) instead of positional children
- Update export detection tests: Swift `internal` (default) is now
correctly treated as exported (module-scoped visibility)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When Swift extensions create multiple Class nodes with the same name
(e.g. Product.swift + ProductMatchableConformance.swift), the call
resolver gets multiple candidates and refuses to emit a CALLS edge.
Add dedup: when all candidates share the same type (Class/Struct) and
differ only by file, prefer the primary definition (shortest filepath).
Note: This fix is partial — some constructor calls inside function
bodies may still be consumed by the type-env constructor binding
scanner before reaching resolveCallTarget. Filed as known limitation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Three changes that together enable cross-file call resolution for Swift:
1. export-detection.ts: Treat internal (default) Swift symbols as exported.
Swift's default access level is `internal` (module-scoped, visible to
all files in the same target). Only private/fileprivate are file-scoped.
Previously all non-public/open symbols were marked unexported.
2. import-processor.ts: Add implicit import edges between all Swift files
in the same module/target. Swift has no file-level imports — all files
see each other automatically. Without these edges, the tiered resolver
can't find cross-file symbols at Tier 2a (import-scoped).
Supports SPM targets via Package.swift; falls back to single-module
for Xcode projects without SPM.
3. call-processor.ts: Add constructor fallback for free-form calls.
Swift constructors look like free function calls (no `new` keyword):
`let ocr = OCRService()`. The call form is inferred as `free`, which
filters out Class/Struct targets. Now retries with `constructor` form
when free-form finds no callable but the name resolves to a type.
Tested on 61-file iOS 26 project (PricePal):
- Before: 0 cross-file CALLS edges
- After: full cross-file resolution (OCRService traced from ScanViewModel)
- 3,099 nodes, 10,449 edges, 246 clusters, 243 flows
Related: #406, #407
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The patch script fails to parse tree-sitter-swift@0.6.0's binding.gyp
because the file contains both Python-style # comments AND trailing
commas in JSON arrays. The existing regex strips # comments but leaves
trailing commas, causing JSON.parse() to fail with:
"Unexpected token ']'"
This silently prevents tree-sitter-swift from building, which means
Swift files are skipped entirely during analysis.
Fix: add a second regex pass to strip trailing commas before ] or }
after comment removal.
Fixes#386, #406
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Allow RFC 1918 private network ranges in the CORS origin allowlist so
users running GitNexus on their home or office LAN can access the web UI
from another device on the same network.
Permitted private ranges:
10.0.0.0/8 (10.x.x.x)
172.16.0.0/12 (172.16.x.x – 172.31.x.x)
192.168.0.0/16 (192.168.x.x)
The origin check is extracted into an exported isAllowedOrigin() helper
so it can be unit-tested in isolation. A new test file covers:
- No origin (curl / server-to-server)
- localhost and 127.0.0.1 variants
- All three RFC 1918 ranges including boundary values
- The deployed gitnexus.vercel.app site
- Public / untrusted origins that must be rejected
The server bind address (127.0.0.1 by default) is unchanged; this PR
only affects which cross-origin browser requests are accepted.
Closes#390
Instead of hardcoding confidence: 1.0, compute it at ingestion time using
the same resolution tier system that CALLS edges already use.
Heritage edges (EXTENDS, IMPLEMENTS):
- resolveHeritageId now returns { id, confidence } using TIER_CONFIDENCE
- Same-file → 0.95, import-scoped → 0.9, global → 0.5
- Edge confidence = geometric mean of source and target confidence
(principled for partially-correlated cross-scope estimates, per
Dillig et al. POPL 2011 and Dempster-Shafer theory)
MRO edges (OVERRIDES):
- MRO-ordered → 0.9, class method wins → 0.95
- Single interface → 0.85, ambiguous/unresolved → 0.5
IMPORTS and CONTAINS intentionally keep 1.0 (deterministic).
Closes#412
When LadybugDB throws a BUSY/lock error (e.g. CLI and server running
concurrently), withLbugDb retries up to 3 times with linear backoff.
Addresses review feedback:
1. **Race condition fix**: Connection cleanup (close + state reset) now
runs inside runWithSessionLock, preventing another operation from
acquiring the lock between cleanup steps and having its connection
closed from under it.
2. **Tests call withLbugDb directly**: Replaced simulateWithRetry helper
with tests that invoke the real withLbugDb implementation, catching
regressions in retry count, backoff, and lock interaction.
Closes#325
Addresses all review items from @magyargergo and Copilot:
1. **Rename --no-git to --skip-git**: Commander.js treats --no-X flags
as negation of --X (stores as options.git = false, not options.noGit).
--skip-git maps correctly to options.skipGit.
2. **Fix false " Already up to date\ on non-git folders**: When
currentCommit is empty string, skip the cache check — we cannot
detect changes without git, so always rebuild.
3. **Replace isGitRepo() with hasGitDir()**: Use filesystem check
(statSync on .git) instead of shelling out to git CLI. Consistent,
faster, and works when git is not installed.
4. **Fix misleading warning**: Message now only fires when .git
directory is actually absent (not when git CLI fails).
5. **Add CLI integration tests**: Verify Commander maps --skip-git
correctly and that non-git folders are rejected without the flag.
- Replace require(" fs\) with ESM-compatible top-level import (statSync)
- Register --no-git option in Commander CLI definition
- Use hasGitDir() instead of isGitRepo() for .gitignore update guard
to match the PR intent (filesystem check vs git CLI invocation)
The cypher tool description and schema resource omit Community and Process
node properties, causing agents to write failing queries on first attempt.
Added property listings sourced from the actual LadybugDB schema definitions:
- Community: heuristicLabel, cohesion, symbolCount, keywords, description, enrichedBy
- Process: heuristicLabel, processType, stepCount, communities, entryPointId, terminalId
Closes#411
Previously gitnexus analyze exited with an error on any directory that
lacked a .git entry, making it impossible to index generated code,
vendored libraries, or monorepo sub-trees that are not git roots.
Changes:
storage/git.ts
- Add hasGitDir(dirPath): boolean — a lightweight synchronous check for
the presence of a .git file or directory. Works for git worktrees
(.git file pointing at the real repo) as well as standard repos.
cli/analyze.ts
- Add noGit?: boolean to AnalyzeOptions.
- When the explicit inputPath resolves to a non-git folder (or the cwd
is not inside any git repo), respect --no-git instead of hard-failing.
- Print an actionable tip pointing at --no-git when git is absent and the
flag was not supplied.
- currentCommit defaults to an empty string for non-git folders so the
up-to-date check still functions (empty string never matches a real
commit hash, so the index is always rebuilt).
- Skip addToGitignore() when no .git is present — there is nothing to
update and the function would create a stale .gitignore at the root.
Git-dependent features that remain disabled for non-git folders:
- Incremental update (always rebuilds from scratch)
- Commit tracking in metadata
- .gitignore update
Closes#384
ORT 1.24.x downloads CUDA provider .so from NuGet at postinstall,
but only for linux/x64. The process.arch guard correctly returns
false on arm64 (safe CPU fallback), but the prior comment implied
arm64 CUDA was supported. Clarify the actual state.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Address PR #300 review findings:
[CRITICAL] hasOrtCudaProvider() was checking the top-level
onnxruntime-node@1.24.3 but @huggingface/transformers loads its own
nested onnxruntime-node@1.21.0 at runtime. The guard inspected the
wrong binary, so the native crash was not prevented.
Fix: resolve onnxruntime-node from transformers' own module scope
(createRequire from transformers' package.json) so the guard always
checks the same binary that will be dlopen'd at runtime.
Also:
- Add npm overrides to force @huggingface/transformers to use our
onnxruntime-node@^1.24.0 (works for global installs where gitnexus
is the root package; npx installs get safety from the resolve fix)
- Replace hardcoded 'x64' with process.arch for arm64 support
- Remove dead napi-v3 path check (ORT 1.21.0 never shipped CUDA .so)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Enhance C++ tree-sitter queries to support inline class method declarations and return types.
- Introduce `importedRawReturnTypes` in `BuildTypeEnvOptions` for cross-file raw return type handling.
- Add `FileTypeEnvBindings` interface to capture file-scope type bindings for exported symbols.
- Implement logic in `parse-worker.ts` to extract and serialize file-scope type bindings for cross-file type resolution.
- Create test fixtures for C++, Go, Ruby, and Rust to validate cross-file binding propagation.
- Update integration tests to verify correct resolution of method calls across files for C++, Go, Ruby, and Rust.
- Document Phase 14: Cross-File Binding Propagation in the type resolution roadmap and system documentation.
Fix server/bridge mode leaving the web UI with 0 nodes and broken
Query/Processes/embeddings by hydrating the worker-side LadybugDB
and BM25 indexes after loading graph data from the backend.
Also fix LadybugDB QueryResult API mismatch where result.getAll()
does not exist in some @ladybugdb/wasm-core versions — falls back
to getAllObjects() or getAllRows().
- Add .env.example with all HTTP embedding env vars documented
- Early return in httpEmbed() for empty text arrays
- Warn once if API returns vectors with different dimensions than
GITNEXUS_EMBEDDING_DIMS — helps catch misconfiguration early
- Extract shared HTTP client (http-client.ts) used by both core and MCP embedders
- Remove module-level httpConfig cache — read env vars fresh on every call
so config set after module load (e.g. via dotenv) takes effect
- Add NaN/non-positive guard on GITNEXUS_EMBEDDING_DIMS in schema.ts
- Include scrubbed URL and batch index in error messages (no API key)
- Wrap fetch rejections (DNS/timeout/connection) with same scrubbed context
- MCP embedder delegates to shared httpEmbedQuery() instead of inline logic
- apiKey confined to http-client.ts internals — not exported in any type or accessor
- Remove HttpEmbeddingConfig from types.ts (replaced by internal HttpConfig)
- All 16 HTTP embedder tests pass, tsc clean
* feat: add markdown file indexing (headings + cross-links)
Parse .md/.mdx files using regex (no tree-sitter dependency) to extract:
- Section nodes from headings (h1-h6) with hierarchy via CONTAINS edges
- Cross-file IMPORTS edges from markdown links to other repo files
Ported from #286 to resolve conflicts with kuzu→lbug rename.
Co-Authored-By: Dennis Palatov <dp-web4@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: add Section to NODE_TABLES and NODE_SCHEMA_QUERIES
The Section schema was defined but not registered in NODE_TABLES or
NODE_SCHEMA_QUERIES, so the table was never created in the database.
Also adds missing FROM File TO Section relation entry.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: update schema test counts for Section node type
NODE_TABLES: 27→28, NODE_SCHEMA_QUERIES: 27→28, SCHEMA_QUERIES: 29→30
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* test: add diagnostic output to skills-e2e idempotency test
Show stdout/stderr in assertion message so CI failures reveal
why the second analyze --skills run exits with code 1.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: add Section COPY query with level column in lbug-adapter
Section table has 8 columns (includes level) but getCopyQuery fell
through to the default 7-column multi-language path. Adds explicit
Section cases to getCopyQuery and insertNodeToLbug/upsertNodeToLbug.
Error was: COPY failed for Section: Number of columns mismatch. Expected 7 but got 8.
---------
Co-authored-by: Dennis Palatov <dp-web4@users.noreply.github.com>
Add MiniMax as a new LLM provider using the Anthropic-compatible API.
Changes:
- Add MiniMax to LLMProvider type and MiniMaxConfig interface
- Add MiniMax chat model creation via ChatAnthropic with custom base URL
- Add MiniMax settings persistence and model list (MiniMax-M2.5, MiniMax-M2.5-highspeed)
- Add MiniMax provider UI in SettingsPanel with API key and model selection
1. Replace fragile regex in buildExportedTypeMapFromGraph with
extractReturnTypeName() for consistent generic unwrapping.
2. Fix gap detection precision: check upstream.has(binding.exportedName)
instead of exportedTypeMap.has(binding.sourcePath) to avoid
over-counting files that import only unrelated symbols.
All 4 test blocks now share one DB lifecycle to avoid cross-block
"Database is closed" errors caused by LadybugDB's shared global DB
in a single vitest fork. Staleness detection (which triggers closeLbug)
runs last to avoid invalidating connections for other blocks.
11/11 tests pass on macOS, Ubuntu, and Windows.
Critical fixes:
- Re-resolution pass now actually re-resolves CALLS edges by calling
processCalls with importedBindingsMap (was building typeEnv but
discarding it without producing edges)
- Worker path populates ExportedTypeMap via buildExportedTypeMapFromGraph
using graph node isExported + SymbolTable returnType/declaredType
(was dead parameter in processCallsFromExtracted)
Important fixes:
- Skip threshold denominator uses totalFiles (was exportedTypeMap.size +
filesWithGaps which made threshold nearly useless)
- processCalls accepts importedBindingsMap parameter to thread cross-file
bindings into buildTypeEnv during re-resolution
All 3454 tests pass.
Add ExportedTypeMap infrastructure to propagate resolved type bindings
across file boundaries. When file A exports `const user = getUser()`
(resolved to `User`), file B importing `user` now gets seeded with
`user → User`, enabling `user.save()` to produce CALLS edges.
Key components:
- `importedBindings` option on BuildTypeEnvOptions with scopeEnv seeding
AFTER walk() to respect first-writer-wins (local declarations win)
- `collectExportedBindings()` in call-processor using graph node
isExported flag (no SymbolDefinition changes needed)
- Inline Kahn's algorithm topological sort with level grouping for
parallel-safe file ordering and cycle detection
- Re-resolution pass in pipeline.ts: topological order, 3% skip
threshold, path validation, per-file export caps (500)
- 32 new tests: 11 topological sort, 6 seeding, 15 integration
(simple cross-file, re-export chain, circular imports)
All 3454 tests pass (32 net new, 0 regressions).
Address code review findings from PR #392 senior compiler review:
- Fix Java "Yes" → "No" in optional-param-arity matrix (Java has no defaults)
- Simplify Kotlin hasDefaultValue while-as-if to direct const/if check
- Update OPTIONAL_PARAM_TYPES comment to include Ruby
- Replace per-declaration Set allocation with size-based Map iteration skip
- Add 11 unit tests for multi-declarator type association and constructorTypeMap
Add requiredParameterCount to SymbolDefinition and MethodSignature,
enabling range-based arity filtering in filterCallableCandidates.
Calls with omitted optional/default arguments now resolve correctly.
Supported: TS, Python, Kotlin, C#, C++, PHP, Ruby (7 languages).
Detection via OPTIONAL_PARAM_TYPES set + hasDefaultValue helper.
9 integration tests added across all 7 languages.
- Watchdog timer now exempts in-flight queries via activeQueryCount,
preventing premature stdout restoration during long queries (>1s)
- Stale detection uses reinitPromises Map to prevent TOCTOU race where
concurrent callers double-close the connection pool
- Throttle meta.json staleness checks to once per 5s per repo
- Add null guard for i.id in IN-clause construction
- Enrichment queries run in parallel on non-arm64 platforms to preserve
performance; sequential only on arm64 macOS where SIGSEGV occurs
- Updated AGENTS.md and CLAUDE.md to reflect new indexing metrics.
- Enhanced call-processor.ts to support cross-file inheritance tracking and improved virtual dispatch resolution.
- Added support for TypeScript overload signatures in tree-sitter queries.
- Improved type extraction for C++, C#, and Kotlin to handle smart pointers and constructor types.
- Introduced inferLiteralType for overload disambiguation across multiple languages.
- Added tests for C++ smart pointer dispatch and Kotlin virtual dispatch scenarios.
- Updated type-resolution-roadmap.md to reflect completion of phases P.1 to P.3 and outline future work on covariant return types.
- Add AbortSignal.timeout(30s) on all fetch calls
- Add retry with backoff for 429/5xx (core: 2 retries, MCP: 1 retry)
- Guard initEmbedder() and getEmbedder() to throw in HTTP mode
- Discard cached embeddings on dimension mismatch during incremental re-index
- Add MCP embedQuery retry for transient failures
- Add 16 unit tests covering both core and MCP HTTP paths
- Fix README: concise, accurate env var docs
Fixes three related issues that cause SIGSEGV crashes and stale data:
1. Impact enrichment queries (Promise.all → sequential await)
The impact() method ran 3 enrichment queries concurrently via
Promise.all against the same LadybugDB connection pool. On arm64
macOS, concurrent native DB access triggers SIGSEGV. Changed to
sequential await. Also caps IN-clause to 100 IDs to prevent
oversized queries. (#285, #290, #292)
2. Silence stdout during query execution
silenceStdout()/restoreStdout() only wrapped createConnection() and
initLbug(). Now also wraps executeQuery() and executeParameterized()
to prevent native stdout writes from corrupting the MCP stdio
stream during all DB operations. (#285)
3. Stale data after re-index
ensureInitialized() checked pool existence but never verified whether
the underlying index was rebuilt. Now reads meta.json's indexedAt
timestamp on each call and closes/re-opens the pool when the index
has changed. (#297)
- schema.ts: FLOAT[${EMBEDDING_DIMS}] reads from GITNEXUS_EMBEDDING_DIMS env
- embedding-pipeline.ts: vector search CAST uses actual query vector length
- mcp/core/embedder.ts: HTTP embedding support for MCP query-time search
Without this, using a 1024d model (e.g. bge-large) fails with
'Expected: 384, Actual: 1024' on LadybugDB vector insert.
Adds support for OpenAI-compatible embedding endpoints as an alternative
to the local transformers.js pipeline. Enables using self-hosted servers
(Infinity, vLLM, TEI, llama.cpp) over Tailscale/VPN, or any cloud
endpoint — with higher-quality models like bge-large-en-v1.5 (1024d).
Configuration via environment variables:
GITNEXUS_EMBEDDING_URL=http://your-server:8080/v1
GITNEXUS_EMBEDDING_MODEL=BAAI/bge-large-en-v1.5
GITNEXUS_EMBEDDING_API_KEY=your-key (default: 'unused')
GITNEXUS_EMBEDDING_DIMS=1024 (auto-detected if omitted)
When env vars are set:
- initEmbedder() skips local model download entirely
- embedText() and embedBatch() call the HTTP endpoint
- Dimensions auto-detected from first response
- Batches in groups of 64
When env vars are NOT set:
- Existing local transformers.js behavior is completely unchanged
Build: tsc clean
Tests: 1776 passed, 0 failed
Add constructorTypeMap to buildTypeEnv — populated during walk when a
declaration has both a type annotation and a constructor initializer.
Add isSubclassOf helper (BFS, depth-5, cycle-safe).
In call-processor, consult constructorTypeMap to override receiver type
when constructor creates a known subclass (same-file only).
Add inferLiteralType to LanguageTypeConfig for Java, Kotlin, C#, C++.
In resolveCallTarget, when multiple candidates survive arity filtering,
lazily infer argument literal types and filter by parameterTypes match.
Worker path falls through gracefully (no AST available).
Add parameterTypes?: string[] to SymbolDefinition and MethodSignature.
Extract per-parameter type names via extractSimpleTypeName during
parsing for overload disambiguation (Java, Kotlin, C#, C++).
Thread through both sequential (parsing-processor) and worker
(parse-worker) paths.
The fileIndex Map stored SymbolDefinition per name, silently dropping
earlier overloads via Map.set(). Changed to SymbolDefinition[] so all
same-name methods (e.g., Java overloads) survive in same-file resolution.
Added lookupExactAll() for resolution-context to pass all same-file
candidates through to candidate filtering.
- Remove dead replayPendingItems array and inert if-block in type-env.ts
- Add nullable_type fallback in extractKotlinDeclaration for val x: User? local vars
- Tighten isCSharpNullableDecl to avoid substring false positives on type names
- Add missing scope boundaries: function_expression (TS), constructor_declaration/
local_function_statement/lambda_expression (C#) in null-check narrowing walkers
- Extend null-check narrowing fixtures and add 4 integration tests covering:
Kotlin local variable nullable, C# constructor + lambda, TS function expression
Adds 17 new fixture directories and 23 new describe blocks covering every
feature in Milestone D (Phases A, B, C) with full cross-language integration
test coverage:
Phase A — Fixpoint Completeness:
- TS/JS object destructuring (const { field } = obj → fieldAccess resolution)
- TS/JS post-fixpoint for-loop replay (iterable var resolved by fixpoint)
- Rust struct_pattern destructuring (let Point { x, y } = p)
Phase B — Inheritance & Receivers:
- Grandparent MRO (depth-2 C→B→A) for all 9 OOP languages:
TS, Kotlin, C#, C++, Java, PHP, Python, Ruby, JS
- Go inc/dec write access (obj.Field++/-- emit ACCESSES write edges)
Phase C — Branch-Sensitive Narrowing:
- Null-check narrowing for TS (!==null, !=null, !==undefined),
C# (!=null, is not null), and Kotlin (!=null)
Bug fix — Kotlin null-check narrowing (3 issues in jvm.ts):
1. patternBindingNodeTypes registered 'comparison_expression' but
tree-sitter-kotlin produces 'equality_expression' for !=
2. Handler checked for 'null_literal' named child but 'null' is an
anonymous node in the Kotlin grammar
3. extractKotlinParameter only searched for 'user_type' direct child,
missing 'nullable_type' wrapper (so x: User? never got a base binding)
17 fixtures, 23 describe blocks, 705 new lines of test code, 0 failures.
Review follow-ups from compiler front-end review (#379):
- Rust extractPendingAssignment now calls unwrapAwait() on value before
type checks, so `let user = get_user().await` resolves correctly
- type-resolution-system.md: removed "no fixpoint inference" from
limitations, updated "Single-pass" to "Walk + fixpoint", replaced
stale single-pass Tier 2 description with fixpoint loop explanation
- type-resolution-roadmap.md: Phase 9 body updated — 9C is delivered,
9B walk-order dependency documented (for-loop Tier 0b runs before
fixpoint, so fixpoint-resolved types can't update loop variables)
- Added this/self/$this fixpoint gap footnote to feature matrix
Replace the sequential Tier 2b/2a propagation with a unified fixpoint
loop that handles four binding kinds: callResult, copy, fieldAccess,
and methodCallResult. The loop iterates until no new bindings are
produced (max 10 iterations), enabling arbitrary-depth mixed chains:
const user = getUser(); // callResult → User
const addr = user.address; // fieldAccess → Address
const city = addr.getCity(); // methodCallResult → City
city.save(); // resolves to City#save
Infrastructure:
- PendingAssignment union extended with fieldAccess and methodCallResult
- resolveFieldType helper: typeName → class nodeId → lookupFieldByOwner
- resolveMethodReturnType helper: typeName → class nodeId → lookupFuzzyCallable filtered by ownerId
- Fixpoint also resolves reverse-order copy chains that single-pass missed
Languages: TS, JS, Java, Kotlin, C#, Go, Rust, Python, PHP, Ruby, C++.
Each gets field access and/or method-call-with-receiver detection in
extractPendingAssignment, plus method-chain-binding test fixtures.
Activate the dormant Tier 2b pendingCallResults infrastructure in
type-env.ts by extending each language's extractPendingAssignment to
emit { kind: 'callResult', lhs, callee } when the RHS of an untyped
variable declaration is a simple function call.
This enables `var user = getUser(); user.save()` to resolve at TypeEnv
build time. Tier 2b now runs before Tier 2a copy-propagation, enabling
mixed chains like `const user = getUser(); const alias = user;
alias.save()`.
Languages: TS, JS, Java, Kotlin, C#, Go, Rust, Python, PHP, Ruby, C++.
Swift excluded. Each language gets a call-result-binding test fixture
and integration tests.
Conservative: only simple calls (no method calls with receivers), only
when exactly one callable matches, first-writer-wins.
* feat: upgrade @ladybugdb/core to 0.15.2 and remove segfault workarounds
The upstream fix (ladybug-nodejs#1) resolves the child QueryResult lifetime
segfault, making .close() safe on all platforms. This removes 6 workaround
sites:
- Remove `dangerouslyIgnoreUnhandledErrors` from vitest config
- Remove platform-conditional .close() guards in global-setup and test helper
- Delete test/setup.ts (process._getActiveHandles unref hack)
- Replace no-op cleanup in test-indexed-db.ts with real adapter close
- Fix pool adapter closeOne() to properly close connections with shared
Database refcount guard and orphaned connection handling in checkin()
- Update segfault-related comments across the codebase
Also bumps @ladybugdb/wasm-core to ^0.15.2 in gitnexus-web for consistency.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: keep dangerouslyIgnoreUnhandledErrors for macOS N-API exit crash
The N-API destructor ordering crash during worker fork exit on macOS is
independent of the QueryResult lifetime fix in 0.15.2. Tests pass, but
the exit triggers a crash. Keep the flag with an updated comment
explaining the actual cause. Can be removed once LadybugDB fixes all
destructor ordering issues upstream.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* ci: unify test run for single-pass coverage
- Update `npm test` to run all tests (unit + integration + lbug-db)
via `vitest run` instead of `vitest run test/unit`
- Add `test:unit` script for running unit tests only
- Remove `ci-integration.yml` — the per-file lbug-db process isolation
is no longer needed with `dangerouslyIgnoreUnhandledErrors` and
`fileParallelism: false` handling fork exit issues
- Update `ci-unit-tests.yml` to run all tests with build + coverage
- Simplify `ci.yml` gate (two jobs: quality + tests)
- Simplify `ci-report.yml` (single coverage artifact, no merge step)
* fix: update cli-commands test for renamed test:all → test:unit script
* fix: set USERPROFILE in setup-skills test for Windows compatibility
os.homedir() checks USERPROFILE on Windows, not HOME.
* fix: add isolate: false to lbug-db project to prevent fork crashes
On macOS, N-API destructors crash fork workers on exit. With
isolate: true (default), vitest recycles the fork between files,
triggering the crash after each file. After several crashes, the
remaining lbug-db files never execute.
isolate: false keeps all 8 lbug-db files in a single fork — the
fork only exits once after all files complete, and that single exit
crash is caught by dangerouslyIgnoreUnhandledErrors.
* fix: add unique sequence.groupOrder to vitest projects
Vitest v4 requires unique groupOrder when projects have different
maxWorkers (lbug-db has fileParallelism: false → maxWorkers: 1).
* fix: await async close() in global-setup and remove isolate: false
global-setup.ts called conn.close() and db.close() without await —
these return Promise<void> in @ladybugdb/core 0.15.2. The setup
function returned before the DB was fully closed, so vitest forks
hit a stale file lock when opening the same DB path, crashing the
lbug-db worker before any test ran.
isolate: false caused native state corruption after 2-3 open/close
cycles in the same fork (vitest-specific, not reproducible in plain
Node.js). Without it, each file gets its own module scope and the
N-API destructor crash at fork exit is caught by
dangerouslyIgnoreUnhandledErrors.
Also fixes fire-and-forget close() calls in the pool adapter —
try/catch around an async close() never catches rejections; changed
to .catch(() => {}) for proper unhandled-rejection prevention.
Before: 0/8 lbug-db files ran on macOS CI (fork crash).
After: 8/8 pass, 84 files, 3077 tests, zero errors.
* fix: update project index references in AGENTS.md and CLAUDE.md to reflect correct symbol counts and relationships
* feat: enhance lbug adapter with external database support and write operation validation
* feat: create ci-tests workflow for comprehensive test coverage across platforms
* ci: move PR report inline to ci.yml, delete ci-report.yml
The old ci-report.yml used workflow_run which always runs code from
the default branch (main). This meant the PR comment used main's
stale report template that still referenced the old unit/integration
split architecture — causing "Merge coverage reports" failures.
Moving the report inline to ci.yml means it runs from the PR branch
and uses the current report template. The report now shows:
- per-platform status (Ubuntu/Windows/macOS columns)
- unified test counts from the single vitest run
- coverage with base branch (main) delta comparison
- commit SHA for traceability
Also removes the save-pr-meta job since the report no longer needs
a separate workflow_run trigger.
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
* feat: Phase 1 ACCESSES edge type — read tracking from chain resolution
Add ACCESSES relationship type to track field read access during call
chain resolution. When walkMixedChain resolves a field access (e.g.,
user.address.save()), an ACCESSES edge with reason 'read' is emitted
from the calling function to the Property node.
Schema: ACCESSES added to RelationshipType, REL_TYPES, VALID_RELATION_TYPES,
context queries, tools/resources descriptions. Excluded from default
impact BFS to prevent traversal explosion.
Implementation: resolveFieldAccessType now returns FieldResolution with
fieldNodeId. walkMixedChain accepts optional onFieldResolved callback.
makeAccessEmitter factory provides Set-based dedup per source node.
Bug fix: Added Java 'field_access' to FIELD_ACCESS_NODE_TYPES — was
missing, causing extractMixedChain to fail for Java member access.
* feat: Phase 2 ACCESSES write edges — assignment detection across 12 languages
Add tree-sitter query patterns for field write detection (obj.field = value)
across all supported languages: TS/JS, Python, Java, Go, C++, C#, Rust,
PHP, Ruby (setter syntax), Kotlin, Swift.
Processing: Sequential path handles assignment captures inline. Worker
path extracts ExtractedAssignment data for deferred resolution via new
processAssignmentsFromExtracted function.
Bug fix: Kotlin/Swift assignment queries used invalid navigation_expression
wrapper — fixed to match actual directly_assignable_expression AST structure.
Tests: Write access integration tests for TS, Java, Python, Go with
dedicated fixtures. All use strict toBe() assertions.
* test: add unit tests for call-routing, shared type extractors, and symbol-table branches
Add 215 new unit tests across 3 files to increase branch coverage toward
the 23% global threshold (was 21.49%):
- call-routing.test.ts (49 tests): Ruby call routing — require/require_relative,
include/extend/prepend heritage, attr_accessor properties with YARD types
- shared-type-extractors.test.ts (108 tests): pure string functions —
extractElementTypeFromString, stripNullable, extractReturnTypeName,
methodToTypeArgPosition, getContainerDescriptor
- symbol-table.test.ts (+29 tests): Property/fieldByOwner index, metadata
spread branches, lazy callable index, lookupExactFull shape
* fix: defer write-access resolution to fix Ruby cross-file property timing
Ruby attr_accessor properties are registered during processCalls (not
the parsing phase), so lookupFieldByOwner fails when service.rb is
processed before models.rb. Fix by collecting pending write-access
edges during the file loop and resolving them after all files are done.
Also adds write-access integration tests and fixtures for 7 languages
(C++, C#, JS, Kotlin, PHP, Ruby, Rust), Ruby compound assignment query,
PHP static property write query, and Kotlin property type extraction.
* fix: address PR #372 review — write-access constructor bindings parity and docs
- Add verified constructor bindings fallback to write-access resolution
in both sequential path (receiverIndex lookup) and worker path
(constructorBindings param for processAssignmentsFromExtracted),
closing the read/write ACCESSES edge asymmetry for factory-returned
receivers
- Clarify inner guard control flow comment in processCalls match loop
- Document Go inc_statement/dec_statement gap in roadmap
- Clarify PHP nullsafe write footnote (invalid syntax, not just untracked)
- Update symbol-table tests for intentional fieldByOwner behavior change
(Properties without declaredType now indexed for dynamic language
write-access tracking)
The SupportedLanguages enum includes Kotlin but the web project's
LANGUAGE_QUERIES and languageFileMap Records were missing it, breaking
the Vercel build with TS2741.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: Phase 8 field/property type resolution — resolve chained member access
Add field/property type extraction to the type resolution system so that
chained member access like `user.address.save()` resolves the intermediate
receiver type (`address → Address`) through Property symbols in SymbolTable.
Key changes:
- SymbolTable: add `declaredType` field, `fieldByOwner` O(1) index,
`lookupFieldByOwner()` method, P0 conditional callableIndex invalidation,
P2 exclude Properties from globalIndex to prevent namespace pollution
- tree-sitter queries: add `definition.property` for TypeScript, Java, Go
- parse-worker: extract declared types for Property nodes via
`extractPropertyDeclaredType()`, capture field-access receiver info
- call-processor: add `resolveFieldAccessType()` helper and field-access
branch in both sequential and worker receiver resolution paths
- Integration tests: new field-types test suite verifying end-to-end
`user.address.save() → Address#save` resolution
* fix: Go tree-sitter query captures field_declaration not field_declaration_list
Post-review fix: the Go struct field query incorrectly put @definition.property
on field_declaration_list (the list container) instead of field_declaration
(the individual field). Also removed unused `language` parameter from
extractPropertyDeclaredType.
* feat: expand field-type tests to 6 languages, fix Go ownerId and Kotlin navigation_expression
- Add integration test fixtures for Java, C#, Go, Kotlin, PHP (alongside existing TS)
- Fix Go: add type_declaration handling in findEnclosingClassId for struct fields
(field_declaration → field_declaration_list → struct_type → type_spec → type_declaration)
- Fix Kotlin: add navigation_expression handling in field-access resolution
(Kotlin uses navigation_expression + navigation_suffix, not member_expression)
- Add extractMemberAccessParts helper in call-processor for cross-language member access
- All 24 field-type tests pass across 6 languages, 181 Go+Kotlin tests pass with no regressions
* refactor: split HAS_METHOD into HAS_METHOD + HAS_PROPERTY edge types
Property nodes now use HAS_PROPERTY edges instead of HAS_METHOD, giving
the graph schema proper semantic separation between methods and fields.
- HAS_METHOD: Method, Constructor, Function (when inside a class)
- HAS_PROPERTY: Property nodes (class fields, struct fields, attributes)
MRO processor only reads HAS_METHOD — properties correctly excluded from
method resolution order. Impact analysis accepts both edge types.
Updated 12 files: graph types, schema, tools docs, parse-worker,
parsing-processor, call-processor, and 6 test files.
* fix(test): update security test to expect 7 VALID_RELATION_TYPES (added HAS_PROPERTY)
* test: add unit tests for Phase 8 SymbolTable features (39 tests, up from 19)
Cover all new branches: declaredType metadata, Property exclusion from
globalIndex, conditional callableIndex invalidation, lookupFieldByOwner
(happy path + edge cases), lookupFuzzyCallable filtering, and clear()
with fieldByOwner. Fixes branch coverage threshold (21.8% → 23%+).
* feat: Phase 8B mixed field+method chain resolution, C++/Rust chain fixes
Unify field and method chain resolution into a single `extractMixedChain`
walker that handles interleaved patterns like `svc.getUser().address.save()`.
Fix C++ chain calls (tree-sitter-cpp `field_expression` uses `argument` not
`object`), Rust unit struct instantiation (`let svc = TypeName;`), and add
stdlib passthrough for `unwrap()`/`clone()`/`expect()` in chain loops.
Key changes:
- Replace `receiverCallChain` + `receiverFieldAccess` with unified
`receiverMixedChain: MixedChainStep[]` on ExtractedCall
- Add `extractMixedChain` in utils.ts (handles both call_expression and
field_expression nodes, including C++ `argument` field)
- Add `TYPE_PRESERVING_METHODS` set for stdlib identity operations
- Add C++ inline method double-indexing guard in parsing-processor.ts
and parse-worker.ts
- Add Rust unit struct recognition in type-extractors/rust.ts
- Split field-types.test.ts into per-language test files
- Add ts-mixed-chain fixture and integration tests
- Resolve rust.test.ts todo: Option<T>.unwrap().save() now works
- Update roadmap: Phases 7+8 complete, Phase 9 is next
* fix: Python declaredType extraction and sequential-path property registration
- Move @definition.property capture from expression_statement to assignment
node in Python queries so Strategy 1 childForFieldName('type') succeeds
- Pass item.declaredType through ctx.symbols.add in sequential call-processor
path, matching worker path behavior (fixes Ruby YARD declaredType drop)
- Add Python chain resolution integration test (user.address.save → Address#save)
- Update Rust/Python status in roadmap and system docs to reflect actual coverage
* fix: Python/Ruby field type disambiguation and Rust chain test
Three fixes from PR #354 third review:
1. Python typed_parameter name extraction: tree-sitter-python's
typed_parameter uses positional children for the name, not a named
field. TypeEnv and extractParameter now fall back to firstNamedChild.
2. Ruby/Python call-step field resolution: Ruby's AST uses `call` nodes
for both property access and method calls. The chain walker now tries
resolveFieldAccessType before resolveCallTarget for call steps, so
attr_accessor properties resolve via declaredType.
3. Rust chain resolution test: added missing integration test asserting
user.address.save() resolves to Address#save.
Also splits C/C++ and TS/JS columns in type-resolution-system.md
language matrix with footnotes for accuracy.
1062 resolver integration tests passing, 0 failures.
* refactor: Phase 8 code review cleanup — extract walkMixedChain, fix MCP agent gaps
- Extract duplicated chain resolution loop into shared walkMixedChain() helper,
eliminating ~60 lines of copy-pasted code between sequential and worker paths
- Add returnType to ResolveResult, removing redundant lookupFuzzy+find per chain step
- Fix context() tool to include HAS_METHOD, HAS_PROPERTY, OVERRIDES in queries
so agents can discover class members
- Fix p.declaredType Cypher example (column doesn't exist) → p.description
- Add HAS_METHOD, HAS_PROPERTY, OVERRIDES to schema resource
- Document HAS_METHOD/HAS_PROPERTY in impact tool description
- Delete dead code extractMemberAccessParts (superseded by extractMixedChain)
- Replace any with SyntaxNode on extractPropertyDeclaredType
- Add Rust deep-field-chain test (5 tests), Java mixed-chain (4), Go mixed-chain (4)
- All 1075 tests pass (13 new, 0 regressions)
* refactor: type SymbolDefinition.type as NodeLabel, add O(1) receiver index
- Change SymbolDefinition.type from string to NodeLabel union (35 members)
across symbol-table.ts, parse-worker.ts, parsing-processor.ts — compiler
now enforces correctness at all comparison/assignment sites
- Replace O(N*M) linear scan in lookupReceiverType with pre-built
ReceiverTypeIndex (Map<funcName, Map<varName, Entry>>) for O(1) lookups
with proper ambiguity handling and file-level fallback
- All 1075 tests pass, 0 regressions
* fix: capture C++ pointer/ref fields, Kotlin data class props, PHP constructor promotion
Add tree-sitter query patterns for three previously missed property declaration
forms: C++ pointer/reference member fields (Address* addr; Address& ref;),
Kotlin primary constructor val/var parameters (data class User(val name: String)),
and PHP 8.0+ constructor property promotion (public Address $address).
Fix "10 languages" off-by-one in docs (Ruby is single-level only, not deep chain).
Update Python feature matrix cell from No* to Yes* after 31b95f0 fix.
11 new integration tests with per-language fixtures verify property capture,
HAS_PROPERTY edge emission, and field-access chain resolution.
* fix: MCP server crashes under parallel tool calls (#326)
* fix: ensure full connection pool is pre-created to avoid race conditions during query execution
* fix: improve graceful shutdown handling with exit codes
* fix: resolve critical concurrency bugs in connection pool init
- Add initPromises dedup map to prevent double-init race when parallel
tool calls trigger initLbug for the same repoId simultaneously
- Move pool.set() after FTS load so concurrent checkout can't grab a
connection mid-async-init (FTS race on available[0])
- Replace lazy createConnection growth path with integrity error — pool
is pre-warmed, lazy creation would silence stdout during active queries
- Add preWarmActive flag so watchdog timer skips stdout restore during
the synchronous pre-warm loop
- Unify stdout capture: server.ts imports realStdoutWrite from
lbug-adapter instead of capturing its own copy
* test: add connection pool parallel stability tests
7 integration tests covering concurrent query safety, waiter queue
overflow, stdout.write restoration, connection leak detection, initLbug
deduplication, atomic pool visibility, and mixed query types.
* fix: run LadybugDB tests sequentially via vitest projects config
Vitest's projects feature splits test files into two groups: lbug-db
(fileParallelism: false) and default (parallel). This prevents native
mmap file-lock conflicts on Windows without requiring the CI shell loop
locally.
* test: add enrichment Promise.all regression test for #292/#316
Verifies that 3 concurrent queries via Promise.all (the exact pattern
from the impact command's enrichment phase at local-backend.ts:1415)
complete without SIGSEGV on a pre-warmed connection pool.
* feat(type-resolution): Phase 7.1+7.2 foundation — ReturnTypeLookup, context object, pendingCallResults
- Move extractReturnTypeName + helpers from call-processor.ts to type-extractors/shared.ts
(breaks circular import risk: call-processor → type-env → type-extractors → call-processor)
- Add SymbolTable.lookupFuzzyCallable(name) — lazy callable-only index, O(1) per call,
invalidated on add(); avoids per-call .filter() on lookupFuzzy results
- Add ReturnTypeLookup interface (conservative: undefined when 0 or 2+ callables match)
- Add ForLoopExtractorContext interface — replaces 4 positional params with context object;
update all 10 language extractor implementations (go, ts, py, jvm×2, cs, rs, rb, php, c-cpp)
- Add PendingAssignment discriminated union (kind: 'copy' | 'callResult');
update PendingAssignmentExtractor in all 9 language extractors that implement it
- Wire buildTypeEnv: build ReturnTypeLookup from optional symbolTable; split pendingAssignments
into pendingCopies + pendingCallResults; add Tier 2b call-result propagation loop
- Update call-processor.test.ts to import extractReturnTypeName from shared.ts
* feat(type-resolution): Phase 7.3 — call_expression iterables in for-loop extractors (7 languages)
Extends for-loop type extraction in all 7 typed-iteration languages to
resolve element types when the iterable is a direct function call.
**New capability**: `for (var u : getUsers())` in Java, `for u in get_users()`
in Python, `for user in getUsers()` in TypeScript, etc. now resolve
`u`/`user` to the callee's return element type via lookupRawReturnType +
extractElementTypeFromString.
Changes per language:
- types.ts: extend ReturnTypeLookup with lookupRawReturnType (raw return
string for container-type extraction); update ForLoopExtractorContext
with returnTypeLookup field
- type-env.ts: implement lookupRawReturnType on the concrete ReturnTypeLookup
built in buildTypeEnv (same guards as lookupReturnType, no extractReturnTypeName)
- go.ts: call_expression branch in range_clause — identifier func or
selector_expression method; existing isChannelType guards updated
- typescript.ts: identifier fn branch inside call_expression handler
- python.ts: identifier fn branch inside call handler
- jvm.ts (Java): method_invocation without object field in enhanced_for_statement
- jvm.ts (Kotlin): simple_identifier callee branch in call_expression node
- csharp.ts: identifier fn branch in invocation_expression handler
- rust.ts: identifier func branch in call_expression handler (alongside
existing field_expression/method-call path)
All branches follow the same conservative pattern:
lookupRawReturnType(callee) → extractElementTypeFromString → bind loop var
* feat(type-resolution): Phase 7.4 — PHP \$this->property iterable via @var class property scan
Adds Strategy C to PHP's extractForLoopBinding for the pattern:
foreach (\$this->property as \$item)
when Strategy A (resolveIterableElementType) and Strategy B (scopeEnv lookup)
both fail to find the element type.
Strategy C: when the iterable is a member_access_expression with object '$this',
walk up the AST to the enclosing class_declaration, scan its declaration_list
for a property_declaration whose variable_name matches the property, and extract
the element type from:
1. PHPDoc @var annotation on a preceding comment sibling (/** @var User[] */)
2. PHP 7.4+ native type field (e.g. UserRepo \$repo — skips generic 'array')
This eliminates the @param workaround that was previously required in the
php-foreach-member-access fixture (which used @param User[] \$users on the method
to populate the method's scopeEnv with a \$users binding).
New helpers in php.ts:
- PHPDOC_VAR_RE: regex for @var extraction
- extractClassPropertyElementType: reads @var or native type from a property_declaration
- findClassPropertyElementType: scans class body for a named property
Tests added (type-env.test.ts):
- PHP: resolves from @var User[] without @param workaround
- PHP: conservative — no binding for unknown property
- PHP: multi-class file — both classes resolve independently
Fixture updated (php-foreach-member-access/App.php):
- Removed the @param User[] \$users workaround from processMembers()
- Test now validates the natural class-property-based resolution path
* docs: mark Phase 7 complete in type-resolution-roadmap.md
Records that 7A (call_expression iterables, 7 languages), 7B (PHP
$this->property via @var scan), and 7C (ReturnTypeLookup + context object)
are all shipped. Adds implementation notes and strikethroughs on resolved
language-specific gaps.
* fix(docs): update project references to feat-phase7-type-resolution in AGENTS.md and CLAUDE.md
* feat(type-resolution): Phase 7.5 — PHP call_expression foreach + integration tests for 7 languages
Add integration test coverage for Phase 7.3's call_expression iterable
resolution across all 7 languages (Go, TypeScript, Python, Java, Kotlin,
PHP, Rust). Each test creates a fixture with competing User/Repo classes
that both define save(), then verifies for-loop iteration over a function
call's return value resolves to the correct class.
PHP was missing function_call_expression support in its for-loop extractor.
Three changes fix this:
- php.ts extractForLoopBinding: handle function_call_expression and
member_call_expression iterables via returnTypeLookup
- php.ts normalizePhpReturnType: preserve array notation (User[]) in
SymbolTable so lookupRawReturnType returns useful container types
- parse-worker.ts + parsing-processor.ts: upgrade uninformative AST
return types (array, iterable) with PHPDoc @return annotations
35 new integration tests (5 per language), 2525 total tests passing.
* fix(type-resolution): address PR #341 review findings — PHP asymmetry + dormant infrastructure docs
- Replace normalizePhpType with extractElementTypeFromString in PHP call-expression
foreach paths, aligning with all 6 other language extractors and preventing
incorrect binding of bare non-container types like User
- Add NOTE comments clarifying pendingCallResults Tier 2b is infrastructure-ready
but no extractor populates it yet
- Expand Go channel-type comments explaining why non-channel assumption is safe
* fix(type-resolution): address verification review — docs accuracy + PHP fallback guard
- Roadmap lines 86/100: correct pendingCallResults from "active" to "dormant infrastructure (Phase 9)"
- type-resolution-system.md line 363: update to reflect Phase 7.3 loop inference is delivered
- type-resolution-system.md line 409: clarify for-loop call-expression resolution (done) vs general assignment propagation (pending)
- php.ts:127: add declaration_list type guard on fallback to prevent silent wrong results
* fix(impact): return structured error + partial results instead of crashing (#321)
- Wrap impact() in try-catch to return structured error JSON instead of
process crash (SIGSEGV/exit 139)
- Extract core logic to _impactImpl() for clean error boundary
- Break out of depth traversal loop on query failure, return partial
results collected so far (previously silently swallowed errors)
- Add 'partial' flag to response when traversal was interrupted
- Add try-catch in CLI impactCommand with structured error output
- Improve formatImpactResult to show suggestion text and partial warning
- Add 3 new unit tests for error/suggestion/partial scenarios
Fixes#321
* fix: address review feedback — 4 bugs from @claude review
Per @claude's review (requested by @magyargergo):
- [BUG 1] Consistent target field shape: error responses now return
{name: string} instead of raw string, matching success response schema
- [BUG 2] Remove misleading partial:true from total-failure responses
(partial is only meaningful when some depth levels succeeded)
- [BUG 3] Move getBackend() inside try-catch in impactCommand so
backend init failures return structured JSON instead of crashing
- [BUG 4] Safe error message extraction: use instanceof Error check
to handle thrown strings correctly (err?.message is undefined for
non-Error thrown values)
- [MINOR] Add radix argument to parseInt (10)
* test: add integration tests for impact error handling (#321)
Per @claude's recommendation (requested by @magyargergo):
- impact: structured error for unknown symbol (no crash)
- impact: error response has consistent {name: string} target shape
- impact: partial:true only set when some results were collected
Tests use existing withTestLbugDB + seeded graph fixture.
- 6 new unit tests for fastStripNullable branches (simple id, nullable union, bare keyword)
- 4 new integration tests for skipGraphPhases pipeline option
- Tests for SKIP_SUBTREE_TYPES and interestingNodeTypes code paths
* feat: Phase 6 type resolution — pattern matching, for-loop Tier 1c, coverage completion
- Add patternBindingNodeTypes gate to LanguageTypeConfig for 50% perf improvement
- Expand ForLoopExtractor signature with optional declarationTypeNodes + scope
- Add extractElementTypeFromString shared utility for container type parsing
- Python match/case: extractPatternBinding for `case User() as u:` pattern
- C# refactor: move is_pattern_expression from extractDeclaration to extractPatternBinding
- Ruby: add extractPendingAssignment for assignment chain propagation
- TS/JS: add for-loop Tier 1c for `for (const user of users)` with User[] inference
- Python: add for-loop Tier 1c for `for user in users:` with type annotation inference
- Go: add for-loop Tier 1c for `for _, user := range users` with []User inference
- Fix 'Property' as any stale cast in call-processor.ts
- Add dual return-type string length cap (2048 pre-cap, 512 post-cap)
- Add chain call integration tests for C#, Go, Rust, Python, JS, C++
- Add Python match/case integration test fixtures
- 27 new extractElementTypeFromString unit tests
- 3 for-loop edge cases skipped (declarationTypeNodes scope key lookup)
* fix: address code review findings for Phase 6
- Add missing patternBindingNodeTypes to C# typeConfig (perf gate)
- Add 2048-char input length guard to extractElementTypeFromString
- Skip Python match/case integration tests (call extraction needs query updates)
* reorganise
* fix: Phase 1 bug fixes — Go range semantics, typed_parameter, bracket depth
- Go single-var range correctly returns early for slices/maps (index, not element)
- Go single-var range on channels correctly resolves element type
- Added map_type and channel_type to extractGoElementTypeFromTypeNode
- Added isChannelType helper for channel detection before skip decision
- Added 'typed_parameter' to TYPED_PARAMETER_TYPES for Python annotated params
- Fixed bracket depth tracking in extractElementTypeFromString — only match
selected closeChar at depth 0, return undefined for mismatched brackets
- Un-skipped 3 prematurely skipped tests (TS local const, Python List/Sequence)
- Added tests for map range, single-var range semantics, bracket edge cases
* refactor: Phase 2 architecture — shared helper, required params, decoupled type nodes
- Extract resolveIterableElementType shared helper in shared.ts implementing
3-strategy fallback (declarationTypeNodes → scopeEnv string → AST walk)
- Refactor TS, Python, Go extractors to use shared helper (eliminates 3x duplication)
- Make ForLoopExtractor params required (aligned with PatternBindingExtractor)
- Update Java, Kotlin, C# extractor signatures to accept required params
- Decouple declarationTypeNodes from scopeEnv — capture raw type annotation
nodes BEFORE extractDeclaration for container types (User[], []User, List[User])
- Hybrid approach: direct name extraction + keysBefore fallback for multi-declarator
- Document declarationTypeNodes invariant change (superset of scopeEnv)
* feat: Phase 3 partial — Rust for-loop + C# var foreach Tier 1c
- Rust: add extractForLoopBinding with for_expression support
- Handles &users, &mut users via reference_expression unwrapping
- extractRustElementTypeFromTypeNode: generic_type, reference_type, slice/array
- findRustParamElementType: AST walk with reference/mut pattern unwrapping
- 4 unit tests (Vec<User>, &[User], range expr negative, no-annotation negative)
- C#: upgrade foreach to handle var (implicit_type) via Tier 1c
- extractCSharpElementTypeFromTypeNode: generic_name, array_type, nullable_type
- findCSharpParamElementType: AST walk to method_declaration parameters
- 3 unit tests (var foreach, explicit type regression, no-annotation negative)
* feat: Phase 3 complete — all language gaps + pattern matching
Kotlin Tier 1c:
- Unannotated for-loop resolves via shared helper
- extractKotlinElementTypeFromTypeNode handles type_projection unwrapping
- findKotlinParamElementType walks to function_declaration
Java Tier 1c:
- var foreach resolves via shared helper
- extractJavaElementTypeFromTypeNode handles generic_type, array_type
- findJavaParamElementType walks to method_declaration
TypeScript:
- readonly User[] unwrapped via readonly_type → array_type recursion
C# switch patterns:
- declaration_pattern added to patternBindingNodeTypes
- extractPatternBinding handles standalone declaration_pattern (switch case/expr)
Rust match arms:
- match_arm added to patternBindingNodeTypes
- extractPatternBinding extended with match_arm → match_expression parent traversal
Python:
- as_pattern tries childForFieldName('alias') before positional fallback
Tests: 237 pass (was 224), 13 new tests added
* feat: Phase 4 — known limitation tests, match arm fix, final verification
- Fix Rust match_arm pattern extraction: unwrap match_pattern to get
tuple_struct_pattern inside (tree-sitter-rust wraps in match_pattern node)
- Add first-writer-wins regression test for match arm scope leakage
- Add 5 documented skip tests for known limitations:
- TS destructured for-of (tuple destructuring)
- Python tuple unpacking in for-loops
- TS instanceof narrowing (block-level scoping)
- Rust for with .iter() (method call iterable)
- Ruby block parameters (closure param inference)
Final: 238 passed, 5 skipped (documented limitations), tsc clean
* test: integration tests for all Phase 6 language gaps + fix Rust param pattern field
Integration test fixtures and tests (30 new tests, all with exact match + negative):
Rust for-loop (5 tests):
- for user in &users with Vec<User> → User#save, negative Repo#save
- for repo in &repos with Vec<Repo> → Repo#save, negative User#save
Rust match arm (5 tests):
- match opt { Some(user) => user.save() } → User#save, negative Repo#save
- if let Ok(repo) = res → Repo#save, negative User#save
C# var foreach (5 tests):
- foreach (var user in users) with List<User> → User#Save, negative Repo#Save
- foreach (var repo in repos) with List<Repo> → Repo#Save
C# switch pattern (4 tests):
- is User user → User#Save, case Repo repo → Repo#Save
Kotlin unannotated for (4 tests):
- for (user in users) with List<User> → user.save, negative repo.save
Go map range (3 tests):
- for _, user := range userMap with map[string]User → User#Save, negative
TypeScript readonly (4 tests):
- for (const user of users) with readonly User[] → user.save, negative
Bug fix: type-env.ts parameter branch now falls back to childForFieldName('pattern')
for Rust parameters (Rust uses 'pattern' not 'name' for parameter names)
* test: add assertion bodies to known limitation skip tests
Convert empty skip test stubs to proper tests with parse/buildTypeEnv/expect
assertions following the codebase convention (e.g., call-processor.test.ts:319).
Each skip test now documents the exact expected behavior, so removing .skip
will cause a meaningful failure when the limitation is eventually fixed.
Also clarify Python integration skip tests as call-extraction issues (not
type-env) and Swift integration skips as build-dep issues (self/super
resolution code already exists in type-env.ts).
* feat: resolve 4 known limitation skip tests + method-aware type arg selection
Unskip 4 of 5 type-env known limitations with full integration test coverage:
1. TS destructured for-of: handle array_pattern by binding last named child
to element type. Fix Map<K,V> to return last generic arg (value type).
2. Python dict.items() loop: handle `call` iterables + `pattern_list` left
side. Fix dict[K,V] extraction via type_parameter with last-arg heuristic.
Unwrap `type` wrapper in extractPyElementTypeFromAnnotation.
3. TS instanceof narrowing: add extractPatternBinding for binary_expression
with positional child access. First-writer-wins (not block-scoped).
4. Rust .iter() for-loops: handle call_expression in for_expression value
node by extracting receiver from field_expression.
Method-aware type arg resolution:
- Add TypeArgPosition ('first'|'last') to resolveIterableElementType
- .keys()/.keySet()/.Keys → first type arg (key); all else → last (value)
- Thread position through all 3 strategy callbacks in TS/Rust/Python
- Add predefined_type to extractSimpleTypeName for TS primitives (string etc)
New fixtures: rust-iter-for-loop, typescript-destructured-for-of,
typescript-instanceof-narrowing, python-dict-items-loop.
248 unit tests pass (6 new), 1 skip (Ruby block params).
* feat: container descriptor table for generic type arg resolution
Replace simple KEY_METHODS heuristic with CONTAINER_DESCRIPTORS table
that maps 30+ container types across all languages to their type parameter
semantics per access method.
Key improvements:
- Container-aware resolution: HashMap.iter() correctly yields V (arity 2),
while Vec.iter() yields T (arity 1) — same method, different semantics
- Cross-language coverage: Map/HashMap/BTreeMap/dict/Dict/Dictionary/
ConcurrentHashMap + List/Vec/Set/HashSet/Queue/Deque/Stack etc.
- Method categorization: keyMethods (keys/keySet/Keys) vs valueMethods
(values/get/pop/iter/first/last) per container type
- Fallback for unknown containers: still uses method name heuristic,
so MyCache<K,V>.keys() correctly returns first arg
- Exported getContainerDescriptor() for future heritage-chain lookups
Each language extractor now passes containerTypeName from scopeEnv to
methodToTypeArgPosition for descriptor-aware resolution.
252 unit tests pass (4 new descriptor tests), 1 skip (Ruby).
* feat: method-aware for-loop extractors + integration tests for all languages
Upgrade 4 existing extractors + create 3 new ones for full cross-language
coverage of call_expression iterables and container descriptor resolution:
Upgraded (add call expr iterable + methodToTypeArgPosition):
- Java: method_invocation (data.keySet(), data.values())
- Kotlin: navigation_expression + call_expression (data.keys, data.values())
- C#: member_access_expression + invocation_expression (data.Keys, data.Values)
- Go: TypeArgPosition threading for Go 1.18+ generics
New for-loop extractors:
- C++: for_range_loop with auto& unwrapping, template_type + qualified_identifier
(std::vector<User>) extraction, explicit vs auto type handling
- PHP: foreach_statement with simple/key-value/by-reference forms, PHPDoc
@param priority over AST array type
- Ruby: for-in with YARD @param type resolution via comment parsing
Integration test fixtures + tests for all 6 languages:
- java-map-keys-values (Map.values() + List iteration)
- kotlin-map-keys-values (HashMap.values + List iteration)
- csharp-dictionary-keys-values (Dictionary.Values foreach)
- cpp-range-for (auto& + const auto& range-based for)
- php-foreach-loop (foreach with PHPDoc @param User[])
- ruby-for-in-loop (for-in with YARD @param Array<User>)
Bugs fixed during integration testing:
- C++: qualified_identifier (std::vector) not unwrapped to template_type
- PHP: extractParameter overwrote PHPDoc-derived types with bare 'array'
252 unit tests pass, 201 integration tests pass across 6 languages.
* fix: update extractElementTypeFromString tests for last-arg default
TypeArgPosition change (default 'last') broke 5 existing tests expecting
first arg from multi-arg generics. Updated expectations and added explicit
pos='first' tests for key type extraction.
* fix: rename C++ fixture files to correct case for case-sensitive CI
On case-sensitive filesystems (Linux/macOS CI), git tracked both the old
lowercase files (app.cpp, user.h) and the new uppercase files (App.cpp,
User.h) as separate files. The pipeline processed both, causing the old
app.cpp (with explicit User& type) to interfere with the new auto& test.
Removes old lowercase entries and re-adds with uppercase casing to match
the #include directives in the fixture.
* feat: PR #318 review findings — pattern bindings, member access iterables, structured bindings
Address all 7 genuine gaps identified in PR #318 deep code review:
- Kotlin: add extractKotlinPatternBinding for when/is (type_test AST node)
with allowPatternBindingOverwrite for smart-cast semantics
- Java: add type_pattern branch for Java 17+ switch pattern variables
- TypeScript: explicit object_pattern skip in for-of (no false bindings)
- Cross-language: member access iterables (self.users, this.users, repo.users)
across all 10 language extractors
- C++: structured_binding_declarator handling in range-for (last-child heuristic)
- Rust: closure_parameter added to TYPED_PARAMETER_TYPES
- PHP: normalizePhpType handles angle-bracket generics (Collection<User>)
Code review fixes applied:
- Remove 4 debug console.log statements (c-cpp.ts, call-processor.ts)
- Hoist KNOWN_CONTAINER_PROPS to module scope (csharp.ts)
- Guard keysBefore allocation behind typeNode check (type-env.ts)
- Add depth limits (50) to 7 recursive type extraction functions
- Add 2048-char length cap to extractSimpleTypeName
- Fix PHP/Ruby missing typeArgPos parameter in resolveIterableElementType
Integration test fixtures: kotlin-when-pattern, java-switch-pattern,
cpp-structured-binding, typescript-member-access-for-loop,
python-member-access-for-loop
* fix: position-indexed when/is bindings, Kotlin param extraction, HashMap.values for-loop
Three root causes for failing Kotlin integration tests:
1. When/is multi-arm resolution: flat scopeEnv stored only the last arm's
type (last-writer-wins). Added PatternOverrides with AST range indexing
so each when arm resolves to its narrowed type independently.
2. HashMap.values for-loop: navigation_expression without call_suffix was
classified as bare property access (iterableName='values' instead of
'data'). Now tries object-as-iterable + property-as-method first, with
fallback to property-as-iterable for this.users patterns.
3. Kotlin parameter extraction: tree-sitter-kotlin parameter nodes use
positional children (simple_identifier, user_type) not named fields
(name, type). Added fallback to findChildByType in both
extractKotlinParameter and extractTypeBinding.
Integration tests added for .keys/.values/Set/MutableMap iteration,
3-arm when/is, multi-call within arms, and when+else branch.
* feat: enhance PHP type resolution for generics and member access in foreach loops
* feat: Phase 6.1 type resolution gap closure — container descriptors, recursive_pattern, class fields
Add 13 missing container type descriptors (Collection, MutableMap, Stream, SortedSet, etc.)
to CONTAINER_DESCRIPTORS for correct element type extraction across C#, Kotlin, and Java.
Extend C# pattern binding to handle recursive_pattern (obj is User { Name: "Alice" } u)
in both is-expression and switch expression contexts.
Add TypeScript class field declaration support (public_field_definition) so for-loop
iteration over this.fieldName resolves element types from class field type annotations.
Includes file-scope fallback in resolveIterableElementType and nested member_expression
handling for this.field.method() patterns.
* docs: add type resolution system documentation with roadmap
Covers the full architecture, resolution tiers (0-2), scope model,
language feature matrix, container descriptors, pipeline integration,
and the Phase 7-9 roadmap for cross-scope propagation, field-type
resolution, and return-type-aware binding.
* feat: Phase 6.2 review findings — C# nested member foreach, C++ deref range-for, Java field_access
Close two gaps found during fourth-pass review of PR #318:
- C# foreach (var user in this.data.Values): nested member_access_expression
now extracts intermediate property name for scopeEnv lookup
- C++ for (auto& user : *ptr): pointer_expression dereference now recognized
as range-for iterable
Root causes fixed in shared infrastructure:
- extractSimpleTypeName: add template_type (C++) and generic_name (C#)
- extractGenericTypeArgs: add generic_name for consistency
- type-env.ts: unwrap variable_declaration wrapper in field_declaration
for declarationTypeNodes capture (zero-allocation manual loop)
Additional review findings addressed:
- Java: add field_access handler for this.data.values() in method_invocation
- C++ pointer_expression: document limitation (*identifier only)
- TypeScript: fix stale comment about property_identifier
All 525 tests pass (278 unit + 247 integration).
* perf: optimize type resolution pipeline — worker threshold, skip graph phases, AST pruning
- Skip worker pool creation for small repos (<15 files or <512KB) — saves 100-400ms
- Add skipGraphPhases option to runPipelineFromRepo to skip MRO/community/process phases
- Add conservative SKIP_SUBTREE_TYPES for leaf-only AST nodes (string, comment, number)
- Pre-compute interestingNodeTypes set — single Set.has() replaces 3 checks per node
- Add fastStripNullable — skip full stripNullable for simple identifiers (90%+ case)
- Replace .children?.find() with manual for loops in extractFunctionName (no array alloc)
- Add hookTimeout: 120000 to vitest.config.ts for CI beforeAll hooks
* fix: review findings — remove template_string from SKIP_SUBTREE_TYPES, handle bare nullable keywords
- Remove template_string and concatenated_string from SKIP_SUBTREE_TYPES
(template literals contain interpolated expressions with typed code)
- Add FAST_NULLABLE_KEYWORDS check to fastStripNullable for behavioral
parity with stripNullable on bare null/undefined/void/None/nil
- Add explanatory comment on extractPendingAssignment scopeEnv guard
* feat: add type resolution system and roadmap documentation
* fix(resolver): prefer same-directory file for Python bare imports
Python's sys.path searches the importing script's own directory first,
so `import user` from services/auth.py should resolve to services/user.py
even if models/user.py was indexed first in the suffix index.
Add a proximity check in resolveImportPath that consults the existing
dirMap index (O(1)) before falling back to global suffix matching, for
single-segment bare Python imports only.
Made-with: Cursor
* refactor(resolver): replace dirMap scan with O(1) allFiles.has() for proximity check
The previous implementation used index.getFilesInDir() + siblings.find()
which had two issues:
- dirMap stores all suffix levels, so getFilesInDir('services') matched
files from every directory named 'services/' across the repo — false
positives in monorepos
- siblings.find() was an O(n) linear scan despite the O(1) claim
Replace with a direct allFiles.has(importerDir + '/' + name + '.py') lookup.
allFiles is a Set<string> of full repo-relative paths, so the lookup is
truly O(1) and exact — no suffix ambiguity possible.
Also fixes: dead code (the '.rb' branch was unreachable since the outer if
gates on Python), and Windows backslash handling via normalize before split.
Made-with: Cursor
* test: remove flag-based demo from unit tests
Made-with: Cursor
* fix(resolver): cover package __init__.py in proximity check and add end-to-end CALLS test
- Also try importerDir/name/__init__.py as a second O(1) candidate so that
`import user` resolves to services/user/__init__.py when the target is a
package rather than a bare module file
- Add unit tests for package proximity, __init__.py fallback, and Windows
backslash path handling
- Add end-to-end CALLS assertion to the bare-import integration test:
svc.execute() must resolve to UserService#execute in services/user.py,
proving the fix propagates correctly through the type inference pipeline
Made-with: Cursor
* refactor: extract Python import resolution into resolvers/python.ts
- Move PEP 328 relative import and proximity-based bare import logic
from standard.ts into a dedicated resolvers/python.ts (resolvePythonImport)
- Dispatch Python imports from resolveLanguageImport in import-processor.ts,
consistent with how Ruby, PHP, and other languages are handled
- standard.ts is now language-agnostic (TS/JS aliases, Rust paths, suffix fallback)
- Add inline comment on __init__.py vs .py resolution order edge case
- Update unit tests to call resolvePythonImport directly
Made-with: Cursor
* docs: add PEP 302/328/451 references to python.ts comments
Made-with: Cursor
* fix(python): address reviewer comments on PEP compliance
- Guard dirParts.pop() against over-traversal: return null when dot
count exceeds directory depth, matching CPython's ImportError for
'attempted relative import beyond top-level package' (PEP 328)
- Swap __init__.py / .py check order to match CPython's finder
precedence (PEP 451 §4); coexistence is physically impossible so
order only matters for spec compliance
- Fix overstated PEP 302 comment: proximity check is a static
heuristic, not a sys.path[0] implementation
- Acknowledge namespace package gap (PEP 420) in docstring
- Add unit test for over-traversal guard
Made-with: Cursor
* test(python): document namespace package resolution behaviour
Add two unit tests for PEP 420 namespace packages (directory with no
__init__.py): bare import returns null (expected — no file exists to
resolve to, CPython sets __file__ = None), while the submodule form
(import user.model) resolves correctly via suffixResolve fallback.
Made-with: Cursor
---------
Co-authored-by: chirag-nighut <chiragnighut@gmail.com>
- Add Codex to Editor Support table
- Add Codex manual config example (~/.codex/config.toml)
- Update editor list in usage table
Fixes#131
Made-with: Cursor
* feat: Phase 5 type resolution — chained calls, pattern matching, class-as-receiver, code review fixes
Phase 5.1: Chained method call resolution (depth-capped at 3)
- resolveChainedReceiver() resolves a.getUser().save() by walking the chain
and looking up intermediate return types from the SymbolTable
- extractReceiverNode() + extractCallChain() shared in utils.ts
- receiverCallChain on ExtractedCall for worker path parity
- MAX_CHAIN_DEPTH=3 enforced in both extraction and resolution
Phase 5.2: Pattern matching binding extractors
- PatternBindingExtractor type added to LanguageTypeConfig
- declarationTypeNodes map tracks original type AST nodes for generic unwrapping
- Rust: if let Some(x)/Ok(x) unwrapping with extractGenericTypeArgs
- Java: instanceof pattern variables (Java 16+)
- C#: is-pattern disambiguation fixture (already working via extractDeclaration)
Phase 5.5d: Python standalone type annotations (name: str)
- expression_statement with type child now captured in DECLARATION_NODE_TYPES
Phase 5.5e: ReceiverKey collision fix for overloaded methods
- receiverKey preserves @startIndex to prevent same-name method collisions
- lookupReceiverType does prefix scan with ambiguity refusal
Class-as-receiver for static method calls (#289)
- UserService.find_user() now resolves via ctx.resolve() tiered lookup
- Respects import scoping — no false positives from unrelated packages
Code review fixes:
- Extracted CALL_EXPRESSION_TYPES + extractCallChain to utils.ts (eliminated duplication)
- Converted resolveChainedReceiver from recursion to loop (no exposed depth param)
- Added depth cap to extractReturnTypeName (defense against nested wrapper types)
- Replaced lookupFuzzy with ctx.resolve for class-as-receiver (architecturally consistent)
Closes#289
Test coverage: 6 new fixtures, 12+ new unit tests, 7 new integration test suites
* fix: Ruby chain calls, Rust Err(x) unwrap, Enum class-as-receiver (#315)
Address three per-language gaps identified in Phase 5 code review:
- Ruby: add `method`/`receiver` field fallbacks to extractCallChain
(tree-sitter-ruby uses different field names than other grammars)
- Rust: handle `Err(e)` pattern binding via typeArgs[1] from Result<T,E>
- Enum: include Enum type in class-as-receiver filter (both paths)
Integration tests added for all three fixes.
* fix: chain base type resolution parity between serial and worker paths (#315)
- Worker path: add typeEnv.lookup for chain base receiver after extraction
(typed parameters like `fn process(svc: &UserService)` were silently lost)
- Serial path: add ctx.resolve class-as-receiver fallback for chain base
(class-name chains like `UserService.find_user().save()` failed)
- Fix misleading comment in parse-worker.ts that described unimplemented logic
- Integration tests: typed-parameter chain, static class-name chain
* fix: Kotlin chain call extraction, createClassNameLookup Enum/Struct (#315)
- Kotlin: extractCallChain now handles navigation_expression → navigation_suffix
AST structure (Kotlin's call_expression has no 'function' field)
- createClassNameLookup: include Enum and Struct alongside Class for consistent
constructor recognition in extractInitializer
- Integration test: kotlin-chain-call fixture verifying svc.getUser().save()
* feat: Phase 4 type resolution — nullable unwrapping, for-loop typing, assignment chains, Kotlin return types
Phase 4.1: Nullable/optional chain unwrapping
- Add stripNullable utility in shared.ts for stripping nullable wrappers
- Apply in lookupInEnv to unwrap User | null → User, User? → User before receiver lookup
- Handles TS union, Kotlin/C#/Swift nullable suffix, Python Union[T, None], Rust Option<T>
- Enables receiver-type disambiguation through ?. optional chaining
Phase 4.2: For-loop element typing (Tier 0 — Java/C#/Kotlin)
- Add ForLoopExtractor type and forLoopNodeTypes to LanguageTypeConfig
- Java enhanced_for_statement, C# foreach_statement, Kotlin for_statement extractors
- Only explicit element types in AST (Tier 0); inference-based languages deferred
Phase 4.3: Assignment chain propagation (single-pass, depth-1)
- Add PendingAssignmentExtractor to LanguageTypeConfig with per-language implementations
- Handles TS/JS variable_declarator, Rust let_declaration, Python assignment,
Go short_var_declaration, C# equals_value_clause, Java/Kotlin variable_declarator
- Single post-walk propagation pass (no fixpoint iteration per Sorbet/Pyright design)
- Resolves const b = a; b.save() when a has known type from Tier 0/1/1b
Phase 4.5: Kotlin return type extraction (bug fix)
- Fix extractMethodSignature to handle Kotlin user_type after function_value_parameters
- Remove lenient test assertions, add strict disambiguation proof
Integration tests across 10+ languages with competing same-name methods
and negative assertions proving disambiguation.
* fix: per-language assignment chain gaps from code review
- Kotlin: new extractKotlinPendingAssignment for property_declaration →
variable_declaration AST (Java's variable_declarator doesn't exist in Kotlin)
- Go: handle var_spec (var b = u) alongside short_var_declaration (:=)
- PHP: add extractPendingAssignment for $alias = $user with $ prefix preserved
Integration tests added for all three languages with competing
same-name methods and negative disambiguation assertions.
* fix: code review fixes — DRY nullable keywords, avoid array allocations, clarify depth comment
Addresses findings from 6-agent code review on PR #310:
- Move stripNullable JSDoc to correct position (was orphaned above NULLABLE_KEYWORDS)
- DRY: reuse NULLABLE_KEYWORDS set in pipe-split filter instead of inline strings
- Replace node.children.find() with findChildByType/manual loops in jvm.ts,
go.ts, csharp.ts to avoid unnecessary array allocations per tree-sitter call
- Clarify "depth-1" comment in type-env.ts: single-pass resolves multi-hop
chains when forward-declared; reverse-order is depth-1 only
- Annotate extractGenericTypeArgs as Phase 5 infrastructure (zero production callers)
- Re-export PendingAssignmentExtractor from index.ts for API consistency
- Add explicit return undefined in Go extractPendingAssignment
- Remove redundant child.text === '=' check in Kotlin extractor
Test coverage:
- 20 new unit tests: stripNullable edge cases, per-language assignment chains,
reverse-order depth limitation, nullable lookup resolution
- 15 new integration tests: multi-hop chains (a→b→c), nullable+chain combined
(User|null + alias), Python User|None through stripNullable path
- 3 new fixtures: ts-multi-hop-chain, ts-nullable-chain, python-nullable-chain
* fix: third-pass review — walrus chain, scanner allocations, Kotlin variable_declaration, C# type guard
Addresses 4 new findings from third-pass CI review:
1. Python walrus operator (:=) now handled by extractPendingAssignment —
named_expression nodes propagate alias chains alongside regular assignment
2. Scanner .namedChildren.find()/.some() in jvm.ts replaced with
findChildByType() — consistent with 98daed4 code review fixes
3. Kotlin extractPendingAssignment extended to handle variable_declaration
nodes in addition to property_declaration (function-local val/var)
4. C# extractPendingAssignment early-returns for is_pattern_expression and
field_declaration nodes (never contain variable_declarator children)
Integration tests:
- Python: walrus chain (alias := u) with disambiguation (5 tests, 1 fixture)
- Kotlin: assignment chain with typed declarations (5 tests, 1 fixture)
- C#: assignment chain + is-pattern coexistence (6 tests, 1 fixture)
- Unit: Python walrus propagation (1 test)
* feat: nullable wrapper unwrapping + C++ assignment chains
Gaps 1, 2, 4 from code review — architectural changes to type resolution:
1. extractSimpleTypeName now unwraps nullable wrapper generics:
- Optional<User> → "User" (Java), Option<User> → "User" (Rust),
Maybe<User> → "User" (Kotlin Arrow/Haskell-style)
- Containers (List, Map) and async wrappers (Promise, Future) are NOT
unwrapped — methods are called on the container, not the inner type
- Uses existing extractGenericTypeArgs (now production-active, was dead code)
- NULLABLE_WRAPPER_TYPES set: Optional, Option, Maybe
2. C++ extractPendingAssignment added for auto alias chains:
- auto alias = user; alias.save() now propagates User type
- Handles pointer/reference declarators, auto/decltype(auto)
3. Updated existing Rust test: Option<User> parameter now correctly
stores "User" instead of "Option" in TypeEnv
Integration tests with fixtures for Java Optional, Rust Option, C++ auto
chain. Full pipeline resolution marked .todo — requires call-processor
enhancement (TypeEnv stores correct types but call-processor needs
additional work to produce CALLS edges for these patterns).
Unit tests: 196 passed (7 new). Integration: all 9 languages green.
* fix: resolve .todo tests — stale dist/ was the root cause
The Rust Option<User> and C++ auto assignment chain integration tests
were marked .todo because the pipeline didn't produce CALLS edges.
Root cause: dist/ was compiled from pre-Phase 4 source and lacked:
- NULLABLE_WRAPPER_TYPES unwrapping in extractSimpleTypeName
- C++ extractPendingAssignment
After npm run build, all tests pass as real assertions:
- Rust: alias.save() resolves to User#save via Option<User> unwrap + chain
- C++: alias.save() and rAlias.save() resolve via auto assignment chain
with correct disambiguation (User vs Repo)
Only remaining .todo: Rust user.unwrap().save() (Phase 5 — chained
return type inference, not a TypeEnv issue).
* feat(ingestion): respect .gitignore and .gitnexusignore during file discovery
Add support for excluding files from indexing based on .gitignore and
.gitnexusignore patterns. Previously, GitNexus used only a hardcoded
ignore list, causing significant index pollution in repositories with
git-ignored directories containing code (e.g., Docker-mounted volumes).
Changes:
- Add `ignore` package for gitignore-spec pattern matching
- Add `loadIgnoreRules()` to parse .gitignore + .gitnexusignore
- Add `createIgnoreFilter()` returning glob-compatible IgnoreLike object
- Integrate filter into glob's `ignore` option for directory-level pruning
- Remove post-glob `.filter()` call (now handled during traversal)
The hardcoded DEFAULT_IGNORE_LIST remains as fallback for non-git repos.
Closes#228
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(ingestion): address review feedback on ignore filtering
- Distinguish ENOENT vs EACCES in loadIgnoreRules (warn on permission errors)
- Add GITNEXUS_NO_GITIGNORE env var to bypass .gitignore parsing
- Fix bare-name pattern matching in childrenIgnored (check both with/without trailing slash)
- Rename isIgnoredDirectory to isHardcodedIgnoredDirectory for clarity
- Add clarifying comments for design decisions (D2 negation, D3 dot:false redundancy)
- Add tests for bare-name patterns, file-glob patterns, EACCES handling, env var
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(ingestion): address second round of review feedback
- G1: Document GITNEXUS_NO_GITIGNORE in `analyze --help` and log when active
- G2: Add comment clarifying path-scurry POSIX normalization contract
- G3: Add IgnoreOptions interface — env var now falls back, callers can
pass `noGitignore` explicitly for testability and future CLI flag
- G4: Add integration test verifying walkRepositoryPaths respects the env var
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* feat(ingestion): gracefully skip files with unavailable tree-sitter grammars
Port unsupported language resilience from PR #301 by @jecanore.
- Make Kotlin import optional (like Swift) in parser-loader and parse-worker
- Add worker-local isLanguageAvailable() with filePath param for tsx distinction
- Track and log skipped files per language in both sequential and worker paths
- Add skippedLanguages to ParseWorkerResult for worker→main aggregation
- Add isLanguageAvailable unit tests
Refs: #301, #155, #228
Co-Authored-By: jecanore <juan@housingbase.io>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* test(e2e): add ignore + language-skip end-to-end test with fixture repo
Add a fixture repo (test/fixtures/ignore-and-skip-repo/) with .gitignore,
.gitnexusignore, TypeScript source files, and a Swift file to exercise
all three features end-to-end:
- File discovery: verifies .gitignore excludes data/ and *.log,
.gitnexusignore excludes vendor/, source files are discovered
- Parsing: verifies TypeScript files produce Function nodes and DEFINES
relationships, Swift files are skipped gracefully when grammar is
unavailable
Add the test to the standalone group in ci-integration.yml and coverage job.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(ci): move ignore-and-skip-e2e test to e2e group per review feedback
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(test): use temp directory instead of fixture for e2e ignore test
The fixture's .gitignore prevented data/seed.json and debug.log from
being committed — these files would be missing after checkout in CI.
Switch to creating the entire test structure in a temp directory via
beforeAll (matching filesystem-walker.test.ts pattern). This ensures
all files exist regardless of git ignore rules.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(test): correct graph API usage in e2e ignore test
Use graph.nodes property getter instead of graph.getNodes(), and check
Function node filePath instead of non-existent File nodes (File nodes
are created by processStructure, not processParsing).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* ci: add workflows permission to ci-integration.yml
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* ci: change workflows permission to write per review
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* ci: move workflows permission from ci-integration.yml to ci.yml caller
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(ci): fix Claude workflows for fork PRs, remove misplaced workflows perm
Three issues prevented Claude from running on fork PRs:
1. claude-code-review.yml lacked workflows:write — push failed when
fork PRs modify .github/workflows/ files
2. claude.yml had no fork PR support — checked out main and couldn't
fetch the fork's branch from origin
3. Cleanup step unconditionally deleted branches even when push failed,
breaking the concurrent claude.yml workflow
Also removes workflows:write from ci.yml's integration job — CI tests
don't need that permission. The permission belongs on the claude
workflows that push fork branches.
Changes:
- Add workflows:write to both claude workflow permissions blocks
- Add fork PR detection + branch push/cleanup to claude.yml
- Add step id to push-fork; cleanup only runs if push succeeded
- Pass branch names via env vars to prevent shell injection (security)
- Add concurrency groups to prevent race conditions between workflows
- Remove misplaced workflows:write from ci.yml integration job
* fix(ci): use GitHub API for fork branch refs instead of git push
GITHUB_TOKEN cannot have 'workflows' permission — it's only valid for
PATs and GitHub Apps. This means git push fails whenever a fork PR
modifies .github/workflows/ files.
Replace git push with the GitHub REST API (POST/PATCH /git/refs) to
create temporary branch refs. The API creates a pointer to the
already-existing PR head commit without triggering the workflow file
push protection. Similarly, cleanup uses DELETE /git/refs instead of
git push --delete.
Also removes the invalid 'workflows: write' from permissions blocks.
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: jecanore <juan@housingbase.io>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
CI was failing because package-lock.json still referenced onnxruntime-node@1.21.0
after package.json was updated to require ^1.24.0.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
onnxruntime-node versions before 1.24.0 are CPU-only and do not ship
libonnxruntime_providers_cuda.so. When system CUDA libraries (e.g.
libcublasLt.so.12) are detected, isCudaAvailable() returns true and
the embedder requests the CUDA execution provider. Since the provider
binary doesn't exist in the package, ONNX Runtime crashes at the native
level (provider_bridge_ort.cc), which is uncatchable by the JS try/catch
fallback — killing the entire process.
This commit:
- Adds onnxruntime-node ^1.24.0 as an explicit dependency (first version
to ship CUDA provider binaries for Linux x64)
- Adds hasOrtCudaProvider() check that verifies the CUDA provider .so
exists in the onnxruntime-node package before attempting CUDA, so
the embedder gracefully falls back to CPU on older ORT versions
Fixes the crash at 92% "Loading embedding model..." on Linux systems
with CUDA toolkit installed. Also related to #165.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: Phase 3 — return type inference, generic args extraction, Ruby YARD type extractor
Three architectural improvements to the type resolution system:
1. Return type inference — wire extractMethodSignature returnType through
SymbolDefinition into call-processor. When var = callee() and callee
has a known return type, bind var to that type. Handles Promise<T>
unwrapping, nullable stripping, pointer/reference removal.
2. Generic type argument extraction — new extractGenericTypeArgs() utility
that extracts type parameters from List<User> → ['User']. Handles
TS/Java/Kotlin/C#/Rust generic syntax. Building block for for-loop
variable typing.
3. Ruby dedicated type extractor — replaces the stub with YARD annotation
parsing (@param name [Type]), handling qualified types, nullable types,
and singleton methods. Ruby now has real type resolution.
Unit tests: 127 → 192+ (type-env) + 65 (symbol-table, call-processor) + 18 (generics)
Integration tests: 8+ new test cases with fixtures across TS/Python/Go/Java/Ruby
* fix: Phase 3 gaps — WRAPPER_GENERICS correctness, Ruby :: qualifier, namespaced constructors
- Remove collection types (List, Array, Vec, Set) from WRAPPER_GENERICS to prevent
false CALLS edges (e.g. List<User> no longer unwraps to User)
- Add :: qualifier handling in extractReturnTypeName for Ruby/C++/Rust namespaced types
- Add Ruby `constant` and `scope_resolution` node types to shared extractors
- Extract shared extractRubyConstructorAssignment helper (dedup type-env.ts + ruby.ts)
- Add integration tests for return type inference: Python, TypeScript, Go, Java, Ruby
- Add Ruby namespaced constructor fixture (Models::UserService.new)
- Add unit tests for collection reclassification and :: qualifiers
* feat: Phase 4 — CONSTRUCTOR_BINDING_SCANNERS for all languages + return type inference tests
Add CONSTRUCTOR_BINDING_SCANNERS for 6 missing languages, completing
return type inference coverage across all 11 supported languages:
- TypeScript/JS: variable_declarator with call_expression, unwraps await
- Go: short_var_declaration single-assignment (skips multi-return, new/make)
- Java: local_variable_declaration with `var` type + method_invocation
- C#: variable_declaration with implicit_type (var) + invocation_expression
- Rust: let_declaration without type annotation, handles mut_pattern
- PHP: assignment_expression with function_call_expression
Also adds property_identifier to extractSimpleTypeName for qualified
member calls (repo.getUser → getUser), fixing namespaced constructor
inference that was previously a known limitation.
Integration tests added for all 11 languages with correct label
assertions (Function vs Method per language's tree-sitter queries).
* refactor: merge CONSTRUCTOR_BINDING_SCANNERS into per-language LanguageTypeConfig
Eliminates the parallel dispatch map in type-env.ts by moving all 11
constructor binding scanners into their respective type-extractors/*.ts
files as `scanConstructorBinding` on LanguageTypeConfig.
- Add ConstructorBindingScanner type to types.ts
- Add shared helpers: hasTypeAnnotation, unwrapAwait, extractCalleeName
- Move scanners to typescript.ts, jvm.ts, python.ts, php.ts, go.ts,
rust.ts, swift.ts, c-cpp.ts, csharp.ts, ruby.ts
- Fix `any` types in C# scanner → SyntaxNode | null
- Delete ~300 lines from type-env.ts (CONSTRUCTOR_BINDING_SCANNERS map)
- Update buildTypeEnv to use config.scanConstructorBinding
All 143 type-env unit tests and all 10 language integration suites pass.
* fix: remove unused import, fix any type in Java scanner, update stale comment
- Remove unused extractCalleeName import from jvm.ts
- Fix (c: any) → (c: SyntaxNode) in Java scanner
- Update stale CONSTRUCTOR_BINDING_SCANNERS reference in ruby.ts comment
* fix: C# and PHP return type inference — scanner fixes, method signature extraction, and cross-file resolution
Addresses code review findings on PR #284:
C# scanner (csharp.ts):
- Fix type node lookup: iterate children instead of childForFieldName('type')
which returns undefined in tree-sitter-c-sharp
- Fix initializer lookup: handle direct invocation_expression children
(no equals_value_clause wrapper in tree-sitter-c-sharp)
C# return type extraction (utils.ts):
- Add 'returns' field check to extractMethodSignature — tree-sitter-c-sharp
uses 'returns', not 'type', for method return types
C# cross-file resolution (call-processor.ts + fixture):
- Add constructor binding verification to sequential processCalls path
(was only in the worker processCallsFromExtracted path)
- Add ReturnType.csproj to csharp-return-type fixture
- Update fixture namespaces to use ReturnType.Models/ReturnType.Services
prefix (matches real C# project conventions)
PHP scanner (php.ts):
- Extend scanConstructorBinding to handle member_call_expression
($this->getUser() patterns), not just function_call_expression
Shared (shared.ts):
- Add member_access_expression to extractSimpleTypeName qualified-names
block (C# method calls like svc.GetUser())
Tests:
- Add Repo.cs/Repo.php disambiguation fixtures (two Save methods)
- Strengthen C# and PHP return type tests with hard disambiguation assertions
- Add C# scanner unit tests and return type extraction test
* feat: per-language ReturnTypeExtractor + doc-comment @param parsing for PHP, JS, Ruby
Add ReturnTypeExtractor to LanguageTypeConfig interface with implementations
for Ruby (YARD @return), PHP (PHPDoc @return), and JS/TS (JSDoc @returns).
The fallback is wired in both parsing-processor and parse-worker paths,
activating only when extractMethodSignature finds no AST-based return type.
Also add doc-comment @param type extraction for PHP and JS/TS, following
Ruby's existing collectYardParams pattern. This enables parameter.method()
resolution in loosely-typed codebases using PHPDoc @param or JSDoc @param.
Additional fixes from PR #284 code review:
- Go: add selector_expression + field_identifier to extractSimpleTypeName
(enables package-qualified factory calls like models.NewUser())
- Ruby: broaden scanConstructorBinding to capture plain call assignments
(user = get_user()) in addition to Class.new patterns
- Ruby: harden return-type fixture with disambiguation (two save methods)
Test coverage: +14 new integration tests across Go, Ruby, PHP, JS/TS
* fix: JSDoc async return type, PHP attribute walkers, and $this receiver disambiguation
Three fixes from fourth-pass code review on PR #284:
1. JSDoc `@returns {Promise<User>}` no longer stripped to `Promise` — extractReturnType
now uses sanitizeReturnType (preserves generics) instead of normalizeJsDocType
(which stripped them before extractReturnTypeName could unwrap WRAPPER_GENERICS).
2. PHP 8+ `#[Attribute]` and JS `@decorator` nodes no longer break doc-comment walkers.
Both extractReturnType and collect*Params functions now skip attribute_list/decorator
nodes instead of breaking on them as named siblings.
3. PHP `$this->method()` now provides receiverClassName for disambiguation.
When two classes define the same method, the enclosing class narrows candidates
via ownerId matching in call-processor, preventing false no-binding results.
* fix: sanitizeReturnType dot corruption, JS test assertions, Ruby constant receiver
- Remove redundant dot-path stripping from sanitizeReturnType that corrupted
qualified names inside generics (e.g. Promise<models.User> → User>)
- Split JS async fixture into separate files and add negative assertions
to properly verify disambiguation (mirroring PHP test pattern)
- Accept 'constant' node type in Ruby scanConstructorBinding for factory
call assignments (SERVICE = build_service())
- Add 'constant' to SIMPLE_RECEIVER_TYPES so extractReceiverName handles
Ruby constant receivers (SERVICE.process)
* fix: nested generic arg splitting, JS/Ruby test false positives
- Replace naive comma split in extractReturnTypeName with bracket-balanced
extractFirstGenericArg so nested types like Future<Result<User, Error>>
unwrap correctly instead of producing malformed "Result<User"
- Add CompletableFuture to WRAPPER_GENERICS for Java async unwrapping
- Split js-jsdoc-return-type fixture models.js into user.js/repo.js and
add negative assertions to prove disambiguation (not just file match)
- Split ruby-constant-factory-call fixture into separate service files
and add negative assertions against AdminService resolution
* fix: review findings — receiverClassName parity, Rust wrappers, Go multi-return, Kotlin/Swift qualified calls
P1: Sequential path now includes receiverClassName narrowing for PHP
$this->method() disambiguation (was missing vs worker path).
P2: Added Rc/Arc/Weak/MutexGuard/Cow + 6 more Rust Deref types to
WRAPPER_GENERICS (Box excluded — Java Swing collision). Extended
Kotlin/Swift scanners to handle navigation_expression callees.
Added Go multi-return support (user, err := f()) with blank/_/err/ok
guard + AST-level first-return extraction in extractMethodSignature.
P3: Extracted shared verifyConstructorBindings() eliminating 60 lines
of duplication between sequential and worker paths. Added return-type
inference integration tests for C++, Rust, Swift with competing
methods and negative disambiguation assertions.
* fix: Swift navigation_suffix unwrapping, Rust lifetime skipping, Kotlin disambiguation tests
- Swift scanConstructorBinding: handle tree-sitter wrapping qualified
identifiers in navigation_suffix nodes
- Add extractFirstTypeArg to skip Rust lifetime parameters ('a, '_)
when unwrapping wrapper generics like Ref<'_, User>
- Kotlin tests: add Repo class fixture with competing save() methods
to prove disambiguation; assert no spurious edges on known gap
- Remove tree-sitter-kotlin from optionalDependencies (now regular dep)
* fix: C# null-conditional calls, Ruby YARD bracket-balanced split, PHPDoc alternate order, escapeValue hardening
- Add C# null-conditional call support (user?.Save()): tree-sitter query for
conditional_access_expression, member_binding_expression in MEMBER_ACCESS_NODE_TYPES,
receiver extraction via conditional_access_expression parent walk
- Fix Ruby YARD type parsing for nested generics (Hash<Symbol, User>): replace
naive split(',') with bracket-balanced splitter respecting <> depth
- Add alternate YARD format (@param [Type] name) alongside standard (@param name [Type])
- Add alternate PHPDoc format (@param $name Type) alongside standard (@param Type $name)
- Harden escapeValue in kuzu-adapter.ts: escape \n and \r to prevent Cypher injection
- Integration tests: C# null-conditional fixture (5 tests), Ruby YARD generics fixture (6 tests)
- Unit tests: PHPDoc alternate order (2 tests), C# null-conditional call-form (updated)
* test: add Python static/classmethod integration tests (issue #289)
Verifies that classes using only @staticmethod/@classmethod have HAS_METHOD
edges connecting them to their child methods. This was the root cause of
issue #289 where context() and impact() returned empty for such classes.
Tests cover: HAS_METHOD edge emission, unique static method resolution
(create_user, delete_user), and ambiguous same-named method handling
(find_user on both UserService and AdminService — safely refused).
* fix: lbug batch escapeValue newline hardening, Rust ::default() scanner exclusion
- Apply \n/\r escaping to batch upsert escapeValue in lbug-adapter.ts:429
(missed instance of the CREATE-path fix from ec4dca4)
- Exclude Rust ::default() from scanConstructorBinding to match
extractInitializer behavior — avoids wasted cross-file lookups on
the broadly-implemented Default trait
- Unit tests: 2 new scanner exclusion tests (::default and ::new)
- Integration tests: 6 new Rust ::default() constructor resolution tests
with disambiguation fixture (User::default vs Repo::default)
* fix: C#/Rust async await unwrap, PHP backslash namespace, fallback escaping
- C# scanConstructorBinding: unwrap await_expression to find invocation_expression
(var user = await svc.GetUserAsync() now produces constructor binding)
- Rust scanConstructorBinding: unwrap .await postfix via shared unwrapAwait helper
(let user = get_user().await now produces constructor binding)
- extractReturnTypeName: handle PHP backslash namespace separator (\App\Models\User → User)
- fallbackRelationshipInserts: match batch escapeValue hardening with \n/\r escaping
Tests: 2 unit (type-env), 3 unit (call-processor), 7 integration (csharp+rust), 7 fixtures
* fix: C#/Rust async-binding test false positives — add competing types and negative assertions
C# fixture: add Order.cs with Order.Save(), change OrderService to return
Task<Order> via GetOrderAsync, add negative assertion proving user.Save()
does not resolve to Order#Save.
Rust fixture: split models.rs into user.rs/repo.rs, make process_user and
process_repo async fn, add bidirectional negative assertions proving no
cross-contamination between User#save and Repo#save.
* fix: C# async-binding broken assertion, bare wrapper type leak, JSDoc optional params
- Split Program.cs Main into ProcessUser/ProcessOrder so negative
assertions use strict toBeUndefined() (matching Rust pattern)
- Guard bare wrapper types (Task, Promise, Option…) in
extractReturnTypeName — return undefined instead of the wrapper name
- Update JSDOC_PARAM_RE to capture @param {Type} [optionalName] syntax
* fix: update symbol and relationship counts in documentation
* refactor: migrate from KuzuDB to LadybugDB v0.15
KuzuDB was archived (Apple acquisition, Oct 2025). LadybugDB is the
community fork with full API compatibility.
- Package swap: kuzu → @ladybugdb/core, kuzu-wasm → @ladybugdb/wasm-core
- Rename all internal paths: kuzu → lbug (adapters, schema, storage)
- Storage path: .gitnexus/kuzu → .gitnexus/lbug (with auto-cleanup)
- Add explicit VECTOR extension loading (required in v0.15)
- Update CI workflow, documentation, and all tests
- 1151 unit + 27 integration tests passing
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: address code review findings (P1-P3)
P1: Fix WASM adapter to use getAll() API, wire cleanupOldKuzuFiles
into analyze command, add symlink path traversal protection.
P2: Cache VECTOR extension load state, batch augmentation engine
queries (20→4), fix web getCopyQuery for multi-language tables,
fix stale KuzuDB references, correct brainstorm package names.
P3: Complete lbug-wasm.d.ts type declarations, batch semantic
search per-label, update stale BM25 comment.
* chore: remove outdated KuzuDB migration brainstorming document
* fix: load FTS extension in MCP pool adapter on init
The read-only pool adapter never loaded the FTS extension, so all
QUERY_FTS_INDEX calls failed silently. This broke search-pool and
augmentation integration tests, and caused empty results in the
web UI server mode.
* feat: implement shared Database caching and connection reference counting
* feat: enhance KuzuDB migration handling and status reporting
* fix: mock cleanupOldKuzuFiles in local backend callTool tests
* fix: update mock for cleanupOldKuzuFiles and adjust imports in callTool tests
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* feat(type-env): constructor-call type inference for TypeEnv (Phase 1)
Add extractInitializer as a Tier 1 fallback in buildTypeEnv: when a
declaration node has no explicit type annotation, infer the type from
constructor-call patterns (new X(), X::new(), X::default(), $x = new X()).
Languages covered: TypeScript/JS, Java (var), Rust, PHP, C++ (auto).
Python/Kotlin/Swift deferred — need symbol-table access to distinguish
class constructors from function calls.
Adds 20 new unit tests covering constructor inference, annotation
precedence, and known limitations across all supported languages.
* fix(type-env): class-aware constructor resolution, multi-declarator fix
- Add collectClassNames pre-scan: walks AST to build Set<string> of
class/struct names defined in the file
- C++ extractInitializer uses classNames.has() to verify identifier is
a known class before inferring (auto x = User() resolves, auto x =
getUser() does not — no false positives)
- Add InitializerExtractor type that receives classNames parameter
- Fix env.size gating: always call extractInitializer when available,
so mixed declarators like const a: A = x, b = new B() resolve both
- Add env.has() guard in Java extractInitializer to skip already-bound vars
- Document Rust new/default whitelist rationale
- Pin all test assertions, add mixed multi-declarator test case
* fix(type-env): resolve Self/self/static/parent to actual type names
- Rust: Self::new()/Self::default() resolves to enclosing impl type
- PHP: new self()/static() resolves to enclosing class, parent() to superclass
- Rust: Tier 0 annotation guard prevents overwrite by constructor inference
- Rust: mut_pattern handling in extractVarName for let mut bindings
- TS: fix misleading comment in extractInitializer
- 58 tests passing (3 new Self/self resolution tests)
* perf(type-env): single-pass AST walk with closure-scoped state
Refactors buildTypeEnv to use closures instead of passing mutable state
as parameters. classNames, env, and config are captured by the inner
walk and extractTypeBinding functions — no parameter mutation.
- Eliminates separate collectClassNames pre-scan (O(2n) → O(n))
- config looked up once per file instead of per-node
- 29 fewer lines
* feat(type-env): constructor-inferred type resolution for all languages
Add cross-file constructor type inference to the ingestion pipeline,
enabling receiver-type disambiguation for member calls like
`user.save()` when the variable is assigned from a constructor without
explicit type annotations.
Pipeline changes:
- Add extractInitializer to Python and Swift type extractors
- Add CONSTRUCTOR_BINDING_SCANNERS for Python, Swift, C/C++ in type-env
- Wire constructorBindings through parse-worker → parsing-processor →
pipeline → processCallsFromExtracted
- Rewrite resolveCallTarget receiver-type filtering (step D) to use
tiered import resolution (same-file → import-scoped → global) before
falling back to fuzzy ownerId matching
- Use collectTieredCandidates for constructor binding verification
instead of raw lookupFuzzy
Bug fixes:
- Fix C++ inline method query: @definition.method was captured on
field_declaration_list instead of function_definition, causing wrong
parameterCount for all inline class methods
- Fix parse-worker accumulated/flush results missing constructorBindings
CI changes:
- Add swift.test.ts to ci-integration pipeline group and coverage job
- Update ci-report to fetch base branch (main) coverage for delta
reporting instead of showing config thresholds
- Add per-suite timing breakdown table (unit/integration/total)
- Add expandable skipped test details section
Tests: 288 passed, 4 skipped (swift — macOS only) across 10 languages
- 36 new constructor-inferred integration tests (4 per language)
- 10 fixture directories with cross-file constructor patterns
- TypeScript, JavaScript, Java, Kotlin, Python, PHP, Rust, Go, C++, Swift
* fix(type-extractors): add type assertion for LanguageTypeConfig
* feat(ruby): constructor-inferred type resolution and self-receiver mapping
Add Ruby User.new constructor binding scanner to type-env, enabling
receiver-type disambiguation for member calls like user.save vs repo.save.
Add self/this → enclosing class resolution in lookupTypeEnv so self.method()
calls resolve to the correct class even when the method name is ambiguous.
* docs: update README with constructor inference and self/this resolution details
* refactor(ingestion): unified ResolutionContext replaces fragmented map passing
Introduce createResolutionContext() as the single resolution API for all
processors. Eliminates duplicated tier-selection logic, fixes heritage
namedImportMap bug, and adds per-file resolution caching.
- NEW resolution-context.ts: closure-factory with resolve(), per-file cache,
TIER_CONFIDENCE constant, and shared ResolutionTier type
- DELETE symbol-resolver.ts: zero production importers, logic now in
resolution-context.ts
- call-processor: all functions take ctx instead of 6 separate maps,
collectTieredCandidates removed (ctx.resolve replaces it),
D4 redundant re-resolve eliminated
- heritage-processor: takes ctx, resolveHeritageId helper extracts
repeated 14-line fallback pattern, namedImportMap now included
- import-processor: takes ctx, dead createImportMap/createPackageMap/
createNamedImportMap factories removed
- pipeline: creates single ctx, wires onProgress to all processors,
logs cache hit rate in dev mode
- Tier renamed: unique-global → global (honest about returning all candidates)
- Tests migrated: 1178 unit + 84 integration passing
* feat(type-env): self/this/super resolution, TypeEnvironment API, and review fixes
Add cross-language receiver keyword resolution:
- self/this/$this → enclosing class name via AST walk
- super/base/parent → parent class name via heritage AST extraction
(8 grammar variants: TS/JS, Java, Python, Ruby, C#, PHP, Kotlin, C++, Swift)
- D-phase widening in resolveCallTarget for super→parent method dispatch
Introduce TypeEnvironment API replacing loose TypeEnvResult + lookupTypeEnv:
- buildTypeEnv() returns TypeEnvironment with .lookup() method
- Single-pass AST walk merges constructor binding scan (was separate traversal)
- ClassNameLookup type replaces over-broad ReadonlySet<string> facade
- Memoized class name lookups to avoid redundant SymbolTable scans
Code review fixes (6 agents, 11 findings):
- Replace ctx.resolve(name, '') hack with direct symbols.lookupFuzzy()
- Extract scope key helpers (extractFuncNameFromScope, receiverKey)
- Simplify D-phase from 5 steps to 4 with deduped typeNodeIds
- Remove C from CONSTRUCTOR_BINDING_SCANNERS (YAGNI — C has no constructors)
- Cache Map reuse in ResolutionContext to reduce GC pressure
- Remove unused TieredCandidates import
Integration tests for self/this, parent, and super resolution across all
12 supported languages with per-language fixture directories.
* fix(type-env): generic parent resolution, TS cast inference, C++ brace-init
Fix generic parent class breaking super resolution:
- extractParentClassFromNode now uses extractSimpleTypeName to strip
generic params (Base<T> → Base) and qualified names (models.Model → Model)
- Affects TS, Java, Python, C# heritage extraction
Fix TypeScript new X() as T / new X()! missed inference:
- Unwrap as_expression and non_null_expression before checking for
new_expression in extractInitializer
Fix C++ brace-init User{} missed inference:
- Handle compound_literal_expression with type_identifier child
in extractInitializer
Clean up deprecated lookupTypeEnv:
- Remove standalone lookupTypeEnv export, migrate all callers to
TypeEnvironment.lookup() method
- Update all 80+ test assertions to use the new API
Integration test fixtures added:
- typescript-cast-constructor-inference (new X() as T, new X()!)
- typescript/java/csharp/kotlin-generic-parent-resolution
- cpp-brace-init-inference (auto x = User{})
* fix(type-extractors): Go &User{}, TS double-cast, Swift .init inference
Fix Go pointer-to-struct literal not inferred:
- Unwrap unary_expression (address-of &) before composite_literal check
- user := &User{} now correctly infers type User
Fix TypeScript double-cast only unwrapping one level:
- Change if to while loop for nested as_expression/non_null_expression
- new User() as unknown as Admin now correctly infers type User
Fix Swift User.init(name:) explicit init call missed:
- Handle navigation_expression callee with .init suffix in extractInitializer
Integration test fixtures:
- go-pointer-constructor-inference (&User{}, &Repo{})
- typescript-double-cast-inference (as unknown as T)
* feat: Rust struct literal, Python qualified ctor, Go new(), Swift .init scanner
- Rust: handle struct_expression in extractInitializer (User { name: "alice" })
- Python: support attribute nodes in extractInitializer (models.User("alice"))
and the cross-file scanner — extractSimpleTypeName handles qualified names
- Go: handle new(User) built-in in extractGoShortVarDeclaration
- Swift: extend CONSTRUCTOR_BINDING_SCANNERS to handle navigation_expression
callee for User.init(name:) cross-file resolution
Unit tests: 87 → 96 (Rust struct literal, Go new(), Python qualified ctor,
Python scanner qualified, plus edge cases)
Integration tests: 4 new describe blocks with fixtures
* fix: Rust Self{} resolution, C++ scoped brace-init, PHP promotion params, Ruby constants
- Rust: resolve Self {} struct literal to enclosing impl type (was stored as "Self")
- C++: replace type_identifier guard with extractSimpleTypeName for compound_literal_expression,
enabling ns::User{} scoped brace-init (closes previously deferred gap)
- PHP: add property_promotion_parameter to TYPED_PARAMETER_TYPES for PHP 8.0+
constructor property promotion (__construct(private Foo $x))
- Ruby: extend extractRubyConstructorBinding to accept constant left-hand side
(REPO = Repo.new)
Unit tests: 96 → 101 (+5: Rust Self{} ×2, C++ ns::User{} ×1, PHP promotion ×1,
Ruby constant ×1)
Integration tests: 4 new describe blocks with fixtures
* feat: Phase 1 type resolution gaps — walrus, PHP properties, nullable, Go make/assert
Phase 1 quick wins from the type resolution gap analysis:
1. Python walrus operator := (named_expression) — extractInitializer + scanner
2. PHP 7.4+ typed class properties — property_declaration in extractDeclaration
3. Nullable union unwrapping — User | null → User in extractSimpleTypeName
4. Go make() builtin — slice/map element type extraction
5. Go type assertions — iface.(User) type extraction
Also: PHP primitive_type handling in extractSimpleTypeName (string, int, etc.)
Unit tests: 101 → 114 (+13)
Integration tests: 8 new describe blocks with fixtures
* feat: Phase 2 type resolution gaps — C++ range-for, Rust if-let, C# pattern matching, Python class annotations
Phase 2 medium-effort improvements:
1. C++ range-for with explicit type — for (User& u : vec) binds u: User
2. Rust if-let/while-let captured_pattern — user @ User { .. } binds user: User
3. C# is-pattern matching — if (obj is User user) binds user: User
4. Python class-level annotations — confirmed already working, added tests
Unit tests: 114 → 127 (+13)
Integration tests: 11 new test cases with fixtures
* fix(ruby): method-level call resolution, HAS_METHOD edges, and dispatch table refactoring
- Replace all `if (language === Ruby)` checks in processors with a
`callRouters` dispatch table in call-routing.ts (renamed from
ruby-call-routing.ts to preserve git history)
- Add Ruby `method` and `singleton_method` to FUNCTION_NODE_TYPES so
findEnclosingFunction produces Method-level CALLS sources
- Add Ruby `class` and `module` to CLASS_CONTAINER_TYPES for HAS_METHOD
edge generation
- Add bare call capture via tree-sitter query `(body_statement (identifier))`
for Ruby methods called without parentheses
- Add Ruby member call detection (`call` node with `receiver` field) to
inferCallForm and extractReceiverName
- Wire resolveRubyImport into resolveLanguageImport
- Add 24 integration tests across 5 suites: heritage/properties, arity
filtering, member calls, ambiguous disambiguation, local shadow
- Add ruby.test.ts to CI integration workflow
* fix(ruby): resolve 6 Ruby resolution gaps from PR review
- Fix singleton_method label mismatch: @definition.function → @definition.method
so CALLS edges from `def self.foo` bodies get correct sourceId
- Add ownerId and HAS_METHOD edges to attr_* Property nodes by calling
findEnclosingClassId in both parse-worker and call-processor property branches
- Distinguish include/extend/prepend heritage: add heritageKind to
RubyHeritageItem, propagate through heritage pipeline as IMPLEMENTS reason
- Document bare call over-capture limitation in tree-sitter query comment
- Add bare `require` (non-relative) import test coverage
- Add prepend/extend test coverage with distinct Loggable/Cacheable modules
31 Ruby integration tests passing, no regressions in other language resolvers.
* fix(ruby): web package parity — heritage reasons, property HAS_METHOD, singleton_method label
- Web call-processor: use item.heritageKind as IMPLEMENTS reason instead of
hardcoded 'trait-impl', add :${kind} suffix to edge ID for uniqueness
- Web call-processor: port findEnclosingClassId, add HAS_METHOD edges for
attr_* Property nodes to match CLI fix
- Web call-processor: singleton_method label 'Function' → 'Method' to match
CLI tree-sitter query fix
- CLI parse-worker: update stale ExtractedHeritage.kind JSDoc to include
'include' | 'extend' | 'prepend'
* fix(web): add HAS_METHOD to RelationshipType union
Web package was missing HAS_METHOD in the RelationshipType union,
causing a type mismatch with the HAS_METHOD edges emitted by the
attr_* property fix in call-processor.ts.
* feat: add Method Resolution Order (MRO) with language-specific rules
Implement full MRO computation for multi-language inheritance hierarchies:
- HAS_METHOD edges: Class→Method ownership edges emitted during parsing
(both worker pool and sequential fallback paths)
- Method signatures: extract parameterCount and returnType from AST nodes
- C# heritage fix: distinguish EXTENDS vs IMPLEMENTS for base_list captures
using symbol table lookup + I[A-Z] naming heuristic fallback
- MRO processor (Phase 4.5): walks inheritance DAG, detects method-name
collisions across parents, applies language-specific resolution:
- C++: leftmost base class in declaration order wins
- C#/Java: class method wins over interface default
- Python: C3 linearization with cycle detection
- Rust: no auto-resolution (requires qualified syntax)
- Default: first definition in BFS order wins
- OVERRIDES edges emitted for resolved method collisions
- KuzuDB schema: Method table extended with parameterCount/returnType;
dedicated CSV writer and COPY query for 10-column Method rows
- MCP tools: updated Cypher examples for HAS_METHOD, OVERRIDES, diamond
72 tests across 5 test files covering MRO resolution, HAS_METHOD edges,
method signature extraction, C# heritage resolution, and integration
tests across C#/Rust/Python/TS/Java/C++.
* feat: add scope-based symbol resolution replacing raw lookupFuzzy
Introduces a shared 3-tier resolveSymbol function used by both
heritage-processor and call-processor:
1. Same-file (lookupExactFull — authoritative)
2. Import-scoped (filtered by ImportMap — high confidence)
3. Global fuzzy (first match — low confidence fallback)
Adds lookupExactFull to SymbolTable returning full SymbolDefinition
with type info needed for heritage Class/Interface disambiguation.
* refactor: tighten symbol resolution — Tier 3 refuses ambiguous matches
- lookupExactFull now O(1) via direct SymbolDefinition storage in fileIndex
(shared object references with globalIndex — zero additional memory)
- Added resolveSymbolInternal() preserving { definition, tier, candidateCount }
for test assertions and logging
- Tier 3 now returns null when multiple global candidates exist instead of
arbitrary allDefs[0] — a wrong edge is worse than no edge
- call-processor: renamed fuzzy-global → unique-global, removed dead branch
- 12 new tests: tier assertions, ambiguous refusal per language family,
heritage false-positive guard, O(1) shared reference verification
* fix: critical language support bugs in import resolution and MRO
Phase 5 critical fixes from all-language analysis:
- Python: add relative_import query capture (PEP 328) — `.models`, `..utils`
were silently dropped, producing zero ImportMap entries
- Rust: extract prefix from grouped imports `crate::module::{A, B}` — brace
groups previously failed resolution entirely
- Swift: use normalizedFileList for Windows path compatibility in module
import resolution (matches Go's resolveGoPackage pattern)
- MRO: fix c_sharp → csharp language name mismatch (enum is 'csharp'),
add Kotlin to C#/Java resolution rules (class method wins over interface)
* feat: add strict multi-language integration tests + fix C/C++ import resolution
Add 32 integration tests across 6 language fixtures (TypeScript, C#, C++,
Java, Python, Rust) with exact toBe/toEqual assertions validating heritage
edges, import resolution, and trait implementations.
Fix C/C++ import resolution bug where dot-to-slash conversion mangled
include paths (e.g. "animal.h" became "animal/h"). Now skips conversion
for C/C++ languages which use actual file paths in #include directives.
* fix: language-gate heritage heuristic, add Swift extension heritage, handle Rust grouped imports
- Gate I[A-Z] naming heuristic to C#/Java only (was firing for all languages)
- Swift unresolved types default to IMPLEMENTS (protocol conformance is the norm)
- Add tree-sitter query for Swift extension protocol conformance (extension Foo: Protocol)
- Handle Rust top-level grouped imports (use {crate::a, crate::b}) in both import loops
- Add 4 new heritage-processor tests (TypeScript refusal, Swift default, Swift Tier 1)
* feat: add Go struct embedding heritage + PackageMap optimization
Add Go struct embedding detection (anonymous fields → EXTENDS edges) via
new tree-sitter heritage query with named-field filtering in both
parse-worker and heritage-processor paths.
Implement PackageMap optimization for Go cross-package resolution:
replace O(N) file-level ImportMap expansion with directory-level suffix
matching (Tier 2b in symbol resolver). Graph IMPORTS edges are preserved
via addImportGraphEdge split.
Remove overly broad @definition.type from GO_QUERIES that was
double-matching structs/interfaces as TypeAlias nodes, breaking Tier 3
unique-global resolution.
Add Go fixture (go-pkg) with Admin→User embedding, cross-package calls,
and 7 integration tests covering structs, functions, imports, calls,
and heritage edges.
* test: add Kotlin heritage integration tests
Adds a kotlin-heritage fixture and 7 integration tests validating
class inheritance, interface implementation, JVM-style import
resolution, and symbol-table-driven EXTENDS/IMPLEMENTS disambiguation
via Kotlin delegation specifiers.
* feat: extract resolvers, add PHP tests, ambiguous tests for all languages
- Extract language-specific resolvers from import-processor.ts into
resolvers/ directory (P7): jvm, go, csharp, php, rust, standard, utils
- import-processor.ts reduced from 1412 to 711 lines (50% reduction)
- Add comprehensive PHP integration tests: PSR-4 imports, traits, enums,
heritage edges, method calls, MRO overrides
- Add ambiguous symbol resolution tests for all 9 languages verifying
correct disambiguation via import chains
- Split monolithic lang-resolution.test.ts (1080 lines) into 9 per-language
files under test/integration/resolvers/ with shared helpers
* feat: update integration tests to include resolver tests for multiple languages
* fix: address code review — schema gap, Rust impl name, Property OVERRIDES
Bugs fixed:
- Add 13 missing FROM/TO pairs in RELATION_SCHEMA for HAS_METHOD edges
(Class/Interface/Struct/Trait/Impl/Record to Method/Constructor/Property)
- Fix findEnclosingClassId to pick implementing type for Rust
impl Trait for Struct blocks (was picking trait name)
- Exclude Property nodes from MRO OVERRIDES collision detection
- Change MRO language fallback from typescript to unknown
Tests added:
- Unit: Property OVERRIDES exclusion (2 tests), Rust impl Trait for
Struct name resolution (2 tests), schema HAS_METHOD pair coverage
- Integration: no OVERRIDES targets Property nodes across all 9 languages
- PHP fixture: added shared $status property to both traits to create
real collision scenario for Property OVERRIDES exclusion test
Documentation:
- OVERRIDES edge direction (Class to Method), Go return type gap,
BFS first-reach heuristic limitation
* feat: harden CALLS-edge resolution — Phase 0 validation
- Fix same-file confidence (0.85 → 0.95) to correctly outrank import-scoped (0.9)
- Fix Tier 1 overload preservation: use globalIndex filter instead of fileIndex lookup
- Add callable-kind guard: refuse CALLS edges to Interface and Enum symbols
- Fix Kotlin countCallArguments: handle call_suffix → value_arguments nesting
- Fix Kotlin extractFunctionName: add simple_identifier to fallback search
- Strictly type findParameterList and countCallArguments (remove all `any`)
- Add arity-based call resolution integration tests for 9 languages
- Add unit regression tests for Interface/Enum CALLS refusal
* chore: remove C# build artifacts from fixtures
* feat: add call-form discrimination and ownerId to symbol table (Phase 1)
Add inferCallForm() and extractReceiverName() to distinguish free/member/constructor
calls at the AST level across all 9 languages. Add ownerId field to SymbolDefinition
linking Method/Constructor/Property to their owning class. Includes 36 unit tests
and member-call integration tests for all 9 languages (132 tests, 0 failures).
* feat: constructor/struct-literal resolution across all languages (Phase 2)
Add constructor discrimination to CALLS-edge resolution: new Foo(),
User{...} struct literals, and C# primary constructors now resolve to
Constructor/Class/Struct/Record nodes instead of being filtered out.
Queries: new_expression (C++), object_creation_expression (PHP),
composite_literal (Go), struct_expression (Rust), primary constructor
and implicit_object_creation_expression (C#).
Relaxes global tier in collectTieredCandidates to pass all candidates
through filterCallableCandidates, allowing kind/arity narrowing to
disambiguate at lower confidence.
* feat: receiver-constrained resolution with integration tests for all 9 languages
Add receiver-type filtering (Phase 3): when a member call like `user.save()`
has a known receiver type from TypeEnv, filter candidates by ownerId to
disambiguate methods with the same name across different classes.
Key changes:
- call-processor: build per-file TypeEnv, pass receiverTypeName to resolveCallTarget
- parse-worker: extract receiverTypeName from TypeEnv in worker thread
- resolveCallTarget: new step D filters by ownerId matching receiver type
- utils: extractReceiverName supports C++ field_expression (argument field)
- utils: findEnclosingClassId extracts Go method receiver types
- type-env: handle Go qualified_type, Kotlin user_type/variable_declaration
- parse-worker + parsing-processor: Function added to needsOwner for
Kotlin/Rust/Python class methods captured as Function nodes
Integration tests added for receiver-constrained resolution across all 9
languages: TypeScript, Java, Python, Go, Rust, C++, C#, Kotlin, PHP.
* feat: NamedImportMap, scoped TypeEnv, broadened signatures + TS rest-param variadic fix
Address all 4 PR #238 review items:
1. Remove redundant lookupFuzzy in processRoutesFromExtracted
2. Add NamedImportMap for TS/Python symbol-level import tracking (Tier 2a)
3. Make TypeEnv scope-aware (Map<scopeKey, Map<varName, type>>) to fix
non-deterministic receiver resolution across functions
4. Broaden extractMethodSignature: Go/Rust/C++ return types, variadic
detection for Go/Java/Python/C++/Kotlin/TypeScript rest params
Discovered and fixed: TS rest params (...args) were not detected as
variadic — added rest_pattern detection inside required_parameter nodes.
Integration tests added: scoped receiver, named import disambiguation,
and variadic call resolution for both TypeScript and Python.
* fix: alias import resolution, Go multi-assign TypeEnv, dead code removal
- NamedImportMap now stores {sourcePath, exportedName} so aliased imports
(import { User as U }) resolve U → User in the source file
- Named binding check moved before empty-allDefs early return in both
call-processor and symbol-resolver, fixing constructor calls via aliases
- Go extractFromGoShortVarDeclaration iterates all LHS/RHS pairs for
multi-assignment (user, repo := User{}, Repo{}) instead of only first
- Remove unused TYPED_DECLARATION_TYPES set (TYPED_PARAMETER_TYPES kept)
- Integration tests for both fixes (go-multi-assign, typescript-alias-imports)
* feat: alias import extraction for Kotlin, Rust, PHP, C# + integration tests
Add named import alias extraction to both pipeline paths
(import-processor.ts and parse-worker.ts) for Kotlin, Rust, PHP,
and C#. Add integration test fixtures and tests for all 5 languages
(Python alias extraction already worked, just needed the test).
Each test verifies: class detection, member call resolution through
aliases to correct target files, and IMPORTS edge emission.
* refactor: use SupportedLanguages enum everywhere instead of raw strings
Replace all raw language string literals and `language: string` types
with the SupportedLanguages enum across 10 files. This ensures
compile-time safety for language dispatch and eliminates dead
`language === 'tsx'` checks (tsx maps to TypeScript in the enum).
* fix: tier-ordering bug, re-export chains, PHP grouped imports, Java named imports
- Fix collectTieredCandidates tier-ordering: same-file now checked before
named bindings, preventing imports from shadowing local definitions
(matches resolveSymbolInternal priority order)
- Add re-export chain resolution for TypeScript/JavaScript barrel files:
export { X } from './base' and export type { X } from './base' now
followed up to 5 hops through NamedImportMap
- Fix PHP grouped import alias extraction: use App\Models\{User, Repo as R}
now correctly handled in both parse-worker and import-processor
- Add Java NamedImportMap support: import com.example.models.User now
records User as a named binding for precise disambiguation
- Add 16 new integration tests across TypeScript, PHP, and Java resolvers
(220 total resolver tests, all passing)
* refactor: consolidate alias extraction + add variadic/constructor/shadow integration tests
- Extract shared named-binding-extraction.ts from duplicate logic in
import-processor.ts and parse-worker.ts (net -200 lines)
- Deduplicate appendKotlinWildcard (now imported from resolvers/index.ts)
- Add integration tests: constructor calls (Kotlin, Python), variadic
resolution (Go, Java, C#, C++, Kotlin), re-export chains (Python),
local definition shadowing (Python, Go)
- Add TODO(stack-graph) for TypeEnv scope key collision
- 225 integration tests passing (was 223)
* fix: PHP non-aliased imports, Python node identity, re-export chain dedup + local-shadow tests
- PHP flat non-aliased imports (use App\Models\User) now stored in NamedImportMap
- PHP grouped non-aliased imports ({User} in {User, Repo as R}) now stored in NamedImportMap
- Python: replace non-public child.id with child.startIndex for node identity
- Extract shared walkBindingChain() from symbol-resolver and call-processor
- Add PHP variadic resolution fixture + test (variadic_parameter already covers PHP)
- Add local-shadow integration tests for Java, C#, Kotlin, Rust, PHP, C++ (6 languages)
* feat: Rust non-aliased use bindings, Kotlin non-aliased imports, re-export chain resolution
Extend NamedImportMap coverage for Rust and Kotlin non-aliased imports:
- Rust: rename collectUseAsClauses → collectRustBindings, extract terminal
scoped_identifier (use crate::models::User) and identifier in use_list
(use crate::models::{User, Repo}) into NamedImportMap. This also enables
pub use re-export chain following via walkBindingChain.
- Kotlin: extend extractKotlinNamedBindings to handle non-aliased imports
(import com.example.User), skipping wildcard imports.
- Add rust-reexport-chain fixture + 3 integration tests verifying Handler{}
resolves through mod.rs pub use to handler.rs.
- Add Kotlin heritage + constructor-calls reason assertions for non-aliased
import-resolved resolution.
- Add C# heritage test documenting namespace import tier behavior.
* fix: skip Kotlin lowercase member imports in NamedImportMap
Member imports like `import util.OneArg.writeAudit` (lowercase last
segment) must not populate NamedImportMap — same-named function imports
from different classes collide, breaking arity-based disambiguation.
Apply the same guard Java already uses: skip lowercase last segments.
* fix: skip spurious path-prefix bindings in Rust grouped imports
collectRustBindings was extracting the path segment (e.g. "models") from
`use crate::models::{User, Repo}` as a spurious NamedImportMap entry.
Skip scoped_identifier nodes that are direct children of scoped_use_list
since they are path prefixes, not importable symbols.
Adds rust-grouped-imports fixture and 4 integration tests verifying both
symbols resolve correctly and no spurious binding leaks through.
* fix: use startIndex in TypeEnv scope key to prevent same-name method collision
Two methods named identically in different classes within the same file
previously shared a scope key, causing non-deterministic type resolution.
Now keys use funcName@startIndex for uniqueness.
Also adds tests documenting destructuring assignment extraction gap.
* test: document C# namespace-level import limitation in named binding extraction
* test: document same-arity overload discrimination limitation in call processor
* perf: parallelize calls/heritage/routes processing in worker path
Worker path now runs processCallsFromExtracted, processHeritageFromExtracted,
and processRoutesFromExtracted via Promise.all instead of sequentially.
Safe because all three only read shared state and write via addRelationship's
dedup guard. Sequential fallback path stays sequential (shared LRU astCache).
Also fixes Rust collectRustBindings spurious path-prefix bindings for 3+ level
grouped imports, and adds @param JSDoc for walkBindingChain's allDefs invariant.
* docs: improve Promise.all safety comment and walkBindingChain JSDoc
Clarify that the parallelization safety comes from disjoint relationship
types + idempotent id-keyed Maps, not from lack of shared state (the
graph is shared). Strengthen allDefs JSDoc to describe silent-miss
consequence of passing pre-filtered results.
* refactor: extract language-specific processing into modular dispatch tables
Phase 1: Extract type binding logic from type-env.ts (635→125 LOC) into
type-extractors/ directory with per-language files and Record<SupportedLanguages,
LanguageTypeConfig> + satisfies dispatch.
Phase 2: Extract 5 config loaders from import-processor.ts into
language-config.ts (removed ~196 LOC of inline loaders).
Phase 3: Convert export-detection.ts switch/case to exhaustive
Record<SupportedLanguages, ExportChecker> + satisfies dispatch table,
fix node: any → SyntaxNode.
Also adds language feature matrix to README.
All 1146 unit tests and 433 integration tests pass.
* refactor: extract type binding logic into type-extractors/ directory (Phase 1)
Extract per-language type extraction from type-env.ts (635→125 LOC) into
type-extractors/ with Record<SupportedLanguages, LanguageTypeConfig> + satisfies
dispatch. 9 per-language files, shared helpers, and barrel index.
* refactor: extract config loaders to language-config.ts (Phase 2)
Move 5 language-specific config loaders and their type interfaces from
import-processor.ts into standalone language-config.ts module.
* calm fix 4 adding skills to repo [ISSUE #140]
* inspect
* unit and integration tests
* fixed hardcoded cohesion miss
* e2e tests for --skills flag for langauge/repo support
* Cohesion test e2e tests
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates CLAUDE.md / AGENTS.md context files.
| `--force`| Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale.
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
Canonical agent instructions: **[AGENTS.md](../AGENTS.md)** (GitNexus MCP rules, monorepo commands, Cursor Cloud notes). **[CLAUDE.md](../CLAUDE.md)** adds Claude Code-specific notes and points back to AGENTS.md for GitNexus.
## Non-negotiables (always apply)
- NEVER edit a function/class/method without running `gitnexus_impact` first.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename`.
- NEVER commit without running `gitnexus_detect_changes()`.
- NEVER ignore HIGH/CRITICAL risk warnings from impact analysis.
- NEVER run `npx gitnexus analyze` without `--embeddings` if `.gitnexus/meta.json` shows stored embeddings.
Full rules: **[AGENTS.md](../AGENTS.md)** (`gitnexus:start` block, Cursor Cloud section).
**Rule architecture:** Prefer this file plus optional `.cursor/rules/*.mdc` globs (YAML `globs` in frontmatter). Legacy `.cursorrules` is deprecated; content lives here.
lines.append(f"- Upstream ABI {upstream_proto_abi}{'is'ifcan_upgradeelse'is NOT'} within target runtime range ({target_abi_range[0]}..{target_abi_range[1]})")
ifcan_upgrade:
lines.append(f"- **Action:** after upgrading to tree-sitter@{TARGET_RUNTIME}, regenerate vendored parser.c from upstream `{upstream_sha}`")
else:
lines.append(f"- **Action:** wait for runtime upgrade beyond {TARGET_RUNTIME} that supports ABI {upstream_proto_abi}")
blockers["vendored-proto-abi"]=f"vendored tree-sitter-proto: upstream ABI {upstream_proto_abi} outside target range"
elifnotin_sync:
lines.append("- **Action:** review upstream changes; vendored copy may need updating")
blockers["vendored-proto-sync"]="vendored tree-sitter-proto: out of sync with upstream"
| **Writes** | Only paths required for the change; keep diffs minimal. Update lockfiles when deps change. |
| **Executes** | `npm`, `npx`, `node` under `gitnexus/` and `gitnexus-web/`; `uv run` for Python under `eval/`; documented CI/dev workflows. |
| **Off-limits** | Real `.env` / secrets, production credentials, unrelated repos, destructive git ops without confirmation. |
## Model Configuration
- **Primary:** Use a named model (e.g. Claude Sonnet 4.x). Avoid `Auto` or unversioned `latest` when reproducibility matters.
- **Notes:** The GitNexus CLI indexer does not call an LLM.
## Execution Sequence (complex tasks)
For multi-step work, state up front:
1. Which rules in this file and **[GUARDRAILS.md](GUARDRAILS.md)** apply (and any relevant Signs).
2. Current **Scope** boundaries.
3. Which **validation commands** you will run (`cd gitnexus && npm test`, `npx tsc --noEmit`).
On long threads, *"Remember: apply all AGENTS.md rules"* re-weights these instructions against context dilution.
## Claude Code hooks
**PreToolUse** hooks can block tools (e.g. `git_commit`) until checks pass. Adapt to this repo: `cd gitnexus && npm test` before commit.
## Context budget
Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.md](CONTRIBUTING.md)**. If always-on rules grow, split into **`.cursor/rules/*.mdc`** (globs). **Cursor:** project-wide rules in `.cursor/index.mdc`. **Claude Code:** load `STANDARDS.md` only when needed.
- **Call-resolution DAG (legacy path):** See ARCHITECTURE.md § Call-Resolution DAG. Typed 6-stage DAG inside the `parse` phase; language-specific behavior behind `inferImplicitReceiver` / `selectDispatch` hooks on `LanguageProvider`. Shared code in `gitnexus/src/core/ingestion/` must not name languages. Types: `gitnexus/src/core/ingestion/call-types.ts`.
- **Scope-resolution pipeline (RFC #909 Ring 3):** See ARCHITECTURE.md § Scope-Resolution Pipeline. Replaces the legacy DAG for languages in `MIGRATED_LANGUAGES` (see `registry-primary-flag.ts`). A language plugs in by implementing `ScopeResolver` (`scope-resolution/contract/scope-resolver.ts`) and registering it in `SCOPE_RESOLVERS`. CI parity gate runs BOTH paths per migrated language on every PR.
This project is indexed by GitNexus as **GitNexus** (1650 symbols, 4291 relationships, 125 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
Indexed as **GitNexus** (4325 symbols, 10556 relationships, 300 execution flows). Use MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
> If any tool warns the index is stale, run `npx gitnexus analyze` first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run`gitnexus_impact({target: "symbolName", direction: "upstream"})`and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
-When exploring unfamiliar code, use`gitnexus_query({query: "concept"})`to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
-When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use`gitnexus_context({name: "symbolName"})`.
- **MUST run impact analysis before editing any symbol.** `gitnexus_impact({target: "symbolName", direction: "upstream"})`— report blast radius to the user.
- **MUST run `gitnexus_detect_changes()` before committing** — verify only expected symbols and flows are affected.
- **MUST warn the user** if impact returns HIGH or CRITICAL risk.
-Explore unfamiliar code with`gitnexus_query({query: "concept"})`(process-grouped, ranked) instead of grepping.
-Full context on a symbol:`gitnexus_context({name: "symbolName"})`.
## When Debugging
1.`gitnexus_query({query: "<error or symptom>"})` — find execution flows related to the issue
2.`gitnexus_context({name: "<suspect function>"})` — see all callers, callees, and process participation
3.`READ gitnexus://repo/GitNexus/process/{processName}` — trace the full execution flow step by step
4.For regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})` — see what your branch changed
1.`gitnexus_query({query: "<error or symptom>"})` — find related execution flows
2.`gitnexus_context({name: "<suspect function>"})` — callers, callees, process participation
3.`READ gitnexus://repo/GitNexus/process/{processName}` — trace flow step by step
- **Renaming**: MUST use`gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Review the preview — graph edits are safe, text_search edits need manual review. Then run with `dry_run: false`.
- **Extracting/Splitting**: MUST run`gitnexus_context({name: "target"})` to see all incoming/outgoing refs, then `gitnexus_impact({target: "target", direction: "upstream"})` to find all external callers before moving code.
- After any refactor: run`gitnexus_detect_changes({scope: "all"})` to verify only expected files changed.
- **Rename:**`gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Graph edits are safe; text_search edits need manual review.
- **Extract/Split:**`gitnexus_context` (incoming/outgoing refs) then `gitnexus_impact` (upstream callers) before moving code.
-**After any refactor:**`gitnexus_detect_changes({scope: "all"})` to verify scope.
## Never Do
-NEVER edit a function, class, or method without first running `gitnexus_impact`on it.
-NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
-NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
-NEVER commit changes without running`gitnexus_detect_changes()` to check affected scope.
-Edit a symbol without running `gitnexus_impact`first.
-Ignore HIGH/CRITICAL risk warnings.
-Rename with find-and-replace — use `gitnexus_rename`.
-Commit without`gitnexus_detect_changes()`.
- Add language-specific behavior to shared ingestion code (`gitnexus/src/core/ingestion/`) — use a `LanguageProvider` hook. Seeing `provider.mroStrategy === 'xxx'` or an import from `languages/xxx.ts` in shared code means stop and add a hook.
> Group mode: pass `repo: "@<groupName>"` to fan out across all member repos, or `repo: "@<groupName>/<memberPath>"` to target a single member (path keys from `group.yaml`). Optional `service: "<monorepo/path>"` filters by service root. Group-level state (contracts, staleness) lives in the resources table below — there are **no** `group_query` / `group_context` / `group_impact` / `group_contracts` / `group_status` MCP tools.
>
> For a full walkthrough of setting up a group across multiple repos that communicate over gRPC, see [docs/guides/microservices-grpc.md](docs/guides/microservices-grpc.md).
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update these |
| d=1 | WILL BREAK — direct callers/importers | MUST update |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
@@ -56,23 +132,83 @@ This project is indexed by GitNexus as **GitNexus** (1650 symbols, 4291 relation
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/context` | Codebase overview, index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
Check `.gitnexus/meta.json``stats.embeddings` (0 = none). A plain `analyze` no longer drops existing vectors — pass `--drop-embeddings` to wipe.
> Claude Code: PostToolUse hook detects a stale index after `git commit` and `git merge` and prompts the agent to run `analyze`. The hook does not invoke `analyze` itself.
## CLI Skills
| Task | Skill file |
|------|-----------|
| Architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
cd gitnexus-web && npm run dev # Web UI: Vite on port 5173
npx gitnexus serve # HTTP API on port 4747 (from any indexed repo)
```
### Testing
**CLI / Core (`gitnexus/`)**
-`npm test` — full vitest suite (~2000 tests)
-`npm run test:unit` — unit tests only
-`npm run test:integration` — integration (~1850 tests). LadybugDB file-locking tests may fail in containers (known env issue).
-`npx tsc --noEmit` — typecheck
**Web UI (`gitnexus-web/`)**
-`npm test` — vitest (~200 tests)
-`npm run test:e2e` — Playwright (7 spec files; requires `gitnexus serve` + `npm run dev`)
-`npx tsc -b --noEmit` — typecheck
**Pre-commit hook** (`.husky/pre-commit`): formatting (prettier via lint-staged) + typecheck for staged packages. Tests do **not** run in pre-commit — CI only.
### Gotchas
-`npm install` in `gitnexus/` triggers `prepare` (builds via `tsc`) and `postinstall` (patches tree-sitter-swift, builds tree-sitter-proto). Native bindings need `python3`, `make`, `g++`.
-`tree-sitter-kotlin` and `tree-sitter-swift` are optional — install warnings expected.
- ESLint configured via `eslint.config.mjs` (TS, React Hooks, unused-imports). No `npm run lint` script; use `npx eslint .`. Prettier runs via lint-staged. CI checks both in `ci-quality.yml`.
1.**Ingestion** — `analyze.ts` → `runFullAnalysis` (`run-analyze.ts`) → `runPipelineFromRepo` (`pipeline.ts`). DAG of 12 phases builds a `KnowledgeGraph` in memory, then loads into LadybugDB under `.gitnexus/`. Repo registered in `~/.gitnexus/registry.json` for MCP discovery.
| `group_list` | List repo groups or details for one group |
| `group_sync` | Rebuild group Contract Registry (`contracts.json`) and bridge graph |
`query`, `context`, and `impact` are group-aware: pass `repo: "@<groupName>"` (or `"@<groupName>/<memberPath>"` to scope to one member) plus optional `service: "<monorepo/path>"`. Group-mode `query` merges per-repo results via Reciprocal Rank Fusion; group-mode `impact` runs the local walk in the chosen member and fans out across boundaries via the Contract Bridge (`gitnexus/src/core/group/cross-impact.ts`). The previously-planned `group_query`, `group_context`, `group_impact`, `group_contracts`, `group_status` MCP tools are intentionally not introduced — group-level state is exposed via resources instead:
-`deps: ReadonlyMap<string, PhaseResult>` — **declared deps only** (runner filters the results map to prevent hidden coupling)
3.**Error handling** — wraps phase errors with the phase name, emits terminal `error` progress event, swallows progress handler errors to preserve the original cause.
4.**Timing** — per-phase `durationMs` in `PhaseResult`, dev-mode console logging.
**Design patterns:**
- **Single graph accumulator** — all phases mutate the same `KnowledgeGraph` in `ctx`; the graph is the primary output.
Typed 6-stage pipeline in `call-processor.ts` (inside the `parse` phase) that resolves method/function calls and emits CALLS edges. Language behavior plugs in at two `LanguageProvider` hook points (stages 3–4); shared code names no languages. Scope: call resolution only — import resolution, type extraction, heritage, and symbol-table population live in other phases.
-`primary: 'owner-scoped'` — MRO walk from receiver's type; used when receiver type is known.
-`fallback: 'free-arity-narrowed'` — after owner-scoped miss, search free-call candidates by arity only (Ruby uses this for implicit-self calls that miss their owner's MRO).
-`ancestryView: 'singleton'` — walk singleton/class ancestry instead of instance ancestry (Ruby `def self.foo` bodies, so `extend`-ed methods are found).
### Adding language behavior
1.**Implicit receivers** — implement `inferImplicitReceiver`: return null if call already has a receiver; otherwise use `findEnclosingClassInfo` (`ast-helpers.ts`) to find the enclosing context, return `ImplicitReceiverOverride` with `receiverSource: 'implicit-self'`, and optionally set `hint` for `selectDispatch`.
2.**Custom dispatch** — implement `selectDispatch`: inspect `receiverSource` and `hint`, return `DispatchDecision` with `primary`, optional `fallback`, optional `ancestryView`; return null to keep shared defaults.
3.**MRO strategy** — confirm `mroStrategy` is `'first-wins'`, `'c3'`, `'ruby-mixin'`, or `'none'`; consumed by `lookupMethodByOwnerWithMRO`.
**Ruby example** (`languages/ruby.ts` + `utils/ruby-self-call.ts`): `inferImplicitReceiver` rewrites bare-identifier calls to `self.method` and sets `hint` to `'instance'`/`'singleton'`; `selectDispatch` uses hint for `ancestryView` and adds `fallback: 'free-arity-narrowed'` for implicit-self calls.
### Code references
| Module | Purpose |
|--------|---------|
| `core/ingestion/call-types.ts` | DAG types: `ReceiverEnriched`, `DispatchDecision`, `ImplicitReceiverOverride` |
| `core/ingestion/model/resolve.ts` | `lookupMethodByOwnerWithMRO`: stage 5 MRO walk |
| `core/ingestion/languages/ruby.ts` | Both hooks + `mroStrategy: 'ruby-mixin'` |
| `core/ingestion/utils/ruby-self-call.ts` | Bare-call rewrite for `inferImplicitReceiver` |
### Coexistence with the scope-resolution pipeline
The Call-Resolution DAG is the **legacy path**. RFC #909 Ring 3 introduces a parallel **scope-resolution pipeline** (next section) that replaces stages 1–6 with a scope-indexed registry lookup. Both paths ship side-by-side and are gated per-language via `MIGRATED_LANGUAGES` + the `REGISTRY_PRIMARY_<LANG>` env var.
- **Unmigrated language** → Call-Resolution DAG runs; scope-resolution phase is a no-op.
- **Migrated language** (currently: Python, C#) → scope-resolution owns CALLS/ACCESSES/USES emission; the legacy DAG gates off for that language via `isRegistryPrimary(lang)` checks in `call-processor.ts` and `import-processor.ts`.
-`import-processor` still populates `importMap` for migrated languages — heritage's `ctx.resolve` reads it to disambiguate parent classes. Only edge emission is gated.
- CI runs BOTH paths for every migrated language on every PR (`.github/workflows/ci-scope-parity.yml`); both must pass.
#### Same-graph guarantee
Edges emitted by the scope-resolution pipeline and edges emitted by the legacy DAG are indistinguishable to downstream consumers (MCP tools, HTTP API, embeddings, group bridge):
- **Node identity** — both paths use `generateId(...)` from `lib/utils.ts`, the same qualified-name keyspace, and the same node labels (`File`, `Folder`, `Class`, `Method`, `Function`, …). Overload disambiguation suffixes `parameterTypes` into the id consistently — see `scope-resolution/graph-bridge/ids.ts` and the legacy emitter in `call-processor.ts`.
- **Edge vocabulary** — both paths emit the same reasons: `'import-resolved' | 'global' | 'local-call' | 'same-file' | 'interface-dispatch' | 'read' | 'write'`. Migrating a language must not change which reasons consumers see for previously-resolved edges.
- **Confidence tier** — both paths attach a numeric `confidence` to each edge using the same scale.
The CI parity workflow (`.github/workflows/ci-scope-parity.yml`) runs both paths against every migrated language's fixture corpus and fails on any divergence.
#### Semantic-model source of truth
Two independent invariants.
**ParsedFile = the AST-level truth.**`ParsedFile` (`gitnexus-shared/src/scope-resolution/parsed-file.ts`) is the single per-file artifact both resolution paths consume. Scope-resolution passes MUST NOT build a parallel parse representation. If a per-language hook needs AST-level facts that `ParsedFile` doesn't expose, it should reuse the orchestrator's `treeCache` (`RunScopeResolutionInput.treeCache`) rather than re-invoking `parser.parse(...)` on its own — the C# `populateNamespaceSiblings` hook is the reference implementation of this pattern.
**SemanticModel = the symbol-level truth.**`SemanticModel` (`gitnexus/src/core/ingestion/model/semantic-model.ts`) is the authoritative store for every symbol-indexed lookup (by `nodeId`, `simpleName`, `qualifiedName`, or `filePath`). Both paths read from here:
- Legacy Call-Resolution DAG → `call-processor` Tier 1/2/3 via `model.symbols.lookupExactAll`, `model.methods.lookupMethodByName`, `model.types.lookupClassByName`, `lookupMethodByOwnerWithMRO`.
Read phase: all resolution passes + MCP + HTTP + embeddings see
SemanticModel (read-only handle); writes are type-errors.
```
`runScopeResolution` narrows `MutableSemanticModel` → `SemanticModel` at the phase boundary so downstream passes physically cannot mutate the model even accidentally.
**Transitional: reconciliation pass.**`reconcileOwnership` (`scope-resolution/pipeline/reconcile-ownership.ts`) is a shim for languages whose legacy extractor doesn't resolve `enclosingClassId` at parse time (Python class-body methods are the canonical case). It walks `parsed.localDefs[i].ownerId` after `populateOwners` and registers any missed methods/fields into the model. Idempotent — safe to re-run, safe alongside languages whose legacy extractor already carries `ownerId` (C#).
The architectural end state is for every language's parse-time extractor to emit the correct `ownerId` directly, making reconciliation a no-op (tracked as a follow-up refactor). The dev-mode validator `validateOwnershipParity` surfaces any drift via `onWarn` under `NODE_ENV !== 'production' && VALIDATE_SEMANTIC_MODEL !== '0'`.
Language-agnostic registry-primary resolver. Replaces the Call-Resolution DAG for migrated languages. Adding a language is one interface implementation (`ScopeResolver`) plus two registrations — no changes to shared code, no new pipeline phase.
Orchestrator: `runScopeResolution(input, provider)` in `scope-resolution/pipeline/run.ts`.
Pipeline phase: `scopeResolutionPhase` in `scope-resolution/pipeline/phase.ts` — iterates `SCOPE_RESOLVERS ∩ MIGRATED_LANGUAGES`, reads per-file Trees from the parse phase's `scopeTreeCache`, disposes the cache at the end.
### `ScopeResolver` contract
Single interface a language implements to plug into the pipeline. Contract fully documented in `scope-resolution/contract/scope-resolver.ts`.
| `hoistTypeBindingsToModule?` | Walk up to Module scope when looking up a method's return-type typeBinding — default off; enable only when bindings are stored at module level |
### Per-language registration
1. Implement `ScopeResolver` in `languages/<lang>/scope-resolver.ts`.
2. Add entry to `SCOPE_RESOLVERS` in `scope-resolution/pipeline/registry.ts`.
3. Add the language to `MIGRATED_LANGUAGES` in `registry-primary-flag.ts` when the shadow-harness corpus parity ≥ 99% fixtures / ≥ 98% corpus.
CI auto-discovers the set via `tsx`. No workflow edit required.
- **Cross-phase Tree cache**: parse phase writes Trees into `scopeTreeCache` (separate from the chunk-local `astCache`) ONLY for languages with `emitScopeCaptures`. Scope-resolution reads from it to skip the second parse. Cleared at end of the phase. Workers leave the cache empty — Trees can't cross MessageChannels; cache miss = fresh parse. `PROF_SCOPE_RESOLUTION=1` emits hit/miss counters and a worker-engaged warning.
- **Typed relationship iteration**: heritage + MRO walk only the EXTENDS / IMPLEMENTS / HAS_METHOD edges via `iterRelationshipsByType`, not the full relationship map.
- **Workspace-resolution-index**: O(1) `findOwnedMember` / `findExportedDef` / `classScopeByDefId` built once per run.
- **SCC-ordered cross-file return-type propagation** (PR #1050): `propagateImportedReturnTypes` walks `indexes.sccs` in reverse-topological order (leaves first), so multi-hop alias chains like `models.User → service.user → app.user` collapse to the terminal class in a single linear pass. Within each importer, the source module's `typeBindings` is chain-followed BEFORE mirroring (so we mirror terminal types, not intermediate refs), and the importer's own `typeBindings` is chain-followed AFTER mirroring (so local `const x = importedFn()` resolves before downstream importers run). Cyclic SCCs reach a partial fixpoint within a single pass without iterating to convergence — see the `ts-circular` cross-file-binding fixture which only asserts pipeline-no-throw. PROF output (`PROF_SCOPE_RESOLUTION=1`) splits `finalize` from `propagate` so quadratic regressions in the chain-follow surface independently.
---
## Language-agnostic graph feeding
16 languages → single unified graph. Four abstraction layers:
| `exportChecker` | Public/exported symbol detection |
| `typeConfig` | Type annotation extraction rules |
| `mroStrategy` | `first-wins` / `c3` / `none` |
16 providers in `languages/index.ts` via `satisfies Record<SupportedLanguages, LanguageProvider>` — missing a language is a compile error.
### Unified capture tags
Per-language tree-sitter queries use different AST node names but produce the **same semantic capture tags**: `@definition.class`, `@definition.function`, `@call.name`, `@import.source`, `@heritage.extends`. Downstream extraction needs no language branching. Defined in `tree-sitter-queries.ts`.
### Import resolution
Per-language import resolution uses the **configs + factory** pattern (like call/method/class extractors). Each language declares an `ImportResolutionConfig` in `import-resolvers/configs/`, listing an ordered chain of `ImportResolverStrategy` functions. `createImportResolver()` (in `resolver-factory.ts`) composes them: first non-null result wins. Low-level helpers shared across strategies live alongside the configs in `import-resolvers/` (e.g. `go.ts`, `rust.ts`, `python.ts`).
Unified 3-tier algorithm (`model/resolution-context.ts`), per-language `importSemantics` controls which tier activates:
| Tier | Confidence | Mechanism |
|------|-----------|-----------|
| 1 — same-file | 0.95 | Symbol table for caller's file |
| 2 — import-scoped | 0.9 | `NamedImportMap` chains (named) or all files in `importMap` (wildcard) |
| 3 — global | 0.5 | O(1) index lookups: class, impl, callable. Fallback only |
All languages emit unified `ExtractedHeritage` (child, parent, `EXTENDS`/`IMPLEMENTS`). MRO phase walks the heritage graph using per-language strategy:
- **`first-wins`** — Java, C#, C++, TS, Ruby, Go
- **`c3`** — Python (C3 linearization)
- **`none`** — single-inheritance languages
Unified walk: `lookupMethodByOwnerWithMRO()` in `model/resolve.ts`.
---
## Full analysis flow
`runFullAnalysis` in `run-analyze.ts` orchestrates everything around the pipeline:
Node IDs use arity suffix (`#<paramCount>`): `Method:file:Class.method#1` vs `#2`.
**Same-arity disambiguation:** type-hash suffix `~type1,type2` when collision detected and type annotations present. Languages without types (Python, Ruby, JS) use arity-only. TS/JS overload signatures excluded (collapse to implementation body). See #651.
**C++ const-qualified:**`$const` suffix after type-hash when non-const collision exists: `Method:file:Container.begin#0$const`.
**Generic/template types:** type-hash uses `rawType` (full AST text including generics): `~vector<int>` vs `~vector<std::string>`.
**ID stability:** collision-only tags mean IDs change when overloads are added. `save#1` becomes `save#1~int` when `save(String)` is added.
**Variadic matching:** confidence 0.7 when one side is variadic and the other has fixed count.
**METHOD_IMPLEMENTS confidence tiering:**
| Match quality | Confidence |
|---|---|
| Exact parameter types match | 1.0 |
| Arity match, types unavailable | 1.0 |
| Variadic vs fixed | 0.7 |
| Insufficient info | 0.7 |
## Related docs
- [MIGRATION.md](MIGRATION.md) — breaking changes and migration guidance
- [RUNBOOK.md](RUNBOOK.md) — operational commands and recovery
- [GUARDRAILS.md](GUARDRAILS.md) — safety boundaries for humans and agents
- [TESTING.md](TESTING.md) — how to run tests
-`AGENTS.md` / `CLAUDE.md` — agent workflows and tool usage
Metadata: version, last reviewed, scope, model policy, reference docs, changelog.
Last updated: 2026-03-22
-->
This project is indexed by GitNexus as **GitNexus** (1650 symbols, 4291 relationships, 125 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
Last reviewed: 2026-04-13
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
**Project:** GitNexus · **Environment:** dev · **Maintainer:** repository maintainers (see GitHub)
## Always Do
Follow **AGENTS.md** for the canonical rules; this file adds Claude Code–specific deltas. Cursor-specific notes live only in `AGENTS.md`.
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Scope
## When Debugging
See the **Scope** table in [AGENTS.md](AGENTS.md) for read/write/execute/off-limits boundaries. Cursor-specific workflow notes also live only in AGENTS.md.
1.`gitnexus_query({query: "<error or symptom>"})` — find execution flows related to the issue
2.`gitnexus_context({name: "<suspect function>"})` — see all callers, callees, and process participation
3.`READ gitnexus://repo/GitNexus/process/{processName}` — trace the full execution flow step by step
4. For regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})` — see what your branch changed
## Model Configuration
## When Refactoring
- **Primary:** Pin per **Claude Code** / Anthropic org policy (explicit model id). Do not rely on an unversioned `latest` alias for governed workflows.
- **Fallback:** As configured in Claude Code (organization default or user override).
- **Notes:** The GitNexus CLI analyzer does not call an LLM.
- **Renaming**: MUST use `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Review the preview — graph edits are safe, text_search edits need manual review. Then run with `dry_run: false`.
- **Extracting/Splitting**: MUST run `gitnexus_context({name: "target"})` to see all incoming/outgoing refs, then `gitnexus_impact({target: "target", direction: "upstream"})` to find all external callers before moving code.
- After any refactor: run `gitnexus_detect_changes({scope: "all"})` to verify only expected files changed.
## Execution Sequence (complex tasks)
## Never Do
Same discipline as [AGENTS.md](AGENTS.md): before large multi-step work, state which **AGENTS.md** / **GUARDRAILS.md** rules apply, current **Scope**, and planned validation commands (`npm test`, `tsc`, etc.). When pausing, summarize progress in the chat or a **local** scratch file (do not add `HANDOFF.md` to the repo), then `/clear` and resume with that summary.
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Claude Code hooks
## Tools Quick Reference
Prefer **PreToolUse** hooks for hard gates (e.g. tests before `git_commit`). Adapt hook commands to `gitnexus/` npm scripts.
If always-on instructions grow, load deep conventions via conditional reads (e.g. *“When writing new code, read STANDARDS.md”*) instead of pasting long blocks here. In Cursor, prefer `.cursor/index.mdc` plus optional `.cursor/rules/*.mdc` globs (see [AGENTS.md](AGENTS.md) § Context budget).
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update these |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
- **Call-resolution DAG:** See ARCHITECTURE.md § Call-Resolution DAG. Shared pipeline code in `gitnexus/src/core/ingestion/` must not name languages — use `LanguageProvider` hooks instead (see AGENTS.md).
- **GitNexus:** `.claude/skills/gitnexus/`; MCP and indexed-repo rules live only in [AGENTS.md](AGENTS.md) (`gitnexus:start` … `gitnexus:end`). See **GitNexus rules** below.
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
Before completing any code modification task, verify:
1.`gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL risk warnings were ignored
3.`gitnexus_detect_changes()` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
---
## CLI
## GitNexus rules
- Re-index: `npx gitnexus analyze`
- Check freshness: `npx gitnexus status`
- Generate docs: `npx gitnexus wiki`
<!-- gitnexus:end -->
See the `<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** for the canonical MCP tools, impact analysis rules, and index instructions.
How to propose changes, run checks locally, and open pull requests.
## License
This project uses the [PolyForm Noncommercial License 1.0.0](https://polyformproject.org/licenses/noncommercial/1.0.0/). By contributing, you agree your contributions are licensed under the same terms unless stated otherwise.
## Where to discuss
- **Issues & feature ideas:** use [GitHub Issues](https://github.com/abhigyanpatwari/GitNexus/issues) for the upstream repo, or your fork’s tracker if you work from a fork.
- **Community:** see the Discord link in the root [README.md](README.md).
4. Run tests as described in [TESTING.md](TESTING.md).
## Branch and pull requests
- Use short-lived branches off the default branch of the repo you are targeting.
- **PR titles MUST follow the conventional-commit format** — `pr-labeler.yml` enforces this on every PR and auto-applies the matching label so release notes group the change correctly.
- **PR description:** what changed, why, how to verify (commands), and any risk or rollback notes.
### Pull request titles
Format: `<type>[(scope)][!]: <subject>`
Allowed types and the release-notes section each one lands in (defined in `.github/release.yml`):
Append `!` to the type (e.g. `feat(api)!: drop /v1 endpoint`) or include `BREAKING CHANGE:` in the PR body to flag a breaking change — the labeler then adds the `breaking` label and the 💥 Breaking Changes section is rendered first.
Commits within a PR may use any style — only the **merged PR title** shows up in release notes, so that's the one the convention applies to.
## Before you open a PR
- [ ] Tests pass for the packages you touched (`gitnexus` and/or `gitnexus-web`).
- [ ] Typecheck passes: `npx tsc --noEmit` in `gitnexus/` and `npx tsc -b --noEmit` in `gitnexus-web/`.
- [ ] No secrets, tokens, or machine-specific paths committed.
- [ ] Documentation updated if behavior or public CLI/MCP contract changes.
- [ ] Pre-commit hook runs clean (`.husky/pre-commit` — formatting via lint-staged + typecheck for staged packages; tests run in CI only).
## Code review
Maintainers may request changes for correctness, tests, performance, or consistency with existing patterns. Keeping diffs focused makes review faster.
## GitHub Actions — Concurrency Convention
Every workflow under `.github/workflows/` MUST declare a top-level `concurrency:` block using this convention:
- **Group key** starts with `${{ github.workflow }}` so no two workflows can collide on the same group name. The discriminator that follows is chosen per event shape:
-`workflow_run` scope (e.g. `ci-report.yml`): `${{ github.workflow }}-${{ github.event.workflow_run.pull_requests[0].number || format('{0}/{1}', github.event.workflow_run.head_repository.full_name, github.event.workflow_run.head_branch) }}` — the fork fallback must be stable across reruns (never `workflow_run.id`, which is per-run-unique and defeats serialization).
- Global single-slot (manual dispatch utilities): `${{ github.workflow }}`
- **Reusable workflows invoked via `workflow_call`:** do NOT use `${{ github.workflow }}` in the group key — in called-workflow context its evaluation is ambiguous and can resolve to the caller's name, which would deadlock against the caller's own group. Use a hardcoded literal prefix and a `github.event_name`-aware expression that falls through to `github.run_id` for reusable invocations (see `ci.yml` for the canonical form). Approved literal prefixes: `CI-` (`ci.yml`) and `docker-build-push-` (`docker.yml`). The `check-workflow-concurrency.py` validation script must be updated whenever a new approved literal prefix is added.
- **Merge queue (`merge_group`)**: when this event is added, use `${{ github.workflow }}-${{ github.event.merge_group.head_ref }}` with `cancel-in-progress: false` (every queue entry is a distinct ref; never cancel).
- **`cancel-in-progress` policy:**
| Event | `cancel-in-progress` | Why |
|-------|----------------------|-----|
| `pull_request` CI run | `true` | New push supersedes old run |
| `push` to `main` | `false` | Every main commit gets validated |
| Tag push (`v*` publish) | `false` | Never cancel mid-publish |
| `push` to `main` for release-candidate | `false` | Never cancel mid-RC publish |
- For workflows that serve multiple events at once (e.g. `ci.yml` handles `pull_request`, `push`, and `workflow_call`), make `cancel-in-progress` event-aware:
- When adding a new workflow, copy the concurrency block from an existing workflow of the same event shape.
## AI-assisted contributions
If you use coding agents, follow project context files (e.g. `AGENTS.md`, `CLAUDE.md`) and avoid drive-by refactors unrelated to the issue. Prefer incremental, test-backed changes.
## Releases
Two publish workflows ship `gitnexus` to npm:
- **Stable** (`.github/workflows/publish.yml`) — triggered by pushing any `v*`
tag. Publishes to the `latest` dist-tag with a changelog-backed GitHub
release. Maintainers are expected to tag from `main` as a convention; the
workflow itself does not enforce branch reachability.
- **Release Candidate** (`.github/workflows/release-candidate.yml`) — runs on
every push to `main` (typically a merged PR) plus manual dispatch. Docs-only
changes are skipped via `paths-ignore`. Publishes to the `rc` dist-tag with
version `X.Y.Z-rc.N` and a GitHub prerelease, where:
- `X.Y.Z` is selected automatically. On push (and on dispatch with
`bump: auto`, the default) the workflow **continues the active rc cycle**:
if the registry already has `X.Y.Z-rc.*` versions with `X.Y.Z` > current
`latest`, it reuses the highest such base; otherwise it patch-bumps
from `latest`. Dispatching with `bump: patch|minor|major` **resets**
the cycle from `latest`.
- `N` is auto-incremented against existing `X.Y.Z-rc.*` entries on the
registry. First rc for a given base is `rc.1`.
- After the npm publish succeeds, the workflow calls `docker.yml` as a
reusable workflow to build and push the corresponding RC Docker images
(e.g. `ghcr.io/abhigyanpatwari/gitnexus:1.7.0-rc.1`, mirrored to
`docker.io/akonlabs/gitnexus:1.7.0-rc.1`). The images are signed
with Cosign; the OIDC identity is `docker.yml@refs/heads/main` (the
caller's ref — see README.md § Docker for the verify command).
Idempotency: the workflow pushes an `rc/<HEAD_SHA>` marker tag and a
`v<RC>` release tag **atomically, before** calling `npm publish`. The guard
refuses to re-run once the marker exists, so a post-publish failure will
not mint a duplicate rc for the same commit. The `v<RC>` tag points at a
detached release commit whose `package.json` matches the npm tarball
exactly (traceable releases). Recovery after a partial failure:
```bash
git push --delete origin rc/<HEAD_SHA> v<RC>
# then redispatch the workflow with force: true
```
**Docker-only partial failure:** if `publish` succeeds (npm tarball + tags
are live) but the `docker` job subsequently fails (e.g. GHCR flakiness),
the npm RC is already published and the `rc/<HEAD_SHA>` marker is in place.
Re-running `release-candidate.yml` with `force: true` will abort at the
"Version already exists on npm" guard. To recover without cutting a new RC:
```bash
# 1. Manually trigger only the docker workflow, passing the existing RC tag:
gh workflow run docker.yml --ref main -f tag=v<RC_VERSION>
# (requires a workflow_dispatch trigger on docker.yml — see note below)
```
Because `docker.yml` intentionally has no `workflow_dispatch` (images are
tag-driven by design), the practical recovery options are:
- Wait for the next commit on `main`, which will cut a new RC that includes
the Docker build.
- Manually run `docker build` + `docker push` locally and sign with Cosign
against the same digest.
- Delete `rc/<HEAD_SHA>` and `v<RC>` tags, then redispatch with `force:
true` to re-run the full RC pipeline (cuts a new RC number).
The rc workflow never moves `latest`. To verify after a change, inspect dist-tags:
This document defines the repo-wide completion bar for production-ready changes in GitNexus. It is the stable baseline. Implementation prompts, agent behavior, and review workflows may add task-specific checks, but they must never weaken this bar.
Use it together with:
-`AGENTS.md` — agent-facing rules of engagement
-`GUARDRAILS.md` — hard safety constraints
-`CONTRIBUTING.md` — contributor workflow
-`TESTING.md` — test strategy and coverage expectations
A change is **Done** when it is correct, safely integrated, appropriately tested, operationally sound, and a net improvement to the codebase — not merely "the code compiles and a test passes."
This DoD applies to:
- CLI, MCP, and HTTP-bridge behavior in `gitnexus/`
- Browser UI in `gitnexus-web/`
- Shared contracts in `gitnexus-shared/`
- CI workflows, release pipelines, and repo-level docs
Out of scope: full agent personas, step-by-step implementation prompts, verbose review formatting rules, repo walkthroughs already covered elsewhere, temporary task-specific acceptance criteria. Those belong in prompts, PR templates, or other repo docs.
## 2. Core Definition of Done
Every change must satisfy **every relevant item** below. If an item does not apply, say so explicitly in the PR description.
### 2.1 Correctness and Completeness
- [ ] The requested behavior is implemented end-to-end in the **real runtime path** for the affected surface — no dead code, partial wiring, test-only shims, or "works in isolation but not in production" seams.
- [ ] Edge cases relevant to the changed surface are handled or explicitly documented as out of scope.
- [ ] Error handling is proportionate: inputs at system boundaries (user input, external APIs, filesystem, process spawn) are validated; internal, framework-guaranteed paths are trusted.
- [ ] The change produces the same result on re-run (idempotent where expected) and does not rely on accidental ordering.
### 2.2 Architecture and Placement
- [ ] The change is placed in the correct package and layer:
-`gitnexus/` for CLI, MCP, HTTP bridge, ingestion, graph, and runtime logic
-`gitnexus-web/` for browser UI (thin client — no WASM workers, all queries via HTTP API)
-`gitnexus-shared/` for shared contracts, types, and constants
- [ ] Pipeline and architecture boundaries remain explicit. Shared ingestion code in `gitnexus/src/core/ingestion/` must not name languages — use `LanguageProvider` hooks (see `AGENTS.md` and `ARCHITECTURE.md` § Call-Resolution DAG).
- [ ] No hidden cross-phase coupling; no leaking of language-specific logic into shared infrastructure without a documented architectural reason.
- [ ] Runtime and graph behavior are consistent — the real source of truth is fixed at the source, not symptom-patched in a downstream layer.
- [ ] Direct imports from `gitnexus-shared` are used. No barrel re-exports introduced to paper over drift between packages.
### 2.3 Design and Readability
- [ ] The implementation is the **smallest correct solution** for the requirement. No speculative abstraction, unnecessary indirection, clever but hard-to-follow control flow, or unrelated cleanup.
- [ ] Naming, control flow, ownership, and extension points are clear enough that the next contributor can extend the code without archaeology.
- [ ] Comments are minimal and useful — they explain intent, invariants, contracts, or non-obvious constraints. No stale comments, placeholder comments, narrated code, commented-out code, or "what" comments where a good name would do.
- [ ] No copy-paste duplication created for convenience; no premature deduplication of three similar lines.
### 2.4 Contracts and Compatibility
- [ ] Existing contracts (types in `gitnexus-shared/`, CLI flags, MCP tools/resources, HTTP routes, graph node/edge shapes, persisted IDs) are preserved unless the task explicitly requires a contract change.
- [ ] Any contract change is intentional, explicit, and reflected in **every direct consumer** in the same change, with types aligned end-to-end.
- [ ] Persisted data changes (graph schema, IDs, embeddings) are backward-compatible or accompanied by a documented migration / reindex path.
- [ ] If user-visible behavior, public usage, CLI help, or README examples change, the relevant docs, examples, help text, or migration notes are updated in the same change.
### 2.5 Security
- [ ] No new injection surfaces (command, path, SQL/Cypher-style, prompt) introduced on paths that consume untrusted input.
- [ ] No secrets, tokens, or credentials committed to the repo, to logs, or to error messages.
- [ ] Filesystem access honors the repo-scope and indexed-repo boundaries documented in `AGENTS.md` and `GUARDRAILS.md`.
- [ ] Third-party dependencies added or bumped are justified, from reputable sources, and do not regress the supply-chain posture.
### 2.6 Performance and Resource Use
- [ ] No repeated avoidable work, unnecessary scans, unnecessary round-trips, unbounded caches, or obvious hot-path regressions.
- [ ] Tree-sitter buffer sizing follows the adaptive 512KB–32MB convention (`getTreeSitterBufferSize`) — do not hard-code new buffer sizes.
- [ ] Memory and handle lifecycles are explicit: database handles (LadybugDB) close cleanly, no dangling process watchers, no leaked tree-sitter parsers.
- [ ] Long-running or large-graph paths remain bounded or are measurably streamed; degradation on large real repos is considered, not assumed benign.
### 2.7 Tests
- [ ] Tests cover the **real changed path** — they would fail if behavior, wiring, or contracts were broken, not only if a mock were misconfigured.
- [ ] Integration tests hit a real database where the production path does; do not introduce mocks that hide migration or schema drift.
- [ ] Assertions are meaningful. Use `toBe` / `toEqual` for exact expectations; avoid `toBeGreaterThanOrEqual` and other bounds-only assertions that mask regressions.
- [ ] Fixtures are realistic enough for the risk of the change — a one-file fixture is not sufficient for a pipeline-wide behavior change.
- [ ] New tests are deterministic and do not depend on network, clock, or host-specific paths without explicit isolation.
### 2.8 Observability and Operability
- [ ] Errors surfaced to users or callers are actionable: they name what failed, what input was involved (without leaking secrets), and how to recover where possible.
- [ ] Logging is proportionate — no noisy debug logs left in hot paths, no silent catches that swallow diagnostics.
- [ ] CLI exit codes and MCP tool responses are correct for each outcome (success, user error, internal error).
- [ ] Progress reporting (`PipelineProgress` and similar shared contracts) remains accurate after the change.
### 2.9 Reversibility and Risk
- [ ] The change has a clear rollback story: revert is safe, or migration is accompanied by a documented rollback / reindex procedure.
- [ ] Residual risks, compatibility impacts, and operational concerns are either resolved or **clearly stated** in the PR description.
- [ ] Destructive or hard-to-reverse operations (graph rebuild, schema change, `git` state manipulation) are opt-in or guarded.
## 3. Agent-Assisted Workflow Guardrails
When the change is produced with or reviewed by an AI agent, the following additional gates apply:
- [ ]**Scope match.** The final diff matches the intended symbols, files, and processes — no speculative refactors, unrelated formatting churn, or collateral edits outside the task scope.
- [ ]**Evidence-based edits.** Claims about repo state are verified against the current code, not trusted from memory or stale documentation.
- [ ]**Impact analysis.** Where GitNexus graph tooling is available and relevant, impact of non-trivial symbol, contract, or runtime-path changes is checked **before** editing.
- [ ]**Embeddings preserved.** If an indexed repo already has embeddings and re-analysis is required, embeddings are preserved — not accidentally dropped by a destructive reindex.
- [ ]**No false-done.** "Done" is claimed only after the Validation Baseline below has been run or any gap is explicitly named. Green tests on an unrelated path do not constitute validation.
Run the commands relevant to the touched area. If something cannot be run in the current environment, state it explicitly in the handoff.
### 4.1 Build ordering
- [ ]`gitnexus-shared/` dist is built before consuming packages are typechecked or tested (CI uses the `setup-gitnexus` action for this — local runs must match).
### 4.2 If `gitnexus/` changed
- [ ]`cd gitnexus && npx tsc --noEmit`
- [ ]`cd gitnexus && npm test`
- [ ]`cd gitnexus && npx prettier --check .` for files in the diff (pre-commit runs the affected-tests subset; do not expand scope)
### 4.3 If `gitnexus-web/` changed
- [ ]`cd gitnexus-web && npx tsc -b --noEmit`
- [ ]`cd gitnexus-web && npm test`
- [ ]`cd gitnexus-web && npm run test:e2e` when browser flows or user-facing UI behavior changed
### 4.4 If `gitnexus-shared/` changed
- [ ] Shared package builds cleanly (`npm run build` in `gitnexus-shared/`)
- [ ] Dependent packages still typecheck and test after the shared change — verify both CLI and web consumers together
### 4.5 If CI workflows or release pipelines changed
- [ ] The workflow passes a dry-run or triggered run before merge; concurrency (`cancel-in-progress`) and the `setup-gitnexus` action remain wired correctly.
- [ ]`CHANGELOG.md` is **not** edited here — it is owned by the release process.
## 5. Review Gates
A reviewer (human or agent) should be able to answer **yes** to each of the following before approving:
1.**Correctness** — Does the change do what it claims on the real runtime path?
2.**Readability** — Will the next contributor understand this in six months without asking?
3.**Architecture** — Is it in the right package, layer, and phase? Are boundaries respected?
4.**Security** — No new injection, leak, or trust-boundary violation?
5.**Performance** — No obvious regression on realistic inputs?
6.**Tests** — Would a regression in the changed behavior fail loudly?
7.**Scope** — Does the diff match the intended change, with no unrelated churn?
## 6. "Not Done" Signals
A change is **not** Done if any of the following is true, even if CI is green:
- The runtime path is not actually exercised by the tests.
- A contract drifted between `gitnexus/`, `gitnexus-web/`, and `gitnexus-shared/` and only one side was updated.
- A language-specific concern leaked into shared ingestion code.
- The diff contains unrelated reformatting, refactors, or cleanup beyond the stated task.
- Logs, comments, or TODOs were added as placeholders for work not done.
- The change depends on a manual step that is not documented.
-`CHANGELOG.md` was edited during PR work.
- Pre-commit, prettier, or typecheck was bypassed without explicit justification.
## 7. Task-Specific DoD Template
Use this in implementation and review prompts. Keep it short and tailor it to the actual change:
```md
# Definition of Done for this implementation
- [ ] Runtime wiring is complete for the affected path.
- [ ] Requested behavior is correct and relevant contracts are preserved or explicitly updated.
- [ ] The design stays scoped, readable, and proportionate to the task.
- [ ] Tests prove the changed behavior and catch broken wiring.
- [ ] Required validation for touched packages has been run, or any gap is explicitly noted.
- [ ] Repo boundaries, security, performance, and operational safety are respected.
- [ ] The diff contains only the intended change — no unrelated churn.
```
## 8. How to Use This File in Claude Review
Reference this file as the repo-wide completion bar. Add a task-specific review instruction such as:
```md
Review this change against `DoD.md` and the repo docs (`AGENTS.md`, `GUARDRAILS.md`,
`CONTRIBUTING.md`, `TESTING.md`, `ARCHITECTURE.md`). Treat `DoD.md` as the minimum
bar for production readiness. Flag anything that is partially wired, contract-unsafe,
under-tested, architecturally misplaced, scope-creeping, or harder to maintain than
necessary. Apply the five-axis review gate: correctness, readability, architecture,
security, performance.
```
## 9. Evolution
This DoD is living. Revisit it when:
- A class of incident slips past it (add a gate).
- A gate becomes consistently ceremonial without catching issues (remove or merge it).
- The architecture evolves in a way that changes what "done" means (update placement, validation, or contracts sections).
Track material updates in the changelog below. Keep the file tight — if it grows past a single read-in-one-sitting, something has drifted into the wrong place.
Rules for **human contributors** and **AI agents**. Complements `AGENTS.md` (workflows) and `CONTRIBUTING.md` (PR process).
## Scope (least privilege)
- **Read:** Source, tests, docs, public config as needed.
- **Write:** Only files required for the fix or feature; no unrelated formatting or refactors.
- **Execute:** Tests, typecheck, documented CLI commands. No destructive commands on user data without approval.
- **Off-limits:** Other people's machines, production deployments you don't own, credentials you lack permission to use.
Maintainer may widen scope per task.
---
## Non-negotiables
1.**Never commit secrets** — API keys, tokens, real `.env` values, private URLs, session cookies. Use `.env.example` with placeholders.
2.**Never rename with find-and-replace** in GitNexus-indexed projects — use `rename` MCP tool with `dry_run: true` first, review `graph` vs `text_search` edits. No separate `gitnexus rename` CLI exists.
3.**Run impact analysis before editing shared symbols** — `impact` (upstream) for functions/classes/methods others call. Do not ignore HIGH/CRITICAL without maintainer sign-off.
4.**Run `detect_changes` before commit** — confirm diffs map to expected symbols/processes when the graph is available.
5.**Preserve embeddings** — plain `npx gitnexus analyze` now preserves any embeddings recorded in `.gitnexus/meta.json` (the previous behavior wiped them). Use `--embeddings` to also generate vectors for new/changed nodes; use `--drop-embeddings` only when an explicit wipe is intended (e.g., model swap).
---
## Signs (recurring failure patterns)
Format: **Trigger → Instruction → Reason**. Append new Signs when the same mistake repeats.
### Stale graph after edits
- **Trigger:** MCP warns index is behind `HEAD`, or search doesn't match latest commit.
- **Do:** `npx gitnexus analyze` (plus `--embeddings` if used).
- **Why:** Tools query LadybugDB from last analyze; git changes are invisible until re-indexed.
### Embeddings vanished after analyze
- **Trigger:** Semantic search quality drops; `stats.embeddings` in `meta.json` is 0 after refresh.
- **Do:** Re-run `npx gitnexus analyze --embeddings` to regenerate. Check the analyze log for a `Warning: could not load cached embeddings` line — if present, the cache restore failed (corrupt DB / schema mismatch) and the rebuild had nothing to preserve. If you intentionally passed `--drop-embeddings`, this is expected.
- **Why:** Plain `analyze` preserves prior vectors by re-inserting them after the rebuild; the only ways to end up at zero are an explicit `--drop-embeddings`, a cache-load failure (now logged), or a model/dimension change that invalidates the cache.
### MCP lists no repos
- **Trigger:** MCP stderr says no indexed repos.
- **Do:** `npx gitnexus analyze` in the target repo; verify `npx gitnexus list` shows it.
- **Why:** MCP discovers repos via `~/.gitnexus/registry.json`, populated by analyze.
### Wrong repo in multi-repo setups
- **Trigger:** Query/impact results belong to another project.
- **Do:** Call `list_repos`, then pass `repo` on subsequent tools.
- **Why:** Default target is ambiguous when multiple repos are registered.
### LadybugDB lock / "database busy"
- **Trigger:** Errors opening `.gitnexus/lbug` while MCP and analyze both run.
- **Do:** Stop overlapping processes (one writer at a time). Retry analyze or restart MCP.
- **Why:** Embedded DB expects single-process ownership.
---
## Publishing & supply chain
- **npm:** Do not publish from unreviewed automation. Bump version intentionally; tag releases to match `package.json`.
- **Dependencies:** Minimal, auditable `package.json` changes; run tests and CI after lockfile updates.
- **License:** PolyForm Noncommercial 1.0.0 — do not relicense without maintainer approval.
---
## Escalation
Stop and ask a **human maintainer** when:
- Impact analysis shows HIGH/CRITICAL risk and the task still requires the change.
- You need to alter CI, release, or security-sensitive config.
- Requirements conflict (e.g. "speed up analyze" vs "must keep all embeddings on huge repo").
- You are unsure whether data loss is acceptable (`clean`, forced migrations, schema changes).
---
## Related docs
- [ARCHITECTURE.md](ARCHITECTURE.md) — components and data flow
- [RUNBOOK.md](RUNBOOK.md) — commands for recovery
- [CONTRIBUTING.md](CONTRIBUTING.md) — PR and commit expectations
⚠️ Important Notice:** GitNexus has NO official cryptocurrency, token, or coin. Any token/coin using the GitNexus name on Pump.fun or any other platform is **not affiliated with, endorsed by, or created by** this project or its maintainers. Do not purchase any cryptocurrency claiming association with GitNexus.
**⚠️ Important Notice:** GitNexus has NO official cryptocurrency, token, or coin. Any token/coin using the GitNexus name on Pump.fun or any other platform is **not affiliated with, endorsed by, or created by** this project or its maintainers. Do not purchase any cryptocurrency claiming association with GitNexus.
<div align="center">
@@ -9,7 +9,7 @@
<h2>Join the official Discord to discuss ideas, issues etc!</h2>
> *Like DeepWiki, but deeper.* DeepWiki helps you *understand* code. GitNexus lets you *analyze* it — because a knowledge graph tracks every relationship, not just descriptions.
**TL;DR:** The **Web UI** is a quick way to chat with any repo. The **CLI + MCP** is how you make your AI agent actually reliable — it gives Cursor, Claude Code, and friends a deep architectural view of your codebase so they stop missing dependencies, breaking call chains, and shipping blind edits. Even smaller models get full architectural clarity, making it compete with goliath models.
**TL;DR:** The **Web UI** is a quick way to chat with any repo. The **CLI + MCP** is how you make your AI agent actually reliable — it gives Cursor, Claude Code, Codex, and friends a deep architectural view of your codebase so they stop missing dependencies, breaking call chains, and shipping blind edits. Even smaller models get full architectural clarity, making it compete with Goliath models.
GitNexus is available as an **enterprise offering** - either as a fully managed **SaaS** or a **self-hosted** deployment. Also available for **commercial use** of the OSS version with proper licensing.
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that enrich searches with graph context + PostToolUse hooks that auto-reindex after commits.
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that enrich searches with graph context + PostToolUse hooks that detect a stale index after commits and prompt the agent to reindex.
gitnexus wiki [path]# Generate repository wiki from knowledge graph
gitnexus wiki --model <model> # Wiki with custom LLM model (default: gpt-4o-mini)
gitnexus wiki --base-url <url> # Wiki with custom LLM API base URL
# Repository groups (multi-repo / monorepo service tracking)
gitnexus group create <name> # Create a repository group
gitnexus group add <group> <groupPath> <registryName> # Add a repo to a group. <groupPath> is a hierarchy path (e.g. hr/hiring/backend); <registryName> is the repo's name from the registry (see `gitnexus list`)
gitnexus group remove <group> <groupPath> # Remove a repo from a group by its hierarchy path
gitnexus group list [name]# List groups, or show one group's config
gitnexus group sync <name> # Extract contracts and match across repos/services
gitnexus group contracts <name> # Inspect extracted contracts and cross-links
gitnexus group query <name> <q> # Search execution flows across all repos in a group
gitnexus group status <name> # Check staleness of repos in a group
```
If `analyze` reports a worker parse timeout on a large or unusual repository, it keeps running and falls back safely. To give slow worker jobs more time, use `gitnexus analyze --worker-timeout 60` or set `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=60000`. For very large files, `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` controls the worker job byte budget.
### What Your AI Agent Gets
**7 tools** exposed via MCP:
**16 tools** exposed via MCP (11 per-repo + 5 group):
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
| `cypher` | Raw Cypher graph queries | Optional |
| `group_list` | List configured repository groups | — |
| `group_sync` | Extract contracts and match across repos/services | — |
| `group_contracts`| Inspect extracted contracts and cross-links | — |
| `group_query` | Search execution flows across all repos in a group | — |
| `group_status` | Check staleness of repos in a group | — |
> When only one repo is indexed, the `repo` parameter is optional. With multiple repos, specify which one: `query({query: "auth", repo: "my-app"})`.
@@ -189,6 +268,10 @@ gitnexus wiki --base-url <url> # Wiki with custom LLM API base URL
- **Impact Analysis** — Analyze blast radius before changes
- **Refactoring** — Plan safe refactors using dependency mapping
**Repo-specific skills** generated with `--skills`:
When you run `gitnexus analyze --skills`, GitNexus detects the functional areas of your codebase (via Leiden community detection) and generates a `SKILL.md` file for each one under `.claude/skills/generated/`. Each skill describes a module's key files, entry points, execution flows, and cross-area connections — so your AI agent gets targeted context for the exact area of code you're working in. Skills are regenerated on each `--skills` run to stay current with the codebase.
---
## Multi-Repo MCP Architecture
@@ -217,8 +300,8 @@ flowchart TD
Server["server.ts"]
Backend["LocalBackend"]
Pool["Connection Pool"]
ConnA["KuzuDB conn A"]
ConnB["KuzuDB conn B"]
ConnA["LadybugDB conn A"]
ConnB["LadybugDB conn B"]
end
Setup -->|"writes global MCP config"| CursorConfig["~/.cursor/mcp.json"]
@@ -235,28 +318,189 @@ flowchart TD
ConnB -->|"queries"| RepoB
```
**How it works:** Each `gitnexus analyze` stores the index in `.gitnexus/` inside the repo (portable, gitignored) and registers a pointer in `~/.gitnexus/registry.json`. When an AI agent starts, the MCP server reads the registry and can serve any indexed repo. KuzuDB connections are opened lazily on first query and evicted after 5 minutes of inactivity (max 5 concurrent). If only one repo is indexed, the `repo` parameter is optional on all tools — agents don't need to change anything.
**How it works:** Each `gitnexus analyze` stores the index in `.gitnexus/` inside the repo (portable, gitignored) and registers a pointer in `~/.gitnexus/registry.json`. When an AI agent starts, the MCP server reads the registry and can serve any indexed repo. LadybugDB connections are opened lazily on first query and evicted after 5 minutes of inactivity (max 5 concurrent). If only one repo is indexed, the `repo` parameter is optional on all tools — agents don't need to change anything.
---
## Web UI (browser-based)
A fully client-side graph explorer and AI chat. No server, no install — your code never leaves the browser.
A client-side graph explorer and AI chat — your code never leaves your machine.
**Try it now:** [gitnexus.vercel.app](https://gitnexus.vercel.app) — drag & drop a ZIP and start exploring.
**Try it now:** [gitnexus.vercel.app](https://gitnexus.vercel.app) — run `npx gitnexus@latest serve` locally and the page auto-connects to your local backend.
cd gitnexus/gitnexus-shared && npm install && npm run build
cd ../gitnexus-web &&npm install
npm run dev
# Then in another terminal, start the backend the frontend connects to:
npx gitnexus@latest serve
```
The web UI uses the same indexing pipeline as the CLI but runs entirely in WebAssembly (Tree-sitter WASM, KuzuDB WASM, in-browser embeddings). It's great for quick exploration but limited by browser memory for larger repos.
## Docker
The official Docker setup ships **two signed images** orchestrated by `docker-compose.yaml`. Each image is published to both **GitHub Container Registry** (GHCR) and **Docker Hub** — same build, same digest, same Cosign signature — so pick whichever registry you prefer:
- [Dockerfile.web](Dockerfile.web) — builds `gitnexus-shared` and `gitnexus-web`, then serves the production frontend.
- [Dockerfile.cli](Dockerfile.cli) — builds the CLI/server (with its native deps) and runs `gitnexus serve --host 0.0.0.0`.
- [docker-compose.yaml](docker-compose.yaml) — starts both signed images side by side.
- [.env.example](.env.example) — overrides for image names, container names, ports, and the workspace mount.
The web UI uses the same indexing pipeline as the CLI but runs entirely in WebAssembly (Tree-sitter WASM, LadybugDB WASM, in-browser embeddings). It's great for quick exploration but limited by browser memory for larger repos.
**Local Backend Mode:** Run `gitnexus serve` and open the web UI locally — it auto-detects the server and shows all your indexed repos, with full AI chat support. No need to re-upload or re-index. The agent's tools (Cypher queries, search, code navigation) route through the backend HTTP API automatically.
@@ -264,7 +508,7 @@ The web UI uses the same indexing pipeline as the CLI but runs entirely in WebAs
## The Problem GitNexus Solves
Tools like **Cursor**, **Claude Code**, **Cline**, **Roo Code**, and **Windsurf** are powerful — but they don't truly know your codebase structure.
Tools like **Cursor**, **Claude Code**,**Codex**,**Cline**, **Roo Code**, and **Windsurf** are powerful — but they don't truly know your codebase structure.
**What happens:**
@@ -313,14 +557,31 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
1.**Structure** — Walks the file tree and maps folder/file relationships
2.**Parsing** — Extracts functions, classes, methods, and interfaces using Tree-sitter ASTs
3.**Resolution** — Resolves imports and function calls across files with language-aware logic
3.**Resolution** — Resolves imports, function calls, heritage, constructor inference, and `self`/`this` receiver types across files with language-aware logic
4.**Clustering** — Groups related symbols into functional communities
5.**Processes** — Traces execution flows from entry points through call chains
6.**Search** — Builds hybrid search indexes for fast retrieval
Short, copy-paste operations for **local development**, **MCP**, and **CI**. Commands assume a Unix shell; on Windows use Git Bash or equivalent paths.
- From repo root, install and build the CLI package:
```bash
cd gitnexus
npm install
npm run build
```
Use `npx gitnexus …` from any path after global/published install, or `node dist/cli/index.js …` when developing from `gitnexus/` with a local build.
---
## Index out of date / “stale” tools
**Symptom:** MCP or resources warn the index is behind `HEAD`, or results don’t reflect recent commits.
**Fix (from the target repo root):**
```bash
npx gitnexus analyze
```
**Force full rebuild** (same commit but suspect corruption or changed ignore rules):
```bash
npx gitnexus analyze --force
```
**Check status:**
```bash
npx gitnexus status
```
**List what MCP knows about:**
```bash
npx gitnexus list
```
---
## Embeddings
**First time with vectors** (slower, more disk/RAM):
```bash
npx gitnexus analyze --embeddings
```
**Important:** If you already had embeddings, **always** pass `--embeddings` on later analyzes, or they can be dropped. See `stats.embeddings` in `.gitnexus/meta.json` (0 means none).
**Large repos:** Analyze may skip or limit embedding work when node counts are very high; watch CLI output.
---
## MCP: no repos / empty tools
**Symptom:**`GitNexus: No indexed repos yet` on stderr when starting MCP.
**Fix:** In each project you want indexed:
```bash
cd /path/to/repo
npx gitnexus analyze
```
Restart the editor MCP session if needed. The server **refreshes the registry lazily**; new analyzes are picked up without necessarily reinstalling MCP.
**Symptom:** Wrong repo when multiple are indexed — pass `repo` on tools or use `list_repos` first.
---
## Clean slate (corrupt or huge `.gitnexus`)
**Current repo only** (prompts for confirmation):
```bash
npx gitnexus clean
```
**Skip confirmation:**
```bash
npx gitnexus clean --force
```
**All registered repos:**
```bash
npx gitnexus clean --all --force
```
Then re-run `npx gitnexus analyze` (and `--embeddings` if you need vectors).
---
## Local bridge for the web UI
```bash
cd gitnexus
npx gitnexus serve
# default http://127.0.0.1:4747 — see serve --help for port/host
```
Use when the browser UI should talk to **local** indexed repos instead of WASM-only mode.
**Note:** Pushes that touch only certain markdown paths may be skipped by `paths-ignore` in CI — see workflow file for exact patterns.
---
## Memory / analyze crashes
Analyze re-execs Node with a **large old-space heap** when needed (`analyze.ts`). If you still OOM on huge repos, close other processes, avoid `--embeddings` for a first pass, or analyze a smaller path if supported by your workflow.
---
## LadybugDB / lock errors
Only one process should open a repo’s `.gitnexus/lbug` store at a time. If MCP and a second `analyze` run conflict, stop one process, then retry `analyze` or restart MCP.
Tests do **not** run in the pre-commit hook — they run in CI (`ci-tests.yml`) only.
Skip with `git commit --no-verify` (use sparingly).
## Test categories
- **Unit** — Pure logic, parsers, graph/query helpers; fast; no network.
- **Integration** — Real combinations (filesystem, MCP wiring, larger pipelines) as already organized under `gitnexus/test/integration`.
- **Eval-style / golden sets** — For agent- or classification-style behavior, keep labeled inputs and expected outputs (JSON or table-driven tests) and run them in CI when relevant.
- **E2E (web)** — Critical user paths only; prefer `data-testid` attributes for stable selectors. Tests run against real backend (`gitnexus serve`) and Vite dev server.
## Performance metrics (targets)
Set targets to match team expectations, then tune to this repo’s CI reality:
- **`ci-e2e.yml`** — Playwright E2E tests, gated on `gitnexus-web/**` changes
Local checks before pushing:
```bash
cd gitnexus && npx tsc --noEmit && npm test
cd ../gitnexus-web && npx tsc -b --noEmit && npm test
```
Or rely on the pre-commit hook which runs these automatically for staged files.
## User acceptance / beta (optional)
For staged releases or UI betas: deploy to a staging environment, collect structured feedback, watch errors and latency, then iterate before a wider release.
GitNexus is a code intelligence tool that builds a knowledge graph from source code using tree-sitter AST parsing across 12 languages and KuzuDB for graph storage. Two packages: `gitnexus/` (CLI/MCP, TypeScript) and `gitnexus-web/` (browser).
- 12 language-specific type extractors in `gitnexus/src/core/ingestion/type-extractors/` must follow identical patterns for: async unwrapping, constructor binding, namespace handling, nullable type stripping, for-loop element typing.
- Past bugs: C#/Rust missing `await_expression` unwrapping that TypeScript handled correctly; PHP backslash namespace splitting inconsistent with other languages' `::` / `.` splitting.
- When reviewing type extractor changes, verify the same pattern exists in ALL applicable language files — asymmetry is the #1 source of bugs.
## Data Integrity (data-integrity-guardian)
- KuzuDB graph operations: schema in `gitnexus/src/core/kuzu/schema.ts`, adapter in `kuzu-adapter.ts`.
- The ingestion pipeline writes symbols and relationships to the graph — changes to node/relation schemas or the ingestion pipeline can corrupt the index.
- Known issue: KuzuDB `close()` hangs on Linux due to C++ destructor — use `detachKuzu()` pattern.
-`lbug-adapter.ts` fallback path needs quote/newline escaping for Cypher injection prevention.
## Security (security-sentinel)
- Cypher query construction in `lbug-adapter.ts` and `kuzu-adapter.ts` — watch for injection via unescaped user-provided symbol names.
- CLI accepts `--repo` parameter and file paths — validate against path traversal.
- MCP server exposes tools to external AI agents — all tool inputs are untrusted.
## Performance (performance-oracle)
- Tree-sitter buffer size is adaptive (512KB–32MB) via `getTreeSitterBufferSize()` in `constants.ts`.
- The ingestion pipeline processes entire repositories — O(n) per file with potential O(n²) in cross-file resolution.
- KuzuDB batch inserts vs individual inserts matter for large repos.
- Shared modules: `export-detection.ts`, `constants.ts`, `utils.ts` — changes here have wide blast radius.
-`gitnexus-web` package drifts behind CLI — flag if a change should be mirrored.
## Voltagent Supplementary Agents
Invoke these via the Agent tool alongside `/ce:review` for deeper specialist analysis. These cover gaps that compound-engineering agents don't:
### voltagent-lang:typescript-pro
**When:** Changes touch type-resolution logic, generics, conditional types, or complex type-level programming in `type-env.ts`, `type-extractors/*.ts`, or `types.ts`.
**Why:** The type resolution system uses advanced TypeScript patterns (discriminated unions, mapped types, recursive generics) that benefit from deep TS type-system review beyond what kieran-typescript-reviewer covers.
### voltagent-qa-sec:security-auditor
**When:** Changes touch MCP tool handlers, Cypher query construction, CLI argument parsing, or any code that processes external input.
**Why:** GitNexus is an MCP server — all tool inputs come from untrusted AI agents. Systematic OWASP-level audit catches injection vectors that spot-checking misses. Past finding: `lbug-adapter.ts` fallback path had unescaped newlines in Cypher queries.
### voltagent-data-ai:database-optimizer
**When:** Changes touch `kuzu-adapter.ts`, `schema.ts`, `lbug-adapter.ts`, or any Cypher query construction/execution.
**Why:** No CE agent specializes in graph database optimization. KuzuDB batch insert patterns, index usage, and query planning directly affect analysis speed on large repos.
## Review Tooling
- Use `gitnexus_impact()` before approving changes to any symbol — check d=1 (WILL BREAK) callers.
- Use `gitnexus_detect_changes({scope: "compare", base_ref: "main"})` to map PR diffs to affected execution flows.
- Use claude-mem to surface past architectural decisions relevant to the code under review.
GitNexus indexes COBOL codebases using a **regex-only extraction** strategy, bypassing tree-sitter entirely. This document explains why, how the pipeline works, and links to detailed sub-documents.
## Why Regex-Only?
The tree-sitter-cobol grammar (v0.0.1) has three critical limitations that make it unusable for production indexing:
| Issue | Impact | Severity |
|-------|--------|----------|
| External scanner hangs on ~5% of files | No timeout mechanism exists for the C scanner; the process blocks indefinitely | **Blocking** |
| Only ~15% of paragraph headers detected | Most procedure-division paragraphs are invisible to the grammar | High |
| Patch markers in cols 1-6 cause parse errors | Enterprise COBOL uses non-standard sequence area content (e.g., `mzADD`, `estero`, `#FIX`) | High |
Because the external scanner hang cannot be interrupted (there is no `setTimeoutMicros` equivalent for tree-sitter), using tree-sitter-cobol would hang the indexing pipeline on a non-trivial fraction of real-world files.
The regex-only approach provides:
- **Speed**: ~1ms per file average extraction time
- **Reliability**: zero hangs, zero crashes across 13,000+ files
- **Coverage**: captures all critical symbols -- program name, paragraphs, sections, CALL, PERFORM, COPY, data items (01-77, 88-level), file declarations, FD entries, EXEC SQL/CICS blocks, ENTRY points, and MOVE statements
## Architecture
```mermaid
flowchart TD
A[Repository Scan] --> B{File Detection}
B -->|Extension match| C[COBOL file]
B -->|GITNEXUS_COBOL_DIRS match| C
B -->|No match| Z[Skip]
C --> D{Copybook?}
D -->|Yes| E[Add to Copybook Map]
D -->|No| F[Source Program]
E --> G[COPY Expansion Engine]
F --> G
G -->|Inline copybook content| H[Expanded Source]
H --> I[Patch Marker Cleanup]
I --> J[Regex State Machine]
J --> K[Extracted Symbols]
K --> L[Graph Model Builder]
L --> M[Knowledge Graph]
subgraph "Per-Chunk Processing"
G
H
I
J
K
L
end
subgraph "Post-Processing"
M --> N[Community Detection]
M --> O[Process Detection]
M --> P[Contract Detection]
end
style J fill:#e8f5e9,stroke:#2e7d32
style G fill:#e3f2fd,stroke:#1565c0
```
## COBOL vs Tree-Sitter Languages
| Feature | COBOL (Regex) | Tree-Sitter Languages |
The COPY statement is COBOL's include mechanism -- analogous to `#include` in C or `import` in modern languages. GitNexus expands COPY statements **before** regex extraction so that symbols defined inside copybooks (data items, paragraphs, etc.) are visible in the program's extracted graph.
## Supported Syntax
### Basic COPY
```cobol
COPY CPSESP.
COPY "WORKGRID.CPY".
```
Inlines the content of the named copybook, replacing the COPY line(s).
Pipeline->>Pipeline: Replace file content with expanded content
end
```
The return type `CopyExpansionResult` contains `expandedContent` and `copyResolutions`. The `expansionDepth` field has been removed from the return type (it was unused by callers).
COPY statement line numbers in `CopyResolution` are 1-based (consistent with the preprocessor's line numbering). The splice operation that replaces COPY lines with expanded content adjusts for 0-based array indexing internally.
## Cycle Detection
Circular COPY references (e.g., copybook A includes copybook B which includes copybook A) are detected and handled:
1. Each expansion chain maintains a `visited` set of resolved copybook paths
2. If a copybook path is already in the visited set, the expansion is skipped
3. A `warnedCircular` set (internal to `expandCopies()`, not a parameter) deduplicates warning messages within a single file expansion
Known circular copybooks in PROJECT-NAME: `ANAZI`, `ANDIP`, `QDIPE` (self-referential includes).
## Max Depth
Nested COPY expansion is limited to **10 levels** (`DEFAULT_MAX_DEPTH`). If a COPY chain exceeds this depth, a warning is logged and the remaining COPY statements are left unexpanded.
## Max Total Expansions
A breadth amplification guard caps the total number of COPY expansions across all branches within a single file to **500** (`MAX_TOTAL_EXPANSIONS`). This prevents exponential blowup from diamond-shaped COPY graphs where N copybooks each include N other copybooks. Once the limit is reached, further COPY statements in that file are left unexpanded and a single warning is logged.
## REPLACING Application Detail
The REPLACING engine works by scanning all COBOL identifiers (matching `\b[A-Z][A-Z0-9-]*\b`) in the copybook content and applying each replacement rule:
```
Original copybook content:
05 ESP-NAME PIC X(30).
05 ESP-CODE PIC X(10).
05 KPSESPL-FLAG PIC X(01).
After REPLACING LEADING "ESP-" BY "LK-ESP-" LEADING "KPSESPL" BY "LK-KPSESPL":
05 LK-ESP-NAME PIC X(30).
05 LK-ESP-CODE PIC X(10).
05 LK-KPSESPL-FLAG PIC X(01).
```
For LEADING replacements, the engine checks if each identifier starts with the `from` prefix (case-insensitive) and replaces only the prefix portion, preserving the rest of the identifier.
For TRAILING replacements, the same logic applies to suffixes.
For EXACT replacements, only identifiers that match the `from` value exactly (case-insensitive) are replaced.
## Copybook Resolution
The resolver tries multiple strategies to match a COPY target name to a copybook file:
1.**Exact match**: `COPY CPSESP` resolves to copybook named `CPSESP`
2.**Strip extension**: `COPY WORKGRID.CPY` strips `.CPY` and resolves to `WORKGRID`
3.**Add extension**: `COPY CPSESP` tries `CPSESP.CPY` and `CPSESP.COPY`
If no match is found, the COPY statement is left in place (unexpanded) and a resolution record with `resolvedPath: null` is created.
## Pipeline Integration
The expansion runs **per chunk**, after file content is read but before dispatch to worker threads:
1. All copybook files are read upfront (they are typically small, collectively under 100MB)
2. Per chunk, the copybook map is merged with chunk content (in case a chunk contains copybooks)
3. Only programs (not copybooks themselves) undergo expansion
4. The expanded content replaces the original content in-place before worker dispatch
## Inline Comment Handling
The copy expander's `stripInlineComment()` helper is quote-aware: pipe characters (`|`) inside single- or double-quoted strings are preserved. This matches the same quote-aware logic used by the preprocessor.
The stack maintains items where each entry's level is strictly less than the next. When a new item arrives with a level <= the top of stack, items are popped until the stack top has a smaller level. A `CONTAINS` edge is created from the stack top to the new item.
For 88-level condition names, the parent is the immediately preceding non-88 data item (found by scanning backwards).
A maximum of **500 data items per file** (`MAX_DATA_ITEMS_PER_FILE`) are processed. Some COBOL programs (especially after COPY expansion) can have 10,000+ data items, which would cause graph bloat and push the V8 relationship Map past its 16.7M entry limit across thousands of files.
The cap applies after extraction: the first 500 items in source order are kept. Since 01-level records appear first, critical top-level structure is preserved.
## EXEC SQL
EXEC SQL blocks are accumulated across lines between `EXEC SQL` and `END-EXEC`, then parsed as a unit.
### Operation Classification
The first SQL keyword determines the operation:
| First Keyword | Operation |
|---------------|-----------|
| `SELECT` | SELECT |
| `INSERT` | INSERT |
| `UPDATE` | UPDATE |
| `DELETE` | DELETE |
| `DECLARE` | DECLARE |
| `OPEN` | OPEN |
| `CLOSE` | CLOSE |
| `FETCH` | FETCH |
| *(anything else)* | OTHER |
### Table Extraction
Tables are extracted from SQL clauses:
| Clause Pattern | Example |
|----------------|---------|
| `FROM <table>` | `SELECT * FROM EMPLOYEES` |
| `INSERT INTO <table>` | `INSERT INTO EMPLOYEES` |
| `UPDATE <table>` | `UPDATE EMPLOYEES SET ...` |
| `JOIN <table>` | `LEFT JOIN DEPARTMENTS ON ...` |
Note: The `INTO` pattern is restricted to `INSERT INTO` to avoid false positives from `FETCH ... INTO :host-var` and `SELECT ... INTO :host-var` statements, where `INTO` introduces host variables rather than table names.
This distinction allows queries to find bulk field-by-field moves separately from simple variable assignments.
## GO TO DEPENDING ON
The `GO TO` statement with multiple targets and a `DEPENDING ON` clause is a computed branch:
```cobol
GOTOPARA-1PARA-2PARA-3
DEPENDINGONWK-SELECTOR.
```
All target paragraph names are extracted and emitted as separate `gotos` entries. Each target produces a `CALLS` edge in the graph (same semantics as PERFORM). The `DEPENDING ON` variable is not currently tracked as a data-flow dependency.
## SORT INPUT/OUTPUT PROCEDURE
SORT and MERGE statements can specify procedural entry points instead of file-based I/O:
```cobol
SORTSORT-FILEONASCENDINGKEYSORT-KEY
INPUTPROCEDUREISPREPARE-INPUT
OUTPUTPROCEDUREISFORMAT-OUTPUT.
```
`INPUT PROCEDURE IS` and `OUTPUT PROCEDURE IS` targets are extracted as control-flow targets (same as PERFORM). They produce `performs` entries and corresponding `CALLS` edges in the graph.
## Fixed-Format Literal Continuation
In fixed-format COBOL, string literals can span multiple lines using the continuation indicator (`-` in column 7). When a continuation line starts with a quote character, the extractor joins it with the predecessor by removing the trailing quote from the previous line and the opening quote from the continuation:
```
Line N: MOVE "THIS IS A LONG STRI
Line N+1 (cont): - "NG VALUE" TO WK-FIELD.
Merged: MOVE "THIS IS A LONG STRING VALUE" TO WK-FIELD.
```
The trailing `"` on line N and the opening `"` on line N+1 are both removed, producing a seamless literal. If no matching quote is found on the predecessor line, the continuation is appended as-is.
## Source Files
-`gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- All extraction logic, clause parsers, EXEC block parsers
GitNexus detects COBOL files through two mechanisms: extension-based mapping and directory-based override for extensionless files. This document covers both, plus the copybook/program classification logic.
| `.open` / `.OPEN` | Copybook | File OPEN fragment |
| `.close` / `.CLOSE` | Copybook | File CLOSE fragment |
| `.ini` / `.INI` | Copybook | Initialization fragment |
| `.def` / `.DEF` | Copybook | Definition fragment |
All extension matching is case-sensitive in `getLanguageFromFilename` (the extensions above are matched as written, including uppercase variants like `.GNM`).
Many enterprise COBOL repositories use extensionless files -- the filename alone identifies the program (e.g., `s/BGTABFL` is the source for program `BGTABFL`). GitNexus handles this via the `GITNEXUS_COBOL_DIRS` environment variable.
### Configuration
Set `GITNEXUS_COBOL_DIRS` to a comma-separated list of directory names:
```bash
# Files in s/, c/, and wfproc/ directories (at any depth) are treated as COBOL
exportGITNEXUS_COBOL_DIRS=s,c,wfproc
```
The matching is **case-insensitive** and checks all path segments:
The path segment check splits the full path on `/` and tests each segment against the cached set.
## Copybook vs Program Classification
After a file is identified as COBOL, it must be classified as either a **program** (to be parsed for symbols) or a **copybook** (to be loaded into the copybook map for COPY expansion).
### Classification Rules
A COBOL file is classified as a **copybook** if ANY of these conditions is true:
1. It has a recognized copybook extension (`.cpy`, `.copy`, `.gnm`, `.fd`, `.wrk`, `.sel`, `.open`, `.close`, `.ini`, `.def`)
2. It is an extensionless file whose path contains a directory segment matching one of: `c`, `copy`, `copybooks`, `copylib`, `cpy`
A file is classified as a **program** if:
1. It has a program extension (`.cbl`, `.cob`, `.cobol`), OR
2. It is extensionless and does NOT match any copybook directory pattern
### Copybook Name Resolution
Copybook names are derived from the filename:
- Strip the extension (if any)
- Convert to uppercase
Examples:
-`c/CPSESP` -- name: `CPSESP`
-`copy/workgrid.cpy` -- name: `WORKGRID`
-`c/ANAZI.GNM` -- name: `ANAZI`
This name is used to resolve `COPY CPSESP.` statements during expansion.
This document describes the graph nodes and edges that GitNexus creates for COBOL codebases. The COBOL graph model is richer than most tree-sitter languages because it captures domain-specific constructs: file declarations, FD entries, data hierarchies, SQL tables, CICS maps, and cross-program contracts.
## Entity-Relationship Diagram
```mermaid
erDiagram
File ||--o{ Module : DEFINES
File ||--o{ Function : DEFINES
File ||--o{ Namespace : DEFINES
File ||--o{ Record : DEFINES
File ||--o{ Property : DEFINES
File ||--o{ Const : DEFINES
File ||--o{ CodeElement : DEFINES
File ||--o{ Constructor : DEFINES
File }o--o{ File : IMPORTS
Module ||--o{ Record : CONTAINS
Module ||--o{ Constructor : CONTAINS
Module }o--o{ CodeElement : ACCESSES
Module }o--o{ Module : CALLS
Module }o--o{ Module : CONTRACTS
Module }o--o{ Property : RECEIVES
Record ||--o{ Property : CONTAINS
Record ||--o{ Const : CONTAINS
Record }o--o{ Record : REDEFINES
Property ||--o{ Property : CONTAINS
Property ||--o{ Const : CONTAINS
Property }o--o{ Property : REDEFINES
Property }o--o{ CodeElement : RECORD_KEY_OF
Property }o--o{ CodeElement : FILE_STATUS_OF
CodeElement ||--o{ CodeElement : CONTAINS
CodeElement ||--o{ Record : CONTAINS
Function }o--o{ Function : CALLS
```
## Node Types
| Node Type | COBOL Concept | Created From | Example |
The worker pool splits each worker's chunk into sub-batches to bound peak memory per `postMessage` serialization. COBOL repos use a smaller sub-batch size than the default:
| Per sub-batch timeout | 120s | 120s (configurable) |
**Why 200?** COBOL regex extraction + preprocessing takes ~1ms per file on average, but with COPY expansion and deep indexing the effective time is ~150ms per file. At sub-batch size 1500, that would be ~225s per sub-batch, exceeding the 120s timeout.
COBOL mode is activated automatically when `GITNEXUS_COBOL_DIRS` is set:
Workers default to `min(8, cpus - 1)`. For COBOL repos, this is usually sufficient since regex extraction is CPU-bound but fast. The bottleneck is typically KuzuDB write, not extraction.
For COBOL-only repos, worker startup is faster because tree-sitter native modules are loaded lazily (skipped entirely if only COBOL files are present).
## Data Item Cap
### Configuration
```typescript
constMAX_DATA_ITEMS_PER_FILE=500;
```
This constant appears in both `parse-worker.ts` (worker path) and `parsing-processor.ts` (sequential fallback).
### Rationale
Some COBOL programs, especially after COPY expansion, can have 10,000+ data items. At that scale:
- The in-memory relationship Map (for CONTAINS, REDEFINES, etc.) approaches the V8 16.7M entry limit across thousands of files
- KuzuDB write time increases linearly with edge count
- Most deep-nested items (level 20+) are rarely queried individually
### Impact
The cap truncates data items beyond the 500th in source order. Since 01-level Records appear first in COBOL source, the cap preserves:
- All 01-level record definitions
- The most important 02-49 level items (those closest to the record root)
- 88-level conditions associated with early items
To increase the cap for specific needs, modify the `MAX_DATA_ITEMS_PER_FILE` constant in both files.
## Memory Management
### COPY Expansion Breadth Guard
A per-file `MAX_TOTAL_EXPANSIONS = 500` limit prevents exponential blowup from diamond-shaped COPY graphs (e.g., N copybooks each containing N COPY statements). Once the limit is reached, further COPY statements in that file are left unexpanded. See [copy-expansion.md](copy-expansion.md) for details.
### COPY Expansion Memory
All copybook content is loaded upfront into a Map before chunk processing begins. For PROJECT-NAME:
- 2,976 copybooks, typically under 100MB total
- The Map is shared (read-only) across chunk iterations
- Per-chunk, the copybook map is merged with chunk file content (in case a chunk contains copybooks not in the pre-loaded set)
- After all chunks are processed, the copybook map is freed (`cobolCopybookContents = undefined`)
### Chunk Budget
Source files are grouped into chunks of max 20MB (`CHUNK_BYTE_BUDGET`). Each chunk's lifecycle:
This ensures only ~20MB of source + ~200-400MB of working memory (ASTs, extracted records, serialization) is active at any time.
### Shared Warning Deduplication
The `warnedCircular` set (used by the COPY expansion engine) is shared across all files in a chunk. This prevents the same circular copybook warning (e.g., `ANAZI includes itself`) from being logged thousands of times.
| tree-sitter-cobol hangs on ~5% of files | Cannot use tree-sitter for COBOL | Regex-only extraction (current approach) |
| Data item cap (500/file) | May miss deeply nested items in large programs | Increase `MAX_DATA_ITEMS_PER_FILE` in source |
| Circular copybooks (ANAZI, ANDIP, QDIPE) | Self-referential includes cannot be expanded | Detected and skipped with warning |
| wfproc/ files may not be pure COBOL | Workflow files may produce extraction noise | Exclude `wfproc` from `GITNEXUS_COBOL_DIRS` if problematic |
| No MOVE DATA_FLOW edges yet | Data flow between variables not in graph | Reserved for future release |
| Continuation line handling | Some complex multi-line continuations (especially in string literals spanning 3+ lines) may not merge correctly | Known edge case; affects <0.1% of lines |
| Single-line EXEC blocks | `EXEC SQL SELECT ... END-EXEC` on one line is handled, but pathological nesting is not | Extremely rare in practice |
| Extension case sensitivity | `.GNM` and `.gnm` are matched differently | Use the exact case from the codebase |
## Troubleshooting
### "COPY expansion failed"
```
[pipeline] COPY expansion failed for s/BGTABFL: Cannot read properties of null
```
**Cause:** A copybook referenced by a COPY statement cannot be found.
**Fix:**
1. Verify `GITNEXUS_COBOL_DIRS` includes the directory containing copybooks (typically `c`)
2. Check that copybook filenames match the COPY target (case-insensitive, after stripping extensions)
3. Ensure copybook files are not in `.gitignore`
### Worker sub-batch timeout
```
Worker 3 sub-batch timed out after 120s (chunk: 200 items)
```
**Cause:** A sub-batch took longer than the timeout. Typically happens when one file is extremely large (50,000+ lines after COPY expansion).
For very large repos (>500MB source), consider `--max-old-space-size=32768`.
### Concurrent analyze corruption
**Rule:** Only ONE `gitnexus analyze` process should run at a time per repository. Concurrent writes to KuzuDB corrupt the database.
If corruption occurs:
```bash
# Remove the KuzuDB directory and re-index
rm -rf .gitnexus/kuzu
gitnexus analyze --force
```
### Slow KuzuDB write phase
The KuzuDB write phase (132s for PROJECT-NAME) is the bottleneck for large COBOL repos. This is proportional to the number of nodes and edges being written. Reducing `MAX_DATA_ITEMS_PER_FILE` or excluding non-essential directories from `GITNEXUS_COBOL_DIRS` can help.
### Verbose output
Enable verbose logging to see per-phase timing and statistics:
The `extractCobolSymbolsWithRegex()` function in `cobol-preprocessor.ts` performs single-pass, state-machine-driven extraction of all COBOL symbols. This document describes the state machine, line processing flow, and every regex pattern used.
## State Machine: Division Tracking
The extractor tracks which COBOL division is currently being processed. Division transitions are detected by the `RE_DIVISION` pattern.
Within the ENVIRONMENT DIVISION, the `currentEnvSection` tracks whether we are in `INPUT-OUTPUT` or `CONFIGURATION` section. SELECT statement accumulation only occurs in `INPUT-OUTPUT`.
## Line Processing Flow
Each raw source line goes through this pipeline:
```
Raw line
|
v
Length < 7? ---------> Skip (flush pending if any)
|
v
Indicator col 7
|
+-- '*' or '/' -----> Comment: skip entirely
|
+-- '-' ------------> Continuation: append to pending line
|
+-- other ----------> Normal: flush pending, strip inline comments (|),
buffer as new pending logical line
```
After all lines are processed, the final pending line is flushed, along with any accumulated SELECT statement, SORT/MERGE accumulator, and any open EXEC block (truncated file without `END-EXEC`).
### Inline Comment Stripping
Enterprise COBOL (particularly Italian dialect) uses the pipe character `|` as an inline comment marker. The `stripInlineComment()` helper is **quote-aware**: it tracks whether the scan position is inside a single- or double-quoted string and only treats `|` as a comment marker when outside quotes. Pipe characters inside string literals are preserved.
Free-format `*>` inline comment stripping uses the same quote-aware approach: the scanner walks character by character, toggling quote state, and only recognizes `*>` as a comment marker when not inside a quoted string.
### Patch Marker Handling
The `preprocessCobolSource()` function (run before extraction in the worker) replaces non-standard content in columns 1-6. Standard COBOL expects spaces or digit sequence numbers in this area. If any letter or `#` character is found, the entire sequence area is replaced with 6 spaces:
```
Before: mzADD MOVE WK-AMT TO WK-TOTAL
After: MOVE WK-AMT TO WK-TOTAL
```
This preserves exact line count for position mapping.
## Regex Pattern Reference
All patterns are compiled once as module-level constants and reused across calls.
SELECT statements are accumulated across multiple lines until a period terminator is found, then parsed for ASSIGN, ORGANIZATION, ACCESS, RECORD KEY, and FILE STATUS clauses.
| `RE_PROC_USING` | `\bPROCEDURE\s+DIVISION\s+USING\s+([\s\S]*?)(?:\.\|$)` | USING parameters | `PROCEDURE DIVISION USING WK-PARAM` |
| `RE_ENTRY` | `\bENTRY\s+"([^"]+)"(?:\s+USING\s+([\s\S]*?))?(?:\.\|$)` | ENTRY point | `ENTRY "SUBPROG" USING WK-DATA` |
| `RE_MOVE` | `\bMOVE\s+((?:CORRESPONDING\|CORR)\s+)?([A-Z][A-Z0-9-]+)\s+TO\s+(.+)` | MOVE statement (supports CORR abbreviation and multi-target) | `MOVE WK-NAME TO OUT-NAME`, `MOVE CORR WK-IN TO WK-OUT` |
The USING parameter list (`RE_PROC_USING`) is split on `\bRETURNING\b` before tokenization -- any RETURNING clause and everything after it is excluded from the parameter list (`.split(/\bRETURNING\b/i)[0]`).
Note: `RE_PROC_SECTION` and `RE_PROC_PARAGRAPH` require exactly 7 spaces of leading indentation (COBOL area A starting at column 8). This is the standard COBOL paragraph indentation.
### All-Division Patterns
These patterns are checked regardless of current division:
| `SORT_CLAUSE_NOISE` | Set of SORT/MERGE clause keywords filtered from USING/GIVING file lists: `ON`, `ASCENDING`, `DESCENDING`, `KEY`, `WITH`, `DUPLICATES`, `IN`, `ORDER`, `COLLATING`, `SEQUENCE`, `IS`, `THROUGH`, `THRU`, `INPUT`, `OUTPUT`, `PROCEDURE` |
SORT and MERGE statements are accumulated across multiple lines (like SELECT) until a period terminator is found, then parsed for USING/GIVING file lists and INPUT/OUTPUT PROCEDURE targets. The `flushSort()` helper encapsulates the flush-and-parse logic, mirroring the existing `flushSelect()` pattern. Both helpers are called at EOF to handle truncated files.
### GO TO Multi-Target
`RE_GOTO` captures all paragraph names in a `GO TO` statement, including the multi-target form `GO TO p1 p2 p3 DEPENDING ON x`. The captured group contains all target names (space-separated), which are split into individual targets. Each target produces a separate `gotos` entry.
### PROGRAM-ID Detection
PROGRAM-ID is detected regardless of the current division state. This handles sibling programs that appear after `END PROGRAM` and omit the `IDENTIFICATION DIVISION` header -- the extractor will still capture the PROGRAM-ID and push a new program boundary.
| `RE_END_EXEC` | `\bEND-EXEC\b` | End of EXEC block | `END-EXEC` |
EXEC blocks accumulate all lines between `EXEC SQL/CICS` and `END-EXEC`, then delegate to `parseExecSqlBlock()` or `parseExecCicsBlock()` for detailed extraction.
## Excluded Paragraph Names
The following names are excluded from paragraph detection to avoid false positives from division/section headers:
```
DECLARATIVES, END, PROCEDURE, IDENTIFICATION,
ENVIRONMENT, DATA, WORKING-STORAGE, LINKAGE,
FILE, LOCAL-STORAGE, COMMUNICATION, REPORT,
SCREEN, INPUT-OUTPUT, CONFIGURATION
```
Additionally, paragraph candidates containing `DIVISION` or `SECTION` as substrings are excluded.
## MOVE Skip List (Figurative Constants)
MOVE statements where the source is a figurative constant are skipped:
```
SPACES, ZEROS, ZEROES, LOW-VALUES, LOW-VALUE,
HIGH-VALUES, HIGH-VALUE, QUOTES, QUOTE, ALL
```
## Source Files
-`gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- `preprocessCobolSource()`, `extractCobolSymbolsWithRegex()`, all regex constants
This guide is for teams whose product lives in **several separate Git repositories** — one per service — and whose services talk to each other over **gRPC** (possibly alongside HTTP and message topics). GitNexus indexes each repo independently, then a _group_ stitches the per-repo indexes into a single cross-repo view that the `impact`, `query`, and `context` tools can traverse. If your services live in one monorepo, much of this still applies — set each service as a member of a group and use the `service` prefix to scope queries — but the walkthrough assumes the harder multi-repo case.
## Mental model
- Each repository has its own `.gitnexus/` index (a LadybugDB graph of symbols, relationships, processes). `gitnexus analyze` in each repo produces that index completely independently.
- A **group** is a higher-level construct stored at `~/.gitnexus/groups/<group>/` that references the per-repo indexes by their registry name.
- Sync-time extractors walk each member repo and emit **contracts** — provider or consumer records keyed by a canonical `contractId` (`grpc::auth.AuthService/Login`, `http::GET::/orders`, etc.).
- The sync step matches providers and consumers that share a `contractId` and writes **cross-links** to `<groupDir>/contracts.json`. Those cross-links are what lets `impact({repo: "@<group>", target: "X"})` hop from one repo into another.
- Contracts come from three places: automatic contract extractors (`grpc-extractor`, `http-route-extractor`, `topic-extractor`), a manifest escape hatch (`config.links` in `group.yaml`), and — for same-name symbol matches where no contract is declared — the exact-match matching cascade in [`matching.ts`](../../gitnexus/src/core/group/matching.ts).
- Each repo stays editable and re-indexable on its own. Re-run `gitnexus analyze` in a repo when it changes, then `gitnexus group sync <group>` to refresh `contracts.json`. `gitnexus group status` reports which members are stale.
## Prerequisites
- GitNexus installed and runnable as `gitnexus` or `npx gitnexus` (see the root [README.md](../../README.md)).
- Each service repository checked out locally. No requirement that they share a parent directory — the group references them by registry name.
- Write access to `~/.gitnexus/` (the default gitnexus home; see `getDefaultGitnexusDir` in [`storage.ts`](../../gitnexus/src/core/group/storage.ts)).
## Step-by-step walkthrough
The example uses three services — a TypeScript API gateway, a Go orders service, and a Python inventory service — with gRPC between them. The gateway is an `orders` consumer; the orders service is both an `orders` provider and an `inventory` consumer; the inventory service is an `inventory` provider.
### 1. Index each repository
Run `analyze` from inside each service repo (or pass the path). The CLI surface lives in [`gitnexus/src/cli/analyze.ts`](../../gitnexus/src/cli/analyze.ts) and is wired in [`gitnexus/src/cli/index.ts`](../../gitnexus/src/cli/index.ts).
```bash
cd ~/code/gateway && npx gitnexus analyze
cd ~/code/orders && npx gitnexus analyze
cd ~/code/inventory && npx gitnexus analyze
```
Useful flags:
-`--force` — reindex even if up to date.
-`--embeddings` — generate embedding vectors (needed only if you want semantic search; the exact-match cross-repo cascade does **not** need them).
-`--name <alias>` — register the repo under a specific alias when two repos share a basename (e.g. two `api/` folders).
-`--skip-git` — index a checkout that isn't a git repo.
Each run writes a `.gitnexus/` folder in the repo and registers the repo in `~/.gitnexus/registry.json`. Confirm with `npx gitnexus list`.
### 2. Author `group.yaml`
Create the group directory and edit the config. Either use the CLI scaffolder or write the file directly — both produce the same shape consumed by [`config-parser.ts`](../../gitnexus/src/core/group/config-parser.ts).
Field notes (schema in [`types.ts`](../../gitnexus/src/core/group/types.ts)):
-`version` — must be `1`. The parser rejects anything else.
-`name` — required; used for the group directory name and all CLI / MCP calls.
-`repos` — a mapping from **group path** (a logical name you choose; can be a hierarchy like `backend/orders`) to **registry name** (the name shown by `npx gitnexus list`). Both sides appear throughout the tooling: contract rows use the group path; `@<group>/<groupPath>` routes tools to a single member.
-`links` — optional manifest escape hatch, one entry per explicit cross-repo contract. Validated by the parser: `from` and `to` must be known repo paths, `type` must be one of `http | grpc | topic | lib | custom`, and `role` must be `provider | consumer`.
-`detect` — toggles per extractor family. Defaults (set in `config-parser.ts`) turn `http`, `grpc`, `topics`, and `shared_libs` on; disable the ones you don't use to speed up sync.
-`matching` — thresholds for the matching cascade. The exact match is always run; other strategies depend on indexer state. Two optional fields reduce false-positive cross-links in large groups:
-`exclude_links_paths` — list of HTTP paths to exclude from cross-link matching (default `[]`). Contracts at these paths are still extracted and visible in the registry, but they don't produce cross-repo links. Useful for health-check endpoints (`/ping`, `/health`) that every service exposes. Trailing slashes are normalized.
-`exclude_links_param_only_paths` — when `true`, exclude routes where every segment is `{param}` (e.g. `/{param}`, `/{param}/{param}`) from cross-link matching (default `false`). Mixed routes like `/users/{param}` are not affected.
### 3. Sync the group
```bash
npx gitnexus group sync payments-platform --verbose
```
What this does (see [`sync.ts`](../../gitnexus/src/core/group/sync.ts)):
1. Opens each member's per-repo LadybugDB.
2. Runs the HTTP, gRPC, and topic extractors against the source files.
3. Applies manifest `links` through [`manifest-extractor.ts`](../../gitnexus/src/core/group/extractors/manifest-extractor.ts).
4. Runs the exact-match cascade, joining providers and consumers that share a normalized `contractId`.
5. Writes `contracts.json` in the group directory.
Flags:
-`--exact-only` — stop after the exact cascade; skip BM25 and embedding fallback.
-`--skip-embeddings` — run exact plus BM25 but not embedding-based matching.
-`--allow-stale` — don't warn if a member's index is stale.
-`--json` — machine-readable output.
The same operation is available over MCP as `group_sync({ name: "payments-platform" })` — see [`tools.ts`](../../gitnexus/src/mcp/tools.ts).
### 4. Inspect the registry
Use `gitnexus group contracts` for the CLI view or read the `gitnexus://group/<name>/contracts` MCP resource for the same data.
```bash
npx gitnexus group contracts payments-platform --type grpc --json
Staleness of the underlying indexes shows up in `npx gitnexus group status payments-platform` or the `gitnexus://group/<name>/status` resource.
### 5. Run cross-repo impact with `@<group>` routing
From any shell (you do **not** have to `cd` into a member repo), the normal `impact` / `query` / `context` tools accept `repo: "@<group>"` to fan out across all members, or `repo: "@<group>/<memberPath>"` to target one member. Routing is implemented in [`resolve-at-member.ts`](../../gitnexus/src/core/group/resolve-at-member.ts) and described in [`tools.ts`](../../gitnexus/src/mcp/tools.ts).
Phase 1 walks within the anchor member; Phase 2 hops across the Contract Bridge wherever a cross-link endpoint matches an impacted symbol. See [`cross-impact.ts`](../../gitnexus/src/core/group/cross-impact.ts) for the bridge query.
## How gRPC extraction works
`GrpcExtractor` ([`grpc-extractor.ts`](../../gitnexus/src/core/group/extractors/grpc-extractor.ts)) runs two passes per member repo:
1.**Proto map.** Every `**/*.proto` file is parsed to enumerate `service Foo { rpc Bar(...) }` blocks and (transitively) resolve the package name. Each RPC method becomes a provider contract with `contractId = grpc::<package>.<Service>/<Method>` and `confidence = 0.85`. Parsing uses the vendored `tree-sitter-proto` grammar when available and falls back to a length-preserving manual parser (`extractServiceBlocks`) otherwise, so `.proto` extraction works on platforms where the grammar fails to build.
2.**Source scan.** Every source file whose extension matches [`GRPC_SCAN_GLOB`](../../gitnexus/src/core/group/extractors/grpc-patterns/index.ts) is parsed by its language plugin:
| Language | Provider signal | Consumer signal |
|----------|-----------------|-----------------|
| Go ([`go.ts`](../../gitnexus/src/core/group/extractors/grpc-patterns/go.ts)) | `pb.RegisterXxxServer(...)`, `pb.UnimplementedXxxServer` embedded in struct | `pb.NewXxxClient(conn)` |
| Java ([`java.ts`](../../gitnexus/src/core/group/extractors/grpc-patterns/java.ts)) | `extends XxxServiceGrpc.XxxServiceImplBase` (with or without `@GrpcService`) | `XxxServiceGrpc.newBlockingStub(...)`, `newStub(...)` |
| Node / TS ([`node.ts`](../../gitnexus/src/core/group/extractors/grpc-patterns/node.ts)) | NestJS `@GrpcMethod('Service','Method')` | `@GrpcClient` field typed `XxxServiceClient`, `client.getService<X>('Service')`, `new XxxServiceClient(...)`, `new foo.bar.XxxService(...)` in files that call `loadPackageDefinition` |
For each source-scan detection the extractor looks up the short service name in the proto map and picks:
-`grpc::<package>.<Service>/<Method>` when a method is named and the service resolves against the proto map,
-`grpc::<package>.<Service>/*` (wildcard) when only the service is known, or
-`grpc::<ServiceName>/*` when no `.proto` is available at all.
Provider detections land at confidence 0.8 (with proto) or 0.65 (without); consumers at 0.75 or 0.55. NestJS `@GrpcMethod` is fixed at 0.8 because the decorator is self-describing.
### Matching
`matching.ts` lowercases the package/service segment before comparing contract ids, so bindings that capitalize names differently (`auth.AuthService` vs `auth.authservice`) still match. Method names are compared case-sensitively because gRPC's wire path is case-sensitive. Service-only wildcards (`grpc::pkg.Svc/*`) match any method on the same service during cross-linking.
### Known limitations
- **Ambiguous proto resolution.** If a short service name exists in more than one `.proto` file and the source-scan hit can't be narrowed down by shared directory segments (`resolveProtoConflict` refuses to guess), the extractor skips contract emission and logs a warning.
- **Proto packages must be resolvable locally.** Transitive imports that point outside the repo produce an empty package segment, which means the contract id collapses to `grpc::<Service>/<Method>`. Cross-repo matches still work as long as both sides agree on the empty package.
- **Rewrite rules are not implemented.** If the provider repo writes `grpc::orders.OrderService/PlaceOrder` and the consumer repo writes `grpc::orderspb.OrderService/PlaceOrder`, they won't cross-link automatically. Use `config.links` to declare the correspondence (see below).
- **One sync = one snapshot.** Contracts are extracted against the indexed snapshot of each repo. Re-index first, then re-sync; the `status` command and resource surface staleness.
## When automatic extraction isn't enough
The escape hatch is the `links` list in `group.yaml`, handled by [`ManifestExtractor`](../../gitnexus/src/core/group/extractors/manifest-extractor.ts). Each entry is a **one-directional** provider/consumer declaration:
```yaml
version:1
name:payments-platform
repos:
gateway:gateway
orders:orders
inventory:inventory
links:
# Explicit gRPC method: use when naming mismatches stop the
# automatic matcher from cross-linking.
- from:gateway
to:orders
type:grpc
contract:OrderService/PlaceOrder
role:consumer
# Service-level link when you don't want to enumerate methods.
- from:orders
to:inventory
type:grpc
contract:InventoryService
role:consumer
# Works for HTTP too — use `METHOD::/path` form for the exact
# handler, or just `/path` for a method-agnostic wildcard.
- from:gateway
to:orders
type:http
contract:POST::/orders
role:consumer
```
What the manifest extractor does (see [`manifest-extractor.ts`](../../gitnexus/src/core/group/extractors/manifest-extractor.ts)):
1. Builds a canonical `contractId` with `buildContractId` — the same canonicalization used by the automatic extractors, so manifest links cross-match automatic contracts on the other side.
2. Tries to resolve each side to a real graph symbol (the `Route` node for HTTP, a `Function|Method` / `Class|Interface` for gRPC, a `Package|Module` for `lib`).
3. If resolution fails, falls back to a deterministic synthetic uid (`manifest::<repo>::<contractId>`) so both sides still line up in cross-impact — name-only links still work when the symbol isn't in the graph.
4. Emits both a provider and a consumer `StoredContract` (confidence `1.0`, `source: "manifest"`) and a `CrossLink` with `matchType: "manifest"`.
Use `links` for exactly the cases the extractor can't infer: different package names across repos (see #701), hand-rolled transports, cases where the provider repo isn't checked out locally but you still want a record, or any contract whose provider and consumer simply don't share a surface the extractors know how to pattern-match.
History: the manifest extractor used to be silently skipped by the sync pipeline; that was fixed in [#827](https://github.com/abhigyanpatwari/GitNexus/pull/827) (tracking issue #826). If you ever see `config.links` with zero cross-links in `contracts.json`, make sure you're on a build that includes that fix, then re-run `group sync`.
## Troubleshooting
1.**`contracts.json` is empty after a sync.** Either no member repo contained a recognizable gRPC pattern, or the extractors are disabled in `detect`. Confirm `detect.grpc: true` and re-run with `--verbose`.
2.**A known provider/consumer pair doesn't cross-link.** Most common cause: the package segment differs. Check the raw contract ids with `gitnexus group contracts <name> --unmatched` — if you see two same-method contracts with different package prefixes, add a manifest `links:` entry to bridge them (no automatic rewrite rules yet).
3.**`matchType: "manifest"` is missing entirely.** The extractor needs `config.links` to be non-empty and the sync pipeline to actually call it — verify you're on a post-#827 build. Empty contract rows for manifest links usually mean `resolveSymbol` couldn't find a graph match; the synthetic uid still lets cross-impact work, it just won't carry a file path.
4.**Ambiguous proto warnings.** Look for `[grpc-extractor] Ambiguous proto resolution` in the sync logs; that means a service name exists in multiple `.proto` files under the same repo and the path-distance heuristic couldn't pick a winner. Resolve by renaming the service or declaring the intended pairing in `config.links`.
5.**Cross-impact says "stale".** Both sides need a fresh per-repo index _and_ a fresh group sync. Order matters: `gitnexus analyze` in each changed repo, then `gitnexus group sync <name>`. Use `gitnexus group status <name>` to see which side is behind.
## Related docs and references
- [AGENTS.md](../../AGENTS.md) — authoritative list of MCP tools and resources, including group-mode routing and the `gitnexus://group/…` resources.
- [ARCHITECTURE.md](../../ARCHITECTURE.md) — overall data flow and the call-resolution DAG that the per-repo indexer uses.
- CALL USING supports mixed modes: `USING BY REFERENCE WS-A BY CONTENT WS-B BY VALUE WS-C`
- CALL USING `ADDRESS OF` and `OMITTED` must be filtered from parameter lists
- EXEC DLI can have multiple SEGMENT levels in hierarchical retrieval (use matchAll)
- DECLARATIVES can have multiple USE sections (one per file + catch-all for INPUT/OUTPUT/I-O/EXTEND)
- INSPECT TALLYING can have multiple counters in a single statement
- STRING/UNSTRING can span multiple lines (need accumulator pattern)
---
# Complete COBOL Language Feature Coverage
## Overview
Implement the remaining 25 unhandled COBOL language features and fix 10 partial features to achieve ~95% coverage (up from 71.9%). The goal is to build the richest possible knowledge graph from COBOL codebases, enabling a future `modernize` MCP command (out of scope for this plan) that would use the graph to assist with COBOL-to-modern-language migration.
## Problem Statement
The COBOL processor currently handles 54 of 89 applicable language features (71.9%). The 25 unhandled features represent real data loss in the knowledge graph:
- **Cross-program data flow** is invisible (CALL ... USING parameters not extracted)
- **IMS/DB programs** produce empty graphs (EXEC DLI not recognized)
- **String transformation logic** is invisible (STRING/UNSTRING/INSPECT not tracked)
- **SQL copybook dependencies** are missing (EXEC SQL INCLUDE not mapped)
- **Error handling flows** are lost (DECLARATIVES/USE AFTER not captured)
## Proposed Solution
Implement features in 4 phases, ordered by graph value density (edges created per LOC of implementation). Each phase is independently shippable and testable.
## Technical Approach
### Phase 1: High-Value Data Flow Edges (~150 LOC, ~8 new edge types)
The highest-ROI features: they create new ACCESSES and IMPORTS edges that directly improve impact analysis.
**Critical research finding**: Multi-line statement accumulation is the dominant challenge. CALL USING, STRING/UNSTRING, and multi-line data item clauses all span multiple lines in production COBOL. The free-format path processes each line independently — these features need statement accumulators (like SORT/SELECT) or the free-format path needs multi-line awareness. Estimated LOC increased from 110 to 150 to account for accumulator infrastructure.
- **What:** After capturing CALL target, scan for USING clause. Extract parameter names (reuse USING_KEYWORDS filter). Store as `calls[].parameters: string[]`
- **Interface:** Add `parameters?: string[]` to calls array type in CobolRegexResults
- Mixed modes: `CALL 'PGM' USING BY REFERENCE WS-A BY CONTENT WS-B BY VALUE WS-C`
- Pointer passing: `CALL 'PGM' USING ADDRESS OF WS-A`
- Placeholder: `CALL 'PGM' USING OMITTED WS-B`
- Filter keywords: add `ADDRESS`, `OMITTED`, `LENGTH` to USING_KEYWORDS (already has BY/VALUE/REFERENCE/CONTENT)
- **Impact tool enhancement:** CALL-USING edges enable BFS traversal through parameter data flow — single most impactful edge type for COBOL impact analysis
#### 1.3 STRING/UNSTRING data flow -> ACCESSES edges
- **File:** `cobol-preprocessor.ts` (new section in extractProcedure)
- **What:** Accumulate multi-line STRING/UNSTRING until period or END-STRING/END-UNSTRING. Extract sources and INTO targets.
- IBM allows `OCCURS 0 TO n DEPENDING ON` (zero minimum) and `OCCURS UNBOUNDED DEPENDING ON` (V6.4)
- Subscripted controlling fields: `DEPENDING ON WS-COUNT(WS-IDX)` — strip subscripts before storing
- **Pre-existing gap**: Multi-line data item clauses without continuation indicator are NOT captured. `05 WS-TABLE\n OCCURS 100\n DEPENDING ON WS-COUNT.` — the current RE_DATA_ITEM only gets the first line, `rest` is empty. Fixing properly requires a data item accumulator (like SELECT). **Defer full fix to Phase 3; implement same-line capture now.**
- KEY IS fields: `ASCENDING KEY IS WS-KEY-1 WS-KEY-2` — capture for SEARCH ALL resolution
- INDEXED BY: `INDEXED BY IDX-1 IDX-2` — capture for SET/SEARCH context
- Numeric with sign/decimal: `VALUE -123.45`, `VALUE +1`
-`VALUE IS` optional — both `VALUE 'A'` and `VALUE IS 'A'` valid
- **Decimal vs period ambiguity**: `VALUE 100.` — is `.` decimal or terminator? `parseDataItemClauses` already strips trailing period, so this is handled
- IBM V6.4: floating-point `VALUE 1.0E5` — extend numeric regex if needed
- Implementation: use a pragmatic `extractValue(rest)` function, not a single complex regex
- **Graph:** CodeElement node + ACCESSES edge to `<ims>:<segmentName>` Record node with reason `dli-{verb}`; ACCESSES edges to INTO/FROM data areas; PSB ACCESSES for SCHD
- **Tests:** `EXEC DLI GU USING PCB(1) SEGMENT(CUSTOMER) INTO(WS-CUST) END-EXEC`
**Research insights (dual IMS interface):**
- **EXEC DLI**: Embedded command interface for CICS-DL/I programs only
- **CBLTDLI CALL**: Batch interface via `CALL 'CBLTDLI' USING function-code PCB io-area SSA1..SSA15`
- CBLTDLI is already captured as a CALL to 'CBLTDLI' — enrich with USING parameter semantics later
- Multiple SEGMENT levels in hierarchical retrieval — use `matchAll` on segment regex
- **Graph:** ACCESSES read on inspected field always; write if REPLACING/CONVERTING. Write edges for tally counters. Reason: `cobol-inspect-read`/`cobol-inspect-write`/`cobol-inspect-tally`
- **Tests:** `INSPECT WS-FIELD TALLYING WS-COUNT FOR ALL 'A'` -> read on WS-FIELD, write on WS-COUNT
**Research insights (INSPECT forms by frequency):**
- REPLACING (~60%): `INSPECT WS-STR REPLACING ALL 'A' BY 'B'`
- TALLYING (~25%): `INSPECT WS-STR TALLYING WS-CNT FOR ALL 'A'` — multiple counters possible
- CONVERTING (~10%): `INSPECT WS-STR CONVERTING 'abc' TO 'ABC'`
- Combined (~5%): TALLYING + REPLACING in single statement
- **Needs multi-line accumulator** — INSPECT frequently spans 3-5 lines in production
- Extract tally counters with `([A-Z][A-Z0-9-]+)\s+FOR\b` matchAll pattern
- Filter figurative constants (SPACES, ZEROS) using existing MOVE_SKIP set
### Phase 3: Completeness Fixes (~60 LOC)
Fix the 10 partial features and small gaps.
#### 3.1 CALL ... RETURNING extraction
- Extend RE_CALL processing to capture RETURNING target after the USING clause
- Store as `calls[].returning?: string`
- Graph: ACCESSES write edge with reason `cobol-call-returning`
#### 3.2 SELECT OPTIONAL flag preservation
- Store `isOptional: boolean` in FileDeclaration interface
- Include in Record node description
#### 3.3 ALTERNATE RECORD KEY extraction
- Add regex in parseSelectStatement: `/\bALTERNATE\s+RECORD\s+KEY\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)/i`
- [ ] Integration test assertions updated with exact counts per phase
- [ ] Benchmark run after each phase to track graph growth
## Dependencies & Risks
### Dependencies
- None. All changes are additive to existing COBOL processor code.
- No LanguageProvider changes needed.
- No graph schema changes needed (all new constructs map to existing node labels + edge types).
### Risks
- **preprocessor.ts size**: Currently 1326 LOC. Phase 1+2 adds ~200 LOC -> 1526 LOC. May need to extract helpers into a separate `cobol-data-flow.ts` module if it exceeds 1500.
- **REPLACE statement** (Phase 3.7) is the most complex feature — requires tracking text substitution state across logical lines. Consider deferring to a separate PR if it takes >100 LOC.
- **EXEC DLI** (Phase 2.1) is only testable against IMS codebases. Need fixture data or synthetic test cases.
## Graph Value Ranking by MCP Tool Impact
Research agent analyzed all 5 MCP tools (query, context, impact, detect_changes, rename) against planned edge types:
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Fix 4 HIGH-priority issues from PR #626 code review before merge.
**Architecture:** Minimal targeted fixes — each task is independent. TDD: tests first, then implementation. No refactoring beyond what's needed.
**Paths:** All file paths are relative to the monorepo root (`GitNexus/`). Git commands run from the root. The `gitnexus/` prefix is a package subdirectory, not a separate repo.
---
### Task 1: Path Traversal — Validate Group Name
**Files:**
- Modify: `gitnexus/src/core/group/storage.ts:17-19` (getGroupDir) and `:63-68` (createGroupDir)
- [ ]**Step 1: Write failing tests for validateGroupName**
In `gitnexus/test/unit/group/storage.test.ts`, add `createGroupDir` and `validateGroupName` to the existing import from `'../../../src/core/group/storage.js'` (line 6-11). Then add these describe blocks at the end of the outer `describe('Group storage', ...)`:
- [ ]**Step 2: Run tests to verify `vendor` and `target` tests fail**
Run: `cd gitnexus && npx vitest run test/unit/group/service-boundary-detector.test.ts`
Expected: `test_detect_skips_vendor_directory` and `test_detect_skips_target_directory` FAIL (vendor/target not excluded). Other new tests may pass since dotfile exclusion already exists.
- [ ]**Step 3: Add EXCLUDED_DIRS constant and update both walking functions**
In `gitnexus/src/core/group/service-boundary-detector.ts`:
- CLI smoke: one integration test that hits `getGroupDir` / `createGroupDir` (e.g. `group create` or `group add`) with an invalid name proves wiring for every subcommand that resolves a group through storage
### CLI/API entry points accepting groupName
All paths flow through `getGroupDir` (which validates), so coverage is implicit. For reference:
**`listGroups`:** Reads directory names from disk without validation. Not a write path, so no traversal risk. May surface manually-created directories with non-conforming names — accepted as-is, not in scope.
**Risk:**`serviceRe = /service\s+(\w+)\s*\{([^}]*)}/gs` stops at first `}`. Proto services with `google.api.http` annotations inside RPCs contain nested `{ }` blocks.
1. Use regex only to find `service <Name> {` start positions (regex consumes the opening `{`)
2. Initialise depth to 1 immediately after the opening `{`
3. Scan forward char by char: `{` -> depth++, `}` -> depth--; collect into body
4. Stop when depth reaches 0 (the matching closing `}`)
5. Return name + body pairs
Inner `rpcRe` regex remains unchanged — it operates on the already-extracted body.
**Malformed input:** If EOF is reached before `depth` returns to 0, skip the incomplete service (do not add to results). Lock this in the test.
**Scope limitation (v1):** Brace-depth only — no lexer for string literals or comments containing `{`/`}`. Sufficient for `google.api.http` annotations. Known false positive: braces inside `//` comments or quoted strings within proto options. Accepted for v1; a proper proto lexer is out of scope.
### Tests
- Proto with single service, no nesting (regression)
- Proto with `google.api.http` nested braces inside RPC options
- Proto with multiple services
- Proto with nested `option` blocks inside RPC (e.g. `google.api.http`)
- Malformed proto with unclosed brace (graceful handling)
---
## Fix 3: Directory Exclusions in Service Boundary Detector
**Risk:** Walks entire repo tree, only skipping dotfiles and `node_modules`. Extremely slow on repos with `vendor/`, `target/`, `__pycache__/`, `.venv/`.
### Solution
Create `EXCLUDED_DIRS` as a `Set<string>` (alongside existing `SERVICE_MARKERS`, `SOURCE_EXTENSIONS`), for example:
```text
node_modules, vendor, target, build, dist,
__pycache__, .venv, venv, .tox, .mypy_cache,
.gradle, .mvn, out, bin
```
(Implement as `new Set([...])` — the list above is the membership, not a string literal.)
Apply in both:
-`walkForBoundaries` (line 77-78) — replace current inline `=== 'node_modules'` check with `EXCLUDED_DIRS.has(entry.name)`
-`hasSourceFilesInSubdirs` (line 130) — replace `entry.name !== 'node_modules'` with `!EXCLUDED_DIRS.has(entry.name)`
Note: remove the old `=== 'node_modules'` literal from both locations — it is covered by `EXCLUDED_DIRS`.
Dotfile exclusion (`.` prefix) remains as a separate check since it's a pattern, not a name.
Exclusions apply only to `isDirectory()` entries — file names are never checked against `EXCLUDED_DIRS`.
**Tradeoff:** Rare layouts that keep source under names like `out/` or `bin/` will be skipped; accepted for performance on typical monorepos.
**Case sensitivity:**`Set.has` is case-sensitive (matches current `=== 'node_modules'` behavior). Windows case-insensitive FS not handled — accepted as-is, consistent with existing code.
Set `GITNEXUS_EVAL_DEBUG=1` to include full Python tracebacks in run summaries and logs. By default, errors are sanitized to avoid leaking host paths or stack traces.
## Quick Start
### Debug a single instance
@@ -148,7 +152,7 @@ Each mode has a `system_{mode}.jinja` + `instance_{mode}.jinja` pair. The agent
1. Docker container starts with SWE-bench instance (repo at specific commit)
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.