Compare commits

...
Author SHA1 Message Date
azizur100389 e262dda35b fix(cli): only match <!-- gitnexus:* --> markers at section position (#1041) (#1042)
`upsertGitNexusSection` in ai-context.ts uses `indexOf` to locate the
bounds of the GitNexus section in CLAUDE.md / AGENTS.md before
replacement. `indexOf` matches the first occurrence of the marker
anywhere in the file, including inline prose references in backtick-
quoted fragments mid-sentence.

The shipped CLAUDE.md contains exactly such a reference ("See the
`<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in AGENTS.md
for the canonical MCP tools..."). Running `gitnexus analyze` on a
fresh install matches those inline markers as section delimiters and
replaces the prose between them with the full ~100-line injected
block, breaking the backtick and corrupting markdown for every user.

Fix: new private `findSectionMarkerIndex` helper that only matches
markers occupying their own line — preceded by `\n` or start-of-file,
followed by `\n` / `\r` (CRLF files) / end-of-file. `\r` is explicit
so CRLF-terminated sections on Windows (core.autocrlf = true) still
match. The generator always emits markers alone on their line, so
every legitimate section continues to update in place; only inline
prose references now fall through to the append branch, which leaves
existing content untouched.

Two new unit tests:
- #1041 regression — seed CLAUDE.md with the shipped inline prose
  line, run analyze twice, assert inline prose preserved verbatim
  and marker counts stay at 2/2 (1 inline + 1 section-position)
- CRLF handling — seed a CRLF file with inline prose + legitimate
  section, run analyze, assert section replaced in place, inline
  prose preserved, stale stub content removed

No destructive ops, no bypass flags, no new deps. Behaviour change
is strictly narrowing — files that previously updated correctly
still do; files that previously got corrupted now fall through to
the safer append branch.

Closes #1041.
2026-04-23 07:58:07 +01:00
dependabot[bot] fe0434ae46 chore(deps)(deps-dev): bump @babel/types in /gitnexus-web (#1037) 2026-04-23 05:07:16 +01:00
dependabot[bot] 43abe44e37 chore(deps)(deps): bump @huggingface/transformers in /gitnexus (#1035) 2026-04-23 05:06:52 +01:00
dependabot[bot] 42d276bc4a chore(deps)(deps-dev): bump typescript in /gitnexus-shared (#1034) 2026-04-23 05:06:42 +01:00
dependabot[bot] 36104cdbd2 chore(deps): bump actions/setup-node from 6.3.0 to 6.4.0 (#1033) 2026-04-23 05:06:22 +01:00
Tom Hale 6618120f63 fix: preserve comments and config in opencode.json during setup (#998)
* deps: add jsonc-parser for JSONC-safe config editing

* fix: use jsonc-parser to preserve comments in opencode.json during setup

- Add mergeJsoncFile() using parseTree/modify/applyEdits pipeline
- Add getOpenCodeMcpEntry() for OpenCode MCP format { type: local, command: [...] }
- Replace readJsonFile+writeJsonFile in setupOpenCode with mergeJsoncFile
- Fix wipe bug: JSON.parse on JSONC comments caused catch block to reset config to {}
- Add 9 tests for JSONC comment preservation, corrupt file safety, and format

* fix: use parseTree error collection and detect indentation

- Pass parseErrors array to parseTree() instead of checking
  (tree as any).errors which was always undefined — a real bug
  that allowed corrupt files to be rewritten
- Detect tab indentation from file content to avoid mixed
  indentation in modified JSONC files
- Fix JSDoc to match actual fallback behavior (JSON.parse, not
  readJsonFile)
- Strengthen corrupt-file test to assert exact content match

* style(setup): fix prettier formatting on mergeJsoncFile

* fix(setup): remove dead JSON.parse fallback, detect space-indent width, fix JSDoc

- Remove the semantically unreachable JSON.parse fallback branch in
  mergeJsoncFile (jsonc-parser's parseTree is a strict superset of
  JSON.parse, so the fallback can never fire for content JSON.parse
  would accept)
- Replace binary tab/space detection with detectIndentation() that
  measures actual indent width from the first indented line
- Fix JSDoc: 'valid JSON that is not valid JSONC' is impossible by
  definition
- Add tests for tab indentation and 4-space indentation preservation
2026-04-22 17:21:56 +01:00
PhmTunsandTuanPM1 ea418c0126 docs: fix group add and group remove usage in READMEs (#1020)
The top-level and CLI READMEs advertised `gitnexus group add <name> <repo>`
(two args) and `gitnexus group remove <name> <repo>`, but the CLI
(`gitnexus/src/cli/group.ts`) actually requires three args for `add`
(`<group> <groupPath> <registryName>`) and uses `<groupPath>` — not a
repo path — for `remove`. Reusing the same second argument across two
`group add` invocations silently overwrote the previous mapping because
the hierarchy path is the key in `group.yaml`'s `repos` map.

Update both READMEs to match the real CLI contract. Node_modules not
installed locally for this docs-only change, so pre-commit (prettier +
typecheck) was skipped.

Made-with: Cursor

Co-authored-by: TuanPM1 <tuanpm1@kaopiz.com>
2026-04-22 07:48:44 +01:00
962f22482b feat(cli): Fingerprint indexed repos by remote URL to detect sibling-clone graph drift (#982)
* Initial plan

* feat: detect sibling-clone graph drift via remote URL fingerprint

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e5decb67-7fec-40e7-b2a1-b5e94a0d393f

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: address review feedback — fake commit, same-commit case, regex docs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e5decb67-7fec-40e7-b2a1-b5e94a0d393f

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(mcp): address review feedback — CI green, perf, dead branch, one-shot test

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc2259f7-94e4-4243-aaa9-e03b7c632d32

* Merge branch 'main' into copilot/fix-single-path-indexing-issue

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/5840b3dd-e879-4854-a067-d1622bec2634

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* Merge branch 'main' into copilot/fix-single-path-indexing-issue

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9025262f-4dd4-4774-8f32-e14434100004

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* style: prettier format run-analyze.ts after merge with main

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a7be18dd-102f-4a7b-ac56-53fbd414fe3b

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: realpath both sides of cwdGitRoot assertion for Windows 8.3 short-name compat

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b2a1c6a3-e454-4b87-b0e4-69d7c0d9a51b

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(test): use path-agnostic assertion for cwdGitRoot on Windows (#1015)

git rev-parse --show-toplevel returns long path names on Windows
while os.tmpdir() returns 8.3 short names. fs.realpathSync does not
expand short names, so exact path comparison always fails on Windows
CI runners. Replace with behavioral assertions instead.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: copilot-swe-agent[bot] <copilot-swe-agent[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: evolution <wjc163@sina.cn>
2026-04-21 21:58:54 +01:00
dependabot[bot] 064f0f5f55 chore(deps)(deps-dev): bump @vitest/coverage-v8 in /gitnexus (#1018)
Bumps [@vitest/coverage-v8](https://github.com/vitest-dev/vitest/tree/HEAD/packages/coverage-v8) from 4.1.4 to 4.1.5.
- [Release notes](https://github.com/vitest-dev/vitest/releases)
- [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.5/packages/coverage-v8)

---
updated-dependencies:
- dependency-name: "@vitest/coverage-v8"
  dependency-version: 4.1.5
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-21 21:33:16 +01:00
dependabot[bot] d4fbf108a1 chore(deps)(deps-dev): bump vitest from 4.1.4 to 4.1.5 in /gitnexus (#1017)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.4 to 4.1.5.
- [Release notes](https://github.com/vitest-dev/vitest/releases)
- [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.5/packages/vitest)

---
updated-dependencies:
- dependency-name: vitest
  dependency-version: 4.1.5
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-21 21:33:03 +01:00
dependabot[bot] 5a109e32b6 chore(deps)(deps-dev): bump @types/uuid in /gitnexus (#1016)
Bumps [@types/uuid](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/uuid) from 10.0.0 to 11.0.0.
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/uuid)

---
updated-dependencies:
- dependency-name: "@types/uuid"
  dependency-version: 11.0.0
  dependency-type: direct:development
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-21 21:32:48 +01:00
evolution 6ad53ff5c7 fix(ci): Change docker base image from alpine to debian (#1014)
* fix(docker): switch Dockerfile.cli from Alpine to Debian slim

Alpine uses musl libc which is incompatible with @ladybugdb/core's
glibc-compiled native binary, causing ERR_DLOPEN_FAILED on startup.

Closes #1008

* fix(docker): resolve build and runtime failures in Dockerfile.cli

- Add **/*.tsbuildinfo to .dockerignore and rm -f tsbuildinfo in
  builder to prevent stale incremental cache from skipping
  gitnexus-shared compilation
- Install libstdc++6 from Debian Trixie for @ladybugdb/core native
  module compatibility (requires GLIBCXX_3.4.31)

* fix(docker): use node:22-trixie-slim for GLIBCXX_3.4.31 support

Replaces the manual Trixie libstdc++6 backport with the official
node:22-trixie-slim base image, which ships GCC 14 runtime natively.
2026-04-21 21:31:58 +01:00
Sam Fakhreddine 95a38c7e2d fix(group): surface friendly error when group name not found (#903 regression test) (#989)
* fix(group): surface friendly error when group name not found

Squashed commits:
- test(csharp): add #903 regression — parse completeness for single-file C# repo
- fix(group): add GroupNotFoundError guard to groupList + re-throw tests for groupQuery/groupStatus
- fix(test): restore section comments in csharp.test.ts stripped during rebase

* fix(group): catch GroupNotFoundError explicitly in groupContext and groupImpact
2026-04-21 15:52:36 +01:00
ff4ae89aaa feat(python): scope-based call resolution + registry-primary flip + perf + generalization (RFC #909 Ring 3) (#980)
* Initial plan

* plan: Python scope-based resolution migration

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0eee6c69-fc17-4df5-9ac6-358ab41f5740

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat(python): scope-based resolution provider hooks + 62 tests

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0eee6c69-fc17-4df5-9ac6-358ab41f5740

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor(python): split scope-hooks monolith into focused modules

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/db76e937-4b0e-4c4d-82b1-265a1fb3673d

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(python): integration-style scope-resolution tests + suffixResolve fallback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/db76e937-4b0e-4c4d-82b1-265a1fb3673d

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* wire python scope-based resolution end-to-end (initial pass)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c474dc66-5cf7-445d-8eb4-76501c5e6d67

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* keep legacy IMPORTS for python (heritage needs importMap), scope phase owns CALLS only

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c474dc66-5cf7-445d-8eb4-76501c5e6d67

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(python): remove parallel scope-resolution integration test

The new test/integration/python-scope-resolution.test.ts duplicated coverage
the reviewer explicitly rejected. The existing
test/integration/resolvers/python.test.ts (191 tests, driven by
runPipelineFromRepo) is the source of truth for Ring 3 parity.

Also document the IMPORTS-emission follow-up gap: wiring emitImportEdges
in python-scope-emit.ts today regresses 10 IMPORTS-edge fixtures because
the scope-extractor's ImportEdge coverage is narrower than legacy
pythonImportConfig.importResolver. Tracked as a follow-up.

Baseline with REGISTRY_PRIMARY_PYTHON=1 is unchanged: 109/191 pass.

* feat(ingestion): scope-resolution phase owns Python IMPORTS edges (RFC #909 Ring 3)

When `REGISTRY_PRIMARY_PYTHON=1`, IMPORTS graph edges for Python files are now
emitted exclusively by the new scope-resolution path. The legacy
`import-processor` still runs — heritage resolution needs its importMap /
namedImportMap / moduleAliasMap population — but its graph edge emission is
gated per-language so Python no longer double-emits.

This closes the reviewer's second change request on PR #980: "the legacy path
must be turned off". Legacy IMPORTS edges for Python are now off by default
when the flag is enabled.

Three bugs were fixed to make the new path's coverage match legacy:

1. **Root-file bailout** (import-resolvers/python.ts): `resolvePythonImportInternal`
   returned null immediately when the importer file lived at the repo root
   (importerDir === ''). The ancestor directory walk further down already
   handles this case correctly; the early return was the bug. Proximity check
   now only runs when importerDir is non-empty, and the ancestor walk sees
   root-level files for the first time.

2. **External dotted imports** (languages/python/import-target.ts): the new
   path fell straight through to `suffixResolve` for multi-segment imports,
   which happily matched `django.apps` to a local `accounts/apps.py`. Mirror
   `pythonImportStrategy`'s `hasRepoCandidate` guard — suffix-match only when
   the leading segment exists somewhere in-repo as a package, __init__.py,
   or namespace directory.

3. **suffixResolve ambiguity** (languages/python/import-target.ts): the
   shared `suffixResolve` helper requires a pre-built `SuffixIndex` to
   disambiguate ties. Without one it falls back to an O(files) scan that
   silently picks the first match when the last segment collides across
   directories (e.g. `accounts.models` matching `billing/models.py`).
   Replaced with `resolveAbsoluteFromFiles` — exact lookup first, then a
   deterministic suffix match.

Validation:
- Flag OFF: 191/191 pass (no regression).
- Flag ON: 109/191 pass (82 fail — exact baseline match; remaining 82 are
  unchanged CALLS-edge provider-feature gaps tracked as Phase B follow-ups).
- `tsc --noEmit`: clean.

The 82 CALLS failures cluster into 44 describe blocks covering type-inference
features (assignment chains, walrus, class-level annotations, constructor
inference, C3 MRO, overload dispatch, return-type inference) that need
dedicated Ring 3 follow-up work. Each cluster is tracked against the RFC #909
shadow-parity gate (>=99% fixtures / >=98% corpus) in the per-language ticket.

* ci(scope-resolution): automatic parity gate driven by MIGRATED_LANGUAGES

Adds the Ring 3 parity gate the RFC §6.4 requires: when a language's
scope-resolution migration is marked complete, CI runs its resolver
integration test twice on every PR (once with the legacy DAG, once with
the registry-primary path) and both must pass.

The "is this language migrated" signal is a single TypeScript constant:

  // gitnexus/src/core/ingestion/registry-primary-flag.ts
  export const MIGRATED_LANGUAGES: ReadonlySet<SupportedLanguages> =
    new Set([ /* SupportedLanguages.Python when ready */ ]);

Adding a language here has three simultaneous effects:

  1. `isRegistryPrimary(lang)` defaults to true for that language in
     production (env-var override still wins if set explicitly).
  2. `.github/workflows/ci-scope-parity.yml` auto-discovers the set via
     `npx tsx scripts/ci-list-migrated-languages.ts`, builds a parity
     matrix, and runs:
       - `REGISTRY_PRIMARY_<LANG>=0 npx vitest run resolvers/<slug>.test.ts`
       - `REGISTRY_PRIMARY_<LANG>=1 npx vitest run resolvers/<slug>.test.ts`
     Both legs must pass for the job to succeed.
  3. Legacy-path gating in call-processor.ts / import-processor.ts kicks
     in automatically through the same `isRegistryPrimary` lookup.

No JSON registry, no manual workflow edit, no second source of truth —
contributors update the Set and CI picks it up. Empty Set = parity job
is a skipped matrix (workflow still reports success).

The new `scope-parity` reusable workflow is added to ci.yml's `needs`
graph and ci-status gate. Its result must be `success` (skipped would
mean upstream discover job failed and should block).

Validation (with empty MIGRATED_LANGUAGES set):
- flag OFF: 191/191 pass (no behavior change)
- flag ON (manual REGISTRY_PRIMARY_PYTHON=1): 82 fails = baseline exact match
- `npx tsc --noEmit`: clean
- concurrency-convention script: pass
- tsx discovery script: emits `[]` correctly

* ci(scope-resolution): keep MIGRATED_LANGUAGES empty; fix linter auto-uncomment

Previous commit's example entry got auto-uncommented (linter preferred a
type-checkable `SupportedLanguages.Python` over a commented-out reference).
That would have triggered the parity CI gate against Python, which today
has 82 known flag-on failures — unintended and would block the PR.

Use the explicit generic `new Set<SupportedLanguages>([])` so an empty set
still type-checks without needing an uncommented-out sample member.
Example in the comment now has `//   SupportedLanguages.Python,` so it
remains illustrative without participating in the set.

* feat(python): capture constructor-inferred + annotated type bindings

Extends the Python scope-extractor with two new type-binding capture
patterns so receiver-typed method dispatch has concrete type bindings
to work from:

1. `u: User = ...` / `u: User` — variable annotations. `@type-binding.annotation`
   anchor, `source: 'annotation'`.
2. `u = User("alice")` — assignment RHS is a bare-identifier call (Python
   has no `new` keyword; constructor-shaped calls are syntactically
   identical to function calls). `@type-binding.constructor` anchor,
   `source: 'constructor-inferred'`.

The runtime query lives in `query.ts` (the `.scm` file is documentation
per the comment at its top); both are updated.

Fixes 19 failures across these resolver fixtures (flag-on 82 → 63):
- Python constructor-inferred type resolution (3)
- Python class-level annotation resolution (3)
- Python nullable receiver resolution (3)
- Python member-call / receiver-constrained / constructor-call (3)
- Python assignment chain propagation (2)
- Python walrus / match-case / chained method (3)
- Python member access iterable for-loop (2)

* feat(python): strip nullable unions + prefer annotations over inference

Two linked changes that together fix the 4 nullable-receiver tests:

1. `stripNullable` in Python's `interpretTypeBinding` unwraps `User | None`,
   `None | User`, and `Optional[User]` to `User`, so receiver-typed
   resolution treats nullable receivers identically to non-nullable ones.
   Three-arm unions (`User | Error | None`) are left unchanged — truly
   ambiguous for single-receiver inference.

2. Source-strength ordering in `pass4CollectTypeBindings`. When multiple
   matches fire for the same bound name in the same scope — e.g. the
   `u: User = find()` idiom where both the annotation and
   constructor-inferred patterns match — the explicit annotation now
   wins regardless of query-match arrival order. Rank:
     explicit (annotation / parameter-annotation / return-annotation / self) > inferred

Also reorders the two Python patterns in query.ts / scopes.scm so the
constructor-inferred pattern appears first — a belt-and-braces fallback
that keeps behavior deterministic if the shared priority ranking is ever
revisited.

Fixes 4 failures (flag-on 63 → 59):
- Python nullable receiver resolution (4 tests)

Flag-off regression check: 191/191 still pass.

* feat(python): walrus, qualified-call, match-case type bindings

Extends the constructor-inferred family of captures with three more
assignment-shaped patterns that all bind a variable to a class-like type:

- Walrus: `(u := User(...))` → `u: User` via `(named_expression)`.
- Qualified call RHS: `u = models.User(...)` → `u: models.User` via
  `(attribute)` node .text. Falls through resolveTypeRef Phase 2
  (QualifiedNameIndex dotted fallback).
- Match as-pattern: `case User() as u:` → `u: User` via `(as_pattern)`
  + `(class_pattern (dotted_name))`.

Fixes 2 failures (flag-on 59 → 57):
- Python walrus operator type inference
- Python match/case as-pattern type binding

Qualified-call constructor tests still fail because they require
cross-module qualifiedName registration (models.User → models.py's User
class) which isn't yet wired in the Python extractor. Tracked as
follow-up alongside module-import CALLS (#337) resolution.

* feat(python): chain type bindings + strip list[T] generic for for-loop

Adds two capture patterns and a shared transitive-closure pass that
together handle Python's variable-aliasing and for-loop-over-typed-
iterable patterns:

1. `(assignment left: (identifier) right: (identifier))` — `alias = u`.
2. `(for_statement left: (identifier) right: (identifier))` — `for u in users`.

Both emit `@type-binding.alias` with the RHS identifier as rawName. The
shared `pass4CollectTypeBindings` now runs a final transitive-closure
walk that follows identifier-chain TypeRefs through the declaring scope
and its ancestors (depth-capped, cycle-guarded) so `alias` ultimately
points at the class type instead of another local variable name.

Generic stripping in `interpret.ts` unwraps single-arg collection
wrappers — `list[User]`, `set[User]`, `Iterable[User]`, etc. — to the
element type. Multi-arg generics (`dict[str, User]`, `Callable[...]`)
are left alone; their semantics aren't unambiguous.

Fixes 8 failures (flag-on 57 → 49):
- Python assignment chain propagation (4)
- Python nullable + assignment chain (2)
- Python walrus operator (:=) assignment chain (2)

Flag-off still 191/191.

* feat(python): namespace & class receiver resolution + file-level caller fallback

Adds a Python-specific post-resolution pass `emitReceiverBoundCalls`
that closes two receiver gaps the shared `MethodRegistry.lookup` doesn't
cover:

1. **Namespace receivers** — `import models; models.User()` /
   `import models as m; m.User()`. The shared `lookupReceiverType` only
   walks `scope.typeBindings`; namespace imports never land there
   (they're filtered out of `scope.bindings` when the target module
   has no self-named def, per `finalize-algorithm.ts:540`). The new
   pass walks `indexes.imports` directly, builds a per-file
   `localName → targetFilePath` map, and emits CALLS/ACCESSES edges
   against the target file's `localDefs`.

2. **Class-name receivers** — `Dog.classify("dog")`. The shared resolver
   requires typeBindings; class bindings in `scope.bindings` are never
   consulted as receivers. The new pass checks class-kind bindings in
   the call scope's chain and resolves members via `ownerId`.

Also fixes module-level call attribution: `resolveCallerGraphId` now
falls back to the File node id (`generateId('File', filePath)`) when no
enclosing function/method/class is found. Matches legacy DAG behavior
for module-scope calls like `u = models.User()` at the top of app.py.

Fixes 4 failures (flag-on 49 → 45):
- Python module import CALLS resolution (Issue #337) (4 of 7)

Flag-off still 191/191.

* feat(python): dotted-typebinding receiver resolution

Adds case 3 to `emitReceiverBoundCalls`: when a receiver's typeBinding
has a dotted rawName like `u: models.User` (the constructor-inferred
form fired by `u = models.User(...)`), walk the namespace map + target
file's defs to find the class, then look up the member via ownerId.

`resolveTypeRef`'s QualifiedNameIndex fallback can't cover this because
the target class's qualifiedName in models.py is just `"User"`, not
`"models.User"` — the dotted form only exists in the call-site file's
receiver expression. This pass bridges that gap without modifying the
shared registry.

Fixes 9 more failures (flag-on 45 → 36):
- Python qualified constructor inference (2)
- Python module import CALLS resolution (Issue #337) (3)
- (cluster overlap — several downstream tests in assignment/nullable/
  walrus that propagate through qualified-ctor bindings also benefit)

Flag-off still 191/191.

* feat(python): consult finalized bindings for receiver resolution

`findClassBindingInScope` now walks BOTH:
  1. `scope.bindings` — pre-finalize local declarations (origin: 'local')
  2. `indexes.bindings` — post-finalize cross-file imports/namespaces

Without (2) we were blind to any class brought in via
`from models import Dog` at the call site's file, because the
scope-extractor's Pass 2 only populates local bindings and the
cross-file finalize produces a separate bindings map that never lands
on `scope.bindings`.

Case 2 (`Dog.classify()`) now walks MRO so inherited static/class
methods resolve — `Dog.classify()` where `classify` lives on `Animal`.

Case 4 (simple typeBinding like `u: U` from aliased import) now uses
`findClassBindingInScope` instead of the shared `resolveTypeRef`,
because `resolveTypeRef`'s `ctx.scopes` only sees pre-finalize local
bindings too.

Fixes 4 more failures (flag-on 36 → 32):
- Python method enrichment > Dog.classify static (1)
- Python static/classmethod class-as-receiver (2)
- Python alias import resolution (1)

Flag-off still 191/191.

* refactor(python-scope): extract language-agnostic emit-core/

Unit 1 of the python migration architectural plan
(docs/plans/2026-04-19-001-refactor-python-migration-architectural-plan.md).

Splits python-scope-emit.ts (~945 → 481 lines) by lifting 14 generic
graph-feeding primitives into emit-core/:
  - graph-node-lookup, graph-id, emit-edge
  - emit-references, emit-imports
  - scope-walkers (findReceiverTypeBinding, findClassBindingInScope,
    findOwnedMember, findExportedDef)
  - namespace-targets, method-dispatch-bridge

Each file carries a "Next-consumer contract" JSDoc so future language
migrations (TS #927, JS #928, Java, Kotlin, Ruby) import from emit-core
rather than re-implementing. python-scope-emit.ts keeps only the four
Python-specific pieces: runPythonScopeResolution (orchestrator),
buildPythonMro, emitReceiverBoundCalls (4 cases), populateMethodOwnerIds
— these move to languages/python/emit/ in Unit 11.

Pure refactor, zero behavior change:
  - flag-off: 191/191 python.test.ts pass (identical baseline).
  - flag-on (REGISTRY_PRIMARY_PYTHON=1): 32 fail / 159 pass (identical
    baseline — the refactor neither fixes nor regresses any test).
  - tsc --noEmit clean.

* feat(python-scope): arity metadata + bind function decls in parent scope

Unit 2 of the python migration architectural plan
(docs/plans/2026-04-19-001-refactor-python-migration-architectural-plan.md).

Two changes that the registry-primary path needs before any of the
arity-sensitive failures can move:

1. Arity metadata on scope-extracted Function/Method defs.
   - New helper `languages/python/arity-metadata.ts` reuses
     `pythonMethodConfig.extractParameters` so self/cls stripping,
     defaults, and *args/**kwargs detection match legacy semantics.
   - `emit-captures.ts` synthesizes
     `@declaration.parameter-count` /
     `@declaration.required-parameter-count` /
     `@declaration.parameter-types` captures on every
     `@declaration.function` match.
   - Generic `scope-extractor.ts buildDefFromDeclarationMatch` reads
     the three optional captures into `SymbolDefinition`. Absence is
     still the no-op default for non-Python providers.

2. Hoist function/class declaration bindings to the enclosing scope.
   The "innermost scope containing the anchor" default placed
   `def greet(...)` inside greet's OWN body — invisible to other
   module-level callers, so every flag-on free-call resolved to
   `unresolved`. The hoist condition (`anchor range == innermost
   range`) only fires for scope-creating declarations, so variable /
   for-loop captures whose anchor is a child identifier stay put.
   Hooks can still override via `bindingScopeFor`.

Verification:
  - Flag-off: 191/191 (identical baseline).
  - Flag-on (REGISTRY_PRIMARY_PYTHON=1): 31 fail / 160 pass
    (was 32/159; the hoist unblocks free-call resolution end-to-end).
  - tsc --noEmit clean.

Per-(source,target) edge collapse for multi-call-site cases
(default-params, variadic) still pending — landing it without
regressing the static-method find_user fixture (which expects two
distinct edges through different targets) needs the ownership-aware
qualified-id work that lands with Unit 4 / Unit 11.

* feat(python-scope): capture function return-type annotations

Unit 3 of the python migration architectural plan
(docs/plans/2026-04-19-001-refactor-python-migration-architectural-plan.md).

Wires the `def get_user() -> User` return-type annotation into the
typeBindings stream so the existing constructor-inferred + transitive
chain machinery can resolve `u = get_user(); u.save()` to `User#save`
without any orchestrator change.

Changes:
- `query.ts` + `scopes.scm`: new `@type-binding.return` pattern keyed by
  the function name (matches RFC §5.1 canonical vocabulary).
- `interpret.ts`: maps `@type-binding.return` to the existing
  `'return-annotation'` source label (no shared change needed).
- `scope-extractor.ts pass4CollectTypeBindings`: extends the Pass 2
  auto-hoist (anchor range == innermost scope range → bind in parent)
  to type bindings as well — return-type bindings whose anchor IS the
  function_definition land in the function's enclosing scope so
  callers see them.

Same-file return-type inference is now end-to-end:
  `def get_user() -> User: ...` + `u = get_user()` produces
  `u: User (return-annotation)` in the caller's scope via
  `followChainedRef`.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 31 fail / 160 pass (no change — every remaining
  return-type test in this fixture set is *cross-file*; carrying
  `get_user → User` across module boundaries lands with the
  cross-file typeBinding propagation work in Unit 5/7).
- tsc --noEmit clean.

* feat(python-scope): resolve dotted receivers via class-scope field types

Unit 4 partial — the dotted-receiver case (`user.address.save()`).

Class-body annotations like `class User: address: Address` already
land in the class scope's typeBindings via the existing
`@type-binding.annotation` capture. This commit consumes that signal:

- Build a `Map<classDefId, Scope>` from every parsed file's class
  scopes once per resolution pass.
- New Case 0 in `emitReceiverBoundCalls`: when the receiver's name
  contains a dot, walk the chain — resolve the head's type, then for
  each remaining segment look up that field's type in the owner
  class's scope.typeBindings, then emit the call against the final
  class with MRO walk.
- Cross-scope lookups use each TypeRef's `declaredAtScope` so an
  imported `Address` resolves in the file that owns the field
  declaration, not the file holding the call site.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 29 fail / 162 pass (was 31/160; both `Field type
  resolution` fixtures now pass — same-file and cross-file disambig).
- tsc --noEmit clean.

Remaining Unit 4 work (write ACCESSES, `self.X` for-loop iteration)
needs Unit 6's tuple/iterable destructuring before it can land —
`for u in self.users` requires the iterable typing path.

* feat(python-scope): chain receiver via call-expression return types

Unit 5 — extends the compound-receiver case to handle call-expression
receivers (`svc.get_user().save()`).

`resolveCompoundReceiverClass` is the single recursive entry point for
all compound receivers. Three shapes:
  - bare identifier — typeBinding chain
  - dotted `obj.field[.field]…` — class-scope field types
  - call `expr.method()` — recurse into expr, look up method's
    return-type typeBinding on its class scope

Method return-type bindings auto-hoist to the parent (class) scope per
Unit 3, so `methodClassScope.typeBindings.get(methodName)` is the
canonical lookup. Free-call return types (`get_user()`) walk the
caller's scope chain.

Depth-capped at 4 hops to bound recursion.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 28 fail / 163 pass (was 29/162; `Python chained method
  call resolution` now passes).
- tsc --noEmit clean.

Two related tests (`city.save() via method chain`, `c.greet().save()
depth-2 MRO`) still fail because the captures yield typeBindings
shaped like `city → user.get_city` (no trailing parens — the capture
grabs the attribute text). Resolving those needs a follow step that
detects the call-shape rawName and feeds it through the compound
recurser. Lands with the chain-typeBinding work in a follow-up.

* feat(python-scope): free-call fallback consults finalized bindings

Unit 7 — closes the cross-file free-call gap.

The shared `MethodRegistry.lookup` walks `scope.bindings` (pre-finalize
local-only) for free-call resolution. Cross-file imports land in
`indexes.bindings` (post-finalize). Without the dual-source lookup,
`from x import f; f()` resolves to "unresolved" and no CALLS edge is
emitted.

Two changes:

- `emit-core/scope-walkers.ts`: new `findCallableBindingInScope` —
  same dual-source pattern as `findClassBindingInScope`, but accepts
  Function/Method/Constructor. Promoted to emit-core because every
  language with cross-file imports needs the same lookup.
- `python-scope-emit.ts emitFreeCallFallback`: post-pass that walks
  every free-call reference site, looks up the callee with the new
  helper, and emits via `tryEmitEdge`. Pre-seeds `seen` from the
  shared resolver's emissions so we never double-count.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 22 fail / 169 pass (was 28/163; +6 tests including
  the Python overload dispatch fixtures, ancestor-directory imports,
  and same-name module-alias collision).
- tsc --noEmit clean.

* feat(python-scope): super() receiver dispatches up the MRO

Unit 8 — `super().method()` inside a class method walks the enclosing
class's MRO chain (skipping self) and resolves to the first ancestor
that owns the method.

New receiver branch in `emitReceiverBoundCalls` recognizes
`super(...)` syntactically (regex-cheap), finds the enclosing class
via a new `findEnclosingClassDef` scope-walk helper, then re-uses
`scopes.methodDispatch.mroFor` + `findOwnedMember` from the existing
class-receiver path. Handled before the compound-receiver case so
`super()` doesn't fall into the bare-identifier branch where `super`
isn't a binding.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 21 fail / 170 pass (was 22/169; `super().save() inside
  User to BaseModel.save` now passes).
- tsc --noEmit clean.

* feat(python-scope): suppress shared resolver on member-call sites

Unit 9 — `app_metrics.get_metrics()` (namespace import alias) was
emitting two CALLS edges: a wrong self-call from the shared
resolver's free-call fallback, plus the correct namespace-receiver
edge from the Python post-pass.

Mechanism:

- `emit-core/emit-references.ts`: new optional `skipSites` parameter
  (`Set<string>` of `${filePath}:${line}:${col}` keys). When supplied,
  references at those positions are skipped — the provider has
  already emitted (or chosen not to emit) for that site.
- `python-scope-emit.ts`: reorders Phase 4 — receiver-bound + free-
  call fallback run FIRST, populating `handledSites`. The shared
  `emitReferencesViaLookup` then runs with that set so the resolver's
  fallback can't fight a precise per-receiver emission. Site keys are
  added only on successful tryEmitEdge (not for sites the post-pass
  saw but couldn't resolve — those still get a chance from the shared
  path).

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 20 fail / 171 pass (was 21/170; same-name module-alias
  collision now resolves correctly).
- tsc --noEmit clean.

* feat(python-scope): propagate return-type bindings across imports

Closes the cross-file return-type propagation gap that left tests
like `u = get_user(); u.save()` (where get_user lives in another
file) with `u` typed as the function name instead of its return type.

The shared finalize pass copies callable bindings (`from x import f`
puts `f` in the importer's bindings) but typeBindings stay file-local
because they live on `Scope.typeBindings`, not on the index. Mutate
post-finalize:

- For each module-scope import binding (`origin: 'import'` or
  `'reexport'`), look up the source file's module-scope typeBinding
  for the def's simple name. If present (return-annotation source),
  mirror it under the importer's local alias. Skip when the importer
  already has its own typeBinding for the name (explicit local always
  wins).
- After propagation, re-run a chain-follow on every scope's
  typeBindings — pass-4 ran before propagation and missed any chain
  whose terminal lived in a foreign file. Same algorithm as
  `followChainedRef` in scope-extractor, but operates on the
  finalized scopes so propagated entries are visible.

Mutating `Scope.typeBindings` is safe — `draftToScope` constructs a
plain `new Map(...)`, not a frozen one.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 16 fail / 175 pass (was 20/171; +4 — both cross-file
  return-type tests, plus two related propagation cases).
- tsc --noEmit clean.

* feat(python-scope): for-loop call-iterable typeBinding

Adds `(for_statement left: (identifier) right: (call function:
(identifier)))` to the typeBinding capture set. Combined with Unit 3's
return-type capture and the cross-file return-type propagation pass,
this makes `for u in get_users(): u.save()` resolve to `User.save`
even when `get_users` is imported from another module.

Captured as `@type-binding.alias` (rawName = function identifier,
without parens) so the existing chain-follow walks the alias to the
function's return-type binding without any new code path.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 12 fail / 179 pass (was 16/175; +4 for-loop call-iterable
  tests across get_users / get_repos fixtures).
- tsc --noEmit clean.

* feat(python-scope): collapse free-call edges per (caller, target)

Free calls (no explicit receiver) now emit a single CALLS edge per
(caller, target) pair regardless of how many call sites the caller
contains. Mirrors the legacy DAG's per-pair dedup contract — what
the `default-params`, `variadic`, and `overload` fixtures expect.

Member calls keep position-based dedup so distinct resolved targets
(e.g. UserService.find_user vs AdminService.find_user from the same
caller) still produce distinct edges.

Implementation: bypass `tryEmitEdge` (which dedupes positionally) and
hand-roll the relationship with a position-independent rel.id
(`rel:CALLS:<caller>-><target>`). Site handling is now unconditional —
even when the dedup-collapse skips the actual emit, we mark the site
handled so the shared `emit-references` doesn't fight us with its
fallback.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 10 fail / 181 pass (was 12/179; +2 — both `default
  parameter arity` tests now pass).
- tsc --noEmit clean.

* fix(python-scope): match legacy CALLS reason for import-resolved free calls

The arity-narrowing test asserts \`rel.reason === 'import-resolved'\`
for cross-file free-call edges. Switch the free-call fallback's
reason to mirror legacy DAG semantics:
  - target-file !== source-file → 'import-resolved'
  - same file                   → 'local-call'

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 9 fail / 182 pass (was 10/181; +1 arity-narrowing test).
- tsc --noEmit clean.

* fix(python-scope): drop dead pre-seeding from receiver-bound pass

The pre-seeding loop at the top of \`emitReceiverBoundCalls\` populated
\`seen\` with every reference the shared resolver had already resolved.
That was useful when emit-references ran FIRST. After Unit 9 reversed
the order (emit-references runs after the Python passes and uses
\`handledSites\` to skip what we processed), the pre-seed only causes
harm: when an MRO walk in Case 0 (compound receiver) and Case 4
(simple typeBinding) both touch the same site at the same position
but resolve to different targets, the pre-seed suppresses the second
emission because the shared resolver had already entered the wrong
target into \`seen\`.

Concrete case: \`c.greet().save()\` — Case 0 emits the outer save edge
to Greeting.save; Case 4 then resolves the inner \`c.greet()\` to
A.greet via MRO walk. With pre-seed both edges should emit (different
targets, different rel.ids); without removing the pre-seed the inner
emission was being deduped against an already-seeded entry and the
A.greet edge was lost.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 8 fail / 183 pass (was 9/182; +1 — \`c.greet() to A#greet
  via MRO walk\` now passes).
- tsc --noEmit clean.

* feat(python-scope): enumerate(X) for-loop tuple destructuring

Adds two new typeBinding capture patterns for the canonical enumerate
pattern:

  for (i, u) in enumerate(users): ...   ; tuple_pattern
  for  i, u  in enumerate(users): ...   ; pattern_list

Both bind the second tuple element (u) to the iterable identifier
(users). The chain-follow then unwraps users → its element type via
the existing generic-strip in interpret.ts (List[User] → User).

The #eq? predicate scopes the pattern to enumerate specifically;
generic tuple destructuring of arbitrary callables is left to a
future iteration once we have a richer signal for "what does this
call yield".

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 7 fail / 184 pass (was 8/183; +1 — `parenthesized tuple:
  for (i, u) in enumerate(users)` now passes).
- tsc --noEmit clean.

* feat(python-scope): dict.items() value-type unwrapping

Two changes that together resolve `for k, v in data.items(): v.save()`:

- `interpret.ts stripGeneric`: extends to `dict[K, V]` /
  `Dict[K, V]` / `Mapping[K, V]` etc., stripping to the value type V.
  Previously only single-arg generics (list[User] → User) were
  stripped; multi-arg ones returned the raw text.
- `query.ts` + `scopes.scm`: new typeBinding patterns for
  `for k, v in X.items()` (both pattern_list and tuple_pattern). The
  second tuple element binds to X; the chain-follow then unwraps X's
  dict annotation to V via the new stripGeneric branch.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 6 fail / 185 pass (was 7/184; +1 — `dict.items() loop`
  test now passes).
- tsc --noEmit clean.

* feat(python-scope): nested tuple destructuring for enumerate(d.items())

Two more for-loop typeBinding patterns:

- `for i, (k, v) in enumerate(d.items())` — nested tuple destructuring
  where v is the value of the dict's items() yield.
- `for v in d.values()` — explicit values() form (companion to items).

Both bind the loop var to the dict identifier; the chain-follow
unwraps via the dict-aware stripGeneric to the value type.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 5 fail / 186 pass (was 6/185; +1 nested tuple test).
- tsc --noEmit clean.

* feat(python-scope): 3-var flat destructuring for enumerate(d.items())

Adds the \`for i, k, v in enumerate(d.items())\` shape — flat
3-variable destructuring of the (i, (k, v)) tuple yielded by
\`enumerate\` over \`items()\`. Binds v (the last identifier in the
pattern_list) to the dict identifier; the existing dict-aware
stripGeneric unwraps to the value type.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 4 fail / 187 pass (was 5/186; +1).
- tsc --noEmit clean.

* feat(python-scope): write ACCESSES edges for attribute assignments

Three changes that together produce ACCESSES (write) edges for
\`obj.field = value\` assignments:

- New \`@reference.write.member\` capture in query.ts and scopes.scm
  matching \`(assignment left: (attribute object: ... attribute: ...))\`.
  Reuses the existing receiver/name capture shape so the
  receiver-bound emit pass can resolve obj's class and look up the
  field.
- \`populateMethodOwnerIds\` now sets ownerId on class-body fields too,
  not only on methods. Previously it only walked Function scopes
  whose parent was Class; class-body annotations like \`name: str\`
  live directly in the Class scope's ownedDefs and were missed, so
  \`findOwnedMember(User, "name")\` returned undefined.
- \`emit-core isLinkableLabel\` extends to Variable and Property so
  field nodes appear in the graph-node lookup (the legacy parser
  emits both kinds for class-body annotations).
- Case 4 in receiver-bound pass now uses the kind word as the edge
  reason for read/write sites — matches the legacy DAG convention
  the test asserts on.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 3 fail / 188 pass (was 4/187; +1 — write-ACCESSES test).
- tsc --noEmit clean.

* feat(python-scope): chain-typebinding + field-fallback method lookup

Reaches the architectural-plan target of >= 189/191 flag-on passing.

Two intertwined changes:

- Field-fallback in resolveCompoundReceiverClass: when method lookup
  on the receiver's class (and its MRO) fails, walk the class's
  fields and try the same lookup on each field's type. Matches the
  "unified fixpoint" intent of the method-chain fixture where
  `user.get_city()` reaches `Address.get_city` through User's
  `address: Address` field.
- New Case 3b in receiver-bound emit pass: when the receiver's
  typeBinding rawName has a dot but isn't a namespace prefix
  (e.g. `city -> user.get_city` from the constructor-inferred capture
  for `city = user.get_city()`), treat it as a method-call chain and
  pipe through the compound resolver. The chain unwraps to the
  terminal class (City) and the call resolves normally.

Verification:
- Flag-off: 191/191 (identical baseline).
- Flag-on: 2 fail / 189 pass (was 3/188; +1 city.save method chain).
- tsc --noEmit clean.

Remaining 2 failures are fixture-driven (self.users / self.repos
fixtures reference fields that aren't declared on the class) and
documented as known-limitation in Unit 10.

* feat(python-scope): flip Python to registry-primary (191/191 parity)

Adds the \`for u in self.X\` heuristic typeBinding capture (binds u to
the attribute name X so the chain-follow can resolve via the enclosing
method's parameter typeBinding) — closes the last two failing
fixtures whose classes reference \`self.X\` for fields that are
actually method parameters.

With 191/191 passing on BOTH legacy and registry-primary paths,
flips \`MIGRATED_LANGUAGES\` to include \`SupportedLanguages.Python\`.

Effects:
- Production default for Python files: registry-primary path.
- CI parity gate auto-discovers Python via the script + workflow
  (\`scripts/ci-list-migrated-languages.ts\` /
  \`.github/workflows/ci-scope-parity.yml\`) and runs the resolver
  integration test BOTH ways on every PR.
- Operators retain the \`REGISTRY_PRIMARY_PYTHON=0\` escape hatch.

Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- Default (unset, post-flip): 191/191 (uses registry).
- tsc --noEmit clean.

This concludes RFC #909 Ring 3 — Python migration.

* refactor(emit-core): EmitProvider interface + promote 5 generic helpers

G-Units 1-2 of the emit-pipeline generalization plan.

Adds:
- emit-core/emit-provider.ts — typed EmitProvider contract (6 required +
  2 optional fields). Will be consumed by the generic orchestrator in
  G-Unit 6. Documents the LanguageProvider vs EmitProvider boundary.
- emit-core/emit-free-call.ts — emitFreeCallFallback promoted as-is
  (drops the unused referenceIndex pre-seed parameter; underscore-prefixed
  to keep the signature compatible).
- emit-core/propagate-return-types.ts — propagateImportedReturnTypes +
  followChainPostFinalize. Documents the mutation contract (Invariant
  I3 + I6 from the plan): runs after finalize, before resolve, mutates
  the non-frozen Scope.typeBindings map.
- emit-core/scope-walkers.ts: + findEnclosingClassDef +
  findExportedDefByName. Both were already generic in the Python
  source.

python-scope-emit.ts shrinks 1055 → 799 lines (–256). Imports the
promoted helpers from emit-core. No behavior change.

Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- tsc --noEmit clean.

* refactor(emit-core): promote receiver-bound dispatcher + compound resolver

G-Unit 3 of the emit-pipeline generalization plan.

- emit-core/emit-compound-receiver.ts — resolveCompoundReceiverClass
  + matchingOpenParen + COMPOUND_RECEIVER_MAX_DEPTH. Field-fallback
  is now an option (default true) so strictly-typed languages can
  opt out via EmitProvider.fieldFallbackOnMethodLookup.
- emit-core/emit-receiver-bound.ts — the 7-case dispatcher (super,
  Cases 0/1/2/3/3b/4). Accepts a ReceiverBoundProviderSubset
  (isSuperReceiver + fieldFallbackOnMethodLookup) so partial wiring
  works during the rest of the migration. Documents Contract
  Invariants I4 (case order) and I5 (no pre-seeding).

python-scope-emit.ts shrinks 799 → 384 lines. The orchestrator now
calls the generic emitReceiverBoundCalls with an inline minimal
provider (pythonEmitProviderInline) — full provider lands in G-Unit 6
when the orchestrator itself moves to languages/python/emit/.

Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- tsc --noEmit clean.

* refactor(emit-core): promote MRO walk + populateClassOwnedMembers

G-Units 4-5 of the emit-pipeline generalization plan.

- emit-core/build-mro.ts — generic buildMro takes a LinearizeStrategy
  hook receiving (classDefId, directParents, parentsByDefId). Three
  shared steps (collect EXTENDS, build defId-by-graphId, walk per
  class) + parametric linearization. Default strategy is BFS-with-
  visited (Python's depth-first first-seen, also correct for
  single-inheritance languages).
- emit-core/scope-walkers.ts: + populateClassOwnedMembers — generic
  OO ownership rule (methods + class-body fields). Both rules ship
  together because every OO language migrated so far (Python; planned
  TS/JS/Java/Kotlin) wants both. Languages that need different rules
  can compose with this as a base step.

python-scope-emit.ts shrinks 384 → 255 lines.

Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- tsc --noEmit clean.

* refactor(scope-resolution): generic orchestrator + language-agnostic phase

G-Units 6-7 of the emit-pipeline generalization plan, plus the
pipeline-phase generalization (the user's observation that the phase
itself is generic once the orchestrator is).

Changes:

- emit-core/orchestrator.ts — runScopeResolution(input, provider).
  The 180 lines of pipeline glue moved here, parametrized by
  EmitProvider. Provider supplies LanguageProvider, importEdgeReason,
  and the 6 emit-side hooks.
- emit-core/emit-provider.ts — EmitProvider gains languageProvider
  and importEdgeReason fields so the orchestrator needs nothing else.
  resolveImportTarget now takes (targetRaw, fromFile, allFilePaths).
- languages/python/emit/index.ts — pythonEmitProvider + thin
  runPythonScopeResolution wrapper. The first reference impl every
  next-language migration copies.
- emit-providers-registry.ts (NEW) — registry of per-language
  EmitProviders keyed by SupportedLanguages. Adding a language is
  one line here + the provider file.
- pipeline-phases/scope-resolution.ts (NEW) — language-agnostic phase
  iterating EMIT_PROVIDERS ∩ MIGRATED_LANGUAGES. Replaces
  pipeline-phases/python-scope.ts (deleted).
- python-scope-emit.ts deleted.
- pipeline.ts swaps pythonScopePhase → scopeResolutionPhase.

The next language migration is now: implement EmitProvider, register
it, add to MIGRATED_LANGUAGES. No new pipeline phase, no orchestrator
copy-paste. The Python migration's 700+ lines of glue collapse to
~80 lines per future language.

Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- Default (post MIGRATED_LANGUAGES flip): 191/191.
- tsc --noEmit clean.

* docs(emit-provider): migration cookbook for next-language porters

* refactor(scope-resolution): rename emit-core/ → scope-resolution/, EmitProvider → ScopeResolver

Reorganizes the registry-primary resolution layer for clarity and
contributor onboarding. Driven by feedback that "emit" was triple-
overloaded (graph-edge emission + tree-sitter capture extraction +
the provider name itself), and the flat 16-file emit-core/ folder
mixed five concerns.

External research (rust-analyzer hir-def/nameres, Pyright analyzer/,
TypeScript binder/checker, Roslyn Binder, IntelliJ Resolver, swc
semantic/, biome semantic/, semgrep naming/, JDT Binding, clangd
Sema) consistently uses **the phase name** for this layer, never an
output verb. "Scope resolution" matches our pipeline-phase name, the
plan, and the RFC.

## Folder rename

  emit-core/                              → scope-resolution/
  ├── (16 flat files)                     → ├── contract/scope-resolver.ts
                                            ├── pipeline/{run,registry,phase}.ts
                                            ├── passes/{receiver-bound-calls,
                                            │           free-call-fallback,
                                            │           compound-receiver,
                                            │           imported-return-types,
                                            │           mro}.ts
                                            ├── graph-bridge/{node-lookup,ids,
                                            │                 edges,references-to-edges,
                                            │                 imports-to-edges,
                                            │                 method-dispatch}.ts
                                            └── scope/{walkers,namespace-targets}.ts

Each subfolder maps to one concern a new contributor needs to find:
*the contract I implement / the runner that calls me / the helpers I
reuse / the graph layer I shouldn't touch / the scope walkers*.

## Symbol renames

  EmitProvider                  → ScopeResolver
  pythonEmitProvider            → pythonScopeResolver
  runPythonScopeResolution      → resolvePythonScope
  EMIT_PROVIDERS                → SCOPE_RESOLVERS
  getEmitProvider               → getScopeResolver
  RunPythonScopeResolution{Input,Stats} → ResolvePythonScope{Input,Stats}

## File renames (per-language)

  languages/python/emit/index.ts → languages/python/scope-resolver.ts
  languages/python/emit-captures.ts → languages/python/captures.ts
                                     (kills the parse-side "emit" collision)

## Mechanics

- Used `git mv` for all files so blame history is preserved.
- Updated ~30 import lines across 18 files plus the pipeline-phases
  barrel and pipeline.ts.
- Updated JSDoc cross-references throughout to match the new vocabulary.

Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- Default (post MIGRATED_LANGUAGES flip): 191/191.
- tsc --noEmit clean.

Migration cookbook in `scope-resolution/contract/scope-resolver.ts`
JSDoc points the next-language porter at all the new names and
folder locations.

* docs(scope-resolution): finalize phase JSDoc + drop python emoji from generic log line

* perf(scope-resolution): O(1) workspace lookup index

Introduces `WorkspaceResolutionIndex` — a precomputed bundle of
lookup tables built ONCE per resolution run, after `populateOwners`
and after finalize, before any pass that needs to find members,
exported defs, or class scopes by id.

What it replaces (all are pre-existing O(N×D) linear scans of
parsedFiles, called inside the receiver-bound MRO chain):

- `findOwnedMember(ownerId, name, parsedFiles)` → `Map.get` via
  `index.memberByOwner.get(ownerId)?.get(name)`. Was the worst
  offender — receiver-bound dispatcher calls this O(sites × MRO
  depth) times.
- `findExportedDef(filePath, name, parsedFiles)` → `Map.get` via
  `index.defsByFileAndName`. Hot for namespace-receiver case.
- `findExportedDefByName` workspace-wide fallback scan → `Map.get`
  via `index.callablesBySimpleName`.
- `classScopeByDefId` (rebuilt inside `emitReceiverBoundCalls` on
  every invocation) — moved to one-shot build during finalize, read
  from `index.classScopeByDefId` everywhere.
- `moduleScopeByFile` (rebuilt inside `propagateImportedReturnTypes`
  on every invocation) — read from `index.moduleScopeByFile`.

Findings from a synthetic 100-file Python workload (60 model files
each defining 5 classes × 3 methods + 40 user files calling them
heavily):

  scope-resolution wall time: 764ms → 710ms (median, 5 iters)

That's a ~7% in-layer win. The smaller-than-expected gain was
informative: profiling the synthetic workload shows scope-resolution
breakdown is `extract=62% resolve=30% emit=4%`; the index touched
the 4% slice (emit + walker calls inside it). Larger O(D) per owner
classes will benefit more.

Profiling the FULL pipeline (49 fixtures × 3 iters) shows
scope-resolution accounts for ~1% of pipeline wall time — the
remaining 99% is parse (tree-sitter), heritage, ORM, MRO, processes,
and DB writes. So further optimization of this specific layer has
marginal pipeline impact; the next-biggest wins live in those
phases. Documented as the "double-parse" finding in the audit
(captures.ts re-parses each Python file even though the parse phase
already produced a tree-sitter Tree) — that's a separate plumbing
project across phase boundaries.

Bonus: opt-in PROF_SCOPE_RESOLUTION=1 env var prints a per-phase
ms breakdown to stderr, so future perf work can measure without
extra code changes.

Verification:
- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- tsc --noEmit clean.

* perf(parse/heritage/mro): typed graph iterator + cross-phase tree cache

Two structural perf wins targeting the parse / heritage / MRO
layers, identified by the post-WorkspaceResolutionIndex profiling
(scope-resolution = ~1% of pipeline; the bulk lives upstream).

## 1. KnowledgeGraph.iterRelationshipsByType (PHM-Units 1-2)

- Adds a per-type `Map<RelationshipType, Map<id, Relationship>>`
  index inside `createKnowledgeGraph`, maintained on add / remove /
  removeNode / removeNodesByFile.
- New `iterRelationshipsByType(type)` returns a typed iterator that
  yields only the requested type. Backwards-compatible: existing
  `iterRelationships()` / `forEachRelationship()` callers untouched.
- Migrated two MRO call sites:
  - `mro-processor.ts buildAdjacency`: split the single
    `forEachRelationship` (which scanned every edge in the graph and
    type-filtered per-iteration) into three typed iterations
    (EXTENDS, IMPLEMENTS, HAS_METHOD).
  - `scope-resolution/passes/mro.ts buildMro`: replaced
    `for (const rel of graph.iterRelationships()) if (rel.type !== 'EXTENDS') continue`
    with `for (const rel of graph.iterRelationshipsByType('EXTENDS'))`.
- Heritage-processor (PHM-Unit 3) was a no-op: it only WRITES
  EXTENDS/IMPLEMENTS edges, never re-reads. Index is still useful
  for the seven other graph-iter consumers (community-processor,
  csv-generator, wildcard-synthesis, process-processor, etc.) — those
  follow-ups can switch to the typed iterator without touching the
  graph layer.
- Adds 5 unit tests for the new method (add/remove/dedupe semantics,
  empty-type fresh iterator, removeNode index sync).

## 2. Cross-phase tree cache (PHM-Units 4-5)

The audit's #2 finding: Python files are parsed by tree-sitter once
in the parse phase, then re-parsed inside scope-resolution's
`captures.ts`. Eliminate the second parse by sharing the Tree across
phases.

- `parse-impl.ts` now maintains TWO ASTCaches with distinct lifetimes:
  - `astCache` (chunk-local, cleared between chunks) — unchanged;
    used by call/heritage/import processors during parse.
  - `scopeTreeCache` (total-parseable-sized, never cleared) — new,
    exposed via `ParseOutput.astCache` for cross-phase consumption.
- `parsing-processor.ts` writes every sequentially-parsed Tree to
  BOTH caches. Worker-mode parses skip the persistent cache too
  (Trees can't cross MessageChannels).
- `LanguageProvider.emitScopeCaptures` gains an optional `cachedTree`
  parameter (typed `unknown` to keep the tree-sitter dep out of the
  contract).
- `captures.ts` short-circuits its own `parser.parse(sourceText)`
  when a cached Tree is supplied. Cache miss falls back to a fresh
  parse — same correctness path as before.
- `runScopeResolution` accepts an optional `treeCache` and forwards
  per-file `cachedTree` to `extractParsedFile`.
- `scope-resolution/pipeline/phase.ts` reads
  `getPhaseOutput<{astCache}>(deps, 'parse')` and passes through.

Verified end-to-end: a small fixture run with PROF_SCOPE_RESOLUTION=1
shows 6/6 cache hits (100% hit rate) on the python-grandparent fixture
that exercises the full pipeline below the worker-pool threshold.

## Verification

- REGISTRY_PRIMARY_PYTHON=0 (legacy): 191/191.
- REGISTRY_PRIMARY_PYTHON=1 (registry): 191/191.
- New graph.test.ts: 25/25 (was 20).
- tsc --noEmit clean.

## Where the win lands

Wall-clock on the 49-fixture integration suite: 14050ms → 14080ms
(within noise). Fixtures are 1-3 files each, dominated by per-fixture
pipeline overhead (worker-pool init, DB writes, fixture startup).
The cache + typed-iterator wins are constant-factor improvements
that scale linearly with workload size and visible only on larger
repos. The dev-mode `PROF_SCOPE_RESOLUTION` instrumentation +
`getPythonCaptureCacheStats()` are kept for future perf work.

## Plan

docs/plans/2026-04-20-002-perf-parse-heritage-mro-plan.md.
PHM-Unit 3 (heritage-processor migration) intentionally collapsed
to a no-op — heritage only writes, never re-reads.

* perf(scope-resolution): bound tree-cache lifetime + gate population

Address P1 residuals from ce:review of 8c6f5cee:

- Dispose scopeTreeCache at end of scopeResolutionPhase via
  astCache.clear(). Trees were previously retained for the full
  pipeline (10-100x memory regression on large repos). Downstream
  phases (mro, community, csv-generator) never read them.
- Gate scopeTreeCache.set on provider.emitScopeCaptures !== undefined.
  Polyglot repos no longer retain Trees for languages with no
  scope-resolution consumer.
- PROF_SCOPE_RESOLUTION=1 now warns when workers engage, since
  Trees can't cross MessageChannels so the cache will be empty for
  worker-parsed files — prevents a silent perf cliff once a repo
  crosses the worker-pool threshold.

Tests: 26/26 graph unit, 299/299 scope-resolution unit, 191/191
python integration both flag paths.

* refactor(scope-resolution): clean up P2/P3 review residuals

P2:
- WASM dual-ownership invariant documented on ASTCache dispose:
  a Tree must live in AT MOST ONE disposing ASTCache. Native
  tree-sitter today is unaffected; WASM adoption would require
  tree.copy() or a non-disposing secondary cache.
- mro-processor C3 ordering test: pins EXTENDS-before-IMPLEMENTS
  parent grouping for classes with interleaved edge additions.
  Asserts exact MRO ['Base', 'Iface'] — a revert to single-loop
  insertion-order iteration would produce ['Iface', 'Base'] and
  fail loudly.
- cached-tree parity test: emitPythonScopeCaptures(src, path, T)
  returns identical CaptureMatch[] to emitPythonScopeCaptures(src,
  path). Pins the cache-hit path's correctness so a regression
  that silently returns stale captures would break the test.

P3:
- Dev-mode cache counters moved from captures.ts to cache-stats.ts.
  Production hot-path module no longer carries the module-global
  export surface; PROF gating behavior preserved.
- ParseOutput field rename astCache → scopeTreeCache. Clarifies
  that the surfaced cache is the persistent cross-phase one, not
  the chunk-local astCache parse-impl clears between chunks.
  Single consumer (scopeResolutionPhase) updated; no other readers.
- ASTCacheReader interface extracted. scopeResolutionPhase now
  reads the phase dep via a shared type instead of a hand-rolled
  inline structural shape that could drift from ASTCache's contract.
- graph.ts dual-index invariant enforced through writeRel/deleteRel
  private helpers instead of duplicated add/delete at 3 mutation
  sites. Adding a new mutation method only needs to call the
  helpers — forgetting to update one index becomes structurally
  impossible.

Tests: 382/382 unit (incl. 2 new), 191/191 python integration both
flag paths. tsc clean.

* fix(ci): prettier formatting + Python-migration test adjustments

CI run 24666612657 failed on three jobs. Fixes:

quality/format:
- Prettier --check flagged 3 files after the accumulated branch work.
  Ran prettier --write from repo root (CI's invocation cwd) to apply:
  simple-hooks.ts, resolve-references.ts, python-hooks.test.ts.

tests/{ubuntu,macos,windows} — 9 assertion failures, all traceable to
Python landing in MIGRATED_LANGUAGES (default-on registry-primary):

  - registry-primary-flag.test.ts (3 tests): the 'returns false by
    default' / 'primaryLanguages empty' / 'Python mid-process
    mutation' assertions were written in Ring 2 when MIGRATED_LANGUAGES
    was empty. Rewrote to assert MIGRATED_LANGUAGES membership is the
    default, use Java (unmigrated) for the no-stale-cache test, and
    verify env overrides work in both directions (migrated-off,
    unmigrated-on).
  - call-processor.test.ts (6 tests in SM-10 + D2-widen blocks):
    these exercise the LEGACY call-resolution DAG on .py fixtures.
    processCalls now gates Python out (isRegistryPrimary === true by
    default), returning 0 edges. Added REGISTRY_PRIMARY_PYTHON=false
    override in the relevant beforeEach + restore in afterEach, so
    the legacy DAG runs for these test-local fixtures without
    affecting the production-default behavior.

Local verification: 4126/4126 unit tests pass, prettier clean.

* docs(python): known-limitation block on scope-resolution public API

Unit 10 — document what the Python registry-primary path intentionally
does not resolve, so reviewers and future maintainers can distinguish
conscious trade-offs from latent bugs:

- Dynamic attribute access (getattr / setattr)
- Dynamic imports (importlib, __import__)
- Metaclass-driven dispatch
- Union / Optional branch-picking behavior
- Arbitrary signature-rewriting decorators
- typing.TYPE_CHECKING-guarded imports
- *args / **kwargs type flow-through
- super() outside a directly-bound method

Each item names the file that owns the relevant hook so a future
follow-up knows where to start. Shadow-harness corpus parity + the
CI parity gate remain the authoritative signal for which of these
matter at fleet scale.

* docs: record scope-resolution pipeline alongside legacy call DAG

Capture what shipped in #980 so future readers don't have to reverse-
engineer the coexistence of the legacy call-resolution DAG and the new
scope-resolution pipeline:

- ARCHITECTURE.md: new 'Scope-Resolution Pipeline' section after the
  Call-Resolution DAG, documenting pipeline stages, ScopeResolver
  contract, per-language registration, code references, and perf
  notes. Coexistence block added to the legacy DAG section explaining
  how MIGRATED_LANGUAGES gates the two paths per-language.
- AGENTS.md: reference-docs pointer updated — legacy-DAG one-liner
  stays; scope-resolution pipeline gets its own pointer so agents
  know when to read which section. Changelog bumped.
- type-resolution-system.md: callout at the 'call-processor.ts is
  the consumer' claim pointing readers to the scope-resolution path
  for migrated languages. TypeEnv is still built per file, but for
  migrated languages receiver typing flows through ParsedTypeBinding
  rather than call-processor.ts.

CHANGELOG.md intentionally not touched — owned by the release process.

* chore: remove obsolete scheduled_tasks.lock file

* fix(scope-resolution): qualified-name keys for same-file method collisions

Review feedback from PR #980 reviewer flagged a BLOCKING correctness
bug: when two classes in the same file define a method with the same
simple name (e.g. class User: def save + class Document: def save),
every d.save() CALLS edge silently resolved to User.save because the
graph node lookup keyed only by (filePath, simpleName) and first-wins
took User's method.

Three-layer fix:

1. populateClassOwnedMembers now promotes a nested def's
   qualifiedName from `save` to `ClassName.save` when the def sits
   inside a class scope. Python's scopes.scm doesn't emit
   @declaration.qualified_name for methods, so without this the
   finalized SymbolDefinition carried only the simple name.
2. buildGraphNodeLookup adds a second key per node:
   (filePath, qualifiedName). For Method/Function nodes the qualifier
   is parsed deterministically out of the node id
   (`Method:file.py:User.save#N` → `User.save`), which is robust to
   Windows-style filePath colons. Simple-name key retained as a
   fallback for callers that don't know the qualifier.
3. resolveDefGraphId now tries the qualified key first, then falls
   back to the simple-name lookup.

Also addresses the non-blocking review items:

- scopeResolutionPhase.deps now includes `crossFile` so the Kahn's
  runner can't schedule scope-resolution before crossFile finishes
  writing heritage edges that buildMro consumes.
- run.ts no longer mutates the finalized ScopeResolutionIndexes via
  `as` cast — spreads into a fresh object with the populated
  methodDispatch field instead.
- Doc nits: scope-resolver.ts registry path + phase.ts Ring number.

Test coverage:
- New fixture test/fixtures/lang-resolution/python-same-file-method-collision
  with User.save + Document.save in one file and app.py calling both
  through typed receivers.
- Three new integration assertions pin that u.save() and d.save()
  target the correct qualified node id. Fail before the fix, pass
  after. Confirmed by running once without populateClassOwnedMembers
  qualifier promotion — reproduces the original User.save-for-both bug.

Verification: 194/194 test/integration/resolvers/python.test.ts pass
both REGISTRY_PRIMARY_PYTHON=0 and =1. 523/523 related unit tests.
tsc --noEmit clean.

* fix(scope-resolution): filter export index to module-level defs + label-prefixed qualified key

Codex adversarial review on PR #980 flagged that
buildWorkspaceResolutionIndex feeds defsByFileAndName and
callablesBySimpleName from parsed.localDefs — the flat set of every
def in the file including methods, fields, and nested functions.
findExportedDef / findExportedDefByName treat those maps as
file-level exports, so `mod.save()` could silently bind to User.save
whenever a method's simple name appeared first in parse order.

Plan: docs/plans/2026-04-21-001-fix-workspace-index-module-scope-only-plan.md

Fix layers:

1. workspace-index.ts: split the single parsed.localDefs loop into
   two passes:
   - Module-export pass: iterate moduleScope.ownedDefs PLUS ownedDefs
     of every child scope whose parent is the module scope. Top-level
     class and function declarations each live in their own scope
     with parent=module, not in moduleScope.ownedDefs directly, so
     the "parent === moduleScope.id" walk is required to reach them.
     Methods (scope.parent === Class scope) and nested functions
     (scope.parent === another Function scope) are excluded.
   - Member-by-owner pass: keeps iterating parsed.localDefs since
     that map is keyed on ownerId and correctly saw class-owned defs
     before this change.

2. graph-bridge/node-lookup.ts: qualified keys now live in a separate
   keyspace (`<q>:filePath::<label>::<qualifiedName>`) and include
   the node label. Without the label prefix, a top-level `def save`
   (Function, qualifier `save`) would collide with a class method
   `User.save` (Method, simple name `save`) in the same simple-key
   slot because the Function's qualifier happens to equal the
   Method's simple name. The label differentiates them.

3. graph-bridge/ids.ts: resolveDefGraphId uses the new
   type-prefixed qualified key when def.type is set. Simple-name
   fallback retained for languages that don't yet synthesize
   qualifiers on their defs.

Test fixture: python-module-export-vs-method-collision places
`class User: def save` BEFORE top-level `def save` — parse order
that exposes the bug (class method enters the index first). Three
new integration assertions:
  - `mod.save(x)` resolves to the module-level Function, not User.save
  - `u.save()` resolves to User.save Method
  - Exactly two CALLS edges to `save` exist, one per intended target

Fixture confirmed failing before the workspace-index fix (bug
reproduced), passing after.

Verification: 197/197 test/integration/resolvers/python.test.ts pass
both REGISTRY_PRIMARY_PYTHON=0 and =1. 523/523 related unit tests.
tsc --noEmit clean.

* fix(scope-resolution): drive module export index from moduleScope.bindings

Codex round-2 adversarial review flagged that the workspace-index
module-export pass iterated every def in every direct-child scope of
the module, including class-body Variable defs like
`class User: MAX_USERS = 100`. `defsByFileAndName[file][MAX_USERS]`
silently aliased to the class attribute. Latent today because Python
doesn't emit ACCESSES edges for `mod.NAME` member access, but the
index-layer leak would surface the moment reference capture widens.

Plan: docs/plans/2026-04-21-002-fix-codex-round2-scope-resolution-plan.md

Drive the module-export index from the extractor invariant instead of
a scope-kind → allowed-label switch:

moduleScope.bindings already contains exactly the names visible at
module level — top-level class/function declarations, module-level
variable assignments, imports. Class methods, class-body attributes,
and nested-function defs bind to their containing (Class or Function)
scope, not the module, so they're naturally excluded.

Filter to `BindingRef.origin === 'local'` so imports and wildcard
re-exports stay out of the index (matches the pre-fix invariant when
the source was `parsed.localDefs`).

No per-kind predicates, no scope-kind / def-kind enumeration, no
two-pass merge between moduleScope.ownedDefs and direct-child scope
walks — one loop, language-agnostic.

Codex also flagged `propagateImportedReturnTypes` as potentially
broken for function-local imports, but scope-dump probing showed the
finalize algorithm puts `from svc import get_user` into the MODULE
scope's finalized bindings even when declared inside a function, so
the existing module-scope propagation already handles the case. The
new python-function-local-import-chain integration test pins that
working behavior as a regression guard; no code change required.

Coverage:
- test/unit/scope-resolution/workspace-index.test.ts (new, 5 tests) —
  directly asserts the index shape. The "excludes class-body Variable
  defs" test fails without this fix and passes after (confirmed via
  stash-pop probe).
- test/integration/resolvers/python.test.ts — 4 new integration
  assertions across two describe blocks (python-class-attr-export-leak,
  python-function-local-import-chain) pin end-to-end invariants.
- Two new fixtures under test/fixtures/lang-resolution/.

Verification: 201/201 test/integration/resolvers/python.test.ts both
REGISTRY_PRIMARY_PYTHON=0 and =1. 528/528 related unit tests (was
523). tsc clean.

* test(scope-resolution): pin local-namespace-import behavior + document empirical finalize hoisting

Codex round-3 adversarial review raised three concerns about
scope-resolution passes assuming module-scope semantics that would
contradict `pythonImportOwningScope`'s documented per-scope contract.
Empirical verification via scope-dump probes resolved each:

Plan: docs/plans/2026-04-21-003-fix-codex-round3-scope-aware-resolution-plan.md

1. Function- and class-local namespace imports: VERIFIED WORKING.
   `def outer(): import svc as s; s.call()` and `class A: import mod;
   def use(self): mod.helper()` both emit CALLS edges with reason
   "scope-resolution: namespace-receiver". finalize-algorithm hoists
   the ImportEdges onto `indexes.imports[moduleScope]` regardless of
   where the `import` statement appears, so collectNamespaceTargets'
   module-scope read finds them.

2. Imported return-type propagation module-scope-only: VERIFIED
   WORKING (already pinned in round 2). `from svc import get_user`
   inside a function body lands in indexes.bindings[moduleScope], so
   propagateImportedReturnTypes' module-scope read still finds it.

3. Nested method-local defs stamped as class members: VERIFIED FALSE.
   The scope extractor creates nested Function scopes for inner
   `def`s; `def helper` inside `def save` inside `class User` lives
   in helper's own Function scope whose parent is save's Function
   scope (NOT the Class scope). populateClassOwnedMembers'
   `parentScope.kind === 'Class'` branch correctly skips it;
   helper.ownerId stays undefined.

Instead of implementing speculative scope-aware refactors that the
tests would pass regardless, this commit:

- Adds regression fixtures and integration assertions that pin each
  working behavior. If finalize routing ever changes to honor the
  hook's per-scope contract, these assertions flip red and signal the
  need for the scope-chain-aware refactor.
- Adds defensive JSDoc to the three flagged call sites
  (collectNamespaceTargets, propagateImportedReturnTypes,
  populateClassOwnedMembers) documenting the empirical invariant so
  future reviewers don't re-derive Codex's theoretical concern
  without the benefit of the probe.

Files:
- Two new fixtures under test/fixtures/lang-resolution/ covering the
  function-local and class-body namespace-import patterns.
- Two new describe blocks in test/integration/resolvers/python.test.ts
  (3 assertions, positive-pin intent).
- Defensive comments in namespace-targets.ts, imported-return-types.ts,
  and scope-resolution/scope/walkers.ts.

Verification: 204/204 test/integration/resolvers/python.test.ts both
REGISTRY_PRIMARY_PYTHON=0 and =1. tsc clean.

* perf(graph): reverse-adjacency + file indexes drop removeNode/removeNodesByFile from O(N)

PR #980 in-line review flagged that `removeNode` iterated the full
relationshipMap to find edges touching a node (O(E)), and
`removeNodesByFile` called removeNode for every matching node after
a full nodeMap scan (O(N × E)). Pre-existing, but worth fixing
properly since the writeRel/deleteRel helpers we just added make the
index-maintenance story coherent.

Two new indexes maintained on every mutation path:

- `edgeIdsByNode: Map<nodeId, Set<relId>>` — reverse adjacency. Every
  edge records both endpoints, so removeNode iterates
  edgeIdsByNode.get(id) instead of every relationship. Self-edges
  skip the duplicate-endpoint write to keep the Set dedup explicit.
- `nodeIdsByFile: Map<filePath, Set<nodeId>>` — file index.
  removeNodesByFile reaches its file's nodes directly.

Complexity:
- removeNode: O(edges-touching-node), was O(total-edges).
- removeNodesByFile: O(file-nodes × avg-edges-per-node + scan of the
  file bucket), was O(total-nodes + file-nodes × total-edges).

Index maintenance is centralized in writeRel/deleteRel + new
addToBucket/removeFromBucket helpers. Empty buckets are pruned to
keep the indexes compact. Existing dual-invariant (relationshipMap ↔
relationshipsByType) preserved.

Nodes without a `filePath` property (e.g. Community/Cluster nodes)
are intentionally NOT indexed in nodeIdsByFile — they can't belong
to any file, so removeNodesByFile correctly leaves them alone.

Coverage: 7 new unit tests (33/33 total, was 26). Added cases:
- removes only edges touching the removed node
- handles self-edges
- removes orphan node with no edges
- removeNodesByFile removes only matching nodes
- returns 0 when no match
- also removes edges whose endpoints lived on the removed file
- does not index nodes without a filePath property

Verification: 204/204 test/integration/resolvers/python.test.ts both
REGISTRY_PRIMARY_PYTHON=0 and =1. 4235/4235 unit tests. tsc clean.

* refactor(ingestion): merge python/ast-utils into utils/ast-helpers; iterative findNodeAtRange

python/ast-utils.ts held three language-agnostic helpers
(nodeToCapture, syntheticCapture, findNodeAtRange) plus two
duplicates of the shared utils version (findChildOfType ==
findChild; findIdentifierChild was unused). Consolidating into
utils/ast-helpers.ts so the next language migrating to the
scope-resolution pipeline imports from one place.

findNodeAtRange rewritten iteratively using an explicit stack.
Previous implementation was recursive — fine for shallow Python
trees today, but a landmine for languages with deeper nesting
(Kotlin sealed-hierarchy decomposition, Rust macro expansion,
etc.) and the task hooks explicitly call out "no recursion".
Children are pushed reverse-index so LIFO pop visits them
left-to-right; row-bound pruning preserves the prior early-skip
optimization (the `break` shortcut is replaced with `continue`
since a stack can't leverage ordered sibling termination).

findChildOfType consumers migrated to the existing findChild
helper. findIdentifierChild deleted — no callers remained.

Coverage: 204/204 test/integration/resolvers/python.test.ts both
REGISTRY_PRIMARY_PYTHON=0 and =1. 339/339 scope-resolution +
graph unit tests. tsc clean.

* refactor(scope-resolution): remove unused shouldShadow / shouldCreateScope hooks

Both LanguageProvider hooks were dead weight:

- `shouldShadow` had zero call sites — the interface declared it,
  Python implemented a trivial always-true no-op, but no consumer
  ever read it. The shadowing decision lives in pythonMergeBindings
  and the central merge algorithm, not in a per-scope predicate.
- `shouldCreateScope` had one call site in pass1BuildScopes but the
  only language implementing it (Python) always returned true. No
  producer ever emits a `@scope.block` for Python, so the hook's
  "declines to create" branch was unreachable. Other languages
  didn't implement it at all.

Removing both:

- Drops the interface declarations in language-provider.ts.
- Drops `shouldCreateScope` from ScopeExtractorHooks Pick and from
  the pass1BuildScopes conditional — the stack-based parent-resolve
  loop becomes unconditional.
- Drops pythonShouldShadow / pythonShouldCreateScope from simple-hooks,
  the Python index barrel, and the python.ts provider wiring.
- Drops the tests that exercised the removed hooks: one block-
  suppression scenario in scope-extractor.test.ts, one shouldCreateScope
  test in parse-worker-scope-integration.test.ts, and the
  pythonShouldShadow / pythonShouldCreateScope always-true assertions
  in python-hooks.test.ts. pythonBindingScopeFor's delegate-to-default
  test is preserved in its own describe block.

Shadowing itself is unchanged: pythonMergeBindings still runs, LEGB
ordering still applies, wildcard transparency is still handled via
the merge precedence rules. The hook API just no longer has a
vestigial per-scope toggle we decided not to use.

Verification: 204/204 test/integration/resolvers/python.test.ts both
REGISTRY_PRIMARY_PYTHON=0 and =1. 335/335 scope-resolution + graph
unit tests (was 339, net -4 after removing the hook-specific
assertions). tsc clean.

* refactor(scope-resolution): drop dead exports surfaced by knip

Knip flagged 44+ dead exports in the PR surface. Cleanup:

Barrel deletion:
- Remove src/core/ingestion/scope-resolution/index.ts entirely.
  It re-exported 30+ symbols but only one file
  (languages/python/scope-resolver.ts) imported from it, and only
  7 symbols. Matches the project's "no barrel re-exports" preference
  and removes a drift surface. scope-resolver.ts now imports from
  concrete files (passes/mro.ts, scope/walkers.ts, contract/...).

Dead functions/interfaces removed:
- resolvePythonScope + ResolvePythonScopeInput + ResolvePythonScopeStats
  in languages/python/scope-resolver.ts — never called. pipelinePhase
  reaches pythonScopeResolver via SCOPE_RESOLVERS, not via a
  per-language entry point.
- getScopeResolver in scope-resolution/pipeline/registry.ts — had zero
  callers. Consumers read SCOPE_RESOLVERS directly.

Exports demoted to module-internal (used only within their own file):
- PYTHON_SCOPE_QUERY (query.ts) + its re-export from python/index.ts
- PROF (cache-stats.ts)
- PythonArityMetadata (arity-metadata.ts)
- ReferenceSiteSkipSet (graph-bridge/references-to-edges.ts)
- ReceiverBoundProviderSubset (passes/receiver-bound-calls.ts)
- ResolveCompoundReceiverOptions interface (passes/compound-receiver.ts)
- matchingOpenParen function (passes/compound-receiver.ts)
- followChainPostFinalize function (passes/imported-return-types.ts)
- RunScopeResolutionInput + RunScopeResolutionStats (pipeline/run.ts)

Also removed:
- Redundant `export type { Scope }` re-export from contract/scope-resolver.ts
  (consumers import Scope directly from gitnexus-shared).

Verification: knip reports zero dead exports in PR-touched files.
204/204 test/integration/resolvers/python.test.ts both flag paths.
335/335 scope-resolution + graph unit tests. tsc clean.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-21 15:50:00 +01:00
azizur100389 bd271da7b7 feat(cli): gitnexus remove <target> to unindex a registered repo by name or path (#664) (#1003)
* feat(cli): gitnexus remove <target> to unindex a registered repo by name or path (#664)

Add a `remove` CLI command that deletes the `.gitnexus/` index AND
unregisters a repo from the global registry (~/.gitnexus/registry.json),
addressing the lifecycle gap flagged in #664: previously users had to
cd into the repo to run `clean`, and there was no path-based or
alias-based remove for an already-deleted working tree.

- New command `gitnexus remove <target> [-f|--force]`. `<target>` is
  alias / basename-derived name / remote-inferred name / absolute path.
- New helper `resolveRegistryEntry(entries, target)` in repo-manager.ts
  with path > name precedence; throws RegistryNotFoundError or
  RegistryAmbiguousTargetError (typed, `kind`-discriminated).
- Atomicity mirrors `clean`: fs.rm first, then unregisterRepo; partial
  failures self-heal on next `listRegisteredRepos({ validate: true })`.
- Idempotent on unknown targets (exit 0 with warning) per the #664
  spec: "behave atomically and idempotently so retries are safe".
- `--force` uses `clean`-style confirmation-skip semantics — distinct
  from `analyze --force` (pipeline re-index); here there is no pipeline
  so no conflation.
- 7 new unit tests cover resolver precedence, case sensitivity,
  ambiguity, and not-found hints; 2 integration tests cover the real
  CLI -> registry -> filesystem chain including the --allow-duplicate-name
  (#829) ambiguity case.

* fix(cli): canonicalize repo paths so remove/register match across platforms (#1003 review)

Address review feedback from @evander-wang and @magyargergo on PR #1003
plus the Windows + macOS CI failure (same root cause).

Problem:
- macOS: /var is a symlink to /private/var. `path.resolve` does NOT
  follow symlinks, so a child running analyze in /var/folders/X stores
  /private/var/folders/X (realpath from OS cwd) but an outer caller
  passing the symlink form misses.
- Windows: GitHub runners surface tmpdirs in 8.3 short-name form
  (RUNNERA~1) while process.cwd() returns the long form (runneradmin).
  Same divergence.

Fix: new `canonicalizePath(p)` helper wraps `path.resolve` plus
`fs.realpathSync.native`, falling back to `path.resolve` when the path
doesn't exist (preserves idempotent-on-missing semantics needed by
`remove <unknown>`). Applied at 3 call-sites — registerRepo,
unregisterRepo, resolveRegistryEntry — canonicalising BOTH the input
and each stored `entry.path` at compare time. That last bit is the
backward-compat story: registries written by older versions
(pre-canonicalisation) still match correctly, so we don't need a
migration script.

Test side: the ambiguous-target integration test now reads the path
from the registry snapshot rather than passing the outer `repoA`
variable directly, so it exercises the registry contract regardless of
which path form the platform stores. 4 new unit tests cover the helper
(idempotent, fallback-on-missing, absolute-for-relative) plus the
backward-compat resolver path.

* fix(cli): store resolved (non-canonical) path, compare via canonicalizePath (#1003 CI)

Follow-up to c5eceba0. The previous commit canonicalised the repo path
at BOTH write-time AND compare-time in registerRepo — that expanded
Windows 8.3 short names (RUNNER~1) to long names (runneradmin) when
storing `entry.path`. Pre-existing #829 unit tests that assert
`path.resolve(err.existingPath) === path.resolve(tmpPath)` then broke
because `tmpPath` is still short-form (path.resolve doesn't expand
8.3) while `entry.path` was long-form (canonicalizePath does).

Fix: split storage from comparison.
- entry.path stores `path.resolve(repoPath)` — whatever form the
  caller passed. `list` output and error messages show the path the
  user typed.
- All compare points (existing-entry lookup in registerRepo, the
  collision guard, unregisterRepo, resolveRegistryEntry path tier)
  canonicalise BOTH sides via `canonicalizePath`. That is where the
  /var ↔ /private/var and RUNNER~1 ↔ runneradmin divergence actually
  matters.

Net effect: storage is tolerant (preserves user input), matching is
strict (canonical-vs-canonical). Pre-existing #829 tests stay green
because `err.existingPath` is unchanged from what `path.resolve` gives
back; the cross-platform CI failure from #1003 stays fixed because
every comparison path goes through `canonicalizePath`.

* fix(cli): refuse destructive fs.rm when registry storagePath isn't <repo>/.gitnexus (#1003 review)

Address @magyargergo's inline review finding on remove.ts:89 and the
sibling vulnerability in clean.ts --all (caught during a pre-commit
safety audit). ~/.gitnexus/registry.json is a user-writable plain-text
file, so a corrupted or hand-edited entry could point storagePath at
the repo root (catastrophic: rm the working tree), an empty string
(→ cwd), a parent dir, or anywhere else. fs.rm(recursive: true,
force: true) on any of those is a runtime disaster.

- New UnsafeStoragePathError + exported assertSafeStoragePath() in
  repo-manager.ts. Pure lexical string check (Windows-case-
  insensitive) asserting entry.storagePath === path.join(entry.path,
  '.gitnexus').
- Guard wired into BOTH destructive registry-trusting sites:
  - remove.ts: exit 1 with actionable hint
  - clean.ts --all: skip the poisoned entry with a warning and
    continue (preserves existing per-repo error tolerance — one bad
    entry doesn't halt the batch)
- clean.ts default path and server/api.ts are safe-by-construction
  (they recompute storagePath from findRepo / getStoragePath rather
  than trusting the registry field).
- 8 unit tests cover the guard (valid, repo-root, parent, empty,
  unrelated, sibling, error payload, Windows case).
- 2 integration tests prove the full CLI path: remove-poisoned exits
  1 without touching the working tree; clean --all with a poisoned
  sibling entry cleans the good entry, skips the bad one, and leaves
  the poisoned repo intact.

* test(cli): assert full remove dry-run + success output shape (#1003 NIT)

Address the one NIT from the senior-reviewer pass on PR #1003: the
integration test was only checking for the "Run with --force" hint in
dry-run output, not verifying that the three actual console.log lines
(alias, repo path, storage path) appear. Same weak check on the
success-branch "Removed" output.

Tighten both assertions to toContain(alias), toContain(entry.path),
toContain(storagePath). Catches silent format regressions — e.g. a
future refactor that drops a console.log line or swaps
entry.name/entry.path in the output.

No code change; +20 test lines. All assertions in the happy-path
integration test now fire for a meaningful reason.
2026-04-21 11:52:59 +01:00
ivkond 0909a908ee fix(group): bubble local-impact phase errors in groupImpact (#1004) (#1007)
When the Phase 1 local-impact leg returned a structured { error: ... }
payload (missing symbol, graph-load failure, or an exception wrapped by
safeLocalImpact), runGroupImpact previously buried it inside a zero-hit
GroupImpactResult with empty cross / outOfScope arrays and risk 'UNKNOWN'.

Callers branch on top-level `error` (CLI, MCP wrapper), so the failure
path surfaced as a silent "no impact across the group" — a false
negative on a safety-critical blast-radius tool.

Fail closed: bubble the error as a top-level { error } prefixed with the
repoPath, matching how runGroupImpact already handles resolveGroupRepo,
config-load, and bridgePrep failures. Chose option 1 (bubble the error)
over option 2 (partial-result discriminant) because runGroupImpact only
runs local impact for a single member repo at this point — cross-repo
fan-out happens later via the bridge, so there is no partial success
data to preserve on the local-phase failure path.

Added two regression tests covering both the port-returned { error }
case and the thrown-exception case (wrapped by safeLocalImpact).

Made-with: Cursor
2026-04-21 11:44:16 +01:00
Copilotandmagyargergo f14068e09b fix(fts): Don't cache failed FTS index ensure; invalidate on pool teardown (#1006)
* Initial plan

* Don't cache failed FTS index ensure; invalidate on pool teardown

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/425d41bd-2cc1-49f6-8cc5-57368f0f238e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-04-21 08:50:33 +01:00
ivkond fb3bc7829e docs(group): add gRPC microservices group guide (#906) (#994)
* docs(group): add gRPC microservices group guide (#906)

Adds `docs/guides/microservices-grpc.md`, a walkthrough for using
GitNexus across multiple repositories whose services communicate over
gRPC. Covers the group mental model, per-repo `gitnexus analyze`, the
`group.yaml` schema, `group sync`, inspecting `contracts.json`,
running cross-repo `impact` with `@<group>` routing, the gRPC
extractor's provider/consumer signals per language, the
`config.links` manifest escape hatch, and a short troubleshooting
list. Wires the new page from the group-mode note in AGENTS.md.

Closes #906.

Made-with: Cursor

* docs(grpc-guide): drop hard line wraps, rely on editor soft wrap

Made-with: Cursor
2026-04-21 08:14:21 +01:00
dependabot[bot] 8fb386d45f chore(deps)(deps-dev): bump @types/node in /gitnexus (#1002) 2026-04-21 06:21:31 +01:00
dependabot[bot] cfeece95f5 chore(deps)(deps): bump graphology from 0.25.4 to 0.26.0 in /gitnexus (#1001) 2026-04-21 06:21:12 +01:00
dependabot[bot] 1e8bacf608 chore(deps)(deps): bump uuid from 13.0.0 to 14.0.0 in /gitnexus (#1000) 2026-04-21 06:20:48 +01:00
evolutionandwangjichao c0ebd160c2 fix(docker): copy gitnexus/package.json into web builder stage (#997)
vite.config.ts reads engines.node from ../gitnexus/package.json,
but Dockerfile.web only copied gitnexus-shared and gitnexus-web,
causing the build to fail with "Cannot find module" during
`npm run build --prefix gitnexus-web`.

Co-authored-by: wangjichao <wangjichao@inke.cn>
2026-04-20 19:13:50 +01:00
evolutionandwangjichao 2ac1baf450 fix(docker): use inputs.tag to detect workflow_call context (#996)
In a reusable workflow, github.event_name inherits the caller's
event (e.g. "push"), not "workflow_call". This caused the
type=raw tag to be disabled when docker.yml was called from
release-candidate.yml, producing no Docker tags at all and
failing the build.

Fix: check `inputs.tag != ''` instead, since inputs.tag is only
populated for workflow_call invocations.

Co-authored-by: wangjichao <wangjichao@inke.cn>
2026-04-20 18:20:57 +01:00
Jonas VanderhaegenandJonas Vanderhaegen 06967e2b66 feat(extractors): add PHP HTTP consumer detection (#993)
Extend the PHP tree-sitter plugin to emit consumer HttpDetections for
three common PHP HTTP call shapes, matching Node plugin parity:

  - Laravel HTTP client:  Http::get/post/put/delete/patch($url)
  - Guzzle / generic:     $client->get/post/...($url)
  - file_get_contents($url) when the URL is absolute http(s)://

String-literal URLs only. Paths built via binary concatenation
(`$base . '/path'`), sprintf, or config lookups are intentionally
deferred — they need constant-folding of the enclosing scope to be
useful and are tracked as follow-up work.

Refs #992

Co-authored-by: Jonas Vanderhaegen <jonasvanderh+claude.ai@gmail.com>
2026-04-20 17:35:31 +01:00
Sam Fakhreddine c24bcc3bf1 fix: expose detect-changes in direct CLI (#892)
Squashed commits:
- test: fix risk_level mock case and prettier formatting in tool-direct-cli.test
- test: add edge-case coverage for detectChangesCommand formatter
2026-04-20 17:12:25 +01:00
jisue0224andjisue0224 8f41a1ba17 fix(bm25): return FTS-matched symbols instead of arbitrary LIMIT 3 nodes (#806)
* fix(bm25): return FTS-matched symbols instead of arbitrary LIMIT 3 nodes

Previously, bm25Search fetched up to 3 arbitrary symbols from the matched
file using MATCH (n) WHERE n.filePath = $filePath LIMIT 3 (no ORDER BY).
This meant the specific function or class that actually scored highest in
the BM25 index could be completely absent from the results.

Fix: propagate nodeId from each FTS hit through searchFTSFromLbug, then
use those nodeIds in bm25Search to look up the exact matched nodes via
WHERE n.id IN $nodeIds. Falls back to the old filePath-based lookup when
nodeIds are unavailable.

Also switches the per-file score aggregation from naive sum-of-all to
sum-of-top-3, which prevents files with many mediocre matches (e.g. test
files) from outranking files with a single highly-relevant symbol.

* test(bm25): add unit tests for top-3 aggregation and nodeIds propagation

Covers the new logic paths added in the previous commit:
- top-3 score aggregation (file with 5+ matches → only top-3 contribute)
- nodeIds propagation through BM25SearchResult
- empty nodeId filtering
- cross-table merge for the same file
- result ranking by aggregated score

Also fixes in-place entries.sort() mutation (bm25-index.ts:125) to use
[...entries].sort() so the Map value is not silently modified.

* style: apply prettier formatting

* fix(test): use importOriginal to avoid missing export errors in vi.mock

* fix(bm25): align queryFTSViaExecutor nodeId extraction to match lbug-adapter

Use node.nodeId || node.id || '' in queryFTSViaExecutor to match the
fallback logic in lbug-adapter.ts:1040. Without this, the MCP pool path
could silently return empty nodeIds if LadybugDB surfaces the node id
under node.nodeId rather than node.id.

---------

Co-authored-by: jisue0224 <>
2026-04-20 17:06:58 +01:00
xiaohaoxingandClaude Sonnet 4.6 d858746476 fix(embeddings): replace recursive AST traversal with iterative DFS (#990)
findFunctionNode and findDeclarationNode had no depth limit, causing
stack overflow on deeply nested or auto-generated ASTs, especially
when --stack-size is not applied (e.g. heap already large enough to
skip ensureHeap re-exec).

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 12:07:19 +01:00
ivkond 00966630c4 feat: cross-repo impact analysis (#794) — @repo MCP routing + group resources (#984) 2026-04-20 11:55:07 +01:00
Andrew Barnes 5c3f56df7d docs: add --skip-git to CLI command list (#750) 2026-04-20 09:26:26 +01:00
Gergő Magyar f53e282026 Update Discord link in README.md 2026-04-20 08:59:50 +01:00
evolutionandwangjichao 2b7cff5fd2 feat(embeddings): structural chunking with data-driven CHUNKING_RULES dispatch (#987)
* feat(embeddings): structural chunking with data-driven CHUNKING_RULES dispatch

Replace hardcoded label comparisons with a CHUNKING_RULES lookup table
that drives chunking strategy and text generation. Key changes:

- Data-driven dispatch: CHUNKING_RULES table maps labels to chunking
  mode (ast-function / ast-declaration), prefix/suffix, field grouping,
  and structural text mode
- Struct support: add Struct to AST declaration chunking with field
  grouping (same as Class)
- Multi-chunk context: preceding chunk tail (prevTail) injected into
  embedding text for cross-chunk coherence
- Version-gated hashes: EMBEDDING_TEXT_VERSION prefix in content hashes
  invalidates stale vectors when text template changes
- Compact container context: first declaration line preserved in every
  structural chunk for identity

* fix(embeddings): address PR review findings for CHUNKING_RULES refactor

- Remove LABEL_ENUM from STRUCTURAL_LABELS to avoid wasted AST parses
- Add maintenance note about extractStructuralNames and EMBEDDING_TEXT_VERSION
- Clarify CHUNK_MODE_CHARACTER is a no-op in CHUNKING_RULES
- Strengthen EMBEDDING_TEXT_VERSION test assertion to exact value

---------

Co-authored-by: wangjichao <wangjichao@inke.cn>
2026-04-20 08:25:31 +01:00
Copilotandmagyargergo d976038dc8 fix: guard RC docker job against empty vtag and add early validation in docker.yml (#983)
* Initial plan

* fix: guard docker job and add tag validation in docker.yml

- Add `&& needs.publish.outputs.vtag != ''` to the `docker` job's
  `if:` in release-candidate.yml so it is skipped when publish
  produces no vtag, preventing an opaque buildx "tag is needed" error.
- Add an early "Validate tag input" step in docker.yml that fails fast
  with a clear ::error:: message when inputs.tag is empty, covering
  direct workflow_call invocations that bypass the release-candidate
  guard.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b9afe2df-85ea-4a87-bf30-77f0e945a64d

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: scope docker.yml tag validation to workflow_call only

Direct tag-push triggers (on: push, tags: v*) populate the tag from
GITHUB_REF and have inputs.tag empty, so the unconditional validation
step would fail every direct tag-push run. Restrict the new step to
workflow_call invocations, which is the only path where an empty tag
is actually a problem.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/4b7e3bfa-15c0-4186-affa-95cd71e50153

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-04-19 09:41:30 +01:00
Copilotandmagyargergo 9926804d75 feat(cli): infer registry name from git remote.origin.url (#981)
* Initial plan

* Plan: smarter index name inference via git remote URL

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/95064d2d-b1da-4c89-9069-5b3e9cc2636a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat(cli): infer registry name from git remote.origin.url (#979)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/95064d2d-b1da-4c89-9069-5b3e9cc2636a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor: skip git subprocess when --name was supplied (review feedback)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/95064d2d-b1da-4c89-9069-5b3e9cc2636a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* style: prettier --write on run-analyze.ts

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a4bf631d-ea6b-4d84-b426-29b1e5c3539f

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-04-19 09:14:14 +01:00
Copilot fa39a4b4a5 fix(docker): build and push Docker images for Release Candidates (#978) 2026-04-19 07:46:21 +01:00
azizur100389 dae7bd3b3f feat(cli): analyze --name <alias> + duplicate-name guard for the repo registry (#955) 2026-04-19 07:23:48 +01:00
Ryanba 363245eb63 fix: detect React component paths before lowercasing (#260) 2026-04-19 07:10:36 +01:00
Gergő Magyar 6222b5be9b feat(ingestion): emit-references drains ReferenceIndex to graph edges (#925, RFC #909 Ring 2 PKG) (#973) 2026-04-18 23:36:10 +01:00
Gergő Magyar e2ba4a04c9 feat(ingestion): shadow-mode parity harness + static dashboard (#923, RFC #909 Ring 2 PKG) (#972)
* feat(ingestion): shadow-mode parity harness + static dashboard (#923, RFC #909 Ring 2 PKG)

Side-car observability for the RFC #909 registry rollout. Callers that
dual-run legacy-DAG + `Registry.lookup` feed their result pairs into
the harness; the harness diffs each pair via shared `diffResolutions`
(#918), aggregates via `aggregateDiffs`, and persists a per-language
parity report that the static dashboard can render offline.

## Shipped

### `gitnexus/src/core/ingestion/shadow-harness.ts` (new)

```ts
createShadowHarness(): ShadowHarness
```

API:
  - `enabled` — `true` iff `GITNEXUS_SHADOW_MODE` is truthy at
    construction. Captured once; later env-var mutations don't flip it.
  - `record({ language, callsite, legacy, newResult, primary })` —
    accumulator. No-op when `enabled === false` (near-zero overhead).
  - `size()` — diagnostic counter.
  - `snapshot(now?)` — deterministic `ShadowParityReport` from the
    accumulated diffs.
  - `persist(outputDir, now?)` — writes BOTH a timestamped
    `<runId>.json` and a `latest.json` pointer. Creates outputDir if
    absent. Returns the per-run file path.
  - `clear()` — resets the accumulator; preserves `enabled`.

Activation: `GITNEXUS_SHADOW_MODE` accepts `'true'` / `'1'` / `'yes'`
(case-insensitive, trimmed); same truthy convention as
`REGISTRY_PRIMARY_<LANG>` from #924. Typos → disabled (fail-safe).

Persisted payload (`PersistedShadowReport`) is schema-versioned (`v1`):

```jsonc
{
  "schemaVersion": 1,
  "runId": "YYYYMMDD-HHMMSS-xxxxxxxx",
  "generatedAt": "ISO 8601",
  "primaryByLanguage": { "python": "legacy", ... },
  "report": { /* ShadowParityReport from #918 aggregateDiffs */ }
}
```

`runId` prefix is the timestamp so files sort chronologically; the
entropy suffix prevents collisions within a clock-second.

### `gitnexus/shadow-parity-dashboard/index.html` (new)

Minimal static dashboard — one HTML file, zero build step, zero runtime
deps. Fetches `./latest.json` and renders:

  - Overall summary cards (total calls, both agree, disagree, overall parity %)
  - Per-language table: language tag ("primary: legacy" / "primary:
    registry" pill) + total / agree / only-legacy / only-new / disagree
    / both-empty / parity%
  - Parity cells colored by threshold: ≥95% green, ≥80% amber, <80% red
  - Light / dark via `prefers-color-scheme`
  - Empty-state message when no records yet

File-serving is static: `cp .gitnexus/shadow-parity/latest.json
gitnexus/shadow-parity-dashboard/` + open in a browser.

## Tests (14, all passing)

  - **Flag detection** (5): default off · truthy variants case-insensitive ·
    falsy / typo → off · record() is no-op when disabled · env flip
    AFTER construction doesn't enable (constructed-once semantics)
  - **Record + snapshot** (4): multi-language accumulation ·
    per-language rows with correct outcomes · snapshot determinism ·
    `clear()` resets accumulator + `primaryByLanguage`
  - **Persistence** (5): mkdir-p on missing outputDir · per-run +
    latest.json match byte-for-byte · schema v1 payload shape ·
    runId timestamp prefix sorts chronologically · empty report
    persists gracefully

Tests use a per-test tmpdir (`fs.mkdtemp`), cleaned in `afterEach`,
so parallel vitest runs don't collide. `GITNEXUS_SHADOW_MODE` is
saved + restored per-test.

## What's deliberately NOT in this PR (call-out in harness docstring)

  - **Dual-run dispatch.** The harness is a side-car — it does NOT
    invoke either resolution path. Call-processor integration that
    actually runs both legacy + registry paths lands as a follow-up.
    Without that integration, `record()` is never called in production
    today. The harness is tested in isolation with synthetic inputs.
  - **CI artifact publishing.** Config work to upload
    `latest.json` + the dashboard HTML per CI run. Tracked separately;
    the harness + dashboard are ready when the CI job wires in.
  - **Fixture-level drill-down.** The issue mentions per-fixture AST
    snippet + evidence trace drill-down. MVP dashboard shows per-language
    rows only; drill-down extends the static JSON format + the dashboard
    JS in a focused follow-up.

## Verification

  - `tsc --noEmit` clean (both `gitnexus-shared` and `gitnexus`)
  - 14/14 new tests pass
  - Full scope-resolution / shadow / model / flag suite: **335/335 pass**

## Part of

  - Parent: #909
  - Depends on (code): #917 (registries), #918 (diff + aggregate)
  - Unblocks Ring 3 language flips: the parity dashboard becomes the
    checkpoint before flipping `REGISTRY_PRIMARY_<LANG>=true` for a
    language — once per-language parity stabilizes, the flip ships.

* chore: prettier format on shadow-parity-dashboard index.html
2026-04-18 21:30:43 +01:00
Gergő Magyar 0c37eda482 feat(ingestion): per-language resolveImportTarget adapter (#922, RFC #909 Ring 2 PKG) (#971)
Bridges the CLI's existing per-language `ImportResolverFn`s (16 languages
already implemented) to the shared `FinalizeHooks.resolveImportTarget`
contract consumed by `finalize()` (#915) and
`finalizeScopeModel` (#921).

No resolver logic is reimplemented — the adapter wraps
`provider.importResolver` from each `LanguageProvider` verbatim.

## Shipped

### `import-target-adapter.ts` (new)

```ts
buildImportTargetWorkspace(providers, resolveCtx): ImportTargetWorkspace
resolveImportTargetAcrossLanguages(targetRaw, fromFile, workspaceIndex): string | null
```

  - `ImportTargetWorkspace` is the opaque `workspaceIndex` shape the
    adapter recognizes: `{ perLanguage: Map<SupportedLanguages,
    { resolver, ctx }> }`. Callers build it once per ingestion run from
    the active language providers.
  - `resolveImportTargetAcrossLanguages` is the `FinalizeHook`
    implementation. It:
      1. Reads `getLanguageFromFilename(fromFile)`.
      2. Looks up the per-language entry.
      3. Calls the existing `ImportResolverFn` — same signature, same
         code path the legacy DAG uses today.
      4. Picks `result.files[0]` (covers both `'files'` and `'package'`
         result kinds; the legacy pipeline's richer multi-file + dirSuffix
         semantics stay accessible through `importResolver` directly).
      5. Returns `null` on any null result, empty files[], unknown
         extension, missing workspace, or resolver exception.
  - Exceptions from resolvers are swallowed — the finalize algorithm
    treats `null` as `linkStatus: 'unresolved'`, which is the right
    fallback for malformed inputs.

### What's deliberately NOT here

  - **Re-implementation of any per-language resolver.** Wraps the
    existing `importResolver` field on each provider.
  - **Dynamic-import handling.** The shared finalize algorithm short-
    circuits `ParsedImport { kind: 'dynamic-unresolved' }` before
    calling `resolveImportTarget`, so the adapter never sees them.
  - **`importPathPreprocessor`.** Preprocessing belongs inside the
    provider's `interpretImport` hook that produces
    `ParsedImport.targetRaw`; the adapter forwards that verbatim.

## Tests (12, all passing)

  - **`buildImportTargetWorkspace`** (3): registers providers with
    importResolver · skips providers without · threads shared ctx
    into every entry
  - **`resolveImportTargetAcrossLanguages`** (9): forwards targetRaw +
    fromFile · dispatches by extension · null resolver result →
    null · `package`-kind takes first file · empty files[] → null ·
    no registered resolver → null · unknown extension → null ·
    undefined/malformed workspace → null · resolver throw → null

Real per-language resolver correctness is covered by the existing
per-language resolver test suites — the adapter is the bridge layer.

## Verification

  - `tsc --noEmit` clean (both `gitnexus-shared` and `gitnexus`)
  - `gitnexus-shared` build clean
  - 12/12 new tests pass
  - Full scope-resolution / shadow / model / flag suite: **333/333 pass**

## Integration flow

```ts
const workspace = buildImportTargetWorkspace(providers, resolveCtx);
const indexes = finalizeScopeModel(parsedFiles, {
  hooks: { resolveImportTarget: resolveImportTargetAcrossLanguages },
  workspaceIndex: workspace,
});
model.attachScopeIndexes(indexes);
```

## Closes part of #909. Unblocks

  - Ring 3 language migrations (#926+): a language flipping to
    `REGISTRY_PRIMARY_<LANG>=true` now has correct import-target
    resolution out of the box via its existing `importResolver`.
  - #923 shadow harness — can run the dual-path comparison knowing
    both sides use the same per-language resolution semantics.
2026-04-18 21:11:24 +01:00
Copilotandmagyargergo 3adb97e993 feat(docker): ship signed UI + CLI/server images via docker-compose (#967)
* Initial plan

* docker: ship signed UI + CLI/server images via docker-compose

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/883bcee1-4a1d-4b3d-bbb9-accd8846da96

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docker: lock image version to npm package + harden cosign verify guidance

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6afd4fcd-5656-4e02-b796-a22b59000bde

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docker: add Sigstore ClusterImagePolicy + k8s admission docs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/aea3dd70-e2a9-443a-b578-cb3eca4093e1

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docker(k8s): collapse redundant image globs in ClusterImagePolicy

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/aea3dd70-e2a9-443a-b578-cb3eca4093e1

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docker(ci): drop deprecated COSIGN_EXPERIMENTAL, dead build-args, and loose verify regex in comment

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/bdf0d2cf-607c-4558-982a-be9b216b2d36

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docker(ci): use ${{ github.repository }} in verify-comment regex for fork portability

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/bdf0d2cf-607c-4558-982a-be9b216b2d36

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* style(deploy): prettier-format cluster-image-policy.yaml (single quotes)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b09a016e-56a1-4b73-bacc-e69084a48782

* ci(docker): drop workflow_dispatch, harden signing loop, fix verify-comment placeholder

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6755064c-7871-4b2e-9b46-b4779eb215ac

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-04-18 20:55:19 +01:00
Gergő Magyar 25520e90a5 feat(ingestion): finalize-orchestrator materializes ScopeResolutionIndexes (#921, RFC #909 Ring 2 PKG) (#970)
Ties the Ring 2 pipeline together. Takes the `ParsedFile[]` produced by
#920's parse-worker integration, feeds them to shared `finalize()`
(#915), and bundles every workspace-wide index for attachment onto
`MutableSemanticModel`. Thin integration glue per issue #884's boundary
— all algorithm lives in `gitnexus-shared`.

## Shipped

### `model/scope-resolution-indexes.ts` (new)

```ts
interface ScopeResolutionIndexes {
  readonly scopeTree: ScopeTree;
  readonly defs: DefIndex;
  readonly qualifiedNames: QualifiedNameIndex;
  readonly moduleScopes: ModuleScopeIndex;
  readonly methodDispatch: MethodDispatchIndex;
  readonly imports: ReadonlyMap<ScopeId, readonly ImportEdge[]>;
  readonly bindings: ReadonlyMap<ScopeId, ReadonlyMap<string, readonly BindingRef[]>>;
  readonly referenceSites: readonly ReferenceSite[];
  readonly sccs: readonly FinalizedScc[];
  readonly stats: FinalizeStats;
}
```

The bundle produced by the orchestrator, consumed by the resolution
phase. `ReferenceIndex` is deliberately NOT here — it's populated in
the next phase (#925).

### `model/semantic-model.ts` — extended

  - `SemanticModel.scopes?: ScopeResolutionIndexes` — undefined until
    attached; once attached, frozen.
  - `MutableSemanticModel.attachScopeIndexes(indexes)` — one-shot write.
    Throws on second call; `Object.freeze`s the bundle on write. `clear()`
    resets the slot back to `undefined` so re-ingestion can re-attach.

### `finalize-orchestrator.ts` (new)

```ts
finalizeScopeModel(parsedFiles, options?): ScopeResolutionIndexes
```

Orchestration steps:

  1. Map `ParsedFile[]` → `FinalizeInput` (`FinalizeFile` is a structural
     subset, so no shape-shifting).
  2. Call shared `finalize()` with provider hooks (defaults provided for
     the zero-provider case today).
  3. Build the four workspace indexes (`DefIndex`, `QualifiedNameIndex`,
     `ModuleScopeIndex`, `ScopeTree`) from per-file unions.
  4. Build an empty `MethodDispatchIndex` as a placeholder (owners=[],
     both callbacks return []). Real MRO wiring lands with the
     per-language adapters in #922.
  5. Bundle + return.

**Empty-input safety.** Zero parsedFiles → valid but empty bundle with
all zero-sized indexes and `stats.totalFiles === 0`. Downstream code
can consult `model.scopes` without branching on presence — only on
`stats`.

**Hook defaults** (`withDefaultHooks`) for missing provider hooks:

  - `resolveImportTarget: () => null` — every import goes `unresolved`
  - `expandsWildcardTo: () => []` — wildcards don't materialize
  - `mergeBindings: (a, b) => [...a, ...b]` — append without precedence

Providers override these in #922 (per-language import adapters).

## Tests (10, all passing)

  - **Empty input** (1): zero parsedFiles → valid empty bundle
  - **Single file** (2): all per-file indexes populated · referenceSites
    aggregated
  - **Cross-file imports** (3): resolveImportTarget threads through +
    links · default-null resolver → unresolved · stats reflect graph
  - **MutableSemanticModel integration** (4): undefined initially · attach
    once · Object.freeze applied · throws on re-attach · clear() resets

## Verification

  - `tsc --noEmit` clean in both packages
  - `gitnexus-shared` build clean
  - 10/10 new tests pass
  - Full scope-resolution / shadow / model / flag suite: **321/321 pass**

## What's deferred (not this PR, per RFC #909 scope)

  - **Per-language hook adapters** (#922): `resolveImportTarget` +
    `expandsWildcardTo` + `mergeBindings` wired per language.
  - **MethodDispatchIndex wiring via HeritageMap**: populate MRO + implements
    via the existing CLI-package HeritageMap strategies. Likely companion
    to #922 or a focused follow-up.
  - **Pipeline invocation**: actually calling `finalizeScopeModel` from
    the real ingestion pipeline. The orchestrator is callable today; the
    ingestion entry point wiring lands with the shadow harness (#923).
  - **`ReferenceIndex` population**: RFC §3.2 Phase 4 / #925.

## Closes part of #909. Unblocks
  - #923 shadow harness — now has a fully materialized `model.scopes` to
    query against the legacy DAG for parity measurement
  - #925 ReferenceIndex → LadybugDB emission — consumes `model.scopes`
  - Ring 3 language migrations (#926+) — a language flipping to
    `REGISTRY_PRIMARY_<LANG>=true` can now expect `model.scopes` to be
    populated when the pipeline wires the orchestrator in
2026-04-18 20:51:19 +01:00
Gergő Magyar 39b5d295c7 feat(ingestion): wire ScopeExtractor into parse-worker + processor (#920, RFC #909 Ring 2 PKG) (#969)
Plumbs the ScopeExtractor (#919) into the real parsing pipeline.
`ParsedFile` artifacts now flow from workers to the parsing-processor
without changing any legacy-DAG behavior.

## Shipped

### `gitnexus/src/core/ingestion/scope-extractor-bridge.ts` (new)

  - `extractParsedFile(provider, sourceText, filePath, onWarn?)`
  - Short-circuits (returns `undefined`) when the provider has not
    implemented `emitScopeCaptures`. True for every language today —
    this is the default no-op path.
  - Invokes the hook + `ScopeExtractor.extract`, returns a `ParsedFile`.
  - **Swallows exceptions on both sides.** Failures route through the
    optional `onWarn` callback (or `console.warn`) and return
    `undefined`. Scope-extraction errors NEVER break legacy parsing on
    the same file.
  - Standalone module (not nested in `parse-worker.ts`) so tests can
    import it directly without triggering the worker's top-level
    `parentPort!.on(...)`.

### `gitnexus/src/core/ingestion/workers/parse-worker.ts`

  - `ParseWorkerResult.parsedFiles: ParsedFile[]` added.
  - `processFileGroup` calls `extractParsedFile` AFTER tree parse,
    BEFORE legacy extraction. Worker provides an `onWarn` callback that
    routes bridge warnings through `parentPort.postMessage({ type:
    'warning', message })`.
  - `mergeResult` includes `parsedFiles` in the sub-batch merge.
  - Initial + reset accumulator templates include `parsedFiles: []`.

### `gitnexus/src/core/ingestion/parsing-processor.ts`

  - `WorkerExtractedData.parsedFiles: ParsedFile[]` added.
  - Empty-result branch and the across-chunk aggregation both include
    `parsedFiles`. Aggregation is tolerant of workers that don't emit
    the field (older builds / partial rollouts).

### Ring 1 tweak: `emitScopeCaptures` sync return

`readonly CaptureMatch[]` (was `Promise<readonly CaptureMatch[]>`).
Tree-sitter and COBOL's regex tagger are both synchronous; no
foreseeable need for async work inside this hook. Sync lets the
already-sync worker pipeline invoke it inline without cascading
`async` up through the batch driver + IPC handler.

## Tests (9 new; full suite 311/311)

`gitnexus/test/unit/scope-resolution/parse-worker-scope-integration.test.ts`:
  - Not-migrated (2): undefined-returning hook · never-invokes-extractor
  - Migrated (3): happy path · argument threading · honors
    `shouldCreateScope` override
  - Error resilience (4): hook throws · extractor throws (no Module) ·
    extractor throws (sibling overlap) · `onWarn` gets routed
    message with filePath + error body

## Verification

  - `tsc --noEmit` clean in both packages
  - `gitnexus-shared` build clean
  - 311/311 combined scope-resolution / shadow / model / flag suite
  - 9/9 new bridge tests

## What's NOT in this PR (still deferred to #921)

  - Actually using the `parsedFiles` — that's the finalize orchestrator.
  - `ModuleScopeIndex.byFilePath` materialization — belongs alongside
    the rest of the SemanticModel indexes in #921.

## Closes part of #909. Unblocks
  - #921 finalize-orchestrator — consumes `WorkerExtractedData.parsedFiles`
2026-04-18 20:27:56 +01:00
Gergő Magyar eece6344fc feat(ingestion): REGISTRY_PRIMARY_<LANG> per-language flag reader (#924, RFC #909 Ring 2 PKG) (#968)
Adds the per-language feature flag primitive that gates the Ring 3
registry-primary rollout. Single source of truth for whether a given
language uses `Registry.lookup` (new) or the legacy DAG (current).

## Shipped

### `gitnexus/src/core/ingestion/registry-primary-flag.ts`

  - `isRegistryPrimary(lang): boolean` — reads
    `REGISTRY_PRIMARY_<UPPER(enum-value)>` from `process.env`.
  - `envVarNameFor(lang): string` — exposed for CI tooling that
    cross-references flag flips (and for test assertions).
  - `primaryLanguages(): ReadonlySet<SupportedLanguages>` — all
    currently-on languages; useful for startup logging + the #923
    shadow dashboard which distinguishes "primary: legacy" vs
    "primary: registry" rows.

### Contract

  - Default: `false` for every language. A language must explicitly
    opt in by setting its env var.
  - Truthy: `'true'`, `'1'`, `'yes'` (case-insensitive, whitespace-
    trimmed). Anything else — typos, empty string, `'off'` — is
    `false`. Fail-safe posture: a misspelled flag doesn't accidentally
    flip a language.
  - No per-process caching. `process.env` is read per call; overhead
    is negligible (one lookup per file at resolution time), and
    test isolation is lexical (no cache-reset coordination).

### Env-var mapping

Uses the enum VALUE, not the TS key, for the env-var suffix:

  - `SupportedLanguages.Python`     → `REGISTRY_PRIMARY_PYTHON`
  - `SupportedLanguages.CPlusPlus`  → `REGISTRY_PRIMARY_CPP`   (value `'cpp'`)
  - `SupportedLanguages.CSharp`     → `REGISTRY_PRIMARY_CSHARP`

Users flip languages by their canonical name, not the TS symbol.

## Tests (16, all passing)

  - `envVarNameFor` (3): upper-casing · enum-VALUE-not-KEY mapping ·
    all-languages uniqueness smoke-test
  - `isRegistryPrimary` (9): default false · `'true'` / `'1'` / `'yes'`
    truthy · mixed-case + whitespace-padded · falsy-looking values ·
    unrecognized tokens (typo-safe) · per-language isolation · no
    stale cache on mid-process mutation · CPlusPlus mapping
  - `primaryLanguages` (3): empty · exact membership · Set instanceof

Tests scrub every `REGISTRY_PRIMARY_*` env var in `beforeEach` +
`afterEach` so parallel vitest runs on the same process don't bleed state.

## What's NOT in this PR (deferred by design)

The actual integration in `call-processor.ts` belongs in #921
(finalize-orchestrator). Reason: the "new path" requires a populated
`SemanticModel` to call `Registry.lookup` against, and the model
becomes accessible only after #921 orchestrates finalize. Wiring a
dead branch now would just get rewritten then.

This PR ships the flag primitive in isolation so #921 has a clean,
tested utility to consult — and so `#923` (shadow harness) has a
stable boolean to read for its "which row is primary?" rendering.

## Closes part of #909. Unblocks
  - #921 finalize-orchestrator — can now consult `isRegistryPrimary`
    at resolution time
  - #923 shadow harness — can distinguish primary-flipped rows
2026-04-18 19:54:54 +01:00
Gergő Magyar c6a291de67 feat(ingestion): ScopeExtractor driver — 5-pass CaptureMatch → ParsedFile (#919, RFC #909 Ring 2 PKG) (#965)
* feat(ingestion): ScopeExtractor driver — 5-pass CaptureMatch → ParsedFile (#919, RFC #909 Ring 2 PKG)

Kicks off Ring 2 PKG. Implements RFC §5.3 + §3.2 Phase 1: the central,
source-agnostic driver that turns a language provider's `CaptureMatch[]`
into a `ParsedFile` — the per-file artifact the finalize orchestrator
(#921) feeds into the shared `finalize()` algorithm (#915).

## Files

### New shared contracts
  - `gitnexus-shared/src/scope-resolution/parsed-file.ts`
    Per-file extraction artifact: scopes, parsedImports, localDefs,
    referenceSites. Structural superset of `FinalizeFile` so the
    finalize orchestrator threads `ParsedFile` through unchanged.
  - `gitnexus-shared/src/scope-resolution/reference-site.ts`
    Pre-resolution usage fact: name, atRange, inScope, kind, optional
    callForm/explicitReceiver/arity. Converted to `Reference` records
    by the resolution phase (populates `ReferenceIndex`).

### Ring 1 collateral tweak
  - `language-provider.ts: emitScopeCaptures` now returns
    `Promise<readonly CaptureMatch[]>` (was `readonly Capture[]`).
    Pre-grouping per tree-sitter match is the provider's job — the
    extractor expects coherent matches, not flat captures. No
    consumers yet (all languages still on legacy DAG), so no breakage.
    Docstring updated.

### New CLI module
  - `gitnexus/src/core/ingestion/scope-extractor.ts`
    Single entry point: `extract(matches, filePath, provider): ParsedFile`.
    Five-pass pipeline:

      Pass 1 — Build scope tree. `@scope.*` → `ScopeDraft[]` via
        range-containment parent derivation. Honors
        `provider.shouldCreateScope` (skip-but-reparent-children) and
        `provider.resolveScopeKind`. Throws `ScopeTreeInvariantError`
        via `buildScopeTree` on malformed input.

      Pass 2 — Attach declarations + local bindings. `@declaration.*`
        → `SymbolDefinition` + `BindingRef { origin: 'local' }`.
        Default attachment: innermost containing scope. Hoisting via
        `provider.bindingScopeFor`.

      Pass 3 — Collect raw imports. `@import.*` → `ParsedImport` via
        `provider.interpretImport`. Attached to ParsedFile
        (finalize resolves owning scope in Phase 2).

      Pass 4 — Collect type bindings. `@type-binding.*` →
        `TypeRef` via `provider.interpretTypeBinding` →
        `scope.typeBindings`. Hoistable via `bindingScopeFor`.

      Pass 5 — Collect reference sites. `@reference.*` →
        `ReferenceSite[]`. Call form from declarative sub-tag
        (`@reference.call.member`) or `provider.classifyCallForm`.

### Tests
  - `gitnexus/test/unit/scope-resolution/scope-extractor.test.ts`
    23 tests organized by pass + one end-to-end fixture exercising
    all 5 passes together. MockProvider emits synthetic
    `CaptureMatch[]` with no AST — extractor is pure given those.

## Design notes

- **Source-agnostic.** No `Tree` / `SyntaxNode` types leak into the
  driver. Works for tree-sitter providers and COBOL's regex tagger.
- **One AST walk per language.** Providers do the walk inside
  `emitScopeCaptures`; this driver does zero traversal.
- **Invariants delegated.** `ScopeTree.buildScopeTree` enforces
  structural rules (non-Module has parent, parent contains child,
  siblings don't overlap). The extractor doesn't try to repair
  malformed captures.
- **Sub-tag whitelist.** `@reference.receiver`, `@declaration.name`,
  `@import.source`, etc. are known sub-tags — excluded from anchor
  selection so the broadest-range heuristic doesn't mis-identify them
  as anchors for their topic. Bug surfaced in the end-to-end fixture
  test (member call with a large-range receiver) and was fixed before
  commit.

## Verification

  - `tsc --noEmit` clean (both `gitnexus-shared` and `gitnexus`)
  - `gitnexus-shared` build clean
  - 23/23 new tests pass
  - Full scope-resolution / model / shadow suite: **285/285 pass**

## Closes part of #909. Unblocks
  - #920 parse-worker integration (emit ParsedFile from the worker)
  - #921 finalize orchestrator (consume ParsedFile[] workspace-wide)
  - #922 per-language import adapters

* chore(ingestion): address #919 review findings on the extractor

Addresses all 5 items from the PR #965 review in-PR.

## Structural changes

- **Extract `ScopeExtractorHooks` as the narrow dependency surface.**
  The extractor now declares its dependency on a `Pick`-narrowed subset
  of `LanguageProvider` (just the 6 scope-resolution hooks it actually
  reads). Test mocks implement exactly that interface — no more
  `as unknown as LanguageProvider` cast hiding missing-field bugs.
  Adding a new hook read becomes a compile error, not a silent test
  pass. (Finding 3.2)

- **Remove dead `ownerDefIdFor` stub + `isOwnerKind` helper.** The
  function always returned `undefined` with `void innermost; void
  drafts;` suppressors — an incomplete-implementation signal. The code
  path was also misleading: creating a clone of the def with
  `ownerId: undefined` is structurally identical to keeping the
  original. Pass 2 now keeps the def as-is. Contract is documented in
  a code comment: providers that need `ownerId` set it from their
  declaration hook; `finalize` (via #914 `MethodDispatchIndex`) fills
  in method/field `ownerId` in a post-extraction pass that has full
  def visibility. (Finding 2.1)

- **Standardize `filePath` threading across passes 4 and 5.** Pass 4
  was reading `drafts[0]!.filePath`; pass 5 was reading
  `anyFilePathFromScopeTree(scopeTree)`. Both equivalent but
  inconsistent. Both now take `filePath` as a parameter from the
  top-level `extract()` call. The `anyFilePathFromScopeTree` helper is
  removed. (Finding 2.2)

## Documentation

- **Snapshot-semantics comment on `scopeTree` + `positionIndex`.** The
  hooks called during Passes 2-5 receive a `scopeTree` built BEFORE any
  bindings/ownedDefs/typeBindings were written. Hooks MUST NOT rely on
  `scope.bindings` etc. being populated — they're for parent/range/kind
  queries only. Added a doc block at the `scopeTree`/`positionIndex`
  construction site so future Ring 3 implementers don't write a
  `classifyCallForm` that reads bindings. (Finding 2.3)

## Tests

- **Regression for the anchor-vs-receiver bug** (Finding 3.1): a
  member-call match where `@reference.receiver` spans columns 0-10
  (wider) and the call name spans 11-15 (narrower). Without the
  `KNOWN_SUB_TAGS` exclusion, the broadest-range heuristic would have
  picked the receiver; the test pins that the call name is the one
  that ends up in `referenceSites[0].name`.

- **Mock provider now types exactly `ScopeExtractorHooks`**, no more
  double-cast. Any future hook added to `extract()` that isn't in
  `ScopeExtractorHooks` is a compile error.

## Verification

- `tsc --noEmit` clean in both `gitnexus-shared` and `gitnexus`
- `gitnexus-shared` build clean
- 24/24 scope-extractor tests pass (+1 regression)
- Full scope-resolution / model / shadow suite: **286/286 pass**
2026-04-18 19:28:51 +01:00
Gergő Magyar e944f90879 chore(shared): apply Ring 2 SHARED review follow-ups in one diff (#964)
* chore(shared): apply Ring 2 SHARED review follow-ups in one diff

Aggregates all actionable follow-ups from the 9 Ring 2 SHARED PRs
(#949–#963) before proceeding to Ring 2 PKG. No behavior changes;
docstring edits, test refinements, and one structural cleanup.

## #913 (DefIndex / ModuleScopeIndex / QualifiedNameIndex)
  - Rename `freezeIndex` → `wrapIndex` across all three index builders.
    The old name implied `Object.freeze` on the wrapper, which we never
    applied; `wrapIndex` more accurately describes the lightweight
    readonly-interface wrap. Safety surface (frozen bucket arrays,
    frozen miss-empty array, readonly Maps) is unchanged.
  - Document in `buildModuleScopeIndex` JSDoc that callers must
    pre-normalize `filePath` keys (no path-separator canonicalization
    happens here). Prevents silent cross-platform misses.
  - Add an explicit hit-path freeze assertion in
    `qualified-name-index.test.ts` (the existing test covered only the
    miss-path `EMPTY` array).

## #914 (MethodDispatchIndex)
  - Differentiate the C3 and BFS test cases: both tests now use
    distinct MRO orderings so they prove the materializer stores
    whatever order the `computeMro` callback produces (not that C3 and
    BFS yield identical output).
  - Add `implementsOfCalls` counter in the first-write-wins test, and
    document the call-count contract in `MethodDispatchInput.implementsOf`
    JSDoc: `implementsOf` fires **per occurrence** in `input.owners`
    (not per unique owner); `computeMro` fires at most once per unique
    owner. Callers with expensive `implementsOf` implementations should
    pre-dedupe `owners`.

## #916 (resolveTypeRef)
  - Document the deliberate exclusion of `'Type'` from `TYPE_KINDS`
    (verified no extractor in `gitnexus/src/core/ingestion/` emits
    `type: 'Type'` for annotation-relevant symbols).
  - Rename the namespace-origin test from `'resolves ...'` to
    `'returns null for a namespace-origin binding whose def is not a
    type kind'`, matching the failure-case intent.

## #918 (shadow diff + aggregate)
  - Remove the partial re-export `export type { ShadowAgreement, ShadowDiff };`
    from `aggregate.ts` — it omitted `ShadowCallsite` and diverged
    from the top-level barrel. Consumers import all three from the
    `gitnexus-shared` entry point.
  - Fix the invalid `'wildcard'` evidence kind in `diff.test.ts` fixture
    (that kind is not a valid `ResolutionEvidence.kind`). Replaced with
    `'global-name'`, a real kind the test treats identically.

## #912 (ScopeTree / PositionIndex / makeScopeId)
  - Document the touching-boundary semantics on `PositionIndex.atPosition`:
    when siblings share a boundary point, the right (later-start) sibling
    wins per the existing innermost-wins sort contract.
  - Resolve the layer-inversion flagged by review: move `ScopeLookup`
    from `resolve-type-ref.ts` to `types.ts` (its natural home in the
    data-model layer). `scope-tree.ts` now imports `ScopeLookup` from
    `types.js` directly; the old re-export from `resolve-type-ref.ts`
    is removed per repo convention (`feedback_no_reexport`). Barrel
    export moved alongside.

## #917 (ClassRegistry / MethodRegistry / FieldRegistry)
  - Replace the dangling "try a name-match among class-like defs"
    comment in `lookupReceiverType` with explicit prose that callers
    must pre-resolve via `resolveTypeRef` if they want richer semantics.
    No behavior change — the function already returned `undefined` on
    ambiguous/missing qnames.
  - Fix `tieBreakKey.origin` default for pure Step-2 candidates.
    Type-binding-only hits no longer falsely inherit `'local'` from
    `ensureCandidate`'s neutral default; they now demote to `'import'`
    on their first type-binding hit, and only a later Step-1 lexical
    hit can upgrade them back to `'local'`. Keeps the Appendix B
    cascade faithful to the true origin.
  - Document `'global-name'` in `evidence.ts`: currently reserved for
    Ring 3's byName global index; `lookupCore` never emits it today.
    The weight stays live so `composeEvidence` remains exhaustive over
    the origin union.
  - Rename the mislabeled Step-7 test from `'confidence DESC is the
    primary key'` (which actually tested hard-shadow baseline) to
    `'inner scope shadows outer, yielding single result'`, and add a
    separate test that actually exercises multi-candidate confidence
    ordering (local vs wildcard at the same scope).

## Verification
  - `tsc --noEmit` clean (both `gitnexus-shared` and `gitnexus`)
  - `gitnexus-shared` build clean
  - Combined scope-resolution / model / shadow suite: **260/260 pass**
    (+1 from the new multi-candidate ordering test in #917)

## Not addressed (non-actionable)
  - #949 CI "failure with zero failing tests": pre-existing Swift Node 22
    grammar flake unrelated to #910 scope.
  - #950: the two non-blocking findings were already addressed in
    follow-up commit `cbac32ba` (ParsedImport discriminated union +
    `ScopeId | null` on the two hooks).
  - #915: the five in-scope findings were already addressed in
    follow-up commit `54515a7e` (dead code, unused params, multi-hop
    docs, cap-hit test, stats granularity).
  - #915 LanguageProvider.resolveImportTarget signature divergence +
    `findDefById` O(F×D) perf: tracked separately as follow-up issues
    for the Ring 3 migration window.

* chore(shared): address ce:review findings on the follow-up diff

ce:review (interactive) on PR #964 surfaced two P2s and several P3s. This
commit applies all `safe_auto` fixes + both manual tests in-line so the
PR ships with a cleaner review trail.

## P2 fixes

- **Complete `freezeIndex` → `wrapIndex` rename.** The prior commit renamed
  3 of 5 sibling index files; `method-dispatch-index.ts` and
  `position-index.ts` still carried the old name. Now all 5 helpers use
  the consistent `wrapIndex` naming.
  (maintainability + project-standards reviewers both flagged this.)

- **Add regression tests for the `recordTypeBindingHit` origin demotion.**
  The prior commit introduced the `tieBreakKey.origin = 'import'`
  demotion for Step-2-only candidates without a direct test. Added:
    - `registries.test.ts`: two Step-2-only siblings under the same
      interface, asserting deterministic DefId.localeCompare tie-break
      AND the stronger invariant that composeEvidence never emits a
      where-found signal for Step-2-only candidates (no `signals.origin`).
    - `position-index.test.ts`: touching-boundary test proving the
      right-sibling-wins rule documented in the new JSDoc.
  (testing + kieran-typescript + api-contract reviewers all flagged these gaps.)

## P3 fixes

- Fix wrong comment in `recordTypeBindingHit` that claimed Step 1 could
  later upgrade a demoted origin. Step 1 runs BEFORE Step 2 — the actual
  upgrade path is Step 3 (`seedFromOwnerScopedContributor`). Comment now
  describes execution order correctly.

- Fix inaccurate "re-exported there" comment in `index.ts`. `types.ts`
  *defines* ScopeLookup natively; it's not a re-export. Phrasing now
  says "defined in types.ts and exported from the type-export block
  above — not from this module."

- Update stale `scope-tree.ts` file-header prose that still referenced
  `ScopeLookup` as living in #916/resolve-type-ref.ts. Now points to
  `./types.js` with a cross-ref to both #916 and #917 consumers.

- Expand `atPosition` touching-boundary JSDoc to name the mechanism
  (backward scan through start-sorted array) so readers can trace the
  binary-search code to the claim.

- Add breadcrumb to `aggregate.ts` module header pointing future readers
  to `./diff.ts` / the top-level barrel for `ShadowAgreement`,
  `ShadowCallsite`, and `ShadowDiff`.

- Remove unnecessary non-null assertion in `recordTypeBindingHit`. Local
  `const existingMroDepth = ...` lets TS narrow to `number` in the
  else-branch, eliminating the `!` without behavior change.

## Verification

- `tsc --noEmit` clean (both `gitnexus-shared` and `gitnexus`)
- `gitnexus-shared` build clean
- Combined scope-resolution / model / shadow suite: **262/262 pass** (+2
  from the new origin-demotion + touching-boundary regression tests)
2026-04-18 18:40:29 +01:00
Gergő Magyar 1bf9fb4ef1 feat(shared): ClassRegistry / MethodRegistry / FieldRegistry + 7-step lookup (#917, RFC #909 Ring 2 SHARED) (#963)
Capstone of Ring 2 SHARED. Implements RFC §4 — the shared, scope-aware
resolution surface the rest of the semantic model feeds into.

## Modules (`gitnexus-shared/src/scope-resolution/registries/`)

  - `context.ts`         — `RegistryContext` bundling ScopeTree / DefIndex
                           / QualifiedNameIndex / ModuleScopeIndex /
                           MethodDispatchIndex + provider hooks.
                           Narrows Ring 1's opaque `RegistryContributor`
                           to concrete `OwnerScopedContributor`.
  - `tie-breaks.ts`      — `compareByConfidenceWithTiebreaks`, the RFC
                           Appendix B cascade: confidence DESC → scope
                           depth ASC → MRO depth ASC → ORIGIN_PRIORITY
                           ASC → DefId.localeCompare.
  - `evidence.ts`        — `composeEvidence(signals)` / `confidenceFromEvidence`.
                           Translates raw walk signals into the typed
                           `ResolutionEvidence[]` using authoritative
                           `EvidenceWeights`. No magic numbers.
  - `lookup-qualified.ts`— RFC §4.5. Qualified-name fast path consumed
                           by `resolveTypeRef` dotted fallback and by
                           Step 6 of lookup-core.
  - `lookup-core.ts`     — The 7-step canonical algorithm. Pure. Param-
                           eterized by `CoreLookupParams`.
  - `{class,method,field}-registry.ts`
                         — Thin wrappers over `lookupCore` that fix
                           `acceptedKinds` + `useReceiverTypeBinding` per
                           kind. `buildClassRegistry` / `buildMethodRegistry`
                           / `buildFieldRegistry` factory functions.

## RFC §4.2 algorithm contract (honored verbatim)

  1. Lexical scope-chain walk. Hard shadow on any `scope.bindings.has(name)`
     regardless of kind survivorship.
  2. Type-binding resolution (methods/fields only, opt-in via
     `useReceiverTypeBinding`). MRO walk via `MethodDispatchIndex.mroFor`.
     MRO-depth-decayed weight via `typeBindingWeightAtDepth`.
  3. Owner-scoped contributor — when the caller knows the receiver owner,
     its direct members merge in as `origin: 'local'`.
  4. Kind filter — `acceptedKinds` per registry; `kind-match` evidence
     at weight 0 is always emitted for debuggability.
  5. Arity filter — `provider.arityCompatibility` per candidate. When at
     least one compatible candidate exists, incompatibles are dropped;
     otherwise the −0.15 penalty alone disambiguates (they stay in the
     result, just ranked lower).
  6. Global fallback — fires only when Steps 1-3 produced NO candidates
     AND the name is dotted. Delegates to `lookupQualified`.
  7. Rank + tie-break — evidence list sorted by the Appendix B cascade.

## §4.7 invariants asserted in tests

  - No tier vocabulary in the return type (`Resolution`, not `TierXResult`).
  - Confidence is per-candidate (not per-tier).
  - Shadowing is a HARD filter; globals are consulted ONLY when lexically
    empty.
  - Caller can read `[0]` for one-shot answers.
  - `Resolution.confidence` is capped at 1.0.
  - `kind-match` is always emitted (weight 0).

## Unresolved-import + dynamic-unresolved evidence shape

  - `BindingRef.via.linkStatus === 'unresolved'` applies the
    `unlinkedImportMultiplier` (0.5×) to the where-found signal only.
    Corroborators (`arity-match`, `owner-match`, `type-binding`) remain
    unaffected — the RFC §4v2 capped-signal rule applies per-signal, not
    per-candidate.
  - `BindingRef.via.kind === 'dynamic-unresolved'` adds a degraded
    `dynamic-import-unresolved` evidence signal at weight 0.02.

## Tests (28 in registries.test.ts, 259/259 combined)

Organized per RFC §4.2 step so a regression localizes to the step it broke:

  - Step 1: local + walk-to-parent + hard-shadow + origin=import
  - Step 2: explicit receiver type-binding + MRO depth decay on ancestor
  - Step 3: owner-scoped contributor + owner-match
  - Step 5: drop-incompatible-when-compatible-exists + soft-penalty-when-all-
            incompatible + unknown-when-no-provider
  - Step 6: global-qualified fires only when lexically empty + never for
            non-dotted names + not consulted when lexical hit exists
  - Step 7: tie-break cascade (inner shadows outer; defId.localeCompare
            final)
  - Corroborators: unresolved-import 0.5× cap per-signal + dynamic-
            unresolved 0.02 degraded signal
  - §4.5: lookupQualified kind filter + empty on miss + deterministic defId
          order for partial classes
  - §4.7: invariants — confidence per-candidate, capped at 1.0, kind-match
          always present, [0]-for-one-shot

## Known follow-up optimizations

`collectOwnedMembers` in `lookup-core.ts` iterates `defs.byId.values()`
for each MRO hop — O(D) per call. Acceptable for Ring 2 fixtures; a
by-owner index should land before Ring 3 migrates large-workspace
languages. Tracked alongside the existing `findDefById` follow-up from
#915 review.

## Module placement

All under `gitnexus-shared/src/scope-resolution/registries/` — consistent
with the Ring 2 SHARED folder layout (#912/#913/#914/#915/#916/#918).
Slight deviation from the issue's `gitnexus-shared/src/registries/`
suggestion for consistency with siblings.

## Part of

- Parent: #909
- Depends on (code): #910, #911, #912, #913, #914, #915, #916, #918.
- Closes the Ring 2 SHARED delivery band. Unblocks Ring 2 PKG (#919–#925
  bridges to the gitnexus/ CLI package) and Ring 3 language migrations.
2026-04-18 17:58:26 +01:00
Gergő Magyar a9a5e1c388 feat(shared): SCC-aware finalize algorithm with bounded fixpoint (#915, RFC #909 Ring 2 SHARED) (#962)
* feat(shared): SCC-aware finalize algorithm with bounded fixpoint (#915, RFC #909 Ring 2 SHARED)

Implements RFC §3.2 Phase 2 as pure logic in `gitnexus-shared`. Takes
per-file parse output and returns linked `ImportEdge[]` + materialized
module-scope bindings, fully language-agnostic (target resolution,
wildcard expansion, and binding precedence all go through caller hooks).

Three-phase algorithm:

  1. Tarjan SCC over the file-level import graph (iterative, deterministic
     node order, O(V+E)). Returns SCCs in reverse-topological order so
     leaves finalize before dependents — and so disjoint SCCs are
     explicitly surfaced for parallel-processing callers.

  2. Per-SCC bounded fixpoint. For each SCC in topo order, iterate up to
     `N = |intra-SCC edges|`; each pass tries to resolve every still-
     unlinked edge by looking up the imported name in the target file's
     local defs. Stops early when no progress. Edges still unlinked after
     the cap get `linkStatus: 'unresolved'` — keeps malformed inputs
     bounded and preserves the RFC §4v2 capped-signal contract for
     unresolved markers.

  3. Wildcard expansion + module-scope binding materialization. For each
     `wildcard` ParsedImport that linked to a module, expand via
     `expandsWildcardTo` into one `wildcard-expanded` ImportEdge per
     exported name. Bindings per module scope are the merge of local defs
     (`origin: 'local'`), named / alias / reexport imports
     (`origin: 'import' | 'reexport'`), namespace imports (`origin:
     'namespace'`), and wildcard expansions (`origin: 'wildcard'`), with
     precedence delegated to `provider.mergeBindings`.

Dynamic imports rule: `kind: 'dynamic-unresolved'` passes through as an
ImportEdge with `targetFile: null` and no BindingRef.

Re-export flattening: reexport edges land with `transitiveVia: [targetFile]`.
Multi-hop chains settle iteratively across the fixpoint.

Types:
  - Adds `'wildcard'` variant to ParsedImport (parse-time signal for
    `import * from M`). The finalize-only `'wildcard-expanded'` ImportEdge
    kind is unchanged and remains finalize output only, as documented.
  - Exports `finalize` + `FinalizeFile` / `FinalizeInput` / `FinalizeHooks`
    / `FinalizeOutput` / `FinalizedScc` / `FinalizeStats`.

Simple-name derivation: `deriveSimpleName` uses `def.qualifiedName` as the
authoritative source (tail after the last `.`). Defs without a
qualifiedName are not name-resolvable by this algorithm — an explicit
design choice that trades strictness for predictability (no heuristic
nodeId parsing).

Tests (20, all passing):
  - Trivial: empty workspace · acyclic resolution · unresolvable target
    (file + name) · dynamic-unresolved passthrough.
  - Cycles: A↔B two-file cycle linked · cycles packed into SCC with
    isCycle=true · disjoint cycles produce disjoint SCCs · mixed
    linked/unresolved edges reported correctly in stats.
  - Wildcards: one ImportEdge per exported name · unresolved wildcards
    survive as single edges · expanded bindings carry origin='wildcard'.
  - Reexports: transitiveVia carries the intermediate file path.
  - Aliased + namespace: alias preserves targetExportedName under its
    local name · namespace links to module scope even without a module-def.
  - Bindings: locals land as origin='local' · imports layer on via
    mergeBindings · mergeBindings can drop existing (last-write-wins
    precedence honored).
  - SCC-DAG: reverse-topological ordering verified (leaf first).

Combined scope-resolution / model / shadow suite: 229/229 pass.
`tsc --noEmit` clean in both `gitnexus-shared` and `gitnexus`.

Closes part of #909. Unblocks #917 (Registry.lookup's import-chain fast
path consumes finalized ImportEdges); unblocks Ring 3 language migrations
(per-language providers supply FinalizeHooks implementations).

* chore(shared): address #915 review findings — dead code, docs, tests

Review thread on PR #962.

Code changes:
  - Remove dead `resolvedTargets` map + `keyFor` + `ParsedImportKey`
    type alias. The map was populated but never read; originally intended
    to cache / dedup resolutions for later phases but that path was never
    wired (finding 1.1).
  - Drop unused params (`_edgeIndex`, `_hooks`, `_workspace`) from
    `tryFinalize`. No planned fixpoint-state consultation; no reason to
    keep them reserved (finding 2.1).

Documentation:
  - `FinalizeFile.localDefs` now documents the multi-hop re-export
    contract explicitly: `finalize` looks names up in the target's
    static `localDefs`; if B only re-exports from C and doesn't surface
    the name in its own localDefs, A's import of that name from B will
    hit the cap and be marked unresolved. Parsers that want multi-hop
    chains to settle end-to-end must include re-exported names in the
    intermediate file's localDefs (finding 1.2).
  - `FinalizeStats` now documents its counting granularity: all edge
    counters are per-`ParsedImport`, not per-materialized-`ImportEdge`.
    A wildcard expanding to N exports counts as one linked edge;
    dynamic-unresolved pass-throughs count as linked. The bindings map
    is the authoritative "has a BindingRef" source (finding 3.2).

Tests (2 added, 22 total in finalize-algorithm.test.ts, 231/231 combined):
  - Explicit cap-hit → `linkStatus: 'unresolved'` assertion for a cycle
    where the name-level lookup never succeeds (distinct from
    `targetFile: null`; cap exhaustion path) (finding 3.1).
  - Multi-hop re-export contract test: demonstrates both variants —
    intermediate B WITHOUT X in localDefs → unresolved; B WITH X in
    localDefs → resolved to the original source DefId (finding 1.2).

Not addressed (filed as follow-up issues):
  - LanguageProvider.resolveImportTarget vs FinalizeHooks signature
    divergence (finding 1.3) — pre-Ring-3 concern.
  - findDefById O(F×D) scan in Phase 5 (finding 4.1) — acceptable for
    Ring 2; optimize before large-workspace Ring 3 migrations.
2026-04-18 17:26:07 +01:00
Gergő Magyar 8cf9ae0e0d feat(shared): ScopeTree + PositionIndex + makeScopeId (#912, RFC #909 Ring 2 SHARED) (#961)
Implements the scope-tree spine and position-indexed lookup as pure logic
in `gitnexus-shared`. Generalizes the `enclosingFunctions` pattern from
closed PR #902 to arbitrary `ScopeKind`s.

Three modules under `gitnexus-shared/src/scope-resolution/`:

1. `scope-id.ts` — `makeScopeId({filePath, range, kind})` builds the
   canonical RFC §2.2 shape
     `scope:{filePath}#{startLine}:{startCol}-{endLine}:{endCol}:{kind}`
   and interns the result through a process-local pool so repeated calls
   with structurally identical inputs return the same string reference.
   `clearScopeIdInternPool()` exported for test isolation.

2. `scope-tree.ts` — `buildScopeTree(scopes)` validates invariants and
   returns an immutable `ScopeTree`:
     - `getScope(id)` / `getParent(id)` / `getChildren(id)` / `getAncestors(id)`
     - Implements the `ScopeLookup` contract from #916, so `resolveTypeRef`
       can consume a `ScopeTree` directly (test included).
   Invariants enforced (throw `ScopeTreeInvariantError` on violation):
     - Non-Module scopes must have a parent.
     - Parent must exist in the supplied set.
     - Parent range STRICTLY contains child range (equal ranges rejected).
     - Sibling ranges under the same parent do not overlap. Ranges that
       merely touch at the boundary (`a.end == b.start`) are accepted.
     - Parent and child live in the same filePath.
     - Duplicate scope ids are rejected.

3. `position-index.ts` — `buildPositionIndex(scopes)` produces a
   `PositionIndex` with `atPosition(filePath, line, col)`. Per-file sorted
   array; binary-search the upper bound of `start ≤ query`, scan backward
   through the prefix, return the first containing hit.
   Complexity: `O(log N_file + D)` typical (D = lexical depth ≤ ~10);
   degrades to `O(N_file)` only under pathological inputs (many scopes
   starting at the same position). "Innermost wins" falls out of the sort
   + backward-scan contract because `ScopeTree`'s invariants guarantee
   that scopes containing a point form an ancestor chain.

Types:
  - `ScopeTree` now exported from `scope-tree.ts`. The Ring 1 opaque
    placeholder in `types.ts` has been removed; LanguageProvider hooks
    that previously took `ScopeTree = unknown` now receive the concrete
    interface (CLI `tsc --noEmit` passes — no existing callers rely on
    the opaque shape).

Tests (39, all passing):
  - scope-id: canonical shape · all six ScopeKinds encoded · identity
    equality (same inputs → same reference) · distinguished by
    filePath / range / kind · purity under repeated calls · intern-pool
    clear preserves canonical shape.
  - scope-tree: empty tree · single module · nested Module→Class→Function
    · multiple siblings input-order preserved · ScopeLookup integration
    with resolveTypeRef · frozen children and ancestor arrays · all six
    invariant violations (non-Module orphan, parent-not-found, parent
    doesn't contain, parent == child, siblings overlap, cross-file parent,
    duplicate id) · boundary-touching siblings accepted.
  - position-index: empty · unindexed filePath · before/after-file
    queries · start/end inclusivity · innermost-wins for nested / co-
    starting / co-ending / same-line scopes · sibling dispatch · multi-
    file isolation · size · id-dedup.

Combined scope-resolution / model / shadow suite: 190/190 pass.
`tsc --noEmit` clean in both `gitnexus-shared` and `gitnexus`.

Closes part of #909. Unblocks #917 (`Registry.lookup` needs the scope
spine); makes `ScopeLookup` in #916 concrete without API churn.
2026-04-18 16:41:38 +01:00
azizur100389 ac148612ab feat(search): per-phase timing instrumentation for the query pipeline (#953)
* feat(search): per-phase timing instrumentation for the query pipeline

The eval harness already measures search-pipeline latency per phase,
but the *product* query() tool has no timing visibility. That leaves
production latency opaque:

 - Is BM25 the tail, or vector search?
 - How much Promise.all overlap do concurrent searches actually save?
 - Does symbol_lookup dominate when per-symbol Cypher round-trips pile up?

None of this is answerable from the outside, which blocks the
latency-quality Pareto work tracked in #546 / #553.

Changes:

* New PhaseTimer class at src/core/search/phase-timer.ts.
  Supports three APIs:
    - start(phase) / stop() for sequential phases (per issue spec)
    - mark(phase, durationMs) for pre-measured durations
    - time(phase, promise) to wrap a promise inside Promise.all

  The issue's original spec was sequential-only, which doesn't work
  for BM25 + vector inside Promise.all — the second start() would
  auto-stop the first and only one phase would get timed. The mark()
  and time() variants resolve that without changing the sequential
  API for the other phases.

* local-backend.ts query() instrumented across seven phase markers:
    bm25, vector   (concurrent via timer.time inside Promise.all)
    merge          (RRF reciprocal-rank-fusion)
    symbol_lookup  (per-symbol process + cohesion + content Cypher)
    ranking        (in-memory priority sort)
    formatting     (response object construction + dedup)
    wall           (end-to-end; separate mark so callers can compare
                   sum(phases) vs wall and see Promise.all savings)

* logQueryTiming() helper next to logQueryError(), same console-based
  pattern (repo has no structured logger). Emits
    GitNexus [query:timing] query="..." totalMs=N phases={...}
  to stdout — greppable prefix, JSON-parseable payload, no new deps.

* timing: Record<string, number> added as a top-level field on the
  query() response. Strict superset of the previous shape — existing
  tests only assert field presence, so no regression. Other MCP tools
  use the same top-level-metadata convention (status, row_count,
  warning) rather than a nested _meta wrapper.

Tests:

 - 6 new unit tests for PhaseTimer covering start/stop, implicit
   stop-on-start, additive mark(), Promise.all-safe time(),
   negative/NaN rejection, and totalMs auto-stop.
 - 3 new assertions on the existing query integration test verifying
   timing.wall is a non-negative number and at least one of
   bm25/vector fired.

Verification:
  npx vitest run test/unit/phase-timer.test.ts       -> 6 pass
  npx vitest run test/unit/calltool-dispatch.test.ts -> 65 pass
  npx vitest run test/integration/local-backend-calltool.test.ts -> 18 pass
  npm run test:unit                                   -> 3777 pass
    (4 pre-existing env failures unchanged: skip-git-cli needs
     built dist/, git-utils tmpdir on Windows worktree)
  npx tsc --noEmit                                    -> clean

Scope declined for v1:

 - In-process histogram aggregation — the log line is enough for
   external tooling
 - Pareto curve generation — issue asks to enable it, not generate it
 - Sub-phases of symbol_lookup (process vs cohesion vs content) —
   issue lists them under one bucket; can split later if demand surfaces

Closes #553

* fix(search): route query:timing log to stderr to preserve stdio MCP contract

CI (#953) failed the `query: JSON appears on stdout, not stderr`
e2e test in test/integration/cli-e2e.test.ts with:

  SyntaxError: Unexpected token 'G', "GitNexus [..." is not valid JSON

Root cause: my initial logQueryTiming() in 63fbdc4 used console.log,
which writes to stdout. The MCP stdio transport uses stdout
exclusively for JSON-RPC responses (#324), and the CLI e2e test
guards that contract by asserting stdout parses as JSON on every
tool invocation. The "GitNexus [query:timing] ..." line was
interleaving with the response JSON and breaking the parse.

Fix: route logQueryTiming through console.error instead. stderr is
the correct channel for human-readable diagnostics and it is what
the sibling logQueryError already uses for the same reason. The log
line format is otherwise unchanged -- still greppable, still
JSON-parseable payload.

Verification (local, with dist built):
  npx vitest run test/integration/cli-e2e.test.ts -t "query: JSON"
    -> now passes (was failing across ubuntu/windows/macos in CI)
  npx tsc --noEmit                                  -> clean
  Two unrelated pre-existing failures on non-git
  directory handling remain (same on upstream/main).

Closes the CI regression introduced in 63fbdc4.
2026-04-18 16:30:07 +01:00
Gergő Magyar 5d76dbcfa2 feat(shared): MethodDispatchIndex materialized view over HeritageMap (#914, RFC #909 Ring 2 SHARED) (#960)
Implements RFC §3.1 `MethodDispatchIndex`: a two-way materialized view
keyed by `DefId` for O(1) method-dispatch resolution:

  - `mroByOwnerDefId`       — owner class → full MRO ancestor chain
                              (excludes self, per-language strategy order)
  - `implsByInterfaceDefId` — interface/trait → classes that implement it

**Not an MRO implementation.** `buildMethodDispatchIndex` is a pure
aggregator that calls back into caller-provided `computeMro` and
`implementsOf` functions. The five existing strategies (Python C3, Ruby
kind-aware, Java/Kotlin linear, Rust qualified-syntax, COBOL none) stay
where they are today (`model/resolve.ts`, `languages/ruby.ts`); this index
does not reimplement them.

Why callbacks rather than a shared registry: the strategies depend on the
CLI's `HeritageMap` + `SemanticModel`. Migrating both to `gitnexus-shared`
is out of scope for #914; callbacks let the shared build stay pure.

Module placement: `gitnexus-shared/src/scope-resolution/method-dispatch-index.ts`
for consistency with the other RFC §3.1 indexes (#913 DefIndex /
ModuleScopeIndex / QualifiedNameIndex; #916 resolveTypeRef).

Safety surface mirrors sibling indexes:
  - First-write-wins on duplicate owners.
  - Repeated (interface, owner) pairs deduplicated.
  - Stored arrays are `Object.freeze`d; caller mutation of the source
    array does not leak into the index.
  - Miss returns a shared frozen empty array.

Tests (19, all passing): empty input, single-inheritance chain, Python
C3 diamond, Java BFS, Ruby kind-aware mixin, Rust qualified-syntax empty,
interface inversion (single, multiple, ordered), dedup within and across
callback calls, frozen miss + bucket arrays, callback-array isolation,
readonly Map iteration.

Closes part of #909.
2026-04-18 16:28:46 +01:00
Gergő Magyar 56e32b310b feat(shared): resolveTypeRef strict single-return type resolver (#916, RFC #909 Ring 2 SHARED) (#959)
Implements RFC §4.6: a strict, pure resolver for `TypeRef`s used by
`Registry.lookup` Step 2 (type-binding propagation) and by any caller that
wants the single best type-target for an annotation without paying for the
full evidence pipeline.

Algorithm (strict):

  1. Walk the scope chain from `ref.declaredAtScope`:
     - Return the first binding for `rawName` whose origin is in
       `{'local','import','namespace','reexport'}` AND whose `def.type` is a
       type-kind (class-like, interface-like, enum-like, alias-like).
     - If bindings exist but none qualify (non-type shadow, wildcard-only
       origin), return null immediately — do NOT fall through to the global
       qualified-name index.
  2. If `rawName` is dotted and the scope walk produced no match, consult
     `QualifiedNameIndex.byQualifiedName`. Only accept a UNIQUE type-kind
     hit; ambiguous or non-type results return null.

`'wildcard'` is deliberately excluded from strict origins — a
wildcard-expanded name is too loose to anchor type resolution.

Module placement: `gitnexus-shared/src/scope-resolution/resolve-type-ref.ts`
(alongside sibling indexes) rather than the issue's suggested
`gitnexus-shared/src/resolve-type-ref.ts`, for consistency with the rest of
the RFC §2/§3 surface.

A minimal `ScopeLookup` interface is declared inline so #916 ships
standalone; #912's `ScopeTree` will satisfy this contract without change.

Closes part of #909.
2026-04-18 16:09:54 +01:00
Gergő Magyar ac2012e5ed feat(shared): DefIndex / ModuleScopeIndex / QualifiedNameIndex (#913, RFC #909 Ring 2 SHARED) (#958)
Three flat O(1) indexes + pure build functions over per-file artifacts.
Contract-only; no runtime behavior change yet — consumers (#917 Registry
lookups, #915 SCC finalize, #919 ScopeExtractor) wire in later.

Each index follows the same shape:
  - build function: flat input list → frozen immutable index
  - public interface: readonly Map + get/has/size accessors
  - first-write-wins on id/filePath collisions (upstream bug signal)
  - pure, side-effect-free, safe to call repeatedly

DefIndex — the global "what is this id?" lookup
  gitnexus-shared/src/scope-resolution/def-index.ts
  buildDefIndex(defs: readonly SymbolDefinition[]): DefIndex
    byId: ReadonlyMap<DefId, SymbolDefinition>
  Consumed by Registry.lookup (#917) to materialize DefId[] hits back to
  full SymbolDefinition records.

ModuleScopeIndex — `filePath → moduleScopeId` for cross-file hops
  gitnexus-shared/src/scope-resolution/module-scope-index.ts
  buildModuleScopeIndex(entries): ModuleScopeIndex
    byFilePath: ReadonlyMap<string, ScopeId>
  Consumed by the SCC finalize link pass (#915) to resolve
  ImportEdge.targetFile to a concrete module scope in constant time.

QualifiedNameIndex — cross-kind qualified-name fast path
  gitnexus-shared/src/scope-resolution/qualified-name-index.ts
  buildQualifiedNameIndex(defs: readonly SymbolDefinition[]): QualifiedNameIndex
    byQualifiedName: ReadonlyMap<string, readonly DefId[]>
  Returns DefId[] (not a single DefId) because partial classes, method
  overloads, and cross-kind collisions can legitimately share a
  qualifiedName. Callers filter by acceptedKinds at the lookup site.
  Consumed by Registry.lookup qualified fast path + resolveTypeRef
  dotted fallback (#916, #917).

Barrel re-exports added to gitnexus-shared/src/index.ts so consumers
import from 'gitnexus-shared' rather than deep paths.

Tests (gitnexus/test/unit/scope-resolution/, 23 total):
  def-index.test.ts (6):
    empty, single def, multiple distinct, first-write-wins collision,
    missing id returns undefined, byId direct iteration
  module-scope-index.test.ts (6):
    empty, single entry, multiple files, first-write-wins on duplicate
    filePath, missing returns undefined, byFilePath direct iteration
  qualified-name-index.test.ts (11):
    empty, single qnamed def, partial classes accumulate, input-order
    preservation, qname separation, skip undefined/empty qname, pair
    dedup, cross-kind indexing, frozen-empty-array on miss, direct
    iteration

Verification:
  - gitnexus-shared + gitnexus build clean (tsc + scripts/build.js)
  - test/unit/scope-resolution: 23/23 pass
  - model + shadow + scope-resolution combined: 129/129 pass
  - No runtime consumer wiring yet — indexes are standalone library
    functions that #915, #917, #919 will import when ready

Depends on #910 (SymbolDefinition, DefId, ScopeId types — already on main).
Unblocks #915 (finalize algorithm), #917 (Registry.lookup), #919
(ScopeExtractor materialization).
2026-04-18 15:59:34 +01:00
Copilotandmagyargergo f73389eac3 fix: ENOBUFS in detect_changes by setting maxBuffer on git/rg execFileSync (#957)
* Initial plan

* Fix ENOBUFS in detect_changes by setting maxBuffer on git/rg execFileSync

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/bb241ed0-3b39-431f-a242-b0c7ced9707b

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-04-18 15:58:31 +01:00
Gergő Magyar 22f0beb057 feat(shared): shadow-mode diff + aggregate — full implementation (#918, RFC #909 Ring 2 SHARED) (#951)
Replaces the scaffold stubs with working pure-logic implementations plus
unit-test coverage for both functions. Unblocks Ring 2 PKG #923 (shadow
harness) to consume a concrete library instead of throwing scaffolds.

gitnexus-shared/src/scope-resolution/shadow/diff.ts
  `diffResolutions(callsite, legacy, newResult): ShadowDiff`
    - [0] on each side is the top match
    - both empty         → 'both-empty',   delta []
    - legacy empty only  → 'only-new',     delta = new top evidence
    - new empty only     → 'only-legacy',  delta = legacy top evidence
    - same top nodeId    → 'both-agree',   delta []
    - different nodeIds  → 'both-disagree',
                           delta = symmetric difference of evidence kinds
                           (legacy-only first in input order, then new-only)
  Evidence identity is `ResolutionEvidence.kind` — weight/note differences
  for the same kind do NOT produce delta entries. Rationale: the aggregator
  wants to know which *signals* explain a disagreement, not fluctuations
  in calibration values.

gitnexus-shared/src/scope-resolution/shadow/aggregate.ts
  `aggregateDiffs(diffs, now?): ShadowParityReport`
    - buckets by `SupportedLanguages`
    - tallies agreements, evidence-breakdown (divergences only — agree and
      empty rows do not contribute)
    - parity = bothAgree / (totalCalls - bothEmpty), yields 0 (not NaN)
      when the denominator is 0
    - perLanguage sorted alphabetically by enum value for stable output
    - evidenceBreakdown internally sorted by kind for stable output
    - overall = column-wise sum across languages
    - `now` parameter makes generatedAt deterministic in tests

gitnexus-shared/src/index.ts
  Re-exports the full shadow API: diffResolutions, aggregateDiffs, and all
  their types (ShadowAgreement, ShadowCallsite, ShadowDiff,
  LanguageParityRow, ShadowParityReport).

gitnexus/test/unit/shadow/diff.test.ts (13 tests)
  - 5 agreement outcomes
  - symmetric-by-kind evidence delta (disjoint, overlapping, fully-overlapping)
  - weight-only differences produce no delta
  - top-match only (ignores indices beyond [0])
  - callsite passthrough
  - delta ordering (legacy-only first, input order preserved)

gitnexus/test/unit/shadow/aggregate.test.ts (9 tests)
  - empty input
  - single language, all agree / mixed / all empty
  - multi-language bucketing + overall sum
  - alphabetical language sort
  - evidence breakdown scope
  - determinism via injected `now` + JSON round-trip identity

Verification:
  - gitnexus-shared + gitnexus build clean (tsc + scripts/build.js)
  - test/unit/shadow: 22/22 pass
  - test/unit/model + test/unit/shadow combined: 106/106 pass
  - No runtime behavior changes (shadow is invoked by #923, not yet wired)

Stacked on main (af1d278a). Depends on types from #910 (merged).
Unblocks: #923 (Ring 2 PKG — shadow harness wiring) — concrete library
to consume instead of scaffold stubs.

Plan: docs/plans/2026-04-18-001-refactor-911-senior-hooks-redesign-plan.md
is about #911; #918's scope is the scaffold+fill-in described in the PR
description of #951.
2026-04-18 15:32:51 +01:00
Gergő Magyar af1d278a7e feat(shared,ingestion): extend LanguageProvider with scope-resolution hooks (#911, RFC #909 Ring 1) (#950)
Adds the 14 optional scope-resolution hooks from RFC #909 §5.2 to
`LanguageProviderConfig` plus the supporting input/output types in
`gitnexus-shared`. Contract-only; no runtime behavior changes.

Review-driven refinements (addresses two non-blocking review comments on #950):

1. `ParsedImport` is now a 5-variant discriminated union, not a flat
   record. Each variant carries only its legal fields so invalid shapes
   are compile errors:
     - 'named', 'alias', 'namespace', 'reexport', 'dynamic-unresolved'
   'wildcard-expanded' is deliberately excluded — finalize materializes
   that kind; a provider must never emit it at parse time.
   'reexport' is a first-class parse-phase variant so syntactically-
   detectable re-exports (TS `export { X } from './y'`, Rust
   `pub use foo::bar`) keep their parse-time signal through to finalize
   rather than being re-derived by the SCC pass.
   `namespace` gains an `importedName` field so `import numpy as np`
   can carry both `localName: 'np'` and `importedName: 'numpy'`.
   `dynamic-unresolved.targetRaw` is `string | null` (was mandatory
   null) so providers can emit the unresolvable expression text for
   diagnostics when available.

2. `bindingScopeFor` and `importOwningScope` return type changed from
   `ScopeId` to `ScopeId | null`, aligning with the X | null convention
   used by the 12 sibling optional hooks (receiverBinding,
   resolveScopeKind, interpretTypeBinding, …). `null` = delegate to the
   central default. Enables partial overrides — a JS provider can
   return a hoisted scope for `var` and `null` for `let`/`const`
   without re-implementing the default lookup.
   Both hooks also gain a purity JSDoc contract: same inputs yield the
   same ScopeId (or null) across invocations; no closure over mutable
   state. Required to keep scope-tree construction deterministic.

   A richer callable-defaults pattern (typed BindingScopeDefaults /
   ImportOwningDefaults helper interfaces on a `defaults` parameter)
   was considered and deferred to Ring 2 PKG #919, where the concrete
   ScopeExtractor will exist to inform the helper shape. Designing that
   pattern before the first consumer would set cross-hook precedent
   based on a single motivating example.

Supporting types added to gitnexus-shared/src/scope-resolution/types.ts:
  - CaptureMatch, ParsedImport, ParsedTypeBinding
  - WorkspaceIndex, ScopeTree (opaque placeholders until Ring 2)
  - Callsite

14 hooks added to LanguageProviderConfig (all optional):
  Parse phase: emitScopeCaptures, interpretImport, receiverBinding,
    interpretTypeBinding, resolveScopeKind, shouldCreateScope,
    bindingScopeFor
  Finalize phase: resolveImportTarget, expandsWildcardTo,
    importOwningScope, mergeBindings
  Reference-extraction phase: classifyCallForm
  Resolution phase: shouldShadow, arityCompatibility

Verification:
  - gitnexus-shared builds clean (tsc)
  - gitnexus builds clean (scripts/build.js)
  - test/unit/model: 84/84 pass — no regressions
  - No provider needs updating (all hooks optional)
  - No BindingScopeDefaults/ImportOwningDefaults/defaults parameter
    introduced (deferred to #919)

Stacked on #910 (merged as afc0a8b6); rebased on main.
Tracking: #909 (meta). Unblocks Ring 2 PKG (#919 ScopeExtractor,
#922 import adapters) and all Ring 3 per-language migrations.

Plan: docs/plans/2026-04-18-001-refactor-911-senior-hooks-redesign-plan.md
2026-04-18 14:54:51 +01:00
Gergő Magyar afc0a8b6c5 feat(shared): add scope-resolution types + constants (#910, RFC #909 Ring 1) (#949)
Lands the authoritative data model and constants for the pure scope-based
resolution RFC (#909) as Ring 1, part 1. No runtime behavior changes —
types + constants only.

New in gitnexus-shared/src/scope-resolution/:
  - types.ts — Scope, ScopeKind, ScopeId, DefId, Range, Capture,
    BindingRef, ImportEdge, TypeRef, Resolution, ResolutionEvidence,
    Reference, ReferenceIndex, LookupParams, RegistryContributor
  - evidence-weights.ts — EvidenceWeights constant map + typeBindingWeightAtDepth
    (RFC Appendix A)
  - origin-priority.ts — ORIGIN_PRIORITY constant map for deterministic
    tie-breaks (RFC Appendix B)
  - language-classification.ts — LanguageClassification type +
    LanguageClassifications map (production × 14, experimental × 2
    for vue/cobol; governs Ring 4 DAG-retirement gate)
  - symbol-definition.ts — SymbolDefinition moved from
    gitnexus/src/core/ingestion/model/symbol-table.ts so scope-resolution
    types can reference it from the shared package

Consumer updates:
  - symbol-table.ts: removes local SymbolDefinition declaration; imports
    from gitnexus-shared
  - model/index.ts: drops SymbolDefinition from barrel re-export per
    "direct imports from gitnexus-shared" convention (see
    gitnexus-shared feedback in project memory)
  - 9 source files + 5 test files: import SymbolDefinition directly
    from 'gitnexus-shared'

Verification:
  - gitnexus-shared builds clean (tsc)
  - gitnexus builds clean (scripts/build.js)
  - 131/132 unit test files pass; 3767 tests green
  - Zero behavior changes; SymbolDefinition shape unchanged

Blocks: #911 (LanguageProvider hook interface extensions) and all of
Ring 2 (#912-#925). Closes part of #909.
2026-04-18 12:55:09 +01:00
Gergő Magyar d9da7d6692 fix(test): isolate cli-e2e from shared mini-repo fixture (#954)
Deterministic fix for the Windows-flaky pipeline-graph-golden test.

Root cause
  cli-e2e.test.ts wrote into the SHARED fixture directory
  (test/fixtures/mini-repo/) — git init, analyze run that creates
  AGENTS.md, CLAUDE.md, .claude/, .gitnexus/. When pipeline-graph-golden
  ran in parallel, its `cpSync` of the source directory could capture
  the mid-flight pollution before cli-e2e's afterAll cleanup fired.
  macOS/Ubuntu won the race often enough that the flake presented as
  Windows-only.

Fix
  cli-e2e now copies mini-repo into a fresh `mkdtemp`'d parent whose
  basename is `mini-repo` (preserving `--repo mini-repo` CLI lookup by
  basename), runs git-init there, and rm's the whole tmpdir in afterAll.
  The shared fixture source is never touched.

  Fallout from the cwd change: bare `--import tsx` specifiers (2
  spawnSync + 1 spawn) can't resolve `tsx` from an os.tmpdir cwd where
  there is no node_modules. Switched them to the already-existing
  `tsxImportUrl` (absolute file:// URL to the tsx loader), matching
  the `runCliOutsideProject` pattern that was already set up for this
  exact case.

  Updated the "MINI_REPO is inside the project tree" comment in the
  `status on non-indexed repo` test — MINI_REPO is now in os.tmpdir,
  so the rationale for using a separate throwaway tmp git repo is
  different (but still valid: previous tests in the suite create
  MINI_REPO/.gitnexus, which findRepo() would pick up).

  Also updated pipeline-graph-golden's comment explaining WHY it
  copies to tmp — it's now defense-in-depth rather than a necessity,
  so a future test that adds files to the source can't silently
  regress the golden.

Verification
  - 5x consecutive `cli-e2e + pipeline-graph-golden` runs: 20/20 pass
    (deterministic)
  - 3x full suite including pipeline.test: 27/27 pass
  - test/fixtures/mini-repo/ post-run contents: only `src/` —
    zero pollution from any test
  - macOS/Ubuntu behavior unchanged (they were passing; tmpdir
    isolation is purely additive)
2026-04-18 12:54:59 +01:00
azizur100389 131d411ae4 feat(mcp): rank context/impact disambiguation candidates and expose kind/file_path hints (#888)
* feat(mcp): rank context/impact disambiguation candidates and expose kind/file_path hints

The `context` MCP tool already returned `{ status: 'ambiguous', candidates }`
when a name hit multiple symbols, but the candidates were returned in
arbitrary DB order and the only hint it accepted was file_path. The
`impact` tool was worse: when its name resolver found multiple viable
matches it silently picked the first one from a priority UNION, with no
signal back to the caller that a different symbol might have been
intended.

Both failure modes were flagged in issue #470 and reconfirmed in the
comments by a second user who described impact as returning "incorrect
parsing results and meaningless tool calls" in the multi-match case.

Changes:

* Add `resolveSymbolCandidates(repo, query, hints)` private helper on
  LocalBackend. Single place that:
   - Short-circuits on direct uid (zero-ambiguity)
   - Runs the same name-or-qualified-id match as before, with LIMIT 20
     (was 10) so the ranker has headroom instead of arbitrary truncation
   - Preserves the #480 Class/Constructor preference -- when the only
     ambiguity is a Class and its own Constructor, the Class wins
     silently
   - Scores each candidate (pure TS, no extra DB round-trip): base 0.50,
     +0.40 for file_path match, +0.20 for kind match, plus a small
     kind-priority tiebreaker (Class > Interface > Function > Method >
     Constructor) when no explicit kind hint is given
   - Sorts desc by score with stable tiebreakers (shorter filePath,
     then lex uid)
   - Promotes to a single confident resolve when the top score is
     >= 0.95 AND beats the runner-up by >= 0.10 -- lets a strong hint
     cut through without forcing the caller through a disambiguation
     round-trip

* Rewire `context()` to use the shared helper. Response shape is a
  strict superset of today's: candidates gain a `score` field, the
  existing `{ uid, name, kind, filePath, line }` keys are preserved so
  every downstream consumer (rename, eval-server formatter, etc.) keeps
  working. New `kind` input hint accepted.

* Rewire `impact()` to use the shared helper. Now emits the same
  `{ status: 'ambiguous', candidates, impactedCount: 0, risk: 'UNKNOWN' }`
  shape instead of silent first-pick. New inputs accepted:
  `target_uid`, `file_path`, `kind`.

* Update tool schemas in mcp/tools.ts to advertise the new inputs and
  describe ranked disambiguation.

Backward compatibility:

The #480 Class/Constructor collapse is preserved and covered by the
existing java-class-impact integration test (still green). The
ambiguous response shape is a strict superset -- `eval-formatters`
unit test that parses the old shape is unchanged and still passes.
`impact` going from silent-first-pick to structured ambiguous is a
semantic improvement that is the entire point of the issue; callers
relying on silent first-pick now get an actionable response.

Scope declined for v1:

module/community hint -- the issue lists it as one of several hints,
but kind + file_path cover the vast majority of disambiguation needs
in practice, and a community-label filter requires an extra graph
query per candidate. Natural v2 follow-up.

Tests: calltool-dispatch.test.ts gains 5 new cases covering file_path
boost, kind hint boost, impact ambiguous shape, impact target_uid
short-circuit, and score field presence on the existing ambiguous
test. Plus the extended assertions on the existing
`context tool returns disambiguation for multiple matches`.

Verification:
  npx vitest run test/unit/calltool-dispatch.test.ts       -> 64 pass
  npx vitest run test/integration/java-class-impact.test.ts -> pass
  npm run test:unit                                         -> 3642 pass
    (4 pre-existing env failures unchanged: skip-git-cli needs built
    dist/, git-utils tmpdir on Windows worktree -- same on main)
  npx tsc --noEmit                                          -> clean

Closes #470

* fix(mcp): enrich labels from UNION when labels(n)[0] is empty; address review findings

CI on PR #888 caught 13 integration-test failures I did not cover locally:
my resolver refactor collected candidates via `labels(n)[0] AS type`, but
LadybugDB returns an empty string for that projection on certain node
types (most importantly Class). With an empty `type`, impact's downstream
`_runImpactBFS` no longer recognised `symType === 'Class' | 'Interface'`
and stopped seeding Constructor + File nodes into the frontier, so the
"impact(upstream) surfaces the file importer" assertion broke across 11
language fixtures plus 2 OVERRIDES filter tests.

The original impact resolver worked around this by running a prioritised
UNION across Class/Interface/Function/Method/Constructor and picking the
first hit. My refactor dropped that. Fix: keep the simple candidate MATCH
but enrich types afterward via a single scoped UNION query, so every
candidate carries an accurate label for both scoring and downstream
BFS seeding. The UID direct-lookup path is patched the same way.

Also addresses the findings from the senior reviewer on PR #888:

* MIGRATION.md: document the `impact` behavioural change (silent first-
  pick → structured `{ status: 'ambiguous', candidates }`) so downstream
  callers know to branch on `result.status` before reading byDepth/
  summary. `context` is unchanged shape-wise (strict superset).

* New test: `context tool promotes top candidate via scoring when
  multiple rows survive DB pre-filter`. The review flagged that the
  existing file_path test works only because the mock ignores WHERE
  parameters -- the scored-promotion path (top ≥ 0.95 AND gap > 0.09)
  wasn't directly exercised. The new test uses two candidates both in
  App.tsx-containing paths plus a kind hint so promotion is decided by
  scoring, not DB pre-filtering. Also tightened the comment on the
  earlier file_path test to describe the mock vs production divergence
  honestly.

* NIT: added a paragraph explaining why `scored.length >= 2` is kept as
  a defensive guard even though the `normalized.length === 1` early
  return already covers the single-candidate path.

* Integration: two tests in `local-backend-calltool.test.ts` targeted
  `'authenticate'`, which now correctly resolves as ambiguous (two
  Method nodes: AuthService.authenticate and BaseService.authenticate).
  Updated both to pass `file_path: 'src/auth.ts'` so they exercise the
  new disambiguation API and still assert the METHOD_OVERRIDES filtering
  they were originally about.

Edge case fix in the promotion gap check: IEEE754 makes 0.50 + 0.40 +
0.20 - 0.90 = 0.09999999999999998 instead of exactly 0.10, which would
otherwise break the "winner clearly dominates" intent for legitimate
1.00 vs 0.90 cases. Changed `>= 0.10` to `> 0.09`; same user-facing
intent, no floating-point sensitivity.

Verification (all from gitnexus/):
  npx vitest run test/integration/class-impact-all-languages.test.ts
    -> 52 pass (was 11 FAIL on CI before this fix)
  npx vitest run test/integration/local-backend-calltool.test.ts
    -> 18 pass (was 2 FAIL on CI before this fix)
  npx vitest run test/integration/java-class-impact.test.ts
    -> 10 pass (regression guard for #480 preserved)
  npx vitest run test/unit/calltool-dispatch.test.ts
    -> 65 pass (1 new test + 4 from original #470 PR)
  npm run test:unit
    -> 3626 pass, 4 pre-existing env failures unchanged
  npx tsc --noEmit
    -> clean
2026-04-18 12:52:42 +01:00
Gergő Magyar b8875b9c80 chore(release): v1.6.2 (#952)
Bumps gitnexus to v1.6.2 and adds the matching CHANGELOG entry.

Highlights since v1.6.1 (61 commits):
  - Docker support (#848)
  - Language-agnostic heritage / call / variable extractors
    (config+factory pattern, #877 #878 #890)
  - AST-aware embedding chunking (#889)
  - jQuery / axios HTTP consumer detection (#887)
  - SemanticModel wired as first-class resolution input, SM-20 (#885)
  - ImportSemantics split into per-strategy hooks (#886)
  - Python dotted-import fix (#899); worker warnings non-terminal
    (#900 / #261); global-install ENOTEMPTY fixes (#843 #846);
    embeddings staleness fix (#831)

See gitnexus/CHANGELOG.md for the full list.

After merge, tag `v1.6.2` triggers publish.yml which runs CI,
verifies tag↔package.json match, publishes to npm with provenance,
and creates the GitHub Release using the extracted CHANGELOG body.
2026-04-18 12:16:51 +01:00
RyanbaandSisyphus 969b4623ca fix: keep worker warnings non-terminal (#900)
* fix: keep worker warnings non-terminal

Treat parse-worker warning messages as informational so a warning can be surfaced without short-circuiting the worker result protocol.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* style: apply prettier formatting

---------

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-04-18 11:09:15 +01:00
Copilotandmagyargergo 018e0e6b14 test(web-e2e): raise status-ready timeout to 45s for parallel-worker stability (#908)
* Initial plan

* plan: stabilize web e2e tests timing out under parallel workers

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/49d09e67-5a8e-4eee-adcd-3d5416675a6b

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(web-e2e): bump status-ready timeout to 45s for parallel-worker stability

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/49d09e67-5a8e-4eee-adcd-3d5416675a6b

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-04-18 10:15:28 +01:00
Kritik Bangeraandkritik.b 040bb7a489 feat: add docker support (#848)
* feat: add docker support

* feat: move docker files to root

* feat: add docker build and push workflow

* fix: pin docker action SHAs to verified commits

Made-with: Cursor

* fix: remove redundant --platform=$TARGETPLATFORM from runtime stage

Made-with: Cursor

* fix: upgrade docker actions to Node.js 24-compatible versions

Made-with: Cursor

* docs: updated readme

* fix: update docker references

* fix(docker-server): reject null bytes in resolvePath

Defensively harden the path traversal guard by returning null
early when the URL contains a null byte, before normalization runs.

Made-with: Cursor

* fix(docker-server): handle createReadStream errors

Attach an error listener before piping so mid-flight read errors
(truncated file, permission change) cleanly destroy the response
instead of being silently swallowed.

Made-with: Cursor

* fix(docker-server): replace existsSync with async stat

Eliminates the TOCTOU race between the initial stat call and the
subsequent existsSync check. Reuses the async stat pattern already
in place and removes the now-unused existsSync import.

Made-with: Cursor

* test(docker-server): add integration tests; fix %00 null-byte bypass

Decode the URL before the null-byte check so percent-encoded null
bytes (%00) are also rejected with 400 instead of falling through
to the SPA fallback. Adds 5 node:test integration tests covering
valid assets, SPA fallback, path traversal, null bytes, and 404.

Made-with: Cursor

* style: fix prettier formatting in docker-server files

Made-with: Cursor

* fix(docker): wire tests into CI, fix resolvePath separator, correct image namespace

- Add `node --test docker-server.test.mjs` step to ci-tests.yml so the
  path-traversal guard tests run in every CI pass instead of being silently skipped.
- Fix resolvePath containment check: `startsWith(root)` would allow sibling
  directories like `/app/dist-evil/`; now guards with `root + sep` or exact match.
- Update docker-compose.yaml default image from `abhigyanpatwari` namespace to
  `brainifii` to match what docker.yml publishes to GHCR.

* fix(docker): update apt-get commands and set user permissions

- Modify Dockerfile and Dockerfile.test to include options for apt-get to bypass validity checks during updates.
- Set ownership of the /app directory to the 'node' user in the runtime stage for improved security and proper permission handling.

* fix(docker): switch to Alpine base images for smaller footprint

- Update Dockerfile to use Alpine-based Node.js images for both builder and runtime stages, reducing image size and improving performance.
- Replace apt-get commands with apk for package installation in the runtime stage.

* fix(docker): update Node.js version in Dockerfile

- Change base image from node:20-alpine to node:22-alpine

* fix(docker): update Node.js version in Dockerfile to 22-alpine for runtime

---------

Co-authored-by: kritik.b <kritik.b@media.net>
2026-04-18 08:39:18 +01:00
dependabot[bot] 725ed3fe66 chore(deps)(deps): bump glob from 11.1.0 to 13.0.6 in /gitnexus (#867) 2026-04-18 07:40:08 +01:00
dependabot[bot] 509185b8f9 chore(deps)(deps): bump commander from 12.1.0 to 14.0.3 in /gitnexus (#868) 2026-04-18 07:28:17 +01:00
dependabot[bot] 4988feec94 chore(deps)(deps-dev): bump wait-on from 8.0.5 to 9.0.5 in /gitnexus-web (#859) 2026-04-18 07:26:56 +01:00
dependabot[bot] ef953beca9 chore(deps)(deps): bump @huggingface/transformers in /gitnexus (#869) 2026-04-18 07:22:34 +01:00
dependabot[bot] 0b5381695f chore(deps)(deps-dev): bump @vitest/coverage-v8 in /gitnexus (#864) 2026-04-18 07:21:02 +01:00
dependabot[bot] 30292d7179 chore(deps)(deps): bump @modelcontextprotocol/sdk in /gitnexus (#866) 2026-04-18 07:20:19 +01:00
dependabot[bot] 7a98a01ad5 chore(deps)(deps): bump lru-cache from 11.2.7 to 11.3.5 in /gitnexus (#870) 2026-04-18 07:19:49 +01:00
dependabot[bot] 94cba48b6d chore(deps)(deps): bump mnemonist from 0.39.8 to 0.40.3 in /gitnexus (#871) 2026-04-18 07:19:24 +01:00
azizur100389 925460ab5b refactor(cli): trim duplicated ai-context CLAUDE.md block (#904) 2026-04-18 07:10:44 +01:00
dependabot[bot] 08b4505197 chore(deps)(deps): bump @ladybugdb/core in /gitnexus (#873) 2026-04-17 21:28:34 +01:00
dependabot[bot] d5225e699f chore(deps)(deps): bump mermaid from 11.12.2 to 11.14.0 in /gitnexus-web (#860) 2026-04-17 21:27:37 +01:00
dependabot[bot] 0a3b9120a0 chore(deps)(deps): bump tailwindcss in /gitnexus-web (#861) 2026-04-17 21:27:07 +01:00
dependabot[bot] 5544350e30 chore(deps)(deps-dev): bump jsdom from 29.0.0 to 29.0.2 in /gitnexus-web (#863) 2026-04-17 21:26:50 +01:00
Copilot dfa449ef41 feat(ingestion): language-agnostic heritage extractor with config+factory pattern (#890) 2026-04-17 17:51:17 +01:00
Yacine Hmito daca8360bf fix(python): avoid local matches for external dotted imports (#899) 2026-04-17 11:35:59 +01:00
317 changed files with 32068 additions and 2223 deletions
+1
View File
@@ -0,0 +1 @@
plans/
+21
View File
@@ -0,0 +1,21 @@
.git
.gitignore
.DS_Store
node_modules
**/node_modules
dist
**/dist
coverage
**/coverage
.env
.env.local
.env.*.local
**/*.tsbuildinfo
.gitnexus
gitnexus-web/playwright-report
gitnexus-web/test-results
+15
View File
@@ -0,0 +1,15 @@
# Images (signed Cosign keyless on every push from main / vX.Y.Z tags)
SERVER_IMAGE=ghcr.io/abhigyanpatwari/gitnexus:latest
WEB_IMAGE=ghcr.io/abhigyanpatwari/gitnexus-web:latest
# Container names
SERVER_CONTAINER_NAME=gitnexus-server
WEB_CONTAINER_NAME=gitnexus-web
# Host ports — the web UI expects the server on http://localhost:4747 by default.
SERVER_HOST_PORT=4747
WEB_HOST_PORT=4173
# Optional read-only mount, exposed to the server as /workspace.
# Override with the directory that contains the repos you want to index.
WORKSPACE_DIR=./
+13 -7
View File
@@ -11,10 +11,14 @@ Rules:
`concurrency:` block.
2. Reusable workflows (on: workflow_call ONLY) do NOT declare one.
3. The `concurrency.group` expression MUST reference either
`${{ github.workflow }}` or a literal `CI-` prefix (the documented
ci.yml reusable-workflow-safe exception). This is checked by substring
containment rather than prefix match because ci.yml's group is a
conditional expression that resolves to a `CI-…` literal at runtime.
`${{ github.workflow }}` or one of the approved hardcoded literal prefixes
for workflows that are simultaneously entry-points AND reusable (on: push/
workflow_call). Two such exceptions are currently approved:
- `CI-` for ci.yml (the original canonical form)
- `docker-build-push-` for docker.yml
This is checked by substring containment rather than prefix match because
the group value is a conditional expression that resolves to a `CI-…` or
`docker-build-push-…` literal at runtime.
We deliberately do not use a YAML library — keeps the script dependency-free
on any vanilla runner. `on:` block parsing is line-based and handles both the
@@ -28,7 +32,7 @@ import re
import sys
REQUIRED_TOKENS = ("${{ github.workflow }}", "CI-")
REQUIRED_TOKENS = ("${{ github.workflow }}", "CI-", "docker-build-push-")
def is_reusable(lines: list[str]) -> bool:
@@ -150,8 +154,10 @@ def check(workflows_dir: pathlib.Path) -> int:
if not any(token in group for token in REQUIRED_TOKENS):
print(
f"::error file={path}::concurrency.group `{group}` must "
f"reference one of {REQUIRED_TOKENS}. See CONTRIBUTING.md -> "
"GitHub Actions — Concurrency Convention."
f"reference one of {REQUIRED_TOKENS} (use ${{{{ github.workflow }}}} "
"for normal entry-point workflows; use an approved literal prefix "
"only for workflows that are both entry-points AND reusable — "
"see CONTRIBUTING.md -> GitHub Actions — Concurrency Convention)."
)
fail = 1
+2 -2
View File
@@ -9,7 +9,7 @@ jobs:
timeout-minutes: 5
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
- uses: actions/setup-node@53b83947a5a98c8d113130e565377fae1a50d02f # v6.3.0
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 20
cache: npm
@@ -22,7 +22,7 @@ jobs:
timeout-minutes: 10
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
- uses: actions/setup-node@53b83947a5a98c8d113130e565377fae1a50d02f # v6.3.0
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 20
cache: npm
+109
View File
@@ -0,0 +1,109 @@
name: Scope Resolution Parity
# Reusable workflow — called from ci.yml. Does NOT declare concurrency;
# it inherits the caller's concurrency group per the convention documented
# in CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
#
# ── Purpose (RFC #909 Ring 3, §6.4 "Observability gates") ──────────────
# For every language in `MIGRATED_LANGUAGES` (exported from
# `gitnexus/src/core/ingestion/registry-primary-flag.ts`), run the
# resolver integration test at `test/integration/resolvers/<slug>.test.ts`
# TWICE on every PR:
#
# 1. `REGISTRY_PRIMARY_<LANG>=0` — legacy DAG path (guarantees we haven't
# broken the old path while migrating).
# 2. `REGISTRY_PRIMARY_<LANG>=1` — registry-primary path (guarantees the
# new path carries the same behavior — the parity gate).
#
# BOTH must pass. The source of truth is the TypeScript constant — adding
# a language to that `Set` is the ONLY contributor action; CI auto-
# discovers it, runs parity, and the language's default production path
# flips to registry-primary in the same change.
#
# When the set is empty (e.g. mid-Ring-3 for every language), the parity
# matrix is skipped and the workflow reports success — no-op until a
# language is explicitly claimed migrated.
on:
workflow_call:
jobs:
discover:
name: Discover migrated languages
runs-on: ubuntu-latest
timeout-minutes: 5
outputs:
languages: ${{ steps.read.outputs.languages }}
has-any: ${{ steps.read.outputs.has-any }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
- uses: ./.github/actions/setup-gitnexus
- name: Extract MIGRATED_LANGUAGES from registry-primary-flag.ts
id: read
shell: bash
working-directory: gitnexus
run: |
set -euo pipefail
# `tsx` evaluates the TS source directly (no build step), imports
# the exported `Set`, and emits a GH-Actions-friendly JSON matrix.
LANGS=$(npx tsx scripts/ci-list-migrated-languages.ts)
COUNT=$(printf '%s' "$LANGS" | jq 'length')
HAS_ANY="false"
if [[ "$COUNT" -gt 0 ]]; then HAS_ANY="true"; fi
echo "languages=$LANGS" >> "$GITHUB_OUTPUT"
echo "has-any=$HAS_ANY" >> "$GITHUB_OUTPUT"
echo "Discovered $COUNT migrated language(s): $LANGS"
echo "Parity matrix will run: $HAS_ANY"
parity:
name: ${{ matrix.lang.slug }} parity
needs: discover
if: needs.discover.outputs.has-any == 'true'
runs-on: ubuntu-latest
timeout-minutes: 20
strategy:
# One language failing must not abort the others — we want the full
# parity matrix result on a single CI run so a reviewer sees every
# regression at once rather than one-at-a-time.
fail-fast: false
matrix:
lang: ${{ fromJSON(needs.discover.outputs.languages) }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
- uses: ./.github/actions/setup-gitnexus
with:
build: 'true'
- name: Verify resolver test file exists
shell: bash
working-directory: gitnexus
run: |
set -euo pipefail
TEST_FILE="test/integration/resolvers/${{ matrix.lang.slug }}.test.ts"
if [[ ! -f "$TEST_FILE" ]]; then
echo "::error title=Missing resolver test::\
Expected $TEST_FILE for '${{ matrix.lang.slug }}' (listed in \
MIGRATED_LANGUAGES). Either fix the slug or add the test file \
before listing this language as migrated."
exit 1
fi
- name: Resolver tests — legacy DAG (REGISTRY_PRIMARY_${{ matrix.lang.envvar }}=0)
shell: bash
working-directory: gitnexus
env:
FLAG_NAME: REGISTRY_PRIMARY_${{ matrix.lang.envvar }}
# Explicitly force the flag to `0` even though it also defaults to
# `MIGRATED_LANGUAGES.has(lang)` — once a language is in the set,
# the default flips to registry-primary, so an unset env var would
# silently re-run the same path as step #2. `env FOO=0 cmd` spawns
# `cmd` with the override scoped to just this invocation.
run: env "$FLAG_NAME=0" npx vitest run "test/integration/resolvers/${{ matrix.lang.slug }}.test.ts"
- name: Resolver tests — registry-primary (REGISTRY_PRIMARY_${{ matrix.lang.envvar }}=1)
shell: bash
working-directory: gitnexus
env:
FLAG_NAME: REGISTRY_PRIMARY_${{ matrix.lang.envvar }}
run: env "$FLAG_NAME=1" npx vitest run "test/integration/resolvers/${{ matrix.lang.slug }}.test.ts"
+3
View File
@@ -41,6 +41,9 @@ jobs:
--outputFile=web-test-results.json
working-directory: gitnexus-web
- name: Run docker-server integration tests
run: node --test docker-server.test.mjs
- name: Upload test reports
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
+28 -6
View File
@@ -29,6 +29,8 @@ concurrency:
# ci-quality.yml — typecheck (tsc --noEmit)
# ci-tests.yml — unit + integration tests with coverage + cross-platform
# ci-e2e.yml — E2E tests (only when gitnexus-web/ changes)
# ci-scope-parity.yml — RFC #909 Ring 3 parity gate: legacy DAG + registry-primary
# both pass, per migrated language in the JSON registry
#
# Shared setup is DRY via .github/actions/setup-gitnexus composite action.
@@ -48,6 +50,11 @@ jobs:
permissions:
contents: read
scope-parity:
uses: ./.github/workflows/ci-scope-parity.yml
permissions:
contents: read
# ── Save PR metadata for the reporting workflow ─────────────────
# The ci-report.yml workflow (triggered by workflow_run) needs the
# PR number and job results to post a comment. We save them as an
@@ -56,7 +63,7 @@ jobs:
save-pr-meta:
name: Save PR Metadata
if: always() && github.event_name == 'pull_request'
needs: [quality, tests, e2e]
needs: [quality, tests, e2e, scope-parity]
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
@@ -67,12 +74,14 @@ jobs:
QUALITY: ${{ needs.quality.result }}
TESTS: ${{ needs.tests.result }}
E2E: ${{ needs.e2e.result }}
SCOPE_PARITY: ${{ needs.scope-parity.result }}
run: |
mkdir -p pr-meta
echo "$PR_NUMBER" > pr-meta/pr_number
echo "$QUALITY" > pr-meta/quality_result
echo "$TESTS" > pr-meta/tests_result
echo "$E2E" > pr-meta/e2e_result
echo "$PR_NUMBER" > pr-meta/pr_number
echo "$QUALITY" > pr-meta/quality_result
echo "$TESTS" > pr-meta/tests_result
echo "$E2E" > pr-meta/e2e_result
echo "$SCOPE_PARITY" > pr-meta/scope_parity_result
# TODO(post-merge): remove backward-compat copies once ci-report.yml
# on main reads underscore names.
# Backward-compat: ci-report.yml on main still reads hyphenated
@@ -95,7 +104,7 @@ jobs:
# Single required check for branch protection.
ci-status:
name: CI Gate
needs: [quality, tests, e2e]
needs: [quality, tests, e2e, scope-parity]
if: always()
runs-on: ubuntu-latest
timeout-minutes: 5
@@ -106,10 +115,12 @@ jobs:
QUALITY: ${{ needs.quality.result }}
TESTS: ${{ needs.tests.result }}
E2E: ${{ needs.e2e.result }}
SCOPE_PARITY: ${{ needs.scope-parity.result }}
run: |
echo "Quality: $QUALITY"
echo "Tests: $TESTS"
echo "E2E: $E2E"
echo "Scope parity: $SCOPE_PARITY"
if [[ "$QUALITY" != "success" ]] ||
[[ "$TESTS" != "success" ]]; then
echo "::error::Quality or test jobs failed"
@@ -119,3 +130,14 @@ jobs:
echo "::error::E2E job failed"
exit 1
fi
# scope-parity is a reusable workflow. With an empty migrated-
# languages list, its parity matrix is skipped and the outer
# workflow still reports `success`. If any entry's legacy-DAG or
# registry-primary run fails, the workflow reports `failure`.
# Accept only `success`; `skipped` would mean the entire
# discover job was skipped too (upstream failure), which should
# still block.
if [[ "$SCOPE_PARITY" != "success" ]]; then
echo "::error::Scope-resolution parity gate failed (RFC #909 Ring 3)"
exit 1
fi
+202
View File
@@ -0,0 +1,202 @@
name: Docker Build & Push
on:
push:
tags:
- 'v*'
# No workflow_dispatch: publishing is exclusively tag-driven so that every
# signed image corresponds 1:1 to a published `gitnexus@X.Y.Z` on npm. A
# manual run from a branch ref would fail the version check below anyway.
workflow_call:
inputs:
tag:
description: >-
The full v-prefixed tag to build (e.g. v1.2.3-rc.1).
The tag must already exist in the repo and its tree must contain
a gitnexus/package.json whose version matches the tag.
required: true
type: string
# Concurrency convention: see CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
# Tag refs are unique per release, so distinct tags run in parallel.
# Re-pushes of the same tag serialize. cancel-in-progress: false — never cancel a publish mid-flight.
# Hardcoded `docker-build-push-` prefix (not `${{ github.workflow }}`) when invoked as a reusable
# workflow: in called-workflow context `github.workflow` is ambiguous and could resolve to the
# caller's name, sharing a concurrency group with the caller → deadlock.
# Direct tag-push invocations use `docker-build-push-<ref>`; workflow_call invocations get a
# per-run-unique group (they are already serialized by the caller's own concurrency group).
concurrency:
group: ${{ (github.event_name == 'push') && format('docker-build-push-{0}', github.ref) || format('docker-build-push-nested-{0}', github.run_id) }}
cancel-in-progress: false
jobs:
build-push:
name: Build & Push ${{ matrix.image.name }}
runs-on: ubuntu-latest
timeout-minutes: 60
permissions:
contents: read
packages: write
# Required for Cosign keyless signing via the OIDC token exchange,
# and for build provenance / SBOM attestations.
id-token: write
attestations: write
strategy:
fail-fast: false
matrix:
image:
# Static UI bundle. Small, fast image. Drop-in replacement for the
# legacy single-image setup at the same `gitnexus` repository slug
# is intentionally avoided — the UI now lives at `gitnexus-web` and
# the CLI/server takes the canonical `gitnexus` slug below.
- name: gitnexus-web
dockerfile: Dockerfile.web
slug: gitnexus-web
# CLI / `gitnexus serve` backend. Heavy native deps (tree-sitter,
# onnxruntime-node) live only in this image.
- name: gitnexus
dockerfile: Dockerfile.cli
slug: gitnexus
steps:
- name: Validate tag input
if: github.event_name == 'workflow_call'
shell: bash
env:
TAG_INPUT: ${{ inputs.tag }}
run: |
if [ -z "${TAG_INPUT}" ]; then
echo "::error::No tag provided to docker.yml — refusing to build/push."
exit 1
fi
# When triggered by workflow_call the caller passes the RC tag as an input;
# we check out that tag so the Dockerfile and package.json match the built image.
# For tag-push events github.ref is already the tag ref — no override needed.
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
ref: ${{ inputs.tag || github.ref }}
# ── Lock the docker image version to the npm package version ──────────
# Mirrors the check in publish.yml: refuse to build unless the git tag
# exactly matches `gitnexus/package.json`'s version. This guarantees
# `ghcr.io/<owner>/gitnexus:X.Y.Z` always corresponds to the same
# `gitnexus@X.Y.Z` published to npm — no drift, no surprises.
- name: Verify tag matches gitnexus/package.json version
id: version
shell: bash
env:
# For workflow_call the tag comes from the caller input; for push events
# it is derived from GITHUB_REF (set to empty so the else-branch fires).
INPUT_TAG: ${{ inputs.tag }}
run: |
if [ -n "$INPUT_TAG" ]; then
TAG_VERSION="${INPUT_TAG#v}"
else
TAG_VERSION="${GITHUB_REF#refs/tags/v}"
fi
if ! [[ "$TAG_VERSION" =~ ^[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$ ]]; then
echo "::error::Tag does not follow semver: v$TAG_VERSION"
exit 1
fi
PKG_VERSION=$(node -p "require('./gitnexus/package.json').version")
if [ "$TAG_VERSION" != "$PKG_VERSION" ]; then
echo "::error::Tag version (v$TAG_VERSION) does not match gitnexus/package.json version ($PKG_VERSION)"
exit 1
fi
echo "version=$PKG_VERSION" >> "$GITHUB_OUTPUT"
echo "Version verified: $PKG_VERSION"
# Required for multi-platform (linux/arm64) emulation.
- name: Set up QEMU
uses: docker/setup-qemu-action@ce360397dd3f832beb865e1373c09c0e9f86d70a # v4.0.0
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@4d04d5d9486b7bd6fa91e7baf45bbb4f8b9deedd # v4.0.0
- name: Install Cosign
uses: sigstore/cosign-installer@cad07c2e89fa2edd6e2d7bab4c1aa38e53f76003 # v4.1.1
- name: Log in to GitHub Container Registry
uses: docker/login-action@4907a6ddec9925e35a0a9e82d7399ccc52663121 # v4.1.0
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
# Computes image tags and labels from the verified semver tag:
# v1.2.3 → :1.2.3, :1.2, :1, :latest (auto, only for non-prerelease)
# v1.2.3-rc.1 → :1.2.3-rc.1 only (prereleases never become :latest)
# `:latest` is only emitted for tag pushes thanks to `flavor: latest=auto`,
# ensuring it always points at a real npm-published version.
#
# For workflow_call invocations github.ref is the caller's branch ref, so
# the type=semver patterns would not match. In that case we add an explicit
# type=raw tag using the version already verified above, so the same
# image-naming rules apply regardless of how the workflow was triggered.
# NOTE: We check `inputs.tag` rather than `github.event_name` because in a
# reusable workflow the github context is inherited from the caller —
# `github.event_name` would still be "push", not "workflow_call".
- name: Extract Docker metadata
id: meta
uses: docker/metadata-action@030e881283bb7a6894de51c315a6bfe6a94e05cf # v6.0.0
with:
images: ghcr.io/${{ github.repository_owner }}/${{ matrix.image.slug }}
flavor: latest=auto
tags: |
type=semver,pattern={{version}}
type=semver,pattern={{major}}.{{minor}}
type=semver,pattern={{major}}
type=raw,value=${{ steps.version.outputs.version }},enable=${{ inputs.tag != '' }}
- name: Build and push
id: build
uses: docker/build-push-action@bcafcacb16a39f128d818304e6c9c0c18556b85f # v7.1.0
with:
context: .
file: ${{ matrix.image.dockerfile }}
platforms: linux/amd64,linux/arm64
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
cache-from: type=gha,scope=${{ matrix.image.slug }}
cache-to: type=gha,mode=max,scope=${{ matrix.image.slug }}
provenance: mode=max
sbom: true
# Cosign keyless signing. Each pushed tag is signed by the workflow's
# OIDC identity, so consumers can verify the image with the strict,
# fully-anchored identity regex (kept in sync with README.md and
# deploy/kubernetes/cluster-image-policy.yaml — update all three together).
# NOTE: `${...}` expression syntax is NOT evaluated inside YAML comments, so
# the example below uses literal `<owner>/<repo>` placeholders that consumers
# substitute themselves; the canonical, fully-rendered command lives in README.md.
# cosign verify ghcr.io/<owner>/<slug>:<tag> \
# --certificate-identity-regexp '^https://github\.com/<owner>/<repo>/\.github/workflows/docker\.yml@refs/tags/v[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$' \
# --certificate-oidc-issuer https://token.actions.githubusercontent.com
# Do NOT relax to `@.*` — that accepts signatures from any ref, including
# unprotected branches and PRs, and defeats the supply-chain guarantee.
- name: Sign image with Cosign (keyless)
env:
# Cosign v2 (installed by sigstore/cosign-installer above) makes
# keyless the default. COSIGN_EXPERIMENTAL is a v1-only opt-in flag
# that is now deprecated/no-op, so it is intentionally omitted.
DIGEST: ${{ steps.build.outputs.digest }}
TAGS: ${{ steps.meta.outputs.tags }}
run: |
# Sign every tag at the same digest so consumers can verify by tag or by digest.
# Use `while read` instead of `for $TAGS` to be robust against tags that
# could ever contain whitespace (the metadata-action output is newline-
# separated, not space-separated).
while IFS= read -r tag; do
[[ -n "$tag" ]] && cosign sign --yes "${tag}@${DIGEST}"
done <<< "$TAGS"
# Attach the SBOM produced by buildx as a verifiable attestation on the digest.
- name: Generate build provenance attestation
uses: actions/attest-build-provenance@a2bbfa25375fe432b6a289bc6b6cd05ecd0c4c32 # v4.1.0
with:
subject-name: ghcr.io/${{ github.repository_owner }}/${{ matrix.image.slug }}
subject-digest: ${{ steps.build.outputs.digest }}
push-to-registry: true
+1 -1
View File
@@ -33,7 +33,7 @@ jobs:
id-token: write
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
- uses: actions/setup-node@53b83947a5a98c8d113130e565377fae1a50d02f # v6.3.0
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 20
registry-url: https://registry.npmjs.org
+23 -1
View File
@@ -125,13 +125,15 @@ jobs:
permissions:
contents: write # push rc tag + marker
id-token: write # npm provenance
outputs:
vtag: ${{ steps.reltag.outputs.vtag }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
fetch-depth: 0
fetch-tags: true
- uses: actions/setup-node@53b83947a5a98c8d113130e565377fae1a50d02f # v6.3.0
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 20
registry-url: https://registry.npmjs.org
@@ -364,3 +366,23 @@ jobs:
Release candidates are pre-stable builds intended for early testing.
Stable releases remain on the `latest` dist-tag.
# ── Build & push RC Docker images ────────────────────────────────────
# Calls docker.yml as a reusable workflow so that the build, signing, and
# attestation logic stays in one place. The publish job exposes `vtag`
# (e.g. `v1.2.3-rc.1`) as an output so we can pass it as the tag input.
# RC images are signed with Cosign keyless signing; the OIDC identity
# will be `docker.yml@refs/heads/main` (the caller's ref) rather than a
# tag ref — see README.md § Docker for the correct verify command for RCs.
docker:
name: Build & Push RC Docker images
needs: [guard, publish]
if: needs.guard.outputs.should_run == 'true' && needs.publish.outputs.vtag != ''
uses: ./.github/workflows/docker.yml
permissions:
contents: read
packages: write
id-token: write
attestations: write
with:
tag: ${{ needs.publish.outputs.vtag }}
+4
View File
@@ -23,6 +23,7 @@ Thumbs.db
.env
.env.local
.env.*.local
docker/.env
# Logs
*.log
@@ -102,3 +103,6 @@ gitnexus/vendor/**/node_modules/
local_docs/
# Local agent scratch / review prompts (never commit)
.tmp/
.agents/
+17 -6
View File
@@ -1,7 +1,7 @@
<!-- version: 1.4.0 -->
<!-- Last updated: 2026-04-16 -->
<!-- version: 1.6.0 -->
<!-- Last updated: 2026-04-20 -->
Last reviewed: 2026-04-16
Last reviewed: 2026-04-20
**Project:** GitNexus · **Environment:** dev · **Maintainer:** repository maintainers (see GitHub)
@@ -39,6 +39,8 @@ Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.
## Reference docs
- **[ARCHITECTURE.md](ARCHITECTURE.md)**, **[CONTRIBUTING.md](CONTRIBUTING.md)**, **[GUARDRAILS.md](GUARDRAILS.md)**
- **Call-resolution DAG (legacy path):** See ARCHITECTURE.md § Call-Resolution DAG. Typed 6-stage DAG inside the `parse` phase; language-specific behavior behind `inferImplicitReceiver` / `selectDispatch` hooks on `LanguageProvider`. Shared code in `gitnexus/src/core/ingestion/` must not name languages. Types: `gitnexus/src/core/ingestion/call-types.ts`.
- **Scope-resolution pipeline (RFC #909 Ring 3):** See ARCHITECTURE.md § Scope-Resolution Pipeline. Replaces the legacy DAG for languages in `MIGRATED_LANGUAGES` (currently Python). A language plugs in by implementing `ScopeResolver` (`scope-resolution/contract/scope-resolver.ts`) and registering it in `SCOPE_RESOLVERS`. CI parity gate runs BOTH paths per migrated language on every PR.
- **Cursor:** `.cursor/index.mdc` (always-on); `.cursor/rules/*.mdc` (glob-scoped). Legacy `.cursorrules` deprecated.
- **GitNexus:** skills in `.claude/skills/gitnexus/`; MCP rules in `gitnexus:start` block below.
@@ -46,6 +48,8 @@ Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.
| Date | Version | Change |
|------|---------|--------|
| 2026-04-20 | 1.6.0 | Added scope-resolution pipeline pointer (RFC #909 Ring 3); Python migrated to registry-primary. |
| 2026-04-19 | 1.5.0 | Cross-repo impact (#794): `impact`/`query`/`context` accept `repo: "@<group>"` + `service`. Removed `group_query`/`group_contracts`/`group_status` MCP tools; added `gitnexus://group/{name}/contracts` and `gitnexus://group/{name}/status` resources. |
| 2026-04-16 | 1.4.0 | Fixed: web UI description, pre-commit behavior, MCP tools (7->16), added gitnexus-shared, removed stale vite-plugin-wasm gotcha. |
| 2026-04-13 | 1.3.0 | Updated GitNexus index stats after DAG refactor. |
| 2026-03-24 | 1.2.0 | Fixed gitnexus:start block duplication. |
@@ -88,6 +92,7 @@ Indexed as **GitNexus** (4325 symbols, 10556 relationships, 300 execution flows)
- Ignore HIGH/CRITICAL risk warnings.
- Rename with find-and-replace — use `gitnexus_rename`.
- Commit without `gitnexus_detect_changes()`.
- Add language-specific behavior to shared ingestion code (`gitnexus/src/core/ingestion/`) — use a `LanguageProvider` hook. Seeing `provider.mroStrategy === 'xxx'` or an import from `languages/xxx.ts` in shared code means stop and add a hook.
## Tools Quick Reference
@@ -105,10 +110,14 @@ Indexed as **GitNexus** (4325 symbols, 10556 relationships, 300 execution flows)
| `tool_map` | MCP/RPC tool definitions | `gitnexus_tool_map({})` |
| `shape_check` | Response shape vs consumer access | `gitnexus_shape_check({route: "/api/users"})` |
| `group_list` | List repo groups | `gitnexus_group_list({})` |
| `group_query` | Cross-repo search in a group | `gitnexus_group_query({name: "myGroup", query: "auth"})` |
| `group_sync` | Rebuild group Contract Registry | `gitnexus_group_sync({name: "myGroup"})` |
| `group_contracts` | Inspect group contracts | `gitnexus_group_contracts({name: "myGroup"})` |
| `group_status` | Group staleness report | `gitnexus_group_status({name: "myGroup"})` |
| `query` (group mode) | Cross-repo search in a group (RRF-merged) | `gitnexus_query({repo: "@myGroup", query: "auth"})` |
| `context` (group mode) | 360° view across all member repos | `gitnexus_context({repo: "@myGroup", name: "validateUser"})` |
| `impact` (group mode) | Cross-repo blast radius via Contract Bridge | `gitnexus_impact({repo: "@myGroup", target: "X", direction: "upstream"})` |
> Group mode: pass `repo: "@<groupName>"` to fan out across all member repos, or `repo: "@<groupName>/<memberPath>"` to target a single member (path keys from `group.yaml`). Optional `service: "<monorepo/path>"` filters by service root. Group-level state (contracts, staleness) lives in the resources table below — there are **no** `group_query` / `group_context` / `group_impact` / `group_contracts` / `group_status` MCP tools.
>
> For a full walkthrough of setting up a group across multiple repos that communicate over gRPC, see [docs/guides/microservices-grpc.md](docs/guides/microservices-grpc.md).
## Impact Risk Levels
@@ -126,6 +135,8 @@ Indexed as **GitNexus** (4325 symbols, 10556 relationships, 300 execution flows)
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
| `gitnexus://group/{name}/contracts` | Group Contract Registry (provider/consumer rows + cross-links) |
| `gitnexus://group/{name}/status` | Per-member index + Contract Registry staleness report |
## Self-Check Before Finishing
+155 -4
View File
@@ -42,10 +42,14 @@ Monorepo: **CLI/MCP** (`gitnexus/`) + **browser UI** (`gitnexus-web/`).
| `tool_map` | MCP/RPC tool definitions and handlers |
| `shape_check` | Response shape vs consumer property access mismatches |
| `group_list` | List repo groups or details for one group |
| `group_query` | Cross-repo search in a group (reciprocal rank fusion) |
| `group_sync` | Rebuild group Contract Registry (`contracts.json`) |
| `group_contracts` | Inspect group contracts and cross-links |
| `group_status` | Index and Contract Registry staleness per repo in a group |
| `group_sync` | Rebuild group Contract Registry (`contracts.json`) and bridge graph |
`query`, `context`, and `impact` are group-aware: pass `repo: "@<groupName>"` (or `"@<groupName>/<memberPath>"` to scope to one member) plus optional `service: "<monorepo/path>"`. Group-mode `query` merges per-repo results via Reciprocal Rank Fusion; group-mode `impact` runs the local walk in the chosen member and fans out across boundaries via the Contract Bridge (`gitnexus/src/core/group/cross-impact.ts`). The previously-planned `group_query`, `group_context`, `group_impact`, `group_contracts`, `group_status` MCP tools are intentionally not introduced — group-level state is exposed via resources instead:
| Resource URI | Purpose |
|--------------|---------|
| `gitnexus://group/{name}/contracts` | Contract Registry (provider/consumer rows + cross-links) |
| `gitnexus://group/{name}/status` | Per-member index + Contract Registry staleness |
## Where to change what
@@ -55,6 +59,7 @@ Monorepo: **CLI/MCP** (`gitnexus/`) + **browser UI** (`gitnexus-web/`).
| Parsing/graph construction | `src/core/ingestion/pipeline-phases/` + `pipeline.ts` |
| Graph schema/DB | `src/core/lbug/` (`schema.ts`, `lbug-adapter.ts`) |
| MCP tools/resources | `src/mcp/server.ts`, `tools.ts`, `resources.ts` |
| Cross-repo groups (sync, contracts, `@<group>` routing) | `src/core/group/` (`service.ts`, `cross-impact.ts`, `sync.ts`, `bridge-db.ts`) |
| Search ranking | `src/core/search/` (BM25, hybrid fusion) |
| Embeddings | `src/core/embeddings/` + `src/core/run-analyze.ts` |
| Wiki generation | `src/core/wiki/` |
@@ -142,6 +147,152 @@ export const myPhase: PipelinePhase<MyPhaseOutput> = {
---
## Call-Resolution DAG
Typed 6-stage pipeline in `call-processor.ts` (inside the `parse` phase) that resolves method/function calls and emits CALLS edges. Language behavior plugs in at two `LanguageProvider` hook points (stages 3–4); shared code names no languages. Scope: call resolution only — import resolution, type extraction, heritage, and symbol-table population live in other phases.
### Stages
```
extract-call ──▶ classify-form ──▶ infer-receiver ──▶ select-dispatch ──▶ resolve-target ──▶ emit-edge
(1) (2) (3) [hook] (4) [hook] (5) (6)
```
| Stage | Produces | Location |
|-------|----------|----------|
| **extract-call** | `ExtractedCallSite` (name, form, receiver, argCount) | `call-extractors/` (per-language); runs in worker |
| **classify-form** | callForm (`free`/`member`/`constructor`) + arity | `call-analysis.ts` → `inferCallForm`; shared, runs in worker |
| **infer-receiver** | `ReceiverEnriched` (receiver type finalized) | `call-processor.ts`; shared default chain, then `inferImplicitReceiver` hook |
| **select-dispatch** | `DispatchDecision` (primary, fallback, ancestryView) | `selectDispatch` hook, falls back to shared default |
| **resolve-target** | `TieredCandidates` | `model/resolve.ts` → `lookupMethodByOwnerWithMRO` (MRO walk) |
| **emit-edge** | CALLS edge in graph | `call-processor.ts`; writes edge with confidence tier |
### Provider hooks
Both hooks are optional on `LanguageProvider`. Ruby is the only current implementer.
**`inferImplicitReceiver`** — called after shared infer-receiver defaults. Returns `ImplicitReceiverOverride | null`.
| | |
|---|---|
| Inputs | `calledName`, `callForm`, `receiverName`, `receiverTypeName`, `callNode` (AST), `filePath` |
| Non-null fields | `callForm`, `receiverName`, `receiverTypeName` (required); `receiverSource: 'implicit-self'` (fixed); `hint?` (opaque, passed to `selectDispatch`) |
| Null | Keep existing `ReceiverEnriched` state |
**`selectDispatch`** — called after infer-receiver (including hook). Returns `DispatchDecision | null`; null uses shared default (constructor → `primary:'constructor'`; typed receiver → `primary:'owner-scoped'`; else → `primary:'free'`).
| | |
|---|---|
| Inputs | `calledName`, `callForm`, `receiverName`, `receiverTypeName`, `receiverSource`, `hint` |
| Non-null fields | `primary: 'owner-scoped' \| 'free' \| 'constructor'`; `fallback?: 'free-arity-narrowed'`; `ancestryView?: 'instance' \| 'singleton'`; `hint?` |
**`DispatchDecision` field semantics:**
- `primary: 'owner-scoped'` — MRO walk from receiver's type; used when receiver type is known.
- `fallback: 'free-arity-narrowed'` — after owner-scoped miss, search free-call candidates by arity only (Ruby uses this for implicit-self calls that miss their owner's MRO).
- `ancestryView: 'singleton'` — walk singleton/class ancestry instead of instance ancestry (Ruby `def self.foo` bodies, so `extend`-ed methods are found).
### Adding language behavior
1. **Implicit receivers** — implement `inferImplicitReceiver`: return null if call already has a receiver; otherwise use `findEnclosingClassInfo` (`ast-helpers.ts`) to find the enclosing context, return `ImplicitReceiverOverride` with `receiverSource: 'implicit-self'`, and optionally set `hint` for `selectDispatch`.
2. **Custom dispatch** — implement `selectDispatch`: inspect `receiverSource` and `hint`, return `DispatchDecision` with `primary`, optional `fallback`, optional `ancestryView`; return null to keep shared defaults.
3. **MRO strategy** — confirm `mroStrategy` is `'first-wins'`, `'c3'`, `'ruby-mixin'`, or `'none'`; consumed by `lookupMethodByOwnerWithMRO`.
**Ruby example** (`languages/ruby.ts` + `utils/ruby-self-call.ts`): `inferImplicitReceiver` rewrites bare-identifier calls to `self.method` and sets `hint` to `'instance'`/`'singleton'`; `selectDispatch` uses hint for `ancestryView` and adds `fallback: 'free-arity-narrowed'` for implicit-self calls.
### Code references
| Module | Purpose |
|--------|---------|
| `core/ingestion/call-types.ts` | DAG types: `ReceiverEnriched`, `DispatchDecision`, `ImplicitReceiverOverride` |
| `core/ingestion/language-provider.ts` | Hook signatures: `inferImplicitReceiver`, `selectDispatch` |
| `core/ingestion/call-processor.ts` | `processCalls`: stages 3–6 |
| `core/ingestion/model/resolve.ts` | `lookupMethodByOwnerWithMRO`: stage 5 MRO walk |
| `core/ingestion/languages/ruby.ts` | Both hooks + `mroStrategy: 'ruby-mixin'` |
| `core/ingestion/utils/ruby-self-call.ts` | Bare-call rewrite for `inferImplicitReceiver` |
### Coexistence with the scope-resolution pipeline
The Call-Resolution DAG is the **legacy path**. RFC #909 Ring 3 introduces a parallel **scope-resolution pipeline** (next section) that replaces stages 1–6 with a scope-indexed registry lookup. Both paths ship side-by-side and are gated per-language via `MIGRATED_LANGUAGES` + the `REGISTRY_PRIMARY_<LANG>` env var.
- **Unmigrated language** → Call-Resolution DAG runs; scope-resolution phase is a no-op.
- **Migrated language** (currently: Python) → scope-resolution owns CALLS/ACCESSES/USES emission; the legacy DAG gates off for that language via `isRegistryPrimary(lang)` checks in `call-processor.ts` and `import-processor.ts`.
- `import-processor` still populates `importMap` for migrated languages — heritage's `ctx.resolve` reads it to disambiguate parent classes. Only edge emission is gated.
- CI runs BOTH paths for every migrated language on every PR (`.github/workflows/ci-scope-parity.yml`); both must pass.
---
## Scope-Resolution Pipeline (RFC #909 Ring 3)
Language-agnostic registry-primary resolver. Replaces the Call-Resolution DAG for migrated languages. Adding a language is one interface implementation (`ScopeResolver`) plus two registrations — no changes to shared code, no new pipeline phase.
### Pipeline stages
```
ParsedFile[] (extractParsedFile per file)
│ finalizeScopeModel (+ provider hooks)
▼
ScopeResolutionIndexes
│ resolveReferenceSites (via MethodRegistry.lookup)
▼
ReferenceIndex
│ emitReceiverBoundCalls ── FIRST
│ emitFreeCallFallback ── THEN
│ emitReferencesViaLookup ── LAST (uses handledSites)
│ emitImportEdges
▼
KnowledgeGraph (IMPORTS / CALLS / ACCESSES / INHERITS / USES)
```
Orchestrator: `runScopeResolution(input, provider)` in `scope-resolution/pipeline/run.ts`.
Pipeline phase: `scopeResolutionPhase` in `scope-resolution/pipeline/phase.ts` — iterates `SCOPE_RESOLVERS ∩ MIGRATED_LANGUAGES`, reads per-file Trees from the parse phase's `scopeTreeCache`, disposes the cache at the end.
### `ScopeResolver` contract
Single interface a language implements to plug into the pipeline. Contract fully documented in `scope-resolution/contract/scope-resolver.ts`.
| Hook | Purpose |
|------|---------|
| `languageProvider` | Base `LanguageProvider` (tree-sitter query, `emitScopeCaptures`, import/binding interpreters, hooks) |
| `populateOwners(parsed)` | Fill deferred `ownerId` fields on method defs (captures can't always know the owning class at parse time) |
| `buildMro(graph, parsed, nodeLookup)` | Produce `mroByClassDefId: Map<DefId, DefId[]>` — C3, Ruby-mixin, or first-wins per language |
| `resolveImportTarget(target, fromFile, allFiles)` | `(rawImportPath, sourceFile) → targetFilePath` (PEP-328 for Python, etc.) |
| `mergeBindings(existing, incoming, scopeId)` | Shadowing / LEGB precedence |
| `arityCompatibility` | Provider consumed by registry during `MethodRegistry.lookup` Step 2 |
| `importEdgeReason` | Confidence-tier string for IMPORTS edge reason field |
| `propagatesReturnTypesAcrossImports?` | Opt out of cross-file return-type propagation (default on) |
### Per-language registration
1. Implement `ScopeResolver` in `languages/<lang>/scope-resolver.ts`.
2. Add entry to `SCOPE_RESOLVERS` in `scope-resolution/pipeline/registry.ts`.
3. Add the language to `MIGRATED_LANGUAGES` in `registry-primary-flag.ts` when the shadow-harness corpus parity ≥ 99% fixtures / ≥ 98% corpus.
CI auto-discovers the set via `tsx`. No workflow edit required.
### Code references
| Module | Purpose |
|--------|---------|
| `scope-resolution/contract/scope-resolver.ts` | `ScopeResolver` interface + shared types |
| `scope-resolution/pipeline/run.ts` | Generic orchestrator |
| `scope-resolution/pipeline/phase.ts` | Pipeline-phase wrapper (deps: `parse`, `structure`) |
| `scope-resolution/pipeline/registry.ts` | `SCOPE_RESOLVERS` map |
| `scope-resolution/passes/*.ts` | Reference-resolution passes (receiver-bound, free-call fallback, compound-receiver, MRO, cross-file return-type propagation) |
| `scope-resolution/graph-bridge/*.ts` | CLI-local translation from resolved references → `KnowledgeGraph` edges |
| `scope-resolution/scope/*.ts` | Generic scope-chain walkers + namespace targets |
| `scope-resolution/workspace-index.ts` | Build-once O(1) lookup index |
| `registry-primary-flag.ts` | `MIGRATED_LANGUAGES` set + `isRegistryPrimary(lang)` |
| `languages/python/index.ts` | Python `ScopeResolver` hooks + known-limitation docs |
| `languages/python/captures.ts` | `emitPythonScopeCaptures` (honors cross-phase Tree cache) |
### Performance notes
- **Cross-phase Tree cache**: parse phase writes Trees into `scopeTreeCache` (separate from the chunk-local `astCache`) ONLY for languages with `emitScopeCaptures`. Scope-resolution reads from it to skip the second parse. Cleared at end of the phase. Workers leave the cache empty — Trees can't cross MessageChannels; cache miss = fresh parse. `PROF_SCOPE_RESOLUTION=1` emits hit/miss counters and a worker-engaged warning.
- **Typed relationship iteration**: heritage + MRO walk only the EXTENDS / IMPLEMENTS / HAS_METHOD edges via `iterRelationshipsByType`, not the full relationship map.
- **Workspace-resolution-index**: O(1) `findOwnedMember` / `findExportedDef` / `classScopeByDefId` built once per run.
---
## Language-agnostic graph feeding
16 languages → single unified graph. Four abstraction layers:
+2 -203
View File
@@ -35,6 +35,7 @@ If always-on instructions grow, load deep conventions via conditional reads (e.g
## Reference Documentation
- **This repository:** [AGENTS.md](AGENTS.md) (Cursor + monorepo notes), [ARCHITECTURE.md](ARCHITECTURE.md), [CONTRIBUTING.md](CONTRIBUTING.md), [GUARDRAILS.md](GUARDRAILS.md).
- **Call-resolution DAG:** See ARCHITECTURE.md § Call-Resolution DAG. Shared pipeline code in `gitnexus/src/core/ingestion/` must not name languages — use `LanguageProvider` hooks instead (see AGENTS.md).
- **GitNexus:** `.claude/skills/gitnexus/`; MCP and indexed-repo rules live only in [AGENTS.md](AGENTS.md) (`gitnexus:start` … `gitnexus:end`). See **GitNexus rules** below.
## Changelog
@@ -50,206 +51,4 @@ If always-on instructions grow, load deep conventions via conditional reads (e.g
## GitNexus rules
GitNexus MCP rules are in the `<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (4325 symbols, 10556 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find execution flows related to the issue
2. `gitnexus_context({name: "<suspect function>"})` — see all callers, callees, and process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace the full execution flow step by step
4. For regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})` — see what your branch changed
## When Refactoring
- **Renaming**: MUST use `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Review the preview — graph edits are safe, text_search edits need manual review. Then run with `dry_run: false`.
- **Extracting/Splitting**: MUST run `gitnexus_context({name: "target"})` to see all incoming/outgoing refs, then `gitnexus_impact({target: "target", direction: "upstream"})` to find all external callers before moving code.
- After any refactor: run `gitnexus_detect_changes({scope: "all"})` to verify only expected files changed.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Tools Quick Reference
| Tool | When to use | Command |
|------|-------------|---------|
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update these |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## Self-Check Before Finishing
Before completing any code modification task, verify:
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL risk warnings were ignored
3. `gitnexus_detect_changes()` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## Keeping the Index Fresh
After committing code changes, the GitNexus index becomes stale. Re-run analyze to update it:
```bash
npx gitnexus analyze
```
If the index previously included embeddings, preserve them by adding `--embeddings`:
```bash
npx gitnexus analyze --embeddings
```
To check whether embeddings exist, inspect `.gitnexus/meta.json` — the `stats.embeddings` field shows the count (0 means no embeddings). **Running analyze without `--embeddings` will delete any previously generated embeddings.**
> Claude Code users: A PostToolUse hook handles this automatically after `git commit` and `git merge`.
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
<!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** — load that section when working with MCP tools or the graph index.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (3298 symbols, 7954 relationships, 185 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find execution flows related to the issue
2. `gitnexus_context({name: "<suspect function>"})` — see all callers, callees, and process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace the full execution flow step by step
4. For regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})` — see what your branch changed
## When Refactoring
- **Renaming**: MUST use `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Review the preview — graph edits are safe, text_search edits need manual review. Then run with `dry_run: false`.
- **Extracting/Splitting**: MUST run `gitnexus_context({name: "target"})` to see all incoming/outgoing refs, then `gitnexus_impact({target: "target", direction: "upstream"})` to find all external callers before moving code.
- After any refactor: run `gitnexus_detect_changes({scope: "all"})` to verify only expected files changed.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Tools Quick Reference
| Tool | When to use | Command |
|------|-------------|---------|
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update these |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## Self-Check Before Finishing
Before completing any code modification task, verify:
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL risk warnings were ignored
3. `gitnexus_detect_changes()` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## Keeping the Index Fresh
After committing code changes, the GitNexus index becomes stale. Re-run analyze to update it:
```bash
npx gitnexus analyze
```
If the index previously included embeddings, preserve them by adding `--embeddings`:
```bash
npx gitnexus analyze --embeddings
```
To check whether embeddings exist, inspect `.gitnexus/meta.json` — the `stats.embeddings` field shows the count (0 means no embeddings). **Running analyze without `--embeddings` will delete any previously generated embeddings.**
> Claude Code users: A PostToolUse hook handles this automatically after `git commit` and `git merge`.
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
<!-- gitnexus:end -->
See the `<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** for the canonical MCP tools, impact analysis rules, and index instructions.
+27 -1
View File
@@ -77,7 +77,7 @@ Every workflow under `.github/workflows/` MUST declare a top-level `concurrency:
- Per-PR scope (for `issue_comment`, `pull_request_review*`, `pull_request` meta events): `${{ github.workflow }}-${{ github.event.pull_request.number || github.event.issue.number }}`
- `workflow_run` scope (e.g. `ci-report.yml`): `${{ github.workflow }}-${{ github.event.workflow_run.pull_requests[0].number || format('{0}/{1}', github.event.workflow_run.head_repository.full_name, github.event.workflow_run.head_branch) }}` — the fork fallback must be stable across reruns (never `workflow_run.id`, which is per-run-unique and defeats serialization).
- Global single-slot (manual dispatch utilities): `${{ github.workflow }}`
- **Reusable workflows invoked via `workflow_call`:** do NOT use `${{ github.workflow }}` in the group key — in called-workflow context its evaluation is ambiguous and can resolve to the caller's name, which would deadlock against the caller's own group. Use a hardcoded literal prefix and a `github.event_name`-aware expression that falls through to `github.run_id` for reusable invocations (see `ci.yml` for the canonical form).
- **Reusable workflows invoked via `workflow_call`:** do NOT use `${{ github.workflow }}` in the group key — in called-workflow context its evaluation is ambiguous and can resolve to the caller's name, which would deadlock against the caller's own group. Use a hardcoded literal prefix and a `github.event_name`-aware expression that falls through to `github.run_id` for reusable invocations (see `ci.yml` for the canonical form). Approved literal prefixes: `CI-` (`ci.yml`) and `docker-build-push-` (`docker.yml`). The `check-workflow-concurrency.py` validation script must be updated whenever a new approved literal prefix is added.
- **Merge queue (`merge_group`)**: when this event is added, use `${{ github.workflow }}-${{ github.event.merge_group.head_ref }}` with `cancel-in-progress: false` (every queue entry is a distinct ref; never cancel).
- **`cancel-in-progress` policy:**
@@ -127,6 +127,11 @@ Two publish workflows ship `gitnexus` to npm:
the cycle from `latest`.
- `N` is auto-incremented against existing `X.Y.Z-rc.*` entries on the
registry. First rc for a given base is `rc.1`.
- After the npm publish succeeds, the workflow calls `docker.yml` as a
reusable workflow to build and push the corresponding RC Docker images
(e.g. `ghcr.io/abhigyanpatwari/gitnexus:1.7.0-rc.1`). The images are
signed with Cosign; the OIDC identity is `docker.yml@refs/heads/main`
(the caller's ref — see README.md § Docker for the verify command).
Idempotency: the workflow pushes an `rc/<HEAD_SHA>` marker tag and a
`v<RC>` release tag **atomically, before** calling `npm publish`. The guard
@@ -140,6 +145,27 @@ Two publish workflows ship `gitnexus` to npm:
# then redispatch the workflow with force: true
```
**Docker-only partial failure:** if `publish` succeeds (npm tarball + tags
are live) but the `docker` job subsequently fails (e.g. GHCR flakiness),
the npm RC is already published and the `rc/<HEAD_SHA>` marker is in place.
Re-running `release-candidate.yml` with `force: true` will abort at the
"Version already exists on npm" guard. To recover without cutting a new RC:
```bash
# 1. Manually trigger only the docker workflow, passing the existing RC tag:
gh workflow run docker.yml --ref main -f tag=v<RC_VERSION>
# (requires a workflow_dispatch trigger on docker.yml — see note below)
```
Because `docker.yml` intentionally has no `workflow_dispatch` (images are
tag-driven by design), the practical recovery options are:
- Wait for the next commit on `main`, which will cut a new RC that includes
the Docker build.
- Manually run `docker build` + `docker push` locally and sign with Cosign
against the same digest.
- Delete `rc/<HEAD_SHA>` and `v<RC>` tags, then redispatch with `force:
true` to re-run the full RC pipeline (cuts a new RC number).
The rc workflow never moves `latest`. To verify after a change, inspect dist-tags:
```bash
+58
View File
@@ -0,0 +1,58 @@
ARG BUILDPLATFORM
ARG TARGETPLATFORM
# ── Builder ────────────────────────────────────────────────────────────
# Native modules (tree-sitter-*, onnxruntime-node, node-gyp builds for
# tree-sitter-proto / tree-sitter-swift) require python3 + a C/C++ toolchain.
FROM node:22-trixie-slim AS builder
WORKDIR /app
# Toolchain for node-gyp / native builds.
RUN apt-get update && apt-get install -y --no-install-recommends python3 make g++ git && rm -rf /var/lib/apt/lists/*
# Build gitnexus-shared first — gitnexus depends on it as a workspace.
COPY gitnexus-shared/package.json gitnexus-shared/package-lock.json ./gitnexus-shared/
RUN npm ci --prefix gitnexus-shared
COPY gitnexus-shared ./gitnexus-shared
RUN rm -f gitnexus-shared/tsconfig.tsbuildinfo
RUN npm run build --prefix gitnexus-shared
# Copy the full gitnexus package before installing — `npm ci` triggers
# `postinstall` (patches tree-sitter-swift, builds the vendored
# tree-sitter-proto) and `prepare` (compiles TypeScript via scripts/build.js),
# both of which need the source tree.
COPY gitnexus ./gitnexus
RUN npm ci --prefix gitnexus
# Drop dev dependencies for a smaller runtime layer.
RUN npm prune --omit=dev --prefix gitnexus
# ── Runtime ────────────────────────────────────────────────────────────
FROM node:22-trixie-slim AS runtime
# curl for the healthcheck; git so `gitnexus` can clone repos at runtime.
RUN apt-get update && apt-get install -y --no-install-recommends curl git && rm -rf /var/lib/apt/lists/*
WORKDIR /app
# Pre-create the data directory and hand it to the unprivileged `node` user
# so the bind-mounted volume is writable without root.
RUN mkdir -p /data/gitnexus && chown -R node:node /data
COPY --from=builder --chown=node:node /app/gitnexus/dist ./gitnexus/dist
COPY --from=builder --chown=node:node /app/gitnexus/node_modules ./gitnexus/node_modules
COPY --from=builder --chown=node:node /app/gitnexus/package.json ./gitnexus/package.json
COPY --from=builder --chown=node:node /app/gitnexus/vendor ./gitnexus/vendor
USER node
# The web UI defaults to http://localhost:4747 — keep that contract.
ENV GITNEXUS_HOME=/data/gitnexus \
NODE_ENV=production \
PORT=4747
EXPOSE 4747
# Bind to 0.0.0.0 so the server is reachable from the host's mapped port.
CMD ["node", "gitnexus/dist/cli/index.js", "serve", "--host", "0.0.0.0", "--port", "4747"]
+37
View File
@@ -0,0 +1,37 @@
ARG BUILDPLATFORM
ARG TARGETPLATFORM
FROM --platform=$BUILDPLATFORM node:22-alpine AS builder
WORKDIR /app
COPY gitnexus-shared/package.json gitnexus-shared/package-lock.json ./gitnexus-shared/
RUN npm ci --prefix gitnexus-shared
COPY gitnexus-shared ./gitnexus-shared
RUN npm run build --prefix gitnexus-shared
COPY gitnexus/package.json ./gitnexus/
COPY gitnexus-web/package.json gitnexus-web/package-lock.json ./gitnexus-web/
RUN npm ci --prefix gitnexus-web
COPY gitnexus-web ./gitnexus-web
RUN npm run build --prefix gitnexus-web
FROM node:22-alpine AS runtime
RUN apk add --no-cache curl
WORKDIR /app
COPY --from=builder /app/gitnexus-web/dist ./dist
COPY docker-server.mjs ./docker-server.mjs
RUN chown -R node:node /app
USER node
EXPOSE 4173
CMD ["node", "docker-server.mjs"]
+44
View File
@@ -1,5 +1,49 @@
# Migration Guide
## `impact` tool may now return `{ status: 'ambiguous' }` (PR #888, issue #470)
Before this change the `impact` MCP tool silently picked the first match
when the `target` name hit multiple symbols (Class → Interface → Function
→ Method → Constructor priority UNION). This often produced analysis for
the wrong symbol with no signal back to the caller.
After this change, when the resolver finds more than one viable match
and the caller supplied none of `target_uid` / `file_path` / `kind`,
`impact` returns a disambiguation response shaped like:
```json
{
"status": "ambiguous",
"message": "Found N symbols matching '<target>'. Use target_uid, file_path, or kind to disambiguate.",
"target": { "name": "<target>" },
"direction": "upstream",
"impactedCount": 0,
"risk": "UNKNOWN",
"candidates": [
{ "uid": "...", "name": "...", "kind": "Function", "filePath": "...", "line": 42, "score": 0.76 }
]
}
```
### Do I need to migrate?
**Probably not, but check for assumptions.** Callers that unconditionally
read `result.byDepth` / `result.summary` / `result.affected_processes`
without first checking `result.status` will now see `undefined` in the
ambiguous case. The fix is to branch on `result.status === 'ambiguous'`
first and follow up with `target_uid` (preferred) or `file_path` / `kind`.
The `context` tool's ambiguous response is a strict superset of the
existing shape — every candidate gains a `score` field, no existing field
has changed. No migration required for `context` callers.
### What happens on re-index?
Nothing — this is an MCP-surface change only. The graph schema, indexer,
and stored data are untouched.
---
## OVERRIDES → METHOD_OVERRIDES (PR #642)
The `OVERRIDES` relationship type has been renamed to `METHOD_OVERRIDES` for
+156 -6
View File
@@ -9,7 +9,7 @@
<h2>Join the official Discord to discuss ideas, issues etc!</h2>
<a href="https://discord.gg/AAsRVT6fGb">
<a href="https://discord.gg/MgJrmsqr62">
<img src="https://img.shields.io/discord/1477255801545429032?color=5865F2&logo=discord&logoColor=white" alt="Discord"/>
</a>
<a href="https://www.npmjs.com/package/gitnexus">
@@ -194,6 +194,7 @@ gitnexus analyze --force # Force full re-index
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
gitnexus analyze --skip-git # Index folders that are not Git repositories
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
@@ -207,11 +208,11 @@ gitnexus wiki --model <model> # Wiki with custom LLM model (default: gpt-4o-m
gitnexus wiki --base-url <url> # Wiki with custom LLM API base URL
# Repository groups (multi-repo / monorepo service tracking)
gitnexus group create <name> # Create a repository group
gitnexus group add <name> <repo> # Add a repo to a group
gitnexus group remove <name> <repo> # Remove a repo from a group
gitnexus group list [name] # List groups, or show one group's config
gitnexus group sync <name> # Extract contracts and match across repos/services
gitnexus group create <name> # Create a repository group
gitnexus group add <group> <groupPath> <registryName> # Add a repo to a group. <groupPath> is a hierarchy path (e.g. hr/hiring/backend); <registryName> is the repo's name from the registry (see `gitnexus list`)
gitnexus group remove <group> <groupPath> # Remove a repo from a group by its hierarchy path
gitnexus group list [name] # List groups, or show one group's config
gitnexus group sync <name> # Extract contracts and match across repos/services
gitnexus group contracts <name> # Inspect extracted contracts and cross-links
gitnexus group query <name> <q> # Search execution flows across all repos in a group
gitnexus group status <name> # Check staleness of repos in a group
@@ -335,6 +336,155 @@ cd ../gitnexus-web && npm install
npm run dev
```
## Docker
The official Docker setup ships **two signed images** orchestrated by `docker-compose.yaml`:
| Image | Purpose |
| -------------------------------------------------- | ---------------------------------------------------------------------- |
| `ghcr.io/abhigyanpatwari/gitnexus:latest` | CLI / `gitnexus serve` backend (HTTP API on port `4747`, MCP, indexer) |
| `ghcr.io/abhigyanpatwari/gitnexus-web:latest` | Static web UI (port `4173`) |
> **Heads-up — image rename.** Earlier releases published the web UI under
> `ghcr.io/abhigyanpatwari/gitnexus`. Starting with the introduction of the
> bundled backend, that slug now hosts the CLI/server image and the UI moved
> to `ghcr.io/abhigyanpatwari/gitnexus-web`. The previous tags remain
> available for pulling, but new versions are only published under the new
> slugs. Update your `docker run` / compose files accordingly (or just adopt
> the bundled compose).
### One-command setup
```bash
docker compose up -d
```
This starts the server on `http://localhost:4747` and the web UI on
`http://localhost:4173`. The UI auto-detects the server because the browser
runs on the host and reaches the container via the mapped port.
A named volume (`gitnexus-data`) persists the global registry, indexes, and
cloned repos at `/data/gitnexus` inside the server container. To make repos on
your host machine indexable, set `WORKSPACE_DIR` before bringing the stack up:
```bash
WORKSPACE_DIR=$HOME/code docker compose up -d
# Inside the server container the directory is mounted read-only at /workspace.
docker compose exec gitnexus-server gitnexus index /workspace/my-repo
```
### Direct `docker run`
```bash
# Server
docker run --rm -d \
--name gitnexus-server \
-p 4747:4747 \
-v gitnexus-data:/data/gitnexus \
ghcr.io/abhigyanpatwari/gitnexus:latest
# Web UI
docker run --rm -d \
--name gitnexus-web \
-p 4173:4173 \
ghcr.io/abhigyanpatwari/gitnexus-web:latest
```
Optional env file (override image tags, container names, ports, workspace dir):
```bash
cp .env.example .env
docker compose --env-file .env up -d
```
### Versioning & supply-chain protection
The Docker images are version-locked to the npm package:
- Stable images are **only published from `vX.Y.Z` git tags** (via `docker.yml`
triggered directly by the tag push), and the workflow refuses to build unless
the tag exactly matches `gitnexus/package.json`'s version. So
`ghcr.io/abhigyanpatwari/gitnexus:1.6.2` is byte-for-byte the same release
as `npm install gitnexus@1.6.2` — no drift, no floating builds from `main`.
- Release-candidate images (e.g. `:1.7.0-rc.1`) are published alongside each
RC npm release. They are built by `release-candidate.yml` calling `docker.yml`
as a reusable workflow after the RC tag is created and pushed.
- `:latest` is auto-promoted only from non-prerelease tags by the Docker
metadata action, so it always points at a real, npm-published version.
Both images are signed with [Cosign keyless signing][cosign-keyless] using the
workflow's GitHub OIDC identity, and shipped with build provenance and SBOM
attestations. **This is your protection against supply-chain attacks**: even if
an attacker republishes a same-named image elsewhere (or somehow pushes to a
typo-squatted registry), they cannot forge a Cosign signature tied to
`abhigyanpatwari/GitNexus`'s `docker.yml`. Always verify before pulling into
sensitive environments:
**Stable releases** — signed from the `v*` tag ref:
```bash
cosign verify ghcr.io/abhigyanpatwari/gitnexus:1.6.2 \
--certificate-identity-regexp '^https://github\.com/abhigyanpatwari/GitNexus/\.github/workflows/docker\.yml@refs/tags/v[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com
```
The regex pins the certificate identity to this repo's `docker.yml` workflow
**run from a `v*` tag** — rejecting unsigned images, images signed by other
workflows, and images signed from unprotected refs.
**Release candidates** — signed from `refs/heads/main` (the caller's ref when
`release-candidate.yml` invokes `docker.yml` as a reusable workflow):
```bash
cosign verify ghcr.io/abhigyanpatwari/gitnexus:1.7.0-rc.1 \
--certificate-identity 'https://github.com/abhigyanpatwari/GitNexus/.github/workflows/docker.yml@refs/heads/main' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com
```
You can also inspect the build provenance and SBOM:
```bash
cosign download attestation ghcr.io/abhigyanpatwari/gitnexus:1.6.2 \
--predicate-type https://slsa.dev/provenance/v1
```
#### Kubernetes: enforce signatures at admission
For Kubernetes deployments, ship the bundled
[`ClusterImagePolicy`](deploy/kubernetes/cluster-image-policy.yaml) so the
[Sigstore policy-controller][policy-controller] rejects any GitNexus pod whose
image is not signed by this repo's `docker.yml` running from a `vX.Y.Z` tag —
the same identity the `cosign verify` snippet above pins.
```bash
# 1. Install the controller (one-time, cluster-wide)
helm repo add sigstore https://sigstore.github.io/helm-charts && helm repo update
helm install policy-controller -n cosign-system --create-namespace \
sigstore/policy-controller
# 2. Opt your namespace in
kubectl label namespace <your-ns> policy.sigstore.dev/include=true
# 3. Apply the policy
kubectl apply -f deploy/kubernetes/cluster-image-policy.yaml
```
After this, attempting to deploy an unsigned image — or one signed by anything
other than `abhigyanpatwari/GitNexus`'s `docker.yml` at a `v*` tag — fails the
admission webhook before a pod is ever created. This turns the verifiable
signature into an enforced policy, which is the supply-chain control most
clusters actually need.
[cosign-keyless]: https://docs.sigstore.dev/cosign/signing/overview/
[policy-controller]: https://docs.sigstore.dev/policy-controller/overview/
### Files
- [Dockerfile.web](Dockerfile.web) — builds `gitnexus-shared` and `gitnexus-web`, then serves the production frontend.
- [Dockerfile.cli](Dockerfile.cli) — builds the CLI/server (with its native deps) and runs `gitnexus serve --host 0.0.0.0`.
- [docker-compose.yaml](docker-compose.yaml) — starts both signed images side by side.
- [.env.example](.env.example) — overrides for image names, container names, ports, and the workspace mount.
The web UI uses the same indexing pipeline as the CLI but runs entirely in WebAssembly (Tree-sitter WASM, LadybugDB WASM, in-browser embeddings). It's great for quick exploration but limited by browser memory for larger repos.
**Local Backend Mode:** Run `gitnexus serve` and open the web UI locally — it auto-detects the server and shows all your indexed repos, with full AI chat support. No need to re-upload or re-index. The agent's tools (Cypher queries, search, code navigation) route through the backend HTTP API automatically.
@@ -0,0 +1,64 @@
# Sigstore policy-controller ClusterImagePolicy for GitNexus container images.
#
# This enforces — at admission time — that every Pod pulling a
# `ghcr.io/abhigyanpatwari/gitnexus` or `gitnexus-web` image is using a build
# that was Cosign-keyless-signed by this repository's `docker.yml` workflow
# running from a `vX.Y.Z` git tag. Unsigned images, images signed by other
# workflows, and images signed from unprotected refs (e.g. `main`, PR branches)
# are rejected.
#
# Prerequisites
# -------------
# 1. Install the Sigstore policy-controller in your cluster (Helm):
#
# helm repo add sigstore https://sigstore.github.io/helm-charts
# helm repo update
# helm install policy-controller -n cosign-system --create-namespace \
# sigstore/policy-controller
#
# 2. Opt namespaces in to verification:
#
# kubectl label namespace <your-ns> policy.sigstore.dev/include=true
#
# 3. Apply this policy:
#
# kubectl apply -f deploy/kubernetes/cluster-image-policy.yaml
#
# After this, `kubectl run --image=ghcr.io/abhigyanpatwari/gitnexus:<tag>` in
# any opted-in namespace will only succeed if the image carries a valid
# Sigstore signature with the pinned identity.
#
# References
# - https://docs.sigstore.dev/policy-controller/overview/
# - https://github.com/sigstore/policy-controller
apiVersion: policy.sigstore.dev/v1beta1
kind: ClusterImagePolicy
metadata:
name: gitnexus-signed-images
spec:
# Apply to both published GitNexus images on GHCR. Image references always
# carry a tag or digest at admission time, so these two globs cover every
# `gitnexus:<tag>`, `gitnexus@sha256:...`, `gitnexus-web:<tag>`, and
# `gitnexus-web@sha256:...` reference.
images:
- glob: 'ghcr.io/abhigyanpatwari/gitnexus*'
authorities:
- name: gitnexus-cosign-keyless
keyless:
# Public-good Sigstore Fulcio root.
url: https://fulcio.sigstore.dev
identities:
# Pin both the OIDC issuer (GitHub Actions) AND the exact workflow
# path running from a `vX.Y.Z` (or `vX.Y.Z-prerelease`) tag. Same
# regex the README's `cosign verify` example uses; it rejects:
# * unsigned images
# * signatures from any other repo / workflow
# * signatures from non-tag refs (main, PRs, release branches)
# * signatures from arbitrary non-semver tags
- issuer: https://token.actions.githubusercontent.com
subjectRegExp: ^https://github\.com/abhigyanpatwari/GitNexus/\.github/workflows/docker\.yml@refs/tags/v[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$
# Cross-check the signature against the public Rekor transparency log,
# so an attacker who briefly compromised Fulcio cannot retroactively
# mint a signature without leaving a public, append-only audit record.
ctlog:
url: https://rekor.sigstore.dev
+45
View File
@@ -0,0 +1,45 @@
services:
gitnexus-server:
image: ${SERVER_IMAGE:-ghcr.io/abhigyanpatwari/gitnexus:latest}
container_name: ${SERVER_CONTAINER_NAME:-gitnexus-server}
# Map the server to the same host port the web UI expects by default
# (http://localhost:4747). The browser runs on the host, so the UI's
# built-in default works without any reconfiguration.
ports:
- '${SERVER_HOST_PORT:-4747}:4747'
volumes:
# Persist the global registry, indexes, and cloned repos across runs.
- gitnexus-data:/data/gitnexus
# Optional: mount a host workspace so `gitnexus index <path>` can see
# repos you already have on disk. The default points at an empty
# `./workspace/` sibling that compose will create on first start —
# it intentionally does NOT bind-mount the repo root, which would
# expose `.git`, `.env`, and CI secrets to the container.
# Override with `WORKSPACE_DIR=/abs/path/to/your/repos`.
- ${WORKSPACE_DIR:-./workspace}:/workspace:ro
restart: unless-stopped
healthcheck:
test: ['CMD', 'curl', '-fsS', 'http://localhost:4747/api/heartbeat']
interval: 30s
timeout: 5s
retries: 3
start_period: 15s
gitnexus-web:
image: ${WEB_IMAGE:-ghcr.io/abhigyanpatwari/gitnexus-web:latest}
container_name: ${WEB_CONTAINER_NAME:-gitnexus-web}
ports:
- '${WEB_HOST_PORT:-4173}:4173'
depends_on:
gitnexus-server:
condition: service_healthy
restart: unless-stopped
healthcheck:
test: ['CMD', 'curl', '-f', 'http://localhost:4173/']
interval: 30s
timeout: 5s
retries: 3
start_period: 10s
volumes:
gitnexus-data:
+81
View File
@@ -0,0 +1,81 @@
import { createReadStream } from 'node:fs';
import { stat } from 'node:fs/promises';
import { createServer } from 'node:http';
import { extname, join, normalize, sep } from 'node:path';
const host = '0.0.0.0';
const port = Number(process.env.PORT || '4173');
const root = join(process.cwd(), 'dist');
const contentTypes = {
'.css': 'text/css; charset=utf-8',
'.html': 'text/html; charset=utf-8',
'.js': 'text/javascript; charset=utf-8',
'.json': 'application/json; charset=utf-8',
'.map': 'application/json; charset=utf-8',
'.png': 'image/png',
'.svg': 'image/svg+xml',
'.txt': 'text/plain; charset=utf-8',
'.woff': 'font/woff',
'.woff2': 'font/woff2',
};
function resolvePath(urlPath) {
let decoded;
try {
decoded = decodeURIComponent(urlPath);
} catch {
return null;
}
if (decoded.includes('\0')) return null;
const cleanPath = normalize(decoded.replace(/^\/+/, ''));
const candidate = join(root, cleanPath);
if (candidate !== root && !candidate.startsWith(root + sep)) return null;
return candidate;
}
const server = createServer(async (req, res) => {
const requestPath = req.url?.split('?')[0] || '/';
let filePath = resolvePath(requestPath);
if (!filePath) {
res.writeHead(400);
res.end('Bad request');
return;
}
try {
const fileStat = await stat(filePath).catch(() => null);
if (fileStat?.isDirectory()) {
filePath = join(filePath, 'index.html');
} else if (!fileStat?.isFile()) {
filePath = join(root, 'index.html');
}
const finalStat = await stat(filePath).catch(() => null);
if (!finalStat?.isFile()) {
res.writeHead(404);
res.end('Not found');
return;
}
res.writeHead(200, {
'Cache-Control': filePath.includes('/assets/')
? 'public, max-age=31536000, immutable'
: 'no-cache',
'Content-Type': contentTypes[extname(filePath)] || 'application/octet-stream',
'Cross-Origin-Opener-Policy': 'same-origin',
'Cross-Origin-Embedder-Policy': 'require-corp',
});
const stream = createReadStream(filePath);
stream.on('error', () => res.destroy());
stream.pipe(res);
} catch (error) {
res.writeHead(500);
res.end(error instanceof Error ? error.message : 'Internal server error');
}
});
server.listen(port, host, () => {
console.log(`gitnexus-web listening on http://${host}:${port}`);
});
+107
View File
@@ -0,0 +1,107 @@
import { mkdir, mkdtemp, rm, unlink, writeFile } from 'node:fs/promises';
import http, { createServer } from 'node:http';
import { tmpdir } from 'node:os';
import { dirname, join } from 'node:path';
import { spawn } from 'node:child_process';
import { fileURLToPath } from 'node:url';
import { after, before, it } from 'node:test';
import assert from 'node:assert/strict';
const __dirname = dirname(fileURLToPath(import.meta.url));
const serverScript = join(__dirname, 'docker-server.mjs');
function getFreePort() {
return new Promise((resolve) => {
const s = createServer();
s.listen(0, '127.0.0.1', () => {
const { port } = s.address();
s.close(() => resolve(port));
});
});
}
function rawGet(port, path) {
return new Promise((resolve, reject) => {
const req = http.request({ host: '127.0.0.1', port, path }, (res) => {
let body = '';
res.setEncoding('utf8');
res.on('data', (chunk) => {
body += chunk;
});
res.on('end', () => resolve({ status: res.statusCode, headers: res.headers, body }));
});
req.on('error', reject);
req.end();
});
}
async function waitForServer(port, retries = 30) {
for (let i = 0; i < retries; i++) {
try {
await rawGet(port, '/');
return;
} catch {
await new Promise((r) => setTimeout(r, 100));
}
}
throw new Error('Server did not start in time');
}
let tmpDir, serverPort, child;
before(async () => {
tmpDir = await mkdtemp(join(tmpdir(), 'gitnexus-docker-test-'));
const distDir = join(tmpDir, 'dist');
const assetsDir = join(distDir, 'assets');
await mkdir(assetsDir, { recursive: true });
await writeFile(join(distDir, 'index.html'), '<html><body>spa</body></html>');
await writeFile(join(assetsDir, 'app.abc123.js'), 'console.log("app")');
serverPort = await getFreePort();
child = spawn(process.execPath, [serverScript], {
cwd: tmpDir,
env: { ...process.env, PORT: String(serverPort) },
stdio: 'pipe',
});
child.on('error', (err) => {
throw err;
});
await waitForServer(serverPort);
});
after(async () => {
child?.kill();
if (tmpDir) await rm(tmpDir, { recursive: true, force: true });
});
it('serves a valid asset with immutable cache header', async () => {
const res = await rawGet(serverPort, '/assets/app.abc123.js');
assert.equal(res.status, 200);
assert.match(res.headers['cache-control'], /immutable/);
assert.equal(res.headers['cross-origin-opener-policy'], 'same-origin');
assert.equal(res.headers['cross-origin-embedder-policy'], 'require-corp');
});
it('serves SPA fallback for unknown routes', async () => {
const res = await rawGet(serverPort, '/some/unknown/route');
assert.equal(res.status, 200);
assert.match(res.body, /spa/);
assert.match(res.headers['cache-control'], /no-cache/);
});
it('rejects path traversal with 400', async () => {
const res = await rawGet(serverPort, '/../../../etc/passwd');
assert.equal(res.status, 400);
});
it('rejects percent-encoded null bytes with 400', async () => {
const res = await rawGet(serverPort, '/foo%00bar');
assert.equal(res.status, 400);
});
it('returns 404 when dist/index.html is missing', async () => {
await unlink(join(tmpDir, 'dist', 'index.html'));
const res = await rawGet(serverPort, '/nonexistent-page');
assert.equal(res.status, 404);
});
+295
View File
@@ -0,0 +1,295 @@
# Using GitNexus across gRPC microservices
## When to use this guide
This guide is for teams whose product lives in **several separate Git repositories** — one per service — and whose services talk to each other over **gRPC** (possibly alongside HTTP and message topics). GitNexus indexes each repo independently, then a _group_ stitches the per-repo indexes into a single cross-repo view that the `impact`, `query`, and `context` tools can traverse. If your services live in one monorepo, much of this still applies — set each service as a member of a group and use the `service` prefix to scope queries — but the walkthrough assumes the harder multi-repo case.
## Mental model
- Each repository has its own `.gitnexus/` index (a LadybugDB graph of symbols, relationships, processes). `gitnexus analyze` in each repo produces that index completely independently.
- A **group** is a higher-level construct stored at `~/.gitnexus/groups/<group>/` that references the per-repo indexes by their registry name.
- Sync-time extractors walk each member repo and emit **contracts** — provider or consumer records keyed by a canonical `contractId` (`grpc::auth.AuthService/Login`, `http::GET::/orders`, etc.).
- The sync step matches providers and consumers that share a `contractId` and writes **cross-links** to `<groupDir>/contracts.json`. Those cross-links are what lets `impact({repo: "@<group>", target: "X"})` hop from one repo into another.
- Contracts come from three places: automatic contract extractors (`grpc-extractor`, `http-route-extractor`, `topic-extractor`), a manifest escape hatch (`config.links` in `group.yaml`), and — for same-name symbol matches where no contract is declared — the exact-match matching cascade in [`matching.ts`](../../gitnexus/src/core/group/matching.ts).
- Each repo stays editable and re-indexable on its own. Re-run `gitnexus analyze` in a repo when it changes, then `gitnexus group sync <group>` to refresh `contracts.json`. `gitnexus group status` reports which members are stale.
## Prerequisites
- GitNexus installed and runnable as `gitnexus` or `npx gitnexus` (see the root [README.md](../../README.md)).
- Each service repository checked out locally. No requirement that they share a parent directory — the group references them by registry name.
- Write access to `~/.gitnexus/` (the default gitnexus home; see `getDefaultGitnexusDir` in [`storage.ts`](../../gitnexus/src/core/group/storage.ts)).
## Step-by-step walkthrough
The example uses three services — a TypeScript API gateway, a Go orders service, and a Python inventory service — with gRPC between them. The gateway is an `orders` consumer; the orders service is both an `orders` provider and an `inventory` consumer; the inventory service is an `inventory` provider.
### 1. Index each repository
Run `analyze` from inside each service repo (or pass the path). The CLI surface lives in [`gitnexus/src/cli/analyze.ts`](../../gitnexus/src/cli/analyze.ts) and is wired in [`gitnexus/src/cli/index.ts`](../../gitnexus/src/cli/index.ts).
```bash
cd ~/code/gateway && npx gitnexus analyze
cd ~/code/orders && npx gitnexus analyze
cd ~/code/inventory && npx gitnexus analyze
```
Useful flags:
- `--force` — reindex even if up to date.
- `--embeddings` — generate embedding vectors (needed only if you want semantic search; the exact-match cross-repo cascade does **not** need them).
- `--name <alias>` — register the repo under a specific alias when two repos share a basename (e.g. two `api/` folders).
- `--skip-git` — index a checkout that isn't a git repo.
Each run writes a `.gitnexus/` folder in the repo and registers the repo in `~/.gitnexus/registry.json`. Confirm with `npx gitnexus list`.
### 2. Author `group.yaml`
Create the group directory and edit the config. Either use the CLI scaffolder or write the file directly — both produce the same shape consumed by [`config-parser.ts`](../../gitnexus/src/core/group/config-parser.ts).
```bash
npx gitnexus group create payments-platform
# or manually:
mkdir -p ~/.gitnexus/groups/payments-platform
$EDITOR ~/.gitnexus/groups/payments-platform/group.yaml
```
Minimal working `group.yaml`:
```yaml
version: 1
name: payments-platform
description: Gateway + orders + inventory (gRPC)
repos:
gateway: gateway
orders: orders
inventory: inventory
# Only add explicit links when the automatic extractors miss something —
# see "When automatic extraction isn't enough" below.
links: []
packages: {}
detect:
http: true
grpc: true
topics: true
shared_libs: true
embedding_fallback: false
matching:
bm25_threshold: 0.7
embedding_threshold: 0.65
max_candidates_per_step: 3
```
Field notes (schema in [`types.ts`](../../gitnexus/src/core/group/types.ts)):
- `version` — must be `1`. The parser rejects anything else.
- `name` — required; used for the group directory name and all CLI / MCP calls.
- `repos` — a mapping from **group path** (a logical name you choose; can be a hierarchy like `backend/orders`) to **registry name** (the name shown by `npx gitnexus list`). Both sides appear throughout the tooling: contract rows use the group path; `@<group>/<groupPath>` routes tools to a single member.
- `links` — optional manifest escape hatch, one entry per explicit cross-repo contract. Validated by the parser: `from` and `to` must be known repo paths, `type` must be one of `http | grpc | topic | lib | custom`, and `role` must be `provider | consumer`.
- `detect` — toggles per extractor family. Defaults (set in `config-parser.ts`) turn `http`, `grpc`, `topics`, and `shared_libs` on; disable the ones you don't use to speed up sync.
- `matching` — thresholds for the matching cascade. The exact match is always run; other strategies depend on indexer state.
### 3. Sync the group
```bash
npx gitnexus group sync payments-platform --verbose
```
What this does (see [`sync.ts`](../../gitnexus/src/core/group/sync.ts)):
1. Opens each member's per-repo LadybugDB.
2. Runs the HTTP, gRPC, and topic extractors against the source files.
3. Applies manifest `links` through [`manifest-extractor.ts`](../../gitnexus/src/core/group/extractors/manifest-extractor.ts).
4. Runs the exact-match cascade, joining providers and consumers that share a normalized `contractId`.
5. Writes `contracts.json` in the group directory.
Flags:
- `--exact-only` — stop after the exact cascade; skip BM25 and embedding fallback.
- `--skip-embeddings` — run exact plus BM25 but not embedding-based matching.
- `--allow-stale` — don't warn if a member's index is stale.
- `--json` — machine-readable output.
The same operation is available over MCP as `group_sync({ name: "payments-platform" })` — see [`tools.ts`](../../gitnexus/src/mcp/tools.ts).
### 4. Inspect the registry
Use `gitnexus group contracts` for the CLI view or read the `gitnexus://group/<name>/contracts` MCP resource for the same data.
```bash
npx gitnexus group contracts payments-platform --type grpc --json
```
A shortened response:
```json
{
"contracts": [
{
"contractId": "grpc::orders.OrderService/PlaceOrder",
"type": "grpc",
"role": "provider",
"repo": "orders",
"symbolRef": { "filePath": "internal/grpc/order_server.go", "name": "RegisterOrderServiceServer" },
"confidence": 0.8,
"meta": { "service": "OrderService", "method": "PlaceOrder", "source": "go_register" }
},
{
"contractId": "grpc::orders.OrderService/PlaceOrder",
"type": "grpc",
"role": "consumer",
"repo": "gateway",
"symbolRef": { "filePath": "src/clients/orders.ts", "name": "OrderServiceClient" },
"confidence": 0.75,
"meta": { "service": "OrderService", "source": "ts_generated_client" }
}
],
"crossLinks": [
{
"from": { "repo": "gateway", "symbolUid": "…", "symbolRef": { "filePath": "src/clients/orders.ts", "name": "OrderServiceClient" } },
"to": { "repo": "orders", "symbolUid": "…", "symbolRef": { "filePath": "internal/grpc/order_server.go", "name": "RegisterOrderServiceServer" } },
"type": "grpc",
"contractId": "grpc::orders.OrderService/PlaceOrder",
"matchType": "exact",
"confidence": 1.0
}
]
}
```
Staleness of the underlying indexes shows up in `npx gitnexus group status payments-platform` or the `gitnexus://group/<name>/status` resource.
### 5. Run cross-repo impact with `@<group>` routing
From any shell (you do **not** have to `cd` into a member repo), the normal `impact` / `query` / `context` tools accept `repo: "@<group>"` to fan out across all members, or `repo: "@<group>/<memberPath>"` to target one member. Routing is implemented in [`resolve-at-member.ts`](../../gitnexus/src/core/group/resolve-at-member.ts) and described in [`tools.ts`](../../gitnexus/src/mcp/tools.ts).
Example MCP calls:
```json
{"tool": "impact", "arguments": {
"repo": "@payments-platform/orders",
"target": "PlaceOrder",
"direction": "upstream",
"crossDepth": 2
}}
```
```json
{"tool": "query", "arguments": {
"repo": "@payments-platform",
"query": "retry logic around PlaceOrder"
}}
```
The CLI equivalents still exist for scripting:
```bash
npx gitnexus group impact payments-platform \
--repo orders --target PlaceOrder --direction upstream --cross-depth 2
```
Phase 1 walks within the anchor member; Phase 2 hops across the Contract Bridge wherever a cross-link endpoint matches an impacted symbol. See [`cross-impact.ts`](../../gitnexus/src/core/group/cross-impact.ts) for the bridge query.
## How gRPC extraction works
`GrpcExtractor` ([`grpc-extractor.ts`](../../gitnexus/src/core/group/extractors/grpc-extractor.ts)) runs two passes per member repo:
1. **Proto map.** Every `**/*.proto` file is parsed to enumerate `service Foo { rpc Bar(...) }` blocks and (transitively) resolve the package name. Each RPC method becomes a provider contract with `contractId = grpc::<package>.<Service>/<Method>` and `confidence = 0.85`. Parsing uses the vendored `tree-sitter-proto` grammar when available and falls back to a length-preserving manual parser (`extractServiceBlocks`) otherwise, so `.proto` extraction works on platforms where the grammar fails to build.
2. **Source scan.** Every source file whose extension matches [`GRPC_SCAN_GLOB`](../../gitnexus/src/core/group/extractors/grpc-patterns/index.ts) is parsed by its language plugin:
| Language | Provider signal | Consumer signal |
|----------|-----------------|-----------------|
| Go ([`go.ts`](../../gitnexus/src/core/group/extractors/grpc-patterns/go.ts)) | `pb.RegisterXxxServer(...)`, `pb.UnimplementedXxxServer` embedded in struct | `pb.NewXxxClient(conn)` |
| Java ([`java.ts`](../../gitnexus/src/core/group/extractors/grpc-patterns/java.ts)) | `extends XxxServiceGrpc.XxxServiceImplBase` (with or without `@GrpcService`) | `XxxServiceGrpc.newBlockingStub(...)`, `newStub(...)` |
| Python ([`python.ts`](../../gitnexus/src/core/group/extractors/grpc-patterns/python.ts)) | `add_XxxServicer_to_server(...)` (bare or `_pb2_grpc.` attribute form) | `XxxStub(channel)` (ignores `Mock`/`Test`/`Fake`/`Stub`) |
| Node / TS ([`node.ts`](../../gitnexus/src/core/group/extractors/grpc-patterns/node.ts)) | NestJS `@GrpcMethod('Service','Method')` | `@GrpcClient` field typed `XxxServiceClient`, `client.getService<X>('Service')`, `new XxxServiceClient(...)`, `new foo.bar.XxxService(...)` in files that call `loadPackageDefinition` |
For each source-scan detection the extractor looks up the short service name in the proto map and picks:
- `grpc::<package>.<Service>/<Method>` when a method is named and the service resolves against the proto map,
- `grpc::<package>.<Service>/*` (wildcard) when only the service is known, or
- `grpc::<ServiceName>/*` when no `.proto` is available at all.
Provider detections land at confidence 0.8 (with proto) or 0.65 (without); consumers at 0.75 or 0.55. NestJS `@GrpcMethod` is fixed at 0.8 because the decorator is self-describing.
### Matching
`matching.ts` lowercases the package/service segment before comparing contract ids, so bindings that capitalize names differently (`auth.AuthService` vs `auth.authservice`) still match. Method names are compared case-sensitively because gRPC's wire path is case-sensitive. Service-only wildcards (`grpc::pkg.Svc/*`) match any method on the same service during cross-linking.
### Known limitations
- **Ambiguous proto resolution.** If a short service name exists in more than one `.proto` file and the source-scan hit can't be narrowed down by shared directory segments (`resolveProtoConflict` refuses to guess), the extractor skips contract emission and logs a warning.
- **Proto packages must be resolvable locally.** Transitive imports that point outside the repo produce an empty package segment, which means the contract id collapses to `grpc::<Service>/<Method>`. Cross-repo matches still work as long as both sides agree on the empty package.
- **Rewrite rules are not implemented.** If the provider repo writes `grpc::orders.OrderService/PlaceOrder` and the consumer repo writes `grpc::orderspb.OrderService/PlaceOrder`, they won't cross-link automatically. Use `config.links` to declare the correspondence (see below).
- **One sync = one snapshot.** Contracts are extracted against the indexed snapshot of each repo. Re-index first, then re-sync; the `status` command and resource surface staleness.
## When automatic extraction isn't enough
The escape hatch is the `links` list in `group.yaml`, handled by [`ManifestExtractor`](../../gitnexus/src/core/group/extractors/manifest-extractor.ts). Each entry is a **one-directional** provider/consumer declaration:
```yaml
version: 1
name: payments-platform
repos:
gateway: gateway
orders: orders
inventory: inventory
links:
# Explicit gRPC method: use when naming mismatches stop the
# automatic matcher from cross-linking.
- from: gateway
to: orders
type: grpc
contract: OrderService/PlaceOrder
role: consumer
# Service-level link when you don't want to enumerate methods.
- from: orders
to: inventory
type: grpc
contract: InventoryService
role: consumer
# Works for HTTP too — use `METHOD::/path` form for the exact
# handler, or just `/path` for a method-agnostic wildcard.
- from: gateway
to: orders
type: http
contract: POST::/orders
role: consumer
```
What the manifest extractor does (see [`manifest-extractor.ts`](../../gitnexus/src/core/group/extractors/manifest-extractor.ts)):
1. Builds a canonical `contractId` with `buildContractId` — the same canonicalization used by the automatic extractors, so manifest links cross-match automatic contracts on the other side.
2. Tries to resolve each side to a real graph symbol (the `Route` node for HTTP, a `Function|Method` / `Class|Interface` for gRPC, a `Package|Module` for `lib`).
3. If resolution fails, falls back to a deterministic synthetic uid (`manifest::<repo>::<contractId>`) so both sides still line up in cross-impact — name-only links still work when the symbol isn't in the graph.
4. Emits both a provider and a consumer `StoredContract` (confidence `1.0`, `source: "manifest"`) and a `CrossLink` with `matchType: "manifest"`.
Use `links` for exactly the cases the extractor can't infer: different package names across repos (see #701), hand-rolled transports, cases where the provider repo isn't checked out locally but you still want a record, or any contract whose provider and consumer simply don't share a surface the extractors know how to pattern-match.
History: the manifest extractor used to be silently skipped by the sync pipeline; that was fixed in [#827](https://github.com/abhigyanpatwari/GitNexus/pull/827) (tracking issue #826). If you ever see `config.links` with zero cross-links in `contracts.json`, make sure you're on a build that includes that fix, then re-run `group sync`.
## Troubleshooting
1. **`contracts.json` is empty after a sync.** Either no member repo contained a recognizable gRPC pattern, or the extractors are disabled in `detect`. Confirm `detect.grpc: true` and re-run with `--verbose`.
2. **A known provider/consumer pair doesn't cross-link.** Most common cause: the package segment differs. Check the raw contract ids with `gitnexus group contracts <name> --unmatched` — if you see two same-method contracts with different package prefixes, add a manifest `links:` entry to bridge them (no automatic rewrite rules yet).
3. **`matchType: "manifest"` is missing entirely.** The extractor needs `config.links` to be non-empty and the sync pipeline to actually call it — verify you're on a post-#827 build. Empty contract rows for manifest links usually mean `resolveSymbol` couldn't find a graph match; the synthetic uid still lets cross-impact work, it just won't carry a file path.
4. **Ambiguous proto warnings.** Look for `[grpc-extractor] Ambiguous proto resolution` in the sync logs; that means a service name exists in multiple `.proto` files under the same repo and the path-distance heuristic couldn't pick a winner. Resolve by renaming the service or declaring the intended pairing in `config.links`.
5. **Cross-impact says "stale".** Both sides need a fresh per-repo index _and_ a fresh group sync. Order matters: `gitnexus analyze` in each changed repo, then `gitnexus group sync <name>`. Use `gitnexus group status <name>` to see which side is behind.
## Related docs and references
- [AGENTS.md](../../AGENTS.md) — authoritative list of MCP tools and resources, including group-mode routing and the `gitnexus://group/…` resources.
- [ARCHITECTURE.md](../../ARCHITECTURE.md) — overall data flow and the call-resolution DAG that the per-repo indexer uses.
- [`gitnexus/src/core/group/`](../../gitnexus/src/core/group/) — `service.ts`, `sync.ts`, `config-parser.ts`, `matching.ts`.
- [`gitnexus/src/core/group/extractors/grpc-extractor.ts`](../../gitnexus/src/core/group/extractors/grpc-extractor.ts) and [`grpc-patterns/`](../../gitnexus/src/core/group/extractors/grpc-patterns/) — gRPC detection.
- [`gitnexus/src/core/group/extractors/manifest-extractor.ts`](../../gitnexus/src/core/group/extractors/manifest-extractor.ts) — the `config.links` escape hatch.
- [`gitnexus/src/mcp/tools.ts`](../../gitnexus/src/mcp/tools.ts) — MCP tool schemas (`group_list`, `group_sync`, plus `@<group>` routing on `impact` / `query` / `context`).
- [`gitnexus/src/cli/group.ts`](../../gitnexus/src/cli/group.ts) — CLI command definitions and flags.
- Upstream issues: [#701](https://github.com/abhigyanpatwari/GitNexus/issues/701), [#826](https://github.com/abhigyanpatwari/GitNexus/issues/826), [#906](https://github.com/abhigyanpatwari/GitNexus/issues/906).
+4 -4
View File
@@ -8,13 +8,13 @@
"name": "gitnexus-shared",
"version": "1.0.0",
"devDependencies": {
"typescript": "^6.0.2"
"typescript": "^6.0.3"
}
},
"node_modules/typescript": {
"version": "6.0.2",
"resolved": "https://registry.npmjs.org/typescript/-/typescript-6.0.2.tgz",
"integrity": "sha512-bGdAIrZ0wiGDo5l8c++HWtbaNCWTS4UTv7RaTH/ThVIgjkveJt83m74bBHMJkuCbslY8ixgLBVZJIOiQlQTjfQ==",
"version": "6.0.3",
"resolved": "https://registry.npmjs.org/typescript/-/typescript-6.0.3.tgz",
"integrity": "sha512-y2TvuxSZPDyQakkFRPZHKFm+KKVqIisdg9/CZwm9ftvKXLP8NRWj38/ODjNbr43SsoXqNuAisEf1GdCxqWcdBw==",
"dev": true,
"license": "Apache-2.0",
"bin": {
+1 -1
View File
@@ -20,6 +20,6 @@
"src"
],
"devDependencies": {
"typescript": "^6.0.2"
"typescript": "^6.0.3"
}
}
+16
View File
@@ -131,4 +131,20 @@ export interface GraphRelationship {
confidence: number;
reason: string;
step?: number;
/**
* Per-signal evidence trace for edges emitted by the scope-based
* resolution pipeline (RFC #909 Ring 2 PKG #925). Populated by
* `emit-references.ts` when draining `ReferenceIndex` into the graph
* so downstream query / audit tools can inspect *why* a given edge
* was emitted with its confidence value.
*
* Optional and additive — every existing edge emitter ignores this
* field, and every existing query continues to work whether or not
* an edge carries it.
*/
evidence?: readonly {
readonly kind: string;
readonly weight: number;
readonly note?: string;
}[];
}
+125
View File
@@ -23,3 +23,128 @@ export type { MroStrategy } from './mro-strategy.js';
// Pipeline progress
export type { PipelinePhase, PipelineProgress } from './pipeline.js';
// ─── Scope-based resolution — RFC #909 (Ring 1 #910) ────────────────────────
// Data model (RFC §2)
export type { SymbolDefinition } from './scope-resolution/symbol-definition.js';
export type {
ScopeId,
DefId,
ScopeKind,
Range,
Capture,
CaptureMatch,
BindingRef,
ImportEdge,
TypeRef,
Scope,
ResolutionEvidence,
Resolution,
Reference,
ReferenceIndex,
LookupParams,
RegistryContributor,
ParsedImport,
ParsedTypeBinding,
WorkspaceIndex,
Callsite,
ScopeLookup,
} from './scope-resolution/types.js';
// Evidence + tie-break constants (RFC Appendix A, Appendix B)
export { EvidenceWeights, typeBindingWeightAtDepth } from './scope-resolution/evidence-weights.js';
export { ORIGIN_PRIORITY } from './scope-resolution/origin-priority.js';
export type { OriginForTieBreak } from './scope-resolution/origin-priority.js';
// Language classification (RFC §6.1 Ring 3/4 governance)
export {
LanguageClassifications,
isProductionLanguage,
} from './scope-resolution/language-classification.js';
export type { LanguageClassification } from './scope-resolution/language-classification.js';
// Core indexes over per-file artifacts (RFC §3.1; Ring 2 SHARED #913)
export { buildDefIndex } from './scope-resolution/def-index.js';
export type { DefIndex } from './scope-resolution/def-index.js';
export { buildModuleScopeIndex } from './scope-resolution/module-scope-index.js';
export type { ModuleScopeIndex, ModuleScopeEntry } from './scope-resolution/module-scope-index.js';
export { buildQualifiedNameIndex } from './scope-resolution/qualified-name-index.js';
export type { QualifiedNameIndex } from './scope-resolution/qualified-name-index.js';
// Strict type-reference resolver (RFC §4.6; Ring 2 SHARED #916)
// `ScopeLookup` is defined in `./scope-resolution/types.js` and exported
// from the type-export block above — not from this module.
export { resolveTypeRef } from './scope-resolution/resolve-type-ref.js';
export type { ResolveTypeRefContext } from './scope-resolution/resolve-type-ref.js';
// ScopeExtractor output contracts (RFC §3.2 Phase 1; Ring 2 PKG #919)
export type { ParsedFile } from './scope-resolution/parsed-file.js';
export type { ReferenceSite, ReferenceKind, CallForm } from './scope-resolution/reference-site.js';
// Method-dispatch materialized view over HeritageMap (RFC §3.1; Ring 2 SHARED #914)
export { buildMethodDispatchIndex } from './scope-resolution/method-dispatch-index.js';
export type {
MethodDispatchIndex,
MethodDispatchInput,
} from './scope-resolution/method-dispatch-index.js';
// SCC-aware cross-file finalize (RFC §3.2 Phase 2; Ring 2 SHARED #915)
export { finalize } from './scope-resolution/finalize-algorithm.js';
export type {
FinalizeInput,
FinalizeFile,
FinalizeHooks,
FinalizeOutput,
FinalizedScc,
FinalizeStats,
} from './scope-resolution/finalize-algorithm.js';
// Scope-aware registries + 7-step lookup (RFC §4; Ring 2 SHARED #917)
export { buildClassRegistry } from './scope-resolution/registries/class-registry.js';
export type { ClassRegistry } from './scope-resolution/registries/class-registry.js';
export { buildMethodRegistry } from './scope-resolution/registries/method-registry.js';
export type {
MethodRegistry,
MethodLookupOptions,
} from './scope-resolution/registries/method-registry.js';
export { buildFieldRegistry } from './scope-resolution/registries/field-registry.js';
export type {
FieldRegistry,
FieldLookupOptions,
} from './scope-resolution/registries/field-registry.js';
export { lookupCore } from './scope-resolution/registries/lookup-core.js';
export type { CoreLookupParams } from './scope-resolution/registries/lookup-core.js';
export { lookupQualified } from './scope-resolution/registries/lookup-qualified.js';
export type { LookupQualifiedParams } from './scope-resolution/registries/lookup-qualified.js';
export { composeEvidence, confidenceFromEvidence } from './scope-resolution/registries/evidence.js';
export type { RawSignals } from './scope-resolution/registries/evidence.js';
export {
compareByConfidenceWithTiebreaks,
CONFIDENCE_EPSILON,
} from './scope-resolution/registries/tie-breaks.js';
export type { TieBreakKey } from './scope-resolution/registries/tie-breaks.js';
export { CLASS_KINDS, METHOD_KINDS, FIELD_KINDS } from './scope-resolution/registries/context.js';
export type {
RegistryContext,
RegistryProviders,
OwnerScopedContributor,
ArityVerdict,
} from './scope-resolution/registries/context.js';
// Scope tree spine + position lookup (RFC §2.2 + §3.1; Ring 2 SHARED #912)
export { makeScopeId, clearScopeIdInternPool } from './scope-resolution/scope-id.js';
export type { ScopeIdInput } from './scope-resolution/scope-id.js';
export { buildScopeTree, ScopeTreeInvariantError } from './scope-resolution/scope-tree.js';
export type { ScopeTree } from './scope-resolution/scope-tree.js';
export { buildPositionIndex } from './scope-resolution/position-index.js';
export type { PositionIndex } from './scope-resolution/position-index.js';
// Shadow-mode diff + aggregation (RFC §6.3; Ring 2 SHARED #918)
export { diffResolutions } from './scope-resolution/shadow/diff.js';
export type {
ShadowAgreement,
ShadowCallsite,
ShadowDiff,
} from './scope-resolution/shadow/diff.js';
export { aggregateDiffs } from './scope-resolution/shadow/aggregate.js';
export type { LanguageParityRow, ShadowParityReport } from './scope-resolution/shadow/aggregate.js';
+37 -14
View File
@@ -1,23 +1,46 @@
/**
* MRO (Method Resolution Order) strategy — shared between CLI and any
* future consumer that reasons about multiple-inheritance semantics.
* MRO (Method Resolution Order) strategy — shared canonical definition.
*
* Lives in `gitnexus-shared` so the low-level resolution module
* (`core/ingestion/model/resolve.ts`) does not need to import from
* `languages/` — keeping the `model/` layer free of language-registry
* coupling.
* Lives in `gitnexus-shared` so `model/resolve.ts` and `mro-processor.ts` share
* the type without importing the language registry (avoids circular coupling).
*
* Strategy semantics:
* - `first-wins`: BFS ancestor walk, first match wins (default).
* - `leftmost-base`: BFS ancestor walk, leftmost base wins (C++).
* - `c3`: C3-linearized ancestor order, first match wins (Python).
* - `implements-split`: BFS walk, first match wins (Java/C#/Kotlin) — full
* interface-default ambiguity is handled at graph level.
* - `qualified-syntax`: No auto-resolution (Rust — requires `<T as Trait>::m`).
* `first-wins` (default, Java/C#/Kotlin/Go/Swift/Dart):
* BFS ancestor walk in declaration order; first match wins.
*
* `leftmost-base` (C++):
* BFS walk; HeritageMap preserves source insertion order, so BFS naturally
* picks the leftmost base in diamond inheritance.
*
* `c3` (Python):
* C3-linearization; falls back to BFS on cyclic/inconsistent hierarchy.
* See model/resolve.ts § c3Linearize.
*
* `implements-split` (Java/C#/Kotlin):
* Low-level lookup is BFS; graph-level mro-processor detects and warns on
* interface-default method ambiguity.
*
* `qualified-syntax` (Rust):
* No auto-resolution — `lookupMethodByOwnerWithMRO` returns undefined immediately.
* Rust requires explicit `<Type as Trait>::method` syntax.
*
* `ruby-mixin` (Ruby):
* Kind-aware walk that does NOT short-circuit on direct owner first (`prepend`
* must beat the class's own method). Walk order:
* 1. Prepend providers (reverse declaration — last-prepended wins)
* 2. Direct owner's own methods
* 3. Include providers (reverse declaration)
* 4. Transitive ancestors (BFS fallback)
* Singleton dispatch: caller passes `ancestryOverride` (extend providers only);
* becomes a simple left-to-right scan. Miss NEVER falls through to file-scoped
* lookup — null-routes or honors `fallback`.
*
* @see model/resolve.ts § lookupMethodByOwnerWithMRO
* @see languages/ruby.ts § selectDispatch
*/
export type MroStrategy =
| 'first-wins'
| 'c3'
| 'leftmost-base'
| 'implements-split'
| 'qualified-syntax';
| 'qualified-syntax'
| 'ruby-mixin';
@@ -0,0 +1,62 @@
/**
* `DefIndex` — O(1) `DefId → SymbolDefinition` materialization.
*
* The global "what is this id?" lookup. Every per-kind registry (ClassRegistry,
* MethodRegistry, FieldRegistry) returns `DefId[]` and resolves them back to
* full `SymbolDefinition` records through this index — one central hash map,
* one allocation per def.
*
* Part of RFC #909 Ring 2 SHARED — #913.
*
* Consumed by: #917 (`Registry.lookup` implementations), #915 (SCC finalize).
*/
import type { SymbolDefinition } from './symbol-definition.js';
import type { DefId } from './types.js';
export interface DefIndex {
readonly byId: ReadonlyMap<DefId, SymbolDefinition>;
readonly size: number;
get(id: DefId): SymbolDefinition | undefined;
has(id: DefId): boolean;
}
/**
* Build a `DefIndex` from a flat list of `SymbolDefinition` records.
*
* **Collision policy: first-write-wins.** `DefId` is meant to be unique
* (`nodeId` is the stable graph identifier), so a collision indicates an
* upstream bug — most likely the same symbol parsed twice or a duplicate
* commit into the pipeline. Rather than silently overwriting with a later
* definition that may be partial or wrong, the first record wins and
* subsequent records for the same id are dropped. Pipeline bugs surface
* later as `has(id) === true` but the def looking older than expected,
* which is easier to debug than a silent overwrite.
*
* Pure function — safe to call repeatedly; no side effects.
*/
export function buildDefIndex(defs: readonly SymbolDefinition[]): DefIndex {
const byId = new Map<DefId, SymbolDefinition>();
for (const def of defs) {
if (byId.has(def.nodeId)) continue; // first-write-wins
byId.set(def.nodeId, def);
}
return wrapIndex(byId);
}
// ─── Internal ───────────────────────────────────────────────────────────────
function wrapIndex(byId: Map<DefId, SymbolDefinition>): DefIndex {
return {
byId,
get size() {
return byId.size;
},
get(id: DefId): SymbolDefinition | undefined {
return byId.get(id);
},
has(id: DefId): boolean {
return byId.has(id);
},
};
}
@@ -0,0 +1,90 @@
/**
* `EvidenceWeights` — RFC Appendix A (authoritative values).
*
* Starting calibration for scope-based resolution. Shadow-first rollout
* tunes these against legacy DAG parity. Every `ResolutionEvidence.weight`
* value in the codebase MUST reference this map; inline magic numbers are a
* lint violation. Extends issue #429 (centralize hardcoded confidence values).
*
* Evidence composes additively inside `composeEvidence`; the sum is capped
* at 1.0 in `Resolution.confidence`.
*/
/**
* Authoritative weight map. Keys are a mix of `ResolutionEvidence.kind`
* values and special modifiers (scope-chain depth, MRO depth decay,
* unlinked-import multiplicative cap).
*/
export const EvidenceWeights = {
// ─── Where-found signals (visibility) ─────────────────────────────────────
/** `BindingRef.origin === 'local'` */
local: 0.55,
/** `BindingRef.origin === 'import'` */
import: 0.45,
/** `BindingRef.origin === 'reexport'` */
reexport: 0.4,
/** `BindingRef.origin === 'namespace'` */
namespace: 0.4,
/** `BindingRef.origin === 'wildcard'` */
wildcard: 0.3,
// ─── Scope-chain deduction (per-hop) ──────────────────────────────────────
/** Deducted per parent-hop taken (depth-0 = 0, depth-1 = −0.02, …). */
scopeChainPerDepth: -0.02,
// ─── Receiver-type-binding signal (decays by MRO depth) ───────────────────
/**
* Weight applied when the receiver's type binding resolves to a class that
* declares the candidate as a method/field. Decays by MRO depth: direct
* class = index 0; 1 parent hop = index 1; etc. Falls back to the last
* value for depths beyond the table.
*/
typeBindingByMroDepth: [0.5, 0.42, 0.36, 0.32, 0.3] as const,
// ─── Corroborating signals ────────────────────────────────────────────────
/** `def.ownerId === resolvedReceiver.def.id` (exact owner match). */
ownerMatch: 0.2,
/** Explanatory only — retained for debuggability. Never discriminates
* because surviving candidates already passed `acceptedKinds`. */
kindMatch: 0.0,
// ─── Arity compatibility (from `provider.arityCompatibility`) ─────────────
/** `provider.arityCompatibility(...) === 'compatible'` */
arityMatchCompatible: 0.1,
/** `provider.arityCompatibility(...) === 'unknown'` */
arityMatchUnknown: 0.0,
/** `provider.arityCompatibility(...) === 'incompatible'` — penalizes;
* candidates filtered only when a compatible candidate exists. */
arityMatchIncompatible: -0.15,
// ─── Global fallback (only when nothing lexically visible) ────────────────
/** Hit via `QualifiedNameIndex.byQualifiedName`. */
globalQualified: 0.35,
/** Fallback hit in a `byName` index (and nothing was lexically visible). */
globalName: 0.1,
// ─── Degraded signals ─────────────────────────────────────────────────────
/** Call/reference flowing through a `dynamic-unresolved` edge. */
dynamicImportUnresolved: 0.02,
// ─── Unresolved-import cap (multiplicative, applied per-signal) ───────────
/**
* Multiplicative cap on the edge-derived evidence signal
* (`import`/`wildcard`/`reexport`/`namespace`) when
* `ImportEdge.linkStatus === 'unresolved'`. Independent corroborating
* signals on the same candidate (`owner-match`, `arity-match`,
* `type-binding`) are NOT penalized.
*/
unlinkedImportMultiplier: 0.5,
} as const;
/**
* Look up the `type-binding` signal weight for a given MRO depth, falling
* back to the last tabulated value for depths beyond the table.
*/
export function typeBindingWeightAtDepth(mroDepth: number): number {
const table = EvidenceWeights.typeBindingByMroDepth;
if (mroDepth < 0) return table[0];
if (mroDepth >= table.length) return table[table.length - 1];
return table[mroDepth];
}
@@ -0,0 +1,663 @@
/**
* `finalize` — cross-file finalize algorithm for the SemanticModel
* (RFC §3.2 Phase 2; Ring 2 SHARED #915).
*
* Pure logic that takes per-file parse output (`ParsedImport[]` +
* `SymbolDefinition[]`) and returns:
*
* - Linked `ImportEdge[]` per module scope, with `targetModuleScope` and
* `targetDefId` filled where resolvable; edges that could not be
* resolved within the hard fixpoint cap are marked
* `linkStatus: 'unresolved'`.
* - Materialized `bindings` per module scope — local defs merged with
* imported / wildcard-expanded / re-exported names via the provider's
* `mergeBindings` precedence.
* - The SCC condensation of the import graph, exposed so disjoint SCCs
* can be processed in parallel by callers that want that.
*
* The algorithm is **SCC-aware**: it runs Tarjan SCC over the file-level
* import graph, processes SCCs in reverse-topological order (leaves
* first), and within each SCC runs a bounded fixpoint link pass capped at
* `N = |edges in SCC|`. Cyclic imports finalize without hanging; malformed
* inputs are bounded by the cap.
*
* **No language-specific logic.** Target resolution, wildcard expansion,
* and binding precedence all go through caller-supplied hooks
* (`resolveImportTarget`, `expandsWildcardTo`, `mergeBindings`) that
* match the LanguageProvider surface from #911.
*
* **Dynamic imports rule.** `kind === 'dynamic-unresolved'` passes through
* as an `ImportEdge { kind: 'dynamic-unresolved', targetFile: null }`
* with no `BindingRef`. They are parse-time signals, not linkable targets.
*/
import type { SymbolDefinition } from './symbol-definition.js';
import type { BindingRef, ImportEdge, ParsedImport, ScopeId, WorkspaceIndex } from './types.js';
// ─── Public contracts ───────────────────────────────────────────────────────
/** Per-file input for the finalize pass. */
export interface FinalizeFile {
readonly filePath: string;
/** The module scope id for this file; owns the finalized imports + bindings. */
readonly moduleScope: ScopeId;
readonly parsedImports: readonly ParsedImport[];
/**
* Defs exported from this file — the "what other files can import by name"
* surface. Typically those with `isExported: true` (the module's own
* declarations) plus, for multi-hop re-export chains, the re-exported
* names the parser chose to surface here.
*
* **Multi-hop re-export contract.** `finalize` resolves an edge
* `A → B (importedName: 'X')` by looking up `X` in `B.localDefs`. If B
* only has `export { X } from './C'` and the parser *does not* include
* `X` in `B.localDefs`, A's edge hits the fixpoint cap and is marked
* `linkStatus: 'unresolved'`. The fixpoint does NOT mutate `localDefs`
* across iterations — it is static input.
*
* Parsers that want multi-hop re-export chains to settle end-to-end must
* include re-exported names in the intermediate file's `localDefs` (with
* the original `DefId` of the source symbol). This keeps the algorithm
* O(1) per lookup and avoids graph-crawl during finalize.
*/
readonly localDefs: readonly SymbolDefinition[];
}
/** Input to `finalize`. */
export interface FinalizeInput {
readonly files: readonly FinalizeFile[];
/** Opaque workspace context forwarded to provider hooks. */
readonly workspaceIndex: WorkspaceIndex;
}
/**
* Provider-supplied hooks. Mirror the optional LanguageProvider scope-
* resolution hooks declared in #911; `finalize` calls them pure-ly and
* expects pure answers.
*/
export interface FinalizeHooks {
/**
* Resolve a raw import target to the concrete file path that owns it.
* Return `null` when no target file is resolvable (e.g., `np.foo` when
* `numpy` is external to the workspace).
*/
resolveImportTarget(
targetRaw: string,
fromFile: string,
workspaceIndex: WorkspaceIndex,
): string | null;
/**
* For a wildcard `import * from M`, return the names visible in the
* exporting module scope `M`. The finalize pass looks each name up in
* `M`'s local defs to produce a concrete `BindingRef`; names with no
* matching export are dropped.
*/
expandsWildcardTo(targetModuleScope: ScopeId, workspaceIndex: WorkspaceIndex): readonly string[];
/**
* Merge `incoming` bindings into `existing` for a given name. Called
* once per name at each scope. Typical rules:
* - Python: local > imported > wildcard (last-write-wins within tier).
* - Rust: explicit `use` > glob; `pub use` overrides.
* Return value replaces the bucket entirely — no implicit append.
*/
mergeBindings(
existing: readonly BindingRef[],
incoming: readonly BindingRef[],
scope: ScopeId,
): readonly BindingRef[];
}
/** One SCC in the file-level import graph. */
export interface FinalizedScc {
readonly files: readonly string[];
/** True iff this SCC has ≥ 2 files OR a single file that self-imports. */
readonly isCycle: boolean;
}
/**
* Counters reported by `finalize`.
*
* **Counting granularity** — all edge counters are **per-`ParsedImport`**,
* not per-materialized-`ImportEdge`. A single `wildcard` ParsedImport that
* expands to N exports counts as one linked edge in these stats; the
* materialized output (`FinalizeOutput.imports`) will have N edges for
* that input. `dynamic-unresolved` ParsedImports count as linked (they
* pass through with no `linkStatus`), so `linkedEdges` ≠ "has a
* BindingRef" — use the `bindings` map for that.
*
* In other words: `totalEdges === input.parsedImports.length` summed
* across files, and `linkedEdges + unresolvedEdges === totalEdges`.
*/
export interface FinalizeStats {
readonly totalFiles: number;
/** Total `ParsedImport` records seen across all files. */
readonly totalEdges: number;
/**
* `ParsedImport`s whose finalized edge does NOT carry
* `linkStatus: 'unresolved'`. Includes `dynamic-unresolved` pass-throughs.
*/
readonly linkedEdges: number;
/** `ParsedImport`s whose finalized edge carries `linkStatus: 'unresolved'`. */
readonly unresolvedEdges: number;
readonly sccCount: number;
readonly largestSccSize: number;
}
export interface FinalizeOutput {
/** Linked `ImportEdge[]` per module scope, in original input order. */
readonly imports: ReadonlyMap<ScopeId, readonly ImportEdge[]>;
/** Materialized bindings per module scope. */
readonly bindings: ReadonlyMap<ScopeId, ReadonlyMap<string, readonly BindingRef[]>>;
/** SCCs in reverse-topological order (leaves first). */
readonly sccs: readonly FinalizedScc[];
readonly stats: FinalizeStats;
}
// ─── Entry point ───────────────────────────────────────────────────────────
export function finalize(input: FinalizeInput, hooks: FinalizeHooks): FinalizeOutput {
const byFilePath = new Map<string, FinalizeFile>();
for (const f of input.files) byFilePath.set(f.filePath, f);
// ── Phase 0: pre-resolve raw import targets (one syscall-equivalent per
// (file, parsedImport)). Edges with no resolvable target become
// `linkStatus: 'unresolved'` or, for dynamic-unresolved, pass through
// with `targetFile: null`.
const edgeIndex = new Map<string, ImportEdgeDraft[]>(); // filePath → drafts
let totalEdges = 0;
for (const file of input.files) {
const drafts: ImportEdgeDraft[] = [];
for (const parsed of file.parsedImports) {
const draft = makeEdgeDraft(parsed, file, hooks, input.workspaceIndex);
drafts.push(draft);
totalEdges++;
}
edgeIndex.set(file.filePath, drafts);
}
// ── Phase 1: build file-level import graph (only resolvable edges form
// graph edges; unresolvable ones are terminal and contribute no
// fixpoint obligation).
const graph = new Map<string, Set<string>>();
for (const file of input.files) {
graph.set(file.filePath, new Set());
}
for (const [fromFile, drafts] of edgeIndex) {
const edges = graph.get(fromFile)!;
for (const d of drafts) {
if (d.targetFile !== null && byFilePath.has(d.targetFile)) {
edges.add(d.targetFile);
}
}
}
// ── Phase 2: Tarjan SCC → reverse-topological list of SCCs.
const sccs = tarjanSccs(graph);
// ── Phase 3: process SCCs in reverse-topological order (leaves first).
// Within each SCC, run a bounded fixpoint that resolves intra-SCC edges.
// Edges leaving the SCC are already resolved (their target SCC is
// already finalized); edges inside the SCC may need multiple passes.
const linkedByScope = new Map<ScopeId, readonly ImportEdge[]>();
let linkedEdges = 0;
for (const scc of sccs) {
const sccFiles = new Set(scc.files);
const capacity = countEdgesWithin(edgeIndex, sccFiles);
// Run the fixpoint up to `capacity` iterations. Each iteration tries to
// resolve every still-unlinked edge in the SCC; stops early if a pass
// makes no progress.
let progressed = true;
let iterations = 0;
while (progressed && iterations < capacity) {
progressed = false;
iterations++;
for (const filePath of scc.files) {
const drafts = edgeIndex.get(filePath)!;
for (const draft of drafts) {
if (draft.finalized !== null) continue;
const finalized = tryFinalize(draft, byFilePath);
if (finalized !== null) {
draft.finalized = finalized;
progressed = true;
}
}
}
}
// Any drafts still not finalized within this SCC hit the cap → unresolved.
for (const filePath of scc.files) {
const drafts = edgeIndex.get(filePath)!;
for (const draft of drafts) {
if (draft.finalized !== null) continue;
draft.finalized = {
...draft.base,
linkStatus: 'unresolved' as const,
};
}
}
}
// ── Phase 4: collect finalized `ImportEdge[]` per module scope, preserving
// input order within each file, and wildcard-expand where applicable.
for (const file of input.files) {
const drafts = edgeIndex.get(file.filePath)!;
const finalized: ImportEdge[] = [];
for (const d of drafts) {
const edge = d.finalized!;
if (d.source.kind === 'wildcard' && edge.linkStatus !== 'unresolved') {
// Produce one `wildcard-expanded` ImportEdge per exported name.
const expanded = expandWildcard(edge, byFilePath, hooks, input.workspaceIndex);
for (const e of expanded) finalized.push(e);
} else {
finalized.push(edge);
}
if (edge.linkStatus !== 'unresolved') linkedEdges++;
}
linkedByScope.set(file.moduleScope, Object.freeze(finalized));
}
// ── Phase 5: materialize module-scope bindings (local + imports + wildcards),
// delegating precedence to `provider.mergeBindings`.
const bindingsByScope = materializeBindings(input.files, linkedByScope, hooks);
// ── Stats.
const sccCount = sccs.length;
let largestSccSize = 0;
for (const scc of sccs) {
if (scc.files.length > largestSccSize) largestSccSize = scc.files.length;
}
const stats: FinalizeStats = {
totalFiles: input.files.length,
totalEdges,
linkedEdges,
unresolvedEdges: totalEdges - linkedEdges,
sccCount,
largestSccSize,
};
return Object.freeze({
imports: linkedByScope,
bindings: bindingsByScope,
sccs,
stats,
});
}
// ─── Internal: edge drafting (phase 0) ──────────────────────────────────────
interface ImportEdgeDraft {
readonly source: ParsedImport;
readonly fromFile: string;
readonly fromScope: ScopeId;
readonly targetFile: string | null;
readonly base: ImportEdge;
finalized: ImportEdge | null;
}
function makeEdgeDraft(
parsed: ParsedImport,
file: FinalizeFile,
hooks: FinalizeHooks,
workspace: WorkspaceIndex,
): ImportEdgeDraft {
// Dynamic-unresolved passes through — no `BindingRef`, no target file.
if (parsed.kind === 'dynamic-unresolved') {
const base: ImportEdge = {
localName: parsed.localName,
targetFile: null,
targetExportedName: '',
kind: 'dynamic-unresolved',
};
return {
source: parsed,
fromFile: file.filePath,
fromScope: file.moduleScope,
targetFile: null,
base,
finalized: base, // already fully finalized
};
}
const targetFile = hooks.resolveImportTarget(parsed.targetRaw ?? '', file.filePath, workspace);
// Edge is unresolvable at the file level — mark unresolved now.
if (targetFile === null) {
const edgeKind = parsed.kind === 'wildcard' ? 'wildcard-expanded' : parsed.kind;
const localName = parsed.kind === 'wildcard' ? '' : parsed.localName;
const targetExportedName = extractExportedName(parsed);
const base: ImportEdge = {
localName,
targetFile: null,
targetExportedName,
kind: edgeKind,
linkStatus: 'unresolved',
};
return {
source: parsed,
fromFile: file.filePath,
fromScope: file.moduleScope,
targetFile: null,
base,
finalized: base,
};
}
// Resolvable at the file level; intra-SCC fixpoint may still fail to fill
// in `targetDefId` (e.g., symbol not exported from target).
const edgeKind = parsed.kind === 'wildcard' ? 'wildcard-expanded' : parsed.kind;
const localName = parsed.kind === 'wildcard' ? '' : parsed.localName;
const targetExportedName = extractExportedName(parsed);
const base: ImportEdge = {
localName,
targetFile,
targetExportedName,
kind: edgeKind,
};
return {
source: parsed,
fromFile: file.filePath,
fromScope: file.moduleScope,
targetFile,
base,
finalized: null,
};
}
function extractExportedName(parsed: ParsedImport): string {
switch (parsed.kind) {
case 'named':
case 'alias':
case 'namespace':
case 'reexport':
return parsed.importedName;
case 'wildcard':
case 'dynamic-unresolved':
return '';
}
}
// ─── Internal: per-edge finalization (phase 3) ─────────────────────────────
function tryFinalize(
draft: ImportEdgeDraft,
byFilePath: Map<string, FinalizeFile>,
): ImportEdge | null {
const targetFile = draft.targetFile;
if (targetFile === null) return draft.base; // already terminal
const targetModule = byFilePath.get(targetFile);
if (targetModule === undefined) return draft.base; // external target — leave as-is
// Wildcards finalize at the file level; their per-name expansion happens
// in phase 4. At this stage we just record the target module scope.
if (draft.source.kind === 'wildcard') {
return {
...draft.base,
targetModuleScope: targetModule.moduleScope,
};
}
// Namespace imports alias the target *module*; they don't name a
// specific export. Link the module scope unconditionally. If the target
// also exposes a def whose simple name matches `importedName` (some
// languages emit a synthetic module-def), pick it up as the `targetDefId`
// so consumers can reach the module as a symbol — but its absence is not
// a failure.
if (draft.source.kind === 'namespace') {
const moduleDef = findExportByName(targetModule.localDefs, extractExportedName(draft.source));
return {
...draft.base,
targetModuleScope: targetModule.moduleScope,
...(moduleDef !== undefined ? { targetDefId: moduleDef.nodeId } : {}),
};
}
// named / alias / reexport: look up the imported name in the target's
// local defs. Multi-hop re-export chains settle iteratively — each hop
// resolves once its prior hop is finalized.
const importedName = extractExportedName(draft.source);
const exported = findExportByName(targetModule.localDefs, importedName);
if (exported === undefined) {
// Target resolvable but the name isn't exported — keep trying in case a
// re-export inside the target's SCC surfaces it in a later iteration.
return null;
}
const transitiveVia = draft.source.kind === 'reexport' ? Object.freeze([targetFile]) : undefined;
return {
...draft.base,
targetModuleScope: targetModule.moduleScope,
targetDefId: exported.nodeId,
...(transitiveVia !== undefined ? { transitiveVia } : {}),
};
}
/**
* The "simple" (unqualified) name of a def, for import-name matching.
*
* Canonical source: `def.qualifiedName` — the tail after the last `.` (or
* the whole string if no dot). Defs without a qualifiedName can't be
* resolved by name here and return `null`; callers treat that as "name
* not exported" and either retry in a later fixpoint iteration or mark
* the edge unresolved.
*/
function deriveSimpleName(def: SymbolDefinition): string | null {
const q = def.qualifiedName;
if (q === undefined || q.length === 0) return null;
const dot = q.lastIndexOf('.');
return dot === -1 ? q : q.slice(dot + 1);
}
function findExportByName(
defs: readonly SymbolDefinition[],
name: string,
): SymbolDefinition | undefined {
for (const d of defs) {
if (deriveSimpleName(d) === name) return d;
}
return undefined;
}
function countEdgesWithin(edgeIndex: Map<string, ImportEdgeDraft[]>, files: Set<string>): number {
let n = 0;
for (const filePath of files) {
const drafts = edgeIndex.get(filePath);
if (drafts === undefined) continue;
for (const d of drafts) {
if (d.targetFile !== null && files.has(d.targetFile)) n++;
}
}
// Guarantee at least one pass even for a trivial SCC (ensures deterministic
// fixpoint termination even when a single-file SCC has zero intra-SCC edges
// but still needs one settle pass).
return Math.max(n, 1);
}
// ─── Internal: wildcard expansion (phase 4) ────────────────────────────────
function expandWildcard(
edge: ImportEdge,
byFilePath: Map<string, FinalizeFile>,
hooks: FinalizeHooks,
workspace: WorkspaceIndex,
): readonly ImportEdge[] {
if (edge.targetModuleScope === undefined || edge.targetFile === null) {
return [edge]; // unresolvable wildcard survives as a single unlinked edge
}
const target = byFilePath.get(edge.targetFile);
if (target === undefined) return [edge];
const names = hooks.expandsWildcardTo(edge.targetModuleScope, workspace);
if (names.length === 0) return [];
const expanded: ImportEdge[] = [];
for (const name of names) {
const def = findExportByName(target.localDefs, name);
if (def === undefined) continue;
expanded.push({
localName: name,
targetFile: edge.targetFile,
targetExportedName: name,
kind: 'wildcard-expanded',
targetModuleScope: edge.targetModuleScope,
targetDefId: def.nodeId,
});
}
return expanded;
}
// ─── Internal: bindings materialization (phase 5) ───────────────────────────
function materializeBindings(
files: readonly FinalizeFile[],
linkedByScope: ReadonlyMap<ScopeId, readonly ImportEdge[]>,
hooks: FinalizeHooks,
): ReadonlyMap<ScopeId, ReadonlyMap<string, readonly BindingRef[]>> {
const out = new Map<ScopeId, ReadonlyMap<string, readonly BindingRef[]>>();
for (const file of files) {
const scopeBindings = new Map<string, readonly BindingRef[]>();
// Start with local defs as `origin: 'local'` bindings.
for (const def of file.localDefs) {
const name = deriveSimpleName(def);
if (name === null) continue;
const incoming: BindingRef[] = [{ def, origin: 'local' }];
const existing = scopeBindings.get(name) ?? [];
scopeBindings.set(name, hooks.mergeBindings(existing, incoming, file.moduleScope));
}
// Layer in finalized imports.
const imports = linkedByScope.get(file.moduleScope) ?? [];
for (const edge of imports) {
if (edge.targetDefId === undefined || edge.linkStatus === 'unresolved') continue;
// Every def the importing file needs to reach is in some other file's
// `localDefs`; walk all files to find it. In practice we could index
// this, but at finalize-time N(files) is small per workspace pass.
const def = findDefById(files, edge.targetDefId);
if (def === undefined) continue;
const origin: BindingRef['origin'] =
edge.kind === 'namespace'
? 'namespace'
: edge.kind === 'wildcard-expanded'
? 'wildcard'
: edge.kind === 'reexport'
? 'reexport'
: 'import';
const fallback = deriveSimpleName(def);
const name = edge.localName.length > 0 ? edge.localName : fallback;
if (name === null) continue;
const incoming: BindingRef[] = [{ def, origin, via: edge }];
const existing = scopeBindings.get(name) ?? [];
scopeBindings.set(name, hooks.mergeBindings(existing, incoming, file.moduleScope));
}
// Freeze nested buckets for immutability.
const frozen = new Map<string, readonly BindingRef[]>();
for (const [name, refs] of scopeBindings) {
frozen.set(name, Object.freeze(refs.slice()));
}
out.set(file.moduleScope, frozen);
}
return out;
}
function findDefById(files: readonly FinalizeFile[], defId: string): SymbolDefinition | undefined {
for (const f of files) {
for (const d of f.localDefs) {
if (d.nodeId === defId) return d;
}
}
return undefined;
}
// ─── Internal: Tarjan SCC ──────────────────────────────────────────────────
/**
* Iterative Tarjan SCC. Returns SCCs in **reverse-topological** order
* (leaves first — a property Tarjan gives for free, and the order
* `finalize` wants so leaves are fully resolved before their dependents).
*/
function tarjanSccs(graph: ReadonlyMap<string, ReadonlySet<string>>): FinalizedScc[] {
const index = new Map<string, number>();
const lowlink = new Map<string, number>();
const onStack = new Set<string>();
const stack: string[] = [];
const sccs: FinalizedScc[] = [];
let idx = 0;
// Iterative DFS to avoid stack overflow on deep import chains.
const allNodes = Array.from(graph.keys()).sort(); // deterministic order
const iterStack: Array<{ node: string; children: Iterator<string>; entered: boolean }> = [];
for (const root of allNodes) {
if (index.has(root)) continue;
iterStack.push({
node: root,
children: (graph.get(root) ?? new Set<string>()).values(),
entered: false,
});
while (iterStack.length > 0) {
const frame = iterStack[iterStack.length - 1]!;
if (!frame.entered) {
frame.entered = true;
index.set(frame.node, idx);
lowlink.set(frame.node, idx);
idx++;
stack.push(frame.node);
onStack.add(frame.node);
}
const nextChild = frame.children.next();
if (nextChild.done) {
// Post-visit: compute SCC membership if frame.node is a root.
if (lowlink.get(frame.node) === index.get(frame.node)) {
const scc: string[] = [];
let selfInCycle = false;
while (true) {
const w = stack.pop()!;
onStack.delete(w);
scc.push(w);
// A single-file self-loop counts as a cycle.
if (w === frame.node) {
selfInCycle = (graph.get(w) ?? new Set()).has(w);
break;
}
}
const isCycle = scc.length > 1 || selfInCycle;
sccs.push({ files: Object.freeze(scc), isCycle });
}
iterStack.pop();
// Propagate lowlink to parent.
if (iterStack.length > 0) {
const parent = iterStack[iterStack.length - 1]!;
lowlink.set(parent.node, Math.min(lowlink.get(parent.node)!, lowlink.get(frame.node)!));
}
continue;
}
const child = nextChild.value;
if (!index.has(child)) {
iterStack.push({
node: child,
children: (graph.get(child) ?? new Set<string>()).values(),
entered: false,
});
} else if (onStack.has(child)) {
lowlink.set(frame.node, Math.min(lowlink.get(frame.node)!, index.get(child)!));
}
}
}
return sccs;
}
@@ -0,0 +1,49 @@
/**
* `LanguageClassification` — RFC §6.1 Ring 3 / Ring 4 governance.
*
* Classifies each `SupportedLanguages` member for the rollout. Ring 4 (DAG
* retirement) is gated on *all production languages* being registry-primary
* and stable for one release cycle; `experimental` and `quarantined`
* languages do not block.
*
* Initial classification (locked in Ring 1 #910):
* - production: javascript, typescript, python, java, c, cpp, csharp, go,
* ruby, rust, php, kotlin, swift, dart
* - experimental: vue (embedded-language / SFC complexity),
* cobol (regex-provider path)
* - quarantined: (none)
*/
import { SupportedLanguages } from '../languages.js';
export type LanguageClassification = 'production' | 'experimental' | 'quarantined';
/**
* The canonical classification for each supported language. Governance
* changes (promote `experimental` → `production`, quarantine a language, …)
* update this map in a dedicated PR.
*/
export const LanguageClassifications: Readonly<Record<SupportedLanguages, LanguageClassification>> =
{
[SupportedLanguages.JavaScript]: 'production',
[SupportedLanguages.TypeScript]: 'production',
[SupportedLanguages.Python]: 'production',
[SupportedLanguages.Java]: 'production',
[SupportedLanguages.C]: 'production',
[SupportedLanguages.CPlusPlus]: 'production',
[SupportedLanguages.CSharp]: 'production',
[SupportedLanguages.Go]: 'production',
[SupportedLanguages.Ruby]: 'production',
[SupportedLanguages.Rust]: 'production',
[SupportedLanguages.PHP]: 'production',
[SupportedLanguages.Kotlin]: 'production',
[SupportedLanguages.Swift]: 'production',
[SupportedLanguages.Dart]: 'production',
[SupportedLanguages.Vue]: 'experimental',
[SupportedLanguages.Cobol]: 'experimental',
};
/** Convenience predicate: is this language gating Ring 4 retirement? */
export function isProductionLanguage(lang: SupportedLanguages): boolean {
return LanguageClassifications[lang] === 'production';
}
@@ -0,0 +1,145 @@
/**
* `MethodDispatchIndex` — materialized view of class hierarchies keyed by
* `DefId` (RFC §3.1; Ring 2 SHARED #914).
*
* Two O(1)-access maps used by `Registry.lookupMethod` and interface-
* dispatch callers:
*
* - `mroByOwnerDefId` : owner class → full MRO ancestor chain
* (excludes the owner itself, in per-language
* strategy order).
* - `implsByInterfaceDefId` : interface/trait → classes that implement it.
*
* **Not an MRO implementation.** The build function is a pure aggregator: it
* asks the caller (via `computeMro` and `implementsOf` callbacks) for the
* per-language answers and materializes the two-way index. MRO strategies
* live where they already do today (`model/resolve.ts § c3Linearize`,
* `languages/ruby.ts § selectDispatch`, etc.) — this index does not
* reimplement them.
*
* Why callbacks and not a shared strategy registry: the five strategies
* (Python C3, Ruby kind-aware, Java/Kotlin linear, Rust qualified-syntax,
* COBOL none) already exist in the CLI package and depend on the CLI's
* `HeritageMap` + `SemanticModel`. Pulling them into `gitnexus-shared` would
* require migrating both — out of scope for #914. Callbacks let the shared
* build stay pure while honoring existing strategies verbatim.
*
* Consumed by: #917 (`Registry.lookupMethod` MRO fast path, interface
* dispatch resolver).
*/
import type { DefId } from './types.js';
// ─── Public contracts ───────────────────────────────────────────────────────
export interface MethodDispatchIndex {
/**
* Full MRO ancestor chain per owner class (excludes the owner itself).
* Order reflects the per-language strategy used by `computeMro`.
*/
readonly mroByOwnerDefId: ReadonlyMap<DefId, readonly DefId[]>;
/** Interfaces / traits → classes that implement them. */
readonly implsByInterfaceDefId: ReadonlyMap<DefId, readonly DefId[]>;
/** `mroByOwnerDefId.get`, with an empty frozen array on miss. */
mroFor(ownerDefId: DefId): readonly DefId[];
/** `implsByInterfaceDefId.get`, with an empty frozen array on miss. */
implementorsOf(interfaceDefId: DefId): readonly DefId[];
}
export interface MethodDispatchInput {
/**
* Owner defs to index (classes, structs, traits, interfaces — any kind
* that can appear on the owner side of a method-dispatch graph).
*/
readonly owners: readonly DefId[];
/**
* Return the full MRO ancestor chain for `ownerDefId`, **excluding the
* owner itself**, in the order dictated by the owner's language-specific
* MRO strategy.
*
* Contract:
* - Pure (no side effects).
* - Deterministic per input.
* - `undefined` not allowed — return `[]` when the owner has no parents.
*/
readonly computeMro: (ownerDefId: DefId) => readonly DefId[];
/**
* Return the set of interface/trait defs that `ownerDefId` implements.
* Transitive inclusion (e.g., `implements` on a parent class) is the
* caller's choice — the build function simply inverts whatever is
* returned.
*
* Repeated IDs in the output are deduplicated automatically.
*
* **Call-count contract.** `implementsOf` is invoked **once per
* occurrence** of an owner in `input.owners`, not once per unique
* owner. Duplicate owners therefore re-invoke it; dedup happens at
* the bucket layer (after the callback returns). Callers with
* expensive `implementsOf` implementations should pass a deduplicated
* `owners` list. `computeMro`, by contrast, is memoized by the first-
* write-wins policy and fires at most once per unique owner.
*/
readonly implementsOf: (ownerDefId: DefId) => readonly DefId[];
}
// ─── Builder ────────────────────────────────────────────────────────────────
export function buildMethodDispatchIndex(input: MethodDispatchInput): MethodDispatchIndex {
const mroByOwnerDefId = new Map<DefId, readonly DefId[]>();
const implsBuilding = new Map<DefId, DefId[]>();
const implsSeen = new Map<DefId, Set<DefId>>();
for (const ownerId of input.owners) {
// First-write-wins on duplicate owner ids: a stable policy consistent
// with sibling indexes (#913 DefIndex / ModuleScopeIndex).
if (!mroByOwnerDefId.has(ownerId)) {
const chain = input.computeMro(ownerId);
mroByOwnerDefId.set(ownerId, Object.freeze(chain.slice()));
}
for (const ifaceId of input.implementsOf(ownerId)) {
let seen = implsSeen.get(ifaceId);
if (seen === undefined) {
seen = new Set<DefId>();
implsSeen.set(ifaceId, seen);
}
if (seen.has(ownerId)) continue;
seen.add(ownerId);
let bucket = implsBuilding.get(ifaceId);
if (bucket === undefined) {
bucket = [];
implsBuilding.set(ifaceId, bucket);
}
bucket.push(ownerId);
}
}
const implsByInterfaceDefId = new Map<DefId, readonly DefId[]>();
for (const [ifaceId, owners] of implsBuilding) {
implsByInterfaceDefId.set(ifaceId, Object.freeze(owners.slice()));
}
return wrapIndex(mroByOwnerDefId, implsByInterfaceDefId);
}
// ─── Internal ───────────────────────────────────────────────────────────────
const EMPTY: readonly DefId[] = Object.freeze([]);
function wrapIndex(
mroByOwnerDefId: Map<DefId, readonly DefId[]>,
implsByInterfaceDefId: Map<DefId, readonly DefId[]>,
): MethodDispatchIndex {
return {
mroByOwnerDefId,
implsByInterfaceDefId,
mroFor(ownerDefId: DefId): readonly DefId[] {
return mroByOwnerDefId.get(ownerDefId) ?? EMPTY;
},
implementorsOf(interfaceDefId: DefId): readonly DefId[] {
return implsByInterfaceDefId.get(interfaceDefId) ?? EMPTY;
},
};
}
@@ -0,0 +1,73 @@
/**
* `ModuleScopeIndex` — O(1) `filePath → moduleScopeId` lookup.
*
* Every file parsed produces exactly one `Module` scope at its root. The
* finalize algorithm needs to resolve `ImportEdge.targetFile` to a concrete
* module scope id in constant time during the link pass; this index is that
* mapping.
*
* Part of RFC #909 Ring 2 SHARED — #913.
*
* Consumed by: #915 (SCC finalize link pass), #923 (shadow harness when
* resolving callsite file → enclosing module).
*/
import type { ScopeId } from './types.js';
export interface ModuleScopeIndex {
readonly byFilePath: ReadonlyMap<string, ScopeId>;
readonly size: number;
get(filePath: string): ScopeId | undefined;
has(filePath: string): boolean;
}
export interface ModuleScopeEntry {
readonly filePath: string;
readonly moduleScopeId: ScopeId;
}
/**
* Build a `ModuleScopeIndex` from a flat list of `{ filePath, moduleScopeId }`
* pairs.
*
* **Collision policy: first-write-wins.** A file should appear exactly once
* in a single ingestion run; collisions indicate the same file was parsed
* twice or a `filePath` normalization bug upstream. Dropping the later
* entry preserves the first-stable id the rest of the pipeline may already
* have registered against.
*
* **Caller contract: filePath keys must be pre-normalized.** This index
* keys on the raw `filePath` string and does NOT canonicalize separators,
* case, or trailing slashes. Callers upstream of this function must agree
* on a canonical form (typically repo-root-relative, POSIX separators,
* no trailing slash) before constructing entries — otherwise `C:\foo\bar.ts`,
* `C:/foo/bar.ts`, and `foo/bar.ts` will all hash to distinct buckets and
* `get()` will miss.
*
* Pure function — safe to call repeatedly; no side effects.
*/
export function buildModuleScopeIndex(entries: readonly ModuleScopeEntry[]): ModuleScopeIndex {
const byFilePath = new Map<string, ScopeId>();
for (const { filePath, moduleScopeId } of entries) {
if (byFilePath.has(filePath)) continue; // first-write-wins
byFilePath.set(filePath, moduleScopeId);
}
return wrapIndex(byFilePath);
}
// ─── Internal ───────────────────────────────────────────────────────────────
function wrapIndex(byFilePath: Map<string, ScopeId>): ModuleScopeIndex {
return {
byFilePath,
get size() {
return byFilePath.size;
},
get(filePath: string): ScopeId | undefined {
return byFilePath.get(filePath);
},
has(filePath: string): boolean {
return byFilePath.has(filePath);
},
};
}
@@ -0,0 +1,30 @@
/**
* `ORIGIN_PRIORITY` — RFC Appendix B (authoritative values).
*
* Tie-break ordering applied inside `Registry.lookup` Step 7 when
* `|Δconfidence| < 0.001` between two `Resolution` candidates. Lower number
* = stronger (wins the tie).
*
* Full tie-break order (§4.2 Step 7):
* confidence DESC → scope depth ASC → MRO depth ASC → ORIGIN_PRIORITY ASC
* → DefId.localeCompare
*/
export type OriginForTieBreak =
| 'local'
| 'import'
| 'reexport'
| 'namespace'
| 'wildcard'
| 'global-qualified'
| 'global-name';
export const ORIGIN_PRIORITY: Readonly<Record<OriginForTieBreak, number>> = {
local: 0,
import: 1,
reexport: 2,
namespace: 3,
wildcard: 4,
'global-qualified': 5,
'global-name': 6,
};
@@ -0,0 +1,65 @@
/**
* `ParsedFile` — the per-file artifact produced by `ScopeExtractor`
* (RFC §3.2 Phase 1; Ring 2 PKG #919).
*
* The boundary between Phase 1 (extraction, per-file, parallelizable) and
* Phase 2 (finalize, cross-file). One `ParsedFile` is emitted per source
* file; the finalize orchestrator (#921) collects them into a workspace-
* wide set and feeds them to the shared `finalize` algorithm (#915).
*
* ## Shape
*
* - `scopes` — every `Scope` created for this file, in tree-
* topological order (module first, then children).
* `Scope.bindings` carry **local-only** bindings at
* this stage; finalize merges imports/wildcards on top.
* - `parsedImports` — raw `ParsedImport[]` for this file; finalize
* resolves each to a concrete `ImportEdge`.
* - `localDefs` — defs structurally declared in this file. A
* superset of every `Scope.ownedDefs` union.
* Listed separately so `finalize` can dedup-index
* without re-walking scopes.
* - `referenceSites` — pre-resolution usage facts; populated by the
* resolution phase into `ReferenceIndex`.
*
* ## What `ParsedFile` deliberately does NOT carry
*
* - Linked `ImportEdge`s. Those are finalize output.
* - A `ScopeTree` instance. Callers build one from `scopes` (cheap —
* `buildScopeTree(parsedFile.scopes)`). Keeping the ParsedFile flat
* makes IPC serialization from worker threads straightforward.
* - Merged module-scope bindings. Finalize owns that materialization.
*
* ## Compatibility with `FinalizeFile`
*
* `FinalizeFile` (defined in `./finalize-algorithm.ts`) is a structural
* subset of `ParsedFile` — `filePath`, `moduleScope`, `parsedImports`,
* `localDefs`. A `ParsedFile` is trivially convertible to a `FinalizeFile`
* by picking those four fields, so the finalize orchestrator threads
* ParsedFile through to the shared algorithm without shape-shifting.
*/
import type { Scope, ScopeId } from './types.js';
import type { ParsedImport } from './types.js';
import type { SymbolDefinition } from './symbol-definition.js';
import type { ReferenceSite } from './reference-site.js';
export interface ParsedFile {
readonly filePath: string;
/** `Scope.id` of the file's root `Module` scope. */
readonly moduleScope: ScopeId;
/**
* All scopes in this file, typically emitted in tree-topological order.
* Caller reconstructs a `ScopeTree` via `buildScopeTree(scopes)` when
* navigation or invariant re-validation is needed.
*/
readonly scopes: readonly Scope[];
readonly parsedImports: readonly ParsedImport[];
/**
* All defs structurally declared in this file (classes, methods, fields,
* variables). Mirrors the union of `Scope.ownedDefs` across `scopes`,
* pre-flattened for O(N) consumption by finalize.
*/
readonly localDefs: readonly SymbolDefinition[];
readonly referenceSites: readonly ReferenceSite[];
}
@@ -0,0 +1,166 @@
/**
* `PositionIndex` — O(log N_file) scope-at-position lookup
* (RFC §3.1; Ring 2 SHARED #912).
*
* Per-file sorted array of `(range, scopeId)` entries, sorted by start
* position ASC (`startLine`, then `startCol`). `atPosition(filePath, line,
* col)` binary-searches for the last entry whose start ≤ (line, col), then
* scans backward through the sorted prefix and returns the first entry
* whose range contains the query position.
*
* **Why this works.** `ScopeTree`'s invariants (parent strictly contains
* child; siblings don't overlap) guarantee that the scopes containing a
* given point form an **ancestor chain**. When scanning backward through
* entries sorted by start position ASC, the first scope we find that
* contains the query is the innermost one — any deeper-starting scope
* that also contained the query would appear *later* in the sorted array,
* but we're only scanning entries with start ≤ query, so anything later
* necessarily starts after the query and can't contain it.
*
* Expected complexity: `O(log N_file + D)` where `D` is the lexical depth
* at the query position (typically ≤ 10). Worst-case degrades to `O(N_file)`
* only under pathological inputs (many scopes starting at the same line).
*
* **Line/column conventions.** Matches `Range` in `types.ts`: lines are
* 1-based, columns are 0-based. Ranges are **inclusive on both ends** —
* a scope whose `endLine:endCol` equals the query position still contains
* it. That matches how tree-sitter captures bodies (closing brace
* included) and how closed PR #902's `enclosingFunctions` behaved.
*/
import type { Range, Scope, ScopeId } from './types.js';
export interface PositionIndex {
/** Total scope entries indexed across all files. */
readonly size: number;
/**
* Innermost scope containing `(line, col)` in `filePath`, or `undefined`
* when nothing contains it (position before file start, after file end,
* or filePath not indexed).
*
* **Touching-boundary semantics.** Ranges are inclusive on both ends.
* When two sibling scopes share a boundary point — e.g.
* `[5:0, 10:0]` and `[10:0, 15:0]`, which is legal under `ScopeTree`'s
* non-overlap invariant — a query at the shared point `(10, 0)` is
* contained by **both**. The innermost-wins tie-break rule applies as
* usual: since neither is nested inside the other, the one that
* **starts latest** wins, i.e. the **right** sibling. The mechanism
* is the backward scan through the start-position-sorted array (see
* `findLastStartLteIndex` below) — both siblings land before the
* upper-bound cursor, and the right sibling is scanned first. Queries at non-boundary positions between them naturally
* fall to the unique containing scope.
*/
atPosition(filePath: string, line: number, col: number): ScopeId | undefined;
}
/**
* Build a `PositionIndex` from a flat list of `Scope` records.
*
* Duplicate `id`s are tolerated and deduplicated — the caller's
* `ScopeTree.buildScopeTree` is the authoritative validator of scope
* identity, and the position index does not need to re-check that
* invariant.
*/
export function buildPositionIndex(scopes: readonly Scope[]): PositionIndex {
const entriesByFile = new Map<string, Entry[]>();
const seen = new Set<ScopeId>();
for (const scope of scopes) {
if (seen.has(scope.id)) continue;
seen.add(scope.id);
let bucket = entriesByFile.get(scope.filePath);
if (bucket === undefined) {
bucket = [];
entriesByFile.set(scope.filePath, bucket);
}
bucket.push({ id: scope.id, range: scope.range });
}
for (const bucket of entriesByFile.values()) {
bucket.sort(compareEntry);
}
return wrapIndex(entriesByFile, seen.size);
}
// ─── Internals ──────────────────────────────────────────────────────────────
interface Entry {
readonly id: ScopeId;
readonly range: Range;
}
/**
* Sort by start position ASC, breaking ties by end position DESC so that
* larger (outer) scopes appear before their smaller (inner) co-starting
* siblings in the array. Makes the backward-scan contract crisp: the
* first containing hit from the end of the scanned prefix is the
* innermost scope.
*/
function compareEntry(a: Entry, b: Entry): number {
if (a.range.startLine !== b.range.startLine) return a.range.startLine - b.range.startLine;
if (a.range.startCol !== b.range.startCol) return a.range.startCol - b.range.startCol;
if (a.range.endLine !== b.range.endLine) return b.range.endLine - a.range.endLine;
return b.range.endCol - a.range.endCol;
}
/** Whether `(line, col)` is at or after `range`'s start. */
function startIsAtOrBefore(range: Range, line: number, col: number): boolean {
if (range.startLine < line) return true;
if (range.startLine > line) return false;
return range.startCol <= col;
}
/** Whether `(line, col)` is at or before `range`'s end (inclusive). */
function endIsAtOrAfter(range: Range, line: number, col: number): boolean {
if (range.endLine > line) return true;
if (range.endLine < line) return false;
return range.endCol >= col;
}
/**
* Return the largest index `i` in `arr` where `arr[i].range` starts at or
* before `(line, col)`. Returns `-1` if no entry starts ≤ the query.
*
* Classic "upper bound - 1" binary search: find the first entry that
* starts *after* the query, then step back one.
*/
function findLastStartLteIndex(arr: readonly Entry[], line: number, col: number): number {
let lo = 0;
let hi = arr.length;
while (lo < hi) {
const mid = (lo + hi) >>> 1;
if (startIsAtOrBefore(arr[mid]!.range, line, col)) {
lo = mid + 1;
} else {
hi = mid;
}
}
return lo - 1;
}
function wrapIndex(entriesByFile: Map<string, Entry[]>, size: number): PositionIndex {
return {
get size() {
return size;
},
atPosition(filePath: string, line: number, col: number): ScopeId | undefined {
const bucket = entriesByFile.get(filePath);
if (bucket === undefined || bucket.length === 0) return undefined;
const endIdx = findLastStartLteIndex(bucket, line, col);
if (endIdx < 0) return undefined;
// Scan backward; first containing hit is innermost (see file header).
for (let i = endIdx; i >= 0; i--) {
const entry = bucket[i]!;
if (endIsAtOrAfter(entry.range, line, col)) {
// `startIsAtOrBefore` is guaranteed true by the binary search.
return entry.id;
}
}
return undefined;
},
};
}
@@ -0,0 +1,92 @@
/**
* `QualifiedNameIndex` — O(1) `qualifiedName → DefId[]` lookup across all kinds.
*
* Cross-kind fast path for qualified-name resolution
* (`lookupQualified(qname, scope, params)` in RFC §4.5). Class, method,
* field, and namespace defs all contribute to a single index here; consumers
* filter the returned `DefId[]` by `p.acceptedKinds` at the call site.
*
* Returns `DefId[]` (not a single `DefId`) because multiple defs can legally
* share a qualified name — partial classes in C#, method overloads, or
* accidental cross-kind collisions. The lookup caller filters to the expected
* kind(s) and ranks the survivors.
*
* Part of RFC #909 Ring 2 SHARED — #913.
*
* Consumed by: #917 (`Registry.lookup` qualified fast path, `resolveTypeRef`
* dotted fallback via #916).
*/
import type { SymbolDefinition } from './symbol-definition.js';
import type { DefId } from './types.js';
export interface QualifiedNameIndex {
readonly byQualifiedName: ReadonlyMap<string, readonly DefId[]>;
readonly size: number;
/** Returns all `DefId`s registered under this qualified name; empty frozen
* array on miss so callers can iterate without null checks. */
get(qualifiedName: string): readonly DefId[];
has(qualifiedName: string): boolean;
}
/**
* Build a `QualifiedNameIndex` from a flat list of `SymbolDefinition` records.
*
* Only defs with a non-empty `qualifiedName` contribute; defs without one are
* silently skipped (not every kind carries a qualified name — anonymous or
* top-level symbols, dynamic-unresolved imports, etc.).
*
* **Duplicate policy: appended in input order.** Each unique `(qname, DefId)`
* pair contributes at most once — repeated entries for the same pair are
* deduplicated. Distinct `DefId`s sharing a `qname` accumulate in insertion
* order (stable output for deterministic lookup ranking at the call site).
*
* Pure function — safe to call repeatedly; no side effects.
*/
export function buildQualifiedNameIndex(defs: readonly SymbolDefinition[]): QualifiedNameIndex {
const byQualifiedName = new Map<string, DefId[]>();
const seenPairs = new Set<string>();
for (const def of defs) {
const qname = def.qualifiedName;
if (qname === undefined || qname.length === 0) continue;
const pairKey = `${qname}\0${def.nodeId}`;
if (seenPairs.has(pairKey)) continue;
seenPairs.add(pairKey);
const bucket = byQualifiedName.get(qname);
if (bucket === undefined) {
byQualifiedName.set(qname, [def.nodeId]);
} else {
bucket.push(def.nodeId);
}
}
// Freeze bucket arrays so consumers can't mutate the index.
const frozen = new Map<string, readonly DefId[]>();
for (const [k, v] of byQualifiedName) {
frozen.set(k, Object.freeze(v.slice()));
}
return wrapIndex(frozen);
}
// ─── Internal ───────────────────────────────────────────────────────────────
const EMPTY: readonly DefId[] = Object.freeze([]);
function wrapIndex(byQualifiedName: Map<string, readonly DefId[]>): QualifiedNameIndex {
return {
byQualifiedName,
get size() {
return byQualifiedName.size;
},
get(qualifiedName: string): readonly DefId[] {
return byQualifiedName.get(qualifiedName) ?? EMPTY;
},
has(qualifiedName: string): boolean {
return byQualifiedName.has(qualifiedName);
},
};
}
@@ -0,0 +1,74 @@
/**
* `ReferenceSite` — a pre-resolution usage fact collected by `ScopeExtractor`
* (RFC §3.2 Phase 1; Ring 2 PKG #919).
*
* One record per `@reference.*` capture. The extractor records:
* - the name being referenced (method/field/class name),
* - the source range,
* - the innermost lexical scope containing the reference,
* - the reference kind (call, read, write, inherits, etc.),
* - optional call-form classification from `provider.classifyCallForm`,
* - optional explicit-receiver hint for dotted calls (`user.save()`),
* - optional arity for call sites.
*
* Reference sites are consumed by the resolution phase (RFC §3.2 Phase 4)
* which routes each through `Registry.lookup` / `resolveTypeRef` and
* emits the final `Reference` record into `ReferenceIndex`.
*
* **Pre-resolution only.** `ReferenceSite` intentionally carries no
* `toDef`, `confidence`, or `evidence`. Those are populated by the
* resolution step that reads this record and produces a `Reference`
* (defined in `./types.ts`).
*/
import type { Range, ScopeId } from './types.js';
/**
* What kind of usage this reference represents — the graph-edge kind
* emitted after resolution (`CALLS`, `READS`, `WRITES`, etc.).
*
* Matches the `kind` field on `Reference` in `./types.ts` so the
* resolution phase can pass it through without re-classification.
*/
export type ReferenceKind =
| 'call'
| 'read'
| 'write'
| 'type-reference'
| 'inherits'
| 'import-use';
/**
* How a call site binds its target. Informs `Registry.lookup` Step 2
* (type-binding path):
* - `'free'` — bare call (no receiver); resolution via lexical chain.
* - `'member'` — dotted call (`x.foo()`); resolution via receiver type.
* - `'constructor'` — `new Foo()`; receiver is the class itself.
* - `'index'` — index expression (`arr[0]`); rare as a dispatch site.
*
* Only meaningful for `kind === 'call'`; ignored for reads/writes.
*/
export type CallForm = 'free' | 'member' | 'constructor' | 'index';
export interface ReferenceSite {
/** The name being referenced (e.g., `'save'`, `'User'`, `'count'`). */
readonly name: string;
/** Source-text range of this reference. */
readonly atRange: Range;
/**
* Innermost lexical scope that contains `atRange`. Resolved by the
* extractor via position lookup and frozen here so the resolution
* phase doesn't re-compute it per call.
*/
readonly inScope: ScopeId;
readonly kind: ReferenceKind;
/** Set when `kind === 'call'`. */
readonly callForm?: CallForm;
/**
* Explicit receiver for dotted calls (`user.save()` → `{ name: 'user' }`).
* Passed through to `Registry.lookup.explicitReceiver`.
*/
readonly explicitReceiver?: { readonly name: string };
/** Argument count at the call site; used by `provider.arityCompatibility`. */
readonly arity?: number;
}
@@ -0,0 +1,41 @@
/**
* `ClassRegistry` — scope-aware lookup for class-like symbols
* (RFC §4.4; Ring 2 SHARED #917).
*
* Thin wrapper over `lookupCore`, specialized for class kinds:
*
* - `acceptedKinds` = Class / Interface / Enum / Struct / Union /
* Trait / TypeAlias / Typedef / Record / Delegate / Annotation /
* Template / Namespace.
* - `useReceiverTypeBinding` is **false** — classes are resolved by
* name through the lexical chain + global qualified fallback, not
* via a receiver type.
* - Arity filter is not applicable (classes are not called with
* argument counts at lookup time).
*/
import type { Resolution, ScopeId } from '../types.js';
import { lookupCore, type CoreLookupParams } from './lookup-core.js';
import { CLASS_KINDS, type RegistryContext } from './context.js';
export interface ClassRegistry {
/**
* Look up a class-like symbol by simple or dotted name anchored at
* `scope`. Returns a confidence-ranked `Resolution[]`; consume `[0]`
* for the best answer.
*/
lookup(name: string, scope: ScopeId): readonly Resolution[];
}
export function buildClassRegistry(ctx: RegistryContext): ClassRegistry {
const params: CoreLookupParams = {
acceptedKinds: CLASS_KINDS,
useReceiverTypeBinding: false,
ownerScopedContributor: null,
};
return {
lookup(name: string, scope: ScopeId) {
return lookupCore(name, scope, params, ctx);
},
};
}
@@ -0,0 +1,110 @@
/**
* `RegistryContext` — the injected state required by the scope-aware
* registry lookups (RFC §4; Ring 2 SHARED #917).
*
* Bundles every Ring 2 index + every provider hook the 7-step algorithm
* might consult. Threaded through `lookupCore` and the three public
* registries unchanged; construction is the caller's responsibility
* (typically once per workspace-indexing pass in Ring 2 PKG).
*
* The design intent is **pure-logic in `gitnexus-shared`, data + hooks
* supplied by the caller**. Nothing here loads files, parses AST, or
* reaches into the CLI package.
*/
import type { NodeLabel } from '../../graph/types.js';
import type { SymbolDefinition } from '../symbol-definition.js';
import type { Callsite, DefId } from '../types.js';
import type { DefIndex } from '../def-index.js';
import type { QualifiedNameIndex } from '../qualified-name-index.js';
import type { ModuleScopeIndex } from '../module-scope-index.js';
import type { ScopeTree } from '../scope-tree.js';
import type { MethodDispatchIndex } from '../method-dispatch-index.js';
// ─── Provider hooks consumed by the registries ─────────────────────────────
export interface RegistryProviders {
/**
* Language-specific arity compatibility between a callsite and a candidate
* `def`. Mirrors `LanguageProvider.arityCompatibility` from #911. Optional:
* when absent, every candidate receives `'unknown'` (neutral signal).
*/
arityCompatibility?(callsite: Callsite, def: SymbolDefinition): ArityVerdict;
}
export type ArityVerdict = 'compatible' | 'unknown' | 'incompatible';
// ─── Owner-scoped contributor (concrete shape for `RegistryContributor`) ────
/**
* Per-owner membership view plugged into `LookupParams.ownerScopedContributor`.
*
* When the caller knows a receiver is of type `Owner` (e.g., after
* resolving an explicit receiver or via `self`), it can supply the
* `Owner`'s own member bucket here. `lookupCore` treats hits from this
* contributor as `origin: 'local'` inside the owner's body scope —
* strongest-visibility evidence, unaffected by the scope-chain hop
* deduction that punishes outer-scope hits.
*
* Ring 1's `RegistryContributor = unknown` opaque placeholder is narrowed
* to this concrete shape here in Ring 2 SHARED (#917).
*/
export interface OwnerScopedContributor {
/** The owner (class/struct/trait/interface) that bounds this view. */
readonly ownerDefId: DefId;
/**
* Methods / fields directly declared on the owner, keyed by simple name.
* Return empty array on miss; implementations should NOT walk the MRO —
* that's `MethodDispatchIndex`'s job, handled in the type-binding step.
*/
byName(name: string): readonly SymbolDefinition[];
}
// ─── Top-level context threaded through every lookup ───────────────────────
export interface RegistryContext {
readonly scopes: ScopeTree;
readonly defs: DefIndex;
readonly qualifiedNames: QualifiedNameIndex;
readonly moduleScopes: ModuleScopeIndex;
/**
* Method-dispatch index; required for method/field registries that
* honor `useReceiverTypeBinding`. Omit for class-only lookups.
*/
readonly methodDispatch?: MethodDispatchIndex;
readonly providers: RegistryProviders;
}
// ─── Per-kind default `acceptedKinds` sets ─────────────────────────────────
//
// Exported so the three public registries stay declarative (each one just
// points at the right constant + passes it to `lookupCore`).
export const CLASS_KINDS: readonly NodeLabel[] = Object.freeze([
'Class',
'Interface',
'Enum',
'Struct',
'Union',
'Trait',
'TypeAlias',
'Typedef',
'Record',
'Delegate',
'Annotation',
'Template',
'Namespace',
]);
export const METHOD_KINDS: readonly NodeLabel[] = Object.freeze([
'Method',
'Function',
'Constructor',
]);
export const FIELD_KINDS: readonly NodeLabel[] = Object.freeze([
'Variable',
'Property',
'Const',
'Static',
]);
@@ -0,0 +1,196 @@
/**
* `composeEvidence` — translate accumulated raw signals per candidate
* into a `ResolutionEvidence[]` using the authoritative `EvidenceWeights`
* map (RFC §4.3 + Appendix A; Ring 2 SHARED #917).
*
* Each `RawSignals` record describes what was observed about a candidate
* during the 7-step walk: where it was found, at what depth, whether
* anything corroborates it. This module turns those raw facts into the
* typed evidence list attached to the outgoing `Resolution`.
*
* **Every weight comes from `EvidenceWeights`.** No inline magic numbers.
* Extends issue #429 (centralize hardcoded confidence values).
*
* **Confidence compose rule.** Signals add; the sum is capped at 1.0 at
* the call site (inside `lookupCore`). This module only emits the list;
* it does NOT compute the capped sum so callers can inspect per-signal
* contributions for debugging.
*/
import type { BindingRef, ResolutionEvidence } from '../types.js';
import { EvidenceWeights, typeBindingWeightAtDepth } from '../evidence-weights.js';
/**
* Raw signals observed for a single candidate during the 7-step walk.
* Optional fields encode "this signal did not fire"; presence encodes
* "emit an evidence record".
*/
export interface RawSignals {
// ── Where-found ────────────────────────────────────────────────────────
/** Visibility origin of the binding that produced this candidate. */
readonly origin?: BindingRef['origin'] | 'global-qualified' | 'global-name';
/** Depth at which the binding was found (hops up from start scope). */
readonly scopeChainDepth?: number;
/** `ImportEdge` that brought the name in; present when origin is a non-local. */
readonly viaUnlinkedImport?: boolean;
// ── Type-binding path ──────────────────────────────────────────────────
/** Set when the candidate came via the receiver's type-binding MRO walk. */
readonly typeBindingMroDepth?: number;
// ── Corroborators ──────────────────────────────────────────────────────
/** `def.ownerId === resolvedReceiver.def.nodeId`. */
readonly ownerMatch?: boolean;
/** Always fires for candidates that pass `acceptedKinds`; weight 0. */
readonly kindMatch: true;
// ── Arity ──────────────────────────────────────────────────────────────
readonly arityVerdict?: 'compatible' | 'unknown' | 'incompatible';
// ── Dynamic-unresolved passthrough ─────────────────────────────────────
/** Candidate flows through a `kind: 'dynamic-unresolved'` ImportEdge. */
readonly dynamicUnresolved?: boolean;
}
/**
* Compose the raw signals into a stable `ResolutionEvidence[]` list.
*
* Emission order mirrors the `EvidenceWeights` layout: where-found →
* type-binding → corroborators → arity → degraded. Stable order makes
* the per-signal contributions easy to reason about in tests and in the
* shadow-mode parity dashboard.
*/
export function composeEvidence(signals: RawSignals): readonly ResolutionEvidence[] {
const out: ResolutionEvidence[] = [];
// ── Where-found visibility ─────────────────────────────────────────────
if (signals.origin !== undefined) {
const baseWeight = getOriginWeight(signals.origin);
const capped = signals.viaUnlinkedImport
? baseWeight * EvidenceWeights.unlinkedImportMultiplier
: baseWeight;
const evidenceKind = whereFoundEvidenceKind(signals.origin);
out.push({
kind: evidenceKind,
weight: capped,
...(signals.viaUnlinkedImport
? { note: `via unresolved import (${EvidenceWeights.unlinkedImportMultiplier}× cap)` }
: {}),
});
}
// ── Scope-chain depth deduction (per-hop, only meaningful for lexical
// hits where scopeChainDepth ≥ 1). Depth 0 = no deduction; depth N ≥ 1
// emits a single `scope-chain` evidence with the accumulated penalty.
if (signals.scopeChainDepth !== undefined && signals.scopeChainDepth > 0) {
out.push({
kind: 'scope-chain',
weight: EvidenceWeights.scopeChainPerDepth * signals.scopeChainDepth,
note: `depth=${signals.scopeChainDepth}`,
});
}
// ── Type-binding / MRO path ────────────────────────────────────────────
if (signals.typeBindingMroDepth !== undefined) {
out.push({
kind: 'type-binding',
weight: typeBindingWeightAtDepth(signals.typeBindingMroDepth),
note: `mroDepth=${signals.typeBindingMroDepth}`,
});
}
// ── Owner match (explanatory for debug) ────────────────────────────────
if (signals.ownerMatch === true) {
out.push({
kind: 'owner-match',
weight: EvidenceWeights.ownerMatch,
});
}
// ── Kind match (always present; weight 0; retained for debuggability) ──
out.push({
kind: 'kind-match',
weight: EvidenceWeights.kindMatch,
});
// ── Arity ──────────────────────────────────────────────────────────────
if (signals.arityVerdict !== undefined) {
const weight =
signals.arityVerdict === 'compatible'
? EvidenceWeights.arityMatchCompatible
: signals.arityVerdict === 'incompatible'
? EvidenceWeights.arityMatchIncompatible
: EvidenceWeights.arityMatchUnknown;
out.push({
kind: 'arity-match',
weight,
note: signals.arityVerdict,
});
}
// ── Dynamic-unresolved (degraded signal) ───────────────────────────────
if (signals.dynamicUnresolved === true) {
out.push({
kind: 'dynamic-import-unresolved',
weight: EvidenceWeights.dynamicImportUnresolved,
});
}
return out;
}
/**
* Sum evidence weights and clamp to `[0, 1]`. Separate from `composeEvidence`
* so tests and the parity dashboard can inspect the raw evidence list.
*/
export function confidenceFromEvidence(evidence: readonly ResolutionEvidence[]): number {
let sum = 0;
for (const e of evidence) sum += e.weight;
if (sum < 0) return 0;
if (sum > 1) return 1;
return sum;
}
// ─── Internal ───────────────────────────────────────────────────────────────
function getOriginWeight(origin: NonNullable<RawSignals['origin']>): number {
switch (origin) {
case 'local':
return EvidenceWeights.local;
case 'import':
return EvidenceWeights.import;
case 'reexport':
return EvidenceWeights.reexport;
case 'namespace':
return EvidenceWeights.namespace;
case 'wildcard':
return EvidenceWeights.wildcard;
case 'global-qualified':
return EvidenceWeights.globalQualified;
case 'global-name':
// Reserved for Ring 3 byName global index. `lookupCore` today only
// emits `'global-qualified'` (via `lookupQualified`, dotted-name
// fallback); no code path constructs `origin: 'global-name'` yet.
// Kept here so the Appendix A weight stays live and `composeEvidence`
// remains exhaustive over the origin union.
return EvidenceWeights.globalName;
}
}
function whereFoundEvidenceKind(
origin: NonNullable<RawSignals['origin']>,
): ResolutionEvidence['kind'] {
switch (origin) {
case 'local':
return 'local';
case 'import':
case 'reexport':
case 'namespace':
case 'wildcard':
return 'import';
case 'global-qualified':
return 'global-qualified';
case 'global-name':
return 'global-name';
}
}
@@ -0,0 +1,43 @@
/**
* `FieldRegistry` — scope-aware lookup for field / property / variable
* access (RFC §4.4; Ring 2 SHARED #917).
*
* Thin wrapper over `lookupCore`, specialized for data-member kinds:
*
* - `acceptedKinds` = Variable / Property / Const / Static.
* - `useReceiverTypeBinding` is **true** — fields are resolved against
* the receiver type's MRO first, then via the lexical chain for
* free variables.
* - `callsite` is not meaningful for field access (no arity), but the
* `explicitReceiver` and `ownerScopedContributor` knobs are.
*/
import type { Resolution, ScopeId } from '../types.js';
import { lookupCore, type CoreLookupParams } from './lookup-core.js';
import type { OwnerScopedContributor, RegistryContext } from './context.js';
import { FIELD_KINDS } from './context.js';
export interface FieldLookupOptions {
readonly explicitReceiver?: { readonly name: string };
readonly ownerScopedContributor?: OwnerScopedContributor;
}
export interface FieldRegistry {
lookup(name: string, scope: ScopeId, options?: FieldLookupOptions): readonly Resolution[];
}
export function buildFieldRegistry(ctx: RegistryContext): FieldRegistry {
return {
lookup(name: string, scope: ScopeId, options: FieldLookupOptions = {}) {
const params: CoreLookupParams = {
acceptedKinds: FIELD_KINDS,
useReceiverTypeBinding: true,
ownerScopedContributor: options.ownerScopedContributor ?? null,
...(options.explicitReceiver !== undefined
? { explicitReceiver: options.explicitReceiver }
: {}),
};
return lookupCore(name, scope, params, ctx);
},
};
}
@@ -0,0 +1,461 @@
/**
* `lookupCore` — the shared 7-step canonical resolution algorithm
* (RFC §4.2; Ring 2 SHARED #917).
*
* Pure function. Given a name, a starting scope, and per-kind parameters,
* walks lexical scopes + optional type-binding MRO + optional owner
* contributor + global qualified-name fallback, and returns a ranked
* `Resolution[]` with per-candidate evidence.
*
* All three public registries (`ClassRegistry` / `MethodRegistry` /
* `FieldRegistry`) dispatch into this function, differing only in the
* parameters they pass. The CHOICE of which steps fire is expressed
* through `LookupParams`, not through different algorithms per kind.
*
* ## Algorithm (RFC §4.2, verbatim names)
*
* **Step 1 — Lexical scope-chain walk.** From `startScope`, walk
* parent-ward. At each scope, consult `scope.bindings.get(name)`:
* - Filter candidates whose `def.type ∈ acceptedKinds`.
* - For each surviving candidate, record a raw signal with the
* binding's origin + the current scope-chain depth.
* - **Hard shadow.** If `bindings.get(name)` is non-empty (including
* non-kind-matching candidates), stop walking. The name is
* lexically bound here; outer scopes are not consulted.
*
* **Step 2 — Type-binding resolution.** When `useReceiverTypeBinding`
* is true, resolve the receiver's type at `startScope` (from
* `scope.typeBindings`), then walk the MRO via
* `MethodDispatchIndex.mroFor(ownerDefId)`. Membership per owner comes
* through `RegistryContext.methodDispatch` + owner lookups into
* `scope.ownedDefs`; each hit records a raw signal with the owner's
* MRO depth.
*
* **Step 3 — Owner-scoped contributor.** When
* `params.ownerScopedContributor` is present, merge its `byName(name)`
* hits with `origin: 'local'` (they are declared directly on the
* receiver). Distinct from Step 2 — Step 2 walks the MRO; Step 3 only
* looks at the directly-declared owner members.
*
* **Step 4 — Kind filter (emit `kind-match` evidence).** Already
* applied during Steps 1-3; this step just adds a `kind-match` signal
* at weight 0 to every candidate for debuggability (so the evidence
* array is self-describing).
*
* **Step 5 — Arity filter.** Call `providers.arityCompatibility(callsite,
* def)` per surviving candidate. Verdicts: `compatible` / `unknown` /
* `incompatible`. If at least one candidate is `compatible`, drop
* `incompatible` ones. Otherwise keep all (the penalty weight alone
* will rank them lower but they remain in the result).
*
* **Step 6 — Global fallback.** When Steps 1-3 produced **no**
* candidates and the name contains a `.`, consult the
* `QualifiedNameIndex` via `lookupQualified` — see §4.5. The `scope`
* argument is NOT passed here because global lookup is scope-agnostic.
*
* **Step 7 — Rank + tie-break.** Compose evidence, compute confidence
* (sum capped at 1.0), sort by the RFC Appendix B cascade.
*
* ## What this module does NOT do
*
* - No AST reads (pure data in, pure data out).
* - No `gitnexus/` imports.
* - No language switches. Language-specific behavior flows exclusively
* through `providers.*` and the `params` object.
* - No caching. Callers that want memoization can wrap this function.
*/
import type { NodeLabel } from '../../graph/types.js';
import type { SymbolDefinition } from '../symbol-definition.js';
import type {
BindingRef,
Callsite,
DefId,
LookupParams,
Resolution,
Scope,
ScopeId,
} from '../types.js';
import type { OriginForTieBreak } from '../origin-priority.js';
import { composeEvidence, confidenceFromEvidence, type RawSignals } from './evidence.js';
import { compareByConfidenceWithTiebreaks, type TieBreakKey } from './tie-breaks.js';
import { lookupQualified } from './lookup-qualified.js';
import type { ArityVerdict, OwnerScopedContributor, RegistryContext } from './context.js';
// ─── Public entry point ─────────────────────────────────────────────────────
/** Extended `LookupParams` narrowing `ownerScopedContributor` to the concrete shape. */
export interface CoreLookupParams extends Omit<LookupParams, 'ownerScopedContributor'> {
readonly ownerScopedContributor: OwnerScopedContributor | null;
/** Call-site description forwarded to `arityCompatibility`. Optional — for non-call lookups. */
readonly callsite?: Callsite;
}
/**
* Run the 7-step lookup. Returns a non-empty `Resolution[]` when any
* candidate was found; an empty array otherwise. Callers consume `[0]`
* for the best answer and optionally inspect the rest for alternates.
*/
export function lookupCore(
name: string,
startScope: ScopeId,
params: CoreLookupParams,
ctx: RegistryContext,
): readonly Resolution[] {
const acceptedKinds = new Set<NodeLabel>(params.acceptedKinds);
const perCandidate = new Map<DefId, CandidateState>();
// ── Step 1: lexical scope-chain walk ──────────────────────────────────
const lexicalShadowed = walkLexicalChain(name, startScope, acceptedKinds, ctx, perCandidate);
// ── Step 2: type-binding / MRO walk (methods/fields) ──────────────────
if (params.useReceiverTypeBinding && ctx.methodDispatch !== undefined) {
walkReceiverTypeBinding(name, startScope, acceptedKinds, params, ctx, perCandidate);
}
// ── Step 3: owner-scoped contributor ──────────────────────────────────
if (params.ownerScopedContributor !== null) {
seedFromOwnerScopedContributor(
name,
params.ownerScopedContributor,
acceptedKinds,
perCandidate,
);
}
// ── Step 4: kind-match evidence (emitted by composeEvidence directly) ──
// Handled inside `composeEvidence`.
// ── Step 5: arity filter ──────────────────────────────────────────────
if (params.callsite !== undefined) {
applyArityFilter(params.callsite, perCandidate, ctx);
}
// ── Step 6: global fallback (only when Steps 1-3 produced nothing) ──
if (perCandidate.size === 0 && !lexicalShadowed && name.includes('.')) {
const globals = lookupQualified(name, { acceptedKinds: params.acceptedKinds }, ctx);
if (globals.length > 0) return globals;
}
if (perCandidate.size === 0) return EMPTY;
// ── Step 7: compose evidence + rank ──────────────────────────────────
return rankCandidates(perCandidate);
}
// ─── Internal state ────────────────────────────────────────────────────────
interface CandidateState {
readonly def: SymbolDefinition;
readonly signals: MutableRawSignals;
readonly tieBreakKey: MutableTieBreakKey;
}
interface MutableRawSignals {
origin?: BindingRef['origin'] | 'global-qualified' | 'global-name';
scopeChainDepth?: number;
viaUnlinkedImport?: boolean;
typeBindingMroDepth?: number;
ownerMatch?: boolean;
kindMatch: true;
arityVerdict?: ArityVerdict;
dynamicUnresolved?: boolean;
}
interface MutableTieBreakKey {
scopeDepth: number;
mroDepth: number;
origin: OriginForTieBreak;
}
function ensureCandidate(
perCandidate: Map<DefId, CandidateState>,
def: SymbolDefinition,
): CandidateState {
const existing = perCandidate.get(def.nodeId);
if (existing !== undefined) return existing;
const fresh: CandidateState = {
def,
signals: { kindMatch: true },
tieBreakKey: { scopeDepth: 0, mroDepth: 0, origin: 'local' },
};
perCandidate.set(def.nodeId, fresh);
return fresh;
}
// ─── Step 1 implementation ─────────────────────────────────────────────────
/**
* Walk the lexical scope chain from `startScope` upward. Returns `true`
* iff a scope with any `bindings.get(name)` entries was found — the
* caller uses this to decide whether to run the global fallback.
*/
function walkLexicalChain(
name: string,
startScope: ScopeId,
acceptedKinds: ReadonlySet<NodeLabel>,
ctx: RegistryContext,
perCandidate: Map<DefId, CandidateState>,
): boolean {
let currentId: ScopeId | null = startScope;
let depth = 0;
const visited = new Set<ScopeId>();
while (currentId !== null) {
if (visited.has(currentId)) return false;
visited.add(currentId);
const scope: Scope | undefined = ctx.scopes.getScope(currentId);
if (scope === undefined) return false;
const bindings = scope.bindings.get(name);
if (bindings !== undefined && bindings.length > 0) {
for (const binding of bindings) {
if (!acceptedKinds.has(binding.def.type)) continue;
recordLexicalHit(perCandidate, binding, depth);
}
return true; // hard shadow regardless of kind-filter survivorship
}
currentId = scope.parent;
depth++;
}
return false;
}
function recordLexicalHit(
perCandidate: Map<DefId, CandidateState>,
binding: BindingRef,
scopeChainDepth: number,
): void {
const state = ensureCandidate(perCandidate, binding.def);
state.signals.origin = binding.origin;
state.signals.scopeChainDepth = scopeChainDepth;
if (binding.via?.linkStatus === 'unresolved') {
state.signals.viaUnlinkedImport = true;
}
if (binding.via?.kind === 'dynamic-unresolved') {
state.signals.dynamicUnresolved = true;
}
state.tieBreakKey.scopeDepth = scopeChainDepth;
state.tieBreakKey.origin = binding.origin as OriginForTieBreak;
}
// ─── Step 2 implementation ─────────────────────────────────────────────────
function walkReceiverTypeBinding(
name: string,
startScope: ScopeId,
acceptedKinds: ReadonlySet<NodeLabel>,
params: CoreLookupParams,
ctx: RegistryContext,
perCandidate: Map<DefId, CandidateState>,
): void {
const ownerDefId = resolveReceiverOwner(startScope, params, ctx);
if (ownerDefId === undefined) return;
if (ctx.methodDispatch === undefined) return;
const ownerDef = ctx.defs.get(ownerDefId);
if (ownerDef === undefined) return;
// Walk the owner itself at depth 0, then its MRO chain.
const walk: DefId[] = [ownerDefId, ...ctx.methodDispatch.mroFor(ownerDefId)];
for (let mroDepth = 0; mroDepth < walk.length; mroDepth++) {
const currentOwnerId = walk[mroDepth]!;
const members = collectOwnedMembers(currentOwnerId, name, ctx);
for (const def of members) {
if (!acceptedKinds.has(def.type)) continue;
recordTypeBindingHit(perCandidate, def, mroDepth, ownerDefId);
}
}
}
function resolveReceiverOwner(
startScope: ScopeId,
params: CoreLookupParams,
ctx: RegistryContext,
): DefId | undefined {
// Explicit receiver: consult the callsite scope's typeBindings for the
// named receiver; the attached TypeRef identifies the owner. Without a
// ready resolveTypeRef call (that module is separate), we do a direct
// lookup and trust the caller to have populated the binding.
if (params.explicitReceiver !== undefined) {
return lookupReceiverType(startScope, params.explicitReceiver.name, ctx);
}
// Implicit `self` / `this` — the scope's typeBindings should carry it.
for (const implicitName of IMPLICIT_RECEIVERS) {
const owner = lookupReceiverType(startScope, implicitName, ctx);
if (owner !== undefined) return owner;
}
return undefined;
}
const IMPLICIT_RECEIVERS: readonly string[] = Object.freeze(['self', 'this']);
function lookupReceiverType(
startScope: ScopeId,
receiverName: string,
ctx: RegistryContext,
): DefId | undefined {
let currentId: ScopeId | null = startScope;
const visited = new Set<ScopeId>();
while (currentId !== null) {
if (visited.has(currentId)) return undefined;
visited.add(currentId);
const scope = ctx.scopes.getScope(currentId);
if (scope === undefined) return undefined;
const typeRef = scope.typeBindings.get(receiverName);
if (typeRef !== undefined) {
// rawName must resolve to a def via qualifiedNames; if it doesn't, we
// can't claim the receiver type. No fallback — that's what
// `resolveTypeRef` would do, but we keep this path lean and let
// callers pre-resolve if they want the richer semantics.
const candidateIds = ctx.qualifiedNames.get(typeRef.rawName);
if (candidateIds.length === 1) return candidateIds[0];
// Ambiguous (≥ 2) or missing (0) — caller must pre-resolve via
// `resolveTypeRef` (#916) if they want the richer semantics. We
// intentionally do NOT re-implement a simple-name fallback here.
return undefined;
}
currentId = scope.parent;
}
return undefined;
}
function collectOwnedMembers(
ownerDefId: DefId,
memberName: string,
ctx: RegistryContext,
): readonly SymbolDefinition[] {
// An owner's members are defs whose `ownerId === ownerDefId` and whose
// simple name matches `memberName`. We iterate `defs.byId` — O(D) per
// call today. A future by-owner index would make this O(K); tracked as
// a follow-up optimization before Ring 3 flips go production.
const out: SymbolDefinition[] = [];
for (const def of ctx.defs.byId.values()) {
if (def.ownerId !== ownerDefId) continue;
if (simpleNameOf(def) !== memberName) continue;
out.push(def);
}
return out;
}
function simpleNameOf(def: SymbolDefinition): string | undefined {
if (def.qualifiedName === undefined || def.qualifiedName.length === 0) return undefined;
const dot = def.qualifiedName.lastIndexOf('.');
return dot === -1 ? def.qualifiedName : def.qualifiedName.slice(dot + 1);
}
function recordTypeBindingHit(
perCandidate: Map<DefId, CandidateState>,
def: SymbolDefinition,
mroDepth: number,
receiverOwner: DefId,
): void {
const state = ensureCandidate(perCandidate, def);
const existingMroDepth = state.signals.typeBindingMroDepth;
const firstHit = existingMroDepth === undefined;
// Only replace if this hit is shallower (smaller MRO depth). The local
// const lets TS narrow to `number` in the `else` branch so no `!`
// assertion is needed.
if (firstHit || mroDepth < existingMroDepth) {
state.signals.typeBindingMroDepth = mroDepth;
state.tieBreakKey.mroDepth = mroDepth;
}
if (def.ownerId === receiverOwner) {
state.signals.ownerMatch = true;
}
// Pure type-binding candidates (no lexical hit) would otherwise keep the
// `ensureCandidate` default `tieBreakKey.origin === 'local'`, making the
// Appendix B cascade lump them with local-origin candidates. Demote them
// to `'import'` — the strongest non-local origin — only when no earlier
// phase set an origin for this candidate. Lexical hits from Step 1 set
// `signals.origin` before Step 2 runs, so the guard skips them; Step 3
// (`seedFromOwnerScopedContributor`) runs AFTER Step 2 and unconditionally
// overrides `tieBreakKey.origin` back to `'local'` for direct-owner
// members, so any same-def overlap still ends up ranked correctly.
if (firstHit && state.signals.origin === undefined) {
state.tieBreakKey.origin = 'import';
}
}
// ─── Step 3 implementation ─────────────────────────────────────────────────
function seedFromOwnerScopedContributor(
name: string,
contributor: OwnerScopedContributor,
acceptedKinds: ReadonlySet<NodeLabel>,
perCandidate: Map<DefId, CandidateState>,
): void {
for (const def of contributor.byName(name)) {
if (!acceptedKinds.has(def.type)) continue;
const state = ensureCandidate(perCandidate, def);
// Treat the contributor's direct membership as `origin: 'local'` —
// strongest visibility, no scope-chain penalty.
state.signals.origin = 'local';
state.signals.scopeChainDepth = 0;
state.signals.ownerMatch = def.ownerId === contributor.ownerDefId;
state.tieBreakKey.origin = 'local';
}
}
// ─── Step 5 implementation ─────────────────────────────────────────────────
function applyArityFilter(
callsite: Callsite,
perCandidate: Map<DefId, CandidateState>,
ctx: RegistryContext,
): void {
const arityFn = ctx.providers.arityCompatibility;
if (arityFn === undefined) {
// No provider → record 'unknown' for every candidate; keeps signal
// shape uniform for composeEvidence.
for (const state of perCandidate.values()) {
state.signals.arityVerdict = 'unknown';
}
return;
}
let anyCompatible = false;
for (const state of perCandidate.values()) {
const verdict = arityFn(callsite, state.def);
state.signals.arityVerdict = verdict;
if (verdict === 'compatible') anyCompatible = true;
}
if (!anyCompatible) return;
// Filter: when at least one compatible candidate exists, drop incompatibles.
for (const [defId, state] of perCandidate) {
if (state.signals.arityVerdict === 'incompatible') {
perCandidate.delete(defId);
}
}
}
// ─── Step 7 implementation ─────────────────────────────────────────────────
function rankCandidates(perCandidate: Map<DefId, CandidateState>): readonly Resolution[] {
const resolutions: Resolution[] = [];
const tieKeys = new Map<string, TieBreakKey>();
for (const state of perCandidate.values()) {
const evidence = composeEvidence(state.signals as RawSignals);
const confidence = confidenceFromEvidence(evidence);
resolutions.push({ def: state.def, confidence, evidence });
tieKeys.set(state.def.nodeId, { ...state.tieBreakKey });
}
resolutions.sort((a, b) => compareByConfidenceWithTiebreaks(a, b, tieKeys));
return Object.freeze(resolutions);
}
// ─── Constants ──────────────────────────────────────────────────────────────
const EMPTY: readonly Resolution[] = Object.freeze([]);
@@ -0,0 +1,71 @@
/**
* `lookupQualified` — qualified-name fast path (RFC §4.5; Ring 2 SHARED #917).
*
* Consults `QualifiedNameIndex` directly, filters by `acceptedKinds`, and
* returns `Resolution[]` with `origin: 'global-qualified'` evidence. Used by:
*
* - `resolveTypeRef` dotted fallback (#916)
* - `Registry.lookup` Step 6 when no lexical candidate survived
* - Explicit dotted identifiers in Cypher / MCP tools where the caller
* knows the target's canonical qualified name
*
* **Strict + deterministic.** No receiver-type resolution, no scope walk.
* Every surviving candidate gets the same base confidence (from
* `EvidenceWeights.globalQualified`), then the tie-break cascade
* disambiguates.
*/
import type { NodeLabel } from '../../graph/types.js';
import type { Resolution } from '../types.js';
import { composeEvidence, confidenceFromEvidence } from './evidence.js';
import { compareByConfidenceWithTiebreaks, type TieBreakKey } from './tie-breaks.js';
import type { RegistryContext } from './context.js';
export interface LookupQualifiedParams {
readonly acceptedKinds: readonly NodeLabel[];
}
/**
* Look up a canonical qualified name (e.g., `app.models.User`) across all
* defs, filtered by `acceptedKinds`. Returns an empty array when the name
* is not indexed or no candidate matches the kind filter.
*
* Callers consume `[0]` for the strict single-return answer; the remainder
* carries alternate candidates (partial classes, overloads, accidental
* cross-kind hits) ordered by the tie-break cascade.
*/
export function lookupQualified(
qualifiedName: string,
params: LookupQualifiedParams,
ctx: RegistryContext,
): readonly Resolution[] {
const defIds = ctx.qualifiedNames.get(qualifiedName);
if (defIds.length === 0) return EMPTY;
const acceptedKinds = new Set<NodeLabel>(params.acceptedKinds);
const resolutions: Resolution[] = [];
const tieKeys = new Map<string, TieBreakKey>();
for (const defId of defIds) {
const def = ctx.defs.get(defId);
if (def === undefined) continue;
if (!acceptedKinds.has(def.type)) continue;
const evidence = composeEvidence({ origin: 'global-qualified', kindMatch: true });
const confidence = confidenceFromEvidence(evidence);
resolutions.push({ def, confidence, evidence });
tieKeys.set(def.nodeId, {
scopeDepth: 0,
mroDepth: 0,
origin: 'global-qualified',
});
}
if (resolutions.length === 0) return EMPTY;
resolutions.sort((a, b) => compareByConfidenceWithTiebreaks(a, b, tieKeys));
return Object.freeze(resolutions);
}
const EMPTY: readonly Resolution[] = Object.freeze([]);
@@ -0,0 +1,54 @@
/**
* `MethodRegistry` — scope-aware lookup for method / function / constructor
* dispatch (RFC §4.4; Ring 2 SHARED #917).
*
* Thin wrapper over `lookupCore`, specialized for callable kinds:
*
* - `acceptedKinds` = Method / Function / Constructor.
* - `useReceiverTypeBinding` is **true** — the type-binding + MRO walk
* (Step 2) is the primary evidence path for receiver-dispatched calls.
* - `callsite.arity` flows through to `provider.arityCompatibility`
* when provided. When the provider is absent, arity evidence is
* `unknown` (neutral signal).
*/
import type { Callsite, Resolution, ScopeId } from '../types.js';
import { lookupCore, type CoreLookupParams } from './lookup-core.js';
import type { OwnerScopedContributor, RegistryContext } from './context.js';
import { METHOD_KINDS } from './context.js';
/**
* Extra per-call parameters that vary across call sites but NOT across
* registries. Kept as a separate shape so `MethodRegistry.lookup` stays
* concise while still exposing the explicit-receiver + owner-contributor +
* arity knobs the RFC algorithm needs.
*/
export interface MethodLookupOptions {
/** Call-site arity for `provider.arityCompatibility`. */
readonly callsite?: Callsite;
/** Explicit receiver (e.g., `user` in `user.save()`). See §4.1. */
readonly explicitReceiver?: { readonly name: string };
/** Optional per-owner contributor (Step 3). */
readonly ownerScopedContributor?: OwnerScopedContributor;
}
export interface MethodRegistry {
lookup(name: string, scope: ScopeId, options?: MethodLookupOptions): readonly Resolution[];
}
export function buildMethodRegistry(ctx: RegistryContext): MethodRegistry {
return {
lookup(name: string, scope: ScopeId, options: MethodLookupOptions = {}) {
const params: CoreLookupParams = {
acceptedKinds: METHOD_KINDS,
useReceiverTypeBinding: true,
ownerScopedContributor: options.ownerScopedContributor ?? null,
...(options.callsite !== undefined ? { callsite: options.callsite } : {}),
...(options.explicitReceiver !== undefined
? { explicitReceiver: options.explicitReceiver }
: {}),
};
return lookupCore(name, scope, params, ctx);
},
};
}
@@ -0,0 +1,76 @@
/**
* `compareByConfidenceWithTiebreaks` — the RFC §4.2 Step 7 total order
* over `Resolution` candidates (Ring 2 SHARED #917).
*
* Primary key is confidence (DESC). Remaining ties within `CONFIDENCE_EPSILON`
* fall through a deterministic cascade so the same inputs always produce
* the same winner, independent of insertion order.
*
* Tie-break cascade (per RFC Appendix B):
*
* 1. confidence DESC (primary)
* 2. scope depth ASC (nearer lexical scope wins)
* 3. MRO depth ASC (nearer class in hierarchy wins)
* 4. `ORIGIN_PRIORITY` ASC (local > import > … > global-name)
* 5. DefId.localeCompare (final deterministic tiebreaker)
*
* The per-candidate inputs needed beyond `Resolution.confidence` —
* `scopeDepth`, `mroDepth`, `origin` — are supplied via a sidecar
* `TieBreakKey` so the comparator stays pure and `Resolution` itself
* doesn't need to carry book-keeping fields.
*/
import { ORIGIN_PRIORITY, type OriginForTieBreak } from '../origin-priority.js';
import type { Resolution } from '../types.js';
export const CONFIDENCE_EPSILON = 0.001;
/** Side-information per candidate used for secondary tie-breaks. */
export interface TieBreakKey {
readonly scopeDepth: number;
readonly mroDepth: number;
readonly origin: OriginForTieBreak;
}
/**
* Pure comparator suitable for `Array.prototype.sort`. Return value follows
* the JavaScript convention: negative → `a` wins, positive → `b` wins.
*
* **Important:** `keys` is keyed by `Resolution.def.nodeId`, not by array
* index — stable across reorderings. Missing keys fall back to neutral
* values (`scopeDepth: 0`, `mroDepth: 0`, `origin: 'local'`), which means
* the tie-break degrades gracefully to defId-lexicographic ordering when
* side-info is unavailable. That keeps the total order deterministic
* even on malformed inputs.
*/
export function compareByConfidenceWithTiebreaks(
a: Resolution,
b: Resolution,
keys: ReadonlyMap<string, TieBreakKey>,
): number {
// Primary: confidence DESC, treating values within epsilon as equal.
const delta = b.confidence - a.confidence;
if (Math.abs(delta) >= CONFIDENCE_EPSILON) return delta < 0 ? -1 : 1;
const ka = keys.get(a.def.nodeId) ?? DEFAULT_KEY;
const kb = keys.get(b.def.nodeId) ?? DEFAULT_KEY;
// Secondary: scope depth ASC.
if (ka.scopeDepth !== kb.scopeDepth) return ka.scopeDepth - kb.scopeDepth;
// Tertiary: MRO depth ASC.
if (ka.mroDepth !== kb.mroDepth) return ka.mroDepth - kb.mroDepth;
// Quaternary: ORIGIN_PRIORITY ASC.
const po = ORIGIN_PRIORITY[ka.origin] - ORIGIN_PRIORITY[kb.origin];
if (po !== 0) return po;
// Final: DefId lexicographic, locale-aware for deterministic cross-platform output.
return a.def.nodeId.localeCompare(b.def.nodeId);
}
const DEFAULT_KEY: TieBreakKey = Object.freeze({
scopeDepth: 0,
mroDepth: 0,
origin: 'local',
});
@@ -0,0 +1,148 @@
/**
* `resolveTypeRef` — strict single-return resolver for `TypeRef`s
* (RFC §4.6; Ring 2 SHARED #916).
*
* Narrower contract than `Registry.lookup`: no name-only global fallback, no
* confidence ranking, no arity check. Used by `Registry.lookup` Step 2 (type-
* binding propagation) and by any caller that wants the single best type-
* target for an annotation without paying for the full evidence pipeline.
*
* **Algorithm (strict).** Walk the scope chain from `ref.declaredAtScope`:
*
* 1. At each scope, inspect `bindings.get(ref.rawName)`:
* - If one of the bindings is a **type-kind** def with a **strict origin**
* (`'local' | 'import' | 'namespace' | 'reexport'`), return it.
* - If any binding for this name exists at this scope but none qualifies
* (e.g., a local variable named `User` shadows an outer import of class
* `User`), return `null`. The nearer binding shadows; we do NOT fall
* through to the global qualified-name index.
* - Otherwise continue to the parent scope.
* 2. If the raw name is a dotted path (e.g., `'models.User'`) and the scope
* walk produced no match, consult `QualifiedNameIndex.byQualifiedName`.
* Only accept **exactly one** type-kind hit — anything ambiguous returns
* `null` rather than a guess.
* 3. Return `null`.
*
* **What `'strict' origins' means.** `'wildcard'` is intentionally excluded.
* A wildcard-expanded name (`from x import *`) is too loose to use as an
* anchor for type resolution — it gives no signal about whether the name was
* actually imported. `Registry.lookup` may accept wildcard bindings at its
* own discretion (with lower evidence weight); `resolveTypeRef` does not.
*
* **What 'type-kind' means.** The subset of `NodeLabel` that a type annotation
* may legitimately reference: class-like, interface-like, enum-like, and
* alias-like kinds. See `TYPE_KINDS` below.
*
* Pure function — safe to call repeatedly; no side effects.
*/
import type { NodeLabel } from '../graph/types.js';
import type { SymbolDefinition } from './symbol-definition.js';
import type { BindingRef, ScopeId, ScopeLookup, TypeRef } from './types.js';
import type { DefIndex } from './def-index.js';
import type { QualifiedNameIndex } from './qualified-name-index.js';
// ─── Public contracts ───────────────────────────────────────────────────────
/**
* All inputs `resolveTypeRef` needs from the semantic model. Bundled into a
* context object so the call site stays short and the interface is stable as
* additional indexes get threaded through in later rings.
*/
export interface ResolveTypeRefContext {
readonly scopes: ScopeLookup;
readonly defIndex: DefIndex;
readonly qualifiedNameIndex: QualifiedNameIndex;
}
// ─── Strict policy constants ────────────────────────────────────────────────
/** `'wildcard'` is deliberately absent. See file header. */
const STRICT_ORIGINS: ReadonlySet<BindingRef['origin']> = new Set<BindingRef['origin']>([
'local',
'import',
'namespace',
'reexport',
]);
/**
* `NodeLabel` values that may appear on the RHS of a type annotation.
*
* Includes the usual class-like and interface-like kinds plus the alias-like
* ones (`TypeAlias`, `Typedef`). `Namespace` is excluded — it is a scope
* container, not a value type. `Function` / `Method` / `Variable` are
* excluded by design: a `rawName` bound to them at a strict origin is a
* *shadowing* binding, which the algorithm short-circuits to `null`.
*
* `'Type'` (the generic `NodeLabel` value) is also excluded — verified
* against `gitnexus/src/core/ingestion/` at the time of writing, no
* production extractor emits `type: 'Type'` for annotation-relevant
* symbols. Should a future extractor start emitting it, add `'Type'`
* here and add a test asserting the new path.
*/
const TYPE_KINDS: ReadonlySet<NodeLabel> = new Set<NodeLabel>([
'Class',
'Interface',
'Enum',
'Struct',
'Union',
'Trait',
'TypeAlias',
'Typedef',
'Record',
'Delegate',
'Annotation',
'Template',
]);
// ─── Main entry point ──────────────────────────────────────────────────────
export function resolveTypeRef(ref: TypeRef, ctx: ResolveTypeRefContext): SymbolDefinition | null {
// Phase 1: scope-chain walk anchored at the declaration site.
let currentId: ScopeId | null = ref.declaredAtScope;
const visited = new Set<ScopeId>();
while (currentId !== null) {
// Cycle guard — a well-formed scope tree never loops, but a bug in the
// construction path should fail fast here rather than hanging.
if (visited.has(currentId)) return null;
visited.add(currentId);
const scope = ctx.scopes.getScope(currentId);
if (scope === undefined) return null; // broken chain = unresolvable
const bindings = scope.bindings.get(ref.rawName);
if (bindings !== undefined && bindings.length > 0) {
// At least one binding exists at this scope → it is the shadowing site.
// Either one of them qualifies, or the name is shadowed by a non-type.
for (const binding of bindings) {
if (!STRICT_ORIGINS.has(binding.origin)) continue;
if (TYPE_KINDS.has(binding.def.type)) {
return binding.def;
}
}
// Shadowed by a non-type / non-strict-origin binding. Fail fast — no
// global fallback, no walk to the parent.
return null;
}
currentId = scope.parent;
}
// Phase 2: dotted fallback via `QualifiedNameIndex`. Only accept a unique
// type-kind hit; anything ambiguous returns null (strict: no guesses).
if (ref.rawName.includes('.')) {
const candidates = ctx.qualifiedNameIndex.get(ref.rawName);
let onlyTypeDef: SymbolDefinition | null = null;
for (const defId of candidates) {
const def = ctx.defIndex.get(defId);
if (def === undefined) continue;
if (!TYPE_KINDS.has(def.type)) continue;
if (onlyTypeDef !== null) return null; // ambiguous
onlyTypeDef = def;
}
if (onlyTypeDef !== null) return onlyTypeDef;
}
return null;
}
@@ -0,0 +1,57 @@
/**
* `ScopeId` canonical constructor + string intern pool
* (RFC §2.2; Ring 2 SHARED #912).
*
* `ScopeId` is a deterministic string derived from the scope's file path,
* byte range, and kind:
*
* scope:{filePath}#{startLine}:{startCol}-{endLine}:{endCol}:{kind}
*
* Two scopes produced by reparsing the same file at the same positions are
* `===`-equal as strings. Beyond the canonical shape, `makeScopeId` also
* **interns** the string through a process-local pool, so repeated calls
* with structurally identical inputs return the same string reference —
* making `Map<ScopeId, ...>` lookups and cache keys identity-fast.
*
* The intern pool is unbounded. The number of distinct `ScopeId`s across a
* single indexing run is O(total scopes in workspace), which is bounded by
* source-text size and already in memory; interning adds no asymptotic
* pressure. `clearScopeIdInternPool` is exported for test isolation.
*/
import type { Range } from './types.js';
import type { ScopeId, ScopeKind } from './types.js';
/** Inputs required to construct a canonical `ScopeId`. */
export interface ScopeIdInput {
readonly filePath: string;
readonly range: Range;
readonly kind: ScopeKind;
}
/**
* Build a canonical `ScopeId` from its structural parts and intern it.
*
* Pure + referentially transparent: given the same input shape, always
* returns the same string reference for the lifetime of the pool.
*/
export function makeScopeId(input: ScopeIdInput): ScopeId {
const raw = `scope:${input.filePath}#${input.range.startLine}:${input.range.startCol}-${input.range.endLine}:${input.range.endCol}:${input.kind}`;
const existing = INTERN_POOL.get(raw);
if (existing !== undefined) return existing;
INTERN_POOL.set(raw, raw);
return raw;
}
/**
* Drop the intern pool. Intended for test setup/teardown — production code
* should not need this, since the pool's memory usage is bounded by the
* number of live scopes and cleaning it mid-run would break identity
* equality for existing scope ids.
*/
export function clearScopeIdInternPool(): void {
INTERN_POOL.clear();
}
/** Internal: shared intern pool (process-local). */
const INTERN_POOL = new Map<string, string>();
@@ -0,0 +1,254 @@
/**
* `ScopeTree` — the lexical-scope spine of the `SemanticModel`
* (RFC §2.2 + §3.1; Ring 2 SHARED #912).
*
* Generalizes the `enclosingFunctions` pattern from closed PR #902 to
* arbitrary `ScopeKind`s. Owns the (parent ↔ children) relationship
* derived from each `Scope.parent` pointer, and validates the structural
* invariants a well-formed scope tree must satisfy.
*
* Invariants enforced at build time (throw on violation):
*
* - Every non-`Module` scope has a non-null parent.
* - Every parent pointer references a scope that was also supplied to
* `buildScopeTree`.
* - Parent range **strictly contains** child range.
* - Sibling ranges under the same parent do not overlap.
* - Parent and child live in the same `filePath`. (Cross-file parent
* pointers would be a category error — a `File` scope is not the
* parent of another file's scopes; imports do that job.)
*
* Satisfies the `ScopeLookup` contract (defined in `./types.js`), so
* `resolveTypeRef` (#916) and the scope-aware registries (#917) can take a
* `ScopeTree` directly without adapters.
*
* Immutable surface: `byId` is a `ReadonlyMap`; children arrays are
* `Object.freeze`d; miss lookups return a shared frozen empty array.
*/
import type { Scope, ScopeId, ScopeLookup, Range } from './types.js';
// ─── Public contract ────────────────────────────────────────────────────────
export interface ScopeTree extends ScopeLookup {
readonly size: number;
readonly byId: ReadonlyMap<ScopeId, Scope>;
getScope(id: ScopeId): Scope | undefined;
getParent(id: ScopeId): Scope | undefined;
/** Child `ScopeId`s of `id`, in input order. Frozen empty array on miss. */
getChildren(id: ScopeId): readonly ScopeId[];
/**
* Ancestor chain from the immediate parent up to (and including) the
* root module scope. Excludes the starting scope itself. Frozen empty
* array on miss / for a root scope.
*/
getAncestors(id: ScopeId): readonly ScopeId[];
has(id: ScopeId): boolean;
}
// ─── Build errors ───────────────────────────────────────────────────────────
/**
* Thrown by `buildScopeTree` when the input violates a structural
* invariant. Carries the offending ids + the invariant name so failed
* extraction pipelines can report actionable diagnostics.
*/
export class ScopeTreeInvariantError extends Error {
constructor(
readonly invariant:
| 'non-module-requires-parent'
| 'parent-not-found'
| 'parent-must-contain-child'
| 'sibling-ranges-overlap'
| 'parent-must-share-filepath'
| 'duplicate-scope-id',
message: string,
) {
super(message);
this.name = 'ScopeTreeInvariantError';
}
}
// ─── Builder ───────────────────────────────────────────────────────────────
/**
* Build an immutable `ScopeTree` from a flat list of `Scope` records.
*
* Throws `ScopeTreeInvariantError` on the first invariant violation; a
* malformed tree is a bug in the extraction pipeline, not a data case for
* consumers to handle, so fail-fast is the correct posture.
*/
export function buildScopeTree(scopes: readonly Scope[]): ScopeTree {
const byId = new Map<ScopeId, Scope>();
const childrenById = new Map<ScopeId, ScopeId[]>();
// ── Pass 1: collect by id + duplicate check ───────────────────────────
for (const scope of scopes) {
if (byId.has(scope.id)) {
throw new ScopeTreeInvariantError(
'duplicate-scope-id',
`Two scopes share id '${scope.id}'. Scope ids must be unique per tree.`,
);
}
byId.set(scope.id, scope);
}
// ── Pass 2: validate parent pointers + build children buckets ─────────
for (const scope of scopes) {
if (scope.parent === null) {
if (scope.kind !== 'Module') {
throw new ScopeTreeInvariantError(
'non-module-requires-parent',
`Scope '${scope.id}' has kind '${scope.kind}' but no parent. Only 'Module' scopes may be root-level.`,
);
}
continue;
}
const parent = byId.get(scope.parent);
if (parent === undefined) {
throw new ScopeTreeInvariantError(
'parent-not-found',
`Scope '${scope.id}' references parent '${scope.parent}' which is not part of this tree.`,
);
}
if (parent.filePath !== scope.filePath) {
throw new ScopeTreeInvariantError(
'parent-must-share-filepath',
`Scope '${scope.id}' (${scope.filePath}) has parent '${parent.id}' in a different file (${parent.filePath}). Parent/child scopes must share filePath.`,
);
}
if (!rangeStrictlyContains(parent.range, scope.range)) {
throw new ScopeTreeInvariantError(
'parent-must-contain-child',
`Parent scope '${parent.id}' at ${formatRange(parent.range)} does not strictly contain child '${scope.id}' at ${formatRange(scope.range)}.`,
);
}
let bucket = childrenById.get(parent.id);
if (bucket === undefined) {
bucket = [];
childrenById.set(parent.id, bucket);
}
bucket.push(scope.id);
}
// ── Pass 3: sibling-overlap check ─────────────────────────────────────
for (const [parentId, childIds] of childrenById) {
if (childIds.length < 2) continue;
// Sort siblings by (startLine, startCol) for an O(n log n) pairwise
// scan instead of O(n²) all-pairs.
const children = childIds.map((id) => byId.get(id)!).slice();
children.sort((a, b) => comparePosition(a.range, b.range));
for (let i = 1; i < children.length; i++) {
const prev = children[i - 1]!;
const curr = children[i]!;
if (rangesOverlap(prev.range, curr.range)) {
throw new ScopeTreeInvariantError(
'sibling-ranges-overlap',
`Sibling scopes under parent '${parentId}' overlap: '${prev.id}' ${formatRange(prev.range)} and '${curr.id}' ${formatRange(curr.range)}.`,
);
}
}
}
// Freeze children arrays so the surface is truly read-only.
const frozenChildren = new Map<ScopeId, readonly ScopeId[]>();
for (const [parentId, childIds] of childrenById) {
frozenChildren.set(parentId, Object.freeze(childIds.slice()));
}
return freezeTree(byId, frozenChildren);
}
// ─── Internals ──────────────────────────────────────────────────────────────
const EMPTY_CHILDREN: readonly ScopeId[] = Object.freeze([]);
function freezeTree(
byId: Map<ScopeId, Scope>,
childrenById: Map<ScopeId, readonly ScopeId[]>,
): ScopeTree {
return {
byId,
get size() {
return byId.size;
},
getScope(id: ScopeId): Scope | undefined {
return byId.get(id);
},
getParent(id: ScopeId): Scope | undefined {
const scope = byId.get(id);
if (scope === undefined || scope.parent === null) return undefined;
return byId.get(scope.parent);
},
getChildren(id: ScopeId): readonly ScopeId[] {
return childrenById.get(id) ?? EMPTY_CHILDREN;
},
getAncestors(id: ScopeId): readonly ScopeId[] {
const start = byId.get(id);
if (start === undefined || start.parent === null) return EMPTY_CHILDREN;
const out: ScopeId[] = [];
const visited = new Set<ScopeId>([id]);
let cursor: ScopeId | null = start.parent;
while (cursor !== null && !visited.has(cursor)) {
visited.add(cursor);
out.push(cursor);
const next = byId.get(cursor);
cursor = next === undefined ? null : next.parent;
}
return Object.freeze(out);
},
has(id: ScopeId): boolean {
return byId.has(id);
},
};
}
/**
* `outer` strictly contains `inner` when `outer`'s start is at or before
* `inner`'s start, `outer`'s end is at or after `inner`'s end, and they are
* not the exact same range. Equal ranges are rejected — a child cannot
* occupy the exact same span as its parent.
*/
function rangeStrictlyContains(outer: Range, inner: Range): boolean {
if (
outer.startLine === inner.startLine &&
outer.startCol === inner.startCol &&
outer.endLine === inner.endLine &&
outer.endCol === inner.endCol
) {
return false;
}
const outerStartsAtOrBefore =
outer.startLine < inner.startLine ||
(outer.startLine === inner.startLine && outer.startCol <= inner.startCol);
const outerEndsAtOrAfter =
outer.endLine > inner.endLine ||
(outer.endLine === inner.endLine && outer.endCol >= inner.endCol);
return outerStartsAtOrBefore && outerEndsAtOrAfter;
}
/**
* Two ranges overlap when neither finishes before the other begins. Ranges
* that merely touch at a single boundary point (`a.end === b.start`) do
* NOT overlap — this matches tree-sitter's half-open-like range semantics
* and the typical "sibling blocks meet but don't overlap" pattern.
*/
function rangesOverlap(a: Range, b: Range): boolean {
const aEndsBeforeB =
a.endLine < b.startLine || (a.endLine === b.startLine && a.endCol <= b.startCol);
const bEndsBeforeA =
b.endLine < a.startLine || (b.endLine === a.startLine && b.endCol <= a.startCol);
return !(aEndsBeforeB || bEndsBeforeA);
}
function comparePosition(a: Range, b: Range): number {
if (a.startLine !== b.startLine) return a.startLine - b.startLine;
return a.startCol - b.startCol;
}
function formatRange(r: Range): string {
return `${r.startLine}:${r.startCol}-${r.endLine}:${r.endCol}`;
}
@@ -0,0 +1,188 @@
/**
* Shadow-mode aggregation — per-language parity %, per-evidence-kind
* breakdown of divergences. Consumed by the parity dashboard (RING2-PKG-5).
*
* Pure functions; no I/O. The harness persists per-run JSON; the dashboard
* reads `.gitnexus/shadow-parity/latest.json` and renders.
*
* Related types — `ShadowAgreement`, `ShadowCallsite`, `ShadowDiff` — are
* defined alongside `diffResolutions` in `./diff.ts` and re-exported
* through the top-level `gitnexus-shared` barrel. Consumers import all
* three from `gitnexus-shared`, not from this module.
*
* Part of RFC #909 Ring 2 SHARED — #918.
*/
import type { SupportedLanguages } from '../../languages.js';
import type { ResolutionEvidence } from '../types.js';
import type { ShadowAgreement, ShadowDiff } from './diff.js';
// ─── Aggregated report shape ────────────────────────────────────────────────
export interface LanguageParityRow {
readonly language: SupportedLanguages;
readonly totalCalls: number;
readonly bothAgree: number;
readonly onlyLegacy: number;
readonly onlyNew: number;
readonly bothDisagree: number;
readonly bothEmpty: number;
/**
* Fraction in [0, 1]. Numerator = `bothAgree`; denominator = "calls where
* at least one side resolved" = `totalCalls - bothEmpty`.
*
* When the denominator is 0 (all calls for this language were
* `both-empty`), returns 0. Callers rendering the dashboard should treat
* a 0 parity alongside `totalCalls === bothEmpty` as "no signal" rather
* than "total disagreement".
*/
readonly parity: number;
/**
* Divergence signals broken down by `ResolutionEvidence.kind`. Sourced
* from `ShadowDiff.evidenceDelta` on non-agreeing rows only — `both-agree`
* and `both-empty` do not contribute.
*/
readonly evidenceBreakdown: ReadonlyMap<ResolutionEvidence['kind'], number>;
}
export interface ShadowParityReport {
readonly generatedAt: string; // ISO 8601
readonly perLanguage: readonly LanguageParityRow[];
readonly overall: Omit<LanguageParityRow, 'language' | 'evidenceBreakdown'>;
}
// ─── Public API ─────────────────────────────────────────────────────────────
/**
* Aggregate a stream of `ShadowDiff` records into a `ShadowParityReport`,
* bucketed by language. Pure function.
*
* - `perLanguage` rows are sorted alphabetically by `SupportedLanguages`
* value for stable JSON output (the dashboard reads
* `.gitnexus/shadow-parity/latest.json` and diffing snapshots is useful).
* - `overall` is the column-wise sum across languages.
* - `generatedAt` is injected via the `now` parameter so tests stay
* deterministic; production callers let it default to `new Date()`.
*/
export function aggregateDiffs(
diffs: readonly { readonly language: SupportedLanguages; readonly diff: ShadowDiff }[],
now: Date = new Date(),
): ShadowParityReport {
const perLanguageMap = new Map<SupportedLanguages, MutableCounts>();
for (const { language, diff } of diffs) {
let counts = perLanguageMap.get(language);
if (!counts) {
counts = makeEmptyCounts();
perLanguageMap.set(language, counts);
}
tallyDiff(counts, diff);
}
const perLanguage: LanguageParityRow[] = Array.from(perLanguageMap.entries())
.map(([language, counts]) => buildRow(language, counts))
.sort((a, b) => a.language.localeCompare(b.language));
const overall = buildOverallRow(perLanguage);
return {
generatedAt: now.toISOString(),
perLanguage,
overall,
};
}
// ─── Internal helpers ───────────────────────────────────────────────────────
interface MutableCounts {
totalCalls: number;
bothAgree: number;
onlyLegacy: number;
onlyNew: number;
bothDisagree: number;
bothEmpty: number;
evidenceBreakdown: Map<ResolutionEvidence['kind'], number>;
}
function makeEmptyCounts(): MutableCounts {
return {
totalCalls: 0,
bothAgree: 0,
onlyLegacy: 0,
onlyNew: 0,
bothDisagree: 0,
bothEmpty: 0,
evidenceBreakdown: new Map(),
};
}
function tallyDiff(counts: MutableCounts, diff: ShadowDiff): void {
counts.totalCalls += 1;
incrementAgreement(counts, diff.agreement);
if (diff.agreement === 'both-agree' || diff.agreement === 'both-empty') return;
for (const ev of diff.evidenceDelta) {
counts.evidenceBreakdown.set(ev.kind, (counts.evidenceBreakdown.get(ev.kind) ?? 0) + 1);
}
}
function incrementAgreement(counts: MutableCounts, agreement: ShadowAgreement): void {
switch (agreement) {
case 'both-agree':
counts.bothAgree += 1;
return;
case 'only-legacy':
counts.onlyLegacy += 1;
return;
case 'only-new':
counts.onlyNew += 1;
return;
case 'both-disagree':
counts.bothDisagree += 1;
return;
case 'both-empty':
counts.bothEmpty += 1;
return;
}
}
function buildRow(language: SupportedLanguages, counts: MutableCounts): LanguageParityRow {
const resolved = counts.totalCalls - counts.bothEmpty;
const parity = resolved > 0 ? counts.bothAgree / resolved : 0;
return {
language,
totalCalls: counts.totalCalls,
bothAgree: counts.bothAgree,
onlyLegacy: counts.onlyLegacy,
onlyNew: counts.onlyNew,
bothDisagree: counts.bothDisagree,
bothEmpty: counts.bothEmpty,
parity,
// Freeze via `new Map` on a sorted-kind copy so downstream consumers
// can't mutate the aggregator's internal state.
evidenceBreakdown: new Map(
Array.from(counts.evidenceBreakdown.entries()).sort(([a], [b]) => a.localeCompare(b)),
),
};
}
function buildOverallRow(
perLanguage: readonly LanguageParityRow[],
): Omit<LanguageParityRow, 'language' | 'evidenceBreakdown'> {
let totalCalls = 0;
let bothAgree = 0;
let onlyLegacy = 0;
let onlyNew = 0;
let bothDisagree = 0;
let bothEmpty = 0;
for (const row of perLanguage) {
totalCalls += row.totalCalls;
bothAgree += row.bothAgree;
onlyLegacy += row.onlyLegacy;
onlyNew += row.onlyNew;
bothDisagree += row.bothDisagree;
bothEmpty += row.bothEmpty;
}
const resolved = totalCalls - bothEmpty;
const parity = resolved > 0 ? bothAgree / resolved : 0;
return { totalCalls, bothAgree, onlyLegacy, onlyNew, bothDisagree, bothEmpty, parity };
}
@@ -0,0 +1,126 @@
/**
* Shadow-mode diff logic — RFC §6.3.
*
* Pure comparison logic for shadow mode. Takes two `Resolution[]` (legacy
* DAG result + new scope-based registry result) and produces a structured
* diff record for the parity dashboard.
*
* Consumed by the Ring 2 PKG shadow harness (#923), which dual-runs each
* call through legacy + new paths, diffs results, and persists per-run JSON
* for the parity dashboard.
*
* Part of RFC #909 Ring 2 SHARED — #918.
*/
import type { Resolution, ResolutionEvidence } from '../types.js';
// ─── Diff record shape ──────────────────────────────────────────────────────
export type ShadowAgreement =
| 'both-agree' // top match identical (same DefId)
| 'only-legacy' // legacy resolved; new did not
| 'only-new' // new resolved; legacy did not
| 'both-disagree' // both resolved, but to different targets
| 'both-empty'; // both returned empty
export interface ShadowDiff {
readonly callsite: ShadowCallsite;
readonly legacy: Resolution | null;
readonly newResult: Resolution | null;
readonly agreement: ShadowAgreement;
/**
* Symmetric difference of the two top resolutions' `evidence` arrays,
* keyed on `ResolutionEvidence.kind`.
*
* - For `'both-agree'` and `'both-empty'` agreements, always empty.
* - For `'both-disagree'`, contains evidence kinds present on exactly one
* side (not in both).
* - For `'only-legacy'`, contains all of legacy's top evidence.
* - For `'only-new'`, contains all of new's top evidence.
*/
readonly evidenceDelta: readonly ResolutionEvidence[];
}
export interface ShadowCallsite {
readonly filePath: string;
readonly line: number;
readonly col: number;
readonly calledName: string;
}
// ─── Public API ─────────────────────────────────────────────────────────────
/**
* Compare two `Resolution[]` arrays (top matches at `[0]`) and produce a
* `ShadowDiff`. Pure function.
*
* Agreement rules:
* - both arrays empty → `'both-empty'`, `evidenceDelta: []`
* - legacy empty, new non-empty → `'only-new'`, `evidenceDelta` = new's top evidence
* - legacy non-empty, new empty → `'only-legacy'`, `evidenceDelta` = legacy's top evidence
* - both non-empty, same top `def.nodeId` → `'both-agree'`, `evidenceDelta: []`
* - both non-empty, different top `def.nodeId` → `'both-disagree'`,
* `evidenceDelta` = symmetric difference by `ResolutionEvidence.kind`
* (first occurrence of a kind-only-on-legacy then kind-only-on-new; order
* preserved from input arrays)
*
* Evidence-delta rationale: callers aggregating divergences want to know
* which signal kinds explain a disagreement. Keying on `kind` (not full
* equality over `weight`/`note`) avoids spurious deltas when the same
* signal fires with slightly different calibration weights on each side.
*/
export function diffResolutions(
callsite: ShadowCallsite,
legacy: readonly Resolution[],
newResult: readonly Resolution[],
): ShadowDiff {
const legacyTop: Resolution | null = legacy.length > 0 ? legacy[0] : null;
const newTop: Resolution | null = newResult.length > 0 ? newResult[0] : null;
const agreement: ShadowAgreement = (() => {
if (legacyTop === null && newTop === null) return 'both-empty';
if (legacyTop === null) return 'only-new';
if (newTop === null) return 'only-legacy';
return legacyTop.def.nodeId === newTop.def.nodeId ? 'both-agree' : 'both-disagree';
})();
const evidenceDelta = computeEvidenceDelta(legacyTop, newTop, agreement);
return {
callsite,
legacy: legacyTop,
newResult: newTop,
agreement,
evidenceDelta,
};
}
// ─── Internal helpers ───────────────────────────────────────────────────────
/**
* Symmetric difference of two evidence arrays, keyed on
* `ResolutionEvidence.kind`. Preserves input order: legacy-only signals
* first (in legacy's original order), then new-only signals (in new's order).
*
* For `'both-agree'` / `'both-empty'` the delta is empty by contract. For
* `'only-legacy'` / `'only-new'` one side's evidence is the delta (nothing to
* subtract against).
*/
function computeEvidenceDelta(
legacy: Resolution | null,
newResult: Resolution | null,
agreement: ShadowAgreement,
): readonly ResolutionEvidence[] {
if (agreement === 'both-agree' || agreement === 'both-empty') return [];
if (agreement === 'only-legacy') return legacy!.evidence;
if (agreement === 'only-new') return newResult!.evidence;
// both-disagree: symmetric difference keyed on `kind`
const legacyKinds = new Set(legacy!.evidence.map((e) => e.kind));
const newKinds = new Set(newResult!.evidence.map((e) => e.kind));
const onlyInLegacy = legacy!.evidence.filter((e) => !newKinds.has(e.kind));
const onlyInNew = newResult!.evidence.filter((e) => !legacyKinds.has(e.kind));
return [...onlyInLegacy, ...onlyInNew];
}
@@ -0,0 +1,35 @@
/**
* `SymbolDefinition` — the canonical shape of an indexed symbol record.
*
* Historically defined in `gitnexus/src/core/ingestion/model/symbol-table.ts`;
* moved into `gitnexus-shared` as part of RFC #909 Ring 1 (#910) so the
* scope-resolution types that reference it can live in the shared package
* alongside their consumers (`gitnexus/` and `gitnexus-web/`).
*
* Shape is unchanged from the prior local definition.
*/
import type { NodeLabel } from '../graph/types.js';
export interface SymbolDefinition {
nodeId: string;
filePath: string;
type: NodeLabel;
/** Canonical dot-separated qualified type name for class-like symbols
* (e.g. `App.Models.User`). Falls back to the simple symbol name when no
* package/namespace/module scope exists or no explicit qualified metadata is provided. */
qualifiedName?: string;
parameterCount?: number;
/** Number of required (non-optional, non-default) parameters.
* Enables range-based arity filtering: argCount >= requiredParameterCount && argCount <= parameterCount. */
requiredParameterCount?: number;
/** Per-parameter type names for overload disambiguation (e.g. ['int', 'String']).
* Populated when parameter types are resolvable from AST (any typed language). */
parameterTypes?: string[];
/** Raw return type text extracted from AST (e.g. 'User', 'Promise<User>') */
returnType?: string;
/** Declared type for non-callable symbols — fields/properties (e.g. 'Address', 'List<User>') */
declaredType?: string;
/** Links Method/Constructor/Property to owning Class/Struct/Trait nodeId */
ownerId?: string;
}
@@ -0,0 +1,432 @@
/**
* Scope-resolution type definitions — RFC §2 data model (authoritative source).
*
* See: https://www.notion.so/346dc50b6ed281cfaacbe480bf231d50
*
* Anti-drift rule: every type, interface, and enum defined here is the single
* source of truth. Later code that references these names must import them
* from `gitnexus-shared`; it must not re-define them locally.
*
* Lifecycle contract (RFC §2.8): scopes are **constructed during extraction,
* linked during finalize, immutable after finalize**. All fields are
* `readonly` at the type level; `Object.freeze` is applied at runtime in dev
* builds. `ReferenceIndex` is the sole structure populated after freeze — by
* resolution, before emission.
*/
import type { NodeLabel } from '../graph/types.js';
import type { SymbolDefinition } from './symbol-definition.js';
// ─── §2.1 Type aliases ──────────────────────────────────────────────────────
/** Stable per-(file, range, kind) scope identifier; interned for identity-fast equality. */
export type ScopeId = string;
/** Stable symbol-definition identifier (graph nodeId). */
export type DefId = string;
/** Kinds of lexical scope a `Scope` node can represent. */
export type ScopeKind =
| 'Module' // file root
| 'Namespace' // C++ namespace, C# namespace, Kotlin package-object, Rust mod
| 'Class' // class/struct/trait/interface body
| 'Function' // function/method/closure/lambda body
| 'Block' // { ... }, if-body, for-body, with-body, match arms
| 'Expression'; // comprehensions, for-init, pattern bindings, lambda param lists
// ─── Range + Capture (parser-agnostic) ──────────────────────────────────────
/** Source-text range. 1-based `startLine`/`endLine`; 0-based `startCol`/`endCol`. */
export interface Range {
readonly startLine: number;
readonly startCol: number;
readonly endLine: number;
readonly endCol: number;
}
/**
* Tagged capture emitted by a LanguageProvider's `emitScopeCaptures` hook.
*
* Parser-agnostic: tree-sitter queries and COBOL's regex tagger both produce
* `Capture[]`. The central `ScopeExtractor` consumes captures without
* knowing which parser produced them.
*/
export interface Capture {
/** Capture name, including leading `@` (e.g., `'@scope.module'`, `'@declaration.class'`). */
readonly name: string;
readonly range: Range;
/** The captured source text. */
readonly text: string;
}
/**
* A grouping of `Capture`s that came from a single query match (e.g., one
* `@import.statement` match carries `@import.source`, `@import.name`,
* `@import.alias?` as child captures). Keyed by capture name for O(1)
* child access.
*/
export type CaptureMatch = Readonly<Record<string, Capture>>;
// ─── Hook input/output types (RFC §5.2) ─────────────────────────────────────
/**
* Provider-interpreted raw import, consumed by finalize (Phase 2) to produce
* linked `ImportEdge[]`. The provider's `interpretImport` hook turns a
* `CaptureMatch` for an `@import.statement` into one of these; the central
* finalize algorithm resolves `targetRaw` to a concrete file via
* `resolveImportTarget` and materializes the final `ImportEdge`.
*
* Discriminated union — each variant carries only the fields that make sense
* for its kind. Invalid shapes (e.g., a `namespace` import with an alias-like
* `importedName` mismatch) are compile errors, not latent bugs. `'wildcard-
* expanded'` is deliberately NOT a variant: that kind is finalize output only,
* produced when `expandsWildcardTo` materializes a wildcard against target
* exports — a provider must never emit it at parse time.
*/
export type ParsedImport =
/**
* Per-name import without rename.
*
* Examples:
* - Python `from foo import X` → `{ kind: 'named', localName: 'X', importedName: 'X', targetRaw: 'foo' }`
* - TS `import { X } from './foo'` → `{ kind: 'named', localName: 'X', importedName: 'X', targetRaw: './foo' }`
* - Java `import foo.bar.X` → `{ kind: 'named', localName: 'X', importedName: 'X', targetRaw: 'foo.bar' }`
*/
| {
readonly kind: 'named';
readonly localName: string;
readonly importedName: string;
readonly targetRaw: string;
}
/**
* Per-name import with rename.
*
* Examples:
* - Python `from foo import X as Y` → `{ kind: 'alias', localName: 'Y', importedName: 'X', alias: 'Y', targetRaw: 'foo' }`
* - TS `import { X as Y } from './foo'` → `{ kind: 'alias', localName: 'Y', importedName: 'X', alias: 'Y', targetRaw: './foo' }`
*/
| {
readonly kind: 'alias';
readonly localName: string;
readonly importedName: string;
readonly alias: string;
readonly targetRaw: string;
}
/**
* Qualified module handle, with or without rename. `importedName` is the
* module being aliased; `localName` is the scope-visible handle (often the
* same unless renamed).
*
* Examples:
* - Python `import numpy` → `{ kind: 'namespace', localName: 'numpy', importedName: 'numpy', targetRaw: 'numpy' }`
* - Python `import numpy as np` → `{ kind: 'namespace', localName: 'np', importedName: 'numpy', targetRaw: 'numpy' }`
* - TS `import * as np from 'numpy'` → `{ kind: 'namespace', localName: 'np', importedName: 'numpy', targetRaw: 'numpy' }`
* - Go `import foo "pkg/bar"` → `{ kind: 'namespace', localName: 'foo', importedName: 'bar', targetRaw: 'pkg/bar' }`
*/
| {
readonly kind: 'namespace';
/** Scope-visible handle (e.g. `np` in `import numpy as np`; `numpy` when unaliased). */
readonly localName: string;
/** Module being aliased (e.g. `numpy` in `import numpy as np`). */
readonly importedName: string;
readonly targetRaw: string;
}
/**
* Syntactically-detectable parse-time re-export. Finalize may still produce
* `ImportEdge { kind: 'reexport', transitiveVia }` when flattening chains;
* this variant preserves the *parse-time* signal so finalize doesn't have
* to re-derive it from scratch.
*
* Examples:
* - TS `export { X } from './y'` → `{ kind: 'reexport', localName: 'X', importedName: 'X', targetRaw: './y' }`
* - TS `export { X as Y } from './y'` → `{ kind: 'reexport', localName: 'Y', importedName: 'X', alias: 'Y', targetRaw: './y' }`
* - Rust `pub use foo::bar` → `{ kind: 'reexport', localName: 'bar', importedName: 'bar', targetRaw: 'foo' }`
*/
| {
readonly kind: 'reexport';
/** Name as re-exported in the current module. */
readonly localName: string;
/** Name in the source module. */
readonly importedName: string;
readonly targetRaw: string;
/** Set when the re-export renames the symbol (e.g. `export { X as Y } from './y'`). */
readonly alias?: string;
}
/**
* Wildcard import — brings every exported name from the target module into
* the importing scope. The finalize algorithm expands this into one
* `BindingRef` per exported name via the provider's `expandsWildcardTo`
* hook, producing the finalize-only `ImportEdge` kind `'wildcard-expanded'`.
*
* Examples:
* - Python `from foo import *` → `{ kind: 'wildcard', targetRaw: 'foo' }`
* - JS `export * from './foo'` → `{ kind: 'wildcard', targetRaw: './foo' }`
* - Rust `pub use foo::*` → `{ kind: 'wildcard', targetRaw: 'foo' }`
*/
| {
readonly kind: 'wildcard';
readonly targetRaw: string;
}
/**
* Runtime-computed target — the import path is not a static literal at
* parse time. Providers SHOULD emit the unresolvable expression's source
* text as `targetRaw` to aid diagnostics; `null` only when no string form
* exists.
*
* Examples:
* - JS `await import(expr)` → `{ kind: 'dynamic-unresolved', localName: '', targetRaw: 'expr' }`
* - Python `importlib.import_module(f'pkg.{name}')` → `{ kind: 'dynamic-unresolved', localName: '', targetRaw: "f'pkg.{name}'" }`
*/
| {
readonly kind: 'dynamic-unresolved';
readonly localName: string;
/** Source text of the unresolved expression when available; `null` otherwise. */
readonly targetRaw: string | null;
};
/**
* Provider-interpreted type binding. The provider's `interpretTypeBinding`
* hook turns a `CaptureMatch` (e.g., `@type-binding.parameter`) into one of
* these; the central extractor attaches the resulting `TypeRef` to the
* appropriate scope's `typeBindings` map.
*/
export interface ParsedTypeBinding {
/** The name being bound (parameter name, `self`, assignment LHS, …). */
readonly boundName: string;
/** The raw type name as written in source (`'User'`, `'models.User'`, …). */
readonly rawTypeName: string;
readonly source: TypeRef['source'];
}
/**
* Cross-file workspace index consumed by finalize-phase hooks
* (`resolveImportTarget`, `expandsWildcardTo`). Opaque placeholder in Ring 1;
* concretely typed in Ring 2 SHARED (#915).
*/
export type WorkspaceIndex = unknown;
// `ScopeTree` is exported from `./scope-tree.js` as of Ring 2 SHARED (#912).
// The former opaque placeholder lived here during Ring 1; removed now that
// the concrete type exists. Consumers import from `gitnexus-shared` directly.
/**
* Minimal scope-lookup contract: map a `ScopeId` back to its `Scope` record.
*
* Lives in the data-model layer so both `ScopeTree` (§3.1) and
* `resolveTypeRef` / `Registry.lookup` (§4) can depend on it without
* inverting each other. `ScopeTree` is the canonical implementation;
* tests and future alternative containers may supply their own.
*/
export interface ScopeLookup {
getScope(id: ScopeId): Scope | undefined;
}
/** Call-site description passed to `arityCompatibility`. */
export interface Callsite {
/** Number of arguments at the call site. */
readonly arity: number;
}
// ─── §2.4 ImportEdge ────────────────────────────────────────────────────────
/**
* A cross-file import edge attached to a module/namespace scope.
*
* Raw (unlinked) edges are emitted during parse (Phase 1); `targetModuleScope`
* and `targetDefId` are filled in during finalize (Phase 2) via SCC-aware
* bounded-fixpoint linking (RFC §3.2).
*/
export interface ImportEdge {
/** How this scope sees the imported name (after alias). */
readonly localName: string;
/** Exporting file; `null` only when `kind === 'dynamic-unresolved'`. */
readonly targetFile: string | null;
/** The name under which the target exports this symbol. */
readonly targetExportedName: string;
/** Pre-resolved at finalize: the module scope of the exporting file. */
readonly targetModuleScope?: ScopeId;
/** Pre-resolved at finalize: the exported symbol's `DefId`. */
readonly targetDefId?: DefId;
readonly kind:
| 'named'
| 'alias'
| 'namespace'
| 'wildcard-expanded'
| 'reexport'
| 'dynamic-unresolved';
/** Re-export chain, for provenance (e.g., `['./y']` when re-exported via `./y`). */
readonly transitiveVia?: readonly string[];
/** Set to `'unresolved'` when the SCC fixpoint could not link this edge. */
readonly linkStatus?: 'unresolved';
}
// ─── §2.3 BindingRef ────────────────────────────────────────────────────────
/**
* A name binding visible at a scope, with provenance.
*
* Provenance stays at the visibility layer — a name being visible because it
* is local vs imported vs wildcard-expanded vs re-exported is a property of
* the binding itself. This keeps evidence emission and `import-use` reference
* stamping first-class instead of reconstructing provenance from a side table.
*/
export interface BindingRef {
readonly def: SymbolDefinition;
readonly origin: 'local' | 'import' | 'namespace' | 'wildcard' | 'reexport';
/** Non-null for non-local origins; carries the `ImportEdge` that brought the name into this scope. */
readonly via?: ImportEdge;
}
// ─── §2.5 TypeRef ───────────────────────────────────────────────────────────
/**
* A reference to a named type, anchored at its declaration site.
*
* Design choice: raw name + declaration-site scope, resolved at lookup time.
* Pre-resolution would invert the extraction/resolution wall. Deferred thunks
* add no capability. Structured type systems are months of work per language.
* This shape keeps V1 tractable while preserving correctness for aliases,
* re-exports, and nested modules. Generics deferred to V2 via `typeArgs`.
*/
export interface TypeRef {
/** The name as written in source (e.g., `'User'`, `'models.User'`, `'List'`). */
readonly rawName: string;
/** Anchor for resolving `rawName` — the scope where the annotation/inference was written. */
readonly declaredAtScope: ScopeId;
readonly source:
| 'annotation'
| 'parameter-annotation'
| 'return-annotation'
| 'self'
| 'assignment-inferred'
| 'constructor-inferred'
| 'receiver-propagated';
/** Reserved for V2+: generic type arguments (`List<User>` → `[TypeRef('User')]`). V1 ignores. */
readonly typeArgs?: readonly TypeRef[];
}
// ─── §2.2 Scope ─────────────────────────────────────────────────────────────
/**
* The canonical lexical-scope node. Forms the spine of the SemanticModel.
*
* ScopeId shape (RFC §2.2): `scope:{filePath}#{startLine}:{startCol}-{endLine}:{endCol}:{kind}`
* — deterministic, stable across reparses of the same source, interned.
*/
export interface Scope {
readonly id: ScopeId;
readonly parent: ScopeId | null;
readonly kind: ScopeKind;
readonly range: Range;
readonly filePath: string;
/** Names visible from this scope. Provenance preserved via `BindingRef.origin`. */
readonly bindings: ReadonlyMap<string, readonly BindingRef[]>;
/** Defs structurally owned by this scope (e.g., methods owned by a class body scope). */
readonly ownedDefs: readonly SymbolDefinition[];
/** Import edges attached to this scope. Mostly module/namespace scopes, but some
* languages allow local imports (Python `def f(): from x import Y`, Rust
* fn-local `use`, TS dynamic `import()`). */
readonly imports: readonly ImportEdge[];
/** Local type facts visible from this scope (parameter annotations, `self` binding, etc.). */
readonly typeBindings: ReadonlyMap<string, TypeRef>;
}
// ─── §2.6 Resolution + ResolutionEvidence ───────────────────────────────────
/**
* One piece of evidence for a `Resolution`. Multiple signals corroborate a
* single match; their weights compose additively to produce `confidence`.
*
* Weights come from `EvidenceWeights` (see `./evidence-weights.ts`).
*/
export interface ResolutionEvidence {
readonly kind:
| 'local'
| 'scope-chain'
| 'import'
| 'type-binding'
| 'owner-match'
| 'kind-match'
| 'arity-match'
| 'global-name'
| 'global-qualified'
| 'dynamic-import-unresolved';
/** Signal weight, sourced from `EvidenceWeights`. Additive; sum capped at 1.0. */
readonly weight: number;
/** Optional debug annotation (e.g., `'matched via self: User'`). */
readonly note?: string;
}
/**
* A ranked resolution candidate returned by `ClassRegistry.lookup` /
* `MethodRegistry.lookup` / `FieldRegistry.lookup`. Evidence composes
* additively; callers read `[0]` for the one-shot answer or inspect the
* evidence trace for debugging.
*/
export interface Resolution {
readonly def: SymbolDefinition;
/** Σ of `evidence[].weight`, capped at 1.0. */
readonly confidence: number;
readonly evidence: readonly ResolutionEvidence[];
/** Optional debug trace: scopes walked to reach `def`. */
readonly path?: readonly ScopeId[];
}
// ─── §2.7 Reference + ReferenceIndex ────────────────────────────────────────
/**
* A post-resolution usage fact: some code at `atRange` inside `fromScope`
* references `toDef` with the given confidence/evidence. Materialized by the
* resolution phase; emitted as graph edges (`CALLS`/`READS`/`WRITES`/etc.)
* during the emit phase.
*/
export interface Reference {
/** Innermost lexical scope containing `atRange`. */
readonly fromScope: ScopeId;
readonly toDef: DefId;
/** Location of the reference in source. */
readonly atRange: Range;
readonly kind: 'call' | 'read' | 'write' | 'type-reference' | 'inherits' | 'import-use';
readonly confidence: number;
readonly evidence: readonly ResolutionEvidence[];
}
/**
* Two-way index over `Reference` records, populated during the resolution
* phase. Scopes stay immutable after finalize; references accumulate here.
*/
export interface ReferenceIndex {
readonly bySourceScope: ReadonlyMap<ScopeId, readonly Reference[]>;
readonly byTargetDef: ReadonlyMap<DefId, readonly Reference[]>;
}
// ─── §4.1 LookupParams ──────────────────────────────────────────────────────
/**
* Opaque placeholder for the per-kind registry passed as the owner-scoped
* contributor. Typed concretely in Ring 2 SHARED (#917); kept as `unknown`
* here so Ring 1 can ship without pulling in the registry implementation.
*/
export type RegistryContributor = unknown;
/**
* Parameters accepted by `Registry.lookup`. Three registries (Class/Method/
* Field) run the same 7-step algorithm with different parameter tuples; see
* RFC §4.4 for per-registry specializations.
*/
export interface LookupParams {
readonly acceptedKinds: readonly NodeLabel[];
/** Class lookups: false. Method/Field lookups: true. */
readonly useReceiverTypeBinding: boolean;
readonly ownerScopedContributor: RegistryContributor | null;
/** Optional arity hint fed to `provider.arityCompatibility`. */
readonly arityHint?: number;
/** Explicit receiver name (e.g., `'user'` in `user.save()`). When present,
* the receiver's type binding at the callsite scope is used; otherwise
* the enclosing method's implicit `self`/`this` is consulted. See §4.1. */
readonly explicitReceiver?: { readonly name: string };
}
+23 -4
View File
@@ -61,13 +61,22 @@ test.beforeAll(async () => {
}
});
// Auto-connect downloads the full graph from the backend; under parallel
// workers in CI the same backend serves multiple downloads concurrently, so
// reaching the "Ready" state can take noticeably longer than a single-worker
// run. Match the 45s budget used by waitForGraphLoaded() in
// server-connect.spec.ts which has been stable on the same backend.
const READY_TIMEOUT_MS = 45_000;
test.describe('Multi-Repo Scoping', () => {
test('auto-connect via ?server= sets ?project= in URL', async ({ page }) => {
// Navigate with ?server= param (the bookmarkable shortcut)
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
// Wait for graph to load
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({
timeout: READY_TIMEOUT_MS,
});
// URL should now contain ?project= with the repo name
const url = new URL(page.url());
@@ -77,8 +86,14 @@ test.describe('Multi-Repo Scoping', () => {
});
test('?server= is preserved in URL for F5 recovery', async ({ page }) => {
// Two sequential auto-connects (initial + reload), each up to READY_TIMEOUT_MS,
// can exceed the default 60s test timeout under parallel workers.
test.slow();
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({
timeout: READY_TIMEOUT_MS,
});
// URL should still have ?server=
const url = new URL(page.url());
@@ -86,12 +101,16 @@ test.describe('Multi-Repo Scoping', () => {
// F5 should reconnect (not show onboarding)
await page.reload();
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({
timeout: READY_TIMEOUT_MS,
});
});
test('node count in status bar matches backend data', async ({ page }) => {
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({
timeout: READY_TIMEOUT_MS,
});
// Fetch expected node count from backend
const res = await fetch(`${BACKEND_URL}/api/repo?repo=${encodeURIComponent(firstRepoName)}`);
+8 -1
View File
@@ -26,7 +26,10 @@ async function enterExploringView(page: import('@playwright/test').Page) {
// Landing screen may not appear (e.g. ?server auto-connect)
}
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// Match the 45s budget used by waitForGraphLoaded() in
// server-connect.spec.ts; under parallel CI workers, downloading the full
// graph can occasionally exceed 30s.
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 45_000 });
}
// ── Flow 1: Onboarding (no server running) ─────────────────────────────────
@@ -244,6 +247,10 @@ test.describe('Flow 3: Analyze form', () => {
test.describe('Flow 4: Repo dropdown in exploring view', () => {
const SKIP_MSG = 'Requires running gitnexus server with indexed repos';
// enterExploringView() can take up to ~45s under parallel CI workers; combined
// with the dropdown interactions this can exceed the default 60s test budget.
test.slow();
test.beforeAll(async () => {
if (process.env.E2E) return;
try {
+26 -5
View File
@@ -84,11 +84,20 @@ test.describe('Hold-queue timeout error', () => {
// ── 2. ?project= URL persistence ─────────────────────────────────────────────
// Auto-connect downloads the full graph from the backend; under parallel
// workers in CI the same backend serves multiple downloads concurrently, so
// reaching the "Ready" state can take noticeably longer than a single-worker
// run. Match the 45s budget used by waitForGraphLoaded() in
// server-connect.spec.ts which has been stable on the same backend.
const READY_TIMEOUT_MS = 45_000;
test.describe('?project= URL persistence', () => {
test('?project= is set in URL after connecting via ?server=', async ({ page }) => {
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({
timeout: READY_TIMEOUT_MS,
});
const url = new URL(page.url());
const project = url.searchParams.get('project');
@@ -98,12 +107,20 @@ test.describe('?project= URL persistence', () => {
});
test('?project= is still present after F5 reload', async ({ page }) => {
// Two sequential auto-connects (initial + reload), each up to READY_TIMEOUT_MS,
// can exceed the default 60s test timeout under parallel workers.
test.slow();
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({
timeout: READY_TIMEOUT_MS,
});
// After connect, URL has ?server=&project= — F5 re-uses both params
await page.reload();
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({
timeout: READY_TIMEOUT_MS,
});
const url = new URL(page.url());
expect(url.searchParams.get('project')).toBeTruthy();
@@ -122,7 +139,9 @@ test.describe('?project= auto-connect', () => {
`/?server=${encodeURIComponent(BACKEND_URL)}&project=${encodeURIComponent(firstRepoName)}`,
);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({
timeout: READY_TIMEOUT_MS,
});
// ?project= in URL should match what we passed in
const url = new URL(page.url());
@@ -155,7 +174,9 @@ test.describe('Windows path normalization', () => {
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({
timeout: READY_TIMEOUT_MS,
});
// URL ?project= must be the short basename, NOT the full Windows path
const url = new URL(page.url());
+166 -397
View File
@@ -29,7 +29,7 @@
"langchain": "^1.2.10",
"lru-cache": "^11.2.4",
"lucide-react": "^0.562.0",
"mermaid": "^11.12.2",
"mermaid": "^11.14.0",
"mnemonist": "^0.39.0",
"pandemonium": "^2.4.0",
"react": "^18.3.1",
@@ -39,12 +39,12 @@
"react-zoom-pan-pinch": "^3.7.0",
"remark-gfm": "^4.0.1",
"sigma": "^3.0.2",
"tailwindcss": "^4.1.18",
"tailwindcss": "^4.2.2",
"uuid": "^13.0.0",
"zod": "^3.25.76"
},
"devDependencies": {
"@babel/types": "^7.28.5",
"@babel/types": "^7.29.0",
"@playwright/test": "^1.58.2",
"@testing-library/jest-dom": "^6.9.1",
"@testing-library/react": "^16.3.2",
@@ -57,12 +57,12 @@
"@vercel/node": "^5.5.16",
"@vitejs/plugin-react": "^5.1.0",
"@vitest/coverage-v8": "^3.2.4",
"jsdom": "^29.0.0",
"jsdom": "^29.0.2",
"tree-sitter-wasms": "^0.1.13",
"typescript": "^5.4.5",
"vite": "^5.2.0",
"vitest": "^3.2.4",
"wait-on": "^8.0.5"
"wait-on": "^9.0.5"
},
"engines": {
"node": ">=20.0.0"
@@ -129,39 +129,49 @@
}
},
"node_modules/@asamuzakjp/css-color": {
"version": "5.0.1",
"resolved": "https://registry.npmjs.org/@asamuzakjp/css-color/-/css-color-5.0.1.tgz",
"integrity": "sha512-2SZFvqMyvboVV1d15lMf7XiI3m7SDqXUuKaTymJYLN6dSGadqp+fVojqJlVoMlbZnlTmu3S0TLwLTJpvBMO1Aw==",
"version": "5.1.11",
"resolved": "https://registry.npmjs.org/@asamuzakjp/css-color/-/css-color-5.1.11.tgz",
"integrity": "sha512-KVw6qIiCTUQhByfTd78h2yD1/00waTmm9uy/R7Ck/ctUyAPj+AEDLkQIdJW0T8+qGgj3j5bpNKK7Q3G+LedJWg==",
"dev": true,
"license": "MIT",
"dependencies": {
"@csstools/css-calc": "^3.1.1",
"@csstools/css-color-parser": "^4.0.2",
"@asamuzakjp/generational-cache": "^1.0.1",
"@csstools/css-calc": "^3.2.0",
"@csstools/css-color-parser": "^4.1.0",
"@csstools/css-parser-algorithms": "^4.0.0",
"@csstools/css-tokenizer": "^4.0.0",
"lru-cache": "^11.2.6"
"@csstools/css-tokenizer": "^4.0.0"
},
"engines": {
"node": "^20.19.0 || ^22.12.0 || >=24.0.0"
}
},
"node_modules/@asamuzakjp/dom-selector": {
"version": "7.0.3",
"resolved": "https://registry.npmjs.org/@asamuzakjp/dom-selector/-/dom-selector-7.0.3.tgz",
"integrity": "sha512-Q6mU0Z6bfj6YvnX2k9n0JxiIwrCFN59x/nWmYQnAqP000ruX/yV+5bp/GRcF5T8ncvfwJQ7fgfP74DlpKExILA==",
"version": "7.0.10",
"resolved": "https://registry.npmjs.org/@asamuzakjp/dom-selector/-/dom-selector-7.0.10.tgz",
"integrity": "sha512-KyOb19eytNSELkmdqzZZUXWCU25byIlOld5qVFg0RYdS0T3tt7jeDByxk9hIAC73frclD8GKrHttr0SUjKCCdQ==",
"dev": true,
"license": "MIT",
"dependencies": {
"@asamuzakjp/generational-cache": "^1.0.1",
"@asamuzakjp/nwsapi": "^2.3.9",
"bidi-js": "^1.0.3",
"css-tree": "^3.2.1",
"is-potential-custom-element-name": "^1.0.1",
"lru-cache": "^11.2.7"
"is-potential-custom-element-name": "^1.0.1"
},
"engines": {
"node": "^20.19.0 || ^22.12.0 || >=24.0.0"
}
},
"node_modules/@asamuzakjp/generational-cache": {
"version": "1.0.1",
"resolved": "https://registry.npmjs.org/@asamuzakjp/generational-cache/-/generational-cache-1.0.1.tgz",
"integrity": "sha512-wajfB8KqzMCN2KGNFdLkReeHncd0AslUSrvHVvvYWuU8ghncRJoA50kT3zP9MVL0+9g4/67H+cdvBskj9THPzg==",
"dev": true,
"license": "MIT",
"engines": {
"node": "^20.19.0 || ^22.12.0 || >=24.0.0"
}
},
"node_modules/@asamuzakjp/nwsapi": {
"version": "2.3.9",
"resolved": "https://registry.npmjs.org/@asamuzakjp/nwsapi/-/nwsapi-2.3.9.tgz",
@@ -484,9 +494,9 @@
}
},
"node_modules/@babel/types": {
"version": "7.28.6",
"resolved": "https://registry.npmjs.org/@babel/types/-/types-7.28.6.tgz",
"integrity": "sha512-0ZrskXVEHSWIqZM/sQZ4EV3jZJXRkio/WCxaqKZP1g//CEWEPSfeZFcms4XeKBCHU0ZKnIkdJeU/kF+eRp5lBg==",
"version": "7.29.0",
"resolved": "https://registry.npmjs.org/@babel/types/-/types-7.29.0.tgz",
"integrity": "sha512-LwdZHpScM4Qz8Xw2iKSzS+cfglZzJGvofQICy7W7v4caru4EaAmyUuO6BGrbyQ2mYV11W0U8j5mBhd14dd3B0A==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -533,54 +543,40 @@
"license": "MIT"
},
"node_modules/@chevrotain/cst-dts-gen": {
"version": "11.0.3",
"resolved": "https://registry.npmjs.org/@chevrotain/cst-dts-gen/-/cst-dts-gen-11.0.3.tgz",
"integrity": "sha512-BvIKpRLeS/8UbfxXxgC33xOumsacaeCKAjAeLyOn7Pcp95HiRbrpl14S+9vaZLolnbssPIUuiUd8IvgkRyt6NQ==",
"version": "12.0.0",
"resolved": "https://registry.npmjs.org/@chevrotain/cst-dts-gen/-/cst-dts-gen-12.0.0.tgz",
"integrity": "sha512-fSL4KXjTl7cDgf0B5Rip9Q05BOrYvkJV/RrBTE/bKDN096E4hN/ySpcBK5B24T76dlQ2i32Zc3PAE27jFnFrKg==",
"license": "Apache-2.0",
"dependencies": {
"@chevrotain/gast": "11.0.3",
"@chevrotain/types": "11.0.3",
"lodash-es": "4.17.21"
"@chevrotain/gast": "12.0.0",
"@chevrotain/types": "12.0.0"
}
},
"node_modules/@chevrotain/cst-dts-gen/node_modules/lodash-es": {
"version": "4.17.21",
"resolved": "https://registry.npmjs.org/lodash-es/-/lodash-es-4.17.21.tgz",
"integrity": "sha512-mKnC+QJ9pWVzv+C4/U3rRsHapFfHvQFoFB92e52xeyGMcX6/OlIl78je1u8vePzYZSkkogMPJ2yjxxsb89cxyw==",
"license": "MIT"
},
"node_modules/@chevrotain/gast": {
"version": "11.0.3",
"resolved": "https://registry.npmjs.org/@chevrotain/gast/-/gast-11.0.3.tgz",
"integrity": "sha512-+qNfcoNk70PyS/uxmj3li5NiECO+2YKZZQMbmjTqRI3Qchu8Hig/Q9vgkHpI3alNjr7M+a2St5pw5w5F6NL5/Q==",
"version": "12.0.0",
"resolved": "https://registry.npmjs.org/@chevrotain/gast/-/gast-12.0.0.tgz",
"integrity": "sha512-1ne/m3XsIT8aEdrvT33so0GUC+wkctpUPK6zU9IlOyJLUbR0rg4G7ZiApiJbggpgPir9ERy3FRjT6T7lpgetnQ==",
"license": "Apache-2.0",
"dependencies": {
"@chevrotain/types": "11.0.3",
"lodash-es": "4.17.21"
"@chevrotain/types": "12.0.0"
}
},
"node_modules/@chevrotain/gast/node_modules/lodash-es": {
"version": "4.17.21",
"resolved": "https://registry.npmjs.org/lodash-es/-/lodash-es-4.17.21.tgz",
"integrity": "sha512-mKnC+QJ9pWVzv+C4/U3rRsHapFfHvQFoFB92e52xeyGMcX6/OlIl78je1u8vePzYZSkkogMPJ2yjxxsb89cxyw==",
"license": "MIT"
},
"node_modules/@chevrotain/regexp-to-ast": {
"version": "11.0.3",
"resolved": "https://registry.npmjs.org/@chevrotain/regexp-to-ast/-/regexp-to-ast-11.0.3.tgz",
"integrity": "sha512-1fMHaBZxLFvWI067AVbGJav1eRY7N8DDvYCTwGBiE/ytKBgP8azTdgyrKyWZ9Mfh09eHWb5PgTSO8wi7U824RA==",
"version": "12.0.0",
"resolved": "https://registry.npmjs.org/@chevrotain/regexp-to-ast/-/regexp-to-ast-12.0.0.tgz",
"integrity": "sha512-p+EW9MaJwgaHguhoqwOtx/FwuGr+DnNn857sXWOi/mClXIkPGl3rn7hGNWvo31HA3vyeQxjqe+H36yZJwYU8cA==",
"license": "Apache-2.0"
},
"node_modules/@chevrotain/types": {
"version": "11.0.3",
"resolved": "https://registry.npmjs.org/@chevrotain/types/-/types-11.0.3.tgz",
"integrity": "sha512-gsiM3G8b58kZC2HaWR50gu6Y1440cHiJ+i3JUvcp/35JchYejb2+5MVeJK0iKThYpAa/P2PYFV4hoi44HD+aHQ==",
"version": "12.0.0",
"resolved": "https://registry.npmjs.org/@chevrotain/types/-/types-12.0.0.tgz",
"integrity": "sha512-S+04vjFQKeuYw0/eW3U52LkAHQsB1ASxsPGsLPUyQgrZ2iNNibQrsidruDzjEX2JYfespXMG0eZmXlhA6z7nWA==",
"license": "Apache-2.0"
},
"node_modules/@chevrotain/utils": {
"version": "11.0.3",
"resolved": "https://registry.npmjs.org/@chevrotain/utils/-/utils-11.0.3.tgz",
"integrity": "sha512-YslZMgtJUyuMbZ+aKvfF3x1f5liK4mWNxghFRv7jqRR9C3R3fAOGTTKvxXDa2Y1s9zSbcpuO0cAxDYsc9SrXoQ==",
"version": "12.0.0",
"resolved": "https://registry.npmjs.org/@chevrotain/utils/-/utils-12.0.0.tgz",
"integrity": "sha512-lB59uJoaGIfOOL9knQqQRfhl9g7x8/wqFkp13zTdkRu1huG9kg6IJs1O8hqj9rs6h7orGxHJUKb+mX3rPbWGhA==",
"license": "Apache-2.0"
},
"node_modules/@cspotcode/source-map-support": {
@@ -628,9 +624,9 @@
}
},
"node_modules/@csstools/css-calc": {
"version": "3.1.1",
"resolved": "https://registry.npmjs.org/@csstools/css-calc/-/css-calc-3.1.1.tgz",
"integrity": "sha512-HJ26Z/vmsZQqs/o3a6bgKslXGFAungXGbinULZO3eMsOyNJHeBBZfup5FiZInOghgoM4Hwnmw+OgbJCNg1wwUQ==",
"version": "3.2.0",
"resolved": "https://registry.npmjs.org/@csstools/css-calc/-/css-calc-3.2.0.tgz",
"integrity": "sha512-bR9e6o2BDB12jzN/gIbjHa5wLJ4UjD1CB9pM7ehlc0ddk6EBz+yYS1EV2MF55/HUxrHcB/hehAyt5vhsA3hx7w==",
"dev": true,
"funding": [
{
@@ -652,9 +648,9 @@
}
},
"node_modules/@csstools/css-color-parser": {
"version": "4.0.2",
"resolved": "https://registry.npmjs.org/@csstools/css-color-parser/-/css-color-parser-4.0.2.tgz",
"integrity": "sha512-0GEfbBLmTFf0dJlpsNU7zwxRIH0/BGEMuXLTCvFYxuL1tNhqzTbtnFICyJLTNK4a+RechKP75e7w42ClXSnJQw==",
"version": "4.1.0",
"resolved": "https://registry.npmjs.org/@csstools/css-color-parser/-/css-color-parser-4.1.0.tgz",
"integrity": "sha512-U0KhLYmy2GVj6q4T3WaAe6NPuFYCPQoE3b0dRGxejWDgcPp8TP7S5rVdM5ZrFaqu4N67X8YaPBw14dQSYx3IyQ==",
"dev": true,
"funding": [
{
@@ -669,7 +665,7 @@
"license": "MIT",
"dependencies": {
"@csstools/color-helpers": "^6.0.2",
"@csstools/css-calc": "^3.1.1"
"@csstools/css-calc": "^3.2.0"
},
"engines": {
"node": ">=20.19.0"
@@ -1764,12 +1760,12 @@
}
},
"node_modules/@mermaid-js/parser": {
"version": "0.6.3",
"resolved": "https://registry.npmjs.org/@mermaid-js/parser/-/parser-0.6.3.tgz",
"integrity": "sha512-lnjOhe7zyHjc+If7yT4zoedx2vo4sHaTmtkl1+or8BRTnCtDmcTpAjpzDSfCZrshM5bCoz0GyidzadJAH1xobA==",
"version": "1.1.0",
"resolved": "https://registry.npmjs.org/@mermaid-js/parser/-/parser-1.1.0.tgz",
"integrity": "sha512-gxK9ZX2+Fex5zu8LhRQoMeMPEHbc73UKZ0FQ54YrQtUxE1VVhMwzeNtKRPAu5aXks4FasbMe4xB4bWrmq6Jlxw==",
"license": "MIT",
"dependencies": {
"langium": "3.3.1"
"langium": "^4.0.0"
}
},
"node_modules/@napi-rs/wasm-runtime": {
@@ -2224,257 +2220,6 @@
"dev": true,
"license": "MIT"
},
"node_modules/@swc/core": {
"version": "1.15.8",
"resolved": "https://registry.npmjs.org/@swc/core/-/core-1.15.8.tgz",
"integrity": "sha512-T8keoJjXaSUoVBCIjgL6wAnhADIb09GOELzKg10CjNg+vLX48P93SME6jTfte9MZIm5m+Il57H3rTSk/0kzDUw==",
"dev": true,
"hasInstallScript": true,
"license": "Apache-2.0",
"optional": true,
"peer": true,
"dependencies": {
"@swc/counter": "^0.1.3",
"@swc/types": "^0.1.25"
},
"engines": {
"node": ">=10"
},
"funding": {
"type": "opencollective",
"url": "https://opencollective.com/swc"
},
"optionalDependencies": {
"@swc/core-darwin-arm64": "1.15.8",
"@swc/core-darwin-x64": "1.15.8",
"@swc/core-linux-arm-gnueabihf": "1.15.8",
"@swc/core-linux-arm64-gnu": "1.15.8",
"@swc/core-linux-arm64-musl": "1.15.8",
"@swc/core-linux-x64-gnu": "1.15.8",
"@swc/core-linux-x64-musl": "1.15.8",
"@swc/core-win32-arm64-msvc": "1.15.8",
"@swc/core-win32-ia32-msvc": "1.15.8",
"@swc/core-win32-x64-msvc": "1.15.8"
},
"peerDependencies": {
"@swc/helpers": ">=0.5.17"
},
"peerDependenciesMeta": {
"@swc/helpers": {
"optional": true
}
}
},
"node_modules/@swc/core-darwin-arm64": {
"version": "1.15.8",
"resolved": "https://registry.npmjs.org/@swc/core-darwin-arm64/-/core-darwin-arm64-1.15.8.tgz",
"integrity": "sha512-M9cK5GwyWWRkRGwwCbREuj6r8jKdES/haCZ3Xckgkl8MUQJZA3XB7IXXK1IXRNeLjg6m7cnoMICpXv1v1hlJOg==",
"cpu": [
"arm64"
],
"dev": true,
"license": "Apache-2.0 AND MIT",
"optional": true,
"os": [
"darwin"
],
"peer": true,
"engines": {
"node": ">=10"
}
},
"node_modules/@swc/core-darwin-x64": {
"version": "1.15.8",
"resolved": "https://registry.npmjs.org/@swc/core-darwin-x64/-/core-darwin-x64-1.15.8.tgz",
"integrity": "sha512-j47DasuOvXl80sKJHSi2X25l44CMc3VDhlJwA7oewC1nV1VsSzwX+KOwE5tLnfORvVJJyeiXgJORNYg4jeIjYQ==",
"cpu": [
"x64"
],
"dev": true,
"license": "Apache-2.0 AND MIT",
"optional": true,
"os": [
"darwin"
],
"peer": true,
"engines": {
"node": ">=10"
}
},
"node_modules/@swc/core-linux-arm-gnueabihf": {
"version": "1.15.8",
"resolved": "https://registry.npmjs.org/@swc/core-linux-arm-gnueabihf/-/core-linux-arm-gnueabihf-1.15.8.tgz",
"integrity": "sha512-siAzDENu2rUbwr9+fayWa26r5A9fol1iORG53HWxQL1J8ym4k7xt9eME0dMPXlYZDytK5r9sW8zEA10F2U3Xwg==",
"cpu": [
"arm"
],
"dev": true,
"license": "Apache-2.0",
"optional": true,
"os": [
"linux"
],
"peer": true,
"engines": {
"node": ">=10"
}
},
"node_modules/@swc/core-linux-arm64-gnu": {
"version": "1.15.8",
"resolved": "https://registry.npmjs.org/@swc/core-linux-arm64-gnu/-/core-linux-arm64-gnu-1.15.8.tgz",
"integrity": "sha512-o+1y5u6k2FfPYbTRUPvurwzNt5qd0NTumCTFscCNuBksycloXY16J8L+SMW5QRX59n4Hp9EmFa3vpvNHRVv1+Q==",
"cpu": [
"arm64"
],
"dev": true,
"license": "Apache-2.0 AND MIT",
"optional": true,
"os": [
"linux"
],
"peer": true,
"engines": {
"node": ">=10"
}
},
"node_modules/@swc/core-linux-arm64-musl": {
"version": "1.15.8",
"resolved": "https://registry.npmjs.org/@swc/core-linux-arm64-musl/-/core-linux-arm64-musl-1.15.8.tgz",
"integrity": "sha512-koiCqL09EwOP1S2RShCI7NbsQuG6r2brTqUYE7pV7kZm9O17wZ0LSz22m6gVibpwEnw8jI3IE1yYsQTVpluALw==",
"cpu": [
"arm64"
],
"dev": true,
"license": "Apache-2.0 AND MIT",
"optional": true,
"os": [
"linux"
],
"peer": true,
"engines": {
"node": ">=10"
}
},
"node_modules/@swc/core-linux-x64-gnu": {
"version": "1.15.8",
"resolved": "https://registry.npmjs.org/@swc/core-linux-x64-gnu/-/core-linux-x64-gnu-1.15.8.tgz",
"integrity": "sha512-4p6lOMU3bC+Vd5ARtKJ/FxpIC5G8v3XLoPEZ5s7mLR8h7411HWC/LmTXDHcrSXRC55zvAVia1eldy6zDLz8iFQ==",
"cpu": [
"x64"
],
"dev": true,
"license": "Apache-2.0 AND MIT",
"optional": true,
"os": [
"linux"
],
"peer": true,
"engines": {
"node": ">=10"
}
},
"node_modules/@swc/core-linux-x64-musl": {
"version": "1.15.8",
"resolved": "https://registry.npmjs.org/@swc/core-linux-x64-musl/-/core-linux-x64-musl-1.15.8.tgz",
"integrity": "sha512-z3XBnbrZAL+6xDGAhJoN4lOueIxC/8rGrJ9tg+fEaeqLEuAtHSW2QHDHxDwkxZMjuF/pZ6MUTjHjbp8wLbuRLA==",
"cpu": [
"x64"
],
"dev": true,
"license": "Apache-2.0 AND MIT",
"optional": true,
"os": [
"linux"
],
"peer": true,
"engines": {
"node": ">=10"
}
},
"node_modules/@swc/core-win32-arm64-msvc": {
"version": "1.15.8",
"resolved": "https://registry.npmjs.org/@swc/core-win32-arm64-msvc/-/core-win32-arm64-msvc-1.15.8.tgz",
"integrity": "sha512-djQPJ9Rh9vP8GTS/Df3hcc6XP6xnG5c8qsngWId/BLA9oX6C7UzCPAn74BG/wGb9a6j4w3RINuoaieJB3t+7iQ==",
"cpu": [
"arm64"
],
"dev": true,
"license": "Apache-2.0 AND MIT",
"optional": true,
"os": [
"win32"
],
"peer": true,
"engines": {
"node": ">=10"
}
},
"node_modules/@swc/core-win32-ia32-msvc": {
"version": "1.15.8",
"resolved": "https://registry.npmjs.org/@swc/core-win32-ia32-msvc/-/core-win32-ia32-msvc-1.15.8.tgz",
"integrity": "sha512-/wfAgxORg2VBaUoFdytcVBVCgf1isWZIEXB9MZEUty4wwK93M/PxAkjifOho9RN3WrM3inPLabICRCEgdHpKKQ==",
"cpu": [
"ia32"
],
"dev": true,
"license": "Apache-2.0 AND MIT",
"optional": true,
"os": [
"win32"
],
"peer": true,
"engines": {
"node": ">=10"
}
},
"node_modules/@swc/core-win32-x64-msvc": {
"version": "1.15.8",
"resolved": "https://registry.npmjs.org/@swc/core-win32-x64-msvc/-/core-win32-x64-msvc-1.15.8.tgz",
"integrity": "sha512-GpMePrh9Sl4d61o4KAHOOv5is5+zt6BEXCOCgs/H0FLGeii7j9bWDE8ExvKFy2GRRZVNR1ugsnzaGWHKM6kuzA==",
"cpu": [
"x64"
],
"dev": true,
"license": "Apache-2.0 AND MIT",
"optional": true,
"os": [
"win32"
],
"peer": true,
"engines": {
"node": ">=10"
}
},
"node_modules/@swc/counter": {
"version": "0.1.3",
"resolved": "https://registry.npmjs.org/@swc/counter/-/counter-0.1.3.tgz",
"integrity": "sha512-e2BR4lsJkkRlKZ/qCHPw9ZaSxc0MVUd7gtbtaB7aMvHeJVYe8sOB8DBZkP2DtISHGSku9sCK6T6cnY0CtXrOCQ==",
"dev": true,
"license": "Apache-2.0",
"optional": true,
"peer": true
},
"node_modules/@swc/types": {
"version": "0.1.25",
"resolved": "https://registry.npmjs.org/@swc/types/-/types-0.1.25.tgz",
"integrity": "sha512-iAoY/qRhNH8a/hBvm3zKj9qQ4oc2+3w1unPJa2XvTK3XjeLXtzcCingVPw/9e5mn1+0yPqxcBGp9Jf0pkfMb1g==",
"dev": true,
"license": "Apache-2.0",
"optional": true,
"peer": true,
"dependencies": {
"@swc/counter": "^0.1.3"
}
},
"node_modules/@swc/wasm": {
"version": "1.15.8",
"resolved": "https://registry.npmjs.org/@swc/wasm/-/wasm-1.15.8.tgz",
"integrity": "sha512-RG2BxGbbsjtddFCo1ghKH6A/BMXbY1eMBfpysV0lJMCpI4DZOjW1BNBnxvBt7YsYmlJtmy5UXIg9/4ekBTFFaQ==",
"dev": true,
"license": "Apache-2.0",
"optional": true,
"peer": true
},
"node_modules/@tailwindcss/node": {
"version": "4.1.18",
"resolved": "https://registry.npmjs.org/@tailwindcss/node/-/node-4.1.18.tgz",
@@ -2490,6 +2235,12 @@
"tailwindcss": "4.1.18"
}
},
"node_modules/@tailwindcss/node/node_modules/tailwindcss": {
"version": "4.1.18",
"resolved": "https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.1.18.tgz",
"integrity": "sha512-4+Z+0yiYyEtUVCScyfHCxOYP06L5Ne+JiHhY2IjR2KWMIWhJOYZKLSGZaP5HkZ8+bY0cxfzwDE5uOmzFXyIwxw==",
"license": "MIT"
},
"node_modules/@tailwindcss/oxide": {
"version": "4.1.18",
"resolved": "https://registry.npmjs.org/@tailwindcss/oxide/-/oxide-4.1.18.tgz",
@@ -2732,6 +2483,12 @@
"vite": "^5.2.0 || ^6 || ^7"
}
},
"node_modules/@tailwindcss/vite/node_modules/tailwindcss": {
"version": "4.1.18",
"resolved": "https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.1.18.tgz",
"integrity": "sha512-4+Z+0yiYyEtUVCScyfHCxOYP06L5Ne+JiHhY2IjR2KWMIWhJOYZKLSGZaP5HkZ8+bY0cxfzwDE5uOmzFXyIwxw==",
"license": "MIT"
},
"node_modules/@testing-library/dom": {
"version": "10.4.1",
"resolved": "https://registry.npmjs.org/@testing-library/dom/-/dom-10.4.1.tgz",
@@ -3358,6 +3115,16 @@
"integrity": "sha512-WmoN8qaIAo7WTYWbAZuG8PYEhn5fkz7dZrqTBZ7dtt//lL2Gwms1IcnQ5yHqjDfX8Ft5j4YzDM23f87zBfDe9g==",
"license": "ISC"
},
"node_modules/@upsetjs/venn.js": {
"version": "2.0.0",
"resolved": "https://registry.npmjs.org/@upsetjs/venn.js/-/venn.js-2.0.0.tgz",
"integrity": "sha512-WbBhLrooyePuQ1VZxrJjtLvTc4NVfpOyKx0sKqioq9bX1C1m7Jgykkn8gLrtwumBioXIqam8DLxp88Adbue6Hw==",
"license": "MIT",
"optionalDependencies": {
"d3-selection": "^3.0.0",
"d3-transition": "^3.0.1"
}
},
"node_modules/@vercel/build-utils": {
"version": "13.2.11",
"resolved": "https://registry.npmjs.org/@vercel/build-utils/-/build-utils-13.2.11.tgz",
@@ -3829,14 +3596,14 @@
"license": "MIT"
},
"node_modules/axios": {
"version": "1.13.2",
"resolved": "https://registry.npmjs.org/axios/-/axios-1.13.2.tgz",
"integrity": "sha512-VPk9ebNqPcy5lRGuSlKx752IlDatOjT9paPlm8A7yOuW2Fbvp4X3JznJtT4f0GzGLLiWE9W8onz51SqLYwzGaA==",
"version": "1.15.0",
"resolved": "https://registry.npmjs.org/axios/-/axios-1.15.0.tgz",
"integrity": "sha512-wWyJDlAatxk30ZJer+GeCWS209sA42X+N5jU2jy6oHTp7ufw8uzUTVFBX9+wTfAlhiJXGS0Bq7X6efruWjuK9Q==",
"license": "MIT",
"dependencies": {
"follow-redirects": "^1.15.6",
"form-data": "^4.0.4",
"proxy-from-env": "^1.1.0"
"follow-redirects": "^1.15.11",
"form-data": "^4.0.5",
"proxy-from-env": "^2.1.0"
}
},
"node_modules/bail": {
@@ -4129,37 +3896,33 @@
}
},
"node_modules/chevrotain": {
"version": "11.0.3",
"resolved": "https://registry.npmjs.org/chevrotain/-/chevrotain-11.0.3.tgz",
"integrity": "sha512-ci2iJH6LeIkvP9eJW6gpueU8cnZhv85ELY8w8WiFtNjMHA5ad6pQLaJo9mEly/9qUyCpvqX8/POVUTf18/HFdw==",
"version": "12.0.0",
"resolved": "https://registry.npmjs.org/chevrotain/-/chevrotain-12.0.0.tgz",
"integrity": "sha512-csJvb+6kEiQaqo1woTdSAuOWdN0WTLIydkKrBnS+V5gZz0oqBrp4kQ35519QgK6TpBThiG3V1vNSHlIkv4AglQ==",
"license": "Apache-2.0",
"dependencies": {
"@chevrotain/cst-dts-gen": "11.0.3",
"@chevrotain/gast": "11.0.3",
"@chevrotain/regexp-to-ast": "11.0.3",
"@chevrotain/types": "11.0.3",
"@chevrotain/utils": "11.0.3",
"lodash-es": "4.17.21"
"@chevrotain/cst-dts-gen": "12.0.0",
"@chevrotain/gast": "12.0.0",
"@chevrotain/regexp-to-ast": "12.0.0",
"@chevrotain/types": "12.0.0",
"@chevrotain/utils": "12.0.0"
},
"engines": {
"node": ">=22.0.0"
}
},
"node_modules/chevrotain-allstar": {
"version": "0.3.1",
"resolved": "https://registry.npmjs.org/chevrotain-allstar/-/chevrotain-allstar-0.3.1.tgz",
"integrity": "sha512-b7g+y9A0v4mxCW1qUhf3BSVPg+/NvGErk/dOkrDaHA0nQIQGAtrOjlX//9OQtRlSCy+x9rfB5N8yC71lH1nvMw==",
"version": "0.4.1",
"resolved": "https://registry.npmjs.org/chevrotain-allstar/-/chevrotain-allstar-0.4.1.tgz",
"integrity": "sha512-PvVJm3oGqrveUVW2Vt/eZGeiAIsJszYweUcYwcskg9e+IubNYKKD+rHHem7A6XVO22eDAL+inxNIGAzZ/VIWlA==",
"license": "MIT",
"dependencies": {
"lodash-es": "^4.17.21"
},
"peerDependencies": {
"chevrotain": "^11.0.0"
"chevrotain": "^12.0.0"
}
},
"node_modules/chevrotain/node_modules/lodash-es": {
"version": "4.17.21",
"resolved": "https://registry.npmjs.org/lodash-es/-/lodash-es-4.17.21.tgz",
"integrity": "sha512-mKnC+QJ9pWVzv+C4/U3rRsHapFfHvQFoFB92e52xeyGMcX6/OlIl78je1u8vePzYZSkkogMPJ2yjxxsb89cxyw==",
"license": "MIT"
},
"node_modules/chownr": {
"version": "3.0.0",
"resolved": "https://registry.npmjs.org/chownr/-/chownr-3.0.0.tgz",
@@ -4830,9 +4593,9 @@
}
},
"node_modules/dagre-d3-es": {
"version": "7.0.13",
"resolved": "https://registry.npmjs.org/dagre-d3-es/-/dagre-d3-es-7.0.13.tgz",
"integrity": "sha512-efEhnxpSuwpYOKRm/L5KbqoZmNNukHa/Flty4Wp62JRvgH2ojwVgPgdYyr4twpieZnyRDdIH7PY2mopX26+j2Q==",
"version": "7.0.14",
"resolved": "https://registry.npmjs.org/dagre-d3-es/-/dagre-d3-es-7.0.14.tgz",
"integrity": "sha512-P4rFMVq9ESWqmOgK+dlXvOtLwYg0i7u0HBGJER0LZDJT2VHIPAMZ/riPxqJceWMStH5+E61QxFra9kIS3AqdMg==",
"license": "MIT",
"dependencies": {
"d3": "^7.9.0",
@@ -6064,9 +5827,9 @@
}
},
"node_modules/joi": {
"version": "18.0.2",
"resolved": "https://registry.npmjs.org/joi/-/joi-18.0.2.tgz",
"integrity": "sha512-RuCOQMIt78LWnktPoeBL0GErkNaJPTBGcYuyaBvUOQSpcpcLfWrHPPihYdOGbV5pam9VTWbeoF7TsGiHugcjGA==",
"version": "18.1.2",
"resolved": "https://registry.npmjs.org/joi/-/joi-18.1.2.tgz",
"integrity": "sha512-rF5MAmps5esSlhCA+N1b6IYHDw9j/btzGaqfgie522jS02Ju/HXBxamlXVlKEHAxoMKQL77HWI8jlqWsFuekZA==",
"dev": true,
"license": "BSD-3-Clause",
"dependencies": {
@@ -6076,7 +5839,7 @@
"@hapi/pinpoint": "^2.0.1",
"@hapi/tlds": "^1.1.1",
"@hapi/topo": "^6.0.2",
"@standard-schema/spec": "^1.0.0"
"@standard-schema/spec": "^1.1.0"
},
"engines": {
"node": ">= 20"
@@ -6098,14 +5861,14 @@
"license": "MIT"
},
"node_modules/jsdom": {
"version": "29.0.0",
"resolved": "https://registry.npmjs.org/jsdom/-/jsdom-29.0.0.tgz",
"integrity": "sha512-9FshNB6OepopZ08unmmGpsF7/qCjxGPbo3NbgfJAnPeHXnsODE9WWffXZtRFRFe0ntzaAOcSKNJFz8wiyvF1jQ==",
"version": "29.0.2",
"resolved": "https://registry.npmjs.org/jsdom/-/jsdom-29.0.2.tgz",
"integrity": "sha512-9VnGEBosc/ZpwyOsJBCQ/3I5p7Q5ngOY14a9bf5btenAORmZfDse1ZEheMiWcJ3h81+Fv7HmJFdS0szo/waF2w==",
"dev": true,
"license": "MIT",
"dependencies": {
"@asamuzakjp/css-color": "^5.0.1",
"@asamuzakjp/dom-selector": "^7.0.2",
"@asamuzakjp/css-color": "^5.1.5",
"@asamuzakjp/dom-selector": "^7.0.6",
"@bramus/specificity": "^2.4.2",
"@csstools/css-syntax-patches-for-csstree": "^1.1.1",
"@exodus/bytes": "^1.15.0",
@@ -6119,7 +5882,7 @@
"saxes": "^6.0.0",
"symbol-tree": "^3.2.4",
"tough-cookie": "^6.0.1",
"undici": "^7.24.3",
"undici": "^7.24.5",
"w3c-xmlserializer": "^5.0.0",
"webidl-conversions": "^8.0.1",
"whatwg-mimetype": "^5.0.0",
@@ -6152,9 +5915,9 @@
}
},
"node_modules/jsdom/node_modules/undici": {
"version": "7.24.3",
"resolved": "https://registry.npmjs.org/undici/-/undici-7.24.3.tgz",
"integrity": "sha512-eJdUmK/Wrx2d+mnWWmwwLRyA7OQCkLap60sk3dOK4ViZR7DKwwptwuIvFBg2HaiP9ESaEdhtpSymQPvytpmkCA==",
"version": "7.25.0",
"resolved": "https://registry.npmjs.org/undici/-/undici-7.25.0.tgz",
"integrity": "sha512-xXnp4kTyor2Zq+J1FfPI6Eq3ew5h6Vl0F/8d9XU5zZQf1tX9s2Su1/3PiMmUANFULpmksxkClamIZcaUqryHsQ==",
"dev": true,
"license": "MIT",
"engines": {
@@ -6295,19 +6058,21 @@
}
},
"node_modules/langium": {
"version": "3.3.1",
"resolved": "https://registry.npmjs.org/langium/-/langium-3.3.1.tgz",
"integrity": "sha512-QJv/h939gDpvT+9SiLVlY7tZC3xB2qK57v0J04Sh9wpMb6MP1q8gB21L3WIo8T5P1MSMg3Ep14L7KkDCFG3y4w==",
"version": "4.2.2",
"resolved": "https://registry.npmjs.org/langium/-/langium-4.2.2.tgz",
"integrity": "sha512-JUshTRAfHI4/MF9dH2WupvjSXyn8JBuUEWazB8ZVJUtXutT0doDlAv1XKbZ1Pb5sMexa8FF4CFBc0iiul7gbUQ==",
"license": "MIT",
"dependencies": {
"chevrotain": "~11.0.3",
"chevrotain-allstar": "~0.3.0",
"@chevrotain/regexp-to-ast": "~12.0.0",
"chevrotain": "~12.0.0",
"chevrotain-allstar": "~0.4.1",
"vscode-languageserver": "~9.0.1",
"vscode-languageserver-textdocument": "~1.0.11",
"vscode-uri": "~3.0.8"
"vscode-uri": "~3.1.0"
},
"engines": {
"node": ">=16.0.0"
"node": ">=20.10.0",
"npm": ">=10.2.3"
}
},
"node_modules/langsmith": {
@@ -6613,16 +6378,16 @@
}
},
"node_modules/lodash": {
"version": "4.17.23",
"resolved": "https://registry.npmjs.org/lodash/-/lodash-4.17.23.tgz",
"integrity": "sha512-LgVTMpQtIopCi79SJeDiP0TfWi5CNEc/L/aRdTh3yIvmZXTnheWpKjSZhnvMl8iXbC1tFg9gdHHDMLoV7CnG+w==",
"version": "4.18.1",
"resolved": "https://registry.npmjs.org/lodash/-/lodash-4.18.1.tgz",
"integrity": "sha512-dMInicTPVE8d1e5otfwmmjlxkZoUpiVLwyeTdUsi/Caj/gfzzblBcCE5sRHV/AsjuCmxWrte2TNGSYuCeCq+0Q==",
"dev": true,
"license": "MIT"
},
"node_modules/lodash-es": {
"version": "4.17.22",
"resolved": "https://registry.npmjs.org/lodash-es/-/lodash-es-4.17.22.tgz",
"integrity": "sha512-XEawp1t0gxSi9x01glktRZ5HDy0HXqrM0x5pXQM98EaI0NxO6jVM7omDOxsuEo5UIASAnm2bRp1Jt/e0a2XU8Q==",
"version": "4.18.1",
"resolved": "https://registry.npmjs.org/lodash-es/-/lodash-es-4.18.1.tgz",
"integrity": "sha512-J8xewKD/Gk22OZbhpOVSwcs60zhd95ESDwezOFuA3/099925PdHJ7OFHNTGtajL3AlZkykD32HykiMo+BIBI8A==",
"license": "MIT"
},
"node_modules/longest-streak": {
@@ -7072,27 +6837,28 @@
}
},
"node_modules/mermaid": {
"version": "11.12.2",
"resolved": "https://registry.npmjs.org/mermaid/-/mermaid-11.12.2.tgz",
"integrity": "sha512-n34QPDPEKmaeCG4WDMGy0OT6PSyxKCfy2pJgShP+Qow2KLrvWjclwbc3yXfSIf4BanqWEhQEpngWwNp/XhZt6w==",
"version": "11.14.0",
"resolved": "https://registry.npmjs.org/mermaid/-/mermaid-11.14.0.tgz",
"integrity": "sha512-GSGloRsBs+JINmmhl0JDwjpuezCsHB4WGI4NASHxL3fHo3o/BRXTxhDLKnln8/Q0lRFRyDdEjmk1/d5Sn1Xz8g==",
"license": "MIT",
"dependencies": {
"@braintree/sanitize-url": "^7.1.1",
"@iconify/utils": "^3.0.1",
"@mermaid-js/parser": "^0.6.3",
"@iconify/utils": "^3.0.2",
"@mermaid-js/parser": "^1.1.0",
"@types/d3": "^7.4.3",
"cytoscape": "^3.29.3",
"@upsetjs/venn.js": "^2.0.0",
"cytoscape": "^3.33.1",
"cytoscape-cose-bilkent": "^4.1.0",
"cytoscape-fcose": "^2.2.0",
"d3": "^7.9.0",
"d3-sankey": "^0.12.3",
"dagre-d3-es": "7.0.13",
"dayjs": "^1.11.18",
"dompurify": "^3.2.5",
"katex": "^0.16.22",
"dagre-d3-es": "7.0.14",
"dayjs": "^1.11.19",
"dompurify": "^3.3.1",
"katex": "^0.16.25",
"khroma": "^2.1.0",
"lodash-es": "^4.17.21",
"marked": "^16.2.1",
"lodash-es": "^4.17.23",
"marked": "^16.3.0",
"roughjs": "^4.6.6",
"stylis": "^4.3.6",
"ts-dedent": "^2.2.0",
@@ -8317,10 +8083,13 @@
}
},
"node_modules/proxy-from-env": {
"version": "1.1.0",
"resolved": "https://registry.npmjs.org/proxy-from-env/-/proxy-from-env-1.1.0.tgz",
"integrity": "sha512-D+zkORCbA9f1tdWRK0RaCR3GPv50cMxcrz4X8k5LTSUD1Dkw47mKJEZQNunItRTkWwgtaUSo1RVFRIG9ZXiFYg==",
"license": "MIT"
"version": "2.1.0",
"resolved": "https://registry.npmjs.org/proxy-from-env/-/proxy-from-env-2.1.0.tgz",
"integrity": "sha512-cJ+oHTW1VAEa8cJslgmUZrc+sjRKgAKl3Zyse6+PV38hZe/V6Z14TbCuXcan9F9ghlz4QrFr2c92TNF82UkYHA==",
"license": "MIT",
"engines": {
"node": ">=10"
}
},
"node_modules/punycode": {
"version": "2.3.1",
@@ -9006,9 +8775,9 @@
"license": "MIT"
},
"node_modules/tailwindcss": {
"version": "4.1.18",
"resolved": "https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.1.18.tgz",
"integrity": "sha512-4+Z+0yiYyEtUVCScyfHCxOYP06L5Ne+JiHhY2IjR2KWMIWhJOYZKLSGZaP5HkZ8+bY0cxfzwDE5uOmzFXyIwxw==",
"version": "4.2.2",
"resolved": "https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.2.2.tgz",
"integrity": "sha512-KWBIxs1Xb6NoLdMVqhbhgwZf2PGBpPEiwOqgI4pFIYbNTfBXiKYyWoTsXgBQ9WFg/OlhnvHaY+AEpW7wSmFo2Q==",
"license": "MIT"
},
"node_modules/tapable": {
@@ -10270,9 +10039,9 @@
"license": "MIT"
},
"node_modules/vscode-uri": {
"version": "3.0.8",
"resolved": "https://registry.npmjs.org/vscode-uri/-/vscode-uri-3.0.8.tgz",
"integrity": "sha512-AyFQ0EVmsOZOlAnxoFOGOq1SQDWAB7C6aqMGS23svWAllfOaxbuFvcT8D1i8z3Gyn8fraVeZNNmN6e9bxxXkKw==",
"version": "3.1.0",
"resolved": "https://registry.npmjs.org/vscode-uri/-/vscode-uri-3.1.0.tgz",
"integrity": "sha512-/BpdSx+yCQGnCvecbyXdxHDkuk55/G3xwnC0GqY4gmQ3j+A+g8kzzgB4Nk/SINjqn6+waqw3EgbVF2QKExkRxQ==",
"license": "MIT"
},
"node_modules/w3c-xmlserializer": {
@@ -10289,15 +10058,15 @@
}
},
"node_modules/wait-on": {
"version": "8.0.5",
"resolved": "https://registry.npmjs.org/wait-on/-/wait-on-8.0.5.tgz",
"integrity": "sha512-J3WlS0txVHkhLRb2FsmRg3dkMTCV1+M6Xra3Ho7HzZDHpE7DCOnoSoCJsZotrmW3uRMhvIJGSKUKrh/MeF4iag==",
"version": "9.0.5",
"resolved": "https://registry.npmjs.org/wait-on/-/wait-on-9.0.5.tgz",
"integrity": "sha512-qgnbHDfDTRIp73ANEJNRW/7kn8CrDUcvZz18xotJQku/P4saTGkbIzvnMZebPmVvVNUiRq1qWAPyqCH+W4H8KA==",
"dev": true,
"license": "MIT",
"dependencies": {
"axios": "^1.12.1",
"joi": "^18.0.1",
"lodash": "^4.17.21",
"axios": "^1.15.0",
"joi": "^18.1.2",
"lodash": "^4.18.1",
"minimist": "^1.2.8",
"rxjs": "^7.8.2"
},
@@ -10305,7 +10074,7 @@
"wait-on": "bin/wait-on"
},
"engines": {
"node": ">=12.0.0"
"node": ">=20.0.0"
}
},
"node_modules/webidl-conversions": {
+5 -5
View File
@@ -39,7 +39,7 @@
"langchain": "^1.2.10",
"lru-cache": "^11.2.4",
"lucide-react": "^0.562.0",
"mermaid": "^11.12.2",
"mermaid": "^11.14.0",
"mnemonist": "^0.39.0",
"pandemonium": "^2.4.0",
"react": "^18.3.1",
@@ -49,12 +49,12 @@
"react-zoom-pan-pinch": "^3.7.0",
"remark-gfm": "^4.0.1",
"sigma": "^3.0.2",
"tailwindcss": "^4.1.18",
"tailwindcss": "^4.2.2",
"uuid": "^13.0.0",
"zod": "^3.25.76"
},
"devDependencies": {
"@babel/types": "^7.28.5",
"@babel/types": "^7.29.0",
"@playwright/test": "^1.58.2",
"@testing-library/jest-dom": "^6.9.1",
"@testing-library/react": "^16.3.2",
@@ -67,11 +67,11 @@
"@vercel/node": "^5.5.16",
"@vitejs/plugin-react": "^5.1.0",
"@vitest/coverage-v8": "^3.2.4",
"jsdom": "^29.0.0",
"jsdom": "^29.0.2",
"tree-sitter-wasms": "^0.1.13",
"typescript": "^5.4.5",
"vite": "^5.2.0",
"vitest": "^3.2.4",
"wait-on": "^8.0.5"
"wait-on": "^9.0.5"
}
}
+1
View File
@@ -16,6 +16,7 @@ export default defineConfig({
alias: {
'@': path.resolve(__dirname, './src'),
'@shared': path.resolve(__dirname, '../shared'),
'gitnexus-shared': path.resolve(__dirname, '../gitnexus-shared/src/index.ts'),
// Fix for Rollup failing to resolve this deep import from @langchain/anthropic
'@anthropic-ai/sdk/lib/transform-json-schema': path.resolve(
__dirname,
+46
View File
@@ -2,6 +2,52 @@
All notable changes to GitNexus will be documented in this file.
## [Unreleased]
### Performance
- **`analyze` ~33% faster** — moved FTS index creation from the analyze pipeline to first-use lazy initialisation. The 5 `CREATE_FTS_INDEX` calls cost ~440 ms each in LadybugDB regardless of table size (≈2 s fixed overhead) and dominated runtime on small repos and slow CI runners. The cost now amortises across the first `query`/`context` call in a session via a new `ensureFTSIndex` helper. Mini-repo `analyze` measured locally on Windows: 6.4 s → 4.0 s warm; on CI Windows runners (≈3× slower) restores comfortable headroom against the 30 s e2e test budget.
## [1.6.2] - 2026-04-18
### Added
- **Docker support** — containerized ingestion and MCP serving for reproducible runs on CI and container platforms (#848)
- **Language-agnostic heritage extractor** — config+factory pattern for class-heritage extraction (EXTENDS / IMPLEMENTS), completing the extractor refactor alongside method/field/call/variable (#890)
- **Language-agnostic call extractor** — config+factory pattern that collapses ~225 lines of inline parse-worker logic into declarative per-language configs (#877)
- **Language-agnostic variable extractor** — structured metadata for `Const` / `Static` / `Variable` nodes via config+factory pattern (#878)
- **AST-aware embedding chunking** — offset-based splitting preserves symbol boundaries, improving semantic search precision on large files (#889)
- **HTTP consumer detection for jQuery and axios object-form** — `$.ajax` / `$.get` / `$.post` and `axios({ url, method })` now recognized as HTTP call sites (#887)
### Fixed
- **Python external dotted imports** — avoid spurious same-file matches when an import path like `foo.bar.baz` refers to a third-party module (#899)
- **Worker warnings no longer terminate ingestion** — non-fatal parser warnings keep the pipeline running instead of aborting the run (#900, #261)
- **Global-install upgrade `ENOTEMPTY`** — devendored `tree-sitter-proto` install lifecycle + preinstall cleanup so `npm i -g gitnexus@latest` succeeds on top of an older install (#843, #846)
- **`env.cacheDir`** now defaults to a user-writable location, unblocking ingestion on systems where the install directory is read-only (#845)
- **Content-hash staleness detection for embeddings** — zero-node rebuilds no longer skip vector-index creation, fixing semantic search after selective re-analysis (#831)
- **`tree-sitter-c-sharp` version pin** — locked to 0.23.1 to avoid a breaking change in a transitive prerelease (#834)
- **`release-drafter` v7 CI** — replaced the removed `disable-releaser` flag with `dry-run` so release-note drafts still work
- **`npm arborist` crash from `tree-sitter-dart`** — switched the dependency URL format so `npm install` no longer crashes on clean installs
- **Service-group `ManifestExtractor`** — `config.links` now wires the manifest extractor properly, restoring cross-link discovery that had silently dropped to zero
### Changed
- **SemanticModel wired as a first-class resolution input (SM-20)** — `call-processor`, `resolution-context`, `type-env`, and `heritage-map` now consult `table.model.*` directly; 37 internal call sites migrated off the SymbolTable wrapper (#885)
- **Per-strategy `ImportSemantics` hooks** — `named` / `wildcard-transitive` / `wildcard-leaf` / `namespace` strategies split into composable hooks, replacing the monolithic conditional (Strategies 1–4 of #886)
- **Class extraction configs moved to `configs/` subdirectory** — per-language class configs now co-locate with the other extractor configs, completing the extractor layer's directory convention (#879)
- **CLI AI-context trimmed** — duplicated CLAUDE.md block removed from the shipped context, reducing token usage in LLM-consuming workflows (#904)
- **LLM context files optimized** — AI-consumed documentation tuned for accuracy and token efficiency (#857)
- **Workflow concurrency standardized** — all CI workflows adopt the consistent concurrency key pattern documented in CONTRIBUTING.md; release-note labeling automated (#837)
- **E2E status-ready timeout raised** — 45s accommodates parallel-worker startup variance on CI (#908)
### Chore / Dependencies
- **tree-sitter 0.25 upgrade readiness** — daily Dependabot monitor for the upcoming major-version bump (#847)
- Dependency bumps: `glob` 11.1.0 → 13.0.6 (#867), `commander` 12.1.0 → 14.0.3 (#868), `@huggingface/transformers` (#869), `@modelcontextprotocol/sdk` (#866), `lru-cache` 11.2.7 → 11.3.5 (#870), `mnemonist` 0.39.8 → 0.40.3 (#871), `@ladybugdb/core` (#873)
- gitnexus-web dependency bumps: `mermaid` 11.12.2 → 11.14.0 (#860), `tailwindcss` (#861), `jsdom` 29.0.0 → 29.0.2 (#863), `wait-on` 8.0.5 → 9.0.5 (#859), `@vitest/coverage-v8` (#864)
- GitHub Actions bumps: `actions/checkout` 4.3.1 → 6.0.2 (#842), `actions/upload-artifact` 4.6.2 → 7.0.1 (#838), `actions/setup-node` 4.4.0 → 6.3.0 (#841), `actions/cache` 5.0.4 → 5.0.5 (#840), `actions/github-script` 7.0.1 → 9.0.0 (#850), `dorny/paths-filter` 3.0.2 → 4.0.1 (#839), `amannn/action-semantic-pull-request` 6.1.1 (#853), `release-drafter/release-drafter` 6.0.0 → 7.2.0 (#852), `marocchino/sticky-pull-request-comment` 3.0.4 (#851), `softprops/action-gh-release` 2.5.0 → 3.0.0 (#849)
## [1.6.1] - 2026-04-13
### Added
+1 -1
View File
@@ -1,6 +1,6 @@
FROM node:20-bookworm
WORKDIR /app
RUN apt-get update && apt-get install -y python3 make g++ && rm -rf /var/lib/apt/lists/*
RUN apt-get -o Acquire::Check-Valid-Until=false -o Acquire::Check-Date=false update && apt-get install -y python3 make g++ && rm -rf /var/lib/apt/lists/*
COPY . .
RUN npm ci --ignore-scripts \
&& node scripts/patch-tree-sitter-swift.cjs \
+5 -5
View File
@@ -166,11 +166,11 @@ gitnexus wiki [path] # Generate LLM-powered docs from knowledge grap
gitnexus wiki --model <model> # Wiki with custom LLM model (default: gpt-4o-mini)
# Repository groups (multi-repo / monorepo service tracking)
gitnexus group create <name> # Create a repository group
gitnexus group add <name> <repo> # Add a repo to a group
gitnexus group remove <name> <repo> # Remove a repo from a group
gitnexus group list [name] # List groups, or show one group's config
gitnexus group sync <name> # Extract contracts and match across repos/services
gitnexus group create <name> # Create a repository group
gitnexus group add <group> <groupPath> <registryName> # Add a repo to a group. <groupPath> is a hierarchy path (e.g. hr/hiring/backend); <registryName> is the repo's name from the registry (see `gitnexus list`)
gitnexus group remove <group> <groupPath> # Remove a repo from a group by its hierarchy path
gitnexus group list [name] # List groups, or show one group's config
gitnexus group sync <name> # Extract contracts and match across repos/services
gitnexus group contracts <name> # Inspect extracted contracts and cross-links
gitnexus group query <name> <q> # Search execution flows across all repos in a group
gitnexus group status <name> # Check staleness of repos in a group
+291 -298
View File
File diff suppressed because it is too large Load Diff
+11 -10
View File
@@ -1,6 +1,6 @@
{
"name": "gitnexus",
"version": "1.6.1",
"version": "1.6.2",
"description": "Graph-powered code intelligence for AI agents. Index any codebase, query via MCP or CLI.",
"author": "Abhigyan Patwari",
"license": "PolyForm-Noncommercial-1.0.0",
@@ -51,22 +51,23 @@
"prepack": "node scripts/build.js"
},
"dependencies": {
"@huggingface/transformers": "^3.0.0",
"@huggingface/transformers": "^4.1.0",
"@ladybugdb/core": "^0.15.2",
"@modelcontextprotocol/sdk": "^1.0.0",
"@scarf/scarf": "^1.4.0",
"cli-progress": "^3.12.0",
"commander": "^12.0.0",
"commander": "^14.0.3",
"cors": "^2.8.5",
"express": "^4.19.2",
"glob": "^11.0.0",
"graphology": "^0.25.4",
"glob": "^13.0.6",
"graphology": "^0.26.0",
"graphology-indices": "^0.17.0",
"graphology-utils": "^2.3.0",
"ignore": "^7.0.5",
"js-yaml": "^4.1.1",
"jsonc-parser": "^3.3.1",
"lru-cache": "^11.0.0",
"mnemonist": "^0.39.0",
"mnemonist": "^0.40.3",
"onnxruntime-node": "^1.24.0",
"pandemonium": "^2.4.0",
"tree-sitter": "^0.21.1",
@@ -81,7 +82,7 @@
"tree-sitter-ruby": "^0.23.1",
"tree-sitter-rust": "0.23.1",
"tree-sitter-typescript": "^0.23.2",
"uuid": "^13.0.0"
"uuid": "^14.0.0"
},
"optionalDependencies": {
"node-addon-api": "^8.0.0",
@@ -92,14 +93,14 @@
"tree-sitter-swift": "^0.6.0"
},
"devDependencies": {
"gitnexus-shared": "file:../gitnexus-shared",
"@types/cli-progress": "^3.11.6",
"@types/cors": "^2.8.17",
"@types/express": "^4.17.21",
"@types/js-yaml": "^4.0.9",
"@types/node": "^20.0.0",
"@types/uuid": "^10.0.0",
"@types/node": "^25.6.0",
"@types/uuid": "^11.0.0",
"@vitest/coverage-v8": "^4.0.18",
"gitnexus-shared": "file:../gitnexus-shared",
"tsx": "^4.0.0",
"typescript": "^5.4.5",
"vitest": "^4.0.18"
+134
View File
@@ -0,0 +1,134 @@
/**
* Synthetic benchmark for scope-resolution. Builds a large in-memory
* Python workspace and times runScopeResolution against it directly,
* isolating the resolution cost from parse / heritage / pipeline
* overhead.
*
* Usage: REGISTRY_PRIMARY_PYTHON=1 npx tsx scripts/bench-scope-resolution.ts
*/
process.env.REGISTRY_PRIMARY_PYTHON = '1';
import { generateId } from '../src/lib/utils.js';
import { createKnowledgeGraph } from '../src/core/graph/graph.js';
import { runScopeResolution } from '../src/core/ingestion/scope-resolution/index.js';
import { pythonScopeResolver } from '../src/core/ingestion/languages/python/scope-resolver.js';
const N_CLASSES = Number(process.env.BENCH_CLASSES ?? '60');
const N_USERS = Number(process.env.BENCH_USERS ?? '40');
const ITERS = Number(process.env.BENCH_ITERS ?? '5');
function buildWorkspace(): { path: string; content: string }[] {
const files: { path: string; content: string }[] = [];
// Build N_CLASSES "model" files, each defining a class with a few methods.
for (let i = 0; i < N_CLASSES; i++) {
const lines: string[] = [];
for (let j = 0; j < 5; j++) {
lines.push(`class Model${i}_${j}:`);
lines.push(` name: str`);
lines.push(` def save(self) -> bool:`);
lines.push(` return True`);
lines.push(` def update(self, name: str) -> "Model${i}_${j}":`);
lines.push(` self.name = name`);
lines.push(` return self`);
lines.push(` def get_other(self) -> "Model${i}_${(j + 1) % 5}":`);
lines.push(` return Model${i}_${(j + 1) % 5}()`);
lines.push('');
}
files.push({ path: `models/m${i}.py`, content: lines.join('\n') });
}
// Build N_USERS "user" files that import from a few model files
// and exercise the receiver-bound dispatcher heavily.
for (let u = 0; u < N_USERS; u++) {
const targets = [u % N_CLASSES, (u + 1) % N_CLASSES, (u + 2) % N_CLASSES];
const imports = targets
.map((t) => `from models.m${t} import Model${t}_0, Model${t}_1, Model${t}_2`)
.join('\n');
const calls: string[] = [];
for (let k = 0; k < 30; k++) {
const t = targets[k % 3]!;
const j = k % 3;
calls.push(` m${k} = Model${t}_${j}()`);
calls.push(` m${k}.save()`);
calls.push(` m${k}.update("x").save()`);
calls.push(` m${k}.get_other().save()`);
}
const content = `${imports}\n\ndef use_${u}() -> None:\n${calls.join('\n')}\n`;
files.push({ path: `app/u${u}.py`, content });
}
return files;
}
function buildGraph(files: { path: string; content: string }[]) {
const graph = createKnowledgeGraph();
// Pre-populate File / Class / Function nodes the resolver expects.
for (const f of files) {
const fileId = generateId('File', f.path);
graph.addNode({
id: fileId,
label: 'File',
properties: { name: f.path, filePath: f.path },
});
// Lightweight regex-extract class & def names so the lookup index
// has something to find. Real pipeline builds these via parse phase;
// for the bench this stand-in is enough to exercise the resolver.
const classRe = /^class (\w+)/gm;
const defRe = /^\s*def (\w+)/gm;
let m: RegExpExecArray | null;
while ((m = classRe.exec(f.content)) !== null) {
const name = m[1]!;
const id = generateId('Class', `${f.path}:${name}`);
graph.addNode({
id,
label: 'Class',
properties: { name, filePath: f.path, qualifiedName: name },
});
}
while ((m = defRe.exec(f.content)) !== null) {
const name = m[1]!;
const id = generateId('Function', `${f.path}:${name}`);
graph.addNode({
id,
label: 'Function',
properties: { name, filePath: f.path, qualifiedName: name },
});
}
}
return graph;
}
async function main() {
const files = buildWorkspace();
console.log(`bench: ${files.length} files (${N_CLASSES} models × 5 classes + ${N_USERS} users)`);
console.log(` × ${ITERS} iterations\n`);
// Warmup
for (let i = 0; i < 2; i++) {
const graph = buildGraph(files);
runScopeResolution({ graph, files, onWarn: () => {} }, pythonScopeResolver);
}
const samples: number[] = [];
for (let i = 0; i < ITERS; i++) {
const graph = buildGraph(files);
const start = process.hrtime.bigint();
runScopeResolution({ graph, files, onWarn: () => {} }, pythonScopeResolver);
const end = process.hrtime.bigint();
const ms = Number(end - start) / 1_000_000;
samples.push(ms);
console.log(` iter ${i + 1}: ${ms.toFixed(0)} ms`);
}
samples.sort((a, b) => a - b);
const median = samples[Math.floor(samples.length / 2)]!;
const min = samples[0]!;
console.log(`\nmin: ${min.toFixed(0)} ms · median: ${median.toFixed(0)} ms`);
}
main().catch((err) => {
console.error(err);
process.exit(1);
});
@@ -0,0 +1,24 @@
/**
* CI helper — emits the `MIGRATED_LANGUAGES` set as a JSON matrix array for
* GitHub Actions (`.github/workflows/ci-scope-parity.yml`).
*
* Consumed by the `discover` job in that workflow. Each entry has:
* - `slug`: lowercase language id, matching `test/integration/resolvers/<slug>.test.ts`.
* - `envvar`: uppercase suffix used to build the `REGISTRY_PRIMARY_<envvar>` toggle.
*
* Run with `npx tsx scripts/ci-list-migrated-languages.ts`. The script
* writes a single JSON array to stdout (no wrapper object) so the
* workflow can pipe it straight into `$GITHUB_OUTPUT`.
*/
import { MIGRATED_LANGUAGES } from '../src/core/ingestion/registry-primary-flag.js';
const entries = [...MIGRATED_LANGUAGES].map((slug) => {
const s = String(slug);
return {
slug: s,
envvar: s.toUpperCase().replace(/-/g, '_'),
};
});
process.stdout.write(JSON.stringify(entries));
+291
View File
@@ -0,0 +1,291 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width,initial-scale=1" />
<title>GitNexus — Shadow Parity Dashboard</title>
<!--
Static dashboard for the RFC #909 shadow-mode parity report.
Reads `latest.json` from this directory and renders a per-language
parity table. Zero build step, zero runtime dependencies — a
single file that any browser or file:// context can open.
Usage:
# from repo root, after a shadow-mode run
cp .gitnexus/shadow-parity/latest.json gitnexus/shadow-parity-dashboard/
open gitnexus/shadow-parity-dashboard/index.html
CI artifact wiring (follow-up): the CI job publishes a snapshot
of this directory + latest.json as a downloadable bundle per run.
-->
<style>
:root {
color-scheme: light dark;
--fg: #1f2937;
--fg-muted: #6b7280;
--bg: #ffffff;
--bg-muted: #f9fafb;
--border: #e5e7eb;
--good: #16a34a;
--warn: #d97706;
--bad: #dc2626;
--primary-tag-legacy: #7c3aed;
--primary-tag-registry: #0ea5e9;
}
@media (prefers-color-scheme: dark) {
:root {
--fg: #e5e7eb;
--fg-muted: #9ca3af;
--bg: #111827;
--bg-muted: #1f2937;
--border: #374151;
}
}
html,
body {
margin: 0;
padding: 0;
background: var(--bg);
color: var(--fg);
font:
14px/1.45 system-ui,
-apple-system,
sans-serif;
}
main {
max-width: 1200px;
margin: 0 auto;
padding: 24px 16px;
}
h1 {
font-size: 20px;
margin: 0 0 4px;
}
.meta {
color: var(--fg-muted);
font-size: 12px;
margin-bottom: 20px;
}
.cards {
display: grid;
grid-template-columns: repeat(auto-fit, minmax(180px, 1fr));
gap: 10px;
margin-bottom: 20px;
}
.card {
border: 1px solid var(--border);
border-radius: 6px;
padding: 10px 12px;
background: var(--bg-muted);
}
.card .k {
color: var(--fg-muted);
font-size: 11px;
text-transform: uppercase;
letter-spacing: 0.04em;
}
.card .v {
font-size: 20px;
font-weight: 600;
}
table {
width: 100%;
border-collapse: collapse;
font-variant-numeric: tabular-nums;
}
th,
td {
padding: 6px 10px;
text-align: right;
border-bottom: 1px solid var(--border);
}
th:first-child,
td:first-child {
text-align: left;
}
thead th {
font-weight: 600;
color: var(--fg-muted);
font-size: 12px;
background: var(--bg-muted);
}
tbody tr:hover {
background: var(--bg-muted);
}
.parity {
font-weight: 600;
}
.parity.good {
color: var(--good);
}
.parity.warn {
color: var(--warn);
}
.parity.bad {
color: var(--bad);
}
.tag {
display: inline-block;
padding: 1px 6px;
border-radius: 10px;
font-size: 10px;
margin-left: 6px;
color: white;
}
.tag.legacy {
background: var(--primary-tag-legacy);
}
.tag.registry {
background: var(--primary-tag-registry);
}
.empty {
padding: 40px;
text-align: center;
color: var(--fg-muted);
}
code {
font-family: ui-monospace, SFMono-Regular, Menlo, monospace;
background: var(--bg-muted);
padding: 1px 4px;
border-radius: 3px;
}
</style>
</head>
<body>
<main>
<h1>Shadow Parity — RFC #909</h1>
<div class="meta" id="meta">loading <code>latest.json</code>…</div>
<div class="cards" id="cards"></div>
<table id="per-language">
<thead>
<tr>
<th>Language</th>
<th>Total</th>
<th>Agree</th>
<th>Only legacy</th>
<th>Only new</th>
<th>Disagree</th>
<th>Both empty</th>
<th>Parity</th>
</tr>
</thead>
<tbody></tbody>
</table>
<div id="empty" class="empty" style="display: none">
No records yet. Enable <code>GITNEXUS_SHADOW_MODE=1</code> and run ingestion to populate.
</div>
</main>
<script>
/* global fetch, document */
(async function () {
const tbody = document.querySelector('#per-language tbody');
const cards = document.getElementById('cards');
const meta = document.getElementById('meta');
const empty = document.getElementById('empty');
const table = document.getElementById('per-language');
let payload;
try {
const r = await fetch('./latest.json', { cache: 'no-store' });
if (!r.ok) throw new Error('HTTP ' + r.status);
payload = await r.json();
} catch (err) {
meta.textContent = 'Failed to load latest.json: ' + err.message;
table.style.display = 'none';
empty.style.display = 'block';
return;
}
const primary = payload.primaryByLanguage || {};
const report = payload.report || {};
const perLang = report.perLanguage || [];
const overall = report.overall || {};
meta.textContent =
'Run ' +
payload.runId +
' — generated ' +
payload.generatedAt +
' (schema v' +
payload.schemaVersion +
')';
// Overall summary cards.
cards.innerHTML = '';
const overallParity = overall.parity !== undefined ? overall.parity : 0;
cards.appendChild(makeCard('Total calls', overall.totalCalls ?? 0));
cards.appendChild(makeCard('Both agree', overall.bothAgree ?? 0));
cards.appendChild(makeCard('Disagree', overall.bothDisagree ?? 0));
cards.appendChild(makeCard('Overall parity', formatPct(overallParity)));
if (!perLang.length) {
table.style.display = 'none';
empty.style.display = 'block';
return;
}
for (const row of perLang) {
const tr = document.createElement('tr');
const primaryTag = primary[row.language];
const tag = primaryTag
? '<span class="tag ' + primaryTag + '">primary: ' + primaryTag + '</span>'
: '';
const parityClass = parityClassFor(row.parity);
tr.innerHTML =
'<td>' +
escape(row.language) +
tag +
'</td>' +
'<td>' +
row.totalCalls +
'</td>' +
'<td>' +
row.bothAgree +
'</td>' +
'<td>' +
row.onlyLegacy +
'</td>' +
'<td>' +
row.onlyNew +
'</td>' +
'<td>' +
row.bothDisagree +
'</td>' +
'<td>' +
row.bothEmpty +
'</td>' +
'<td class="parity ' +
parityClass +
'">' +
formatPct(row.parity) +
'</td>';
tbody.appendChild(tr);
}
function makeCard(k, v) {
const div = document.createElement('div');
div.className = 'card';
div.innerHTML =
'<div class="k">' + escape(k) + '</div><div class="v">' + escape(String(v)) + '</div>';
return div;
}
function formatPct(x) {
if (typeof x !== 'number' || !isFinite(x)) return '—';
return (x * 100).toFixed(1) + '%';
}
function parityClassFor(x) {
if (typeof x !== 'number') return '';
if (x >= 0.95) return 'good';
if (x >= 0.8) return 'warn';
return 'bad';
}
function escape(s) {
return String(s).replace(/[&<>"']/g, function (c) {
return { '&': '&amp;', '<': '&lt;', '>': '&gt;', '"': '&quot;', "'": '&#39;' }[c];
});
}
})();
</script>
</body>
</html>
+40 -62
View File
@@ -32,6 +32,33 @@ export interface AIContextOptions {
const GITNEXUS_START_MARKER = '<!-- gitnexus:start -->';
const GITNEXUS_END_MARKER = '<!-- gitnexus:end -->';
/**
* Find the index of a section marker that occupies its own line.
* Unlike `indexOf`, this rejects inline prose references like
* `` See the `<!-- gitnexus:start -->` block `` that appear
* mid-sentence (#1041). A marker counts as section-position only when:
* - preceded by newline or start-of-file, AND
* - followed by newline, `\r` (CRLF files), or end-of-file.
* The generator always emits each marker alone on its line, so this
* matches every legitimate section and none of the inline mentions.
*
* `startFrom` lets the end-marker lookup start after the already-found
* start marker, avoiding a scan from 0 and guaranteeing we never pick
* up an end marker that appears earlier in the file than the start.
*/
function findSectionMarkerIndex(content: string, marker: string, startFrom = 0): number {
let idx = content.indexOf(marker, startFrom);
while (idx !== -1) {
const atLineStart = idx === 0 || content[idx - 1] === '\n';
const endPos = idx + marker.length;
const atLineEnd =
endPos === content.length || content[endPos] === '\n' || content[endPos] === '\r';
if (atLineStart && atLineEnd) return idx;
idx = content.indexOf(marker, idx + 1);
}
return -1;
}
/**
* Generate the full GitNexus context content.
*
@@ -101,19 +128,6 @@ This project is indexed by GitNexus as **${projectName}**${noStats ? '' : ` (${s
- When exploring unfamiliar code, use \`gitnexus_query({query: "concept"})\` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use \`gitnexus_context({name: "symbolName"})\`.
## When Debugging
1. \`gitnexus_query({query: "<error or symptom>"})\` — find execution flows related to the issue
2. \`gitnexus_context({name: "<suspect function>"})\` — see all callers, callees, and process participation
3. \`READ gitnexus://repo/${projectName}/process/{processName}\` — trace the full execution flow step by step
4. For regressions: \`gitnexus_detect_changes({scope: "compare", base_ref: "main"})\` — see what your branch changed
## When Refactoring
- **Renaming**: MUST use \`gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})\` first. Review the preview — graph edits are safe, text_search edits need manual review. Then run with \`dry_run: false\`.
- **Extracting/Splitting**: MUST run \`gitnexus_context({name: "target"})\` to see all incoming/outgoing refs, then \`gitnexus_impact({target: "target", direction: "upstream"})\` to find all external callers before moving code.
- After any refactor: run \`gitnexus_detect_changes({scope: "all"})\` to verify only expected files changed.
## Never Do
- NEVER edit a function, class, or method without first running \`gitnexus_impact\` on it.
@@ -121,25 +135,6 @@ This project is indexed by GitNexus as **${projectName}**${noStats ? '' : ` (${s
- NEVER rename symbols with find-and-replace — use \`gitnexus_rename\` which understands the call graph.
- NEVER commit changes without running \`gitnexus_detect_changes()\` to check affected scope.
## Tools Quick Reference
| Tool | When to use | Command |
|------|-------------|---------|
| \`query\` | Find code by concept | \`gitnexus_query({query: "auth validation"})\` |
| \`context\` | 360-degree view of one symbol | \`gitnexus_context({name: "validateUser"})\` |
| \`impact\` | Blast radius before editing | \`gitnexus_impact({target: "X", direction: "upstream"})\` |
| \`detect_changes\` | Pre-commit scope check | \`gitnexus_detect_changes({scope: "staged"})\` |
| \`rename\` | Safe multi-file rename | \`gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})\` |
| \`cypher\` | Custom graph queries | \`gitnexus_cypher({query: "MATCH ..."})\` |
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update these |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
## Resources
| Resource | Use for |
@@ -149,37 +144,11 @@ This project is indexed by GitNexus as **${projectName}**${noStats ? '' : ` (${s
| \`gitnexus://repo/${projectName}/processes\` | All execution flows |
| \`gitnexus://repo/${projectName}/process/{name}\` | Step-by-step execution trace |
## Self-Check Before Finishing
Before completing any code modification task, verify:
1. \`gitnexus_impact\` was run for all modified symbols
2. No HIGH/CRITICAL risk warnings were ignored
3. \`gitnexus_detect_changes()\` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## Keeping the Index Fresh
After committing code changes, the GitNexus index becomes stale. Re-run analyze to update it:
\`\`\`bash
npx gitnexus analyze
\`\`\`
If the index previously included embeddings, preserve them by adding \`--embeddings\`:
\`\`\`bash
npx gitnexus analyze --embeddings
\`\`\`
To check whether embeddings exist, inspect \`.gitnexus/meta.json\` — the \`stats.embeddings\` field shows the count (0 means no embeddings). **Running analyze without \`--embeddings\` will delete any previously generated embeddings.**
> Claude Code users: A PostToolUse hook handles this automatically after \`git commit\` and \`git merge\`.
${
groupNames && groupNames.length > 0
? `## Cross-Repo Groups
This repository is listed under GitNexus **group(s): ${groupNames.join(', ')}** (see \`~/.gitnexus/groups/\`). For blast radius across repository boundaries, use MCP tools \`group_impact\`, \`group_sync\`, \`group_query\`, \`group_contracts\`, \`group_status\`, and \`group_list\`. From the terminal: \`npx gitnexus group list\`, \`npx gitnexus group sync <name>\`, \`npx gitnexus group impact <name> --target <symbol> --repo <group-path>\`.
This repository is listed under GitNexus **group(s): ${groupNames.join(', ')}** (see \`~/.gitnexus/groups/\`). For cross-repo analysis, use MCP tools \`impact\`, \`query\`, and \`context\` with \`repo\` set to \`@<groupName>\` or \`@<groupName>/<memberPath>\` (paths match keys in that group’s \`group.yaml\`). Use \`group_list\` / \`group_sync\` for membership and sync. From the terminal: \`npx gitnexus group list\`, \`npx gitnexus group sync <name>\`, \`npx gitnexus group impact <name> --target <symbol> --repo <group-path>\`.
`
: ''
@@ -221,9 +190,18 @@ async function upsertGitNexusSection(
const existingContent = await fs.readFile(filePath, 'utf-8');
// Check if GitNexus section already exists
const startIdx = existingContent.indexOf(GITNEXUS_START_MARKER);
const endIdx = existingContent.indexOf(GITNEXUS_END_MARKER);
// Check if GitNexus section already exists. Matching is restricted
// to markers that occupy their own line so that inline prose
// references (e.g. `` See the `<!-- gitnexus:start -->` block `` in
// the shipped CLAUDE.md) are NOT treated as section delimiters
// (#1041). The end-marker scan starts after the start-marker so it
// can never pick up an earlier end in the file.
const startIdx = findSectionMarkerIndex(existingContent, GITNEXUS_START_MARKER);
const endIdx = findSectionMarkerIndex(
existingContent,
GITNEXUS_END_MARKER,
startIdx === -1 ? 0 : startIdx,
);
if (startIdx !== -1 && endIdx !== -1 && endIdx > startIdx) {
// Replace existing section
+45 -1
View File
@@ -13,7 +13,11 @@ import { execFileSync } from 'child_process';
import v8 from 'v8';
import cliProgress from 'cli-progress';
import { closeLbug } from '../core/lbug/lbug-adapter.js';
import { getStoragePaths, getGlobalRegistryPath } from '../storage/repo-manager.js';
import {
getStoragePaths,
getGlobalRegistryPath,
RegistryNameCollisionError,
} from '../storage/repo-manager.js';
import { getGitRoot, hasGitDir } from '../storage/git.js';
import { runFullAnalysis } from '../core/run-analyze.js';
import fs from 'fs/promises';
@@ -59,6 +63,21 @@ export interface AnalyzeOptions {
noStats?: boolean;
/** Index the folder even when no .git directory is present. */
skipGit?: boolean;
/**
* Override the default basename-derived registry `name` with a
* user-supplied alias (#829). Disambiguates repos whose paths share a
* basename. Persisted — subsequent re-analyses of the same path without
* `--name` preserve the alias.
*/
name?: string;
/**
* Allow registration even when another path already uses the same
* `--name` alias (#829). Intentionally a distinct flag from `--force`
* because the user may want to coexist under the same name WITHOUT
* paying the cost of a pipeline re-index. Maps to registerRepo's
* `allowDuplicateName` option end-to-end.
*/
allowDuplicateName?: boolean;
}
export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOptions) => {
@@ -186,11 +205,20 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
const result = await runFullAnalysis(
repoPath,
{
// Pipeline re-index — OR'd with --skills because skill generation
// needs a fresh pipelineResult. Has no bearing on the registry
// collision guard (see allowDuplicateName below).
force: options?.force || options?.skills,
embeddings: options?.embeddings,
skipGit: options?.skipGit,
skipAgentsMd: options?.skipAgentsMd,
noStats: options?.noStats,
registryName: options?.name,
// Registry-collision bypass — its own CLI flag, intentionally NOT
// overloading --force. A user who hits the collision guard should
// be able to accept the duplicate name without also paying the
// cost of a full pipeline re-index. See #829 review round 2.
allowDuplicateName: options?.allowDuplicateName,
},
{
onProgress: (_phase, percent, message) => {
@@ -298,6 +326,22 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
bar.stop();
const msg = err.message || String(err);
// Registry name-collision from --name (#829) — surface as an
// actionable error rather than a generic stack-trace.
if (err instanceof RegistryNameCollisionError) {
console.error(`\n Registry name collision:\n`);
console.error(` "${err.registryName}" is already used by "${err.existingPath}".\n`);
console.error(` Options:`);
console.error(` • Pick a different alias: gitnexus analyze --name <alias>`);
console.error(
` • Allow the duplicate: gitnexus analyze --allow-duplicate-name (leaves "-r ${err.registryName}" ambiguous)`,
);
console.error('');
process.exitCode = 1;
return;
}
console.error(`\n Analysis failed: ${msg}\n`);
// Provide helpful guidance for known failure modes
+25 -1
View File
@@ -6,7 +6,13 @@
*/
import fs from 'fs/promises';
import { findRepo, unregisterRepo, listRegisteredRepos } from '../storage/repo-manager.js';
import {
findRepo,
unregisterRepo,
listRegisteredRepos,
assertSafeStoragePath,
UnsafeStoragePathError,
} from '../storage/repo-manager.js';
export const cleanCommand = async (options?: { force?: boolean; all?: boolean }) => {
// --all flag: clean all indexed repos
@@ -27,6 +33,24 @@ export const cleanCommand = async (options?: { force?: boolean; all?: boolean })
const entries = await listRegisteredRepos();
for (const entry of entries) {
// Safety guard (#1003 review — @magyargergo): same rationale as
// remove.ts. `~/.gitnexus/registry.json` is user-writable, so a
// corrupted or hand-edited entry could point storagePath at the
// repo root, an empty string, or anywhere else — and
// fs.rm(recursive: true) on any of those would be catastrophic.
// Skip poisoned entries without touching disk, but keep going
// through the rest of the registry (preserves the existing
// per-repo error-tolerance semantics of `clean --all`).
try {
assertSafeStoragePath(entry);
} catch (err) {
if (err instanceof UnsafeStoragePathError) {
console.error(`Refusing to clean ${entry.name}: ${err.message}`);
continue;
}
throw err;
}
try {
await fs.rm(entry.storagePath, { recursive: true, force: true });
await unregisterRepo(entry.path);
+77
View File
@@ -184,6 +184,83 @@ export function registerGroupCommands(program: Command): void {
}
});
group
.command('impact <name>')
.description('Cross-repo impact for a symbol in one member repo of a group')
.requiredOption('--target <symbol>', 'Symbol or file name to analyze')
.requiredOption(
'--repo <groupPath>',
'Member path from group.yaml (e.g. app/backend), not the indexed repo name',
)
.option('--direction <dir>', 'upstream or downstream', 'upstream')
.option('--service <path>', 'Optional monorepo service directory prefix (path filter)')
.option(
'--subgroup <path>',
'Optional prefix limiting which group repos participate in cross fan-out',
)
.option('--max-depth <n>', 'Max graph traversal depth')
.option('--cross-depth <n>', 'Cross-repository hop depth')
.option('--min-confidence <n>', 'Minimum relation confidence (0–1)')
.option('--include-tests', 'Include test files in traversal', false)
.option('--timeout-ms <n>', 'Phase-1 local impact wall time in milliseconds')
.option('--json', 'JSON output')
.action(async (name: string, opts: Record<string, string | boolean | undefined>) => {
const { LocalBackend } = await import('../mcp/local/local-backend.js');
const backend = new LocalBackend();
try {
await backend.init();
const payload: Record<string, unknown> = {
name,
repo: opts.repo,
target: opts.target,
direction: (opts.direction as string) || 'upstream',
};
if (opts.service) payload.service = opts.service;
if (opts.subgroup) payload.subgroup = opts.subgroup;
if (opts.maxDepth !== undefined && opts.maxDepth !== '') {
const n = parseInt(String(opts.maxDepth), 10);
if (!Number.isNaN(n)) payload.maxDepth = n;
}
if (opts.crossDepth !== undefined && opts.crossDepth !== '') {
const n = parseInt(String(opts.crossDepth), 10);
if (!Number.isNaN(n)) payload.crossDepth = n;
}
if (opts.minConfidence !== undefined && opts.minConfidence !== '') {
const n = parseFloat(String(opts.minConfidence));
if (!Number.isNaN(n)) payload.minConfidence = n;
}
if (opts.timeoutMs !== undefined && opts.timeoutMs !== '') {
const n = parseInt(String(opts.timeoutMs), 10);
if (!Number.isNaN(n)) payload.timeoutMs = n;
}
if (opts.includeTests) payload.includeTests = true;
const raw = await backend.getGroupService().groupImpact(payload);
if (raw && typeof raw === 'object' && 'error' in raw) {
console.error(String((raw as { error: string }).error));
process.exitCode = 1;
return;
}
if (opts.json) {
console.log(JSON.stringify(raw, null, 2));
} else {
const summary = (raw as { summary?: Record<string, number> })?.summary;
const risk = (raw as { risk?: string })?.risk;
console.log(`Group impact for "${name}" (${String(opts.repo)}): risk=${risk ?? '?'}`);
if (summary) {
console.log(
` direct=${summary.direct ?? 0} processes=${summary.processes_affected ?? 0} cross=${summary.cross_repo_hits ?? 0}`,
);
}
}
} finally {
await backend.dispose().catch(() => {});
}
});
group
.command('query <name> <query>')
.description('Search execution flows across all repos in a group')
+8 -1
View File
@@ -17,7 +17,7 @@ import {
addToGitignore,
registerRepo,
} from '../storage/repo-manager.js';
import { getGitRoot, isGitRepo } from '../storage/git.js';
import { getGitRoot, getRemoteUrl, isGitRepo } from '../storage/git.js';
export interface IndexOptions {
force?: boolean;
@@ -107,6 +107,13 @@ export const indexCommand = async (inputPathParts?: string[], options?: IndexOpt
}
// ── Register in global registry ───────────────────────────────────
// Refresh the on-disk meta with a freshly captured `remoteUrl` if
// it's missing, so an `index` of an older `.gitnexus/` still gets
// sibling-clone fingerprinting on subsequent use without forcing a
// full re-analyze.
if (!meta.remoteUrl && isGitRepo(repoPath)) {
meta.remoteUrl = getRemoteUrl(repoPath);
}
await registerRepo(repoPath, meta);
await addToGitignore(repoPath);
+28
View File
@@ -28,6 +28,16 @@ program
.option('--skip-agents-md', 'Skip updating the gitnexus section in AGENTS.md and CLAUDE.md')
.option('--no-stats', 'Omit volatile file/symbol counts from AGENTS.md and CLAUDE.md')
.option('--skip-git', 'Index a folder without requiring a .git directory')
.option(
'--name <alias>',
'Register this repo under a custom name in ~/.gitnexus/registry.json ' +
'(disambiguates repos whose paths share a basename, e.g. two different .../app folders)',
)
.option(
'--allow-duplicate-name',
'Register this repo even if another path already uses the same --name alias. ' +
'Leaves `-r <name>` ambiguous for the two paths; use -r <path> to disambiguate.',
)
.option('-v, --verbose', 'Enable verbose ingestion warnings (default: false)')
.addHelpText(
'after',
@@ -73,6 +83,15 @@ program
.option('--all', 'Clean all indexed repos')
.action(createLazyAction(() => import('./clean.js'), 'cleanCommand'));
program
.command('remove <target>')
.description(
'Delete the GitNexus index for a registered repo (by alias, name, or absolute path). ' +
'Unlike `clean`, does not require being inside the repo. Idempotent on unknown targets.',
)
.option('-f, --force', 'Skip confirmation prompt')
.action(createLazyAction(() => import('./remove.js'), 'removeCommand'));
program
.command('wiki [path]')
.description('Generate repository wiki from knowledge graph')
@@ -141,6 +160,15 @@ program
.option('-r, --repo <name>', 'Target repository')
.action(createLazyAction(() => import('./tool.js'), 'cypherCommand'));
program
.command('detect-changes')
.alias('detect_changes')
.description('Map git diff hunks to indexed symbols and affected execution flows')
.option('-s, --scope <scope>', 'What to analyze: unstaged, staged, all, or compare', 'unstaged')
.option('-b, --base-ref <ref>', 'Branch/commit for compare scope (e.g. main)')
.option('-r, --repo <name>', 'Target repository')
.action(createLazyAction(() => import('./tool.js'), 'detectChangesCommand'));
// ─── Eval Server (persistent daemon for SWE-bench) ─────────────────
program
+12 -1
View File
@@ -17,12 +17,23 @@ export const listCommand = async () => {
console.log(`\n Indexed Repositories (${entries.length})\n`);
// Count occurrences of each name so colliding entries can be
// disambiguated in the header (#829). Unique-name entries render
// identically to pre-#829 output; only collisions gain a suffix.
const nameCounts = new Map<string, number>();
for (const e of entries) {
const key = e.name.toLowerCase();
nameCounts.set(key, (nameCounts.get(key) ?? 0) + 1);
}
for (const entry of entries) {
const indexedDate = new Date(entry.indexedAt).toLocaleString();
const stats = entry.stats || {};
const commitShort = entry.lastCommit?.slice(0, 7) || 'unknown';
const hasCollision = (nameCounts.get(entry.name.toLowerCase()) ?? 0) > 1;
const header = hasCollision ? `${entry.name} (${entry.path})` : entry.name;
console.log(` ${entry.name}`);
console.log(` ${header}`);
console.log(` Path: ${entry.path}`);
console.log(` Indexed: ${indexedDate}`);
console.log(` Commit: ${commitShort}`);
+110
View File
@@ -0,0 +1,110 @@
/**
* Remove Command (#664)
*
* Delete the `.gitnexus/` index for a registered repo and unregister it
* from the global registry (~/.gitnexus/registry.json). The target is
* identified by alias / basename-derived name / remote-inferred name /
* absolute path — no `--repo` flag, just a positional argument so the
* destructive-command ergonomics match `clean` (which is also
* destructive but scoped to `process.cwd()`).
*
* Compared to `clean`:
* - `clean` acts on the repo discovered by walking up from cwd.
* - `remove` acts on any registered repo identified by name or path.
*
* Behaviour notes:
* - Idempotent on unknown targets: exits 0 with a warning so that
* `remove X && analyze Y` keeps working in scripts. Per #664:
* "behave atomically and idempotently so retries are safe".
* - Atomic order mirrors `clean`: fs.rm FIRST, then unregister. A
* partial failure leaves the registry pointing at a missing dir
* (recoverable by `listRegisteredRepos({ validate: true })` on
* next read) rather than the opposite, which would orphan
* .gitnexus/ directories on disk.
* - `-f` / `--force` matches the confirmation-skip semantics of
* `clean -f`. (Distinct from `analyze --force`, which re-indexes;
* here there is no pipeline, so no conflation.)
*/
import fs from 'fs/promises';
import {
readRegistry,
resolveRegistryEntry,
assertSafeStoragePath,
unregisterRepo,
RegistryNotFoundError,
RegistryAmbiguousTargetError,
UnsafeStoragePathError,
} from '../storage/repo-manager.js';
export const removeCommand = async (target: string, options?: { force?: boolean }) => {
// Read the registry snapshot once and pass it to the resolver — this
// lets us render the "before" state in the dry-run path without a
// second disk read.
const entries = await readRegistry();
let entry;
try {
entry = resolveRegistryEntry(entries, target);
} catch (err) {
if (err instanceof RegistryNotFoundError) {
// Idempotent: missing target is a no-op warning, not an error.
// The `availableNames` hint comes from the error itself so users
// can see what they might have meant.
console.warn(`Nothing to remove: ${err.message}`);
return;
}
if (err instanceof RegistryAmbiguousTargetError) {
// Duplicate aliases are allowed via --allow-duplicate-name (#829);
// refuse to guess which one the user meant — surface the full list
// and exit non-zero so scripts don't silently pick the wrong repo.
console.error(`Error: ${err.message}`);
process.exit(1);
}
throw err;
}
// Confirmation gate — same shape as `clean`. Default is a dry-run
// that describes what would be deleted; `--force` actually deletes.
if (!options?.force) {
console.log(`This will delete the GitNexus index for: ${entry.name}`);
console.log(` Path: ${entry.path}`);
console.log(` Storage: ${entry.storagePath}`);
console.log('\nRun with --force to confirm deletion.');
return;
}
// Safety guard (#1003 review — @magyargergo): refuse to proceed if
// the registry entry's `storagePath` isn't the canonical
// `<entry.path>/.gitnexus` subfolder. `~/.gitnexus/registry.json` is
// user-writable, so a corrupted or hand-edited entry could point
// storagePath at the repo root, an empty string (→ cwd), a parent
// dir, or anywhere else; `fs.rm(recursive: true, force: true)` on
// any of those would be a runtime disaster. Bail before touching
// disk, with an actionable hint for recovering a broken registry.
try {
assertSafeStoragePath(entry);
} catch (err) {
if (err instanceof UnsafeStoragePathError) {
console.error(`Error: ${err.message}`);
process.exit(1);
}
throw err;
}
// Deletion order: fs.rm first, then unregister. If fs.rm fails mid-way,
// the registry entry stays so the user can retry. If fs.rm succeeds but
// unregister throws (e.g. ENOSPC on registry write), the entry becomes
// orphaned — `listRegisteredRepos({ validate: true })` prunes those on
// next read, so the failure is self-healing.
try {
await fs.rm(entry.storagePath, { recursive: true, force: true });
await unregisterRepo(entry.path);
console.log(`Removed: ${entry.name}`);
console.log(` Path: ${entry.path}`);
console.log(` Storage: ${entry.storagePath}`);
} catch (err) {
console.error(`Failed to remove ${entry.name}:`, err);
process.exit(1);
}
};
+82 -6
View File
@@ -13,6 +13,7 @@ import { execFile, execFileSync } from 'child_process';
import { promisify } from 'util';
import { fileURLToPath } from 'url';
import { glob } from 'glob';
import { parseTree, modify, applyEdits, ParseError } from 'jsonc-parser';
import { getGlobalDir } from '../storage/repo-manager.js';
const __filename = fileURLToPath(import.meta.url);
@@ -75,6 +76,23 @@ function getMcpEntry() {
};
}
/**
* OpenCode uses a different MCP format: { type: "local", command: [...] }
* where command is a flat array (command + args combined).
*/
function getOpenCodeMcpEntry() {
const bin = resolveGitnexusBin();
if (bin) {
return { type: 'local', command: [bin, 'mcp'] };
}
if (process.platform === 'win32') {
return { type: 'local', command: ['cmd', '/c', 'npx', '-y', 'gitnexus@latest', 'mcp'] };
}
return { type: 'local', command: ['npx', '-y', 'gitnexus@latest', 'mcp'] };
}
/**
* Merge gitnexus entry into an existing MCP config JSON object.
* Returns the updated config.
@@ -110,6 +128,62 @@ async function writeJsonFile(filePath: string, data: any): Promise<void> {
await fs.writeFile(filePath, JSON.stringify(data, null, 2) + '\n', 'utf-8');
}
/**
* Detect indentation style from file content.
* Returns formatting options matching the file's existing style.
*/
function detectIndentation(raw: string): { tabSize: number; insertSpaces: boolean } {
const firstIndented = raw.match(/^( +|\t)/m);
if (!firstIndented) return { tabSize: 2, insertSpaces: true };
if (firstIndented[1] === '\t') return { tabSize: 1, insertSpaces: false };
return { tabSize: firstIndented[1].length, insertSpaces: true };
}
/**
* Merge a key/value pair into a JSONC config file, preserving comments and formatting.
* If the file is genuinely corrupt (not valid JSONC), leaves it untouched.
*/
async function mergeJsoncFile(
filePath: string,
keyPath: string[],
value: unknown,
): Promise<boolean> {
let raw: string;
try {
raw = await fs.readFile(filePath, 'utf-8');
} catch {
raw = '';
}
if (raw.trim().length === 0) {
const config: any = {};
let parent: any = config;
for (let i = 0; i < keyPath.length; i++) {
if (i === keyPath.length - 1) {
parent[keyPath[i]] = value;
} else {
parent[keyPath[i]] = {};
parent = parent[keyPath[i]];
}
}
await writeJsonFile(filePath, config);
return true;
}
const parseErrors: ParseError[] = [];
const tree = parseTree(raw, parseErrors);
if (tree && tree.type === 'object' && parseErrors.length === 0) {
const formattingOptions = detectIndentation(raw);
const edits = modify(raw, keyPath, value, { formattingOptions });
const result = applyEdits(raw, edits);
await fs.writeFile(filePath, result, 'utf-8');
return true;
}
return false;
}
/**
* Check if a directory exists
*/
@@ -267,12 +341,14 @@ async function setupOpenCode(result: SetupResult): Promise<void> {
const configPath = path.join(opencodeDir, 'opencode.json');
try {
const existing = await readJsonFile(configPath);
const config = existing || {};
if (!config.mcp) config.mcp = {};
config.mcp.gitnexus = getMcpEntry();
await writeJsonFile(configPath, config);
result.configured.push('OpenCode');
const ok = await mergeJsoncFile(configPath, ['mcp', 'gitnexus'], getOpenCodeMcpEntry());
if (ok) {
result.configured.push('OpenCode');
} else {
result.errors.push(
'OpenCode: opencode.json is corrupt — skipping to preserve existing content',
);
}
} catch (err: any) {
result.errors.push(`OpenCode: ${err.message}`);
}
+52
View File
@@ -164,3 +164,55 @@ export async function cypherCommand(
});
output(result);
}
function formatDetectChangesResult(result: any): string {
if (result?.error) return `Error: ${result.error}`;
const summary = result?.summary || {};
if ((summary.changed_count || 0) === 0) {
return 'No changes detected.';
}
const lines: string[] = [];
lines.push(`Changes: ${summary.changed_files || 0} files, ${summary.changed_count || 0} symbols`);
lines.push(`Affected processes: ${summary.affected_count || 0}`);
lines.push(`Risk level: ${summary.risk_level || 'unknown'}`);
lines.push('');
const changed = result?.changed_symbols || [];
if (changed.length > 0) {
lines.push('Changed symbols:');
for (const symbol of changed.slice(0, 15)) {
lines.push(` ${symbol.type} ${symbol.name} → ${symbol.filePath}`);
}
if (changed.length > 15) {
lines.push(` ... and ${changed.length - 15} more`);
}
lines.push('');
}
const affected = result?.affected_processes || [];
if (affected.length > 0) {
lines.push('Affected execution flows:');
for (const processInfo of affected.slice(0, 10)) {
const steps = (processInfo.changed_steps || []).map((s: any) => s.symbol).join(', ');
lines.push(` • ${processInfo.name} (${processInfo.step_count} steps) — changed: ${steps}`);
}
}
return lines.join('\n').trim();
}
export async function detectChangesCommand(options?: {
scope?: string;
baseRef?: string;
repo?: string;
}): Promise<void> {
const backend = await getBackend();
const result = await backend.callTool('detect_changes', {
scope: options?.scope || 'unstaged',
base_ref: options?.baseRef,
repo: options?.repo,
});
output(formatDetectChangesResult(result));
}
+18 -18
View File
@@ -64,16 +64,16 @@ const FUNCTION_LIKE_TYPES = new Set([
* numbers don't apply.
*/
export const findFunctionNode = (root: any): any | null => {
if (FUNCTION_LIKE_TYPES.has(root.type)) return root;
for (let i = 0; i < root.namedChildCount; i++) {
const child = root.namedChild(i);
if (!child) continue;
if (FUNCTION_LIKE_TYPES.has(child.type)) return child;
const found = findFunctionNode(child);
if (found) return found;
// Iterative DFS — avoids stack overflow on deeply nested ASTs.
const stack = [root];
while (stack.length > 0) {
const node = stack.pop()!;
if (FUNCTION_LIKE_TYPES.has(node.type)) return node;
for (let i = node.namedChildCount - 1; i >= 0; i--) {
const child = node.namedChild(i);
if (child) stack.push(child);
}
}
return null;
};
@@ -98,15 +98,15 @@ export const findDeclarationNode = (root: any): any | null => {
'impl_item', // Rust: impl
]);
if (CLASS_LIKE_TYPES.has(root.type)) return root;
for (let i = 0; i < root.namedChildCount; i++) {
const child = root.namedChild(i);
if (!child) continue;
if (CLASS_LIKE_TYPES.has(child.type)) return child;
const found = findDeclarationNode(child);
if (found) return found;
// Iterative DFS — avoids stack overflow on deeply nested ASTs.
const stack = [root];
while (stack.length > 0) {
const node = stack.pop()!;
if (CLASS_LIKE_TYPES.has(node.type)) return node;
for (let i = node.namedChildCount - 1; i >= 0; i--) {
const child = node.namedChild(i);
if (child) stack.push(child);
}
}
return null;
};
+50 -27
View File
@@ -13,6 +13,12 @@ import { characterChunk } from './character-chunk.js';
import type { Chunk } from './character-chunk.js';
import { ensureAndParse, findDeclarationNode, findFunctionNode } from './ast-utils.js';
import { buildLineIndex, resolveChunkLines } from './line-index.js';
import {
CHUNKING_RULES,
CHUNK_MODE_AST_DECLARATION,
CHUNK_MODE_AST_FUNCTION,
type ChunkingRule,
} from './types.js';
/**
* Main chunkNode function: dispatches by label
@@ -40,31 +46,39 @@ export const chunkNode = async (
];
}
// Only function-like labels get AST chunking
if (label === 'Function' || label === 'Method' || label === 'Constructor') {
try {
const astChunks = await astChunk(content, filePath, startLine, endLine, chunkSize, overlap);
if (astChunks.length > 0) return astChunks;
} catch {
// AST parsing failed — fall through to character fallback
}
const rule = CHUNKING_RULES[label];
if (!rule) {
return characterChunk(content, startLine, endLine, chunkSize, overlap);
}
if (label === 'Class' || label === 'Interface') {
try {
const declarationChunks = await declarationChunk(
label,
try {
if (rule.mode === CHUNK_MODE_AST_FUNCTION) {
const astChunks = await astChunk(
content,
filePath,
startLine,
endLine,
chunkSize,
overlap,
rule,
);
if (astChunks.length > 0) return astChunks;
}
if (rule.mode === CHUNK_MODE_AST_DECLARATION) {
const declarationChunks = await declarationChunk(
content,
filePath,
startLine,
endLine,
chunkSize,
overlap,
rule,
);
if (declarationChunks.length > 0) return declarationChunks;
} catch {
// AST parsing failed — fall through to character fallback
}
} catch {
// AST parsing failed — fall through to character fallback
}
// Character-based fallback for everything else
@@ -83,6 +97,7 @@ const astChunk = async (
endLine: number,
chunkSize: number,
overlap: number,
rule: ChunkingRule,
): Promise<Chunk[]> => {
const tree = await ensureAndParse(content, filePath);
if (!tree) return [];
@@ -121,8 +136,8 @@ const astChunk = async (
statements,
targetNode.startIndex,
targetNode.endIndex,
true,
true,
rule.includePrefix,
rule.includeSuffix,
);
};
@@ -145,13 +160,13 @@ const FIELD_LIKE_MEMBER_TYPES = new Set([
]);
const declarationChunk = async (
label: 'Class' | 'Interface',
content: string,
filePath: string,
startLine: number,
endLine: number,
chunkSize: number,
overlap: number,
rule: ChunkingRule,
): Promise<Chunk[]> => {
const tree = await ensureAndParse(content, filePath);
if (!tree) return [];
@@ -162,7 +177,7 @@ const declarationChunk = async (
const bodyNode = getDeclarationBodyNode(targetNode);
if (!bodyNode) return [];
const members = collectDeclarationUnits(bodyNode, label);
const members = collectDeclarationUnits(bodyNode, rule.groupFields);
if (members.length === 0) return [];
return chunkByUnits(
@@ -174,8 +189,8 @@ const declarationChunk = async (
members,
targetNode.startIndex,
targetNode.endIndex,
false,
false,
rule.includePrefix,
rule.includeSuffix,
);
};
@@ -237,14 +252,22 @@ const chunkByUnits = (
if (candidateEndOffset - chunkStartOffset > chunkSize) {
const oversizedUnit = units[chunkStartUnitIdx];
const oversizedStartOffset =
chunkStartUnitIdx === 0 && includeContainerPrefixOnFirstChunk
? containerStartOffset
: oversizedUnit.startIndex;
const oversizedEndOffset =
chunkStartUnitIdx === units.length - 1 && includeContainerSuffixOnLastChunk
? containerEndOffset
: oversizedUnit.endIndex;
const oversizedLineRange = resolveChunkLines(
lineOffsets,
oversizedUnit.startIndex,
oversizedUnit.endIndex,
oversizedStartOffset,
oversizedEndOffset,
baseStartLine,
);
const oversizedChunks = characterChunk(
content.slice(oversizedUnit.startIndex, oversizedUnit.endIndex),
content.slice(oversizedStartOffset, oversizedEndOffset),
oversizedLineRange.startLine,
oversizedLineRange.endLine,
chunkSize,
@@ -252,8 +275,8 @@ const chunkByUnits = (
).map((chunk, offsetIdx) => ({
...chunk,
chunkIndex: chunks.length + offsetIdx,
startOffset: chunk.startOffset + oversizedUnit.startIndex,
endOffset: chunk.endOffset + oversizedUnit.startIndex,
startOffset: chunk.startOffset + oversizedStartOffset,
endOffset: chunk.endOffset + oversizedStartOffset,
}));
chunks.push(...oversizedChunks);
chunkStartUnitIdx += 1;
@@ -325,7 +348,7 @@ const getDeclarationBodyNode = (node: any): any | null => {
const collectDeclarationUnits = (
bodyNode: any,
label: 'Class' | 'Interface',
groupFields: boolean,
): Array<{ startIndex: number; endIndex: number }> => {
const members: Array<{ startIndex: number; endIndex: number; groupable: boolean }> = [];
@@ -335,7 +358,7 @@ const collectDeclarationUnits = (
members.push({
startIndex: child.startIndex,
endIndex: child.endIndex,
groupable: label === 'Class' && FIELD_LIKE_MEMBER_TYPES.has(child.type),
groupable: groupFields && FIELD_LIKE_MEMBER_TYPES.has(child.type),
});
}
@@ -30,6 +30,7 @@ import {
DEFAULT_EMBEDDING_CONFIG,
EMBEDDABLE_LABELS,
isShortLabel,
LABEL_METHOD,
LABELS_WITH_EXPORTED,
STRUCTURAL_LABELS,
collectBestChunks,
@@ -43,6 +44,12 @@ import {
import { loadVectorExtension } from '../lbug/lbug-adapter.js';
const isDev = process.env.NODE_ENV === 'development';
/**
* Bump this when the embedding text template changes in a way that should
* invalidate existing vectors, such as metadata/header shape changes,
* structural container context changes, or preceding-context formatting rules.
*/
export const EMBEDDING_TEXT_VERSION = 'v2';
/**
* Compute a stable content fingerprint for an embeddable node.
@@ -57,12 +64,13 @@ export const contentHashForNode = (
// Hash must be deterministic across runs, so exclude methodNames/fieldNames
// which are populated during the batch loop via AST extraction.
// Using only node.content ensures the hash stays stable.
// NOTE: A change to extractStructuralNames behavior requires bumping EMBEDDING_TEXT_VERSION.
const text = generateEmbeddingText(
{ ...node, methodNames: undefined, fieldNames: undefined },
node.content,
config,
);
return createHash('sha1').update(text).digest('hex');
return createHash('sha1').update(EMBEDDING_TEXT_VERSION).update('\n').update(text).digest('hex');
};
/**
@@ -83,7 +91,7 @@ const queryEmbeddableNodes = async (
try {
let query: string;
if (label === 'Method') {
if (label === LABEL_METHOD) {
// Method has parameterCount and returnType
query = `
MATCH (n:Method)
@@ -115,7 +123,7 @@ const queryEmbeddableNodes = async (
const rows = await executeQuery(query);
for (const row of rows) {
const hasExportedColumn = label === 'Method' || LABELS_WITH_EXPORTED.has(label);
const hasExportedColumn = label === LABEL_METHOD || LABELS_WITH_EXPORTED.has(label);
allNodes.push({
id: row.id ?? row[0],
name: row.name ?? row[1],
@@ -126,7 +134,7 @@ const queryEmbeddableNodes = async (
endLine: row.endLine ?? row[6],
isExported: hasExportedColumn ? (row.isExported ?? row[7]) : undefined,
description: row.description ?? (hasExportedColumn ? row[8] : row[7]),
...(label === 'Method'
...(label === LABEL_METHOD
? {
parameterCount: row.parameterCount ?? row[9],
returnType: row.returnType ?? row[10],
@@ -415,8 +423,15 @@ export const runEmbeddingPipeline = async (
}
}
let prevTail = '';
for (const chunk of chunks) {
const text = generateEmbeddingText(node, chunk.text, finalConfig);
const text = generateEmbeddingText(
node,
chunk.text,
finalConfig,
chunk.chunkIndex,
prevTail,
);
allTexts.push(text);
allUpdates.push({
nodeId: node.id,
@@ -425,6 +440,7 @@ export const runEmbeddingPipeline = async (
endLine: chunk.endLine,
contentHash: hash,
});
prevTail = overlap > 0 ? chunk.text.slice(-overlap) : '';
}
}
+45 -26
View File
@@ -10,7 +10,12 @@
*/
import type { EmbeddableNode, EmbeddingConfig } from './types.js';
import { DEFAULT_EMBEDDING_CONFIG, isShortLabel } from './types.js';
import {
CHUNKING_RULES,
DEFAULT_EMBEDDING_CONFIG,
STRUCTURAL_TEXT_MODE_DECLARATION,
isShortLabel,
} from './types.js';
/**
* Truncate description to max length at sentence/word boundary
@@ -95,47 +100,62 @@ const generateCodeBodyText = (
node: EmbeddableNode,
codeBody: string,
config: Partial<EmbeddingConfig>,
prevTail?: string,
): string => {
const header = buildMetadataHeader(node, config);
const cleaned = cleanContent(codeBody);
return `${header}\n\n${cleaned}`;
const parts = [header];
if (prevTail) {
parts.push(`[preceding context]: ...${cleanContent(prevTail)}`);
}
parts.push('', cleanContent(codeBody));
return parts.join('\n');
};
/**
* Generate embedding text for Class nodes
* Signature + properties + method name list only (no method bodies)
* Method/field names come from AST extractors via node.methodNames/node.fieldNames.
*/
const generateClassText = (
node: EmbeddableNode,
codeBody: string,
config: Partial<EmbeddingConfig>,
): string => {
return generateStructuralTypeText(node, codeBody, config);
const getCompactContainerContext = (
cleanedContent: string,
declarationOnly: string,
): string | undefined => {
const source = declarationOnly || cleanedContent;
const nlIdx = source.indexOf('\n');
const firstLine = (nlIdx === -1 ? source : source.substring(0, nlIdx)).trim();
return firstLine ? `Container: ${firstLine}` : undefined;
};
const generateStructuralTypeText = (
node: EmbeddableNode,
codeBody: string,
config: Partial<EmbeddingConfig>,
chunkIndex?: number,
prevTail?: string,
): string => {
const header = buildMetadataHeader(node, config);
const parts: string[] = [header];
const isFirstChunk = chunkIndex === undefined || chunkIndex === 0;
const cleanedContent = cleanContent(node.content);
const declarationOnly = extractDeclarationOnly(cleanedContent);
const compactContainerContext = getCompactContainerContext(cleanedContent, declarationOnly);
if (node.methodNames?.length) {
if (compactContainerContext) {
parts.push(compactContainerContext);
}
if (prevTail) {
parts.push(`[preceding context]: ...${cleanContent(prevTail)}`);
}
if (isFirstChunk && node.methodNames?.length) {
parts.push(`Methods: ${node.methodNames.join(', ')}`);
}
if (node.fieldNames?.length) {
if (isFirstChunk && node.fieldNames?.length) {
parts.push(`Properties: ${node.fieldNames.join(', ')}`);
}
const declarationOnly = extractDeclarationOnly(cleanContent(node.content));
if (declarationOnly) {
if (isFirstChunk && declarationOnly) {
parts.push('', declarationOnly);
}
const cleanedChunk = cleanContent(codeBody);
if (cleanedChunk && cleanedChunk !== cleanContent(node.content)) {
if (cleanedChunk && cleanedChunk !== cleanedContent) {
parts.push('', cleanedChunk);
}
@@ -229,6 +249,8 @@ export const generateEmbeddingText = (
node: EmbeddableNode,
codeBody: string,
config: Partial<EmbeddingConfig> = {},
chunkIndex?: number,
prevTail?: string,
): string => {
if (isShortLabel(node.label)) {
const header = buildMetadataHeader(node, config);
@@ -236,15 +258,12 @@ export const generateEmbeddingText = (
return `${header}\n\n${cleaned}`;
}
if (node.label === 'Class') {
return generateClassText(node, codeBody, config);
const chunkingRule = CHUNKING_RULES[node.label];
if (chunkingRule?.structuralTextMode === STRUCTURAL_TEXT_MODE_DECLARATION) {
return generateStructuralTypeText(node, codeBody, config, chunkIndex, prevTail);
}
if (node.label === 'Interface') {
return generateStructuralTypeText(node, codeBody, config);
}
return generateCodeBodyText(node, codeBody, config);
return generateCodeBodyText(node, codeBody, config, prevTail);
};
/**
+122 -29
View File
@@ -4,35 +4,76 @@
* Type definitions for the embedding generation and semantic search system.
*/
export const LABEL_FUNCTION = 'Function' as const;
export const LABEL_METHOD = 'Method' as const;
export const LABEL_CONSTRUCTOR = 'Constructor' as const;
export const LABEL_CLASS = 'Class' as const;
export const LABEL_INTERFACE = 'Interface' as const;
export const LABEL_STRUCT = 'Struct' as const;
export const LABEL_ENUM = 'Enum' as const;
export const LABEL_TRAIT = 'Trait' as const;
export const LABEL_IMPL = 'Impl' as const;
export const LABEL_MACRO = 'Macro' as const;
export const LABEL_NAMESPACE = 'Namespace' as const;
export const LABEL_TYPE_ALIAS = 'TypeAlias' as const;
export const LABEL_TYPEDEF = 'Typedef' as const;
export const LABEL_CONST = 'Const' as const;
export const LABEL_PROPERTY = 'Property' as const;
export const LABEL_RECORD = 'Record' as const;
export const LABEL_UNION = 'Union' as const;
export const LABEL_STATIC = 'Static' as const;
export const LABEL_VARIABLE = 'Variable' as const;
export const LABEL_CODE_ELEMENT = 'CodeElement' as const;
export const CHUNK_MODE_AST_FUNCTION = 'ast-function' as const;
export const CHUNK_MODE_AST_DECLARATION = 'ast-declaration' as const;
// CHUNK_MODE_CHARACTER exists for type completeness but is a no-op in CHUNKING_RULES —
// omit the entry entirely to get character fallback via chunker.ts dispatch.
export const CHUNK_MODE_CHARACTER = 'character' as const;
export const STRUCTURAL_TEXT_MODE_NONE = 'none' as const;
export const STRUCTURAL_TEXT_MODE_DECLARATION = 'declaration' as const;
export interface ChunkingRule {
mode:
| typeof CHUNK_MODE_AST_FUNCTION
| typeof CHUNK_MODE_AST_DECLARATION
| typeof CHUNK_MODE_CHARACTER;
includePrefix: boolean;
includeSuffix: boolean;
groupFields: boolean;
structuralTextMode: typeof STRUCTURAL_TEXT_MODE_NONE | typeof STRUCTURAL_TEXT_MODE_DECLARATION;
}
/**
* Node labels that need chunking (have code body, potentially long)
*/
export const CHUNKABLE_LABELS = [
'Function',
'Method',
'Constructor',
'Class',
'Interface',
'Struct',
'Enum',
'Trait',
'Impl',
'Macro',
'Namespace',
LABEL_FUNCTION,
LABEL_METHOD,
LABEL_CONSTRUCTOR,
LABEL_CLASS,
LABEL_INTERFACE,
LABEL_STRUCT,
LABEL_ENUM,
LABEL_TRAIT,
LABEL_IMPL,
LABEL_MACRO,
LABEL_NAMESPACE,
] as const;
/**
* Node labels that are short (no chunking needed, embed directly)
*/
export const SHORT_LABELS = [
'TypeAlias',
'Typedef',
'Const',
'Property',
'Record',
'Union',
'Static',
'Variable',
LABEL_TYPE_ALIAS,
LABEL_TYPEDEF,
LABEL_CONST,
LABEL_PROPERTY,
LABEL_RECORD,
LABEL_UNION,
LABEL_STATIC,
LABEL_VARIABLE,
] as const;
/**
@@ -61,26 +102,78 @@ export const isShortLabel = (label: string): boolean =>
(SHORT_LABELS as readonly string[]).includes(label);
/**
* Node labels that have structural names (methods/fields) extractable via AST
* Node labels that have structural names (methods/fields) extractable via AST.
* Only labels that consume methodNames/fieldNames in their embedding text should
* be listed here — extra entries trigger wasted AST parses with no effect on output.
*/
export const STRUCTURAL_LABELS: ReadonlySet<string> = new Set([
'Class',
'Struct',
'Interface',
'Enum',
LABEL_CLASS,
LABEL_STRUCT,
LABEL_INTERFACE,
]);
/**
* Node labels that have isExported column in their schema
*/
export const LABELS_WITH_EXPORTED = new Set([
'Function',
'Class',
'Interface',
'Method',
'CodeElement',
LABEL_FUNCTION,
LABEL_CLASS,
LABEL_INTERFACE,
LABEL_METHOD,
LABEL_CODE_ELEMENT,
]) as ReadonlySet<string>;
/**
* Labels that need special chunking and/or structural text semantics.
* Any chunkable label omitted here intentionally falls back to characterChunk
* plus generateCodeBodyText (for example Enum/Trait/Impl/Macro/Namespace).
*/
type ChunkableLabel = (typeof CHUNKABLE_LABELS)[number];
export const CHUNKING_RULES: Readonly<Partial<Record<ChunkableLabel, ChunkingRule>>> = {
[LABEL_FUNCTION]: {
mode: CHUNK_MODE_AST_FUNCTION,
includePrefix: true,
includeSuffix: true,
groupFields: false,
structuralTextMode: STRUCTURAL_TEXT_MODE_NONE,
},
[LABEL_METHOD]: {
mode: CHUNK_MODE_AST_FUNCTION,
includePrefix: true,
includeSuffix: true,
groupFields: false,
structuralTextMode: STRUCTURAL_TEXT_MODE_NONE,
},
[LABEL_CONSTRUCTOR]: {
mode: CHUNK_MODE_AST_FUNCTION,
includePrefix: true,
includeSuffix: true,
groupFields: false,
structuralTextMode: STRUCTURAL_TEXT_MODE_NONE,
},
[LABEL_CLASS]: {
mode: CHUNK_MODE_AST_DECLARATION,
includePrefix: true,
includeSuffix: false,
groupFields: true,
structuralTextMode: STRUCTURAL_TEXT_MODE_DECLARATION,
},
[LABEL_INTERFACE]: {
mode: CHUNK_MODE_AST_DECLARATION,
includePrefix: true,
includeSuffix: false,
groupFields: false,
structuralTextMode: STRUCTURAL_TEXT_MODE_DECLARATION,
},
[LABEL_STRUCT]: {
mode: CHUNK_MODE_AST_DECLARATION,
includePrefix: true,
includeSuffix: false,
groupFields: true,
structuralTextMode: STRUCTURAL_TEXT_MODE_DECLARATION,
},
};
/**
* Embedding pipeline phases
*/
+111
View File
@@ -4,6 +4,9 @@
*/
import { execFileSync } from 'node:child_process';
import path from 'path';
import { readRegistry, type RegistryEntry, type CwdMatch } from '../storage/repo-manager.js';
import { getGitRoot, getCurrentCommit, getRemoteUrl } from '../storage/git.js';
export interface StalenessInfo {
isStale: boolean;
@@ -37,3 +40,111 @@ export function checkStaleness(repoPath: string, lastCommit: string): StalenessI
return { isStale: false, commitsBehind: 0 };
}
}
/**
* Compare a sibling-clone HEAD against an indexed `lastCommit`. Returns
* `undefined` when the indexed commit is not reachable from the sibling
* (e.g. divergent branches, shallow clone, missing ref). The caller
* should treat `undefined` as "drift unknown" rather than "no drift".
*/
function commitsAheadOfIndexed(siblingPath: string, indexedCommit: string): number | undefined {
if (!indexedCommit) return undefined;
try {
const result = execFileSync('git', ['rev-list', '--count', `${indexedCommit}..HEAD`], {
cwd: siblingPath,
encoding: 'utf-8',
stdio: ['pipe', 'pipe', 'pipe'],
}).trim();
return parseInt(result, 10) || 0;
} catch {
return undefined;
}
}
/**
* Resolve a working directory against the global registry. Returns:
* - `match: 'path'` when `cwd` is inside a registered entry's path
* - `match: 'sibling-by-remote'` when `cwd` lives in a different on-disk clone
* of the same repo (same `remoteUrl`)
* - `match: 'none'` when neither match applies
*
* For sibling-by-remote matches, the caller's HEAD and the drift vs the
* indexed `lastCommit` are also returned so the MCP layer can warn
* before serving silently-stale answers (issue: silent graph drift
* across sibling clones).
*
* `path` matches deliberately use the longest-prefix rule so a cwd
* inside a sub-path of a registered repo still matches that repo, not
* a coincidentally-aliased shorter entry.
*/
export async function checkCwdMatch(cwd: string): Promise<CwdMatch> {
const entries = await readRegistry();
if (entries.length === 0) return { match: 'none' };
const isWin = process.platform === 'win32';
const norm = (p: string) => (isWin ? path.resolve(p).toLowerCase() : path.resolve(p));
const sep = path.sep;
const cwdResolved = path.resolve(cwd);
const cwdNorm = norm(cwdResolved);
// 1) Path-based match (longest prefix wins, boundary-safe).
let bestPath: RegistryEntry | undefined;
let bestLen = -1;
for (const e of entries) {
const p = norm(e.path);
if (cwdNorm === p || cwdNorm.startsWith(p + sep)) {
if (p.length > bestLen) {
bestPath = e;
bestLen = p.length;
}
}
}
if (bestPath) return { match: 'path', entry: bestPath };
// 2) Sibling-by-remote: locate the cwd's git root, get its remote
// URL, and look for any registered entry with the same fingerprint.
const cwdGitRoot = getGitRoot(cwdResolved);
if (!cwdGitRoot) return { match: 'none' };
const cwdRemote = getRemoteUrl(cwdGitRoot);
if (!cwdRemote) return { match: 'none' };
const sibling = entries.find(
(e) => e.remoteUrl === cwdRemote && norm(e.path) !== norm(cwdGitRoot),
);
if (!sibling) return { match: 'none' };
const cwdHead = getCurrentCommit(cwdGitRoot) || undefined;
const drift = commitsAheadOfIndexed(cwdGitRoot, sibling.lastCommit);
// Same commit on both clones → still report match=sibling-by-remote
// (the relationship is real and useful to callers like list_repos /
// future tooling) but leave `hint` unset: there's nothing to warn
// about, and `maybeWarnSiblingDrift` already short-circuits this
// case independently. Surfacing a no-op hint would force callers
// to second-guess whether they need to display it.
let hint: string | undefined;
if (cwdHead && cwdHead === sibling.lastCommit) {
hint = undefined;
} else if (drift && drift > 0) {
hint =
`⚠️ Index for "${sibling.name}" was built at ${sibling.path}; ` +
`your cwd (${cwdGitRoot}) is a sibling clone that is ${drift} commit${drift > 1 ? 's' : ''} ` +
`ahead of the indexed commit. Results may be stale or incorrect — re-run \`gitnexus analyze\` ` +
`to refresh the index.`;
} else {
hint =
`⚠️ Index for "${sibling.name}" was built at ${sibling.path}; ` +
`your cwd (${cwdGitRoot}) is a sibling clone whose HEAD differs from the indexed commit. ` +
`Results may be stale or incorrect — re-run \`gitnexus analyze\` to refresh the index.`;
}
return {
match: 'sibling-by-remote',
entry: sibling,
cwdGitRoot,
cwdHead,
drift,
hint,
};
}
+110 -21
View File
@@ -1,35 +1,117 @@
import type { GraphNode, GraphRelationship } from 'gitnexus-shared';
import type { GraphNode, GraphRelationship, RelationshipType } from 'gitnexus-shared';
import { KnowledgeGraph } from './types.js';
/** Fresh empty iterator per call — `[].values()` returns a new
* exhausted iterator each invocation, so empty-type lookups don't
* share a single already-exhausted iterator across callers. */
function emptyRelIter(): IterableIterator<GraphRelationship> {
return ([] as GraphRelationship[]).values();
}
export const createKnowledgeGraph = (): KnowledgeGraph => {
const nodeMap = new Map<string, GraphNode>();
const relationshipMap = new Map<string, GraphRelationship>();
// Per-type index maintained alongside `relationshipMap`. Bucket
// values are `Map<id, Relationship>` so per-type iteration is cheap
// and per-edge removal is O(1). See plan
// docs/plans/2026-04-20-002-perf-parse-heritage-mro-plan.md (Unit 1).
const relationshipsByType = new Map<RelationshipType, Map<string, GraphRelationship>>();
// Reverse-adjacency index: nodeId → Set<relId> of every edge where
// this node appears as source OR target. Maintained on writeRel /
// deleteRel so `removeNode` can delete a node's edges in
// O(edges-touching-node) instead of O(total-edges).
const edgeIdsByNode = new Map<string, Set<string>>();
// File index: filePath → Set<nodeId>. Maintained on addNode /
// removeNode so `removeNodesByFile` reaches its file's nodes
// directly instead of scanning the whole node map.
const nodeIdsByFile = new Map<string, Set<string>>();
// Private helpers that encode the dual-index invariants in one
// place. All mutation paths go through these — adding a new
// mutation method only needs to call the helper, not remember to
// touch every index.
const addToBucket = <K, V>(map: Map<K, Set<V>>, key: K, value: V): void => {
let bucket = map.get(key);
if (bucket === undefined) {
bucket = new Set();
map.set(key, bucket);
}
bucket.add(value);
};
const removeFromBucket = <K, V>(map: Map<K, Set<V>>, key: K, value: V): void => {
const bucket = map.get(key);
if (bucket === undefined) return;
bucket.delete(value);
if (bucket.size === 0) map.delete(key);
};
const writeRel = (rel: GraphRelationship): void => {
relationshipMap.set(rel.id, rel);
let typeBucket = relationshipsByType.get(rel.type);
if (typeBucket === undefined) {
typeBucket = new Map();
relationshipsByType.set(rel.type, typeBucket);
}
typeBucket.set(rel.id, rel);
addToBucket(edgeIdsByNode, rel.sourceId, rel.id);
// Guard against a self-edge writing the same rel.id into the
// same Set twice — Set dedup handles it, but we skip explicitly
// for clarity.
if (rel.targetId !== rel.sourceId) {
addToBucket(edgeIdsByNode, rel.targetId, rel.id);
}
};
const deleteRel = (rel: GraphRelationship): void => {
relationshipMap.delete(rel.id);
const typeBucket = relationshipsByType.get(rel.type);
if (typeBucket !== undefined) {
typeBucket.delete(rel.id);
if (typeBucket.size === 0) relationshipsByType.delete(rel.type);
}
removeFromBucket(edgeIdsByNode, rel.sourceId, rel.id);
if (rel.targetId !== rel.sourceId) {
removeFromBucket(edgeIdsByNode, rel.targetId, rel.id);
}
};
const addNode = (node: GraphNode) => {
if (!nodeMap.has(node.id)) {
nodeMap.set(node.id, node);
if (nodeMap.has(node.id)) return;
nodeMap.set(node.id, node);
const filePath = node.properties?.filePath;
if (typeof filePath === 'string' && filePath.length > 0) {
addToBucket(nodeIdsByFile, filePath, node.id);
}
};
const addRelationship = (relationship: GraphRelationship) => {
if (!relationshipMap.has(relationship.id)) {
relationshipMap.set(relationship.id, relationship);
}
if (relationshipMap.has(relationship.id)) return;
writeRel(relationship);
};
/**
* Remove a single node and all relationships involving it
* Remove a single node and all relationships involving it.
* O(edges-touching-node) via the reverse-adjacency index — no full
* relationshipMap scan.
*/
const removeNode = (nodeId: string): boolean => {
if (!nodeMap.has(nodeId)) return false;
const node = nodeMap.get(nodeId);
if (node === undefined) return false;
nodeMap.delete(nodeId);
const filePath = node.properties?.filePath;
if (typeof filePath === 'string' && filePath.length > 0) {
removeFromBucket(nodeIdsByFile, filePath, nodeId);
}
// Remove all relationships involving this node
for (const [relId, rel] of relationshipMap) {
if (rel.sourceId === nodeId || rel.targetId === nodeId) {
relationshipMap.delete(relId);
const touchingEdgeIds = edgeIdsByNode.get(nodeId);
if (touchingEdgeIds !== undefined) {
// Snapshot the ids before iterating — deleteRel mutates the same
// Set via removeFromBucket, which would break mid-loop iteration.
for (const relId of [...touchingEdgeIds]) {
const rel = relationshipMap.get(relId);
if (rel !== undefined) deleteRel(rel);
}
edgeIdsByNode.delete(nodeId);
}
return true;
};
@@ -39,21 +121,24 @@ export const createKnowledgeGraph = (): KnowledgeGraph => {
* Returns true if the relationship existed and was removed, false otherwise.
*/
const removeRelationship = (relationshipId: string): boolean => {
return relationshipMap.delete(relationshipId);
const rel = relationshipMap.get(relationshipId);
if (rel === undefined) return false;
deleteRel(rel);
return true;
};
/**
* Remove all nodes (and their relationships) belonging to a file.
* O(file-nodes × avg-edges-per-node) via the file index — no full
* node-map scan.
*/
const removeNodesByFile = (filePath: string): number => {
let removed = 0;
for (const [nodeId, node] of nodeMap) {
if (node.properties?.filePath === filePath) {
removeNode(nodeId);
removed++;
}
}
return removed;
const nodeIds = nodeIdsByFile.get(filePath);
if (nodeIds === undefined) return 0;
// Snapshot before iterating — removeNode mutates nodeIdsByFile.
const snapshot = [...nodeIds];
for (const nodeId of snapshot) removeNode(nodeId);
return snapshot.length;
};
return {
@@ -67,6 +152,10 @@ export const createKnowledgeGraph = (): KnowledgeGraph => {
iterNodes: () => nodeMap.values(),
iterRelationships: () => relationshipMap.values(),
iterRelationshipsByType: (type: RelationshipType) => {
const bucket = relationshipsByType.get(type);
return bucket === undefined ? emptyRelIter() : bucket.values();
},
forEachNode(fn: (node: GraphNode) => void) {
nodeMap.forEach(fn);
},
+12 -1
View File
@@ -6,7 +6,7 @@
*
* This file only defines the CLI's KnowledgeGraph with mutation methods.
*/
import type { GraphNode, GraphRelationship } from 'gitnexus-shared';
import type { GraphNode, GraphRelationship, RelationshipType } from 'gitnexus-shared';
// CLI-specific: full KnowledgeGraph with mutation methods for incremental updates
export interface KnowledgeGraph {
@@ -14,6 +14,17 @@ export interface KnowledgeGraph {
relationships: GraphRelationship[];
iterNodes: () => IterableIterator<GraphNode>;
iterRelationships: () => IterableIterator<GraphRelationship>;
/**
* Iterate ONLY relationships of the given type, backed by a per-type
* index maintained in `addRelationship` / `removeRelationship` /
* `removeNode` / `removeNodesByFile`. Returns an empty iterator when
* the graph contains no relationships of that type.
*
* Prefer this over `iterRelationships()` + per-edge type filtering
* for hot paths (MRO setup, heritage walks). Backwards-compatible:
* existing `iterRelationships()` callers keep working.
*/
iterRelationshipsByType: (type: RelationshipType) => IterableIterator<GraphRelationship>;
forEachNode: (fn: (node: GraphNode) => void) => void;
forEachRelationship: (fn: (rel: GraphRelationship) => void) => void;
getNode: (id: string) => GraphNode | undefined;
+16 -1
View File
@@ -89,10 +89,25 @@ export function parseGroupConfig(yamlContent: string): GroupConfig {
};
}
export class GroupNotFoundError extends Error {
constructor(public readonly groupName: string) {
super(`Group "${groupName}" not found`);
this.name = 'GroupNotFoundError';
}
}
export async function loadGroupConfig(groupDir: string): Promise<GroupConfig> {
const fsp = await import('node:fs/promises');
const path = await import('node:path');
const yamlPath = path.join(groupDir, 'group.yaml');
const content = await fsp.readFile(yamlPath, 'utf-8');
let content: string;
try {
content = await fsp.readFile(yamlPath, 'utf-8');
} catch (err) {
if ((err as NodeJS.ErrnoException).code === 'ENOENT') {
throw new GroupNotFoundError(path.basename(groupDir));
}
throw err;
}
return parseGroupConfig(content);
}
+549
View File
@@ -0,0 +1,549 @@
/**
* Cross-repo impact (Phase 1 local walk + Phase 2 bridge fan-out).
* All bridge Cypher for this feature lives in this module.
*/
import fsp from 'node:fs/promises';
import path from 'node:path';
import type {
BridgeHandle,
ContractType,
CrossRepoImpact,
GroupConfig,
GroupImpactResult,
MatchType,
OutOfScopeLink,
} from './types.js';
import type { GroupRepoHandle, GroupToolPort } from './service.js';
import { GroupNotFoundError, loadGroupConfig } from './config-parser.js';
import {
fileMatchesServicePrefix,
normalizeServicePrefix,
repoInSubgroup,
} from './group-path-utils.js';
import { getGroupDir } from './storage.js';
import { closeBridgeDb, openBridgeDbReadOnly, queryBridge, readBridgeMeta } from './bridge-db.js';
import { BRIDGE_SCHEMA_VERSION } from './bridge-schema.js';
/** Cross-boundary hops beyond this value are clamped (multi-hop reserved for future work). */
export const MAX_SUPPORTED_CROSS_DEPTH = 1;
/** Default wall-clock budget for the Phase 1 `impact` leg when callers omit `timeoutMs`. */
export const DEFAULT_LOCAL_IMPACT_TIMEOUT_MS = 30_000;
const CY_NEIGHBORS_UPSTREAM = `
MATCH (consumer:Contract)-[l:ContractLink]->(provider:Contract)
WHERE provider.repo = $localRepo
AND provider.symbolUid IN $uids
AND provider.role = 'provider'
RETURN consumer.repo AS neighborRepo,
consumer.symbolUid AS neighborUid,
consumer.filePath AS neighborFilePath,
l.matchType AS matchType,
l.confidence AS confidence,
l.contractId AS contractId,
consumer.type AS contractType
`;
const CY_NEIGHBORS_DOWNSTREAM = `
MATCH (consumer:Contract)-[l:ContractLink]->(provider:Contract)
WHERE consumer.repo = $localRepo
AND consumer.symbolUid IN $uids
AND consumer.role = 'consumer'
RETURN provider.repo AS neighborRepo,
provider.symbolUid AS neighborUid,
provider.filePath AS neighborFilePath,
l.matchType AS matchType,
l.confidence AS confidence,
l.contractId AS contractId,
provider.type AS contractType
`;
type BridgeNeighborRow = {
neighborRepo: string;
neighborUid: string;
neighborFilePath?: string;
matchType: string;
confidence: number;
contractId: string;
contractType: string;
};
export interface RunGroupImpactDeps {
port: GroupToolPort;
gitnexusDir: string;
}
function parseDirection(raw: unknown): 'upstream' | 'downstream' | null {
if (raw === 'upstream' || raw === 'downstream') return raw;
return null;
}
function clampCrossDepth(raw: unknown): { depth: number; warning?: string } {
const n = typeof raw === 'number' && Number.isFinite(raw) ? Math.floor(raw) : 1;
const d = n < 1 ? 1 : n;
if (d > MAX_SUPPORTED_CROSS_DEPTH) {
return {
depth: MAX_SUPPORTED_CROSS_DEPTH,
warning: `crossDepth was ${d}; multi-hop cross-boundary traversal beyond ${MAX_SUPPORTED_CROSS_DEPTH} is not implemented yet. Using crossDepth ${MAX_SUPPORTED_CROSS_DEPTH}.`,
};
}
return { depth: d };
}
export function validateGroupImpactParams(params: Record<string, unknown>):
| {
ok: true;
name: string;
repoPath: string;
target: string;
direction: 'upstream' | 'downstream';
maxDepth: number;
crossDepth: number;
crossDepthWarning?: string;
relationTypes?: string[];
includeTests: boolean;
minConfidence: number;
service?: string;
subgroup?: string;
timeoutMs: number;
}
| { ok: false; error: string } {
const name = String(params.name ?? '').trim();
const repoPath = String(params.repo ?? '').trim();
const target = String(params.target ?? '').trim();
if (!name) return { ok: false, error: 'name is required' };
if (!repoPath)
return { ok: false, error: 'repo is required (group repo path, e.g. app/backend)' };
if (!target) return { ok: false, error: 'target is required' };
if (
params.service !== undefined &&
params.service !== null &&
String(params.service).trim() === ''
) {
return { ok: false, error: 'service must not be an empty string' };
}
const direction = parseDirection(params.direction);
if (!direction) return { ok: false, error: 'direction must be upstream or downstream' };
let maxDepth = typeof params.maxDepth === 'number' && params.maxDepth > 0 ? params.maxDepth : 3;
if (maxDepth > 32) maxDepth = 32;
const { depth: crossDepth, warning: crossDepthWarning } = clampCrossDepth(params.crossDepth);
const relationTypes = Array.isArray(params.relationTypes)
? params.relationTypes.filter((t): t is string => typeof t === 'string')
: undefined;
const includeTests = Boolean(params.includeTests);
let minConfidence = typeof params.minConfidence === 'number' ? params.minConfidence : 0;
if (minConfidence < 0) minConfidence = 0;
if (minConfidence > 1) minConfidence = 1;
const service = normalizeServicePrefix(params.service);
const subgroup = typeof params.subgroup === 'string' ? params.subgroup : undefined;
let timeoutMs =
typeof params.timeoutMs === 'number' && params.timeoutMs > 0
? params.timeoutMs
: typeof params.timeout === 'number' && params.timeout > 0
? params.timeout
: DEFAULT_LOCAL_IMPACT_TIMEOUT_MS;
if (timeoutMs > 3_600_000) timeoutMs = 3_600_000;
return {
ok: true,
name,
repoPath,
target,
direction,
maxDepth,
crossDepth,
crossDepthWarning,
relationTypes,
includeTests,
minConfidence,
service,
subgroup,
timeoutMs,
};
}
async function resolveGroupRepo(
port: GroupToolPort,
config: GroupConfig,
repoPath: string,
): Promise<GroupRepoHandle | { error: string }> {
const registryName = config.repos[repoPath];
if (!registryName) {
return { error: `Unknown repo path "${repoPath}" in this group.` };
}
try {
return await port.resolveRepo(registryName);
} catch (e) {
return { error: e instanceof Error ? e.message : String(e) };
}
}
async function safeLocalImpact(
port: GroupToolPort,
repo: GroupRepoHandle,
impactParams: Parameters<GroupToolPort['impact']>[1],
timeoutMs: number,
): Promise<{ value: unknown; timedOut: boolean }> {
let timer: ReturnType<typeof setTimeout> | undefined;
const impactP = port.impact(repo, impactParams).catch((err) => ({
error: err instanceof Error ? err.message : String(err),
}));
const timeoutP = new Promise<'timeout'>((resolve) => {
timer = setTimeout(() => resolve('timeout'), timeoutMs);
});
const won = await Promise.race([
impactP.then((v) => ({ tag: 'impact' as const, v })),
timeoutP.then(() => ({ tag: 'timeout' as const })),
]);
if (timer !== undefined) clearTimeout(timer);
if (won.tag === 'timeout') {
return {
value: { error: 'Local impact timed out', partial: true },
timedOut: true,
};
}
return { value: won.v, timedOut: false };
}
export function collectImpactSymbolUids(
local: unknown,
servicePrefix: string | undefined,
): { uids: string[]; targetFilePath?: string } {
const uids = new Set<string>();
let targetFilePath: string | undefined;
const obj = local as Record<string, unknown> | null;
if (!obj || typeof obj !== 'object') return { uids: [], targetFilePath };
const target = obj.target as { id?: string; filePath?: string } | undefined;
if (target?.id) {
targetFilePath = typeof target.filePath === 'string' ? target.filePath : undefined;
if (fileMatchesServicePrefix(targetFilePath, servicePrefix)) {
uids.add(String(target.id));
}
}
const byDepth = obj.byDepth as Record<string | number, unknown> | undefined;
if (byDepth && typeof byDepth === 'object') {
for (const items of Object.values(byDepth)) {
if (!Array.isArray(items)) continue;
for (const it of items) {
const row = it as { id?: string; filePath?: string };
if (row?.id && fileMatchesServicePrefix(row.filePath, servicePrefix)) {
uids.add(String(row.id));
}
}
}
}
return { uids: [...uids], targetFilePath };
}
function extractProcessNames(impact: unknown): string[] {
const o = impact as { affected_processes?: Array<{ name?: string }> };
if (!o?.affected_processes) return [];
return o.affected_processes.map((p) => String(p.name ?? '')).filter(Boolean);
}
function mergeRisk(localRisk: string, cross: CrossRepoImpact[]): string {
const highConf = cross.some((c) => c.contract.confidence >= 0.85);
if (localRisk === 'CRITICAL') return 'CRITICAL';
if (cross.length >= 3) return 'CRITICAL';
if (highConf) return 'HIGH';
if (cross.length > 0 && (localRisk === 'LOW' || localRisk === 'UNKNOWN')) return 'MEDIUM';
return localRisk;
}
async function ensureBridgeReady(
groupDir: string,
): Promise<{ handle: BridgeHandle } | { error: string }> {
const meta = await readBridgeMeta(groupDir);
if (meta.version > 0 && meta.version !== BRIDGE_SCHEMA_VERSION) {
return {
error: `Bridge schema version mismatch (meta.json has ${meta.version}, expected ${BRIDGE_SCHEMA_VERSION}). Run gitnexus group sync for this group.`,
};
}
const dbPath = path.join(groupDir, 'bridge.lbug');
try {
await fsp.access(dbPath);
} catch {
return {
error: `No bridge.lbug in this group directory. Run gitnexus group sync (schema ${BRIDGE_SCHEMA_VERSION}).`,
};
}
const handle = await openBridgeDbReadOnly(groupDir);
if (!handle) {
return {
error: `Could not open bridge.lbug read-only (schema ${BRIDGE_SCHEMA_VERSION}). Run gitnexus group sync.`,
};
}
return { handle };
}
function rowToNeighbor(r: Record<string, unknown>): BridgeNeighborRow | null {
const neighborRepo = String(r.neighborRepo ?? r[0] ?? '');
const neighborUid = String(r.neighborUid ?? r[1] ?? '');
if (!neighborRepo || !neighborUid) return null;
return {
neighborRepo,
neighborUid,
neighborFilePath:
r.neighborFilePath !== undefined ? String(r.neighborFilePath) : String(r[2] ?? ''),
matchType: String(r.matchType ?? r[3] ?? 'exact'),
confidence: Number(r.confidence ?? r[4] ?? 0),
contractId: String(r.contractId ?? r[5] ?? ''),
contractType: String(r.contractType ?? r[6] ?? 'custom'),
};
}
export async function runGroupImpact(
deps: RunGroupImpactDeps,
params: Record<string, unknown>,
): Promise<GroupImpactResult | { error: string }> {
const parsed = validateGroupImpactParams(params);
if (parsed.ok === false) return { error: parsed.error };
const {
name,
repoPath,
target,
direction,
maxDepth,
crossDepth: _crossDepth,
crossDepthWarning,
relationTypes,
includeTests,
minConfidence,
service: servicePrefix,
subgroup,
timeoutMs,
} = parsed;
const groupDir = getGroupDir(deps.gitnexusDir, name);
let config: GroupConfig;
try {
config = await loadGroupConfig(groupDir);
} catch (e) {
if (e instanceof GroupNotFoundError)
return { error: `Group "${name}" not found. Run group_list to see configured groups.` };
return { error: e instanceof Error ? e.message : String(e) };
}
const resolved = await resolveGroupRepo(deps.port, config, repoPath);
if ('error' in resolved) return { error: resolved.error };
const impactParams: Parameters<GroupToolPort['impact']>[1] = {
target,
direction,
maxDepth,
relationTypes: relationTypes && relationTypes.length > 0 ? relationTypes : undefined,
includeTests,
minConfidence,
};
const deadline = Date.now() + Math.max(0, timeoutMs);
const { value: local, timedOut: localTimedOut } = await safeLocalImpact(
deps.port,
resolved,
impactParams,
timeoutMs,
);
if (localTimedOut) {
const _base = local as Record<string, unknown>;
return {
local,
group: name,
cross: [],
outOfScope: [],
truncated: true,
truncatedRepos: [],
summary: {
direct: 0,
processes_affected: 0,
modules_affected: 0,
cross_repo_hits: 0,
},
risk: 'UNKNOWN',
timeoutMs,
truncationReason: 'timeout',
crossDepthWarning,
};
}
const localObj = local as Record<string, unknown> | null;
if (localObj?.error && typeof localObj.error === 'string') {
// Fail closed: the local-impact phase errored (missing symbol, graph-load
// failure, thrown exception wrapped by safeLocalImpact, or port-returned
// `{ error }`). Do NOT wrap it into a zero-hit success payload — callers
// branch on top-level `error`, and a blast-radius tool reporting "no
// impact" on the failure path is a false negative on a safety-critical
// signal. Bubble the error so consumers treat it as a failure.
return { error: `Local impact failed for ${repoPath}: ${localObj.error}` };
}
if (servicePrefix) {
const tf = (localObj?.target as { filePath?: string } | undefined)?.filePath;
if (!fileMatchesServicePrefix(tf, servicePrefix)) {
return {
local: {},
group: name,
cross: [],
outOfScope: [],
truncated: false,
truncatedRepos: [],
summary: {
direct: 0,
processes_affected: 0,
modules_affected: 0,
cross_repo_hits: 0,
},
risk: 'LOW',
timeoutMs,
crossDepthWarning,
};
}
}
const { uids } = collectImpactSymbolUids(local, servicePrefix);
if (uids.length === 0) {
const s = (local as { summary?: Record<string, number> })?.summary || {};
return {
local,
group: name,
cross: [],
outOfScope: [],
truncated: Boolean((local as { partial?: boolean }).partial),
truncatedRepos: [],
summary: {
direct: s.direct ?? 0,
processes_affected: s.processes_affected ?? 0,
modules_affected: s.modules_affected ?? 0,
cross_repo_hits: 0,
},
risk: String((local as { risk?: string }).risk ?? 'LOW'),
timeoutMs,
truncationReason: (local as { partial?: boolean }).partial ? 'partial' : undefined,
crossDepthWarning,
};
}
const bridgePrep = await ensureBridgeReady(groupDir);
if ('error' in bridgePrep) return { error: bridgePrep.error };
const handle = bridgePrep.handle;
const cross: CrossRepoImpact[] = [];
const outOfScope: OutOfScopeLink[] = [];
const truncatedRepos: string[] = [];
try {
const cypher = direction === 'upstream' ? CY_NEIGHBORS_UPSTREAM : CY_NEIGHBORS_DOWNSTREAM;
const rows = await queryBridge<Record<string, unknown>>(handle, cypher, {
localRepo: repoPath,
uids,
});
const neighbors: BridgeNeighborRow[] = [];
for (const raw of rows) {
const n = rowToNeighbor(raw);
if (n) neighbors.push(n);
}
neighbors.sort((a, b) => b.confidence - a.confidence);
const seen = new Set<string>();
for (const n of neighbors) {
if (servicePrefix && !fileMatchesServicePrefix(n.neighborFilePath, servicePrefix)) {
continue;
}
if (!repoInSubgroup(n.neighborRepo, subgroup)) {
outOfScope.push({
from: direction === 'upstream' ? n.neighborRepo : repoPath,
to: direction === 'upstream' ? repoPath : n.neighborRepo,
contractId: n.contractId,
confidence: n.confidence,
});
continue;
}
const key = `${n.neighborRepo}\0${n.neighborUid}\0${n.contractId}`;
if (seen.has(key)) continue;
seen.add(key);
if (Date.now() > deadline) {
truncatedRepos.push(n.neighborRepo);
continue;
}
const regName = config.repos[n.neighborRepo];
if (!regName) continue;
let neighborHandle: GroupRepoHandle;
try {
neighborHandle = await deps.port.resolveRepo(regName);
} catch {
truncatedRepos.push(n.neighborRepo);
continue;
}
const fan = await deps.port.impactByUid(neighborHandle.id, n.neighborUid, direction, {
maxDepth,
relationTypes: relationTypes ?? [],
minConfidence,
includeTests,
});
if (fan == null) {
truncatedRepos.push(n.neighborRepo);
continue;
}
cross.push({
repo: regName,
repo_path: n.neighborRepo,
contract: {
id: n.contractId,
type: n.contractType as ContractType,
match_type: (n.matchType as MatchType) || 'exact',
confidence: n.confidence,
},
by_depth: ((fan as { byDepth?: unknown }).byDepth ?? {}) as Record<string, unknown[]>,
affected_processes: extractProcessNames(fan),
});
}
} finally {
await closeBridgeDb(handle);
}
const localSum = (local as { summary?: Record<string, number> })?.summary || {};
const localRisk = String((local as { risk?: string }).risk ?? 'LOW');
const localPartial = Boolean((local as { partial?: boolean }).partial);
const truncated = truncatedRepos.length > 0 || localPartial;
const result: GroupImpactResult = {
local,
group: name,
cross,
outOfScope,
truncated,
truncatedRepos: [...new Set(truncatedRepos)],
summary: {
direct: localSum.direct ?? 0,
processes_affected: localSum.processes_affected ?? 0,
modules_affected: localSum.modules_affected ?? 0,
cross_repo_hits: cross.length,
},
risk: mergeRisk(localRisk, cross),
timeoutMs,
truncationReason: truncated ? 'partial' : undefined,
crossDepthWarning,
};
return result;
}
export { normalizeServicePrefix, fileMatchesServicePrefix } from './group-path-utils.js';
@@ -3,33 +3,92 @@ import {
compilePatterns,
runCompiledPatterns,
unquoteLiteral,
type CompiledPatterns,
type LanguagePatterns,
type PatternSpec,
} from '../tree-sitter-scanner.js';
import type { HttpDetection, HttpLanguagePlugin } from './types.js';
/**
* PHP HTTP plugin — Laravel `Route::get/post/...` declarations.
* PHP HTTP plugin.
*
* Providers:
* - Laravel `Route::get/post/...`
*
* Consumers (string-literal URLs only):
* - Laravel HTTP client: `Http::get/post/put/delete/patch($url)`
* - Guzzle / generic object method: `$client->get/post/...($url)`
* - `file_get_contents($url)`
*
* The pipeline already uses `PHP.php_only` for ingesting plain `.php`
* files (see `core/tree-sitter/parser-loader.ts`), and we do the same
* here so Laravel route files are parsed with the right grammar dialect.
*
* Scope notes: consumer patterns match string literals only. URLs built
* via binary concatenation (`$base . '/path'`), `sprintf`, or config
* lookup (`config('services.foo.base').'/path'`) are intentionally left
* for a follow-up — they require constant-folding the surrounding
* scope to be meaningful.
*/
const LARAVEL_PATTERNS = compilePatterns({
name: 'php-laravel',
language: PHP.php_only,
patterns: [
{
meta: {},
query: `
(scoped_call_expression
scope: (name) @scope (#eq? @scope "Route")
name: (name) @method (#match? @method "^(get|post|put|delete|patch)$")
arguments: (arguments . (argument (string) @path)))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
const LARAVEL_ROUTE_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(scoped_call_expression
scope: (name) @scope (#eq? @scope "Route")
name: (name) @method (#match? @method "^(get|post|put|delete|patch)$")
arguments: (arguments . (argument (string) @path)))
`,
};
const HTTP_FACADE_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(scoped_call_expression
scope: (name) @scope (#eq? @scope "Http")
name: (name) @method (#match? @method "^(get|post|put|delete|patch)$")
arguments: (arguments . (argument (string) @path)))
`,
};
const GUZZLE_MEMBER_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(member_call_expression
name: (name) @method (#match? @method "^(get|post|put|delete|patch)$")
arguments: (arguments . (argument (string) @path)))
`,
};
const FILE_GET_CONTENTS_SPEC: PatternSpec<Record<string, never>> = {
meta: {},
query: `
(function_call_expression
function: (name) @fn (#eq? @fn "file_get_contents")
arguments: (arguments . (argument (string) @path)))
`,
};
interface PhpPatternBundle {
laravelRoute: CompiledPatterns<Record<string, never>>;
httpFacade: CompiledPatterns<Record<string, never>>;
guzzleMember: CompiledPatterns<Record<string, never>>;
fileGetContents: CompiledPatterns<Record<string, never>>;
}
const mk = (spec: PatternSpec<Record<string, never>>, suffix: string) =>
compilePatterns({
name: `php-${suffix}`,
language: PHP.php_only,
patterns: [spec],
} satisfies LanguagePatterns<Record<string, never>>);
const PHP_PATTERNS: PhpPatternBundle = {
laravelRoute: mk(LARAVEL_ROUTE_SPEC, 'laravel-route'),
httpFacade: mk(HTTP_FACADE_SPEC, 'http-facade'),
guzzleMember: mk(GUZZLE_MEMBER_SPEC, 'guzzle-member'),
fileGetContents: mk(FILE_GET_CONTENTS_SPEC, 'file-get-contents'),
};
/**
* Extract the inner text of a PHP `string` node. The tree-sitter-php
@@ -39,11 +98,8 @@ const LARAVEL_PATTERNS = compilePatterns({
* child nodes.
*/
function phpStringText(node: import('tree-sitter').SyntaxNode): string | null {
// Most single-quoted strings expose their inner content through the
// full node text (including quotes), which unquoteLiteral strips.
const direct = unquoteLiteral(node.text);
if (direct !== null && direct !== node.text) return direct;
// Fall back to child string_content / string_value node if present.
for (const child of node.children) {
if (child.type === 'string_content' || child.type === 'string_value') {
return child.text;
@@ -52,13 +108,32 @@ function phpStringText(node: import('tree-sitter').SyntaxNode): string | null {
return direct;
}
/**
* HTTP client helpers (`Http::`, Guzzle) are almost always called with
* a path relative to a configured base URL, or a full URL. File paths
* are rare. Accept both relative (`/api/...`) and absolute (`http(s)://`).
*/
function isHttpClientPath(path: string): boolean {
return path.startsWith('/') || path.startsWith('http://') || path.startsWith('https://');
}
/**
* `file_get_contents` is used for both HTTP and filesystem reads. Only
* emit a consumer contract when the URL is an absolute HTTP(S) URL to
* avoid false positives for local file paths and stream wrappers
* (`php://input`, `file://`, `data:`, ...).
*/
function isHttpUrlLiteral(path: string): boolean {
return path.startsWith('http://') || path.startsWith('https://');
}
export const PHP_HTTP_PLUGIN: HttpLanguagePlugin = {
name: 'php-http',
language: PHP.php_only,
scan(tree) {
const out: HttpDetection[] = [];
for (const match of runCompiledPatterns(LARAVEL_PATTERNS, tree)) {
for (const match of runCompiledPatterns(PHP_PATTERNS.laravelRoute, tree)) {
const methodNode = match.captures.method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
@@ -74,6 +149,53 @@ export const PHP_HTTP_PLUGIN: HttpLanguagePlugin = {
});
}
for (const match of runCompiledPatterns(PHP_PATTERNS.httpFacade, tree)) {
const methodNode = match.captures.method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const path = phpStringText(pathNode);
if (path === null || !isHttpClientPath(path)) continue;
out.push({
role: 'consumer',
framework: 'laravel-http',
method: methodNode.text.toUpperCase(),
path,
name: null,
confidence: 0.7,
});
}
for (const match of runCompiledPatterns(PHP_PATTERNS.guzzleMember, tree)) {
const methodNode = match.captures.method;
const pathNode = match.captures.path;
if (!methodNode || !pathNode) continue;
const path = phpStringText(pathNode);
if (path === null || !isHttpClientPath(path)) continue;
out.push({
role: 'consumer',
framework: 'guzzle',
method: methodNode.text.toUpperCase(),
path,
name: null,
confidence: 0.7,
});
}
for (const match of runCompiledPatterns(PHP_PATTERNS.fileGetContents, tree)) {
const pathNode = match.captures.path;
if (!pathNode) continue;
const path = phpStringText(pathNode);
if (path === null || !isHttpUrlLiteral(path)) continue;
out.push({
role: 'consumer',
framework: 'file-get-contents',
method: 'GET',
path,
name: null,
confidence: 0.7,
});
}
return out;
},
};
@@ -0,0 +1,42 @@
/**
* Shared service-path normalization for group tools (`service` monorepo filter)
* and subgroup membership checks.
*
* Inputs may originate from tree-sitter, the OS file API, or user-supplied
* MCP arguments, so both `\` and `/` separators are accepted. Internally we
* normalize to POSIX-style `/` for case-sensitive segment comparisons.
*/
function toPosix(p: string): string {
return p.replace(/\\/g, '/');
}
export function normalizeServicePrefix(service: unknown): string | undefined {
if (service === undefined || service === null) return undefined;
const s = toPosix(String(service)).trim().replace(/\/+$/, '');
return s.length > 0 ? s : undefined;
}
export function fileMatchesServicePrefix(
filePath: string | undefined,
prefix: string | undefined,
): boolean {
if (!prefix) return true;
if (!filePath) return false;
const normalized = toPosix(filePath);
return normalized === prefix || normalized.startsWith(`${prefix}/`);
}
/**
* True if `repoPath` is at or beneath `subgroup` (member-path prefix in
* `group.yaml`). Empty / missing `subgroup` matches every repo.
*
* @param exact When set, requires an exact equality match (no descendant repos).
*/
export function repoInSubgroup(repoPath: string, subgroup?: string, exact?: boolean): boolean {
if (!subgroup?.trim()) return true;
const s = toPosix(subgroup).replace(/\/+$/, '');
const r = toPosix(repoPath);
if (exact) return r === s;
return r === s || r.startsWith(`${s}/`);
}
@@ -0,0 +1,34 @@
/**
* Map MCP/CLI `@groupName` or `@groupName/memberPath` to a concrete member path in group.yaml.
*/
import { loadGroupConfig } from './config-parser.js';
import { getDefaultGitnexusDir, getGroupDir } from './storage.js';
export async function resolveAtGroupMemberRepoPath(
groupName: string,
explicitMemberPath: string | undefined,
): Promise<{ ok: true; repoPath: string } | { ok: false; error: string }> {
const trimmed = groupName.trim();
if (!trimmed) return { ok: false, error: 'Group name is empty.' };
try {
const groupDir = getGroupDir(getDefaultGitnexusDir(), trimmed);
const config = await loadGroupConfig(groupDir);
const keys = Object.keys(config.repos).sort((a, b) => a.localeCompare(b));
if (keys.length === 0) {
return { ok: false, error: `Group "${trimmed}" has no repos in group.yaml.` };
}
if (explicitMemberPath !== undefined && explicitMemberPath !== '') {
if (!(explicitMemberPath in config.repos)) {
return {
ok: false,
error: `Unknown member path "${explicitMemberPath}" in group "${trimmed}". Known paths: ${keys.join(', ')}`,
};
}
return { ok: true, repoPath: explicitMemberPath };
}
return { ok: true, repoPath: keys[0]! };
} catch (e) {
return { ok: false, error: e instanceof Error ? e.message : String(e) };
}
}
+332 -39
View File
@@ -3,10 +3,24 @@
* DB access is injected via GroupToolPort so this module stays free of LocalBackend private API.
*/
import fsp from 'node:fs/promises';
import path from 'node:path';
import { checkStaleness } from '../git-staleness.js';
import { loadGroupConfig } from './config-parser.js';
import { GroupNotFoundError, loadGroupConfig } from './config-parser.js';
import {
fileMatchesServicePrefix,
normalizeServicePrefix,
repoInSubgroup,
} from './group-path-utils.js';
import { getDefaultGitnexusDir, getGroupDir, listGroups, readContractRegistry } from './storage.js';
import { syncGroup } from './sync.js';
import type {
ContractRegistry,
CrossLink,
GroupConfig,
GroupContextResult,
StoredContract,
} from './types.js';
export interface GroupRepoHandle {
id: string;
@@ -52,12 +66,149 @@ export interface GroupToolPort {
includeTests: boolean;
},
): Promise<unknown | null>;
context(
repo: GroupRepoHandle,
params: {
name?: string;
uid?: string;
file_path?: string;
include_content?: boolean;
},
): Promise<unknown>;
}
function repoInSubgroup(repoPath: string, subgroup?: string): boolean {
if (!subgroup?.trim()) return true;
const s = subgroup.replace(/\/+$/, '');
return repoPath === s || repoPath.startsWith(`${s}/`);
function isStoredContract(raw: unknown): raw is StoredContract {
if (!raw || typeof raw !== 'object') return false;
const o = raw as Record<string, unknown>;
return (
typeof o.contractId === 'string' &&
typeof o.type === 'string' &&
typeof o.repo === 'string' &&
typeof o.role === 'string' &&
(o.role === 'provider' || o.role === 'consumer') &&
typeof o.symbolUid === 'string' &&
typeof o.symbolName === 'string' &&
typeof o.confidence === 'number' &&
o.meta !== undefined &&
typeof o.meta === 'object' &&
o.meta !== null &&
o.symbolRef !== undefined &&
typeof o.symbolRef === 'object' &&
o.symbolRef !== null &&
typeof (o.symbolRef as Record<string, unknown>).filePath === 'string' &&
typeof (o.symbolRef as Record<string, unknown>).name === 'string'
);
}
function filterQueryByServicePrefix(
queryResult: {
processes?: Array<Record<string, unknown>>;
process_symbols?: Array<Record<string, unknown>>;
},
servicePrefix: string,
): { processes: Array<Record<string, unknown>>; process_symbols: Array<Record<string, unknown>> } {
const symbols = (queryResult.process_symbols || []).filter((s) =>
fileMatchesServicePrefix(
typeof s.filePath === 'string' ? s.filePath : undefined,
servicePrefix,
),
);
const allowed = new Set(
symbols.map((s) => String((s as { process_id?: string }).process_id ?? '')).filter(Boolean),
);
const processes = (queryResult.processes || []).filter((p) => allowed.has(String(p.id)));
return { processes, process_symbols: symbols };
}
function isCrossLink(raw: unknown): raw is CrossLink {
if (!raw || typeof raw !== 'object') return false;
const o = raw as Record<string, unknown>;
const from = o.from as Record<string, unknown> | undefined;
const to = o.to as Record<string, unknown> | undefined;
if (!from || !to) return false;
if (typeof from.repo !== 'string' || typeof to.repo !== 'string') return false;
return typeof o.contractId === 'string' && typeof o.type === 'string';
}
async function loadContractRegistryResilient(
groupDir: string,
): Promise<
{ ok: true; registry: ContractRegistry; skippedCorrupt: number } | { ok: false; error: string }
> {
const filePath = path.join(groupDir, 'contracts.json');
let raw: string;
try {
raw = await fsp.readFile(filePath, 'utf-8');
} catch (e) {
if ((e as NodeJS.ErrnoException).code === 'ENOENT') {
return { ok: false, error: `No contracts.json for this group. Run group_sync first.` };
}
return { ok: false, error: e instanceof Error ? e.message : String(e) };
}
let root: unknown;
try {
root = JSON.parse(raw);
} catch {
return { ok: false, error: 'contracts.json is not valid JSON' };
}
if (!root || typeof root !== 'object' || Array.isArray(root)) {
return { ok: false, error: 'contracts.json has an invalid root object' };
}
const base = root as Record<string, unknown>;
const contractsRaw = base.contracts;
const crossRaw = base.crossLinks;
let skippedCorrupt = 0;
const contracts: StoredContract[] = [];
if (Array.isArray(contractsRaw)) {
for (const row of contractsRaw) {
try {
if (isStoredContract(row)) {
contracts.push(row);
} else {
skippedCorrupt++;
console.warn('[group] skipping corrupt contract row in contracts.json');
}
} catch {
skippedCorrupt++;
console.warn('[group] skipping corrupt contract row in contracts.json');
}
}
}
const crossLinks: CrossLink[] = [];
if (Array.isArray(crossRaw)) {
for (const row of crossRaw) {
try {
if (isCrossLink(row)) {
crossLinks.push(row);
} else {
skippedCorrupt++;
console.warn('[group] skipping corrupt crossLinks row in contracts.json');
}
} catch {
skippedCorrupt++;
console.warn('[group] skipping corrupt crossLinks row in contracts.json');
}
}
}
const registry: ContractRegistry = {
version: typeof base.version === 'number' ? base.version : 0,
generatedAt: typeof base.generatedAt === 'string' ? base.generatedAt : '',
repoSnapshots:
base.repoSnapshots && typeof base.repoSnapshots === 'object' && base.repoSnapshots !== null
? (base.repoSnapshots as Record<string, { indexedAt: string; lastCommit: string }>)
: {},
missingRepos: Array.isArray(base.missingRepos) ? (base.missingRepos as string[]) : [],
contracts,
crossLinks,
};
return { ok: true, registry, skippedCorrupt };
}
export class GroupService {
@@ -70,7 +221,14 @@ export class GroupService {
return { groups };
}
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
let config: GroupConfig;
try {
config = await loadGroupConfig(groupDir);
} catch (err) {
if (err instanceof GroupNotFoundError)
return { error: `Group "${name}" not found. Run group_list to see configured groups.` };
throw err;
}
return {
name: config.name,
description: config.description,
@@ -83,7 +241,14 @@ export class GroupService {
const name = String(params.name ?? '').trim();
if (!name) return { error: 'name is required' };
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
let config: GroupConfig;
try {
config = await loadGroupConfig(groupDir);
} catch (err) {
if (err instanceof GroupNotFoundError)
return { error: `Group "${name}" not found. Run group_list to see configured groups.` };
throw err;
}
const result = await syncGroup(config, {
groupDir,
exactOnly: Boolean(params.exactOnly),
@@ -103,10 +268,14 @@ export class GroupService {
const name = String(params.name ?? '').trim();
if (!name) return { error: 'name is required' };
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const registry = await readContractRegistry(groupDir);
if (!registry) {
return { error: `No contracts.json for group "${name}". Run group_sync first.` };
const loaded = await loadContractRegistryResilient(groupDir);
if (loaded.ok === false) {
if (loaded.error.includes('No contracts.json')) {
return { error: `No contracts.json for group "${name}". Run group_sync first.` };
}
return { error: loaded.error };
}
const { registry, skippedCorrupt } = loaded;
let contracts = registry.contracts;
if (params.type) contracts = contracts.filter((c) => c.type === params.type);
if (params.repo) contracts = contracts.filter((c) => c.repo === params.repo);
@@ -119,42 +288,162 @@ export class GroupService {
);
contracts = contracts.filter((c) => !matchedIds.has(`${c.repo}::${c.contractId}`));
}
return { contracts, crossLinks: registry.crossLinks };
const out: Record<string, unknown> = { contracts, crossLinks: registry.crossLinks };
if (skippedCorrupt > 0) out.skippedCorrupt = skippedCorrupt;
return out;
}
async groupImpact(params: Record<string, unknown>): Promise<unknown> {
const { runGroupImpact } = await import('./cross-impact.js');
return runGroupImpact({ port: this.port, gitnexusDir: getDefaultGitnexusDir() }, params);
}
async groupContext(params: Record<string, unknown>): Promise<GroupContextResult> {
const name = String(params.name ?? '').trim();
const target = typeof params.target === 'string' ? params.target.trim() : '';
const uid = typeof params.uid === 'string' ? params.uid.trim() : undefined;
const file_path = typeof params.file_path === 'string' ? params.file_path : undefined;
const include_content = Boolean(params.include_content);
if (
params.service !== undefined &&
params.service !== null &&
String(params.service).trim() === ''
) {
return { group: name || '', error: 'service must not be an empty string', results: [] };
}
const servicePrefix = normalizeServicePrefix(params.service);
const subgroup = typeof params.subgroup === 'string' ? params.subgroup : undefined;
const subgroupExact = params.subgroupExact === true;
if (!name) {
return { group: '', error: 'name is required', results: [] };
}
if (!uid && !target) {
return { group: name, error: 'target or uid is required', results: [] };
}
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
let config: GroupConfig;
try {
config = await loadGroupConfig(groupDir);
} catch (e) {
if (e instanceof GroupNotFoundError)
return {
group: name,
target: target || uid,
service: servicePrefix,
error: `Group "${name}" not found. Run group_list to see configured groups.`,
results: [],
};
return {
group: name,
target: target || uid,
service: servicePrefix,
error: e instanceof Error ? e.message : String(e),
results: [],
};
}
const memberEntries = Object.entries(config.repos).filter(([repoPath]) =>
repoInSubgroup(repoPath, subgroup, subgroupExact),
);
const results: GroupContextResult['results'] = await Promise.all(
memberEntries.map(async ([repoPath, registryName]) => {
try {
const repoObj = await this.port.resolveRepo(registryName);
const payload = await this.port.context(repoObj, {
name: target || undefined,
uid,
file_path,
include_content,
});
if (servicePrefix) {
const st = (payload as { status?: string })?.status;
const sym = (payload as { symbol?: { filePath?: string } })?.symbol;
if (st === 'found' && !fileMatchesServicePrefix(sym?.filePath, servicePrefix)) {
return { repoPath, registryName, payload: {} };
}
}
return { repoPath, registryName, payload };
} catch (e) {
return {
repoPath,
registryName,
payload: { error: e instanceof Error ? e.message : String(e) },
};
}
}),
);
return {
group: name,
target: target || uid,
service: servicePrefix,
results,
};
}
async groupQuery(params: Record<string, unknown>): Promise<unknown> {
const name = String(params.name ?? '').trim();
const queryText = String(params.query ?? '').trim();
if (!name || !queryText) return { error: 'name and query are required' };
if (
params.service !== undefined &&
params.service !== null &&
String(params.service).trim() === ''
) {
return { error: 'service must not be an empty string' };
}
const servicePrefix = normalizeServicePrefix(params.service);
const limit = typeof params.limit === 'number' && params.limit > 0 ? params.limit : 5;
const subgroup = typeof params.subgroup === 'string' ? params.subgroup : undefined;
const subgroupExact = params.subgroupExact === true;
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
const perRepo: Array<{ repo: string; score: number; processes: unknown[] }> = [];
for (const [repoPath, registryName] of Object.entries(config.repos)) {
if (!repoInSubgroup(repoPath, subgroup)) continue;
try {
const repoObj = await this.port.resolveRepo(registryName);
const queryResult = (await this.port.query(repoObj, {
query: queryText,
limit,
max_symbols: 10,
include_content: false,
})) as { processes?: Array<Record<string, unknown>> };
const processes = queryResult.processes || [];
const scored = processes.map((p, idx) => ({
...p,
_rrf_score: 1 / (idx + 1 + 60),
_repo: repoPath,
}));
perRepo.push({ repo: repoPath, score: 0, processes: scored });
} catch {
perRepo.push({ repo: repoPath, score: 0, processes: [] });
}
let config: GroupConfig;
try {
config = await loadGroupConfig(groupDir);
} catch (err) {
if (err instanceof GroupNotFoundError)
return { error: `Group "${name}" not found. Run group_list to see configured groups.` };
throw err;
}
const memberEntries = Object.entries(config.repos).filter(([repoPath]) =>
repoInSubgroup(repoPath, subgroup, subgroupExact),
);
const perRepo = await Promise.all(
memberEntries.map(async ([repoPath, registryName]) => {
try {
const repoObj = await this.port.resolveRepo(registryName);
const queryResult = (await this.port.query(repoObj, {
query: queryText,
limit,
max_symbols: 10,
include_content: false,
})) as {
processes?: Array<Record<string, unknown>>;
process_symbols?: Array<Record<string, unknown>>;
};
const processes = servicePrefix
? filterQueryByServicePrefix(queryResult, servicePrefix).processes
: queryResult.processes || [];
const scored = processes.map((p, idx) => ({
...p,
_rrf_score: 1 / (idx + 1 + 60),
_repo: repoPath,
}));
return { repo: repoPath, score: 0, processes: scored as unknown[] };
} catch {
return { repo: repoPath, score: 0, processes: [] as unknown[] };
}
}),
);
const allProcesses = perRepo.flatMap((r) => r.processes as Array<Record<string, unknown>>);
allProcesses.sort((a, b) => (b._rrf_score as number) - (a._rrf_score as number));
const topN = allProcesses.slice(0, limit);
@@ -171,7 +460,14 @@ export class GroupService {
const name = String(params.name ?? '').trim();
if (!name) return { error: 'name is required' };
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
let config: GroupConfig;
try {
config = await loadGroupConfig(groupDir);
} catch (err) {
if (err instanceof GroupNotFoundError)
return { error: `Group "${name}" not found. Run group_list to see configured groups.` };
throw err;
}
const registry = await readContractRegistry(groupDir);
const repoStatuses: Record<
@@ -184,13 +480,10 @@ export class GroupService {
}
> = {};
const fsp = await import('node:fs/promises');
const pathMod = await import('node:path');
for (const [repoPath, registryName] of Object.entries(config.repos)) {
try {
const repoObj = await this.port.resolveRepo(registryName);
const metaPath = pathMod.join(repoObj.storagePath, 'meta.json');
const metaPath = path.join(repoObj.storagePath, 'meta.json');
const metaRaw = await fsp.readFile(metaPath, 'utf-8').catch(() => '{}');
const meta = JSON.parse(metaRaw) as { lastCommit?: string; indexedAt?: string };
+33
View File
@@ -96,6 +96,9 @@ export interface RepoHandle {
storagePath: string;
}
/** Why local impact or fan-out stopped early (e.g. wall-clock budget exhausted). */
export type GroupImpactTruncationReason = 'timeout' | 'partial';
export interface GroupImpactResult {
local: unknown;
group: string;
@@ -110,6 +113,36 @@ export interface GroupImpactResult {
cross_repo_hits: number;
};
risk: string;
/**
* Milliseconds budget applied to the **Phase 1 local impact** leg (`safeLocalImpact`).
* If the walk hits this wall first, expect `truncationReason: 'timeout'` and a partial `local` payload.
*/
timeoutMs?: number;
/** Present when local impact or fan-out stopped early (timeout, graph cap, etc.). */
truncationReason?: GroupImpactTruncationReason;
/**
* Human-readable note when `crossDepth` was clamped (e.g. multi-hop not implemented yet).
*/
crossDepthWarning?: string;
}
/** One repo’s `context` tool payload in a group-scoped context run. */
export interface GroupContextRepoEntry {
repoPath: string;
registryName: string;
payload: unknown;
}
/**
* Aggregated group `context`: explicit per-repo rows (no merged symbol payloads).
* Use top-level `error` only for unrecoverable failures, not for “no matches” or service scope misses.
*/
export interface GroupContextResult {
group: string;
target?: string;
service?: string;
error?: string;
results: GroupContextRepoEntry[];
}
export interface CrossRepoImpact {
+31 -3
View File
@@ -1,8 +1,24 @@
import { LRUCache } from 'lru-cache';
import Parser from 'tree-sitter';
/**
* Minimal structural shape consumers need when reading Trees back
* through a phase-dependency boundary. Declared here so phases that
* receive ASTCache via `getPhaseOutput<...>` don't hand-roll their
* own inline structural types that silently drift when ASTCache's
* contract changes.
*
* Typed as `unknown` at the Tree boundary because consumers on the
* other side of the phase-output map don't share tree-sitter's type
* graph (e.g. COBOL's standalone processor).
*/
export interface ASTCacheReader {
get(filePath: string): unknown;
clear(): void;
}
// Define the interface for the Cache
export interface ASTCache {
export interface ASTCache extends ASTCacheReader {
get: (filePath: string) => Parser.Tree | undefined;
set: (filePath: string, tree: Parser.Tree) => void;
clear: () => void;
@@ -17,8 +33,20 @@ export const createASTCache = (maxSize: number = 50): ASTCache => {
max: effectiveMax,
dispose: (tree) => {
try {
// NOTE: web-tree-sitter has tree.delete(); native tree-sitter trees are GC-managed.
// Keep this try/catch so we don't crash on either runtime.
// NOTE: web-tree-sitter has tree.delete(); native tree-sitter
// trees are GC-managed and .delete is absent (no-op here).
//
// Single-owner invariant (load-bearing under WASM): a given
// Parser.Tree reference must live in AT MOST ONE ASTCache
// that disposes. The parse-phase chunk-local cache clears
// between chunks; the cross-phase `scopeTreeCache` (also an
// ASTCache today) holds the same Tree by reference. Under
// native tree-sitter this is benign (dispose is a no-op).
// If/when GitNexus adopts web-tree-sitter for sequential
// parsing, the cross-phase cache must either (a) skip
// writing Trees that are already owned by a disposing cache,
// or (b) use tree.copy() per entry. Failing to pick one
// will hand freed memory to scope-resolution.
(tree as unknown as { delete?: () => void }).delete?.();
} catch (e) {
console.warn('Failed to delete tree from WASM memory', e);
+209 -50
View File
@@ -1,12 +1,34 @@
import { KnowledgeGraph } from '../graph/types.js';
import { ASTCache } from './ast-cache.js';
import type {
SymbolDefinition,
SymbolTableReader,
HeritageMap,
ExtractedHeritage,
} from './model/index.js';
import type { SymbolDefinition } from 'gitnexus-shared';
import type { SymbolTableReader, HeritageMap, ExtractedHeritage } from './model/index.js';
import { CLASS_TYPES, CALL_TARGET_TYPES, lookupMethodByOwnerWithMRO } from './model/index.js';
import type { DispatchDecision, ReceiverEnriched } from './call-types.js';
/** Shorthand for the receiver-source discriminant shared across the DAG. */
type ReceiverSource = ReceiverEnriched['receiverSource'];
/**
* DAG stage 4 fallback: used when `selectDispatch` is absent or returns null.
* Preserves pre-DAG dispatch semantics:
* - 'constructor' → constructor branch
* - 'free' → free branch (admits Swift/Kotlin class-target fast path)
* - 'member' or undefined → owner-scoped branch
*
* `undefined` callForm MUST route through owner-scoped (not free) so bare
* identifiers without a classified shape do NOT trigger `resolveFreeCall`'s
* class-target fast path. Without a `receiverTypeName`, the owner-scoped
* branch falls through to `resolveModuleAliasedCall` + `singleCandidate`,
* matching legacy behavior where non-callable symbols (Class, Interface)
* null-route instead of producing spurious Constructor edges.
*/
const defaultDispatchDecision = (
callForm: 'free' | 'member' | 'constructor' | undefined,
): DispatchDecision => {
if (callForm === 'constructor') return { primary: 'constructor' };
if (callForm === 'free') return { primary: 'free' };
return { primary: 'owner-scoped' };
};
import Parser from 'tree-sitter';
import type { ResolutionContext } from './model/resolution-context.js';
import { TIER_CONFIDENCE, type ResolutionTier } from './model/resolution-context.js';
@@ -15,6 +37,7 @@ import { isLanguageAvailable, loadParser, loadLanguage } from '../tree-sitter/pa
import { getProvider } from './languages/index.js';
import { generateId } from '../../lib/utils.js';
import { getLanguageFromFilename, SupportedLanguages } from 'gitnexus-shared';
import { isRegistryPrimary } from './registry-primary-flag.js';
import { isVerboseIngestionEnabled } from './utils/verbose.js';
import { yieldToEventLoop } from './utils/event-loop.js';
import {
@@ -728,6 +751,8 @@ export const processCalls = async (
const language = getLanguageFromFilename(file.path);
if (!language) continue;
// Registry-primary gate: scope-based phase owns CALLS for this lang.
if (isRegistryPrimary(language)) continue;
if (!isLanguageAvailable(language)) {
if (skippedByLang) {
skippedByLang.set(language, (skippedByLang.get(language) ?? 0) + 1);
@@ -766,22 +791,26 @@ export const processCalls = async (
// Extract heritage from query matches to build parentMap for buildTypeEnv.
// Heritage-processor runs in PARALLEL, so graph edges don't exist when buildTypeEnv runs.
const fileParentMap = new Map<string, string[]>();
for (const match of matches) {
const captureMap: Record<string, any> = {};
match.captures.forEach((c) => (captureMap[c.name] = c.node));
if (captureMap['heritage.class'] && captureMap['heritage.extends']) {
const className: string = captureMap['heritage.class'].text;
const parentName: string = captureMap['heritage.extends'].text;
const extendsNode = captureMap['heritage.extends'];
const fieldDecl = extendsNode.parent;
if (fieldDecl?.type === 'field_declaration' && fieldDecl.childForFieldName('name'))
continue;
let parents = fileParentMap.get(className);
if (!parents) {
parents = [];
fileParentMap.set(className, parents);
if (provider.heritageExtractor) {
for (const match of matches) {
const captureMap: Record<string, any> = {};
match.captures.forEach((c) => (captureMap[c.name] = c.node));
if (captureMap['heritage.class']) {
const heritageItems = provider.heritageExtractor.extract(captureMap, {
filePath: file.path,
language,
});
for (const item of heritageItems) {
if (item.kind === 'extends') {
let parents = fileParentMap.get(item.className);
if (!parents) {
parents = [];
fileParentMap.set(item.className, parents);
}
if (!parents.includes(item.parentName)) parents.push(item.parentName);
}
}
}
if (!parents.includes(parentName)) parents.push(parentName);
}
}
const parentMap: ReadonlyMap<string, readonly string[]> = fileParentMap;
@@ -991,6 +1020,28 @@ export const processCalls = async (
const calledName = nameNode.text;
// Check heritage extractor for call-based heritage (e.g., Ruby include/extend/prepend)
if (provider.heritageExtractor?.extractFromCall) {
const heritageItems = provider.heritageExtractor.extractFromCall(
calledName,
captureMap['call'],
{ filePath: file.path, language },
);
if (heritageItems !== null) {
for (const item of heritageItems) {
collectedHeritage.push({
filePath: file.path,
className: item.className,
parentName: item.parentName,
kind: item.kind,
});
}
return;
}
}
// Dispatch: route language-specific calls (properties, imports)
// Heritage routing is handled by heritageExtractor.extractFromCall above.
const routed = callRouter?.(calledName, captureMap['call']);
if (routed) {
switch (routed.kind) {
@@ -998,17 +1049,6 @@ export const processCalls = async (
case 'import':
return;
case 'heritage':
for (const item of routed.items) {
collectedHeritage.push({
filePath: file.path,
className: item.enclosingClass,
parentName: item.mixinName,
kind: item.heritageKind,
});
}
return;
case 'properties': {
const fileId = generateId('File', file.path);
const propEnclosingClassId = findEnclosingClassId(captureMap['call'], file.path);
@@ -1061,10 +1101,17 @@ export const processCalls = async (
if (provider.isBuiltInName(calledName)) return;
const callForm = inferCallForm(callNode, nameNode);
const receiverName = callForm === 'member' ? extractReceiverName(nameNode) : undefined;
// --- DAG stage 2-3: classify-form + infer-receiver (shared defaults) ---
// These stages run the shared inference chain. Language providers can
// customize infer-receiver (stage 3) via the inferImplicitReceiver hook
// which runs AFTER this default chain (typed-binding → constructor-map →
// module-alias → class-as-receiver → mixed-chain), and selectDispatch
// (stage 4) which picks the resolver branch.
let callForm = inferCallForm(callNode, nameNode);
let receiverName = callForm === 'member' ? extractReceiverName(nameNode) : undefined;
let receiverTypeName =
receiverName && typeEnv ? typeEnv.lookup(receiverName, callNode) : undefined;
let receiverSource: ReceiverSource = receiverTypeName ? 'typed-binding' : 'none';
// Phase P: virtual dispatch override — when the declared type is a base class but
// the constructor created a known subclass, prefer the more specific type.
// Checks per-file parentMap first, then falls back to globalParentMap for
@@ -1103,6 +1150,7 @@ export const processCalls = async (
ctx.model.types.lookupClassByName(receiverTypeName).length > 0)
) {
receiverTypeName = ctorType;
receiverSource = 'constructor-map';
}
}
}
@@ -1111,10 +1159,14 @@ export const processCalls = async (
const enclosingFunc = findEnclosingFunction(callNode, file.path, ctx, provider);
const funcName = enclosingFunc ? extractFuncNameFromSourceId(enclosingFunc) : '';
receiverTypeName = lookupReceiverType(receiverIndex, funcName, receiverName);
if (receiverTypeName) receiverSource = 'constructor-map';
}
// Fall back to class-as-receiver for static method calls (e.g. UserService.find_user()).
// When the receiver name is not a variable in TypeEnv but resolves to a Class/Struct/Interface
// through the standard tiered resolution, use it directly as the receiver type.
// Fall back to class-as-receiver for static method calls (e.g. UserService.find_user(),
// Greetable.format()). When the receiver name is not a variable in TypeEnv but
// resolves to a class-like symbol (Class / Interface / Struct / Enum / Trait) via
// tiered resolution, use it directly as the receiver type. `Trait` is included so
// Ruby module class-method calls flow through the class-as-receiver path and reach
// the `selectDispatch` hook's singleton branch.
if (!receiverTypeName && receiverName && callForm === 'member') {
const typeResolved = ctx.resolve(receiverName, file.path);
if (
@@ -1124,10 +1176,12 @@ export const processCalls = async (
d.type === 'Class' ||
d.type === 'Interface' ||
d.type === 'Struct' ||
d.type === 'Enum',
d.type === 'Enum' ||
d.type === 'Trait',
)
) {
receiverTypeName = receiverName;
receiverSource = 'class-as-receiver';
}
}
// Hoist sourceId so it's available for ACCESSES edge emission during chain walk.
@@ -1173,11 +1227,51 @@ export const processCalls = async (
makeAccessEmitter(graph, sourceId),
heritageMap,
);
if (receiverTypeName) receiverSource = 'mixed-chain';
}
}
}
}
// --- DAG stage 3: infer-receiver (provider hook) ---
// Synthesize implicit receivers for languages that omit them (e.g., Ruby bare-call).
// This hook runs AFTER the shared inference chain so explicit receivers /
// typed bindings always take precedence. Output (if non-null) overlays onto
// the ReceiverEnriched for the next stage.
let dispatchHint: string | undefined;
if (provider.inferImplicitReceiver) {
const override = provider.inferImplicitReceiver({
calledName,
callForm,
receiverName,
receiverTypeName,
callNode,
filePath: file.path,
});
if (override) {
callForm = override.callForm;
receiverName = override.receiverName;
receiverTypeName = override.receiverTypeName;
receiverSource = override.receiverSource;
dispatchHint = override.hint;
}
}
// --- DAG stage 4: select-dispatch (provider hook + default fallback) ---
// Decide which resolver path to try first (primary) and fallback strategy.
// Language providers can customize dispatch via selectDispatch hook; all
// others use the shared defaultDispatchDecision. Always non-null after this
// block so downstream resolvers are table-driven.
const dispatchDecision: DispatchDecision =
provider.selectDispatch?.({
calledName,
callForm,
receiverName,
receiverTypeName,
receiverSource,
hint: dispatchHint,
}) ?? defaultDispatchDecision(callForm);
// Build overload hints for languages with inferLiteralType (Java/Kotlin/C#/C++).
// Only used when multiple candidates survive arity filtering — ~1-3% of calls.
const langConfig = provider.typeConfig;
@@ -1199,6 +1293,7 @@ export const processCalls = async (
widenCache,
undefined,
heritageMap,
dispatchDecision,
);
if (!resolved) return;
@@ -1737,11 +1832,20 @@ const resolveCallTarget = (
widenCache?: WidenCache,
preComputedArgTypes?: (string | undefined)[],
heritageMap?: HeritageMap,
dispatchDecision?: DispatchDecision,
): ResolveResult | null => {
const tiered = ctx.resolve(call.calledName, currentFile);
if (!tiered) return null;
if (call.callForm === 'free') {
// DAG dispatch: use decision.primary to pick the resolver branch.
// Callers that own the DAG (processCalls + crossFile deferred paths)
// pass a decision; other callers use the shared default ladder.
// Language-specific primary / fallback / ancestryView overrides come from
// the provider's `selectDispatch` hook.
const decision = dispatchDecision ?? defaultDispatchDecision(call.callForm);
const primary = decision.primary;
if (primary === 'free') {
return resolveFreeCall(
call.calledName,
currentFile,
@@ -1752,7 +1856,7 @@ const resolveCallTarget = (
preComputedArgTypes,
);
}
if (call.callForm === 'constructor') {
if (primary === 'constructor') {
return (
resolveStaticCall(
call.calledName,
@@ -1765,6 +1869,7 @@ const resolveCallTarget = (
) ?? singleCandidate(tiered, call.argCount, 'constructor')
);
}
// primary === 'owner-scoped'
if (call.receiverTypeName) {
// Skip the owner-scoped MRO path when the tiered pool has genuine
// overload ambiguity that needs D1-D4+E handling, not D0.
@@ -1772,6 +1877,15 @@ const resolveCallTarget = (
(!!overloadHints || !!preComputedArgTypes) &&
countCallableCandidates(tiered.candidates, call.argCount, call.callForm) > 1;
// Try owner-scoped (resolveMemberCall) then file-scoped (resolveMemberCallByFile).
// DAG: dispatchDecision.ancestryView selects instance vs singleton ancestry
// for kind-aware MRO strategies. Ruby `Account.log` flows via 'singleton'.
//
// Singleton-ancestry miss MUST NOT degrade to the file-scoped fallback:
// resolveMemberCallByFile matches by ownerId and would happily pick an
// instance method defined on the same class, leaking instance dispatch
// onto what was declared a class-method call. For singleton dispatch,
// a miss either null-routes or falls through to `decision.fallback`.
const singletonDispatch = decision.ancestryView === 'singleton';
const memberResult =
(!skipMember
? resolveMemberCall(
@@ -1781,18 +1895,21 @@ const resolveCallTarget = (
ctx,
heritageMap,
call.argCount,
decision.ancestryView,
)
: null) ??
resolveMemberCallByFile(
call.calledName,
call.receiverTypeName,
currentFile,
ctx,
call.argCount,
call.callForm,
overloadHints,
preComputedArgTypes,
);
(singletonDispatch
? null
: resolveMemberCallByFile(
call.calledName,
call.receiverTypeName,
currentFile,
ctx,
call.argCount,
call.callForm,
overloadHints,
preComputedArgTypes,
));
if (memberResult) return memberResult;
// Module-alias narrowing runs as a FALLBACK, after owner/file-scoped
@@ -1828,7 +1945,26 @@ const resolveCallTarget = (
// hierarchy. When the type is NOT in the index (PHP `mixed`, dynamic
// types, unresolvable aliases), the scoped resolvers had nothing to
// work with and singleCandidate is the correct last resort.
//
// DAG fallback override: when `select-dispatch` returned
// `fallback: 'free-arity-narrowed'` (today: Ruby implicit-self bare
// calls whose enclosing class doesn't define the method), fall through
// to free-call resolution instead of null-routing. This preserves
// existing free-call arity-narrowing heuristics for bare calls that
// happen to target methods on unrelated classes.
if (typeResolves && typeResolves.candidates.length > 0) {
if (decision.fallback === 'free-arity-narrowed') {
const free = resolveFreeCall(
call.calledName,
currentFile,
ctx,
call.argCount,
tiered,
overloadHints,
preComputedArgTypes,
);
if (free) return free;
}
return null; // null-route: type resolved, no candidate matched
}
return singleCandidate(tiered, call.argCount, call.callForm);
@@ -2024,6 +2160,13 @@ const resolveMethodByOwner = (
ctx: ResolutionContext,
heritageMap?: HeritageMap,
argCount?: number,
/**
* DAG-sourced ancestry selector. `'singleton'` routes through
* `heritageMap.getSingletonAncestry(owner)` for class-method dispatch
* (Ruby `Account.log` via `extend LoggerMixin`). Default / undefined
* uses the walker's instance-dispatch behavior.
*/
ancestryView?: 'instance' | 'singleton',
): { def: SymbolDefinition; tier: ResolutionTier } | undefined => {
const typeResolved = ctx.resolve(receiverTypeName, filePath);
if (!typeResolved) return undefined;
@@ -2052,6 +2195,14 @@ const resolveMethodByOwner = (
let ambiguous = false;
for (const candidate of typeResolved.candidates) {
if (!CLASS_LIKE_TYPES.has(candidate.type)) continue;
// Singleton dispatch: when the DAG decision requested the singleton
// ancestry view, pass `heritageMap.getSingletonAncestry` as the walker's
// ancestry override. Kind-aware strategies (e.g. MroStrategy 'ruby-mixin')
// honor the override by scanning it linearly in place of their default walk.
const singletonOverride =
ancestryView === 'singleton' && canWalkMRO && heritageMap
? heritageMap.getSingletonAncestry(candidate.nodeId).map((e) => e.parentId)
: undefined;
const def = canWalkMRO
? lookupMethodByOwnerWithMRO(
candidate.nodeId,
@@ -2060,6 +2211,7 @@ const resolveMethodByOwner = (
ctx.model,
mroStrategy,
argCount,
singletonOverride,
)
: ctx.model.methods.lookupMethodByOwner(candidate.nodeId, methodName, argCount);
if (!def) continue;
@@ -2114,6 +2266,7 @@ export const resolveMemberCall = (
ctx: ResolutionContext,
heritageMap?: HeritageMap,
argCount?: number,
ancestryView?: 'instance' | 'singleton',
): ResolveResult | null => {
const resolved = resolveMethodByOwner(
ownerType,
@@ -2122,6 +2275,7 @@ export const resolveMemberCall = (
ctx,
heritageMap,
argCount,
ancestryView,
);
if (!resolved) return null;
return toResolveResult(resolved.def, resolved.tier);
@@ -2580,6 +2734,11 @@ export const processCallsFromExtracted = async (
await yieldToEventLoop();
}
// Registry-primary gate: skip Python (etc.) entirely when the
// scope-based phase owns CALLS for this language.
const fileLanguage = getLanguageFromFilename(filePath);
if (fileLanguage && isRegistryPrimary(fileLanguage)) continue;
ctx.enableCache(filePath);
const widenCache: WidenCache = new Map();
const receiverMap = fileReceiverTypes.get(filePath);
+13 -42
View File
@@ -1,10 +1,14 @@
/**
* Shared Ruby call routing logic.
*
* Ruby expresses imports, heritage (mixins), and property definitions as
* method calls rather than syntax-level constructs. This module provides a
* routing function used by the CLI call-processor, CLI parse-worker, and
* the web call-processor so that the classification logic lives in one place.
* Ruby expresses imports and property definitions as method calls rather
* than syntax-level constructs. This module provides a routing function
* used by the CLI call-processor, CLI parse-worker, and the web
* call-processor so that the classification logic lives in one place.
*
* Heritage (mixins: include/extend/prepend) was previously routed here
* but is now handled by heritageExtractor.extractFromCall before the
* call router runs. The router still returns 'skip' for these calls.
*
* NOTE: This file is intentionally duplicated in gitnexus-web/ because the
* two packages have separate build targets (Node native vs WASM/browser).
@@ -30,17 +34,10 @@ export type CallRouter = (calledName: string, callNode: SyntaxNode) => CallRouti
export type RubyCallRouting =
| { kind: 'import'; importPath: string; isRelative: boolean }
| { kind: 'heritage'; items: RubyHeritageItem[] }
| { kind: 'properties'; items: RubyPropertyItem[] }
| { kind: 'call' }
| { kind: 'skip' };
export interface RubyHeritageItem {
enclosingClass: string;
mixinName: string;
heritageKind: 'include' | 'extend' | 'prepend';
}
export type RubyAccessorType = 'attr_accessor' | 'attr_reader' | 'attr_writer';
export interface RubyPropertyItem {
@@ -56,9 +53,6 @@ export interface RubyPropertyItem {
const CALL_RESULT: RubyCallRouting = { kind: 'call' };
const SKIP_RESULT: RubyCallRouting = { kind: 'skip' };
/** Max depth for parent-walking loops to prevent pathological AST traversals */
const MAX_PARENT_DEPTH = 50;
// ── Routing function ────────────────────────────────────────────────────────
/**
@@ -88,35 +82,12 @@ export function routeRubyCall(calledName: string, callNode: SyntaxNode): RubyCal
return { kind: 'import', importPath, isRelative };
}
// ── include / extend / prepend → heritage (mixin) ──────────────────────
// ── include / extend / prepend — heritage (now handled by heritageExtractor) ─
// Call-based heritage is intercepted by heritageExtractor.extractFromCall
// before the call router runs. Return SKIP_RESULT so these calls don't
// fall through to normal call processing.
if (calledName === 'include' || calledName === 'extend' || calledName === 'prepend') {
let enclosingClass: string | null = null;
let current = callNode.parent;
let depth = 0;
while (current && ++depth <= MAX_PARENT_DEPTH) {
if (current.type === 'class' || current.type === 'module') {
const nameNode = current.childForFieldName?.('name');
if (nameNode) {
enclosingClass = nameNode.text;
break;
}
}
current = current.parent;
}
if (!enclosingClass) return SKIP_RESULT;
const items: RubyHeritageItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of argList?.children ?? []) {
if (arg.type === 'constant' || arg.type === 'scope_resolution') {
items.push({
enclosingClass,
mixinName: arg.text,
heritageKind: calledName as 'include' | 'extend' | 'prepend',
});
}
}
return items.length > 0 ? { kind: 'heritage', items } : SKIP_RESULT;
return SKIP_RESULT;
}
// ── attr_accessor / attr_reader / attr_writer → property definitions ───
+97
View File
@@ -78,3 +78,100 @@ export interface CallExtractionConfig {
*/
typeAsReceiverHeuristic?: boolean;
}
// ---------------------------------------------------------------------------
// Call-resolution DAG types
// ---------------------------------------------------------------------------
//
// The call-resolution pipeline is a typed DAG:
//
// extract-call ──▶ classify-form ──▶ infer-receiver ──▶ select-dispatch ──▶ resolve-target ──▶ emit-edge
//
// Provider hooks plug in at infer-receiver and select-dispatch; shared stages
// stay language-agnostic. Stages 1-2 run in the parse worker; stages 3-6 run
// on the main thread. DAG-internal types below are main-thread-only and never
// serialize to the graph.
/**
* DAG stage 3 output: call record with receiver type and source discriminant.
*
* `receiverTypeName` is resolved via TypeEnv → constructor-map → class-as-receiver →
* mixed-chain, or synthesized by `inferImplicitReceiver`. `receiverSource` tags
* which path won and drives MRO strategy selection in stage 4.
*
* Invariants:
* - `receiverSource` MUST match how `receiverTypeName` was resolved; every
* discriminant must have a live reader and writer.
* - `hint` is opaque to shared stages; only the same provider's `selectDispatch` reads it.
*
* @see language-provider.ts § inferImplicitReceiver, selectDispatch
*/
export interface ReceiverEnriched {
readonly calledName: string;
readonly callForm: 'free' | 'member' | 'constructor' | undefined;
readonly receiverName: string | undefined;
readonly receiverTypeName: string | undefined;
readonly receiverSource:
| 'none'
| 'typed-binding'
| 'constructor-map'
| 'class-as-receiver'
| 'mixed-chain'
| 'implicit-self';
/** Free-form hint from the provider hook; opaque to shared stages. */
readonly hint?: string;
}
/**
* Provider hook output for `LanguageProvider.inferImplicitReceiver` (DAG stage 3).
*
* Overlay applied to `ReceiverEnriched` when an implicit receiver is synthesized.
* Ruby example: bare `serialize` inside `Account#call_serialize` →
* `{ callForm: 'member', receiverName: 'self', receiverTypeName: 'Account',
* receiverSource: 'implicit-self', hint: 'instance' }`
*
* Invariants:
* - `receiverSource` is always `'implicit-self'` — the only variant this type produces.
* - `callForm` is always `'member'` — the rewrite converts bare-call to method invocation.
* - `hint` is opaque to shared stages; consumed by the same language's `selectDispatch`.
*/
export interface ImplicitReceiverOverride {
readonly callForm: 'free' | 'member' | 'constructor';
readonly receiverName: string;
readonly receiverTypeName: string;
readonly receiverSource: Extract<ReceiverEnriched['receiverSource'], 'implicit-self'>;
/** Free-form language tag (e.g. Ruby sets 'singleton' for `def self.foo`
* method bodies). Consumed by the same language's `selectDispatch` hook. */
readonly hint?: string;
}
/**
* DAG stage 4 output: dispatch strategy for resolving the target method.
*
* Encodes which resolver branch to try first and an optional fallback.
* Stage 5 delegates to `resolveMemberCall`, `resolveFreeCall`, or
* `resolveStaticCall` based on `primary`.
*
* - `primary`: `'owner-scoped'` = MRO walk, `'free'` = arity-tiered global lookup,
* `'constructor'` = type instantiation.
* - `fallback`: Only `'free-arity-narrowed'` exists; used by Ruby implicit-self
* to degrade to arity-tiered free lookup when the MRO walk misses.
* - `ancestryView`: Ruby `'ruby-mixin'` only. `'singleton'` walks extend providers
* only; a miss NEVER falls through to file-scoped lookup (enforced in
* resolveCallTarget). `'instance'` is the default.
*
* Common patterns:
* - `{primary: 'constructor'}` — constructor call
* - `{primary: 'owner-scoped'}` — member call with known type
* - `{primary: 'owner-scoped', fallback: 'free-arity-narrowed', ancestryView: 'instance'}` — Ruby implicit-self
* - `{primary: 'owner-scoped', ancestryView: 'singleton'}` — Ruby class-method call
*
* @see language-provider.ts § selectDispatch
* @see call-processor.ts § defaultDispatchDecision, resolveCallTarget
*/
export interface DispatchDecision {
readonly primary: 'owner-scoped' | 'free' | 'constructor';
readonly fallback?: 'free-arity-narrowed';
readonly ancestryView?: 'instance' | 'singleton';
readonly hint?: string;
}
@@ -0,0 +1,299 @@
/**
* Phase 5 of the RFC #909 ingestion lifecycle: drain `ReferenceIndex`
* into the knowledge graph as labeled edges with `confidence` and
* `evidence` properties (Ring 2 PKG #925).
*
* The resolution phase (future PR) writes `Reference` records into
* `model.scopes.referenceSites`-derived `ReferenceIndex`; this module
* materializes those records as `GraphRelationship`s via
* `graph.addRelationship`. Every emitted edge carries:
*
* - `type`: one of `'CALLS' | 'ACCESSES' | 'INHERITS' | 'USES'`
* (mapped from `Reference.kind` — `'read'` and `'write'` both route
* to `ACCESSES`; `'type-reference'` and `'import-use'` route to
* `USES`; `'call'` stays `CALLS`; `'inherits'` stays `INHERITS`).
* - `confidence`: the pre-computed confidence from the Reference record.
* - `reason`: human-readable summary (`"scope-resolution: call | confidence 0.75"`).
* - `evidence`: the full `ResolutionEvidence[]` trace — additive graph
* property (see `GraphRelationship.evidence` in gitnexus-shared),
* so queries that don't know about it are unaffected.
* - `step`: carries the reference's access-kind discriminant when
* available (`1` for read, `2` for write) so `ACCESSES` edges retain
* the read/write distinction without forcing a new edge type.
*
* ## Optional scope-tree flush
*
* When `INGESTION_EMIT_SCOPES=1` is set, this module also emits:
*
* - `Scope` nodes for every `Scope` in the tree
* - `CONTAINS` edges from parent scope to child scope
* - `DEFINES` edges from scope to its `ownedDefs` members
* - `IMPORTS` edges from scope to `targetModuleScope` of each finalized
* `ImportEdge` that carries one
*
* Off by default — existing queries that don't know about `Scope` nodes
* continue to work, and the storage cost is opt-in.
*
* ## Source-of-truth: the caller def for a reference
*
* A `Reference` says "some code inside `fromScope` references `toDef`".
* The graph wants `(callerNodeId, calleeNodeId)`. We resolve the caller
* by walking up the scope tree from `fromScope` until we find a scope
* whose `ownedDefs` contains a Function-like def. If no such ancestor
* exists, the edge is attributed to the first def owned by the innermost
* ancestor scope, and if THAT produces nothing either the edge is
* skipped (with a count returned in `EmitStats.skippedNoCaller`).
*/
import type {
NodeLabel,
RelationshipType,
Reference,
ReferenceIndex,
ResolutionEvidence,
Scope,
ScopeId,
SymbolDefinition,
} from 'gitnexus-shared';
import type { KnowledgeGraph } from '../graph/types.js';
import type { ScopeResolutionIndexes } from './model/scope-resolution-indexes.js';
// ─── Public API ─────────────────────────────────────────────────────────────
export interface EmitStats {
readonly edgesEmitted: number;
/** References dropped because no caller def could be resolved. */
readonly skippedNoCaller: number;
/** References dropped because `toDef` was not found in the DefIndex. */
readonly skippedMissingTarget: number;
/** Scope nodes emitted — `0` unless `INGESTION_EMIT_SCOPES=1`. */
readonly scopeNodesEmitted: number;
/** Scope-tree structural edges emitted — `0` unless `INGESTION_EMIT_SCOPES=1`. */
readonly scopeEdgesEmitted: number;
}
export interface EmitReferencesInput {
readonly graph: KnowledgeGraph;
readonly scopes: ScopeResolutionIndexes;
readonly referenceIndex: ReferenceIndex;
/** Human-consumable label for the `reason` prefix. Defaults to `'scope-resolution'`. */
readonly sourceLabel?: string;
}
/**
* Drain `referenceIndex.bySourceScope` into graph edges.
*
* The scope-tree flush is controlled separately by
* `INGESTION_EMIT_SCOPES` — callers can run `emitReferencesToGraph`
* without scope-node emission or layer the two calls as needed.
*/
export function emitReferencesToGraph(input: EmitReferencesInput): EmitStats {
const { graph, scopes, referenceIndex } = input;
const sourceLabel = input.sourceLabel ?? 'scope-resolution';
let edgesEmitted = 0;
let skippedNoCaller = 0;
let skippedMissingTarget = 0;
for (const [fromScope, refs] of referenceIndex.bySourceScope) {
for (const ref of refs) {
const targetDef = scopes.defs.get(ref.toDef);
if (targetDef === undefined) {
skippedMissingTarget++;
continue;
}
const callerId = resolveCallerNodeId(fromScope, scopes);
if (callerId === undefined) {
skippedNoCaller++;
continue;
}
graph.addRelationship(buildRelationship(ref, callerId, targetDef, sourceLabel));
edgesEmitted++;
}
}
const scopeStats = isScopeEmissionEnabled()
? emitScopeGraph({ graph, scopes })
: { scopeNodesEmitted: 0, scopeEdgesEmitted: 0 };
return { edgesEmitted, skippedNoCaller, skippedMissingTarget, ...scopeStats };
}
/**
* Emit `Scope` nodes + `CONTAINS`/`DEFINES`/`IMPORTS` edges representing
* the lexical scope tree itself. Skipped unless `INGESTION_EMIT_SCOPES=1`
* at the public entry point; exported here for tests that want to
* exercise the path directly.
*/
export function emitScopeGraph(input: {
readonly graph: KnowledgeGraph;
readonly scopes: ScopeResolutionIndexes;
}): { readonly scopeNodesEmitted: number; readonly scopeEdgesEmitted: number } {
const { graph, scopes } = input;
let scopeNodesEmitted = 0;
let scopeEdgesEmitted = 0;
for (const scope of scopes.scopeTree.byId.values()) {
graph.addNode({
id: scope.id,
label: 'CodeElement' as NodeLabel, // the generic bucket for non-symbol graph nodes
properties: {
name: scope.kind,
filePath: scope.filePath,
startLine: scope.range.startLine,
endLine: scope.range.endLine,
description: `Scope: ${scope.kind}`,
} as unknown as Parameters<KnowledgeGraph['addNode']>[0]['properties'],
});
scopeNodesEmitted++;
if (scope.parent !== null) {
graph.addRelationship({
id: `rel:contains:${scope.parent}->${scope.id}`,
sourceId: scope.parent,
targetId: scope.id,
type: 'CONTAINS',
confidence: 1,
reason: 'scope-tree parent/child',
});
scopeEdgesEmitted++;
}
for (const def of scope.ownedDefs) {
graph.addRelationship({
id: `rel:defines:${scope.id}->${def.nodeId}`,
sourceId: scope.id,
targetId: def.nodeId,
type: 'DEFINES',
confidence: 1,
reason: 'scope.ownedDefs',
});
scopeEdgesEmitted++;
}
}
for (const [scopeId, edges] of scopes.imports) {
for (const edge of edges) {
if (edge.targetModuleScope === undefined) continue;
graph.addRelationship({
id: `rel:imports:${scopeId}->${edge.targetModuleScope}:${edge.localName}`,
sourceId: scopeId,
targetId: edge.targetModuleScope,
type: 'IMPORTS',
confidence: edge.linkStatus === 'unresolved' ? 0.5 : 1,
reason: `import ${edge.kind} ${edge.localName}`,
});
scopeEdgesEmitted++;
}
}
return { scopeNodesEmitted, scopeEdgesEmitted };
}
// ─── Internal ───────────────────────────────────────────────────────────────
/** Accepted truthy values for `INGESTION_EMIT_SCOPES`. */
const TRUTHY: ReadonlySet<string> = new Set(['true', '1', 'yes']);
function isScopeEmissionEnabled(): boolean {
const raw = process.env['INGESTION_EMIT_SCOPES'];
if (raw === undefined) return false;
return TRUTHY.has(raw.trim().toLowerCase());
}
/**
* Walk up from `startScope` looking for the first ancestor scope whose
* `ownedDefs` contains a Function-like def (Function / Method /
* Constructor). Fall back to the innermost ancestor's first `ownedDef`
* if none is found; return `undefined` if all ancestors have no defs.
*/
function resolveCallerNodeId(
startScope: ScopeId,
scopes: ScopeResolutionIndexes,
): string | undefined {
const tree = scopes.scopeTree;
let current: ScopeId | null = startScope;
const visited = new Set<ScopeId>();
let firstOwnedFallback: string | undefined;
while (current !== null) {
if (visited.has(current)) break;
visited.add(current);
const scope: Scope | undefined = tree.getScope(current);
if (scope === undefined) break;
// Prefer a Function-like owner.
const fnDef = scope.ownedDefs.find((d) => isFunctionLike(d.type));
if (fnDef !== undefined) return fnDef.nodeId;
// Stash the first owned def we see as a conservative fallback.
if (firstOwnedFallback === undefined && scope.ownedDefs.length > 0) {
firstOwnedFallback = scope.ownedDefs[0]!.nodeId;
}
current = scope.parent;
}
return firstOwnedFallback;
}
function isFunctionLike(type: NodeLabel): boolean {
return type === 'Function' || type === 'Method' || type === 'Constructor';
}
function buildRelationship(
ref: Reference,
callerId: string,
targetDef: SymbolDefinition,
sourceLabel: string,
): Parameters<KnowledgeGraph['addRelationship']>[0] {
const type = mapKindToType(ref.kind);
const reason = `${sourceLabel}: ${ref.kind} | confidence ${ref.confidence.toFixed(3)}`;
// `step` encodes read/write discriminator for ACCESSES edges (1=read, 2=write).
// Other kinds omit `step`.
const step = ref.kind === 'read' ? 1 : ref.kind === 'write' ? 2 : undefined;
return {
id: `rel:${type}:${callerId}->${targetDef.nodeId}:${ref.atRange.startLine}:${ref.atRange.startCol}`,
sourceId: callerId,
targetId: targetDef.nodeId,
type,
confidence: ref.confidence,
reason,
evidence: ref.evidence.map(serializeEvidence),
...(step !== undefined ? { step } : {}),
};
}
/**
* Map a `Reference.kind` to an existing `RelationshipType`. Read/write
* both fold into `ACCESSES`; `type-reference` + `import-use` both fold
* into `USES`. This keeps the graph schema additive — no new
* RelationshipType values are introduced by this module.
*/
function mapKindToType(kind: Reference['kind']): RelationshipType {
switch (kind) {
case 'call':
return 'CALLS';
case 'read':
case 'write':
return 'ACCESSES';
case 'inherits':
return 'INHERITS';
case 'type-reference':
case 'import-use':
return 'USES';
}
}
function serializeEvidence(e: ResolutionEvidence): {
readonly kind: string;
readonly weight: number;
readonly note?: string;
} {
return {
kind: e.kind,
weight: e.weight,
...(e.note !== undefined ? { note: e.note } : {}),
};
}

Some files were not shown because too many files have changed in this diff Show More