Compare commits

...
Author SHA1 Message Date
copilot-swe-agent[bot] 011339f777 docs: confirm RING2-PKG-7 implementation already shipped
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2ca478a0-4386-40fa-a190-964c60227c86
2026-04-19 06:01:03 +00:00
copilot-swe-agent[bot] 9531d01a43 Initial plan 2026-04-19 05:56:40 +00:00
Gergő Magyar 6222b5be9b feat(ingestion): emit-references drains ReferenceIndex to graph edges (#925, RFC #909 Ring 2 PKG) (#973) 2026-04-18 23:36:10 +01:00
Gergő Magyar e2ba4a04c9 feat(ingestion): shadow-mode parity harness + static dashboard (#923, RFC #909 Ring 2 PKG) (#972)
* feat(ingestion): shadow-mode parity harness + static dashboard (#923, RFC #909 Ring 2 PKG)

Side-car observability for the RFC #909 registry rollout. Callers that
dual-run legacy-DAG + `Registry.lookup` feed their result pairs into
the harness; the harness diffs each pair via shared `diffResolutions`
(#918), aggregates via `aggregateDiffs`, and persists a per-language
parity report that the static dashboard can render offline.

## Shipped

### `gitnexus/src/core/ingestion/shadow-harness.ts` (new)

```ts
createShadowHarness(): ShadowHarness
```

API:
  - `enabled` — `true` iff `GITNEXUS_SHADOW_MODE` is truthy at
    construction. Captured once; later env-var mutations don't flip it.
  - `record({ language, callsite, legacy, newResult, primary })` —
    accumulator. No-op when `enabled === false` (near-zero overhead).
  - `size()` — diagnostic counter.
  - `snapshot(now?)` — deterministic `ShadowParityReport` from the
    accumulated diffs.
  - `persist(outputDir, now?)` — writes BOTH a timestamped
    `<runId>.json` and a `latest.json` pointer. Creates outputDir if
    absent. Returns the per-run file path.
  - `clear()` — resets the accumulator; preserves `enabled`.

Activation: `GITNEXUS_SHADOW_MODE` accepts `'true'` / `'1'` / `'yes'`
(case-insensitive, trimmed); same truthy convention as
`REGISTRY_PRIMARY_<LANG>` from #924. Typos → disabled (fail-safe).

Persisted payload (`PersistedShadowReport`) is schema-versioned (`v1`):

```jsonc
{
  "schemaVersion": 1,
  "runId": "YYYYMMDD-HHMMSS-xxxxxxxx",
  "generatedAt": "ISO 8601",
  "primaryByLanguage": { "python": "legacy", ... },
  "report": { /* ShadowParityReport from #918 aggregateDiffs */ }
}
```

`runId` prefix is the timestamp so files sort chronologically; the
entropy suffix prevents collisions within a clock-second.

### `gitnexus/shadow-parity-dashboard/index.html` (new)

Minimal static dashboard — one HTML file, zero build step, zero runtime
deps. Fetches `./latest.json` and renders:

  - Overall summary cards (total calls, both agree, disagree, overall parity %)
  - Per-language table: language tag ("primary: legacy" / "primary:
    registry" pill) + total / agree / only-legacy / only-new / disagree
    / both-empty / parity%
  - Parity cells colored by threshold: ≥95% green, ≥80% amber, <80% red
  - Light / dark via `prefers-color-scheme`
  - Empty-state message when no records yet

File-serving is static: `cp .gitnexus/shadow-parity/latest.json
gitnexus/shadow-parity-dashboard/` + open in a browser.

## Tests (14, all passing)

  - **Flag detection** (5): default off · truthy variants case-insensitive ·
    falsy / typo → off · record() is no-op when disabled · env flip
    AFTER construction doesn't enable (constructed-once semantics)
  - **Record + snapshot** (4): multi-language accumulation ·
    per-language rows with correct outcomes · snapshot determinism ·
    `clear()` resets accumulator + `primaryByLanguage`
  - **Persistence** (5): mkdir-p on missing outputDir · per-run +
    latest.json match byte-for-byte · schema v1 payload shape ·
    runId timestamp prefix sorts chronologically · empty report
    persists gracefully

Tests use a per-test tmpdir (`fs.mkdtemp`), cleaned in `afterEach`,
so parallel vitest runs don't collide. `GITNEXUS_SHADOW_MODE` is
saved + restored per-test.

## What's deliberately NOT in this PR (call-out in harness docstring)

  - **Dual-run dispatch.** The harness is a side-car — it does NOT
    invoke either resolution path. Call-processor integration that
    actually runs both legacy + registry paths lands as a follow-up.
    Without that integration, `record()` is never called in production
    today. The harness is tested in isolation with synthetic inputs.
  - **CI artifact publishing.** Config work to upload
    `latest.json` + the dashboard HTML per CI run. Tracked separately;
    the harness + dashboard are ready when the CI job wires in.
  - **Fixture-level drill-down.** The issue mentions per-fixture AST
    snippet + evidence trace drill-down. MVP dashboard shows per-language
    rows only; drill-down extends the static JSON format + the dashboard
    JS in a focused follow-up.

## Verification

  - `tsc --noEmit` clean (both `gitnexus-shared` and `gitnexus`)
  - 14/14 new tests pass
  - Full scope-resolution / shadow / model / flag suite: **335/335 pass**

## Part of

  - Parent: #909
  - Depends on (code): #917 (registries), #918 (diff + aggregate)
  - Unblocks Ring 3 language flips: the parity dashboard becomes the
    checkpoint before flipping `REGISTRY_PRIMARY_<LANG>=true` for a
    language — once per-language parity stabilizes, the flip ships.

* chore: prettier format on shadow-parity-dashboard index.html
2026-04-18 21:30:43 +01:00
Gergő Magyar 0c37eda482 feat(ingestion): per-language resolveImportTarget adapter (#922, RFC #909 Ring 2 PKG) (#971)
Bridges the CLI's existing per-language `ImportResolverFn`s (16 languages
already implemented) to the shared `FinalizeHooks.resolveImportTarget`
contract consumed by `finalize()` (#915) and
`finalizeScopeModel` (#921).

No resolver logic is reimplemented — the adapter wraps
`provider.importResolver` from each `LanguageProvider` verbatim.

## Shipped

### `import-target-adapter.ts` (new)

```ts
buildImportTargetWorkspace(providers, resolveCtx): ImportTargetWorkspace
resolveImportTargetAcrossLanguages(targetRaw, fromFile, workspaceIndex): string | null
```

  - `ImportTargetWorkspace` is the opaque `workspaceIndex` shape the
    adapter recognizes: `{ perLanguage: Map<SupportedLanguages,
    { resolver, ctx }> }`. Callers build it once per ingestion run from
    the active language providers.
  - `resolveImportTargetAcrossLanguages` is the `FinalizeHook`
    implementation. It:
      1. Reads `getLanguageFromFilename(fromFile)`.
      2. Looks up the per-language entry.
      3. Calls the existing `ImportResolverFn` — same signature, same
         code path the legacy DAG uses today.
      4. Picks `result.files[0]` (covers both `'files'` and `'package'`
         result kinds; the legacy pipeline's richer multi-file + dirSuffix
         semantics stay accessible through `importResolver` directly).
      5. Returns `null` on any null result, empty files[], unknown
         extension, missing workspace, or resolver exception.
  - Exceptions from resolvers are swallowed — the finalize algorithm
    treats `null` as `linkStatus: 'unresolved'`, which is the right
    fallback for malformed inputs.

### What's deliberately NOT here

  - **Re-implementation of any per-language resolver.** Wraps the
    existing `importResolver` field on each provider.
  - **Dynamic-import handling.** The shared finalize algorithm short-
    circuits `ParsedImport { kind: 'dynamic-unresolved' }` before
    calling `resolveImportTarget`, so the adapter never sees them.
  - **`importPathPreprocessor`.** Preprocessing belongs inside the
    provider's `interpretImport` hook that produces
    `ParsedImport.targetRaw`; the adapter forwards that verbatim.

## Tests (12, all passing)

  - **`buildImportTargetWorkspace`** (3): registers providers with
    importResolver · skips providers without · threads shared ctx
    into every entry
  - **`resolveImportTargetAcrossLanguages`** (9): forwards targetRaw +
    fromFile · dispatches by extension · null resolver result →
    null · `package`-kind takes first file · empty files[] → null ·
    no registered resolver → null · unknown extension → null ·
    undefined/malformed workspace → null · resolver throw → null

Real per-language resolver correctness is covered by the existing
per-language resolver test suites — the adapter is the bridge layer.

## Verification

  - `tsc --noEmit` clean (both `gitnexus-shared` and `gitnexus`)
  - `gitnexus-shared` build clean
  - 12/12 new tests pass
  - Full scope-resolution / shadow / model / flag suite: **333/333 pass**

## Integration flow

```ts
const workspace = buildImportTargetWorkspace(providers, resolveCtx);
const indexes = finalizeScopeModel(parsedFiles, {
  hooks: { resolveImportTarget: resolveImportTargetAcrossLanguages },
  workspaceIndex: workspace,
});
model.attachScopeIndexes(indexes);
```

## Closes part of #909. Unblocks

  - Ring 3 language migrations (#926+): a language flipping to
    `REGISTRY_PRIMARY_<LANG>=true` now has correct import-target
    resolution out of the box via its existing `importResolver`.
  - #923 shadow harness — can run the dual-path comparison knowing
    both sides use the same per-language resolution semantics.
2026-04-18 21:11:24 +01:00
Copilotandmagyargergo 3adb97e993 feat(docker): ship signed UI + CLI/server images via docker-compose (#967)
* Initial plan

* docker: ship signed UI + CLI/server images via docker-compose

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/883bcee1-4a1d-4b3d-bbb9-accd8846da96

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docker: lock image version to npm package + harden cosign verify guidance

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6afd4fcd-5656-4e02-b796-a22b59000bde

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docker: add Sigstore ClusterImagePolicy + k8s admission docs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/aea3dd70-e2a9-443a-b578-cb3eca4093e1

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docker(k8s): collapse redundant image globs in ClusterImagePolicy

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/aea3dd70-e2a9-443a-b578-cb3eca4093e1

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docker(ci): drop deprecated COSIGN_EXPERIMENTAL, dead build-args, and loose verify regex in comment

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/bdf0d2cf-607c-4558-982a-be9b216b2d36

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docker(ci): use ${{ github.repository }} in verify-comment regex for fork portability

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/bdf0d2cf-607c-4558-982a-be9b216b2d36

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* style(deploy): prettier-format cluster-image-policy.yaml (single quotes)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b09a016e-56a1-4b73-bacc-e69084a48782

* ci(docker): drop workflow_dispatch, harden signing loop, fix verify-comment placeholder

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6755064c-7871-4b2e-9b46-b4779eb215ac

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-04-18 20:55:19 +01:00
Gergő Magyar 25520e90a5 feat(ingestion): finalize-orchestrator materializes ScopeResolutionIndexes (#921, RFC #909 Ring 2 PKG) (#970)
Ties the Ring 2 pipeline together. Takes the `ParsedFile[]` produced by
#920's parse-worker integration, feeds them to shared `finalize()`
(#915), and bundles every workspace-wide index for attachment onto
`MutableSemanticModel`. Thin integration glue per issue #884's boundary
— all algorithm lives in `gitnexus-shared`.

## Shipped

### `model/scope-resolution-indexes.ts` (new)

```ts
interface ScopeResolutionIndexes {
  readonly scopeTree: ScopeTree;
  readonly defs: DefIndex;
  readonly qualifiedNames: QualifiedNameIndex;
  readonly moduleScopes: ModuleScopeIndex;
  readonly methodDispatch: MethodDispatchIndex;
  readonly imports: ReadonlyMap<ScopeId, readonly ImportEdge[]>;
  readonly bindings: ReadonlyMap<ScopeId, ReadonlyMap<string, readonly BindingRef[]>>;
  readonly referenceSites: readonly ReferenceSite[];
  readonly sccs: readonly FinalizedScc[];
  readonly stats: FinalizeStats;
}
```

The bundle produced by the orchestrator, consumed by the resolution
phase. `ReferenceIndex` is deliberately NOT here — it's populated in
the next phase (#925).

### `model/semantic-model.ts` — extended

  - `SemanticModel.scopes?: ScopeResolutionIndexes` — undefined until
    attached; once attached, frozen.
  - `MutableSemanticModel.attachScopeIndexes(indexes)` — one-shot write.
    Throws on second call; `Object.freeze`s the bundle on write. `clear()`
    resets the slot back to `undefined` so re-ingestion can re-attach.

### `finalize-orchestrator.ts` (new)

```ts
finalizeScopeModel(parsedFiles, options?): ScopeResolutionIndexes
```

Orchestration steps:

  1. Map `ParsedFile[]` → `FinalizeInput` (`FinalizeFile` is a structural
     subset, so no shape-shifting).
  2. Call shared `finalize()` with provider hooks (defaults provided for
     the zero-provider case today).
  3. Build the four workspace indexes (`DefIndex`, `QualifiedNameIndex`,
     `ModuleScopeIndex`, `ScopeTree`) from per-file unions.
  4. Build an empty `MethodDispatchIndex` as a placeholder (owners=[],
     both callbacks return []). Real MRO wiring lands with the
     per-language adapters in #922.
  5. Bundle + return.

**Empty-input safety.** Zero parsedFiles → valid but empty bundle with
all zero-sized indexes and `stats.totalFiles === 0`. Downstream code
can consult `model.scopes` without branching on presence — only on
`stats`.

**Hook defaults** (`withDefaultHooks`) for missing provider hooks:

  - `resolveImportTarget: () => null` — every import goes `unresolved`
  - `expandsWildcardTo: () => []` — wildcards don't materialize
  - `mergeBindings: (a, b) => [...a, ...b]` — append without precedence

Providers override these in #922 (per-language import adapters).

## Tests (10, all passing)

  - **Empty input** (1): zero parsedFiles → valid empty bundle
  - **Single file** (2): all per-file indexes populated · referenceSites
    aggregated
  - **Cross-file imports** (3): resolveImportTarget threads through +
    links · default-null resolver → unresolved · stats reflect graph
  - **MutableSemanticModel integration** (4): undefined initially · attach
    once · Object.freeze applied · throws on re-attach · clear() resets

## Verification

  - `tsc --noEmit` clean in both packages
  - `gitnexus-shared` build clean
  - 10/10 new tests pass
  - Full scope-resolution / shadow / model / flag suite: **321/321 pass**

## What's deferred (not this PR, per RFC #909 scope)

  - **Per-language hook adapters** (#922): `resolveImportTarget` +
    `expandsWildcardTo` + `mergeBindings` wired per language.
  - **MethodDispatchIndex wiring via HeritageMap**: populate MRO + implements
    via the existing CLI-package HeritageMap strategies. Likely companion
    to #922 or a focused follow-up.
  - **Pipeline invocation**: actually calling `finalizeScopeModel` from
    the real ingestion pipeline. The orchestrator is callable today; the
    ingestion entry point wiring lands with the shadow harness (#923).
  - **`ReferenceIndex` population**: RFC §3.2 Phase 4 / #925.

## Closes part of #909. Unblocks
  - #923 shadow harness — now has a fully materialized `model.scopes` to
    query against the legacy DAG for parity measurement
  - #925 ReferenceIndex → LadybugDB emission — consumes `model.scopes`
  - Ring 3 language migrations (#926+) — a language flipping to
    `REGISTRY_PRIMARY_<LANG>=true` can now expect `model.scopes` to be
    populated when the pipeline wires the orchestrator in
2026-04-18 20:51:19 +01:00
Gergő Magyar 39b5d295c7 feat(ingestion): wire ScopeExtractor into parse-worker + processor (#920, RFC #909 Ring 2 PKG) (#969)
Plumbs the ScopeExtractor (#919) into the real parsing pipeline.
`ParsedFile` artifacts now flow from workers to the parsing-processor
without changing any legacy-DAG behavior.

## Shipped

### `gitnexus/src/core/ingestion/scope-extractor-bridge.ts` (new)

  - `extractParsedFile(provider, sourceText, filePath, onWarn?)`
  - Short-circuits (returns `undefined`) when the provider has not
    implemented `emitScopeCaptures`. True for every language today —
    this is the default no-op path.
  - Invokes the hook + `ScopeExtractor.extract`, returns a `ParsedFile`.
  - **Swallows exceptions on both sides.** Failures route through the
    optional `onWarn` callback (or `console.warn`) and return
    `undefined`. Scope-extraction errors NEVER break legacy parsing on
    the same file.
  - Standalone module (not nested in `parse-worker.ts`) so tests can
    import it directly without triggering the worker's top-level
    `parentPort!.on(...)`.

### `gitnexus/src/core/ingestion/workers/parse-worker.ts`

  - `ParseWorkerResult.parsedFiles: ParsedFile[]` added.
  - `processFileGroup` calls `extractParsedFile` AFTER tree parse,
    BEFORE legacy extraction. Worker provides an `onWarn` callback that
    routes bridge warnings through `parentPort.postMessage({ type:
    'warning', message })`.
  - `mergeResult` includes `parsedFiles` in the sub-batch merge.
  - Initial + reset accumulator templates include `parsedFiles: []`.

### `gitnexus/src/core/ingestion/parsing-processor.ts`

  - `WorkerExtractedData.parsedFiles: ParsedFile[]` added.
  - Empty-result branch and the across-chunk aggregation both include
    `parsedFiles`. Aggregation is tolerant of workers that don't emit
    the field (older builds / partial rollouts).

### Ring 1 tweak: `emitScopeCaptures` sync return

`readonly CaptureMatch[]` (was `Promise<readonly CaptureMatch[]>`).
Tree-sitter and COBOL's regex tagger are both synchronous; no
foreseeable need for async work inside this hook. Sync lets the
already-sync worker pipeline invoke it inline without cascading
`async` up through the batch driver + IPC handler.

## Tests (9 new; full suite 311/311)

`gitnexus/test/unit/scope-resolution/parse-worker-scope-integration.test.ts`:
  - Not-migrated (2): undefined-returning hook · never-invokes-extractor
  - Migrated (3): happy path · argument threading · honors
    `shouldCreateScope` override
  - Error resilience (4): hook throws · extractor throws (no Module) ·
    extractor throws (sibling overlap) · `onWarn` gets routed
    message with filePath + error body

## Verification

  - `tsc --noEmit` clean in both packages
  - `gitnexus-shared` build clean
  - 311/311 combined scope-resolution / shadow / model / flag suite
  - 9/9 new bridge tests

## What's NOT in this PR (still deferred to #921)

  - Actually using the `parsedFiles` — that's the finalize orchestrator.
  - `ModuleScopeIndex.byFilePath` materialization — belongs alongside
    the rest of the SemanticModel indexes in #921.

## Closes part of #909. Unblocks
  - #921 finalize-orchestrator — consumes `WorkerExtractedData.parsedFiles`
2026-04-18 20:27:56 +01:00
25 changed files with 3089 additions and 53 deletions
+15 -3
View File
@@ -1,3 +1,15 @@
IMAGE_NAME=ghcr.io/abhigyanpatwari/gitnexus:latest
CONTAINER_NAME=gitnexus
HOST_PORT=4173
# Images (signed Cosign keyless on every push from main / vX.Y.Z tags)
SERVER_IMAGE=ghcr.io/abhigyanpatwari/gitnexus:latest
WEB_IMAGE=ghcr.io/abhigyanpatwari/gitnexus-web:latest
# Container names
SERVER_CONTAINER_NAME=gitnexus-server
WEB_CONTAINER_NAME=gitnexus-web
# Host ports — the web UI expects the server on http://localhost:4747 by default.
SERVER_HOST_PORT=4747
WEB_HOST_PORT=4173
# Optional read-only mount, exposed to the server as /workspace.
# Override with the directory that contains the repos you want to index.
WORKSPACE_DIR=./
+103 -19
View File
@@ -4,30 +4,72 @@ on:
push:
tags:
- 'v*'
branches:
- main
paths-ignore: ['**.md', 'docs/**', 'LICENSE']
workflow_dispatch:
# No workflow_dispatch: publishing is exclusively tag-driven so that every
# signed image corresponds 1:1 to a published `gitnexus@X.Y.Z` on npm. A
# manual run from a branch ref would fail the version check below anyway.
# Concurrency convention: see CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
# Tag refs are unique per release — distinct tags run in parallel.
# Pushes to main serialize; cancel superseded runs.
# Tag refs are unique per release, so distinct tags run in parallel.
# Re-pushes of the same tag serialize. cancel-in-progress: false — never cancel a publish mid-flight.
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: ${{ github.ref == 'refs/heads/main' }}
cancel-in-progress: false
jobs:
build-push:
name: Build & Push image
name: Build & Push ${{ matrix.image.name }}
runs-on: ubuntu-latest
timeout-minutes: 30
timeout-minutes: 60
permissions:
contents: read
packages: write
# Required for Cosign keyless signing via the OIDC token exchange,
# and for build provenance / SBOM attestations.
id-token: write
attestations: write
strategy:
fail-fast: false
matrix:
image:
# Static UI bundle. Small, fast image. Drop-in replacement for the
# legacy single-image setup at the same `gitnexus` repository slug
# is intentionally avoided — the UI now lives at `gitnexus-web` and
# the CLI/server takes the canonical `gitnexus` slug below.
- name: gitnexus-web
dockerfile: Dockerfile.web
slug: gitnexus-web
# CLI / `gitnexus serve` backend. Heavy native deps (tree-sitter,
# onnxruntime-node) live only in this image.
- name: gitnexus
dockerfile: Dockerfile.cli
slug: gitnexus
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
# ── Lock the docker image version to the npm package version ──────────
# Mirrors the check in publish.yml: refuse to build unless the git tag
# exactly matches `gitnexus/package.json`'s version. This guarantees
# `ghcr.io/<owner>/gitnexus:X.Y.Z` always corresponds to the same
# `gitnexus@X.Y.Z` published to npm — no drift, no surprises.
- name: Verify tag matches gitnexus/package.json version
id: version
shell: bash
run: |
TAG_VERSION="${GITHUB_REF#refs/tags/v}"
if ! [[ "$TAG_VERSION" =~ ^[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$ ]]; then
echo "::error::Tag does not follow semver: v$TAG_VERSION"
exit 1
fi
PKG_VERSION=$(node -p "require('./gitnexus/package.json').version")
if [ "$TAG_VERSION" != "$PKG_VERSION" ]; then
echo "::error::Tag version (v$TAG_VERSION) does not match gitnexus/package.json version ($PKG_VERSION)"
exit 1
fi
echo "version=$PKG_VERSION" >> "$GITHUB_OUTPUT"
echo "Version verified: $PKG_VERSION"
# Required for multi-platform (linux/arm64) emulation.
- name: Set up QEMU
uses: docker/setup-qemu-action@ce360397dd3f832beb865e1373c09c0e9f86d70a # v4.0.0
@@ -35,6 +77,9 @@ jobs:
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@4d04d5d9486b7bd6fa91e7baf45bbb4f8b9deedd # v4.0.0
- name: Install Cosign
uses: sigstore/cosign-installer@cad07c2e89fa2edd6e2d7bab4c1aa38e53f76003 # v4.1.1
- name: Log in to GitHub Container Registry
uses: docker/login-action@4907a6ddec9925e35a0a9e82d7399ccc52663121 # v4.1.0
with:
@@ -42,30 +87,69 @@ jobs:
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
# Computes image tags and labels from Git metadata:
# v* tag → ghcr.io/<owner>/<repo>:<semver> (e.g. 1.2.3, 1.2, 1)
# main push → ghcr.io/<owner>/<repo>:latest
# Computes image tags and labels from the verified semver tag:
# v1.2.3 → :1.2.3, :1.2, :1, :latest (auto, only for non-prerelease)
# v1.2.3-rc.1 → :1.2.3-rc.1 only (prereleases never become :latest)
# `:latest` is only emitted for tag pushes thanks to `flavor: latest=auto`,
# ensuring it always points at a real npm-published version.
- name: Extract Docker metadata
id: meta
uses: docker/metadata-action@030e881283bb7a6894de51c315a6bfe6a94e05cf # v6.0.0
with:
images: ghcr.io/${{ github.repository }}
images: ghcr.io/${{ github.repository_owner }}/${{ matrix.image.slug }}
flavor: latest=auto
tags: |
type=semver,pattern={{version}}
type=semver,pattern={{major}}.{{minor}}
type=semver,pattern={{major}}
type=raw,value=latest,enable={{is_default_branch}}
type=sha,prefix=sha-,format=short
- name: Build and push
id: build
uses: docker/build-push-action@bcafcacb16a39f128d818304e6c9c0c18556b85f # v7.1.0
with:
context: .
file: ${{ matrix.image.dockerfile }}
platforms: linux/amd64,linux/arm64
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
cache-from: type=gha
cache-to: type=gha,mode=max
build-args: |
BUILDPLATFORM=${{ runner.os == 'Linux' && 'linux/amd64' || 'linux/amd64' }}
cache-from: type=gha,scope=${{ matrix.image.slug }}
cache-to: type=gha,mode=max,scope=${{ matrix.image.slug }}
provenance: mode=max
sbom: true
# Cosign keyless signing. Each pushed tag is signed by the workflow's
# OIDC identity, so consumers can verify the image with the strict,
# fully-anchored identity regex (kept in sync with README.md and
# deploy/kubernetes/cluster-image-policy.yaml — update all three together).
# NOTE: `${...}` expression syntax is NOT evaluated inside YAML comments, so
# the example below uses literal `<owner>/<repo>` placeholders that consumers
# substitute themselves; the canonical, fully-rendered command lives in README.md.
# cosign verify ghcr.io/<owner>/<slug>:<tag> \
# --certificate-identity-regexp '^https://github\.com/<owner>/<repo>/\.github/workflows/docker\.yml@refs/tags/v[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$' \
# --certificate-oidc-issuer https://token.actions.githubusercontent.com
# Do NOT relax to `@.*` — that accepts signatures from any ref, including
# unprotected branches and PRs, and defeats the supply-chain guarantee.
- name: Sign image with Cosign (keyless)
env:
# Cosign v2 (installed by sigstore/cosign-installer above) makes
# keyless the default. COSIGN_EXPERIMENTAL is a v1-only opt-in flag
# that is now deprecated/no-op, so it is intentionally omitted.
DIGEST: ${{ steps.build.outputs.digest }}
TAGS: ${{ steps.meta.outputs.tags }}
run: |
# Sign every tag at the same digest so consumers can verify by tag or by digest.
# Use `while read` instead of `for $TAGS` to be robust against tags that
# could ever contain whitespace (the metadata-action output is newline-
# separated, not space-separated).
while IFS= read -r tag; do
[[ -n "$tag" ]] && cosign sign --yes "${tag}@${DIGEST}"
done <<< "$TAGS"
# Attach the SBOM produced by buildx as a verifiable attestation on the digest.
- name: Generate build provenance attestation
uses: actions/attest-build-provenance@a2bbfa25375fe432b6a289bc6b6cd05ecd0c4c32 # v4.1.0
with:
subject-name: ghcr.io/${{ github.repository_owner }}/${{ matrix.image.slug }}
subject-digest: ${{ steps.build.outputs.digest }}
push-to-registry: true
+57
View File
@@ -0,0 +1,57 @@
ARG BUILDPLATFORM
ARG TARGETPLATFORM
# ── Builder ────────────────────────────────────────────────────────────
# Native modules (tree-sitter-*, onnxruntime-node, node-gyp builds for
# tree-sitter-proto / tree-sitter-swift) require python3 + a C/C++ toolchain.
FROM node:22-alpine AS builder
WORKDIR /app
# Toolchain for node-gyp / native builds.
RUN apk add --no-cache python3 make g++ git
# Build gitnexus-shared first — gitnexus depends on it as a workspace.
COPY gitnexus-shared/package.json gitnexus-shared/package-lock.json ./gitnexus-shared/
RUN npm ci --prefix gitnexus-shared
COPY gitnexus-shared ./gitnexus-shared
RUN npm run build --prefix gitnexus-shared
# Copy the full gitnexus package before installing — `npm ci` triggers
# `postinstall` (patches tree-sitter-swift, builds the vendored
# tree-sitter-proto) and `prepare` (compiles TypeScript via scripts/build.js),
# both of which need the source tree.
COPY gitnexus ./gitnexus
RUN npm ci --prefix gitnexus
# Drop dev dependencies for a smaller runtime layer.
RUN npm prune --omit=dev --prefix gitnexus
# ── Runtime ────────────────────────────────────────────────────────────
FROM node:22-alpine AS runtime
# curl for the healthcheck; git so `gitnexus` can clone repos at runtime.
RUN apk add --no-cache curl git
WORKDIR /app
# Pre-create the data directory and hand it to the unprivileged `node` user
# so the bind-mounted volume is writable without root.
RUN mkdir -p /data/gitnexus && chown -R node:node /data
COPY --from=builder --chown=node:node /app/gitnexus/dist ./gitnexus/dist
COPY --from=builder --chown=node:node /app/gitnexus/node_modules ./gitnexus/node_modules
COPY --from=builder --chown=node:node /app/gitnexus/package.json ./gitnexus/package.json
COPY --from=builder --chown=node:node /app/gitnexus/vendor ./gitnexus/vendor
USER node
# The web UI defaults to http://localhost:4747 — keep that contract.
ENV GITNEXUS_HOME=/data/gitnexus \
NODE_ENV=production \
PORT=4747
EXPOSE 4747
# Bind to 0.0.0.0 so the server is reachable from the host's mapped port.
CMD ["node", "gitnexus/dist/cli/index.js", "serve", "--host", "0.0.0.0", "--port", "4747"]
-1
View File
@@ -11,7 +11,6 @@ RUN npm ci --prefix gitnexus-shared
COPY gitnexus-shared ./gitnexus-shared
RUN npm run build --prefix gitnexus-shared
COPY gitnexus/package.json ./gitnexus/package.json
COPY gitnexus-web/package.json gitnexus-web/package-lock.json ./gitnexus-web/
RUN npm ci --prefix gitnexus-web
+118 -19
View File
@@ -337,39 +337,138 @@ npm run dev
## Docker
```bash
docker run --rm \
--name gitnexus \
-p 4173:4173 \
ghcr.io/abhigyanpatwari/gitnexus:latest
```
The official Docker setup ships **two signed images** orchestrated by `docker-compose.yaml`:
Or with Docker Compose:
| Image | Purpose |
| -------------------------------------------------- | ---------------------------------------------------------------------- |
| `ghcr.io/abhigyanpatwari/gitnexus:latest` | CLI / `gitnexus serve` backend (HTTP API on port `4747`, MCP, indexer) |
| `ghcr.io/abhigyanpatwari/gitnexus-web:latest` | Static web UI (port `4173`) |
> **Heads-up — image rename.** Earlier releases published the web UI under
> `ghcr.io/abhigyanpatwari/gitnexus`. Starting with the introduction of the
> bundled backend, that slug now hosts the CLI/server image and the UI moved
> to `ghcr.io/abhigyanpatwari/gitnexus-web`. The previous tags remain
> available for pulling, but new versions are only published under the new
> slugs. Update your `docker run` / compose files accordingly (or just adopt
> the bundled compose).
### One-command setup
```bash
docker compose up -d
```
Optional env file:
This starts the server on `http://localhost:4747` and the web UI on
`http://localhost:4173`. The UI auto-detects the server because the browser
runs on the host and reaches the container via the mapped port.
A named volume (`gitnexus-data`) persists the global registry, indexes, and
cloned repos at `/data/gitnexus` inside the server container. To make repos on
your host machine indexable, set `WORKSPACE_DIR` before bringing the stack up:
```bash
WORKSPACE_DIR=$HOME/code docker compose up -d
# Inside the server container the directory is mounted read-only at /workspace.
docker compose exec gitnexus-server gitnexus index /workspace/my-repo
```
### Direct `docker run`
```bash
# Server
docker run --rm -d \
--name gitnexus-server \
-p 4747:4747 \
-v gitnexus-data:/data/gitnexus \
ghcr.io/abhigyanpatwari/gitnexus:latest
# Web UI
docker run --rm -d \
--name gitnexus-web \
-p 4173:4173 \
ghcr.io/abhigyanpatwari/gitnexus-web:latest
```
Optional env file (override image tags, container names, ports, workspace dir):
```bash
cp .env.example .env
set -a
source .env
set +a
docker compose --env-file .env up -d
```
Docker files:
### Versioning & supply-chain protection
- [Dockerfile](Dockerfile) is the source for the published `gitnexus` image. It builds `gitnexus-shared` and `gitnexus-web`, then serves the production frontend.
- [docker-compose.yaml](docker-compose.yaml) starts the published image with Docker Compose.
- [.env.example](.env.example) sets the image name, container name, and exposed port for the example commands.
The Docker images are version-locked to the npm package:
Notes:
- Both images are **only published from `vX.Y.Z` git tags**, and the workflow
refuses to build unless the tag exactly matches `gitnexus/package.json`'s
version. So `ghcr.io/abhigyanpatwari/gitnexus:1.6.2` is byte-for-byte the
same release as `npm install gitnexus@1.6.2` — no drift, no floating
builds from `main`.
- `:latest` is auto-promoted only from non-prerelease tags by the Docker
metadata action, so it always points at a real, npm-published version.
- The published image serves the production frontend only. It does not start `gitnexus serve`.
- In backend mode, the app still defaults to `http://localhost:4747` unless you change the server URL in the UI.
- If you do not want an env file, the defaults are `ghcr.io/abhigyanpatwari/gitnexus:latest`, container name `gitnexus`, and port `4173`.
Both images are signed with [Cosign keyless signing][cosign-keyless] using the
workflow's GitHub OIDC identity, and shipped with build provenance and SBOM
attestations. **This is your protection against supply-chain attacks**: even if
an attacker republishes a same-named image elsewhere (or somehow pushes to a
typo-squatted registry), they cannot forge a Cosign signature tied to
`abhigyanpatwari/GitNexus`'s `docker.yml`. Always verify before pulling into
sensitive environments:
```bash
cosign verify ghcr.io/abhigyanpatwari/gitnexus:1.6.2 \
--certificate-identity-regexp '^https://github\.com/abhigyanpatwari/GitNexus/\.github/workflows/docker\.yml@refs/tags/v[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com
```
The regex pins the certificate identity to this repo's `docker.yml` workflow
**run from a `v*` tag** — rejecting unsigned images, images signed by other
workflows, and images signed from unprotected refs.
You can also inspect the build provenance and SBOM:
```bash
cosign download attestation ghcr.io/abhigyanpatwari/gitnexus:1.6.2 \
--predicate-type https://slsa.dev/provenance/v1
```
#### Kubernetes: enforce signatures at admission
For Kubernetes deployments, ship the bundled
[`ClusterImagePolicy`](deploy/kubernetes/cluster-image-policy.yaml) so the
[Sigstore policy-controller][policy-controller] rejects any GitNexus pod whose
image is not signed by this repo's `docker.yml` running from a `vX.Y.Z` tag —
the same identity the `cosign verify` snippet above pins.
```bash
# 1. Install the controller (one-time, cluster-wide)
helm repo add sigstore https://sigstore.github.io/helm-charts && helm repo update
helm install policy-controller -n cosign-system --create-namespace \
sigstore/policy-controller
# 2. Opt your namespace in
kubectl label namespace <your-ns> policy.sigstore.dev/include=true
# 3. Apply the policy
kubectl apply -f deploy/kubernetes/cluster-image-policy.yaml
```
After this, attempting to deploy an unsigned image — or one signed by anything
other than `abhigyanpatwari/GitNexus`'s `docker.yml` at a `v*` tag — fails the
admission webhook before a pod is ever created. This turns the verifiable
signature into an enforced policy, which is the supply-chain control most
clusters actually need.
[cosign-keyless]: https://docs.sigstore.dev/cosign/signing/overview/
[policy-controller]: https://docs.sigstore.dev/policy-controller/overview/
### Files
- [Dockerfile.web](Dockerfile.web) — builds `gitnexus-shared` and `gitnexus-web`, then serves the production frontend.
- [Dockerfile.cli](Dockerfile.cli) — builds the CLI/server (with its native deps) and runs `gitnexus serve --host 0.0.0.0`.
- [docker-compose.yaml](docker-compose.yaml) — starts both signed images side by side.
- [.env.example](.env.example) — overrides for image names, container names, ports, and the workspace mount.
The web UI uses the same indexing pipeline as the CLI but runs entirely in WebAssembly (Tree-sitter WASM, LadybugDB WASM, in-browser embeddings). It's great for quick exploration but limited by browser memory for larger repos.
@@ -0,0 +1,64 @@
# Sigstore policy-controller ClusterImagePolicy for GitNexus container images.
#
# This enforces — at admission time — that every Pod pulling a
# `ghcr.io/abhigyanpatwari/gitnexus` or `gitnexus-web` image is using a build
# that was Cosign-keyless-signed by this repository's `docker.yml` workflow
# running from a `vX.Y.Z` git tag. Unsigned images, images signed by other
# workflows, and images signed from unprotected refs (e.g. `main`, PR branches)
# are rejected.
#
# Prerequisites
# -------------
# 1. Install the Sigstore policy-controller in your cluster (Helm):
#
# helm repo add sigstore https://sigstore.github.io/helm-charts
# helm repo update
# helm install policy-controller -n cosign-system --create-namespace \
# sigstore/policy-controller
#
# 2. Opt namespaces in to verification:
#
# kubectl label namespace <your-ns> policy.sigstore.dev/include=true
#
# 3. Apply this policy:
#
# kubectl apply -f deploy/kubernetes/cluster-image-policy.yaml
#
# After this, `kubectl run --image=ghcr.io/abhigyanpatwari/gitnexus:<tag>` in
# any opted-in namespace will only succeed if the image carries a valid
# Sigstore signature with the pinned identity.
#
# References
# - https://docs.sigstore.dev/policy-controller/overview/
# - https://github.com/sigstore/policy-controller
apiVersion: policy.sigstore.dev/v1beta1
kind: ClusterImagePolicy
metadata:
name: gitnexus-signed-images
spec:
# Apply to both published GitNexus images on GHCR. Image references always
# carry a tag or digest at admission time, so these two globs cover every
# `gitnexus:<tag>`, `gitnexus@sha256:...`, `gitnexus-web:<tag>`, and
# `gitnexus-web@sha256:...` reference.
images:
- glob: 'ghcr.io/abhigyanpatwari/gitnexus*'
authorities:
- name: gitnexus-cosign-keyless
keyless:
# Public-good Sigstore Fulcio root.
url: https://fulcio.sigstore.dev
identities:
# Pin both the OIDC issuer (GitHub Actions) AND the exact workflow
# path running from a `vX.Y.Z` (or `vX.Y.Z-prerelease`) tag. Same
# regex the README's `cosign verify` example uses; it rejects:
# * unsigned images
# * signatures from any other repo / workflow
# * signatures from non-tag refs (main, PRs, release branches)
# * signatures from arbitrary non-semver tags
- issuer: https://token.actions.githubusercontent.com
subjectRegExp: ^https://github\.com/abhigyanpatwari/GitNexus/\.github/workflows/docker\.yml@refs/tags/v[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$
# Cross-check the signature against the public Rekor transparency log,
# so an attacker who briefly compromised Fulcio cannot retroactively
# mint a signature without leaving a public, append-only audit record.
ctlog:
url: https://rekor.sigstore.dev
+36 -4
View File
@@ -1,9 +1,38 @@
services:
gitnexus:
image: ${IMAGE_NAME:-ghcr.io/brainifii/gitnexus:latest}
container_name: ${CONTAINER_NAME:-gitnexus}
gitnexus-server:
image: ${SERVER_IMAGE:-ghcr.io/abhigyanpatwari/gitnexus:latest}
container_name: ${SERVER_CONTAINER_NAME:-gitnexus-server}
# Map the server to the same host port the web UI expects by default
# (http://localhost:4747). The browser runs on the host, so the UI's
# built-in default works without any reconfiguration.
ports:
- '${HOST_PORT:-4173}:4173'
- '${SERVER_HOST_PORT:-4747}:4747'
volumes:
# Persist the global registry, indexes, and cloned repos across runs.
- gitnexus-data:/data/gitnexus
# Optional: mount a host workspace so `gitnexus index <path>` can see
# repos you already have on disk. The default points at an empty
# `./workspace/` sibling that compose will create on first start —
# it intentionally does NOT bind-mount the repo root, which would
# expose `.git`, `.env`, and CI secrets to the container.
# Override with `WORKSPACE_DIR=/abs/path/to/your/repos`.
- ${WORKSPACE_DIR:-./workspace}:/workspace:ro
restart: unless-stopped
healthcheck:
test: ['CMD', 'curl', '-fsS', 'http://localhost:4747/api/heartbeat']
interval: 30s
timeout: 5s
retries: 3
start_period: 15s
gitnexus-web:
image: ${WEB_IMAGE:-ghcr.io/abhigyanpatwari/gitnexus-web:latest}
container_name: ${WEB_CONTAINER_NAME:-gitnexus-web}
ports:
- '${WEB_HOST_PORT:-4173}:4173'
depends_on:
gitnexus-server:
condition: service_healthy
restart: unless-stopped
healthcheck:
test: ['CMD', 'curl', '-f', 'http://localhost:4173/']
@@ -11,3 +40,6 @@ services:
timeout: 5s
retries: 3
start_period: 10s
volumes:
gitnexus-data:
+16
View File
@@ -131,4 +131,20 @@ export interface GraphRelationship {
confidence: number;
reason: string;
step?: number;
/**
* Per-signal evidence trace for edges emitted by the scope-based
* resolution pipeline (RFC #909 Ring 2 PKG #925). Populated by
* `emit-references.ts` when draining `ReferenceIndex` into the graph
* so downstream query / audit tools can inspect *why* a given edge
* was emitted with its confidence value.
*
* Optional and additive — every existing edge emitter ignores this
* field, and every existing query continues to work whether or not
* an edge carries it.
*/
evidence?: readonly {
readonly kind: string;
readonly weight: number;
readonly note?: string;
}[];
}
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "gitnexus",
"version": "1.6.1",
"version": "1.6.2",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "gitnexus",
"version": "1.6.1",
"version": "1.6.2",
"hasInstallScript": true,
"license": "PolyForm-Noncommercial-1.0.0",
"dependencies": {
+291
View File
@@ -0,0 +1,291 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width,initial-scale=1" />
<title>GitNexus — Shadow Parity Dashboard</title>
<!--
Static dashboard for the RFC #909 shadow-mode parity report.
Reads `latest.json` from this directory and renders a per-language
parity table. Zero build step, zero runtime dependencies — a
single file that any browser or file:// context can open.
Usage:
# from repo root, after a shadow-mode run
cp .gitnexus/shadow-parity/latest.json gitnexus/shadow-parity-dashboard/
open gitnexus/shadow-parity-dashboard/index.html
CI artifact wiring (follow-up): the CI job publishes a snapshot
of this directory + latest.json as a downloadable bundle per run.
-->
<style>
:root {
color-scheme: light dark;
--fg: #1f2937;
--fg-muted: #6b7280;
--bg: #ffffff;
--bg-muted: #f9fafb;
--border: #e5e7eb;
--good: #16a34a;
--warn: #d97706;
--bad: #dc2626;
--primary-tag-legacy: #7c3aed;
--primary-tag-registry: #0ea5e9;
}
@media (prefers-color-scheme: dark) {
:root {
--fg: #e5e7eb;
--fg-muted: #9ca3af;
--bg: #111827;
--bg-muted: #1f2937;
--border: #374151;
}
}
html,
body {
margin: 0;
padding: 0;
background: var(--bg);
color: var(--fg);
font:
14px/1.45 system-ui,
-apple-system,
sans-serif;
}
main {
max-width: 1200px;
margin: 0 auto;
padding: 24px 16px;
}
h1 {
font-size: 20px;
margin: 0 0 4px;
}
.meta {
color: var(--fg-muted);
font-size: 12px;
margin-bottom: 20px;
}
.cards {
display: grid;
grid-template-columns: repeat(auto-fit, minmax(180px, 1fr));
gap: 10px;
margin-bottom: 20px;
}
.card {
border: 1px solid var(--border);
border-radius: 6px;
padding: 10px 12px;
background: var(--bg-muted);
}
.card .k {
color: var(--fg-muted);
font-size: 11px;
text-transform: uppercase;
letter-spacing: 0.04em;
}
.card .v {
font-size: 20px;
font-weight: 600;
}
table {
width: 100%;
border-collapse: collapse;
font-variant-numeric: tabular-nums;
}
th,
td {
padding: 6px 10px;
text-align: right;
border-bottom: 1px solid var(--border);
}
th:first-child,
td:first-child {
text-align: left;
}
thead th {
font-weight: 600;
color: var(--fg-muted);
font-size: 12px;
background: var(--bg-muted);
}
tbody tr:hover {
background: var(--bg-muted);
}
.parity {
font-weight: 600;
}
.parity.good {
color: var(--good);
}
.parity.warn {
color: var(--warn);
}
.parity.bad {
color: var(--bad);
}
.tag {
display: inline-block;
padding: 1px 6px;
border-radius: 10px;
font-size: 10px;
margin-left: 6px;
color: white;
}
.tag.legacy {
background: var(--primary-tag-legacy);
}
.tag.registry {
background: var(--primary-tag-registry);
}
.empty {
padding: 40px;
text-align: center;
color: var(--fg-muted);
}
code {
font-family: ui-monospace, SFMono-Regular, Menlo, monospace;
background: var(--bg-muted);
padding: 1px 4px;
border-radius: 3px;
}
</style>
</head>
<body>
<main>
<h1>Shadow Parity — RFC #909</h1>
<div class="meta" id="meta">loading <code>latest.json</code>…</div>
<div class="cards" id="cards"></div>
<table id="per-language">
<thead>
<tr>
<th>Language</th>
<th>Total</th>
<th>Agree</th>
<th>Only legacy</th>
<th>Only new</th>
<th>Disagree</th>
<th>Both empty</th>
<th>Parity</th>
</tr>
</thead>
<tbody></tbody>
</table>
<div id="empty" class="empty" style="display: none">
No records yet. Enable <code>GITNEXUS_SHADOW_MODE=1</code> and run ingestion to populate.
</div>
</main>
<script>
/* global fetch, document */
(async function () {
const tbody = document.querySelector('#per-language tbody');
const cards = document.getElementById('cards');
const meta = document.getElementById('meta');
const empty = document.getElementById('empty');
const table = document.getElementById('per-language');
let payload;
try {
const r = await fetch('./latest.json', { cache: 'no-store' });
if (!r.ok) throw new Error('HTTP ' + r.status);
payload = await r.json();
} catch (err) {
meta.textContent = 'Failed to load latest.json: ' + err.message;
table.style.display = 'none';
empty.style.display = 'block';
return;
}
const primary = payload.primaryByLanguage || {};
const report = payload.report || {};
const perLang = report.perLanguage || [];
const overall = report.overall || {};
meta.textContent =
'Run ' +
payload.runId +
' — generated ' +
payload.generatedAt +
' (schema v' +
payload.schemaVersion +
')';
// Overall summary cards.
cards.innerHTML = '';
const overallParity = overall.parity !== undefined ? overall.parity : 0;
cards.appendChild(makeCard('Total calls', overall.totalCalls ?? 0));
cards.appendChild(makeCard('Both agree', overall.bothAgree ?? 0));
cards.appendChild(makeCard('Disagree', overall.bothDisagree ?? 0));
cards.appendChild(makeCard('Overall parity', formatPct(overallParity)));
if (!perLang.length) {
table.style.display = 'none';
empty.style.display = 'block';
return;
}
for (const row of perLang) {
const tr = document.createElement('tr');
const primaryTag = primary[row.language];
const tag = primaryTag
? '<span class="tag ' + primaryTag + '">primary: ' + primaryTag + '</span>'
: '';
const parityClass = parityClassFor(row.parity);
tr.innerHTML =
'<td>' +
escape(row.language) +
tag +
'</td>' +
'<td>' +
row.totalCalls +
'</td>' +
'<td>' +
row.bothAgree +
'</td>' +
'<td>' +
row.onlyLegacy +
'</td>' +
'<td>' +
row.onlyNew +
'</td>' +
'<td>' +
row.bothDisagree +
'</td>' +
'<td>' +
row.bothEmpty +
'</td>' +
'<td class="parity ' +
parityClass +
'">' +
formatPct(row.parity) +
'</td>';
tbody.appendChild(tr);
}
function makeCard(k, v) {
const div = document.createElement('div');
div.className = 'card';
div.innerHTML =
'<div class="k">' + escape(k) + '</div><div class="v">' + escape(String(v)) + '</div>';
return div;
}
function formatPct(x) {
if (typeof x !== 'number' || !isFinite(x)) return '—';
return (x * 100).toFixed(1) + '%';
}
function parityClassFor(x) {
if (typeof x !== 'number') return '';
if (x >= 0.95) return 'good';
if (x >= 0.8) return 'warn';
return 'bad';
}
function escape(s) {
return String(s).replace(/[&<>"']/g, function (c) {
return { '&': '&amp;', '<': '&lt;', '>': '&gt;', '"': '&quot;', "'": '&#39;' }[c];
});
}
})();
</script>
</body>
</html>
@@ -0,0 +1,299 @@
/**
* Phase 5 of the RFC #909 ingestion lifecycle: drain `ReferenceIndex`
* into the knowledge graph as labeled edges with `confidence` and
* `evidence` properties (Ring 2 PKG #925).
*
* The resolution phase (future PR) writes `Reference` records into
* `model.scopes.referenceSites`-derived `ReferenceIndex`; this module
* materializes those records as `GraphRelationship`s via
* `graph.addRelationship`. Every emitted edge carries:
*
* - `type`: one of `'CALLS' | 'ACCESSES' | 'INHERITS' | 'USES'`
* (mapped from `Reference.kind` — `'read'` and `'write'` both route
* to `ACCESSES`; `'type-reference'` and `'import-use'` route to
* `USES`; `'call'` stays `CALLS`; `'inherits'` stays `INHERITS`).
* - `confidence`: the pre-computed confidence from the Reference record.
* - `reason`: human-readable summary (`"scope-resolution: call | confidence 0.75"`).
* - `evidence`: the full `ResolutionEvidence[]` trace — additive graph
* property (see `GraphRelationship.evidence` in gitnexus-shared),
* so queries that don't know about it are unaffected.
* - `step`: carries the reference's access-kind discriminant when
* available (`1` for read, `2` for write) so `ACCESSES` edges retain
* the read/write distinction without forcing a new edge type.
*
* ## Optional scope-tree flush
*
* When `INGESTION_EMIT_SCOPES=1` is set, this module also emits:
*
* - `Scope` nodes for every `Scope` in the tree
* - `CONTAINS` edges from parent scope to child scope
* - `DEFINES` edges from scope to its `ownedDefs` members
* - `IMPORTS` edges from scope to `targetModuleScope` of each finalized
* `ImportEdge` that carries one
*
* Off by default — existing queries that don't know about `Scope` nodes
* continue to work, and the storage cost is opt-in.
*
* ## Source-of-truth: the caller def for a reference
*
* A `Reference` says "some code inside `fromScope` references `toDef`".
* The graph wants `(callerNodeId, calleeNodeId)`. We resolve the caller
* by walking up the scope tree from `fromScope` until we find a scope
* whose `ownedDefs` contains a Function-like def. If no such ancestor
* exists, the edge is attributed to the first def owned by the innermost
* ancestor scope, and if THAT produces nothing either the edge is
* skipped (with a count returned in `EmitStats.skippedNoCaller`).
*/
import type {
NodeLabel,
RelationshipType,
Reference,
ReferenceIndex,
ResolutionEvidence,
Scope,
ScopeId,
SymbolDefinition,
} from 'gitnexus-shared';
import type { KnowledgeGraph } from '../graph/types.js';
import type { ScopeResolutionIndexes } from './model/scope-resolution-indexes.js';
// ─── Public API ─────────────────────────────────────────────────────────────
export interface EmitStats {
readonly edgesEmitted: number;
/** References dropped because no caller def could be resolved. */
readonly skippedNoCaller: number;
/** References dropped because `toDef` was not found in the DefIndex. */
readonly skippedMissingTarget: number;
/** Scope nodes emitted — `0` unless `INGESTION_EMIT_SCOPES=1`. */
readonly scopeNodesEmitted: number;
/** Scope-tree structural edges emitted — `0` unless `INGESTION_EMIT_SCOPES=1`. */
readonly scopeEdgesEmitted: number;
}
export interface EmitReferencesInput {
readonly graph: KnowledgeGraph;
readonly scopes: ScopeResolutionIndexes;
readonly referenceIndex: ReferenceIndex;
/** Human-consumable label for the `reason` prefix. Defaults to `'scope-resolution'`. */
readonly sourceLabel?: string;
}
/**
* Drain `referenceIndex.bySourceScope` into graph edges.
*
* The scope-tree flush is controlled separately by
* `INGESTION_EMIT_SCOPES` — callers can run `emitReferencesToGraph`
* without scope-node emission or layer the two calls as needed.
*/
export function emitReferencesToGraph(input: EmitReferencesInput): EmitStats {
const { graph, scopes, referenceIndex } = input;
const sourceLabel = input.sourceLabel ?? 'scope-resolution';
let edgesEmitted = 0;
let skippedNoCaller = 0;
let skippedMissingTarget = 0;
for (const [fromScope, refs] of referenceIndex.bySourceScope) {
for (const ref of refs) {
const targetDef = scopes.defs.get(ref.toDef);
if (targetDef === undefined) {
skippedMissingTarget++;
continue;
}
const callerId = resolveCallerNodeId(fromScope, scopes);
if (callerId === undefined) {
skippedNoCaller++;
continue;
}
graph.addRelationship(buildRelationship(ref, callerId, targetDef, sourceLabel));
edgesEmitted++;
}
}
const scopeStats = isScopeEmissionEnabled()
? emitScopeGraph({ graph, scopes })
: { scopeNodesEmitted: 0, scopeEdgesEmitted: 0 };
return { edgesEmitted, skippedNoCaller, skippedMissingTarget, ...scopeStats };
}
/**
* Emit `Scope` nodes + `CONTAINS`/`DEFINES`/`IMPORTS` edges representing
* the lexical scope tree itself. Skipped unless `INGESTION_EMIT_SCOPES=1`
* at the public entry point; exported here for tests that want to
* exercise the path directly.
*/
export function emitScopeGraph(input: {
readonly graph: KnowledgeGraph;
readonly scopes: ScopeResolutionIndexes;
}): { readonly scopeNodesEmitted: number; readonly scopeEdgesEmitted: number } {
const { graph, scopes } = input;
let scopeNodesEmitted = 0;
let scopeEdgesEmitted = 0;
for (const scope of scopes.scopeTree.byId.values()) {
graph.addNode({
id: scope.id,
label: 'CodeElement' as NodeLabel, // the generic bucket for non-symbol graph nodes
properties: {
name: scope.kind,
filePath: scope.filePath,
startLine: scope.range.startLine,
endLine: scope.range.endLine,
description: `Scope: ${scope.kind}`,
} as unknown as Parameters<KnowledgeGraph['addNode']>[0]['properties'],
});
scopeNodesEmitted++;
if (scope.parent !== null) {
graph.addRelationship({
id: `rel:contains:${scope.parent}->${scope.id}`,
sourceId: scope.parent,
targetId: scope.id,
type: 'CONTAINS',
confidence: 1,
reason: 'scope-tree parent/child',
});
scopeEdgesEmitted++;
}
for (const def of scope.ownedDefs) {
graph.addRelationship({
id: `rel:defines:${scope.id}->${def.nodeId}`,
sourceId: scope.id,
targetId: def.nodeId,
type: 'DEFINES',
confidence: 1,
reason: 'scope.ownedDefs',
});
scopeEdgesEmitted++;
}
}
for (const [scopeId, edges] of scopes.imports) {
for (const edge of edges) {
if (edge.targetModuleScope === undefined) continue;
graph.addRelationship({
id: `rel:imports:${scopeId}->${edge.targetModuleScope}:${edge.localName}`,
sourceId: scopeId,
targetId: edge.targetModuleScope,
type: 'IMPORTS',
confidence: edge.linkStatus === 'unresolved' ? 0.5 : 1,
reason: `import ${edge.kind} ${edge.localName}`,
});
scopeEdgesEmitted++;
}
}
return { scopeNodesEmitted, scopeEdgesEmitted };
}
// ─── Internal ───────────────────────────────────────────────────────────────
/** Accepted truthy values for `INGESTION_EMIT_SCOPES`. */
const TRUTHY: ReadonlySet<string> = new Set(['true', '1', 'yes']);
function isScopeEmissionEnabled(): boolean {
const raw = process.env['INGESTION_EMIT_SCOPES'];
if (raw === undefined) return false;
return TRUTHY.has(raw.trim().toLowerCase());
}
/**
* Walk up from `startScope` looking for the first ancestor scope whose
* `ownedDefs` contains a Function-like def (Function / Method /
* Constructor). Fall back to the innermost ancestor's first `ownedDef`
* if none is found; return `undefined` if all ancestors have no defs.
*/
function resolveCallerNodeId(
startScope: ScopeId,
scopes: ScopeResolutionIndexes,
): string | undefined {
const tree = scopes.scopeTree;
let current: ScopeId | null = startScope;
const visited = new Set<ScopeId>();
let firstOwnedFallback: string | undefined;
while (current !== null) {
if (visited.has(current)) break;
visited.add(current);
const scope: Scope | undefined = tree.getScope(current);
if (scope === undefined) break;
// Prefer a Function-like owner.
const fnDef = scope.ownedDefs.find((d) => isFunctionLike(d.type));
if (fnDef !== undefined) return fnDef.nodeId;
// Stash the first owned def we see as a conservative fallback.
if (firstOwnedFallback === undefined && scope.ownedDefs.length > 0) {
firstOwnedFallback = scope.ownedDefs[0]!.nodeId;
}
current = scope.parent;
}
return firstOwnedFallback;
}
function isFunctionLike(type: NodeLabel): boolean {
return type === 'Function' || type === 'Method' || type === 'Constructor';
}
function buildRelationship(
ref: Reference,
callerId: string,
targetDef: SymbolDefinition,
sourceLabel: string,
): Parameters<KnowledgeGraph['addRelationship']>[0] {
const type = mapKindToType(ref.kind);
const reason = `${sourceLabel}: ${ref.kind} | confidence ${ref.confidence.toFixed(3)}`;
// `step` encodes read/write discriminator for ACCESSES edges (1=read, 2=write).
// Other kinds omit `step`.
const step = ref.kind === 'read' ? 1 : ref.kind === 'write' ? 2 : undefined;
return {
id: `rel:${type}:${callerId}->${targetDef.nodeId}:${ref.atRange.startLine}:${ref.atRange.startCol}`,
sourceId: callerId,
targetId: targetDef.nodeId,
type,
confidence: ref.confidence,
reason,
evidence: ref.evidence.map(serializeEvidence),
...(step !== undefined ? { step } : {}),
};
}
/**
* Map a `Reference.kind` to an existing `RelationshipType`. Read/write
* both fold into `ACCESSES`; `type-reference` + `import-use` both fold
* into `USES`. This keeps the graph schema additive — no new
* RelationshipType values are introduced by this module.
*/
function mapKindToType(kind: Reference['kind']): RelationshipType {
switch (kind) {
case 'call':
return 'CALLS';
case 'read':
case 'write':
return 'ACCESSES';
case 'inherits':
return 'INHERITS';
case 'type-reference':
case 'import-use':
return 'USES';
}
}
function serializeEvidence(e: ResolutionEvidence): {
readonly kind: string;
readonly weight: number;
readonly note?: string;
} {
return {
kind: e.kind,
weight: e.weight,
...(e.note !== undefined ? { note: e.note } : {}),
};
}
@@ -0,0 +1,196 @@
/**
* `finalizeScopeModel` — turn a workspace's `ParsedFile[]` into a
* materialized `ScopeResolutionIndexes` (RFC §3.2 Phase 2; Ring 2 PKG #921).
*
* Thin integration glue, per issue #884's boundary: all algorithmic logic
* lives in `gitnexus-shared` (finalize algorithm #915, the four per-file
* indexes #913, the method-dispatch materialization #914, the scope tree
* #912). This file does three things only:
*
* 1. Map `ParsedFile[]` → `FinalizeInput` and call shared `finalize()`.
* 2. Build the four workspace-wide indexes from the union of per-file
* defs/scopes/modules/qualified-names.
* 3. Bundle the results into `ScopeResolutionIndexes` for
* `MutableSemanticModel.attachScopeIndexes(...)`.
*
* ## What this module is NOT responsible for
*
* - Invoking tree-sitter or running AST walks. That's the extractor (#919).
* - Per-language import-target resolution. Hooks are plumbed through
* but default to "unresolved" when no provider supplies them — the
* real adapters land with #922.
* - Populating `ReferenceIndex`. That's the resolution phase (#925).
* - Deciding which language uses registry-primary lookup. That's the
* flag reader (#924).
*
* ## Empty-input behavior
*
* When `parsedFiles` is empty (the common case today — no language has
* migrated yet), the orchestrator produces a valid but empty bundle: all
* indexes are zero-sized, the scope tree is empty, and
* `finalize.stats.totalFiles === 0`. This lets downstream consumers
* safely consult `model.scopes` without branching on presence.
*/
import type {
BindingRef,
FinalizeFile,
FinalizeHooks,
ParsedFile,
Scope,
ScopeId,
SymbolDefinition,
WorkspaceIndex,
} from 'gitnexus-shared';
import {
buildDefIndex,
buildMethodDispatchIndex,
buildModuleScopeIndex,
buildQualifiedNameIndex,
buildScopeTree,
finalize,
} from 'gitnexus-shared';
import type { ScopeResolutionIndexes } from './model/scope-resolution-indexes.js';
// ─── Public entry point ─────────────────────────────────────────────────────
/**
* Options forwarded to the orchestrator. All fields optional so callers
* that don't yet have per-language hooks (today) get sensible defaults;
* #922 will populate `hooks.resolveImportTarget` + friends per language.
*/
export interface FinalizeOrchestratorOptions {
/**
* Hooks forwarded to shared `finalize()`. Any omitted field gets a
* no-op default: unresolved targets, empty wildcard expansion, append
* merge for bindings.
*/
readonly hooks?: Partial<FinalizeHooks>;
/**
* Opaque workspace context forwarded to hooks. `undefined` today; Ring
* 2 PKG #922 populates this with a real cross-file index for the
* per-language resolvers.
*/
readonly workspaceIndex?: WorkspaceIndex;
}
/**
* Produce a fully materialized `ScopeResolutionIndexes` from the
* workspace's per-file artifacts.
*
* Pure function (given pure hooks). No I/O, no globals consulted. The
* pipeline calls this once per ingestion run and hands the result to
* `MutableSemanticModel.attachScopeIndexes`.
*/
export function finalizeScopeModel(
parsedFiles: readonly ParsedFile[],
options: FinalizeOrchestratorOptions = {},
): ScopeResolutionIndexes {
const hooks = withDefaultHooks(options.hooks ?? {});
const workspaceIndex: WorkspaceIndex = options.workspaceIndex ?? undefined;
// ── Step 1: Shared finalize — runs SCC-aware cross-file link + binding
// materialization. Returns linked imports + merged bindings per module
// scope + SCC condensation + stats.
const finalizeInput = {
files: parsedFiles.map(toFinalizeFile),
workspaceIndex,
};
const finalizeOut = finalize(finalizeInput, hooks);
// ── Step 2: Workspace-wide indexes built from the per-file unions.
// These are pure aggregations — no algorithm beyond what the builders
// in gitnexus-shared already encapsulate (first-write-wins, qname
// collision buckets, etc.).
const allScopes: Scope[] = [];
const allDefs: SymbolDefinition[] = [];
const moduleEntries: { filePath: string; moduleScopeId: ScopeId }[] = [];
const allReferenceSites = [] as ReturnType<typeof collectReferenceSites>;
for (const file of parsedFiles) {
for (const s of file.scopes) allScopes.push(s);
for (const d of file.localDefs) allDefs.push(d);
moduleEntries.push({ filePath: file.filePath, moduleScopeId: file.moduleScope });
}
// References kept out of the loop above to centralize list-init.
allReferenceSites.push(...collectReferenceSites(parsedFiles));
const scopeTree = buildScopeTree(allScopes);
const defs = buildDefIndex(allDefs);
const qualifiedNames = buildQualifiedNameIndex(allDefs);
const moduleScopes = buildModuleScopeIndex(moduleEntries);
// ── Step 3: MethodDispatchIndex. Today we lack per-language MRO
// strategies wired into this orchestrator (that belongs with the
// HeritageMap bridge, a separate piece of work). Ship an EMPTY index
// so the bundle shape is consistent; the callbacks return `[]` for
// every owner and `implementsOf` returns `[]`. Populating this
// properly is tracked alongside the per-language provider hooks.
const methodDispatch = buildMethodDispatchIndex({
owners: [], // empty → no MRO entries; `mroFor(x)` returns the frozen empty array
computeMro: () => [],
implementsOf: () => [],
});
return {
scopeTree,
defs,
qualifiedNames,
moduleScopes,
methodDispatch,
imports: finalizeOut.imports,
bindings: finalizeOut.bindings,
referenceSites: Object.freeze([...allReferenceSites]),
sccs: finalizeOut.sccs,
stats: finalizeOut.stats,
};
}
// ─── Internal ───────────────────────────────────────────────────────────────
/** Shape-reduce a `ParsedFile` to the narrower `FinalizeFile` the shared
* algorithm reads. The subset is stable — `FinalizeFile` is a proper
* subset of `ParsedFile`. */
function toFinalizeFile(file: ParsedFile): FinalizeFile {
return {
filePath: file.filePath,
moduleScope: file.moduleScope,
parsedImports: file.parsedImports,
localDefs: file.localDefs,
};
}
/** Flatten every file's reference sites into one list. Order reflects
* input-file order, then capture order inside each file. Deterministic. */
function collectReferenceSites(parsedFiles: readonly ParsedFile[]) {
const out: ParsedFile['referenceSites'][number][] = [];
for (const file of parsedFiles) {
for (const site of file.referenceSites) out.push(site);
}
return out;
}
/**
* Fill in no-op defaults for any omitted hook. Keeps `finalize()`
* behavior well-defined for the zero-provider case today:
*
* - `resolveImportTarget: () => null` — every import edge ends up
* `linkStatus: 'unresolved'` (or dynamic-unresolved pass-through).
* - `expandsWildcardTo: () => []` — wildcards don't materialize.
* - `mergeBindings: (existing, incoming) => [...existing, ...incoming]`
* — append without precedence; providers override to implement local-
* shadows-import and similar rules.
*/
function withDefaultHooks(partial: Partial<FinalizeHooks>): FinalizeHooks {
return {
resolveImportTarget: partial.resolveImportTarget ?? (() => null),
expandsWildcardTo: partial.expandsWildcardTo ?? (() => []),
mergeBindings:
partial.mergeBindings ??
((
existing: readonly BindingRef[],
incoming: readonly BindingRef[],
): readonly BindingRef[] => [...existing, ...incoming]),
};
}
@@ -0,0 +1,124 @@
/**
* Bridge between CLI-package per-language `ImportResolverFn`s and the
* shared `FinalizeHooks.resolveImportTarget` contract
* (RFC §5.2; Ring 2 PKG #922).
*
* The shared finalize algorithm (#915) asks one question:
*
* resolveImportTarget(targetRaw, fromFile, workspaceIndex): string | null
*
* The CLI already has 16 language-specific resolvers satisfying a
* richer signature:
*
* ImportResolverFn(rawImportPath, filePath, resolveCtx): ImportResult
*
* This module builds a dispatch adapter — one FinalizeHook implementation
* that looks up the file's language from its path and delegates to the
* right per-language resolver. Callers package per-language resolvers +
* a shared `ResolveCtx` into an opaque `ImportTargetWorkspace` and pass
* it as `workspaceIndex` to `finalizeScopeModel`.
*
* ## What's deliberately NOT here
*
* - **Re-implementation of any per-language resolver.** We wrap the
* existing `importResolver` field on each `LanguageProvider` — the
* same code path the legacy DAG uses today.
* - **Dynamic-import handling.** The shared finalize algorithm short-
* circuits `ParsedImport { kind: 'dynamic-unresolved' }` before
* calling `resolveImportTarget`, so the adapter never sees those.
* - **`importPathPreprocessor`.** Preprocessing belongs inside the
* provider's `interpretImport` hook (which writes the final
* `ParsedImport.targetRaw`). By the time finalize passes a
* `targetRaw` to this adapter, it is the string the provider wants
* resolved verbatim.
*/
import {
getLanguageFromFilename,
type SupportedLanguages,
type WorkspaceIndex,
} from 'gitnexus-shared';
import type { ImportResolverFn, ImportResult, ResolveCtx } from './import-resolvers/types.js';
import type { LanguageProvider } from './language-provider.js';
/** A single language's resolver bundled with the context it needs. */
export interface LanguageResolverEntry {
readonly resolver: ImportResolverFn;
readonly ctx: ResolveCtx;
}
/**
* The opaque `workspaceIndex` shape recognized by
* `resolveImportTargetAcrossLanguages`. Built once per ingestion run via
* `buildImportTargetWorkspace`, threaded through `finalizeScopeModel`.
*/
export interface ImportTargetWorkspace {
readonly perLanguage: ReadonlyMap<SupportedLanguages, LanguageResolverEntry>;
}
/**
* Build the workspace index from a map of language → provider. Providers
* whose `importResolver` is absent are silently skipped (no language will
* ever hit that branch at dispatch time).
*
* The `resolveCtx` is shared across all languages. Callers assemble it
* once per run (the existing pipeline already does this for the legacy
* DAG) and hand it to both the legacy resolution path and this factory.
*/
export function buildImportTargetWorkspace(
providers: ReadonlyMap<SupportedLanguages, LanguageProvider>,
resolveCtx: ResolveCtx,
): ImportTargetWorkspace {
const perLanguage = new Map<SupportedLanguages, LanguageResolverEntry>();
for (const [lang, provider] of providers) {
if (provider.importResolver === undefined) continue;
perLanguage.set(lang, { resolver: provider.importResolver, ctx: resolveCtx });
}
return { perLanguage };
}
/**
* The FinalizeHooks-compatible implementation. Dispatches on `fromFile`'s
* extension → per-language resolver. Returns the first resolved file,
* or `null` if the resolver returns `null` or doesn't know about the
* language.
*
* Picks the first entry of `files[]` for both `'files'` and `'package'`
* result kinds — the legacy pipeline uses the whole array, but the
* shared `finalize()` hook contract is single-file. If the workspace
* later needs richer semantics (split-target packages), this is the
* single site to extend.
*/
export function resolveImportTargetAcrossLanguages(
targetRaw: string,
fromFile: string,
workspaceIndex: WorkspaceIndex,
): string | null {
const workspace = workspaceIndex as ImportTargetWorkspace | undefined;
if (workspace === undefined || workspace.perLanguage === undefined) return null;
const lang = getLanguageFromFilename(fromFile);
if (lang === null) return null;
const entry = workspace.perLanguage.get(lang);
if (entry === undefined) return null;
let result: ImportResult;
try {
result = entry.resolver(targetRaw, fromFile, entry.ctx);
} catch {
// Existing resolvers can throw on malformed inputs (e.g., Python
// relative paths above the workspace root). Swallow — the shared
// algorithm treats a null here as `linkStatus: 'unresolved'`, which
// is the right fallback.
return null;
}
if (result === null) return null;
// Both `files` and `package` variants expose a `files` array; the
// package variant also carries `dirSuffix` which we ignore at the
// FinalizeHook boundary (single-file contract). Legacy consumers
// continue to see the full result via `importResolver` directly.
const first = result.files[0];
return first ?? null;
}
@@ -321,12 +321,15 @@ interface LanguageProviderConfig {
* Providers that have not yet migrated continue to run through the
* legacy DAG path (feature-flagged per `REGISTRY_PRIMARY_<LANG>`).
*
* **Sync return.** Tree-sitter query execution and COBOL's regex
* tagger are both synchronous; no current or foreseeable provider
* needs async work inside this hook. The sync signature lets
* `parse-worker.ts` (#920) invoke it inline in its already-sync
* per-file loop without cascading `async` through the batch pipeline.
*
* Default: undefined (language continues to use legacy DAG).
*/
readonly emitScopeCaptures?: (
sourceText: string,
filePath: string,
) => Promise<readonly CaptureMatch[]>;
readonly emitScopeCaptures?: (sourceText: string, filePath: string) => readonly CaptureMatch[];
/**
* Interpret a raw `@import.statement` capture group into a `ParsedImport`.
@@ -0,0 +1,73 @@
/**
* `ScopeResolutionIndexes` — the bundle of materialized indexes produced
* by the finalize-orchestrator (RFC #909 Ring 2 PKG #921) and attached
* to `MutableSemanticModel`.
*
* Produced by `finalizeScopeModel(parsedFiles, hooks)` in
* `finalize-orchestrator.ts`. Consumed by the resolution phase (future
* tickets) where `Registry.lookup` / `resolveTypeRef` query this bundle
* to answer call-resolution questions without re-walking any AST.
*
* ## Lifecycle
*
* 1. Pipeline collects `ParsedFile[]` from the parsing-processor (#920).
* 2. Pipeline invokes `finalizeScopeModel(parsedFiles, hooks)` →
* returns a `ScopeResolutionIndexes` (this interface).
* 3. Pipeline calls `model.attachScopeIndexes(indexes)` to stamp them
* onto the `MutableSemanticModel`. This is a **one-shot write**;
* subsequent calls throw. After attachment, the indexes are frozen
* at the type level (everything is `readonly`) and at runtime via
* `Object.freeze` on the bundle.
* 4. Resolution callers hold a `SemanticModel` reference and read
* `model.scopes` to query.
*
* ## Content
*
* - `scopeTree` / `moduleScopes` / `defs` / `qualifiedNames` — the
* four Ring 2 SHARED indexes built over per-file artifacts.
* - `methodDispatch` — MRO + implements materialized view (#914).
* - `imports` — finalized `ImportEdge[]` per module scope (`parsedImports`
* resolved through cross-file link + wildcard expansion).
* - `bindings` — merged bindings per module scope (local + import +
* wildcard + re-export), with the provider's precedence applied.
* - `referenceSites` — union of every file's pre-resolution usage
* facts. Consumed by the resolution phase (future) to emit
* `Reference` records into `ReferenceIndex`.
* - `stats` — coarse-grained counts from the shared finalize algorithm
* (total files/edges, linked vs unresolved, SCC topology).
*
* `ReferenceIndex` is deliberately NOT here — it is populated in a later
* phase (RFC §3.2 Phase 4 / Ring 2 PKG #925) and owned separately.
*/
import type {
BindingRef,
DefIndex,
FinalizedScc,
FinalizeStats,
ImportEdge,
MethodDispatchIndex,
ModuleScopeIndex,
QualifiedNameIndex,
ReferenceSite,
ScopeId,
ScopeTree,
} from 'gitnexus-shared';
export interface ScopeResolutionIndexes {
readonly scopeTree: ScopeTree;
readonly defs: DefIndex;
readonly qualifiedNames: QualifiedNameIndex;
readonly moduleScopes: ModuleScopeIndex;
readonly methodDispatch: MethodDispatchIndex;
/** Finalized `ImportEdge[]` per module scope. */
readonly imports: ReadonlyMap<ScopeId, readonly ImportEdge[]>;
/** Merged bindings (local + imports + wildcards) per module scope. */
readonly bindings: ReadonlyMap<ScopeId, ReadonlyMap<string, readonly BindingRef[]>>;
/** Pre-resolution usage facts; consumed by the resolution phase. */
readonly referenceSites: readonly ReferenceSite[];
/** SCC condensation of the file-level import graph — callers that want
* parallel per-SCC processing in the resolution phase read this. */
readonly sccs: readonly FinalizedScc[];
readonly stats: FinalizeStats;
}
@@ -57,6 +57,7 @@ import type { SymbolDefinition } from 'gitnexus-shared';
import type { SymbolTableReader, SymbolTableWriter, AddMetadata } from './symbol-table.js';
import { createSymbolTable } from './symbol-table.js';
import { createRegistrationTable } from './registration-table.js';
import type { ScopeResolutionIndexes } from './scope-resolution-indexes.js';
// ---------------------------------------------------------------------------
// Public read-only interface
@@ -83,6 +84,19 @@ export interface SemanticModel {
readonly methods: MethodRegistry;
readonly fields: FieldRegistry;
readonly symbols: SymbolTableReader;
/**
* Materialized scope-resolution indexes from RFC #909 Ring 2 PKG #921.
*
* `undefined` until the finalize-orchestrator attaches them. While
* `undefined`, the legacy DAG is the sole resolution surface; once set,
* resolvers whose language has `REGISTRY_PRIMARY_<LANG>=true` consult
* these indexes instead.
*
* The attach is a one-shot write (see `MutableSemanticModel`). Callers
* holding a read-only `SemanticModel` handle see either `undefined` or
* the final frozen bundle — never a half-populated view.
*/
readonly scopes?: ScopeResolutionIndexes;
}
// ---------------------------------------------------------------------------
@@ -100,6 +114,17 @@ export interface MutableSemanticModel extends SemanticModel {
readonly symbols: SymbolTableWriter;
/** Clear all registries AND the nested SymbolTable. */
clear(): void;
/**
* Stamp the finalize-orchestrator's output onto this model.
*
* **One-shot write.** Throws when called a second time — the indexes are
* meant to be materialized once per ingestion run. `Object.freeze` is
* applied to the attached bundle so consumers cannot mutate after attach.
*
* `clear()` resets the attached bundle back to `undefined`, enabling a
* fresh re-ingestion to attach a new bundle.
*/
attachScopeIndexes(indexes: ScopeResolutionIndexes): void;
}
// ---------------------------------------------------------------------------
@@ -152,6 +177,21 @@ export const createSemanticModel = (): MutableSemanticModel => {
return def;
};
// Scope-resolution bundle slot. Starts `undefined`; populated by a
// one-shot `attachScopeIndexes(...)` from the finalize-orchestrator.
// Held inside the factory closure so the returned `SemanticModel`
// surface exposes it as a plain `readonly` property without a setter.
let attachedScopes: ScopeResolutionIndexes | undefined;
const attachScopeIndexes = (indexes: ScopeResolutionIndexes): void => {
if (attachedScopes !== undefined) {
throw new Error(
'SemanticModel: scope indexes already attached. ' + 'Call `clear()` before re-attaching.',
);
}
attachedScopes = Object.freeze(indexes);
};
// Cascade clear: single source of truth for "reset the entire model".
// Wired into both `model.clear()` AND `model.symbols.clear()` so that a
// caller holding only a SymbolTable reference can't leave the
@@ -162,6 +202,7 @@ export const createSemanticModel = (): MutableSemanticModel => {
methods.clear();
fields.clear();
rawSymbols.clear();
attachedScopes = undefined;
};
// Writer-typed facade: exposes reads + add, but NO `clear` field.
@@ -184,6 +225,10 @@ export const createSemanticModel = (): MutableSemanticModel => {
methods,
fields,
symbols,
get scopes() {
return attachedScopes;
},
clear: cascadeClear,
attachScopeIndexes,
};
};
@@ -31,6 +31,7 @@ import {
buildCollisionGroups,
} from './utils/method-props.js';
import type { LanguageProvider } from './language-provider.js';
import type { ParsedFile } from 'gitnexus-shared';
import { WorkerPool } from './workers/worker-pool.js';
import type {
ParseWorkerResult,
@@ -62,6 +63,14 @@ export interface WorkerExtractedData {
ormQueries: ExtractedORMQuery[];
constructorBindings: FileConstructorBindings[];
fileScopeBindings: FileScopeBindings[];
/**
* Per-file `ParsedFile` artifacts from the new scope-based resolution
* pipeline (RFC #909 Ring 2). Empty until a provider implements
* `emitScopeCaptures` — additive to the legacy DAG path. Aggregated
* from every worker chunk; consumed downstream by #921's
* finalize-orchestrator.
*/
parsedFiles: ParsedFile[];
}
// ============================================================================
@@ -96,6 +105,7 @@ const processParsingWithWorkers = async (
ormQueries: [],
constructorBindings: [],
fileScopeBindings: [],
parsedFiles: [],
};
const total = files.length;
@@ -120,6 +130,7 @@ const processParsingWithWorkers = async (
const allORMQueries: ExtractedORMQuery[] = [];
const allConstructorBindings: FileConstructorBindings[] = [];
const fileScopeBindingsByFile: FileScopeBindings[] = [];
const allParsedFiles: ParsedFile[] = [];
for (const result of chunkResults) {
for (const node of result.nodes) {
graph.addNode({
@@ -157,6 +168,11 @@ const processParsingWithWorkers = async (
for (const item of result.constructorBindings) allConstructorBindings.push(item);
if (result.fileScopeBindings)
for (const item of result.fileScopeBindings) fileScopeBindingsByFile.push(item);
// RFC #909 Ring 2: aggregate per-file scope artifacts. Tolerant of
// workers that don't emit the field yet (older worker builds or
// partial rollouts), since the additive contract means undefined =
// "this worker produced no ParsedFiles for this chunk".
if (result.parsedFiles) for (const item of result.parsedFiles) allParsedFiles.push(item);
}
// Merge and log skipped languages from workers
@@ -187,6 +203,7 @@ const processParsingWithWorkers = async (
ormQueries: allORMQueries,
constructorBindings: allConstructorBindings,
fileScopeBindings: fileScopeBindingsByFile,
parsedFiles: allParsedFiles,
};
};
@@ -0,0 +1,54 @@
/**
* Bridge between a language provider's `emitScopeCaptures` hook and the
* `ScopeExtractor` (RFC #909 Ring 2 PKG #920).
*
* Extracted into its own module so it can be imported by test code
* without pulling in `parse-worker.ts` — which has a top-level
* `parentPort!.on('message', ...)` call that assumes a worker-thread
* context and throws on direct import.
*
* The bridge:
*
* 1. Short-circuits when the provider has NOT implemented
* `emitScopeCaptures`. Returns `undefined`; zero work done. This is
* the state of every language today — `ParsedFile` production stays
* dormant until a language migrates.
* 2. Invokes the hook + feeds its output to `ScopeExtractor.extract`.
* 3. **Swallows exceptions from either side.** A failure here returns
* `undefined` and emits a warning via `onWarn`; legacy parsing on
* the same file continues unaffected by the scope-extraction miss.
* Scope-based resolution is the new path under construction — it
* must not destabilize the legacy DAG.
*/
import type { ParsedFile } from 'gitnexus-shared';
import { extract as extractScope } from './scope-extractor.js';
import type { LanguageProvider } from './language-provider.js';
/** Callback used to report scope-extraction warnings to the host (worker or direct). */
export type ScopeBridgeWarn = (message: string) => void;
/**
* Produce a `ParsedFile` for the given file, or `undefined` when the
* provider hasn't migrated / the extractor throws. Never propagates
* exceptions.
*/
export function extractParsedFile(
provider: LanguageProvider,
sourceText: string,
filePath: string,
onWarn?: ScopeBridgeWarn,
): ParsedFile | undefined {
if (provider.emitScopeCaptures === undefined) return undefined;
try {
const captures = provider.emitScopeCaptures(sourceText, filePath);
return extractScope(captures, filePath, provider);
} catch (err) {
const message = `scope extraction failed for ${filePath}: ${
err instanceof Error ? err.message : String(err)
}`;
if (onWarn !== undefined) onWarn(message);
else console.warn(message);
return undefined;
}
}
@@ -0,0 +1,222 @@
/**
* Shadow-mode parity harness — dual-run observability for the RFC #909
* registry rollout (RFC §6.3; Ring 2 PKG #923).
*
* ## What it does
*
* - Exposes `record({ language, callsite, legacy, newResult })` for
* every call site where the caller has BOTH a legacy-DAG resolution
* and a new `Registry.lookup` resolution.
* - Computes a `ShadowDiff` per record via shared `diffResolutions`
* (#918) and accumulates them in a per-language bucket.
* - At the end of a run, aggregates into a `ShadowParityReport` via
* shared `aggregateDiffs` (#918) — per-language parity %,
* evidence-kind breakdown of divergences, grand-total overall row.
* - Optionally persists the report as JSON under
* `.gitnexus/shadow-parity/` so the static dashboard at
* `gitnexus/shadow-parity-dashboard/` can render it offline.
*
* ## What it does NOT do
*
* - **Invoke either resolution path itself.** The caller must run
* legacy + `Registry.lookup` and pass results in. The harness is a
* side-car, not a dispatcher — this keeps call-processor integration
* surgical when it lands (tracked as a follow-up; the shared model
* doesn't dual-invoke on its own).
* - **Flip anything.** `REGISTRY_PRIMARY_<LANG>` lives in
* `registry-primary-flag.ts` (#924); the harness records the
* caller-supplied "which side is primary" bit for each record so the
* dashboard can label rows, but it does not consult the flag itself.
*
* ## Activation
*
* `GITNEXUS_SHADOW_MODE=1` (or `'true'`, `'yes'`, case-insensitive,
* trimmed) enables the harness. When disabled, `record()` is a cheap
* no-op: no accumulation, no allocation beyond the harness object
* itself. Callers can always construct a harness and hand it through;
* the "off" overhead is near-zero.
*
* ## Persistence shape
*
* When `persist()` is called, the harness writes TWO files:
*
* - `<outputDir>/<runId>.json` — the timestamped snapshot (immutable)
* - `<outputDir>/latest.json` — a pointer that the dashboard reads
*
* Both files contain the same `PersistedShadowReport` payload:
*
* {
* schemaVersion: 1,
* runId: "<iso-8601>-<rand>",
* generatedAt: "<iso-8601>",
* primaryByLanguage: { [lang]: "legacy" | "registry" },
* report: <ShadowParityReport>
* }
*
* Schema-version-gated so future format changes don't silently confuse
* older dashboards. The dashboard renders `report.perLanguage` rows and
* annotates each with `primaryByLanguage[lang]`.
*/
import * as fs from 'node:fs/promises';
import * as path from 'node:path';
import {
aggregateDiffs,
diffResolutions,
type Resolution,
type ShadowCallsite,
type ShadowDiff,
type ShadowParityReport,
type SupportedLanguages,
} from 'gitnexus-shared';
// ─── Public API ────────────────────────────────────────────────────────────
/** Which side of the dual-run is considered authoritative for this language. */
export type PrimarySide = 'legacy' | 'registry';
/** One record per call site the caller dual-runs. */
export interface ShadowRecordInput {
readonly language: SupportedLanguages;
readonly callsite: ShadowCallsite;
readonly legacy: readonly Resolution[];
readonly newResult: readonly Resolution[];
/**
* Which side drove the actual runtime answer for this record. Lets the
* dashboard distinguish "registry-primary, legacy is shadow" from the
* default "legacy-primary, registry is shadow" without re-reading
* `REGISTRY_PRIMARY_<LANG>` env vars at render time.
*/
readonly primary: PrimarySide;
}
/** Persisted JSON shape. Schema-versioned for future migrations. */
export interface PersistedShadowReport {
readonly schemaVersion: 1;
readonly runId: string;
readonly generatedAt: string;
readonly primaryByLanguage: Readonly<Partial<Record<SupportedLanguages, PrimarySide>>>;
readonly report: ShadowParityReport;
}
export interface ShadowHarness {
/** `true` iff `GITNEXUS_SHADOW_MODE` is truthy. When `false`, `record()` is a no-op. */
readonly enabled: boolean;
/** Accumulate a dual-run observation. No-op when `enabled === false`. */
record(input: ShadowRecordInput): void;
/** Number of records accumulated so far. Useful for diagnostics / tests. */
size(): number;
/**
* Aggregate the accumulated records into a `ShadowParityReport`
* without persisting. Returns a deterministic snapshot each call;
* idempotent with respect to `record()` ordering.
*/
snapshot(now?: Date): ShadowParityReport;
/**
* Write the aggregated snapshot to JSON. Resolves to the path of the
* per-run file. Also writes/overwrites `latest.json` alongside.
*
* Creates `outputDir` if it doesn't exist.
*/
persist(outputDir: string, now?: Date): Promise<string>;
/** Reset the accumulator. Preserves `enabled`. */
clear(): void;
}
/**
* Construct a harness. Reads `GITNEXUS_SHADOW_MODE` at construction time
* (not per-`record()` call) so repeated no-op records don't re-check the
* env var in the hot path.
*/
export function createShadowHarness(): ShadowHarness {
const enabled = parseShadowModeEnv(process.env['GITNEXUS_SHADOW_MODE']);
interface Accumulated {
readonly language: SupportedLanguages;
readonly diff: ShadowDiff;
}
const records: Accumulated[] = [];
const primaryByLanguage: Partial<Record<SupportedLanguages, PrimarySide>> = {};
const recordImpl = (input: ShadowRecordInput): void => {
if (!enabled) return;
const diff = diffResolutions(input.callsite, input.legacy, input.newResult);
records.push({ language: input.language, diff });
// Primary per-language is resolved by last-write. In practice a run
// is single-threaded with respect to flag readings, so this is
// deterministic; a language's primary cannot change mid-run.
primaryByLanguage[input.language] = input.primary;
};
const snapshotImpl = (now: Date = new Date()): ShadowParityReport => {
return aggregateDiffs(records, now);
};
const persistImpl = async (outputDir: string, now: Date = new Date()): Promise<string> => {
await fs.mkdir(outputDir, { recursive: true });
const report = snapshotImpl(now);
const runId = makeRunId(now);
const payload: PersistedShadowReport = {
schemaVersion: 1,
runId,
generatedAt: now.toISOString(),
primaryByLanguage,
report,
};
const json = JSON.stringify(payload, null, 2);
const perRunPath = path.join(outputDir, `${runId}.json`);
const latestPath = path.join(outputDir, 'latest.json');
await fs.writeFile(perRunPath, json, 'utf8');
await fs.writeFile(latestPath, json, 'utf8');
return perRunPath;
};
const clearImpl = (): void => {
records.length = 0;
for (const key of Object.keys(primaryByLanguage)) {
delete primaryByLanguage[key as SupportedLanguages];
}
};
return {
enabled,
record: recordImpl,
size: () => records.length,
snapshot: snapshotImpl,
persist: persistImpl,
clear: clearImpl,
};
}
// ─── Internal helpers ─────────────────────────────────────────────────────
/**
* Env-var parser for `GITNEXUS_SHADOW_MODE`. Accepts the same truthy
* conventions as `REGISTRY_PRIMARY_<LANG>` from #924: `'true'` / `'1'` /
* `'yes'`, case-insensitive, whitespace-trimmed. Anything else — including
* `undefined`, `''`, `'false'`, `'off'`, typos — is false.
*/
function parseShadowModeEnv(raw: string | undefined): boolean {
if (raw === undefined) return false;
const normalized = raw.trim().toLowerCase();
return normalized === 'true' || normalized === '1' || normalized === 'yes';
}
/**
* Deterministic run id derived from the timestamp plus 4 random bytes
* of entropy. The timestamp comes first so files sort chronologically;
* the entropy suffix prevents collisions when multiple runs share a
* clock-second. Shape: `YYYYMMDD-HHMMSS-xxxxxxxx`.
*/
function makeRunId(now: Date): string {
const y = now.getUTCFullYear().toString().padStart(4, '0');
const m = (now.getUTCMonth() + 1).toString().padStart(2, '0');
const d = now.getUTCDate().toString().padStart(2, '0');
const h = now.getUTCHours().toString().padStart(2, '0');
const min = now.getUTCMinutes().toString().padStart(2, '0');
const s = now.getUTCSeconds().toString().padStart(2, '0');
const entropy = Math.floor(Math.random() * 0xffffffff)
.toString(16)
.padStart(8, '0');
return `${y}${m}${d}-${h}${min}${s}-${entropy}`;
}
@@ -77,6 +77,8 @@ import {
buildCollisionGroups,
} from '../utils/method-props.js';
import type { LanguageProvider } from '../language-provider.js';
import type { ParsedFile } from 'gitnexus-shared';
import { extractParsedFile } from '../scope-extractor-bridge.js';
// ============================================================================
// Types for serializable results
@@ -269,6 +271,14 @@ export interface ParseWorkerResult {
constructorBindings: FileConstructorBindings[];
/** All-scope type bindings from TypeEnv for BindingAccumulator (includes function-local). */
fileScopeBindings: FileScopeBindings[];
/**
* Per-file `ParsedFile` artifacts from the new scope-based resolution
* pipeline (RFC #909 Ring 2). Empty unless the file's provider implements
* `emitScopeCaptures` — default for every language today, so this is
* additive and leaves the legacy DAG untouched. Consumed by #921's
* finalize-orchestrator.
*/
parsedFiles: ParsedFile[];
skippedLanguages: Record<string, number>;
fileCount: number;
}
@@ -711,6 +721,7 @@ const processBatch = (
ormQueries: [],
constructorBindings: [],
fileScopeBindings: [],
parsedFiles: [],
skippedLanguages: {},
fileCount: 0,
};
@@ -1396,11 +1407,24 @@ const processFileGroup = (
continue;
}
const provider = getProvider(language);
// RFC #909 Ring 2: produce a `ParsedFile` for the new scope-based
// resolution pipeline. No-op (returns undefined) for every language
// today — only fires once a provider implements `emitScopeCaptures`.
// Runs BEFORE legacy extraction and its result is independent: a
// failure here is caught inside `extractParsedFile` and does NOT
// affect the legacy DAG path that follows.
const parsedFile = extractParsedFile(provider, parseContent, file.path, (message) => {
if (parentPort) parentPort.postMessage({ type: 'warning', message });
else console.warn(message);
});
if (parsedFile !== undefined) result.parsedFiles.push(parsedFile);
// Pre-pass: extract heritage from query matches to build parentMap for buildTypeEnv.
// Heritage edges (EXTENDS/IMPLEMENTS) are created by heritage-processor which runs
// in PARALLEL with call-processor, so the graph edges don't exist when buildTypeEnv
// runs. This pre-pass makes parent class information available for type resolution.
const provider = getProvider(language);
const fileParentMap = new Map<string, string[]>();
if (provider.heritageExtractor) {
for (const match of matches) {
@@ -2282,6 +2306,7 @@ let accumulated: ParseWorkerResult = {
ormQueries: [],
constructorBindings: [],
fileScopeBindings: [],
parsedFiles: [],
skippedLanguages: {},
fileCount: 0,
};
@@ -2309,6 +2334,7 @@ const mergeResult = (target: ParseWorkerResult, src: ParseWorkerResult) => {
appendAll(target.ormQueries, src.ormQueries);
appendAll(target.constructorBindings, src.constructorBindings);
appendAll(target.fileScopeBindings, src.fileScopeBindings);
appendAll(target.parsedFiles, src.parsedFiles);
for (const [lang, count] of Object.entries(src.skippedLanguages)) {
target.skippedLanguages[lang] = (target.skippedLanguages[lang] || 0) + count;
}
@@ -2360,6 +2386,7 @@ parentPort!.on('message', (msg: WorkerIncomingMessage) => {
ormQueries: [],
constructorBindings: [],
fileScopeBindings: [],
parsedFiles: [],
skippedLanguages: {},
fileCount: 0,
};
@@ -0,0 +1,495 @@
/**
* Unit tests for `emit-references` (RFC #909 Ring 2 PKG #925).
*
* Covers kind → RelationshipType mapping, enclosing-def resolution
* through the scope tree, evidence serialization onto emitted edges,
* skip counts, and the optional `INGESTION_EMIT_SCOPES` scope-node
* flush.
*/
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import {
buildDefIndex,
buildMethodDispatchIndex,
buildModuleScopeIndex,
buildQualifiedNameIndex,
buildScopeTree,
type BindingRef,
type DefId,
type Range,
type Reference,
type ReferenceIndex,
type Scope,
type ScopeId,
type SymbolDefinition,
} from 'gitnexus-shared';
import { createKnowledgeGraph } from '../../../src/core/graph/graph.js';
import {
emitReferencesToGraph,
emitScopeGraph,
} from '../../../src/core/ingestion/emit-references.js';
import type { ScopeResolutionIndexes } from '../../../src/core/ingestion/model/scope-resolution-indexes.js';
// ─── Env isolation ────────────────────────────────────────────────────────
let savedEnv: string | undefined;
beforeEach(() => {
savedEnv = process.env['INGESTION_EMIT_SCOPES'];
delete process.env['INGESTION_EMIT_SCOPES'];
});
afterEach(() => {
if (savedEnv === undefined) delete process.env['INGESTION_EMIT_SCOPES'];
else process.env['INGESTION_EMIT_SCOPES'] = savedEnv;
});
// ─── Fixture builders ─────────────────────────────────────────────────────
const range = (sl = 1, sc = 0, el = 100, ec = 0): Range => ({
startLine: sl,
startCol: sc,
endLine: el,
endCol: ec,
});
const def = (
nodeId: string,
type: SymbolDefinition['type'] = 'Method',
qname?: string,
): SymbolDefinition => ({
nodeId,
filePath: 'x.ts',
type,
...(qname !== undefined ? { qualifiedName: qname } : {}),
});
const scope = (
id: ScopeId,
parent: ScopeId | null,
kind: Scope['kind'],
ownedDefs: readonly SymbolDefinition[] = [],
r: Range = range(),
filePath = 'x.ts',
bindings: Record<string, readonly BindingRef[]> = {},
): Scope => ({
id,
parent,
kind,
range: r,
filePath,
bindings: new Map(Object.entries(bindings)),
ownedDefs,
imports: [],
typeBindings: new Map(),
});
function makeIndexes(scopes: Scope[], allDefs: SymbolDefinition[]): ScopeResolutionIndexes {
return {
scopeTree: buildScopeTree(scopes),
defs: buildDefIndex(allDefs),
qualifiedNames: buildQualifiedNameIndex(allDefs),
moduleScopes: buildModuleScopeIndex(
scopes
.filter((s) => s.kind === 'Module')
.map((s) => ({ filePath: s.filePath, moduleScopeId: s.id })),
),
methodDispatch: buildMethodDispatchIndex({
owners: [],
computeMro: () => [],
implementsOf: () => [],
}),
imports: new Map(),
bindings: new Map(),
referenceSites: [],
sccs: [],
stats: {
totalFiles: 0,
totalEdges: 0,
linkedEdges: 0,
unresolvedEdges: 0,
sccCount: 0,
largestSccSize: 0,
},
};
}
function buildRefIndex(sourceScope: ScopeId, refs: readonly Reference[]): ReferenceIndex {
const bySource = new Map<ScopeId, readonly Reference[]>();
bySource.set(sourceScope, refs);
const byTarget = new Map<DefId, Reference[]>();
for (const ref of refs) {
const bucket = byTarget.get(ref.toDef) ?? [];
bucket.push(ref);
byTarget.set(ref.toDef, bucket);
}
return {
bySourceScope: bySource,
byTargetDef: new Map(
Array.from(byTarget.entries()).map(([k, v]) => [k, Object.freeze([...v])]),
),
};
}
// ─── Kind mapping + basic emission ────────────────────────────────────────
describe('emitReferencesToGraph: kind mapping', () => {
it('maps call → CALLS and carries confidence + evidence onto the edge', () => {
const callerFn = def('def:saveUser', 'Function', 'saveUser');
const targetFn = def('def:User.save', 'Method', 'User.save');
const mod = scope('scope:m', null, 'Module', [callerFn, targetFn]);
const indexes = makeIndexes([mod], [callerFn, targetFn]);
const ref: Reference = {
fromScope: 'scope:m',
toDef: 'def:User.save',
atRange: range(10, 4, 10, 8),
kind: 'call',
confidence: 0.75,
evidence: [
{ kind: 'local', weight: 0.55 },
{ kind: 'arity-match', weight: 0.1, note: 'compatible' },
],
};
const graph = createKnowledgeGraph();
const stats = emitReferencesToGraph({
graph,
scopes: indexes,
referenceIndex: buildRefIndex('scope:m', [ref]),
});
expect(stats.edgesEmitted).toBe(1);
expect(graph.relationships).toHaveLength(1);
const edge = graph.relationships[0]!;
expect(edge.type).toBe('CALLS');
expect(edge.sourceId).toBe('def:saveUser');
expect(edge.targetId).toBe('def:User.save');
expect(edge.confidence).toBe(0.75);
expect(edge.evidence).toEqual([
{ kind: 'local', weight: 0.55 },
{ kind: 'arity-match', weight: 0.1, note: 'compatible' },
]);
expect(edge.reason).toContain('call');
expect(edge.reason).toContain('0.750');
});
it('maps read / write → ACCESSES and stamps step=1 / step=2 for discrimination', () => {
const fn = def('def:render', 'Function');
const field = def('def:User.name', 'Property');
const mod = scope('scope:m', null, 'Module', [fn, field]);
const indexes = makeIndexes([mod], [fn, field]);
const readRef: Reference = {
fromScope: 'scope:m',
toDef: 'def:User.name',
atRange: range(5, 0, 5, 4),
kind: 'read',
confidence: 0.55,
evidence: [{ kind: 'local', weight: 0.55 }],
};
const writeRef: Reference = {
fromScope: 'scope:m',
toDef: 'def:User.name',
atRange: range(6, 0, 6, 4),
kind: 'write',
confidence: 0.55,
evidence: [{ kind: 'local', weight: 0.55 }],
};
const graph = createKnowledgeGraph();
emitReferencesToGraph({
graph,
scopes: indexes,
referenceIndex: buildRefIndex('scope:m', [readRef, writeRef]),
});
const edges = graph.relationships;
expect(edges).toHaveLength(2);
expect(edges.every((e) => e.type === 'ACCESSES')).toBe(true);
const readEdge = edges.find((e) => e.step === 1)!;
const writeEdge = edges.find((e) => e.step === 2)!;
expect(readEdge).toBeDefined();
expect(writeEdge).toBeDefined();
});
it('maps inherits → INHERITS and type-reference/import-use → USES', () => {
const hostFn = def('def:host', 'Function');
const base = def('def:Base', 'Class');
const mixin = def('def:Mixin', 'Class');
const module = def('def:SomeModule', 'Namespace');
const mod = scope('scope:m', null, 'Module', [hostFn, base, mixin, module]);
const indexes = makeIndexes([mod], [hostFn, base, mixin, module]);
const refs: Reference[] = [
{
fromScope: 'scope:m',
toDef: 'def:Base',
atRange: range(1, 0, 1, 4),
kind: 'inherits',
confidence: 0.9,
evidence: [],
},
{
fromScope: 'scope:m',
toDef: 'def:Mixin',
atRange: range(2, 0, 2, 4),
kind: 'type-reference',
confidence: 0.7,
evidence: [],
},
{
fromScope: 'scope:m',
toDef: 'def:SomeModule',
atRange: range(3, 0, 3, 4),
kind: 'import-use',
confidence: 0.5,
evidence: [],
},
];
const graph = createKnowledgeGraph();
emitReferencesToGraph({
graph,
scopes: indexes,
referenceIndex: buildRefIndex('scope:m', refs),
});
const types = graph.relationships.map((r) => r.type).sort();
expect(types).toEqual(['INHERITS', 'USES', 'USES']);
});
});
// ─── Enclosing-def resolution ─────────────────────────────────────────────
describe('enclosing-def resolution', () => {
it('uses the innermost Function/Method ancestor as the caller', () => {
const method = def('def:User.save', 'Method');
const classScope = scope('scope:c', 'scope:m', 'Class', [], range(5, 0, 40, 0));
const methodScope = scope('scope:f', 'scope:c', 'Function', [method], range(10, 0, 30, 0));
const mod = scope('scope:m', null, 'Module', [], range(1, 0, 100, 0));
const target = def('def:Logger.log', 'Method');
const indexes = makeIndexes([mod, classScope, methodScope], [method, target]);
// Reference fires from a block inside the method scope.
const ref: Reference = {
fromScope: 'scope:f',
toDef: 'def:Logger.log',
atRange: range(20, 4, 20, 8),
kind: 'call',
confidence: 0.75,
evidence: [],
};
const graph = createKnowledgeGraph();
emitReferencesToGraph({
graph,
scopes: indexes,
referenceIndex: buildRefIndex('scope:f', [ref]),
});
expect(graph.relationships[0]!.sourceId).toBe('def:User.save');
});
it('walks up to the parent Class if the immediate scope has no Function def', () => {
const classDef = def('def:User', 'Class');
const targetFn = def('def:Logger.log', 'Method');
const mod = scope('scope:m', null, 'Module', [], range(1, 0, 100, 0));
// Class scope owns the Class def but no Function/Method; fallback
// walks into the Class's owned defs.
const classScope = scope('scope:c', 'scope:m', 'Class', [classDef], range(5, 0, 50, 0));
const indexes = makeIndexes([mod, classScope], [classDef, targetFn]);
const ref: Reference = {
fromScope: 'scope:c',
toDef: 'def:Logger.log',
atRange: range(7, 0, 7, 4),
kind: 'call',
confidence: 0.55,
evidence: [],
};
const graph = createKnowledgeGraph();
emitReferencesToGraph({
graph,
scopes: indexes,
referenceIndex: buildRefIndex('scope:c', [ref]),
});
// No Function ancestor — falls back to the first owned def: the class itself.
expect(graph.relationships[0]!.sourceId).toBe('def:User');
});
it('increments skippedNoCaller when no ancestor has any owned defs', () => {
// Module scope is empty; the lone child scope references something
// but neither it nor its ancestors own anything.
const target = def('def:someClass', 'Class');
const mod = scope('scope:m', null, 'Module', [], range(1, 0, 100, 0));
const child = scope('scope:c', 'scope:m', 'Function', [], range(5, 0, 10, 0));
const indexes = makeIndexes([mod, child], [target]);
const ref: Reference = {
fromScope: 'scope:c',
toDef: 'def:someClass',
atRange: range(7, 0, 7, 4),
kind: 'type-reference',
confidence: 0.3,
evidence: [],
};
const graph = createKnowledgeGraph();
const stats = emitReferencesToGraph({
graph,
scopes: indexes,
referenceIndex: buildRefIndex('scope:c', [ref]),
});
expect(stats.edgesEmitted).toBe(0);
expect(stats.skippedNoCaller).toBe(1);
expect(graph.relationships).toHaveLength(0);
});
});
// ─── Missing target ──────────────────────────────────────────────────────
describe('missing target', () => {
it('skips references whose toDef is not in the DefIndex', () => {
const callerFn = def('def:caller', 'Function');
const mod = scope('scope:m', null, 'Module', [callerFn]);
const indexes = makeIndexes([mod], [callerFn]); // target def missing
const ref: Reference = {
fromScope: 'scope:m',
toDef: 'def:ghost',
atRange: range(5, 0, 5, 4),
kind: 'call',
confidence: 0.3,
evidence: [],
};
const graph = createKnowledgeGraph();
const stats = emitReferencesToGraph({
graph,
scopes: indexes,
referenceIndex: buildRefIndex('scope:m', [ref]),
});
expect(stats.edgesEmitted).toBe(0);
expect(stats.skippedMissingTarget).toBe(1);
expect(graph.relationships).toHaveLength(0);
});
});
// ─── Scope-graph emission (INGESTION_EMIT_SCOPES) ─────────────────────────
describe('scope-graph emission', () => {
it('stays off by default — no scope nodes emitted', () => {
const callerFn = def('def:caller', 'Function');
const targetFn = def('def:target', 'Method');
const mod = scope('scope:m', null, 'Module', [callerFn, targetFn]);
const indexes = makeIndexes([mod], [callerFn, targetFn]);
const ref: Reference = {
fromScope: 'scope:m',
toDef: 'def:target',
atRange: range(5, 0, 5, 4),
kind: 'call',
confidence: 0.5,
evidence: [],
};
const graph = createKnowledgeGraph();
const stats = emitReferencesToGraph({
graph,
scopes: indexes,
referenceIndex: buildRefIndex('scope:m', [ref]),
});
expect(stats.scopeNodesEmitted).toBe(0);
expect(stats.scopeEdgesEmitted).toBe(0);
// No scope nodes in the graph either.
expect(graph.nodes.filter((n) => n.id.startsWith('scope:')).length).toBe(0);
});
it('emits Scope nodes + CONTAINS + DEFINES when INGESTION_EMIT_SCOPES=1', () => {
process.env['INGESTION_EMIT_SCOPES'] = '1';
const fn = def('def:fn', 'Function');
const childScope = scope('scope:f', 'scope:m', 'Function', [fn], range(5, 0, 10, 0));
const mod = scope('scope:m', null, 'Module', [], range(1, 0, 100, 0));
const indexes = makeIndexes([mod, childScope], [fn]);
const graph = createKnowledgeGraph();
const stats = emitReferencesToGraph({
graph,
scopes: indexes,
referenceIndex: buildRefIndex('scope:f', []),
});
expect(stats.scopeNodesEmitted).toBe(2); // module + function scope
// 1 CONTAINS (module→function) + 1 DEFINES (function→fn def) = 2
expect(stats.scopeEdgesEmitted).toBe(2);
const containsEdge = graph.relationships.find((e) => e.type === 'CONTAINS');
const definesEdge = graph.relationships.find((e) => e.type === 'DEFINES');
expect(containsEdge).toBeDefined();
expect(containsEdge!.sourceId).toBe('scope:m');
expect(containsEdge!.targetId).toBe('scope:f');
expect(definesEdge).toBeDefined();
expect(definesEdge!.targetId).toBe('def:fn');
});
it("treats 'true', 'yes' (case-insensitive) as enabled; anything else as disabled", () => {
const fn = def('def:fn', 'Function');
const mod = scope('scope:m', null, 'Module', [fn]);
const indexes = makeIndexes([mod], [fn]);
for (const value of ['true', 'TRUE', 'yes', '1']) {
process.env['INGESTION_EMIT_SCOPES'] = value;
const g = createKnowledgeGraph();
const stats = emitReferencesToGraph({
graph: g,
scopes: indexes,
referenceIndex: buildRefIndex('scope:m', []),
});
expect(stats.scopeNodesEmitted).toBeGreaterThan(0);
}
for (const value of ['false', '0', '', 'off', 'tru']) {
process.env['INGESTION_EMIT_SCOPES'] = value;
const g = createKnowledgeGraph();
const stats = emitReferencesToGraph({
graph: g,
scopes: indexes,
referenceIndex: buildRefIndex('scope:m', []),
});
expect(stats.scopeNodesEmitted).toBe(0);
}
});
it('emitScopeGraph can be called directly (bypasses env flag)', () => {
const fn = def('def:fn', 'Function');
const mod = scope('scope:m', null, 'Module', [fn]);
const indexes = makeIndexes([mod], [fn]);
const graph = createKnowledgeGraph();
const stats = emitScopeGraph({ graph, scopes: indexes });
expect(stats.scopeNodesEmitted).toBe(1);
expect(stats.scopeEdgesEmitted).toBe(1); // only the DEFINES edge; no parent scope
});
});
// ─── Empty input ──────────────────────────────────────────────────────────
describe('empty input', () => {
it('returns zeroed stats and mutates nothing when ReferenceIndex is empty', () => {
const mod = scope('scope:m', null, 'Module', []);
const indexes = makeIndexes([mod], []);
const graph = createKnowledgeGraph();
const stats = emitReferencesToGraph({
graph,
scopes: indexes,
referenceIndex: { bySourceScope: new Map(), byTargetDef: new Map() },
});
expect(stats).toEqual({
edgesEmitted: 0,
skippedNoCaller: 0,
skippedMissingTarget: 0,
scopeNodesEmitted: 0,
scopeEdgesEmitted: 0,
});
expect(graph.nodes).toHaveLength(0);
expect(graph.relationships).toHaveLength(0);
});
});
@@ -0,0 +1,219 @@
/**
* Unit tests for `finalize-orchestrator` (RFC #909 Ring 2 PKG #921).
*
* Covers empty-input, single-file, multi-file-with-imports, and the
* `MutableSemanticModel.attachScopeIndexes` one-shot contract.
*
* Builds synthetic `ParsedFile` inputs directly — the orchestrator is
* below the extraction layer and independent of tree-sitter, so the
* tests don't need a real parser.
*/
import { describe, it, expect } from 'vitest';
import type {
BindingRef,
ParsedFile,
ParsedImport,
Scope,
ScopeId,
SymbolDefinition,
} from 'gitnexus-shared';
import { finalizeScopeModel } from '../../../src/core/ingestion/finalize-orchestrator.js';
import { createSemanticModel } from '../../../src/core/ingestion/model/semantic-model.js';
import type { ScopeResolutionIndexes } from '../../../src/core/ingestion/model/scope-resolution-indexes.js';
// ─── Fixture helpers ────────────────────────────────────────────────────────
const mkScope = (
id: ScopeId,
parent: ScopeId | null,
filePath: string,
bindings: Record<string, readonly BindingRef[]> = {},
): Scope => ({
id,
parent,
kind: parent === null ? 'Module' : 'Class',
range: { startLine: 1, startCol: 0, endLine: 100, endCol: 0 },
filePath,
bindings: new Map(Object.entries(bindings)),
ownedDefs: [],
imports: [],
typeBindings: new Map(),
});
const mkFile = (filePath: string, overrides: Partial<ParsedFile> = {}): ParsedFile => ({
filePath,
moduleScope: `scope:${filePath}#module`,
scopes: overrides.scopes ?? [mkScope(`scope:${filePath}#module`, null, filePath)],
parsedImports: overrides.parsedImports ?? [],
localDefs: overrides.localDefs ?? [],
referenceSites: overrides.referenceSites ?? [],
});
const mkDef = (nodeId: string, filePath: string, qname: string): SymbolDefinition => ({
nodeId,
filePath,
type: 'Class',
qualifiedName: qname,
});
// ─── Empty input ───────────────────────────────────────────────────────────
describe('finalizeScopeModel: empty input', () => {
it('produces a valid but empty bundle for zero parsedFiles', () => {
const out = finalizeScopeModel([]);
expect(out.scopeTree.size).toBe(0);
expect(out.defs.size).toBe(0);
expect(out.qualifiedNames.size).toBe(0);
expect(out.moduleScopes.size).toBe(0);
expect(out.methodDispatch.mroByOwnerDefId.size).toBe(0);
expect(out.imports.size).toBe(0);
expect(out.bindings.size).toBe(0);
expect(out.referenceSites).toEqual([]);
expect(out.sccs).toEqual([]);
expect(out.stats.totalFiles).toBe(0);
expect(out.stats.totalEdges).toBe(0);
});
});
// ─── Single file ───────────────────────────────────────────────────────────
describe('finalizeScopeModel: single file', () => {
it('builds all per-file indexes from a single ParsedFile', () => {
const userClass = mkDef('def:User', 'models.ts', 'models.User');
const file = mkFile('models.ts', {
localDefs: [userClass],
});
const out = finalizeScopeModel([file]);
expect(out.scopeTree.size).toBe(1);
expect(out.defs.get('def:User')).toBe(userClass);
expect(out.qualifiedNames.get('models.User')).toEqual(['def:User']);
expect(out.moduleScopes.get('models.ts')).toBe(file.moduleScope);
expect(out.stats.totalFiles).toBe(1);
});
it('forwards per-file referenceSites into the aggregated list', () => {
const file = mkFile('a.ts', {
referenceSites: [
{
name: 'save',
atRange: { startLine: 5, startCol: 0, endLine: 5, endCol: 4 },
inScope: 'scope:a.ts#module',
kind: 'call',
},
],
});
const out = finalizeScopeModel([file]);
expect(out.referenceSites).toHaveLength(1);
expect(out.referenceSites[0]!.name).toBe('save');
});
});
// ─── Multi-file with cross-file imports ────────────────────────────────────
describe('finalizeScopeModel: cross-file imports', () => {
it('links a named import when the caller provides resolveImportTarget', () => {
const userClass = mkDef('def:User', 'models.ts', 'models.User');
const modelsFile = mkFile('models.ts', { localDefs: [userClass] });
const importOfUser: ParsedImport = {
kind: 'named',
localName: 'User',
importedName: 'User',
targetRaw: 'models.ts',
};
const appFile = mkFile('app.ts', { parsedImports: [importOfUser] });
const out = finalizeScopeModel([appFile, modelsFile], {
hooks: {
resolveImportTarget: (targetRaw) => (targetRaw === 'models.ts' ? 'models.ts' : null),
},
});
const appImports = out.imports.get(appFile.moduleScope) ?? [];
expect(appImports).toHaveLength(1);
expect(appImports[0]!.linkStatus).toBeUndefined();
expect(appImports[0]!.targetFile).toBe('models.ts');
expect(appImports[0]!.targetDefId).toBe('def:User');
});
it('leaves imports unresolved when no resolveImportTarget is supplied (default hook)', () => {
// Default `resolveImportTarget: () => null` — every import ends up
// with `linkStatus: 'unresolved'`. This is the zero-provider case
// today; behavior is well-defined, not a crash.
const importOfUser: ParsedImport = {
kind: 'named',
localName: 'User',
importedName: 'User',
targetRaw: 'models.ts',
};
const appFile = mkFile('app.ts', { parsedImports: [importOfUser] });
const out = finalizeScopeModel([appFile]);
const appImports = out.imports.get(appFile.moduleScope) ?? [];
expect(appImports).toHaveLength(1);
expect(appImports[0]!.linkStatus).toBe('unresolved');
});
it('surfaces FinalizeStats for observability', () => {
const userClass = mkDef('def:User', 'models.ts', 'models.User');
const modelsFile = mkFile('models.ts', { localDefs: [userClass] });
const appFile = mkFile('app.ts', {
parsedImports: [
{
kind: 'named',
localName: 'User',
importedName: 'User',
targetRaw: 'models.ts',
},
],
});
const out = finalizeScopeModel([appFile, modelsFile], {
hooks: { resolveImportTarget: () => 'models.ts' },
});
expect(out.stats.totalFiles).toBe(2);
expect(out.stats.totalEdges).toBe(1);
expect(out.stats.linkedEdges).toBe(1);
expect(out.stats.unresolvedEdges).toBe(0);
});
});
// ─── Integration with MutableSemanticModel ─────────────────────────────────
describe('MutableSemanticModel.attachScopeIndexes', () => {
it('starts as undefined and accepts a one-shot attach', () => {
const model = createSemanticModel();
expect(model.scopes).toBeUndefined();
const indexes = finalizeScopeModel([]);
model.attachScopeIndexes(indexes);
expect(model.scopes).toBe(indexes);
expect(model.scopes!.stats.totalFiles).toBe(0);
});
it('freezes the attached bundle (callers cannot mutate after attach)', () => {
const model = createSemanticModel();
const indexes: ScopeResolutionIndexes = finalizeScopeModel([]);
model.attachScopeIndexes(indexes);
expect(Object.isFrozen(model.scopes)).toBe(true);
});
it('throws on a second attach without clear()', () => {
const model = createSemanticModel();
model.attachScopeIndexes(finalizeScopeModel([]));
expect(() => model.attachScopeIndexes(finalizeScopeModel([]))).toThrowError(/already attached/);
});
it('clear() resets the bundle, enabling re-attach', () => {
const model = createSemanticModel();
model.attachScopeIndexes(finalizeScopeModel([]));
model.clear();
expect(model.scopes).toBeUndefined();
// Second attach now succeeds.
model.attachScopeIndexes(finalizeScopeModel([]));
expect(model.scopes).toBeDefined();
});
});
@@ -0,0 +1,152 @@
/**
* Unit tests for `import-target-adapter` (RFC #909 Ring 2 PKG #922).
*
* Exercises the language-dispatching FinalizeHook. We don't need the
* real per-language resolvers here — mock `ImportResolverFn`s let each
* branch be tested in isolation. Real-resolver integration is covered
* by the existing per-language import-resolver test suites.
*/
import { describe, it, expect } from 'vitest';
import { SupportedLanguages } from 'gitnexus-shared';
import {
buildImportTargetWorkspace,
resolveImportTargetAcrossLanguages,
type ImportTargetWorkspace,
} from '../../../src/core/ingestion/import-target-adapter.js';
import type {
ImportResolverFn,
ResolveCtx,
} from '../../../src/core/ingestion/import-resolvers/types.js';
import type { LanguageProvider } from '../../../src/core/ingestion/language-provider.js';
// ─── Helpers ───────────────────────────────────────────────────────────────
const emptyCtx: ResolveCtx = {
allFilePaths: new Set(),
allFileList: [],
normalizedFileList: [],
index: { bySuffix: new Map() } as unknown as ResolveCtx['index'],
resolveCache: new Map(),
configs: {
tsconfigPaths: null,
goModule: null,
composerConfig: null,
swiftPackageConfig: null,
csharpConfigs: [],
},
};
function fakeProvider(importResolver: ImportResolverFn | undefined): LanguageProvider {
return { importResolver } as unknown as LanguageProvider;
}
function workspace(
entries: Array<[SupportedLanguages, ImportResolverFn | undefined]>,
): ImportTargetWorkspace {
const providers = new Map<SupportedLanguages, LanguageProvider>();
for (const [lang, resolver] of entries) providers.set(lang, fakeProvider(resolver));
return buildImportTargetWorkspace(providers, emptyCtx);
}
// ─── buildImportTargetWorkspace ────────────────────────────────────────────
describe('buildImportTargetWorkspace', () => {
it('registers languages that expose an importResolver', () => {
const pyResolver: ImportResolverFn = () => ({ kind: 'files', files: ['resolved.py'] });
const ws = workspace([[SupportedLanguages.Python, pyResolver]]);
expect(ws.perLanguage.has(SupportedLanguages.Python)).toBe(true);
});
it("skips providers whose importResolver is absent (defensive — shouldn't happen in practice)", () => {
const ws = workspace([[SupportedLanguages.Python, undefined]]);
expect(ws.perLanguage.size).toBe(0);
});
it('threads the shared ResolveCtx into every entry', () => {
const pyResolver: ImportResolverFn = () => ({ kind: 'files', files: ['x.py'] });
const tsResolver: ImportResolverFn = () => ({ kind: 'files', files: ['x.ts'] });
const ws = workspace([
[SupportedLanguages.Python, pyResolver],
[SupportedLanguages.TypeScript, tsResolver],
]);
expect(ws.perLanguage.get(SupportedLanguages.Python)!.ctx).toBe(emptyCtx);
expect(ws.perLanguage.get(SupportedLanguages.TypeScript)!.ctx).toBe(emptyCtx);
});
});
// ─── resolveImportTargetAcrossLanguages ────────────────────────────────────
describe('resolveImportTargetAcrossLanguages', () => {
it('dispatches to the resolver for the fromFile extension', () => {
let seenPath: string | undefined;
const pyResolver: ImportResolverFn = (raw, _file) => {
seenPath = raw;
return { kind: 'files', files: ['models/user.py'] };
};
const ws = workspace([[SupportedLanguages.Python, pyResolver]]);
const result = resolveImportTargetAcrossLanguages('models.user', 'src/app.py', ws);
expect(seenPath).toBe('models.user');
expect(result).toBe('models/user.py');
});
it('routes to different resolvers based on the fromFile extension', () => {
const pyResolver: ImportResolverFn = () => ({ kind: 'files', files: ['resolved.py'] });
const tsResolver: ImportResolverFn = () => ({ kind: 'files', files: ['resolved.ts'] });
const ws = workspace([
[SupportedLanguages.Python, pyResolver],
[SupportedLanguages.TypeScript, tsResolver],
]);
expect(resolveImportTargetAcrossLanguages('x', 'a.py', ws)).toBe('resolved.py');
expect(resolveImportTargetAcrossLanguages('x', 'a.ts', ws)).toBe('resolved.ts');
});
it('returns null when the resolver returns null', () => {
const pyResolver: ImportResolverFn = () => null;
const ws = workspace([[SupportedLanguages.Python, pyResolver]]);
expect(resolveImportTargetAcrossLanguages('external_pkg', 'app.py', ws)).toBeNull();
});
it('takes the first file from a package-kind result', () => {
const resolver: ImportResolverFn = () => ({
kind: 'package',
files: ['pkg/index.py', 'pkg/other.py'],
dirSuffix: 'pkg',
});
const ws = workspace([[SupportedLanguages.Python, resolver]]);
expect(resolveImportTargetAcrossLanguages('pkg', 'app.py', ws)).toBe('pkg/index.py');
});
it('returns null when a result has kind=files but an empty files[]', () => {
// Defensive: resolvers shouldn't return this shape, but tolerate it.
const resolver: ImportResolverFn = () => ({ kind: 'files', files: [] });
const ws = workspace([[SupportedLanguages.Python, resolver]]);
expect(resolveImportTargetAcrossLanguages('x', 'a.py', ws)).toBeNull();
});
it('returns null when no resolver is registered for the language', () => {
const ws = workspace([]); // empty
// .py file but no Python resolver registered
expect(resolveImportTargetAcrossLanguages('x', 'a.py', ws)).toBeNull();
});
it('returns null when fromFile has an unknown extension', () => {
const pyResolver: ImportResolverFn = () => ({ kind: 'files', files: ['resolved.py'] });
const ws = workspace([[SupportedLanguages.Python, pyResolver]]);
expect(resolveImportTargetAcrossLanguages('x', 'README.xyz', ws)).toBeNull();
});
it('returns null when workspaceIndex is undefined / malformed', () => {
expect(resolveImportTargetAcrossLanguages('x', 'a.py', undefined)).toBeNull();
// Cast to exercise the runtime guard against caller misuse.
expect(resolveImportTargetAcrossLanguages('x', 'a.py', {} as unknown)).toBeNull();
});
it('swallows resolver exceptions and returns null (treated upstream as unresolved)', () => {
const throwingResolver: ImportResolverFn = () => {
throw new Error('resolver boom');
};
const ws = workspace([[SupportedLanguages.Python, throwingResolver]]);
expect(resolveImportTargetAcrossLanguages('x', 'a.py', ws)).toBeNull();
});
});
@@ -0,0 +1,166 @@
/**
* Unit tests for `extractParsedFile` — the parse-worker → ScopeExtractor
* bridge (RFC #909 Ring 2 PKG #920).
*
* The goal is to pin three invariants:
*
* 1. When a provider does NOT implement `emitScopeCaptures`, the helper
* returns `undefined` silently. This is the state of every language
* today — `ParseWorkerResult.parsedFiles` stays empty and the legacy
* DAG continues unaffected.
* 2. When a provider DOES implement the hook, the helper threads its
* output through `ScopeExtractor.extract` and returns a `ParsedFile`.
* 3. Exceptions from either the hook or the extractor are caught
* locally. The helper returns `undefined` — scope-extraction
* failures must NEVER break legacy parsing on the same file.
*/
import { describe, it, expect } from 'vitest';
import type { Capture, CaptureMatch } from 'gitnexus-shared';
import { extractParsedFile } from '../../../src/core/ingestion/scope-extractor-bridge.js';
import type { LanguageProvider } from '../../../src/core/ingestion/language-provider.js';
// ─── Capture helpers ────────────────────────────────────────────────────────
const cap = (
name: string,
startLine: number,
startCol: number,
endLine: number,
endCol: number,
text = '',
): Capture => ({ name, range: { startLine, startCol, endLine, endCol }, text });
const moduleScopeMatch = (): CaptureMatch => ({
'@scope.module': cap('@scope.module', 1, 0, 100, 0),
});
/**
* Build a `LanguageProvider` whose shape is only as narrow as
* `extractParsedFile` reads. Tests cast to the full provider type since
* `extractParsedFile` is typed against `LanguageProvider` (not the narrow
* `ScopeExtractorHooks`); the real worker always has a full provider.
*/
function fakeProvider(
hooks: Partial<
Pick<LanguageProvider, 'emitScopeCaptures' | 'shouldCreateScope' | 'resolveScopeKind'>
>,
): LanguageProvider {
return hooks as unknown as LanguageProvider;
}
// ─── Tests ─────────────────────────────────────────────────────────────────
describe('extractParsedFile', () => {
describe('provider has NOT migrated (no emitScopeCaptures)', () => {
it('returns undefined — silent no-op for legacy languages', () => {
const provider = fakeProvider({}); // no hook
const result = extractParsedFile(provider, 'source text', 'src/file.ts');
expect(result).toBeUndefined();
});
it('never calls the scope extractor when the hook is absent — cannot throw', () => {
// If the extractor was wrongly invoked, it would complain about the
// missing Module scope for empty captures. This test proves the
// short-circuit actually fires.
const provider = fakeProvider({});
expect(() => extractParsedFile(provider, '', 'x.ts')).not.toThrow();
});
});
describe('provider HAS migrated', () => {
it('threads emitScopeCaptures output through ScopeExtractor', () => {
const provider = fakeProvider({
emitScopeCaptures: () => [moduleScopeMatch()],
});
const result = extractParsedFile(provider, 'source text', 'src/file.ts');
expect(result).toBeDefined();
expect(result!.filePath).toBe('src/file.ts');
expect(result!.scopes).toHaveLength(1);
expect(result!.scopes[0]!.kind).toBe('Module');
});
it('forwards the correct arguments to emitScopeCaptures', () => {
let seenText: string | undefined;
let seenPath: string | undefined;
const provider = fakeProvider({
emitScopeCaptures: (text, path) => {
seenText = text;
seenPath = path;
return [moduleScopeMatch()];
},
});
extractParsedFile(provider, 'the real text', 'deep/path/file.ts');
expect(seenText).toBe('the real text');
expect(seenPath).toBe('deep/path/file.ts');
});
it('honors provider hooks beyond emitScopeCaptures (shouldCreateScope)', () => {
// A Block scope the provider declines to create — the resulting
// ParsedFile should have only the Module scope, not the Block.
const provider = fakeProvider({
emitScopeCaptures: () => [
moduleScopeMatch(),
{ '@scope.block': cap('@scope.block', 10, 0, 20, 0) },
],
shouldCreateScope: (match) => match['@scope.block'] === undefined,
});
const result = extractParsedFile(provider, 'src', 'a.ts');
expect(result!.scopes).toHaveLength(1);
expect(result!.scopes[0]!.kind).toBe('Module');
});
});
describe('error resilience — never breaks legacy parsing', () => {
it('returns undefined when emitScopeCaptures throws', () => {
const provider = fakeProvider({
emitScopeCaptures: () => {
throw new Error('provider boom');
},
});
const result = extractParsedFile(provider, 'src', 'a.ts');
expect(result).toBeUndefined();
});
it('routes errors through the onWarn callback when provided', () => {
const warnings: string[] = [];
const provider = fakeProvider({
emitScopeCaptures: () => {
throw new Error('provider boom');
},
});
const result = extractParsedFile(provider, 'src', 'path/to/file.ts', (msg) => {
warnings.push(msg);
});
expect(result).toBeUndefined();
expect(warnings).toHaveLength(1);
expect(warnings[0]).toContain('path/to/file.ts');
expect(warnings[0]).toContain('provider boom');
});
it('returns undefined when ScopeExtractor throws (missing Module scope)', () => {
// Emits a Class scope but no Module — extractor throws; helper
// swallows and returns undefined. Legacy parsing on the same file
// continues unaffected by this failure.
const provider = fakeProvider({
emitScopeCaptures: () => [{ '@scope.class': cap('@scope.class', 5, 0, 10, 0) }],
});
const result = extractParsedFile(provider, 'src', 'a.ts');
expect(result).toBeUndefined();
});
it('returns undefined when ScopeExtractor throws on malformed captures (overlap)', () => {
// Siblings with overlapping ranges trip the ScopeTreeInvariantError
// from #912. The helper catches it and returns undefined.
const provider = fakeProvider({
emitScopeCaptures: () => [
moduleScopeMatch(),
{ '@scope.function': cap('@scope.function', 10, 0, 20, 0) },
{ '@scope.function': cap('@scope.function', 15, 0, 25, 0) }, // overlap
],
});
const result = extractParsedFile(provider, 'src', 'a.ts');
expect(result).toBeUndefined();
});
});
});
@@ -0,0 +1,290 @@
/**
* Unit tests for `shadow-harness` (RFC #909 Ring 2 PKG #923).
*
* Covers flag detection, record accumulation, aggregation, and JSON
* persistence (real fs in a per-test tmpdir).
*/
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import * as fs from 'node:fs';
import * as fsp from 'node:fs/promises';
import * as os from 'node:os';
import * as path from 'node:path';
import {
EvidenceWeights,
SupportedLanguages,
type Resolution,
type ShadowCallsite,
type SymbolDefinition,
} from 'gitnexus-shared';
import {
createShadowHarness,
type PersistedShadowReport,
type ShadowHarness,
} from '../../../src/core/ingestion/shadow-harness.js';
// ─── Env isolation — GITNEXUS_SHADOW_MODE bleeds between tests otherwise ──
let savedEnv: string | undefined;
beforeEach(() => {
savedEnv = process.env['GITNEXUS_SHADOW_MODE'];
delete process.env['GITNEXUS_SHADOW_MODE'];
});
afterEach(() => {
if (savedEnv === undefined) delete process.env['GITNEXUS_SHADOW_MODE'];
else process.env['GITNEXUS_SHADOW_MODE'] = savedEnv;
});
// ─── Fixture helpers ──────────────────────────────────────────────────────
const callsite = (filePath = 'a.ts', line = 1): ShadowCallsite => ({
filePath,
range: { startLine: line, startCol: 0, endLine: line, endCol: 10 },
});
const def = (nodeId: string): SymbolDefinition => ({
nodeId,
filePath: 'x.ts',
type: 'Class',
});
const resolution = (nodeId: string): Resolution => ({
def: def(nodeId),
confidence: EvidenceWeights.local,
evidence: [{ kind: 'local', weight: EvidenceWeights.local }],
});
function enable(): void {
process.env['GITNEXUS_SHADOW_MODE'] = 'true';
}
function freshHarness(): ShadowHarness {
return createShadowHarness();
}
// ─── Flag detection ───────────────────────────────────────────────────────
describe('createShadowHarness: enabled flag', () => {
it('is disabled by default (no env var set)', () => {
expect(freshHarness().enabled).toBe(false);
});
it("is enabled when GITNEXUS_SHADOW_MODE is 'true' / '1' / 'yes' / case-insensitive", () => {
for (const value of ['true', '1', 'yes', 'TRUE', ' Yes ']) {
process.env['GITNEXUS_SHADOW_MODE'] = value;
expect(freshHarness().enabled).toBe(true);
}
});
it('stays disabled for falsy-looking or typo values', () => {
for (const value of ['', 'false', '0', 'off', 'tru']) {
process.env['GITNEXUS_SHADOW_MODE'] = value;
expect(freshHarness().enabled).toBe(false);
}
});
it('record() is a no-op when disabled', () => {
const h = freshHarness(); // disabled
h.record({
language: SupportedLanguages.Python,
callsite: callsite(),
legacy: [resolution('def:a')],
newResult: [resolution('def:a')],
primary: 'legacy',
});
expect(h.size()).toBe(0);
});
it('does NOT re-check the env var per call (constructed-once semantics)', () => {
const h = freshHarness(); // disabled at construction
process.env['GITNEXUS_SHADOW_MODE'] = 'true'; // flip AFTER construction
h.record({
language: SupportedLanguages.Python,
callsite: callsite(),
legacy: [resolution('def:a')],
newResult: [resolution('def:a')],
primary: 'legacy',
});
// Still disabled — the harness captured its `enabled` at construction.
expect(h.size()).toBe(0);
});
});
// ─── Record + snapshot ────────────────────────────────────────────────────
describe('record + snapshot', () => {
it('accumulates records across languages', () => {
enable();
const h = freshHarness();
h.record({
language: SupportedLanguages.Python,
callsite: callsite(),
legacy: [resolution('def:a')],
newResult: [resolution('def:a')],
primary: 'legacy',
});
h.record({
language: SupportedLanguages.TypeScript,
callsite: callsite('b.ts'),
legacy: [resolution('def:b')],
newResult: [],
primary: 'registry',
});
expect(h.size()).toBe(2);
});
it('snapshot reports per-language rows with correct outcomes', () => {
enable();
const h = freshHarness();
h.record({
language: SupportedLanguages.Python,
callsite: callsite(),
legacy: [resolution('def:a')],
newResult: [resolution('def:a')],
primary: 'legacy',
});
h.record({
language: SupportedLanguages.Python,
callsite: callsite('a.py', 2),
legacy: [resolution('def:b')],
newResult: [],
primary: 'legacy',
});
const report = h.snapshot(new Date('2026-04-18T00:00:00Z'));
expect(report.perLanguage).toHaveLength(1);
const py = report.perLanguage[0]!;
expect(py.language).toBe(SupportedLanguages.Python);
expect(py.totalCalls).toBe(2);
expect(py.bothAgree).toBe(1);
expect(py.onlyLegacy).toBe(1);
});
it('snapshot is deterministic across repeated calls', () => {
enable();
const h = freshHarness();
h.record({
language: SupportedLanguages.Python,
callsite: callsite(),
legacy: [resolution('def:a')],
newResult: [resolution('def:a')],
primary: 'legacy',
});
const now = new Date('2026-04-18T12:00:00Z');
const a = h.snapshot(now);
const b = h.snapshot(now);
expect(JSON.stringify(a)).toBe(JSON.stringify(b));
});
it('clear() resets the accumulator and primaryByLanguage', async () => {
enable();
const h = freshHarness();
h.record({
language: SupportedLanguages.Python,
callsite: callsite(),
legacy: [resolution('def:a')],
newResult: [resolution('def:a')],
primary: 'registry',
});
expect(h.size()).toBe(1);
h.clear();
expect(h.size()).toBe(0);
// Verify primary is also cleared: persist after a fresh record with a
// different primary should reflect the new value only.
h.record({
language: SupportedLanguages.Python,
callsite: callsite('a.py'),
legacy: [resolution('def:a')],
newResult: [resolution('def:a')],
primary: 'legacy',
});
const dir = await fsp.mkdtemp(path.join(os.tmpdir(), 'gn-sh-clear-'));
try {
await h.persist(dir);
const payload = JSON.parse(fs.readFileSync(path.join(dir, 'latest.json'), 'utf8'));
expect(payload.primaryByLanguage.python).toBe('legacy');
} finally {
await fsp.rm(dir, { recursive: true, force: true });
}
});
});
// ─── Persistence ──────────────────────────────────────────────────────────
describe('persist', () => {
let tmpDir: string;
beforeEach(async () => {
tmpDir = await fsp.mkdtemp(path.join(os.tmpdir(), 'gn-shadow-harness-'));
});
afterEach(async () => {
await fsp.rm(tmpDir, { recursive: true, force: true });
});
it('creates outputDir if it does not exist', async () => {
enable();
const h = freshHarness();
const nested = path.join(tmpDir, 'nested', 'a', 'b');
await h.persist(nested);
expect(fs.existsSync(nested)).toBe(true);
expect(fs.existsSync(path.join(nested, 'latest.json'))).toBe(true);
});
it('writes BOTH a timestamped file and latest.json with the same payload', async () => {
enable();
const h = freshHarness();
h.record({
language: SupportedLanguages.Python,
callsite: callsite(),
legacy: [resolution('def:a')],
newResult: [resolution('def:a')],
primary: 'legacy',
});
const perRunPath = await h.persist(tmpDir, new Date('2026-04-18T12:34:56Z'));
const latestPath = path.join(tmpDir, 'latest.json');
expect(fs.existsSync(perRunPath)).toBe(true);
expect(fs.existsSync(latestPath)).toBe(true);
expect(fs.readFileSync(perRunPath, 'utf8')).toBe(fs.readFileSync(latestPath, 'utf8'));
});
it('persisted payload matches the schema v1 shape', async () => {
enable();
const h = freshHarness();
h.record({
language: SupportedLanguages.TypeScript,
callsite: callsite('a.ts'),
legacy: [resolution('def:a')],
newResult: [resolution('def:a')],
primary: 'registry',
});
const now = new Date('2026-04-18T00:00:00Z');
await h.persist(tmpDir, now);
const payload: PersistedShadowReport = JSON.parse(
fs.readFileSync(path.join(tmpDir, 'latest.json'), 'utf8'),
);
expect(payload.schemaVersion).toBe(1);
expect(payload.runId).toMatch(/^\d{8}-\d{6}-[0-9a-f]{8}$/);
expect(payload.generatedAt).toBe('2026-04-18T00:00:00.000Z');
expect(payload.primaryByLanguage.typescript).toBe('registry');
expect(payload.report.overall.totalCalls).toBe(1);
expect(payload.report.overall.bothAgree).toBe(1);
});
it('runId embeds the run timestamp for chronological sorting', async () => {
enable();
const h1 = freshHarness();
const h2 = freshHarness();
const p1 = await h1.persist(tmpDir, new Date('2026-04-18T00:00:00Z'));
const p2 = await h2.persist(tmpDir, new Date('2026-04-18T01:00:00Z'));
// Timestamp prefix means the second file sorts after the first.
expect(path.basename(p2) > path.basename(p1)).toBe(true);
});
it('persists an empty report gracefully (no records, no error)', async () => {
enable();
const h = freshHarness();
await h.persist(tmpDir);
const payload = JSON.parse(fs.readFileSync(path.join(tmpDir, 'latest.json'), 'utf8'));
expect(payload.report.overall.totalCalls).toBe(0);
expect(payload.report.perLanguage).toEqual([]);
});
});