Compare commits

...
Author SHA1 Message Date
Copilot be3833d9c9 chore(security): upgrade @vercel/node in gitnexus-web and remediate transitive advisories (#1705) 2026-05-22 06:26:09 +01:00
dependabot[bot] dc96bb048a chore(deps)(deps-dev): bump tsx from 4.21.1 to 4.22.0 in /gitnexus (#1768) 2026-05-22 05:27:51 +01:00
Abhigyan Patwari 5a0f5e81db ci(web): use npm ci for deterministic Vercel installs (#1764) 2026-05-22 05:08:42 +01:00
dependabot[bot] 8c1983a8bf chore(deps)(deps-dev): bump @types/node in /gitnexus (#1767) 2026-05-22 04:41:57 +01:00
231ad71d40 fix(mcp): disambiguate duplicate-name repo resolution for worktrees (#1753)
* fix(mcp): disambiguate duplicate-name repo resolution for worktrees

When multiple indexed repos share the same registry name (main checkout plus linked worktrees), MCP tools no longer silently pick the first sibling. Resolution prefers the repo matching process.cwd()'s git root, throws RegistryAmbiguousTargetError when still ambiguous, and uses canonical path matching aligned with the CLI registry.

Fixes #1658. Complements worktree detect_changes fixes in #1654/#1691.

* fix(mcp): refresh registry on duplicate-name ambiguity before failing

resolveRepo now retries resolveRepoFromCache after RegistryAmbiguousTargetError so stale in-memory siblings clear when the registry changes. Adds detect_changes callTool ambiguity test, registry-refresh regression test, pickRepoHandleForCwd MCP cwd doc, and temp-dir cleanup in #1658 fixtures.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(mcp): PR #1753 review follow-ups + collision-id case bug

Address Findings 3-6 from the production-readiness review on PR #1753,
plus a latent bug surfaced while writing the F5 regression test:

- F3: drop the no-op `try { ... } catch (err) { throw err; }` wrapper
  around the miss-path retry in `resolveRepo`; the catch only re-threw.
- F4: rewrite the misleading "child/repo" example on the relative-path
  tier — `child/repo` would be classified as path-like and never reach
  this branch. Comment now describes bare, separator-free names
  resolved against `process.cwd()`.
- F5: add regression test for the stable hashed-id tier so a duplicate
  sibling can be reached by its `<name>-<hash>` id. Writing this test
  exposed that `repoId()` produced a mixed-case base64url suffix while
  `resolveRepoFromCache` lowercased the param before the Map lookup, so
  collision ids with any uppercase byte in the hash were unreachable.
  Fix: lowercase the hash in `repoId` so it survives `paramLower`.
- F6: add regression test asserting two repos sharing a name prefix
  (`project-a`, `project-b`) cause `resolveRepo("project")` to reject
  as not-found rather than silently returning the first partial match.

* refactor(mcp): tighten PR #1753 follow-up tests + pin hash length

Address three P2 maintainability findings from the ce-code-review pass
on commit aa7f2050:

- Export `REPO_ID_HASH_LENGTH` from local-backend.ts and use it in both
  `repoId()` and the hashed-id test. Closes the silent-drift hole where
  the test's inline formula could fall out of sync with the source
  without any signal.
- Extract `makeSharedPrefixFixture(nameA, nameB)` next to
  `makeDuplicateNameFixture`. Centralises the temp-dir + `.gitnexus`
  scaffolding + `duplicateFixtureDirs.push()` cleanup contract so
  future callers can't drop the cleanup step.
- Reorder the hashed-id test's comment block so the intentional-coupling
  rationale leads, before the description of the formula being mirrored.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* chore: re-run CI

---------

Co-authored-by: Test <test@example.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-21 19:21:25 +01:00
df2ed009ce fix(group): detect httpx AsyncClient alias imports (#1687)
* fix(group): detect httpx AsyncClient alias imports

* fix(group): anchor httpx dotted imports and skip shadowed aliases

Addresses Findings 1-3 of the production-readiness review on PR #1687.

- F1: the `(dotted_name (identifier) @module)` capture matches every
  segment of a dotted module path, so `import package.httpx as hx` and
  `from package.httpx import AsyncClient` would falsely populate the
  alias sets. Anchor the check on `moduleNode.parent?.text === 'httpx'`
  so the full dotted_name must equal `httpx`.

- F2: `moduleAliases` and `asyncClientAliases` were file-global and
  unaware of Python scope. A function-local rebind like
  `AsyncClient = lambda: MockClient()` left the alias entry intact and
  any subsequent `client = AsyncClient(); client.get(...)` emitted a
  false-positive consumer contract. Walk every
  `(assignment left: (identifier) @name)` whose name matches an alias,
  record the enclosing function/class scope as poisoned, and skip
  direct- and module-attribute matches when the call site is inside
  that scope chain.

- F3: extend the existing fixture with dotted-package look-alikes and
  three local-shadow cases (`shadow_direct_alias`, `shadow_module_alias`,
  `shadow_direct_context`) and assert the would-be FP contractIds are
  not emitted.

- F6: refresh the module-level docstring to mention the supported
  import-alias forms and the shadow-exclusion behavior.

* refactor(group): tighten httpx alias shadow detection and broaden tests

Follow-up addressing the residual review findings on PR #1687.

- Replace inline scope-key construction in isAliasShadowed with a
  getScopeKey call so the two helpers cannot drift apart (M1).
- Collapse the double tree traversal in collectHttpxAsyncClients: build
  one combined alias set and pass it to a single
  collectAliasShadowScopes call (perf, P2).
- Add a `shadowScopeKey` helper that returns the scope a rebind actually
  shadows under Python LEGB rules: function scope for in-function
  rebinds, 'module' for top-level rebinds, and `null` for class-body
  rebinds (class attributes do not shadow bare-name lookups in methods).
  Removes the previous blanket `scopeKey === 'module'` skip and now
  correctly poisons module-level rebinds (correctness #1).
- Extend `ALIAS_SHADOW_PATTERNS` to cover tuple, list, and pattern_list
  destructuring targets (correctness #2).
- Rename `ALIAS_REBIND_PATTERNS` to `ALIAS_SHADOW_PATTERNS` and update
  the block comment to say "shadowed" rather than "poisoned" (M4).
- Collapse `callScopeKeys` to a single-line return; the dead Set wrap
  was misleading future readers (M2).

Tests:
- New negative fixtures for 3-segment dotted import
  (`import a.b.c.httpx as deep_evil`), relative import
  (`from .httpx import AsyncClient as rel_evil_async`), tuple
  destructuring rebind, and an isolated file exercising the module-level
  rebind path (T1, correctness #2, expanded F2).
- New positive fixture confirming that a class-body assignment of
  `AsyncClient` does NOT poison the surrounding methods.
- Add a positive control assertion for `module_direct_client` so the
  dotted-package negative assertions cannot pass vacuously (T3).

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Test <test@example.com>
2026-05-21 18:24:27 +01:00
luyua9andGergő Magyar dd3527327d feat(ingestion): Link object literal methods to exported bindings (#1718)
* fix: link object literal methods to exported bindings

* fix(ingestion): bridge object-literal value receivers in scope-resolution (PR #1718 review)

Addresses adversarial production-readiness review on PR #1718 / issue #1358:
- F1 (caller resolution) — setting `ownerId` on object-literal method symbols
  alone is not sufficient; the scope-resolution receiver-bound resolver only
  consults class-like or type-annotated bindings, so lowercase value receivers
  (`export const fooService = {...}; fooService.getUser(...)`) never reach the
  owner-indexed lookup. Adds a Case 5 value-receiver bridge in
  receiver-bound-calls.ts that resolves the receiver name as a Const/Variable
  binding, translates its def to the canonical graph node id, and emits the
  CALLS edge via the owner-indexed method registry.
- F2 (boundary guard) — rewrites findObjectLiteralBindingInfo as an explicit
  two-phase AST walk: Phase A tracks object-literal depth (returns null for
  nested literals and pre-declarator function/class boundaries — IIFE
  patterns); Phase B walks the declarator's ancestors and rejects function,
  class, and block-statement containers (if / for / while / try / catch /
  switch / etc.) before reaching program/export_statement. Prevents false
  HAS_METHOD edges for locally-scoped or block-scoped object literals.
- F4 — drops the dead `ownerName` field from ObjectLiteralBindingInfo.

Constraint: TS/JS are scope-resolution migrated per RFC #909; the legacy
Call-Resolution DAG (call-processor.ts) is intentionally left untouched.

Tests:
- test/integration/ast-helpers-object-literal-binding.test.ts (13 cases) —
  pins helper semantics: happy paths, function/arrow/class-ctor boundaries,
  nested literals, block scope (if / for-of / try), IIFE, assignment
  expressions without declarator.
- test/integration/object-literal-owner-resolution.test.ts (9 cases) —
  drives the full pipeline against an on-disk fixture: sequential CALLS edge
  emission (issue #1358 proof), worker-mode parity, negative local binding,
  and nested-literal attribution boundary.

Full sweep: 2958/2958 integration + 6056/6056 unit tests pass.

* refactor(ingestion): address code-review findings on object-literal owner resolution

Multi-agent code review on the prior commit surfaced 7 actionable findings,
all walked through and applied here. None change observable behavior for
issue #1358's fix; all harden correctness, predicate stability, and test
signal.

- #1 (P1 / 3-reviewer corroboration): Case 5 in receiver-bound-calls.ts no
  longer hand-builds graph.addRelationship + a dedup key. New
  tryEmitEdgeWithExplicitTargetId in edges.ts takes a pre-resolved target
  id (the canonical Method nodeId from the parser) and reuses every
  invariant of tryEmitEdge: dedup-key format, collapse-flag honoring,
  caller-id resolution, rel-id shape, mapReferenceKindToEdgeType for
  read/write ACCESSES. This also lands the adversarial reviewer's "F2"
  follow-up (hardcoded type: 'CALLS' for non-call sites) for free.

- #2 (P2 cross-reviewer): findValueBindingInScope's predicate inverted
  from denylist ("not class-like and not callable") to explicit allowlist
  matching reconcileOwnership's registration set:
  Const | Variable | Property | Static. Extracted as isOwnableValueLabel
  so future NodeLabel additions require an explicit opt-in.

- #6 (P2): walkScopeChain<T>() extracted; both findClassBindingInScope
  and findValueBindingInScope now route through it. Local scope.bindings
  are exhausted BEFORE lookupBindingsAt (imported/augmented) at every
  scope level — preserves JavaScript lexical scoping where a local const
  shadows an imported binding of the same name. Behavior was already
  correct in findClassBindingInScope but was implicit; now it is the
  walker's explicit, documented contract.

- #7 (P2): scope-walker duplication closed. findClassBindingInScope and
  findValueBindingInScope reduce to thin wrappers over walkScopeChain
  with their respective predicate. findClassBindingInScope keeps its
  qualifiedNames + dotted-name fallback tail.

- #3 (P2): parse-worker.ts hoists `const ownerId = enclosingClassId ??
  objectLiteralOwnerInfo?.ownerId` once before the symbol push, dropping
  the duplicated coalesce + `as string` cast. Matches the cast-free
  pattern at parsing-processor.ts:793. HAS_METHOD emit site reuses the
  same hoisted local.

- #4 (P2): object-literal-owner-resolution.test.ts Test A's CALLS-edge
  assertion no longer matches by name alone. .toEqual now pins the
  canonical target id (Method:src/service.ts:getUser#1 via generateId),
  confidence (0.85), and reason ('import-resolved'). A regression that
  emits the edge at confidence=0, with the wrong reason, or against a
  phantom Method node now fails the test.

- #5 (P2): worker-parity test adds a CI tripwire — when CI=1 and
  dist/parse-worker.js is missing, throw at module top with a clear
  message. Locally, skipIf(!hasDistWorker) keeps the fast-iteration
  experience; CI cannot pass with U3 (worker-path ownerId) unverified.

Verification: tsc --noEmit clean. Targeted regression sweep on
ast-helpers-object-literal-binding (13), object-literal-owner-resolution
(9), has-method (60), cross-file-binding (40) — 122/122 pass. Full unit
sweep: 6056/6056. Integration suite: 1 pre-existing Windows-flake in
worker-pool.test.ts (passes 28/28 in isolation) unrelated to this diff.

* refactor(scope-resolution): align Const label emission with legacy DAG (PR #1718 review F1)

Eliminates the architectural fragility surfaced by PR #1718's adversarial review
Finding 1. Previously, normalizeNodeLabel('const') returned 'Variable' while
the legacy DAG parse phase emits 'Const' graph nodes (via @definition.const
capture for lexical_declaration). PR #1718's Case 5 value-receiver bridge
resolved correctly only because resolveDefGraphId happened to fall back to
simpleKey after the qualified-key miss — accidental correctness.

After this change, scope-resolution defs for `const x = ...` declarations
report def.type === 'Const', matching the graph node label. resolveDefGraphId's
qualified-key path now hits on the first try; the simple-key fallback is no
longer load-bearing for value receivers and can be tightened in future without
silently breaking Case 5.

Audit completeness verification:
- Grep `\bVariable\b` across src/core/ingestion/scope-resolution/ surfaced two
  consumer sites that already accept both labels: reconcile-ownership.ts:101+168
  (`def.type === 'Variable' || def.type === 'Const' || ...`) and
  walkers.ts:207 isOwnableValueLabel (`Const | Variable | Property | Static`).
  No language hook in src/core/ingestion/languages/ branches on
  `def.type === 'Variable'` for what's actually a const declaration.
- Sentinel stress test (the full unit + integration suite run with the
  renamed label in place): 6137/6137 unit tests pass; 2967/2967 integration
  tests pass. One pre-existing Windows-only flake on worker-pool.test.ts when
  run alongside the full integration suite (passes 28/28 in isolation,
  unrelated to scope-extractor — same flake observed before this diff).

The variable mapping (`'variable' → 'Variable'`) is preserved for `var`
declarations, matching the legacy DAG's `@definition.variable` capture for
variable_declaration. The split now mirrors the parse-phase capture
distinction exactly.

Per plan docs/plans/2026-05-21-002-feat-pr1718-followups-class-instance-and-label-normalization-plan.md
U4 + U5. T1 (class-instance singleton resolution from issue #1358's second
sub-case) is deferred to a standalone pre-plan investigation, not shipped
here.

* test(ingestion): add regression coverage for issue #1358 singleton sub-cases

Closes the remaining sub-cases of issue #1358 surfaced by PR #1718's
adversarial review (Finding 4, NOTED): the class-instance singleton
(`export const fooService = new FooService();`) and the factory-pattern
singleton (`export const fooService = makeFooService();`).

Pre-plan investigation (per docs/plans/2026-05-21-002 § "Pre-Plan
Investigation Task (T1)") confirmed Outcome A for both patterns — they
already resolve end-to-end through scope-resolution's
`@type-binding.constructor` capture (languages/typescript/query.ts:489-511)
+ `propagateImportedReturnTypes` chain-follow
(scope-resolution/passes/imported-return-types.ts:114) + receiver-bound
Case 4 simple typeBinding lookup (receiver-bound-calls.ts:625). The
mechanism was wired correctly before this session; the regression-net
wasn't.

This test pins the behavior:
- Pattern 1: `caller → FooService.getUser` CALLS edge with
  confidence 0.85 and reason 'import-resolved'
- Pattern 2: same edge shape via factory chain-follow (the
  `@type-binding.alias` capture for `const u = find()` style)

Both assertions use exact `.toEqual([{...}])` shape pinning so a future
regression that targets a phantom Method node, emits at lower confidence,
or drops the cross-file import-resolved reason fails loudly.

Verification: 5/5 pass, 127/127 in targeted regression sweep including
object-literal-owner-resolution.test.ts, ast-helpers-object-literal-
binding.test.ts, has-method.test.ts, and cross-file-binding.test.ts.

No production code change. The class methods get a class-qualified node id
(`Method:src/service.ts:FooService.getUser#1`) distinguishing them from
same-name methods on other classes — distinct from the bare-name node id
shape PR #1718's object-literal case uses.

* test(resolvers): add class-instance + factory-pattern singleton coverage for TS/JS (issue #1358)

Closes the remaining sub-cases of issue #1358 surfaced by PR #1718's
adversarial review (Finding 4). PR #1718 fixed object-literal-shorthand
singletons (`export const fooService = { getUser() {} }`); this commit adds
parallel coverage for the two other singleton shapes that resolve through
the existing scope-resolution chain:

  // Pattern 1 — class-instance singleton
  export class FooService { getUser(id) { ... } }
  export const fooService = new FooService();

  // Pattern 2 — factory-pattern singleton
  export class FooService { getUser(id) { ... } }
  export function makeFooService() { return new FooService(); }
  export const fooService = makeFooService();

Pre-plan investigation (per local plan docs/plans/2026-05-21-002 § "Pre-Plan
Investigation Task (T1)") confirmed Outcome A — both patterns already
resolve end-to-end through:
  - `@type-binding.constructor` capture (languages/{typescript,javascript}/
    query.ts) seeds `fooService → FooService` at parse time
  - `propagateImportedReturnTypes` (scope-resolution/passes/
    imported-return-types.ts:114) mirrors the typeBinding cross-file
  - Receiver-bound Case 4 simple typeBinding lookup
    (scope-resolution/passes/receiver-bound-calls.ts:625) MRO-walks
    FooService and emits the CALLS edge to getUser

Tests added per language × pattern (5 each, 10 total):
- node existence (Class, Method, Function, Const, plus Function for the
  factory pattern's `makeFooService`)
- HAS_METHOD edge from class to method (class-instance variant)
- CALLS edge from caller to `getUser` with `targetFilePath: 'src/service.{ts,js}'`,
  `reason: 'import-resolved'`, `confidence: 0.85` — exact `.toEqual([{...}])`
  shape pinning so a regression that emits at lower confidence or drops the
  cross-file reason fails loudly

Fixtures placed under the existing `test/fixtures/lang-resolution/` convention.
Tests appended to `test/integration/resolvers/{typescript,javascript}.test.ts`,
matching the in-file pattern of every other resolver scenario.

Also supersedes and removes the standalone
`test/integration/class-instance-and-factory-singleton-resolution.test.ts`
introduced earlier in this PR session (`0df91b77`) — the proper home for
language-resolver scenarios is the per-language resolver test file alongside
similar fixtures (`javascript-self-this-resolution`, `javascript-cross-file`,
`typescript-tsconfig-paths`, etc.). One canonical location for the scenario,
not two.

Verification: 10/10 new singleton tests pass; 297/297 full TS+JS resolver
suite pass (no regression in any existing resolver test).

* test(resolvers): gate TS/JS singleton tests behind scope-resolution parity (CI run 26223603426)

The class-instance and factory-pattern singleton CALLS-edge resolution
tests added in c8e573bc rely on scope-resolution-only mechanisms
(`@type-binding.constructor` capture + `propagateImportedReturnTypes`
mirror + receiver-bound Case 4). The `scope-parity / typescript parity`
and `scope-parity / javascript parity` CI jobs run with
`REGISTRY_PRIMARY_TYPESCRIPT=0` / `REGISTRY_PRIMARY_JAVASCRIPT=0` and
exercise the legacy DAG path, which has no cross-file constructor-derived
typeBinding propagation. Verified by job 77202610819 (TS parity) and
77202610869 (JS parity) failing with:

  × resolves caller.fooService.getUser() to FooService.getUser via constructor-inferred typeBinding
  × resolves caller.fooService.getUser() through the factory chain to FooService.getUser

Note: my local Windows shell-prefix env-var invocation did not propagate
the flag into vitest workers correctly (the cpp parity gate's 47-skipped
behavior masked the issue when I ran an ad-hoc comparison), so the
empirical "both modes pass" finding I posted earlier was wrong. CI is the
source of truth.

Changes:
- test/integration/resolvers/helpers.ts: add `typescript` and `javascript`
  entries to `LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES` for the 2 CALLS-edge
  resolution tests in each language. Node-existence and HAS_METHOD
  assertions are NOT excluded — those pass under legacy DAG (parser-level
  emission is intact).
- test/integration/resolvers/typescript.test.ts: drop the `it` import from
  vitest; replace with `const it = createResolverParityIt('typescript');`
  shadow (matches the c/cpp/csharp/go pattern at the top of those files).
- test/integration/resolvers/javascript.test.ts: same shadow with
  `createResolverParityIt('javascript')`.

Verification:
- Default mode (registry-primary): 297/297 TS+JS resolver tests pass.
- Legacy DAG mode: the 4 listed singleton CALLS-edge tests will skip; all
  other singleton assertions (node existence + HAS_METHOD edge) continue
  to run and pass under both modes.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-21 17:18:27 +01:00
Gergő MagyarandCursor d3de5fa5d5 fix(install): materialize vendored grammars to fix Windows EPERM (#1728) (#1729)
* fix(install): materialize vendored grammars to fix Windows EPERM (#1728)

Stop using file: optionalDependencies for tree-sitter-dart/proto/swift,
which made npm symlink vendor paths on install and fail on Windows without
symlink privileges. Copy vendor trees into node_modules at postinstall
instead; keep native builds and #836 vendor hygiene.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(install): atomic materialize swap + fail-soft tests (#1728, #836)

Hardens PR #1729 against two issues the original implementation could
still hit:

1. Torn-state on rmSync→cpSync. The previous loop deleted the
   destination before copying. If cpSync threw — the exact Windows EPERM
   scenario this PR targets — a previously-working grammar was silently
   wiped. Now we copy to {dest}.materialize-tmp first and renameSync into
   place, so an interrupted copy leaves the prior materialization intact.

2. Fail-soft try/catch had no test coverage. Adds two POSIX-only tests
   (chmod 0o555 to deterministically force cpSync to throw) that verify
   (a) a single grammar failure does not abort the other two, and (b) an
   existing materialization survives a partial-copy failure. Skipped on
   Windows where chmod doesn't enforce write restriction; runs on Linux
   CI.

Other test improvements locking in the install-hygiene invariants:

- All three vendored grammars (dart/proto/swift) checked, not just dart.
- GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1 short-circuit is exercised.
- Vendor cleanliness (#836): no node_modules/build under vendor/.
- Idempotent re-runs (clean overwrite verified via sentinel file).
- Missing-vendor warn+continue path now has explicit coverage.
- Vendored package manifests asserted to carry no install script or
  runtime dependencies.
- package.json optionalDependencies asserted free of vendored grammars.
- package-lock.json assertion tightened from `if (entry !== undefined)
  { expect(entry.link).not.toBe(true); }` (vacuous when entry is absent,
  i.e. the expected post-fix state) to `expect(...).toBeUndefined()`.

Verified locally:
- npx tsc --noEmit: clean
- vitest test/unit/materialize-vendor-grammars.test.ts: 8 pass + 2
  POSIX-only skipped on Windows
- npm pack tarball: no vendor/*/node_modules or vendor/*/build entries
- Isolated global install (clean + upgrade + SKIP env) into temp prefix:
  succeeds; gitnexus --version → 1.6.5; vendor stays clean post-install.

* fix(install): address review feedback — Swift parity, atomicity, CI smoke

Resolves all findings from the automated production-readiness review on
verify/issue-1728-symlink.

Swift warning parity (review #2):
  Add tree-sitter-swift to OPTIONAL_GRAMMARS in src/cli/optional-grammars.ts
  alongside Dart and Proto. Before this commit, Swift was materialized at
  postinstall and probed by build-tree-sitter-swift.cjs but the runtime
  warnMissingOptionalGrammars() never warned when it failed to load —
  users got silent Swift degradation from the optional-grammars surface
  (parser-loader's separate unavailableNote only fires on demand). Now
  the warning path matches the materialize path.

README env-var table (review #1):
  Update the GITNEXUS_SKIP_OPTIONAL_GRAMMARS row at README.md line 248 to
  list all three vendored grammars (dart, proto, swift). The quick note
  earlier in the README already mentioned all three; only the table row
  was stale.

Atomicity hardening (review #3):
  materialize-vendor-grammars.cjs now copies to {dest}.materialize-tmp,
  renames the existing dest to {dest}.materialize-bak (if present), then
  renames the partial into dest, then removes the backup. If the
  partial→dest rename fails (e.g. Windows AV scanner racing the swap),
  the catch block restores from backup so the previously-materialized
  grammar is preserved. Closes the narrow torn-state window where the
  prior implementation could leave dest deleted after rmSync succeeded
  but renameSync failed.

Swift probe docs (review #4):
  build-tree-sitter-swift.cjs script header rewritten to describe what
  the script actually does — probe node-gyp-build at install time so
  missing-prebuild failures surface as install-time warnings instead of
  first-parse runtime errors. The script does not "activate" anything;
  the runtime require() in parser-loader does the actual load. Console
  warning text updated to match ("prebuild probe" not "activation").

Windows packaged-install smoke test (review #5):
  New CI job `packaged-install-smoke` in .github/workflows/ci-tests.yml
  matrices on windows-latest and ubuntu-latest. Runs npm pack, installs
  the produced tarball globally into RUNNER_TEMP, then asserts:
    * no vendor/*/node_modules or vendor/*/build (#836 invariant)
    * tree-sitter-{dart,proto,swift} in node_modules are real
      directories, not junctions/symlinks (#1728 invariant)
    * gitnexus --version runs against the installed CLI
  Closes the coverage gap where the existing windows-latest job only
  ran `npm ci` in the source checkout — exercising postinstall but not
  the tarball reify step that historically tripped EPERM.

Verified locally:
  npx tsc --noEmit: clean
  vitest test/unit/materialize-vendor-grammars.test.ts test/unit/cli-commands.test.ts:
    18 pass + 2 POSIX-only skipped on Windows
  prettier + eslint on all changed files: clean

* fix(ci): disable credential persistence on packaged-install-smoke checkout

GitHub Advanced Security (zizmor artipacked) flagged the new
packaged-install-smoke job's actions/checkout step as a potential
credential-persistence risk. The job runs `npm pack` + global install
and never pushes back, so the GITHUB_TOKEN that checkout would persist
in .git/config provides no value and only widens the leak surface (any
future artifact-upload step in this job would carry the token).

Disable persistence explicitly via `persist-credentials: false` on this
job's checkout. Scoped to the new job — pre-existing checkouts above
are left unchanged.

* fix(ci): use find instead of ls for tarball lookup (SC2012)

actionlint shellcheck SC2012 flagged `TARBALL=$(ls gitnexus-*.tgz | head -n1)`.
Switch to `find . -maxdepth 1 -name 'gitnexus-*.tgz' -print -quit` which
handles non-alphanumeric filenames safely. Also add an explicit
empty-result check so the failure mode is a clear error message instead
of a silent `npm install -g ""` later.

* fix(tests): sabotage vendor src (not partial path) in POSIX fail-soft tests

The fail-soft tests in materialize-vendor-grammars.test.ts pre-chmod'd
the destination's .materialize-tmp partial directory to 0o555 to force
cpSync to throw. After the atomicity rewrite (`fix(install): atomic
materialize swap + fail-soft tests`), the materialize script now starts
each grammar's loop with `fs.rmSync(partial, { force: true })`, which
deletes the chmod'd sabotage before cpSync runs — so cpSync succeeds and
the partial is then renamed into dest, leaving the test's `finally`
block with no path to chmod back (ENOENT) and the assertion that proto
remained unmaterialized failing because it materialized cleanly.

Fix: sabotage the *vendor source* directory (which the script reads from
but never modifies) by chmod'ing it to 0o000. cpSync then fails on
readdir, the catch block fires per-grammar, dart and swift still
materialize from their unaffected sources, and the existing-dest
preservation test verifies that a sabotaged second-run leaves the prior
materialization (and its sentinel file) intact.

Tests now pass locally (8 pass + 2 POSIX-only skipped on Windows) and
should pass on macOS/Ubuntu CI where the sabotage runs.

* fix(tests): restrict fail-soft tests to Linux (macOS Node cpSync abort)

Node 22 on macOS aborts the process with `libc++abi: terminating due
to uncaught exception filesystem_error` when fs.cpSync hits a source
directory it can't read — the abort happens at the C++ filesystem layer
and bypasses Node's JS try/catch entirely (nodejs/node#51399). My
chmod-0o000-the-source sabotage strategy triggers this SIGABRT on
macOS CI before the production script's `try { cpSync } catch` ever
runs, so the test sees a child-process crash instead of the fail-soft
warning it's verifying.

The production script's fail-soft is correct on Linux (where EACCES
surfaces as a normal JS exception) and effectively untestable on macOS
via permission sabotage. Real installs don't hit this — npm always
ships vendor/ with readable permissions — so the macOS gap is a test
artifact, not a behavior gap.

Restrict the two chmod-based tests to Linux only by replacing
`skipOnWin` with `linuxOnly`. Linux CI continues to verify both the
one-grammar-fails-others-succeed and existing-materialization-preserved
invariants. macOS and Windows runs skip these two scenarios; the other
8 tests still run on every platform.

* fix(tests): remove materialize unit tests, rely on CI smoke job

The materialize-vendor-grammars.test.ts file has been a recurring source
of platform-specific CI noise:

  - Windows: chmod doesn't enforce read/write restrictions the way POSIX
    does, so the fail-soft tests had to be skipped there.
  - macOS Node 22: cpSync against an unreadable source aborts the process
    with a libc++ filesystem_error (nodejs/node#51399) that bypasses JS
    try/catch entirely — making the chmod-based fail-soft tests
    unrunnable on macOS too.
  - The "vendor-cleanliness" and "idempotency" tests on Windows
    intermittently flake due to fs.cpSync timing on the GitHub runner.

The invariants these tests verified are now covered by stronger,
more realistic surfaces:

  - packaged-install-smoke (ci-tests.yml): runs `npm pack` then
    `npm install -g ./gitnexus-*.tgz` on windows-latest and
    ubuntu-latest, then asserts no vendor/*/node_modules,
    no vendor/*/build (#836), no junctions/symlinks on the
    materialized grammar directories (#1728), and a working
    `gitnexus --version`. This is the actual end-user install path.

  - cli-commands.test.ts (kept, unmodified): asserts package.json
    declares no `file:` optionalDependencies for vendored grammars,
    the Swift vendor manifest carries no install script or
    dependencies, and the postinstall chain runs
    materialize-vendor-grammars.cjs + build-tree-sitter-swift.cjs.
    These are static manifest checks — deterministic, fast, no
    flake risk.

Removing the dynamic script-execution tests trades unit-level coverage
for end-to-end smoke coverage that actually exercises the
`file:` → cpSync change against a real npm install lifecycle, on
the platform the fix targets (windows-latest).

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-21 16:47:22 +01:00
4c06d64a3b chore(deps)(deps): bump zod from 4.3.6 to 4.4.3 in /gitnexus-web (#1736)
Bumps [zod](https://github.com/colinhacks/zod) from 4.3.6 to 4.4.3.
- [Release notes](https://github.com/colinhacks/zod/releases)
- [Commits](https://github.com/colinhacks/zod/compare/v4.3.6...v4.4.3)

---
updated-dependencies:
- dependency-name: zod
  dependency-version: 4.4.3
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-21 16:25:35 +01:00
ChamHerryandwangxc 2a3d14057a fix(analyze): prevent cache-hit native workers from aborting (#1751)
* fix(analyze): prevent cache-hit native workers from aborting

Delay parse worker startup until a cache miss requires it, fall back to sequential parsing when initial worker readiness fails, and preserve analyzer diagnostics/progress when heap respawn captures child output.

Constraint: Node 25 and tree-sitter/N-API worker initialization can abort before ready, while warm-cache analysis should not start workers at all.

Rejected: Treating status-134/SIGABRT as heap OOM unconditionally | native worker aborts require distinct recovery guidance and stderr/stdout evidence.

Rejected: cli-progress noTTYOutput for respawn progress | it appends newline frames instead of preserving one-line redraw UX.

Confidence: high

Scope-risk: moderate

Directive: Keep parse-worker creation behind confirmed cache misses and preserve TTY-style progress when respawn pipes stderr for crash classification.

Tested: GitNexus impact analysis for ensureHeap, runChunkedParseAndResolve, createWorkerPool, WorkerPool, walkRepositoryPaths; GitNexus detect_changes scoped to staged worktree; targeted vitest for analyze respawn, parse lazy cache, filesystem walker, worker pool; npx tsc --noEmit; npm run build; NODE_OPTIONS='--max-old-space-size=8192' npm test.

Not-tested: Windows terminal rendering and published npm package install path.

* ci(docker): tolerate slower arm64 TypeScript builds

Docker PR builds run gitnexus prepare under QEMU for linux/arm64, where the fixed 120s TypeScript timeout can kill otherwise healthy builds. Increase the default timeout and allow GITNEXUS_BUILD_TIMEOUT_MS to tune slower environments without changing the build steps.

Constraint: PR #1751 Docker Build & Push gitnexus failed with spawnSync /bin/sh ETIMEDOUT while running node_modules/.bin/tsc in scripts/build.js.\nRejected: Rerunning CI only | the failure was the build script's deterministic timeout boundary under arm64 emulation, not a code assertion.\nConfidence: high\nScope-risk: narrow\nDirective: Keep build timeout changes in scripts/build.js configurable; do not hide real compiler failures, only allow slower successful compiles to finish.\nTested: GitNexus impact for gitnexus/scripts/build.js reported LOW; gitnexus detect_changes reported 1 changed file, 0 affected processes, low risk; git diff --check; gitnexus npm run build.\nNot-tested: GitHub Docker arm64 build rerun before pushing; local Docker multi-platform build under QEMU.

* fix(analyze): truncate respawn progress safely

Preserve complete ANSI escape sequences and grapheme boundaries when the respawn progress terminal shim truncates wrapped output, so the shim does not emit dangling escape bytes or split surrogate pairs while keeping raw writes untouched.

Constraint: Claude review on PR #1751 flagged `s.slice(0, width)` in createAnsiPipeTerminal.write() as a latent terminal-corruption risk.
Rejected: Adding a display-width dependency | a local helper is sufficient for this narrow respawn terminal shim and avoids new dependency churn.
Rejected: Changing silent status-134 classification | current tests already document the output-less 134 fallback as heap guidance.
Confidence: high
Scope-risk: narrow
Directive: Keep respawn terminal writes ANSI-aware and preserve rawWrite bypass semantics for callers that intentionally write control sequences.
Tested: GitNexus impact for createAnsiPipeTerminal reported LOW; GitNexus detect_changes reported 2 changed files, 3 affected processes, medium risk; targeted vitest for analyze respawn progress and heap respawn; gitnexus npx tsc --noEmit; prettier check for changed files; eslint for changed files.
Not-tested: Full npm test suite; manual terminal rendering on Windows.

---------

Co-authored-by: wangxc <wangxc_a_bj@si-tech.com.cn>
2026-05-21 16:17:02 +01:00
a9fef2c68d fix(lbug): keep serve stable when sidecars are missing (#1747)
* fix(lbug): keep serve stable when sidecars are missing

Shared missing-shadow WAL recovery prevents repeated read-only open warnings when LadybugDB sidecars are absent, while the Express preflight fix keeps `gitnexus serve` compatible with Express 5 route parsing.

Constraint: LadybugDB read-only replay can require a `.shadow` sidecar that may be absent after interrupted writes or checkpoint edge cases.
Rejected: keep reactive WARN-only quarantine in each adapter | it leaves repeated user-visible warnings and duplicate recovery behavior.
Confidence: high
Scope-risk: broad
Directive: Do not silently delete large orphan WALs; only quarantine tiny orphan WALs before open and keep large WALs for explicit recovery.
Tested: cd gitnexus && npx vitest run test/unit/sidecar-recovery.test.ts test/unit/lbug-adapter-wal-schema.test.ts test/unit/pool-wal-recovery.test.ts test/unit/web-ui-serving.test.ts && npx tsc --noEmit
Not-tested: full npm test in this split branch; full unit suite passed on the source branch before PR split.

Co-authored-by: OmX <omx@oh-my-codex.dev>

* fix(lbug): pool-caller ENOENT guard, symmetric size gate, permission-aware errors (PR #1747 review)

Addresses the production-readiness review of PR #1747 (Findings 1, 2, 3 of 6).
Findings 4, 5, 6 are deferred to follow-ups per the plan.

1. ENOENT-tolerance scoped to pool-adapter callers only
   - `quarantineWalForMissingShadow` stays strict in `sidecar-recovery.ts`.
     The direct adapter calls it inside `acquireInitLock` (cross-process
     file lock) — ENOENT there means the file vanished under lock and
     remains a real bug to surface.
   - New `tryQuarantineForMissingShadow` local helper in `pool-adapter.ts`
     returns a discriminated union { kind: 'quarantined', path } |
     { kind: 'peer-handled' }. Catches ENOENT, re-verifies via
     statIfExists, and converts to 'peer-handled' only when WAL really
     is gone. Defensive: if ENOENT but WAL still present, throws as
     classified error rather than silently returning success.

2. Symmetric WAL-size gate on both recovery paths
   - `refuseLargeWalQuarantine` applied in both
     `reopenReadOnlyAfterMissingShadow` and
     `reopenWritableAfterMissingShadow`. Closes the read-only data-loss
     vector (large orphan WAL silently discarded would never be replayed
     by a later writable open).

3. Permission-aware error classifier
   - New `renameFailureMessage` and `isPermissionRenameError` in
     `sidecar-recovery.ts`. EACCES / EPERM / EBUSY now surface a
     permission-specific message pointing at ACLs, AV exclusions, and
     file-locks. Other codes (ENOSPC, EROFS, EIO, ENOENT) fall through
     to `shadowSidecarRecoveryMessage`.
   - Used at both pool-adapter and direct-adapter caller catches around
     `quarantineWalForMissingShadow`.
   - `doInitLbug`'s pass-through classifier extended to include the new
     permission message. The lock-retry substring match tightened so
     "file-lock error" in the permission message is not mistaken for a
     LadybugDB lock-retry trigger.

Tests
   - sidecar-recovery.test.ts: 7 new tests for `renameFailureMessage` and
     `isPermissionRenameError`.
   - pool-wal-recovery.test.ts: 6 new tests covering ENOENT race,
     EACCES/EPERM/EBUSY classification, ENOSPC fallthrough, and the
     defensive "WAL still present after ENOENT" branch.
   - lbug-adapter-wal-schema.test.ts: 5 new tests covering the symmetric
     size gate on both recovery paths, including the boundary at exactly
     TINY_ORPHAN_WAL_BYTES (4096) and the off-by-one at 4097.

Deferred (tracked as follow-up work)
   - Brittle LadybugDB error-string matching (Finding 4).
   - PNA header end-to-end coverage gap (Finding 5).
   - warnedKeys module-global persistence (Finding 6).
   - Cross-process init lock for pool-adapter.

* fix(lbug): dedup shadow-replay predicate + counter-based warn anti-spam (PR #1747 review, Findings 4 & 6)

Smallest viable response to the two remaining non-blocking findings from the
production-readiness review of PR #1747. An earlier-revision plan proposed
regex widening + a near-miss detector + per-dbPath warn scoping; an
adversarial doc-review found those defended against hypothetical strings
LadybugDB does not produce, added observability theater with no recovery
behavior change, and did not actually fix the long-running gitnexus serve
case for hot dbPaths (where finalizeLbugSidecarsAfterClose rarely fires).
Scope shrunk to dedup + counter-based — strictly behavior-changing and
fully testable.

Finding 4 — dedup + version-coupling markers
   - `isReadOnlyShadowReplayError` was inlined in both `lbug-adapter.ts:451`
     and `pool-adapter.ts:317`. Centralized as an export from
     `sidecar-recovery.ts`. The two local copies are removed; both adapters
     now import from the shared module.
   - Both LadybugDB-coupled predicates (`isMissingShadowSidecarError` and
     `isReadOnlyShadowReplayError`) gain a `// LADYBUGDB-CONTRACT:` marker
     comment citing `@ladybugdb/core ^0.16.1`. When bumping LadybugDB,
     `git grep "LADYBUGDB-CONTRACT"` enumerates every version-coupled spot.
   - Strict matcher unchanged — when LadybugDB actually changes the error
     format, the failure mode stays loud (raw native error propagates) and
     the markers make every affected predicate trivially greppable.

Finding 6 — counter-based warn anti-spam
   - `warnedKeys: Set<string>` → `warnedKeyCounts: Map<string, number>`.
     `warnOnce` keeps its signature `(logger, key, message)` and keying
     convention unchanged — the swap is internal.
   - `WARN_MILESTONES = [1, 10, 100, 1000, 10000]`. Logarithmic spacing
     gives O(log N) warns for a condition that fires N times. Past the
     first occurrence the warn message is suffixed with "(Nth occurrence
     of this condition)" so persistence is visible in the log line itself.
   - Solves the long-running serve case: a hot dbPath hitting the same
     condition 100 times now fires 3 warns (occurrences 1, 10, 100)
     instead of 1 warn + 99 silent debug lines.

Tests (10 new in sidecar-recovery.test.ts, all green)
   - Centralized isReadOnlyShadowReplayError: positive match, false-positive
     guard, structural assertion that the duplicate regex is gone from both
     adapter files, LADYBUGDB-CONTRACT marker count.
   - Counter-based warnOnce: milestone-at-10 with suffix, milestone-at-100,
     key isolation across dbPaths, reset zeroes the counter, first-occurrence
     message does NOT carry the suffix.

Deferred (tracked separately)
   - Finding 5 — PNA header end-to-end coverage gap (CORS boundary is sound).
   - LadybugDB structured error codes (if/when the library exposes them).
   - Per-call milestone configurability — re-open if tuning is needed.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* ci: trigger CI rebuild

---------

Co-authored-by: wangxc <wangxc_a_bj@si-tech.com.cn>
Co-authored-by: OmX <omx@oh-my-codex.dev>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-21 12:35:43 +01:00
dependabot[bot] 8d71847791 chore(deps)(deps): bump @tailwindcss/vite in /gitnexus-web (#1734)
Bumps [@tailwindcss/vite](https://github.com/tailwindlabs/tailwindcss/tree/HEAD/packages/@tailwindcss-vite) from 4.2.4 to 4.3.0.
- [Release notes](https://github.com/tailwindlabs/tailwindcss/releases)
- [Changelog](https://github.com/tailwindlabs/tailwindcss/blob/main/CHANGELOG.md)
- [Commits](https://github.com/tailwindlabs/tailwindcss/commits/v4.3.0/packages/@tailwindcss-vite)

---
updated-dependencies:
- dependency-name: "@tailwindcss/vite"
  dependency-version: 4.3.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:07:41 +01:00
dependabot[bot] 061e123d72 chore(deps): bump actions/dependency-review-action from 4.9.0 to 5.0.0 (#1739)
Bumps [actions/dependency-review-action](https://github.com/actions/dependency-review-action) from 4.9.0 to 5.0.0.
- [Release notes](https://github.com/actions/dependency-review-action/releases)
- [Commits](https://github.com/actions/dependency-review-action/compare/2031cfc080254a8a887f58cffee85186f0e49e48...a1d282b36b6f3519aa1f3fc636f609c47dddb294)

---
updated-dependencies:
- dependency-name: actions/dependency-review-action
  dependency-version: 5.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:07:22 +01:00
dependabot[bot] a4954368ad chore(deps): bump release-drafter/release-drafter from 7.2.1 to 7.3.0 (#1740)
Bumps [release-drafter/release-drafter](https://github.com/release-drafter/release-drafter) from 7.2.1 to 7.3.0.
- [Release notes](https://github.com/release-drafter/release-drafter/releases)
- [Commits](https://github.com/release-drafter/release-drafter/compare/563bf132657a13ded0b01fcb723c5a58cdd824e2...c2e2804cc59f45f57076a99af580d0fedb697927)

---
updated-dependencies:
- dependency-name: release-drafter/release-drafter
  dependency-version: 7.3.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:07:07 +01:00
dependabot[bot] be1071143a chore(deps)(deps): bump react-syntax-highlighter in /gitnexus-web (#1731)
Bumps [react-syntax-highlighter](https://github.com/react-syntax-highlighter/react-syntax-highlighter) from 16.1.0 to 16.1.1.
- [Release notes](https://github.com/react-syntax-highlighter/react-syntax-highlighter/releases)
- [Changelog](https://github.com/react-syntax-highlighter/react-syntax-highlighter/blob/master/CHANGELOG.MD)
- [Commits](https://github.com/react-syntax-highlighter/react-syntax-highlighter/compare/v16.1.0...v16.1.1)

---
updated-dependencies:
- dependency-name: react-syntax-highlighter
  dependency-version: 16.1.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:06:41 +01:00
dependabot[bot] 3d8aa7f435 chore(deps): bump github/codeql-action from 4.35.3 to 4.35.4 (#1738)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.35.3 to 4.35.4.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/e46ed2cbd01164d986452f91f178727624ae40d7...68bde559dea0fdcac2102bfdf6230c5f70eb485e)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.35.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:06:15 +01:00
dependabot[bot] 4606e24f25 chore(deps)(deps): bump dompurify from 3.4.2 to 3.4.3 in /gitnexus-web (#1735)
Bumps [dompurify](https://github.com/cure53/DOMPurify) from 3.4.2 to 3.4.3.
- [Release notes](https://github.com/cure53/DOMPurify/releases)
- [Commits](https://github.com/cure53/DOMPurify/compare/3.4.2...3.4.3)

---
updated-dependencies:
- dependency-name: dompurify
  dependency-version: 3.4.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:06:02 +01:00
74653a8ffc feat(web): Support GitLab repository urls. (#1565)
Add GitLab URL input mode alongside existing GitHub and local modes:
- GitLab URL validation for gitlab.com and self-hosted instances
- GitLab icon component (custom SVG, matching existing GitHub icon pattern)
- Mode tab UI with GitLab option
- Backend API integration for GitLab HTTPS URLs

No token configuration included — public repositories supported only.

Close: #378


AI-model: kimi-for-coding/k2p6

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-21 10:52:01 +01:00
Gergő MagyarandCursor 8db51184ab fix(server): restore gitnexus serve startup under Express 5 (#1749)
* fix(server): restore gitnexus serve startup under Express 5

Express 5 rejects app.options('*'), which broke CI e2e when the backend
failed to start. Move PNA middleware before cors so preflight responses
include Access-Control-Allow-Private-Network, and add regression tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(server): address PR review — prettier, ephemeral port, cleanup

- Format integration and rate-limit test files for CI quality/format
- Use OS-assigned port instead of random 47xxx range
- Remove per-test GITNEXUS_HOME temp dir in afterEach
- Use regex for PNA-before-cors structural guard (indent-agnostic)

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-21 10:18:09 +01:00
1b5c6e5b6a feat(ingestion): add Kotlin scope resolver (#1727)
* feat(ingestion): add Kotlin scope resolver

* fix(ingestion): tighten Kotlin scope captures

---------

Co-authored-by: Shining <xuenning@qiyi.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-21 08:52:23 +01:00
CopilotandGergő Magyar c34c36036f fix(workers): resilient + zero-copy ingestion worker pool — prevent analyze hangs on TS-root-scale loads (#1693)
* Initial plan

* fix: skip worker-timeout files in sequential fallback and optimize TS capture node lookup

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0e53743e-0600-4690-bd0d-198894daef58

* refactor: clarify TS capture helpers after validation feedback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0e53743e-0600-4690-bd0d-198894daef58

* fix(workers): exclude in-flight file on worker error/exit, not just singleton timeout

WorkerPoolDispatchError previously surfaced the stalled path only for the
singleton-timeout final-fail branch. Worker `error` and `exit` events (and
the msg-channel `error` reply) fell back to plain `Error`, so the sequential
fallback re-attempted every file in the active job — re-hanging on the same
pathological file when the worker crashed mid-parse.

Lift the in-flight-file inference into `inFlightExcludePath(job, lastProgress)`
and wire it into the three remaining in-pool failure sites. `lastProgress` is
already in `runWorker` scope, so `items[lastProgress]` (the next file the
worker was about to acknowledge) is the best single guess at the culprit;
earlier files are still re-tried sequentially. Returns `[]` when no path is
determinable (`lastProgress >= items.length`, or path missing/non-string) so
sequential retries the whole job.

Replacement-worker startup failures stay plain `Error` (no job context); the
result-before-flush protocol bug stays plain `Error` (code fault, not file).

Tests cover the three new exclusion paths plus a negative test confirming
non-WorkerPoolDispatchError throws fall through to full sequential retry.

* fix(review): apply autofix feedback

- Use cause-neutral "worker-excluded" label in skip messages and tests now
  that worker error/exit paths share the same exclusion contract as
  singleton-timeout (correctness + maintainability reviewers).
- Add JSDoc to findSelfOrAncestorOfType{s} explaining the parent-walk
  short-circuit vs root-DFS fallback (maintainability reviewer).

* feat(workers): resilient + scalable worker pool

Restructures `createWorkerPool` so a single bad file no longer kills the
pool for the rest of an analyze run. Five interlocking layers:

1. **Auto-respawn on error/exit** — worker death triggers `replaceWorker`
   on the same slot, bounded by `maxRespawnsPerSlot` (default 3). The slot
   is dropped from rotation when the budget is exhausted; other slots
   keep running.

2. **Circuit breaker** — replaces the permanent `poolBroken=true` with a
   consecutive-failure counter. The pool only trips after
   `consecutiveFailureThreshold` deaths (default `max(3, poolSize)`) with
   no successful job in between. A successful job resets the counter so
   transient bursts of bad files don't escalate.

3. **Session-scoped file quarantine** — paths identified as the in-flight
   file at the moment of a worker death are added to a `Set<string>` on
   the pool. `dispatch()` filters quarantined items up front (they never
   reach a worker again this pool lifetime). Exposed via the new
   `WorkerPool.getQuarantinedPaths()` so callers can log/route them.
   `processParsing` surfaces the per-chunk quarantine summary alongside
   the existing fallback-exclusion log.

4. **Authoritative in-flight tracking** — `parse-worker.ts` emits
   `{type:'starting-file', path}` before each file. The pool tracks this
   per slot and uses it for crash attribution, falling back to the
   `items[lastProgress]` heuristic only when no starting-file has been
   observed (very-early crash, older worker build). Closes the
   reorder/race concerns raised by reviewers C1 and R3 in the earlier
   review run.

5. **Per-job cumulative timeout budget** — each `WorkerJob` tracks the
   total wall time spent across attempts/splits/retries. When the budget
   is exhausted (default 5x `subBatchIdleTimeoutMs`), the pool surfaces
   the in-flight path instead of letting exponential backoff balloon
   into multi-hour stalls.

Cross-layer wiring: a new `wakeIdleSlots` helper kicks any non-busy live
slot when items are requeued (after a death or split-retry), so a dropped
slot doesn't strand work in the queue. `recoverAndResume` consolidates
the per-job teardown shared by the three in-pool death sites (`error`,
`exit`, msg-channel `error`).

New env knobs: `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT`,
`GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS`,
`GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD`.
New `WorkerPoolOptions.workerFactory` injection point for unit tests.

Tests: 12 new unit tests using a FakeWorker mock cover quarantine
seeding, slot-respawn, slot-drop after budget, breaker trip + reset,
and quarantine filtering. Plus option-resolution tests for the three
new env vars. All 19 worker-pool/-fallback/-options tests pass; full
unit suite 6040 passed / 30 skipped / 0 failed.

* fix(workers): apply code-review fixes (12 findings)

Walks through every finding from ce-code-review run
20260519-094648-3549cf5e. All 12 picked Apply.

Critical:
- F1 — Layer 5 cumulative-timeout exhaustion no longer silently drops
  the rest of the job. `requeueRemainder` is now invoked before
  `handleWorkerDeath` in both Layer 5 and singleton-final-fail give-up
  paths so non-quarantined items get re-tried by another worker.
- F2 — idle-timer recovery overhaul. `!shouldContinue` branch no
  longer calls `replaceWorker` (double-spawn race with the
  `handleWorkerDeath` inside `requeueAfterTimeout`). `shouldContinue`
  branch now enforces `maxRespawnsPerSlot` before respawning, closing
  the budget-bypass for the timeout-retry path. Also fixes premature
  `maybeDone` by simplifying the bookkeeping.
- F3 — `requeueRemainder` no longer pre-charges `cumulativeTimeoutMs`
  by `job.timeoutMs`. The death itself consumed no budget, so the
  next `requeueAfterTimeout` was double-billing the first attempt.
- F4 — `WorkerPool.getQuarantinedPaths` is now optional on the
  interface, matching the defensive `?.()` call site and the existing
  mocks. Removes the contract-vs-callsite contradiction.
- F5 — per-job unattributed-death tracking. When a worker dies with
  no exclusion attribution, `requeueRemainder` tracks death count per
  `startIndex`. First time: re-queue intact. Second time: quarantine
  items[0] as best guess, or drop the job entirely when items lack
  paths. Bounds the death loop the original design admitted to.
- F6 — per-slot consecutive-failure counter. Replaces the pool-wide
  scalar so a chronically-failing slot trips the breaker on its own
  streak instead of being masked by another slot's successes.

Smaller:
- F7 — exhaustiveness `never` check on `WorkerOutgoingMessage` union.
- F8 — recursive `runWorker` on fully-quarantined jobs converted to
  a while-loop.
- F9 — `tripBreaker` calls `reject(err)` BEFORE awaiting
  `worker.terminate()`. A stuck terminate no longer blocks the caller.
- F10 — `parsing-processor.ts` quarantine log de-duplicates per pool
  instance via a `WeakMap`. Only newly-quarantined paths are logged
  in each chunk; the per-chunk count still surfaces via progress.
- F11 — extract `firstPath` local in `requeueAfterTimeout`; eliminates
  double `itemPath` call and the `unknown as string` cast.

Tests (F12, 6 new):
- crash-error event path (errorHandler).
- F5 drop-branch coverage via items without `.path`.
- Common-case unattributable crash falling back to items[0] heuristic.
- `replaceWorker` startup failure (workerFactory emits 'exit' before
  'online').
- All-slots-dropped breaker trip.
- `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` env override.

Residual gap (deferred): no unit test exercises the Layer 5
cumulative-budget runtime path — requires fake-timer interleaving
with FakeWorker that's too brittle for this iteration. Tracked.

Unit suite: 257 files / 6056 passed / 30 skipped / 0 failed.

* test(workers): integration tests for resilience layers + fix requeue-after-timeout flow

Adds 6 new real-worker integration tests covering the PR #1693
resilience layers + fixes 3 follow-on bugs surfaced while writing them.

New integration coverage (real worker threads + temp fixture scripts):

- `respawns the slot after worker process.exit and finishes the work on
  the replacement` — exercises Layer 1 auto-respawn + Layer 3 quarantine
  through real IPC.
- `attributes exactly via authoritative starting-file message on worker
  crash` — Layer 4 end-to-end: starting-file message → exact quarantine
  attribution (not the items[0] heuristic).
- `quarantine filters subsequent dispatches without sending to a worker`
  — second dispatch's sub-batch payload audited via filesystem; the
  quarantined path is never sent across the message channel.
- `drops a slot after maxRespawnsPerSlot and continues on the survivor`
  — 2-slot pool, slot dies twice past budget, survivor finishes
  re-queued remainder.
- `trips the circuit breaker on cascading per-slot consecutive failures`
  — single-slot pool, dies on every job, breaker trips after
  consecutiveFailureThreshold with WorkerPoolDispatchError carrying
  the cumulative quarantine.
- `survives a worker error event (uncaught throw) the same as a
  process.exit` — validates recoverAndResume on the errorHandler path
  via a real worker `throw` (not just process.exit).

Bug fixes uncovered while writing these tests:

1. **Stack-overflow recursion in runWorker's no-worker branch** —
   `if (!worker) { ...; wakeIdleSlots(); maybeDone(); }` recursed
   indefinitely when multiple slots were mid-respawn simultaneously
   (wakeIdleSlots → runWorker → no worker → wakeIdleSlots → …).
   Removed the wakeIdleSlots call: the slot's own respawn IIFE owns
   runWorker post-respawn, and other slots will pick up work via
   finishJob's runWorker.

2. **requeueAfterTimeout dispatched work before respawn completed** —
   the F2 fix had `requeueAfterTimeout` `void`-discarding
   `handleWorkerDeath`, so the `!shouldContinue` IIFE had no way to
   know when the respawn finished. New design: `requeueAfterTimeout`
   returns a `TimeoutDecision` discriminated union; the IIFE owns
   the death-and-respawn-and-dispatch orchestration in an async
   closure so it can `await handleWorkerDeath` and then call
   `runWorker` deterministically.

3. **Stalled-singleton + protocol-error + replacement-startup-crash
   tests** had stale contracts predating the resilience refactor. The
   stalled-singleton no longer rejects (it quarantines + resolves
   `[]`); the protocol-error rejection message now mentions
   "circuit breaker tripped"; the replacement-startup-crash test
   documents the known `waitForWorkerOnline` race (online fires
   before the worker's main script runs, so a top-level throw looks
   like a successful spawn) — the test asserts the file is
   quarantined via the second-idle-timeout give-up path.

Full suite: 334 files / 8982 passed / 43 skipped / 0 failed (second
run; first run had a Vitest-reported flake from an uncaught worker
exception bleeding into the test report — repeated runs are clean).

* perf(workers): raise pool cap to cores-1 + defer per-chunk extraction to keep workers busy

User reported 4-5% CPU utilization on a multi-core machine during
ingestion. Two structural reasons:

1. **Pool cap.** `createWorkerPool` resolved size as
   `Math.min(8, max(1, os.cpus().length - 1))` — a 16-core box got 8
   workers (50% theoretical max). U1 lifts the default to
   `min(16, max(1, cores - 1))`, exposes `GITNEXUS_WORKER_POOL_SIZE`
   env override, and adds `--workers <N>` CLI flag (`0` disables the
   pool for sequential fallback).

2. **Per-chunk extraction serialized the loop.** Per chunk:
   dispatch → await workers → main-thread `processImportsFromExtracted`
   + `processHeritageFromExtracted` + `processRoutesFromExtracted`
   + `synthesizeWildcardImportBindings` + `seedCrossFileReceiverTypes`
   → next chunk dispatch. Workers sat idle through every extraction
   block. U2 (revised from the plan's pipelined-chunks design) defers
   these passes to a single end-of-loop batch. Chunk loop becomes
   parse + merge + accumulate. Resolution sees strictly-more-info
   (full repo graph) so cross-chunk import/heritage targets resolve at
   least as well as before. Memory cost: `deferredWorkerImports`
   accumulates across chunks; bounded by total file count, acceptable.

Plan deviation note: the plan called for an in-flight chunk pipeline
(N concurrent dispatches with bounded memory). That design needed
either a `processParsing` API refactor or duplicating its catch-block
fallback in `parse-impl`. The deferred-extraction approach delivers
the same "workers stay busy" outcome with much smaller surface area
and zero changes to `processParsing`. The `GITNEXUS_PARSE_CHUNK_CONCURRENCY`
env var documented in U2 of the plan is therefore not implemented in
this commit; if memory growth from `deferredWorkerImports` becomes
a problem at very-large-repo scale, a bounded sliding-window variant
can land as a follow-up.

Tests:
- New `test/unit/analyze-worker-pool-size.test.ts` covers --workers
  validation (5 invalid inputs rejected with exit code 1 + clear
  error; valid integers set the env var; `--workers 0` routes to
  sequential).
- Extended `worker-pool-resilience.test.ts` with `resolveAutoPoolSize`
  scenarios: env override, env=0, env above cap, invalid env fallback,
  auto-formula match, integer return type.
- Full unit suite: 6097 / 6127 passed / 30 skipped / 0 failed.
- Full integration suite (second run): 77 / 78 passed / 1 skipped /
  0 failed. First run had a known cosmetic flake from an uncaught
  worker exception bleeding into the test reporter.

Resilience contract from PR #1693 preserved: per-slot respawn budget,
circuit breaker, quarantine, authoritative in-flight tracking,
cumulative timeout budget — all unchanged.

New env vars surfaced in --help: GITNEXUS_WORKER_POOL_SIZE,
GITNEXUS_PARSE_CHUNK_CONCURRENCY (reserved for future bounded
pipelining).

* docs(readme): document --workers CLI flag

* feat(workers): add getStats() and per-chunk throughput logging

* test(workers): cleanup leaked temp-dirs and drop duplicate option-resolution block

- Add afterEach to worker-pool-resilience.test.ts cleaning up the per-test temp
  directory created by beforeEach (~25 stale dirs per CI run previously).
- Delete the duplicated describe('worker pool option resolution', ...) block.
  Verified the first block (lines 490-532) is a strict superset (includes the
  GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS env test the second block omitted),
  so deletion loses no test coverage.

Addresses PR #1693 review findings L2 (temp-dir leak) and L3 (duplicate block).

* feat(cli): thread --workers via PipelineOptions + snapshot/restore CLI env

Resolves PR #1693 review B2 (env-var leak in long-running hosts):

- --workers is now threaded through AnalyzeOptions -> runFullAnalysis
  -> PipelineOptions.workerPoolSize -> createWorkerPool's explicit
  poolSize arg, bypassing the GITNEXUS_WORKER_POOL_SIZE env channel.
  The env var remains as a back-compat fallback inside resolveAutoPoolSize
  for operators who set it directly.
- analyzeCommand and wikiCommand snapshot the GITNEXUS_* env vars they
  mutate at function entry and restore them in finally. Inner *Impl
  extraction keeps the diff surgical (no body re-indent). process.exit(0)
  on the CLI success path still terminates the process; restoration
  matters for programmatic callers (tests, long-running hosts) reaching
  early-return paths or the alreadyUpToDate fast path.
- Tests updated to assert the new behavior:
    analyze-worker-pool-size.test.ts: workerPoolSize flows through
      runFullAnalysis options; env is not mutated; back-to-back calls
      see their own values, not the previous call's leak.
    analyze-worker-timeout.test.ts: env IS set during the runFullAnalysis
      call (captured via mockImplementation) and restored after, proving
      the timeout reaches downstream while the leak fix holds.
- Also addresses L4: afterEach NODE_OPTIONS restore so back-to-back test
  runs don't accumulate --max-old-space-size=8192 tokens.

Addresses PR #1693 review B2 (blocker) and L4 (test polish).

* feat(workers): harden worker lifecycle (messageerror + availableParallelism + ready handshake)

Resolves PR #1693 review H1, H2, M4:

H1 - messageerror handler at every dispatch site
  V8 deserialization failure on postMessage previously left the message
  silently lost; the pool would wait out the idle timeout (default 30s)
  instead of treating it as worker death. The dispatch loop now wires
  worker.once('messageerror', ...) alongside error/exit and routes through
  recoverAndResume so the existing per-slot respawn budget, in-flight
  file attribution, and circuit-breaker layers fire as designed.

H2 - resolveAutoPoolSize uses os.availableParallelism()
  Mirrors the pattern at capabilities.ts:85 (defaultEmbeddingThreads).
  os.cpus().length returns the host CPU count, which over-sizes the pool
  on cgroup-limited containers, taskset-restricted runtimes, and CI
  runners with explicit CPU quotas. Falls back to os.cpus().length on
  Node < 18.14.

M4 - worker-side ready handshake replaces online-trust
  parse-worker.ts now emits {type: 'ready'} after all top-of-script
  initialization completes, BEFORE the message handler is attached. The
  pool's renamed waitForWorkerReady listens for this message under a
  bounded WORKER_READY_TIMEOUT_MS (5s) budget instead of trusting Node's
  online event - which fires when the worker thread starts, BEFORE the
  script body runs, letting init crashes slip past pool startup. ready
  is added to WorkerOutgoingMessage with an exhaustiveness-checked
  no-op branch in the dispatch handler (defensive: the message is
  consumed by waitForWorkerReady before dispatch handlers attach).
  messageerror is wired into waitForWorkerReady the same way.

Test scaffolding:
  - FakeWorker emits {type: 'ready'} in addition to 'online' so
    replacement workers in unit tests don't hit the 5s budget.
  - Integration test ad-hoc worker scripts go through a writeReadyWorker
    helper that prepends the ready handshake. Tests intending to script
    "crash BEFORE ready" can bypass the helper.

61/61 worker-pool unit tests pass; 28/28 integration tests pass.

* feat(parse-impl): monotonic progress + verbose-gated throughput log + seed-before-build

Resolves PR #1693 review M2, M3, L1, L5 in a single parse-impl.ts pass:

M2 - Monotonic progress through deferred phase (no more "stuck at 82%")
  Previously the deferred resolution stages (imports, heritage, routes,
  calls) all emitted percent: 82 — the UI looked frozen for the duration
  of the deferred work, which on large repos is several seconds to minutes
  and visually identical to the hang PR #1693 set out to fix.
  Redistributed:
    parse phase:  20-70 (was 20-82)
    imports:      70-75
    heritage:     75-80
    routes:       80-85
    calls:        85-95
  Each deferred stage now advances through its own band via the existing
  per-batch progress callback. Skipped stages (zero deferred input) leave
  their band as a no-op jump - the next stage still starts at its own
  band, preserving strict monotonicity. The "no parseable files" early
  return now jumps to 95 (was 82), and the duplicate "Parsing N files..."
  announcement is suppressed when totalParseable === 0 to avoid a
  non-monotonic 95 -> 20 regression that pre-existed (uncovered by the
  new monotonic test).

M3 - Throughput log gated on `--verbose`, not just NODE_ENV=development
  The per-chunk files/s log was gated on `isDev`, so operators running
  `gitnexus analyze --verbose` in a production install never saw it.
  Now fires when (isDev || isVerboseIngestionEnabled()) — matches the
  documented promise that `--verbose` shows tuning observability.

L1 - Typo rename: `chunkChunkStartMs` -> `chunkStartMs`

L5 - `buildExportedTypeMapFromGraph` runs BEFORE `seedCrossFileReceiverTypes`
  Previously the seeding branch was reached with `exportedTypeMap.size === 0`
  in the worker path (the map was only built far below, AFTER the seeding
  branch), so the seed dead-coded itself silently and call resolution
  never got the cross-file receiver-type enrichment. Now the map is
  populated from the in-progress graph before the seed call; the
  post-parse builder remains as a defensive sequential-path fallback,
  guarded by `size === 0` so we don't pay the cost twice on the worker
  path. Net win: cross-file CALLS edges that previously had no receiver
  type now get enriched.

New test: parse-impl-progress-monotonic.test.ts
  Asserts the emitted percent stream is strictly non-decreasing across
  the parse + deferred phases, and that the deferred band (>=70) is
  actually reached. Also pins the "no parseable files" path to exactly
  [95] so the 95 -> 20 regression we just fixed can't re-emerge.

* feat(parse-impl): bounded chunk concurrency via file-pre-fetch pipeline

Resolves PR #1693 review B1 (GITNEXUS_PARSE_CHUNK_CONCURRENCY documented
in --help but unimplemented).

The chunk loop now pre-fetches chunk file contents up to
`parseChunkConcurrency` chunks ahead of the worker-dispatch cursor so
disk I/O overlaps with worker compute. Worker dispatch itself stays
serial because WorkerPool.dispatch is not reentrant — concurrent calls
would race on the shared per-slot busy/in-flight state, regressing the
hang/resilience work this PR is built on. The pre-fetch path is the
honest interpretation of "concurrent in-flight parse chunks" that the
help text advertises: I/O overlap, not parallel worker dispatch.

Concurrency value resolution:
  1. PipelineOptions.parseChunkConcurrency (threaded from CLI)
  2. GITNEXUS_PARSE_CHUNK_CONCURRENCY env var
  3. Default 2 (matches the help text)

F4 (wildcard-synthesis ordering) is preserved: deferred-state
aggregation runs in chunkIdx order because the for-loop iterates
sequentially after awaiting each chunk's pre-fetched contents.
Cross-chunk processors (processImportsFromExtracted,
synthesizeWildcardImportBindings, etc.) still run only after all
chunks complete — they see deterministic input regardless of
file-read completion order.

Concurrency=1 produces behavior identical to the pure-serial loop;
that's the regression baseline.

New test: parse-impl-chunk-concurrency.test.ts
  - Asserts graph output is identical (nodeCount + relationshipCount)
    between parseChunkConcurrency=1 and =2 — the critical correctness
    invariant. Exact .toBe(N) comparisons per DoD §2.7 (the second run's
    counts must equal the first run's exactly).
  - Pins specific fixture symbols (foo/bar/Baz) under both
    parseChunkConcurrency=1 and the env-fallback (3) path.
  - Env-fallback test confirms GITNEXUS_PARSE_CHUNK_CONCURRENCY is
    honored when the option is undefined.

* test(workers): pin cumulative-timeout exhaustion behavior

Resolves PR #1693 review M6: the existing resilience suite asserts only
the *default value* of maxCumulativeTimeoutMs (5x subBatchIdleTimeoutMs),
not that dispatch actually aborts the offending job when the cumulative
wall-clock budget is exhausted. Without this test, a future refactor
could remove the exhaustion branch in requeueAfterTimeout and the suite
would stay green while the pool sat in retry loops for an hour on a
real production stall.

Scenario:
  subBatchIdleTimeoutMs    = 100ms
  timeoutBackoffFactor     = 10
  maxCumulativeTimeoutMs   = 300ms

Single file, HangingWorker that never responds. First attempt times
out at 100ms (cumulative=100). The next backoff (1000ms, cumulative
1100ms) exceeds the 300ms cap, so requeueAfterTimeout returns
give-up on the first timeout retry and the file goes to the session
quarantine. Asserts:
  - pool.getQuarantinedPaths() includes 'src/stuck.ts' after dispatch
  - if dispatch rejected, the error is a WorkerPoolDispatchError
    (the typed surface that routes to sequential fallback)

Uses a local minimal HangingWorker double rather than the full
action-scripted FakeWorker from worker-pool-resilience.test.ts —
the inverse pattern (always hang) doesn't need the scripted-action
machinery and keeps the test file focused on the one behavior.

* docs(readme): add environment-variables reference table

Resolves PR #1693 review L6: operator-facing env vars were either
mentioned inline (GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS) or only
documented via `gitnexus --help`, with no single place to look up
the full set. The new "Environment variables" subsection under the
Quick Start CLI block lists every operator-facing knob with default,
effect, and tuning guidance, matching the names in cli/index.ts
addHelpText post-U2 / U1.

Covers:
  GITNEXUS_WORKER_POOL_SIZE           (--workers)
  GITNEXUS_PARSE_CHUNK_CONCURRENCY    (newly real per U1)
  GITNEXUS_VERBOSE                    (--verbose)
  GITNEXUS_MAX_FILE_SIZE              (--max-file-size)
  GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS (--worker-timeout × 1000)
  GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES
  GITNEXUS_CHUNK_BYTE_BUDGET
  GITNEXUS_NO_GITIGNORE
  GITNEXUS_SKIP_OPTIONAL_GRAMMARS

CLI flag vs env-var precedence is stated explicitly (CLI > env > default)
so operators running long-lived hosts (MCP server, eval-server) know
which channel wins.

* test(workers): pin quarantine path round-trip and non-normalization contract

Resolves PR #1693 review M5 (Windows quarantine path-normalization
coverage). worker-pool.ts quarantines paths via a Set<string> keyed by
exact string equality. The existing suite never asserted this contract,
which lets a future "helpfully normalizing" refactor on one side of the
pipeline (caller, worker, or pool) silently break quarantine filtering
on Windows.

This file pins the contract from both directions:

1. Round-trip: a path the caller dispatches with backslashes
   (src\bad.ts) flows through starting-file -> death -> quarantine ->
   next-dispatch filter verbatim. The replacement worker never sees the
   re-dispatched bad path because the pool's pre-dispatch filter
   short-circuits it.

2. Non-normalization: quarantining src\poison.ts does NOT filter
   src/poison.ts. Whoever changes that contract has to update this test
   alongside (the load-bearing assertion catches accidental
   path.normalize() calls in the quarantine path).

Runs on every platform — the path strings are test-injected, so the
test exercises the same code path regardless of the host's path.sep.
Used a self-contained FakeWorker that emits {type:'ready'} for U3's
waitForWorkerReady handshake, so the test doesn't depend on the larger
worker-pool-resilience.test.ts harness.

* test(typescript): pin capture-anchor rewrite invariants (B5 regression)

Resolves PR #1693 review B5: the captures.ts ancestor-walk rewrite
(findSelfOrAncestorOfType[s] + pickFirstNode replacing the prior
findNodeAtRange-from-root path) was semantically equivalent to its
predecessor per Lane 4 of the production-readiness review, but the
existing typescript-captures.test.ts didn't pin the specific sharp
edges where an over-aggressive walk would silently break captures.
This file does.

Each test exercises a capture class whose anchor type is one the
rewrite explicitly handles:

  - member call obj.foo() -> @reference.call.member (call_expression
    anchor walks to self)
  - dynamic import import("./helper") -> raw @import.dynamic gets
    decomposed by splitImportStatement into @import.statement with
    @import.kind=dynamic + @import.source stripped of quotes
  - JSX <Foo /> in .tsx -> @reference.call.free emitted (TSX query
    pattern, query.ts:899-905) but @declaration.parameter-count is
    NOT synthesized because findSelfOrAncestorOfType('call_expression')
    returns null on a jsx_self_closing_element anchor. Pre-rewrite the
    range lookup also returned null. Pinning this contract catches
    accidental "walk JSX -> outer call" refactors.
  - constructor `new Foo(1,2)` -> @reference.call.constructor (new_expression
    anchor walks to self)
  - named/namespace import + re-export -> @import.statement (one each)
  - class method override -> @declaration.method per class, no collapse
  - member read obj.foo (no call) -> @reference.read.member

All assertions use exact .toBe(N) per DoD §2.7.

* test(parse-impl): pin multi-chunk graph equivalence under deferred extraction

Resolves PR #1693 review B4: the deferred-extraction reorder (moving
processImportsFromExtracted / Heritage / Routes / Wildcard /
ReceiverTypes from per-chunk to end-of-loop) was proven observably
equivalent by Lane 4 of the production-readiness review. Until now,
the existing suite never asserted cross-chunk graph equivalence,
which lets a future refactor that accidentally tightens the per-chunk
vs end-of-loop coupling silently break cross-chunk resolution.

This test forces multi-chunk parsing on a small fixture by setting
GITNEXUS_CHUNK_BYTE_BUDGET=64 BEFORE the parse-impl module loads
(the budget is captured at module load via vi.resetModules — a future
move to function-scope env reads is U14 in Phase 2). Then runs the
same fixture under a 10MB budget (single chunk) and asserts the two
graphs are byte-identical: same nodeCount, same relationshipCount,
exact .toBe(N) per DoD §2.7.

Fixture: 3-file class hierarchy with cross-file inheritance — Animal
(a.ts) -> Dog extends Animal (b.ts) -> makeDog returns Dog (c.ts).
Forces the resolver to chain imports + heritage across chunks. A
second test pins specific symbol names (Animal, Dog, makeDog, speak,
bark) in the multi-chunk graph so a regression in chunk-boundary
resolution surfaces as a missing-symbol failure with a specific
diagnostic instead of a bare count mismatch.

* test(parse-impl): wall-clock integration pinning multi-chunk pipeline (B3)

Resolves PR #1693 review B3 — the final P0/P1 merge blocker. With this
test, all five doc-review blockers (B1-B5) are pinned by regression
coverage.

The PR's headline claim is "analyze no longer hangs on TS-root-shaped
loads". The existing suite pins each resilience layer (worker-pool-
resilience.test.ts), the deferred-extraction equivalence (U7), and
the chunk-concurrency contract (U1). What was missing: a single
end-to-end run that exercises the full chunked parse-and-resolve
path on a multi-chunk fixture, BOUNDED by a wall-clock budget so a
regression that re-introduces the hang fails this test loudly via
timeout rather than slipping past as a count drift.

Implementation:
  - 17-file synthetic fixture: 15 small modules (one function each),
    one "realistic dense" complex.ts (30 functions + class + interface),
    and an index.ts re-exporting them. Forces cross-chunk import
    chains.
  - GITNEXUS_CHUNK_BYTE_BUDGET=64 via vi.resetModules forces multi-chunk
    parsing on the small fixture.
  - Promise.race with 30s timeout: a hang fails as
    "exceeded WALL_CLOCK_BUDGET_MS — likely the hang B3 was meant to
    prevent", not as a bounds-only inequality (DoD §2.7 distinction —
    hang-detector via exception, not regression-mask via inequality).
  - Exact .toBe(true) assertions on specific expected symbols
    (fn0..fn14, Service, Config, configure, describe, complex0/15/29)
    so a silent mid-chunk crash that exits 0 without producing graph
    data also fails this test, not just the hang case.

Scope: runs the sequential-fallback path (skipWorkers: true) because
the full real-worker scenario requires a built dist/parse-worker.js
and ~60s wall-clock per run — appropriate for a CI-integration job,
not vitest. The load-bearing invariants pinned here catch the bulk
of B3's concern; the dist-worker swap is a Phase 2 follow-up
documented in the file header.

* refactor(parse-impl): move chunk-byte-budget env read to function scope

Resolves PR #1693 review F7 / U14: pre-U14, `CHUNK_BYTE_BUDGET` was a
module-load IIFE constant that captured `GITNEXUS_CHUNK_BYTE_BUDGET`
once and froze the value for the module's lifetime. That defeated
per-call option threading (a future
`PipelineOptions.chunkByteBudget` was silently no-op'd because the
function body read the frozen module-level constant) AND forced tests
to use `vi.resetModules` to vary chunk layout. The U7
deferred-extraction test and the U6 multi-chunk integration test
both used the workaround.

After this change:

  - `DEFAULT_CHUNK_BYTE_BUDGET = 2 * 1024 * 1024` stays as a
    module-level constant — purely a default, no env access.
  - `resolveChunkByteBudget(options)` runs per call: option wins,
    then env, then default. Same options-first/env-fallback/default
    pattern as resolveAutoPoolSize and the U1 parseChunkConcurrency
    resolver — keeps the ingestion code's configuration model uniform.
  - `PipelineOptions.chunkByteBudget?` added with documentation that
    threading through options lets long-running hosts (eval-server,
    MCP daemon) size per-call without leaking process.env state
    across analyze invocations.

New test (parse-impl-env-reads.test.ts) pins all four behaviors:
  1. option-first: option present + env present -> option wins
  2. env-fallback: option absent + env present -> env wins
  3. default-fallback: both absent -> 2 MB default
  4. per-call: two back-to-back runs in the same vitest worker with
     different chunkByteBudget option values observe their OWN values,
     proving the module-load freeze is gone (no vi.resetModules in
     this test — that's the invariant being verified).

All four assertions use exact `.toBe(N)` per DoD §2.7. The chunk
count is observed by parsing the `Parsing chunk X/Y` progress message
stream — a stable proxy that doesn't require exposing internal
parse-impl counter state.

Note: U7 and U6 tests still use `vi.resetModules` because they were
written before this change. A follow-up cleanup could simplify those
tests (drop the resetModules dance, pass chunkByteBudget via options),
but they pass as-is so this commit doesn't touch them.

* feat(workers): per-slot generation counter for late-event protection (U12)

Adds a monotonic per-slot generation counter to createWorkerPool's
state. Each successful worker replacement (replaceWorker) bumps the
slot's counter exactly once — atomically with the workers[slotIndex]
swap, so observers (getStats) see the new (worker, generation) pair
consistently. Handler closures in the dispatch loop capture the
slot's generation at attach time and short-circuit when they fire
on a stale generation.

In the current implementation, cleanup() synchronously removes
listeners on a Worker instance the moment a death is observed, so
no listener naturally fires on a stale generation — the guard is a
defensive layer protecting against any future refactor that loosens
cleanup() ordering or re-attaches handlers across the swap. The
load-bearing observable is the slotGenerations[] array exposed via
WorkerPoolStats so operators (and tests) can confirm a slot was
actually replaced and not just the same worker recycled.

Implementation:
  - const slotGenerations: number[] = new Array(size).fill(0) in
    createWorkerPool's per-pool state, alongside respawnCount and
    consecutiveFailuresPerSlot.
  - replaceWorker: slotGenerations[workerIndex]++ AFTER the
    workers[workerIndex] = replacement swap (only on the success
    branch — drop-slot paths leave the counter unchanged).
  - runWorker dispatch loop: const slotGen = slotGenerations[workerIndex]
    captured before handler attachment; every handler (handler /
    errorHandler / exitHandler / messageErrorHandler) starts with
    `if (slotGenerations[workerIndex] !== slotGen) return`.
  - WorkerPoolStats gains `readonly slotGenerations: readonly number[]`.
  - getStats() returns slotGenerations.slice() so callers can't mutate
    pool state by writing to the returned array.

Two existing toEqual snapshots in worker-pool-resilience.test.ts
extended with the new slotGenerations field (both expect all-zeros —
neither test scenario triggers a respawn).

New test file (worker-pool-slot-generation.test.ts, 4 tests):
  1. Fresh pool: every slot at generation 0.
  2. Successful crash + respawn: generation bumps to 1 exactly once.
  3. Crash that drops the slot (maxRespawnsPerSlot:0): generation
     stays at 0 because no successful respawn happened. The dispatch
     rejection on breaker trip is the expected outcome here; the
     load-bearing assertion is the post-rejection stats.
  4. Multi-slot independence: one slot crashing bumps only that
     slot's generation, not the other. Order-independent via sort()
     because the round-robin assignment isn't pinned by contract.

All assertions exact .toEqual / .toBe per DoD §2.7.

* docs(bench): add parse-throughput benchmark scaffold (R13)

Resolves PR #1693 review R13 (benchmark artifact requirement).

Creates `gitnexus/bench/parse-throughput.md` documenting:

- Synthetic fixture spec (same shape as the U6 integration test, so
  CI smoke baseline and ad-hoc benchmark exercise the same paths).
- What to measure (wall-clock, peak heap, chunk count, getStats
  snapshot) and the hardware-shape metadata to record alongside.
- Harness recipe — vitest + env-var overrides to exercise sequential
  fallback vs worker-pool paths.
- Latest-measurement table with placeholder rows for the three paths
  (sequential, workers+concurrency, workers single-threaded) and an
  explicit "Status: scaffold — fill in before merging" callout. The
  U6 test's observed ~6 s wall-clock is captured as a smoke-baseline.
- Operator-tuning quick reference cross-linked to the README env-var
  section (U11) so the doc is actionable without re-reading the PR.
- "What this benchmark does NOT measure" section explicitly scoping
  the artifact's limits (synthetic ≠ real-repo, throughput-only ≠
  resilience-tested, Phase 3 IPC repack row reserved for U16-U17).

Mitigates the doc-review SG5 "static doc drift" concern via:
  1. Explicit "regenerate this file before merging" callout at the top.
  2. Self-contained methodology so anyone can re-run the numbers.
  3. Cross-links to the U6 integration test that already bounds the
     wall-clock as part of the CI suite — so "is it still completing?"
     is regression-tested even if the numbers in this doc drift.

The standalone harness script (`bench/scripts/parse-throughput.ts`)
remains a stretch goal per the original plan. The U6 vitest with
verbose ingestion logs covers the primary observability gap until
the standalone harness lands.

* perf(parse-impl): free deferred-extraction arrays after consumption (U15 lightweight M1)

PR #1693 review M1 noted that the deferred-extraction accumulator
arrays (`deferredWorkerImports`, `deferredWorkerCalls`,
`deferredWorkerHeritage`, `deferredConstructorBindings`,
`deferredAssignments`) were retained until function return, making
peak accumulator memory O(repo) instead of O(in-flight stage).

This commit implements the LIGHTWEIGHT version: free each array
immediately after its last consumer drains/reads it, dropping peak
accumulator memory progressively through the deferred-extraction
stages. The structural per-chunk streaming variant (the original
U15 framing) is deliberately deferred — the doc-review's adversarial
reviewer (A4) flagged it as defending unmeasured memory pressure,
and the simpler array-clearing captures the bulk of the benefit
without committing to a scheduling-strategy decision (microtask vs
parallel extractor task vs worker-side) that profile data should
inform.

Clears added:

  1. After `processImportsFromExtracted` (the sole consumer of
     `deferredWorkerImports`): clear the imports array before
     the heavier heritage/calls stages run.
  2. After `buildHeritageMap` (the LAST consumer of the raw
     `deferredWorkerHeritage` records — processCallsFromExtracted
     reads from the derived `fullWorkerHeritageMap` instead):
     clear the heritage array before the call-resolution stage.
  3. After `processAssignmentsFromExtracted` (the joint last
     consumer with processCallsFromExtracted for the calls/
     bindings/assignments triple): clear all three before
     downstream graph-build / scope-resolution uses its own
     working memory.

Arrays returned in the function result object (allFetchCalls,
allExtractedRoutes, allDecoratorRoutes, allToolDefs, allORMQueries,
allParsedFiles) intentionally stay live — downstream consumers
need them.

Graph-output equivalence is preserved (U7 multi-chunk equivalence
test passes — the clears happen AFTER each array's last consumer
has copied data into the graph or derived structures).

* feat(workers): introduce protocol.ts wire-format module (U16, IPC scaffold)

Defines the binary frame for worker-thread IPC as an isolated, fully-tested
module. Production wiring is deferred to U17 — shipping the wire-format
contract first de-risks the migration by establishing a single source of
truth for the byte layout. Resolves the scaffold half of PR #1693 review
R12.

Wire layout (per message, single buffer):

  +---------+-----------+---------------------+
  | tag     | length    | payload bytes …     |
  | 1 byte  | 4 bytes   |                     |
  +---------+-----------+---------------------+

  tag    : MessageTag enum value (0x01 DispatchJob ... 0x08 Ready)
  length : little-endian uint32 byte count for the payload region
  payload: UTF-8 JSON-encoded value, possibly "null"

Why JSON for the body (rather than per-shape binary encoders): the
doc-review adversarial reviewer (A2) flagged that a true per-shape
binary encoder for the result message — which carries nested
heterogeneous extracted-call / import / heritage / route arrays —
would be 500-1500 LOC and a substantial maintenance burden. The
honest perf win the IPC repack targets is moving file CONTENTS via
ArrayBuffer transferList (zero-copy ownership transfer for the
largest single piece of state in any message). That win is captured
by U17 layering transferList over the bulk file-content payload while
keeping this module's framing for the surrounding metadata. If U18
benchmark data shows the JSON body is itself a bottleneck after U17
lands, a follow-up unit can swap to per-shape binary encoding behind
the same encodeMessage / decodeMessage surface without changing the
frame.

API:
  - MessageTag (const object): stable byte tags 0x01..0x08
  - PROTOCOL_HEADER_BYTES = 5
  - ProtocolDecodeError extends Error: distinct class so U17's
    pool-side handler can route protocol violations through the
    existing messageerror recovery layer (U3 H1) distinctly from
    other failure classes
  - encodeMessage(tag, payload): Buffer
  - decodeMessage(buf): { tag, payload }
  - Uses Buffer#subarray instead of the deprecated Buffer#slice

Tests (18, all exact-equality per DoD §2.7):
  - byte layout (tag at offset 0, length LE uint32 at offset 1)
  - empty/null payload encodes to 5-byte header + 4-byte "null" body
  - round-trip for every MessageTag with representative payloads
  - non-ASCII path string (UTF-8 byte-length boundary)
  - 9 MB payload (well past the existing 8 MB sub-batch budget)
  - decode errors surface as ProtocolDecodeError, not generic Error:
      * buffer < header size
      * tag outside valid range
      * declared length exceeds buffer
      * payload bytes are not valid JSON
  - error class name is preserved through prototype chain so callers
    can `err instanceof ProtocolDecodeError` reliably

* refactor(workers): extract quarantine into its own module (U13 partial)

Honest partial U13: extract the quarantine resilience layer (Layer 3
of the 5-layer model) into a dedicated module with a small explicit
interface. The full 5-module split that the original plan named was
flagged by doc-review A10 as abstraction-without-multi-consumer-demand
("Each has exactly one consumer: worker-pool.ts. None of these layers
is imported elsewhere in the codebase pre-extraction, and the plan
doesn't identify any future consumer.") This commit ships the smallest
self-contained layer as a named module to validate the factory +
interface pattern with minimal risk. The remaining four layers
(respawn-budget, cumulative-timeout, circuit-breaker, slot-attribution)
stay inline until a real second consumer emerges (e.g., a non-parse
worker pool that reuses the same resilience layers).

Module shape (`workers/quarantine.ts`, ~30 LOC):

  interface Quarantine {
    add(path: string): void;
    has(path: string): boolean;
    snapshot(): string[];   // defensive copy
    readonly size: number;  // getter, reflects state at access time
  }
  function createQuarantine(): Quarantine

Replaces in `worker-pool.ts`:
  - `const quarantined: Set<string> = new Set()` -> `createQuarantine()`
  - `quarantined.has(p)`            -> `quarantine.has(p)` (2 sites)
  - `quarantined.add(p)`            -> `quarantine.add(p)` (2 sites)
  - `quarantined.size`              -> `quarantine.size` (2 sites)
  - `Array.from(quarantined)`       -> `quarantine.snapshot()` (6 sites)

Public worker-pool.ts API is unchanged — `getQuarantinedPaths()` still
returns the same defensive `string[]` copy. The behavioral contract is
preserved: paths are quarantined as opaque strings (the U9 / M5
non-normalization contract still holds — see the new dedicated test).

Tests:
  - 8 isolated unit tests for the quarantine module — pins the
    interface contract (empty start, add/has/size, dedup on repeated
    add, no separator normalization, snapshot defensive copy + freshness,
    size-getter live behavior).
  - All 86 existing worker-pool tests pass unchanged — they exercise
    the quarantine through the pool and act as the regression net for
    behavior preservation.

Why not the full 5-module extraction in this commit: doc-review A10's
concern is real — a single-consumer abstraction adds module-boundary
overhead (5 sets of imports, 5 dedicated test files, 5 interfaces to
keep in sync with worker-pool) without any structural benefit until a
second consumer materializes. Extracting one validates the pattern;
the remaining four can be moved on demand.

* feat(workers): wire protocol.ts encoded IPC into parse-worker + pool (U17)

Production worker IPC now uses the U16 binary wire format (1-byte tag +
4-byte LE length + UTF-8 JSON body) end-to-end. The pool encodes every
outgoing `sub-batch` / `flush` dispatch via `encodeMessage`; the worker
decodes incoming frames via `decodeMessage` and encodes its `ready`,
`starting-file`, `progress`, `sub-batch-done`, `result`, `warning`, and
`error` outputs the same way.

The load-bearing correctness fix is making `decodeMessage` accept
`Uint8Array` rather than only `Buffer`: Node's `worker_threads`
`postMessage` structured-clones the payload, which strips the `Buffer`
prototype on the receive side. A frame sent as `Buffer` arrives as a
plain `Uint8Array`, and `Buffer.isBuffer(raw)` returns false — so the
first attempt at U17 (gating decode on `Buffer.isBuffer`) silently
treated every incoming frame as POJO and the worker never responded.
The fix adopts the underlying memory zero-copy via
`Buffer.from(view.buffer, view.byteOffset, view.byteLength)` and uses
`raw instanceof Uint8Array` at every call site (parse-worker decode,
pool dispatch handler, pool ready-handshake handler, FakeWorker test
mocks, and the integration-test worker preamble).

The pool stays tolerant of POJO incoming so unit-test FakeWorkers
don't need rewriting — only the new outgoing encoded dispatches require
the test scaffolding to decode on receive, which the test FakeWorkers
and the integration test's inline `parentPort.on` wrapper now do.

The slot-drop integration test was rewritten from a shared-counter-file
race (which pre-U17 timing happened to land on the assertion-friendly
counter==2 endpoint, but post-U17 protocol decoding latency shifted to
counter==1 and produced 3 quarantines instead of 2) to a deterministic
path-based crash trigger: slot 0 crashes on a.ts, respawns, crashes on
the requeued b.ts, slot is dropped after budget exhausted; slot 1
handles [c.ts, d.ts] normally. Outcome no longer depends on inter-worker
file-write ordering.

Protocol coverage adds two regression tests pinning the Uint8Array
decode path: structured-clone-stripped frames decode identically to
their Buffer originals, and Uint8Array views with non-zero byteOffset
into a wider ArrayBuffer also decode correctly (catches `Buffer.from(uint8)`
copying semantics if a future refactor loses the zero-copy adoption).

All 94 worker-pool tests (9 files, unit + integration) pass; the full
unit suite (6128 tests across 268 files) passes unchanged.

* perf(workers): zero-copy file content transfer via transferList (U19)

Pool dispatch now hoists `{path, content: string}[]` file contents OUT
of the U17 JSON envelope into separately-allocated `Uint8Array`s whose
ArrayBuffers are passed to `worker.postMessage`'s `transferList` for
zero-copy ownership transfer. The envelope itself carries only
lightweight metadata (`{path, byteLength}` per file) and is structure-
cloned the same as before.

What this saves vs U17 baseline:

- **JSON.stringify of file contents on main thread** drops to zero —
  the envelope is now O(paths + sizes), not O(total bytes). For a 200-
  file sub-batch of 10 KB TS files, that's ~2 MB of escape processing
  per dispatch that disappears. JSON.stringify's per-character branch
  on quotes/backslashes/control chars is roughly 2x slower than
  UTF-8 transcode in TextEncoder, so the replacement is a CPU win
  even though it adds a single TextEncoder.encode per file.
- **Structured-clone memcpy of file contents** drops to zero — the
  contents' backing ArrayBuffers are ownership-transferred, not copied
  into the worker's heap. The envelope's struct-clone cost is now
  proportional to metadata size only.
- **JSON.parse on worker thread** likewise no longer scales with
  content size. Worker decodes each `Uint8Array` to string via
  `TextDecoder` lazily at the parse boundary — runs on the worker
  thread, parallel with continued main-thread work, vs U17's
  sequential JSON.parse blocking the worker before processBatch can
  start.

Pipelining: TextEncoder.encode (main) and TextDecoder.decode (worker)
can both run while the OTHER side is doing useful work. Under U17,
struct-clone was a synchronous main-thread blocker.

The ArrayBuffer ownership contract is load-bearing:

- File-content `Uint8Array`s are allocated via `TextEncoder.encode`,
  NOT `Buffer.from(str, 'utf8')`. TextEncoder produces a dedicated
  ArrayBuffer per call; `Buffer.from(str)` carves from Node's shared
  `Buffer.poolSize` slab for small strings, so transferring one
  pool-backed Buffer's ArrayBuffer would detach every other Buffer
  that shares that slab — silent data corruption.
- The envelope itself is NOT transferred. It MAY be pool-backed by
  `encodeMessage`, and at ~30-80 bytes/file the struct-clone cost is
  negligible. Not transferring avoids the same detach-collateral risk
  the contents path is careful to dodge.

Detection is strict: every input element must have both `path: string`
and `content: string`. A single non-conforming element disqualifies
the whole batch from the transfer path and falls back to the legacy
single-Uint8Array `encodeMessage` envelope. Safer than partial
transfer (which would split a sub-batch into mixed-shape messages
the worker can't reassemble).

`parse-worker.ts` `decodeIncomingMessage` recognizes the hybrid
`{envelope, contents}` shape, decodes the envelope, zips metadata
positionally with the contents array, decodes UTF-8 → string per file,
and hands the reassembled `ParseWorkerInput[]` to the existing
`processBatch`. Identical downstream behavior to U17 — the IPC
optimization is invisible above this line.

Test scaffolding (3 FakeWorkers + 1 integration-test preamble) gain a
`decodeDispatchedMessage` helper that tolerates BOTH shapes (legacy
single-frame Uint8Array AND the new hybrid envelope+contents) so the
in-process unit mocks keep their existing action-scripting API and the
9 ad-hoc integration test workers keep their `msg.type === 'sub-batch'`
handlers unchanged.

`buildDispatchMessage` is now exported from worker-pool.ts so its
contract can be tested in isolation. A new
`test/unit/worker-pool-transferlist.test.ts` pins:
  - hybrid shape produced for parse-worker inputs
  - transferList carries one ArrayBuffer per file in input order
  - envelope decodes to metadata only (no `content` field)
  - content bytes round-trip byte-for-byte through UTF-8 (ASCII,
    multi-byte, surrogate-pair emoji)
  - each content's ArrayBuffer is independently allocated (no pool
    sharing) — the load-bearing transfer-safety invariant
  - non-parse shapes, empty arrays, and mixed-conformance arrays all
    fall back to the legacy single-frame path

All 271 test files (6166 unit + integration tests) pass.

* fix(workers,tests,docs): apply ce-code-review findings (16 items)

Walks the full set of findings from a multi-agent code review (11
reviewers, 1 maintainability dispatch lost to tool-permission denial)
of the PR #1693 branch. All 16 actionable findings — 4 P1, 4 P2,
8 P3 — applied in a single pass against a consistent tree. Tests
pass (269/269 unit files, 29/29 integration).

P1 — bounds-only / disguised-bounds assertions across 4 test files
(per user-memory DoD §2.7):
  - worker-pool.test.ts: 5 sites — `nodes.length > 0` dropped (redundant
    after `.toContain('validateInput')`); `files.length >= 4` pinned to
    `.toBe(7)` (mini-repo/src has exactly 7 .ts files); `results.length
    > 0` pinned to `.toHaveLength(1)` (default sub-batch absorbs all 7);
    `result.fileCount >= 0` pinned to `.toBe(1)` (empty file is still
    "processed"); `warnRecords.length > 0` replaced with content-
    predicate `/respawn|dropping|replacement|did not report ready/`
    (catches silenced warnings); `fallbackExcludePaths.length > 0`
    pinned to exact `['one.ts', 'two.ts']` (deterministic given the
    single-slot pool + 2 items + per-item starting-file).
  - parse-impl-fallback.test.ts: 3 sites — `astCacheClearCalls >= 1`
    pinned to exact 4 (per-chunk × 2 + finally × 2); the two error-path
    delta checks pinned to exact +2 and +3 (verified empirically).
  - parse-impl-progress-monotonic.test.ts: `percents.length > 0` →
    `.not.toEqual([])`; per-element `Math.max(prev, cur)` tautology
    replaced with direct `if (cur < prev) throw`; final-percent
    `Math.min(last, 95)` tautology pinned to exact `.toBe(70)` (3-file
    skipWorkers fixture's deferred band lands at the band start).
  - parse-impl-large-fixture.test.ts: `Math.min(elapsedMs, BUDGET)`
    tautology removed; Promise.race rejection is the load-bearing
    wall-clock check.

P1 — terminate() lacks `.catch` mask:
  - worker-pool.ts terminate() now matches the `.catch(() => undefined)`
    pattern used at every other internal terminate site. Prevents a
    hung/OOM worker's terminate rejection from masking the original
    pipeline error when called from parse-impl.ts's finally block, and
    guarantees `workers.length = 0` / `activeSlots.clear()` always run.

P1 — hybrid envelope length-mismatch + null-payload silent data loss:
  - parse-worker.ts decodeIncomingMessage: explicit non-null-and-typed
    check before `.type` access (decodeMessage permits null payloads
    per encodeMessage contract); explicit length-equality assertion
    between `decoded.files` and `contents` before zipping. Without
    these, `TextDecoder.decode(undefined)` silently returns "" and
    produces empty-content graph nodes — a contract violation that
    used to be undetectable. Both throws route through the outer
    try/catch → worker `error` reply → pool's recoverAndResume.

P1 — unsafe casts at the IPC boundary:
  - buildDispatchMessage now uses a properly-typed `isParseWorkerItemArray`
    type guard. The narrowed branch accesses `item.path` and
    `item.content` as statically-typed strings — a future rename of
    `ParseWorkerInput.content` would fail to compile inside the branch
    instead of silently mismatching at runtime. The remaining
    decodeMessage payload casts are bounded by the F3/F6 runtime
    guards.

P2 — idle-timeout retry bypasses circuit breaker:
  - worker-pool.ts timeout-retry IIFE now increments
    `consecutiveFailuresPerSlot[workerIndex]` alongside `respawnCount`.
    A slot that consistently times out (vs crashes) now trips the
    per-slot breaker, instead of consuming its full respawn budget
    over potentially tens of minutes without the breaker firing.

P2 — null/non-object worker message crashes pool handler:
  - Dispatch handler in worker-pool.ts now guards `null /
    non-object / no string type discriminant` before `msg.type` access
    and routes through recoverAndResume on violation. Previously a
    legitimate `null` payload would throw TypeError out of the
    EventEmitter listener → uncaughtException on main, crashing the
    analyze.

P2 — workerPoolSize === 0 creates unusable pool:
  - parse-impl.ts now treats `workerPoolSize === 0` as `skipWorkers`
    at the gate. Matches the PipelineOptions docstring contract ("0
    disables the pool entirely — equivalent to skipWorkers"); avoids
    constructing a pool that rejects every dispatch and logs
    "Worker pool parsing stopped" per chunk.

P2 — encodeMessage 2-buffer allocation per frame:
  - protocol.ts encodeMessage coalesced to a single
    `Buffer.allocUnsafe + writeUInt8 + writeUInt32LE + buf.write
    (string, offset, 'utf8')`. Drops the intermediate
    `Buffer.from(JSON.stringify(...), 'utf8')` allocation + memcpy.
    Length pre-check via `Buffer.byteLength(string, 'utf8')` surfaces
    the uint32 cap before any allocation.

P3 — slotGenerations made optional on WorkerPoolStats so external
  implementations of getStats() that predate U12 don't compile-break;
  in-repo callers already use optional chaining.

P3 — buildDispatchMessage marked `@internal` so it isn't surfaced as
  public API by typedoc / api-extractor (it's a test-only export).

P3 — verboseThroughputLog hoisted above the chunk loop (env vars can't
  change mid-run; one O(env-read) per analyze, not per chunk).

P3 — corrected the messageerror routing comment in worker-pool.ts
  dispatch handler. `ProtocolDecodeError` is caught by the surrounding
  try/catch — distinct from `messageerror`, which fires for V8
  structured-clone failures before the message body would reach the
  handler.

P3 — initial pool spawn now uses a `Promise.allSettled` ready-handshake
  gate symmetric with `replaceWorker`. Dispatch awaits this gate before
  selecting slots, so an init-crashing initial worker is dropped from
  `activeSlots` and a downstream OOM/missing-native-binding failure
  surfaces in seconds (bounded by WORKER_READY_TIMEOUT_MS) rather than
  waiting for the first idle timeout (30s default).

P3 — `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT`,
  `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS`,
  `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` added to:
    - CLI `--help` text in src/cli/index.ts
    - Root README env-var table
    - gitnexus/README troubleshooting section (new "Worker pool
      resilience tuning" subsection)

P3 — CLI `catch (e: any)` / `catch (err: any)` in analyze.ts replaced
  with `catch (err: unknown)` + narrowed access; matches modern TS
  best practice and the codebase pattern at other catch sites.

P3 — `WorkerPoolStats.terminated: boolean` field added (optional, for
  backward compatibility). `terminate()` sets it true; `getStats()`
  surfaces it. Distinguishes graceful shutdown from a circuit-breaker
  trip in observability surfaces.

Coverage / advisory items not addressed in this commit (kept in the
report only):
  - maintainability reviewer failed (Read/Bash denied) — god-module
    audit on worker-pool.ts (~1400 LOC) carried as residual risk
  - quarantine case-sensitivity contract unpinned (adversarial #8)
  - WORKER_READY_TIMEOUT_MS env-configurability (adversarial #2)
  - chunk-byte-budget × parseChunkConcurrency memory multiplier doc
    (adversarial #5)
  - MCP discoverability gaps for env vars / verbose (agent-native W1/W2)
  - bench/parse-throughput.md scaffold-with-TBD-rows (PS RR-003)

* fix(parsing): sequential gap-fill for worker-quarantined chunk files (U20.U1)

When the worker pool's Layer 3 quarantine filters one or more files
out of a chunk's dispatch, the worker results returned to
processParsing are silently narrower than the input chunk. Without
this reparse, the graph for this run would be missing every quarantined
file's symbols/imports/calls/heritage with no failure signal.

After the existing per-chunk quarantine log emits in
processParsing's worker-path try-block, run processParsingSequential
on JUST the quarantined-in-chunk files. The sequential path writes
directly to the graph, so symbols for those files land alongside
worker output for the surviving files.

Mirrors the WorkerPoolDispatchError catch-block's processParsingSequential
call shape — same signature, same args, same scopeTreeCache wiring.
Emits a structured warn naming `reparsedPaths` so operators can
observe the sequential fall-through.

This fixes the in-run side of the corruption Codex's adversarial
review of PR #1693 flagged. The cross-run side (chunk-cache
poisoning) is closed by U20.U2 in a follow-up commit.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* fix(parse-impl): suppress chunk-cache write when any chunk file was quarantined (U20.U2)

The chunk hash at parse-impl.ts:424-428 is computed from every file
in the chunk. The worker pool's Layer 3 quarantine
(worker-pool.ts createQuarantine) filters quarantined files out of
dispatch, so `rawResults` reflects only the surviving files. Before
this commit, the write at line 500-507 stored that partial result
under the full-coverage chunk hash — and on the next analyze with
unchanged content, the cache HIT branch (line 439-464) silently
replayed the incomplete result. Symbols from the quarantined file
were missing from the graph for as long as the cache survived.

Codex's adversarial review of PR #1693 flagged this as a silent-
corruption class because there's no failure signal: no warn log
during the replay, no graph-equivalence check, no exit code change.
The corruption only surfaces if an operator notices a missing symbol
in `gitnexus_query` output.

Guard the write with `chunkFiles.some(f => quarantineSet.has(f.path))`.
When any chunk file is in the worker pool's cumulative quarantine
snapshot, skip the `parseCache.entries.set` call. Emits a verbose-
only info log so operators investigating "why aren't my chunks
caching" have a diagnostic trail.

Skipping the write means the next analyze gets a cache miss for this
chunk and re-dispatches it. Quarantine is session-scoped (a fresh
createWorkerPool starts with an empty quarantine), so the new pool
gives the quarantined file another chance. If quarantine fires again,
U20.U1's sequential gap-fill still produces a complete graph for that
run; the cache stays empty for the chunk until a fully-clean
dispatch lands.

The cache-hit replay branch at parse-impl.ts:439-464 is unchanged.
Its contract strengthens: "cache entries are complete" becomes true
post-fix, but the replay code doesn't need to know that.

Closes the cross-run side of the Codex finding. U20.U3 adds the
regression test.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* test(parse-impl): integration regression for quarantine + chunk-cache (U20.U3)

Pins the U20 fix end-to-end via REAL `worker_threads` + `createWorkerPool`.
Mirrors the writeReadyWorker pattern from `test/integration/worker-pool.test.ts`
— inline READY_PREAMBLE + custom test worker script that:

  1. Decodes the U17/U19 IPC protocol (Buffer frame OR hybrid envelope/
     contents shape) the same way the production parse-worker does.
  2. Emits a `{type:'ready'}` handshake so the pool's
     `waitForWorkerReady` resolves promptly.
  3. On a sub-batch containing `poison.ts`, emits starting-file +
     `process.exit(134)`. The pool attributes the death to `poison.ts`
     via the in-flight signal and adds it to the session-scoped
     quarantine.
  4. On a sub-batch without poison, synthesizes a minimal valid
     `ParseWorkerResult` with one `Function` node per file (no
     tree-sitter dep in the test worker — the synthesized nodes give
     `mergeChunkResults` deterministic content for the graph).

Assertions exercise both fix layers:

  - U1 (sequential gap-fill in processParsing): the graph contains a
    `Function` node named `poison` AFTER the run. The custom worker
    never emits anything for `poison.ts`, so the only path for that
    symbol to reach the graph is `processParsing`'s sequential
    reparse of the quarantined-in-chunk file using the real
    tree-sitter parser against the actual source.
  - U2 (cache-write suppression in runChunkedParseAndResolve):
    `parseCache.entries` does NOT contain the chunk hash after the
    run; `parseCache.usedKeys` DOES contain it (chunk processed,
    cache write specifically skipped).
  - Cross-run: a second pass over the same fixture with the same
    parseCache and a fresh worker pool re-dispatches the chunk
    (cache empty), the worker crashes again, sequential gap-fill
    runs again, and the cache stays empty. Pins the round-trip
    contract.

Adds `workerUrlForTest?: URL` to PipelineOptions — same `@internal`
test-only injection precedent as `workerThresholdsForTest` (already
in PipelineOptions for thresholds). When set, parse-impl uses the
provided URL instead of the src/ → dist/ resolution dance. Production
call sites never set this field; the only consumer today is this
integration test.

Why integration over unit:
  - The fix lives at the boundary between parsing-processor.ts and
    parse-impl.ts under a real WorkerPool. Unit-mocking the
    worker-pool module bypasses the structured-clone boundary, the
    dispatch lifecycle, and the actual quarantine flow — it verifies
    the test setup rather than the contract. The real worker thread
    executing through the U17/U19 IPC protocol IS the load-bearing
    surface.
  - User-explicit preference (saved as
    feedback_integration_over_vimock.md memory). For worker-pool /
    parse-impl / IPC-touching code: write integration tests under
    test/integration/ using writeReadyWorker patterns; avoid
    vi.mock on worker-pool.js.

Test wall-clock: under 2s; both `it` blocks together complete in
~1.8s under the existing CI conditions.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* refactor(parsing): remove sequential-parser fallback (U20 design pivot)

The worker pool's resilience layers — respawn budget, circuit breaker,
quarantine, slot-attribution, cumulative timeout — are now the SOLE
contract for handling worker failures. Two sequential-reparse paths
are removed from processParsing:

1. **U20.U1 sequential gap-fill for quarantined chunk files** (just
   added in commit 7dd489e9, now reverted). The pre-emptive rescue
   would re-run processParsingSequential on the file that ALREADY
   killed a worker — which for the most common quarantine cause
   (tree-sitter native SIGSEGV on a pathological file) re-triggers
   the same native crash on the main thread, killing the entire
   analyze. The "rescue" turned silent missing-symbols into a louder
   analyze-wide crash. Drop the rescue; accept the per-run gap.

2. **Pre-existing WorkerPoolDispatchError catch-block sequential
   fallback** (in production since PR #1693's resilience layer
   landed). Same risk class — when the pool exhausts its respawn
   budget / trips the circuit breaker, the failing files are
   precisely the ones likely to crash a sequential parser too. The
   "graceful degradation" hid pool failures behind degraded-but-
   completing analyze runs, making operational issues harder to
   surface and diagnose. Drop the catch-block; WorkerPoolDispatchError
   propagates to the analyze entry point where the user sees a clear
   hard signal.

What stays:
- The `skipWorkers: true` / small-repo path that uses
  `processParsingSequential` as the EXPLICIT primary path (not a
  fallback). Caller-driven opt-out and tiny-repo perf optimization
  are different intents.
- U2's chunk-cache write suppression in parse-impl.ts (commit
  7c9c9556). When quarantine fires, the chunk stays uncached so the
  next analyze with a fresh pool retries the file cleanly. That's
  the cross-run correctness Codex's adversarial review actually
  asked for.
- The per-chunk quarantine warn log (parsing-processor.ts) — operators
  see which files were skipped, both immediately and across runs.

What changed:
- `processParsing` worker-path try-block: unwrapped. The
  `processParsingWithWorkers` call is now direct (no try/catch
  wrapping); errors propagate to the chunk-loop caller.
- `parsing-worker-fallback.test.ts` rewritten: the previous 5 tests
  asserted graceful sequential-fallback behavior. Replaced with 3
  tests pinning the new contract — raw Error propagates, WorkerPool-
  DispatchError propagates with fallbackExcludePaths intact, normal
  quarantine signal does NOT throw and surfaces via progress detail.
- `parse-impl-quarantine-cache-skip.test.ts` (U20 integration test)
  updated: poison.ts is NOT in the post-run graph; surviving files
  are; chunk-cache stays empty; second pass re-dispatches and leaves
  cache empty.
- Plan doc updated to mark R1 as dropped and explain the U20 pivot
  in the Summary.

User decision: explicit directive ("let's remove the sequential
fallback entirely we must rely on entirely that the parallel process
is resilient enough to work itself through the code base"). The pool's
resilience layers are designed for this — respawn budget, circuit
breaker, quarantine, slot-generation, cumulative-timeout cap — and
adding a layer below them was redundant insurance with real downside.

Tests: 269/269 unit files (6135 tests) green. 31/31 worker-pool +
parse-impl integration tests green. The 2 reported "errors" in the
integration run are the pre-existing intentional-process.exit unhandled-
exception leaks from test workers — unchanged by U20.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* fix(workers,tests,docs): address ce-ultrareview findings F1/F2/F3/F4

Multi-lane review run on the PR #1693 branch surfaced four addressable
items beyond the blocking three.

F1 (minor, CodeQL): unused `findMatch` helper in
test/unit/scope-resolution/typescript/typescript-captures-anchor.test.ts:28
removed. `countMatchesTsx` flagged by the same CodeQL pass is a false
positive — it's called at line 88 by the JSX-anchor regression tests
so the rewrite case actually fires under TSX, not just TS.

F2 (medium, docs): bench/parse-throughput.md retitled as
"(scaffold)" with an explicit "no measurement data has been collected
yet" note above the table. The self-contradictory "Regenerate this
file before merging any PR that touches the ingestion pipeline"
instruction is dropped — the file ships intentionally without
numbers; the load-bearing perf-regression protection lives in
test/integration/parse-impl-large-fixture.test.ts (U6, 30s
Promise.race wall-clock budget). The Latest measurement section now
preserves the ~6s sequential observation as a smoke reference, not as
a regression target.

F3 (low, API hygiene): `WorkerPoolDispatchError.fallbackExcludePaths`
renamed to `quarantinedPaths`. The "fallback" terminology was
load-bearing under the pre-U20 design when `processParsing`'s
sequential-fallback catch-block consumed it to filter the fallback
file list. After commit be1f65c removed that catch-block, no
production code reads the field — but it stays populated by the pool
because the snapshot is genuinely useful operator diagnostics when
the breaker trips. The rename clarifies the field's actual semantics
(here are the files the pool quarantined before it tripped) without
changing wire behavior. Definition + the lone surviving in-pool
comment reference + both test assertions updated.

F4 (low → real fix, reliability): timeout-retry IIFE in
worker-pool.ts now consults `consecutiveFailureThreshold` and trips
the circuit breaker when the per-slot consecutive-failure count
crosses it. Closes a gap left by ce-code-review's REL-02 patch — that
fix added the `consecutiveFailuresPerSlot[workerIndex]++` increment
in the timeout-retry path but did NOT add the corresponding
threshold-check + tripBreaker call. Result: chronic pure-timeout
deaths accumulated counts that never tripped the breaker until the
slot also hit `respawnCount > maxRespawnsPerSlot`. Now timeouts and
crashes are structurally treated the same way by the breaker, which
is what the REL-02 increment was meant to enable. Test coverage:
worker-pool-resilience.test.ts already exercises the breaker via the
shared handleWorkerDeath path; this new branch traces the same
trip semantics with a different entry point, so the breaker-tripped
state is observable via the same `getStats().poolBroken` and
`WorkerPoolDispatchError.quarantinedPaths` surface.

Out of scope here (caller actions or future PRs):
  - F5 (info): cumulative-quarantine cache check is safe in practice
    because chunks are alphabetically deterministic; no action.
  - F6 (low): exit-code-0 quarantine exemption — pre-existing P2
    residual, bounded by quarantine + respawn budget; deferred.
  - F7 (info): dispatch non-reentrancy contract documented but not
    enforced; no production caller violates it; deferred.
  - PR title `[WIP]` removal — happens on GitHub side.

Tests: 274/274 test files (6185 passing, 30 skipped). The single
"error" in the integration runner is the pre-existing intentional-
process.exit unhandled-exception leak from the deliberate startup-
crash test worker, unchanged by these fixes.

* fix(workers): swap protocol body from JSON to V8 serialize/deserialize

CI scope-parity tests on Ubuntu surfaced silent data loss in the
worker IPC: `Phase 'scopeResolution' failed: scope.typeBindings is not
iterable` (Python, Go) and `importerModule.typeBindings.has is not a
function` (Python). Plus three #1066 large-file regression tests
(Python / C# / TypeScript) failed because call relationships weren't
resolving from the worker output.

**Root cause:** U17 introduced `JSON.stringify`/`JSON.parse` as the
protocol body codec. JSON has no representation for `Map`, `Set`,
`Date`, `RegExp`, `BigInt`, `TypedArray`, `undefined` values, or
circular refs — `JSON.stringify(someMap)` returns `"{}"`. Production
scope-resolution code keys data structures on Maps throughout
(`ParsedFile.scopes[*].typeBindings: ReadonlyMap<string, TypeRef>`,
plus `bindings`, `bySourceScope`, `byTargetDef`, the finalize-algorithm
edge indexes, etc.). The JSON round-trip silently turned every Map
into an empty object, manifesting downstream as iteration / `.has`
calls failing on the decoded payload.

**Fix:** replace the JSON body with `node:v8`'s `serialize` /
`deserialize`. That's the same structured-clone algorithm Node's
`worker.postMessage` uses natively — bit-for-bit compatible with the
pre-U17 implicit-clone path. Full type fidelity for Map, Set, Date,
RegExp, BigInt, TypedArray, undefined values, and circular refs. No
external dependency.

A previous iteration of this fix attempted to bolt a Map/Set
replacer+reviver onto the JSON path. Rejected in favor of V8
serialization because:
  - the JSON tag-marker approach requires per-type registration
    (Map, Set; then Date, RegExp, BigInt would each need their own
    sentinels); V8 handles them all uniformly
  - keys to JSON-encode would still need handling for nested types
    (and the marker approach doesn't survive nested Maps-in-Maps
    cleanly without recursive replacer logic)
  - V8 is faster than JSON for object-heavy payloads anyway (binary
    format, no string escaping pass)
  - the user-explicit ask was "a much more generic solution that will
    work for everything" — V8 serialization IS the generic solution

Trade-offs documented in the module header:
  - body bytes are opaque (binary, not human-readable) — debugging
    requires `v8.deserialize` ad-hoc; protocol.test.ts exercises every
    supported MessageTag including the new type-fidelity cases as a
    regression net.
  - format is tied to the running Node major. Pool always spawns
    workers on the same Node instance the main thread runs, so this is
    moot in production. Would matter if frames ever persisted to disk
    (nothing does today).

Protocol test file rewritten:
  - drops the JSON-specific byte-layout assertions (e.g. `body must
    equal "null" string`) — replaced with V8-derived expected lengths
  - adds a "structured-clone type fidelity" describe block that pins
    Map, nested Map, Set, Date, RegExp, BigInt, TypedArray, undefined
    values, and circular-ref round-trips. These are the load-bearing
    regression tests preventing a future "optimize" PR from quietly
    swapping V8 back to JSON.
  - the bad-body decode-error test now uses arbitrary non-V8 bytes
    instead of `{not-json}` — same intent.

Integration test READY_PREAMBLEs (worker-pool.test.ts and
parse-impl-quarantine-cache-skip.test.ts) update their inline
decoders to use `v8.deserialize` matching the production codec.
Both files have a standalone CJS worker preamble that can't import
dist/protocol.js by relative path, so the V8 dependency is required
via `node:v8` directly.

Tests: 271/271 unit files (6163 tests + 30 skipped). 28/28
worker-pool integration. 3/3 parse-impl integration. 791/791
scope-parity tests (the four CI-failing files: python.test.ts,
go.test.ts, typescript.test.ts, csharp.test.ts) all green again.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* refactor(workers): drop protocol.ts; use native postMessage + transferList

The protocol.ts framing layer was redundant — Node's `worker.postMessage`
already runs V8 structured-clone internally, the same algorithm that
backed `v8.serialize`. Wrapping V8.serialize → Buffer →
postMessage(struct-clone-Buffer) was a double-walk: one full
structured-clone pass to produce the Buffer, then another pass when
postMessage cloned that Buffer across threads. This commit cuts the
wrapper layer; workers and pool exchange POJO directly via
`worker.postMessage(value, transferList)`, with file-content
`ArrayBuffer`s in `transferList` for zero-copy ownership transfer.

What changes:

- **Deleted** `src/core/ingestion/workers/protocol.ts` (~180 LOC) +
  `test/unit/workers/protocol.test.ts` (~250 LOC). The MessageTag
  enum / ProtocolDecodeError / encodeMessage / decodeMessage surface
  is gone. Tag-based routing is replaced by the `msg.type`
  discriminant that every receive site already checks. Protocol-decode
  errors map to Node's `messageerror` event (V8 deserialization
  failures during postMessage), which the pool already wires to
  `recoverAndResume`.
- **`worker-pool.ts`**: `decodeIncomingWorkerMessage` removed; handlers
  receive POJO directly. `buildDispatchMessage` now returns
  `{message: {type:'sub-batch', files: [{path, content: Uint8Array}]},
  transferList: ArrayBuffer[]}`. The Uint8Array-per-content allocation
  via `TextEncoder.encode` is preserved (it's the load-bearing
  transfer-safety contract that keeps content out of Node's shared
  `Buffer.poolSize` slab). Flush dispatch is now plain
  `worker.postMessage({type:'flush'})`.
- **`parse-worker.ts`**: `decodeIncomingMessage` removed. The message
  handler receives POJO directly; the only conversion is
  `Uint8Array → string` for sub-batch file contents at the
  `decodeSubBatchFiles` boundary, before handing to `processBatch`.
  Outgoing messages are emitted as POJO via plain
  `parentPort.postMessage({type:'starting-file', ...})` etc. The
  `sharedHybridDecoder` is now `sharedContentDecoder` (same intent,
  clearer name for the simpler shape).
- **Test scaffolding**: FakeWorkers in `worker-pool-resilience`,
  `worker-pool-windows-quarantine`, and `worker-pool-slot-generation`
  drop their `decodeMessage` import + `decodeDispatchedMessage` helper.
  The helpers stay (still convert `files[i].content` Uint8Array →
  string for test-action introspection) but no longer touch any
  protocol framing — just shape-check for sub-batch.
- **Integration READY_PREAMBLEs** (worker-pool.test.ts and
  parse-impl-quarantine-cache-skip.test.ts): drop the inline
  v8.deserialize + envelope-unzip logic; the preamble is now just
  the ready handshake + a `parentPort.on` wrapper that converts
  `files[i].content` Uint8Array → string for the ad-hoc test worker
  scripts.
- **`worker-pool-transferlist.test.ts`**: contract tests updated for
  the new buildDispatchMessage shape — no `envelope` field anymore;
  `message.files[i].content` is Uint8Array; transferList holds each
  content.buffer in input order. Pool-slab independence still pinned.

What stays the same:

- Zero-copy file-content transfer via transferList — every file's
  ArrayBuffer is ownership-transferred to the worker (no copy).
- Full structured-clone type fidelity — Map / Set / Date / RegExp /
  BigInt / TypedArray / undefined / circular refs all preserved by
  Node's native postMessage. The V8 fix from commit 06f6957e is
  inherent in this path; there's no JSON layer to lose them.
- TextEncoder-per-content allocation — keeps content buffers out of
  the shared `Buffer.poolSize` slab so transferring one cannot detach
  another.
- The pool's resilience layers (respawn, breaker, quarantine,
  starting-file attribution, cumulative timeout, ready handshake,
  slot-generation guard) — unchanged.
- U20 chunk-cache write suppression on quarantine — unchanged.

Net: ~430 LOC removed (protocol.ts + tests + inline decoders + helpers),
~120 LOC simplified in worker-pool.ts and parse-worker.ts. One less
serialization pass per message on the hot path.

Tests: 270/270 unit files (6133 + 30 skipped). 822/822 integration
tests including the four CI-failing scope-parity files (Python, Go,
TypeScript, C#) — the V8-fidelity contract holds via native
postMessage with no explicit serializer. The single "error" reported
in worker-pool.test.ts is the pre-existing intentional
process.exit unhandled-exception artifact from the deliberate
startup-crash test, unchanged by this commit.

* refactor(parse-worker): drop legacy single-message dispatch mode

The `parentPort.on('message', ...)` handler had an `Array.isArray(msg)`
branch left over from a pre-sub-batch dispatch shape — the pool used
to send the items array directly, before the worker pool added
sub-batching and the `{type:'sub-batch', files: ...}` envelope.

No production caller has dispatched that shape since the sub-batching
refactor landed; verified by grepping the repo for `postMessage([`
patterns (zero matches). The `ParseWorkerInput[]` arm in the
`WorkerIncomingMessage` discriminated union also blocked
exhaustiveness narrowing — flagged by the kieran-typescript code
review (RR-01) as "if a future unit removes the legacy array path,
this arm should be dropped." Dropping it now.

What changes:
  - Remove the `Array.isArray(msg)` branch from the message handler.
  - Drop `ParseWorkerInput[]` from the `WorkerIncomingMessage` union;
    it's now a clean `{type:'sub-batch'} | {type:'flush'}` discriminated
    union, so the dispatch switch is exhaustive over `msg.type`.

Tests: 71/71 worker-pool unit + integration tests green (resilience,
slot-generation, windows-quarantine, transferlist, parsing-worker-
fallback, worker-pool integration, parse-impl-quarantine-cache-skip).

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 20:39:35 +01:00
azizur100389andGergő Magyar aa8f4d6efe fix(group): Union HTTP graph and source contracts (#1709)
* Union HTTP graph and source contracts

* test(group): Document HTTP source union follow-ups

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 17:44:07 +01:00
Shane Thurston Wijaya 4d2ed0e525 fix(eval-server): localhost now doesn't normalize into IPv4 instead lets OS decide which to bind (#1722)
* fix(eval-server): localhost now doesn't normalize into IPv4 instead lets OS decide which to bind

* fix(eval-server): EADDRNOTAVAIL now treats as potential IPv6

* test(eval-server): new integration test for --host localhost

* docs(eval-server): updated eval/README.md based on latest update

* fix(eval-server): clarify EADDRNOTAVAIL diagnostic, guard server.address(), and soften localhost docs
2026-05-20 16:14:13 +01:00
df1882d36b fix(ingestion): surface skipped large-file paths by default (#1659) (#1661)
* fix(ingestion): surface skipped large-file paths by default (#1659)

The 512 KB skip threshold in filesystem-walker is necessary, but the
existing warning only said "Skipped N large files" with no paths unless
GITNEXUS_VERBOSE=1 was set. In a repo with one or two oversized first-
party source files (e.g. a 17K-line cron handler), every IMPORTS/CALLS
edge from that file silently disappeared and the surface looked like a
Python resolver bug. Issue #1659 was filed against the resolver for
exactly that reason, but the resolver was fine; the file was being
dropped before parse.

Changes:
  * Always print up to 5 skipped paths after the count line.
  * If more than 5 were skipped, append "...and N more" with a hint to
    set GITNEXUS_VERBOSE=1 for the full list.
  * When running at the default threshold, emit a one-line hint about
    GITNEXUS_MAX_FILE_SIZE=<KB> so operators know how to widen it.
  * Cover the new behavior with three additional tests in the existing
    filesystem-walker integration suite, plus a new describe block for
    the >5 preview-cap case.

Verified end-to-end on a 680-file Python repo that hit #1659: before
the patch, "Skipped 3 large files (>512KB, ...)" was the only signal
and impact upstream of a function called from cron.py returned 1 of 5
real callers; after the patch the cron file is listed by name with the
hint, and running with GITNEXUS_MAX_FILE_SIZE=1024 brings the missing
callers back (impactedCount 1 -> 9).

* fix(ingestion): address #1661 adversarial review follow-ups (F1/F2/F3)

Three non-blocking nits flagged by the adversarial review on #1661:

F1 (output stability) — skippedLargePaths was populated by concurrent
fs.stat callbacks in batches of 32, so push order within a batch was
completion-order rather than input-order. The default preview's "first
5" could vary across runs on the same repo. Fix: sort the array before
slicing. New test asserts the verbose output is in sorted order.

F2 (boundary coverage) — the preview-cap describe block created 8
large files, so the SKIPPED_PREVIEW_CAP = 5 comparison was never
exercised at the exact <= boundary. A future off-by-one (<= → <) would
not fail the suite. Fix: add two tests, one with exactly 5 files (all
listed, no truncation) and one with exactly 6 files (5 listed plus
"...and 1 more").

F3 (hint accuracy) — isDefault compared effective bytes, so an
operator who explicitly set GITNEXUS_MAX_FILE_SIZE=512 (the same KB as
the default) would still see the "Set GITNEXUS_MAX_FILE_SIZE=<KB>..."
hint. Fix: gate the hint on whether the env var is unset, not on the
resulting byte value. New test pins the explicit-default-value case.

All 34 filesystem-walker tests pass (was 30; +4 new). Prettier clean,
typecheck clean for the changed files.

---------

Co-authored-by: scotjelinski <58397194+scotjelinski@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 14:37:58 +01:00
f350ae278a feat: Add analyze --repair-fts, enforce FTS verification, and harden repair safeguards (#1720)
* Initial plan

* feat(analyze): add --repair-fts and verify FTS index rebuilds

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/dccb3673-af86-43aa-aede-2e1449399775

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor(fts): tighten repair/verify messaging and option naming

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/dccb3673-af86-43aa-aede-2e1449399775

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: highlight analyze --repair-fts vs --force in READMEs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/61edc967-debc-419f-9f51-aebf2ef08d22

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(analyze): guard repair mode against missing graph store

* fix(cli): reject --repair-fts with --force

* test(analyze): document repair-store fixture intent

* test(analyze): tidy repair failure fixtures and constants

* test(analyze): clarify mock constants in repair tests

* test(analyze): rename simulated missing-index constant

* test(analyze): clarify mocked graph shape in full-verify test

* refactor(analyze): finalize flag validation and test clarity

* test(skip-git): avoid hard failing when FTS extension is unavailable

* test(skip-git): log visible FTS-unavailable test skips

* test(skip-git): tighten FTS-unavailable error detection

* test(skip-git): simplify FTS-unavailable message checks

* test(skip-git): avoid HOME pointing at parent repo in fixture env

* fix(analyze): address Claude follow-up findings for repair guardrails

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* fix(repair-fts): clarify invalid graph-store preflight errors

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* test(analyze): strengthen assertions for conflict and missing-store errors

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* fix(repair-fts): make invalid graph-store type errors explicit

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* fix(repair-fts): improve graph-store type diagnostics

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 13:37:04 +01:00
azizur100389andGergő Magyar dae70a26ea feat(cpp): Add pointer nullptr ellipsis conversion ranks (#1708)
* Add C++ pointer null ellipsis ranks

* test(cpp): Strengthen pointer overload assertions

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 12:06:51 +01:00
CopilotandGergő Magyar b4a2a4b91e fix(ingestion): Prioritize same-module Java type resolution for duplicate FQNs across modules (#1712)
* Initial plan

* Fix Java same-name type resolution with same-module priority

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/df0843e3-e244-4e0f-a94a-311df3899bd0

* Refine Java ambiguity fallback safety check

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/df0843e3-e244-4e0f-a94a-311df3899bd0

* Remove Java-specific fallback from shared scope walkers

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e882906a-2c96-411e-94a5-123a345421a9

* Harden Java module key and ambiguous owner fallback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e882906a-2c96-411e-94a5-123a345421a9

* Add negative assertions for duplicate-FQN module edges

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e882906a-2c96-411e-94a5-123a345421a9

* Make Java same-module ordering path-agnostic

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Refine generic Java path-affinity ordering safeguards

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Polish Java path-affinity ordering clarity

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Simplify Java path-affinity ordering logic

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Revert legacy DAG Java ambiguity ordering changes

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/94e50cf2-9733-4e69-a0eb-9fd38cbdb589

* Skip duplicate-FQN Java assertions in legacy parity mode

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/1b560efa-1b3b-4697-b590-c6ef447f431e

* Tighten duplicate-FQN Java CALLS edge cardinality assertions

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/67c18f93-5e56-4b15-8404-cdf1be9b4485

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 08:00:03 +01:00
Nilotpal KashyapandGergő Magyar d7e1815aa3 fix(detect-changes): guard resolveWorktreeCwd against overriding a separately-indexed worktree (#1691)
* fix(detect-changes): guard resolveWorktreeCwd against overriding a separately-indexed worktree

When the repo registry entry points to a linked worktree (both main
checkout and worktree indexed separately), resolveWorktreeCwd was
incorrectly replacing the correct worktree repoPath with the server's
main-checkout launch directory. Both share the same canonical root so
the existing same-repo check passed, causing git diff to run from the
wrong directory and return 0 changes (issue #1659).

Fix: early-exit guard — if tryRealpath(repoPath) differs from
tryRealpath(getCanonicalRepoRoot(repoPath)), repoPath is itself a
linked worktree and is returned unchanged. Auto-detection only fires
when repoPath equals the canonical main-checkout root.

Also normalises the launchCanonical comparison in the auto-detect path
to use tryRealpath for cross-platform consistency.

Regression test: 'returns worktreeDir unchanged when repoPath IS a
linked worktree and launchCwd is the main checkout'.

* test(detect-changes): add worktreeA→worktreeB case and assumption comment

Cover the missing case from the production-readiness review:
repoPath = wt-A (indexed), launchCwd = wt-B (server on a different
linked worktree). The guard fires on repoPath being a worktree
regardless of launchCwd, so wt-A is returned unchanged.

Also add an inline comment documenting the assumption that repoPath
is a git root or linked-worktree root (not an arbitrary subdirectory),
as noted in Finding 2 of the review.

* refactor(detect-changes): validate repoPath is a git root before canonical comparison

Instead of relying on a comment asserting repoPath is always a git
root, call getGitRoot(repoPath) first. Only if the result matches
repoPath itself do we call getCanonicalRepoRoot and apply the guard.

This eliminates the over-classification risk for subdirectory repoPath
values and makes the assumption explicit in code. repoCanonical is
shared across both the guard and the auto-detect block.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 06:46:16 +01:00
dependabot[bot] 92ad0f5491 chore(deps): bump idna in /eval in the uv group across 1 directory (#1713) 2026-05-20 05:38:26 +01:00
LocallyInsaneDBandGergő Magyar 803f0bed5f fix(lbug): probe-then-load FTS extension on Windows (#1690) (#1692)
* fix(lbug): probe-then-load FTS extension on Windows (#1690)

The Windows skip-on-process.platform==='win32' guard in pool-adapter.ts
hard-skipped loadFTSExtension() for every Windows host, even when the
FTS extension binary was already present locally at
~/.lbdb/extension/<version>/win_amd64/fts/libfts.lbug_extension.

That left BM25 silently degraded on Windows hosts that had a working
extension on disk, with no error path — `gitnexus doctor` still reported
FTS as available, but query returned 0 BM25 hits.

This patch adds hasLocalWinFtsExtension() which probes
~/.lbdb/extension/*/win_amd64/fts/ before the Windows skip. When a binary
is on disk we call loadFTSExtension(..., { policy: 'load-only' }); the
crashing install path documented in #1199 / #1217 is never exercised at
query time, and LadybugDB's version-specific resolution combined with
the ExtensionManager's tryLoad try/catch handles stale or zero-byte
sibling version dirs cleanly (no dlopen attempted on a stale binary).
When no binary is on disk at all, we fall back to the upstream skip so
install-time SIGSEGV continues to be avoided.

Verified on Windows 10 + Node 22.19.0 + gitnexus 1.6.5 +
@ladybugdb/core 0.16.1 with the FTS extension cached at 0.16.0:

  * BM25 timing goes from 0 → ~250-326ms on previously-zero queries
  * gitnexus context / impact / cypher unaffected
  * Adversarial-mixed-state run (real 0.16.0 binary + zero-byte stubs at
    0.15.0, 0.16.1, 0.17.0): exits 0, no SIGSEGV, FTS resolves to the
    real 0.16.0 binary, BM25 returns real hits
  * Stub-only state at the resolution path (0.16.0, zero-byte): exits 0,
    emits "FTS extension unavailable; load-only policy: extension not
    pre-installed", FTS marked unavailable cleanly via markUnavailable
    in extension-loader.ts — no silent greenlight

Closes #1690

* test(lbug): cover hasLocalWinFtsExtension probe + format pool-adapter

- Export hasLocalWinFtsExtension and add lbug-pool-win-fts-probe.test.ts
  with 7 cases against a real tmpdir + os.homedir spy:
    * missing ~/.lbdb/extension dir -> false
    * extension root present but no version dirs -> false
    * one version dir with binary present -> true
    * zero-byte stub at probe path -> true (LOAD failure handled downstream)
    * multi-version with binary only in a non-first dir -> true
    * multi-version with no binary anywhere (Nix/Bazel/MDM tree) -> false
    * fs.readdir throws (EACCES) -> false

  The Windows conditional in doInitLbug / initLbugWithDb is intentionally
  not unit-isolated: it reduces to `probe ? load : true` over a fully
  constructed lbug.Database + Connection pool, which the
  test/integration/lbug-pool*.test.ts suites already exercise on the
  windows-latest CI matrix.

- Apply prettier format to the fs.stat() call in pool-adapter.ts,
  resolving the quality/format CI failure surfaced by gitnexus/autofix.

Addresses DoD §2.7 test-coverage blocker raised in the production-
readiness review on #1692, and the dir-exists-no-file regression case
raised on #1690.

Refs #1690.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-19 12:09:21 +01:00
55f8d442f6 fix(mcp): setup fallback on Windows when global gitnexus resolves to a non-spawnable shim (#1694)
* Initial plan

* fix: avoid invalid Windows MCP shim paths

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a052306e-483a-42d0-b65a-2646906457c7

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: cover .ps1 windows mcp fallback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/5aaed570-2a0b-4ed9-a0ac-ca099ce5675e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: assert windows fallback for cursor and codex

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/5aaed570-2a0b-4ed9-a0ac-ca099ce5675e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-19 08:17:26 +01:00
18167400c4 chore(deps)(deps): bump express and @types/express in /gitnexus (#872)
* chore(deps)(deps): bump express and @types/express in /gitnexus

Bumps [express](https://github.com/expressjs/express) and [@types/express](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/express). These dependencies needed to be updated together.

Updates `express` from 4.22.1 to 5.2.1
- [Release notes](https://github.com/expressjs/express/releases)
- [Changelog](https://github.com/expressjs/express/blob/master/History.md)
- [Commits](https://github.com/expressjs/express/compare/v4.22.1...v5.2.1)

Updates `@types/express` from 4.17.25 to 5.0.6
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/express)

---
updated-dependencies:
- dependency-name: "@types/express"
  dependency-version: 5.0.6
  dependency-type: direct:development
  update-type: version-update:semver-major
- dependency-name: express
  dependency-version: 5.2.1
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

* fix(server): normalize jobId param for Express 5 SSE routes

Express 5 types req.params values as string | string[]. mountSSEProgress uses a dynamic route path so TypeScript cannot narrow jobId; assert it once with assertString and reuse in the SSE progress callback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(ci): retrigger CI

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-19 07:43:07 +01:00
dad1ca7ab5 chore(deps)(deps): bump zod from 3.25.76 to 4.3.6 in /gitnexus-web (#1464)
Bumps [zod](https://github.com/colinhacks/zod) from 3.25.76 to 4.3.6.
- [Release notes](https://github.com/colinhacks/zod/releases)
- [Commits](https://github.com/colinhacks/zod/compare/v3.25.76...v4.3.6)

---
updated-dependencies:
- dependency-name: zod
  dependency-version: 4.3.6
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:56:04 +01:00
c746f30c90 chore(deps)(deps): bump langsmith (#1552)
Bumps the npm_and_yarn group with 1 update in the /gitnexus-web directory: [langsmith](https://github.com/langchain-ai/langsmith-sdk).


Updates `langsmith` from 0.5.23 to 0.6.3
- [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases)
- [Commits](https://github.com/langchain-ai/langsmith-sdk/commits)

---
updated-dependencies:
- dependency-name: langsmith
  dependency-version: 0.6.3
  dependency-type: indirect
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:55:44 +01:00
15a667ae5e chore(deps)(deps-dev): bump @vitest/coverage-v8 in /gitnexus (#1604)
Bumps [@vitest/coverage-v8](https://github.com/vitest-dev/vitest/tree/HEAD/packages/coverage-v8) from 4.1.5 to 4.1.6.
- [Release notes](https://github.com/vitest-dev/vitest/releases)
- [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.6/packages/coverage-v8)

---
updated-dependencies:
- dependency-name: "@vitest/coverage-v8"
  dependency-version: 4.1.6
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:55:24 +01:00
6210d80f1e chore(deps)(deps-dev): bump tsx from 4.21.0 to 4.21.1 in /gitnexus (#1698)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.21.0 to 4.21.1.
- [Release notes](https://github.com/privatenumber/tsx/releases)
- [Changelog](https://github.com/privatenumber/tsx/blob/master/release.config.cjs)
- [Commits](https://github.com/privatenumber/tsx/compare/v4.21.0...v4.21.1)

---
updated-dependencies:
- dependency-name: tsx
  dependency-version: 4.21.1
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:55:04 +01:00
637cfca39c chore(deps)(deps): bump express-rate-limit in /gitnexus (#1697)
Bumps [express-rate-limit](https://github.com/express-rate-limit/express-rate-limit) from 8.5.1 to 8.5.2.
- [Release notes](https://github.com/express-rate-limit/express-rate-limit/releases)
- [Commits](https://github.com/express-rate-limit/express-rate-limit/compare/v8.5.1...v8.5.2)

---
updated-dependencies:
- dependency-name: express-rate-limit
  dependency-version: 8.5.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:54:33 +01:00
dependabot[bot]andGergő Magyar 73543a4714 chore(deps)(deps-dev): bump @types/node in /gitnexus (#1696)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 25.6.2 to 25.7.0.
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node)

---
updated-dependencies:
- dependency-name: "@types/node"
  dependency-version: 25.7.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-19 06:54:14 +01:00
DuduPhudu b37974fdac feat(javascript): migrate JavaScript to scope-based resolution (RFC #909 Ring 3, issue #928) (#1640) 2026-05-19 06:23:13 +01:00
dependabot[bot] ade2069633 chore(deps)(deps): bump brace-expansion from 5.0.5 to 5.0.6 in /gitnexus (#1689) 2026-05-19 05:35:27 +01:00
azizur100389 5f0c0eba0e feat(cpp): Expand type_traits constraint registry (#1648) 2026-05-18 21:10:18 +01:00
Gergő Magyar 2632bcccc0 fix(api): open lbug read-only for /api/graph, /api/search, /api/grep (#1686) 2026-05-18 19:57:35 +01:00
Gergő Magyar c9199b654f fix(test): retry Windows temp cleanup in cli-e2e teardown (#1688) 2026-05-18 18:17:54 +01:00
Shane Thurston Wijaya 33f18ceaa2 feat(eval-server): added --host for user configured host IP instead of system hardcoded IP (127.0.0.1) (#1667)
* feat(eval-server): added --host for user configured host IP instead of system hardcoded IP (127.0.0.1)

* fix(eval-server): localhost value in --host now returns 127.0.0.1 instead of the raw input to fix wrong address, handled error for ipv6 disabled containers

* feat(eval-server): add --host flag with validation and error handling

  Co-Authored-By: Val Vladescu <val.vladescu@thirdbridge.com>

* fix(eval-server): bracketed IPv6 addresses to remove ambiguity

* docs(eval-server): document --host flag, READY signal format, and parser migration note

* fix(eval-server): use actual bound port in READY signal; strengthen --host e2e tests

  Co-Authored-By: Val Vladescu <val.vladescu@thirdbridge.com>

* feat(eval): wire eval-server --host through gitnexus_docker.py

* docs(eval): added guidance for docker user

* docs(eval): revise the imprecise documentation

* fix(e2e): updated original stdout for new format
2026-05-18 16:00:42 +01:00
c30833fad3 perf(scope-resolution): use owner-keyed lookup for Step 2 member resolution (#1657)
* perf(scope-resolution): use owner-keyed lookup for Step 2 member resolution (#1656)

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(scope-resolution): index Const/Static in FieldRegistry for Step 2 lookup

Extend FieldRegistry to hold multiple defs per (owner, name), reconcile Const and Static into the owner-keyed index, and wire lookupAllByOwner through the production hook so Step 2 does not drop field kinds the registry never indexed. Pass explicitReceiver on read/write reference sites and document undefined-vs-empty hook semantics for defs fallback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* perf(scope-resolution): centralize O(1) owned-member hook and guard hot path

Extract lookupOwnedMembersByOwner for the production Step 2 hook so merges stay O(1) per registry with no defs.byId scan. Add a perf-contract unit test that throws if byId.values runs when the hook is wired. Reuse a frozen empty sentinel on double miss to avoid per-probe allocations.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: drop unused buildFieldRegistry import

* chore(scope-resolution): apply ce-code-review safe_auto fixes

- Drop unreachable return + unused values() capture in perf-contract trap (Finding #7)
- Type lookupOwnedMembersByOwner ownerDefId as DefId (Finding #9)
- Add Static-kind Step 2 lookup test mirroring the Const case (Finding #11)

* docs(field-registry): document lookupFieldByOwner first-wins semantics

Audit of all 6 production callers (call-processor.ts:2279, walkers.ts:535,
receiver-bound-calls.ts:380+730, type-env.ts:627+631) confirms none depends
on last-wins precedence — all treat the return as a generic 'field with
this name owned by this class'. Clarify the JSDoc to surface the semantic
change introduced when FieldRegistry moved from last-wins to append-order
storage (ce-code-review finding #2).

* test(scope-resolution): extend Step 2 perf contract to implicit-self, MRO, field paths

Adds three sibling tests under the Step 2 perf contract describe block, each
asserting defs.byId.values() does NOT execute when ownedMembersByOwner is wired:

- implicit-self receiver via typeBindings.self (no explicitReceiver branch)
- 2-level MRO chain (Child extends Parent, save resolves on Parent at depth 1)
- FieldRegistry read via Step 2 (property lookup, separate registry path)

Pins the perf invariant on every distinct entry into walkReceiverTypeBinding
so a regression bypassing the hook on any sub-path now fails CI immediately
(ce-code-review finding #8).

* test(resolve-references): cover arity-overload filtering via resolveReferenceSites

Pins the orchestration-layer wiring of providers.arityCompatibility:
hook returns [save(arity 1), save(arity 2)], referenceSite.arity = 1,
arityCompatibility verdicts 'compatible'/'incompatible' by parameterCount,
exactly one reference emitted with toDef = the arity-1 overload.

registries.test.ts already covered arity at the buildMethodRegistry level;
this adds the missing entry-point check that resolveReferenceSites threads
providers correctly through to lookupCore.Step5 (ce-code-review finding #10).

* test(resolve-references): add hook-on vs hook-off parity test

Runs resolveReferenceSites twice on the same fixture (Parent.save method
hit + Child.name field hit, Child extends Parent MRO chain) — once with
ownedMembersByOwner wired to a synthetic registry, once with the hook
absent so collectOwnedMembers takes the defs.byId fallback. Asserts:

- stats are identical (sitesProcessed / referencesEmitted / unresolved)
- referenceIndex.bySourceScope entries have equal length
- toDef sets are equal
- each per-site reference (including evidence and depth) is .toEqual

Locks the semantic-parity claim in code while both paths still exist.
Will be removed alongside the fallback in finding #1 (ce-code-review #3).

* test(typescript): probe Step 2 MRO walk against ambient (declare class) base

Adds typescript-ambient-base-class fixture with an export declare class
AmbientBase + Derived extends AmbientBase and a call site d.ambientMethod().
Integration assertions:

- Both classes are detected
- EXTENDS edge Derived → AmbientBase emitted
- CALLS edge to ambient.ts:ambientMethod resolved via MRO walk

Probes the ce-code-review #6 concern that ambient-only owners (whose
bodies are never parsed) might be silently skipped by Step 2 after the
owner-keyed lookup change. Result: the call resolves correctly — the
method signature inside the declare class body still flows through
reconcileOwnership into model.methods, so the hook returns the right
ancestor hits. Residual risk is empirically closed.

* feat(scope-resolution): route nested types via owner-keyed TypeRegistry

Closes the Step 2 contract footgun where 'hook returns [] = authoritative
miss' silently dropped any owned def whose NodeLabel was outside the
method/field if-chain in reconcileOwnership.

- TypeRegistry: add nestedByOwner Map + lookupAllByOwner(owner, simple)
  + registerByOwner(owner, simple, def). Mirrors MethodRegistry/
  FieldRegistry shape; cleared with the rest on cascade clear.
- reconcileOwnership: route class-like NodeLabels (Class/Interface/Enum/
  Struct/Union/Trait/TypeAlias/Typedef/Record/Delegate/Annotation/
  Template/Namespace) via types.registerByOwner. New nestedTypesRegistered
  stat. Idempotent skip via nodeId match.
- validateOwnershipParity: extend the I9 invariant check to nested types.
- lookupOwnedMembersByOwner: merge methods + fields + nested-type hits;
  short-circuit when any one source contributes the full result.

Unblocks future receiver-MRO registries that need to resolve 'Outer.Inner'
through the receiver's type-binding chain (ce-code-review finding #5a).

* refactor(scope-resolution): make ownedMembersByOwner required; delete byId fallback

Per ce-code-review finding #1, the optional-hook design encoded a silent
O(|defs|) perf cliff into the type system: any RegistryContext built
without the hook regressed Step 2 to scanning every def per probe with
no warning. Production wires the hook unconditionally; the fallback was
exercised only by tests.

- RegistryContext.ownedMembersByOwner: required, returns readonly
  SymbolDefinition[] (no | undefined). Implementations MUST return [] on
  authoritative miss.
- collectOwnedMembers in lookup-core.ts collapses to a one-line forward
  to the hook; the defs.byId.values() scan and simpleNameOf helper are
  deleted (simpleNameOf had no other consumers).
- ResolveReferencesInput.ownedMembersByOwner: required to match.
- Tests: drop three fallback-path tests (registries Const fallback,
  resolveReferenceSites no-hook fallback, resolveReferenceSites Const-
  undefined fallback) and the hook-vs-fallback parity test added by
  finding #3. makeCtx in registries.test.ts now defaults to a real
  owner-keyed scan over the test fixture defs so tests that don't care
  about the hook keep working.

* perf(free-call-fallback): cache global callables by simple name once per pass

pickUniqueGlobalCallable scanned scopes.defs.byId.values() on every
free-call fallback site. After PR #1656 fixed Step 2, this scan became
the dominant remaining O(|defs|) hot path on large repos (ce-code-review
finding #4).

- buildGlobalCallableIndex builds a Map<simpleName, SymbolDefinition[]>
  over scopes.defs once at the top of emitFreeCallFallback. Same filter
  the per-site scan applied: Function / Method / Constructor, keyed by
  the last .-segment of qualifiedName.
- pickUniqueGlobalCallable consumes the prebuilt index via O(1) Map.get
  instead of iterating every def. Per-site complexity drops from
  O(|defs|) to O(|defs with this simple name|).
- Cost: O(|defs|) once per pass instead of O(|defs| * |free-call sites|).

Subsequent narrowing (arity, conversion-rank) and the model-side fallback
(model.symbols.lookupCallableByName + model.methods.lookupMethodByName)
are unchanged.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* ci: trigger build

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Gergo Magyar <abhigyan1.patwari@gmail.com>
2026-05-18 13:14:27 +01:00
Copilot 7d500390b9 fix: Use Ladybug native read-only enforcement and prepared statement execution for Cypher query paths (#1655) 2026-05-18 06:54:24 +01:00
Nilotpal Kashyap bdc0439a10 feat(detect-changes): support git worktrees (#1654) 2026-05-17 20:54:41 +01:00
Shane Thurston Wijaya 105efd0f7c feat(wiki): added --lang <lang> flags to gitnexus wiki for multilanguage wiki generation support (#1613) 2026-05-17 19:54:02 +01:00
Copilot 493827222d fix(ingestion): Raise analyze auto-heap to 16GB and tighten cross-platform OOM guidance for UE5-scale repositories (#1652) 2026-05-17 16:28:07 +01:00
Copilot ed50a6729f fix(wiki): Remove the hidden 60s default timeout, validate gitnexus wiki timeout/retry flags, and surface timeout errors (#1651) 2026-05-17 12:03:54 +01:00
Nilotpal Kashyap dfbe68ad24 fix(lbug): issue #1647, detect WAL corruption in schema init and surface recovery (#1650) 2026-05-17 10:46:45 +01:00
azizur100389andGergő Magyar 2376912ca7 feat(ingestion): Add C++ parameter type class sidecar (#1642)
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-16 21:44:26 +01:00
Zander Raycraft a4dfebd073 feat(cpp): sfinae filter (#1623)
* feat(cpp): SFINAE-aware overload filter — drops candidates whose enable_if_t / requires constraints fail (#1579)

* fix(cpp):  SFINAE follow-ups for is_integral_v/is_arithmetic_v bool and char support, an unqualified F1 test fixture, and parameter-lookup gap documentation (#1579) -> claude feedback

* revert: reverting all changes to .md files
2026-05-16 20:23:13 +01:00
237 changed files with 23333 additions and 2764 deletions
@@ -17,11 +17,11 @@ npx gitnexus analyze
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates CLAUDE.md / AGENTS.md context files.
| Flag | Effect |
| ------------------- | ------------------------------------------------------------------------------------------------------- |
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
| Flag | Effect |
| -------------- | ---------------------------------------------------------------- |
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
+102
View File
@@ -75,3 +75,105 @@ jobs:
build: 'true'
- run: npx vitest run
working-directory: gitnexus
# End-to-end smoke test for the #1728 packaging fix: pack the published
# tarball, install it globally into a temp prefix, and assert no junction
# creation (the EPERM root cause) plus working CLI plus vendor cleanliness
# (#836). Runs on windows-latest because that is the platform the fix
# targets; the in-repo `npm ci` job above only exercises the dev-tree path
# and skips the tarball reify step where the historical EPERM occurred.
packaged-install-smoke:
name: packaged install smoke (${{ matrix.os }})
strategy:
fail-fast: false
matrix:
os: [windows-latest, ubuntu-latest]
runs-on: ${{ matrix.os }}
timeout-minutes: 15
steps:
# persist-credentials: false — this job runs npm pack + npm install -g
# from a tarball and never pushes back; the token in .git/config would
# be at risk of leaking through any future artifact-upload step
# (zizmor artipacked audit). Disable upfront.
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- uses: ./.github/actions/setup-gitnexus
with:
build: 'true'
- name: Pack gitnexus tarball
shell: bash
run: npm pack
working-directory: gitnexus
- name: Install gitnexus tarball into isolated prefix
shell: bash
run: |
set -euo pipefail
PREFIX="$RUNNER_TEMP/gitnexus-smoke"
mkdir -p "$PREFIX"
TARBALL=$(find . -maxdepth 1 -name 'gitnexus-*.tgz' -print -quit)
if [ -z "$TARBALL" ]; then
echo "ERROR: no gitnexus-*.tgz tarball found in $(pwd)" >&2
exit 1
fi
echo "Installing $TARBALL into $PREFIX"
npm install -g --prefix "$PREFIX" "./$TARBALL" --no-audit --no-fund
echo "PREFIX=$PREFIX" >> "$GITHUB_ENV"
working-directory: gitnexus
- name: Assert no junctions or vendor build artifacts
shell: bash
run: |
set -euo pipefail
# Locate the installed gitnexus package across npm prefix layouts
# (lib/node_modules on POSIX, node_modules on Windows).
for candidate in "$PREFIX/lib/node_modules/gitnexus" "$PREFIX/node_modules/gitnexus"; do
if [ -d "$candidate" ]; then
INSTALLED="$candidate"
break
fi
done
if [ -z "${INSTALLED:-}" ]; then
echo "ERROR: installed gitnexus package not found under $PREFIX" >&2
ls -la "$PREFIX" || true
exit 1
fi
echo "Installed package at: $INSTALLED"
# #836 invariant: no node_modules/ or build/ under any vendor/*.
BAD=$(find "$INSTALLED/vendor" \( -name node_modules -o -name build \) -print 2>/dev/null || true)
if [ -n "$BAD" ]; then
echo "ERROR: vendor tree contains forbidden build artifacts (#836):" >&2
echo "$BAD" >&2
exit 1
fi
# #1728 invariant: materialized grammar dirs are real directories,
# not junctions/symlinks (which is what the EPERM regression created).
for name in tree-sitter-dart tree-sitter-proto tree-sitter-swift; do
entry="$INSTALLED/node_modules/$name"
if [ ! -e "$entry" ]; then
echo "WARN: $name not materialized (toolchain/prebuild may be unavailable on $RUNNER_OS)"
continue
fi
if [ -L "$entry" ]; then
echo "ERROR: $entry is a symlink/junction — #1728 regression" >&2
exit 1
fi
if [ ! -d "$entry" ]; then
echo "ERROR: $entry is not a directory" >&2
exit 1
fi
done
- name: Assert gitnexus --version works
shell: bash
run: |
set -euo pipefail
if [ "$RUNNER_OS" = "Windows" ]; then
"$PREFIX/gitnexus.cmd" --version
else
"$PREFIX/bin/gitnexus" --version
fi
+2 -2
View File
@@ -48,7 +48,7 @@ jobs:
persist-credentials: false
- name: Initialize CodeQL
uses: github/codeql-action/init@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/init@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
languages: ${{ matrix.language }}
queries: security-and-quality
@@ -69,6 +69,6 @@ jobs:
- '**/test/fixtures/**'
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/analyze@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
category: '/language:${{ matrix.language }}'
+1 -1
View File
@@ -33,7 +33,7 @@ jobs:
persist-credentials: false
- name: Dependency Review
uses: actions/dependency-review-action@2031cfc080254a8a887f58cffee85186f0e49e48 # v4.9.0
uses: actions/dependency-review-action@a1d282b36b6f3519aa1f3fc636f609c47dddb294 # v5.0.0
with:
fail-on-severity: high
comment-summary-in-pr: on-failure
+1 -1
View File
@@ -108,7 +108,7 @@ jobs:
# Pinned to v7.2.0. Verify SHA via:
# gh api repos/release-drafter/release-drafter/git/refs/tags/v7.2.0
# v7 removed `disable-releaser`; use `dry-run: true` to only autolabel.
- uses: release-drafter/release-drafter@563bf132657a13ded0b01fcb723c5a58cdd824e2 # v7.2.1
- uses: release-drafter/release-drafter@c2e2804cc59f45f57076a99af580d0fedb697927 # v7.3.0
with:
config-name: release-drafter.yml
dry-run: true
+1 -1
View File
@@ -53,6 +53,6 @@ jobs:
retention-days: 5
- name: Upload to Security tab
uses: github/codeql-action/upload-sarif@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/upload-sarif@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
sarif_file: results.sarif
+1 -1
View File
@@ -76,7 +76,7 @@ jobs:
exit-code: '0'
- name: Upload to Security tab
uses: github/codeql-action/upload-sarif@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/upload-sarif@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
sarif_file: trivy-${{ matrix.image.name }}.sarif
category: trivy-${{ matrix.image.name }}
+1 -1
View File
@@ -76,7 +76,7 @@ jobs:
continue-on-error: true
- name: Upload SARIF
uses: github/codeql-action/upload-sarif@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/upload-sarif@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
sarif_file: zizmor.sarif
category: zizmor
+41 -108
View File
@@ -62,131 +62,64 @@ Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
Indexed as **GitNexus** (4325 symbols, 10556 relationships, 300 execution flows). Use MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any tool warns the index is stale, run `npx gitnexus analyze` first.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** `gitnexus_impact({target: "symbolName", direction: "upstream"})` — report blast radius to the user.
- **MUST run `gitnexus_detect_changes()` before committing** — verify only expected symbols and flows are affected.
- **MUST warn the user** if impact returns HIGH or CRITICAL risk.
- Explore unfamiliar code with `gitnexus_query({query: "concept"})` (process-grouped, ranked) instead of grepping.
- Full context on a symbol: `gitnexus_context({name: "symbolName"})`.
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find related execution flows
2. `gitnexus_context({name: "<suspect function>"})` — callers, callees, process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace flow step by step
4. Regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})`
## When Refactoring
- **Rename:** `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Graph edits are safe; text_search edits need manual review.
- **Extract/Split:** `gitnexus_context` (incoming/outgoing refs) then `gitnexus_impact` (upstream callers) before moving code.
- **After any refactor:** `gitnexus_detect_changes({scope: "all"})` to verify scope.
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Never Do
- Edit a symbol without running `gitnexus_impact` first.
- Ignore HIGH/CRITICAL risk warnings.
- Rename with find-and-replace — use `gitnexus_rename`.
- Commit without `gitnexus_detect_changes()`.
- Add language-specific behavior to shared ingestion code (`gitnexus/src/core/ingestion/`) — use a `LanguageProvider` hook. Seeing `provider.mroStrategy === 'xxx'` or an import from `languages/xxx.ts` in shared code means stop and add a hook.
## Tools Quick Reference
| Tool | When to use | Example |
|------|-------------|---------|
| `list_repos` | Discover indexed repos | `gitnexus_list_repos({})` |
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
| `api_impact` | Pre-change API route impact | `gitnexus_api_impact({route: "/api/users", method: "GET"})` |
| `route_map` | Route → handler → consumer map | `gitnexus_route_map({})` |
| `tool_map` | MCP/RPC tool definitions | `gitnexus_tool_map({})` |
| `shape_check` | Response shape vs consumer access | `gitnexus_shape_check({route: "/api/users"})` |
| `group_list` | List repo groups | `gitnexus_group_list({})` |
| `group_sync` | Rebuild group Contract Registry | `gitnexus_group_sync({name: "myGroup"})` |
| `query` (group mode) | Cross-repo search in a group (RRF-merged) | `gitnexus_query({repo: "@myGroup", query: "auth"})` |
| `context` (group mode) | 360° view across all member repos | `gitnexus_context({repo: "@myGroup", name: "validateUser"})` |
| `impact` (group mode) | Cross-repo blast radius via Contract Bridge | `gitnexus_impact({repo: "@myGroup", target: "X", direction: "upstream"})` |
> Group mode: pass `repo: "@<groupName>"` to fan out across all member repos, or `repo: "@<groupName>/<memberPath>"` to target a single member (path keys from `group.yaml`). Optional `service: "<monorepo/path>"` filters by service root. Group-level state (contracts, staleness) lives in the resources table below — there are **no** `group_query` / `group_context` / `group_impact` / `group_contracts` / `group_status` MCP tools.
>
> For a full walkthrough of setting up a group across multiple repos that communicate over gRPC, see [docs/guides/microservices-grpc.md](docs/guides/microservices-grpc.md).
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, index freshness |
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
| `gitnexus://group/{name}/contracts` | Group Contract Registry (provider/consumer rows + cross-links) |
| `gitnexus://group/{name}/status` | Per-member index + Contract Registry staleness report |
## Self-Check Before Finishing
## CLI
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL warnings were ignored
3. `gitnexus_detect_changes()` confirms expected scope
4. All d=1 dependents were updated
## Keeping the Index Fresh
```bash
npx gitnexus analyze # incremental by default; preserves embeddings
npx gitnexus analyze --force # full rebuild from scratch (opt out of incremental)
npx gitnexus analyze --embeddings # also generate embeddings for new/changed nodes
npx gitnexus analyze --drop-embeddings # explicit opt-in to wipe existing embeddings
```
`analyze` runs **incrementally by default**. The pipeline still parses every file every run (cross-file resolution requires it), but tree-sitter parsing is **served from a content-addressed cache** under `.gitnexus/parse-cache/` (per-chunk JSON shards plus `index.json`) for chunks whose file contents haven't changed since the last run. Older installs may still have a legacy single file `.gitnexus/parse-cache.json`, which is read for backward compatibility but no longer written. Only changed-file rows (and their importers) are rewritten in LadybugDB; unchanged-file rows are preserved. Output is byte-equivalent to a full rebuild. Pass `--force` to wipe and re-index from scratch (e.g., to recover from a corrupt index, or after upgrading GitNexus).
The parse cache key is **content-addressed and version-tagged**: it survives `--force` runs, and is automatically invalidated by a `gitnexus` package upgrade (so a new tree-sitter grammar doesn't silently replay stale parse output). Safe to delete the whole `.gitnexus/parse-cache/` directory (and remove any legacy `.gitnexus/parse-cache.json` if present) at any time — it'll be rebuilt on the next analyze.
Check `.gitnexus/meta.json` `stats.embeddings` (0 = none). A plain `analyze` no longer drops existing vectors — pass `--drop-embeddings` to wipe.
> Claude Code: PostToolUse hook detects a stale index after `git commit` and `git merge` and prompts the agent to run `analyze`. The hook does not invoke `analyze` itself.
## CLI Skills
| Task | Skill file |
|------|-----------|
| Architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Debugging / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Refactoring | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools/resources/schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| CLI commands (index, status, clean, wiki) | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
## Hook env knobs
The Claude Code hook (`gitnexus/hooks/claude/gitnexus-hook.cjs` and the mirrored plugin copy under `gitnexus-claude-plugin/hooks/`) honours these env vars. Defaults work for normal installations; set them only to override resolution. All path overrides ignore values that do not exist on disk and fall through to the standard resolution chain.
| Env var | Type | Default | Purpose |
|---------|------|---------|---------|
| `GITNEXUS_HOOK_CLI_PATH` | path | resolved via package layout / `require.resolve` | Override path to the `gitnexus` CLI entry the hook spawns for `augment`. |
| `GITNEXUS_HOOK_LSOF_PATH` | path | `lsof` on `PATH` (with `/usr/bin/lsof`, `/usr/sbin/lsof`, `/sbin/lsof` fallbacks) | Override POSIX `lsof` location for the DB-lock probe. |
| `GITNEXUS_HOOK_PS_PATH` | path | `ps` on `PATH` (with `/bin/ps`, `/usr/bin/ps` fallbacks) | Override POSIX `ps` location. |
| `GITNEXUS_HOOK_POWERSHELL_PATH` | path | `%SystemRoot%\System32\WindowsPowerShell\v1.0\powershell.exe` (then `SysWOW64`, then `powershell.exe` on `PATH`) | Override Windows PowerShell location used by the Restart-Manager probe. |
| `GITNEXUS_HOOK_LINUX_PROC_BUDGET_MS` | integer ms | `1200` | Max wall-clock for the Linux `/proc` fd scan before bailing out to the `lsof` fallback. |
| `GITNEXUS_HOOK_RM_TARGET` | path | derived | Restart-Manager target file (the LadybugDB path under `.gitnexus/`). Set internally by the hook; rarely overridden manually. |
| `GITNEXUS_DEBUG` | boolean (`1`/`true`) | unset | Verbose stderr from the hook: prints discarded augment-stderr prefixes and one-shot `.ps1` load-failure warnings. |
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (239 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Extractors area (135 symbols) | `.claude/skills/generated/extractors/SKILL.md` |
| Work in the Components area (112 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Lbug area (96 symbols) | `.claude/skills/generated/lbug/SKILL.md` |
| Work in the Group area (94 symbols) | `.claude/skills/generated/group/SKILL.md` |
| Work in the Cli area (92 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Configs area (92 symbols) | `.claude/skills/generated/configs/SKILL.md` |
| Work in the Type-extractors area (90 symbols) | `.claude/skills/generated/type-extractors/SKILL.md` |
| Work in the Hooks area (88 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Unit area (80 symbols) | `.claude/skills/generated/unit/SKILL.md` |
| Work in the Cpp area (73 symbols) | `.claude/skills/generated/cpp/SKILL.md` |
| Work in the Scope-resolution area (72 symbols) | `.claude/skills/generated/scope-resolution/SKILL.md` |
| Work in the Server area (66 symbols) | `.claude/skills/generated/server/SKILL.md` |
| Work in the Local area (61 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Wiki area (60 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Workers area (57 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Embeddings area (56 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Typescript area (53 symbols) | `.claude/skills/generated/typescript/SKILL.md` |
| Work in the Storage area (51 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Php area (48 symbols) | `.claude/skills/generated/php/SKILL.md` |
<!-- gitnexus:end -->
+64
View File
@@ -52,3 +52,67 @@ If always-on instructions grow, load deep conventions via conditional reads (e.g
## GitNexus rules
See the `<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** for the canonical MCP tools, impact analysis rules, and index instructions.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (239 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Extractors area (135 symbols) | `.claude/skills/generated/extractors/SKILL.md` |
| Work in the Components area (112 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Lbug area (96 symbols) | `.claude/skills/generated/lbug/SKILL.md` |
| Work in the Group area (94 symbols) | `.claude/skills/generated/group/SKILL.md` |
| Work in the Cli area (92 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Configs area (92 symbols) | `.claude/skills/generated/configs/SKILL.md` |
| Work in the Type-extractors area (90 symbols) | `.claude/skills/generated/type-extractors/SKILL.md` |
| Work in the Hooks area (88 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Unit area (80 symbols) | `.claude/skills/generated/unit/SKILL.md` |
| Work in the Cpp area (73 symbols) | `.claude/skills/generated/cpp/SKILL.md` |
| Work in the Scope-resolution area (72 symbols) | `.claude/skills/generated/scope-resolution/SKILL.md` |
| Work in the Server area (66 symbols) | `.claude/skills/generated/server/SKILL.md` |
| Work in the Local area (61 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Wiki area (60 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Workers area (57 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Embeddings area (56 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Typescript area (53 symbols) | `.claude/skills/generated/typescript/SKILL.md` |
| Work in the Storage area (51 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Php area (48 symbols) | `.claude/skills/generated/php/SKILL.md` |
<!-- gitnexus:end -->
+100 -80
View File
@@ -1,4 +1,5 @@
# GitNexus
**⚠️ Important Notice:** GitNexus has NO official cryptocurrency, token, or coin. Any token/coin using the GitNexus name on Pump.fun or any other platform is **not affiliated with, endorsed by, or created by** this project or its maintainers. Do not purchase any cryptocurrency claiming association with GitNexus.
<div align="center">
@@ -30,14 +31,9 @@
Indexes any codebase into a knowledge graph — every dependency, call chain, cluster, and execution flow — then exposes it through smart tools so AI agents never miss code.
https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
> *Like DeepWiki, but deeper.* DeepWiki helps you *understand* code. GitNexus lets you *analyze* it — because a knowledge graph tracks every relationship, not just descriptions.
> _Like DeepWiki, but deeper._ DeepWiki helps you _understand_ code. GitNexus lets you _analyze_ it — because a knowledge graph tracks every relationship, not just descriptions.
**TL;DR:** The **Web UI** is a quick way to chat with any repo. The **CLI + MCP** is how you make your AI agent actually reliable — it gives Cursor, Claude Code, Codex, and friends a deep architectural view of your codebase so they stop missing dependencies, breaking call chains, and shipping blind edits. Even smaller models get full architectural clarity, making it compete with Goliath models.
@@ -47,18 +43,17 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
[![Star History Chart](https://api.star-history.com/svg?repos=abhigyanpatwari/GitNexus&type=date&legend=top-left)](https://www.star-history.com/#abhigyanpatwari/GitNexus&type=date&legend=top-left)
## Two Ways to Use GitNexus
| | **CLI + MCP** | **Web UI** |
| ----------------- | -------------------------------------------------------------- | ------------------------------------------------------------ |
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
| **For** | Daily development with Cursor, Claude Code, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Privacy** | Everything local, no network | Everything in-browser, no server |
| | **CLI + MCP** | **Web UI** |
| ----------- | --------------------------------------------------------------------- | -------------------------------------------------------------------- |
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
| **For** | Daily development with Cursor, Claude Code, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Privacy** | Everything local, no network | Everything in-browser, no server |
> **Bridge mode:** `gitnexus serve` connects the two — the web UI auto-detects the local server and can browse all your CLI-indexed repos without re-uploading or re-indexing.
@@ -69,6 +64,7 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
GitNexus is available as an **enterprise offering** - either as a fully managed **SaaS** or a **self-hosted** deployment. Also available for **commercial use** of the OSS version with proper licensing.
Enterprise includes:
- **PR Review** - automated blast radius analysis on pull requests
- **Auto-updating Code Wiki** - always up-to-date documentation (Code Wiki is also available in OSS)
- **Auto-reindexing** - knowledge graph stays fresh automatically
@@ -77,6 +73,7 @@ Enterprise includes:
- **Priority feature/language support** - request new languages or features
**Upcoming:**
- Auto regression forensics
- End-to-end test generation
@@ -109,7 +106,7 @@ That's it. This indexes the codebase, installs agent skills, registers Claude Co
To configure MCP for your editor, run `npx gitnexus setup` once — or set it up manually below.
> **Faster install (no C++ toolchain needed):** set `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` before `npm install -g gitnexus` to skip the native `tree-sitter-dart` and `tree-sitter-proto` builds. Dart/Proto files won't be parsed, but install completes in seconds without `python3`/`make`/`g++`. Strict `=1` only — any other value falls through to the rebuild.
> **Faster install (no C++ toolchain needed):** set `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` before `npm install -g gitnexus` to skip vendored grammar materialize/build (`tree-sitter-dart`, `tree-sitter-proto`, `tree-sitter-swift`). Dart/Proto/Swift files won't be parsed, but install completes in seconds without `python3`/`make`/`g++`. Strict `=1` only — any other value falls through to the rebuild.
### MCP Setup
@@ -117,13 +114,13 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
### Editor Support
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
| --------------------- | --- | ------ | -------------------- | -------------- |
| **Claude Code** | Yes | Yes | Yes (PreToolUse + PostToolUse) | **Full** |
| **Cursor** | Yes | Yes | Yes (postToolUse, [manual install](gitnexus-cursor-integration/README.md#hook-install)) | **Full** |
| **Codex** | Yes | Yes | — | MCP + Skills |
| **Windsurf** | Yes | — | — | MCP |
| **OpenCode** | Yes | Yes | — | MCP + Skills |
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
| --------------- | --- | ------ | --------------------------------------------------------------------------------------- | ------------ |
| **Claude Code** | Yes | Yes | Yes (PreToolUse + PostToolUse) | **Full** |
| **Cursor** | Yes | Yes | Yes (postToolUse, [manual install](gitnexus-cursor-integration/README.md#hook-install)) | **Full** |
| **Codex** | Yes | Yes | — | MCP + Skills |
| **Windsurf** | Yes | — | — | MCP |
| **OpenCode** | Yes | Yes | — | MCP + Skills |
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that enrich searches with graph context + PostToolUse hooks that detect a stale index after commits and prompt the agent to reindex.
@@ -131,10 +128,10 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
Built by the community — not officially maintained, but worth checking out.
| Project | Author | Description |
|---------|--------|-------------|
| [pi-gitnexus](https://github.com/tintinweb/pi-gitnexus) | [@tintinweb](https://github.com/tintinweb) | GitNexus plugin for [pi](https://pi.dev) — `pi install npm:pi-gitnexus` |
| [gitnexus-stable-ops](https://github.com/ShunsukeHayashi/gitnexus-stable-ops) | [@ShunsukeHayashi](https://github.com/ShunsukeHayashi) | Stable ops & deployment workflows (Miyabi ecosystem) |
| Project | Author | Description |
| ----------------------------------------------------------------------------- | ------------------------------------------------------ | ----------------------------------------------------------------------- |
| [pi-gitnexus](https://github.com/tintinweb/pi-gitnexus) | [@tintinweb](https://github.com/tintinweb) | GitNexus plugin for [pi](https://pi.dev) — `pi install npm:pi-gitnexus` |
| [gitnexus-stable-ops](https://github.com/ShunsukeHayashi/gitnexus-stable-ops) | [@ShunsukeHayashi](https://github.com/ShunsukeHayashi) | Stable ops & deployment workflows (Miyabi ecosystem) |
> Have a project built on GitNexus? Open a PR to add it here!
@@ -197,7 +194,8 @@ args = ["-y", "gitnexus@latest", "mcp"]
```bash
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --repair-fts # Fast path: rebuild/verify only FTS indexes on existing index data
gitnexus analyze --force # Full rebuild: re-parse + graph rebuild + FTS rebuild
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
@@ -205,6 +203,7 @@ gitnexus analyze --skip-git # Index folders that are not Git repositories
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus analyze --worker-timeout 60 # Increase worker idle timeout for slow parses
gitnexus analyze --workers <n> # Parse worker pool size (default: cores-1, capped at 16; 0 = sequential)
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
gitnexus serve # Start local HTTP server (multi-repo) for web UI connection
gitnexus list # List all indexed repositories
@@ -229,6 +228,25 @@ gitnexus group status <name> # Check staleness of repos in a group
If `analyze` reports a worker parse timeout on a large or unusual repository, it keeps running and falls back safely. To give slow worker jobs more time, use `gitnexus analyze --worker-timeout 60` or set `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=60000`. For very large files, `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` controls the worker job byte budget.
#### Environment variables
Most `analyze` knobs are also CLI flags (`--workers`, `--worker-timeout`, `--max-file-size`, `--verbose`). Use the env-var form when you'd otherwise repeat the same flag every run, or when invoking GitNexus from a long-running host (MCP server, eval-server, CI shell) that already manages its own environment. CLI flags take precedence over env vars; env vars take precedence over built-in defaults.
| Variable | Default | Effect | Tune when… |
| -------------------------------------- | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
| `GITNEXUS_WORKER_POOL_SIZE` | `cores - 1`, capped at 16 | Parse worker pool size. `0` disables the pool (sequential fallback). Equivalent to `--workers <n>`. | Constrained containers (cgroup CPU limits), CI runners with explicit quotas, or debugging a worker-only crash via `0`. |
| `GITNEXUS_PARSE_CHUNK_CONCURRENCY` | `2` | Number of chunks whose file contents may be read into memory in parallel while the pool dispatches the current chunk. Worker dispatch itself stays serial. | Repos large enough to chunk (multi-MB total source) where disk I/O is a measurable fraction of analyze wall-clock. |
| `GITNEXUS_VERBOSE` | unset | When `1`, enables verbose ingestion logs (skipped-file warnings, per-chunk throughput, parse-cache stats). Equivalent to `--verbose`. | Debugging an analyze that "completed" but seems to have missed files; tuning `--workers` / chunk concurrency against observable throughput. |
| `GITNEXUS_MAX_FILE_SIZE` | `512` (KB) | Walker skip threshold in KB. Hard cap is `32768` (tree-sitter buffer ceiling). Equivalent to `--max-file-size <kb>`. | Indexing repos with intentionally-large source files (generated parsers, vendored bundles) that should still be parsed. |
| `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS` | `30000` | Worker idle timeout in milliseconds before retry/fallback. Equivalent to `--worker-timeout <seconds>` × 1000. | Slow-parsing files (large minified JS, deeply-nested TS types) that legitimately need more than 30s. |
| `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` | `8388608` (8 MB) | Per-job byte budget the pool will send to a worker in one `postMessage`. | Very large individual files; mostly diagnostic — bumping past 8 MB risks structured-clone memory pressure. |
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per worker slot before the slot is dropped from the active rotation. Bounds respawn loops on a chronically-crashing slot. | Hosts where a flaky worker should retry more (raise) or fail-fast (lower) before the slot is dropped. |
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Combined with `timeoutBackoffFactor`, prevents exponentially-growing retries from stalling for hours. | Slow files that legitimately need long total retry windows; lower to fail-fast on stalls. |
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD`| `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, every subsequent dispatch rejects until a fresh pool is created. | Hosts where a SIGSEGV-prone native grammar should trip the breaker sooner; CI runners that should fail loudly. |
| `GITNEXUS_CHUNK_BYTE_BUDGET` | `2097152` (2 MB) | Chunk boundary used for cache-key composition and dispatch. Smaller = finer-grained cache hits but more dispatch overhead. | Tuning incremental-analyze cache behavior on monorepos. |
| `GITNEXUS_NO_GITIGNORE` | unset | When set, skips `.gitignore` parsing. `.gitnexusignore` is still honored. | Indexing a repo whose `.gitignore` excludes files you actually want indexed (e.g., generated code committed for cross-repo lookup). |
| `GITNEXUS_SKIP_OPTIONAL_GRAMMARS` | unset | When `=1` strictly, skips vendored grammar materialize/build for `tree-sitter-dart`, `tree-sitter-proto`, and `tree-sitter-swift` at install time. | Installing on a host without a C++ toolchain or where Swift prebuilds don't match; you're willing to skip Dart/Proto/Swift parsing. |
#### Publishing to understand-quickly (opt-in)
[`looptech-ai/understand-quickly`](https://github.com/looptech-ai/understand-quickly) is a public registry of code-knowledge graphs that lists `gitnexus@1` as a first-class format. After registering your repo once (`npx @understand-quickly/cli add` or the [wizard](https://looptech-ai.github.io/understand-quickly/add.html)), `gitnexus publish` fires a single `repository_dispatch` event so the registry resyncs your entry on demand instead of waiting for the nightly job.
@@ -239,27 +257,27 @@ It is opt-in and a no-op without `UNDERSTAND_QUICKLY_TOKEN` — a fine-grained G
**16 tools** exposed via MCP (11 per-repo + 5 group):
| Tool | What It Does | `repo` Param |
| ------------------ | ----------------------------------------------------------------- | -------------- |
| `list_repos` | Discover all indexed repositories | — |
| `query` | Process-grouped hybrid search (BM25 + semantic + RRF) | Optional |
| `context` | 360-degree symbol view — categorized refs, process participation | Optional |
| `impact` | Blast radius analysis with depth grouping and confidence | Optional |
| `detect_changes` | Git-diff impact — maps changed lines to affected processes | Optional |
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
| `cypher` | Raw Cypher graph queries | Optional |
| `group_list` | List configured repository groups | — |
| `group_sync` | Extract contracts and match across repos/services | — |
| `group_contracts`| Inspect extracted contracts and cross-links | — |
| `group_query` | Search execution flows across all repos in a group | — |
| `group_status` | Check staleness of repos in a group | — |
| Tool | What It Does | `repo` Param |
| ----------------- | ---------------------------------------------------------------- | ------------ |
| `list_repos` | Discover all indexed repositories | — |
| `query` | Process-grouped hybrid search (BM25 + semantic + RRF) | Optional |
| `context` | 360-degree symbol view — categorized refs, process participation | Optional |
| `impact` | Blast radius analysis with depth grouping and confidence | Optional |
| `detect_changes` | Git-diff impact — maps changed lines to affected processes | Optional |
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
| `cypher` | Raw Cypher graph queries | Optional |
| `group_list` | List configured repository groups | — |
| `group_sync` | Extract contracts and match across repos/services | — |
| `group_contracts` | Inspect extracted contracts and cross-links | — |
| `group_query` | Search execution flows across all repos in a group | — |
| `group_status` | Check staleness of repos in a group | — |
> When only one repo is indexed, the `repo` parameter is optional. With multiple repos, specify which one: `query({query: "auth", repo: "my-app"})`.
**Resources** for instant context:
| Resource | Purpose |
| ----------------------------------------- | ---------------------------------------------------- |
| Resource | Purpose |
| --------------------------------------- | ---------------------------------------------------- |
| `gitnexus://repos` | List all indexed repositories (read this first) |
| `gitnexus://repo/{name}/context` | Codebase stats, staleness check, and available tools |
| `gitnexus://repo/{name}/clusters` | All functional clusters with cohesion scores |
@@ -270,9 +288,9 @@ It is opt-in and a no-op without `UNDERSTAND_QUICKLY_TOKEN` — a fine-grained G
**2 MCP prompts** for guided workflows:
| Prompt | What It Does |
| ----------------- | ------------------------------------------------------------------------- |
| `detect_impact` | Pre-commit change analysis — scope, affected processes, risk level |
| Prompt | What It Does |
| --------------- | ------------------------------------------------------------------------- |
| `detect_impact` | Pre-commit change analysis — scope, affected processes, risk level |
| `generate_map` | Architecture documentation from the knowledge graph with mermaid diagrams |
**4 agent skills** installed to `.claude/skills/` automatically:
@@ -359,10 +377,10 @@ npx gitnexus@latest serve
The official Docker setup ships **two signed images** orchestrated by `docker-compose.yaml`. Each image is published to both **GitHub Container Registry** (GHCR) and **Docker Hub** — same build, same digest, same Cosign signature — so pick whichever registry you prefer:
| Purpose | GHCR (default in `docker-compose.yaml`) | Docker Hub mirror |
| ---------------------------------------------------------------------- | --------------------------------------------- | ------------------------------------------- |
| CLI / `gitnexus serve` backend (HTTP API on port `4747`, MCP, indexer) | `ghcr.io/abhigyanpatwari/gitnexus:latest` | `akonlabs/gitnexus:latest` |
| Static web UI (port `4173`) | `ghcr.io/abhigyanpatwari/gitnexus-web:latest` | `akonlabs/gitnexus-web:latest` |
| Purpose | GHCR (default in `docker-compose.yaml`) | Docker Hub mirror |
| ---------------------------------------------------------------------- | --------------------------------------------- | ------------------------------ |
| CLI / `gitnexus serve` backend (HTTP API on port `4747`, MCP, indexer) | `ghcr.io/abhigyanpatwari/gitnexus:latest` | `akonlabs/gitnexus:latest` |
| Static web UI (port `4173`) | `ghcr.io/abhigyanpatwari/gitnexus-web:latest` | `akonlabs/gitnexus-web:latest` |
> **Heads-up — image rename.** Earlier releases published the web UI under
> `ghcr.io/abhigyanpatwari/gitnexus`. Starting with the introduction of the
@@ -578,22 +596,22 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
### Supported Languages
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
|----------|---------|----------------|---------|----------|-----------------|---------------------|--------|------------|-------------|
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
| ---------- | ------- | -------------- | ------- | -------- | ---------------- | --------------------- | ------ | ---------- | ------------ |
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
@@ -725,9 +743,11 @@ gitnexus wiki --force
# Increase the timeout or retries for large codebase or slow LLM providers
gitnexus wiki --timeout <seconds> # Per-attempt LLM request timeout in seconds (default: 60)
gitnexus wiki --timeout <seconds> # LLM request timeout in seconds (default: disabled)
gitnexus wiki --retries <n> # Max LLM retry attempts per request (default: 3)
# Change the language generation for wiki
gitnexus wiki --lang <lang> # Output language for generated documentation (e.g. english, chinese, spanish, japanese)
```
The wiki generator reads the indexed graph structure, groups files into modules via LLM, generates per-module documentation pages, and creates an overview page — all with cross-references to the knowledge graph.
@@ -736,16 +756,16 @@ The wiki generator reads the indexed graph structure, groups files into modules
## Tech Stack
| Layer | CLI | Web |
| ------------------------- | ------------------------------------- | --------------------------------------- |
| Layer | CLI | Web |
| ------------------- | ------------------------------------- | --------------------------------------- |
| **Runtime** | Node.js (native) | Browser (WASM) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Database** | LadybugDB native | LadybugDB WASM |
| **Database** | LadybugDB native | LadybugDB WASM |
| **Embeddings** | HuggingFace transformers.js (GPU/CPU) | transformers.js (WebGPU/WASM) |
| **Search** | BM25 + semantic + RRF | BM25 + semantic + RRF |
| **Agent Interface** | MCP (stdio) | LangChain ReAct agent |
| **Visualization** | — | Sigma.js + Graphology (WebGL) |
| **Frontend** | — | React 18, TypeScript, Vite, Tailwind v4 |
| **Visualization** | — | Sigma.js + Graphology (WebGL) |
| **Frontend** | — | React 18, TypeScript, Vite, Tailwind v4 |
| **Clustering** | Graphology | Graphology |
| **Concurrency** | Worker threads + async | Web Workers + Comlink |
@@ -761,12 +781,12 @@ The wiki generator reads the indexed graph structure, groups files into modules
### Recently Completed
- [X] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
- [X] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
- [X] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
- [X] Multi-Repo MCP, Zero-Config Setup, 14 Language Support
- [X] Community Detection, Process Detection, Confidence Scoring
- [X] Hybrid Search, Vector Index
- [x] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
- [x] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
- [x] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
- [x] Multi-Repo MCP, Zero-Config Setup, 14 Language Support
- [x] Community Detection, Process Detection, Confidence Scoring
- [x] Hybrid Search, Vector Index
---
+57 -2
View File
@@ -162,8 +162,8 @@ Each mode has a `system_{mode}.jinja` + `instance_{mode}.jinja` pair. The agent
```
Agent → bash command → /usr/local/bin/gitnexus-query
→ curl localhost:4848/tool/query (fast path: eval-server, ~100ms)
→ npx gitnexus query (fallback: cold CLI, ~5-10s)
→ curl http://127.0.0.1:4848/tool/query (fast path: eval-server, ~100ms)
→ npx gitnexus query (fallback: cold CLI, ~5-10s)
```
Each tool script in `/usr/local/bin/` is standalone — no sourcing, no env inheritance needed. This is critical because mini-swe-agent runs every command via `subprocess.run` in a fresh subshell.
@@ -176,6 +176,61 @@ The eval-server is a lightweight HTTP daemon that:
- Includes next-step hints to guide tool chaining (query → context → impact → fix)
- Auto-shuts down after idle timeout
**CLI flags:**
| Flag | Default | Purpose |
|------|---------|---------|
| `--port <port>` | `4848` | Port to listen on |
| `--host <host>` | `127.0.0.1` | Bind address — use `0.0.0.0` for cross-container access |
| `--idle-timeout <seconds>` | `0` (disabled) | Auto-shutdown after N seconds of inactivity |
**READY signal:**
When the server is ready, it writes to stdout:
```
# IPv4
GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848
# IPv6 (bracketed to avoid colon ambiguity)
GITNEXUS_EVAL_SERVER_READY:[::1]:4848
```
Parse the port as the last colon-segment (`split(':').pop()`) — not `split(':')[1]`, which breaks for IPv6 and for non-loopback IPv4 hosts added in this release.
### Custom port and host
`run_eval.py` does not expose `--port` or `--host` as CLI flags. Configure them in your mode YAML under the `environment:` key:
```yaml
# configs/modes/native_augment.yaml (or whichever mode you're running)
environment:
eval_server_port: 4849 # change if 4848 is already in use on the host
eval_server_host: "0.0.0.0" # bind all interfaces — needed for cross-container setups
```
Defaults are `port: 4848` and `host: 127.0.0.1` (loopback only). Use `0.0.0.0` only when the agent container needs to reach the eval-server from a separate network namespace. The health probe and tool scripts connect via the configured bind host (defaulting to `127.0.0.1`), which is reachable for both loopback and all-interface binds.
`"localhost"` is also a valid `eval_server_host` value. The OS resolves it at bind time — typically `127.0.0.1` on dual-stack or IPv4-only systems, and `::1` on IPv6-only systems. The exact result depends on your `/etc/hosts` and `gai.conf`. The READY signal will reflect the actual bound address (e.g. `GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848` or `GITNEXUS_EVAL_SERVER_READY:[::1]:4848`), not the literal string `localhost`. Use this when you want the server to bind to whichever loopback address the OS prefers rather than forcing IPv4.
**Running eval-server directly in Docker / Docker Compose:**
```bash
# Bind to all interfaces so sibling containers can reach it
gitnexus eval-server --host 0.0.0.0 --port 4848
# Then probe from a sibling container via its service hostname
curl http://eval-container:4848/health
```
If you need a non-default port (e.g. to avoid conflicts), pass `--port <port>` alongside `--host`. The READY signal will reflect both:
```
GITNEXUS_EVAL_SERVER_READY:0.0.0.0:5000
```
Parse the port as the last colon-segment (`split(':').pop()`) — safe for both IPv4 and bracketed IPv6 forms.
### Index caching
SWE-bench repos repeat (Django has 200+ instances at different commits). The harness caches GitNexus indexes per `(repo, commit)` hash in `~/.gitnexus-eval-cache/` to avoid redundant re-indexing.
+18 -5
View File
@@ -39,6 +39,7 @@ logger = logging.getLogger("gitnexus_docker")
DEFAULT_CACHE_DIR = Path.home() / ".gitnexus-eval-cache"
EVAL_SERVER_PORT = 4848
EVAL_SERVER_HOST = "127.0.0.1"
class GitNexusDockerEnvironment(DockerEnvironment):
@@ -62,6 +63,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
skip_embeddings: bool = True,
gitnexus_timeout: int = 120,
eval_server_port: int = EVAL_SERVER_PORT,
eval_server_host: str = EVAL_SERVER_HOST,
**kwargs,
):
super().__init__(**kwargs)
@@ -70,6 +72,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
self.skip_embeddings = skip_embeddings
self.gitnexus_timeout = gitnexus_timeout
self.eval_server_port = eval_server_port
self.eval_server_host = eval_server_host
self.index_time: float = 0.0
self._gitnexus_ready = False
@@ -165,22 +168,29 @@ class GitNexusDockerEnvironment(DockerEnvironment):
def _start_eval_server(self):
"""Start the GitNexus eval-server daemon in the background."""
logger.info(f"Starting eval-server on port {self.eval_server_port}...")
logger.info(
f"Starting eval-server on {self.eval_server_host}:{self.eval_server_port}..."
)
self.execute({
"command": (
f"nohup npx gitnexus eval-server --port {self.eval_server_port} "
f"--host {self.eval_server_host} "
f"--idle-timeout 600 "
f"> /tmp/gitnexus-eval-server.log 2>&1 &"
),
"timeout": 5,
})
# Use 127.0.0.1 for the health probe — reachable whether server binds
# loopback or all interfaces (0.0.0.0), avoiding DNS resolution issues.
health_host = "127.0.0.1"
# Wait for the server to be ready (up to ~15s for KuzuDB init)
for i in range(EVAL_SERVER_HEALTH_RETRIES):
time.sleep(EVAL_SERVER_HEALTH_INTERVAL_SECONDS)
health = self.execute({
"command": f"curl -sf http://127.0.0.1:{self.eval_server_port}/health 2>/dev/null || echo 'NOT_READY'",
"command": f"curl -sf http://{health_host}:{self.eval_server_port}/health 2>/dev/null || echo 'NOT_READY'",
"timeout": EVAL_SERVER_HEALTH_TIMEOUT_SECONDS,
})
output = health.get("output", "").strip()
@@ -201,7 +211,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
)
@staticmethod
def _render_tool_script(spec: ToolScriptSpec, port: str) -> str:
def _render_tool_script(spec: ToolScriptSpec, port: str, host: str = EVAL_SERVER_HOST) -> str:
"""
Render a standalone bash script for a GitNexus tool.
@@ -212,6 +222,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
if spec.endpoint:
lines.append(f'PORT="${{GITNEXUS_EVAL_PORT:-{port}}}"')
lines.append(f'HOST="${{GITNEXUS_EVAL_HOST:-{host}}}"')
if spec.header:
lines.append(spec.header.strip())
@@ -221,7 +232,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
if spec.endpoint:
lines.append(
f'result=$(curl -sf -X POST "http://127.0.0.1:${{PORT}}{spec.endpoint}" '
f'result=$(curl -sf -X POST "http://${{HOST}}:${{PORT}}{spec.endpoint}" '
'-H "Content-Type: application/json" -d "$payload" 2>/dev/null)'
)
lines.append('if [ $? -eq 0 ] && [ -n "$result" ]; then echo "$result"; exit 0; fi')
@@ -244,9 +255,10 @@ class GitNexusDockerEnvironment(DockerEnvironment):
Uses heredocs with quoted delimiter to avoid all quoting/escaping issues.
"""
port = str(self.eval_server_port)
host = self.eval_server_host
for spec in TOOL_SPECS.values():
script_content = self._render_tool_script(spec, port).strip()
script_content = self._render_tool_script(spec, port, host).strip()
# Use heredoc with quoted delimiter — prevents all variable expansion and quoting issues
self.execute({
"command": (
@@ -387,5 +399,6 @@ class GitNexusDockerEnvironment(DockerEnvironment):
"index_time_seconds": round(self.index_time, 2),
"skip_embeddings": self.skip_embeddings,
"eval_server_port": self.eval_server_port,
"eval_server_host": self.eval_server_host,
}
return base
Generated
+3 -3
View File
@@ -760,11 +760,11 @@ wheels = [
[[package]]
name = "idna"
version = "3.11"
version = "3.15"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/6f/6d/0703ccc57f3a7233505399edb88de3cbd678da106337b9fcde432b65ed60/idna-3.11.tar.gz", hash = "sha256:795dafcc9c04ed0c1fb032c2aa73654d8e8c5023a7df64a53f39190ada629902", size = 194582, upload-time = "2025-10-12T14:55:20.501Z" }
sdist = { url = "https://files.pythonhosted.org/packages/82/77/7b3966d0b9d1d31a36ddf1746926a11dface89a83409bf1483f0237aa758/idna-3.15.tar.gz", hash = "sha256:ca962446ea538f7092a95e057da437618e886f4d349216d2b1e294abfdb65fdc", size = 199245, upload-time = "2026-05-12T22:45:57.011Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/0e/61/66938bbb5fc52dbdf84594873d5b51fb1f7c7794e9c0f5bd885f30bc507b/idna-3.11-py3-none-any.whl", hash = "sha256:771a87f49d9defaf64091e6e6fe9c18d4833f140bd19464795bc32d966ca37ea", size = 71008, upload-time = "2025-10-12T14:55:18.883Z" },
{ url = "https://files.pythonhosted.org/packages/d2/23/408243171aa9aaba178d3e2559159c24c1171a641aa83b67bdd3394ead8e/idna-3.15-py3-none-any.whl", hash = "sha256:048adeaf8c2d788c40fee287673ccaa74c24ffd8dcf09ffa555a2fbb59f10ac8", size = 72340, upload-time = "2026-05-12T22:45:55.733Z" },
]
[[package]]
@@ -56,15 +56,15 @@ Generates repository documentation from the knowledge graph using an LLM. Requir
| Flag | Effect |
|------|--------|
| `--force` | Force full regeneration |
| `--force` | Force full regeneration, also required to re-gerenate an existing wiki in a different language |
| `--model <model>` | LLM model (default: minimax/minimax-m2.5) |
| `--base-url <url>` | LLM API base URL |
| `--api-key <key>` | LLM API key |
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
| `--gist` | Publish wiki as a public GitHub Gist |
| `--timeout <seconds>` | Per-attempt LLM request timeout in seconds (default: 60) |
| `--timeout <seconds>` | LLM request timeout in seconds (default: disabled) |
| `--retries <n>` | Max LLM retry attempts per request (default: 3) |
| `--lang <lang>` | Output language for generated documentation (e.g. english, chinese, spanish, japanese)|
### list — Show all indexed repos
```bash
+3 -1
View File
@@ -26,7 +26,7 @@ export type { PipelinePhase, PipelineProgress } from './pipeline.js';
// ─── Scope-based resolution — RFC #909 (Ring 1 #910) ────────────────────────
// Data model (RFC §2)
export type { SymbolDefinition } from './scope-resolution/symbol-definition.js';
export type { ParameterTypeClass, SymbolDefinition } from './scope-resolution/symbol-definition.js';
export type {
ScopeId,
DefId,
@@ -127,8 +127,10 @@ export { CLASS_KINDS, METHOD_KINDS, FIELD_KINDS } from './scope-resolution/regis
export type {
RegistryContext,
RegistryProviders,
OwnedMembersByOwnerLookup,
OwnerScopedContributor,
ArityVerdict,
ConstraintContext,
} from './scope-resolution/registries/context.js';
// Scope tree spine + position lookup (RFC §2.2 + §3.1; Ring 2 SHARED #912)
@@ -21,6 +21,7 @@
* (defined in `./types.ts`).
*/
import type { ParameterTypeClass } from './symbol-definition.js';
import type { Range, ScopeId } from './types.js';
/**
@@ -79,4 +80,11 @@ export interface ReferenceSite {
* (C#: `42` → `'int'`, `"alice"` → `'string'`).
*/
readonly argumentTypes?: readonly string[];
/**
* Optional per-argument type-shape sidecar for languages that need
* cv/ref/pointer distinctions during constraint filtering. This is
* intentionally separate from `argumentTypes`, which stays normalized
* for existing overload narrowing and conversion-rank logic.
*/
readonly argumentTypeClasses?: readonly ParameterTypeClass[];
}
@@ -13,7 +13,7 @@
*/
import type { NodeLabel } from '../../graph/types.js';
import type { SymbolDefinition } from '../symbol-definition.js';
import type { ParameterTypeClass, SymbolDefinition } from '../symbol-definition.js';
import type { Callsite, DefId } from '../types.js';
import type { DefIndex } from '../def-index.js';
import type { QualifiedNameIndex } from '../qualified-name-index.js';
@@ -30,10 +30,50 @@ export interface RegistryProviders {
* when absent, every candidate receives `'unknown'` (neutral signal).
*/
arityCompatibility?(callsite: Callsite, def: SymbolDefinition): ArityVerdict;
/**
* Language-specific constraint compatibility between a callsite and a
* candidate `def`. Mirrors `arityCompatibility` and shares its three-valued
* verdict shape; the third value `'unknown'` MUST keep the candidate
* (monotonicity: adding a predicate can only narrow correctly, never
* produce a wrong edge). Consulted by `narrowOverloadCandidates` after
* arity + type filters when a candidate carries `templateConstraints`.
*
* Optional; when absent the constraint filter is a pass-through. Languages
* with no constrained-overload semantics leave this undefined.
*/
constraintCompatibility?(
callsite: Callsite,
def: SymbolDefinition,
ctx: ConstraintContext,
): ArityVerdict;
}
export type ArityVerdict = 'compatible' | 'unknown' | 'incompatible';
/**
* Context threaded into `constraintCompatibility`. Kept minimal in the
* Tier-A scope (only `argumentTypes`, riding here until a separate
* `Callsite`-widening refactor moves them onto the call site directly).
* Future Tier-B graph-aware predicates (`is_base_of_v`, etc.) will widen
* this interface with `lookupTypeByName` and similar helpers.
*/
export interface ConstraintContext {
/**
* Per-slot argument types at the call site, normalized per the language
* adapter. Empty string means unknown. Same convention as
* `narrowOverloadCandidates`' `argTypes` parameter.
*/
readonly argumentTypes?: readonly string[];
/**
* Optional shape-preserving sidecar aligned with `argumentTypes`.
* Unknown or unsupported slots should be omitted by producers or
* marked with `indirection: 'unknown'`; consumers must preserve the
* monotonic fallback and return 'unknown' instead of guessing.
*/
readonly argumentTypeClasses?: readonly ParameterTypeClass[];
}
// ─── Owner-scoped contributor (concrete shape for `RegistryContributor`) ────
/**
@@ -60,6 +100,19 @@ export interface OwnerScopedContributor {
byName(name: string): readonly SymbolDefinition[];
}
/**
* Required owner-keyed lookup hook for Step 2 receiver/MRO member walks.
* Production callers wire this to the SemanticModel's authoritative
* method/field/nested-type registries so each `(ownerDefId, memberName)`
* probe is O(1). Implementations MUST return `[]` on an indexed miss —
* Step 2 treats `[]` as authoritative and does not consult `defs` for a
* fallback scan.
*/
export type OwnedMembersByOwnerLookup = (
ownerDefId: DefId,
memberName: string,
) => readonly SymbolDefinition[];
// ─── Top-level context threaded through every lookup ───────────────────────
export interface RegistryContext {
@@ -67,6 +120,7 @@ export interface RegistryContext {
readonly defs: DefIndex;
readonly qualifiedNames: QualifiedNameIndex;
readonly moduleScopes: ModuleScopeIndex;
readonly ownedMembersByOwner: OwnedMembersByOwnerLookup;
/**
* Method-dispatch index; required for method/field registries that
* honor `useReceiverTypeBinding`. Omit for class-only lookups.
@@ -27,8 +27,10 @@
* is true, resolve the receiver's type at `startScope` (from
* `scope.typeBindings`), then walk the MRO via
* `MethodDispatchIndex.mroFor(ownerDefId)`. Membership per owner comes
* through `RegistryContext.methodDispatch` + owner lookups into
* `scope.ownedDefs`; each hit records a raw signal with the owner's
* through an optional `RegistryContext.ownedMembersByOwner` hook when
* supplied (`undefined` → fall back to `defs.byId`; `[]` → indexed
* miss), otherwise via the compatibility fallback scan over
* `defs.byId`; each hit records a raw signal with the owner's
* MRO depth.
*
* **Step 3 — Owner-scoped contributor.** When
@@ -263,13 +265,14 @@ function walkReceiverTypeBinding(
// Walk the owner itself at depth 0, then its MRO chain.
const walk: DefId[] = [ownerDefId, ...ctx.methodDispatch.mroFor(ownerDefId)];
for (let mroDepth = 0; mroDepth < walk.length; mroDepth++) {
const currentOwnerId = walk[mroDepth]!;
let mroDepth = 0;
for (const currentOwnerId of walk) {
const members = collectOwnedMembers(currentOwnerId, name, ctx);
for (const def of members) {
if (!acceptedKinds.has(def.type)) continue;
recordTypeBindingHit(perCandidate, def, mroDepth, ownerDefId);
}
mroDepth++;
}
}
@@ -333,23 +336,7 @@ function collectOwnedMembers(
memberName: string,
ctx: RegistryContext,
): readonly SymbolDefinition[] {
// An owner's members are defs whose `ownerId === ownerDefId` and whose
// simple name matches `memberName`. We iterate `defs.byId` — O(D) per
// call today. A future by-owner index would make this O(K); tracked as
// a follow-up optimization before Ring 3 flips go production.
const out: SymbolDefinition[] = [];
for (const def of ctx.defs.byId.values()) {
if (def.ownerId !== ownerDefId) continue;
if (simpleNameOf(def) !== memberName) continue;
out.push(def);
}
return out;
}
function simpleNameOf(def: SymbolDefinition): string | undefined {
if (def.qualifiedName === undefined || def.qualifiedName.length === 0) return undefined;
const dot = def.qualifiedName.lastIndexOf('.');
return dot === -1 ? def.qualifiedName : def.qualifiedName.slice(dot + 1);
return ctx.ownedMembersByOwner(ownerDefId, memberName);
}
function recordTypeBindingHit(
@@ -11,6 +11,17 @@
import type { NodeLabel } from '../graph/types.js';
export interface ParameterTypeClass {
/** Normalized base type, matching the coarse `parameterTypes` vocabulary when known. */
base: string;
/** Top-level cv signal preserved from the original C++ parameter spelling. */
cv: 'none' | 'const' | 'volatile' | 'const volatile' | 'unknown';
/** Coarse value/reference/pointer shape. */
indirection: 'value' | 'lvalue-ref' | 'rvalue-ref' | 'pointer' | 'unknown';
/** Number of pointer markers when indirection is `pointer`; otherwise 0. */
pointerDepth: number;
}
export interface SymbolDefinition {
nodeId: string;
filePath: string;
@@ -26,12 +37,22 @@ export interface SymbolDefinition {
/** Per-parameter type names for overload disambiguation (e.g. ['int', 'String']).
* Populated when parameter types are resolvable from AST (any typed language). */
parameterTypes?: string[];
/** Additive per-parameter type shape sidecar for languages that need cv/ref/pointer distinctions.
* Does not participate in graph node identity unless a resolver explicitly opts in. */
parameterTypeClasses?: ParameterTypeClass[];
/** Raw return type text extracted from AST (e.g. 'User', 'Promise<User>') */
returnType?: string;
/** Declared type for non-callable symbols — fields/properties (e.g. 'Address', 'List<User>') */
declaredType?: string;
/** Generic/template specialization arguments for class-like symbols (e.g. ['User'], ['T*']). */
templateArguments?: string[];
/** Per-language constraint payload for template / generic overloads
* (e.g. C++ `enable_if_t<P, T>` predicate trees, C++20 `requires` clauses).
* Opaque to shared code — the producing language adapter owns the shape
* and is the only consumer. Read via the optional
* `ScopeResolver.constraintCompatibility` hook during overload narrowing.
* Absent for symbols that have no constraints (the common case). */
templateConstraints?: unknown;
/** Links Method/Constructor/Property to owning Class/Struct/Trait nodeId */
ownerId?: string;
}
+446 -397
View File
File diff suppressed because it is too large Load Diff
+19 -6
View File
@@ -18,7 +18,6 @@
"test:e2e:report": "playwright show-report"
},
"dependencies": {
"gitnexus-shared": "file:../gitnexus-shared",
"@langchain/anthropic": "^1.3.29",
"@langchain/core": "^1.1.44",
"@langchain/google-genai": "^2.1.30",
@@ -26,10 +25,11 @@
"@langchain/ollama": "^1.2.6",
"@langchain/openai": "^1.4.5",
"@sigma/edge-curve": "^3.1.0",
"@tailwindcss/vite": "^4.2.4",
"@tailwindcss/vite": "^4.3.0",
"axios": "^1.16.0",
"d3": "^7.9.0",
"dompurify": "^3.4.2",
"dompurify": "^3.4.3",
"gitnexus-shared": "file:../gitnexus-shared",
"graphology": "^0.26.0",
"graphology-indices": "^0.17.0",
"graphology-layout-force": "^0.2.4",
@@ -45,13 +45,13 @@
"react": "^19.2.5",
"react-dom": "^19.2.6",
"react-markdown": "^10.1.0",
"react-syntax-highlighter": "^16.1.0",
"react-syntax-highlighter": "^16.1.1",
"react-zoom-pan-pinch": "^4.0.3",
"remark-gfm": "^4.0.1",
"sigma": "^3.0.2",
"tailwindcss": "^4.2.4",
"uuid": "^14.0.0",
"zod": "^3.25.76"
"zod": "^4.4.3"
},
"devDependencies": {
"@babel/types": "^7.29.0",
@@ -64,7 +64,7 @@
"@types/react": "^19.2.14",
"@types/react-dom": "^19.2.3",
"@types/react-syntax-highlighter": "^15.5.13",
"@vercel/node": "^5.5.16",
"@vercel/node": "^5.8.2",
"@vitejs/plugin-react": "^5.1.4",
"@vitest/coverage-v8": "^4.1.5",
"jsdom": "^29.1.1",
@@ -73,5 +73,18 @@
"vite": "^8.0.11",
"vitest": "^4.1.5",
"wait-on": "^9.0.5"
},
"overrides": {
"@vercel/static-config": {
"ajv": "8.18.0"
},
"@vercel/node": {
"path-to-regexp": "6.3.0",
"undici": "6.24.0"
},
"@vercel/python-analysis": {
"minimatch": "10.2.3",
"smol-toml": "1.6.1"
}
}
}
+96 -4
View File
@@ -9,6 +9,7 @@
import { useState, useRef, useEffect, useId } from 'react';
import {
Github,
Gitlab,
FolderOpen,
Loader2,
Check,
@@ -26,15 +27,20 @@ import { AnalyzeProgress } from './AnalyzeProgress';
// ── Helpers ──────────────────────────────────────────────────────────────────
type InputMode = 'github' | 'local';
type InputMode = 'github' | 'gitlab' | 'local';
const GITHUB_RE = /^https?:\/\/(www\.)?github\.com\/[^/\s]+\/[^/\s]+/i;
const GITLAB_RE = /^https?:\/\/[^/\s]+\/[^/\s]+\/[^/\s]+(\/.*)?$/i;
const IS_WINDOWS = navigator.userAgent.toLowerCase().includes('win');
function isValidGithubUrl(value: string): boolean {
return GITHUB_RE.test(value.trim());
}
function isValidGitlabUrl(value: string): boolean {
return GITLAB_RE.test(value.trim());
}
// ── Mode tabs ────────────────────────────────────────────────────────────────
function ModeTabs({ mode, onChange }: { mode: InputMode; onChange: (m: InputMode) => void }) {
@@ -53,6 +59,19 @@ function ModeTabs({ mode, onChange }: { mode: InputMode; onChange: (m: InputMode
<Github className="h-3 w-3" />
GitHub URL
</button>
<button
role="tab"
aria-selected={mode === 'gitlab'}
onClick={() => onChange('gitlab')}
className={`flex flex-1 cursor-pointer items-center justify-center gap-1.5 rounded-md px-3 py-1.5 text-xs font-medium transition-all duration-150 ${
mode === 'gitlab'
? 'bg-accent text-white shadow-sm'
: 'text-text-muted hover:text-text-secondary'
} `}
>
<Gitlab className="h-3 w-3" />
GitLab URL
</button>
<button
role="tab"
aria-selected={mode === 'local'}
@@ -138,6 +157,7 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
const folderInputRef = useRef<HTMLInputElement>(null);
const [mode, setMode] = useState<InputMode>('github');
const [githubUrl, setGithubUrl] = useState('');
const [gitlabUrl, setGitlabUrl] = useState('');
const [localPath, setLocalPath] = useState('');
const [phase, setPhase] = useState<InternalPhase>('input');
const [validationError, setValidationError] = useState<string | null>(null);
@@ -162,6 +182,7 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
const handleModeChange = (m: InputMode) => {
setMode(m);
setGithubUrl('');
setGitlabUrl('');
setLocalPath('');
setValidationError(null);
};
@@ -175,13 +196,19 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
const canSubmit =
mode === 'github'
? isValidGithubUrl(githubUrl) && (phase === 'input' || phase === 'error')
: localPath.trim().length > 1 && (phase === 'input' || phase === 'error');
: mode === 'gitlab'
? isValidGitlabUrl(gitlabUrl) && (phase === 'input' || phase === 'error')
: localPath.trim().length > 1 && (phase === 'input' || phase === 'error');
const handleAnalyze = async () => {
if (mode === 'github' && !isValidGithubUrl(githubUrl)) {
setValidationError('Please enter a valid GitHub repository URL.');
return;
}
if (mode === 'gitlab' && !isValidGitlabUrl(gitlabUrl)) {
setValidationError('Please enter a valid GitLab repository URL.');
return;
}
if (mode === 'local' && localPath.trim().length < 2) {
setValidationError('Please enter a folder path.');
return;
@@ -191,12 +218,22 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
setPhase('starting');
try {
const request = mode === 'github' ? { url: githubUrl.trim() } : { path: localPath.trim() };
const request =
mode === 'github'
? { url: githubUrl.trim() }
: mode === 'gitlab'
? { url: gitlabUrl.trim() }
: { path: localPath.trim() };
const { jobId } = await startAnalyze(request);
jobIdRef.current = jobId;
setPhase('analyzing');
const nameSource = mode === 'github' ? githubUrl.trim() : localPath.trim();
const nameSource =
mode === 'github'
? githubUrl.trim()
: mode === 'gitlab'
? gitlabUrl.trim()
: localPath.trim();
const controller = streamAnalyzeProgress(
jobId,
(p) => setProgress(p),
@@ -297,6 +334,61 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
</div>
)}
{/* GitLab URL input */}
{showInput && mode === 'gitlab' && (
<div className="space-y-2">
<label
htmlFor={inputId}
className="block text-xs font-medium tracking-wider text-text-secondary uppercase"
>
GitLab Repository URL
</label>
<div
className={`flex items-center gap-3 rounded-xl border bg-void px-4 py-3.5 transition-all duration-200 ${
validationError && phase === 'error'
? 'border-red-500/50'
: isValidGitlabUrl(gitlabUrl)
? 'border-accent/50 shadow-[0_0_0_3px_rgba(124,58,237,0.08)]'
: 'border-border-default focus-within:border-accent/40'
} `}
>
<Gitlab className="h-4 w-4 shrink-0 text-text-muted" />
<input
id={inputId}
type="url"
value={gitlabUrl}
onChange={(e) => {
setGitlabUrl(e.target.value);
if (validationError) setValidationError(null);
}}
onKeyDown={(e) => {
if (e.key === 'Enter' && canSubmit && !isLoading) {
e.preventDefault();
handleAnalyze();
}
}}
disabled={isLoading}
placeholder="https://gitlab.com/owner/repo"
autoComplete="url"
spellCheck={false}
className="flex-1 border-none bg-transparent font-mono text-sm text-text-primary outline-none placeholder:text-text-muted disabled:opacity-50"
/>
{gitlabUrl.length > 10 && (
<div className="shrink-0">
{isValidGitlabUrl(gitlabUrl) ? (
<Check className="h-3.5 w-3.5 text-emerald-400" />
) : (
<AlertCircle className="h-3.5 w-3.5 text-text-muted" />
)}
</div>
)}
</div>
<p className="text-xs text-text-muted">
Supports GitLab.com and self-hosted GitLab instances.
</p>
</div>
)}
{/* Local folder input */}
{showInput && mode === 'local' && (
<div className="space-y-2">
+61
View File
@@ -123,6 +123,67 @@ export {
* defaults to `currentColor`, so Tailwind `text-*` utilities work the same as
* with any other icon in this module.
*/
/**
* GitLab tanuki mark — SVG path data from simple-icons (CC0-1.0).
*
* GitLab's logo (the tanuki/fox-head) is a registered trademark of GitLab Inc.
* We use it here only to indicate GitLab source-repo integration.
*
* API-compatible with `lucide-react` icons (`LucideProps`).
*/
export const Gitlab = forwardRef<SVGSVGElement, LucideProps>(function Gitlab(
{
size = 24,
color = 'currentColor',
className,
strokeWidth: _strokeWidth,
absoluteStrokeWidth: _absoluteStrokeWidth,
...rest
},
ref,
) {
const numericSize = typeof size === 'string' ? Number.parseFloat(size) : size;
const useSmallVariant = Number.isFinite(numericSize) && (numericSize as number) <= 16;
if (useSmallVariant) {
return (
<svg
ref={ref}
xmlns="http://www.w3.org/2000/svg"
width={size}
height={size}
viewBox="0 0 16 16"
fill={color}
className={className}
{...rest}
>
<path d="M8 15.282l1.855-5.717H6.145L8 15.282z" />
<path d="M8 15.282L6.145 9.565H2.333L8 15.282z" />
<path d="M2.333 9.565l-.944-2.942c-.09-.267.067-.553.333-.553h3.153L2.333 9.565z" />
<path d="M4.875 6.07L6.145 9.565H2.333l2.542-3.495z" />
<path d="M13.667 9.565l.944-2.942c.09-.267-.067-.553-.333-.553h-3.153l2.542 3.495z" />
<path d="M11.125 6.07L9.855 9.565h3.812l-2.542-3.495z" />
<path d="M8 15.282l1.855-5.717H6.145L8 15.282z" />
</svg>
);
}
return (
<svg
ref={ref}
xmlns="http://www.w3.org/2000/svg"
width={size}
height={size}
viewBox="0 0 24 24"
fill={color}
className={className}
{...rest}
>
<path d="m23.6004 9.5927-.0337-.0862L20.3.9814a.851.851 0 0 0-.3362-.405.8748.8748 0 0 0-.9997.0539.8748.8748 0 0 0-.29.4399l-2.2055 6.748H7.5375l-2.2057-6.748a.8573.8573 0 0 0-.29-.4412.8748.8748 0 0 0-.9997-.0537.8585.8585 0 0 0-.3362.4049L.4332 9.5015l-.0325.0862a6.0657 6.0657 0 0 0 2.0119 7.0105l.0113.0087.03.0213 4.976 3.7264 2.462 1.8633 1.4995 1.1321a1.0085 1.0085 0 0 0 1.2197 0l1.4995-1.1321 2.4619-1.8633 5.006-3.7489.0125-.01a6.0682 6.0682 0 0 0 2.0094-7.003z" />
</svg>
);
});
export const Github = forwardRef<SVGSVGElement, LucideProps>(function Github(
{
size = 24,
+1 -1
View File
@@ -1,3 +1,3 @@
{
"installCommand": "cd ../gitnexus-shared && npm install && npm run build && cd ../gitnexus-web && npm install"
"installCommand": "cd ../gitnexus-shared && npm install && npm run build && cd ../gitnexus-web && npm ci --include=dev"
}
+15 -2
View File
@@ -1,5 +1,18 @@
{
"permissions": {
"allow": ["mcp__plugin_claude-mem_mcp-search__get_observations"]
}
"allow": [
"mcp__plugin_claude-mem_mcp-search__get_observations",
"Skill(gitnexus-exploring)",
"Bash(npx gitnexus *)",
"mcp__obsidian-memory__search_nodes",
"mcp__obsidian-memory__add_observations",
"WebSearch",
"WebFetch(domain:cppreference.net)",
"Bash(xargs grep -l \"templateArguments\\\\|parameterTypes\")",
"Bash(gh issue *)",
"Bash(gh pr *)"
]
},
"enableAllProjectMcpServers": true,
"enabledMcpjsonServers": ["gitnexus"]
}
+12 -1
View File
@@ -151,7 +151,8 @@ Your AI agent gets these tools automatically:
```bash
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --repair-fts # Fast path: rebuild/verify only FTS indexes on existing index data
gitnexus analyze --force # Full rebuild: re-parse + graph rebuild + FTS rebuild
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
@@ -358,6 +359,16 @@ npx gitnexus analyze
For repositories with very large source files, `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` controls the worker job byte budget. The default is **8388608 bytes (8 MB)**.
### Worker pool resilience tuning
Three env vars expose the pool's resilience layers (respawn budget, cumulative-timeout cap, circuit breaker). Defaults are tuned for typical repos; bump them when an analyze legitimately needs more retries, or lower them to fail-fast on a known-bad shape.
| Variable | Default | Effect |
| ------------------------------------------------- | ------------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per slot before the slot is dropped from the active rotation. |
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Bounds exponentially-growing retry waits. |
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` | `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, dispatches require a fresh pool. |
## Privacy
- All processing happens locally on your machine
+175
View File
@@ -0,0 +1,175 @@
# Parse-throughput benchmark (scaffold)
> **Status: methodology + harness scaffold, no measurement data yet.**
> The Latest measurement table below contains `_TBD_` placeholders.
> This file ships intentionally without numbers — populating it
> requires a dedicated bench-pass against the U6 fixture (and ideally
> a real-world TS-root-scale repo) on consistent hardware, which is
> tracked as future work rather than gated on PR #1693's merge.
> Until the table is populated, the load-bearing perf-regression
> protection lives in `gitnexus/test/integration/parse-impl-large-fixture.test.ts`
> (U6, 30 s wall-clock budget via `Promise.race`).
Tracks `runChunkedParseAndResolve` wall-clock + peak heap on a synthetic
fixture so PR #1693's "analyze no longer hangs on TS-root-shaped loads"
claim is measurable, not just asserted by smoke tests. The harness
recipe below is deliberately small enough to re-run in a few minutes
when the bench-pass is undertaken.
---
## Methodology
### Fixture
Synthetic TypeScript repo, _not_ a clone of microsoft/TypeScript. CI cost
of cloning real-world repos is prohibitive; the synthetic shape exercises
the same pipeline paths (chunking, deferred extraction, cross-chunk
imports + heritage) without the disk-I/O overhead. Larger numbers can be
manually captured against real repos and cross-referenced here, but the
authoritative regression-tracking shape is the synthetic fixture so runs
are reproducible across hardware.
The fixture matches the structure pinned by
`gitnexus/test/integration/parse-impl-large-fixture.test.ts` (U6):
- 15 small modules (`mod0.ts` … `mod14.ts`), one exported function each.
- 1 dense `complex.ts` with 30 functions + 1 class + 1 interface.
- 1 `index.ts` re-exporting every symbol from every module.
`GITNEXUS_CHUNK_BYTE_BUDGET=64` forces multi-chunk parsing on this small
fixture — without that override the whole thing fits in one chunk and
the deferred-extraction path is not exercised end-to-end.
### What to measure
| Metric | How |
| --------------------------- | -------------------------------------------------------------------------------- |
| Wall-clock total | `Date.now()` delta around `runChunkedParseAndResolve` |
| Peak heap | Sample `process.memoryUsage().heapUsed` every 50 ms during the run; keep the max |
| Chunks observed | Count distinct `Parsing chunk X/Y` progress messages |
| `getStats()` final snapshot | Quarantined paths, dropped slots, breaker state |
### Hardware shape (record alongside each measurement)
- OS + version
- CPU model + logical core count
- RAM
- Node version
- gitnexus commit SHA (so the snapshot is anchored to a tree, not "main")
---
## Harness recipe
The U6 test (`test/integration/parse-impl-large-fixture.test.ts`) is the
checked-in mini-benchmark — it exercises the same fixture and bounds the
wall-clock at 30 s via `Promise.race`. To produce a richer snapshot for
this doc, run it under instrumentation:
```bash
# From the gitnexus/ subdir:
cd gitnexus
# Single-threaded baseline (sequential fallback):
npx vitest run test/integration/parse-impl-large-fixture.test.ts --reporter=verbose
# Worker-pool path (requires built dist/ — pre-built by `npm run build`):
npm run build && \
GITNEXUS_WORKER_POOL_SIZE=4 \
GITNEXUS_PARSE_CHUNK_CONCURRENCY=2 \
GITNEXUS_VERBOSE=1 \
npx vitest run test/integration/parse-impl-large-fixture.test.ts --reporter=verbose
```
For peak-heap sampling, wrap the dispatch call in a Node script that
polls `process.memoryUsage()`. A future helper at
`gitnexus/bench/scripts/parse-throughput.ts` would automate this — the
plan's stretch goal. Until that lands, capture peak heap manually via:
```bash
node --inspect=0 \
--require ./scripts/heap-sampler.js \
./node_modules/.bin/vitest run test/integration/parse-impl-large-fixture.test.ts
```
---
## Latest measurement
> _No measurement data has been collected yet — this file is the
> methodology + harness scaffold. The single recorded data point is the
> U6 wall-clock smoke baseline below; the worker-pool rows are
> placeholders for future bench-pass output._
The U6 integration test (`gitnexus/test/integration/parse-impl-large-fixture.test.ts`)
was observed completing the synthetic fixture in **~6 seconds** under
the sequential path (`skipWorkers: true`) on the development machine,
well under the 30 s `Promise.race` wall-clock budget. That number is a
smoke baseline only — recorded here for reference, not as a regression
target.
| Path | files/s | wall-clock | peak heap | chunks | quarantined |
| ------------------------------------------ | ------- | -------------------- | --------- | ------ | ----------- |
| Sequential (`skipWorkers: true`, U6 smoke) | _TBD_ | ~6 s _(observation)_ | _TBD_ | 17 | 0 |
| Worker pool, `--workers 4`, concurrency 2 | _TBD_ | _TBD_ | _TBD_ | _TBD_ | 0 |
| Worker pool, `--workers 1`, concurrency 1 | _TBD_ | _TBD_ | _TBD_ | _TBD_ | 0 |
**Hardware:** _TBD — record OS, CPU, RAM, Node version, gitnexus SHA at
the time of the bench-pass that populates the table above._
---
## Operator-tuning quick reference
Cross-links to the env vars documented in the [README](../../README.md#environment-variables).
Use this section as a starting point when the benchmark numbers above
suggest a tuning opportunity for your hardware shape.
- **CPU-bound, big repo, lots of cores:** raise `GITNEXUS_WORKER_POOL_SIZE`
past the default cap of 16. The 16-worker cap exists because past that
point main-thread merge / extraction dominates; if you've measurably
ruled that out, the env var lifts the cap explicitly. (See
`worker-pool.ts` `DEFAULT_POOL_SIZE_CAP`.)
- **Slow files (large minified JS, deep TS types):** raise
`GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS` past 30 000 ms. The cumulative
budget is 5× this value (U10 pins this) so a 60 s idle timeout permits
300 s of total retry-and-split wall-clock before quarantining the file.
- **Constrained container (cgroup CPU limit):** the pool now uses
`os.availableParallelism()` (U3 H2), which honors cgroup limits — no
manual `GITNEXUS_WORKER_POOL_SIZE` override needed unless the auto-
resolved value is too aggressive for your I/O budget.
- **Long-running host (eval-server, MCP daemon) running back-to-back
analyzes:** `--workers` is now threaded through `AnalyzeOptions`
(U2 B2), so per-invocation sizing is honored without `process.env`
state leaking across calls. `GITNEXUS_VERBOSE` is similarly snapshot/
restore-bracketed.
---
## What this benchmark does NOT measure
- **Real-repo performance.** The synthetic fixture is sized for CI; it
doesn't exercise the cumulative-load shape (50k files, occasional
pathological file) that drove the original PR #1693 hang report. Real-
repo numbers should be captured ad-hoc against the user's target repo
and cross-referenced here only as supplementary evidence.
- **Worker-pool resilience under real crashes.** That's verified by the
`worker-pool.test.ts` integration tests (real `process.exit`, real
`error` events, real protocol violations) and the unit suite. The
benchmark cares about throughput on the happy path.
- **IPC repack throughput.** Phase 3 of the PR #1693 plan introduces a
transferList + binary wire-format IPC repack (U16-U17). Once that
lands, an `IPC repack` row should be added to the "Latest measurement"
table above with before/after numbers on the same hardware.
---
## Related artifacts
- Plan: `docs/plans/2026-05-20-001-feat-pr1693-resilience-hardening-and-ipc-repack-plan.md`
- Integration test (mini-benchmark with wall-clock guard): `gitnexus/test/integration/parse-impl-large-fixture.test.ts` (U6)
- Operator env-var reference: `README.md` → Environment variables
- Resilience layer tests: `gitnexus/test/unit/worker-pool-resilience.test.ts`,
`worker-pool-cumulative-timeout.test.ts`,
`worker-pool-windows-quarantine.test.ts`,
`worker-pool-slot-generation.test.ts`
+318 -768
View File
File diff suppressed because it is too large Load Diff
+4 -7
View File
@@ -48,7 +48,7 @@
"test:integration": "vitest run test/integration",
"test:watch": "vitest",
"test:coverage": "vitest run --coverage",
"postinstall": "node scripts/build-tree-sitter-dart.cjs && node scripts/build-tree-sitter-proto.cjs",
"postinstall": "node scripts/materialize-vendor-grammars.cjs && node scripts/build-tree-sitter-dart.cjs && node scripts/build-tree-sitter-proto.cjs && node scripts/build-tree-sitter-swift.cjs",
"prepare": "node scripts/build.js",
"prepack": "node scripts/build.js"
},
@@ -60,7 +60,7 @@
"cli-progress": "^3.12.0",
"commander": "^14.0.3",
"cors": "^2.8.5",
"express": "^4.19.2",
"express": "^5.2.1",
"express-rate-limit": "^8.4.1",
"glob": "^13.0.6",
"graphology": "^0.26.0",
@@ -92,15 +92,12 @@
"optionalDependencies": {
"node-addon-api": "^8.0.0",
"node-gyp-build": "^4.8.0",
"tree-sitter-dart": "file:./vendor/tree-sitter-dart",
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-proto": "file:./vendor/tree-sitter-proto",
"tree-sitter-swift": "file:./vendor/tree-sitter-swift"
"tree-sitter-kotlin": "^0.3.8"
},
"devDependencies": {
"@types/cli-progress": "^3.11.6",
"@types/cors": "^2.8.17",
"@types/express": "^4.17.21",
"@types/express": "^5.0.6",
"@types/js-yaml": "^4.0.9",
"@types/node": "^25.6.0",
"@types/uuid": "^11.0.0",
@@ -1,4 +1,8 @@
#!/usr/bin/env node
/**
* Build tree-sitter-dart native binding in node_modules/ after materialize-vendor-grammars.cjs.
* Vendored source lives in vendor/ only; see #836 and #1728.
*/
const fs = require('fs');
const path = require('path');
const { execSync } = require('child_process');
+3 -4
View File
@@ -4,7 +4,7 @@
*
* Why this script exists:
* tree-sitter-proto is vendored under gitnexus/vendor/tree-sitter-proto/
* and declared as a `file:` optionalDependency. Previously, the vendored
* and copied into node_modules/ by materialize-vendor-grammars.cjs. Previously, the vendored
* package had its own `dependencies` and `install` script, which caused
* npm to create `vendor/tree-sitter-proto/node_modules/` and
* `vendor/tree-sitter-proto/build/` during install. Those directories
@@ -20,9 +20,8 @@
* gitnexus's own optionalDependencies, and moved native compilation here.
*
* What this does:
* Runs `npx node-gyp rebuild` inside `node_modules/tree-sitter-proto/`
* (which npm creates as a copy of vendor/tree-sitter-proto/ when
* resolving the file: dep). Build output lands in
* Runs `npx node-gyp rebuild` inside `node_modules/tree-sitter-proto/`.
* Build output lands in
* `node_modules/tree-sitter-proto/build/Release/tree_sitter_proto_binding.node`
* — under npm-managed territory, safe on upgrade.
*
@@ -0,0 +1,39 @@
#!/usr/bin/env node
/**
* Probe tree-sitter-swift prebuild availability at install time.
*
* The vendored package ships platform prebuilds; node-gyp-build selects the
* correct binary at require time. This script calls node-gyp-build once
* against the materialized package so a missing-prebuild failure surfaces
* as an install-time warning (with the rest of the gitnexus install
* succeeding) rather than as a runtime error the first time Swift parsing
* is requested. The result is discarded — it does not copy, register, or
* mutate anything; the runtime require() path in parser-loader does the
* actual load. Running this probe here instead of an npm `install` script
* on the vendored package preserves the #836 hygiene (no scripts.install
* inside vendor/).
*/
const fs = require('fs');
const path = require('path');
if (process.env.GITNEXUS_SKIP_OPTIONAL_GRAMMARS === '1') {
console.warn('[tree-sitter-swift] Skipping prebuild probe (GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1).');
process.exit(0);
}
const swiftDir = path.join(__dirname, '..', 'node_modules', 'tree-sitter-swift');
try {
if (!fs.existsSync(path.join(swiftDir, 'bindings', 'node', 'index.js'))) {
process.exit(0);
}
const nodeGypBuild = require('node-gyp-build');
nodeGypBuild(swiftDir);
} catch (err) {
console.warn('[tree-sitter-swift] Prebuild probe failed:', err.message);
console.warn(
'[tree-sitter-swift] Swift parsing will be unavailable. Non-Swift functionality is unaffected.',
);
process.exit(0);
}
+20 -4
View File
@@ -18,6 +18,22 @@ const ROOT = path.resolve(__dirname, '..');
const SHARED_ROOT = path.resolve(ROOT, '..', 'gitnexus-shared');
const DIST = path.join(ROOT, 'dist');
const SHARED_DEST = path.join(DIST, '_shared');
const DEFAULT_BUILD_TIMEOUT_MS = 300_000;
function getBuildTimeoutMs() {
const raw = process.env.GITNEXUS_BUILD_TIMEOUT_MS;
if (raw === undefined || raw.trim() === '') return DEFAULT_BUILD_TIMEOUT_MS;
const parsed = Number.parseInt(raw, 10);
if (Number.isFinite(parsed) && parsed > 0) return parsed;
console.warn(
`[build] ignoring invalid GITNEXUS_BUILD_TIMEOUT_MS=${JSON.stringify(raw)}; using ${DEFAULT_BUILD_TIMEOUT_MS}ms`,
);
return DEFAULT_BUILD_TIMEOUT_MS;
}
const BUILD_TIMEOUT_MS = getBuildTimeoutMs();
// ── 1. Build gitnexus-shared ───────────────────────────────────────
console.log('[build] compiling gitnexus-shared…');
@@ -25,11 +41,11 @@ const tscCmd =
process.platform === 'win32'
? path.join('node_modules', '.bin', 'tsc.cmd')
: path.join('node_modules', '.bin', 'tsc');
execSync(tscCmd, { cwd: SHARED_ROOT, stdio: 'inherit', timeout: 120_000 });
execSync(tscCmd, { cwd: SHARED_ROOT, stdio: 'inherit', timeout: BUILD_TIMEOUT_MS });
// ── 2. Build gitnexus ──────────────────────────────────────────────
console.log('[build] compiling gitnexus…');
execSync(tscCmd, { cwd: ROOT, stdio: 'inherit', timeout: 120_000 });
execSync(tscCmd, { cwd: ROOT, stdio: 'inherit', timeout: BUILD_TIMEOUT_MS });
// ── 3. Copy shared dist ────────────────────────────────────────────
console.log('[build] copying shared module into dist/_shared…');
@@ -82,9 +98,9 @@ if (fs.existsSync(path.join(WEB_ROOT, 'package.json'))) {
console.log('[build] building gitnexus-web…');
if (!fs.existsSync(path.join(WEB_ROOT, 'node_modules'))) {
console.log('[build] installing gitnexus-web dependencies…');
execSync('npm ci', { cwd: WEB_ROOT, stdio: 'inherit', timeout: 120_000 });
execSync('npm ci', { cwd: WEB_ROOT, stdio: 'inherit', timeout: BUILD_TIMEOUT_MS });
}
execSync('npm run build', { cwd: WEB_ROOT, stdio: 'inherit', timeout: 120_000 });
execSync('npm run build', { cwd: WEB_ROOT, stdio: 'inherit', timeout: BUILD_TIMEOUT_MS });
// Copy dist → gitnexus/web/ (shipped in the npm package)
fs.rmSync(WEB_DEST, { recursive: true, force: true });
@@ -0,0 +1,72 @@
#!/usr/bin/env node
/**
* Copy vendored tree-sitter grammars into node_modules/ using real files (fs.cpSync).
*
* Published gitnexus used to declare these as optionalDependencies with
* `file:./vendor/...`, which makes npm symlink/junction vendor → node_modules on
* install. Windows without Developer Mode often fails with EPERM (#1728).
*
* Vendor trees stay read-only in gitnexus/vendor/; build artifacts must only
* land under node_modules/ (see #836).
*/
const fs = require('fs');
const path = require('path');
const ROOT = path.join(__dirname, '..');
const VENDORED_GRAMMARS = ['tree-sitter-dart', 'tree-sitter-proto', 'tree-sitter-swift'];
if (process.env.GITNEXUS_SKIP_OPTIONAL_GRAMMARS === '1') {
console.warn(
'[gitnexus] Skipping vendored grammar materialize (GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1). Dart/Proto/Swift parsing will be unavailable.',
);
process.exit(0);
}
for (const name of VENDORED_GRAMMARS) {
const src = path.join(ROOT, 'vendor', name);
const dest = path.join(ROOT, 'node_modules', name);
if (!fs.existsSync(src)) {
console.warn(`[gitnexus] vendor/${name} missing; skipping materialize.`);
continue;
}
// Sequence: copy src → partial; rename dest → backup; rename partial → dest;
// remove backup. If any step fails, restore from backup so a previously-
// materialized grammar is never lost. Targets the #1728 EPERM scenario plus
// narrower failure modes (Windows AV scanner racing on rename, EBUSY mid-swap).
const partial = `${dest}.materialize-tmp`;
const backup = `${dest}.materialize-bak`;
try {
fs.mkdirSync(path.join(ROOT, 'node_modules'), { recursive: true });
fs.rmSync(partial, { recursive: true, force: true });
fs.rmSync(backup, { recursive: true, force: true });
fs.cpSync(src, partial, { recursive: true, verbatim: true });
if (fs.existsSync(dest)) {
fs.renameSync(dest, backup);
}
try {
fs.renameSync(partial, dest);
} catch (renameErr) {
// Best-effort rollback: restore the previous dest from backup.
if (fs.existsSync(backup)) {
try {
fs.renameSync(backup, dest);
} catch {
// If rollback also fails, the prior backup directory still exists on
// disk — the catch block below surfaces both errors via the warning.
}
}
throw renameErr;
}
fs.rmSync(backup, { recursive: true, force: true });
} catch (err) {
// Fail-soft: a single locked/inaccessible file (common on Windows) must not
// abort the whole gitnexus install. Matches build-tree-sitter-*.cjs pattern.
fs.rmSync(partial, { recursive: true, force: true });
console.warn(`[gitnexus] Could not materialize vendor/${name}: ${err.message}`);
console.warn(
`[gitnexus] ${name} parsing will be unavailable. Other functionality is unaffected.`,
);
}
}
+519 -29
View File
@@ -9,10 +9,11 @@
*/
import path from 'path';
import { execFileSync } from 'child_process';
import { spawn } from 'child_process';
import v8 from 'v8';
import cliProgress from 'cli-progress';
import { closeLbug } from '../core/lbug/lbug-adapter.js';
import { isWalCorruptionError, WAL_RECOVERY_SUGGESTION } from '../core/lbug/lbug-config.js';
import {
getStoragePaths,
getGlobalRegistryPath,
@@ -36,6 +37,7 @@ import { isHfDownloadFailure } from '../core/embeddings/hf-env.js';
// previous behaviour silently swallowed stack traces and made #1169
// indistinguishable from a no-op success on Windows.
const realStderrWrite = process.stderr.write.bind(process.stderr);
const realStdoutWrite = process.stdout.write.bind(process.stdout);
const writeFatalToStderr = (label: string, err: unknown): void => {
const isErr = err instanceof Error;
@@ -67,14 +69,354 @@ const installFatalHandlers = (): void => {
});
};
const HEAP_MB = 8192;
const HEAP_FLAG = `--max-old-space-size=${HEAP_MB}`;
const HEAP_MB = 16384;
const TEST_RESPAWN_HEAP_MB = Number(process.env.GITNEXUS_TEST_RESPAWN_HEAP_MB);
const RESPAWN_HEAP_MB =
Number.isFinite(TEST_RESPAWN_HEAP_MB) && TEST_RESPAWN_HEAP_MB > 0
? Math.floor(TEST_RESPAWN_HEAP_MB)
: HEAP_MB;
const HEAP_FLAG = `--max-old-space-size=${RESPAWN_HEAP_MB}`;
/** Increase default stack size (KB) to prevent stack overflow on deep class hierarchies. */
const STACK_KB = 4096;
const STACK_FLAG = `--stack-size=${STACK_KB}`;
const RESPAWN_OUTPUT_TAIL_CHARS = 1024 * 1024;
const RESPAWN_PROGRESS_ENV = 'GITNEXUS_RESPAWN_PROGRESS_TTY';
/** Re-exec the process with an 8GB heap and larger stack if we're currently below that. */
function ensureHeap(): boolean {
interface CliProgressTerminal {
cursorSave(): void;
cursorRestore(): void;
cursor(enabled: boolean): void;
lineWrapping(enabled: boolean): void;
cursorTo(x?: number | null, y?: number | null): void;
cursorRelative(dx?: number | null, dy?: number | null): void;
cursorRelativeReset(): void;
clearRight(): void;
clearLine(): void;
clearBottom(): void;
newline(): void;
write(s: string, rawWrite?: boolean): void;
isTTY(): boolean;
getWidth(): number;
}
const terminalColumns = (): number => {
const parsed = Number(process.env.COLUMNS);
return Number.isFinite(parsed) && parsed > 0 ? Math.floor(parsed) : 80;
};
const ANSI_ESCAPE_PATTERN =
/\x1B(?:\[[0-?]*[ -/]*[@-~]|\][^\x07]*(?:\x07|\x1B\\)|[PX^_][\s\S]*?\x1B\\|[78]|[@-Z\\-_])/y;
interface IntlSegmenterLike {
segment(input: string): Iterable<{ segment: string }>;
}
type IntlWithOptionalSegmenter = typeof Intl & {
Segmenter?: new (
locales?: string | string[],
options?: { granularity?: 'grapheme' },
) => IntlSegmenterLike;
};
const splitGraphemes = (text: string): string[] => {
const Segmenter = (Intl as IntlWithOptionalSegmenter).Segmenter;
if (Segmenter) {
return Array.from(
new Segmenter(undefined, { granularity: 'grapheme' }).segment(text),
(s) => s.segment,
);
}
return Array.from(text);
};
const isZeroWidthCodePoint = (codePoint: number): boolean =>
codePoint === 0x200d ||
(codePoint >= 0x0300 && codePoint <= 0x036f) ||
(codePoint >= 0x1ab0 && codePoint <= 0x1aff) ||
(codePoint >= 0x1dc0 && codePoint <= 0x1dff) ||
(codePoint >= 0x20d0 && codePoint <= 0x20ff) ||
(codePoint >= 0xfe00 && codePoint <= 0xfe0f) ||
(codePoint >= 0xfe20 && codePoint <= 0xfe2f);
const isWideCodePoint = (codePoint: number): boolean =>
codePoint >= 0x1100 &&
(codePoint <= 0x115f ||
codePoint === 0x2329 ||
codePoint === 0x232a ||
(codePoint >= 0x2e80 && codePoint <= 0xa4cf && codePoint !== 0x303f) ||
(codePoint >= 0xac00 && codePoint <= 0xd7a3) ||
(codePoint >= 0xf900 && codePoint <= 0xfaff) ||
(codePoint >= 0xfe10 && codePoint <= 0xfe19) ||
(codePoint >= 0xfe30 && codePoint <= 0xfe6f) ||
(codePoint >= 0xff00 && codePoint <= 0xff60) ||
(codePoint >= 0xffe0 && codePoint <= 0xffe6) ||
(codePoint >= 0x1f300 && codePoint <= 0x1faff) ||
(codePoint >= 0x20000 && codePoint <= 0x3fffd));
const visibleColumns = (text: string): number => {
let columns = 0;
for (const char of Array.from(text)) {
const codePoint = char.codePointAt(0);
if (codePoint === undefined || isZeroWidthCodePoint(codePoint)) continue;
columns += isWideCodePoint(codePoint) ? 2 : 1;
}
return columns;
};
const readAnsiEscapeAt = (text: string, index: number): string | undefined => {
ANSI_ESCAPE_PATTERN.lastIndex = index;
return ANSI_ESCAPE_PATTERN.exec(text)?.[0];
};
const truncateAnsiToColumns = (text: string, maxColumns: number): string => {
if (!Number.isFinite(maxColumns) || maxColumns <= 0) return '';
let output = '';
let columns = 0;
let index = 0;
while (index < text.length) {
const escape = readAnsiEscapeAt(text, index);
if (escape) {
output += escape;
index += escape.length;
continue;
}
const nextEscapeIndex = text.indexOf('\x1B', index);
const plainEnd = nextEscapeIndex === -1 ? text.length : nextEscapeIndex;
const plainText = text.slice(index, plainEnd);
for (const segment of splitGraphemes(plainText)) {
const width = visibleColumns(segment);
if (width > 0 && columns + width > maxColumns) return output;
output += segment;
columns += width;
}
index = plainEnd;
}
return output;
};
const createAnsiPipeTerminal = (stream: NodeJS.WriteStream): CliProgressTerminal => {
let linewrap = true;
let dy = 0;
const write = (s: string): void => {
stream.write(s);
};
const moveVertical = (delta: number): void => {
if (delta > 0) write(`\x1B[${delta}B`);
else if (delta < 0) write(`\x1B[${Math.abs(delta)}A`);
};
return {
cursorSave: () => write('\x1B7'),
cursorRestore: () => write('\x1B8'),
cursor: (enabled) => write(enabled ? '\x1B[?25h' : '\x1B[?25l'),
lineWrapping: (enabled) => {
linewrap = enabled;
write(enabled ? '\x1B[?7h' : '\x1B[?7l');
},
cursorTo: (x = null, y = null) => {
if (typeof y === 'number' && typeof x === 'number') {
write(`\x1B[${y + 1};${x + 1}H`);
return;
}
if (typeof x === 'number') {
write(x === 0 ? '\r' : `\x1B[${x + 1}G`);
}
},
cursorRelative: (dx = null, nextDy = null) => {
if (typeof dx === 'number' && dx !== 0) {
write(dx > 0 ? `\x1B[${dx}C` : `\x1B[${Math.abs(dx)}D`);
}
if (typeof nextDy === 'number' && nextDy !== 0) {
dy += nextDy;
moveVertical(nextDy);
}
},
cursorRelativeReset: () => {
moveVertical(-dy);
write('\r');
dy = 0;
},
clearRight: () => write('\x1B[0K'),
clearLine: () => write('\x1B[2K'),
clearBottom: () => write('\x1B[0J'),
newline: () => {
write('\n');
dy++;
},
write: (s, rawWrite = false) => {
const width = terminalColumns();
write(linewrap && rawWrite === false ? truncateAnsiToColumns(s, width) : s);
},
isTTY: () => true,
getWidth: terminalColumns,
};
};
const shouldBridgeRespawnProgressTty = (): boolean =>
process.stderr.isTTY === true || process.stdout.isTTY === true;
interface RespawnExit {
status?: number | null;
signal?: NodeJS.Signals | null;
stdout?: string;
stderr?: string;
message?: string;
}
const appendOutputTail = (tail: string, chunk: unknown): string => {
const text = Buffer.isBuffer(chunk)
? chunk.toString('utf8')
: typeof chunk === 'string'
? chunk
: String(chunk ?? '');
if (!text) return tail;
const next = tail + text;
return next.length > RESPAWN_OUTPUT_TAIL_CHARS ? next.slice(-RESPAWN_OUTPUT_TAIL_CHARS) : next;
};
/**
* Run the respawned analyzer while teeing child output through to the parent
* and keeping a bounded tail for crash classification.
*
* `execFileSync(..., { stdio: 'inherit' })` preserved live progress but hid
* stderr/stdout from the parent on abnormal exits. That made every
* SIGABRT/status-134 child look like an output-less V8 heap OOM, even when the
* terminal had already shown a native crash such as
* `libc++abi: ... Napi::Error`. Piped streams plus an explicit tee keeps the UX
* and gives `childProcessLikelyOom` the evidence it needs.
*/
const runRespawnedAnalyze = (
args: readonly string[],
env: NodeJS.ProcessEnv,
): Promise<RespawnExit> =>
new Promise((resolve) => {
let stdout = '';
let stderr = '';
let settled = false;
const finish = (exit: RespawnExit): void => {
if (settled) return;
settled = true;
resolve(exit);
};
const child = spawn(process.execPath, [...args], {
stdio: ['inherit', 'pipe', 'pipe'],
env,
});
child.stdout?.on('data', (chunk) => {
stdout = appendOutputTail(stdout, chunk);
realStdoutWrite(chunk);
});
child.stderr?.on('data', (chunk) => {
stderr = appendOutputTail(stderr, chunk);
realStderrWrite(chunk);
});
child.on('error', (err) => {
finish({
status: 1,
signal: null,
stdout,
stderr,
message: err instanceof Error ? err.message : String(err),
});
});
child.on('close', (status, signal) => {
finish({
status,
signal,
stdout,
stderr,
message: `Command failed: ${process.execPath} ${args.join(' ')}`,
});
});
});
/**
* Heuristic for "child re-exec likely died from V8 OOM".
*
* Platform-independent detection is best-effort: V8/Node usually emit stable
* heap-exhaustion phrases in stderr/message across Linux/macOS/Windows (for
* example "JavaScript heap out of memory" or "Reached heap limit"). When the
* child produced no output at all, we still treat status 134/SIGABRT as likely
* heap OOM. If stderr/stdout contains a native crash diagnostic, the output
* evidence wins and we do not print heap guidance.
*/
const childProcessLikelyOom = (err: unknown): boolean => {
if (!err || typeof err !== 'object') return false;
const e = err as {
status?: unknown;
signal?: unknown;
stderr?: unknown;
stdout?: unknown;
message?: unknown;
};
const hasHeapOomSignature = (v: unknown): boolean => {
const text = (
Buffer.isBuffer(v) ? v.toString('utf8') : typeof v === 'string' ? v : ''
).toLowerCase();
if (!text) return false;
return (
text.includes('javascript heap out of memory') ||
text.includes('reached heap limit') ||
text.includes('allocation failed - javascript heap out of memory') ||
text.includes('fatalprocessoutofmemory')
);
};
const fields = [e.message, e.stderr, e.stdout];
if (fields.some((v) => hasHeapOomSignature(v))) return true;
const hasAnyChildOutput = [e.stderr, e.stdout].some(
(v) => (Buffer.isBuffer(v) && v.length > 0) || (typeof v === 'string' && v.length > 0),
);
if (hasAnyChildOutput) return false;
return e.status === 134 || e.signal === 'SIGABRT';
};
const childProcessLikelyNativeAbort = (err: unknown): boolean => {
if (!err || typeof err !== 'object') return false;
const e = err as {
stderr?: unknown;
stdout?: unknown;
message?: unknown;
};
const hasNativeAbortSignature = (v: unknown): boolean => {
const text = (
Buffer.isBuffer(v) ? v.toString('utf8') : typeof v === 'string' ? v : ''
).toLowerCase();
if (!text) return false;
return (
text.includes('napi::error') ||
text.includes('libc++abi: terminating') ||
text.includes('abort trap') ||
text.includes('native stack') ||
text.includes('native worker') ||
text.includes('native binding')
);
};
return [e.message, e.stderr, e.stdout].some((v) => hasNativeAbortSignature(v));
};
const forceHeapOOMForTestIfEnabled = (): void => {
if (process.env.GITNEXUS_TEST_FORCE_HEAP_OOM !== '1') return;
// Allocate JS strings (not Buffers) so pressure lands on V8 heap itself.
// Buffers can allocate off-heap, which makes OOM triggering less reliable.
const chunks: string[] = [];
for (;;) chunks.push('x'.repeat(1024 * 1024));
};
/** Re-exec the process with a 16GB heap and larger stack if we're currently below that. */
async function ensureHeap(): Promise<boolean> {
const nodeOpts = process.env.NODE_OPTIONS || '';
if (nodeOpts.includes('--max-old-space-size')) return false;
@@ -86,19 +428,79 @@ function ensureHeap(): boolean {
const cliFlags = [HEAP_FLAG];
if (!nodeOpts.includes('--stack-size')) cliFlags.push(STACK_FLAG);
try {
execFileSync(process.execPath, [...cliFlags, ...process.argv.slice(1)], {
stdio: 'inherit',
env: { ...process.env, NODE_OPTIONS: `${nodeOpts} ${HEAP_FLAG}`.trim() },
});
} catch (e: any) {
process.exitCode = e.status ?? 1;
const childArgs = [...cliFlags, ...process.argv.slice(1)];
const childEnv = {
...process.env,
NODE_OPTIONS: `${nodeOpts} ${HEAP_FLAG}`.trim(),
};
if (shouldBridgeRespawnProgressTty()) childEnv[RESPAWN_PROGRESS_ENV] = '1';
const childExit = await runRespawnedAnalyze(childArgs, childEnv);
if (childExit.status !== 0 || childExit.signal) {
if (childProcessLikelyOom(childExit)) {
cliError(
` Analysis likely ran out of memory.\n` +
` Retry with a larger heap if your machine allows it:\n` +
` NODE_OPTIONS="--max-old-space-size=24576" gitnexus analyze [your-args]\n` +
` (Windows: set NODE_OPTIONS=--max-old-space-size=24576 && gitnexus analyze [your-args])\n` +
` If this persists, it may be a native crash unrelated to heap size.\n`,
{ recoveryHint: 'heap-oom-respawn' },
);
} else if (childProcessLikelyNativeAbort(childExit)) {
cliError(
` Analysis aborted in a native worker or native binding path.\n` +
` Try one of these recovery paths:\n` +
` gitnexus analyze --workers 0\n` +
` npm uninstall -g gitnexus && npm install -g gitnexus@latest\n` +
` Use Node 22 LTS if you are on a newer non-LTS runtime.\n`,
{ recoveryHint: 'native-worker-abort' },
);
}
const status =
typeof childExit.status === 'number' && childExit.status !== 0 ? childExit.status : 1;
process.exitCode = status;
}
return true;
}
/**
* GITNEXUS_* env vars that `analyzeCommand` writes for backward-compatible
* downstream consumption. Snapshotted at function entry and restored in the
* finally block so that programmatic callers (tests, long-running hosts)
* don't see leaked state across invocations. `GITNEXUS_WORKER_POOL_SIZE` is
* NOT in this list: that knob is threaded through `runFullAnalysis` options
* (see `workerPoolSize` plumbing) so the CLI never has to mutate `process.env`
* for it in the first place.
*/
const ANALYZE_CLI_ENV_KEYS = [
'GITNEXUS_VERBOSE',
'GITNEXUS_MAX_FILE_SIZE',
'GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS',
'GITNEXUS_EMBEDDING_THREADS',
'GITNEXUS_EMBEDDING_BATCH_SIZE',
'GITNEXUS_EMBEDDING_SUB_BATCH_SIZE',
'GITNEXUS_EMBEDDING_DEVICE',
'GITNEXUS_ANALYZE_PROGRESS_ACTIVE',
] as const;
type AnalyzeEnvSnapshot = Record<(typeof ANALYZE_CLI_ENV_KEYS)[number], string | undefined>;
const snapshotAnalyzeEnv = (): AnalyzeEnvSnapshot => {
const snap = {} as AnalyzeEnvSnapshot;
for (const k of ANALYZE_CLI_ENV_KEYS) snap[k] = process.env[k];
return snap;
};
const restoreAnalyzeEnv = (snap: AnalyzeEnvSnapshot): void => {
for (const k of ANALYZE_CLI_ENV_KEYS) {
const v = snap[k];
if (v === undefined) delete process.env[k];
else process.env[k] = v;
}
};
export interface AnalyzeOptions {
force?: boolean;
repairFts?: boolean;
/**
* Embedding generation toggle. Commander parses `--embeddings [limit]` as:
* - `undefined` when the flag is omitted
@@ -158,6 +560,8 @@ export interface AnalyzeOptions {
maxFileSize?: string;
/** Override worker sub-batch idle timeout in seconds. */
workerTimeout?: string;
/** Parse worker pool size; 0 disables workers (sequential fallback). */
workers?: string;
embeddingThreads?: string;
embeddingBatchSize?: string;
embeddingSubBatchSize?: string;
@@ -183,13 +587,30 @@ export const shouldGenerateCommunitySkillFiles = (
): boolean => Boolean(options?.skills && pipelineResult && !options?.indexOnly);
export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOptions) => {
if (ensureHeap()) return;
if (await ensureHeap()) return;
forceHeapOOMForTestIfEnabled();
// Install fatal handlers immediately after re-exec resolution so any
// async error that escapes the try/catch below (#1169) surfaces with
// a stack trace and a non-zero exit code instead of a silent exit 0.
installFatalHandlers();
// Snapshot the GITNEXUS_* env vars that the impl writes for downstream
// consumption, so they don't leak across `analyzeCommand` invocations in
// programmatic callers (tests, long-running hosts). `process.exit(0)` on
// the success path bypasses `finally` — intentional: when the process is
// exiting, restoration is moot. For early-return paths (validation
// errors) and the alreadyUpToDate fast path the finally restores the
// pre-call values.
const envSnap = snapshotAnalyzeEnv();
try {
await analyzeCommandImpl(inputPath, options);
} finally {
restoreAnalyzeEnv(envSnap);
}
};
const analyzeCommandImpl = async (inputPath?: string, options?: AnalyzeOptions): Promise<void> => {
if (options?.verbose) {
process.env.GITNEXUS_VERBOSE = '1';
}
@@ -210,6 +631,26 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
);
}
// `--workers` is threaded through `runFullAnalysis` options → PipelineOptions
// → createWorkerPool, intentionally bypassing the GITNEXUS_WORKER_POOL_SIZE
// env channel so this CLI surface never mutates `process.env` for pool size.
// Tests can therefore re-invoke analyzeCommand with different --workers
// values back-to-back and observe the value they passed, not whatever the
// previous call leaked.
let workerPoolSize: number | undefined;
if (options?.workers !== undefined) {
const parsedWorkers = Number(options.workers);
if (!Number.isInteger(parsedWorkers) || parsedWorkers < 0) {
cliError(
' --workers must be a non-negative integer. ' +
'Pass 0 to disable the worker pool (sequential fallback).\n',
);
process.exitCode = 1;
return;
}
workerPoolSize = parsedWorkers;
}
// Parse `--embeddings [limit]`: `true` → default cap, string → numeric cap
// (0 disables the cap entirely). Validated up here so failures match the
// sibling-validation pattern (exit before bar.start() — otherwise
@@ -275,6 +716,15 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
process.env.GITNEXUS_EMBEDDING_DEVICE = options.embeddingDevice;
}
if (options?.repairFts && options?.force) {
cliError(
' Cannot combine `--repair-fts` with `--force`. ' +
'Use `--repair-fts` for fast FTS-only repair, or `--force` for a full rebuild.\n',
);
process.exitCode = 1;
return;
}
console.log('\n GitNexus Analyzer\n');
// `--index-only` is the stronger contract — it suppresses every form of file
@@ -360,19 +810,25 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
}
// ── CLI progress bar setup ─────────────────────────────────────────
const bar = new cliProgress.SingleBar(
{
format: ' {bar} {percentage}% | {phase}',
barCompleteChar: '\u2588',
barIncompleteChar: '\u2591',
hideCursor: true,
barGlue: '',
autopadding: true,
clearOnComplete: false,
stopOnComplete: false,
},
cliProgress.Presets.shades_grey,
);
const barOptions: cliProgress.Options & { terminal?: CliProgressTerminal } = {
format: ' {bar} {percentage}% | {phase}',
barCompleteChar: '\u2588',
barIncompleteChar: '\u2591',
hideCursor: true,
barGlue: '',
autopadding: true,
clearOnComplete: false,
stopOnComplete: false,
};
if (process.env[RESPAWN_PROGRESS_ENV] === '1' && process.stderr.isTTY !== true) {
// Heap respawn pipes stderr so the parent can classify native/OOM crashes.
// The parent was a real TTY when it opted into this env var, so forward
// ANSI cursor controls through the pipe instead of cli-progress' non-TTY
// newline mode. That keeps one-line redraw UX while retaining stderr tail
// capture for diagnostics.
barOptions.terminal = createAnsiPipeTerminal(process.stderr);
}
const bar = new cliProgress.SingleBar(barOptions, cliProgress.Presets.shades_grey);
bar.start(100, 0, { phase: 'Initializing...' });
@@ -406,7 +862,7 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
// eslint-disable-next-line no-console -- intentional console-routing for progress bar UX
const origError = console.error.bind(console);
let barCurrentValue = 0;
const barLog = (...args: any[]) => {
const barLog = (...args: unknown[]) => {
process.stdout.write('\x1b[2K\r');
origLog(args.map((a) => (typeof a === 'string' ? a : String(a))).join(' '));
bar.update(barCurrentValue);
@@ -416,6 +872,7 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
console.warn = barLog;
// eslint-disable-next-line no-console -- intentional console-routing for progress bar UX
console.error = barLog;
process.env.GITNEXUS_ANALYZE_PROGRESS_ACTIVE = '1';
// Track elapsed time per phase
let lastPhaseLabel = 'Initializing...';
@@ -453,9 +910,11 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
// needs a fresh pipelineResult. Has no bearing on the registry
// collision guard (see allowDuplicateName below).
force: options?.force || options?.skills,
repairFts: options?.repairFts,
embeddings: embeddingsEnabled,
embeddingsNodeLimit,
dropEmbeddings: options?.dropEmbeddings,
verbose: options?.verbose,
skipGit: options?.skipGit,
skipAgentsMd,
skipSkills,
@@ -471,6 +930,10 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
// be able to accept the duplicate name without also paying the
// cost of a full pipeline re-index. See #829 review round 2.
allowDuplicateName: options?.allowDuplicateName,
// Worker pool size threaded from --workers, replacing the previous
// GITNEXUS_WORKER_POOL_SIZE env mutation. `undefined` defers to the
// env / auto-formula fallback inside the pipeline.
workerPoolSize,
},
{
onProgress: (_phase, percent, message) => {
@@ -500,6 +963,19 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
return;
}
if (result.ftsRepairedOnly) {
clearInterval(elapsedTimer);
process.removeListener('SIGINT', sigintHandler);
console.log = origLog;
// eslint-disable-next-line no-console -- restoring after intentional progress-bar routing
console.warn = origWarn;
// eslint-disable-next-line no-console -- restoring after intentional progress-bar routing
console.error = origError;
bar.stop();
console.log(' FTS indexes repaired successfully\n');
return;
}
// Post-finalize invariant (#1169): runFullAnalysis nominally writes
// meta.json and registers the repo, but on Windows it has been
// observed to return successfully with neither artifact present
@@ -595,7 +1071,7 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
}
console.log('');
} catch (err: any) {
} catch (err: unknown) {
clearInterval(elapsedTimer);
process.removeListener('SIGINT', sigintHandler);
console.log = origLog;
@@ -605,7 +1081,7 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
console.error = origError;
bar.stop();
const msg = err.message || String(err);
const msg = err instanceof Error ? err.message : String(err);
// Registry name-collision from --name (#829) — surface as an
// actionable error rather than a generic stack-trace.
@@ -638,6 +1114,20 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
return;
}
// WAL corruption — the index file is unreadable. Give a clear recovery
// path without a confusing stack trace (the native error message alone
// is enough signal).
if (isWalCorruptionError(err) || msg.includes('LadybugDB WAL corruption')) {
cliError(
` The GitNexus index has a corrupted WAL file.\n` +
` This usually happens when a previous analysis was interrupted mid-write.\n` +
` ${WAL_RECOVERY_SUGGESTION}\n`,
{ recoveryHint: 'wal-corruption' },
);
process.exitCode = 1;
return;
}
// HF download failure — show clean guidance without the raw stack trace.
// Checked before writeFatalToStderr so the user sees one focused message
// rather than a stack-trace dump followed by a second remediation block.
+104 -9
View File
@@ -14,9 +14,14 @@
* Agent bash cmd → curl localhost:PORT/tool/query → eval-server → LocalBackend → format → text
*
* Usage:
* gitnexus eval-server # default port 4848
* gitnexus eval-server --port 4848 # explicit port
* gitnexus eval-server --idle-timeout 300 # auto-shutdown after 300s idle
* gitnexus eval-server # default port 4848, binds 127.0.0.1
* gitnexus eval-server --port 4848 # explicit port
* gitnexus eval-server --host 0.0.0.0 # reachable from other VMs / containers
* gitnexus eval-server --idle-timeout 300 # auto-shutdown after 300s idle
*
* READY signal format: GITNEXUS_EVAL_SERVER_READY:<host>:<port>
* IPv4: GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848
* IPv6: GITNEXUS_EVAL_SERVER_READY:[::1]:4848
*
* API:
* POST /tool/:name — Call a tool. Body is JSON arguments. Returns formatted text.
@@ -25,16 +30,30 @@
*/
import http from 'http';
import { isIPv4, isIPv6 } from 'node:net';
import { writeSync } from 'node:fs';
import { LocalBackend } from '../mcp/local/local-backend.js';
import { logger } from '../core/logger.js';
import { cliInfo, cliWarn } from './cli-message.js';
import { cliInfo, cliWarn, cliError } from './cli-message.js';
export interface EvalServerOptions {
port?: string;
host?: string;
idleTimeout?: string;
}
/**
* Validate the --host value. Accepts IPv4, IPv6, or "localhost".
* Returns the host string unchanged, or null if invalid.
* "localhost" is passed through so the OS resolves it to the correct loopback
* address (127.0.0.1 or ::1) at bind time rather than forcing IPv4.
*/
export function validateHost(raw: string): string | null {
if (raw === 'localhost') return raw;
if (isIPv4(raw) || isIPv6(raw)) return raw;
return null;
}
// ─── Text Formatters ──────────────────────────────────────────────────
// Convert structured JSON results into compact, LLM-friendly text.
// Design: minimize tokens, maximize actionability.
@@ -330,6 +349,22 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
const port = parseInt(options?.port || '4848');
const idleTimeoutSec = parseInt(options?.idleTimeout || '0');
const rawHost = options?.host ?? '127.0.0.1';
const host = validateHost(rawHost);
if (!host) {
cliError(
`Invalid --host value "${rawHost}":\n` +
` Must be an IP address or "localhost".\n\n` +
` Examples:\n` +
` gitnexus eval-server --host 127.0.0.1 (loopback only, default)\n` +
` gitnexus eval-server --host 0.0.0.0 (all network interfaces)\n` +
` gitnexus eval-server --host 192.168.1.5 (specific interface)\n` +
` gitnexus eval-server --host localhost (OS-resolved loopback)\n`,
{ flag: '--host', value: rawHost },
);
process.exit(1);
}
const backend = new LocalBackend();
const ok = await backend.init();
@@ -426,12 +461,72 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
}
});
server.listen(port, '127.0.0.1', () => {
server.on('error', (err: NodeJS.ErrnoException) => {
if (err.code === 'EADDRINUSE') {
cliError(
`\nGitNexus eval-server failed to start:\n` +
` Port ${port} is already in use.\n\n` +
` Either:\n` +
` 1. Stop the process already using port ${port}\n` +
` 2. Use a different port: gitnexus eval-server --port 4849\n`,
{ code: err.code, port, host },
);
} else if (err.code === 'EADDRNOTAVAIL') {
// "localhost" may resolve to ::1 on IPv6-only systems; treat it as
// potentially IPv6 so the user gets the right diagnostic hint.
const isIPv6Host = isIPv6(host) || host === 'localhost';
cliError(
`\nGitNexus eval-server failed to start:\n` +
` Address ${host} is not available on this machine.\n\n` +
(isIPv6Host
? ` Address ${host} resolved but is not reachable — IPv6 may be disabled, or the loopback interface may be unavailable.\n` +
` Docker containers and many CI environments disable IPv6 by default.\n\n`
: ` The --host value must be an IP assigned to a local network interface.\n` +
` Run \`ip addr\` (Linux) or \`ipconfig\` (Windows) to list available addresses.\n\n`) +
` Common fixes:\n` +
` gitnexus eval-server --host 127.0.0.1 (loopback, this machine only)\n` +
` gitnexus eval-server --host 0.0.0.0 (all interfaces, reachable from other VMs)\n`,
{ code: err.code, port, host },
);
} else if (err.code === 'EACCES') {
cliError(
`\nGitNexus eval-server failed to start:\n` +
` Permission denied binding to port ${port}.\n\n` +
` Ports below 1024 require elevated privileges.\n` +
` Use a port above 1024: gitnexus eval-server --port 4848\n`,
{ code: err.code, port, host },
);
} else {
cliError(`\nGitNexus eval-server failed to start:\n ${err.message}\n`, {
code: err.code,
port,
host,
});
}
process.exit(1);
});
server.listen(port, host, () => {
// Plain-text banner for the human watching stderr; structured record
// for log aggregation (split into two so the user sees a real banner
// not `{"level":30,"msg":"...","port":4747,"endpoints":[...]}`).
// Use server.address() so the banner and READY signal reflect what the OS
// actually bound to, not the input host string. This matters when "localhost"
// is passed: the OS may resolve it to ::1 on some systems.
const addr = server.address();
// server.listen callback only fires after a successful TCP bind, so
// server.address() is guaranteed to return an AddressInfo object here.
if (typeof addr !== 'object' || addr === null) {
cliError(
`\nGitNexus eval-server: unexpected server.address() value after bind: ${JSON.stringify(addr)}\n`,
);
process.exit(1);
}
const boundPort = addr.port;
const boundAddress = addr.address;
const displayHost = boundAddress.includes(':') ? `[${boundAddress}]` : boundAddress;
const bannerLines = [
`GitNexus eval-server: listening on http://127.0.0.1:${port}`,
`GitNexus eval-server: listening on http://${displayHost}:${boundPort}`,
` POST /tool/query — search execution flows`,
` POST /tool/context — 360-degree symbol view`,
` POST /tool/impact — blast radius analysis`,
@@ -443,8 +538,8 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
bannerLines.push(` Auto-shutdown after ${idleTimeoutSec}s idle`);
}
cliInfo(bannerLines.join('\n'), {
port,
host: '127.0.0.1',
port: boundPort,
host,
idleTimeoutSec: idleTimeoutSec > 0 ? idleTimeoutSec : undefined,
endpoints: [
'POST /tool/query',
@@ -457,7 +552,7 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
});
try {
// Use fd 1 directly — LadybugDB captures process.stdout (#324)
writeSync(1, `GITNEXUS_EVAL_SERVER_READY:${port}\n`);
writeSync(1, `GITNEXUS_EVAL_SERVER_READY:${displayHost}:${boundPort}\n`);
} catch {
// stdout may not be available (e.g., broken pipe)
}
+19 -1
View File
@@ -23,6 +23,7 @@ program
.command('analyze [path]')
.description('Index a repository (full analysis)')
.option('-f, --force', 'Force full re-index even if up to date')
.option('--repair-fts', 'Repair/rebuild search FTS indexes without full re-analysis')
.option(
'--embeddings [limit]',
'Enable embedding generation for semantic search (off by default). ' +
@@ -70,6 +71,10 @@ program
'--worker-timeout <seconds>',
'Worker sub-batch idle timeout before retry/fallback. Default: 30.',
)
.option(
'--workers <n>',
'Parse worker pool size. Default: cores-1 capped at 16. Pass 0 to disable workers (sequential).',
)
.option('--embedding-threads <n>', 'Limit local ONNX embedding CPU threads')
.option('--embedding-batch-size <n>', 'Number of nodes per embedding batch')
.option('--embedding-sub-batch-size <n>', 'Number of chunks per embedding model call')
@@ -81,6 +86,11 @@ program
' GITNEXUS_MAX_FILE_SIZE=N Override large-file skip threshold (KB). Default 512, max 32768.\n' +
' GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=N Worker idle timeout in milliseconds. Default 30000.\n' +
' GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES=N Worker job byte budget. Default 8388608.\n' +
' GITNEXUS_WORKER_POOL_SIZE=N Parse worker count override. Default cores-1 capped at 16.\n' +
' GITNEXUS_PARSE_CHUNK_CONCURRENCY=N Concurrent in-flight parse chunks. Default 2.\n' +
' GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT=N Max replacement spawns per slot before drop. Default 3.\n' +
' GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS=N Total retry wall-time per job. Default 5x sub-batch timeout.\n' +
' GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD=N Per-slot deaths to trip circuit breaker. Default max(3, poolSize).\n' +
' GITNEXUS_EMBEDDING_THREADS=N Limit local ONNX CPU threads for --embeddings.\n' +
' GITNEXUS_SEMANTIC_EXACT_SCAN_LIMIT=N Max embedding chunks for exact-scan fallback. Default 10000.\n' +
'\nTip: `.gitnexusignore` supports `.gitignore`-style negation. Add e.g.\n' +
@@ -161,11 +171,15 @@ program
)
.option('--no-reasoning-model', 'Disable reasoning model mode (overrides saved config)')
.option('--concurrency <n>', 'Parallel LLM calls (default: 3)', '3')
.option('--timeout <seconds>', 'Per-attempt LLM request timeout in seconds (default: 60)')
.option('--timeout <seconds>', 'LLM request timeout in seconds (default: disabled)')
.option('--retries <n>', 'Max LLM retry attempts per request (default: 3)')
.option('--gist', 'Publish wiki as a public GitHub Gist after generation')
.option('-v, --verbose', 'Enable verbose output (show LLM commands and responses)')
.option('--review', 'Stop after grouping to review module structure before generating pages')
.option(
'--lang <lang>',
'Output language for generated documentation (e.g. english, chinese, spanish, japanese)',
)
.action(createLazyAction(() => import('./wiki.js'), 'wikiCommand'));
program
@@ -237,6 +251,10 @@ program
.command('eval-server')
.description('Start lightweight HTTP server for fast tool calls during evaluation')
.option('-p, --port <port>', 'Port number', '4848')
.option(
'--host <host>',
'Bind address (default: 127.0.0.1, use 0.0.0.0 to expose to all interfaces)',
)
.option('--idle-timeout <seconds>', 'Auto-shutdown after N seconds idle (0 = disabled)', '0')
.action(createLazyAction(() => import('./eval-server.js'), 'evalServerCommand'));
+12 -8
View File
@@ -1,15 +1,18 @@
/**
* Optional grammar availability check.
*
* tree-sitter-dart and tree-sitter-proto are optionalDependencies that
* require a `node-gyp rebuild` at install time. The build can be skipped
* via GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1 (postinstall scripts), or it can
* silently soft-fail when the C++ toolchain is missing.
* tree-sitter-dart, tree-sitter-proto, and tree-sitter-swift are vendored
* under vendor/ and materialized into node_modules/ at postinstall. Dart
* and Proto are built from source with node-gyp; Swift ships platform
* prebuilds activated via node-gyp-build. All three can be skipped via
* GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1 (postinstall scripts), or can silently
* soft-fail when the toolchain is missing (Dart/Proto) or no prebuild
* matches the host platform (Swift).
*
* Either path produces the same observable: the .node binding is absent
* at runtime. This helper detects that condition and surfaces a single
* stderr line per missing grammar so users learn why .dart/.proto support
* is unavailable instead of silently getting a degraded index.
* stderr line per missing grammar so users learn why .dart/.proto/.swift
* support is unavailable instead of silently getting a degraded index.
*/
import { createRequire } from 'module';
@@ -29,6 +32,7 @@ interface OptionalGrammar {
const OPTIONAL_GRAMMARS: OptionalGrammar[] = [
{ name: 'tree-sitter-dart', pkg: 'tree-sitter-dart', extensions: ['.dart'] },
{ name: 'tree-sitter-proto', pkg: 'tree-sitter-proto', extensions: ['.proto'] },
{ name: 'tree-sitter-swift', pkg: 'tree-sitter-swift', extensions: ['.swift'] },
];
export interface MissingGrammar {
@@ -40,8 +44,8 @@ export interface MissingGrammar {
* Returns the list of optional grammars whose native binding cannot be
* loaded. Actually `require()`s the package — `require.resolve` would
* locate the entry path even when the `.node` binding is absent (the
* `file:` package directory is installed regardless of postinstall
* outcome), giving false negatives for the exact users we want to warn:
* package directory exists without a working `.node` binding), giving false
* negatives for the exact users we want to warn:
* those who installed with `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` or whose
* native rebuild soft-failed for missing toolchain.
*
+8 -1
View File
@@ -1,6 +1,7 @@
import { createServer } from '../server/api.js';
import { logger, flushLoggerSync } from '../core/logger.js';
import { cliError } from './cli-message.js';
import { isWalCorruptionError, WAL_RECOVERY_SUGGESTION } from '../core/lbug/lbug-config.js';
// Catch anything that would cause a silent exit. Pino v10's default
// destination is `sync: false` (SonicBoom buffered) — call
@@ -34,7 +35,13 @@ export const serveCommand = async (options?: { port?: string; host?: string }) =
try {
await createServer(port, host);
} catch (err: any) {
if (err.code === 'EADDRINUSE') {
if (isWalCorruptionError(err)) {
cliError(
`\nGitNexus server could not start: the index has a corrupted WAL file.\n` +
` ${WAL_RECOVERY_SUGGESTION}\n`,
{ recoveryHint: 'wal-corruption' },
);
} else if (err.code === 'EADDRINUSE') {
cliError(
`\nFailed to start GitNexus server:\n` +
` ${err.message || err}\n\n` +
+6 -4
View File
@@ -61,11 +61,13 @@ function resolveGitnexusBin(): string | null {
.filter(Boolean);
if (isWin) {
// On Windows, `where` returns multiple entries (e.g. the POSIX shell
// script AND the .cmd/.bat wrapper). Prefer the wrapper because
// child_process.spawn() cannot execute a shell script directly.
// On Windows, npm global installs can surface multiple launchers for the
// same package (e.g. a POSIX shell shim plus .cmd/.bat wrappers). Claude
// and the other MCP hosts need a directly spawnable command path, so only
// accept the Windows wrapper. If it is missing, fall back to the slower
// npx entry instead of persisting a non-spawnable shim path.
const cmdLine = lines.find((l) => /\.(cmd|bat)$/i.test(l));
return cmdLine || lines[0] || null;
return cmdLine || null;
}
return lines[0] || null;
+54 -6
View File
@@ -35,6 +35,24 @@ export interface WikiCommandOptions {
review?: boolean;
timeout?: string;
retries?: string;
lang?: string;
}
function parsePositiveIntegerOption(
value: string | undefined,
flag: string,
multiplier = 1,
): number | undefined {
if (value === undefined) return undefined;
const trimmed = value.trim();
if (!/^[1-9]\d*$/.test(trimmed)) {
throw new Error(`${flag} must be a positive integer`);
}
const parsed = parseInt(trimmed, 10);
if (parsed > Math.floor(Number.MAX_SAFE_INTEGER / multiplier)) {
throw new Error(`${flag} is too large`);
}
return parsed;
}
/**
@@ -89,6 +107,24 @@ function prompt(question: string, hide = false): Promise<string> {
}
export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptions) => {
// Snapshot GITNEXUS_VERBOSE at entry — wikiCommand mutates it (the impl
// below) so cursor-client (process.env-driven) sees the right value during
// this run. Restored in finally so back-to-back wiki calls in long-running
// hosts don't leak verbose state from one invocation to the next. Pairs
// with the same snapshot/restore pattern in `analyzeCommand`.
const originalVerbose = process.env.GITNEXUS_VERBOSE;
try {
await wikiCommandImpl(inputPath, options);
} finally {
if (originalVerbose === undefined) {
delete process.env.GITNEXUS_VERBOSE;
} else {
process.env.GITNEXUS_VERBOSE = originalVerbose;
}
}
};
const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions): Promise<void> => {
// Set verbose mode globally for cursor-client to pick up
if (options?.verbose) {
process.env.GITNEXUS_VERBOSE = '1';
@@ -127,6 +163,17 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
return;
}
let timeoutSeconds: number | undefined;
let retries: number | undefined;
try {
timeoutSeconds = parsePositiveIntegerOption(options?.timeout, '--timeout', 1000);
retries = parsePositiveIntegerOption(options?.retries, '--retries');
} catch (error) {
console.log(` Error: ${(error as Error).message}\n`);
process.exitCode = 1;
return;
}
// ── Resolve LLM config (with interactive fallback) ─────────────────
// Save any CLI overrides immediately
if (
@@ -350,13 +397,11 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
}
// ── Apply per-run overrides not saved to config ────────────────────
if (options?.timeout) {
const secs = parseInt(options.timeout, 10);
if (!isNaN(secs) && secs > 0) llmConfig.requestTimeoutMs = secs * 1000;
if (timeoutSeconds !== undefined) {
llmConfig.requestTimeoutMs = timeoutSeconds * 1000;
}
if (options?.retries) {
const n = parseInt(options.retries, 10);
if (!isNaN(n) && n > 0) llmConfig.maxAttempts = n;
if (retries !== undefined) {
llmConfig.maxAttempts = retries;
}
// ── Setup progress bar with elapsed timer ──────────────────────────
@@ -395,6 +440,7 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
force: options?.force,
concurrency: options?.concurrency ? parseInt(options.concurrency, 10) : undefined,
reviewOnly: options?.review,
lang: options?.lang,
};
const generator = new WikiGenerator(
@@ -563,6 +609,8 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
if (err.message?.includes('No source files')) {
console.log(`\n ${err.message}\n`);
} else if (err.message?.includes('LLM request timed out after')) {
console.log(`\n Timeout: ${err.message}\n`);
} else if (err.message?.includes('content filter')) {
// Content filter block — actionable message
console.log(`\n Content Filter: ${err.message}\n`);
@@ -13,7 +13,12 @@ import type { HttpDetection, HttpLanguagePlugin } from './types.js';
* - FastAPI `@app.get("/path")` provider decorators
* - `requests.get/post/...("url")` consumer calls
* - Generic `requests.request("METHOD", "url")` consumer calls
* - `httpx.AsyncClient` instances calling `.get/.post/...("url")`
* - `httpx.AsyncClient` instances calling `.get/.post/...("url")`, including
* aliased imports such as `import httpx as hx`,
* `from httpx import AsyncClient`, and
* `from httpx import AsyncClient as HttpxAsyncClient`.
* Locally rebound names (e.g. `AsyncClient = mock_factory()` inside a
* function) are excluded to avoid false-positive consumer contracts.
*/
const FASTAPI_VERBS: Record<string, string> = {
@@ -80,12 +85,50 @@ const REQUESTS_GENERIC_PATTERNS = compilePatterns({
} satisfies LanguagePatterns<Record<string, never>>);
// ─── Consumer: httpx.AsyncClient assignments ────────────────────────
// NOTE: This targeted detector only tracks explicit `httpx.AsyncClient(...)`
// construction. Direct imports (`from httpx import AsyncClient`) and module
// aliases (`import httpx as hx`) and annotated assignments (`client: httpx.AsyncClient = ...`)
// are intentionally left for a follow-up. Module-scope clients are only matched
// Module-scope clients are only matched
// at module scope; calls inside functions require a function/class-local tracked
// client to avoid false positives from same-name local variables.
const HTTPX_MODULE_IMPORT_PATTERNS = compilePatterns({
name: 'python-httpx-module-imports',
language: Python,
patterns: [
{
meta: {},
query: `
(import_statement
name: (aliased_import
name: (dotted_name (identifier) @module)
alias: (identifier) @alias))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
const HTTPX_ASYNC_CLIENT_IMPORT_PATTERNS = compilePatterns({
name: 'python-httpx-async-client-imports',
language: Python,
patterns: [
{
meta: {},
query: `
(import_from_statement
module_name: (dotted_name (identifier) @module)
name: (dotted_name (identifier) @client_class))
`,
},
{
meta: {},
query: `
(import_from_statement
module_name: (dotted_name (identifier) @module)
name: (aliased_import
name: (dotted_name (identifier) @client_class)
alias: (identifier) @alias))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
const HTTPX_ASYNC_CLIENT_ASSIGN_PATTERNS = compilePatterns({
name: 'python-httpx-async-client-assign',
language: Python,
@@ -97,8 +140,24 @@ const HTTPX_ASYNC_CLIENT_ASSIGN_PATTERNS = compilePatterns({
left: (_) @client
right: (call
function: (attribute
object: (identifier) @module (#eq? @module "httpx")
attribute: (identifier) @client_class (#eq? @client_class "AsyncClient"))))
object: (identifier) @module
attribute: (identifier) @client_class)))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
const HTTPX_ASYNC_CLIENT_DIRECT_ASSIGN_PATTERNS = compilePatterns({
name: 'python-httpx-async-client-direct-assign',
language: Python,
patterns: [
{
meta: {},
query: `
(assignment
left: (_) @client
right: (call
function: (identifier) @client_class))
`,
},
],
@@ -115,8 +174,24 @@ const HTTPX_ASYNC_CLIENT_WITH_ALIAS_PATTERNS = compilePatterns({
(as_pattern
(call
function: (attribute
object: (identifier) @module (#eq? @module "httpx")
attribute: (identifier) @client_class (#eq? @client_class "AsyncClient")))
object: (identifier) @module
attribute: (identifier) @client_class))
(as_pattern_target (identifier) @client))
`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
const HTTPX_ASYNC_CLIENT_DIRECT_WITH_ALIAS_PATTERNS = compilePatterns({
name: 'python-httpx-async-client-direct-with-alias',
language: Python,
patterns: [
{
meta: {},
query: `
(as_pattern
(call
function: (identifier) @client_class)
(as_pattern_target (identifier) @client))
`,
},
@@ -150,17 +225,137 @@ function trackedClientScopeKey(clientNode: Parser.SyntaxNode): string {
}
function callScopeKeys(clientNode: Parser.SyntaxNode): string[] {
const keys = new Set<string>();
const preferClass = clientNode.text.includes('.');
const nearestScope = getScopeKey(clientNode.parent, preferClass);
return [getScopeKey(clientNode.parent, clientNode.text.includes('.'))];
}
keys.add(nearestScope);
// Returns the scope key that a rebind of an imported alias would shadow under
// Python LEGB rules, or `null` when the rebind does not shadow anything that
// could produce a false-positive consumer detection.
// - Rebind inside a function/method → that function's scope.
// - Rebind at module top level → 'module' (shadows the whole file).
// - Rebind in a class body without an enclosing function → null. Python
// class attributes do not shadow bare-name lookups inside methods (methods
// see the module binding, not the class attribute), so we must not poison
// them.
function shadowScopeKey(node: Parser.SyntaxNode | null): string | null {
let current = node;
let passedThroughClass = false;
while (current) {
if (current.type === 'function_definition') {
// Reuse getScopeKey's key format so the two helpers cannot drift apart.
return getScopeKey(current);
}
if (current.type === 'class_definition') {
passedThroughClass = true;
}
current = current.parent;
}
return passedThroughClass ? null : 'module';
}
return [...keys];
function collectHttpxImportAliases(tree: Parser.Tree): {
moduleAliases: Set<string>;
asyncClientAliases: Set<string>;
} {
const moduleAliases = new Set<string>(['httpx']);
const asyncClientAliases = new Set<string>();
// The @module capture is a single identifier inside a `dotted_name`, so for
// `import package.httpx as hx` the pattern would match the inner `httpx`
// segment. Check the full `dotted_name` text via `parent` to anchor the match.
for (const match of runCompiledPatterns(HTTPX_MODULE_IMPORT_PATTERNS, tree)) {
const moduleNode = match.captures.module;
const aliasNode = match.captures.alias;
if (moduleNode?.parent?.text === 'httpx' && aliasNode) moduleAliases.add(aliasNode.text);
}
for (const match of runCompiledPatterns(HTTPX_ASYNC_CLIENT_IMPORT_PATTERNS, tree)) {
const moduleNode = match.captures.module;
const classNode = match.captures.client_class;
if (moduleNode?.parent?.text !== 'httpx' || classNode?.text !== 'AsyncClient') continue;
asyncClientAliases.add(match.captures.alias?.text ?? classNode.text);
}
return { moduleAliases, asyncClientAliases };
}
// Tracks local rebindings (`AsyncClient = ...`, `hx = ...`) that shadow an
// imported alias. We treat the whole enclosing scope (module, class, or
// function) as shadowed for that alias name, so subsequent constructions in
// that scope are not falsely detected as httpx consumers. Covers bare-identifier
// targets and the common tuple / list destructuring shapes.
const ALIAS_SHADOW_PATTERNS = compilePatterns({
name: 'python-httpx-alias-shadow',
language: Python,
patterns: [
{
meta: {},
query: `(assignment left: (identifier) @name)`,
},
{
meta: {},
query: `(assignment left: (pattern_list (identifier) @name))`,
},
{
meta: {},
query: `(assignment left: (tuple_pattern (identifier) @name))`,
},
{
meta: {},
query: `(assignment left: (list_pattern (identifier) @name))`,
},
],
} satisfies LanguagePatterns<Record<string, never>>);
function collectAliasShadowScopes(
tree: Parser.Tree,
aliases: Set<string>,
): Map<string, Set<string>> {
const shadowed = new Map<string, Set<string>>();
if (aliases.size === 0) return shadowed;
for (const match of runCompiledPatterns(ALIAS_SHADOW_PATTERNS, tree)) {
const nameNode = match.captures.name;
if (!nameNode || !aliases.has(nameNode.text)) continue;
const scopeKey = shadowScopeKey(nameNode.parent);
if (scopeKey === null) continue;
const set = shadowed.get(nameNode.text) ?? new Set<string>();
set.add(scopeKey);
shadowed.set(nameNode.text, set);
}
return shadowed;
}
function isAliasShadowed(
shadowed: Map<string, Set<string>>,
aliasName: string,
node: Parser.SyntaxNode,
): boolean {
const scopes = shadowed.get(aliasName);
if (!scopes || scopes.size === 0) return false;
let current: Parser.SyntaxNode | null = node.parent;
while (current) {
if (current.type === 'function_definition') {
// Reuse getScopeKey's key format so the two helpers cannot drift apart.
if (scopes.has(getScopeKey(current))) return true;
}
current = current.parent;
}
// A module-level rebind shadows the alias for the entire file.
return scopes.has('module');
}
function collectHttpxAsyncClients(tree: Parser.Tree): Map<string, Set<string>> {
const clients = new Map<string, Set<string>>();
const { moduleAliases, asyncClientAliases } = collectHttpxImportAliases(tree);
// Module aliases (`hx`) and AsyncClient aliases (`AsyncClient`,
// `HttpxAsyncClient`) share disjoint name spaces, so one shadow map keyed by
// alias name serves both lookups and we only walk the tree for rebinds once.
const shadowed = collectAliasShadowScopes(
tree,
new Set([...moduleAliases, ...asyncClientAliases]),
);
const addClient = (clientNode: Parser.SyntaxNode | undefined) => {
if (!clientNode) return;
@@ -172,10 +367,34 @@ function collectHttpxAsyncClients(tree: Parser.Tree): Map<string, Set<string>> {
};
for (const match of runCompiledPatterns(HTTPX_ASYNC_CLIENT_ASSIGN_PATTERNS, tree)) {
const moduleNode = match.captures.module;
const classNode = match.captures.client_class;
if (!moduleNode || !classNode) continue;
if (!moduleAliases.has(moduleNode.text) || classNode.text !== 'AsyncClient') continue;
if (isAliasShadowed(shadowed, moduleNode.text, moduleNode)) continue;
addClient(match.captures.client);
}
for (const match of runCompiledPatterns(HTTPX_ASYNC_CLIENT_DIRECT_ASSIGN_PATTERNS, tree)) {
const classNode = match.captures.client_class;
if (!classNode || !asyncClientAliases.has(classNode.text)) continue;
if (isAliasShadowed(shadowed, classNode.text, classNode)) continue;
addClient(match.captures.client);
}
for (const match of runCompiledPatterns(HTTPX_ASYNC_CLIENT_WITH_ALIAS_PATTERNS, tree)) {
const moduleNode = match.captures.module;
const classNode = match.captures.client_class;
if (!moduleNode || !classNode) continue;
if (!moduleAliases.has(moduleNode.text) || classNode.text !== 'AsyncClient') continue;
if (isAliasShadowed(shadowed, moduleNode.text, moduleNode)) continue;
addClient(match.captures.client);
}
for (const match of runCompiledPatterns(HTTPX_ASYNC_CLIENT_DIRECT_WITH_ALIAS_PATTERNS, tree)) {
const classNode = match.captures.client_class;
if (!classNode || !asyncClientAliases.has(classNode.text)) continue;
if (isAliasShadowed(shadowed, classNode.text, classNode)) continue;
addClient(match.captures.client);
}
@@ -18,12 +18,14 @@ import { getPluginForFile, HTTP_SCAN_GLOB, type HttpDetection } from './http-pat
* the preferred path because the graph has richer symbol metadata
* (real uids, class/method structure, etc.).
*
* 2. **Source-scan fallback (Strategy B)** — parse files directly with
* the per-language plugin registry in `./http-patterns/`. Used when
* the graph has no routes/fetches for this repo (e.g. a repo that
* hasn't been indexed yet, or whose indexer doesn't know the
* framework). Each plugin owns its tree-sitter grammar and query
* sources — this orchestrator imports NO grammars or query strings.
* 2. **Source-scan supplement (Strategy B)** — parse files directly with
* the per-language plugin registry in `./http-patterns/`. Used to
* fill gaps when graph extraction only covers part of a polyglot repo
* (e.g. Java graph routes plus Go source-scan routes). Graph entries
* remain authoritative for duplicate contract IDs because they carry
* richer symbol metadata. Each plugin owns its tree-sitter grammar
* and query sources — this orchestrator imports NO grammars or query
* strings.
*
* Adding a new language for Strategy B is a one-file edit in
* `http-patterns/index.ts`: register a new `HttpLanguagePlugin` and
@@ -194,17 +196,19 @@ export class HttpRouteExtractor implements ContractExtractor {
const graphProviders =
dbExecutor != null ? await this.extractProvidersGraph(dbExecutor, getDetections) : [];
const providers =
graphProviders.length > 0
? graphProviders
: this.extractProvidersSourceScan(await getScannedFiles(), getDetections);
// Source scan always runs to capture routes in languages/files not covered
// by graph edges; the glob and per-file parse results are cached above.
const providers = this.mergeGraphAndSourceContracts(
graphProviders,
this.extractProvidersSourceScan(await getScannedFiles(), getDetections),
);
const graphConsumers =
dbExecutor != null ? await this.extractConsumersGraph(dbExecutor, getDetections) : [];
const consumers =
graphConsumers.length > 0
? graphConsumers
: this.extractConsumersSourceScan(await getScannedFiles(), getDetections);
const consumers = this.mergeGraphAndSourceContracts(
graphConsumers,
this.extractConsumersSourceScan(await getScannedFiles(), getDetections),
);
return [...providers, ...consumers];
}
@@ -473,4 +477,18 @@ export class HttpRouteExtractor implements ContractExtractor {
}
return out;
}
private mergeGraphAndSourceContracts(
graphContracts: ExtractedContract[],
sourceContracts: ExtractedContract[],
): ExtractedContract[] {
const seenContractIds = new Set(graphContracts.map((c) => c.contractId));
const out = [...graphContracts];
for (const contract of sourceContracts) {
if (seenContractIds.has(contract.contractId)) continue;
seenContractIds.add(contract.contractId);
out.push(contract);
}
return out;
}
}
@@ -23,6 +23,21 @@ export interface FilePath {
}
const READ_CONCURRENCY = 32;
const ANALYZE_PROGRESS_ACTIVE_ENV = 'GITNEXUS_ANALYZE_PROGRESS_ACTIVE';
const warnLargeFileSkip = (message: string): void => {
if (process.env[ANALYZE_PROGRESS_ACTIVE_ENV] === '1') {
// analyze.ts routes console.warn through the progress bar logger while
// the bar is active. Emitting the operator-facing large-file notice there
// avoids raw pino NDJSON corrupting the one-line progress display in the
// heap-respawn child, whose stderr is intentionally piped for crash
// classification.
// eslint-disable-next-line no-console -- intentionally routed by analyze progress UI
console.warn(message);
return;
}
logger.warn(message);
};
/**
* Phase 1: Scan repository — stat files to get paths + sizes, no content loaded.
@@ -74,12 +89,35 @@ export const walkRepositoryPaths = async (
if (skippedLarge > 0) {
const isDefault = maxFileSizeBytes === DEFAULT_MAX_FILE_SIZE_BYTES;
const isOverrideUnset = !process.env.GITNEXUS_MAX_FILE_SIZE;
const suffix = isDefault ? ', likely generated/vendored' : '';
logger.warn(` Skipped ${skippedLarge} large files (>${maxFileSizeBytes / 1024}KB${suffix})`);
if (isVerboseIngestionEnabled()) {
for (const p of skippedLargePaths) {
logger.warn(` - ${p}`);
}
warnLargeFileSkip(
` Skipped ${skippedLarge} large files (>${maxFileSizeBytes / 1024}KB${suffix})`,
);
// Always show at least the first few paths so users can diagnose why
// edges are missing from a specific file (issue #1659). The full list is
// gated behind GITNEXUS_VERBOSE=1 to avoid flooding output on repos with
// many generated/vendored blobs. Sort before slicing so the preview is
// stable across runs (fs.stat callbacks race within each batch).
skippedLargePaths.sort();
const SKIPPED_PREVIEW_CAP = 5;
const showAll = isVerboseIngestionEnabled() || skippedLargePaths.length <= SKIPPED_PREVIEW_CAP;
const preview = showAll ? skippedLargePaths : skippedLargePaths.slice(0, SKIPPED_PREVIEW_CAP);
for (const p of preview) {
warnLargeFileSkip(` - ${p}`);
}
if (!showAll) {
const remaining = skippedLargePaths.length - SKIPPED_PREVIEW_CAP;
warnLargeFileSkip(` ...and ${remaining} more (set GITNEXUS_VERBOSE=1 to list them all)`);
}
// Only hint about the env var when the user has not set it at all. An
// explicit GITNEXUS_MAX_FILE_SIZE=512 happens to resolve to the same
// bytes as the default but the operator clearly already knows the knob.
if (isDefault && isOverrideUnset) {
warnLargeFileSkip(
` Set GITNEXUS_MAX_FILE_SIZE=<KB> to include files above the default cap.`,
);
}
}
@@ -210,6 +210,37 @@ interface LanguageProviderConfig {
ancestorNode: SyntaxNode,
) => { funcName: string; label: NodeLabel } | null;
// ── Template constraint extraction (SFINAE / `requires`) ────────────
/**
* Extract a per-language template-constraint payload for a templated
* function / method definition. Used by `parsing-processor` to
* disambiguate same-name same-arity overloads whose distinguishing
* signal is their template constraints rather than their parameter
* types — the canonical C++ SFINAE case (issue #1579):
*
* template<class T, std::enable_if_t<is_integral_v<T>, int> = 0>
* void process(T); // overload A
*
* template<class T, std::enable_if_t<is_floating_point_v<T>, int> = 0>
* void process(T); // overload B
*
* Both overloads' `parameterTypes` collapse to `['T']`, so without a
* constraint fingerprint in the graph node ID they merge into one
* Function node and the resolver only ever sees one candidate to
* narrow. The hook's return value is stamped onto the node's ID via
* `templateConstraintsIdTag()` AND stored on the node's
* `templateConstraints` property so `resolveDefGraphId` can look up
* the right overload by re-hashing the def's constraints at resolve
* time.
*
* Returns the opaque payload (any JSON-serializable shape — the
* producing adapter owns it; shared code MUST NOT inspect) or
* `undefined` when no constraints exist / the node isn't a templated
* function. Languages without SFINAE / concept semantics leave this
* undefined and the disambiguation is a pass-through.
*/
readonly extractTemplateConstraints?: (definitionNode: SyntaxNode) => unknown;
// ── Labels ────────────────────────────────────────────────────────
/** Override the default node label for definition.function captures.
* Return null to skip (C/C++ duplicate), a different label to reclassify
@@ -64,6 +64,7 @@ import {
cppImportOwningScope,
cppReceiverBinding,
} from './cpp/index.js';
import { extractCppTemplateConstraints } from './cpp/constraint-extractor.js';
const C_BUILT_INS: ReadonlySet<string> = new Set([
'printf',
@@ -463,6 +464,7 @@ export const cppProvider = defineLanguage({
heritageExtractor: createHeritageExtractor(SupportedLanguages.CPlusPlus),
labelOverride: cppLabelOverride,
builtInNames: C_BUILT_INS,
extractTemplateConstraints: extractCppTemplateConstraintsForProvider,
// ── RFC #909 Ring 3: scope-based resolution hooks (RFC §5) ──────────
emitScopeCaptures: emitCppScopeCaptures,
@@ -474,3 +476,46 @@ export const cppProvider = defineLanguage({
arityCompatibility: cppArityCompatibility,
// mergeBindings + resolveImportTarget live on ScopeResolver (see cpp/scope-resolver.ts).
});
/**
* LanguageProvider hook: walk from a function definition node up to its
* enclosing `template_declaration` and extract the SFINAE / `requires`-
* clause constraint payload. Used by `parsing-processor` to fingerprint
* the graph node ID so two SFINAE overloads with identical
* `parameterTypes` get distinct nodes (issue #1579).
*
* Returns `undefined` for non-templated functions and for templated
* functions whose constraints the extractor can't model — both cases
* result in no constraint suffix on the node ID.
*/
function extractCppTemplateConstraintsForProvider(definitionNode: SyntaxNode): unknown {
// Walk up to the enclosing template_declaration. Bound the walk so we
// can't accidentally land on a far-ancestor template_declaration that
// wraps an unrelated function.
let cur: SyntaxNode | null = definitionNode.parent;
let hops = 8;
let templateDecl: SyntaxNode | null = null;
while (cur !== null && hops-- > 0) {
if (cur.type === 'template_declaration') {
templateDecl = cur;
break;
}
if (cur.type === 'translation_unit') break;
cur = cur.parent;
}
if (templateDecl === null) return undefined;
// Find the function_declarator inside the function definition so the
// extractor can map template params to function-argument indices.
let declarator: SyntaxNode | null = definitionNode.childForFieldName('declarator');
let walk = 8;
while (declarator !== null && walk-- > 0) {
if (declarator.type === 'function_declarator') break;
if (declarator.type === 'pointer_declarator' || declarator.type === 'reference_declarator') {
declarator = declarator.childForFieldName('declarator');
continue;
}
break;
}
return extractCppTemplateConstraints(templateDecl, declarator);
}
@@ -1,9 +1,11 @@
import type { SyntaxNode } from '../../utils/ast-helpers.js';
import type { ParameterTypeClass } from 'gitnexus-shared';
export interface CppArityInfo {
parameterCount?: number;
requiredParameterCount?: number;
parameterTypes?: string[];
parameterTypeClasses?: ParameterTypeClass[];
}
/**
@@ -73,26 +75,35 @@ export function computeCppDeclarationArity(node: SyntaxNode): CppArityInfo {
const totalNonVariadic = requiredCount + optionalCount;
const types: string[] = [];
const typeClasses: ParameterTypeClass[] = [];
for (const p of params) {
if (p.type === 'variadic_parameter') {
types.push('...');
typeClasses.push(unknownTypeClass('...'));
} else if (p.type === 'variadic_parameter_declaration') {
// Parameter pack: treated as variadic
types.push('...');
typeClasses.push(unknownTypeClass('...'));
} else {
const typeNode = p.childForFieldName('type');
types.push(normalizeCppParamType(typeNode?.text ?? 'unknown'));
const rawType = typeNode?.text ?? 'unknown';
types.push(normalizeCppParamType(rawType));
typeClasses.push(
classifyCppParameterType(rawType, p.childForFieldName('declarator')?.text, p.text),
);
}
}
// Append '...' for C-style variadic if not already in types
if (hasEllipsis && !types.includes('...')) {
types.push('...');
typeClasses.push(unknownTypeClass('...'));
}
return {
parameterCount: isVariadic ? undefined : totalNonVariadic,
requiredParameterCount: requiredCount,
parameterTypes: types,
parameterTypeClasses: typeClasses,
};
}
@@ -120,8 +131,14 @@ export function computeCppCallArity(node: SyntaxNode): number {
* so that `narrowOverloadCandidates` can match against literal-inferred
* argument types (e.g. `inferCppLiteralType` returns `'string'` for
* string literals, not `'std::string'`).
*
* This intentionally remains coarse and graph-ID-stable: cv-qualifiers,
* reference markers, and pointer markers are stripped here. C++ callers
* that need those distinctions should read `parameterTypeClasses`, which
* is an additive sidecar and does not participate in overload node ID
* hashing.
*/
function normalizeCppParamType(raw: string): string {
export function normalizeCppParamType(raw: string): string {
let t = raw.trim();
// Strip const, volatile, etc.
t = t.replace(/\b(const|volatile|restrict|mutable|constexpr)\b/g, '').trim();
@@ -158,6 +175,52 @@ function normalizeCppParamType(raw: string): string {
return STD_MAP[t] ?? t;
}
export function classifyCppParameterType(
rawType: string,
declaratorText?: string,
fullParameterText?: string,
): ParameterTypeClass {
const source = fullParameterText ?? `${rawType} ${declaratorText ?? ''}`.trim();
if (rawType === 'unknown') return unknownTypeClass('unknown');
const hasConst = /\bconst\b/.test(source);
const hasVolatile = /\bvolatile\b/.test(source);
const cv: ParameterTypeClass['cv'] =
hasConst && hasVolatile
? 'const volatile'
: hasConst
? 'const'
: hasVolatile
? 'volatile'
: 'none';
const pointerDepth = (source.match(/\*/g) ?? []).length;
const indirection: ParameterTypeClass['indirection'] =
pointerDepth > 0
? 'pointer'
: /&&/.test(source)
? 'rvalue-ref'
: /&/.test(source)
? 'lvalue-ref'
: 'value';
return {
base: normalizeCppParamType(rawType),
cv,
indirection,
pointerDepth,
};
}
function unknownTypeClass(base: string): ParameterTypeClass {
return {
base,
cv: 'unknown',
indirection: 'unknown',
pointerDepth: 0,
};
}
function findFuncDeclarator(node: SyntaxNode): SyntaxNode | null {
let decl = node.childForFieldName('declarator');
if (decl === null) {
@@ -8,7 +8,10 @@ import type { Callsite, SymbolDefinition } from 'gitnexus-shared';
* - Default parameters (requiredParameterCount < parameterCount)
* - Variadic functions (C-style `...`)
* - Parameter packs (V1: treated as variadic)
* - Templates (V1: generic-ignored, arity check on non-template params)
* - Templates: arity check on non-template params; SFINAE / `requires`
* constraints are filtered separately via `constraintCompatibility`
* (see `constraint-filter.ts` and issue #1579). Type-argument generic
* substitution (`List<T>` ≡ `List<U>`) remains out of V1 scope.
*
* Verdict:
* - 'compatible': callsite.arity fits within [required, total] range
@@ -1,4 +1,4 @@
import type { Capture, CaptureMatch } from 'gitnexus-shared';
import type { Capture, CaptureMatch, ParameterTypeClass } from 'gitnexus-shared';
import {
findNodeAtRange,
nodeToCapture,
@@ -9,11 +9,16 @@ import { getCppParser, getCppScopeQuery } from './query.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
import { splitCppInclude, splitCppUsingDecl } from './import-decomposer.js';
import { computeCppDeclarationArity, computeCppCallArity } from './arity-metadata.js';
import {
classifyCppParameterType,
computeCppDeclarationArity,
computeCppCallArity,
} from './arity-metadata.js';
import { markCppAnonymousNamespaceRange, markFileLocal } from './file-local-linkage.js';
import { markCppDependentBase } from './two-phase-lookup.js';
import { markCppAdlSiteArgs, markCppAdlSiteNoAdl, type CppAdlArgInfo } from './adl.js';
import { markCppInlineNamespaceRange } from './inline-namespaces.js';
import { extractCppTemplateConstraints } from './constraint-extractor.js';
export function emitCppScopeCaptures(
sourceText: string,
@@ -114,6 +119,13 @@ export function emitCppScopeCaptures(
JSON.stringify(arity.parameterTypes),
);
}
if (arity.parameterTypeClasses !== undefined) {
grouped['@declaration.parameter-type-classes'] = syntheticCapture(
'@declaration.parameter-type-classes',
fnNode,
JSON.stringify(arity.parameterTypeClasses),
);
}
// Detect static storage class (file-local linkage)
if (hasStaticStorageClass(fnNode)) {
@@ -130,6 +142,24 @@ export function emitCppScopeCaptures(
markFileLocal(filePath, nameText);
}
}
// SFINAE / `requires`-clause aware constraints for overload
// narrowing (issue #1579). Walk from the enclosing
// `template_declaration` — not the inner `function_definition` —
// so inline method templates (`template<...> class C { template<...> void f(); }`)
// pick up the correct outer constraint scope.
const templateDecl = findEnclosingTemplateDeclaration(fnNode);
if (templateDecl !== null) {
const funcDeclarator = findFunctionDeclarator(fnNode);
const constraints = extractCppTemplateConstraints(templateDecl, funcDeclarator);
if (constraints !== undefined) {
grouped['@declaration.template-constraints'] = syntheticCapture(
'@declaration.template-constraints',
fnNode,
JSON.stringify(constraints),
);
}
}
}
}
@@ -191,6 +221,14 @@ export function emitCppScopeCaptures(
JSON.stringify(argTypes),
);
}
const argTypeClasses = inferCppCallArgTypeClasses(cNode);
if (argTypeClasses !== undefined && argTypeClasses.length > 0) {
grouped['@reference.parameter-type-classes'] = syntheticCapture(
'@reference.parameter-type-classes',
cNode,
JSON.stringify(argTypeClasses),
);
}
}
}
@@ -552,6 +590,52 @@ function extractBaseLookupName(baseNode: SyntaxNode): string {
return '';
}
/**
* Walk parent chain from a function_definition / declaration / field_declaration
* to find the enclosing `template_declaration`. Returns null when the function
* isn't templated. The walk only ascends through wrapper nodes the C++
* grammar inserts between `template_declaration` and the function — direct
* parent in the common case, two hops for member templates whose outer
* class is also templated (we return the INNERMOST template_declaration,
* which carries this function's own template parameters).
*/
function findEnclosingTemplateDeclaration(fnNode: SyntaxNode): SyntaxNode | null {
let cur: SyntaxNode | null = fnNode.parent;
// Cap the walk — `template_declaration` is typically the immediate parent
// or one wrapper away. Anything deeper is an inline-method-in-template
// shape and we still want the innermost templates_declaration whose body
// wraps `fnNode`.
let hops = 8;
while (cur !== null && hops-- > 0) {
if (cur.type === 'template_declaration') return cur;
// Don't ascend past structural boundaries that should reset template scope.
if (cur.type === 'translation_unit') return null;
cur = cur.parent;
}
return null;
}
/**
* Locate the `function_declarator` AST node within a function definition
* or declaration. Unwraps pointer/reference declarator wrappers. Returns
* null when no function_declarator is found (e.g. variable declaration
* mis-classified upstream).
*/
function findFunctionDeclarator(fnNode: SyntaxNode): SyntaxNode | null {
const direct = fnNode.childForFieldName('declarator');
let cur: SyntaxNode | null = direct;
let hops = 8;
while (cur !== null && hops-- > 0) {
if (cur.type === 'function_declarator') return cur;
if (cur.type === 'pointer_declarator' || cur.type === 'reference_declarator') {
cur = cur.childForFieldName('declarator');
continue;
}
break;
}
return findFirstDescendantOfType(fnNode, 'function_declarator');
}
/** Find the first direct child matching one of the given types. */
function findChildOfType(node: SyntaxNode, types: readonly string[]): SyntaxNode | null {
for (let i = 0; i < node.childCount; i++) {
@@ -611,6 +695,35 @@ function inferCppCallArgTypes(node: SyntaxNode): string[] | undefined {
return types.length > 0 ? types : undefined;
}
function inferCppCallArgTypeClasses(node: SyntaxNode): ParameterTypeClass[] | undefined {
const argList = node.childForFieldName('arguments');
if (argList === null) return undefined;
const classes: ParameterTypeClass[] = [];
for (let i = 0; i < argList.childCount; i++) {
const child = argList.child(i);
if (child === null) continue;
if (child.type === ',' || child.type === '(' || child.type === ')') continue;
const litType = inferCppLiteralType(child);
if (litType !== '') {
classes.push(valueTypeClass(litType));
} else if (child.type === 'identifier') {
classes.push(lookupDeclaredTypeClassForIdentifier(child));
} else {
classes.push(unknownTypeClass('unknown'));
}
}
return classes.length > 0 ? classes : undefined;
}
function valueTypeClass(base: string): ParameterTypeClass {
return { base, cv: 'none', indirection: 'value', pointerDepth: 0 };
}
function unknownTypeClass(base: string): ParameterTypeClass {
return { base, cv: 'unknown', indirection: 'unknown', pointerDepth: 0 };
}
/**
* Infer the canonical type name of a C++ literal AST node.
* Returns empty string for non-literal / unknown nodes.
@@ -655,6 +768,15 @@ function inferCppLiteralType(node: SyntaxNode): string {
* - `int n = ...` → 'int'
* - `const int n = ...` → 'int'
* Returns empty string if no declaration found or type is auto/placeholder.
*
* Limitation: only `declaration` siblings inside the enclosing
* `compound_statement` are inspected. Function parameters live in the
* `function_declarator`'s `parameter_list` and are NOT resolved here, so
* `void run(int n) { process(n); }`
* infers `''` for `n` and the constraint filter falls through to
* `'unknown'` → ambiguity suppression → 0 CALLS edges. This is a
* "degrade not lie" gap (no wrong edges, just missing ones); extending
* the scan to `parameter_list` is tracked under #1579 as a follow-up.
*/
function lookupDeclaredTypeForIdentifier(identNode: SyntaxNode): string {
const varName = identNode.text;
@@ -669,6 +791,9 @@ function lookupDeclaredTypeForIdentifier(identNode: SyntaxNode): string {
}
if (scope === null) return '';
const paramType = lookupFunctionParameterType(scope, varName);
if (paramType !== '') return paramType;
// Scan declarations in the scope for a matching variable name
for (let i = 0; i < scope.childCount; i++) {
const stmt = scope.child(i);
@@ -682,18 +807,118 @@ function lookupDeclaredTypeForIdentifier(identNode: SyntaxNode): string {
// Check init_declarator children for the variable name
const declarator = stmt.childForFieldName('declarator');
if (declarator === null) continue;
if (declarator.type === 'init_declarator') {
const nameChild = declarator.childForFieldName('declarator');
if (nameChild !== null && nameChild.text === varName) {
return normalizeCppTypeText(typeNode.text);
}
} else if (declarator.text === varName) {
const nameChild = declaredNameNode(declarator);
if (nameChild !== null && extractDeclaratorLeafName(nameChild) === varName) {
return normalizeCppTypeText(typeNode.text);
}
}
return '';
}
function lookupDeclaredTypeClassForIdentifier(identNode: SyntaxNode): ParameterTypeClass {
const varName = identNode.text;
let scope: SyntaxNode | null = identNode.parent;
while (
scope !== null &&
scope.type !== 'compound_statement' &&
scope.type !== 'translation_unit'
) {
scope = scope.parent;
}
if (scope === null) return unknownTypeClass('unknown');
const paramTypeClass = lookupFunctionParameterTypeClass(scope, varName, identNode);
if (paramTypeClass !== undefined) return paramTypeClass;
for (let i = 0; i < scope.childCount; i++) {
const stmt = scope.child(i);
if (stmt === null || stmt.type !== 'declaration') continue;
const typeNode = stmt.childForFieldName('type');
if (typeNode === null) continue;
if (typeNode.type === 'placeholder_type_specifier') continue;
const declarator = stmt.childForFieldName('declarator');
if (declarator === null) continue;
const nameChild = declaredNameNode(declarator);
if (nameChild === null || extractDeclaratorLeafName(nameChild) !== varName) continue;
const typeClass = classifyCppParameterType(
typeNode.text,
nameChild.text,
stmt.text.replace(/;\s*$/, ''),
);
if (isKnownEnumName(identNode, typeClass.base)) {
return { ...typeClass, base: `enum:${typeClass.base}` };
}
return typeClass;
}
return unknownTypeClass('unknown');
}
function lookupFunctionParameterType(scope: SyntaxNode, varName: string): string {
const param = findEnclosingFunctionParameter(scope, varName);
if (param === null) return '';
const typeNode = param.childForFieldName('type');
if (typeNode === null) return '';
return normalizeCppTypeText(typeNode.text);
}
function lookupFunctionParameterTypeClass(
scope: SyntaxNode,
varName: string,
identNode: SyntaxNode,
): ParameterTypeClass | undefined {
const param = findEnclosingFunctionParameter(scope, varName);
if (param === null) return undefined;
const typeNode = param.childForFieldName('type');
if (typeNode === null) return undefined;
const declarator = param.childForFieldName('declarator');
if (declarator === null) return undefined;
const typeClass = classifyCppParameterType(typeNode.text, declarator.text, param.text);
if (isKnownEnumName(identNode, typeClass.base)) {
return { ...typeClass, base: `enum:${typeClass.base}` };
}
return typeClass;
}
function findEnclosingFunctionParameter(scope: SyntaxNode, varName: string): SyntaxNode | null {
let node: SyntaxNode | null = scope.parent;
while (node !== null) {
if (node.type === 'function_definition' || node.type === 'function_declarator') {
const fnDecl =
node.type === 'function_declarator'
? node
: findFirstDescendantOfType(node, 'function_declarator');
const params = fnDecl?.childForFieldName('parameters') ?? null;
if (params !== null) {
for (let i = 0; i < params.namedChildCount; i++) {
const param = params.namedChild(i);
if (param === null || param.type !== 'parameter_declaration') continue;
const declarator = param.childForFieldName('declarator');
if (declarator !== null && extractDeclaratorLeafName(declarator) === varName) {
return param;
}
}
}
return null;
}
node = node.parent;
}
return null;
}
function declaredNameNode(declarator: SyntaxNode): SyntaxNode | null {
if (declarator.type !== 'init_declarator') return declarator;
for (let i = 0; i < declarator.namedChildCount; i++) {
const child = declarator.namedChild(i);
if (child === null) continue;
if (child.type === 'identifier') return child;
if (child.type.endsWith('_declarator')) return child;
}
return declarator.childForFieldName('declarator');
}
/** Normalize a type-specifier text for argument type matching.
* Strips qualifiers (const, volatile), namespace prefixes (std::),
* and pointer/reference markers. */
@@ -705,6 +930,25 @@ function normalizeCppTypeText(text: string): string {
return t;
}
function isKnownEnumName(node: SyntaxNode, typeName: string): boolean {
if (typeName === '' || typeName === 'unknown') return false;
let root: SyntaxNode = node;
while (root.parent !== null) root = root.parent;
const stack: SyntaxNode[] = [root];
while (stack.length > 0) {
const cur = stack.pop()!;
if (cur.type === 'enum_specifier') {
const name = cur.childForFieldName('name');
if (name?.text === typeName) return true;
}
for (let i = 0; i < cur.childCount; i++) {
const child = cur.child(i);
if (child !== null) stack.push(child);
}
}
return false;
}
/**
* Detect whether a `namespace_definition` AST node is inline.
* Tree-sitter-cpp exposes the `inline` keyword as an anonymous child
@@ -1166,7 +1410,9 @@ function extractDeclaratorLeafName(node: SyntaxNode): string | null {
const next =
cur.childForFieldName('declarator') ??
// parenthesized_declarator: single named child
(cur.type === 'parenthesized_declarator' ? cur.namedChild(0) : null);
(cur.type === 'parenthesized_declarator' || cur.type.endsWith('_declarator')
? cur.namedChild(0)
: null);
if (next === null) return null;
cur = next;
}
@@ -0,0 +1,335 @@
/**
* Extract C++ template constraint expressions for SFINAE-aware overload
* narrowing (issue #1579). Recognizes 3 AST shapes:
*
* F1 — unqualified non-type template param default:
* `template<class T, enable_if_t<P, int> = 0> void f(T);`
* F2 — `std::`-qualified variant (canonical ticket form):
* `template<class T, std::enable_if_t<P, int> = 0> void f(T);`
* F4 — C++20 leading requires-clause:
* `template<class T> requires P void f(T);`
*
* Deferred (return `{kind:'unknown'}`):
* F3 — void-default `typename = enable_if_t<P>` (cppref labels this
* `/* WRONG *\/` because adjacent overloads collapse to redeclarations)
* F5 — trailing requires (`void f(T) requires P;`)
* `requires_expression` blocks (`requires { typename T::U; }`)
* `decltype(...)`, fold-expressions, user-defined `_v` aliases.
*
* The output payload is opaque to shared code — only
* `constraint-filter.ts` consumes it. See ISO `[temp.constr.normal]` /
* `<https://en.cppreference.com/w/cpp/language/constraints>` for the
* normalization the Kleene 3-valued evaluator implements.
*/
import type { SyntaxNode } from '../../utils/ast-helpers.js';
export type ConstraintExpr =
| { readonly kind: 'atomic'; readonly name: string; readonly args: readonly string[] }
| { readonly kind: 'and'; readonly children: readonly ConstraintExpr[] }
| { readonly kind: 'or'; readonly children: readonly ConstraintExpr[] }
| { readonly kind: 'not'; readonly child: ConstraintExpr }
| { readonly kind: 'unknown' };
export interface CppConstraintPayload {
/** Ordered template parameter names (type-params only — non-type defaults
* carrying enable_if predicates are folded into `expr`). */
readonly templateParams: readonly string[];
/**
* Mapping from each template parameter name to the call-site argument
* index where its deduced type lives. Computed by scanning the function's
* parameter list for the first parameter whose type is the bare template
* parameter name (or template-typed by it). Missing entries → 'unknown'
* verdict at evaluation time.
*/
readonly paramArgIndex: { readonly [paramName: string]: number };
/** Root constraint expression. When multiple constraints (multiple
* enable_if defaults, requires clause, etc.) are present they are
* implicitly conjoined under a top-level `and` node. */
readonly expr: ConstraintExpr;
}
/**
* Walk a `template_declaration` AST node and extract its constraint
* payload. Caller is responsible for passing the OUTER `template_declaration`
* — for class-member template functions, that means the enclosing
* template_declaration of the class OR of the method, whichever
* directly precedes the function definition.
*
* Returns `undefined` when the template_declaration declares no
* constraints worth tracking (no enable_if default, no requires clause).
* Returns a payload whose `expr.kind === 'unknown'` when constraints are
* present but the extractor cannot model them — monotonicity guarantees
* the filter keeps the candidate in that case.
*/
export function extractCppTemplateConstraints(
templateDecl: SyntaxNode,
funcDeclarator: SyntaxNode | null,
): CppConstraintPayload | undefined {
const paramList = childOfType(templateDecl, 'template_parameter_list');
if (paramList === null) return undefined;
const templateParams: string[] = [];
const exprs: ConstraintExpr[] = [];
for (let i = 0; i < paramList.namedChildCount; i++) {
const param = paramList.namedChild(i);
if (param === null) continue;
if (
param.type === 'type_parameter_declaration' ||
param.type === 'optional_type_parameter_declaration' ||
param.type === 'variadic_type_parameter_declaration'
) {
const id = firstDescendantOfType(param, 'type_identifier');
if (id !== null) templateParams.push(id.text);
continue;
}
// Non-type parameter — F1 / F2 default-value carries the enable_if
// predicate. Shape: `optional_parameter_declaration` with field
// `default_value`, whose value is a `template_type` named
// `enable_if_t` (F1) or a qualified version (F2).
if (param.type === 'optional_parameter_declaration') {
const defaultVal = param.childForFieldName('default_value');
const typeNode = param.childForFieldName('type');
const candidate = extractEnableIfPredicate(typeNode);
if (candidate !== undefined) {
exprs.push(candidate);
} else if (defaultVal !== null) {
// Default-value-as-predicate not yet supported. Bail conservatively.
exprs.push({ kind: 'unknown' });
}
}
}
// F4 — C++20 leading `requires` clause. Tree-sitter-cpp exposes it as a
// `requires_clause` child of `template_declaration` (sibling of the
// template_parameter_list).
const requiresClause = childOfType(templateDecl, 'requires_clause');
if (requiresClause !== null) {
const parsed = parseRequiresClause(requiresClause);
if (parsed !== undefined) exprs.push(parsed);
}
if (templateParams.length === 0 && exprs.length === 0) return undefined;
const paramArgIndex = buildParamArgIndex(templateParams, funcDeclarator);
const expr: ConstraintExpr =
exprs.length === 0
? { kind: 'unknown' }
: exprs.length === 1
? exprs[0]
: { kind: 'and', children: exprs };
return { templateParams, paramArgIndex, expr };
}
/**
* Inspect a non-type template parameter's declared type to see whether
* it's `enable_if_t<P, T>` (F1) or `std::enable_if_t<P, T>` (F2). When
* matched, extract the predicate `P` and return it as a `ConstraintExpr`.
*
* Returns undefined when the parameter's type is not enable_if (so the
* caller can decide whether to bail or ignore).
*/
function extractEnableIfPredicate(typeNode: SyntaxNode | null): ConstraintExpr | undefined {
if (typeNode === null) return undefined;
// Unwrap a type_descriptor wrapper (when present).
let t: SyntaxNode | null = typeNode;
if (t.type === 'type_descriptor') {
t = t.childForFieldName('type') ?? firstDescendantOfType(t, 'template_type');
}
// F2 shape: tree-sitter-cpp models `std::enable_if_t<...>` as
// `qualified_identifier` whose `name` field is the `template_type`.
// F1 shape (unqualified `enable_if_t<...>`) is `template_type` directly.
if (t !== null && t.type === 'qualified_identifier') {
const inner = t.childForFieldName('name') ?? firstDescendantOfType(t, 'template_type');
if (inner !== null && inner.type === 'template_type') {
t = inner;
}
}
if (t === null || t.type !== 'template_type') return undefined;
const nameNode = t.childForFieldName('name');
if (nameNode === null) return undefined;
const tail = stripQualifiedPrefix(nameNode.text);
if (tail !== 'enable_if_t' && tail !== 'enable_if') return undefined;
// Predicate is the first template argument of enable_if_t.
const argList = t.childForFieldName('arguments') ?? childOfType(t, 'template_argument_list');
if (argList === null) return { kind: 'unknown' };
for (let i = 0; i < argList.namedChildCount; i++) {
const arg = argList.namedChild(i);
if (arg === null) continue;
if (arg.type !== 'type_descriptor') continue;
const inner = arg.childForFieldName('type') ?? arg.namedChild(0);
if (inner === null) continue;
return parseAtomicOrBoolean(inner);
}
return { kind: 'unknown' };
}
/** Parse a requires-clause body. The body is a binary or unary expression
* over atomic predicates (variable templates like `is_integral_v<T>`). */
function parseRequiresClause(requiresClause: SyntaxNode): ConstraintExpr | undefined {
// tree-sitter-cpp exposes the expression as a named child or via a
// `constraint` field. Probe both.
let expr: SyntaxNode | null = requiresClause.childForFieldName('constraint');
if (expr === null) {
for (let i = 0; i < requiresClause.namedChildCount; i++) {
const c = requiresClause.namedChild(i);
if (c === null) continue;
// Skip the `requires` keyword token.
if (c.type === 'requires') continue;
expr = c;
break;
}
}
if (expr === null) return undefined;
return parseAtomicOrBoolean(expr);
}
/**
* Recursively parse a constraint sub-expression. Recognizes:
* - `template_type` / `template_function` named `<predicate>_v` → atomic
* - binary_expression with `&&` / `||` → conjunction / disjunction
* - unary_expression with `!` → negation
* - parenthesized_expression → unwrap
* - anything else → `{kind:'unknown'}` (monotonicity-safe)
*
* `requires_expression` blocks intentionally fall through to 'unknown'
* — they need substitution semantics we don't model in V1.
*/
function parseAtomicOrBoolean(node: SyntaxNode): ConstraintExpr {
// Unwrap parentheses.
if (node.type === 'parenthesized_expression') {
const inner = node.namedChild(0);
return inner === null ? { kind: 'unknown' } : parseAtomicOrBoolean(inner);
}
// Boolean composition.
if (node.type === 'binary_expression') {
const left = node.childForFieldName('left');
const right = node.childForFieldName('right');
const opNode = node.childForFieldName('operator');
if (left !== null && right !== null && opNode !== null) {
const op = opNode.text;
const l = parseAtomicOrBoolean(left);
const r = parseAtomicOrBoolean(right);
if (op === '&&') return { kind: 'and', children: [l, r] };
if (op === '||') return { kind: 'or', children: [l, r] };
}
return { kind: 'unknown' };
}
if (node.type === 'unary_expression') {
const opNode = node.childForFieldName('operator') ?? node.namedChild(0);
const arg = node.childForFieldName('argument') ?? node.namedChild(1) ?? node.namedChild(0);
if (opNode !== null && opNode.text === '!' && arg !== null && arg !== opNode) {
return { kind: 'not', child: parseAtomicOrBoolean(arg) };
}
return { kind: 'unknown' };
}
// Atomic predicate — `template_type` is the typical shape for variable
// templates like `is_integral_v<T>`. Some grammar variants surface it as
// `template_function` or via a `qualified_identifier` wrapper.
if (node.type === 'template_type' || node.type === 'template_function') {
return parseAtomicTemplate(node);
}
if (node.type === 'qualified_identifier') {
// `std::is_integral_v<T>` shape (without template_type wrapping).
const inner = node.childForFieldName('name');
if (inner !== null && (inner.type === 'template_type' || inner.type === 'template_function')) {
return parseAtomicTemplate(inner);
}
return { kind: 'unknown' };
}
// `requires { typename T::U; }` blocks and decltype: out of V1 scope.
return { kind: 'unknown' };
}
function parseAtomicTemplate(t: SyntaxNode): ConstraintExpr {
const nameNode = t.childForFieldName('name');
if (nameNode === null) return { kind: 'unknown' };
const name = stripQualifiedPrefix(nameNode.text);
const argList = t.childForFieldName('arguments') ?? childOfType(t, 'template_argument_list');
const args: string[] = [];
if (argList !== null) {
for (let i = 0; i < argList.namedChildCount; i++) {
const arg = argList.namedChild(i);
if (arg === null) continue;
if (arg.type !== 'type_descriptor') continue;
const inner = arg.childForFieldName('type') ?? arg.namedChild(0);
if (inner === null) continue;
// For Tier-A predicates the args are bare template-parameter names
// (`T`, `U`). Anything more elaborate is bailed via 'unknown' at the
// top level if needed; here we just record the textual identifier.
const id =
inner.type === 'type_identifier' ? inner : firstDescendantOfType(inner, 'type_identifier');
args.push(id !== null ? id.text : inner.text);
}
}
return { kind: 'atomic', name, args };
}
/** Build a `paramName → call-site argument index` map by scanning the
* function's parameter list for parameters typed by each template param. */
function buildParamArgIndex(
templateParams: readonly string[],
funcDeclarator: SyntaxNode | null,
): { [paramName: string]: number } {
const out: { [paramName: string]: number } = {};
if (funcDeclarator === null || templateParams.length === 0) return out;
const paramList = funcDeclarator.childForFieldName('parameters');
if (paramList === null) return out;
let argIdx = 0;
for (let i = 0; i < paramList.childCount; i++) {
const p = paramList.child(i);
if (p === null) continue;
if (
p.type !== 'parameter_declaration' &&
p.type !== 'optional_parameter_declaration' &&
p.type !== 'variadic_parameter_declaration'
) {
continue;
}
const typeNode = p.childForFieldName('type');
if (typeNode !== null) {
const tname = bareTypeIdentifier(typeNode);
if (tname !== null && templateParams.includes(tname) && !(tname in out)) {
out[tname] = argIdx;
}
}
argIdx++;
}
return out;
}
function bareTypeIdentifier(typeNode: SyntaxNode): string | null {
if (typeNode.type === 'type_identifier') return typeNode.text;
// Allow `T const`, `T&`, `T*` shapes — the inner type_identifier still wins.
const id = firstDescendantOfType(typeNode, 'type_identifier');
return id !== null ? id.text : null;
}
function stripQualifiedPrefix(text: string): string {
const idx = text.lastIndexOf('::');
return idx >= 0 ? text.slice(idx + 2) : text;
}
function childOfType(node: SyntaxNode, type: string): SyntaxNode | null {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c !== null && c.type === type) return c;
}
return null;
}
function firstDescendantOfType(node: SyntaxNode, type: string): SyntaxNode | null {
if (node.type === type) return node;
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c === null) continue;
const hit = firstDescendantOfType(c, type);
if (hit !== null) return hit;
}
return null;
}
@@ -0,0 +1,261 @@
/**
* Kleene 3-valued evaluator + curated 4-predicate registry +
* `cppConstraintCompatibility` hook export for SFINAE / `requires`-clause
* filtering (issue #1579).
*
* Semantics:
* - `'incompatible'` → predicate provably fails for these argumentTypes
* (ISO `[temp.constr.atomic]` "not satisfied")
* - `'compatible'` → predicate provably holds
* - `'unknown'` → cannot decide (missing arg-type info, predicate
* not in registry, AST shape bailed during extraction). The shared
* filter keeps the candidate on `'unknown'` — monotonicity guarantee.
*
* Kleene rules (extension of ISO's 2-valued short-circuit conjunction in
* `<https://en.cppreference.com/w/cpp/language/constraints>`):
* AND: incompatible if any child incompatible; compatible iff all
* children compatible; otherwise unknown.
* OR: compatible if any child compatible; incompatible iff all
* children incompatible; otherwise unknown.
* NOT: flip compatible↔incompatible; pass through unknown.
*/
import type {
ArityVerdict,
Callsite,
ConstraintContext,
ParameterTypeClass,
SymbolDefinition,
} from 'gitnexus-shared';
import { classifyType, type TypeClass } from './type-classifier.js';
import type { ConstraintExpr, CppConstraintPayload } from './constraint-extractor.js';
interface ConstraintArgClass {
readonly typeClass: TypeClass;
readonly shape?: ParameterTypeClass;
}
type AtomicEvaluator = (args: readonly ConstraintArgClass[]) => ArityVerdict;
/**
* Curated Tier-A predicate registry. Predicates that depend on pointer,
* reference, or cv shape consult `ConstraintContext.argumentTypeClasses`.
* Missing or unsupported shape returns 'unknown' to preserve monotonicity.
*/
// ISO `<type_traits>` treats `bool`, `char`, and the signed/unsigned char
// variants as integral types (§21.3.4 Table 48), so `is_integral_v<bool>`
// and `is_integral_v<char>` must both yield `true`. We keep the `TypeClass`
// enum precise (separate `'bool'` / `'char'` buckets) so that
// `is_same_v<bool, int>` still resolves to `'incompatible'`; the integral-
// family widening lives here in the predicate evaluators instead.
function isIntegralClass(c: TypeClass | undefined): boolean {
return c === 'integral' || c === 'bool' || c === 'char';
}
const REGISTRY = new Map<string, AtomicEvaluator>([
[
'is_void_v',
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && arg.typeClass === 'void'),
],
[
'is_integral_v',
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && isIntegralClass(arg.typeClass)),
],
[
'is_floating_point_v',
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && arg.typeClass === 'floating'),
],
[
'is_arithmetic_v',
(args) =>
unaryVerdict(
args,
(arg) =>
isPlainValue(arg) && (isIntegralClass(arg.typeClass) || arg.typeClass === 'floating'),
),
],
[
'is_enum_v',
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && arg.typeClass === 'enum'),
],
[
'is_class_v',
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && arg.typeClass === 'class'),
],
[
'is_pointer_v',
(args) =>
unaryShapeVerdict(args, (shape) => shape.indirection === 'pointer' && shape.pointerDepth > 0),
],
[
'is_reference_v',
(args) =>
unaryShapeVerdict(
args,
(shape) => shape.indirection === 'lvalue-ref' || shape.indirection === 'rvalue-ref',
),
],
[
'is_const_v',
(args) =>
unaryShapeVerdict(args, (shape) => shape.cv === 'const' || shape.cv === 'const volatile', {
requireTopLevelCv: true,
}),
],
[
'is_volatile_v',
(args) =>
unaryShapeVerdict(args, (shape) => shape.cv === 'volatile' || shape.cv === 'const volatile', {
requireTopLevelCv: true,
}),
],
[
'is_same_v',
(args) => {
if (args.length < 2 || args[0].typeClass === 'unknown' || args[1].typeClass === 'unknown') {
return 'unknown';
}
return args[0].typeClass === args[1].typeClass ? 'compatible' : 'incompatible';
},
],
]);
function unaryVerdict(
args: readonly ConstraintArgClass[],
predicate: (arg: ConstraintArgClass) => boolean,
): ArityVerdict {
const arg = args[0];
if (arg === undefined || arg.typeClass === 'unknown') return 'unknown';
return predicate(arg) ? 'compatible' : 'incompatible';
}
function unaryShapeVerdict(
args: readonly ConstraintArgClass[],
predicate: (shape: ParameterTypeClass) => boolean,
options: { readonly requireTopLevelCv?: boolean } = {},
): ArityVerdict {
const arg = args[0];
if (arg === undefined || arg.typeClass === 'unknown') return 'unknown';
const shape = arg.shape;
if (shape === undefined || shape.indirection === 'unknown' || shape.cv === 'unknown') {
return 'unknown';
}
if (options.requireTopLevelCv === true && shape.indirection === 'pointer') {
return 'unknown';
}
return predicate(shape) ? 'compatible' : 'incompatible';
}
function isPlainValue(arg: ConstraintArgClass): boolean {
const shape = arg.shape;
if (shape === undefined) return true;
return shape.indirection === 'value';
}
function classifyConstraintArg(
token: string | undefined,
shape?: ParameterTypeClass,
): ConstraintArgClass {
if (shape !== undefined && shape.base.startsWith('enum:')) {
return { typeClass: 'enum', shape };
}
const typeClass = token === undefined || token === '' ? 'unknown' : classifyType(token);
return { typeClass, ...(shape !== undefined ? { shape } : {}) };
}
function tokenForArg(ctx: ConstraintContext, argIdx: number): string | undefined {
const shape = ctx.argumentTypeClasses?.[argIdx];
if (shape?.base.startsWith('enum:')) return shape.base;
return ctx.argumentTypes?.[argIdx];
}
function shapeForTemplateParam(
ctx: ConstraintContext,
paramName: string,
argIdx: number,
def?: SymbolDefinition,
): ParameterTypeClass | undefined {
const argShape = ctx.argumentTypeClasses?.[argIdx];
if (argShape === undefined) return undefined;
const paramShape = def?.parameterTypeClasses?.[argIdx];
if (paramShape === undefined) return argShape;
if (paramShape.base === paramName && paramShape.indirection === 'value') return argShape;
return undefined;
}
/** Public surface — registered as `ScopeResolver.constraintCompatibility`. */
export function cppConstraintCompatibility(
_callsite: Callsite,
def: SymbolDefinition,
ctx: ConstraintContext,
): ArityVerdict {
const payload = def.templateConstraints as CppConstraintPayload | undefined;
if (payload === undefined) return 'unknown';
return evaluate(payload.expr, payload, ctx, def);
}
function evaluate(
expr: ConstraintExpr,
payload: CppConstraintPayload,
ctx: ConstraintContext,
def?: SymbolDefinition,
): ArityVerdict {
switch (expr.kind) {
case 'unknown':
return 'unknown';
case 'atomic': {
const evaluator = REGISTRY.get(expr.name);
if (evaluator === undefined) return 'unknown';
const classes = expr.args.map((paramName) => {
const argIdx = payload.paramArgIndex[paramName];
if (argIdx === undefined) return { typeClass: 'unknown' as TypeClass };
return classifyConstraintArg(
tokenForArg(ctx, argIdx),
shapeForTemplateParam(ctx, paramName, argIdx, def),
);
});
return evaluator(classes);
}
case 'and': {
let result: ArityVerdict = 'compatible';
for (const child of expr.children) {
const v = evaluate(child, payload, ctx, def);
if (v === 'incompatible') return 'incompatible';
if (v === 'unknown') result = 'unknown';
}
return result;
}
case 'or': {
let result: ArityVerdict = 'incompatible';
for (const child of expr.children) {
const v = evaluate(child, payload, ctx, def);
if (v === 'compatible') return 'compatible';
if (v === 'unknown') result = 'unknown';
}
return result;
}
case 'not': {
const v = evaluate(expr.child, payload, ctx, def);
if (v === 'compatible') return 'incompatible';
if (v === 'incompatible') return 'compatible';
return 'unknown';
}
}
}
/** Exposed for unit tests — lets `cpp-constraint.test.ts` assert
* `expect(getRegistrySize()).toBe(4)` without exporting the Map itself. */
export function getRegistrySize(): number {
return REGISTRY.size;
}
/** Exposed for unit tests covering the Kleene 3-valued truth table
* directly, without an AST round-trip. */
export function evaluateForTest(
expr: ConstraintExpr,
payload: CppConstraintPayload,
ctx: ConstraintContext,
): ArityVerdict {
return evaluate(expr, payload, ctx);
}
@@ -1,32 +1,30 @@
/**
* C++ conversion-rank scoring for overload resolution (#1578).
* C++ conversion-rank scoring for overload resolution (#1578, #1637).
*
* Operates on **normalized** type strings (output of
* `normalizeCppParamType` in `arity-metadata.ts`). After normalization:
* - int/long/short/unsigned → 'int'
* - float/double → 'double'
* - char → 'char', bool → 'bool'
*
* Because the normalizer collapses promotion pairs (int↔long,
* float↔double) to the same string, those promotions are invisible at
* this layer — they appear as exact matches (rank 0).
* Operates on normalized type strings (output of `normalizeCppParamType`
* in `arity-metadata.ts`) plus optional shape sidecars from #1630.
* Normalization intentionally collapses cv/ref/pointer spelling for stable
* graph IDs, so pointer/nullptr rules must consult `ParameterTypeClass`.
*
* Post-normalization ranking:
* - rank 0 — exact (same normalized type)
* - rank 1 — integral promotion (char→int, bool→int)
* - rank 2 — standard arithmetic conversion (int↔double, char→double,
* bool→double)
* - Infinity — mismatch (string↔int, user types, pointers, etc.)
* - rank 0: exact (same normalized type)
* - rank 1: integral promotion (char -> int, bool -> int)
* - rank 2: standard conversion (arithmetic, nullptr -> T*, T* -> bool,
* T* -> void*)
* - rank 3: nullptr -> bool (kept worse than nullptr -> T*)
* - rank 4: ellipsis conversion (worst viable)
* - Infinity: mismatch (string -> int, user types, unsupported shapes)
*
* This function is intentionally C++-specific (issue #1578 pitfall:
* keep conversion-rank tables out of shared overload-narrowing). Other
* languages may define their own `ConversionRankFn` in the future.
* This function is intentionally C++-specific. Other languages may define
* their own `ConversionRankFn` in the future.
*/
import type { ParameterTypeClass } from 'gitnexus-shared';
/** Set of normalized arithmetic types that support implicit conversion. */
const ARITHMETIC = new Set(['int', 'double', 'char', 'bool']);
/** Integral promotion targets: char→int and bool→int are rank 1. */
/** Integral promotion targets: char -> int and bool -> int are rank 1. */
const INTEGRAL_PROMOTION = new Map([
['char', 'int'],
['bool', 'int'],
@@ -35,13 +33,40 @@ const INTEGRAL_PROMOTION = new Map([
/**
* Return the conversion rank from `argType` to `paramType`.
*
* @returns 0 for exact match, 1 for integral promotion (char/bool→int),
* 2 for standard arithmetic conversion, Infinity for mismatch.
* @returns 0 for exact match, 1 for integral promotion, 2 for standard
* conversion, 3 for nullptr -> bool, 4 for ellipsis, Infinity
* for mismatch.
*/
export function cppConversionRank(argType: string, paramType: string): number {
if (argType === paramType) return 0;
// Integral promotions: char→int, bool→int (ISO C++ [conv.prom])
export function cppConversionRank(
argType: string,
paramType: string,
argTypeClass?: ParameterTypeClass,
paramTypeClass?: ParameterTypeClass,
): number {
if (argType === paramType) {
return exactShapeCompatible(argTypeClass, paramTypeClass) ? 0 : Infinity;
}
if (paramType === '...') return 4;
if (INTEGRAL_PROMOTION.get(argType) === paramType) return 1;
if (ARITHMETIC.has(argType) && ARITHMETIC.has(paramType)) return 2;
if (argType === 'null' && isPointer(paramTypeClass)) return 2;
if (argType === 'null' && paramType === 'bool') return 3;
if (isPointer(argTypeClass) && paramType === 'bool') return 2;
if (isPointer(argTypeClass) && isPointer(paramTypeClass) && paramType === 'void') return 2;
return Infinity;
}
function isPointer(typeClass: ParameterTypeClass | undefined): boolean {
return typeClass?.indirection === 'pointer' && typeClass.pointerDepth > 0;
}
function exactShapeCompatible(
argTypeClass: ParameterTypeClass | undefined,
paramTypeClass: ParameterTypeClass | undefined,
): boolean {
if (argTypeClass === undefined || paramTypeClass === undefined) return true;
if (argTypeClass.indirection === 'unknown' || paramTypeClass.indirection === 'unknown') {
return true;
}
return isPointer(argTypeClass) === isPointer(paramTypeClass);
}
@@ -33,6 +33,7 @@ import {
resolveCppQualifiedNamespaceMember,
} from './inline-namespaces.js';
import { populateCppRangeBindings } from './range-bindings.js';
import { cppConstraintCompatibility } from './constraint-filter.js';
/**
* C++ `ScopeResolver` registered in `SCOPE_RESOLVERS` and consumed by
@@ -85,6 +86,12 @@ export const cppScopeResolver: ScopeResolver = {
// (def, callsite). ScopeResolver contract is (callsite, def).
arityCompatibility: (callsite, def) => cppArityCompatibility(def, callsite),
// SFINAE / `requires`-clause aware overload filter (issue #1579).
// Drops candidates whose template constraints (`enable_if_t<P, T>`,
// C++20 `requires P`) provably fail at the call site. Three-valued —
// `'unknown'` keeps the candidate, preserving "degrade not lie".
constraintCompatibility: cppConstraintCompatibility,
buildMro: (graph, parsedFiles, nodeLookup) =>
buildMro(graph, parsedFiles, nodeLookup, defaultLinearize),
@@ -0,0 +1,66 @@
/**
* Coarse-grained type classifier for C++ constraint evaluation
* (`<https://en.cppreference.com/w/cpp/types/is_integral>`,
* `<https://en.cppreference.com/w/cpp/types/is_floating_point>`).
*
* Maps a normalized type token (as produced by `normalizeCppParamType` /
* the call-site inference in `captures.ts`) to one of the categories
* the `<type_traits>` predicate registry uses for SFINAE filtering.
*
* `argumentTypes` remain normalized for overload narrowing, while
* constraint predicates that need cv/ref/pointer shape read the parallel
* `argumentTypeClasses` sidecar. Unknown shapes must stay unknown rather
* than being guessed as incompatible.
*/
export type TypeClass =
| 'integral'
| 'floating'
| 'bool'
| 'char'
| 'string'
| 'null'
| 'void'
| 'enum'
| 'class'
| 'pointer'
| 'reference'
| 'unknown';
/**
* Classify a normalized C++ type token. The mapping mirrors the literal-
* inference table in `captures.ts:inferCppLiteralType` plus the std::
* normalization in `arity-metadata.ts:normalizeCppParamType`.
*
* Caller note: token should be normalized for overload matching. Enum
* tokens produced by the C++ adapter use the internal `enum:<Name>`
* prefix so `is_enum_v` does not have to guess that every user token is
* class-like.
*/
export function classifyType(token: string): TypeClass {
if (token.length === 0) return 'unknown';
if (token.startsWith('enum:')) return 'enum';
switch (token) {
case 'void':
return 'void';
case 'int':
return 'integral';
case 'double':
case 'float':
return 'floating';
case 'bool':
return 'bool';
case 'char':
return 'char';
case 'string':
return 'string';
case 'null':
return 'null';
default:
// After normalization, anything that isn't a recognized primitive
// is assumed to be a class-like type. The Tier-A predicate registry
// doesn't introspect class types — `is_integral_v` etc. simply
// returns `false` for `'class'`, matching ISO behavior.
return 'class';
}
}
@@ -27,6 +27,7 @@ import { javaMethodConfig } from '../method-extractors/configs/jvm.js';
import { createVariableExtractor } from '../variable-extractors/generic.js';
import { javaVariableConfig } from '../variable-extractors/configs/jvm.js';
import { createHeritageExtractor } from '../heritage-extractors/generic.js';
import type { SymbolDefinition } from 'gitnexus-shared';
import {
emitJavaScopeCaptures,
interpretJavaImport,
@@ -39,6 +40,48 @@ import {
resolveJavaImportTarget,
} from './java/index.js';
const orderJavaSameNameTypeCandidates = ({
callSiteFilePath,
candidates,
}: {
readonly typeName: string;
readonly callSiteFilePath: string;
readonly candidates: readonly SymbolDefinition[];
}): readonly SymbolDefinition[] | null => {
if (!callSiteFilePath.endsWith('.java')) return null;
if (candidates.length <= 1) return null;
const callerDir = splitDirectorySegments(callSiteFilePath);
const scored = candidates.map((candidate, index) => ({
candidate,
index,
score: sharedPrefixLength(callerDir, splitDirectorySegments(candidate.filePath)),
}));
const bestScore = Math.max(...scored.map((entry) => entry.score));
// When all candidates tie, we have no structural signal to prefer one path.
// Returning null keeps downstream ambiguity handling conservative.
if (scored.every((entry) => entry.score === bestScore)) return null;
const ordered = [...scored]
.sort((a, b) => b.score - a.score || a.index - b.index)
.map((entry) => entry.candidate);
return ordered;
};
const splitDirectorySegments = (filePath: string): string[] => {
const normalized = filePath.replace(/\\/g, '/');
// Remove empty segments from leading/trailing/multiple slashes, then drop filename.
const segments = normalized.split('/').filter(Boolean);
return segments.slice(0, -1);
};
const sharedPrefixLength = (left: readonly string[], right: readonly string[]): number => {
const max = Math.min(left.length, right.length);
let idx = 0;
while (idx < max && left[idx] === right[idx]) idx += 1;
return idx;
};
export const javaProvider = defineLanguage({
id: SupportedLanguages.Java,
extensions: ['.java'],
@@ -87,4 +130,5 @@ export const javaProvider = defineLanguage({
receiverBinding: javaReceiverBinding,
arityCompatibility: javaArityCompatibility,
resolveImportTarget: resolveJavaImportTarget,
orderSameNameTypeCandidates: orderJavaSameNameTypeCandidates,
});
@@ -0,0 +1,12 @@
/**
* Arity compatibility for JavaScript.
*
* Delegates to `typescriptArityCompatibility` unchanged — JavaScript
* supports the same arity constructs (rest parameters `...args`, default
* parameters `p = v`) and the metadata shape (`parameterCount`,
* `requiredParameterCount`, `parameterTypes`) is synthesized by the same
* `computeTsArityMetadata` function (which understands both TS and JS
* parameter node types via `extractTsJsParameters`).
*/
export { typescriptArityCompatibility as jsArityCompatibility } from '../typescript/arity.js';
@@ -0,0 +1,722 @@
/**
* `emitScopeCaptures` for JavaScript.
*
* Adapts `emitTsScopeCaptures` for the JavaScript grammar:
*
* 1. **JS grammar** — uses `tree-sitter-javascript` instead of
* `tree-sitter-typescript`. The JS scope query is a subset of the
* TypeScript one (TypeScript-only node types dropped).
*
* 2. **CJS `require()` decomposition** — `const { X } = require('./m')`
* and `const X = require('./m')` are walked in a post-query pass and
* synthesized as `@import.kind/name/alias/source` markers so that
* `interpretJsImport` can recover a `ParsedImport` using the same
* shape as the TypeScript ESM decomposer.
*
* 3. **JSDoc type bindings** — JavaScript has no static type annotations
* so `@type-binding.parameter` / `@type-binding.return` must be
* inferred from leading JSDoc comments. A lightweight regex scanner
* (`parseJsDocParams` / `parseJsDocReturn`) extracts `@param {T} n`
* and `@returns {T}` tags and emits synthetic captures positioned on
* the annotated function node.
*
* 4. **Shared synthesis passes** — destructuring, for-of map-tuple, and
* instanceof narrowing passes are duplicated from `typescript/captures.ts`
* (they are pure AST operations with no grammar-specific logic).
*
* Pure given the input source text. No I/O, no globals consulted.
*/
import type { Capture, CaptureMatch } from 'gitnexus-shared';
import {
findNodeAtRange,
nodeToCapture,
syntheticCapture,
type SyntaxNode,
} from '../../utils/ast-helpers.js';
import { splitImportStatement } from '../typescript/import-decomposer.js';
import { getJsParser, getJsScopeQuery, jsCachedTreeMatchesGrammar } from './query.js';
import { computeTsArityMetadata } from '../typescript/arity-metadata.js';
import { synthesizeTsReceiverBinding } from '../typescript/receiver-binding.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
/** JS function-like node types that may carry a synthesized `this` binding.
* Kept in sync with the `@scope.function` patterns in `query.ts`. */
const FUNCTION_NODE_TYPES = [
'method_definition',
'arrow_function',
'function_expression',
'function_declaration',
'generator_function_declaration',
] as const;
/** Declaration anchors that carry function-like arity metadata. */
const FUNCTION_DECL_TAGS = ['@declaration.method', '@declaration.function'] as const;
/** Callsite anchors that should carry `@reference.arity` + param types. */
const CALL_TAGS = [
'@reference.call.free',
'@reference.call.member',
'@reference.call.constructor',
] as const;
function pickFirstDefined(grouped: CaptureMatch, tags: readonly string[]): Capture | undefined {
for (const tag of tags) {
const cap = grouped[tag];
if (cap !== undefined) return cap;
}
return undefined;
}
/** Filter `@reference.read.member` in non-read contexts (same logic as TS). */
function shouldEmitReadMember(memberNode: SyntaxNode): boolean {
const parent = memberNode.parent;
if (parent === null) return true;
switch (parent.type) {
case 'call_expression':
return parent.childForFieldName('function')?.id !== memberNode.id;
case 'new_expression':
return parent.childForFieldName('constructor')?.id !== memberNode.id;
case 'assignment_expression':
case 'augmented_assignment_expression':
return parent.childForFieldName('left')?.id !== memberNode.id;
case 'jsx_self_closing_element':
case 'jsx_opening_element':
return parent.childForFieldName('name')?.id !== memberNode.id;
default:
return true;
}
}
/** Find the first JS function-like node at the given range. */
function findFunctionNode(rootNode: SyntaxNode, range: Capture['range']): SyntaxNode | null {
for (const nodeType of FUNCTION_NODE_TYPES) {
const n = findNodeAtRange(rootNode, range, nodeType);
if (n !== null) return n;
}
return null;
}
/** Infer a callsite argument's static type from literal shapes. */
function inferArgType(argNode: SyntaxNode): string {
switch (argNode.type) {
case 'number':
return 'number';
case 'string':
case 'template_string':
return 'string';
case 'true':
case 'false':
return 'boolean';
case 'null':
return 'null';
case 'undefined':
return 'undefined';
case 'array':
return 'Array';
case 'object':
return 'object';
case 'regex':
return 'RegExp';
case 'new_expression': {
const ctor = argNode.childForFieldName('constructor');
return ctor?.text ?? '';
}
default:
return '';
}
}
// ─── CJS require() decomposition ─────────────────────────────────────────
/**
* Walk the AST and synthesize `@import.*` captures for CJS `require()` calls:
*
* - `const { X, Y } = require('./m')` → one match per destructured name,
* `@import.kind = 'named'`, `@import.name = X / Y`.
* - `const X = require('./m')` → `@import.kind = 'namespace'`,
* `@import.alias = X` (the whole module is bound to X).
* - `require('./m')` as a bare expression-statement → side-effect.
*
* CJS named-alias form (`const { X: alias } = require('./m')`) emits
* `@import.kind = 'named-alias'` with `@import.name = X` and
* `@import.alias = alias`.
*
* The synthesized markers are identical to those produced by
* `splitImportStatement` for ESM, so `interpretJsImport` can delegate
* unchanged to `interpretTsImport` for all cases.
*/
function synthesizeCjsImports(root: SyntaxNode, out: CaptureMatch[]): void {
const stack: SyntaxNode[] = [root];
for (;;) {
const node = stack.pop();
if (node === undefined) break;
for (const child of node.namedChildren) {
if (child !== null) stack.push(child);
}
if (node.type !== 'call_expression') continue;
// Require call: function must be bare identifier "require".
const fn = node.childForFieldName('function');
if (fn === null || fn.type !== 'identifier' || fn.text !== 'require') continue;
const argsNode = node.childForFieldName('arguments');
if (argsNode === null) continue;
// Source must be a string literal.
const firstArg = argsNode.namedChild(0);
if (firstArg === null || firstArg.type !== 'string') continue;
const rawSource = firstArg.text; // includes surrounding quotes
const source = firstArg.namedChild(0)?.text ?? rawSource.slice(1, -1);
const parent = node.parent;
// Case 1: const { X } = require('./m') OR const X = require('./m')
if (parent?.type === 'variable_declarator') {
const nameNode = parent.childForFieldName('name');
if (nameNode === null) continue;
if (nameNode.type === 'object_pattern') {
// Destructured: emit one match per specifier.
for (const field of nameNode.namedChildren) {
if (field === null) continue;
if (field.type === 'shorthand_property_identifier_pattern') {
const name = field.text;
out.push({
'@import.statement': syntheticCapture('@import.statement', node, rawSource),
'@import.kind': syntheticCapture('@import.kind', node, 'named'),
'@import.name': syntheticCapture('@import.name', field, name),
'@import.source': syntheticCapture('@import.source', firstArg, source),
});
} else if (field.type === 'pair_pattern') {
const key = field.childForFieldName('key');
const value = field.childForFieldName('value');
if (key === null || value === null || value.type !== 'identifier') continue;
out.push({
'@import.statement': syntheticCapture('@import.statement', node, rawSource),
'@import.kind': syntheticCapture('@import.kind', node, 'named-alias'),
'@import.name': syntheticCapture('@import.name', key, key.text),
'@import.alias': syntheticCapture('@import.alias', value, value.text),
'@import.source': syntheticCapture('@import.source', firstArg, source),
});
}
}
} else if (nameNode.type === 'identifier') {
// Namespace-style: const X = require('./m') → bind whole module to X.
out.push({
'@import.statement': syntheticCapture('@import.statement', node, rawSource),
'@import.kind': syntheticCapture('@import.kind', node, 'namespace'),
'@import.alias': syntheticCapture('@import.alias', nameNode, nameNode.text),
'@import.source': syntheticCapture('@import.source', firstArg, source),
});
}
continue;
}
// Case 2: bare require('./m') — side-effect import.
if (parent?.type === 'expression_statement') {
out.push({
'@import.statement': syntheticCapture('@import.statement', node, rawSource),
'@import.kind': syntheticCapture('@import.kind', node, 'side-effect'),
'@import.source': syntheticCapture('@import.source', firstArg, source),
});
}
}
}
// ─── JSDoc type binding synthesis ────────────────────────────────────────
interface JsDocParam {
readonly name: string;
readonly type: string;
}
/** Extract `@param {Type} name` entries from a JSDoc comment block. */
function parseJsDocParams(text: string): readonly JsDocParam[] {
const results: JsDocParam[] = [];
// Match @param {Type} name or @param {Type} [name] (optional)
const re = /@param\s+\{([^}]+)\}\s+\[?(\w+)\]?/g;
let m: RegExpExecArray | null;
while ((m = re.exec(text)) !== null) {
results.push({ type: m[1].trim(), name: m[2].trim() });
}
return results;
}
/** Extract `@returns {Type}` or `@return {Type}` from a JSDoc comment. */
function parseJsDocReturn(text: string): string | null {
const m = /@returns?\s+\{([^}]+)\}/.exec(text);
return m ? m[1].trim() : null;
}
/** Extract `@type {Type}` from a JSDoc comment (variable-level annotation). */
function parseJsDocType(text: string): string | null {
const m = /@type\s+\{([^}]+)\}/.exec(text);
return m ? m[1].trim() : null;
}
/**
* Walk the AST and synthesize `@type-binding.*` captures from JSDoc
* comments immediately preceding function declarations / expressions.
*
* Only `/** … *​/` block comments are scanned. Line comments (`//`) are
* intentionally excluded — JSDoc lives in block comments.
*
* Emits:
* - `@type-binding.parameter` for each `@param {T} n` tag.
* - `@type-binding.return` for `@returns {T}` / `@return {T}`.
* - `@type-binding.annotation` for `@type {T}` on `let`/`const`/`var`
* declarations — covers the common `/** @type {User} *​/ const u = …`
* pattern (ECMA-262 §14.3.1/§14.3.2 variable declarations).
*
* The binding is anchored on the function node so `tsBindingScopeFor`
* can hoist method return-type bindings to Module scope (matching the
* TypeScript path where `hoistTypeBindingsToModule: true`).
*/
function synthesizeJsDocBindings(root: SyntaxNode, out: CaptureMatch[]): void {
const stack: SyntaxNode[] = [root];
for (;;) {
const node = stack.pop();
if (node === undefined) break;
for (const child of node.namedChildren) {
if (child !== null) stack.push(child);
}
const isFnDecl =
node.type === 'function_declaration' || node.type === 'generator_function_declaration';
const isMethodDef = node.type === 'method_definition';
// Also check lexical_declaration containing an arrow/fn-expression
const isLexDecl = node.type === 'lexical_declaration' || node.type === 'variable_declaration';
if (!isFnDecl && !isMethodDef && !isLexDecl) continue;
// For `export function foo() { ... }`, the JSDoc comment precedes the
// wrapping export_statement, not the inner function_declaration.
// Walk up to the export_statement so the preceding-sibling search finds it.
const lookupNode =
(isFnDecl || isLexDecl) && node.parent?.type === 'export_statement' ? node.parent : node;
// Find the preceding sibling comment.
let sibling = lookupNode.previousNamedSibling;
while (sibling !== null && sibling.type === 'comment') {
const text = sibling.text;
if (text.startsWith('/**')) {
// Found a JSDoc block.
const params = parseJsDocParams(text);
const retType = parseJsDocReturn(text);
const varType = isLexDecl ? parseJsDocType(text) : null;
// Determine the anchor node (the function-like node, for hoisting).
const anchor = node;
for (const p of params) {
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', anchor, p.name),
'@type-binding.type': syntheticCapture('@type-binding.type', anchor, p.type),
'@type-binding.parameter': syntheticCapture('@type-binding.parameter', anchor, '1'),
});
}
if (retType !== null) {
// For named functions, use the function name as the binding name so
// `hoistTypeBindingsToModule` knows which function's return type this is.
let fnName: string | null = null;
if (isFnDecl) {
fnName = node.childForFieldName('name')?.text ?? null;
} else if (isMethodDef) {
// method_definition uses `name:` field for the method name
const nameNode = node.childForFieldName('name');
if (nameNode?.type === 'property_identifier') fnName = nameNode.text;
} else if (isLexDecl) {
const declarator = node.namedChild(0);
const nameNode = declarator?.childForFieldName('name');
if (nameNode?.type === 'identifier') fnName = nameNode.text;
}
if (fnName !== null) {
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', anchor, fnName),
'@type-binding.type': syntheticCapture('@type-binding.type', anchor, retType),
'@type-binding.return': syntheticCapture('@type-binding.return', anchor, '1'),
});
}
}
// @type {T} on let/const/var: `/** @type {User} */ const u = getUser()`.
// Emits annotation-strength binding (source = 'annotation') so it
// overrides any weaker constructor/alias inference on the same name.
if (varType !== null) {
for (const declarator of node.namedChildren) {
if (declarator === null || declarator.type !== 'variable_declarator') continue;
const nameNode = declarator.childForFieldName('name');
if (nameNode === null || nameNode.type !== 'identifier') continue;
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', nameNode, nameNode.text),
'@type-binding.type': syntheticCapture('@type-binding.type', nameNode, varType),
'@type-binding.annotation': syntheticCapture(
'@type-binding.annotation',
nameNode,
'1',
),
});
}
}
break;
}
sibling = sibling.previousNamedSibling;
}
}
}
// ─── Destructuring / for-of / instanceof (shared with TS captures) ───────
function synthesizeDestructuringBindings(root: SyntaxNode, out: CaptureMatch[]): void {
const stack: SyntaxNode[] = [root];
for (;;) {
const node = stack.pop();
if (node === undefined) break;
for (const child of node.namedChildren) {
if (child !== null) stack.push(child);
}
if (node.type !== 'variable_declarator') continue;
const nameNode = node.childForFieldName('name');
const valueNode = node.childForFieldName('value');
if (nameNode === null || valueNode === null) continue;
if (nameNode.type !== 'object_pattern') continue;
if (valueNode.type !== 'identifier') continue;
const rhsName = valueNode.text;
for (const fieldNode of nameNode.namedChildren) {
if (fieldNode === null) continue;
if (fieldNode.type === 'shorthand_property_identifier_pattern') {
const localName = fieldNode.text;
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', fieldNode, localName),
'@type-binding.type': syntheticCapture(
'@type-binding.type',
fieldNode,
`${rhsName}.${localName}`,
),
'@type-binding.destructured': syntheticCapture(
'@type-binding.destructured',
fieldNode,
fieldNode.text,
),
});
} else if (fieldNode.type === 'pair_pattern') {
const key = fieldNode.childForFieldName('key');
const value = fieldNode.childForFieldName('value');
if (key === null || value === null || value.type !== 'identifier') continue;
const fieldName = key.text;
const localName = value.text;
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', value, localName),
'@type-binding.type': syntheticCapture(
'@type-binding.type',
fieldNode,
`${rhsName}.${fieldName}`,
),
'@type-binding.destructured': syntheticCapture(
'@type-binding.destructured',
fieldNode,
fieldNode.text,
),
});
}
}
}
}
function synthesizeForOfMapTupleBindings(root: SyntaxNode, out: CaptureMatch[]): void {
const stack: SyntaxNode[] = [root];
for (;;) {
const node = stack.pop();
if (node === undefined) break;
for (const child of node.namedChildren) {
if (child !== null) stack.push(child);
}
if (node.type !== 'for_in_statement') continue;
const left = node.childForFieldName('left');
const right = node.childForFieldName('right');
if (left === null || right === null) continue;
if (left.type !== 'array_pattern' || right.type !== 'identifier') continue;
const rhs = right.text;
let slot = 0;
for (const child of left.namedChildren) {
if (child === null || child.type !== 'identifier') continue;
const localName = child.text;
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', child, localName),
'@type-binding.type': syntheticCapture(
'@type-binding.type',
child,
`__MAP_TUPLE_${slot}__:${rhs}`,
),
'@type-binding.map-tuple-entry': syntheticCapture(
'@type-binding.map-tuple-entry',
child,
String(slot),
),
});
slot++;
}
}
}
function synthesizeInstanceofNarrowings(root: SyntaxNode, out: CaptureMatch[]): void {
const stack: SyntaxNode[] = [root];
for (;;) {
const node = stack.pop();
if (node === undefined) break;
for (const child of node.namedChildren) {
if (child !== null) stack.push(child);
}
if (node.type !== 'if_statement') continue;
const cond = node.childForFieldName('condition');
if (cond === null) continue;
const inner = cond.type === 'parenthesized_expression' ? cond.namedChildren[0] : cond;
if (inner === null || inner.type !== 'binary_expression') continue;
const op = inner.childForFieldName('operator');
const left = inner.childForFieldName('left');
const right = inner.childForFieldName('right');
if (op === null || left === null || right === null) continue;
if (op.type !== 'instanceof') continue;
if (left.type !== 'identifier') continue;
if (right.type !== 'identifier') continue;
const varName = left.text;
const typeName = right.text;
const cons = node.childForFieldName('consequence');
if (cons === null) continue;
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', cons, varName),
'@type-binding.type': syntheticCapture('@type-binding.type', right, typeName),
'@type-binding.instanceof-narrow': syntheticCapture(
'@type-binding.instanceof-narrow',
cons,
'1',
),
});
}
}
// ─── Constructor field type bindings ─────────────────────────────────────
/**
* Synthesize class-scope type bindings from `this.X = new Y()` assignments
* inside constructor method bodies. Covers the traditional ES5+ OOP pattern:
*
* class User {
* constructor() {
* /** @type {Address} *\/
* this.address = new Address();
* }
* }
*
* The emitted `@type-binding.class-field` is hoisted to the Class scope by
* `tsBindingScopeFor` so that compound-receiver resolution can look up
* `User.address → Address` when resolving `user.address.save()`.
*
* Type source priority:
* 1. JSDoc `@type {T}` comment immediately preceding the statement
* 2. `new Y()` constructor inference
*/
function synthesizeConstructorFieldBindings(root: SyntaxNode, out: CaptureMatch[]): void {
const stack: SyntaxNode[] = [root];
for (;;) {
const node = stack.pop();
if (node === undefined) break;
for (const child of node.namedChildren) {
if (child !== null) stack.push(child);
}
// Only process constructor method definitions
if (node.type !== 'method_definition') continue;
const nameNode = node.childForFieldName('name');
if (nameNode?.text !== 'constructor') continue;
const body = node.childForFieldName('body');
if (body === null) continue;
for (const stmt of body.namedChildren) {
if (stmt === null || stmt.type !== 'expression_statement') continue;
const expr = stmt.namedChild(0);
if (expr === null || expr.type !== 'assignment_expression') continue;
const left = expr.childForFieldName('left');
const right = expr.childForFieldName('right');
if (left === null || right === null) continue;
if (left.type !== 'member_expression') continue;
const obj = left.childForFieldName('object');
const prop = left.childForFieldName('property');
if (obj === null || prop === null) continue;
if (obj.text !== 'this' || prop.type !== 'property_identifier') continue;
const fieldName = prop.text;
// Prefer JSDoc @type annotation on the preceding sibling comment.
let typeName: string | null = null;
const prevSib: SyntaxNode | null = stmt.previousNamedSibling;
if (prevSib !== null && prevSib.type === 'comment') {
const m = /@type\s*\{([^}]+)\}/.exec(prevSib.text);
if (m?.[1]) typeName = m[1].trim();
}
// Fall back to constructor inference from `new Y()`.
if (typeName === null && right.type === 'new_expression') {
const ctor = right.childForFieldName('constructor');
if (ctor !== null && ctor.type === 'identifier') typeName = ctor.text;
}
if (typeName === null) continue;
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', prop, fieldName),
'@type-binding.type': syntheticCapture('@type-binding.type', prop, typeName),
// Anchor: positioned inside the constructor body so tsBindingScopeFor
// can walk up from the Function (constructor) scope to the Class scope.
'@type-binding.class-field': syntheticCapture('@type-binding.class-field', stmt, '1'),
});
}
}
}
// ─── Main emitter ──────────────────────────────────────────────────────────
export function emitJsScopeCaptures(
sourceText: string,
filePath: string,
cachedTree?: unknown,
): readonly CaptureMatch[] {
let tree = cachedTree as ReturnType<ReturnType<typeof getJsParser>['parse']> | undefined;
if (tree !== undefined && !jsCachedTreeMatchesGrammar(tree)) {
tree = undefined;
}
if (tree === undefined) {
tree = parseSourceSafe(getJsParser(filePath), sourceText, undefined, {
bufferSize: getTreeSitterBufferSize(sourceText),
});
}
const rawMatches = getJsScopeQuery(filePath).matches(tree.rootNode);
const out: CaptureMatch[] = [];
for (const m of rawMatches) {
const grouped: Record<string, Capture> = {};
for (const c of m.captures) {
const tag = '@' + c.name;
grouped[tag] = nodeToCapture(tag, c.node);
}
if (Object.keys(grouped).length === 0) continue;
// Decompose ESM import_statement / re-export export_statement.
if (grouped['@import.statement'] !== undefined) {
const stmtCapture = grouped['@import.statement'];
const stmtNode =
findNodeAtRange(tree.rootNode, stmtCapture.range, 'import_statement') ??
findNodeAtRange(tree.rootNode, stmtCapture.range, 'export_statement');
if (stmtNode !== null) {
const decomposed = splitImportStatement(stmtNode);
for (const d of decomposed) out.push(d);
}
continue;
}
// Decompose dynamic import() calls.
if (grouped['@import.dynamic'] !== undefined) {
const dynCapture = grouped['@import.dynamic'];
const callNode = findNodeAtRange(tree.rootNode, dynCapture.range, 'call_expression');
if (callNode !== null) {
const decomposed = splitImportStatement(callNode);
for (const d of decomposed) out.push(d);
}
continue;
}
// Filter @reference.read.member false-positives.
if (grouped['@reference.read.member'] !== undefined) {
const anchor = grouped['@reference.read.member'];
const memberNode = findNodeAtRange(tree.rootNode, anchor.range, 'member_expression');
if (memberNode === null || !shouldEmitReadMember(memberNode)) {
continue;
}
}
// Synthesize arity metadata on function-like declarations.
const declAnchor = pickFirstDefined(grouped, FUNCTION_DECL_TAGS);
if (declAnchor !== undefined) {
const fnNode = findFunctionNode(tree.rootNode, declAnchor.range);
if (fnNode !== null) {
const arity = computeTsArityMetadata(fnNode);
if (arity.parameterCount !== undefined) {
grouped['@declaration.parameter-count'] = syntheticCapture(
'@declaration.parameter-count',
fnNode,
String(arity.parameterCount),
);
}
if (arity.requiredParameterCount !== undefined) {
grouped['@declaration.required-parameter-count'] = syntheticCapture(
'@declaration.required-parameter-count',
fnNode,
String(arity.requiredParameterCount),
);
}
if (arity.parameterTypes !== undefined) {
grouped['@declaration.parameter-types'] = syntheticCapture(
'@declaration.parameter-types',
fnNode,
JSON.stringify(arity.parameterTypes),
);
}
}
}
// Synthesize @reference.arity on callsites.
const callAnchor = pickFirstDefined(grouped, CALL_TAGS);
if (callAnchor !== undefined && grouped['@reference.arity'] === undefined) {
const callNode =
findNodeAtRange(tree.rootNode, callAnchor.range, 'call_expression') ??
findNodeAtRange(tree.rootNode, callAnchor.range, 'new_expression');
if (callNode !== null) {
const argList = callNode.childForFieldName('arguments');
const args: SyntaxNode[] =
argList === null
? []
: argList.namedChildren.filter(
(c): c is SyntaxNode => c !== null && c.type !== 'comment',
);
grouped['@reference.arity'] = syntheticCapture(
'@reference.arity',
callNode,
String(args.length),
);
grouped['@reference.parameter-types'] = syntheticCapture(
'@reference.parameter-types',
callNode,
JSON.stringify(args.map(inferArgType)),
);
}
}
out.push(grouped);
// Synthesize `this` receiver type-bindings on class member functions.
const scopeFnAnchor = grouped['@scope.function'];
if (scopeFnAnchor !== undefined) {
const fnNode = findFunctionNode(tree.rootNode, scopeFnAnchor.range);
if (fnNode !== null) {
const synth = synthesizeTsReceiverBinding(fnNode);
if (synth !== null) out.push(synth);
}
}
}
// Post-query synthesis passes.
synthesizeCjsImports(tree.rootNode, out);
synthesizeJsDocBindings(tree.rootNode, out);
synthesizeConstructorFieldBindings(tree.rootNode, out);
synthesizeDestructuringBindings(tree.rootNode, out);
synthesizeForOfMapTupleBindings(tree.rootNode, out);
synthesizeInstanceofNarrowings(tree.rootNode, out);
return out;
}
@@ -0,0 +1,72 @@
/**
* Import-target resolver for JavaScript.
*
* Delegates to the TypeScript `resolveTsTarget` standard-strategy resolver
* with `language: SupportedLanguages.JavaScript` so the resolver tries
* `.js` / `.jsx` extensions in addition to (or instead of) `.ts` / `.tsx`.
*
* The `TsResolveContext.language` flag already exists in `import-target.ts`
* and the resolver (`resolveImportPath`) already branches on it — this
* adapter just wires the right value in.
*
* CJS `require()` calls reference the same module-path strings as ESM
* `import` statements, so the resolver handles them uniformly without any
* CJS-specific logic here.
*
* No `tsconfig.json` path-alias support (JavaScript projects don't use
* `tsconfig.json` compilerOptions.paths in general). Projects that DO use
* tsconfig-based aliases alongside JavaScript can still resolve via the
* standard extension-suffix fallback; the alias branch is a no-op when
* `tsconfigPaths` is null.
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { resolveTsTarget, type TsResolveContext } from '../typescript/import-target.js';
export type JsResolveContext = TsResolveContext;
type PassCache = {
readonly key: ReadonlySet<string>;
readonly allFilePaths: Set<string>;
readonly allFileList: readonly string[];
readonly normalizedFileList: readonly string[];
readonly resolveCache: Map<string, string | null>;
};
/**
* Build a memoized `resolveImportTarget` adapter for JavaScript.
* Caches the derived arrays and per-pass resolve cache across
* `resolveImportTarget` calls within a single workspace pass.
*/
export function makeJsResolveImportTarget(): (
targetRaw: string,
fromFile: string,
allFilePaths: ReadonlySet<string>,
resolutionConfig?: unknown,
) => string | readonly string[] | null {
let cached: PassCache | null = null;
return (targetRaw, fromFile, allFilePaths) => {
if (cached === null || cached.key !== allFilePaths) {
const allFileList = Array.from(allFilePaths);
cached = {
key: allFilePaths,
allFilePaths: new Set(allFilePaths),
allFileList,
normalizedFileList: allFileList.map((f) => f.toLowerCase()),
resolveCache: new Map(),
};
}
const ws: JsResolveContext = {
fromFile,
language: SupportedLanguages.JavaScript,
allFilePaths: cached.allFilePaths,
allFileList: cached.allFileList,
normalizedFileList: cached.normalizedFileList,
resolveCache: cached.resolveCache,
tsconfigPaths: null,
};
return resolveTsTarget(targetRaw, ws);
};
}
@@ -0,0 +1,49 @@
/**
* JavaScript scope-resolution hooks (RFC #909 Ring 3, issue #928).
*
* Public API barrel. Consumers should import from this file rather
* than the individual modules.
*
* Module layout (each file is a single concern):
*
* - `query.ts` — JS scope query string + lazy parser/query
* singletons (`getJsParser`, `getJsScopeQuery`)
* - `captures.ts` — `emitJsScopeCaptures` — runs the JS scope query,
* synthesizes CJS require() imports and JSDoc-
* derived type bindings, delegates arity synthesis
* and destructuring/instanceof passes to shared
* or TypeScript utilities
* - `interpret.ts` — `interpretJsImport` / `interpretJsTypeBinding`
* (delegate to TypeScript interpreters — same
* capture-marker vocabulary)
* - `simple-hooks.ts` — `jsBindingScopeFor` (var hoisting),
* `jsImportOwningScope`, `jsReceiverBinding`
* (all delegate to TypeScript counterparts)
* - `merge-bindings.ts` — `jsMergeBindings` (LEGB via typescriptMergeBindings)
* - `arity.ts` — `jsArityCompatibility` (delegates to TS function)
* - `import-target.ts` — `makeJsResolveImportTarget` (memoized adapter)
* - `scope-resolver.ts` — `javascriptScopeResolver` wiring object
*
* ## Known limitations
*
* 1. **JSDoc coverage** — `@param {T} name`, `@returns {T}` / `@return {T}`,
* and `@type {T}` on variable declarations are synthesized. `@typedef`
* is not yet synthesized (tracked in #1646).
* 2. **CJS chained destructuring** — `const { X: { Y } } = require(...)`
* (nested destructuring) emits only the outer `X` binding; `Y` is not
* resolved.
* 3. **Dynamic require** — `require(computedPath)` is skipped (non-literal
* argument — cannot statically resolve the target).
* 4. **`module.exports` / `exports.X`** — CJS export forms are not yet
* modeled as re-exports. The finalize algorithm treats the exporting
* module as a namespace; importers that do `const X = require('./m')`
* bind the module namespace, and member-call resolution walks the
* class graph from there.
*/
export { emitJsScopeCaptures } from './captures.js';
export { interpretJsImport, interpretJsTypeBinding } from './interpret.js';
export { jsMergeBindings } from './merge-bindings.js';
export { jsArityCompatibility } from './arity.js';
export { makeJsResolveImportTarget } from './import-target.js';
export { jsBindingScopeFor, jsImportOwningScope, jsReceiverBinding } from './simple-hooks.js';
@@ -0,0 +1,45 @@
/**
* Capture-match → semantic-shape interpreters for JavaScript.
*
* `interpretJsImport` delegates to `interpretTsImport` for all cases
* because `emitJsScopeCaptures` synthesizes the same
* `@import.kind/name/alias/source` markers for both ESM and CJS imports.
*
* The `@import.kind` values emitted for CJS by `captures.ts`:
*
* - `'named'` : `const { X } = require('./m')` → named import
* - `'named-alias'` : `const { X: Y } = require('./m')` → aliased import
* - `'namespace'` : `const X = require('./m')` → namespace import
* - `'side-effect'` : `require('./m')` bare expression → side-effect
*
* These match the kinds `interpretTsImport` already handles for ESM
* (`import { X }`, `import { X as Y }`, `import * as X`, `import './m'`),
* so no new branch is needed here.
*
* `interpretJsTypeBinding` handles the JS-only `@type-binding.class-field`
* tag before delegating to `interpretTsTypeBinding`. The class-field tag
* is emitted by `synthesizeConstructorFieldBindings` and should produce
* `source = 'annotation'` — the same strength as an explicit type
* annotation. Remapping it to `@type-binding.annotation` achieves this
* without adding a JS-specific branch to the shared TS interpreter
* (DoD.md §2.2).
*/
import type { CaptureMatch, ParsedImport, ParsedTypeBinding } from 'gitnexus-shared';
import { interpretTsImport, interpretTsTypeBinding } from '../typescript/interpret.js';
export function interpretJsImport(captures: CaptureMatch): ParsedImport | null {
return interpretTsImport(captures);
}
export function interpretJsTypeBinding(captures: CaptureMatch): ParsedTypeBinding | null {
// @type-binding.class-field is a JS-only tag emitted by
// synthesizeConstructorFieldBindings. Remap it to the standard
// @type-binding.annotation tag so interpretTsTypeBinding assigns
// source = 'annotation' without a JS-specific branch in shared code.
if (captures['@type-binding.class-field'] !== undefined) {
const { '@type-binding.class-field': classField, ...rest } = captures;
return interpretTsTypeBinding({ ...rest, '@type-binding.annotation': classField });
}
return interpretTsTypeBinding(captures);
}
@@ -0,0 +1,21 @@
/**
* Binding-merge precedence for JavaScript.
*
* JavaScript has no TypeScript declaration-merging (no `interface + class`
* coexisting in the same scope, no `namespace + class` dual-space declarations).
* However, `typescriptMergeBindings` handles these by falling back to
* `['value']` for any `NodeLabel` not explicitly mapped to multiple spaces —
* which is what every JavaScript declaration produces. The result is pure
* LEGB precedence without any cross-space logic, which is exactly what
* JavaScript needs.
*
* Reuse rather than reimplementing to keep the single source of truth for
* the tier (local 0 / import-namespace-reexport 1 / wildcard 2) ordering.
*/
import type { BindingRef } from 'gitnexus-shared';
import { typescriptMergeBindings } from '../typescript/merge-bindings.js';
export function jsMergeBindings(bindings: readonly BindingRef[]): readonly BindingRef[] {
return typescriptMergeBindings(bindings);
}
@@ -0,0 +1,421 @@
/**
* Tree-sitter query for JavaScript scope captures (RFC §5.1, Ring 3).
*
* Subset of the TypeScript scope query (`languages/typescript/query.ts`)
* compiled against `tree-sitter-javascript`. TypeScript-only node types
* (`interface_declaration`, `type_alias_declaration`, `enum_declaration`,
* `internal_module`, `abstract_class_declaration`, `function_signature`,
* `method_signature`, `abstract_method_signature`, `type_annotation`,
* `public_field_definition`) are dropped because:
*
* 1. The JS grammar doesn't define them — the query compiler would
* throw `InvalidNodeType` if they were included.
* 2. JavaScript has no static type annotations, so the `@type-binding.*`
* patterns derived from TS annotation nodes don't apply.
*
* What IS shared with the TypeScript query:
*
* - Scope patterns: `program`, `class_declaration`, `(class)` (the JS
* grammar node for class expressions — NOT `class_expression`, which
* does not exist in `tree-sitter-javascript`), `function_declaration`,
* `generator_function_declaration`, `function_expression`,
* `arrow_function`, `method_definition`.
* - Declaration patterns for functions, classes, const/let/var,
* object-property arrows (Zustand, TanStack, etc.), and HOC-wrapped
* variable declarations (forwardRef / memo / useCallback / useMemo).
* - Import patterns: `import_statement`, `export_statement` re-exports,
* and dynamic `import()` (represented as `call_expression(import)` in
* both grammars — the `import` leaf node exists in tree-sitter-javascript
* as well as tree-sitter-typescript).
* - Type-binding patterns that work without static annotations:
* constructor inference (`new User()`), call-result alias
* (`const u = getUser()`), member-access alias (`const a = u.addr`),
* identifier alias, assignment rebind, and for-of element bindings.
* JSDoc-derived type bindings (`@param {User} u`, `@returns {User}`)
* are handled separately in `captures.ts` via comment-node scanning.
* - Reference patterns: free calls, member calls, constructor calls,
* write-access, read-access, and dynamic import.
*
* CJS `require()` is NOT captured here; it is handled in `captures.ts`
* by scanning parent context (destructured vs. namespace) of `call_expression`
* nodes whose callee is the identifier `require`.
*
* Grammar version: `tree-sitter-javascript` pinned in gitnexus/package.json.
*
* Exposes lazy `Parser` and `Query` singletons so callers don't pay
* tree-sitter init cost per file.
*/
import Parser from 'tree-sitter';
import JS from 'tree-sitter-javascript';
const JS_GRAMMAR = JS as Parameters<Parser['setLanguage']>[0];
/** True when the file should be parsed with the JSX-extended query. */
function isJsxFile(filePath: string): boolean {
return filePath.endsWith('.jsx');
}
const JAVASCRIPT_SCOPE_QUERY = `
;; Scopes — module / class-likes / function-likes
(program) @scope.module
(class_declaration) @scope.class
(class) @scope.class
(function_declaration) @scope.function
(generator_function_declaration) @scope.function
(function_expression) @scope.function
(arrow_function) @scope.function
(method_definition) @scope.function
;; Declarations — classes
(class_declaration
name: (identifier) @declaration.name) @declaration.class
;; Declarations — methods (inside class bodies)
(method_definition
name: (property_identifier) @declaration.name) @declaration.method
;; Declarations — class fields (JS uses field_definition, not public_field_definition)
(field_definition
property: (property_identifier) @declaration.name) @declaration.property
;; Declarations — free functions
(function_declaration
name: (identifier) @declaration.name) @declaration.function
(generator_function_declaration
name: (identifier) @declaration.name) @declaration.function
;; Arrow / function-expression assigned to a const/let/var.
;; Anchor discipline: @declaration.function sits on the INNER arrow or
;; function_expression, NOT on the lexical_declaration wrapper. This
;; aligns anchor.range with the @scope.function range so
;; pass2AttachDeclarations resolves the innermost scope correctly and
;; resolveCallerGraphId walks up to the right caller anchor.
(lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (arrow_function) @declaration.function))
(lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (function_expression) @declaration.function))
(export_statement
declaration: (lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (arrow_function) @declaration.function)))
(export_statement
declaration: (lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (function_expression) @declaration.function)))
(variable_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (arrow_function) @declaration.function))
(variable_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (function_expression) @declaration.function))
;; Object-property arrows / function expressions named by their pair key.
;; Same anchor discipline as the lexical_declaration block above: the
;; @declaration.function capture must sit on the INNER arrow/fn-expression.
(pair
key: (property_identifier) @declaration.name
value: (arrow_function) @declaration.function)
(pair
key: (property_identifier) @declaration.name
value: (function_expression) @declaration.function)
(pair
key: (string (string_fragment) @declaration.name)
value: (arrow_function) @declaration.function)
(pair
key: (string (string_fragment) @declaration.name)
value: (function_expression) @declaration.function)
;; HOC-wrapped variable declarations: const X = HOC((args) => { ... }).
;; Covers React.forwardRef, memo, useCallback, useMemo, observer,
;; debounce, and any user-defined HOC factory.
(lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (call_expression
arguments: (arguments
(arrow_function) @declaration.function))))
(lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (call_expression
arguments: (arguments
(function_expression) @declaration.function))))
(export_statement
declaration: (lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (call_expression
arguments: (arguments
(arrow_function) @declaration.function)))))
(export_statement
declaration: (lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (call_expression
arguments: (arguments
(function_expression) @declaration.function)))))
(variable_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (call_expression
arguments: (arguments
(arrow_function) @declaration.function))))
(variable_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (call_expression
arguments: (arguments
(function_expression) @declaration.function))))
;; Variable / constant declarations (non-function values).
(lexical_declaration
(variable_declarator
name: (identifier) @declaration.name)) @declaration.const
(export_statement
declaration: (lexical_declaration
(variable_declarator
name: (identifier) @declaration.name))) @declaration.const
(variable_declaration
(variable_declarator
name: (identifier) @declaration.name)) @declaration.variable
;; Imports (ESM) — single anchor per statement; decomposer emits per-specifier markers.
(import_statement) @import.statement
;; Re-exports with a source clause.
(export_statement
source: (string)) @import.statement
;; Dynamic imports: import('./m') — tree-sitter-javascript represents this
;; as call_expression with a named import leaf as the function field,
;; identical to tree-sitter-typescript.
(call_expression
function: (import)) @import.dynamic
;; ── Type bindings (no static annotations in JS; inferred from AST shape) ──
;; Constructor-inferred: const u = new User()
(variable_declarator
name: (identifier) @type-binding.name
value: (new_expression
constructor: (identifier) @type-binding.type)) @type-binding.constructor
;; Qualified constructor: const u = new models.User()
(variable_declarator
name: (identifier) @type-binding.name
value: (new_expression
constructor: (member_expression) @type-binding.type)) @type-binding.constructor
;; Call-result alias: const u = getUser()
(variable_declarator
name: (identifier) @type-binding.name
value: (call_expression
function: (identifier) @type-binding.type)) @type-binding.alias
;; Member-call alias: const u = svc.getUser()
(variable_declarator
name: (identifier) @type-binding.name
value: (call_expression
function: (member_expression) @type-binding.type)) @type-binding.alias
;; Await chain: const u = await getUser() / await svc.getUser()
(variable_declarator
name: (identifier) @type-binding.name
value: (await_expression
(call_expression
function: (identifier) @type-binding.type))) @type-binding.alias
(variable_declarator
name: (identifier) @type-binding.name
value: (await_expression
(call_expression
function: (member_expression) @type-binding.type))) @type-binding.alias
;; Member-access alias: const addr = user.address
(variable_declarator
name: (identifier) @type-binding.name
value: (member_expression) @type-binding.type) @type-binding.member-alias
;; Identifier alias: const alias = user
(variable_declarator
name: (identifier) @type-binding.name
value: (identifier) @type-binding.type) @type-binding.alias
;; Assignment rebind: u = new User() / u = getUser()
(assignment_expression
left: (identifier) @type-binding.name
right: (new_expression
constructor: (identifier) @type-binding.type)) @type-binding.constructor
(assignment_expression
left: (identifier) @type-binding.name
right: (call_expression
function: (identifier) @type-binding.type)) @type-binding.alias
(assignment_expression
left: (identifier) @type-binding.name
right: (identifier) @type-binding.type) @type-binding.alias
;; For-of element: for (const u of users) / for (const u of getUsers())
(for_in_statement
left: (identifier) @type-binding.name
right: (identifier) @type-binding.type) @type-binding.alias
(for_in_statement
left: (identifier) @type-binding.name
right: (call_expression
function: (identifier) @type-binding.type)) @type-binding.alias
(for_in_statement
left: (identifier) @type-binding.name
right: (call_expression
function: (member_expression) @type-binding.type)) @type-binding.alias
(for_in_statement
left: (identifier) @type-binding.name
right: (member_expression
property: (property_identifier) @type-binding.type)) @type-binding.alias
;; ── References ────────────────────────────────────────────────────────────
;; Free calls: fn(args). The dynamic-import filter runs in captures.ts.
(call_expression
function: (identifier) @reference.name) @reference.call.free
;; Awaited free call: await fn<T>(...) re-associated by tree-sitter.
(call_expression
function: (await_expression
(identifier) @reference.name)) @reference.call.free
;; Member calls: obj.method() (includes optional chain).
(call_expression
function: (member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name)) @reference.call.member
;; Awaited member call: await svc.m<T>(...)
(call_expression
function: (await_expression
(member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name))) @reference.call.member
;; Constructor calls: new User() / new ns.User()
(new_expression
constructor: (identifier) @reference.name) @reference.call.constructor
(new_expression
constructor: (member_expression) @reference.call.constructor.qualified) @reference.call.constructor
;; Write access: obj.field = value
(assignment_expression
left: (member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name)) @reference.write.member
(augmented_assignment_expression
left: (member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name)) @reference.write.member
;; Read access: obj.field (in read context; captures.ts filters non-reads).
(member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name) @reference.read.member
`;
/** JSX-only suffix — appended when compiling against the JSX grammar for .jsx files. */
const JSX_QUERY_SUFFIX = `
;; <Foo />
((jsx_self_closing_element
name: (identifier) @reference.name) @reference.call.free
(#match? @reference.name "^[A-Z]"))
;; <Foo> ... </Foo>
((jsx_opening_element
name: (identifier) @reference.name) @reference.call.free
(#match? @reference.name "^[A-Z]"))
;; <Foo.Bar />
(jsx_self_closing_element
name: (member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name)) @reference.call.member
(jsx_opening_element
name: (member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name)) @reference.call.member
`;
let _jsParser: Parser | null = null;
let _jsQuery: Parser.Query | null = null;
let _jsxParser: Parser | null = null;
let _jsxQuery: Parser.Query | null = null;
export function getJsParser(filePath?: string): Parser {
// JSX files use the same JavaScript grammar in tree-sitter-javascript;
// both .js and .jsx parse with the same grammar object. We keep separate
// singletons only to mirror the TypeScript pattern and in case a future
// version of the grammar diverges.
if (filePath !== undefined && isJsxFile(filePath)) {
if (_jsxParser === null) {
_jsxParser = new Parser();
_jsxParser.setLanguage(JS_GRAMMAR);
}
return _jsxParser;
}
if (_jsParser === null) {
_jsParser = new Parser();
_jsParser.setLanguage(JS_GRAMMAR);
}
return _jsParser;
}
export function getJsScopeQuery(filePath?: string): Parser.Query {
if (filePath !== undefined && isJsxFile(filePath)) {
if (_jsxQuery === null) {
_jsxQuery = new Parser.Query(JS_GRAMMAR, JAVASCRIPT_SCOPE_QUERY + JSX_QUERY_SUFFIX);
}
return _jsxQuery;
}
if (_jsQuery === null) {
_jsQuery = new Parser.Query(JS_GRAMMAR, JAVASCRIPT_SCOPE_QUERY);
}
return _jsQuery;
}
/** Validate that a cached Tree was produced by the JS grammar. */
export function jsCachedTreeMatchesGrammar(tree: unknown): boolean {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const lang = (tree as any)?.getLanguage?.();
if (lang === undefined || lang === null) return true;
return lang === JS_GRAMMAR;
}
@@ -0,0 +1,89 @@
/**
* JavaScript `ScopeResolver` registered in `SCOPE_RESOLVERS` and
* consumed by the generic `runScopeResolution` orchestrator
* (RFC #909 Ring 3, issue #928).
*
* Follows the same minimal wiring-only pattern as TypeScript (the third
* migration). Per-hook logic lives in sibling modules:
*
* - `query.ts` — JS scope query + parser/query singletons
* - `captures.ts` — `emitJsScopeCaptures` (JS grammar, CJS, JSDoc)
* - `interpret.ts` — `interpretJsImport` (delegates to TS interpreter)
* - `simple-hooks.ts` — `jsBindingScopeFor`, `jsImportOwningScope`,
* `jsReceiverBinding` (all delegate to TS hooks)
* - `merge-bindings.ts` — `jsMergeBindings` (delegates to TS function)
* - `arity.ts` — `jsArityCompatibility` (delegates to TS function)
* - `import-target.ts` — `makeJsResolveImportTarget` (TS resolver, JS extensions)
*
* See `./index.ts` for the full per-module rationale.
*
* ## Key differences from TypeScript resolver
*
* - `fieldFallbackOnMethodLookup: true` — JavaScript is dynamically typed;
* the field-fallback heuristic is ENABLED (unlike TypeScript, which
* disables it because the type-binding layer is precise).
* - `allowGlobalFreeCallFallback: true` — CJS `require` patterns and
* global helpers (e.g. `process`, `console`) benefit from workspace-
* wide unique-name fallback. TypeScript uses explicit imports.
* - `loadResolutionConfig` is omitted — JavaScript projects don't use
* `tsconfig.json` path aliases in general. `tsconfigPaths: null` is
* threaded through the resolver adapter.
* - `hoistTypeBindingsToModule: true` — JSDoc `@returns {T}` bindings are
* synthesized on the function scope and hoisted, matching TypeScript's
* method return-type hoisting strategy for cross-file chain resolution.
*/
import type { ParsedFile } from 'gitnexus-shared';
import { SupportedLanguages } from 'gitnexus-shared';
import { buildMro, defaultLinearize } from '../../scope-resolution/passes/mro.js';
import { populateClassOwnedMembers } from '../../scope-resolution/scope/walkers.js';
import type { ScopeResolver } from '../../scope-resolution/contract/scope-resolver.js';
import { javascriptProvider } from '../typescript.js';
import { jsMergeBindings } from './merge-bindings.js';
import { jsArityCompatibility } from './arity.js';
import { makeJsResolveImportTarget } from './import-target.js';
const javascriptScopeResolver: ScopeResolver = {
language: SupportedLanguages.JavaScript,
languageProvider: javascriptProvider,
importEdgeReason: 'javascript-scope: import',
resolveImportTarget: makeJsResolveImportTarget(),
// JavaScript LEGB — same tier ordering as TypeScript; no declaration-
// merging across type/value/namespace spaces.
mergeBindings: (existing, incoming) => [...jsMergeBindings([...existing, ...incoming])],
// Adapter: jsArityCompatibility uses (def, callsite); contract is (callsite, def).
arityCompatibility: (callsite, def) => jsArityCompatibility(def, callsite),
buildMro: (graph, parsedFiles, nodeLookup) =>
buildMro(graph, parsedFiles, nodeLookup, defaultLinearize),
populateOwners: (parsed: ParsedFile) => populateClassOwnedMembers(parsed),
// JavaScript `super` keyword: same pattern as TypeScript.
isSuperReceiver: (text) => /^super(\s*\(|\s*\.|\s*\[|\s*$)/.test(text.trim()),
// JavaScript is dynamically typed — enable the field-fallback heuristic
// so member-call receivers without type annotations can still resolve
// through declared class fields (e.g. JSDoc-typed fields).
fieldFallbackOnMethodLookup: true,
// Return-type propagation (across ESM imports) mirrors TypeScript's
// default behavior. JSDoc @returns bindings are hoisted to Module scope
// and propagated to importers via the standard mechanism.
propagatesReturnTypesAcrossImports: true,
// JSDoc @returns bindings are synthesized on the function/method node
// and hoisted to Module scope by `jsBindingScopeFor` (identical to the
// TypeScript `tsBindingScopeFor` `@type-binding.return` branch).
hoistTypeBindingsToModule: true,
// CJS-heavy codebases often have utility functions exported without
// explicit imports at the call site. Workspace-wide unique-name fallback
// recovers these edges.
allowGlobalFreeCallFallback: true,
};
export { javascriptScopeResolver };
@@ -0,0 +1,48 @@
/**
* Simple hooks for the JavaScript scope-resolution provider.
*
* `jsBindingScopeFor` wraps `tsBindingScopeFor` and adds the JS-only
* `@type-binding.class-field` hoisting rule. The other two hooks
* (`jsImportOwningScope`, `jsReceiverBinding`) are identical to their
* TypeScript counterparts and are re-exported directly.
*
* ## Why class-field hoisting lives here (not in `tsBindingScopeFor`)
*
* `@type-binding.class-field` is emitted exclusively by
* `synthesizeConstructorFieldBindings` in `captures.ts`, which is a
* JavaScript-only synthesis pass. TypeScript uses
* `@type-binding.parameter-property` for constructor parameter
* properties instead. Keeping the JS-only rule in the JS hook file
* prevents language-specific logic from leaking into shared TypeScript
* infrastructure (DoD.md §2.2).
*/
import type { CaptureMatch, Scope, ScopeId, ScopeTree } from 'gitnexus-shared';
import { tsBindingScopeFor, walkToScope } from '../typescript/simple-hooks.js';
export {
tsImportOwningScope as jsImportOwningScope,
tsReceiverBinding as jsReceiverBinding,
} from '../typescript/simple-hooks.js';
/**
* Like `tsBindingScopeFor` but additionally hoists
* `@type-binding.class-field` captures to the enclosing Class scope.
*
* `@type-binding.class-field` is anchored inside the constructor body
* (by `synthesizeConstructorFieldBindings`) so that `walkToScope` can
* walk up from the Function (constructor) scope to the Class scope.
* This puts `User.address → Address` in the class's typeBindings so
* compound-receiver resolution finds it when resolving
* `user.address.save()`.
*/
export function jsBindingScopeFor(
decl: CaptureMatch,
innermost: Scope,
tree: ScopeTree,
): ScopeId | null {
if (decl['@type-binding.class-field'] !== undefined) {
return walkToScope(innermost, tree, 'Class');
}
return tsBindingScopeFor(decl, innermost, tree);
}
@@ -29,6 +29,16 @@ import { kotlinMethodConfig } from '../method-extractors/configs/jvm.js';
import { createVariableExtractor } from '../variable-extractors/generic.js';
import { kotlinVariableConfig } from '../variable-extractors/configs/jvm.js';
import { createHeritageExtractor } from '../heritage-extractors/generic.js';
import {
emitKotlinScopeCaptures,
interpretKotlinImport,
interpretKotlinTypeBinding,
kotlinArityCompatibility,
kotlinBindingScopeFor,
kotlinImportOwningScope,
kotlinMergeBindings,
kotlinReceiverBinding,
} from './kotlin/index.js';
/** Check if a Kotlin function_declaration capture is inside a class_body (i.e., a method).
* Kotlin grammar uses function_declaration for both top-level functions and class methods.
@@ -166,4 +176,14 @@ export const kotlinProvider = defineLanguage({
if (isKotlinClassMethod(functionNode)) return 'Method';
return defaultLabel;
},
// ── RFC #909 Ring 3: scope-based resolution hooks ──
emitScopeCaptures: emitKotlinScopeCaptures,
interpretImport: interpretKotlinImport,
interpretTypeBinding: interpretKotlinTypeBinding,
bindingScopeFor: kotlinBindingScopeFor,
importOwningScope: kotlinImportOwningScope,
mergeBindings: (_scope, bindings) => kotlinMergeBindings(bindings),
receiverBinding: kotlinReceiverBinding,
arityCompatibility: kotlinArityCompatibility,
});
@@ -0,0 +1,26 @@
import type { SyntaxNode } from '../../utils/ast-helpers.js';
import { kotlinMethodConfig } from '../../method-extractors/configs/jvm.js';
export interface KotlinArityMetadata {
readonly parameterCount: number | undefined;
readonly requiredParameterCount: number | undefined;
readonly parameterTypes: readonly string[] | undefined;
}
export function computeKotlinArityMetadata(fnNode: SyntaxNode): KotlinArityMetadata {
const params = kotlinMethodConfig.extractParameters?.(fnNode) ?? [];
let hasVararg = false;
const parameterTypes: string[] = [];
for (const param of params) {
if (param.isVariadic) hasVararg = true;
if (param.type !== null) parameterTypes.push(param.type);
}
if (hasVararg) parameterTypes.push('vararg');
const required = params.filter((p) => !p.isOptional && !p.isVariadic).length;
return {
parameterCount: hasVararg ? undefined : params.length,
requiredParameterCount: required,
parameterTypes: parameterTypes.length > 0 ? parameterTypes : undefined,
};
}
@@ -0,0 +1,18 @@
import type { Callsite, SymbolDefinition } from 'gitnexus-shared';
export function kotlinArityCompatibility(
def: SymbolDefinition,
callsite: Callsite,
): 'compatible' | 'unknown' | 'incompatible' {
const min = def.requiredParameterCount;
const max = def.parameterCount;
if (min === undefined && max === undefined) return 'unknown';
const argCount = callsite.arity;
if (!Number.isFinite(argCount) || argCount < 0) return 'unknown';
const hasVararg = def.parameterTypes?.some((t) => t === 'vararg') ?? false;
if (min !== undefined && argCount < min) return 'incompatible';
if (max !== undefined && argCount > max && !hasVararg) return 'incompatible';
return 'compatible';
}
@@ -0,0 +1,19 @@
let hits = 0;
let misses = 0;
export function recordKotlinCacheHit(): void {
hits += 1;
}
export function recordKotlinCacheMiss(): void {
misses += 1;
}
export function getKotlinCaptureCacheStats(): { readonly hits: number; readonly misses: number } {
return { hits, misses };
}
export function resetKotlinCaptureCacheStats(): void {
hits = 0;
misses = 0;
}
@@ -0,0 +1,484 @@
import type { Capture, CaptureMatch } from 'gitnexus-shared';
import {
findNodeAtRange,
nodeToCapture,
syntheticCapture,
type SyntaxNode,
} from '../../utils/ast-helpers.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
import { computeKotlinArityMetadata } from './arity-metadata.js';
import { splitKotlinImportHeader } from './import-decomposer.js';
import { recordKotlinCacheHit, recordKotlinCacheMiss } from './cache-stats.js';
import { normalizeKotlinType } from './interpret.js';
import { synthesizeKotlinReceiverBinding } from './receiver-binding.js';
import { getKotlinParser, getKotlinScopeQuery } from './query.js';
const FUNCTION_DECL_TAGS = ['@declaration.function'] as const;
export function emitKotlinScopeCaptures(
sourceText: string,
_filePath: string,
cachedTree?: unknown,
): readonly CaptureMatch[] {
let tree = cachedTree as ReturnType<ReturnType<typeof getKotlinParser>['parse']> | undefined;
if (tree === undefined) {
tree = parseSourceSafe(getKotlinParser(), sourceText, undefined, {
bufferSize: getTreeSitterBufferSize(sourceText),
});
recordKotlinCacheMiss();
} else {
recordKotlinCacheHit();
}
const out: CaptureMatch[] = [];
const returnTypes = collectKotlinReturnTypeTexts(tree.rootNode);
out.push(...synthesizeKotlinLocalAssignmentBindings(tree.rootNode, returnTypes));
out.push(...synthesizeKotlinLoopBindings(tree.rootNode, returnTypes));
for (const match of getKotlinScopeQuery().matches(tree.rootNode)) {
const grouped: Record<string, Capture> = {};
for (const capture of match.captures) {
const tag = '@' + capture.name;
grouped[tag] = nodeToCapture(tag, capture.node);
}
if (Object.keys(grouped).length === 0) continue;
if (grouped['@import.statement'] !== undefined) {
const importNode = findNodeAtRange(
tree.rootNode,
grouped['@import.statement']!.range,
'import_header',
);
if (importNode !== null) {
const decomposed = splitKotlinImportHeader(importNode);
if (decomposed !== null) {
out.push(decomposed);
continue;
}
}
}
if (
grouped['@reference.call.free'] !== undefined &&
grouped['@reference.receiver'] !== undefined
) {
continue;
}
if (grouped['@reference.read.member'] !== undefined) {
const anchor = grouped['@reference.read.member']!;
const navNode = findNodeAtRange(tree.rootNode, anchor.range, 'navigation_expression');
if (navNode === null || !shouldEmitReadMember(navNode)) continue;
}
if (grouped['@scope.function'] !== undefined) {
out.push(grouped);
const fnNode = findNodeAtRange(
tree.rootNode,
grouped['@scope.function']!.range,
'function_declaration',
);
if (fnNode !== null) {
out.push(...synthesizeKotlinReceiverBinding(fnNode));
}
continue;
}
const declTag = FUNCTION_DECL_TAGS.find((tag) => grouped[tag] !== undefined);
if (declTag !== undefined) {
const fnNode = findNodeAtRange(
tree.rootNode,
grouped[declTag]!.range,
'function_declaration',
);
if (fnNode !== null) {
const arity = computeKotlinArityMetadata(fnNode);
if (arity.parameterCount !== undefined) {
grouped['@declaration.parameter-count'] = syntheticCapture(
'@declaration.parameter-count',
fnNode,
String(arity.parameterCount),
);
}
if (arity.requiredParameterCount !== undefined) {
grouped['@declaration.required-parameter-count'] = syntheticCapture(
'@declaration.required-parameter-count',
fnNode,
String(arity.requiredParameterCount),
);
}
if (arity.parameterTypes !== undefined) {
grouped['@declaration.parameter-types'] = syntheticCapture(
'@declaration.parameter-types',
fnNode,
JSON.stringify(arity.parameterTypes),
);
}
}
}
const callTag = (
['@reference.call.free', '@reference.call.member', '@reference.call.constructor'] as const
).find((tag) => grouped[tag] !== undefined);
if (callTag !== undefined && grouped['@reference.arity'] === undefined) {
const callNode = findNodeAtRange(tree.rootNode, grouped[callTag]!.range, 'call_expression');
if (callNode !== null) {
const args = callArguments(callNode);
grouped['@reference.arity'] = syntheticCapture(
'@reference.arity',
callNode,
String(args.length),
);
grouped['@reference.parameter-types'] = syntheticCapture(
'@reference.parameter-types',
callNode,
JSON.stringify(args.map(inferArgType)),
);
}
}
out.push(grouped);
const extensionFallback = extensionFreeCallFallback(grouped, tree.rootNode);
if (extensionFallback !== null) out.push(extensionFallback);
}
return out;
}
function synthesizeKotlinLoopBindings(
rootNode: SyntaxNode,
returnTypes: ReadonlyMap<string, string>,
): CaptureMatch[] {
const out: CaptureMatch[] = [];
for (const fnNode of descendantsOfType(rootNode, 'function_declaration')) {
const localTypes = collectKotlinLocalTypeTexts(fnNode, returnTypes);
for (const forNode of descendantsOfType(fnNode, 'for_statement')) {
const variable = forNode.namedChildren.find((child) => child.type === 'variable_declaration');
const name = variable?.namedChildren.find((child) => child.type === 'simple_identifier');
if (variable === undefined || name === undefined) continue;
const explicitType = variable.namedChildren.find((child) => isKotlinTypeNode(child));
const iterable = forNode.namedChildren.find(
(child) => child.id !== variable.id && child.type !== 'control_structure_body',
);
const rawType =
explicitType?.text ??
(iterable === undefined
? null
: inferKotlinIterableElementType(iterable, localTypes, returnTypes));
if (rawType === null || rawType.trim() === '') continue;
const anchor =
forNode.namedChildren.find((child) => child.type === 'control_structure_body') ?? forNode;
out.push({
'@type-binding.annotation': nodeToCapture('@type-binding.annotation', anchor),
'@type-binding.name': syntheticCapture('@type-binding.name', name, name.text),
'@type-binding.type': syntheticCapture(
'@type-binding.type',
explicitType ?? iterable ?? name,
normalizeKotlinType(rawType),
),
});
}
}
return out;
}
function synthesizeKotlinLocalAssignmentBindings(
rootNode: SyntaxNode,
returnTypes: ReadonlyMap<string, string>,
): CaptureMatch[] {
const out: CaptureMatch[] = [];
for (const fnNode of descendantsOfType(rootNode, 'function_declaration')) {
const localTypes = new Map<string, string>();
for (const prop of descendantsOfType(fnNode, 'property_declaration')) {
const inferred = inferKotlinPropertyType(prop, localTypes, returnTypes);
if (inferred === null) continue;
localTypes.set(inferred.name.text, inferred.rawType);
if (inferred.synthetic) {
out.push({
'@type-binding.annotation': nodeToCapture('@type-binding.annotation', prop),
'@type-binding.name': syntheticCapture(
'@type-binding.name',
inferred.name,
inferred.name.text,
),
'@type-binding.type': syntheticCapture(
'@type-binding.type',
inferred.source,
normalizeKotlinType(inferred.rawType),
),
});
}
}
}
return out;
}
function collectKotlinLocalTypeTexts(
fnNode: SyntaxNode,
returnTypes: ReadonlyMap<string, string>,
): Map<string, string> {
const out = new Map<string, string>();
for (const node of descendants(fnNode)) {
if (node.type === 'parameter') {
const name = descendantsOfType(node, 'simple_identifier')[0];
const type = node.namedChildren.find((child) => isKotlinTypeNode(child));
if (name !== undefined && type !== undefined) out.set(name.text, type.text);
continue;
}
if (node.type === 'property_declaration') {
const inferred = inferKotlinPropertyType(node, out, returnTypes);
if (inferred !== null) out.set(inferred.name.text, inferred.rawType);
}
}
return out;
}
function collectKotlinReturnTypeTexts(rootNode: SyntaxNode): Map<string, string> {
const out = new Map<string, string>();
for (const fnNode of descendantsOfType(rootNode, 'function_declaration')) {
const name = fnNode.namedChildren.find((child) => child.type === 'simple_identifier');
const paramsIndex = fnNode.namedChildren.findIndex(
(child) => child.type === 'function_value_parameters',
);
const type =
paramsIndex < 0
? undefined
: fnNode.namedChildren.slice(paramsIndex + 1).find((child) => isKotlinTypeNode(child));
if (name !== undefined && type !== undefined) out.set(name.text, type.text);
}
return out;
}
function inferKotlinPropertyType(
prop: SyntaxNode,
localTypes: ReadonlyMap<string, string>,
returnTypes: ReadonlyMap<string, string>,
): { name: SyntaxNode; rawType: string; source: SyntaxNode; synthetic: boolean } | null {
const variable = prop.namedChildren.find((child) => child.type === 'variable_declaration');
const name = variable?.namedChildren.find((child) => child.type === 'simple_identifier');
if (variable === undefined || name === undefined) return null;
const explicitType = variable.namedChildren.find((child) => isKotlinTypeNode(child));
if (explicitType !== undefined) {
return { name, rawType: explicitType.text, source: explicitType, synthetic: false };
}
const value = prop.namedChildren.find(
(child) => child.id !== variable.id && child.type !== 'binding_pattern_kind',
);
if (value?.type === 'simple_identifier') {
const rawType = localTypes.get(value.text);
return rawType === undefined ? null : { name, rawType, source: value, synthetic: true };
}
if (value?.type === 'call_expression') {
const callee = value.namedChildren.find((child) => child.type === 'simple_identifier');
if (callee === undefined) return null;
const rawType =
returnTypes.get(callee.text) ?? (isUppercaseName(callee.text) ? callee.text : null);
if (rawType === null) return null;
return { name, rawType, source: callee, synthetic: true };
}
return null;
}
function inferKotlinIterableElementType(
iterable: SyntaxNode,
localTypes: ReadonlyMap<string, string>,
returnTypes: ReadonlyMap<string, string>,
): string | null {
if (iterable.type === 'simple_identifier') {
const raw = localTypes.get(iterable.text);
return raw === undefined ? null : kotlinContainerElementType(raw, 'values');
}
if (iterable.type === 'navigation_expression') {
const receiver = iterable.namedChildren[0];
const member = iterable.namedChildren
.find((child) => child.type === 'navigation_suffix')
?.namedChildren.find((child) => child.type === 'simple_identifier')?.text;
if (receiver?.type !== 'simple_identifier') return null;
const raw = localTypes.get(receiver.text);
return raw === undefined ? null : kotlinContainerElementType(raw, member ?? 'values');
}
if (iterable.type === 'call_expression') {
const callee = iterable.namedChildren.find((child) => child.type === 'simple_identifier');
if (callee === undefined) return null;
const raw = returnTypes.get(callee.text);
return raw === undefined ? null : kotlinContainerElementType(raw, 'values');
}
return null;
}
function isUppercaseName(text: string): boolean {
return /^[A-Z]/.test(text);
}
function kotlinContainerElementType(rawType: string, member: string): string | null {
const parsed = parseKotlinGeneric(rawType);
if (parsed === null) return normalizeKotlinType(rawType);
const base = parsed.base.split('.').pop() ?? parsed.base;
if (isKotlinMapType(base)) {
if (member === 'keys') return parsed.args[0] ?? null;
return parsed.args[1] ?? null;
}
if (isKotlinIterableType(base)) return parsed.args[0] ?? null;
return normalizeKotlinType(rawType);
}
function parseKotlinGeneric(text: string): { base: string; args: string[] } | null {
const trimmed = text.trim().replace(/\?$/, '');
const open = trimmed.indexOf('<');
const close = trimmed.lastIndexOf('>');
if (open < 0 || close < open) return null;
return {
base: trimmed.slice(0, open).trim(),
args: splitTopLevelKotlinArgs(trimmed.slice(open + 1, close)),
};
}
function splitTopLevelKotlinArgs(text: string): string[] {
const out: string[] = [];
let depth = 0;
let start = 0;
for (let i = 0; i < text.length; i++) {
const ch = text[i];
if (ch === '<') depth++;
else if (ch === '>') depth--;
else if (ch === ',' && depth === 0) {
out.push(text.slice(start, i).trim());
start = i + 1;
}
}
out.push(text.slice(start).trim());
return out.filter((arg) => arg.length > 0);
}
function isKotlinMapType(base: string): boolean {
return ['Map', 'MutableMap', 'HashMap', 'LinkedHashMap'].includes(base);
}
function isKotlinIterableType(base: string): boolean {
return [
'List',
'MutableList',
'ArrayList',
'Set',
'MutableSet',
'Collection',
'Iterable',
'Sequence',
'Array',
].includes(base);
}
function isKotlinTypeNode(node: SyntaxNode): boolean {
return (
node.type === 'user_type' || node.type === 'nullable_type' || node.type === 'function_type'
);
}
function descendantsOfType(node: SyntaxNode, type: string): SyntaxNode[] {
return descendants(node).filter((child) => child.type === type);
}
function descendants(node: SyntaxNode): SyntaxNode[] {
const out: SyntaxNode[] = [];
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child === null) continue;
out.push(child, ...descendants(child));
}
return out;
}
function shouldEmitReadMember(navNode: SyntaxNode): boolean {
const parent = navNode.parent;
if (parent === null) return true;
if (parent.type === 'call_expression') return false;
if (parent.type === 'directly_assignable_expression') return false;
return true;
}
function callArguments(callNode: SyntaxNode): SyntaxNode[] {
const suffix = callNode.namedChildren.find((child) => child.type === 'call_suffix');
if (suffix === undefined) return [];
const valueArgs = suffix?.namedChildren.find((child) => child.type === 'value_arguments');
const args = valueArgs?.namedChildren.filter((child) => child.type === 'value_argument') ?? [];
const trailingLambdas = suffix.namedChildren.filter((child) => child.type === 'annotated_lambda');
return [...args, ...trailingLambdas];
}
function inferArgType(argNode: SyntaxNode): string {
const value = argNode.namedChild(0) ?? argNode;
switch (value.type) {
case 'integer_literal':
case 'long_literal':
return 'Int';
case 'real_literal':
return 'Double';
case 'string_literal':
case 'line_string_literal':
case 'multi_line_string_literal':
return 'String';
case 'character_literal':
return 'Char';
case 'boolean_literal':
return 'Boolean';
case 'call_expression': {
const first = value.namedChild(0);
return first?.type === 'simple_identifier' ? first.text : '';
}
default:
return '';
}
}
function extensionFreeCallFallback(
grouped: Record<string, Capture>,
rootNode: SyntaxNode,
): CaptureMatch | null {
const member = grouped['@reference.call.member'];
const receiver = grouped['@reference.receiver'];
const name = grouped['@reference.name'];
if (member === undefined || receiver === undefined || name === undefined) return null;
const callNode = findNodeAtRange(rootNode, member.range, 'call_expression');
if (callNode === null) return null;
const receiverNode = findNodeAtRange(rootNode, receiver.range);
if (receiverNode === null || !isLiteralReceiver(receiverNode)) return null;
const out: Record<string, Capture> = {
'@reference.call.free': syntheticCapture('@reference.call.free', callNode, callNode.text),
'@reference.name': syntheticCapture('@reference.name', callNode, name.text),
};
if (grouped['@reference.arity'] !== undefined)
out['@reference.arity'] = grouped['@reference.arity'];
if (grouped['@reference.parameter-types'] !== undefined) {
out['@reference.parameter-types'] = grouped['@reference.parameter-types'];
}
return out;
}
function isLiteralReceiver(node: SyntaxNode): boolean {
return [
'integer_literal',
'long_literal',
'real_literal',
'string_literal',
'line_string_literal',
'multi_line_string_literal',
'character_literal',
'boolean_literal',
].includes(node.type);
}
@@ -0,0 +1,49 @@
import type { Capture, CaptureMatch } from 'gitnexus-shared';
import { nodeToCapture, syntheticCapture, type SyntaxNode } from '../../utils/ast-helpers.js';
type KotlinImportKind = 'named' | 'alias' | 'wildcard';
interface KotlinImportSpec {
readonly kind: KotlinImportKind;
readonly source: string;
readonly name: string;
readonly alias?: string;
readonly atNode: SyntaxNode;
}
export function splitKotlinImportHeader(importNode: SyntaxNode): CaptureMatch | null {
if (importNode.type !== 'import_header') return null;
const spec = parseKotlinImport(importNode);
if (spec === null) return null;
const out: Record<string, Capture> = {
'@import.statement': nodeToCapture('@import.statement', importNode),
'@import.kind': syntheticCapture('@import.kind', spec.atNode, spec.kind),
'@import.source': syntheticCapture('@import.source', spec.atNode, spec.source),
'@import.name': syntheticCapture('@import.name', spec.atNode, spec.name),
};
if (spec.alias !== undefined) {
out['@import.alias'] = syntheticCapture('@import.alias', spec.atNode, spec.alias);
}
return out;
}
function parseKotlinImport(node: SyntaxNode): KotlinImportSpec | null {
const identifier = node.namedChildren.find((child) => child.type === 'identifier');
if (identifier === undefined) return null;
const source = identifier.text.trim();
if (source.length === 0) return null;
const hasWildcard = node.namedChildren.some((child) => child.type === 'wildcard_import');
if (hasWildcard) {
return { kind: 'wildcard', source, name: '*', atNode: node };
}
const aliasNode = node.namedChildren.find((child) => child.type === 'import_alias');
const alias = aliasNode?.namedChildren.find((child) => child.type === 'type_identifier')?.text;
const importedName = source.split('.').pop() ?? source;
if (alias !== undefined && alias.length > 0) {
return { kind: 'alias', source, name: importedName, alias, atNode: node };
}
return { kind: 'named', source, name: importedName, atNode: node };
}
@@ -0,0 +1,76 @@
import type { ParsedImport, WorkspaceIndex } from 'gitnexus-shared';
export interface KotlinResolveContext {
readonly fromFile: string;
readonly allFilePaths: ReadonlySet<string>;
}
export function resolveKotlinImportTarget(
parsedImport: ParsedImport,
workspaceIndex: WorkspaceIndex,
): string | null {
const ctx = workspaceIndex as KotlinResolveContext | undefined;
if (
ctx === undefined ||
typeof (ctx as { fromFile?: unknown }).fromFile !== 'string' ||
!((ctx as { allFilePaths?: unknown }).allFilePaths instanceof Set)
) {
return null;
}
if (parsedImport.kind === 'dynamic-unresolved') return null;
if (parsedImport.targetRaw === null || parsedImport.targetRaw === '') return null;
const target = parsedImport.targetRaw.endsWith('.*')
? parsedImport.targetRaw.slice(0, -2)
: parsedImport.targetRaw;
const pathLike = target.replace(/\./g, '/');
return (
findKotlinFile(ctx.allFilePaths, pathLike) ??
findKotlinFile(ctx.allFilePaths, pathLike.split('/').slice(0, -1).join('/')) ??
findByProgressivePrefixStrip(ctx.allFilePaths, pathLike)
);
}
function findKotlinFile(allFilePaths: ReadonlySet<string>, pathLike: string): string | null {
if (pathLike === '') return null;
const extensions = ['.kt', '.kts'];
const suffix = `/${pathLike}`;
const dirPrefix = `${pathLike}/`;
const suffixDirPrefix = `/${dirPrefix}`;
let suffixFile: string | null = null;
let directoryChild: string | null = null;
for (const raw of allFilePaths) {
const file = raw.replace(/\\/g, '/');
if (!extensions.some((ext) => file.endsWith(ext))) continue;
for (const ext of extensions) {
if (file === `${pathLike}${ext}`) return raw;
if (suffixFile === null && file.endsWith(`${suffix}${ext}`)) suffixFile = raw;
}
if (directoryChild === null) {
const atRoot = file.startsWith(dirPrefix);
const atNested = file.includes(suffixDirPrefix);
if (atRoot || atNested) {
const idx = atRoot ? 0 : file.indexOf(suffixDirPrefix) + 1;
const after = file.slice(idx + dirPrefix.length);
if (after.length > 0 && !after.includes('/')) directoryChild = raw;
}
}
}
return suffixFile ?? directoryChild;
}
function findByProgressivePrefixStrip(
allFilePaths: ReadonlySet<string>,
pathLike: string,
): string | null {
const segments = pathLike.split('/').filter(Boolean);
for (let skip = 1; skip < segments.length; skip++) {
const found = findKotlinFile(allFilePaths, segments.slice(skip).join('/'));
if (found !== null) return found;
}
return null;
}
@@ -0,0 +1,12 @@
export { emitKotlinScopeCaptures } from './captures.js';
export { getKotlinCaptureCacheStats, resetKotlinCaptureCacheStats } from './cache-stats.js';
export { interpretKotlinImport, interpretKotlinTypeBinding } from './interpret.js';
export { kotlinArityCompatibility } from './arity.js';
export { resolveKotlinImportTarget, type KotlinResolveContext } from './import-target.js';
export { kotlinMergeBindings } from './merge-bindings.js';
export { populateKotlinOwners } from './owners.js';
export {
kotlinBindingScopeFor,
kotlinImportOwningScope,
kotlinReceiverBinding,
} from './simple-hooks.js';
@@ -0,0 +1,71 @@
import type { CaptureMatch, ParsedImport, ParsedTypeBinding, TypeRef } from 'gitnexus-shared';
export function interpretKotlinImport(captures: CaptureMatch): ParsedImport | null {
const kind = captures['@import.kind']?.text;
const source = captures['@import.source']?.text;
const name = captures['@import.name']?.text;
if (kind === undefined || source === undefined) return null;
switch (kind) {
case 'named':
return {
kind: 'named',
localName: name ?? source.split('.').pop() ?? source,
importedName: name ?? source.split('.').pop() ?? source,
targetRaw: source,
};
case 'alias': {
const alias = captures['@import.alias']?.text;
if (alias === undefined || name === undefined) return null;
return {
kind: 'alias',
localName: alias,
importedName: name,
alias,
targetRaw: source,
};
}
case 'wildcard':
return { kind: 'wildcard', targetRaw: source.endsWith('.*') ? source : `${source}.*` };
default:
return null;
}
}
export function interpretKotlinTypeBinding(captures: CaptureMatch): ParsedTypeBinding | null {
const nameCap = captures['@type-binding.name'];
const typeCap = captures['@type-binding.type'];
if (nameCap === undefined || typeCap === undefined) return null;
let source: TypeRef['source'] = 'annotation';
if (captures['@type-binding.self'] !== undefined) source = 'self';
else if (captures['@type-binding.parameter'] !== undefined) source = 'parameter-annotation';
else if (captures['@type-binding.return'] !== undefined) source = 'return-annotation';
else if (captures['@type-binding.constructor'] !== undefined) source = 'constructor-inferred';
return {
boundName: nameCap.text,
rawTypeName: normalizeKotlinType(typeCap.text),
source,
};
}
export function normalizeKotlinType(text: string): string {
let out = text.trim();
while (out.endsWith('?')) out = out.slice(0, -1).trim();
const lastDot = out.lastIndexOf('.');
if (lastDot >= 0) out = out.slice(lastDot + 1);
const collection = out.match(
/^(?:List|MutableList|ArrayList|Set|MutableSet|Collection|Iterable|Sequence|Array)<([^,<>]+)>$/,
);
if (collection !== null) return normalizeKotlinType(collection[1]!);
const map = out.match(/^(?:Map|MutableMap|HashMap|LinkedHashMap)<[^,<>]+,\s*([^,<>]+)>$/);
if (map !== null) return normalizeKotlinType(map[1]!);
const erased = out.match(/^([A-Za-z_][A-Za-z0-9_]*)<.+>$/s);
if (erased !== null) return erased[1]!;
return out;
}
@@ -0,0 +1,26 @@
import type { BindingRef } from 'gitnexus-shared';
function tierOf(binding: BindingRef): number {
switch (binding.origin) {
case 'local':
return 0;
case 'import':
case 'namespace':
case 'reexport':
return 1;
case 'wildcard':
return 2;
default:
return 3;
}
}
export function kotlinMergeBindings(bindings: readonly BindingRef[]): readonly BindingRef[] {
if (bindings.length === 0) return bindings;
const best = Math.min(...bindings.map(tierOf));
const seen = new Map<string, BindingRef>();
for (const binding of bindings) {
if (tierOf(binding) === best) seen.set(binding.def.nodeId, binding);
}
return [...seen.values()];
}
@@ -0,0 +1,50 @@
import type { ParsedFile, ScopeId, SymbolDefinition } from 'gitnexus-shared';
import { isClassLike, populateClassOwnedMembers } from '../../scope-resolution/scope/walkers.js';
export function populateKotlinOwners(parsed: ParsedFile): void {
populateClassOwnedMembers(parsed);
populateCompanionMembersOnEnclosingClass(parsed);
}
function populateCompanionMembersOnEnclosingClass(parsed: ParsedFile): void {
const scopesById = new Map<ScopeId, ParsedFile['scopes'][number]>();
for (const scope of parsed.scopes) scopesById.set(scope.id, scope);
for (const scope of parsed.scopes) {
if (scope.kind !== 'Function' || scope.parent === null) continue;
const parent = scopesById.get(scope.parent);
if (parent === undefined || parent.kind !== 'Class') continue;
if (parent.ownedDefs.some((def) => isClassLike(def.type))) continue;
const enclosing = findEnclosingClassWithDef(parent.parent, scopesById);
if (enclosing === undefined) continue;
for (const def of scope.ownedDefs) {
if (def.ownerId !== undefined) continue;
(def as { ownerId?: string }).ownerId = enclosing.nodeId;
qualify(def, enclosing);
}
}
}
function findEnclosingClassWithDef(
start: ScopeId | null,
scopesById: ReadonlyMap<ScopeId, ParsedFile['scopes'][number]>,
): SymbolDefinition | undefined {
let current = start;
while (current !== null) {
const scope = scopesById.get(current);
if (scope === undefined) return undefined;
if (scope.kind === 'Class') {
const classDef = scope.ownedDefs.find((def) => isClassLike(def.type));
if (classDef !== undefined) return classDef;
}
current = scope.parent;
}
return undefined;
}
function qualify(def: SymbolDefinition, owner: SymbolDefinition): void {
if (def.qualifiedName === undefined || def.qualifiedName.includes('.')) return;
if (owner.qualifiedName === undefined || owner.qualifiedName.length === 0) return;
(def as { qualifiedName: string }).qualifiedName = `${owner.qualifiedName}.${def.qualifiedName}`;
}
@@ -0,0 +1,116 @@
import Parser from 'tree-sitter';
import Kotlin from 'tree-sitter-kotlin';
const KOTLIN_SCOPE_QUERY = `
;; Scopes
(source_file) @scope.module
(class_declaration) @scope.class
(object_declaration) @scope.class
(companion_object) @scope.class
(function_declaration) @scope.function
;; Declarations — types
(class_declaration
"interface"
(type_identifier) @declaration.name) @declaration.interface
(class_declaration
"class"
(type_identifier) @declaration.name) @declaration.class
(object_declaration
(type_identifier) @declaration.name) @declaration.class
(companion_object
(type_identifier) @declaration.name) @declaration.class
(type_alias
(type_identifier) @declaration.name) @declaration.type_alias
;; Declarations — functions / methods / properties
(function_declaration
(simple_identifier) @declaration.name) @declaration.function
(property_declaration
(variable_declaration
(simple_identifier) @declaration.name)) @declaration.property
(class_parameter
(binding_pattern_kind)
(simple_identifier) @declaration.name) @declaration.property
;; Imports
(import_header) @import.statement
;; Type bindings — parameters
(parameter
(simple_identifier) @type-binding.name
[(user_type) (nullable_type) (function_type)] @type-binding.type) @type-binding.parameter
;; Type bindings — property / local annotations
(property_declaration
(variable_declaration
(simple_identifier) @type-binding.name
[(user_type) (nullable_type) (function_type)] @type-binding.type)) @type-binding.annotation
(class_parameter
(binding_pattern_kind)
(simple_identifier) @type-binding.name
[(user_type) (nullable_type) (function_type)] @type-binding.type) @type-binding.annotation
;; Type bindings — constructor-inferred val user = User(...)
(property_declaration
(variable_declaration
(simple_identifier) @type-binding.name)
(call_expression
(simple_identifier) @type-binding.type)) @type-binding.constructor
;; Type bindings — return annotations after function parameters
(function_declaration
(simple_identifier) @type-binding.name
(function_value_parameters)
[(user_type) (nullable_type) (function_type)] @type-binding.type) @type-binding.return
;; References — direct calls / constructor syntax
(call_expression
(simple_identifier) @reference.name) @reference.call.free
;; References — member calls: obj.method()
(call_expression
(navigation_expression
(_) @reference.receiver
(navigation_suffix
(simple_identifier) @reference.name))) @reference.call.member
;; References — property writes
(assignment
(directly_assignable_expression
(_) @reference.receiver
(navigation_suffix
(simple_identifier) @reference.name))
(_)) @reference.write.member
;; References — property reads
(navigation_expression
(_) @reference.receiver
(navigation_suffix
(simple_identifier) @reference.name)) @reference.read.member
`;
let parser: Parser | null = null;
let query: Parser.Query | null = null;
export function getKotlinParser(): Parser {
if (parser === null) {
parser = new Parser();
parser.setLanguage(Kotlin as Parameters<Parser['setLanguage']>[0]);
}
return parser;
}
export function getKotlinScopeQuery(): Parser.Query {
if (query === null) {
query = new Parser.Query(Kotlin as Parameters<Parser['setLanguage']>[0], KOTLIN_SCOPE_QUERY);
}
return query;
}
@@ -0,0 +1,106 @@
import type { Capture, CaptureMatch } from 'gitnexus-shared';
import { nodeToCapture, syntheticCapture, type SyntaxNode } from '../../utils/ast-helpers.js';
import { normalizeKotlinType } from './interpret.js';
const TYPE_DECL_NODE_TYPES = new Set([
'class_declaration',
'object_declaration',
'companion_object',
]);
export function synthesizeKotlinReceiverBinding(fnNode: SyntaxNode): CaptureMatch[] {
if (fnNode.type !== 'function_declaration') return [];
const anchorNode = findFunctionBody(fnNode);
if (anchorNode === null) return [];
const extensionReceiver = extensionReceiverType(fnNode);
if (extensionReceiver !== null) {
return [buildReceiverMatch(anchorNode, 'this', extensionReceiver)];
}
const enclosingType = findEnclosingTypeDeclaration(fnNode);
if (enclosingType === null) return [];
const enclosingName = typeDeclarationName(enclosingType);
if (enclosingName === null) return [];
const out = [buildReceiverMatch(anchorNode, 'this', enclosingName)];
const superName = firstSuperclassText(enclosingType);
if (superName !== null) out.push(buildReceiverMatch(anchorNode, 'super', superName));
return out;
}
function findFunctionBody(fnNode: SyntaxNode): SyntaxNode | null {
for (let i = 0; i < fnNode.namedChildCount; i++) {
const child = fnNode.namedChild(i);
if (child?.type === 'function_body') return child;
}
return fnNode;
}
function extensionReceiverType(fnNode: SyntaxNode): string | null {
for (let i = 0; i < fnNode.namedChildCount; i++) {
const child = fnNode.namedChild(i);
if (child === null) continue;
if (child.type === 'simple_identifier') return null;
if (child.type === 'user_type' || child.type === 'nullable_type') {
return normalizeKotlinType(child.text);
}
}
return null;
}
function findEnclosingTypeDeclaration(node: SyntaxNode): SyntaxNode | null {
let current = node.parent;
while (current !== null) {
if (TYPE_DECL_NODE_TYPES.has(current.type)) return current;
current = current.parent;
}
return null;
}
function typeDeclarationName(typeNode: SyntaxNode): string | null {
if (typeNode.type === 'companion_object') {
return (
typeNode.namedChildren.find((child) => child.type === 'type_identifier')?.text ??
enclosingNonCompanionTypeName(typeNode) ??
'Companion'
);
}
return typeNode.namedChildren.find((child) => child.type === 'type_identifier')?.text ?? null;
}
function enclosingNonCompanionTypeName(node: SyntaxNode): string | null {
let current = node.parent;
while (current !== null) {
if (current.type === 'class_declaration' || current.type === 'object_declaration') {
return current.namedChildren.find((child) => child.type === 'type_identifier')?.text ?? null;
}
current = current.parent;
}
return null;
}
function firstSuperclassText(typeNode: SyntaxNode): string | null {
if (typeNode.type !== 'class_declaration') return null;
for (const child of typeNode.namedChildren) {
if (child.type !== 'delegation_specifier') continue;
const ctor = child.namedChildren.find((n) => n.type === 'constructor_invocation');
const userType =
ctor?.namedChildren.find((n) => n.type === 'user_type') ??
child.namedChildren.find((n) => n.type === 'user_type');
const name = userType?.namedChildren.find((n) => n.type === 'type_identifier')?.text;
if (name !== undefined) return normalizeKotlinType(name);
}
return null;
}
function buildReceiverMatch(anchorNode: SyntaxNode, name: string, typeText: string): CaptureMatch {
const out: Record<string, Capture> = {
'@type-binding.self': nodeToCapture('@type-binding.self', anchorNode),
'@type-binding.name': syntheticCapture('@type-binding.name', anchorNode, name),
'@type-binding.type': syntheticCapture('@type-binding.type', anchorNode, typeText),
};
return out;
}
@@ -0,0 +1,56 @@
import { SupportedLanguages, type ParsedFile } from 'gitnexus-shared';
import { buildMro, defaultLinearize } from '../../scope-resolution/passes/mro.js';
import type { ScopeResolver } from '../../scope-resolution/contract/scope-resolver.js';
import { kotlinProvider } from '../kotlin.js';
import {
kotlinArityCompatibility,
kotlinMergeBindings,
populateKotlinOwners,
resolveKotlinImportTarget,
type KotlinResolveContext,
} from './index.js';
/**
* Kotlin scope resolver for RFC #909 Ring 3.
*
* Kotlin is intentionally registered but not yet listed in
* `MIGRATED_LANGUAGES`, matching the Java migration pattern from #1482:
* the resolver can run in shadow/forced mode, while production default
* stays on the legacy DAG until registry-primary parity reaches the
* RFC threshold. Forced mode currently passes 154/175 fixtures (88%),
* including core import, receiver, companion, default-param, vararg,
* constructor, local assignment-chain, and collection-iteration fixtures.
* Remaining gaps are advanced TypeEnv behaviors such as smart casts,
* cross-file iterable return propagation, method-chain fixpoint cases,
* overload target-id selection, virtual dispatch, and interface default
* method dispatch.
*/
export const kotlinScopeResolver: ScopeResolver = {
language: SupportedLanguages.Kotlin,
languageProvider: kotlinProvider,
importEdgeReason: 'kotlin-scope: import',
resolveImportTarget: (targetRaw, fromFile, allFilePaths) => {
const ws: KotlinResolveContext = { fromFile, allFilePaths };
return resolveKotlinImportTarget(
{ kind: 'named', localName: '_', importedName: '_', targetRaw },
ws,
);
},
mergeBindings: (existing, incoming) => [...kotlinMergeBindings([...existing, ...incoming])],
arityCompatibility: (callsite, def) => kotlinArityCompatibility(def, callsite),
buildMro: (graph, parsedFiles, nodeLookup) =>
buildMro(graph, parsedFiles, nodeLookup, defaultLinearize),
populateOwners: (parsed: ParsedFile) => populateKotlinOwners(parsed),
isSuperReceiver: (text) => text.trim() === 'super',
fieldFallbackOnMethodLookup: false,
propagatesReturnTypesAcrossImports: true,
collapseMemberCallsByCallerTarget: false,
hoistTypeBindingsToModule: true,
};
@@ -0,0 +1,36 @@
import type {
CaptureMatch,
ParsedImport,
Scope,
ScopeId,
ScopeTree,
TypeRef,
} from 'gitnexus-shared';
export function kotlinBindingScopeFor(
decl: CaptureMatch,
innermost: Scope,
tree: ScopeTree,
): ScopeId | null {
if (decl['@type-binding.return'] === undefined) return null;
let current: Scope | undefined = innermost;
while (current !== undefined && current.kind !== 'Module') {
if (current.parent === null) break;
current = tree.getScope(current.parent);
}
return current?.kind === 'Module' ? current.id : null;
}
export function kotlinImportOwningScope(
_imp: ParsedImport,
_innermost: Scope,
_tree: ScopeTree,
): ScopeId | null {
return null;
}
export function kotlinReceiverBinding(functionScope: Scope): TypeRef | null {
if (functionScope.kind !== 'Function') return null;
return functionScope.typeBindings.get('this') ?? functionScope.typeBindings.get('super') ?? null;
}
@@ -56,6 +56,16 @@ import {
typescriptArityCompatibility,
resolveTsImportTarget,
} from './typescript/index.js';
import {
emitJsScopeCaptures,
interpretJsImport,
interpretJsTypeBinding,
jsBindingScopeFor,
jsImportOwningScope,
jsReceiverBinding,
jsMergeBindings,
jsArityCompatibility,
} from './javascript/index.js';
/**
* TypeScript/JavaScript: arrow_function and function_expression are
@@ -359,4 +369,19 @@ export const javascriptProvider = defineLanguage({
classExtractor: createClassExtractor(javascriptClassConfig),
heritageExtractor: createHeritageExtractor(SupportedLanguages.JavaScript),
builtInNames: BUILT_INS,
// ── RFC #909 Ring 3: scope-based resolution hooks (RFC §5) ──────────
// JavaScript is the fourth migration after Python, C#, and TypeScript.
// Hooks are thin wrappers over the TypeScript implementations where
// semantics are identical; JS-specific additions (CJS require(),
// JSDoc type bindings) live in ./javascript/captures.ts.
// See ./javascript/index.ts for the full per-module rationale.
emitScopeCaptures: emitJsScopeCaptures,
interpretImport: interpretJsImport,
interpretTypeBinding: interpretJsTypeBinding,
bindingScopeFor: jsBindingScopeFor,
importOwningScope: jsImportOwningScope,
mergeBindings: (_scope, bindings) => jsMergeBindings(bindings),
receiverBinding: jsReceiverBinding,
arityCompatibility: jsArityCompatibility,
});
@@ -64,7 +64,7 @@ const CALL_TAGS = [
'@reference.call.constructor',
] as const;
function pickFirstDefined(grouped: CaptureMatch, tags: readonly string[]): Capture | undefined {
function pickFirstCapture(grouped: CaptureMatch, tags: readonly string[]): Capture | undefined {
for (const tag of tags) {
const cap = grouped[tag];
if (cap !== undefined) return cap;
@@ -72,6 +72,17 @@ function pickFirstDefined(grouped: CaptureMatch, tags: readonly string[]): Captu
return undefined;
}
function pickFirstNode(
grouped: Record<string, SyntaxNode | undefined>,
tags: readonly string[],
): SyntaxNode | undefined {
for (const tag of tags) {
const node = grouped[tag];
if (node !== undefined) return node;
}
return undefined;
}
/**
* Drop `@reference.read.member` matches whose underlying `member_expression`
* is NOT actually a read context:
@@ -113,6 +124,34 @@ function shouldEmitReadMember(memberNode: SyntaxNode): boolean {
}
}
/** Walks the parent chain from `node` (inclusive), returning the first node
* whose type matches, or null. Faster than `findNodeAtRange` when the caller
* already holds the anchor node — avoids re-scanning the tree from the root. */
function findSelfOrAncestorOfType(node: SyntaxNode | undefined, type: string): SyntaxNode | null {
if (node === undefined) return null;
let current: SyntaxNode | null = node;
while (current !== null) {
if (current.type === type) return current;
current = current.parent;
}
return null;
}
/** Walks the parent chain from `node` (inclusive), returning the first node
* whose type is in the set, or null. Plural form of {@link findSelfOrAncestorOfType}. */
function findSelfOrAncestorOfTypes(
node: SyntaxNode | undefined,
types: readonly string[],
): SyntaxNode | null {
if (node === undefined) return null;
let current: SyntaxNode | null = node;
while (current !== null) {
if (types.includes(current.type)) return current;
current = current.parent;
}
return null;
}
export function emitTsScopeCaptures(
sourceText: string,
filePath: string,
@@ -151,9 +190,11 @@ export function emitTsScopeCaptures(
// `@`; we put it back so the central extractor's prefix lookups
// (`@scope.`, `@declaration.`, …) work.
const grouped: Record<string, Capture> = {};
const groupedNodes: Record<string, SyntaxNode> = {};
for (const c of m.captures) {
const tag = '@' + c.name;
grouped[tag] = nodeToCapture(tag, c.node);
groupedNodes[tag] = c.node;
}
if (Object.keys(grouped).length === 0) continue;
@@ -165,6 +206,10 @@ export function emitTsScopeCaptures(
if (grouped['@import.statement'] !== undefined) {
const stmtCapture = grouped['@import.statement'];
const stmtNode =
findSelfOrAncestorOfTypes(groupedNodes['@import.statement'], [
'import_statement',
'export_statement',
]) ??
findNodeAtRange(tree.rootNode, stmtCapture.range, 'import_statement') ??
findNodeAtRange(tree.rootNode, stmtCapture.range, 'export_statement');
if (stmtNode !== null) {
@@ -183,7 +228,9 @@ export function emitTsScopeCaptures(
// `splitDynamicImport` branch consumes.
if (grouped['@import.dynamic'] !== undefined) {
const dynCapture = grouped['@import.dynamic'];
const callNode = findNodeAtRange(tree.rootNode, dynCapture.range, 'call_expression');
const callNode =
findSelfOrAncestorOfType(groupedNodes['@import.dynamic'], 'call_expression') ??
findNodeAtRange(tree.rootNode, dynCapture.range, 'call_expression');
if (callNode !== null) {
const decomposed = splitImportStatement(callNode);
for (const d of decomposed) out.push(d);
@@ -197,7 +244,9 @@ export function emitTsScopeCaptures(
// we rely on this emit-side filter so the query stays simple.
if (grouped['@reference.read.member'] !== undefined) {
const anchor = grouped['@reference.read.member'];
const memberNode = findNodeAtRange(tree.rootNode, anchor.range, 'member_expression');
const memberNode =
findSelfOrAncestorOfType(groupedNodes['@reference.read.member'], 'member_expression') ??
findNodeAtRange(tree.rootNode, anchor.range, 'member_expression');
if (memberNode === null || !shouldEmitReadMember(memberNode)) {
continue;
}
@@ -208,9 +257,10 @@ export function emitTsScopeCaptures(
// overloads — TypeScript supports overload signatures via
// function_signature, so `parameterTypes` is populated when
// available.
const declAnchor = pickFirstDefined(grouped, FUNCTION_DECL_TAGS);
const declAnchor = pickFirstCapture(grouped, FUNCTION_DECL_TAGS);
const declAnchorNode = pickFirstNode(groupedNodes, FUNCTION_DECL_TAGS);
if (declAnchor !== undefined) {
const fnNode = findFunctionNode(tree.rootNode, declAnchor.range);
const fnNode = findFunctionNode(tree.rootNode, declAnchor.range, declAnchorNode);
if (fnNode !== null) {
const arity = computeTsArityMetadata(fnNode);
if (arity.parameterCount !== undefined) {
@@ -255,9 +305,11 @@ export function emitTsScopeCaptures(
// calls to disambiguate by props-arity, a JSX-aware arity
// synthesizer would need to count `jsx_attribute` children of the
// opening tag instead of `arguments`.
const callAnchor = pickFirstDefined(grouped, CALL_TAGS);
const callAnchor = pickFirstCapture(grouped, CALL_TAGS);
const callAnchorNode = pickFirstNode(groupedNodes, CALL_TAGS);
if (callAnchor !== undefined && grouped['@reference.arity'] === undefined) {
const callNode =
findSelfOrAncestorOfTypes(callAnchorNode, ['call_expression', 'new_expression']) ??
findNodeAtRange(tree.rootNode, callAnchor.range, 'call_expression') ??
findNodeAtRange(tree.rootNode, callAnchor.range, 'new_expression');
if (callNode !== null) {
@@ -293,7 +345,11 @@ export function emitTsScopeCaptures(
// lookup instead of synthesis — covered by `tsReceiverBinding`.
const scopeFnAnchor = grouped['@scope.function'];
if (scopeFnAnchor !== undefined) {
const fnNode = findFunctionNode(tree.rootNode, scopeFnAnchor.range);
const fnNode = findFunctionNode(
tree.rootNode,
scopeFnAnchor.range,
groupedNodes['@scope.function'],
);
if (fnNode !== null) {
const synth = synthesizeTsReceiverBinding(fnNode);
if (synth !== null) out.push(synth);
@@ -518,7 +574,13 @@ function inferArgType(argNode: SyntaxNode): string {
* The `@scope.function` anchor range covers the whole node, but the
* tag alone doesn't identify which node type among the many TS
* function-likes. */
function findFunctionNode(rootNode: SyntaxNode, range: Capture['range']): SyntaxNode | null {
function findFunctionNode(
rootNode: SyntaxNode,
range: Capture['range'],
anchorNode?: SyntaxNode,
): SyntaxNode | null {
const fromAnchor = findSelfOrAncestorOfTypes(anchorNode, FUNCTION_NODE_TYPES);
if (fromAnchor !== null) return fromAnchor;
for (const nodeType of FUNCTION_NODE_TYPES) {
const n = findNodeAtRange(rootNode, range, nodeType);
if (n !== null) return n;
@@ -75,8 +75,11 @@ export function tsBindingScopeFor(
* any of `kinds`. Returns the matching scope's id or `null` when no
* ancestor matches (e.g., a return type binding emitted outside any
* Module scope — shouldn't happen in well-formed input).
*
* Exported so language-specific hook wrappers (e.g. `jsBindingScopeFor`)
* can reuse it without duplicating the traversal logic.
*/
function walkToScope(
export function walkToScope(
from: Scope,
tree: ScopeTree,
...kinds: readonly Scope['kind'][]
@@ -2,18 +2,35 @@
* Field Registry
*
* Owner-scoped field/property index extracted from SymbolTable.
* Stores Property symbols keyed by `ownerNodeId\0fieldName` for O(1) lookup.
* Stores Property / Variable / Const / Static symbols keyed by
* `ownerNodeId\0fieldName` for O(1) lookup. Supports multiple defs
* under the same (owner, name) — e.g. legacy Property plus a
* scope-resolution Variable reconciliation entry.
*/
import type { SymbolDefinition } from 'gitnexus-shared';
const EMPTY: readonly SymbolDefinition[] = Object.freeze([]);
// ---------------------------------------------------------------------------
// Public read-only interface
// ---------------------------------------------------------------------------
export interface FieldRegistry {
/** Look up a field/property by its owning class nodeId and field name. */
/**
* First field registered under `(ownerNodeId, fieldName)`, if any.
* Registration order is first-wins: when a Property and a Variable share
* an `(owner, simpleName)` key, the earlier `register(...)` call's def is
* returned. Prefer `lookupAllByOwner` when overloads or duplicate-kind
* entries under the same name must all be visible.
*/
lookupFieldByOwner(ownerNodeId: string, fieldName: string): SymbolDefinition | undefined;
/**
* Every field registered under `(ownerNodeId, fieldName)` in registration
* order. Returns `[]` on miss.
*/
lookupAllByOwner(ownerNodeId: string, fieldName: string): readonly SymbolDefinition[];
}
// ---------------------------------------------------------------------------
@@ -21,7 +38,7 @@ export interface FieldRegistry {
// ---------------------------------------------------------------------------
export interface MutableFieldRegistry extends FieldRegistry {
/** Register a field/property under its owner. */
/** Register a field under its owner. Appends when the key already exists. */
register(ownerNodeId: string, fieldName: string, def: SymbolDefinition): void;
/** Clear all entries. */
clear(): void;
@@ -32,22 +49,36 @@ export interface MutableFieldRegistry extends FieldRegistry {
// ---------------------------------------------------------------------------
export const createFieldRegistry = (): MutableFieldRegistry => {
const fieldByOwner = new Map<string, SymbolDefinition>();
const fieldByOwner = new Map<string, SymbolDefinition[]>();
const lookupAllByOwner = (
ownerNodeId: string,
fieldName: string,
): readonly SymbolDefinition[] => {
return fieldByOwner.get(`${ownerNodeId}\0${fieldName}`) ?? EMPTY;
};
const lookupFieldByOwner = (
ownerNodeId: string,
fieldName: string,
): SymbolDefinition | undefined => {
return fieldByOwner.get(`${ownerNodeId}\0${fieldName}`);
const pool = lookupAllByOwner(ownerNodeId, fieldName);
return pool.length === 0 ? undefined : pool[0];
};
const register = (ownerNodeId: string, fieldName: string, def: SymbolDefinition): void => {
fieldByOwner.set(`${ownerNodeId}\0${fieldName}`, def);
const key = `${ownerNodeId}\0${fieldName}`;
const existing = fieldByOwner.get(key);
if (existing) {
existing.push(def);
} else {
fieldByOwner.set(key, [def]);
}
};
const clear = (): void => {
fieldByOwner.clear();
};
return { lookupFieldByOwner, register, clear };
return { lookupFieldByOwner, lookupAllByOwner, register, clear };
};
@@ -0,0 +1,45 @@
/**
* Owner-keyed member lookup for Step 2 (RFC #909 / PR #1656).
*
* Merges MethodRegistry + FieldRegistry hits for `(ownerDefId, memberName)`
* in O(1) map time per registry — no `defs.byId` scan. Callers that omit
* this helper and leave `ownedMembersByOwner` unset fall back to an O(|defs|)
* compatibility scan inside `lookupCore.collectOwnedMembers`.
*/
import type { DefId, SymbolDefinition } from 'gitnexus-shared';
import type { SemanticModel } from './semantic-model.js';
const EMPTY: readonly SymbolDefinition[] = Object.freeze([]);
/**
* Production hook for `RegistryContext.ownedMembersByOwner`.
* Returns `[]` on miss (authoritative indexed empty) — never `undefined`.
*
* Merges hits from all three owner-keyed registries (methods, fields,
* nested types) under the same `(ownerDefId, memberName)` key. The
* caller's `acceptedKinds` filter in `lookupCore` picks the right subset.
*/
export function lookupOwnedMembersByOwner(
model: Pick<SemanticModel, 'methods' | 'fields' | 'types'>,
ownerDefId: DefId,
memberName: string,
): readonly SymbolDefinition[] {
const methods = model.methods.lookupAllByOwner(ownerDefId, memberName);
const fields = model.fields.lookupAllByOwner(ownerDefId, memberName);
const nestedTypes = model.types.lookupAllByOwner(ownerDefId, memberName);
const methodCount = methods.length;
const fieldCount = fields.length;
const typeCount = nestedTypes.length;
const total = methodCount + fieldCount + typeCount;
if (total === 0) return EMPTY;
if (methodCount === total) return methods;
if (fieldCount === total) return fields;
if (typeCount === total) return nestedTypes;
const merged = new Array<SymbolDefinition>(total);
let i = 0;
for (let j = 0; j < methodCount; j++) merged[i++] = methods[j]!;
for (let j = 0; j < fieldCount; j++) merged[i++] = fields[j]!;
for (let j = 0; j < typeCount; j++) merged[i++] = nestedTypes[j]!;
return merged;
}
@@ -34,7 +34,7 @@
* logic up the dependency chain instead.
*/
import type { NodeLabel, SymbolDefinition } from 'gitnexus-shared';
import type { NodeLabel, ParameterTypeClass, SymbolDefinition } from 'gitnexus-shared';
/**
* Class-like NodeLabels — used for qualifiedName fallback inside
@@ -126,6 +126,7 @@ export interface AddMetadata {
parameterCount?: number;
requiredParameterCount?: number;
parameterTypes?: string[];
parameterTypeClasses?: ParameterTypeClass[];
returnType?: string;
declaredType?: string;
templateArguments?: string[];
@@ -276,6 +277,9 @@ export const createSymbolTable = (): InternalSymbolTable => {
...(metadata?.parameterTypes !== undefined
? { parameterTypes: metadata.parameterTypes }
: {}),
...(metadata?.parameterTypeClasses !== undefined
? { parameterTypeClasses: metadata.parameterTypeClasses }
: {}),
...(metadata?.returnType !== undefined ? { returnType: metadata.returnType } : {}),
...(metadata?.declaredType !== undefined ? { declaredType: metadata.declaredType } : {}),
...(metadata?.templateArguments !== undefined
@@ -8,6 +8,8 @@
import type { SymbolDefinition } from 'gitnexus-shared';
const EMPTY: readonly SymbolDefinition[] = Object.freeze([]);
// ---------------------------------------------------------------------------
// Public read-only interface
// ---------------------------------------------------------------------------
@@ -35,6 +37,14 @@ export interface TypeRegistry {
* Returned array is a view into the live index — do not mutate.
*/
lookupImplByName(name: string): readonly SymbolDefinition[];
/**
* Look up nested-type defs registered under `(ownerNodeId, simpleName)`
* in registration order. Returns `[]` on miss. Used by Step 2 Receiver/MRO
* resolution when the receiver's owner declares nested classes/structs/
* enums/typedefs/etc. that the caller's `acceptedKinds` includes.
*/
lookupAllByOwner(ownerNodeId: string, simpleName: string): readonly SymbolDefinition[];
}
// ---------------------------------------------------------------------------
@@ -46,6 +56,8 @@ export interface MutableTypeRegistry extends TypeRegistry {
registerClass(name: string, qualifiedName: string, def: SymbolDefinition): void;
/** Register a Rust Impl block by name. */
registerImpl(name: string, def: SymbolDefinition): void;
/** Register a nested type under its owner. Appends when the key already exists. */
registerByOwner(ownerNodeId: string, simpleName: string, def: SymbolDefinition): void;
/** Clear all entries. */
clear(): void;
}
@@ -58,6 +70,7 @@ export const createTypeRegistry = (): MutableTypeRegistry => {
const classByName = new Map<string, SymbolDefinition[]>();
const classByQualifiedName = new Map<string, SymbolDefinition[]>();
const implByName = new Map<string, SymbolDefinition[]>();
const nestedByOwner = new Map<string, SymbolDefinition[]>();
const lookupClassByName = (name: string): SymbolDefinition[] => {
return classByName.get(name) ?? [];
@@ -71,6 +84,13 @@ export const createTypeRegistry = (): MutableTypeRegistry => {
return implByName.get(name) ?? [];
};
const lookupAllByOwner = (
ownerNodeId: string,
simpleName: string,
): readonly SymbolDefinition[] => {
return nestedByOwner.get(`${ownerNodeId}\0${simpleName}`) ?? EMPTY;
};
const registerClass = (name: string, qualifiedName: string, def: SymbolDefinition): void => {
const existing = classByName.get(name);
if (existing) {
@@ -96,18 +116,35 @@ export const createTypeRegistry = (): MutableTypeRegistry => {
}
};
const registerByOwner = (
ownerNodeId: string,
simpleName: string,
def: SymbolDefinition,
): void => {
const key = `${ownerNodeId}\0${simpleName}`;
const existing = nestedByOwner.get(key);
if (existing) {
existing.push(def);
} else {
nestedByOwner.set(key, [def]);
}
};
const clear = (): void => {
classByName.clear();
classByQualifiedName.clear();
implByName.clear();
nestedByOwner.clear();
};
return {
lookupClassByName,
lookupClassByQualifiedName,
lookupImplByName,
lookupAllByOwner,
registerClass,
registerImpl,
registerByOwner,
clear,
};
};
+130 -26
View File
@@ -1,4 +1,4 @@
import type { GraphNode, GraphRelationship, NodeLabel } from 'gitnexus-shared';
import type { GraphNode, GraphRelationship, NodeLabel, ParameterTypeClass } from 'gitnexus-shared';
import { KnowledgeGraph } from '../graph/types.js';
import Parser from 'tree-sitter';
import { loadParser, loadLanguage, isLanguageAvailable } from '../tree-sitter/parser-loader.js';
@@ -14,6 +14,7 @@ import { isVerboseIngestionEnabled } from './utils/verbose.js';
import {
getDefinitionNodeFromCaptures,
findEnclosingClassInfo,
findObjectLiteralBindingInfo,
getLabelFromCaptures,
CLASS_CONTAINER_TYPES,
type SyntaxNode,
@@ -30,7 +31,11 @@ import {
constTagForId,
buildCollisionGroups,
} from './utils/method-props.js';
import { extractTemplateArguments, templateArgumentsIdTag } from './utils/template-arguments.js';
import {
extractTemplateArguments,
templateArgumentsIdTag,
templateConstraintsIdTag,
} from './utils/template-arguments.js';
import type { LanguageProvider } from './language-provider.js';
import type { ParsedFile } from 'gitnexus-shared';
import { WorkerPool } from './workers/worker-pool.js';
@@ -128,6 +133,7 @@ export const mergeChunkResults = (
parameterCount: sym.parameterCount,
requiredParameterCount: sym.requiredParameterCount,
parameterTypes: sym.parameterTypes,
parameterTypeClasses: sym.parameterTypeClasses,
returnType: sym.returnType,
declaredType: sym.declaredType,
templateArguments: sym.templateArguments,
@@ -526,6 +532,10 @@ const processParsingSequential = async (
)
: null;
const enclosingClassId = enclosingClassInfo?.classId ?? null;
const objectLiteralOwnerInfo =
!enclosingClassId && nodeLabel === 'Method' && definitionNode
? findObjectLiteralBindingInfo(definitionNode, file.path)
: null;
// Qualify method/property IDs with enclosing class name to avoid collisions
// e.g. "Method:animal.dart:Animal.speak" vs "Method:animal.dart:Dog.speak"
@@ -650,9 +660,38 @@ const processParsingSequential = async (
classTemplateArguments.length > 0
? templateArgumentsIdTag(classTemplateArguments)
: '';
// SFINAE / `requires`-clause aware ID disambiguation (issue #1579).
// Function-template overloads with identical parameterTypes but
// mutually-exclusive constraints (e.g. `enable_if_t<is_integral_v<T>>`
// vs `enable_if_t<is_floating_point_v<T>>`) need distinct graph
// nodes so the constraint-filter step in `narrowOverloadCandidates`
// has two candidates to narrow between. Without this tag they
// collapse to a single Function node and the SFINAE call resolves
// to only one edge regardless of which overload's constraint holds.
// The provider hook is the right invocation point — parsing-processor
// sees raw tree-sitter matches without the `@`-prefixed synthetic
// captures `scope-extractor` consumes, so we delegate extraction to
// the language adapter (C++ implements this; other languages opt out).
let parsedTemplateConstraints: unknown = undefined;
let constraintsTag = '';
if (
(nodeLabel === 'Function' || nodeLabel === 'Method') &&
provider.extractTemplateConstraints !== undefined &&
definitionNode !== null
) {
try {
parsedTemplateConstraints = provider.extractTemplateConstraints(definitionNode);
if (parsedTemplateConstraints !== undefined) {
constraintsTag = templateConstraintsIdTag(parsedTemplateConstraints);
}
} catch {
parsedTemplateConstraints = undefined;
constraintsTag = '';
}
}
const nodeId = generateId(
nodeLabel,
`${file.path}:${qualifiedName}${classTemplateTag}${arityTag}`,
`${file.path}:${qualifiedName}${classTemplateTag}${arityTag}${constraintsTag}`,
);
const classNodeForSymbol = definitionNodeForRange || definitionNode || nameNode;
const qualifiedTypeName =
@@ -689,6 +728,9 @@ const processParsingSequential = async (
...(classTemplateArguments !== undefined && classTemplateArguments.length > 0
? { templateArguments: classTemplateArguments }
: {}),
...(parsedTemplateConstraints !== undefined
? { templateConstraints: parsedTemplateConstraints }
: {}),
...(frameworkHint
? {
astFrameworkMultiplier: frameworkHint.entryPointMultiplier,
@@ -744,10 +786,11 @@ const processParsingSequential = async (
parameterCount: methodProps.parameterCount as number | undefined,
requiredParameterCount: methodProps.requiredParameterCount as number | undefined,
parameterTypes: methodProps.parameterTypes as string[] | undefined,
parameterTypeClasses: methodProps.parameterTypeClasses as ParameterTypeClass[] | undefined,
returnType: methodProps.returnType as string | undefined,
declaredType,
templateArguments: classTemplateArguments,
ownerId: enclosingClassId ?? undefined,
ownerId: enclosingClassId ?? objectLiteralOwnerInfo?.ownerId ?? undefined,
qualifiedName: qualifiedTypeName,
});
@@ -767,15 +810,18 @@ const processParsingSequential = async (
graph.addRelationship(relationship);
// ── HAS_METHOD / HAS_PROPERTY: link member to enclosing class ──
if (enclosingClassId) {
const ownerIdForMemberEdge = enclosingClassId ?? objectLiteralOwnerInfo?.ownerId ?? null;
if (ownerIdForMemberEdge) {
const memberEdgeType = nodeLabel === 'Property' ? 'HAS_PROPERTY' : 'HAS_METHOD';
graph.addRelationship({
id: generateId(memberEdgeType, `${enclosingClassId}->${nodeId}`),
sourceId: enclosingClassId,
id: generateId(memberEdgeType, `${ownerIdForMemberEdge}->${nodeId}`),
sourceId: ownerIdForMemberEdge,
targetId: nodeId,
type: memberEdgeType,
confidence: 1.0,
reason: '',
reason: objectLiteralOwnerInfo
? 'object literal method belongs to exported object binding'
: '',
});
}
});
@@ -794,6 +840,14 @@ const processParsingSequential = async (
// Public API
// ============================================================================
/**
* Per-`WorkerPool` log-dedup state for quarantine reporting. Keyed on the
* pool instance so multiple concurrent pools (test fixtures, future
* multi-pool callers) each get their own seen-set. WeakMap entries vanish
* when the pool is garbage-collected.
*/
const loggedQuarantineByPool = new WeakMap<WorkerPool, Set<string>>();
export const processParsing = async (
graph: KnowledgeGraph,
files: { path: string; content: string }[],
@@ -836,25 +890,75 @@ export const processParsing = async (
`[scope-resolution prof] worker pool engaged for ${files.length} files — cross-phase tree cache will be empty; scope-resolution re-parses.`,
);
}
try {
return await processParsingWithWorkers(
graph,
files,
symbolTable,
astCache,
workerPool,
reportProgress,
outRawResults,
);
} catch (err) {
const message = err instanceof Error ? err.message : String(err);
logger.warn({ message }, 'Worker pool parsing stopped; continuing with sequential parser:');
reportProgress?.(
lastProgress,
files.length,
`Sequential fallback after worker issue: ${message}`,
);
// U20 design pivot: the worker pool's resilience layers
// (respawn budget, circuit breaker, quarantine, slot-attribution,
// cumulative timeout) are the SOLE contract for handling worker
// failures. There is no sequential-parser fallback for either
// partial quarantine or full pool failure — the operator must see
// a clear hard signal when workers can't recover, instead of a
// silently-degraded graph from a possibly-crashing main-thread
// sequential parser. A failing tree-sitter native binding that
// quarantined a worker would, under the previous design, re-trigger
// the same SIGSEGV on the main thread; we avoid that risk entirely.
//
// - Partial quarantine: the file is missing from this run's graph;
// the per-chunk warn log below surfaces it; U2's chunk-cache
// write-guard in parse-impl.ts keeps the chunk uncached so the
// next analyze gets a cache miss and a fresh pool retries.
// - Full pool failure: `WorkerPoolDispatchError` propagates from
// `processParsingWithWorkers` up through this function. The
// analyze run errors out instead of falling back to sequential.
const data = await processParsingWithWorkers(
graph,
files,
symbolTable,
astCache,
workerPool,
reportProgress,
outRawResults,
);
// Session-scoped quarantine (worker-pool resilience Layer 3): surface
// any files this pool has decided are unsafe for workers so the
// operator can see what was skipped. The pool already filtered them
// out of dispatch; we only need to log + progress-report. Quarantine
// is session-scoped per pool instance — a fresh `createWorkerPool`
// call clears it.
//
// Dedup: log full path list only for entries newly quarantined since
// the previous dispatch on the same pool. The per-chunk progress
// message still surfaces the count for UX continuity, but the
// structured `quarantinedFiles` payload is only emitted when there
// is new signal — prevents O(quarantine × chunks) log spam.
const quarantineSnapshot = workerPool.getQuarantinedPaths?.() ?? [];
const quarantineSet = new Set(quarantineSnapshot);
if (quarantineSet.size > 0) {
const quarantinedInChunk = files.filter((file) => quarantineSet.has(file.path));
if (quarantinedInChunk.length > 0) {
const seenForPool = loggedQuarantineByPool.get(workerPool) ?? new Set<string>();
const newlyQuarantined = quarantinedInChunk
.map((file) => file.path)
.filter((p) => !seenForPool.has(p));
for (const p of newlyQuarantined) seenForPool.add(p);
loggedQuarantineByPool.set(workerPool, seenForPool);
if (newlyQuarantined.length > 0) {
logger.warn(
{
newlyQuarantined,
cumulativeQuarantine: quarantineSet.size,
chunkSkipped: quarantinedInChunk.length,
},
`Worker quarantine: ${newlyQuarantined.length} new file(s) skipped this chunk ` +
`(${quarantinedInChunk.length} skipped total, ${quarantineSet.size} cumulative).`,
);
}
reportProgress?.(
lastProgress,
files.length,
`${quarantinedInChunk.length} worker-quarantined file(s) skipped`,
);
}
}
return data;
}
// Fallback: sequential parsing (no pre-extracted data)
@@ -48,13 +48,14 @@ import { ASTCache, createASTCache } from '../ast-cache.js';
import { type PipelineProgress, getLanguageFromFilename } from 'gitnexus-shared';
import { readFileContents } from '../filesystem-walker.js';
import { isLanguageAvailable } from '../../tree-sitter/parser-loader.js';
import { createWorkerPool } from '../workers/worker-pool.js';
import { createWorkerPool, WorkerPoolInitializationError } from '../workers/worker-pool.js';
import type { WorkerPool } from '../workers/worker-pool.js';
import type {
ExtractedAssignment,
ExtractedCall,
ExtractedDecoratorRoute,
ExtractedFetchCall,
ExtractedImport,
ExtractedORMQuery,
ExtractedRoute,
ExtractedToolDef,
@@ -69,6 +70,7 @@ import path from 'node:path';
import { fileURLToPath, pathToFileURL } from 'node:url';
import { isDev } from '../utils/env.js';
import { isVerboseIngestionEnabled } from '../utils/verbose.js';
import { synthesizeWildcardImportBindings, needsSynthesis } from './wildcard-synthesis.js';
import { extractORMQueriesInline } from './orm-extraction.js';
@@ -85,11 +87,24 @@ import { logger } from '../../logger.js';
* gives a useful invalidation floor (~1/N chunks on a multi-MB repo)
* while keeping worker dispatch overhead under 5% on cold runs.
*/
const CHUNK_BYTE_BUDGET = (() => {
/**
* Built-in chunk byte budget when neither `PipelineOptions.chunkByteBudget`
* nor `GITNEXUS_CHUNK_BYTE_BUDGET` is set. Tuned to give a useful
* cache-invalidation floor (~1/N chunks on a multi-MB repo) while keeping
* worker dispatch overhead under 5% on cold runs. Resolution happens at
* call time inside `runChunkedParseAndResolve` (U14 from PR #1693 review)
* — previously this was a module-load IIFE, which froze the env value at
* import time and meant per-call option threading silently no-op'd.
*/
const DEFAULT_CHUNK_BYTE_BUDGET = 2 * 1024 * 1024;
function resolveChunkByteBudget(options?: PipelineOptions): number {
const opt = options?.chunkByteBudget;
if (typeof opt === 'number' && Number.isFinite(opt) && opt > 0) return opt;
const env = Number(process.env.GITNEXUS_CHUNK_BYTE_BUDGET);
if (Number.isFinite(env) && env > 0) return env;
return 2 * 1024 * 1024;
})();
return DEFAULT_CHUNK_BYTE_BUDGET;
}
// ── Main parse + resolve function ──────────────────────────────────────────
@@ -177,18 +192,28 @@ export async function runChunkedParseAndResolve(
if (totalParseable === 0) {
onProgress({
phase: 'parsing',
percent: 82,
// Skip directly to the end of the parse-phase progress band (M2 from PR
// #1693 review). Parse 20-70%, deferred 70-95%; nothing in either runs
// when there's no parseable file, so jump to 95.
percent: 95,
message: 'No parseable files found — skipping parsing phase',
stats: { filesProcessed: 0, totalFiles: 0, nodesCreated: graph.nodeCount },
});
}
// Build byte-budget chunks
// Build byte-budget chunks. The budget is resolved per-call (U14): options
// first, then env, then the built-in default. Pre-U14 this was a
// module-load IIFE constant, which froze the env value at import time
// and made `PipelineOptions.chunkByteBudget` silently no-op on warm test
// runs. Resolving in the function body restores per-call configurability
// and matches the pattern used by resolveAutoPoolSize and the U1
// parseChunkConcurrency resolver.
const chunkByteBudget = resolveChunkByteBudget(options);
const chunks: string[][] = [];
let currentChunk: string[] = [];
let currentBytes = 0;
for (const file of parseableScanned) {
if (currentChunk.length > 0 && currentBytes + file.size > CHUNK_BYTE_BUDGET) {
if (currentChunk.length > 0 && currentBytes + file.size > chunkByteBudget) {
chunks.push(currentChunk);
currentChunk = [];
currentBytes = 0;
@@ -203,16 +228,22 @@ export async function runChunkedParseAndResolve(
if (isDev) {
const totalMB = parseableScanned.reduce((s, f) => s + f.size, 0) / (1024 * 1024);
logger.info(
`📂 Scan: ${totalFiles} paths, ${totalParseable} parseable (${totalMB.toFixed(0)}MB), ${numChunks} chunks @ ${CHUNK_BYTE_BUDGET / (1024 * 1024)}MB budget`,
`📂 Scan: ${totalFiles} paths, ${totalParseable} parseable (${totalMB.toFixed(0)}MB), ${numChunks} chunks @ ${chunkByteBudget / (1024 * 1024)}MB budget`,
);
}
onProgress({
phase: 'parsing',
percent: 20,
message: `Parsing ${totalParseable} files in ${numChunks} chunk${numChunks !== 1 ? 's' : ''}...`,
stats: { filesProcessed: 0, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
});
// Skip the "Parsing N files..." announcement when there's nothing to parse
// — the early-return branch above already emitted percent 95 ("skipping
// parsing phase"), and emitting percent 20 here would regress the
// progress stream non-monotonically (M2 from PR #1693 review).
if (totalParseable > 0) {
onProgress({
phase: 'parsing',
percent: 20,
message: `Parsing ${totalParseable} files in ${numChunks} chunk${numChunks !== 1 ? 's' : ''}...`,
stats: { filesProcessed: 0, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
});
}
// Don't spawn workers for tiny repos — overhead exceeds benefit.
// Test suites may lower the thresholds via `options.workerThresholdsForTest`
@@ -221,18 +252,36 @@ export async function runChunkedParseAndResolve(
const MIN_BYTES_FOR_WORKERS = options?.workerThresholdsForTest?.minBytes ?? 512 * 1024;
const totalBytes = parseableScanned.reduce((s, f) => s + f.size, 0);
// Create worker pool once, reuse across chunks
let workerPool: WorkerPool | undefined;
if (
// Create worker pool lazily, reuse across cache-miss chunks.
//
// `workerPoolSize === 0` is a programmatic equivalent of `skipWorkers:
// true` per the `PipelineOptions.workerPoolSize` contract. Short-
// circuiting here avoids constructing a useless pool. The pool is
// intentionally NOT created before parse-cache lookup: a warm-cache
// all-hit run should replay cached worker output without loading
// parse-worker.js or any tree-sitter/N-API native bindings.
const shouldUseWorkers =
!options?.skipWorkers &&
(totalParseable >= MIN_FILES_FOR_WORKERS || totalBytes >= MIN_BYTES_FOR_WORKERS)
) {
options?.workerPoolSize !== 0 &&
(totalParseable >= MIN_FILES_FOR_WORKERS || totalBytes >= MIN_BYTES_FOR_WORKERS);
let workerPool: WorkerPool | undefined;
let workerPoolDisabled = false;
const getOrCreateWorkerPool = (): WorkerPool | undefined => {
if (!shouldUseWorkers || workerPoolDisabled) return undefined;
if (workerPool) return workerPool;
try {
let workerUrl = new URL('../workers/parse-worker.js', import.meta.url);
// U20.U3 test-only injection: integration tests pass a custom
// worker script URL via `workerUrlForTest` (mirrors the
// `workerThresholdsForTest` precedent) so they can drive the
// chunk-loop with deterministically-misbehaving workers without
// mocking the module import graph. When unset, the normal src/
// → dist/ resolution runs.
let workerUrl =
options?.workerUrlForTest ?? new URL('../workers/parse-worker.js', import.meta.url);
// When running under vitest, import.meta.url points to src/ where no .js exists.
// Fall back to the compiled dist/ worker so the pool can spawn real worker threads.
const thisDir = fileURLToPath(new URL('.', import.meta.url));
if (!fs.existsSync(fileURLToPath(workerUrl))) {
if (!options?.workerUrlForTest && !fs.existsSync(fileURLToPath(workerUrl))) {
const distWorker = path.resolve(
thisDir,
'..',
@@ -249,14 +298,17 @@ export async function runChunkedParseAndResolve(
workerUrl = pathToFileURL(distWorker);
}
}
workerPool = createWorkerPool(workerUrl);
workerPool = createWorkerPool(workerUrl, options?.workerPoolSize);
return workerPool;
} catch (err) {
workerPoolDisabled = true;
logger.warn(
{ err: (err as Error).message },
'Worker pool creation failed, using sequential fallback:',
);
return undefined;
}
}
};
let filesParsedSoFar = 0;
@@ -301,6 +353,16 @@ export async function runChunkedParseAndResolve(
const deferredWorkerHeritage: ExtractedHeritage[] = [];
const deferredConstructorBindings: FileConstructorBindings[] = [];
const deferredAssignments: ExtractedAssignment[] = [];
// Imports accumulated across chunks. Previously processed per-chunk
// via `processImportsFromExtracted` inside the chunk loop, which
// forced workers to sit idle on the main thread's extraction pass
// between chunk dispatches (4-5% CPU utilization symptom). Deferring
// to a single end-of-loop pass lets the worker pool start chunk N+1
// immediately after chunk N's worker dispatch returns. Resolution is
// strictly-more-information at end-of-loop because graph now has
// every chunk's symbols — improves cross-chunk import targets.
const deferredWorkerImports: ExtractedImport[] = [];
let anyChunkNeedsWildcardSynth = false;
// Aggregated per-file ParsedFile artifacts produced by workers' calls
// to `extractParsedFile`. Threaded through to the scope-resolution
// phase so it can SKIP its own re-extraction on cache hits — this is
@@ -317,13 +379,63 @@ export async function runChunkedParseAndResolve(
let chunkCacheMisses = 0;
try {
// U1 — bounded chunk concurrency (B1 from PR #1693 review): pre-fetch
// chunk file contents up to `parseChunkConcurrency` chunks ahead of the
// dispatch cursor so file I/O overlaps with worker compute. Worker
// dispatch itself stays serial because `WorkerPool.dispatch` is not
// reentrant (concurrent calls would race on the shared per-slot
// busy/in-flight state). With concurrency=1 behavior is identical to
// the pure-serial loop. F4: deferred-state aggregation still happens
// in chunkIdx order (the for-loop below iterates sequentially), so
// cross-chunk processors see deterministic input regardless of
// file-read completion order. Honors options.parseChunkConcurrency
// (threaded from the CLI), then GITNEXUS_PARSE_CHUNK_CONCURRENCY env
// (default 2 — matches the help text the CLI advertises).
const parseChunkConcurrency = ((): number => {
const opt = options?.parseChunkConcurrency;
if (typeof opt === 'number' && Number.isInteger(opt) && opt >= 1) return opt;
const env = Number(process.env.GITNEXUS_PARSE_CHUNK_CONCURRENCY);
if (Number.isInteger(env) && env >= 1) return env;
return 2;
})();
const chunkContentPromises = new Array<Promise<Map<string, string>> | undefined>(numChunks);
const startChunkPrefetch = (i: number): void => {
if (i >= numChunks || chunkContentPromises[i] !== undefined) return;
chunkContentPromises[i] = readFileContents(repoPath, chunks[i]);
};
for (let i = 0; i < Math.min(parseChunkConcurrency, numChunks); i++) {
startChunkPrefetch(i);
}
// Hoisted loop-invariant: GITNEXUS_VERBOSE / NODE_ENV are read once
// (not on every chunk). Previously evaluated at the top of the loop
// body, which re-read process.env on every iteration even though
// the env can't change mid-run.
const verboseThroughputLog = isDev || isVerboseIngestionEnabled();
for (let chunkIdx = 0; chunkIdx < numChunks; chunkIdx++) {
const chunkPaths = chunks[chunkIdx];
// Start wall-clock for the per-chunk throughput log emitted at end
// of this iteration. The gate is computed once above; here we just
// sample the clock if the gate is on. Computed when either
// NODE_ENV=development OR the operator passed `--verbose`
// (GITNEXUS_VERBOSE) — the previous `isDev`-only gate meant
// operators running `gitnexus analyze --verbose` in production
// never saw the log (M3 from PR #1693 review).
const chunkStartMs: number | null = verboseThroughputLog ? Date.now() : null;
const chunkContents = await readFileContents(repoPath, chunkPaths);
const chunkFiles = chunkPaths
.filter((p) => chunkContents.has(p))
.map((p) => ({ path: p, content: chunkContents.get(p)! }));
const chunkContentPromise = chunkContentPromises[chunkIdx];
if (!chunkContentPromise) {
throw new Error(`Missing prefetched parse chunk ${chunkIdx + 1}/${numChunks}`);
}
const chunkContents = await chunkContentPromise;
chunkContentPromises[chunkIdx] = undefined; // release the in-memory copy
startChunkPrefetch(chunkIdx + parseChunkConcurrency);
const chunkFiles: Array<{ path: string; content: string }> = [];
for (const p of chunkPaths) {
const content = chunkContents.get(p);
if (content !== undefined) chunkFiles.push({ path: p, content });
}
// Compute the chunk's content-hash signature (if cache available).
let chunkHash: string | null = null;
@@ -336,7 +448,7 @@ export async function runChunkedParseAndResolve(
}
let chunkWorkerData: WorkerExtractedData | null;
const cachedRaw = chunkHash ? parseCache!.entries.get(chunkHash) : undefined;
const cachedRaw = chunkHash && parseCache ? parseCache.entries.get(chunkHash) : undefined;
// Track every chunk hash we touched so the orchestrator can
// prune stale entries (chunks whose composition no longer
@@ -350,14 +462,18 @@ export async function runChunkedParseAndResolve(
chunkWorkerData = mergeChunkResults(graph, symbolTable, cachedRaw);
if (isDev) {
logger.info(
`📦 parse-cache HIT: chunk ${chunkIdx + 1}/${numChunks} (${chunkFiles.length} files, ${chunkHash!.slice(0, 8)})`,
`📦 parse-cache HIT: chunk ${chunkIdx + 1}/${numChunks} (${chunkFiles.length} files, ${chunkHash?.slice(0, 8) ?? 'unknown'})`,
);
}
// Progress update so UI advances even on a cache hit.
const cachedFiles = chunkFiles.length;
onProgress({
phase: 'parsing',
percent: Math.round(20 + ((filesParsedSoFar + cachedFiles) / totalParseable) * 62),
// Parse phase covers 20-70 (50 points). Deferred extraction below
// takes 70-95 so the UI advances through the (potentially long)
// resolution stages instead of holding at 82 (M2 from PR #1693
// review).
percent: Math.round(20 + ((filesParsedSoFar + cachedFiles) / totalParseable) * 50),
message: `Parsing chunk ${chunkIdx + 1}/${numChunks} (cache)...`,
stats: {
filesProcessed: filesParsedSoFar + cachedFiles,
@@ -370,85 +486,121 @@ export async function runChunkedParseAndResolve(
// them under the chunk hash for the next run.
chunkCacheMisses++;
const rawResults: ParseWorkerResult[] = [];
chunkWorkerData = await processParsing(
graph,
chunkFiles,
symbolTable,
astCache,
scopeTreeCache,
(current, _total, filePath) => {
const globalCurrent = filesParsedSoFar + current;
const parsingProgress = 20 + (globalCurrent / totalParseable) * 62;
onProgress({
phase: 'parsing',
percent: Math.round(parsingProgress),
message: `Parsing chunk ${chunkIdx + 1}/${numChunks}...`,
detail: filePath,
stats: {
filesProcessed: globalCurrent,
totalFiles: totalParseable,
nodesCreated: graph.nodeCount,
},
});
},
workerPool,
// Capture raw results only when we have a cache to write to —
// otherwise we'd retain extra arrays for nothing.
parseCache && chunkHash ? rawResults : undefined,
);
const progressForChunk = (current: number, _total: number, filePath: string) => {
const globalCurrent = filesParsedSoFar + current;
// Parse phase covers 20-70 (M2). Deferred extraction handles 70-95.
const parsingProgress = 20 + (globalCurrent / totalParseable) * 50;
onProgress({
phase: 'parsing',
percent: Math.round(parsingProgress),
message: `Parsing chunk ${chunkIdx + 1}/${numChunks}...`,
detail: filePath,
stats: {
filesProcessed: globalCurrent,
totalFiles: totalParseable,
nodesCreated: graph.nodeCount,
},
});
};
const activeWorkerPool = getOrCreateWorkerPool();
try {
chunkWorkerData = await processParsing(
graph,
chunkFiles,
symbolTable,
astCache,
scopeTreeCache,
progressForChunk,
activeWorkerPool,
// Capture raw results only when we have a cache to write to —
// otherwise we'd retain extra arrays for nothing.
parseCache && chunkHash && activeWorkerPool ? rawResults : undefined,
);
} catch (err) {
if (!(err instanceof WorkerPoolInitializationError)) throw err;
logger.warn(
{
err: err.message,
readinessFailures: err.readinessFailures,
},
'Worker pool initialization failed, using sequential fallback:',
);
rawResults.length = 0;
workerPoolDisabled = true;
const failedPool = workerPool;
workerPool = undefined;
await failedPool?.terminate().catch(() => undefined);
chunkWorkerData = await processParsing(
graph,
chunkFiles,
symbolTable,
astCache,
scopeTreeCache,
progressForChunk,
undefined,
undefined,
);
}
// Persist the raw results for this chunk hash. Sequential path
// doesn't populate rawResults (it writes directly to graph), so
// small repos without worker pool simply don't cache. That's fine.
//
// U20.U2: refuse the write when any chunk file is in the
// worker pool's cumulative quarantine snapshot. The chunkHash
// is computed from EVERY file in the chunk, but the pool's
// Layer 3 quarantine filters quarantined files out of dispatch
// — so `rawResults` is narrower than the chunkHash key implies.
// Caching it would silently replay incomplete results on the
// next run with unchanged content (the corruption class Codex's
// adversarial review of PR #1693 flagged).
//
// Skipping the write means the next analyze gets a cache miss
// for this chunk and re-dispatches against a fresh worker pool
// (quarantine is session-scoped — `createQuarantine` is called
// per-pool at worker-pool.ts), giving the quarantined file
// another chance. If quarantine fires again, U20.U1's
// sequential gap-fill still produces a complete graph for this
// run; the cache just stays empty for this chunk until a fully-
// clean dispatch lands.
if (parseCache && chunkHash && rawResults.length > 0) {
parseCache.entries.set(chunkHash, rawResults);
if (isDev) {
logger.info(
`📦 parse-cache MISS+store: chunk ${chunkIdx + 1}/${numChunks} (${chunkFiles.length} files, ${chunkHash.slice(0, 8)})`,
);
const quarantineSnapshot = workerPool?.getQuarantinedPaths?.() ?? [];
const quarantineSet = new Set(quarantineSnapshot);
const chunkHadQuarantine = chunkFiles.some((f) => quarantineSet.has(f.path));
if (chunkHadQuarantine) {
if (isDev) {
const quarantinedInChunk = chunkFiles.filter((f) => quarantineSet.has(f.path)).length;
logger.info(
`📦 parse-cache SKIP: chunk ${chunkIdx + 1}/${numChunks} ` +
`had ${quarantinedInChunk} worker-quarantined file(s); ` +
`next run will rediscover (${chunkHash.slice(0, 8)})`,
);
}
} else {
parseCache.entries.set(chunkHash, rawResults);
if (isDev) {
logger.info(
`📦 parse-cache MISS+store: chunk ${chunkIdx + 1}/${numChunks} (${chunkFiles.length} files, ${chunkHash.slice(0, 8)})`,
);
}
}
}
}
const chunkBasePercent = 20 + (filesParsedSoFar / totalParseable) * 62;
// Per-chunk extraction passes (processImportsFromExtracted,
// processHeritageFromExtracted, processRoutesFromExtracted,
// synthesizeWildcardImportBindings, seedCrossFileReceiverTypes)
// moved out of the chunk loop into a single end-of-loop pass below.
// Reason: per-chunk extraction blocked the chunk loop on
// main-thread work between worker dispatches — workers sat idle
// and total CPU utilization plateaued at 4-5% on multi-core boxes.
// Deferring keeps workers busy chunk-after-chunk; resolution sees
// strictly-more-information (full repo graph) so cross-chunk import
// and heritage targets resolve at least as well as before.
if (chunkWorkerData) {
await processImportsFromExtracted(
graph,
allPathObjects,
chunkWorkerData.imports,
ctx,
(current, total) => {
onProgress({
phase: 'parsing',
percent: Math.round(chunkBasePercent),
message: `Resolving imports (chunk ${chunkIdx + 1}/${numChunks})...`,
detail: `${current}/${total} files`,
stats: {
filesProcessed: filesParsedSoFar,
totalFiles: totalParseable,
nodesCreated: graph.nodeCount,
},
});
},
repoPath,
importCtx,
);
if (chunkNeedsSynthesis[chunkIdx]) {
synthesizeWildcardImportBindings(graph, ctx);
hasSynthesized = true;
}
if (exportedTypeMap.size > 0 && ctx.namedImportMap.size > 0) {
const { enrichedCount } = seedCrossFileReceiverTypes(
chunkWorkerData.calls,
ctx.namedImportMap,
exportedTypeMap,
);
if (isDev && enrichedCount > 0) {
logger.info(
`🔗 E1: Seeded ${enrichedCount} cross-file receiver types (chunk ${chunkIdx + 1})`,
);
}
anyChunkNeedsWildcardSynth = true;
}
for (const item of chunkWorkerData.imports) deferredWorkerImports.push(item);
for (const item of chunkWorkerData.calls) deferredWorkerCalls.push(item);
for (const item of chunkWorkerData.heritage) deferredWorkerHeritage.push(item);
for (const item of chunkWorkerData.constructorBindings)
@@ -463,35 +615,6 @@ export async function runChunkedParseAndResolve(
for (const item of chunkWorkerData.assignments) deferredAssignments.push(item);
}
await Promise.all([
processHeritageFromExtracted(graph, chunkWorkerData.heritage, ctx, (current, total) => {
onProgress({
phase: 'parsing',
percent: Math.round(chunkBasePercent),
message: `Resolving heritage (chunk ${chunkIdx + 1}/${numChunks})...`,
detail: `${current}/${total} records`,
stats: {
filesProcessed: filesParsedSoFar,
totalFiles: totalParseable,
nodesCreated: graph.nodeCount,
},
});
}),
processRoutesFromExtracted(graph, chunkWorkerData.routes ?? [], ctx, (current, total) => {
onProgress({
phase: 'parsing',
percent: Math.round(chunkBasePercent),
message: `Resolving routes (chunk ${chunkIdx + 1}/${numChunks})...`,
detail: `${current}/${total} routes`,
stats: {
filesProcessed: filesParsedSoFar,
totalFiles: totalParseable,
nodesCreated: graph.nodeCount,
},
});
}),
]);
if (chunkWorkerData.fileScopeBindings?.length) {
for (const { filePath, bindings } of chunkWorkerData.fileScopeBindings) {
if (typeof filePath !== 'string' || filePath.length === 0) continue;
@@ -530,6 +653,24 @@ export async function runChunkedParseAndResolve(
filesParsedSoFar += chunkFiles.length;
astCache.clear();
// Throughput observability (U3): emit a per-chunk metrics line
// under verbose ingestion mode so operators can verify CPU
// utilization moved + tune `--workers` / batch sizes without
// guessing. Cheap snapshot — just reads pool closure state.
if (verboseThroughputLog && chunkStartMs !== null) {
const elapsedMs = Date.now() - chunkStartMs;
const filesPerSec = elapsedMs > 0 ? (chunkFiles.length * 1000) / elapsedMs : 0;
const stats = workerPool?.getStats?.();
const poolFrag = stats
? ` pool: ${stats.activeSlots}/${stats.size} active, ` +
`${stats.quarantined} quarantined${stats.poolBroken ? ', BROKEN' : ''}`
: ' (sequential)';
logger.info(
`📊 chunk ${chunkIdx + 1}/${numChunks}: ${chunkFiles.length} files in ${elapsedMs}ms ` +
`(${filesPerSec.toFixed(1)} files/s)${poolFrag}`,
);
}
}
if (isDev && parseCache && (chunkCacheHits > 0 || chunkCacheMisses > 0)) {
@@ -538,10 +679,129 @@ export async function runChunkedParseAndResolve(
);
}
// Deferred end-of-loop extraction (moved out of the per-chunk block):
// 1. processImportsFromExtracted on all chunks' imports
// 2. synthesizeWildcardImportBindings (if any chunk had wildcards)
// 3. seedCrossFileReceiverTypes on deferred calls (depends on
// namedImportMap populated by step 1)
// 4. processHeritageFromExtracted on all chunks' heritage
// 5. processRoutesFromExtracted on all chunks' routes
// Same logic as the prior per-chunk passes, just batched — resolution
// sees the full repo graph instead of just current-and-earlier chunks.
// Deferred extraction band (M2 from PR #1693 review): the 4 stages below
// each get their own 5-10 point slice of the 70-95 range so percent
// advances monotonically through the (potentially long) resolution work
// instead of holding flat at 82. Stages that are skipped (zero-length
// input) leave their band as a no-op jump — the next stage still starts
// at its own band, preserving monotonicity.
// imports: 70 -> 75 (5)
// heritage: 75 -> 80 (5)
// routes: 80 -> 85 (5)
// calls: 85 -> 95 (10)
if (deferredWorkerImports.length > 0) {
await processImportsFromExtracted(
graph,
allPathObjects,
deferredWorkerImports,
ctx,
(current, total) => {
const ratio = total > 0 ? current / total : 1;
onProgress({
phase: 'parsing',
percent: 70 + Math.round(ratio * 5),
message: 'Resolving imports (all chunks)...',
detail: `${current}/${total} files`,
stats: {
filesProcessed: filesParsedSoFar,
totalFiles: totalParseable,
nodesCreated: graph.nodeCount,
},
});
},
repoPath,
importCtx,
);
// U15 (lightweight M1): processImportsFromExtracted is the sole
// consumer of `deferredWorkerImports`. Free the array now so the
// GC can reclaim the per-file ExtractedImport records before the
// heavier downstream stages run (heritage, routes, calls). Peak
// accumulator memory drops from O(repo) to O(repo - imports) for
// the remainder of the deferred phase. The future per-chunk
// streaming upgrade can rewrite this with the same correctness
// contract once profile data shows it's warranted.
deferredWorkerImports.length = 0;
}
if (anyChunkNeedsWildcardSynth) {
synthesizeWildcardImportBindings(graph, ctx);
hasSynthesized = true;
}
// L5 from PR #1693 review: populate `exportedTypeMap` from the in-progress
// graph BEFORE `seedCrossFileReceiverTypes` runs. Previously the seeding
// branch below was reached with `exportedTypeMap.size === 0` in the
// worker path (the map was only built at the post-parse block far below,
// AFTER the seeding branch), so the seed dead-coded itself silently and
// call resolution never got the cross-file receiver-type enrichment.
// The post-parse builder still runs as a defensive fallback on the
// sequential path; its `size === 0` guard means we don't pay the cost
// twice on the worker path.
if (exportedTypeMap.size === 0 && graph.nodeCount > 0) {
const graphExports = buildExportedTypeMapFromGraph(graph, ctx.model.symbols);
for (const [fp, exports] of graphExports) exportedTypeMap.set(fp, exports);
}
if (exportedTypeMap.size > 0 && ctx.namedImportMap.size > 0 && deferredWorkerCalls.length > 0) {
const { enrichedCount } = seedCrossFileReceiverTypes(
deferredWorkerCalls,
ctx.namedImportMap,
exportedTypeMap,
);
if (isDev && enrichedCount > 0) {
logger.info(`🔗 E1: Seeded ${enrichedCount} cross-file receiver types (all chunks)`);
}
}
if (deferredWorkerHeritage.length > 0) {
await processHeritageFromExtracted(graph, deferredWorkerHeritage, ctx, (current, total) => {
const ratio = total > 0 ? current / total : 1;
onProgress({
phase: 'parsing',
percent: 75 + Math.round(ratio * 5),
message: 'Resolving heritage (all chunks)...',
detail: `${current}/${total} records`,
stats: {
filesProcessed: filesParsedSoFar,
totalFiles: totalParseable,
nodesCreated: graph.nodeCount,
},
});
});
}
if (allExtractedRoutes.length > 0) {
await processRoutesFromExtracted(graph, allExtractedRoutes, ctx, (current, total) => {
const ratio = total > 0 ? current / total : 1;
onProgress({
phase: 'parsing',
percent: 80 + Math.round(ratio * 5),
message: 'Resolving routes (all chunks)...',
detail: `${current}/${total} routes`,
stats: {
filesProcessed: filesParsedSoFar,
totalFiles: totalParseable,
nodesCreated: graph.nodeCount,
},
});
});
}
const fullWorkerHeritageMap =
deferredWorkerHeritage.length > 0
? buildHeritageMap(deferredWorkerHeritage, ctx, getHeritageStrategyForLanguage)
: undefined;
// U15 (lightweight M1): buildHeritageMap is the LAST consumer of the
// raw `deferredWorkerHeritage` records — processCallsFromExtracted
// below reads from the derived `fullWorkerHeritageMap` instead. Free
// the raw heritage array now so the GC can reclaim it before the
// (potentially long) call-resolution stage. processHeritageFromExtracted
// earlier was a read-only consumer (pushed to graph, didn't drain).
deferredWorkerHeritage.length = 0;
if (deferredWorkerCalls.length > 0) {
await processCallsFromExtracted(
@@ -549,9 +809,13 @@ export async function runChunkedParseAndResolve(
deferredWorkerCalls,
ctx,
(current, total) => {
const ratio = total > 0 ? current / total : 1;
onProgress({
phase: 'parsing',
percent: 82,
// Calls is the longest deferred stage on real repos — give it the
// 10-point tail 85-95 so the progress bar visibly advances during
// call resolution instead of holding at 82 (M2).
percent: 85 + Math.round(ratio * 10),
message: 'Resolving calls (all chunks)...',
detail: `${current}/${total} files`,
stats: {
@@ -576,6 +840,20 @@ export async function runChunkedParseAndResolve(
bindingAccumulator,
);
}
// U15 (lightweight M1): all three arrays have had their last consumer
// by the time we reach this point — processCallsFromExtracted drained
// `deferredWorkerCalls` and read `deferredConstructorBindings`;
// processAssignmentsFromExtracted drained `deferredAssignments` and
// also read `deferredConstructorBindings`. Free them now so the
// function-scope references die before downstream graph-build /
// scope-resolution starts using its own working memory. Note: arrays
// returned in the function result object (allFetchCalls,
// allExtractedRoutes, allDecoratorRoutes, allToolDefs, allORMQueries,
// allParsedFiles) intentionally stay live — downstream consumers
// need them.
deferredWorkerCalls.length = 0;
deferredConstructorBindings.length = 0;
deferredAssignments.length = 0;
} finally {
await workerPool?.terminate();
}
@@ -605,9 +883,11 @@ export async function runChunkedParseAndResolve(
const cachedSequentialChunkFiles: Array<Array<{ path: string; content: string }>> = [];
for (const chunkPaths of sequentialChunkPaths) {
const chunkContents = await readFileContents(repoPath, chunkPaths);
const chunkFiles = chunkPaths
.filter((p) => chunkContents.has(p))
.map((p) => ({ path: p, content: chunkContents.get(p)! }));
const chunkFiles: Array<{ path: string; content: string }> = [];
for (const p of chunkPaths) {
const content = chunkContents.get(p);
if (content !== undefined) chunkFiles.push({ path: p, content });
}
cachedSequentialChunkFiles.push(chunkFiles);
astCache = createASTCache(chunkFiles.length);
const sequentialHeritage = await extractExtractedHeritageFromFiles(chunkFiles, astCache);
+50
View File
@@ -55,6 +55,16 @@ export interface PipelineOptions {
minFiles?: number;
minBytes?: number;
};
/**
* @internal Test-only override for the worker script URL the pool
* spawns. When unset, parse-impl resolves `parse-worker.js` from the
* adjacent `workers/` directory (or the compiled `dist/` fallback
* under vitest). Integration tests use this to inject a custom
* worker script that deterministically triggers worker-pool
* resilience paths (e.g., crash-on-poison-file) — same precedent as
* `workerThresholdsForTest`. Do not use from production call sites.
*/
workerUrlForTest?: URL;
/**
* Incremental-indexing parse cache. When provided:
* - The parse phase looks up each chunk's content hash in
@@ -68,6 +78,46 @@ export interface PipelineOptions {
* See `gitnexus/src/storage/parse-cache.ts`.
*/
parseCache?: import('../../storage/parse-cache.js').ParseCache;
/**
* Worker pool size override, threaded from the CLI `--workers` flag
* via `AnalyzeOptions`. When set, parse-impl passes this directly to
* `createWorkerPool` so the pool sizing bypasses the env-var fallback
* in `resolveAutoPoolSize`. The env-var channel
* (`GITNEXUS_WORKER_POOL_SIZE`) remains as a back-compat fallback when
* this field is undefined. Setting `workerPoolSize: 0` disables the
* pool entirely (sequential fallback) — equivalent to `skipWorkers`
* but expressed in the same units as `--workers <N>` so long-running
* hosts (eval-server, MCP daemon) can size per-call without leaking
* `process.env` state across analyze invocations.
*/
workerPoolSize?: number;
/**
* Number of chunks whose file contents may be read into memory in
* parallel while the worker pool is busy dispatching the current
* chunk. Pre-fetching overlaps disk I/O for chunk N+1..N+K with the
* worker compute on chunk N — modest but real wall-clock win on
* repos large enough to chunk. Worker dispatch itself remains serial
* because `WorkerPool.dispatch` is not reentrant (concurrent calls
* would race on the shared per-slot busy/in-flight state).
*
* `1` matches today's pure-serial behavior; `2` is the documented
* default (`GITNEXUS_PARSE_CHUNK_CONCURRENCY`). Falls back to the
* env var when undefined; defaults to 2 when neither is set.
*/
parseChunkConcurrency?: number;
/**
* Byte budget per parse chunk (in bytes). When set, parse-impl uses
* this instead of the `GITNEXUS_CHUNK_BYTE_BUDGET` env var or the
* built-in 2 MB default. Smaller values produce more chunks (finer
* cache-hit granularity, more worker dispatches); larger values
* batch more files per dispatch.
*
* Threading the value through options instead of the env var lets
* tests vary the chunk layout per-call without `vi.resetModules` and
* lets long-running hosts (eval-server, MCP daemon) size per-call
* without leaking `process.env` state across invocations.
*/
chunkByteBudget?: number;
}
// ── Phase registry ─────────────────────────────────────────────────────────
@@ -74,6 +74,7 @@ export const MIGRATED_LANGUAGES: ReadonlySet<SupportedLanguages> = new Set<Suppo
SupportedLanguages.C,
SupportedLanguages.CPlusPlus,
SupportedLanguages.PHP,
SupportedLanguages.JavaScript,
]);
/**
@@ -64,6 +64,8 @@ export interface ResolveReferencesInput {
readonly scopes: ScopeResolutionIndexes;
/** Provider hooks consumed by the registries (e.g. `arityCompatibility`). */
readonly providers?: RegistryProviders;
/** Required owner-keyed member lookup used by Step 2 receiver/MRO walks. */
readonly ownedMembersByOwner: RegistryContext['ownedMembersByOwner'];
}
export interface ResolveStats {
@@ -92,6 +94,7 @@ export function resolveReferenceSites(input: ResolveReferencesInput): ResolveRef
defs: scopes.defs,
qualifiedNames: scopes.qualifiedNames,
moduleScopes: scopes.moduleScopes,
ownedMembersByOwner: input.ownedMembersByOwner,
methodDispatch: scopes.methodDispatch,
providers,
};
@@ -191,7 +194,10 @@ function lookupForSite(
case 'write': {
// Try field first; fall through to method then class so bare-name
// reads of a function (e.g. `cb = save`) still resolve.
const fieldHits = fieldRegistry.lookup(site.name, site.inScope);
const fieldOpts: Parameters<FieldRegistry['lookup']>[2] = {
...(site.explicitReceiver !== undefined ? { explicitReceiver: site.explicitReceiver } : {}),
};
const fieldHits = fieldRegistry.lookup(site.name, site.inScope, fieldOpts);
if (fieldHits.length > 0) return fieldHits;
const methodHits = methodRegistry.lookup(site.name, site.inScope);
if (methodHits.length > 0) return methodHits;
+80 -1
View File
@@ -63,6 +63,7 @@ import type {
BindingRef,
CaptureMatch,
ImportEdge,
ParameterTypeClass,
ParsedFile,
ParsedImport,
ReferenceSite,
@@ -545,8 +546,12 @@ function buildDefFromDeclarationMatch(
const parameterCount = parseIntCapture(match['@declaration.parameter-count']);
const requiredParameterCount = parseIntCapture(match['@declaration.required-parameter-count']);
const parameterTypes = parseJsonStringArrayCapture(match['@declaration.parameter-types']);
const parameterTypeClasses = parseJsonParameterTypeClassesCapture(
match['@declaration.parameter-type-classes'],
);
const declaredType = match['@declaration.field-type']?.text;
const returnType = match['@declaration.return-type']?.text;
const templateConstraints = parseJsonCapture(match['@declaration.template-constraints']);
return {
nodeId: makeDefId(filePath, anchor.range, type, nameCap.text),
@@ -556,18 +561,79 @@ function buildDefFromDeclarationMatch(
...(parameterCount !== undefined ? { parameterCount } : {}),
...(requiredParameterCount !== undefined ? { requiredParameterCount } : {}),
...(parameterTypes !== undefined ? { parameterTypes } : {}),
...(parameterTypeClasses !== undefined ? { parameterTypeClasses } : {}),
...(declaredType !== undefined ? { declaredType } : {}),
...(returnType !== undefined ? { returnType } : {}),
...(templateArguments !== undefined ? { templateArguments } : {}),
...(templateConstraints !== undefined ? { templateConstraints } : {}),
};
}
/** Parse an opaque JSON payload synthesized by per-language captures
* (e.g. C++ `@declaration.template-constraints`). Producer owns the
* shape; shared code threads it through as `unknown` per the
* `SymbolDefinition.templateConstraints` contract. */
function parseJsonCapture(cap: { readonly text: string } | undefined): unknown {
if (cap === undefined) return undefined;
try {
return JSON.parse(cap.text);
} catch {
return undefined;
}
}
function parseIntCapture(cap: { readonly text: string } | undefined): number | undefined {
if (cap === undefined) return undefined;
const n = Number.parseInt(cap.text, 10);
return Number.isFinite(n) ? n : undefined;
}
function parseJsonParameterTypeClassesCapture(
cap: { readonly text: string } | undefined,
): ParameterTypeClass[] | undefined {
if (cap === undefined) return undefined;
try {
const parsed = JSON.parse(cap.text);
if (!Array.isArray(parsed)) return undefined;
const out: ParameterTypeClass[] = [];
for (const item of parsed) {
if (item === null || typeof item !== 'object') return undefined;
const o = item as Record<string, unknown>;
if (typeof o.base !== 'string') return undefined;
if (
o.cv !== 'none' &&
o.cv !== 'const' &&
o.cv !== 'volatile' &&
o.cv !== 'const volatile' &&
o.cv !== 'unknown'
) {
return undefined;
}
if (
o.indirection !== 'value' &&
o.indirection !== 'lvalue-ref' &&
o.indirection !== 'rvalue-ref' &&
o.indirection !== 'pointer' &&
o.indirection !== 'unknown'
) {
return undefined;
}
if (typeof o.pointerDepth !== 'number' || !Number.isFinite(o.pointerDepth)) {
return undefined;
}
out.push({
base: o.base,
cv: o.cv,
indirection: o.indirection,
pointerDepth: o.pointerDepth,
});
}
return out;
} catch {
return undefined;
}
}
function parseJsonStringArrayCapture(
cap: { readonly text: string } | undefined,
): string[] | undefined {
@@ -627,8 +693,14 @@ function normalizeNodeLabel(kindStr: string): SymbolDefinition['type'] | undefin
case 'property':
return 'Property';
case 'variable':
case 'const':
return 'Variable';
// `const` / `let` declarations align with the legacy DAG parse phase,
// which emits `Const` graph nodes via `@definition.const` capture for
// `lexical_declaration`. Returning `'Const'` here lets resolveDefGraphId's
// qualified-key path succeed for value receivers without relying on the
// simple-key fallback (PR #1718 review Finding 1 / 2026-05-21-002 U4).
case 'const':
return 'Const';
case 'typealias':
case 'type_alias':
return 'TypeAlias';
@@ -847,6 +919,9 @@ function pass5CollectReferences(
const explicitReceiver = extractExplicitReceiver(match);
const arity = extractArity(match);
const argumentTypes = extractArgumentTypes(match);
const argumentTypeClasses = parseJsonParameterTypeClassesCapture(
match['@reference.parameter-type-classes'],
);
const site: ReferenceSite = {
name: nameCap.text,
@@ -857,6 +932,7 @@ function pass5CollectReferences(
...(explicitReceiver !== undefined ? { explicitReceiver } : {}),
...(arity !== undefined ? { arity } : {}),
...(argumentTypes !== undefined ? { argumentTypes } : {}),
...(argumentTypeClasses !== undefined ? { argumentTypeClasses } : {}),
};
referenceSites.push(site);
}
@@ -974,9 +1050,12 @@ const KNOWN_SUB_TAGS: ReadonlySet<string> = new Set<string>([
'@reference.receiver',
'@reference.arity',
'@reference.parameter-types',
'@reference.parameter-type-classes',
'@declaration.parameter-count',
'@declaration.required-parameter-count',
'@declaration.parameter-types',
'@declaration.parameter-type-classes',
'@declaration.template-constraints',
]);
/**
@@ -254,6 +254,7 @@
import type {
BindingRef,
Callsite,
ConstraintContext,
ParsedFile,
ScopeId,
SupportedLanguages,
@@ -279,6 +280,10 @@ export type LinearizeStrategy = (
/** Result of `ScopeResolver.arityCompatibility` — mirrors `RegistryProviders.arityCompatibility`. */
export type ArityVerdict = 'compatible' | 'unknown' | 'incompatible';
/** Re-exported for ScopeResolver consumers — same shape as
* `RegistryProviders.constraintCompatibility`'s third parameter. */
export type { ConstraintContext } from 'gitnexus-shared';
export interface ScopeResolver {
/** Identity for telemetry + per-language flag check. */
readonly language: SupportedLanguages;
@@ -374,6 +379,28 @@ export interface ScopeResolver {
*/
arityCompatibility(callsite: Callsite, def: SymbolDefinition): ArityVerdict;
/**
* Per-language constraint compatibility between a callsite and a
* candidate `def` that carries `templateConstraints` metadata.
* Mirrors `arityCompatibility` semantics: the three-valued verdict
* MUST treat `'unknown'` as keep-candidate (monotonicity — adding
* a predicate can only narrow correctly, never produce a wrong
* edge). Consulted by `narrowOverloadCandidates` after the arity
* and parameter-type filters.
*
* Optional. Languages without constrained-overload semantics
* (SFINAE, `requires` clauses, trait bounds, conditional types)
* leave this undefined and the constraint filter is a pass-through.
*
* C++ is the first consumer; see `languages/cpp/constraint-filter.ts`
* for the Tier-A predicate registry and Kleene 3-valued evaluator.
*/
readonly constraintCompatibility?: (
callsite: Callsite,
def: SymbolDefinition,
ctx: ConstraintContext,
) => ArityVerdict;
// ─── Per-language strategies ───────────────────────────────────────────────
/**
@@ -98,3 +98,54 @@ export function tryEmitEdge(
});
return true;
}
/**
* Variant of `tryEmitEdge` that takes a pre-resolved target graph id
* instead of resolving it from a `SymbolDefinition`. Used by the
* value-receiver-owner bridge (`receiver-bound-calls.ts` Case 5) where
* the picked owner-indexed method def carries no `qualifiedName` (object
* literals have no class owner to seed it) and therefore cannot
* round-trip through `resolveDefGraphId`. The def's `nodeId` IS the
* canonical graph node id (written by the parse phase), so the caller
* passes it directly.
*
* All other invariants of `tryEmitEdge` apply: dedup key shape, collapse
* flag honoring, edge-type mapping, caller-id resolution.
*/
export function tryEmitEdgeWithExplicitTargetId(
graph: KnowledgeGraph,
scopes: ScopeResolutionIndexes,
nodeLookup: GraphNodeLookup,
site: {
readonly inScope: ScopeId;
readonly atRange: { startLine: number; startCol: number };
readonly kind: string;
},
targetGraphId: string,
reason: string,
seen: Set<string>,
confidence = 0.85,
collapseByCallerTarget = false,
): boolean {
const callerGraphId = resolveCallerGraphId(site.inScope, scopes, nodeLookup);
const edgeType = mapReferenceKindToEdgeType(site.kind as Reference['kind']);
if (callerGraphId === undefined) return false;
if (edgeType === undefined) return false;
const useCollapsed = collapseByCallerTarget && edgeType === 'CALLS';
const dedupKey = useCollapsed
? `${edgeType}:${callerGraphId}->${targetGraphId}`
: `${edgeType}:${callerGraphId}->${targetGraphId}:${site.atRange.startLine}:${site.atRange.startCol}`;
if (seen.has(dedupKey)) return false;
seen.add(dedupKey);
graph.addRelationship({
id: `rel:${dedupKey}`,
sourceId: callerGraphId,
targetId: targetGraphId,
type: edgeType,
confidence,
reason,
});
return true;
}
@@ -21,6 +21,7 @@ import type { NodeLabel, ScopeId, SymbolDefinition } from 'gitnexus-shared';
import type { ScopeResolutionIndexes } from '../../model/scope-resolution-indexes.js';
import { generateId } from '../../../../lib/utils.js';
import { qualifiedKey, simpleKey, type GraphNodeLookup } from '../graph-bridge/node-lookup.js';
import { templateConstraintsIdTag } from '../../utils/template-arguments.js';
/**
* Labels that may legitimately ANCHOR a CALLS/ACCESSES edge as the
* source ("caller"). A Variable / Property can be the TARGET of an
@@ -76,12 +77,31 @@ export function resolveDefGraphId(
type?: NodeLabel;
parameterTypes?: readonly string[];
templateArguments?: readonly string[];
templateConstraints?: unknown;
},
nodeLookup: GraphNodeLookup,
): string | undefined {
const qn = def.qualifiedName;
if (qn === undefined || qn.length === 0) return undefined;
if (def.type !== undefined) {
// SFINAE / `requires`-clause disambiguation (issue #1579) — try the
// constraint-fingerprinted key FIRST. Two function-template overloads
// with identical `parameterTypes` but mutually-exclusive SFINAE
// constraints route to their distinct graph nodes via this key.
// Must run before the parameter-types key because both overloads
// share the latter.
if (
(def.type === 'Function' || def.type === 'Method') &&
def.templateConstraints !== undefined
) {
const cKey = qualifiedKey(
filePath,
def.type,
`${qn}${templateConstraintsIdTag(def.templateConstraints)}`,
);
const cHit = nodeLookup.get(cKey);
if (cHit !== undefined) return cHit;
}
// Overload disambiguation: when the def carries parameter types,
// try the parameter-typed key first so same-name same-arity
// overloads route to their distinct graph nodes.
@@ -20,6 +20,7 @@
import type { NodeLabel } from 'gitnexus-shared';
import type { KnowledgeGraph } from '../../../graph/types.js';
import { templateConstraintsIdTag } from '../../utils/template-arguments.js';
export type GraphNodeLookup = ReadonlyMap<string, string>;
@@ -97,6 +98,21 @@ export function buildGraphNodeLookup(graph: KnowledgeGraph): GraphNodeLookup {
// Each overload is unique — set unconditionally.
lookup.set(pKey, node.id);
}
// SFINAE / `requires`-clause disambiguation (issue #1579) — register
// a constraint-fingerprinted key so resolveDefGraphId can locate the
// correct overload by hashing the def's `templateConstraints`. Mirrors
// the parameter-types key but keys on the opaque constraint payload
// instead, separating two `process<T>` overloads whose
// `parameterTypes=['T']` would otherwise collide.
const tConstraints = (props as { templateConstraints?: unknown }).templateConstraints;
if (tConstraints !== undefined && (node.label === 'Function' || node.label === 'Method')) {
const cKey = qualifiedKey(
props.filePath,
node.label,
`${qualified}${templateConstraintsIdTag(tConstraints)}`,
);
lookup.set(cKey, node.id);
}
if (
(node.label === 'Class' ||
node.label === 'Struct' ||
@@ -143,6 +159,12 @@ export function isLinkableLabel(label: NodeLabel): boolean {
// ACCESSES edges target field nodes (e.g. `user.name = "x"` →
// ACCESSES edge to User's `name` Variable/Property node).
label === 'Variable' ||
label === 'Property'
label === 'Property' ||
// Const is linkable so the value-receiver-owner bridge in
// `receiver-bound-calls.ts` Case 5 can translate the scope-resolution
// `Variable` def for `export const fooService = {...}` to the canonical
// `Const:filePath:name` graph node id, against which object-literal
// method symbols register their `ownerId` (PR #1718 / issue #1358).
label === 'Const'
);
}
@@ -17,12 +17,19 @@
* generalization plan.
*/
import type { ParsedFile, Reference, ScopeId, SymbolDefinition } from 'gitnexus-shared';
import type {
ParameterTypeClass,
ParsedFile,
Reference,
ScopeId,
SymbolDefinition,
} from 'gitnexus-shared';
import type { KnowledgeGraph } from '../../../graph/types.js';
import type { ScopeResolutionIndexes } from '../../model/scope-resolution-indexes.js';
import type { SemanticModel } from '../../model/semantic-model.js';
import type { WorkspaceResolutionIndex } from '../workspace-index.js';
import type { GraphNodeLookup } from '../graph-bridge/node-lookup.js';
import type { ScopeResolver } from '../contract/scope-resolver.js';
import { resolveCallerGraphId, resolveDefGraphId } from '../graph-bridge/ids.js';
import {
findAllCallableBindingsInScope,
@@ -66,11 +73,23 @@ export function emitFreeCallFallback(
parsedFiles: readonly ParsedFile[],
) => readonly SymbolDefinition[] | undefined;
readonly conversionRankFn?: ConversionRankFn;
/** Optional per-language constraint hook threaded into
* `narrowOverloadCandidates`. Drops candidates whose template
* constraints (e.g. C++ `enable_if_t`, C++20 `requires`) provably
* fail at the call site. Three-valued; `'unknown'` keeps the
* candidate (monotonicity). */
readonly constraintCompatibility?: ScopeResolver['constraintCompatibility'];
} = {},
): number {
let emitted = 0;
const seen = new Set<string>();
// Build an O(1) simple-name -> callable defs index over scopes.defs once
// per pass so pickUniqueGlobalCallable doesn't re-scan defs.byId.values()
// per call site. Same name + callable-kind filter that the previous scan
// applied (see pickUniqueGlobalCallable JSDoc). Cost: O(|defs|) once.
const globalCallablesBySimpleName = buildGlobalCallableIndex(scopes);
for (const parsed of parsedFiles) {
for (const site of parsed.referenceSites) {
if (site.kind !== 'call') continue;
@@ -93,13 +112,10 @@ export function emitFreeCallFallback(
// the same name in a single class, choose the best match by
// arity + argument types.
if (fnDef === undefined) {
fnDef = pickImplicitThisOverload(
site,
scopes,
workspaceIndex,
model,
options.conversionRankFn,
);
fnDef = pickImplicitThisOverload(site, scopes, workspaceIndex, model, {
conversionRankFn: options.conversionRankFn,
constraintCompatibility: options.constraintCompatibility,
});
}
// Scope-chain callable lookup. First-match preserves scope-chain
// precedence (local shadows import). When a conversion-rank function
@@ -121,7 +137,11 @@ export function emitFreeCallFallback(
allCallables,
site.arity,
site.argumentTypes,
options.conversionRankFn,
{
argumentTypeClasses: site.argumentTypeClasses,
conversionRankFn: options.conversionRankFn,
constraintCompatibility: options.constraintCompatibility,
},
);
if (narrowed.length === 1) {
fnDef = narrowed[0];
@@ -166,37 +186,46 @@ export function emitFreeCallFallback(
parsedFiles,
);
// When ADL contributed no candidates, narrow ordinary candidates
// with conversion-rank scoring when multiple overloads exist.
// Single candidate or empty falls through to first-match.
const siteKey = `${parsed.filePath}:${site.atRange.startLine}:${site.atRange.startCol}`;
if (adl === undefined || adl.length === 0) {
if (ordinary.length <= 1 || options.conversionRankFn === undefined) {
// No ADL contribution. Default behavior: `ordinary[0]` —
// scope-chain walk preserves local-shadows-import precedence.
//
// Narrowing kicks in when either disambiguation signal is
// present: any candidate carries `templateConstraints`
// (SFINAE / `requires`-clause guarded templates, #1579), OR
// a conversion-rank function is provided (#1606 / #1578).
// Both hooks are threaded into `narrowOverloadCandidates`
// via the unified `OverloadNarrowingHookCtx`.
const hasConstraints = ordinary.some((d) => d.templateConstraints !== undefined);
const canNarrow = hasConstraints || options.conversionRankFn !== undefined;
if (ordinary.length <= 1 || !canNarrow) {
fnDef = ordinary[0];
} else {
const siteKey = `${parsed.filePath}:${site.atRange.startLine}:${site.atRange.startCol}`;
const narrowed = narrowOverloadCandidates(
ordinary,
site.arity,
site.argumentTypes,
options.conversionRankFn,
);
const narrowed = narrowOverloadCandidates(ordinary, site.arity, site.argumentTypes, {
argumentTypeClasses: site.argumentTypeClasses,
conversionRankFn: options.conversionRankFn,
constraintCompatibility: options.constraintCompatibility,
});
if (narrowed.length === 1) {
fnDef = narrowed[0];
} else if (narrowed.length > 1) {
// Multiple survivors — suppress when same-file (true
// overloads), mirrors ADL merged-candidate behavior.
} else if (narrowed.length === 0) {
handledSites.add(siteKey);
continue;
} else {
// >1 survivors: same-file → suppress (true overloads,
// "degrade not lie" — no edge beats a wrong one, and
// SFINAE-ambiguous calls land here). Cross-file →
// first-match (shadowing semantics).
const sameFile = narrowed.every((d) => d.filePath === narrowed[0]!.filePath);
if (sameFile) {
handledSites.add(siteKey);
continue;
}
fnDef = ordinary[0]; // cross-file shadowing → first-match
} else {
fnDef = ordinary[0]; // narrowed empty → first-match
fnDef = ordinary[0];
}
}
} else {
const siteKey = `${parsed.filePath}:${site.atRange.startLine}:${site.atRange.startCol}`;
const merged: SymbolDefinition[] = [];
const seenMerge = new Set<string>();
const push = (defs: readonly SymbolDefinition[]): void => {
@@ -209,12 +238,11 @@ export function emitFreeCallFallback(
push(ordinary);
push(adl);
const narrowed = narrowOverloadCandidates(
merged,
site.arity,
site.argumentTypes,
options.conversionRankFn,
);
const narrowed = narrowOverloadCandidates(merged, site.arity, site.argumentTypes, {
argumentTypeClasses: site.argumentTypeClasses,
conversionRankFn: options.conversionRankFn,
constraintCompatibility: options.constraintCompatibility,
});
if (narrowed.length === 1) {
fnDef = narrowed[0];
} else if (narrowed.length === 0) {
@@ -241,7 +269,7 @@ export function emitFreeCallFallback(
fnDef = pickUniqueGlobalCallable(
site.name,
model,
scopes,
globalCallablesBySimpleName,
parsed.filePath,
options.isFileLocalDef,
site.arity,
@@ -255,6 +283,7 @@ export function emitFreeCallFallback(
})
: undefined,
site.argumentTypes,
site.argumentTypeClasses,
options.conversionRankFn,
);
}
@@ -286,23 +315,46 @@ export function emitFreeCallFallback(
return emitted;
}
/**
* Build a `simpleName -> callable defs` index from `scopes.defs` once per
* pass. Mirrors the filter the old per-site scan applied: Function /
* Method / Constructor, keyed by the last `.`-segment of `qualifiedName`
* (falling back to the qualifiedName itself when undotted). Used by
* `pickUniqueGlobalCallable` so every free-call fallback site is O(1)
* instead of O(|defs|).
*/
function buildGlobalCallableIndex(
scopes: ScopeResolutionIndexes,
): ReadonlyMap<string, readonly SymbolDefinition[]> {
const out = new Map<string, SymbolDefinition[]>();
for (const def of scopes.defs.byId.values()) {
if (def.type !== 'Function' && def.type !== 'Method' && def.type !== 'Constructor') continue;
const qualified = def.qualifiedName;
if (qualified === undefined || qualified.length === 0) continue;
const dot = qualified.lastIndexOf('.');
const simple = dot === -1 ? qualified : qualified.slice(dot + 1);
const bucket = out.get(simple);
if (bucket) bucket.push(def);
else out.set(simple, [def]);
}
return out;
}
function pickUniqueGlobalCallable(
name: string,
model: SemanticModel,
scopes: ScopeResolutionIndexes,
globalCallablesBySimpleName: ReadonlyMap<string, readonly SymbolDefinition[]>,
callerFilePath: string,
isFileLocalDef?: (def: SymbolDefinition) => boolean,
callArity?: number,
isCallerVisible?: (candidate: SymbolDefinition) => boolean,
callArgTypes?: readonly string[],
callArgTypeClasses?: readonly ParameterTypeClass[],
conversionRankFn?: ConversionRankFn,
): SymbolDefinition | undefined {
const scopeDefs: SymbolDefinition[] = [];
const scopeSeen = new Set<string>();
for (const def of scopes.defs.byId.values()) {
const simple = def.qualifiedName?.split('.').pop() ?? def.qualifiedName;
if (simple !== name) continue;
if (def.type !== 'Function' && def.type !== 'Method' && def.type !== 'Constructor') continue;
for (const def of globalCallablesBySimpleName.get(name) ?? []) {
// Skip file-local defs (e.g. C `static` functions) that live in a
// different file from the caller — they are logically invisible.
if (isFileLocalDef !== undefined && def.filePath !== callerFilePath && isFileLocalDef(def)) {
@@ -335,7 +387,10 @@ function pickUniqueGlobalCallable(
// best-rank candidate when exact-type or conversion-rank scoring can
// disambiguate (e.g., `f(int)` vs `f(double)` called with `f(2.5)`).
if (scopeDefs.length > 1) {
const narrowed = narrowOverloadCandidates(scopeDefs, callArity, callArgTypes, conversionRankFn);
const narrowed = narrowOverloadCandidates(scopeDefs, callArity, callArgTypes, {
argumentTypeClasses: callArgTypeClasses,
conversionRankFn,
});
if (narrowed.length === 1) return narrowed[0];
}
@@ -373,7 +428,10 @@ function pickUniqueGlobalCallable(
}
// Same argument-type + conversion-rank narrowing for the model pool.
if (defs.length > 1) {
const narrowed = narrowOverloadCandidates(defs, callArity, callArgTypes, conversionRankFn);
const narrowed = narrowOverloadCandidates(defs, callArity, callArgTypes, {
argumentTypeClasses: callArgTypeClasses,
conversionRankFn,
});
if (narrowed.length === 1) return narrowed[0];
}
@@ -445,11 +503,15 @@ export function pickImplicitThisOverload(
readonly name: string;
readonly arity?: number;
readonly argumentTypes?: readonly string[];
readonly argumentTypeClasses?: readonly import('gitnexus-shared').ParameterTypeClass[];
},
scopes: ScopeResolutionIndexes,
workspaceIndex: WorkspaceResolutionIndex,
model: SemanticModel,
conversionRankFn?: ConversionRankFn,
hookCtx?: {
readonly conversionRankFn?: ConversionRankFn;
readonly constraintCompatibility?: ScopeResolver['constraintCompatibility'];
},
): SymbolDefinition | undefined {
// Find the enclosing Class scope by walking parents.
let curId: ScopeId | null = site.inScope;
@@ -477,12 +539,11 @@ export function pickImplicitThisOverload(
// ambiguous narrowing (multiple compatible candidates with no
// disambiguating signal) leaves the call unresolved rather than
// routing to an arbitrary first overload by registration order.
const candidates = narrowOverloadCandidates(
overloads,
site.arity,
site.argumentTypes,
conversionRankFn,
);
const candidates = narrowOverloadCandidates(overloads, site.arity, site.argumentTypes, {
argumentTypeClasses: site.argumentTypeClasses,
conversionRankFn: hookCtx?.conversionRankFn,
constraintCompatibility: hookCtx?.constraintCompatibility,
});
if (candidates.length !== 1) return undefined;
return candidates[0];
}
@@ -25,15 +25,26 @@
* counts as a match. Mismatches disqualify. A non-empty typed
* result wins; otherwise return the arity-filtered candidates.
* 4b. When the exact-type filter from step 4 returns empty AND a
* `conversionRankFn` is provided, rank candidates via pairwise
* dominance comparison (ISO C++ [over.ics.rank]): F1 beats F2
* only when F1 is not worse for every arg and better for at
* least one. Non-dominated candidates are returned; multiple
* survivors are genuinely ambiguous.
* `conversionRankFn` is provided (via `hookCtx`), rank candidates
* via pairwise dominance comparison (ISO C++ [over.ics.rank]):
* F1 beats F2 only when F1 is not worse for every arg and better
* for at least one. Non-dominated candidates are returned;
* multiple survivors are genuinely ambiguous.
* 4c. Final per-candidate constraint filter (SFINAE / `requires`).
* When `constraintCompatibility` is provided via `hookCtx`, drop
* candidates whose template constraints provably fail at the
* call site. Three-valued; `'unknown'` keeps the candidate
* (monotonicity).
* 5. Empty input returns empty output.
*/
import type { SymbolDefinition } from 'gitnexus-shared';
import type {
ArityVerdict,
Callsite,
ConstraintContext,
ParameterTypeClass,
SymbolDefinition,
} from 'gitnexus-shared';
/**
* Per-slot conversion-rank function. Returns a numeric cost for
@@ -46,13 +57,43 @@ import type { SymbolDefinition } from 'gitnexus-shared';
* Each language provides its own implementation. The function operates
* on normalized type strings (output of the language's type normalizer).
*/
export type ConversionRankFn = (argType: string, paramType: string) => number;
export type ConversionRankFn = (
argType: string,
paramType: string,
argTypeClass?: ParameterTypeClass,
paramTypeClass?: ParameterTypeClass,
) => number;
/**
* Optional hook bundle for narrowing extension points. Threaded in
* from `pickOverload` / `pickImplicitThisOverload` so per-language
* narrowing can layer in conversion-rank scoring (#1606) and
* constraint filtering (#1579) without changing the call signature
* at every site. Each hook is independently optional — leaving both
* undefined preserves the legacy arity + exact-type behavior.
*/
export interface OverloadNarrowingHookCtx {
/** Shape-preserving per-argument sidecar aligned with `argTypes`. */
readonly argumentTypeClasses?: ConstraintContext['argumentTypeClasses'];
/** Conversion-rank scoring fallback (step 4b). Engages when the
* exact-type filter rejects every candidate. */
readonly conversionRankFn?: ConversionRankFn;
/** Constraint filter (step 4c). Drops candidates whose template
* guards (SFINAE `enable_if_t`, C++20 `requires`, future Rust
* trait bounds, etc.) provably fail at the call site. Three-valued
* — `'unknown'` keeps the candidate (monotonicity). */
readonly constraintCompatibility?: (
callsite: Callsite,
def: SymbolDefinition,
ctx: ConstraintContext,
) => ArityVerdict;
}
export function narrowOverloadCandidates(
overloads: readonly SymbolDefinition[],
argCount: number | undefined,
argTypes: readonly string[] | undefined,
conversionRankFn?: ConversionRankFn,
hookCtx?: OverloadNarrowingHookCtx,
): readonly SymbolDefinition[] {
if (overloads.length === 0) return [];
@@ -93,31 +134,99 @@ export function narrowOverloadCandidates(
const candidates: readonly SymbolDefinition[] =
arityMatches.length > 0 ? arityMatches : anyUnknownBounds ? overloads : [];
let result: readonly SymbolDefinition[] = candidates;
if (argTypes !== undefined && argTypes.length > 0) {
const typed = candidates.filter((d) => {
const params = d.parameterTypes;
if (params === undefined) return false;
for (let i = 0; i < argTypes.length && i < params.length; i++) {
if (argTypes[i] === '') continue;
if (argTypes[i] !== params[i]) return false;
if (
!exactTypeSlotMatches(
argTypes[i],
params[i],
hookCtx?.argumentTypeClasses?.[i],
d.parameterTypeClasses?.[i],
)
) {
return false;
}
}
return true;
});
if (typed.length > 0) return typed;
// ── Conversion-rank scoring (step 4b) ──────────────────────────
// The exact-type filter above rejected every candidate. When a
// per-language conversion-rank function is available, rank via
// pairwise dominance: F1 beats F2 only when F1 is not worse for
// every arg and better for at least one. Non-dominated candidates
// are returned; multiple survivors are genuinely ambiguous.
if (conversionRankFn !== undefined) {
const ranked = rankByConversion(candidates, argTypes, conversionRankFn);
if (ranked.length > 0) return ranked;
if (typed.length > 0) {
result = typed;
} else if (hookCtx?.conversionRankFn !== undefined) {
// ── Conversion-rank scoring (step 4b) ──────────────────────────
// The exact-type filter rejected every candidate. Rank via
// pairwise dominance: F1 beats F2 only when F1 is not worse for
// every arg and better for at least one. Non-dominated candidates
// are returned; multiple survivors are genuinely ambiguous. When
// ranking also yields empty, fall through to the arity-filtered
// `candidates` set — matches pre-#1606 behavior.
const ranked = rankByConversion(
candidates,
argTypes,
hookCtx.conversionRankFn,
hookCtx.argumentTypeClasses,
);
if (ranked.length > 0) result = ranked;
}
}
return candidates;
// Constraint filter (step 4c; Tier-A — SFINAE / `requires` clauses).
// Runs after arity, exact-type, and conversion-rank filters so the
// hook only sees candidates already viable on the other axes.
// Three-valued: `'compatible'` and `'unknown'` keep the candidate
// (monotonicity — adding a predicate must never cause a wrong edge);
// only `'incompatible'` drops it. Candidates without
// `templateConstraints` are always kept.
//
// No fallback to the unconstrained set when this filter empties the
// candidate list: a fully-`'incompatible'` verdict is authoritative.
// The downstream `OVERLOAD_AMBIGUOUS` sentinel still guards the empty
// case, so a buggy hook that wrongly returns `'incompatible'` for
// every candidate degrades to today's "suppress edge" behavior rather
// than emitting a wrong edge.
if (hookCtx?.constraintCompatibility !== undefined && argCount !== undefined) {
const callsite: Callsite = { arity: argCount };
const ctx: ConstraintContext =
argTypes !== undefined
? {
argumentTypes: argTypes,
...(hookCtx.argumentTypeClasses !== undefined
? { argumentTypeClasses: hookCtx.argumentTypeClasses }
: {}),
}
: {};
result = result.filter((def) => {
if (def.templateConstraints === undefined) return true;
return hookCtx.constraintCompatibility!(callsite, def, ctx) !== 'incompatible';
});
}
return result;
}
function exactTypeSlotMatches(
argType: string,
paramType: string,
argTypeClass?: ParameterTypeClass,
paramTypeClass?: ParameterTypeClass,
): boolean {
if (argType !== paramType) return false;
// C++ normalizes away pointer markers (`int*` -> `int`). When both sides
// provide shape sidecars, do not let that collapse make `int` exactly match
// `int*`. Unknown sidecar evidence preserves the previous string-only path.
if (argTypeClass === undefined || paramTypeClass === undefined) return true;
if (argTypeClass.indirection === 'unknown' || paramTypeClass.indirection === 'unknown') {
return true;
}
return isPointerShape(argTypeClass) === isPointerShape(paramTypeClass);
}
function isPointerShape(typeClass: ParameterTypeClass): boolean {
return typeClass.indirection === 'pointer' && typeClass.pointerDepth > 0;
}
/**
@@ -136,6 +245,7 @@ function rankByConversion(
candidates: readonly SymbolDefinition[],
argTypes: readonly string[],
rankFn: ConversionRankFn,
argTypeClasses?: readonly ParameterTypeClass[],
): readonly SymbolDefinition[] {
// Step 1: compute per-slot ranks and exclude non-viable candidates.
const viable: Array<{ def: SymbolDefinition; ranks: number[] }> = [];
@@ -144,12 +254,22 @@ function rankByConversion(
if (params === undefined) continue;
const ranks: number[] = [];
let ok = true;
for (let i = 0; i < argTypes.length && i < params.length; i++) {
for (let i = 0; i < argTypes.length; i++) {
const paramType = parameterTypeAt(params, i);
if (paramType === undefined) {
ok = false;
break;
}
if (argTypes[i] === '') {
ranks.push(0); // unknown arg → any-match (rank 0)
continue;
}
const r = rankFn(argTypes[i], params[i]);
const r = rankFn(
argTypes[i],
paramType,
argTypeClasses?.[i],
parameterTypeClassAt(d.parameterTypeClasses, i),
);
if (!isFinite(r)) {
ok = false;
break;
@@ -176,6 +296,20 @@ function rankByConversion(
return viable.filter((_, idx) => !dominated.has(idx)).map((v) => v.def);
}
function parameterTypeAt(params: readonly string[], argIndex: number): string | undefined {
if (argIndex < params.length) return params[argIndex];
return params[params.length - 1] === '...' ? '...' : undefined;
}
function parameterTypeClassAt(
params: readonly ParameterTypeClass[] | undefined,
argIndex: number,
): ParameterTypeClass | undefined {
if (params === undefined) return undefined;
if (argIndex < params.length) return params[argIndex];
return params[params.length - 1]?.base === '...' ? params[params.length - 1] : undefined;
}
/**
* Compare two per-slot rank vectors.
* Returns -1 if `a` dominates `b` (not worse everywhere, better somewhere),
@@ -21,6 +21,11 @@
* but not a namespace prefix → compound resolver
* 7. **Case 4 (simple typeBinding)** — `typeRef.rawName` has no dot →
* MRO walk + `findOwnedMember`
* 8. **Case 5 (value-receiver bridge)** — receiver is a `Const`/`Variable`
* whose `nodeId` is referenced as an `ownerId` in `model.methods`
* (object-literal services). Last-resort fallback for lowercase
* receivers with no class-like or type-binding match. Mirrors
* the legacy DAG bridge in `call-processor.ts`.
*
* Reordering or merging cases changes resolution semantics.
*
@@ -46,9 +51,10 @@ import {
findExportedDef,
findOwnedMember,
findReceiverTypeBinding,
findValueBindingInScope,
isClassLike,
} from '../scope/walkers.js';
import { tryEmitEdge } from '../graph-bridge/edges.js';
import { tryEmitEdge, tryEmitEdgeWithExplicitTargetId } from '../graph-bridge/edges.js';
import { resolveCompoundReceiverClass } from '../passes/compound-receiver.js';
import { resolveDefGraphId } from '../graph-bridge/ids.js';
import {
@@ -74,6 +80,7 @@ type ReceiverBoundProviderSubset = Pick<
| 'resolveQualifiedReceiverMember'
| 'resolveThisViaEnclosingClass'
| 'conversionRankFn'
| 'constraintCompatibility'
>;
function normalizeTemplateArgToken(value: string): string {
@@ -344,7 +351,11 @@ export function emitReceiverBoundCalls(
methodOverloads,
site.arity,
site.argumentTypes,
provider.conversionRankFn,
{
argumentTypeClasses: site.argumentTypeClasses,
conversionRankFn: provider.conversionRankFn,
constraintCompatibility: provider.constraintCompatibility,
},
);
if (isOverloadAmbiguousAfterNormalization(narrowed, site.arity)) {
ambiguous = true;
@@ -648,13 +659,7 @@ export function emitReceiverBoundCalls(
let memberDef: SymbolDefinition | undefined;
let ambiguous = false;
for (const ownerId of chain) {
const picked = pickOverload(
ownerId,
memberName,
site,
model,
provider.conversionRankFn,
);
const picked = pickOverload(ownerId, memberName, site, model, provider);
if (picked === OVERLOAD_AMBIGUOUS) {
ambiguous = true;
break;
@@ -707,6 +712,61 @@ export function emitReceiverBoundCalls(
}
}
}
// ── Case 5: value-receiver bridge (object-literal services) ──
// When prior cases couldn't resolve the receiver as a class or
// type binding, fall back to value-binding resolution. Covers:
//
// export const fooService = { getUser(id) {...} };
// import { fooService } from './service';
// fooService.getUser(id); // ← resolve here
//
// `fooService` is a `Const`/`Variable` (not class-like, no typeBinding
// for unannotated literals), so Cases 2-4 skip it. Scope-resolution
// defs for non-class values carry a synthetic id, so we translate to
// the canonical graph node ID via `resolveDefGraphId` before owner-
// indexed lookup — the parser writes the graph node ID as `ownerId`
// on the method symbol-table entry to match.
//
// Object-literal methods do not carry a `qualifiedName` (no class
// owner to seed it), so the picked def cannot round-trip through
// `tryEmitEdge` → `resolveDefGraphId`. We use
// `tryEmitEdgeWithExplicitTargetId` instead, passing `picked.nodeId`
// directly — same dedup-key shape, collapse-flag honoring, and
// caller resolution as `tryEmitEdge`.
const valueDef = findValueBindingInScope(site.inScope, receiverName, scopes);
if (valueDef !== undefined) {
const ownerGraphId =
resolveDefGraphId(valueDef.filePath, valueDef, nodeLookup) ?? valueDef.nodeId;
const picked = pickOverload(ownerGraphId, memberName, site, model, provider);
if (picked === OVERLOAD_AMBIGUOUS) {
handledSites.add(siteKey);
continue;
}
if (picked !== undefined) {
const reason =
site.kind === 'write' || site.kind === 'read'
? site.kind
: picked.filePath !== parsed.filePath
? 'import-resolved'
: 'global';
const confidence = site.kind === 'write' || site.kind === 'read' ? 1.0 : 0.85;
const ok = tryEmitEdgeWithExplicitTargetId(
graph,
scopes,
nodeLookup,
site,
picked.nodeId,
reason,
seen,
confidence,
collapse,
);
if (ok) emitted++;
handledSites.add(siteKey);
continue;
}
}
}
}
@@ -722,7 +782,7 @@ function pickOverload(
memberName: string,
site: ParsedFile['referenceSites'][number],
model: SemanticModel,
conversionRankFn?: (argType: string, paramType: string) => number,
provider: ReceiverBoundProviderSubset,
): SymbolDefinition | typeof OVERLOAD_AMBIGUOUS | undefined {
const overloads = model.methods.lookupAllByOwner(ownerId, memberName);
if (overloads.length === 0) {
@@ -733,12 +793,11 @@ function pickOverload(
}
if (overloads.length === 1) return overloads[0];
const candidates = narrowOverloadCandidates(
overloads,
site.arity,
site.argumentTypes,
conversionRankFn,
);
const candidates = narrowOverloadCandidates(overloads, site.arity, site.argumentTypes, {
argumentTypeClasses: site.argumentTypeClasses,
conversionRankFn: provider.conversionRankFn,
constraintCompatibility: provider.constraintCompatibility,
});
// When narrowing leaves >1 candidate that share identical normalized
// parameter-types (e.g., C++ `f(int)` vs `f(long)` both collapsed to
// `['int']` by `normalizeCppParamType`), suppress the edge entirely.

Some files were not shown because too many files have changed in this diff Show More