Compare commits

..
Author SHA1 Message Date
gitnexus-release-bot[bot] db9d749d73 release: v1.6.6-rc.26 2026-05-20 17:03:09 +00:00
azizur100389andGergő Magyar aa8f4d6efe fix(group): Union HTTP graph and source contracts (#1709)
* Union HTTP graph and source contracts

* test(group): Document HTTP source union follow-ups

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 17:44:07 +01:00
Shane Thurston Wijaya 4d2ed0e525 fix(eval-server): localhost now doesn't normalize into IPv4 instead lets OS decide which to bind (#1722)
* fix(eval-server): localhost now doesn't normalize into IPv4 instead lets OS decide which to bind

* fix(eval-server): EADDRNOTAVAIL now treats as potential IPv6

* test(eval-server): new integration test for --host localhost

* docs(eval-server): updated eval/README.md based on latest update

* fix(eval-server): clarify EADDRNOTAVAIL diagnostic, guard server.address(), and soften localhost docs
2026-05-20 16:14:13 +01:00
df1882d36b fix(ingestion): surface skipped large-file paths by default (#1659) (#1661)
* fix(ingestion): surface skipped large-file paths by default (#1659)

The 512 KB skip threshold in filesystem-walker is necessary, but the
existing warning only said "Skipped N large files" with no paths unless
GITNEXUS_VERBOSE=1 was set. In a repo with one or two oversized first-
party source files (e.g. a 17K-line cron handler), every IMPORTS/CALLS
edge from that file silently disappeared and the surface looked like a
Python resolver bug. Issue #1659 was filed against the resolver for
exactly that reason, but the resolver was fine; the file was being
dropped before parse.

Changes:
  * Always print up to 5 skipped paths after the count line.
  * If more than 5 were skipped, append "...and N more" with a hint to
    set GITNEXUS_VERBOSE=1 for the full list.
  * When running at the default threshold, emit a one-line hint about
    GITNEXUS_MAX_FILE_SIZE=<KB> so operators know how to widen it.
  * Cover the new behavior with three additional tests in the existing
    filesystem-walker integration suite, plus a new describe block for
    the >5 preview-cap case.

Verified end-to-end on a 680-file Python repo that hit #1659: before
the patch, "Skipped 3 large files (>512KB, ...)" was the only signal
and impact upstream of a function called from cron.py returned 1 of 5
real callers; after the patch the cron file is listed by name with the
hint, and running with GITNEXUS_MAX_FILE_SIZE=1024 brings the missing
callers back (impactedCount 1 -> 9).

* fix(ingestion): address #1661 adversarial review follow-ups (F1/F2/F3)

Three non-blocking nits flagged by the adversarial review on #1661:

F1 (output stability) — skippedLargePaths was populated by concurrent
fs.stat callbacks in batches of 32, so push order within a batch was
completion-order rather than input-order. The default preview's "first
5" could vary across runs on the same repo. Fix: sort the array before
slicing. New test asserts the verbose output is in sorted order.

F2 (boundary coverage) — the preview-cap describe block created 8
large files, so the SKIPPED_PREVIEW_CAP = 5 comparison was never
exercised at the exact <= boundary. A future off-by-one (<= → <) would
not fail the suite. Fix: add two tests, one with exactly 5 files (all
listed, no truncation) and one with exactly 6 files (5 listed plus
"...and 1 more").

F3 (hint accuracy) — isDefault compared effective bytes, so an
operator who explicitly set GITNEXUS_MAX_FILE_SIZE=512 (the same KB as
the default) would still see the "Set GITNEXUS_MAX_FILE_SIZE=<KB>..."
hint. Fix: gate the hint on whether the env var is unset, not on the
resulting byte value. New test pins the explicit-default-value case.

All 34 filesystem-walker tests pass (was 30; +4 new). Prettier clean,
typecheck clean for the changed files.

---------

Co-authored-by: scotjelinski <58397194+scotjelinski@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 14:37:58 +01:00
f350ae278a feat: Add analyze --repair-fts, enforce FTS verification, and harden repair safeguards (#1720)
* Initial plan

* feat(analyze): add --repair-fts and verify FTS index rebuilds

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/dccb3673-af86-43aa-aede-2e1449399775

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor(fts): tighten repair/verify messaging and option naming

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/dccb3673-af86-43aa-aede-2e1449399775

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: highlight analyze --repair-fts vs --force in READMEs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/61edc967-debc-419f-9f51-aebf2ef08d22

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(analyze): guard repair mode against missing graph store

* fix(cli): reject --repair-fts with --force

* test(analyze): document repair-store fixture intent

* test(analyze): tidy repair failure fixtures and constants

* test(analyze): clarify mock constants in repair tests

* test(analyze): rename simulated missing-index constant

* test(analyze): clarify mocked graph shape in full-verify test

* refactor(analyze): finalize flag validation and test clarity

* test(skip-git): avoid hard failing when FTS extension is unavailable

* test(skip-git): log visible FTS-unavailable test skips

* test(skip-git): tighten FTS-unavailable error detection

* test(skip-git): simplify FTS-unavailable message checks

* test(skip-git): avoid HOME pointing at parent repo in fixture env

* fix(analyze): address Claude follow-up findings for repair guardrails

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* fix(repair-fts): clarify invalid graph-store preflight errors

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* test(analyze): strengthen assertions for conflict and missing-store errors

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* fix(repair-fts): make invalid graph-store type errors explicit

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* fix(repair-fts): improve graph-store type diagnostics

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 13:37:04 +01:00
azizur100389andGergő Magyar dae70a26ea feat(cpp): Add pointer nullptr ellipsis conversion ranks (#1708)
* Add C++ pointer null ellipsis ranks

* test(cpp): Strengthen pointer overload assertions

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 12:06:51 +01:00
CopilotandGergő Magyar b4a2a4b91e fix(ingestion): Prioritize same-module Java type resolution for duplicate FQNs across modules (#1712)
* Initial plan

* Fix Java same-name type resolution with same-module priority

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/df0843e3-e244-4e0f-a94a-311df3899bd0

* Refine Java ambiguity fallback safety check

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/df0843e3-e244-4e0f-a94a-311df3899bd0

* Remove Java-specific fallback from shared scope walkers

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e882906a-2c96-411e-94a5-123a345421a9

* Harden Java module key and ambiguous owner fallback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e882906a-2c96-411e-94a5-123a345421a9

* Add negative assertions for duplicate-FQN module edges

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e882906a-2c96-411e-94a5-123a345421a9

* Make Java same-module ordering path-agnostic

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Refine generic Java path-affinity ordering safeguards

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Polish Java path-affinity ordering clarity

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Simplify Java path-affinity ordering logic

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Revert legacy DAG Java ambiguity ordering changes

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/94e50cf2-9733-4e69-a0eb-9fd38cbdb589

* Skip duplicate-FQN Java assertions in legacy parity mode

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/1b560efa-1b3b-4697-b590-c6ef447f431e

* Tighten duplicate-FQN Java CALLS edge cardinality assertions

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/67c18f93-5e56-4b15-8404-cdf1be9b4485

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 08:00:03 +01:00
Nilotpal KashyapandGergő Magyar d7e1815aa3 fix(detect-changes): guard resolveWorktreeCwd against overriding a separately-indexed worktree (#1691)
* fix(detect-changes): guard resolveWorktreeCwd against overriding a separately-indexed worktree

When the repo registry entry points to a linked worktree (both main
checkout and worktree indexed separately), resolveWorktreeCwd was
incorrectly replacing the correct worktree repoPath with the server's
main-checkout launch directory. Both share the same canonical root so
the existing same-repo check passed, causing git diff to run from the
wrong directory and return 0 changes (issue #1659).

Fix: early-exit guard — if tryRealpath(repoPath) differs from
tryRealpath(getCanonicalRepoRoot(repoPath)), repoPath is itself a
linked worktree and is returned unchanged. Auto-detection only fires
when repoPath equals the canonical main-checkout root.

Also normalises the launchCanonical comparison in the auto-detect path
to use tryRealpath for cross-platform consistency.

Regression test: 'returns worktreeDir unchanged when repoPath IS a
linked worktree and launchCwd is the main checkout'.

* test(detect-changes): add worktreeA→worktreeB case and assumption comment

Cover the missing case from the production-readiness review:
repoPath = wt-A (indexed), launchCwd = wt-B (server on a different
linked worktree). The guard fires on repoPath being a worktree
regardless of launchCwd, so wt-A is returned unchanged.

Also add an inline comment documenting the assumption that repoPath
is a git root or linked-worktree root (not an arbitrary subdirectory),
as noted in Finding 2 of the review.

* refactor(detect-changes): validate repoPath is a git root before canonical comparison

Instead of relying on a comment asserting repoPath is always a git
root, call getGitRoot(repoPath) first. Only if the result matches
repoPath itself do we call getCanonicalRepoRoot and apply the guard.

This eliminates the over-classification risk for subdirectory repoPath
values and makes the assumption explicit in code. repoCanonical is
shared across both the guard and the auto-detect block.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 06:46:16 +01:00
dependabot[bot] 92ad0f5491 chore(deps): bump idna in /eval in the uv group across 1 directory (#1713) 2026-05-20 05:38:26 +01:00
LocallyInsaneDBandGergő Magyar 803f0bed5f fix(lbug): probe-then-load FTS extension on Windows (#1690) (#1692)
* fix(lbug): probe-then-load FTS extension on Windows (#1690)

The Windows skip-on-process.platform==='win32' guard in pool-adapter.ts
hard-skipped loadFTSExtension() for every Windows host, even when the
FTS extension binary was already present locally at
~/.lbdb/extension/<version>/win_amd64/fts/libfts.lbug_extension.

That left BM25 silently degraded on Windows hosts that had a working
extension on disk, with no error path — `gitnexus doctor` still reported
FTS as available, but query returned 0 BM25 hits.

This patch adds hasLocalWinFtsExtension() which probes
~/.lbdb/extension/*/win_amd64/fts/ before the Windows skip. When a binary
is on disk we call loadFTSExtension(..., { policy: 'load-only' }); the
crashing install path documented in #1199 / #1217 is never exercised at
query time, and LadybugDB's version-specific resolution combined with
the ExtensionManager's tryLoad try/catch handles stale or zero-byte
sibling version dirs cleanly (no dlopen attempted on a stale binary).
When no binary is on disk at all, we fall back to the upstream skip so
install-time SIGSEGV continues to be avoided.

Verified on Windows 10 + Node 22.19.0 + gitnexus 1.6.5 +
@ladybugdb/core 0.16.1 with the FTS extension cached at 0.16.0:

  * BM25 timing goes from 0 → ~250-326ms on previously-zero queries
  * gitnexus context / impact / cypher unaffected
  * Adversarial-mixed-state run (real 0.16.0 binary + zero-byte stubs at
    0.15.0, 0.16.1, 0.17.0): exits 0, no SIGSEGV, FTS resolves to the
    real 0.16.0 binary, BM25 returns real hits
  * Stub-only state at the resolution path (0.16.0, zero-byte): exits 0,
    emits "FTS extension unavailable; load-only policy: extension not
    pre-installed", FTS marked unavailable cleanly via markUnavailable
    in extension-loader.ts — no silent greenlight

Closes #1690

* test(lbug): cover hasLocalWinFtsExtension probe + format pool-adapter

- Export hasLocalWinFtsExtension and add lbug-pool-win-fts-probe.test.ts
  with 7 cases against a real tmpdir + os.homedir spy:
    * missing ~/.lbdb/extension dir -> false
    * extension root present but no version dirs -> false
    * one version dir with binary present -> true
    * zero-byte stub at probe path -> true (LOAD failure handled downstream)
    * multi-version with binary only in a non-first dir -> true
    * multi-version with no binary anywhere (Nix/Bazel/MDM tree) -> false
    * fs.readdir throws (EACCES) -> false

  The Windows conditional in doInitLbug / initLbugWithDb is intentionally
  not unit-isolated: it reduces to `probe ? load : true` over a fully
  constructed lbug.Database + Connection pool, which the
  test/integration/lbug-pool*.test.ts suites already exercise on the
  windows-latest CI matrix.

- Apply prettier format to the fs.stat() call in pool-adapter.ts,
  resolving the quality/format CI failure surfaced by gitnexus/autofix.

Addresses DoD §2.7 test-coverage blocker raised in the production-
readiness review on #1692, and the dir-exists-no-file regression case
raised on #1690.

Refs #1690.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-19 12:09:21 +01:00
55f8d442f6 fix(mcp): setup fallback on Windows when global gitnexus resolves to a non-spawnable shim (#1694)
* Initial plan

* fix: avoid invalid Windows MCP shim paths

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a052306e-483a-42d0-b65a-2646906457c7

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: cover .ps1 windows mcp fallback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/5aaed570-2a0b-4ed9-a0ac-ca099ce5675e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: assert windows fallback for cursor and codex

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/5aaed570-2a0b-4ed9-a0ac-ca099ce5675e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-19 08:17:26 +01:00
18167400c4 chore(deps)(deps): bump express and @types/express in /gitnexus (#872)
* chore(deps)(deps): bump express and @types/express in /gitnexus

Bumps [express](https://github.com/expressjs/express) and [@types/express](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/express). These dependencies needed to be updated together.

Updates `express` from 4.22.1 to 5.2.1
- [Release notes](https://github.com/expressjs/express/releases)
- [Changelog](https://github.com/expressjs/express/blob/master/History.md)
- [Commits](https://github.com/expressjs/express/compare/v4.22.1...v5.2.1)

Updates `@types/express` from 4.17.25 to 5.0.6
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/express)

---
updated-dependencies:
- dependency-name: "@types/express"
  dependency-version: 5.0.6
  dependency-type: direct:development
  update-type: version-update:semver-major
- dependency-name: express
  dependency-version: 5.2.1
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

* fix(server): normalize jobId param for Express 5 SSE routes

Express 5 types req.params values as string | string[]. mountSSEProgress uses a dynamic route path so TypeScript cannot narrow jobId; assert it once with assertString and reuse in the SSE progress callback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(ci): retrigger CI

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-19 07:43:07 +01:00
dad1ca7ab5 chore(deps)(deps): bump zod from 3.25.76 to 4.3.6 in /gitnexus-web (#1464)
Bumps [zod](https://github.com/colinhacks/zod) from 3.25.76 to 4.3.6.
- [Release notes](https://github.com/colinhacks/zod/releases)
- [Commits](https://github.com/colinhacks/zod/compare/v3.25.76...v4.3.6)

---
updated-dependencies:
- dependency-name: zod
  dependency-version: 4.3.6
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:56:04 +01:00
c746f30c90 chore(deps)(deps): bump langsmith (#1552)
Bumps the npm_and_yarn group with 1 update in the /gitnexus-web directory: [langsmith](https://github.com/langchain-ai/langsmith-sdk).


Updates `langsmith` from 0.5.23 to 0.6.3
- [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases)
- [Commits](https://github.com/langchain-ai/langsmith-sdk/commits)

---
updated-dependencies:
- dependency-name: langsmith
  dependency-version: 0.6.3
  dependency-type: indirect
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:55:44 +01:00
15a667ae5e chore(deps)(deps-dev): bump @vitest/coverage-v8 in /gitnexus (#1604)
Bumps [@vitest/coverage-v8](https://github.com/vitest-dev/vitest/tree/HEAD/packages/coverage-v8) from 4.1.5 to 4.1.6.
- [Release notes](https://github.com/vitest-dev/vitest/releases)
- [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.6/packages/coverage-v8)

---
updated-dependencies:
- dependency-name: "@vitest/coverage-v8"
  dependency-version: 4.1.6
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:55:24 +01:00
6210d80f1e chore(deps)(deps-dev): bump tsx from 4.21.0 to 4.21.1 in /gitnexus (#1698)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.21.0 to 4.21.1.
- [Release notes](https://github.com/privatenumber/tsx/releases)
- [Changelog](https://github.com/privatenumber/tsx/blob/master/release.config.cjs)
- [Commits](https://github.com/privatenumber/tsx/compare/v4.21.0...v4.21.1)

---
updated-dependencies:
- dependency-name: tsx
  dependency-version: 4.21.1
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:55:04 +01:00
637cfca39c chore(deps)(deps): bump express-rate-limit in /gitnexus (#1697)
Bumps [express-rate-limit](https://github.com/express-rate-limit/express-rate-limit) from 8.5.1 to 8.5.2.
- [Release notes](https://github.com/express-rate-limit/express-rate-limit/releases)
- [Commits](https://github.com/express-rate-limit/express-rate-limit/compare/v8.5.1...v8.5.2)

---
updated-dependencies:
- dependency-name: express-rate-limit
  dependency-version: 8.5.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:54:33 +01:00
dependabot[bot]andGergő Magyar 73543a4714 chore(deps)(deps-dev): bump @types/node in /gitnexus (#1696)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 25.6.2 to 25.7.0.
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node)

---
updated-dependencies:
- dependency-name: "@types/node"
  dependency-version: 25.7.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-19 06:54:14 +01:00
DuduPhudu b37974fdac feat(javascript): migrate JavaScript to scope-based resolution (RFC #909 Ring 3, issue #928) (#1640) 2026-05-19 06:23:13 +01:00
dependabot[bot] ade2069633 chore(deps)(deps): bump brace-expansion from 5.0.5 to 5.0.6 in /gitnexus (#1689) 2026-05-19 05:35:27 +01:00
azizur100389 5f0c0eba0e feat(cpp): Expand type_traits constraint registry (#1648) 2026-05-18 21:10:18 +01:00
Gergő Magyar 2632bcccc0 fix(api): open lbug read-only for /api/graph, /api/search, /api/grep (#1686) 2026-05-18 19:57:35 +01:00
Gergő Magyar c9199b654f fix(test): retry Windows temp cleanup in cli-e2e teardown (#1688) 2026-05-18 18:17:54 +01:00
Shane Thurston Wijaya 33f18ceaa2 feat(eval-server): added --host for user configured host IP instead of system hardcoded IP (127.0.0.1) (#1667)
* feat(eval-server): added --host for user configured host IP instead of system hardcoded IP (127.0.0.1)

* fix(eval-server): localhost value in --host now returns 127.0.0.1 instead of the raw input to fix wrong address, handled error for ipv6 disabled containers

* feat(eval-server): add --host flag with validation and error handling

  Co-Authored-By: Val Vladescu <val.vladescu@thirdbridge.com>

* fix(eval-server): bracketed IPv6 addresses to remove ambiguity

* docs(eval-server): document --host flag, READY signal format, and parser migration note

* fix(eval-server): use actual bound port in READY signal; strengthen --host e2e tests

  Co-Authored-By: Val Vladescu <val.vladescu@thirdbridge.com>

* feat(eval): wire eval-server --host through gitnexus_docker.py

* docs(eval): added guidance for docker user

* docs(eval): revise the imprecise documentation

* fix(e2e): updated original stdout for new format
2026-05-18 16:00:42 +01:00
c30833fad3 perf(scope-resolution): use owner-keyed lookup for Step 2 member resolution (#1657)
* perf(scope-resolution): use owner-keyed lookup for Step 2 member resolution (#1656)

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(scope-resolution): index Const/Static in FieldRegistry for Step 2 lookup

Extend FieldRegistry to hold multiple defs per (owner, name), reconcile Const and Static into the owner-keyed index, and wire lookupAllByOwner through the production hook so Step 2 does not drop field kinds the registry never indexed. Pass explicitReceiver on read/write reference sites and document undefined-vs-empty hook semantics for defs fallback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* perf(scope-resolution): centralize O(1) owned-member hook and guard hot path

Extract lookupOwnedMembersByOwner for the production Step 2 hook so merges stay O(1) per registry with no defs.byId scan. Add a perf-contract unit test that throws if byId.values runs when the hook is wired. Reuse a frozen empty sentinel on double miss to avoid per-probe allocations.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: drop unused buildFieldRegistry import

* chore(scope-resolution): apply ce-code-review safe_auto fixes

- Drop unreachable return + unused values() capture in perf-contract trap (Finding #7)
- Type lookupOwnedMembersByOwner ownerDefId as DefId (Finding #9)
- Add Static-kind Step 2 lookup test mirroring the Const case (Finding #11)

* docs(field-registry): document lookupFieldByOwner first-wins semantics

Audit of all 6 production callers (call-processor.ts:2279, walkers.ts:535,
receiver-bound-calls.ts:380+730, type-env.ts:627+631) confirms none depends
on last-wins precedence — all treat the return as a generic 'field with
this name owned by this class'. Clarify the JSDoc to surface the semantic
change introduced when FieldRegistry moved from last-wins to append-order
storage (ce-code-review finding #2).

* test(scope-resolution): extend Step 2 perf contract to implicit-self, MRO, field paths

Adds three sibling tests under the Step 2 perf contract describe block, each
asserting defs.byId.values() does NOT execute when ownedMembersByOwner is wired:

- implicit-self receiver via typeBindings.self (no explicitReceiver branch)
- 2-level MRO chain (Child extends Parent, save resolves on Parent at depth 1)
- FieldRegistry read via Step 2 (property lookup, separate registry path)

Pins the perf invariant on every distinct entry into walkReceiverTypeBinding
so a regression bypassing the hook on any sub-path now fails CI immediately
(ce-code-review finding #8).

* test(resolve-references): cover arity-overload filtering via resolveReferenceSites

Pins the orchestration-layer wiring of providers.arityCompatibility:
hook returns [save(arity 1), save(arity 2)], referenceSite.arity = 1,
arityCompatibility verdicts 'compatible'/'incompatible' by parameterCount,
exactly one reference emitted with toDef = the arity-1 overload.

registries.test.ts already covered arity at the buildMethodRegistry level;
this adds the missing entry-point check that resolveReferenceSites threads
providers correctly through to lookupCore.Step5 (ce-code-review finding #10).

* test(resolve-references): add hook-on vs hook-off parity test

Runs resolveReferenceSites twice on the same fixture (Parent.save method
hit + Child.name field hit, Child extends Parent MRO chain) — once with
ownedMembersByOwner wired to a synthetic registry, once with the hook
absent so collectOwnedMembers takes the defs.byId fallback. Asserts:

- stats are identical (sitesProcessed / referencesEmitted / unresolved)
- referenceIndex.bySourceScope entries have equal length
- toDef sets are equal
- each per-site reference (including evidence and depth) is .toEqual

Locks the semantic-parity claim in code while both paths still exist.
Will be removed alongside the fallback in finding #1 (ce-code-review #3).

* test(typescript): probe Step 2 MRO walk against ambient (declare class) base

Adds typescript-ambient-base-class fixture with an export declare class
AmbientBase + Derived extends AmbientBase and a call site d.ambientMethod().
Integration assertions:

- Both classes are detected
- EXTENDS edge Derived → AmbientBase emitted
- CALLS edge to ambient.ts:ambientMethod resolved via MRO walk

Probes the ce-code-review #6 concern that ambient-only owners (whose
bodies are never parsed) might be silently skipped by Step 2 after the
owner-keyed lookup change. Result: the call resolves correctly — the
method signature inside the declare class body still flows through
reconcileOwnership into model.methods, so the hook returns the right
ancestor hits. Residual risk is empirically closed.

* feat(scope-resolution): route nested types via owner-keyed TypeRegistry

Closes the Step 2 contract footgun where 'hook returns [] = authoritative
miss' silently dropped any owned def whose NodeLabel was outside the
method/field if-chain in reconcileOwnership.

- TypeRegistry: add nestedByOwner Map + lookupAllByOwner(owner, simple)
  + registerByOwner(owner, simple, def). Mirrors MethodRegistry/
  FieldRegistry shape; cleared with the rest on cascade clear.
- reconcileOwnership: route class-like NodeLabels (Class/Interface/Enum/
  Struct/Union/Trait/TypeAlias/Typedef/Record/Delegate/Annotation/
  Template/Namespace) via types.registerByOwner. New nestedTypesRegistered
  stat. Idempotent skip via nodeId match.
- validateOwnershipParity: extend the I9 invariant check to nested types.
- lookupOwnedMembersByOwner: merge methods + fields + nested-type hits;
  short-circuit when any one source contributes the full result.

Unblocks future receiver-MRO registries that need to resolve 'Outer.Inner'
through the receiver's type-binding chain (ce-code-review finding #5a).

* refactor(scope-resolution): make ownedMembersByOwner required; delete byId fallback

Per ce-code-review finding #1, the optional-hook design encoded a silent
O(|defs|) perf cliff into the type system: any RegistryContext built
without the hook regressed Step 2 to scanning every def per probe with
no warning. Production wires the hook unconditionally; the fallback was
exercised only by tests.

- RegistryContext.ownedMembersByOwner: required, returns readonly
  SymbolDefinition[] (no | undefined). Implementations MUST return [] on
  authoritative miss.
- collectOwnedMembers in lookup-core.ts collapses to a one-line forward
  to the hook; the defs.byId.values() scan and simpleNameOf helper are
  deleted (simpleNameOf had no other consumers).
- ResolveReferencesInput.ownedMembersByOwner: required to match.
- Tests: drop three fallback-path tests (registries Const fallback,
  resolveReferenceSites no-hook fallback, resolveReferenceSites Const-
  undefined fallback) and the hook-vs-fallback parity test added by
  finding #3. makeCtx in registries.test.ts now defaults to a real
  owner-keyed scan over the test fixture defs so tests that don't care
  about the hook keep working.

* perf(free-call-fallback): cache global callables by simple name once per pass

pickUniqueGlobalCallable scanned scopes.defs.byId.values() on every
free-call fallback site. After PR #1656 fixed Step 2, this scan became
the dominant remaining O(|defs|) hot path on large repos (ce-code-review
finding #4).

- buildGlobalCallableIndex builds a Map<simpleName, SymbolDefinition[]>
  over scopes.defs once at the top of emitFreeCallFallback. Same filter
  the per-site scan applied: Function / Method / Constructor, keyed by
  the last .-segment of qualifiedName.
- pickUniqueGlobalCallable consumes the prebuilt index via O(1) Map.get
  instead of iterating every def. Per-site complexity drops from
  O(|defs|) to O(|defs with this simple name|).
- Cost: O(|defs|) once per pass instead of O(|defs| * |free-call sites|).

Subsequent narrowing (arity, conversion-rank) and the model-side fallback
(model.symbols.lookupCallableByName + model.methods.lookupMethodByName)
are unchanged.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* ci: trigger build

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Gergo Magyar <abhigyan1.patwari@gmail.com>
2026-05-18 13:14:27 +01:00
Copilot 7d500390b9 fix: Use Ladybug native read-only enforcement and prepared statement execution for Cypher query paths (#1655) 2026-05-18 06:54:24 +01:00
Nilotpal Kashyap bdc0439a10 feat(detect-changes): support git worktrees (#1654) 2026-05-17 20:54:41 +01:00
Shane Thurston Wijaya 105efd0f7c feat(wiki): added --lang <lang> flags to gitnexus wiki for multilanguage wiki generation support (#1613) 2026-05-17 19:54:02 +01:00
Copilot 493827222d fix(ingestion): Raise analyze auto-heap to 16GB and tighten cross-platform OOM guidance for UE5-scale repositories (#1652) 2026-05-17 16:28:07 +01:00
Copilot ed50a6729f fix(wiki): Remove the hidden 60s default timeout, validate gitnexus wiki timeout/retry flags, and surface timeout errors (#1651) 2026-05-17 12:03:54 +01:00
125 changed files with 7990 additions and 1612 deletions
@@ -17,11 +17,11 @@ npx gitnexus analyze
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates CLAUDE.md / AGENTS.md context files.
| Flag | Effect |
| ------------------- | ------------------------------------------------------------------------------------------------------- |
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
| Flag | Effect |
| -------------- | ---------------------------------------------------------------- |
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
+41 -108
View File
@@ -62,131 +62,64 @@ Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
Indexed as **GitNexus** (4325 symbols, 10556 relationships, 300 execution flows). Use MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any tool warns the index is stale, run `npx gitnexus analyze` first.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** `gitnexus_impact({target: "symbolName", direction: "upstream"})` — report blast radius to the user.
- **MUST run `gitnexus_detect_changes()` before committing** — verify only expected symbols and flows are affected.
- **MUST warn the user** if impact returns HIGH or CRITICAL risk.
- Explore unfamiliar code with `gitnexus_query({query: "concept"})` (process-grouped, ranked) instead of grepping.
- Full context on a symbol: `gitnexus_context({name: "symbolName"})`.
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find related execution flows
2. `gitnexus_context({name: "<suspect function>"})` — callers, callees, process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace flow step by step
4. Regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})`
## When Refactoring
- **Rename:** `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Graph edits are safe; text_search edits need manual review.
- **Extract/Split:** `gitnexus_context` (incoming/outgoing refs) then `gitnexus_impact` (upstream callers) before moving code.
- **After any refactor:** `gitnexus_detect_changes({scope: "all"})` to verify scope.
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Never Do
- Edit a symbol without running `gitnexus_impact` first.
- Ignore HIGH/CRITICAL risk warnings.
- Rename with find-and-replace — use `gitnexus_rename`.
- Commit without `gitnexus_detect_changes()`.
- Add language-specific behavior to shared ingestion code (`gitnexus/src/core/ingestion/`) — use a `LanguageProvider` hook. Seeing `provider.mroStrategy === 'xxx'` or an import from `languages/xxx.ts` in shared code means stop and add a hook.
## Tools Quick Reference
| Tool | When to use | Example |
|------|-------------|---------|
| `list_repos` | Discover indexed repos | `gitnexus_list_repos({})` |
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
| `api_impact` | Pre-change API route impact | `gitnexus_api_impact({route: "/api/users", method: "GET"})` |
| `route_map` | Route → handler → consumer map | `gitnexus_route_map({})` |
| `tool_map` | MCP/RPC tool definitions | `gitnexus_tool_map({})` |
| `shape_check` | Response shape vs consumer access | `gitnexus_shape_check({route: "/api/users"})` |
| `group_list` | List repo groups | `gitnexus_group_list({})` |
| `group_sync` | Rebuild group Contract Registry | `gitnexus_group_sync({name: "myGroup"})` |
| `query` (group mode) | Cross-repo search in a group (RRF-merged) | `gitnexus_query({repo: "@myGroup", query: "auth"})` |
| `context` (group mode) | 360° view across all member repos | `gitnexus_context({repo: "@myGroup", name: "validateUser"})` |
| `impact` (group mode) | Cross-repo blast radius via Contract Bridge | `gitnexus_impact({repo: "@myGroup", target: "X", direction: "upstream"})` |
> Group mode: pass `repo: "@<groupName>"` to fan out across all member repos, or `repo: "@<groupName>/<memberPath>"` to target a single member (path keys from `group.yaml`). Optional `service: "<monorepo/path>"` filters by service root. Group-level state (contracts, staleness) lives in the resources table below — there are **no** `group_query` / `group_context` / `group_impact` / `group_contracts` / `group_status` MCP tools.
>
> For a full walkthrough of setting up a group across multiple repos that communicate over gRPC, see [docs/guides/microservices-grpc.md](docs/guides/microservices-grpc.md).
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, index freshness |
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
| `gitnexus://group/{name}/contracts` | Group Contract Registry (provider/consumer rows + cross-links) |
| `gitnexus://group/{name}/status` | Per-member index + Contract Registry staleness report |
## Self-Check Before Finishing
## CLI
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL warnings were ignored
3. `gitnexus_detect_changes()` confirms expected scope
4. All d=1 dependents were updated
## Keeping the Index Fresh
```bash
npx gitnexus analyze # incremental by default; preserves embeddings
npx gitnexus analyze --force # full rebuild from scratch (opt out of incremental)
npx gitnexus analyze --embeddings # also generate embeddings for new/changed nodes
npx gitnexus analyze --drop-embeddings # explicit opt-in to wipe existing embeddings
```
`analyze` runs **incrementally by default**. The pipeline still parses every file every run (cross-file resolution requires it), but tree-sitter parsing is **served from a content-addressed cache** under `.gitnexus/parse-cache/` (per-chunk JSON shards plus `index.json`) for chunks whose file contents haven't changed since the last run. Older installs may still have a legacy single file `.gitnexus/parse-cache.json`, which is read for backward compatibility but no longer written. Only changed-file rows (and their importers) are rewritten in LadybugDB; unchanged-file rows are preserved. Output is byte-equivalent to a full rebuild. Pass `--force` to wipe and re-index from scratch (e.g., to recover from a corrupt index, or after upgrading GitNexus).
The parse cache key is **content-addressed and version-tagged**: it survives `--force` runs, and is automatically invalidated by a `gitnexus` package upgrade (so a new tree-sitter grammar doesn't silently replay stale parse output). Safe to delete the whole `.gitnexus/parse-cache/` directory (and remove any legacy `.gitnexus/parse-cache.json` if present) at any time — it'll be rebuilt on the next analyze.
Check `.gitnexus/meta.json` `stats.embeddings` (0 = none). A plain `analyze` no longer drops existing vectors — pass `--drop-embeddings` to wipe.
> Claude Code: PostToolUse hook detects a stale index after `git commit` and `git merge` and prompts the agent to run `analyze`. The hook does not invoke `analyze` itself.
## CLI Skills
| Task | Skill file |
|------|-----------|
| Architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Debugging / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Refactoring | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools/resources/schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| CLI commands (index, status, clean, wiki) | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
## Hook env knobs
The Claude Code hook (`gitnexus/hooks/claude/gitnexus-hook.cjs` and the mirrored plugin copy under `gitnexus-claude-plugin/hooks/`) honours these env vars. Defaults work for normal installations; set them only to override resolution. All path overrides ignore values that do not exist on disk and fall through to the standard resolution chain.
| Env var | Type | Default | Purpose |
|---------|------|---------|---------|
| `GITNEXUS_HOOK_CLI_PATH` | path | resolved via package layout / `require.resolve` | Override path to the `gitnexus` CLI entry the hook spawns for `augment`. |
| `GITNEXUS_HOOK_LSOF_PATH` | path | `lsof` on `PATH` (with `/usr/bin/lsof`, `/usr/sbin/lsof`, `/sbin/lsof` fallbacks) | Override POSIX `lsof` location for the DB-lock probe. |
| `GITNEXUS_HOOK_PS_PATH` | path | `ps` on `PATH` (with `/bin/ps`, `/usr/bin/ps` fallbacks) | Override POSIX `ps` location. |
| `GITNEXUS_HOOK_POWERSHELL_PATH` | path | `%SystemRoot%\System32\WindowsPowerShell\v1.0\powershell.exe` (then `SysWOW64`, then `powershell.exe` on `PATH`) | Override Windows PowerShell location used by the Restart-Manager probe. |
| `GITNEXUS_HOOK_LINUX_PROC_BUDGET_MS` | integer ms | `1200` | Max wall-clock for the Linux `/proc` fd scan before bailing out to the `lsof` fallback. |
| `GITNEXUS_HOOK_RM_TARGET` | path | derived | Restart-Manager target file (the LadybugDB path under `.gitnexus/`). Set internally by the hook; rarely overridden manually. |
| `GITNEXUS_DEBUG` | boolean (`1`/`true`) | unset | Verbose stderr from the hook: prints discarded augment-stderr prefixes and one-shot `.ps1` load-failure warnings. |
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (239 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Extractors area (135 symbols) | `.claude/skills/generated/extractors/SKILL.md` |
| Work in the Components area (112 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Lbug area (96 symbols) | `.claude/skills/generated/lbug/SKILL.md` |
| Work in the Group area (94 symbols) | `.claude/skills/generated/group/SKILL.md` |
| Work in the Cli area (92 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Configs area (92 symbols) | `.claude/skills/generated/configs/SKILL.md` |
| Work in the Type-extractors area (90 symbols) | `.claude/skills/generated/type-extractors/SKILL.md` |
| Work in the Hooks area (88 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Unit area (80 symbols) | `.claude/skills/generated/unit/SKILL.md` |
| Work in the Cpp area (73 symbols) | `.claude/skills/generated/cpp/SKILL.md` |
| Work in the Scope-resolution area (72 symbols) | `.claude/skills/generated/scope-resolution/SKILL.md` |
| Work in the Server area (66 symbols) | `.claude/skills/generated/server/SKILL.md` |
| Work in the Local area (61 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Wiki area (60 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Workers area (57 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Embeddings area (56 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Typescript area (53 symbols) | `.claude/skills/generated/typescript/SKILL.md` |
| Work in the Storage area (51 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Php area (48 symbols) | `.claude/skills/generated/php/SKILL.md` |
<!-- gitnexus:end -->
+64
View File
@@ -52,3 +52,67 @@ If always-on instructions grow, load deep conventions via conditional reads (e.g
## GitNexus rules
See the `<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** for the canonical MCP tools, impact analysis rules, and index instructions.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (239 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Extractors area (135 symbols) | `.claude/skills/generated/extractors/SKILL.md` |
| Work in the Components area (112 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Lbug area (96 symbols) | `.claude/skills/generated/lbug/SKILL.md` |
| Work in the Group area (94 symbols) | `.claude/skills/generated/group/SKILL.md` |
| Work in the Cli area (92 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Configs area (92 symbols) | `.claude/skills/generated/configs/SKILL.md` |
| Work in the Type-extractors area (90 symbols) | `.claude/skills/generated/type-extractors/SKILL.md` |
| Work in the Hooks area (88 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Unit area (80 symbols) | `.claude/skills/generated/unit/SKILL.md` |
| Work in the Cpp area (73 symbols) | `.claude/skills/generated/cpp/SKILL.md` |
| Work in the Scope-resolution area (72 symbols) | `.claude/skills/generated/scope-resolution/SKILL.md` |
| Work in the Server area (66 symbols) | `.claude/skills/generated/server/SKILL.md` |
| Work in the Local area (61 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Wiki area (60 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Workers area (57 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Embeddings area (56 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Typescript area (53 symbols) | `.claude/skills/generated/typescript/SKILL.md` |
| Work in the Storage area (51 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Php area (48 symbols) | `.claude/skills/generated/php/SKILL.md` |
<!-- gitnexus:end -->
+4 -1
View File
@@ -197,7 +197,8 @@ args = ["-y", "gitnexus@latest", "mcp"]
```bash
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --repair-fts # Fast path: rebuild/verify only FTS indexes on existing index data
gitnexus analyze --force # Full rebuild: re-parse + graph rebuild + FTS rebuild
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
@@ -728,6 +729,8 @@ gitnexus wiki --force
gitnexus wiki --timeout <seconds> # LLM request timeout in seconds (default: disabled)
gitnexus wiki --retries <n> # Max LLM retry attempts per request (default: 3)
# Change the language generation for wiki
gitnexus wiki --lang <lang> # Output language for generated documentation (e.g. english, chinese, spanish, japanese)
```
The wiki generator reads the indexed graph structure, groups files into modules via LLM, generates per-module documentation pages, and creates an overview page — all with cross-references to the knowledge graph.
+57 -2
View File
@@ -162,8 +162,8 @@ Each mode has a `system_{mode}.jinja` + `instance_{mode}.jinja` pair. The agent
```
Agent → bash command → /usr/local/bin/gitnexus-query
→ curl localhost:4848/tool/query (fast path: eval-server, ~100ms)
→ npx gitnexus query (fallback: cold CLI, ~5-10s)
→ curl http://127.0.0.1:4848/tool/query (fast path: eval-server, ~100ms)
→ npx gitnexus query (fallback: cold CLI, ~5-10s)
```
Each tool script in `/usr/local/bin/` is standalone — no sourcing, no env inheritance needed. This is critical because mini-swe-agent runs every command via `subprocess.run` in a fresh subshell.
@@ -176,6 +176,61 @@ The eval-server is a lightweight HTTP daemon that:
- Includes next-step hints to guide tool chaining (query → context → impact → fix)
- Auto-shuts down after idle timeout
**CLI flags:**
| Flag | Default | Purpose |
|------|---------|---------|
| `--port <port>` | `4848` | Port to listen on |
| `--host <host>` | `127.0.0.1` | Bind address — use `0.0.0.0` for cross-container access |
| `--idle-timeout <seconds>` | `0` (disabled) | Auto-shutdown after N seconds of inactivity |
**READY signal:**
When the server is ready, it writes to stdout:
```
# IPv4
GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848
# IPv6 (bracketed to avoid colon ambiguity)
GITNEXUS_EVAL_SERVER_READY:[::1]:4848
```
Parse the port as the last colon-segment (`split(':').pop()`) — not `split(':')[1]`, which breaks for IPv6 and for non-loopback IPv4 hosts added in this release.
### Custom port and host
`run_eval.py` does not expose `--port` or `--host` as CLI flags. Configure them in your mode YAML under the `environment:` key:
```yaml
# configs/modes/native_augment.yaml (or whichever mode you're running)
environment:
eval_server_port: 4849 # change if 4848 is already in use on the host
eval_server_host: "0.0.0.0" # bind all interfaces — needed for cross-container setups
```
Defaults are `port: 4848` and `host: 127.0.0.1` (loopback only). Use `0.0.0.0` only when the agent container needs to reach the eval-server from a separate network namespace. The health probe and tool scripts connect via the configured bind host (defaulting to `127.0.0.1`), which is reachable for both loopback and all-interface binds.
`"localhost"` is also a valid `eval_server_host` value. The OS resolves it at bind time — typically `127.0.0.1` on dual-stack or IPv4-only systems, and `::1` on IPv6-only systems. The exact result depends on your `/etc/hosts` and `gai.conf`. The READY signal will reflect the actual bound address (e.g. `GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848` or `GITNEXUS_EVAL_SERVER_READY:[::1]:4848`), not the literal string `localhost`. Use this when you want the server to bind to whichever loopback address the OS prefers rather than forcing IPv4.
**Running eval-server directly in Docker / Docker Compose:**
```bash
# Bind to all interfaces so sibling containers can reach it
gitnexus eval-server --host 0.0.0.0 --port 4848
# Then probe from a sibling container via its service hostname
curl http://eval-container:4848/health
```
If you need a non-default port (e.g. to avoid conflicts), pass `--port <port>` alongside `--host`. The READY signal will reflect both:
```
GITNEXUS_EVAL_SERVER_READY:0.0.0.0:5000
```
Parse the port as the last colon-segment (`split(':').pop()`) — safe for both IPv4 and bracketed IPv6 forms.
### Index caching
SWE-bench repos repeat (Django has 200+ instances at different commits). The harness caches GitNexus indexes per `(repo, commit)` hash in `~/.gitnexus-eval-cache/` to avoid redundant re-indexing.
+18 -5
View File
@@ -39,6 +39,7 @@ logger = logging.getLogger("gitnexus_docker")
DEFAULT_CACHE_DIR = Path.home() / ".gitnexus-eval-cache"
EVAL_SERVER_PORT = 4848
EVAL_SERVER_HOST = "127.0.0.1"
class GitNexusDockerEnvironment(DockerEnvironment):
@@ -62,6 +63,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
skip_embeddings: bool = True,
gitnexus_timeout: int = 120,
eval_server_port: int = EVAL_SERVER_PORT,
eval_server_host: str = EVAL_SERVER_HOST,
**kwargs,
):
super().__init__(**kwargs)
@@ -70,6 +72,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
self.skip_embeddings = skip_embeddings
self.gitnexus_timeout = gitnexus_timeout
self.eval_server_port = eval_server_port
self.eval_server_host = eval_server_host
self.index_time: float = 0.0
self._gitnexus_ready = False
@@ -165,22 +168,29 @@ class GitNexusDockerEnvironment(DockerEnvironment):
def _start_eval_server(self):
"""Start the GitNexus eval-server daemon in the background."""
logger.info(f"Starting eval-server on port {self.eval_server_port}...")
logger.info(
f"Starting eval-server on {self.eval_server_host}:{self.eval_server_port}..."
)
self.execute({
"command": (
f"nohup npx gitnexus eval-server --port {self.eval_server_port} "
f"--host {self.eval_server_host} "
f"--idle-timeout 600 "
f"> /tmp/gitnexus-eval-server.log 2>&1 &"
),
"timeout": 5,
})
# Use 127.0.0.1 for the health probe — reachable whether server binds
# loopback or all interfaces (0.0.0.0), avoiding DNS resolution issues.
health_host = "127.0.0.1"
# Wait for the server to be ready (up to ~15s for KuzuDB init)
for i in range(EVAL_SERVER_HEALTH_RETRIES):
time.sleep(EVAL_SERVER_HEALTH_INTERVAL_SECONDS)
health = self.execute({
"command": f"curl -sf http://127.0.0.1:{self.eval_server_port}/health 2>/dev/null || echo 'NOT_READY'",
"command": f"curl -sf http://{health_host}:{self.eval_server_port}/health 2>/dev/null || echo 'NOT_READY'",
"timeout": EVAL_SERVER_HEALTH_TIMEOUT_SECONDS,
})
output = health.get("output", "").strip()
@@ -201,7 +211,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
)
@staticmethod
def _render_tool_script(spec: ToolScriptSpec, port: str) -> str:
def _render_tool_script(spec: ToolScriptSpec, port: str, host: str = EVAL_SERVER_HOST) -> str:
"""
Render a standalone bash script for a GitNexus tool.
@@ -212,6 +222,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
if spec.endpoint:
lines.append(f'PORT="${{GITNEXUS_EVAL_PORT:-{port}}}"')
lines.append(f'HOST="${{GITNEXUS_EVAL_HOST:-{host}}}"')
if spec.header:
lines.append(spec.header.strip())
@@ -221,7 +232,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
if spec.endpoint:
lines.append(
f'result=$(curl -sf -X POST "http://127.0.0.1:${{PORT}}{spec.endpoint}" '
f'result=$(curl -sf -X POST "http://${{HOST}}:${{PORT}}{spec.endpoint}" '
'-H "Content-Type: application/json" -d "$payload" 2>/dev/null)'
)
lines.append('if [ $? -eq 0 ] && [ -n "$result" ]; then echo "$result"; exit 0; fi')
@@ -244,9 +255,10 @@ class GitNexusDockerEnvironment(DockerEnvironment):
Uses heredocs with quoted delimiter to avoid all quoting/escaping issues.
"""
port = str(self.eval_server_port)
host = self.eval_server_host
for spec in TOOL_SPECS.values():
script_content = self._render_tool_script(spec, port).strip()
script_content = self._render_tool_script(spec, port, host).strip()
# Use heredoc with quoted delimiter — prevents all variable expansion and quoting issues
self.execute({
"command": (
@@ -387,5 +399,6 @@ class GitNexusDockerEnvironment(DockerEnvironment):
"index_time_seconds": round(self.index_time, 2),
"skip_embeddings": self.skip_embeddings,
"eval_server_port": self.eval_server_port,
"eval_server_host": self.eval_server_host,
}
return base
Generated
+3 -3
View File
@@ -760,11 +760,11 @@ wheels = [
[[package]]
name = "idna"
version = "3.11"
version = "3.15"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/6f/6d/0703ccc57f3a7233505399edb88de3cbd678da106337b9fcde432b65ed60/idna-3.11.tar.gz", hash = "sha256:795dafcc9c04ed0c1fb032c2aa73654d8e8c5023a7df64a53f39190ada629902", size = 194582, upload-time = "2025-10-12T14:55:20.501Z" }
sdist = { url = "https://files.pythonhosted.org/packages/82/77/7b3966d0b9d1d31a36ddf1746926a11dface89a83409bf1483f0237aa758/idna-3.15.tar.gz", hash = "sha256:ca962446ea538f7092a95e057da437618e886f4d349216d2b1e294abfdb65fdc", size = 199245, upload-time = "2026-05-12T22:45:57.011Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/0e/61/66938bbb5fc52dbdf84594873d5b51fb1f7c7794e9c0f5bd885f30bc507b/idna-3.11-py3-none-any.whl", hash = "sha256:771a87f49d9defaf64091e6e6fe9c18d4833f140bd19464795bc32d966ca37ea", size = 71008, upload-time = "2025-10-12T14:55:18.883Z" },
{ url = "https://files.pythonhosted.org/packages/d2/23/408243171aa9aaba178d3e2559159c24c1171a641aa83b67bdd3394ead8e/idna-3.15-py3-none-any.whl", hash = "sha256:048adeaf8c2d788c40fee287673ccaa74c24ffd8dcf09ffa555a2fbb59f10ac8", size = 72340, upload-time = "2026-05-12T22:45:55.733Z" },
]
[[package]]
@@ -56,7 +56,7 @@ Generates repository documentation from the knowledge graph using an LLM. Requir
| Flag | Effect |
|------|--------|
| `--force` | Force full regeneration |
| `--force` | Force full regeneration, also required to re-gerenate an existing wiki in a different language |
| `--model <model>` | LLM model (default: minimax/minimax-m2.5) |
| `--base-url <url>` | LLM API base URL |
| `--api-key <key>` | LLM API key |
@@ -64,7 +64,7 @@ Generates repository documentation from the knowledge graph using an LLM. Requir
| `--gist` | Publish wiki as a public GitHub Gist |
| `--timeout <seconds>` | LLM request timeout in seconds (default: disabled) |
| `--retries <n>` | Max LLM retry attempts per request (default: 3) |
| `--lang <lang>` | Output language for generated documentation (e.g. english, chinese, spanish, japanese)|
### list — Show all indexed repos
```bash
+1
View File
@@ -127,6 +127,7 @@ export { CLASS_KINDS, METHOD_KINDS, FIELD_KINDS } from './scope-resolution/regis
export type {
RegistryContext,
RegistryProviders,
OwnedMembersByOwnerLookup,
OwnerScopedContributor,
ArityVerdict,
ConstraintContext,
@@ -21,6 +21,7 @@
* (defined in `./types.ts`).
*/
import type { ParameterTypeClass } from './symbol-definition.js';
import type { Range, ScopeId } from './types.js';
/**
@@ -79,4 +80,11 @@ export interface ReferenceSite {
* (C#: `42` → `'int'`, `"alice"` → `'string'`).
*/
readonly argumentTypes?: readonly string[];
/**
* Optional per-argument type-shape sidecar for languages that need
* cv/ref/pointer distinctions during constraint filtering. This is
* intentionally separate from `argumentTypes`, which stays normalized
* for existing overload narrowing and conversion-rank logic.
*/
readonly argumentTypeClasses?: readonly ParameterTypeClass[];
}
@@ -13,7 +13,7 @@
*/
import type { NodeLabel } from '../../graph/types.js';
import type { SymbolDefinition } from '../symbol-definition.js';
import type { ParameterTypeClass, SymbolDefinition } from '../symbol-definition.js';
import type { Callsite, DefId } from '../types.js';
import type { DefIndex } from '../def-index.js';
import type { QualifiedNameIndex } from '../qualified-name-index.js';
@@ -65,6 +65,13 @@ export interface ConstraintContext {
* `narrowOverloadCandidates`' `argTypes` parameter.
*/
readonly argumentTypes?: readonly string[];
/**
* Optional shape-preserving sidecar aligned with `argumentTypes`.
* Unknown or unsupported slots should be omitted by producers or
* marked with `indirection: 'unknown'`; consumers must preserve the
* monotonic fallback and return 'unknown' instead of guessing.
*/
readonly argumentTypeClasses?: readonly ParameterTypeClass[];
}
// ─── Owner-scoped contributor (concrete shape for `RegistryContributor`) ────
@@ -93,6 +100,19 @@ export interface OwnerScopedContributor {
byName(name: string): readonly SymbolDefinition[];
}
/**
* Required owner-keyed lookup hook for Step 2 receiver/MRO member walks.
* Production callers wire this to the SemanticModel's authoritative
* method/field/nested-type registries so each `(ownerDefId, memberName)`
* probe is O(1). Implementations MUST return `[]` on an indexed miss —
* Step 2 treats `[]` as authoritative and does not consult `defs` for a
* fallback scan.
*/
export type OwnedMembersByOwnerLookup = (
ownerDefId: DefId,
memberName: string,
) => readonly SymbolDefinition[];
// ─── Top-level context threaded through every lookup ───────────────────────
export interface RegistryContext {
@@ -100,6 +120,7 @@ export interface RegistryContext {
readonly defs: DefIndex;
readonly qualifiedNames: QualifiedNameIndex;
readonly moduleScopes: ModuleScopeIndex;
readonly ownedMembersByOwner: OwnedMembersByOwnerLookup;
/**
* Method-dispatch index; required for method/field registries that
* honor `useReceiverTypeBinding`. Omit for class-only lookups.
@@ -27,8 +27,10 @@
* is true, resolve the receiver's type at `startScope` (from
* `scope.typeBindings`), then walk the MRO via
* `MethodDispatchIndex.mroFor(ownerDefId)`. Membership per owner comes
* through `RegistryContext.methodDispatch` + owner lookups into
* `scope.ownedDefs`; each hit records a raw signal with the owner's
* through an optional `RegistryContext.ownedMembersByOwner` hook when
* supplied (`undefined` → fall back to `defs.byId`; `[]` → indexed
* miss), otherwise via the compatibility fallback scan over
* `defs.byId`; each hit records a raw signal with the owner's
* MRO depth.
*
* **Step 3 — Owner-scoped contributor.** When
@@ -263,13 +265,14 @@ function walkReceiverTypeBinding(
// Walk the owner itself at depth 0, then its MRO chain.
const walk: DefId[] = [ownerDefId, ...ctx.methodDispatch.mroFor(ownerDefId)];
for (let mroDepth = 0; mroDepth < walk.length; mroDepth++) {
const currentOwnerId = walk[mroDepth]!;
let mroDepth = 0;
for (const currentOwnerId of walk) {
const members = collectOwnedMembers(currentOwnerId, name, ctx);
for (const def of members) {
if (!acceptedKinds.has(def.type)) continue;
recordTypeBindingHit(perCandidate, def, mroDepth, ownerDefId);
}
mroDepth++;
}
}
@@ -333,23 +336,7 @@ function collectOwnedMembers(
memberName: string,
ctx: RegistryContext,
): readonly SymbolDefinition[] {
// An owner's members are defs whose `ownerId === ownerDefId` and whose
// simple name matches `memberName`. We iterate `defs.byId` — O(D) per
// call today. A future by-owner index would make this O(K); tracked as
// a follow-up optimization before Ring 3 flips go production.
const out: SymbolDefinition[] = [];
for (const def of ctx.defs.byId.values()) {
if (def.ownerId !== ownerDefId) continue;
if (simpleNameOf(def) !== memberName) continue;
out.push(def);
}
return out;
}
function simpleNameOf(def: SymbolDefinition): string | undefined {
if (def.qualifiedName === undefined || def.qualifiedName.length === 0) return undefined;
const dot = def.qualifiedName.lastIndexOf('.');
return dot === -1 ? def.qualifiedName : def.qualifiedName.slice(dot + 1);
return ctx.ownedMembersByOwner(ownerDefId, memberName);
}
function recordTypeBindingHit(
+8 -22
View File
@@ -41,7 +41,7 @@
"sigma": "^3.0.2",
"tailwindcss": "^4.2.4",
"uuid": "^14.0.0",
"zod": "^3.25.76"
"zod": "^4.3.6"
},
"devDependencies": {
"@babel/types": "^7.29.0",
@@ -5599,13 +5599,12 @@
}
},
"node_modules/langsmith": {
"version": "0.5.23",
"resolved": "https://registry.npmjs.org/langsmith/-/langsmith-0.5.23.tgz",
"integrity": "sha512-dE/M/2Gg2S2R8ygDdkWGJVO3JstijvsNvPXsy9V8WGbpb88Zn8xF/aTjPx4mIy5gIoo02T6FssOgYyLf51Dv1Q==",
"version": "0.6.3",
"resolved": "https://registry.npmjs.org/langsmith/-/langsmith-0.6.3.tgz",
"integrity": "sha512-pXrQ4/4myQvjFFOAUmt5pWRrLEZR20gzIJD7MNdUH+5/S5nLI4ZRBo/SYKC6coaYj9pYTfQdBIzcs+3kfJ5uDA==",
"license": "MIT",
"dependencies": {
"p-queue": "6.6.2",
"uuid": "10.0.0"
"p-queue": "6.6.2"
},
"peerDependencies": {
"@opentelemetry/api": "*",
@@ -5632,19 +5631,6 @@
}
}
},
"node_modules/langsmith/node_modules/uuid": {
"version": "10.0.0",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-10.0.0.tgz",
"integrity": "sha512-8XkAphELsDnEGrDxUOHB3RGvXz6TeuYSGEZBOjtTtPm2lwhGBjLgOzLHB63IUWfBpNucQjND6d3AOudO+H3RWQ==",
"funding": [
"https://github.com/sponsors/broofa",
"https://github.com/sponsors/ctavan"
],
"license": "MIT",
"bin": {
"uuid": "dist/bin/uuid"
}
},
"node_modules/layout-base": {
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/layout-base/-/layout-base-1.0.2.tgz",
@@ -8903,9 +8889,9 @@
}
},
"node_modules/zod": {
"version": "3.25.76",
"resolved": "https://registry.npmjs.org/zod/-/zod-3.25.76.tgz",
"integrity": "sha512-gzUt/qt81nXsFGKIFcC3YnfEAx5NkunCfnDlvuBSSFS02bcXu4Lmea0AFIUwbLWxWPx3d9p8S5QoaujKcNQxcQ==",
"version": "4.3.6",
"resolved": "https://registry.npmjs.org/zod/-/zod-4.3.6.tgz",
"integrity": "sha512-rftlrkhHZOcjDwkGlnUtZZkvaPHCsDATp4pGpuOOMDaTdDDXF91wuVDJoWoPsKX/3YPQ5fHuF3STjcYyKr+Qhg==",
"license": "MIT",
"funding": {
"url": "https://github.com/sponsors/colinhacks"
+1 -1
View File
@@ -51,7 +51,7 @@
"sigma": "^3.0.2",
"tailwindcss": "^4.2.4",
"uuid": "^14.0.0",
"zod": "^3.25.76"
"zod": "^4.3.6"
},
"devDependencies": {
"@babel/types": "^7.29.0",
+2 -1
View File
@@ -151,7 +151,8 @@ Your AI agent gets these tools automatically:
```bash
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --repair-fts # Fast path: rebuild/verify only FTS indexes on existing index data
gitnexus analyze --force # Full rebuild: re-parse + graph rebuild + FTS rebuild
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
+209 -639
View File
File diff suppressed because it is too large Load Diff
+3 -3
View File
@@ -1,6 +1,6 @@
{
"name": "gitnexus",
"version": "1.6.5",
"version": "1.6.6-rc.26",
"description": "Graph-powered code intelligence for AI agents. Index any codebase, query via MCP or CLI.",
"author": "Abhigyan Patwari",
"license": "PolyForm-Noncommercial-1.0.0",
@@ -60,7 +60,7 @@
"cli-progress": "^3.12.0",
"commander": "^14.0.3",
"cors": "^2.8.5",
"express": "^4.19.2",
"express": "^5.2.1",
"express-rate-limit": "^8.4.1",
"glob": "^13.0.6",
"graphology": "^0.26.0",
@@ -100,7 +100,7 @@
"devDependencies": {
"@types/cli-progress": "^3.11.6",
"@types/cors": "^2.8.17",
"@types/express": "^4.17.21",
"@types/express": "^5.0.6",
"@types/js-yaml": "^4.0.9",
"@types/node": "^25.6.0",
"@types/uuid": "^11.0.0",
+95 -3
View File
@@ -68,13 +68,69 @@ const installFatalHandlers = (): void => {
});
};
const HEAP_MB = 8192;
const HEAP_FLAG = `--max-old-space-size=${HEAP_MB}`;
const HEAP_MB = 16384;
const TEST_RESPAWN_HEAP_MB = Number(process.env.GITNEXUS_TEST_RESPAWN_HEAP_MB);
const RESPAWN_HEAP_MB =
Number.isFinite(TEST_RESPAWN_HEAP_MB) && TEST_RESPAWN_HEAP_MB > 0
? Math.floor(TEST_RESPAWN_HEAP_MB)
: HEAP_MB;
const HEAP_FLAG = `--max-old-space-size=${RESPAWN_HEAP_MB}`;
/** Increase default stack size (KB) to prevent stack overflow on deep class hierarchies. */
const STACK_KB = 4096;
const STACK_FLAG = `--stack-size=${STACK_KB}`;
/** Re-exec the process with an 8GB heap and larger stack if we're currently below that. */
/**
* Heuristic for "child re-exec likely died from V8 OOM".
*
* Platform-independent detection is best-effort: V8/Node usually emit
* stable heap-exhaustion phrases in stderr/message across Linux/macOS/Windows
* (for example "JavaScript heap out of memory" or "Reached heap limit"),
* while some environments only expose status/signal (e.g. 134/SIGABRT).
* We combine both text signatures and process-exit signatures.
*/
const childProcessLikelyOom = (err: unknown): boolean => {
if (!err || typeof err !== 'object') return false;
const e = err as {
status?: unknown;
signal?: unknown;
stderr?: unknown;
stdout?: unknown;
message?: unknown;
};
const hasHeapOomSignature = (v: unknown): boolean => {
const text = (
Buffer.isBuffer(v) ? v.toString('utf8') : typeof v === 'string' ? v : ''
).toLowerCase();
if (!text) return false;
return (
text.includes('javascript heap out of memory') ||
text.includes('reached heap limit') ||
text.includes('allocation failed - javascript heap out of memory') ||
text.includes('fatalprocessoutofmemory')
);
};
const fields = [e.message, e.stderr, e.stdout];
if (fields.some((v) => hasHeapOomSignature(v))) return true;
const hasAnyChildOutput = [e.stderr, e.stdout].some(
(v) => (Buffer.isBuffer(v) && v.length > 0) || (typeof v === 'string' && v.length > 0),
);
if (hasAnyChildOutput) return false;
return e.status === 134 || e.signal === 'SIGABRT';
};
const forceHeapOOMForTestIfEnabled = (): void => {
if (process.env.GITNEXUS_TEST_FORCE_HEAP_OOM !== '1') return;
// Allocate JS strings (not Buffers) so pressure lands on V8 heap itself.
// Buffers can allocate off-heap, which makes OOM triggering less reliable.
const chunks: string[] = [];
for (;;) chunks.push('x'.repeat(1024 * 1024));
};
/** Re-exec the process with a 16GB heap and larger stack if we're currently below that. */
function ensureHeap(): boolean {
const nodeOpts = process.env.NODE_OPTIONS || '';
if (nodeOpts.includes('--max-old-space-size')) return false;
@@ -93,6 +149,16 @@ function ensureHeap(): boolean {
env: { ...process.env, NODE_OPTIONS: `${nodeOpts} ${HEAP_FLAG}`.trim() },
});
} catch (e: any) {
if (childProcessLikelyOom(e)) {
cliError(
` Analysis likely ran out of memory.\n` +
` Retry with a larger heap if your machine allows it:\n` +
` NODE_OPTIONS="--max-old-space-size=24576" gitnexus analyze [your-args]\n` +
` (Windows: set NODE_OPTIONS=--max-old-space-size=24576 && gitnexus analyze [your-args])\n` +
` If this persists, it may be a native crash unrelated to heap size.\n`,
{ recoveryHint: 'heap-oom-respawn' },
);
}
process.exitCode = e.status ?? 1;
}
return true;
@@ -100,6 +166,7 @@ function ensureHeap(): boolean {
export interface AnalyzeOptions {
force?: boolean;
repairFts?: boolean;
/**
* Embedding generation toggle. Commander parses `--embeddings [limit]` as:
* - `undefined` when the flag is omitted
@@ -185,6 +252,7 @@ export const shouldGenerateCommunitySkillFiles = (
export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOptions) => {
if (ensureHeap()) return;
forceHeapOOMForTestIfEnabled();
// Install fatal handlers immediately after re-exec resolution so any
// async error that escapes the try/catch below (#1169) surfaces with
@@ -276,6 +344,15 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
process.env.GITNEXUS_EMBEDDING_DEVICE = options.embeddingDevice;
}
if (options?.repairFts && options?.force) {
cliError(
' Cannot combine `--repair-fts` with `--force`. ' +
'Use `--repair-fts` for fast FTS-only repair, or `--force` for a full rebuild.\n',
);
process.exitCode = 1;
return;
}
console.log('\n GitNexus Analyzer\n');
// `--index-only` is the stronger contract — it suppresses every form of file
@@ -454,9 +531,11 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
// needs a fresh pipelineResult. Has no bearing on the registry
// collision guard (see allowDuplicateName below).
force: options?.force || options?.skills,
repairFts: options?.repairFts,
embeddings: embeddingsEnabled,
embeddingsNodeLimit,
dropEmbeddings: options?.dropEmbeddings,
verbose: options?.verbose,
skipGit: options?.skipGit,
skipAgentsMd,
skipSkills,
@@ -501,6 +580,19 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
return;
}
if (result.ftsRepairedOnly) {
clearInterval(elapsedTimer);
process.removeListener('SIGINT', sigintHandler);
console.log = origLog;
// eslint-disable-next-line no-console -- restoring after intentional progress-bar routing
console.warn = origWarn;
// eslint-disable-next-line no-console -- restoring after intentional progress-bar routing
console.error = origError;
bar.stop();
console.log(' FTS indexes repaired successfully\n');
return;
}
// Post-finalize invariant (#1169): runFullAnalysis nominally writes
// meta.json and registers the repo, but on Windows it has been
// observed to return successfully with neither artifact present
+104 -9
View File
@@ -14,9 +14,14 @@
* Agent bash cmd → curl localhost:PORT/tool/query → eval-server → LocalBackend → format → text
*
* Usage:
* gitnexus eval-server # default port 4848
* gitnexus eval-server --port 4848 # explicit port
* gitnexus eval-server --idle-timeout 300 # auto-shutdown after 300s idle
* gitnexus eval-server # default port 4848, binds 127.0.0.1
* gitnexus eval-server --port 4848 # explicit port
* gitnexus eval-server --host 0.0.0.0 # reachable from other VMs / containers
* gitnexus eval-server --idle-timeout 300 # auto-shutdown after 300s idle
*
* READY signal format: GITNEXUS_EVAL_SERVER_READY:<host>:<port>
* IPv4: GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848
* IPv6: GITNEXUS_EVAL_SERVER_READY:[::1]:4848
*
* API:
* POST /tool/:name — Call a tool. Body is JSON arguments. Returns formatted text.
@@ -25,16 +30,30 @@
*/
import http from 'http';
import { isIPv4, isIPv6 } from 'node:net';
import { writeSync } from 'node:fs';
import { LocalBackend } from '../mcp/local/local-backend.js';
import { logger } from '../core/logger.js';
import { cliInfo, cliWarn } from './cli-message.js';
import { cliInfo, cliWarn, cliError } from './cli-message.js';
export interface EvalServerOptions {
port?: string;
host?: string;
idleTimeout?: string;
}
/**
* Validate the --host value. Accepts IPv4, IPv6, or "localhost".
* Returns the host string unchanged, or null if invalid.
* "localhost" is passed through so the OS resolves it to the correct loopback
* address (127.0.0.1 or ::1) at bind time rather than forcing IPv4.
*/
export function validateHost(raw: string): string | null {
if (raw === 'localhost') return raw;
if (isIPv4(raw) || isIPv6(raw)) return raw;
return null;
}
// ─── Text Formatters ──────────────────────────────────────────────────
// Convert structured JSON results into compact, LLM-friendly text.
// Design: minimize tokens, maximize actionability.
@@ -330,6 +349,22 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
const port = parseInt(options?.port || '4848');
const idleTimeoutSec = parseInt(options?.idleTimeout || '0');
const rawHost = options?.host ?? '127.0.0.1';
const host = validateHost(rawHost);
if (!host) {
cliError(
`Invalid --host value "${rawHost}":\n` +
` Must be an IP address or "localhost".\n\n` +
` Examples:\n` +
` gitnexus eval-server --host 127.0.0.1 (loopback only, default)\n` +
` gitnexus eval-server --host 0.0.0.0 (all network interfaces)\n` +
` gitnexus eval-server --host 192.168.1.5 (specific interface)\n` +
` gitnexus eval-server --host localhost (OS-resolved loopback)\n`,
{ flag: '--host', value: rawHost },
);
process.exit(1);
}
const backend = new LocalBackend();
const ok = await backend.init();
@@ -426,12 +461,72 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
}
});
server.listen(port, '127.0.0.1', () => {
server.on('error', (err: NodeJS.ErrnoException) => {
if (err.code === 'EADDRINUSE') {
cliError(
`\nGitNexus eval-server failed to start:\n` +
` Port ${port} is already in use.\n\n` +
` Either:\n` +
` 1. Stop the process already using port ${port}\n` +
` 2. Use a different port: gitnexus eval-server --port 4849\n`,
{ code: err.code, port, host },
);
} else if (err.code === 'EADDRNOTAVAIL') {
// "localhost" may resolve to ::1 on IPv6-only systems; treat it as
// potentially IPv6 so the user gets the right diagnostic hint.
const isIPv6Host = isIPv6(host) || host === 'localhost';
cliError(
`\nGitNexus eval-server failed to start:\n` +
` Address ${host} is not available on this machine.\n\n` +
(isIPv6Host
? ` Address ${host} resolved but is not reachable — IPv6 may be disabled, or the loopback interface may be unavailable.\n` +
` Docker containers and many CI environments disable IPv6 by default.\n\n`
: ` The --host value must be an IP assigned to a local network interface.\n` +
` Run \`ip addr\` (Linux) or \`ipconfig\` (Windows) to list available addresses.\n\n`) +
` Common fixes:\n` +
` gitnexus eval-server --host 127.0.0.1 (loopback, this machine only)\n` +
` gitnexus eval-server --host 0.0.0.0 (all interfaces, reachable from other VMs)\n`,
{ code: err.code, port, host },
);
} else if (err.code === 'EACCES') {
cliError(
`\nGitNexus eval-server failed to start:\n` +
` Permission denied binding to port ${port}.\n\n` +
` Ports below 1024 require elevated privileges.\n` +
` Use a port above 1024: gitnexus eval-server --port 4848\n`,
{ code: err.code, port, host },
);
} else {
cliError(`\nGitNexus eval-server failed to start:\n ${err.message}\n`, {
code: err.code,
port,
host,
});
}
process.exit(1);
});
server.listen(port, host, () => {
// Plain-text banner for the human watching stderr; structured record
// for log aggregation (split into two so the user sees a real banner
// not `{"level":30,"msg":"...","port":4747,"endpoints":[...]}`).
// Use server.address() so the banner and READY signal reflect what the OS
// actually bound to, not the input host string. This matters when "localhost"
// is passed: the OS may resolve it to ::1 on some systems.
const addr = server.address();
// server.listen callback only fires after a successful TCP bind, so
// server.address() is guaranteed to return an AddressInfo object here.
if (typeof addr !== 'object' || addr === null) {
cliError(
`\nGitNexus eval-server: unexpected server.address() value after bind: ${JSON.stringify(addr)}\n`,
);
process.exit(1);
}
const boundPort = addr.port;
const boundAddress = addr.address;
const displayHost = boundAddress.includes(':') ? `[${boundAddress}]` : boundAddress;
const bannerLines = [
`GitNexus eval-server: listening on http://127.0.0.1:${port}`,
`GitNexus eval-server: listening on http://${displayHost}:${boundPort}`,
` POST /tool/query — search execution flows`,
` POST /tool/context — 360-degree symbol view`,
` POST /tool/impact — blast radius analysis`,
@@ -443,8 +538,8 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
bannerLines.push(` Auto-shutdown after ${idleTimeoutSec}s idle`);
}
cliInfo(bannerLines.join('\n'), {
port,
host: '127.0.0.1',
port: boundPort,
host,
idleTimeoutSec: idleTimeoutSec > 0 ? idleTimeoutSec : undefined,
endpoints: [
'POST /tool/query',
@@ -457,7 +552,7 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
});
try {
// Use fd 1 directly — LadybugDB captures process.stdout (#324)
writeSync(1, `GITNEXUS_EVAL_SERVER_READY:${port}\n`);
writeSync(1, `GITNEXUS_EVAL_SERVER_READY:${displayHost}:${boundPort}\n`);
} catch {
// stdout may not be available (e.g., broken pipe)
}
+9
View File
@@ -23,6 +23,7 @@ program
.command('analyze [path]')
.description('Index a repository (full analysis)')
.option('-f, --force', 'Force full re-index even if up to date')
.option('--repair-fts', 'Repair/rebuild search FTS indexes without full re-analysis')
.option(
'--embeddings [limit]',
'Enable embedding generation for semantic search (off by default). ' +
@@ -166,6 +167,10 @@ program
.option('--gist', 'Publish wiki as a public GitHub Gist after generation')
.option('-v, --verbose', 'Enable verbose output (show LLM commands and responses)')
.option('--review', 'Stop after grouping to review module structure before generating pages')
.option(
'--lang <lang>',
'Output language for generated documentation (e.g. english, chinese, spanish, japanese)',
)
.action(createLazyAction(() => import('./wiki.js'), 'wikiCommand'));
program
@@ -237,6 +242,10 @@ program
.command('eval-server')
.description('Start lightweight HTTP server for fast tool calls during evaluation')
.option('-p, --port <port>', 'Port number', '4848')
.option(
'--host <host>',
'Bind address (default: 127.0.0.1, use 0.0.0.0 to expose to all interfaces)',
)
.option('--idle-timeout <seconds>', 'Auto-shutdown after N seconds idle (0 = disabled)', '0')
.action(createLazyAction(() => import('./eval-server.js'), 'evalServerCommand'));
+6 -4
View File
@@ -61,11 +61,13 @@ function resolveGitnexusBin(): string | null {
.filter(Boolean);
if (isWin) {
// On Windows, `where` returns multiple entries (e.g. the POSIX shell
// script AND the .cmd/.bat wrapper). Prefer the wrapper because
// child_process.spawn() cannot execute a shell script directly.
// On Windows, npm global installs can surface multiple launchers for the
// same package (e.g. a POSIX shell shim plus .cmd/.bat wrappers). Claude
// and the other MCP hosts need a directly spawnable command path, so only
// accept the Windows wrapper. If it is missing, fall back to the slower
// npx entry instead of persisting a non-spawnable shim path.
const cmdLine = lines.find((l) => /\.(cmd|bat)$/i.test(l));
return cmdLine || lines[0] || null;
return cmdLine || null;
}
return lines[0] || null;
+2
View File
@@ -35,6 +35,7 @@ export interface WikiCommandOptions {
review?: boolean;
timeout?: string;
retries?: string;
lang?: string;
}
function parsePositiveIntegerOption(
@@ -421,6 +422,7 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
force: options?.force,
concurrency: options?.concurrency ? parseInt(options.concurrency, 10) : undefined,
reviewOnly: options?.review,
lang: options?.lang,
};
const generator = new WikiGenerator(
@@ -18,12 +18,14 @@ import { getPluginForFile, HTTP_SCAN_GLOB, type HttpDetection } from './http-pat
* the preferred path because the graph has richer symbol metadata
* (real uids, class/method structure, etc.).
*
* 2. **Source-scan fallback (Strategy B)** — parse files directly with
* the per-language plugin registry in `./http-patterns/`. Used when
* the graph has no routes/fetches for this repo (e.g. a repo that
* hasn't been indexed yet, or whose indexer doesn't know the
* framework). Each plugin owns its tree-sitter grammar and query
* sources — this orchestrator imports NO grammars or query strings.
* 2. **Source-scan supplement (Strategy B)** — parse files directly with
* the per-language plugin registry in `./http-patterns/`. Used to
* fill gaps when graph extraction only covers part of a polyglot repo
* (e.g. Java graph routes plus Go source-scan routes). Graph entries
* remain authoritative for duplicate contract IDs because they carry
* richer symbol metadata. Each plugin owns its tree-sitter grammar
* and query sources — this orchestrator imports NO grammars or query
* strings.
*
* Adding a new language for Strategy B is a one-file edit in
* `http-patterns/index.ts`: register a new `HttpLanguagePlugin` and
@@ -194,17 +196,19 @@ export class HttpRouteExtractor implements ContractExtractor {
const graphProviders =
dbExecutor != null ? await this.extractProvidersGraph(dbExecutor, getDetections) : [];
const providers =
graphProviders.length > 0
? graphProviders
: this.extractProvidersSourceScan(await getScannedFiles(), getDetections);
// Source scan always runs to capture routes in languages/files not covered
// by graph edges; the glob and per-file parse results are cached above.
const providers = this.mergeGraphAndSourceContracts(
graphProviders,
this.extractProvidersSourceScan(await getScannedFiles(), getDetections),
);
const graphConsumers =
dbExecutor != null ? await this.extractConsumersGraph(dbExecutor, getDetections) : [];
const consumers =
graphConsumers.length > 0
? graphConsumers
: this.extractConsumersSourceScan(await getScannedFiles(), getDetections);
const consumers = this.mergeGraphAndSourceContracts(
graphConsumers,
this.extractConsumersSourceScan(await getScannedFiles(), getDetections),
);
return [...providers, ...consumers];
}
@@ -473,4 +477,18 @@ export class HttpRouteExtractor implements ContractExtractor {
}
return out;
}
private mergeGraphAndSourceContracts(
graphContracts: ExtractedContract[],
sourceContracts: ExtractedContract[],
): ExtractedContract[] {
const seenContractIds = new Set(graphContracts.map((c) => c.contractId));
const out = [...graphContracts];
for (const contract of sourceContracts) {
if (seenContractIds.has(contract.contractId)) continue;
seenContractIds.add(contract.contractId);
out.push(contract);
}
return out;
}
}
@@ -74,12 +74,31 @@ export const walkRepositoryPaths = async (
if (skippedLarge > 0) {
const isDefault = maxFileSizeBytes === DEFAULT_MAX_FILE_SIZE_BYTES;
const isOverrideUnset = !process.env.GITNEXUS_MAX_FILE_SIZE;
const suffix = isDefault ? ', likely generated/vendored' : '';
logger.warn(` Skipped ${skippedLarge} large files (>${maxFileSizeBytes / 1024}KB${suffix})`);
if (isVerboseIngestionEnabled()) {
for (const p of skippedLargePaths) {
logger.warn(` - ${p}`);
}
// Always show at least the first few paths so users can diagnose why
// edges are missing from a specific file (issue #1659). The full list is
// gated behind GITNEXUS_VERBOSE=1 to avoid flooding output on repos with
// many generated/vendored blobs. Sort before slicing so the preview is
// stable across runs (fs.stat callbacks race within each batch).
skippedLargePaths.sort();
const SKIPPED_PREVIEW_CAP = 5;
const showAll = isVerboseIngestionEnabled() || skippedLargePaths.length <= SKIPPED_PREVIEW_CAP;
const preview = showAll ? skippedLargePaths : skippedLargePaths.slice(0, SKIPPED_PREVIEW_CAP);
for (const p of preview) {
logger.warn(` - ${p}`);
}
if (!showAll) {
const remaining = skippedLargePaths.length - SKIPPED_PREVIEW_CAP;
logger.warn(` ...and ${remaining} more (set GITNEXUS_VERBOSE=1 to list them all)`);
}
// Only hint about the env var when the user has not set it at all. An
// explicit GITNEXUS_MAX_FILE_SIZE=512 happens to resolve to the same
// bytes as the default but the operator clearly already knows the knob.
if (isDefault && isOverrideUnset) {
logger.warn(` Set GITNEXUS_MAX_FILE_SIZE=<KB> to include files above the default cap.`);
}
}
@@ -1,4 +1,4 @@
import type { Capture, CaptureMatch } from 'gitnexus-shared';
import type { Capture, CaptureMatch, ParameterTypeClass } from 'gitnexus-shared';
import {
findNodeAtRange,
nodeToCapture,
@@ -9,7 +9,11 @@ import { getCppParser, getCppScopeQuery } from './query.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
import { splitCppInclude, splitCppUsingDecl } from './import-decomposer.js';
import { computeCppDeclarationArity, computeCppCallArity } from './arity-metadata.js';
import {
classifyCppParameterType,
computeCppDeclarationArity,
computeCppCallArity,
} from './arity-metadata.js';
import { markCppAnonymousNamespaceRange, markFileLocal } from './file-local-linkage.js';
import { markCppDependentBase } from './two-phase-lookup.js';
import { markCppAdlSiteArgs, markCppAdlSiteNoAdl, type CppAdlArgInfo } from './adl.js';
@@ -217,6 +221,14 @@ export function emitCppScopeCaptures(
JSON.stringify(argTypes),
);
}
const argTypeClasses = inferCppCallArgTypeClasses(cNode);
if (argTypeClasses !== undefined && argTypeClasses.length > 0) {
grouped['@reference.parameter-type-classes'] = syntheticCapture(
'@reference.parameter-type-classes',
cNode,
JSON.stringify(argTypeClasses),
);
}
}
}
@@ -683,6 +695,35 @@ function inferCppCallArgTypes(node: SyntaxNode): string[] | undefined {
return types.length > 0 ? types : undefined;
}
function inferCppCallArgTypeClasses(node: SyntaxNode): ParameterTypeClass[] | undefined {
const argList = node.childForFieldName('arguments');
if (argList === null) return undefined;
const classes: ParameterTypeClass[] = [];
for (let i = 0; i < argList.childCount; i++) {
const child = argList.child(i);
if (child === null) continue;
if (child.type === ',' || child.type === '(' || child.type === ')') continue;
const litType = inferCppLiteralType(child);
if (litType !== '') {
classes.push(valueTypeClass(litType));
} else if (child.type === 'identifier') {
classes.push(lookupDeclaredTypeClassForIdentifier(child));
} else {
classes.push(unknownTypeClass('unknown'));
}
}
return classes.length > 0 ? classes : undefined;
}
function valueTypeClass(base: string): ParameterTypeClass {
return { base, cv: 'none', indirection: 'value', pointerDepth: 0 };
}
function unknownTypeClass(base: string): ParameterTypeClass {
return { base, cv: 'unknown', indirection: 'unknown', pointerDepth: 0 };
}
/**
* Infer the canonical type name of a C++ literal AST node.
* Returns empty string for non-literal / unknown nodes.
@@ -750,6 +791,9 @@ function lookupDeclaredTypeForIdentifier(identNode: SyntaxNode): string {
}
if (scope === null) return '';
const paramType = lookupFunctionParameterType(scope, varName);
if (paramType !== '') return paramType;
// Scan declarations in the scope for a matching variable name
for (let i = 0; i < scope.childCount; i++) {
const stmt = scope.child(i);
@@ -763,18 +807,118 @@ function lookupDeclaredTypeForIdentifier(identNode: SyntaxNode): string {
// Check init_declarator children for the variable name
const declarator = stmt.childForFieldName('declarator');
if (declarator === null) continue;
if (declarator.type === 'init_declarator') {
const nameChild = declarator.childForFieldName('declarator');
if (nameChild !== null && nameChild.text === varName) {
return normalizeCppTypeText(typeNode.text);
}
} else if (declarator.text === varName) {
const nameChild = declaredNameNode(declarator);
if (nameChild !== null && extractDeclaratorLeafName(nameChild) === varName) {
return normalizeCppTypeText(typeNode.text);
}
}
return '';
}
function lookupDeclaredTypeClassForIdentifier(identNode: SyntaxNode): ParameterTypeClass {
const varName = identNode.text;
let scope: SyntaxNode | null = identNode.parent;
while (
scope !== null &&
scope.type !== 'compound_statement' &&
scope.type !== 'translation_unit'
) {
scope = scope.parent;
}
if (scope === null) return unknownTypeClass('unknown');
const paramTypeClass = lookupFunctionParameterTypeClass(scope, varName, identNode);
if (paramTypeClass !== undefined) return paramTypeClass;
for (let i = 0; i < scope.childCount; i++) {
const stmt = scope.child(i);
if (stmt === null || stmt.type !== 'declaration') continue;
const typeNode = stmt.childForFieldName('type');
if (typeNode === null) continue;
if (typeNode.type === 'placeholder_type_specifier') continue;
const declarator = stmt.childForFieldName('declarator');
if (declarator === null) continue;
const nameChild = declaredNameNode(declarator);
if (nameChild === null || extractDeclaratorLeafName(nameChild) !== varName) continue;
const typeClass = classifyCppParameterType(
typeNode.text,
nameChild.text,
stmt.text.replace(/;\s*$/, ''),
);
if (isKnownEnumName(identNode, typeClass.base)) {
return { ...typeClass, base: `enum:${typeClass.base}` };
}
return typeClass;
}
return unknownTypeClass('unknown');
}
function lookupFunctionParameterType(scope: SyntaxNode, varName: string): string {
const param = findEnclosingFunctionParameter(scope, varName);
if (param === null) return '';
const typeNode = param.childForFieldName('type');
if (typeNode === null) return '';
return normalizeCppTypeText(typeNode.text);
}
function lookupFunctionParameterTypeClass(
scope: SyntaxNode,
varName: string,
identNode: SyntaxNode,
): ParameterTypeClass | undefined {
const param = findEnclosingFunctionParameter(scope, varName);
if (param === null) return undefined;
const typeNode = param.childForFieldName('type');
if (typeNode === null) return undefined;
const declarator = param.childForFieldName('declarator');
if (declarator === null) return undefined;
const typeClass = classifyCppParameterType(typeNode.text, declarator.text, param.text);
if (isKnownEnumName(identNode, typeClass.base)) {
return { ...typeClass, base: `enum:${typeClass.base}` };
}
return typeClass;
}
function findEnclosingFunctionParameter(scope: SyntaxNode, varName: string): SyntaxNode | null {
let node: SyntaxNode | null = scope.parent;
while (node !== null) {
if (node.type === 'function_definition' || node.type === 'function_declarator') {
const fnDecl =
node.type === 'function_declarator'
? node
: findFirstDescendantOfType(node, 'function_declarator');
const params = fnDecl?.childForFieldName('parameters') ?? null;
if (params !== null) {
for (let i = 0; i < params.namedChildCount; i++) {
const param = params.namedChild(i);
if (param === null || param.type !== 'parameter_declaration') continue;
const declarator = param.childForFieldName('declarator');
if (declarator !== null && extractDeclaratorLeafName(declarator) === varName) {
return param;
}
}
}
return null;
}
node = node.parent;
}
return null;
}
function declaredNameNode(declarator: SyntaxNode): SyntaxNode | null {
if (declarator.type !== 'init_declarator') return declarator;
for (let i = 0; i < declarator.namedChildCount; i++) {
const child = declarator.namedChild(i);
if (child === null) continue;
if (child.type === 'identifier') return child;
if (child.type.endsWith('_declarator')) return child;
}
return declarator.childForFieldName('declarator');
}
/** Normalize a type-specifier text for argument type matching.
* Strips qualifiers (const, volatile), namespace prefixes (std::),
* and pointer/reference markers. */
@@ -786,6 +930,25 @@ function normalizeCppTypeText(text: string): string {
return t;
}
function isKnownEnumName(node: SyntaxNode, typeName: string): boolean {
if (typeName === '' || typeName === 'unknown') return false;
let root: SyntaxNode = node;
while (root.parent !== null) root = root.parent;
const stack: SyntaxNode[] = [root];
while (stack.length > 0) {
const cur = stack.pop()!;
if (cur.type === 'enum_specifier') {
const name = cur.childForFieldName('name');
if (name?.text === typeName) return true;
}
for (let i = 0; i < cur.childCount; i++) {
const child = cur.child(i);
if (child !== null) stack.push(child);
}
}
return false;
}
/**
* Detect whether a `namespace_definition` AST node is inline.
* Tree-sitter-cpp exposes the `inline` keyword as an anonymous child
@@ -1247,7 +1410,9 @@ function extractDeclaratorLeafName(node: SyntaxNode): string | null {
const next =
cur.childForFieldName('declarator') ??
// parenthesized_declarator: single named child
(cur.type === 'parenthesized_declarator' ? cur.namedChild(0) : null);
(cur.type === 'parenthesized_declarator' || cur.type.endsWith('_declarator')
? cur.namedChild(0)
: null);
if (next === null) return null;
cur = next;
}
@@ -20,20 +20,27 @@
* NOT: flip compatible↔incompatible; pass through unknown.
*/
import type { ArityVerdict, Callsite, ConstraintContext, SymbolDefinition } from 'gitnexus-shared';
import type {
ArityVerdict,
Callsite,
ConstraintContext,
ParameterTypeClass,
SymbolDefinition,
} from 'gitnexus-shared';
import { classifyType, type TypeClass } from './type-classifier.js';
import type { ConstraintExpr, CppConstraintPayload } from './constraint-extractor.js';
type AtomicEvaluator = (argClasses: readonly TypeClass[]) => ArityVerdict;
interface ConstraintArgClass {
readonly typeClass: TypeClass;
readonly shape?: ParameterTypeClass;
}
type AtomicEvaluator = (args: readonly ConstraintArgClass[]) => ArityVerdict;
/**
* Curated Tier-A predicate registry — the four canonical
* `<type_traits>` variable templates whose truth tables are closed-form
* over our coarse `TypeClass` enum.
*
* Deferred predicates that need a cv/ref/pointer sidecar on
* `normalizeCppParamType` (today the normalizer strips those markers
* before storage) live in #1579 as one-line follow-up adds.
* Curated Tier-A predicate registry. Predicates that depend on pointer,
* reference, or cv shape consult `ConstraintContext.argumentTypeClasses`.
* Missing or unsupported shape returns 'unknown' to preserve monotonicity.
*/
// ISO `<type_traits>` treats `bool`, `char`, and the signed/unsigned char
// variants as integral types (§21.3.4 Table 48), so `is_integral_v<bool>`
@@ -46,30 +53,135 @@ function isIntegralClass(c: TypeClass | undefined): boolean {
}
const REGISTRY = new Map<string, AtomicEvaluator>([
['is_integral_v', (cls) => verdictFromBool(isIntegralClass(cls[0]), cls)],
['is_floating_point_v', (cls) => verdictFromBool(cls[0] === 'floating', cls)],
[
'is_void_v',
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && arg.typeClass === 'void'),
],
[
'is_integral_v',
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && isIntegralClass(arg.typeClass)),
],
[
'is_floating_point_v',
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && arg.typeClass === 'floating'),
],
[
'is_arithmetic_v',
(cls) => verdictFromBool(isIntegralClass(cls[0]) || cls[0] === 'floating', cls),
(args) =>
unaryVerdict(
args,
(arg) =>
isPlainValue(arg) && (isIntegralClass(arg.typeClass) || arg.typeClass === 'floating'),
),
],
[
'is_enum_v',
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && arg.typeClass === 'enum'),
],
[
'is_class_v',
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && arg.typeClass === 'class'),
],
[
'is_pointer_v',
(args) =>
unaryShapeVerdict(args, (shape) => shape.indirection === 'pointer' && shape.pointerDepth > 0),
],
[
'is_reference_v',
(args) =>
unaryShapeVerdict(
args,
(shape) => shape.indirection === 'lvalue-ref' || shape.indirection === 'rvalue-ref',
),
],
[
'is_const_v',
(args) =>
unaryShapeVerdict(args, (shape) => shape.cv === 'const' || shape.cv === 'const volatile', {
requireTopLevelCv: true,
}),
],
[
'is_volatile_v',
(args) =>
unaryShapeVerdict(args, (shape) => shape.cv === 'volatile' || shape.cv === 'const volatile', {
requireTopLevelCv: true,
}),
],
// NOTE: cv-qualifiers are stripped by `normalizeCppParamType` before the
// type token reaches `classifyType`, so `is_same_v<const T, T>` returns
// `'compatible'` instead of the ISO-correct `false`. Tracked under the
// cv-sidecar refactor in #1579's "Out of scope" list; until that lands
// this approximation matches the common `is_same_v<T, ConcreteType>`
// dispatch idiom and silently degrades on cv-distinct compares.
[
'is_same_v',
(cls) => {
if (cls.length < 2 || cls[0] === 'unknown' || cls[1] === 'unknown') return 'unknown';
return cls[0] === cls[1] ? 'compatible' : 'incompatible';
(args) => {
if (args.length < 2 || args[0].typeClass === 'unknown' || args[1].typeClass === 'unknown') {
return 'unknown';
}
return args[0].typeClass === args[1].typeClass ? 'compatible' : 'incompatible';
},
],
]);
function verdictFromBool(predicate: boolean, cls: readonly TypeClass[]): ArityVerdict {
if (cls[0] === 'unknown') return 'unknown';
return predicate ? 'compatible' : 'incompatible';
function unaryVerdict(
args: readonly ConstraintArgClass[],
predicate: (arg: ConstraintArgClass) => boolean,
): ArityVerdict {
const arg = args[0];
if (arg === undefined || arg.typeClass === 'unknown') return 'unknown';
return predicate(arg) ? 'compatible' : 'incompatible';
}
function unaryShapeVerdict(
args: readonly ConstraintArgClass[],
predicate: (shape: ParameterTypeClass) => boolean,
options: { readonly requireTopLevelCv?: boolean } = {},
): ArityVerdict {
const arg = args[0];
if (arg === undefined || arg.typeClass === 'unknown') return 'unknown';
const shape = arg.shape;
if (shape === undefined || shape.indirection === 'unknown' || shape.cv === 'unknown') {
return 'unknown';
}
if (options.requireTopLevelCv === true && shape.indirection === 'pointer') {
return 'unknown';
}
return predicate(shape) ? 'compatible' : 'incompatible';
}
function isPlainValue(arg: ConstraintArgClass): boolean {
const shape = arg.shape;
if (shape === undefined) return true;
return shape.indirection === 'value';
}
function classifyConstraintArg(
token: string | undefined,
shape?: ParameterTypeClass,
): ConstraintArgClass {
if (shape !== undefined && shape.base.startsWith('enum:')) {
return { typeClass: 'enum', shape };
}
const typeClass = token === undefined || token === '' ? 'unknown' : classifyType(token);
return { typeClass, ...(shape !== undefined ? { shape } : {}) };
}
function tokenForArg(ctx: ConstraintContext, argIdx: number): string | undefined {
const shape = ctx.argumentTypeClasses?.[argIdx];
if (shape?.base.startsWith('enum:')) return shape.base;
return ctx.argumentTypes?.[argIdx];
}
function shapeForTemplateParam(
ctx: ConstraintContext,
paramName: string,
argIdx: number,
def?: SymbolDefinition,
): ParameterTypeClass | undefined {
const argShape = ctx.argumentTypeClasses?.[argIdx];
if (argShape === undefined) return undefined;
const paramShape = def?.parameterTypeClasses?.[argIdx];
if (paramShape === undefined) return argShape;
if (paramShape.base === paramName && paramShape.indirection === 'value') return argShape;
return undefined;
}
/** Public surface — registered as `ScopeResolver.constraintCompatibility`. */
@@ -80,13 +192,14 @@ export function cppConstraintCompatibility(
): ArityVerdict {
const payload = def.templateConstraints as CppConstraintPayload | undefined;
if (payload === undefined) return 'unknown';
return evaluate(payload.expr, payload, ctx);
return evaluate(payload.expr, payload, ctx, def);
}
function evaluate(
expr: ConstraintExpr,
payload: CppConstraintPayload,
ctx: ConstraintContext,
def?: SymbolDefinition,
): ArityVerdict {
switch (expr.kind) {
case 'unknown':
@@ -96,17 +209,18 @@ function evaluate(
if (evaluator === undefined) return 'unknown';
const classes = expr.args.map((paramName) => {
const argIdx = payload.paramArgIndex[paramName];
if (argIdx === undefined) return 'unknown' as TypeClass;
const token = ctx.argumentTypes?.[argIdx];
if (token === undefined || token === '') return 'unknown' as TypeClass;
return classifyType(token);
if (argIdx === undefined) return { typeClass: 'unknown' as TypeClass };
return classifyConstraintArg(
tokenForArg(ctx, argIdx),
shapeForTemplateParam(ctx, paramName, argIdx, def),
);
});
return evaluator(classes);
}
case 'and': {
let result: ArityVerdict = 'compatible';
for (const child of expr.children) {
const v = evaluate(child, payload, ctx);
const v = evaluate(child, payload, ctx, def);
if (v === 'incompatible') return 'incompatible';
if (v === 'unknown') result = 'unknown';
}
@@ -115,14 +229,14 @@ function evaluate(
case 'or': {
let result: ArityVerdict = 'incompatible';
for (const child of expr.children) {
const v = evaluate(child, payload, ctx);
const v = evaluate(child, payload, ctx, def);
if (v === 'compatible') return 'compatible';
if (v === 'unknown') result = 'unknown';
}
return result;
}
case 'not': {
const v = evaluate(expr.child, payload, ctx);
const v = evaluate(expr.child, payload, ctx, def);
if (v === 'compatible') return 'incompatible';
if (v === 'incompatible') return 'compatible';
return 'unknown';
@@ -1,32 +1,30 @@
/**
* C++ conversion-rank scoring for overload resolution (#1578).
* C++ conversion-rank scoring for overload resolution (#1578, #1637).
*
* Operates on **normalized** type strings (output of
* `normalizeCppParamType` in `arity-metadata.ts`). After normalization:
* - int/long/short/unsigned → 'int'
* - float/double → 'double'
* - char → 'char', bool → 'bool'
*
* Because the normalizer collapses promotion pairs (int↔long,
* float↔double) to the same string, those promotions are invisible at
* this layer — they appear as exact matches (rank 0).
* Operates on normalized type strings (output of `normalizeCppParamType`
* in `arity-metadata.ts`) plus optional shape sidecars from #1630.
* Normalization intentionally collapses cv/ref/pointer spelling for stable
* graph IDs, so pointer/nullptr rules must consult `ParameterTypeClass`.
*
* Post-normalization ranking:
* - rank 0 — exact (same normalized type)
* - rank 1 — integral promotion (char→int, bool→int)
* - rank 2 — standard arithmetic conversion (int↔double, char→double,
* bool→double)
* - Infinity — mismatch (string↔int, user types, pointers, etc.)
* - rank 0: exact (same normalized type)
* - rank 1: integral promotion (char -> int, bool -> int)
* - rank 2: standard conversion (arithmetic, nullptr -> T*, T* -> bool,
* T* -> void*)
* - rank 3: nullptr -> bool (kept worse than nullptr -> T*)
* - rank 4: ellipsis conversion (worst viable)
* - Infinity: mismatch (string -> int, user types, unsupported shapes)
*
* This function is intentionally C++-specific (issue #1578 pitfall:
* keep conversion-rank tables out of shared overload-narrowing). Other
* languages may define their own `ConversionRankFn` in the future.
* This function is intentionally C++-specific. Other languages may define
* their own `ConversionRankFn` in the future.
*/
import type { ParameterTypeClass } from 'gitnexus-shared';
/** Set of normalized arithmetic types that support implicit conversion. */
const ARITHMETIC = new Set(['int', 'double', 'char', 'bool']);
/** Integral promotion targets: char→int and bool→int are rank 1. */
/** Integral promotion targets: char -> int and bool -> int are rank 1. */
const INTEGRAL_PROMOTION = new Map([
['char', 'int'],
['bool', 'int'],
@@ -35,13 +33,40 @@ const INTEGRAL_PROMOTION = new Map([
/**
* Return the conversion rank from `argType` to `paramType`.
*
* @returns 0 for exact match, 1 for integral promotion (char/bool→int),
* 2 for standard arithmetic conversion, Infinity for mismatch.
* @returns 0 for exact match, 1 for integral promotion, 2 for standard
* conversion, 3 for nullptr -> bool, 4 for ellipsis, Infinity
* for mismatch.
*/
export function cppConversionRank(argType: string, paramType: string): number {
if (argType === paramType) return 0;
// Integral promotions: char→int, bool→int (ISO C++ [conv.prom])
export function cppConversionRank(
argType: string,
paramType: string,
argTypeClass?: ParameterTypeClass,
paramTypeClass?: ParameterTypeClass,
): number {
if (argType === paramType) {
return exactShapeCompatible(argTypeClass, paramTypeClass) ? 0 : Infinity;
}
if (paramType === '...') return 4;
if (INTEGRAL_PROMOTION.get(argType) === paramType) return 1;
if (ARITHMETIC.has(argType) && ARITHMETIC.has(paramType)) return 2;
if (argType === 'null' && isPointer(paramTypeClass)) return 2;
if (argType === 'null' && paramType === 'bool') return 3;
if (isPointer(argTypeClass) && paramType === 'bool') return 2;
if (isPointer(argTypeClass) && isPointer(paramTypeClass) && paramType === 'void') return 2;
return Infinity;
}
function isPointer(typeClass: ParameterTypeClass | undefined): boolean {
return typeClass?.indirection === 'pointer' && typeClass.pointerDepth > 0;
}
function exactShapeCompatible(
argTypeClass: ParameterTypeClass | undefined,
paramTypeClass: ParameterTypeClass | undefined,
): boolean {
if (argTypeClass === undefined || paramTypeClass === undefined) return true;
if (argTypeClass.indirection === 'unknown' || paramTypeClass.indirection === 'unknown') {
return true;
}
return isPointer(argTypeClass) === isPointer(paramTypeClass);
}
@@ -7,11 +7,10 @@
* the call-site inference in `captures.ts`) to one of the categories
* the `<type_traits>` predicate registry uses for SFINAE filtering.
*
* Intentionally coarse: cv / pointer / reference qualifiers are stripped
* upstream by `normalizeCppParamType`. Tier-A predicates
* (`is_integral_v`, `is_floating_point_v`, `is_arithmetic_v`, `is_same_v`)
* are insensitive to those modifiers per ISO `<type_traits>` semantics
* ("including any cv-qualified variants").
* `argumentTypes` remain normalized for overload narrowing, while
* constraint predicates that need cv/ref/pointer shape read the parallel
* `argumentTypeClasses` sidecar. Unknown shapes must stay unknown rather
* than being guessed as incompatible.
*/
export type TypeClass =
@@ -21,7 +20,11 @@ export type TypeClass =
| 'char'
| 'string'
| 'null'
| 'void'
| 'enum'
| 'class'
| 'pointer'
| 'reference'
| 'unknown';
/**
@@ -29,13 +32,17 @@ export type TypeClass =
* inference table in `captures.ts:inferCppLiteralType` plus the std::
* normalization in `arity-metadata.ts:normalizeCppParamType`.
*
* Caller note: token must already be normalized (no `const`, no `&` / `*`,
* no `std::` prefix). Tokens passed via `ConstraintContext.argumentTypes`
* coming from `inferCppCallArgTypes` satisfy this.
* Caller note: token should be normalized for overload matching. Enum
* tokens produced by the C++ adapter use the internal `enum:<Name>`
* prefix so `is_enum_v` does not have to guess that every user token is
* class-like.
*/
export function classifyType(token: string): TypeClass {
if (token.length === 0) return 'unknown';
if (token.startsWith('enum:')) return 'enum';
switch (token) {
case 'void':
return 'void';
case 'int':
return 'integral';
case 'double':
@@ -27,6 +27,7 @@ import { javaMethodConfig } from '../method-extractors/configs/jvm.js';
import { createVariableExtractor } from '../variable-extractors/generic.js';
import { javaVariableConfig } from '../variable-extractors/configs/jvm.js';
import { createHeritageExtractor } from '../heritage-extractors/generic.js';
import type { SymbolDefinition } from 'gitnexus-shared';
import {
emitJavaScopeCaptures,
interpretJavaImport,
@@ -39,6 +40,48 @@ import {
resolveJavaImportTarget,
} from './java/index.js';
const orderJavaSameNameTypeCandidates = ({
callSiteFilePath,
candidates,
}: {
readonly typeName: string;
readonly callSiteFilePath: string;
readonly candidates: readonly SymbolDefinition[];
}): readonly SymbolDefinition[] | null => {
if (!callSiteFilePath.endsWith('.java')) return null;
if (candidates.length <= 1) return null;
const callerDir = splitDirectorySegments(callSiteFilePath);
const scored = candidates.map((candidate, index) => ({
candidate,
index,
score: sharedPrefixLength(callerDir, splitDirectorySegments(candidate.filePath)),
}));
const bestScore = Math.max(...scored.map((entry) => entry.score));
// When all candidates tie, we have no structural signal to prefer one path.
// Returning null keeps downstream ambiguity handling conservative.
if (scored.every((entry) => entry.score === bestScore)) return null;
const ordered = [...scored]
.sort((a, b) => b.score - a.score || a.index - b.index)
.map((entry) => entry.candidate);
return ordered;
};
const splitDirectorySegments = (filePath: string): string[] => {
const normalized = filePath.replace(/\\/g, '/');
// Remove empty segments from leading/trailing/multiple slashes, then drop filename.
const segments = normalized.split('/').filter(Boolean);
return segments.slice(0, -1);
};
const sharedPrefixLength = (left: readonly string[], right: readonly string[]): number => {
const max = Math.min(left.length, right.length);
let idx = 0;
while (idx < max && left[idx] === right[idx]) idx += 1;
return idx;
};
export const javaProvider = defineLanguage({
id: SupportedLanguages.Java,
extensions: ['.java'],
@@ -87,4 +130,5 @@ export const javaProvider = defineLanguage({
receiverBinding: javaReceiverBinding,
arityCompatibility: javaArityCompatibility,
resolveImportTarget: resolveJavaImportTarget,
orderSameNameTypeCandidates: orderJavaSameNameTypeCandidates,
});
@@ -0,0 +1,12 @@
/**
* Arity compatibility for JavaScript.
*
* Delegates to `typescriptArityCompatibility` unchanged — JavaScript
* supports the same arity constructs (rest parameters `...args`, default
* parameters `p = v`) and the metadata shape (`parameterCount`,
* `requiredParameterCount`, `parameterTypes`) is synthesized by the same
* `computeTsArityMetadata` function (which understands both TS and JS
* parameter node types via `extractTsJsParameters`).
*/
export { typescriptArityCompatibility as jsArityCompatibility } from '../typescript/arity.js';
@@ -0,0 +1,722 @@
/**
* `emitScopeCaptures` for JavaScript.
*
* Adapts `emitTsScopeCaptures` for the JavaScript grammar:
*
* 1. **JS grammar** — uses `tree-sitter-javascript` instead of
* `tree-sitter-typescript`. The JS scope query is a subset of the
* TypeScript one (TypeScript-only node types dropped).
*
* 2. **CJS `require()` decomposition** — `const { X } = require('./m')`
* and `const X = require('./m')` are walked in a post-query pass and
* synthesized as `@import.kind/name/alias/source` markers so that
* `interpretJsImport` can recover a `ParsedImport` using the same
* shape as the TypeScript ESM decomposer.
*
* 3. **JSDoc type bindings** — JavaScript has no static type annotations
* so `@type-binding.parameter` / `@type-binding.return` must be
* inferred from leading JSDoc comments. A lightweight regex scanner
* (`parseJsDocParams` / `parseJsDocReturn`) extracts `@param {T} n`
* and `@returns {T}` tags and emits synthetic captures positioned on
* the annotated function node.
*
* 4. **Shared synthesis passes** — destructuring, for-of map-tuple, and
* instanceof narrowing passes are duplicated from `typescript/captures.ts`
* (they are pure AST operations with no grammar-specific logic).
*
* Pure given the input source text. No I/O, no globals consulted.
*/
import type { Capture, CaptureMatch } from 'gitnexus-shared';
import {
findNodeAtRange,
nodeToCapture,
syntheticCapture,
type SyntaxNode,
} from '../../utils/ast-helpers.js';
import { splitImportStatement } from '../typescript/import-decomposer.js';
import { getJsParser, getJsScopeQuery, jsCachedTreeMatchesGrammar } from './query.js';
import { computeTsArityMetadata } from '../typescript/arity-metadata.js';
import { synthesizeTsReceiverBinding } from '../typescript/receiver-binding.js';
import { getTreeSitterBufferSize } from '../../constants.js';
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
/** JS function-like node types that may carry a synthesized `this` binding.
* Kept in sync with the `@scope.function` patterns in `query.ts`. */
const FUNCTION_NODE_TYPES = [
'method_definition',
'arrow_function',
'function_expression',
'function_declaration',
'generator_function_declaration',
] as const;
/** Declaration anchors that carry function-like arity metadata. */
const FUNCTION_DECL_TAGS = ['@declaration.method', '@declaration.function'] as const;
/** Callsite anchors that should carry `@reference.arity` + param types. */
const CALL_TAGS = [
'@reference.call.free',
'@reference.call.member',
'@reference.call.constructor',
] as const;
function pickFirstDefined(grouped: CaptureMatch, tags: readonly string[]): Capture | undefined {
for (const tag of tags) {
const cap = grouped[tag];
if (cap !== undefined) return cap;
}
return undefined;
}
/** Filter `@reference.read.member` in non-read contexts (same logic as TS). */
function shouldEmitReadMember(memberNode: SyntaxNode): boolean {
const parent = memberNode.parent;
if (parent === null) return true;
switch (parent.type) {
case 'call_expression':
return parent.childForFieldName('function')?.id !== memberNode.id;
case 'new_expression':
return parent.childForFieldName('constructor')?.id !== memberNode.id;
case 'assignment_expression':
case 'augmented_assignment_expression':
return parent.childForFieldName('left')?.id !== memberNode.id;
case 'jsx_self_closing_element':
case 'jsx_opening_element':
return parent.childForFieldName('name')?.id !== memberNode.id;
default:
return true;
}
}
/** Find the first JS function-like node at the given range. */
function findFunctionNode(rootNode: SyntaxNode, range: Capture['range']): SyntaxNode | null {
for (const nodeType of FUNCTION_NODE_TYPES) {
const n = findNodeAtRange(rootNode, range, nodeType);
if (n !== null) return n;
}
return null;
}
/** Infer a callsite argument's static type from literal shapes. */
function inferArgType(argNode: SyntaxNode): string {
switch (argNode.type) {
case 'number':
return 'number';
case 'string':
case 'template_string':
return 'string';
case 'true':
case 'false':
return 'boolean';
case 'null':
return 'null';
case 'undefined':
return 'undefined';
case 'array':
return 'Array';
case 'object':
return 'object';
case 'regex':
return 'RegExp';
case 'new_expression': {
const ctor = argNode.childForFieldName('constructor');
return ctor?.text ?? '';
}
default:
return '';
}
}
// ─── CJS require() decomposition ─────────────────────────────────────────
/**
* Walk the AST and synthesize `@import.*` captures for CJS `require()` calls:
*
* - `const { X, Y } = require('./m')` → one match per destructured name,
* `@import.kind = 'named'`, `@import.name = X / Y`.
* - `const X = require('./m')` → `@import.kind = 'namespace'`,
* `@import.alias = X` (the whole module is bound to X).
* - `require('./m')` as a bare expression-statement → side-effect.
*
* CJS named-alias form (`const { X: alias } = require('./m')`) emits
* `@import.kind = 'named-alias'` with `@import.name = X` and
* `@import.alias = alias`.
*
* The synthesized markers are identical to those produced by
* `splitImportStatement` for ESM, so `interpretJsImport` can delegate
* unchanged to `interpretTsImport` for all cases.
*/
function synthesizeCjsImports(root: SyntaxNode, out: CaptureMatch[]): void {
const stack: SyntaxNode[] = [root];
for (;;) {
const node = stack.pop();
if (node === undefined) break;
for (const child of node.namedChildren) {
if (child !== null) stack.push(child);
}
if (node.type !== 'call_expression') continue;
// Require call: function must be bare identifier "require".
const fn = node.childForFieldName('function');
if (fn === null || fn.type !== 'identifier' || fn.text !== 'require') continue;
const argsNode = node.childForFieldName('arguments');
if (argsNode === null) continue;
// Source must be a string literal.
const firstArg = argsNode.namedChild(0);
if (firstArg === null || firstArg.type !== 'string') continue;
const rawSource = firstArg.text; // includes surrounding quotes
const source = firstArg.namedChild(0)?.text ?? rawSource.slice(1, -1);
const parent = node.parent;
// Case 1: const { X } = require('./m') OR const X = require('./m')
if (parent?.type === 'variable_declarator') {
const nameNode = parent.childForFieldName('name');
if (nameNode === null) continue;
if (nameNode.type === 'object_pattern') {
// Destructured: emit one match per specifier.
for (const field of nameNode.namedChildren) {
if (field === null) continue;
if (field.type === 'shorthand_property_identifier_pattern') {
const name = field.text;
out.push({
'@import.statement': syntheticCapture('@import.statement', node, rawSource),
'@import.kind': syntheticCapture('@import.kind', node, 'named'),
'@import.name': syntheticCapture('@import.name', field, name),
'@import.source': syntheticCapture('@import.source', firstArg, source),
});
} else if (field.type === 'pair_pattern') {
const key = field.childForFieldName('key');
const value = field.childForFieldName('value');
if (key === null || value === null || value.type !== 'identifier') continue;
out.push({
'@import.statement': syntheticCapture('@import.statement', node, rawSource),
'@import.kind': syntheticCapture('@import.kind', node, 'named-alias'),
'@import.name': syntheticCapture('@import.name', key, key.text),
'@import.alias': syntheticCapture('@import.alias', value, value.text),
'@import.source': syntheticCapture('@import.source', firstArg, source),
});
}
}
} else if (nameNode.type === 'identifier') {
// Namespace-style: const X = require('./m') → bind whole module to X.
out.push({
'@import.statement': syntheticCapture('@import.statement', node, rawSource),
'@import.kind': syntheticCapture('@import.kind', node, 'namespace'),
'@import.alias': syntheticCapture('@import.alias', nameNode, nameNode.text),
'@import.source': syntheticCapture('@import.source', firstArg, source),
});
}
continue;
}
// Case 2: bare require('./m') — side-effect import.
if (parent?.type === 'expression_statement') {
out.push({
'@import.statement': syntheticCapture('@import.statement', node, rawSource),
'@import.kind': syntheticCapture('@import.kind', node, 'side-effect'),
'@import.source': syntheticCapture('@import.source', firstArg, source),
});
}
}
}
// ─── JSDoc type binding synthesis ────────────────────────────────────────
interface JsDocParam {
readonly name: string;
readonly type: string;
}
/** Extract `@param {Type} name` entries from a JSDoc comment block. */
function parseJsDocParams(text: string): readonly JsDocParam[] {
const results: JsDocParam[] = [];
// Match @param {Type} name or @param {Type} [name] (optional)
const re = /@param\s+\{([^}]+)\}\s+\[?(\w+)\]?/g;
let m: RegExpExecArray | null;
while ((m = re.exec(text)) !== null) {
results.push({ type: m[1].trim(), name: m[2].trim() });
}
return results;
}
/** Extract `@returns {Type}` or `@return {Type}` from a JSDoc comment. */
function parseJsDocReturn(text: string): string | null {
const m = /@returns?\s+\{([^}]+)\}/.exec(text);
return m ? m[1].trim() : null;
}
/** Extract `@type {Type}` from a JSDoc comment (variable-level annotation). */
function parseJsDocType(text: string): string | null {
const m = /@type\s+\{([^}]+)\}/.exec(text);
return m ? m[1].trim() : null;
}
/**
* Walk the AST and synthesize `@type-binding.*` captures from JSDoc
* comments immediately preceding function declarations / expressions.
*
* Only `/** … *​/` block comments are scanned. Line comments (`//`) are
* intentionally excluded — JSDoc lives in block comments.
*
* Emits:
* - `@type-binding.parameter` for each `@param {T} n` tag.
* - `@type-binding.return` for `@returns {T}` / `@return {T}`.
* - `@type-binding.annotation` for `@type {T}` on `let`/`const`/`var`
* declarations — covers the common `/** @type {User} *​/ const u = …`
* pattern (ECMA-262 §14.3.1/§14.3.2 variable declarations).
*
* The binding is anchored on the function node so `tsBindingScopeFor`
* can hoist method return-type bindings to Module scope (matching the
* TypeScript path where `hoistTypeBindingsToModule: true`).
*/
function synthesizeJsDocBindings(root: SyntaxNode, out: CaptureMatch[]): void {
const stack: SyntaxNode[] = [root];
for (;;) {
const node = stack.pop();
if (node === undefined) break;
for (const child of node.namedChildren) {
if (child !== null) stack.push(child);
}
const isFnDecl =
node.type === 'function_declaration' || node.type === 'generator_function_declaration';
const isMethodDef = node.type === 'method_definition';
// Also check lexical_declaration containing an arrow/fn-expression
const isLexDecl = node.type === 'lexical_declaration' || node.type === 'variable_declaration';
if (!isFnDecl && !isMethodDef && !isLexDecl) continue;
// For `export function foo() { ... }`, the JSDoc comment precedes the
// wrapping export_statement, not the inner function_declaration.
// Walk up to the export_statement so the preceding-sibling search finds it.
const lookupNode =
(isFnDecl || isLexDecl) && node.parent?.type === 'export_statement' ? node.parent : node;
// Find the preceding sibling comment.
let sibling = lookupNode.previousNamedSibling;
while (sibling !== null && sibling.type === 'comment') {
const text = sibling.text;
if (text.startsWith('/**')) {
// Found a JSDoc block.
const params = parseJsDocParams(text);
const retType = parseJsDocReturn(text);
const varType = isLexDecl ? parseJsDocType(text) : null;
// Determine the anchor node (the function-like node, for hoisting).
const anchor = node;
for (const p of params) {
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', anchor, p.name),
'@type-binding.type': syntheticCapture('@type-binding.type', anchor, p.type),
'@type-binding.parameter': syntheticCapture('@type-binding.parameter', anchor, '1'),
});
}
if (retType !== null) {
// For named functions, use the function name as the binding name so
// `hoistTypeBindingsToModule` knows which function's return type this is.
let fnName: string | null = null;
if (isFnDecl) {
fnName = node.childForFieldName('name')?.text ?? null;
} else if (isMethodDef) {
// method_definition uses `name:` field for the method name
const nameNode = node.childForFieldName('name');
if (nameNode?.type === 'property_identifier') fnName = nameNode.text;
} else if (isLexDecl) {
const declarator = node.namedChild(0);
const nameNode = declarator?.childForFieldName('name');
if (nameNode?.type === 'identifier') fnName = nameNode.text;
}
if (fnName !== null) {
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', anchor, fnName),
'@type-binding.type': syntheticCapture('@type-binding.type', anchor, retType),
'@type-binding.return': syntheticCapture('@type-binding.return', anchor, '1'),
});
}
}
// @type {T} on let/const/var: `/** @type {User} */ const u = getUser()`.
// Emits annotation-strength binding (source = 'annotation') so it
// overrides any weaker constructor/alias inference on the same name.
if (varType !== null) {
for (const declarator of node.namedChildren) {
if (declarator === null || declarator.type !== 'variable_declarator') continue;
const nameNode = declarator.childForFieldName('name');
if (nameNode === null || nameNode.type !== 'identifier') continue;
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', nameNode, nameNode.text),
'@type-binding.type': syntheticCapture('@type-binding.type', nameNode, varType),
'@type-binding.annotation': syntheticCapture(
'@type-binding.annotation',
nameNode,
'1',
),
});
}
}
break;
}
sibling = sibling.previousNamedSibling;
}
}
}
// ─── Destructuring / for-of / instanceof (shared with TS captures) ───────
function synthesizeDestructuringBindings(root: SyntaxNode, out: CaptureMatch[]): void {
const stack: SyntaxNode[] = [root];
for (;;) {
const node = stack.pop();
if (node === undefined) break;
for (const child of node.namedChildren) {
if (child !== null) stack.push(child);
}
if (node.type !== 'variable_declarator') continue;
const nameNode = node.childForFieldName('name');
const valueNode = node.childForFieldName('value');
if (nameNode === null || valueNode === null) continue;
if (nameNode.type !== 'object_pattern') continue;
if (valueNode.type !== 'identifier') continue;
const rhsName = valueNode.text;
for (const fieldNode of nameNode.namedChildren) {
if (fieldNode === null) continue;
if (fieldNode.type === 'shorthand_property_identifier_pattern') {
const localName = fieldNode.text;
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', fieldNode, localName),
'@type-binding.type': syntheticCapture(
'@type-binding.type',
fieldNode,
`${rhsName}.${localName}`,
),
'@type-binding.destructured': syntheticCapture(
'@type-binding.destructured',
fieldNode,
fieldNode.text,
),
});
} else if (fieldNode.type === 'pair_pattern') {
const key = fieldNode.childForFieldName('key');
const value = fieldNode.childForFieldName('value');
if (key === null || value === null || value.type !== 'identifier') continue;
const fieldName = key.text;
const localName = value.text;
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', value, localName),
'@type-binding.type': syntheticCapture(
'@type-binding.type',
fieldNode,
`${rhsName}.${fieldName}`,
),
'@type-binding.destructured': syntheticCapture(
'@type-binding.destructured',
fieldNode,
fieldNode.text,
),
});
}
}
}
}
function synthesizeForOfMapTupleBindings(root: SyntaxNode, out: CaptureMatch[]): void {
const stack: SyntaxNode[] = [root];
for (;;) {
const node = stack.pop();
if (node === undefined) break;
for (const child of node.namedChildren) {
if (child !== null) stack.push(child);
}
if (node.type !== 'for_in_statement') continue;
const left = node.childForFieldName('left');
const right = node.childForFieldName('right');
if (left === null || right === null) continue;
if (left.type !== 'array_pattern' || right.type !== 'identifier') continue;
const rhs = right.text;
let slot = 0;
for (const child of left.namedChildren) {
if (child === null || child.type !== 'identifier') continue;
const localName = child.text;
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', child, localName),
'@type-binding.type': syntheticCapture(
'@type-binding.type',
child,
`__MAP_TUPLE_${slot}__:${rhs}`,
),
'@type-binding.map-tuple-entry': syntheticCapture(
'@type-binding.map-tuple-entry',
child,
String(slot),
),
});
slot++;
}
}
}
function synthesizeInstanceofNarrowings(root: SyntaxNode, out: CaptureMatch[]): void {
const stack: SyntaxNode[] = [root];
for (;;) {
const node = stack.pop();
if (node === undefined) break;
for (const child of node.namedChildren) {
if (child !== null) stack.push(child);
}
if (node.type !== 'if_statement') continue;
const cond = node.childForFieldName('condition');
if (cond === null) continue;
const inner = cond.type === 'parenthesized_expression' ? cond.namedChildren[0] : cond;
if (inner === null || inner.type !== 'binary_expression') continue;
const op = inner.childForFieldName('operator');
const left = inner.childForFieldName('left');
const right = inner.childForFieldName('right');
if (op === null || left === null || right === null) continue;
if (op.type !== 'instanceof') continue;
if (left.type !== 'identifier') continue;
if (right.type !== 'identifier') continue;
const varName = left.text;
const typeName = right.text;
const cons = node.childForFieldName('consequence');
if (cons === null) continue;
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', cons, varName),
'@type-binding.type': syntheticCapture('@type-binding.type', right, typeName),
'@type-binding.instanceof-narrow': syntheticCapture(
'@type-binding.instanceof-narrow',
cons,
'1',
),
});
}
}
// ─── Constructor field type bindings ─────────────────────────────────────
/**
* Synthesize class-scope type bindings from `this.X = new Y()` assignments
* inside constructor method bodies. Covers the traditional ES5+ OOP pattern:
*
* class User {
* constructor() {
* /** @type {Address} *\/
* this.address = new Address();
* }
* }
*
* The emitted `@type-binding.class-field` is hoisted to the Class scope by
* `tsBindingScopeFor` so that compound-receiver resolution can look up
* `User.address → Address` when resolving `user.address.save()`.
*
* Type source priority:
* 1. JSDoc `@type {T}` comment immediately preceding the statement
* 2. `new Y()` constructor inference
*/
function synthesizeConstructorFieldBindings(root: SyntaxNode, out: CaptureMatch[]): void {
const stack: SyntaxNode[] = [root];
for (;;) {
const node = stack.pop();
if (node === undefined) break;
for (const child of node.namedChildren) {
if (child !== null) stack.push(child);
}
// Only process constructor method definitions
if (node.type !== 'method_definition') continue;
const nameNode = node.childForFieldName('name');
if (nameNode?.text !== 'constructor') continue;
const body = node.childForFieldName('body');
if (body === null) continue;
for (const stmt of body.namedChildren) {
if (stmt === null || stmt.type !== 'expression_statement') continue;
const expr = stmt.namedChild(0);
if (expr === null || expr.type !== 'assignment_expression') continue;
const left = expr.childForFieldName('left');
const right = expr.childForFieldName('right');
if (left === null || right === null) continue;
if (left.type !== 'member_expression') continue;
const obj = left.childForFieldName('object');
const prop = left.childForFieldName('property');
if (obj === null || prop === null) continue;
if (obj.text !== 'this' || prop.type !== 'property_identifier') continue;
const fieldName = prop.text;
// Prefer JSDoc @type annotation on the preceding sibling comment.
let typeName: string | null = null;
const prevSib: SyntaxNode | null = stmt.previousNamedSibling;
if (prevSib !== null && prevSib.type === 'comment') {
const m = /@type\s*\{([^}]+)\}/.exec(prevSib.text);
if (m?.[1]) typeName = m[1].trim();
}
// Fall back to constructor inference from `new Y()`.
if (typeName === null && right.type === 'new_expression') {
const ctor = right.childForFieldName('constructor');
if (ctor !== null && ctor.type === 'identifier') typeName = ctor.text;
}
if (typeName === null) continue;
out.push({
'@type-binding.name': syntheticCapture('@type-binding.name', prop, fieldName),
'@type-binding.type': syntheticCapture('@type-binding.type', prop, typeName),
// Anchor: positioned inside the constructor body so tsBindingScopeFor
// can walk up from the Function (constructor) scope to the Class scope.
'@type-binding.class-field': syntheticCapture('@type-binding.class-field', stmt, '1'),
});
}
}
}
// ─── Main emitter ──────────────────────────────────────────────────────────
export function emitJsScopeCaptures(
sourceText: string,
filePath: string,
cachedTree?: unknown,
): readonly CaptureMatch[] {
let tree = cachedTree as ReturnType<ReturnType<typeof getJsParser>['parse']> | undefined;
if (tree !== undefined && !jsCachedTreeMatchesGrammar(tree)) {
tree = undefined;
}
if (tree === undefined) {
tree = parseSourceSafe(getJsParser(filePath), sourceText, undefined, {
bufferSize: getTreeSitterBufferSize(sourceText),
});
}
const rawMatches = getJsScopeQuery(filePath).matches(tree.rootNode);
const out: CaptureMatch[] = [];
for (const m of rawMatches) {
const grouped: Record<string, Capture> = {};
for (const c of m.captures) {
const tag = '@' + c.name;
grouped[tag] = nodeToCapture(tag, c.node);
}
if (Object.keys(grouped).length === 0) continue;
// Decompose ESM import_statement / re-export export_statement.
if (grouped['@import.statement'] !== undefined) {
const stmtCapture = grouped['@import.statement'];
const stmtNode =
findNodeAtRange(tree.rootNode, stmtCapture.range, 'import_statement') ??
findNodeAtRange(tree.rootNode, stmtCapture.range, 'export_statement');
if (stmtNode !== null) {
const decomposed = splitImportStatement(stmtNode);
for (const d of decomposed) out.push(d);
}
continue;
}
// Decompose dynamic import() calls.
if (grouped['@import.dynamic'] !== undefined) {
const dynCapture = grouped['@import.dynamic'];
const callNode = findNodeAtRange(tree.rootNode, dynCapture.range, 'call_expression');
if (callNode !== null) {
const decomposed = splitImportStatement(callNode);
for (const d of decomposed) out.push(d);
}
continue;
}
// Filter @reference.read.member false-positives.
if (grouped['@reference.read.member'] !== undefined) {
const anchor = grouped['@reference.read.member'];
const memberNode = findNodeAtRange(tree.rootNode, anchor.range, 'member_expression');
if (memberNode === null || !shouldEmitReadMember(memberNode)) {
continue;
}
}
// Synthesize arity metadata on function-like declarations.
const declAnchor = pickFirstDefined(grouped, FUNCTION_DECL_TAGS);
if (declAnchor !== undefined) {
const fnNode = findFunctionNode(tree.rootNode, declAnchor.range);
if (fnNode !== null) {
const arity = computeTsArityMetadata(fnNode);
if (arity.parameterCount !== undefined) {
grouped['@declaration.parameter-count'] = syntheticCapture(
'@declaration.parameter-count',
fnNode,
String(arity.parameterCount),
);
}
if (arity.requiredParameterCount !== undefined) {
grouped['@declaration.required-parameter-count'] = syntheticCapture(
'@declaration.required-parameter-count',
fnNode,
String(arity.requiredParameterCount),
);
}
if (arity.parameterTypes !== undefined) {
grouped['@declaration.parameter-types'] = syntheticCapture(
'@declaration.parameter-types',
fnNode,
JSON.stringify(arity.parameterTypes),
);
}
}
}
// Synthesize @reference.arity on callsites.
const callAnchor = pickFirstDefined(grouped, CALL_TAGS);
if (callAnchor !== undefined && grouped['@reference.arity'] === undefined) {
const callNode =
findNodeAtRange(tree.rootNode, callAnchor.range, 'call_expression') ??
findNodeAtRange(tree.rootNode, callAnchor.range, 'new_expression');
if (callNode !== null) {
const argList = callNode.childForFieldName('arguments');
const args: SyntaxNode[] =
argList === null
? []
: argList.namedChildren.filter(
(c): c is SyntaxNode => c !== null && c.type !== 'comment',
);
grouped['@reference.arity'] = syntheticCapture(
'@reference.arity',
callNode,
String(args.length),
);
grouped['@reference.parameter-types'] = syntheticCapture(
'@reference.parameter-types',
callNode,
JSON.stringify(args.map(inferArgType)),
);
}
}
out.push(grouped);
// Synthesize `this` receiver type-bindings on class member functions.
const scopeFnAnchor = grouped['@scope.function'];
if (scopeFnAnchor !== undefined) {
const fnNode = findFunctionNode(tree.rootNode, scopeFnAnchor.range);
if (fnNode !== null) {
const synth = synthesizeTsReceiverBinding(fnNode);
if (synth !== null) out.push(synth);
}
}
}
// Post-query synthesis passes.
synthesizeCjsImports(tree.rootNode, out);
synthesizeJsDocBindings(tree.rootNode, out);
synthesizeConstructorFieldBindings(tree.rootNode, out);
synthesizeDestructuringBindings(tree.rootNode, out);
synthesizeForOfMapTupleBindings(tree.rootNode, out);
synthesizeInstanceofNarrowings(tree.rootNode, out);
return out;
}
@@ -0,0 +1,72 @@
/**
* Import-target resolver for JavaScript.
*
* Delegates to the TypeScript `resolveTsTarget` standard-strategy resolver
* with `language: SupportedLanguages.JavaScript` so the resolver tries
* `.js` / `.jsx` extensions in addition to (or instead of) `.ts` / `.tsx`.
*
* The `TsResolveContext.language` flag already exists in `import-target.ts`
* and the resolver (`resolveImportPath`) already branches on it — this
* adapter just wires the right value in.
*
* CJS `require()` calls reference the same module-path strings as ESM
* `import` statements, so the resolver handles them uniformly without any
* CJS-specific logic here.
*
* No `tsconfig.json` path-alias support (JavaScript projects don't use
* `tsconfig.json` compilerOptions.paths in general). Projects that DO use
* tsconfig-based aliases alongside JavaScript can still resolve via the
* standard extension-suffix fallback; the alias branch is a no-op when
* `tsconfigPaths` is null.
*/
import { SupportedLanguages } from 'gitnexus-shared';
import { resolveTsTarget, type TsResolveContext } from '../typescript/import-target.js';
export type JsResolveContext = TsResolveContext;
type PassCache = {
readonly key: ReadonlySet<string>;
readonly allFilePaths: Set<string>;
readonly allFileList: readonly string[];
readonly normalizedFileList: readonly string[];
readonly resolveCache: Map<string, string | null>;
};
/**
* Build a memoized `resolveImportTarget` adapter for JavaScript.
* Caches the derived arrays and per-pass resolve cache across
* `resolveImportTarget` calls within a single workspace pass.
*/
export function makeJsResolveImportTarget(): (
targetRaw: string,
fromFile: string,
allFilePaths: ReadonlySet<string>,
resolutionConfig?: unknown,
) => string | readonly string[] | null {
let cached: PassCache | null = null;
return (targetRaw, fromFile, allFilePaths) => {
if (cached === null || cached.key !== allFilePaths) {
const allFileList = Array.from(allFilePaths);
cached = {
key: allFilePaths,
allFilePaths: new Set(allFilePaths),
allFileList,
normalizedFileList: allFileList.map((f) => f.toLowerCase()),
resolveCache: new Map(),
};
}
const ws: JsResolveContext = {
fromFile,
language: SupportedLanguages.JavaScript,
allFilePaths: cached.allFilePaths,
allFileList: cached.allFileList,
normalizedFileList: cached.normalizedFileList,
resolveCache: cached.resolveCache,
tsconfigPaths: null,
};
return resolveTsTarget(targetRaw, ws);
};
}
@@ -0,0 +1,49 @@
/**
* JavaScript scope-resolution hooks (RFC #909 Ring 3, issue #928).
*
* Public API barrel. Consumers should import from this file rather
* than the individual modules.
*
* Module layout (each file is a single concern):
*
* - `query.ts` — JS scope query string + lazy parser/query
* singletons (`getJsParser`, `getJsScopeQuery`)
* - `captures.ts` — `emitJsScopeCaptures` — runs the JS scope query,
* synthesizes CJS require() imports and JSDoc-
* derived type bindings, delegates arity synthesis
* and destructuring/instanceof passes to shared
* or TypeScript utilities
* - `interpret.ts` — `interpretJsImport` / `interpretJsTypeBinding`
* (delegate to TypeScript interpreters — same
* capture-marker vocabulary)
* - `simple-hooks.ts` — `jsBindingScopeFor` (var hoisting),
* `jsImportOwningScope`, `jsReceiverBinding`
* (all delegate to TypeScript counterparts)
* - `merge-bindings.ts` — `jsMergeBindings` (LEGB via typescriptMergeBindings)
* - `arity.ts` — `jsArityCompatibility` (delegates to TS function)
* - `import-target.ts` — `makeJsResolveImportTarget` (memoized adapter)
* - `scope-resolver.ts` — `javascriptScopeResolver` wiring object
*
* ## Known limitations
*
* 1. **JSDoc coverage** — `@param {T} name`, `@returns {T}` / `@return {T}`,
* and `@type {T}` on variable declarations are synthesized. `@typedef`
* is not yet synthesized (tracked in #1646).
* 2. **CJS chained destructuring** — `const { X: { Y } } = require(...)`
* (nested destructuring) emits only the outer `X` binding; `Y` is not
* resolved.
* 3. **Dynamic require** — `require(computedPath)` is skipped (non-literal
* argument — cannot statically resolve the target).
* 4. **`module.exports` / `exports.X`** — CJS export forms are not yet
* modeled as re-exports. The finalize algorithm treats the exporting
* module as a namespace; importers that do `const X = require('./m')`
* bind the module namespace, and member-call resolution walks the
* class graph from there.
*/
export { emitJsScopeCaptures } from './captures.js';
export { interpretJsImport, interpretJsTypeBinding } from './interpret.js';
export { jsMergeBindings } from './merge-bindings.js';
export { jsArityCompatibility } from './arity.js';
export { makeJsResolveImportTarget } from './import-target.js';
export { jsBindingScopeFor, jsImportOwningScope, jsReceiverBinding } from './simple-hooks.js';
@@ -0,0 +1,45 @@
/**
* Capture-match → semantic-shape interpreters for JavaScript.
*
* `interpretJsImport` delegates to `interpretTsImport` for all cases
* because `emitJsScopeCaptures` synthesizes the same
* `@import.kind/name/alias/source` markers for both ESM and CJS imports.
*
* The `@import.kind` values emitted for CJS by `captures.ts`:
*
* - `'named'` : `const { X } = require('./m')` → named import
* - `'named-alias'` : `const { X: Y } = require('./m')` → aliased import
* - `'namespace'` : `const X = require('./m')` → namespace import
* - `'side-effect'` : `require('./m')` bare expression → side-effect
*
* These match the kinds `interpretTsImport` already handles for ESM
* (`import { X }`, `import { X as Y }`, `import * as X`, `import './m'`),
* so no new branch is needed here.
*
* `interpretJsTypeBinding` handles the JS-only `@type-binding.class-field`
* tag before delegating to `interpretTsTypeBinding`. The class-field tag
* is emitted by `synthesizeConstructorFieldBindings` and should produce
* `source = 'annotation'` — the same strength as an explicit type
* annotation. Remapping it to `@type-binding.annotation` achieves this
* without adding a JS-specific branch to the shared TS interpreter
* (DoD.md §2.2).
*/
import type { CaptureMatch, ParsedImport, ParsedTypeBinding } from 'gitnexus-shared';
import { interpretTsImport, interpretTsTypeBinding } from '../typescript/interpret.js';
export function interpretJsImport(captures: CaptureMatch): ParsedImport | null {
return interpretTsImport(captures);
}
export function interpretJsTypeBinding(captures: CaptureMatch): ParsedTypeBinding | null {
// @type-binding.class-field is a JS-only tag emitted by
// synthesizeConstructorFieldBindings. Remap it to the standard
// @type-binding.annotation tag so interpretTsTypeBinding assigns
// source = 'annotation' without a JS-specific branch in shared code.
if (captures['@type-binding.class-field'] !== undefined) {
const { '@type-binding.class-field': classField, ...rest } = captures;
return interpretTsTypeBinding({ ...rest, '@type-binding.annotation': classField });
}
return interpretTsTypeBinding(captures);
}
@@ -0,0 +1,21 @@
/**
* Binding-merge precedence for JavaScript.
*
* JavaScript has no TypeScript declaration-merging (no `interface + class`
* coexisting in the same scope, no `namespace + class` dual-space declarations).
* However, `typescriptMergeBindings` handles these by falling back to
* `['value']` for any `NodeLabel` not explicitly mapped to multiple spaces —
* which is what every JavaScript declaration produces. The result is pure
* LEGB precedence without any cross-space logic, which is exactly what
* JavaScript needs.
*
* Reuse rather than reimplementing to keep the single source of truth for
* the tier (local 0 / import-namespace-reexport 1 / wildcard 2) ordering.
*/
import type { BindingRef } from 'gitnexus-shared';
import { typescriptMergeBindings } from '../typescript/merge-bindings.js';
export function jsMergeBindings(bindings: readonly BindingRef[]): readonly BindingRef[] {
return typescriptMergeBindings(bindings);
}
@@ -0,0 +1,421 @@
/**
* Tree-sitter query for JavaScript scope captures (RFC §5.1, Ring 3).
*
* Subset of the TypeScript scope query (`languages/typescript/query.ts`)
* compiled against `tree-sitter-javascript`. TypeScript-only node types
* (`interface_declaration`, `type_alias_declaration`, `enum_declaration`,
* `internal_module`, `abstract_class_declaration`, `function_signature`,
* `method_signature`, `abstract_method_signature`, `type_annotation`,
* `public_field_definition`) are dropped because:
*
* 1. The JS grammar doesn't define them — the query compiler would
* throw `InvalidNodeType` if they were included.
* 2. JavaScript has no static type annotations, so the `@type-binding.*`
* patterns derived from TS annotation nodes don't apply.
*
* What IS shared with the TypeScript query:
*
* - Scope patterns: `program`, `class_declaration`, `(class)` (the JS
* grammar node for class expressions — NOT `class_expression`, which
* does not exist in `tree-sitter-javascript`), `function_declaration`,
* `generator_function_declaration`, `function_expression`,
* `arrow_function`, `method_definition`.
* - Declaration patterns for functions, classes, const/let/var,
* object-property arrows (Zustand, TanStack, etc.), and HOC-wrapped
* variable declarations (forwardRef / memo / useCallback / useMemo).
* - Import patterns: `import_statement`, `export_statement` re-exports,
* and dynamic `import()` (represented as `call_expression(import)` in
* both grammars — the `import` leaf node exists in tree-sitter-javascript
* as well as tree-sitter-typescript).
* - Type-binding patterns that work without static annotations:
* constructor inference (`new User()`), call-result alias
* (`const u = getUser()`), member-access alias (`const a = u.addr`),
* identifier alias, assignment rebind, and for-of element bindings.
* JSDoc-derived type bindings (`@param {User} u`, `@returns {User}`)
* are handled separately in `captures.ts` via comment-node scanning.
* - Reference patterns: free calls, member calls, constructor calls,
* write-access, read-access, and dynamic import.
*
* CJS `require()` is NOT captured here; it is handled in `captures.ts`
* by scanning parent context (destructured vs. namespace) of `call_expression`
* nodes whose callee is the identifier `require`.
*
* Grammar version: `tree-sitter-javascript` pinned in gitnexus/package.json.
*
* Exposes lazy `Parser` and `Query` singletons so callers don't pay
* tree-sitter init cost per file.
*/
import Parser from 'tree-sitter';
import JS from 'tree-sitter-javascript';
const JS_GRAMMAR = JS as Parameters<Parser['setLanguage']>[0];
/** True when the file should be parsed with the JSX-extended query. */
function isJsxFile(filePath: string): boolean {
return filePath.endsWith('.jsx');
}
const JAVASCRIPT_SCOPE_QUERY = `
;; Scopes — module / class-likes / function-likes
(program) @scope.module
(class_declaration) @scope.class
(class) @scope.class
(function_declaration) @scope.function
(generator_function_declaration) @scope.function
(function_expression) @scope.function
(arrow_function) @scope.function
(method_definition) @scope.function
;; Declarations — classes
(class_declaration
name: (identifier) @declaration.name) @declaration.class
;; Declarations — methods (inside class bodies)
(method_definition
name: (property_identifier) @declaration.name) @declaration.method
;; Declarations — class fields (JS uses field_definition, not public_field_definition)
(field_definition
property: (property_identifier) @declaration.name) @declaration.property
;; Declarations — free functions
(function_declaration
name: (identifier) @declaration.name) @declaration.function
(generator_function_declaration
name: (identifier) @declaration.name) @declaration.function
;; Arrow / function-expression assigned to a const/let/var.
;; Anchor discipline: @declaration.function sits on the INNER arrow or
;; function_expression, NOT on the lexical_declaration wrapper. This
;; aligns anchor.range with the @scope.function range so
;; pass2AttachDeclarations resolves the innermost scope correctly and
;; resolveCallerGraphId walks up to the right caller anchor.
(lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (arrow_function) @declaration.function))
(lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (function_expression) @declaration.function))
(export_statement
declaration: (lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (arrow_function) @declaration.function)))
(export_statement
declaration: (lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (function_expression) @declaration.function)))
(variable_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (arrow_function) @declaration.function))
(variable_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (function_expression) @declaration.function))
;; Object-property arrows / function expressions named by their pair key.
;; Same anchor discipline as the lexical_declaration block above: the
;; @declaration.function capture must sit on the INNER arrow/fn-expression.
(pair
key: (property_identifier) @declaration.name
value: (arrow_function) @declaration.function)
(pair
key: (property_identifier) @declaration.name
value: (function_expression) @declaration.function)
(pair
key: (string (string_fragment) @declaration.name)
value: (arrow_function) @declaration.function)
(pair
key: (string (string_fragment) @declaration.name)
value: (function_expression) @declaration.function)
;; HOC-wrapped variable declarations: const X = HOC((args) => { ... }).
;; Covers React.forwardRef, memo, useCallback, useMemo, observer,
;; debounce, and any user-defined HOC factory.
(lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (call_expression
arguments: (arguments
(arrow_function) @declaration.function))))
(lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (call_expression
arguments: (arguments
(function_expression) @declaration.function))))
(export_statement
declaration: (lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (call_expression
arguments: (arguments
(arrow_function) @declaration.function)))))
(export_statement
declaration: (lexical_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (call_expression
arguments: (arguments
(function_expression) @declaration.function)))))
(variable_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (call_expression
arguments: (arguments
(arrow_function) @declaration.function))))
(variable_declaration
(variable_declarator
name: (identifier) @declaration.name
value: (call_expression
arguments: (arguments
(function_expression) @declaration.function))))
;; Variable / constant declarations (non-function values).
(lexical_declaration
(variable_declarator
name: (identifier) @declaration.name)) @declaration.const
(export_statement
declaration: (lexical_declaration
(variable_declarator
name: (identifier) @declaration.name))) @declaration.const
(variable_declaration
(variable_declarator
name: (identifier) @declaration.name)) @declaration.variable
;; Imports (ESM) — single anchor per statement; decomposer emits per-specifier markers.
(import_statement) @import.statement
;; Re-exports with a source clause.
(export_statement
source: (string)) @import.statement
;; Dynamic imports: import('./m') — tree-sitter-javascript represents this
;; as call_expression with a named import leaf as the function field,
;; identical to tree-sitter-typescript.
(call_expression
function: (import)) @import.dynamic
;; ── Type bindings (no static annotations in JS; inferred from AST shape) ──
;; Constructor-inferred: const u = new User()
(variable_declarator
name: (identifier) @type-binding.name
value: (new_expression
constructor: (identifier) @type-binding.type)) @type-binding.constructor
;; Qualified constructor: const u = new models.User()
(variable_declarator
name: (identifier) @type-binding.name
value: (new_expression
constructor: (member_expression) @type-binding.type)) @type-binding.constructor
;; Call-result alias: const u = getUser()
(variable_declarator
name: (identifier) @type-binding.name
value: (call_expression
function: (identifier) @type-binding.type)) @type-binding.alias
;; Member-call alias: const u = svc.getUser()
(variable_declarator
name: (identifier) @type-binding.name
value: (call_expression
function: (member_expression) @type-binding.type)) @type-binding.alias
;; Await chain: const u = await getUser() / await svc.getUser()
(variable_declarator
name: (identifier) @type-binding.name
value: (await_expression
(call_expression
function: (identifier) @type-binding.type))) @type-binding.alias
(variable_declarator
name: (identifier) @type-binding.name
value: (await_expression
(call_expression
function: (member_expression) @type-binding.type))) @type-binding.alias
;; Member-access alias: const addr = user.address
(variable_declarator
name: (identifier) @type-binding.name
value: (member_expression) @type-binding.type) @type-binding.member-alias
;; Identifier alias: const alias = user
(variable_declarator
name: (identifier) @type-binding.name
value: (identifier) @type-binding.type) @type-binding.alias
;; Assignment rebind: u = new User() / u = getUser()
(assignment_expression
left: (identifier) @type-binding.name
right: (new_expression
constructor: (identifier) @type-binding.type)) @type-binding.constructor
(assignment_expression
left: (identifier) @type-binding.name
right: (call_expression
function: (identifier) @type-binding.type)) @type-binding.alias
(assignment_expression
left: (identifier) @type-binding.name
right: (identifier) @type-binding.type) @type-binding.alias
;; For-of element: for (const u of users) / for (const u of getUsers())
(for_in_statement
left: (identifier) @type-binding.name
right: (identifier) @type-binding.type) @type-binding.alias
(for_in_statement
left: (identifier) @type-binding.name
right: (call_expression
function: (identifier) @type-binding.type)) @type-binding.alias
(for_in_statement
left: (identifier) @type-binding.name
right: (call_expression
function: (member_expression) @type-binding.type)) @type-binding.alias
(for_in_statement
left: (identifier) @type-binding.name
right: (member_expression
property: (property_identifier) @type-binding.type)) @type-binding.alias
;; ── References ────────────────────────────────────────────────────────────
;; Free calls: fn(args). The dynamic-import filter runs in captures.ts.
(call_expression
function: (identifier) @reference.name) @reference.call.free
;; Awaited free call: await fn<T>(...) re-associated by tree-sitter.
(call_expression
function: (await_expression
(identifier) @reference.name)) @reference.call.free
;; Member calls: obj.method() (includes optional chain).
(call_expression
function: (member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name)) @reference.call.member
;; Awaited member call: await svc.m<T>(...)
(call_expression
function: (await_expression
(member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name))) @reference.call.member
;; Constructor calls: new User() / new ns.User()
(new_expression
constructor: (identifier) @reference.name) @reference.call.constructor
(new_expression
constructor: (member_expression) @reference.call.constructor.qualified) @reference.call.constructor
;; Write access: obj.field = value
(assignment_expression
left: (member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name)) @reference.write.member
(augmented_assignment_expression
left: (member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name)) @reference.write.member
;; Read access: obj.field (in read context; captures.ts filters non-reads).
(member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name) @reference.read.member
`;
/** JSX-only suffix — appended when compiling against the JSX grammar for .jsx files. */
const JSX_QUERY_SUFFIX = `
;; <Foo />
((jsx_self_closing_element
name: (identifier) @reference.name) @reference.call.free
(#match? @reference.name "^[A-Z]"))
;; <Foo> ... </Foo>
((jsx_opening_element
name: (identifier) @reference.name) @reference.call.free
(#match? @reference.name "^[A-Z]"))
;; <Foo.Bar />
(jsx_self_closing_element
name: (member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name)) @reference.call.member
(jsx_opening_element
name: (member_expression
object: (_) @reference.receiver
property: (property_identifier) @reference.name)) @reference.call.member
`;
let _jsParser: Parser | null = null;
let _jsQuery: Parser.Query | null = null;
let _jsxParser: Parser | null = null;
let _jsxQuery: Parser.Query | null = null;
export function getJsParser(filePath?: string): Parser {
// JSX files use the same JavaScript grammar in tree-sitter-javascript;
// both .js and .jsx parse with the same grammar object. We keep separate
// singletons only to mirror the TypeScript pattern and in case a future
// version of the grammar diverges.
if (filePath !== undefined && isJsxFile(filePath)) {
if (_jsxParser === null) {
_jsxParser = new Parser();
_jsxParser.setLanguage(JS_GRAMMAR);
}
return _jsxParser;
}
if (_jsParser === null) {
_jsParser = new Parser();
_jsParser.setLanguage(JS_GRAMMAR);
}
return _jsParser;
}
export function getJsScopeQuery(filePath?: string): Parser.Query {
if (filePath !== undefined && isJsxFile(filePath)) {
if (_jsxQuery === null) {
_jsxQuery = new Parser.Query(JS_GRAMMAR, JAVASCRIPT_SCOPE_QUERY + JSX_QUERY_SUFFIX);
}
return _jsxQuery;
}
if (_jsQuery === null) {
_jsQuery = new Parser.Query(JS_GRAMMAR, JAVASCRIPT_SCOPE_QUERY);
}
return _jsQuery;
}
/** Validate that a cached Tree was produced by the JS grammar. */
export function jsCachedTreeMatchesGrammar(tree: unknown): boolean {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const lang = (tree as any)?.getLanguage?.();
if (lang === undefined || lang === null) return true;
return lang === JS_GRAMMAR;
}
@@ -0,0 +1,89 @@
/**
* JavaScript `ScopeResolver` registered in `SCOPE_RESOLVERS` and
* consumed by the generic `runScopeResolution` orchestrator
* (RFC #909 Ring 3, issue #928).
*
* Follows the same minimal wiring-only pattern as TypeScript (the third
* migration). Per-hook logic lives in sibling modules:
*
* - `query.ts` — JS scope query + parser/query singletons
* - `captures.ts` — `emitJsScopeCaptures` (JS grammar, CJS, JSDoc)
* - `interpret.ts` — `interpretJsImport` (delegates to TS interpreter)
* - `simple-hooks.ts` — `jsBindingScopeFor`, `jsImportOwningScope`,
* `jsReceiverBinding` (all delegate to TS hooks)
* - `merge-bindings.ts` — `jsMergeBindings` (delegates to TS function)
* - `arity.ts` — `jsArityCompatibility` (delegates to TS function)
* - `import-target.ts` — `makeJsResolveImportTarget` (TS resolver, JS extensions)
*
* See `./index.ts` for the full per-module rationale.
*
* ## Key differences from TypeScript resolver
*
* - `fieldFallbackOnMethodLookup: true` — JavaScript is dynamically typed;
* the field-fallback heuristic is ENABLED (unlike TypeScript, which
* disables it because the type-binding layer is precise).
* - `allowGlobalFreeCallFallback: true` — CJS `require` patterns and
* global helpers (e.g. `process`, `console`) benefit from workspace-
* wide unique-name fallback. TypeScript uses explicit imports.
* - `loadResolutionConfig` is omitted — JavaScript projects don't use
* `tsconfig.json` path aliases in general. `tsconfigPaths: null` is
* threaded through the resolver adapter.
* - `hoistTypeBindingsToModule: true` — JSDoc `@returns {T}` bindings are
* synthesized on the function scope and hoisted, matching TypeScript's
* method return-type hoisting strategy for cross-file chain resolution.
*/
import type { ParsedFile } from 'gitnexus-shared';
import { SupportedLanguages } from 'gitnexus-shared';
import { buildMro, defaultLinearize } from '../../scope-resolution/passes/mro.js';
import { populateClassOwnedMembers } from '../../scope-resolution/scope/walkers.js';
import type { ScopeResolver } from '../../scope-resolution/contract/scope-resolver.js';
import { javascriptProvider } from '../typescript.js';
import { jsMergeBindings } from './merge-bindings.js';
import { jsArityCompatibility } from './arity.js';
import { makeJsResolveImportTarget } from './import-target.js';
const javascriptScopeResolver: ScopeResolver = {
language: SupportedLanguages.JavaScript,
languageProvider: javascriptProvider,
importEdgeReason: 'javascript-scope: import',
resolveImportTarget: makeJsResolveImportTarget(),
// JavaScript LEGB — same tier ordering as TypeScript; no declaration-
// merging across type/value/namespace spaces.
mergeBindings: (existing, incoming) => [...jsMergeBindings([...existing, ...incoming])],
// Adapter: jsArityCompatibility uses (def, callsite); contract is (callsite, def).
arityCompatibility: (callsite, def) => jsArityCompatibility(def, callsite),
buildMro: (graph, parsedFiles, nodeLookup) =>
buildMro(graph, parsedFiles, nodeLookup, defaultLinearize),
populateOwners: (parsed: ParsedFile) => populateClassOwnedMembers(parsed),
// JavaScript `super` keyword: same pattern as TypeScript.
isSuperReceiver: (text) => /^super(\s*\(|\s*\.|\s*\[|\s*$)/.test(text.trim()),
// JavaScript is dynamically typed — enable the field-fallback heuristic
// so member-call receivers without type annotations can still resolve
// through declared class fields (e.g. JSDoc-typed fields).
fieldFallbackOnMethodLookup: true,
// Return-type propagation (across ESM imports) mirrors TypeScript's
// default behavior. JSDoc @returns bindings are hoisted to Module scope
// and propagated to importers via the standard mechanism.
propagatesReturnTypesAcrossImports: true,
// JSDoc @returns bindings are synthesized on the function/method node
// and hoisted to Module scope by `jsBindingScopeFor` (identical to the
// TypeScript `tsBindingScopeFor` `@type-binding.return` branch).
hoistTypeBindingsToModule: true,
// CJS-heavy codebases often have utility functions exported without
// explicit imports at the call site. Workspace-wide unique-name fallback
// recovers these edges.
allowGlobalFreeCallFallback: true,
};
export { javascriptScopeResolver };
@@ -0,0 +1,48 @@
/**
* Simple hooks for the JavaScript scope-resolution provider.
*
* `jsBindingScopeFor` wraps `tsBindingScopeFor` and adds the JS-only
* `@type-binding.class-field` hoisting rule. The other two hooks
* (`jsImportOwningScope`, `jsReceiverBinding`) are identical to their
* TypeScript counterparts and are re-exported directly.
*
* ## Why class-field hoisting lives here (not in `tsBindingScopeFor`)
*
* `@type-binding.class-field` is emitted exclusively by
* `synthesizeConstructorFieldBindings` in `captures.ts`, which is a
* JavaScript-only synthesis pass. TypeScript uses
* `@type-binding.parameter-property` for constructor parameter
* properties instead. Keeping the JS-only rule in the JS hook file
* prevents language-specific logic from leaking into shared TypeScript
* infrastructure (DoD.md §2.2).
*/
import type { CaptureMatch, Scope, ScopeId, ScopeTree } from 'gitnexus-shared';
import { tsBindingScopeFor, walkToScope } from '../typescript/simple-hooks.js';
export {
tsImportOwningScope as jsImportOwningScope,
tsReceiverBinding as jsReceiverBinding,
} from '../typescript/simple-hooks.js';
/**
* Like `tsBindingScopeFor` but additionally hoists
* `@type-binding.class-field` captures to the enclosing Class scope.
*
* `@type-binding.class-field` is anchored inside the constructor body
* (by `synthesizeConstructorFieldBindings`) so that `walkToScope` can
* walk up from the Function (constructor) scope to the Class scope.
* This puts `User.address → Address` in the class's typeBindings so
* compound-receiver resolution finds it when resolving
* `user.address.save()`.
*/
export function jsBindingScopeFor(
decl: CaptureMatch,
innermost: Scope,
tree: ScopeTree,
): ScopeId | null {
if (decl['@type-binding.class-field'] !== undefined) {
return walkToScope(innermost, tree, 'Class');
}
return tsBindingScopeFor(decl, innermost, tree);
}
@@ -56,6 +56,16 @@ import {
typescriptArityCompatibility,
resolveTsImportTarget,
} from './typescript/index.js';
import {
emitJsScopeCaptures,
interpretJsImport,
interpretJsTypeBinding,
jsBindingScopeFor,
jsImportOwningScope,
jsReceiverBinding,
jsMergeBindings,
jsArityCompatibility,
} from './javascript/index.js';
/**
* TypeScript/JavaScript: arrow_function and function_expression are
@@ -359,4 +369,19 @@ export const javascriptProvider = defineLanguage({
classExtractor: createClassExtractor(javascriptClassConfig),
heritageExtractor: createHeritageExtractor(SupportedLanguages.JavaScript),
builtInNames: BUILT_INS,
// ── RFC #909 Ring 3: scope-based resolution hooks (RFC §5) ──────────
// JavaScript is the fourth migration after Python, C#, and TypeScript.
// Hooks are thin wrappers over the TypeScript implementations where
// semantics are identical; JS-specific additions (CJS require(),
// JSDoc type bindings) live in ./javascript/captures.ts.
// See ./javascript/index.ts for the full per-module rationale.
emitScopeCaptures: emitJsScopeCaptures,
interpretImport: interpretJsImport,
interpretTypeBinding: interpretJsTypeBinding,
bindingScopeFor: jsBindingScopeFor,
importOwningScope: jsImportOwningScope,
mergeBindings: (_scope, bindings) => jsMergeBindings(bindings),
receiverBinding: jsReceiverBinding,
arityCompatibility: jsArityCompatibility,
});
@@ -75,8 +75,11 @@ export function tsBindingScopeFor(
* any of `kinds`. Returns the matching scope's id or `null` when no
* ancestor matches (e.g., a return type binding emitted outside any
* Module scope — shouldn't happen in well-formed input).
*
* Exported so language-specific hook wrappers (e.g. `jsBindingScopeFor`)
* can reuse it without duplicating the traversal logic.
*/
function walkToScope(
export function walkToScope(
from: Scope,
tree: ScopeTree,
...kinds: readonly Scope['kind'][]
@@ -2,18 +2,35 @@
* Field Registry
*
* Owner-scoped field/property index extracted from SymbolTable.
* Stores Property symbols keyed by `ownerNodeId\0fieldName` for O(1) lookup.
* Stores Property / Variable / Const / Static symbols keyed by
* `ownerNodeId\0fieldName` for O(1) lookup. Supports multiple defs
* under the same (owner, name) — e.g. legacy Property plus a
* scope-resolution Variable reconciliation entry.
*/
import type { SymbolDefinition } from 'gitnexus-shared';
const EMPTY: readonly SymbolDefinition[] = Object.freeze([]);
// ---------------------------------------------------------------------------
// Public read-only interface
// ---------------------------------------------------------------------------
export interface FieldRegistry {
/** Look up a field/property by its owning class nodeId and field name. */
/**
* First field registered under `(ownerNodeId, fieldName)`, if any.
* Registration order is first-wins: when a Property and a Variable share
* an `(owner, simpleName)` key, the earlier `register(...)` call's def is
* returned. Prefer `lookupAllByOwner` when overloads or duplicate-kind
* entries under the same name must all be visible.
*/
lookupFieldByOwner(ownerNodeId: string, fieldName: string): SymbolDefinition | undefined;
/**
* Every field registered under `(ownerNodeId, fieldName)` in registration
* order. Returns `[]` on miss.
*/
lookupAllByOwner(ownerNodeId: string, fieldName: string): readonly SymbolDefinition[];
}
// ---------------------------------------------------------------------------
@@ -21,7 +38,7 @@ export interface FieldRegistry {
// ---------------------------------------------------------------------------
export interface MutableFieldRegistry extends FieldRegistry {
/** Register a field/property under its owner. */
/** Register a field under its owner. Appends when the key already exists. */
register(ownerNodeId: string, fieldName: string, def: SymbolDefinition): void;
/** Clear all entries. */
clear(): void;
@@ -32,22 +49,36 @@ export interface MutableFieldRegistry extends FieldRegistry {
// ---------------------------------------------------------------------------
export const createFieldRegistry = (): MutableFieldRegistry => {
const fieldByOwner = new Map<string, SymbolDefinition>();
const fieldByOwner = new Map<string, SymbolDefinition[]>();
const lookupAllByOwner = (
ownerNodeId: string,
fieldName: string,
): readonly SymbolDefinition[] => {
return fieldByOwner.get(`${ownerNodeId}\0${fieldName}`) ?? EMPTY;
};
const lookupFieldByOwner = (
ownerNodeId: string,
fieldName: string,
): SymbolDefinition | undefined => {
return fieldByOwner.get(`${ownerNodeId}\0${fieldName}`);
const pool = lookupAllByOwner(ownerNodeId, fieldName);
return pool.length === 0 ? undefined : pool[0];
};
const register = (ownerNodeId: string, fieldName: string, def: SymbolDefinition): void => {
fieldByOwner.set(`${ownerNodeId}\0${fieldName}`, def);
const key = `${ownerNodeId}\0${fieldName}`;
const existing = fieldByOwner.get(key);
if (existing) {
existing.push(def);
} else {
fieldByOwner.set(key, [def]);
}
};
const clear = (): void => {
fieldByOwner.clear();
};
return { lookupFieldByOwner, register, clear };
return { lookupFieldByOwner, lookupAllByOwner, register, clear };
};
@@ -0,0 +1,45 @@
/**
* Owner-keyed member lookup for Step 2 (RFC #909 / PR #1656).
*
* Merges MethodRegistry + FieldRegistry hits for `(ownerDefId, memberName)`
* in O(1) map time per registry — no `defs.byId` scan. Callers that omit
* this helper and leave `ownedMembersByOwner` unset fall back to an O(|defs|)
* compatibility scan inside `lookupCore.collectOwnedMembers`.
*/
import type { DefId, SymbolDefinition } from 'gitnexus-shared';
import type { SemanticModel } from './semantic-model.js';
const EMPTY: readonly SymbolDefinition[] = Object.freeze([]);
/**
* Production hook for `RegistryContext.ownedMembersByOwner`.
* Returns `[]` on miss (authoritative indexed empty) — never `undefined`.
*
* Merges hits from all three owner-keyed registries (methods, fields,
* nested types) under the same `(ownerDefId, memberName)` key. The
* caller's `acceptedKinds` filter in `lookupCore` picks the right subset.
*/
export function lookupOwnedMembersByOwner(
model: Pick<SemanticModel, 'methods' | 'fields' | 'types'>,
ownerDefId: DefId,
memberName: string,
): readonly SymbolDefinition[] {
const methods = model.methods.lookupAllByOwner(ownerDefId, memberName);
const fields = model.fields.lookupAllByOwner(ownerDefId, memberName);
const nestedTypes = model.types.lookupAllByOwner(ownerDefId, memberName);
const methodCount = methods.length;
const fieldCount = fields.length;
const typeCount = nestedTypes.length;
const total = methodCount + fieldCount + typeCount;
if (total === 0) return EMPTY;
if (methodCount === total) return methods;
if (fieldCount === total) return fields;
if (typeCount === total) return nestedTypes;
const merged = new Array<SymbolDefinition>(total);
let i = 0;
for (let j = 0; j < methodCount; j++) merged[i++] = methods[j]!;
for (let j = 0; j < fieldCount; j++) merged[i++] = fields[j]!;
for (let j = 0; j < typeCount; j++) merged[i++] = nestedTypes[j]!;
return merged;
}
@@ -8,6 +8,8 @@
import type { SymbolDefinition } from 'gitnexus-shared';
const EMPTY: readonly SymbolDefinition[] = Object.freeze([]);
// ---------------------------------------------------------------------------
// Public read-only interface
// ---------------------------------------------------------------------------
@@ -35,6 +37,14 @@ export interface TypeRegistry {
* Returned array is a view into the live index — do not mutate.
*/
lookupImplByName(name: string): readonly SymbolDefinition[];
/**
* Look up nested-type defs registered under `(ownerNodeId, simpleName)`
* in registration order. Returns `[]` on miss. Used by Step 2 Receiver/MRO
* resolution when the receiver's owner declares nested classes/structs/
* enums/typedefs/etc. that the caller's `acceptedKinds` includes.
*/
lookupAllByOwner(ownerNodeId: string, simpleName: string): readonly SymbolDefinition[];
}
// ---------------------------------------------------------------------------
@@ -46,6 +56,8 @@ export interface MutableTypeRegistry extends TypeRegistry {
registerClass(name: string, qualifiedName: string, def: SymbolDefinition): void;
/** Register a Rust Impl block by name. */
registerImpl(name: string, def: SymbolDefinition): void;
/** Register a nested type under its owner. Appends when the key already exists. */
registerByOwner(ownerNodeId: string, simpleName: string, def: SymbolDefinition): void;
/** Clear all entries. */
clear(): void;
}
@@ -58,6 +70,7 @@ export const createTypeRegistry = (): MutableTypeRegistry => {
const classByName = new Map<string, SymbolDefinition[]>();
const classByQualifiedName = new Map<string, SymbolDefinition[]>();
const implByName = new Map<string, SymbolDefinition[]>();
const nestedByOwner = new Map<string, SymbolDefinition[]>();
const lookupClassByName = (name: string): SymbolDefinition[] => {
return classByName.get(name) ?? [];
@@ -71,6 +84,13 @@ export const createTypeRegistry = (): MutableTypeRegistry => {
return implByName.get(name) ?? [];
};
const lookupAllByOwner = (
ownerNodeId: string,
simpleName: string,
): readonly SymbolDefinition[] => {
return nestedByOwner.get(`${ownerNodeId}\0${simpleName}`) ?? EMPTY;
};
const registerClass = (name: string, qualifiedName: string, def: SymbolDefinition): void => {
const existing = classByName.get(name);
if (existing) {
@@ -96,18 +116,35 @@ export const createTypeRegistry = (): MutableTypeRegistry => {
}
};
const registerByOwner = (
ownerNodeId: string,
simpleName: string,
def: SymbolDefinition,
): void => {
const key = `${ownerNodeId}\0${simpleName}`;
const existing = nestedByOwner.get(key);
if (existing) {
existing.push(def);
} else {
nestedByOwner.set(key, [def]);
}
};
const clear = (): void => {
classByName.clear();
classByQualifiedName.clear();
implByName.clear();
nestedByOwner.clear();
};
return {
lookupClassByName,
lookupClassByQualifiedName,
lookupImplByName,
lookupAllByOwner,
registerClass,
registerImpl,
registerByOwner,
clear,
};
};
@@ -74,6 +74,7 @@ export const MIGRATED_LANGUAGES: ReadonlySet<SupportedLanguages> = new Set<Suppo
SupportedLanguages.C,
SupportedLanguages.CPlusPlus,
SupportedLanguages.PHP,
SupportedLanguages.JavaScript,
]);
/**
@@ -64,6 +64,8 @@ export interface ResolveReferencesInput {
readonly scopes: ScopeResolutionIndexes;
/** Provider hooks consumed by the registries (e.g. `arityCompatibility`). */
readonly providers?: RegistryProviders;
/** Required owner-keyed member lookup used by Step 2 receiver/MRO walks. */
readonly ownedMembersByOwner: RegistryContext['ownedMembersByOwner'];
}
export interface ResolveStats {
@@ -92,6 +94,7 @@ export function resolveReferenceSites(input: ResolveReferencesInput): ResolveRef
defs: scopes.defs,
qualifiedNames: scopes.qualifiedNames,
moduleScopes: scopes.moduleScopes,
ownedMembersByOwner: input.ownedMembersByOwner,
methodDispatch: scopes.methodDispatch,
providers,
};
@@ -191,7 +194,10 @@ function lookupForSite(
case 'write': {
// Try field first; fall through to method then class so bare-name
// reads of a function (e.g. `cb = save`) still resolve.
const fieldHits = fieldRegistry.lookup(site.name, site.inScope);
const fieldOpts: Parameters<FieldRegistry['lookup']>[2] = {
...(site.explicitReceiver !== undefined ? { explicitReceiver: site.explicitReceiver } : {}),
};
const fieldHits = fieldRegistry.lookup(site.name, site.inScope, fieldOpts);
if (fieldHits.length > 0) return fieldHits;
const methodHits = methodRegistry.lookup(site.name, site.inScope);
if (methodHits.length > 0) return methodHits;
@@ -913,6 +913,9 @@ function pass5CollectReferences(
const explicitReceiver = extractExplicitReceiver(match);
const arity = extractArity(match);
const argumentTypes = extractArgumentTypes(match);
const argumentTypeClasses = parseJsonParameterTypeClassesCapture(
match['@reference.parameter-type-classes'],
);
const site: ReferenceSite = {
name: nameCap.text,
@@ -923,6 +926,7 @@ function pass5CollectReferences(
...(explicitReceiver !== undefined ? { explicitReceiver } : {}),
...(arity !== undefined ? { arity } : {}),
...(argumentTypes !== undefined ? { argumentTypes } : {}),
...(argumentTypeClasses !== undefined ? { argumentTypeClasses } : {}),
};
referenceSites.push(site);
}
@@ -1040,9 +1044,11 @@ const KNOWN_SUB_TAGS: ReadonlySet<string> = new Set<string>([
'@reference.receiver',
'@reference.arity',
'@reference.parameter-types',
'@reference.parameter-type-classes',
'@declaration.parameter-count',
'@declaration.required-parameter-count',
'@declaration.parameter-types',
'@declaration.parameter-type-classes',
'@declaration.template-constraints',
]);
@@ -17,7 +17,13 @@
* generalization plan.
*/
import type { ParsedFile, Reference, ScopeId, SymbolDefinition } from 'gitnexus-shared';
import type {
ParameterTypeClass,
ParsedFile,
Reference,
ScopeId,
SymbolDefinition,
} from 'gitnexus-shared';
import type { KnowledgeGraph } from '../../../graph/types.js';
import type { ScopeResolutionIndexes } from '../../model/scope-resolution-indexes.js';
import type { SemanticModel } from '../../model/semantic-model.js';
@@ -78,6 +84,12 @@ export function emitFreeCallFallback(
let emitted = 0;
const seen = new Set<string>();
// Build an O(1) simple-name -> callable defs index over scopes.defs once
// per pass so pickUniqueGlobalCallable doesn't re-scan defs.byId.values()
// per call site. Same name + callable-kind filter that the previous scan
// applied (see pickUniqueGlobalCallable JSDoc). Cost: O(|defs|) once.
const globalCallablesBySimpleName = buildGlobalCallableIndex(scopes);
for (const parsed of parsedFiles) {
for (const site of parsed.referenceSites) {
if (site.kind !== 'call') continue;
@@ -126,6 +138,7 @@ export function emitFreeCallFallback(
site.arity,
site.argumentTypes,
{
argumentTypeClasses: site.argumentTypeClasses,
conversionRankFn: options.conversionRankFn,
constraintCompatibility: options.constraintCompatibility,
},
@@ -190,6 +203,7 @@ export function emitFreeCallFallback(
fnDef = ordinary[0];
} else {
const narrowed = narrowOverloadCandidates(ordinary, site.arity, site.argumentTypes, {
argumentTypeClasses: site.argumentTypeClasses,
conversionRankFn: options.conversionRankFn,
constraintCompatibility: options.constraintCompatibility,
});
@@ -225,6 +239,7 @@ export function emitFreeCallFallback(
push(adl);
const narrowed = narrowOverloadCandidates(merged, site.arity, site.argumentTypes, {
argumentTypeClasses: site.argumentTypeClasses,
conversionRankFn: options.conversionRankFn,
constraintCompatibility: options.constraintCompatibility,
});
@@ -254,7 +269,7 @@ export function emitFreeCallFallback(
fnDef = pickUniqueGlobalCallable(
site.name,
model,
scopes,
globalCallablesBySimpleName,
parsed.filePath,
options.isFileLocalDef,
site.arity,
@@ -268,6 +283,7 @@ export function emitFreeCallFallback(
})
: undefined,
site.argumentTypes,
site.argumentTypeClasses,
options.conversionRankFn,
);
}
@@ -299,23 +315,46 @@ export function emitFreeCallFallback(
return emitted;
}
/**
* Build a `simpleName -> callable defs` index from `scopes.defs` once per
* pass. Mirrors the filter the old per-site scan applied: Function /
* Method / Constructor, keyed by the last `.`-segment of `qualifiedName`
* (falling back to the qualifiedName itself when undotted). Used by
* `pickUniqueGlobalCallable` so every free-call fallback site is O(1)
* instead of O(|defs|).
*/
function buildGlobalCallableIndex(
scopes: ScopeResolutionIndexes,
): ReadonlyMap<string, readonly SymbolDefinition[]> {
const out = new Map<string, SymbolDefinition[]>();
for (const def of scopes.defs.byId.values()) {
if (def.type !== 'Function' && def.type !== 'Method' && def.type !== 'Constructor') continue;
const qualified = def.qualifiedName;
if (qualified === undefined || qualified.length === 0) continue;
const dot = qualified.lastIndexOf('.');
const simple = dot === -1 ? qualified : qualified.slice(dot + 1);
const bucket = out.get(simple);
if (bucket) bucket.push(def);
else out.set(simple, [def]);
}
return out;
}
function pickUniqueGlobalCallable(
name: string,
model: SemanticModel,
scopes: ScopeResolutionIndexes,
globalCallablesBySimpleName: ReadonlyMap<string, readonly SymbolDefinition[]>,
callerFilePath: string,
isFileLocalDef?: (def: SymbolDefinition) => boolean,
callArity?: number,
isCallerVisible?: (candidate: SymbolDefinition) => boolean,
callArgTypes?: readonly string[],
callArgTypeClasses?: readonly ParameterTypeClass[],
conversionRankFn?: ConversionRankFn,
): SymbolDefinition | undefined {
const scopeDefs: SymbolDefinition[] = [];
const scopeSeen = new Set<string>();
for (const def of scopes.defs.byId.values()) {
const simple = def.qualifiedName?.split('.').pop() ?? def.qualifiedName;
if (simple !== name) continue;
if (def.type !== 'Function' && def.type !== 'Method' && def.type !== 'Constructor') continue;
for (const def of globalCallablesBySimpleName.get(name) ?? []) {
// Skip file-local defs (e.g. C `static` functions) that live in a
// different file from the caller — they are logically invisible.
if (isFileLocalDef !== undefined && def.filePath !== callerFilePath && isFileLocalDef(def)) {
@@ -349,6 +388,7 @@ function pickUniqueGlobalCallable(
// disambiguate (e.g., `f(int)` vs `f(double)` called with `f(2.5)`).
if (scopeDefs.length > 1) {
const narrowed = narrowOverloadCandidates(scopeDefs, callArity, callArgTypes, {
argumentTypeClasses: callArgTypeClasses,
conversionRankFn,
});
if (narrowed.length === 1) return narrowed[0];
@@ -389,6 +429,7 @@ function pickUniqueGlobalCallable(
// Same argument-type + conversion-rank narrowing for the model pool.
if (defs.length > 1) {
const narrowed = narrowOverloadCandidates(defs, callArity, callArgTypes, {
argumentTypeClasses: callArgTypeClasses,
conversionRankFn,
});
if (narrowed.length === 1) return narrowed[0];
@@ -462,6 +503,7 @@ export function pickImplicitThisOverload(
readonly name: string;
readonly arity?: number;
readonly argumentTypes?: readonly string[];
readonly argumentTypeClasses?: readonly import('gitnexus-shared').ParameterTypeClass[];
},
scopes: ScopeResolutionIndexes,
workspaceIndex: WorkspaceResolutionIndex,
@@ -498,6 +540,7 @@ export function pickImplicitThisOverload(
// disambiguating signal) leaves the call unresolved rather than
// routing to an arbitrary first overload by registration order.
const candidates = narrowOverloadCandidates(overloads, site.arity, site.argumentTypes, {
argumentTypeClasses: site.argumentTypeClasses,
conversionRankFn: hookCtx?.conversionRankFn,
constraintCompatibility: hookCtx?.constraintCompatibility,
});
@@ -38,7 +38,13 @@
* 5. Empty input returns empty output.
*/
import type { ArityVerdict, Callsite, ConstraintContext, SymbolDefinition } from 'gitnexus-shared';
import type {
ArityVerdict,
Callsite,
ConstraintContext,
ParameterTypeClass,
SymbolDefinition,
} from 'gitnexus-shared';
/**
* Per-slot conversion-rank function. Returns a numeric cost for
@@ -51,7 +57,12 @@ import type { ArityVerdict, Callsite, ConstraintContext, SymbolDefinition } from
* Each language provides its own implementation. The function operates
* on normalized type strings (output of the language's type normalizer).
*/
export type ConversionRankFn = (argType: string, paramType: string) => number;
export type ConversionRankFn = (
argType: string,
paramType: string,
argTypeClass?: ParameterTypeClass,
paramTypeClass?: ParameterTypeClass,
) => number;
/**
* Optional hook bundle for narrowing extension points. Threaded in
@@ -62,6 +73,8 @@ export type ConversionRankFn = (argType: string, paramType: string) => number;
* undefined preserves the legacy arity + exact-type behavior.
*/
export interface OverloadNarrowingHookCtx {
/** Shape-preserving per-argument sidecar aligned with `argTypes`. */
readonly argumentTypeClasses?: ConstraintContext['argumentTypeClasses'];
/** Conversion-rank scoring fallback (step 4b). Engages when the
* exact-type filter rejects every candidate. */
readonly conversionRankFn?: ConversionRankFn;
@@ -128,7 +141,16 @@ export function narrowOverloadCandidates(
if (params === undefined) return false;
for (let i = 0; i < argTypes.length && i < params.length; i++) {
if (argTypes[i] === '') continue;
if (argTypes[i] !== params[i]) return false;
if (
!exactTypeSlotMatches(
argTypes[i],
params[i],
hookCtx?.argumentTypeClasses?.[i],
d.parameterTypeClasses?.[i],
)
) {
return false;
}
}
return true;
});
@@ -142,7 +164,12 @@ export function narrowOverloadCandidates(
// are returned; multiple survivors are genuinely ambiguous. When
// ranking also yields empty, fall through to the arity-filtered
// `candidates` set — matches pre-#1606 behavior.
const ranked = rankByConversion(candidates, argTypes, hookCtx.conversionRankFn);
const ranked = rankByConversion(
candidates,
argTypes,
hookCtx.conversionRankFn,
hookCtx.argumentTypeClasses,
);
if (ranked.length > 0) result = ranked;
}
}
@@ -163,7 +190,15 @@ export function narrowOverloadCandidates(
// than emitting a wrong edge.
if (hookCtx?.constraintCompatibility !== undefined && argCount !== undefined) {
const callsite: Callsite = { arity: argCount };
const ctx: ConstraintContext = argTypes !== undefined ? { argumentTypes: argTypes } : {};
const ctx: ConstraintContext =
argTypes !== undefined
? {
argumentTypes: argTypes,
...(hookCtx.argumentTypeClasses !== undefined
? { argumentTypeClasses: hookCtx.argumentTypeClasses }
: {}),
}
: {};
result = result.filter((def) => {
if (def.templateConstraints === undefined) return true;
return hookCtx.constraintCompatibility!(callsite, def, ctx) !== 'incompatible';
@@ -173,6 +208,27 @@ export function narrowOverloadCandidates(
return result;
}
function exactTypeSlotMatches(
argType: string,
paramType: string,
argTypeClass?: ParameterTypeClass,
paramTypeClass?: ParameterTypeClass,
): boolean {
if (argType !== paramType) return false;
// C++ normalizes away pointer markers (`int*` -> `int`). When both sides
// provide shape sidecars, do not let that collapse make `int` exactly match
// `int*`. Unknown sidecar evidence preserves the previous string-only path.
if (argTypeClass === undefined || paramTypeClass === undefined) return true;
if (argTypeClass.indirection === 'unknown' || paramTypeClass.indirection === 'unknown') {
return true;
}
return isPointerShape(argTypeClass) === isPointerShape(paramTypeClass);
}
function isPointerShape(typeClass: ParameterTypeClass): boolean {
return typeClass.indirection === 'pointer' && typeClass.pointerDepth > 0;
}
/**
* Pairwise dominance comparison (ISO C++ [over.ics.rank]).
*
@@ -189,6 +245,7 @@ function rankByConversion(
candidates: readonly SymbolDefinition[],
argTypes: readonly string[],
rankFn: ConversionRankFn,
argTypeClasses?: readonly ParameterTypeClass[],
): readonly SymbolDefinition[] {
// Step 1: compute per-slot ranks and exclude non-viable candidates.
const viable: Array<{ def: SymbolDefinition; ranks: number[] }> = [];
@@ -197,12 +254,22 @@ function rankByConversion(
if (params === undefined) continue;
const ranks: number[] = [];
let ok = true;
for (let i = 0; i < argTypes.length && i < params.length; i++) {
for (let i = 0; i < argTypes.length; i++) {
const paramType = parameterTypeAt(params, i);
if (paramType === undefined) {
ok = false;
break;
}
if (argTypes[i] === '') {
ranks.push(0); // unknown arg → any-match (rank 0)
continue;
}
const r = rankFn(argTypes[i], params[i]);
const r = rankFn(
argTypes[i],
paramType,
argTypeClasses?.[i],
parameterTypeClassAt(d.parameterTypeClasses, i),
);
if (!isFinite(r)) {
ok = false;
break;
@@ -229,6 +296,20 @@ function rankByConversion(
return viable.filter((_, idx) => !dominated.has(idx)).map((v) => v.def);
}
function parameterTypeAt(params: readonly string[], argIndex: number): string | undefined {
if (argIndex < params.length) return params[argIndex];
return params[params.length - 1] === '...' ? '...' : undefined;
}
function parameterTypeClassAt(
params: readonly ParameterTypeClass[] | undefined,
argIndex: number,
): ParameterTypeClass | undefined {
if (params === undefined) return undefined;
if (argIndex < params.length) return params[argIndex];
return params[params.length - 1]?.base === '...' ? params[params.length - 1] : undefined;
}
/**
* Compare two per-slot rank vectors.
* Returns -1 if `a` dominates `b` (not worse everywhere, better somewhere),
@@ -346,6 +346,7 @@ export function emitReceiverBoundCalls(
site.arity,
site.argumentTypes,
{
argumentTypeClasses: site.argumentTypeClasses,
conversionRankFn: provider.conversionRankFn,
constraintCompatibility: provider.constraintCompatibility,
},
@@ -732,6 +733,7 @@ function pickOverload(
if (overloads.length === 1) return overloads[0];
const candidates = narrowOverloadCandidates(overloads, site.arity, site.argumentTypes, {
argumentTypeClasses: site.argumentTypeClasses,
conversionRankFn: provider.conversionRankFn,
constraintCompatibility: provider.constraintCompatibility,
});
@@ -18,8 +18,8 @@
* undefined `ownerId` is reachable via either:
* - `model.methods.lookupAllByOwner(ownerId, simpleName)` — if the
* def is a Method / Function / Constructor, OR
* - `model.fields.lookupFieldByOwner(ownerId, simpleName)` — if the
* def is a Property / Variable.
* - `model.fields.lookupAllByOwner(ownerId, simpleName)` — if the
* def is a Property / Variable / Const / Static.
*
* This invariant is the foundation of Contract Invariant I9
* (`contract/scope-resolver.ts`): scope-resolution passes MUST read
@@ -45,11 +45,29 @@ import type { ParsedFile } from 'gitnexus-shared';
import type { MutableSemanticModel, SemanticModel } from '../../model/semantic-model.js';
import { simpleQualifiedName } from '../graph-bridge/ids.js';
const NESTED_TYPE_KINDS = new Set<string>([
'Class',
'Interface',
'Enum',
'Struct',
'Union',
'Trait',
'TypeAlias',
'Typedef',
'Record',
'Delegate',
'Annotation',
'Template',
'Namespace',
]);
export interface ReconcileStats {
/** Method/Function/Constructor defs registered into MethodRegistry. */
readonly methodsRegistered: number;
/** Property/Variable defs registered into FieldRegistry. */
readonly fieldsRegistered: number;
/** Class-like nested type defs registered into TypeRegistry by owner. */
readonly nestedTypesRegistered: number;
/** Defs already present (idempotent skip). */
readonly skippedAlreadyPresent: number;
}
@@ -60,6 +78,7 @@ export function reconcileOwnership(
): ReconcileStats {
let methodsRegistered = 0;
let fieldsRegistered = 0;
let nestedTypesRegistered = 0;
let skippedAlreadyPresent = 0;
for (const parsed of parsedFiles) {
@@ -77,19 +96,32 @@ export function reconcileOwnership(
}
model.methods.register(ownerId, simple, def);
methodsRegistered++;
} else if (def.type === 'Property' || def.type === 'Variable') {
const existing = model.fields.lookupFieldByOwner(ownerId, simple);
if (existing !== undefined && existing.nodeId === def.nodeId) {
} else if (
def.type === 'Property' ||
def.type === 'Variable' ||
def.type === 'Const' ||
def.type === 'Static'
) {
const existing = model.fields.lookupAllByOwner(ownerId, simple);
if (existing.some((e) => e.nodeId === def.nodeId)) {
skippedAlreadyPresent++;
continue;
}
model.fields.register(ownerId, simple, def);
fieldsRegistered++;
} else if (NESTED_TYPE_KINDS.has(def.type)) {
const existing = model.types.lookupAllByOwner(ownerId, simple);
if (existing.some((e) => e.nodeId === def.nodeId)) {
skippedAlreadyPresent++;
continue;
}
model.types.registerByOwner(ownerId, simple, def);
nestedTypesRegistered++;
}
}
}
return { methodsRegistered, fieldsRegistered, skippedAlreadyPresent };
return { methodsRegistered, fieldsRegistered, nestedTypesRegistered, skippedAlreadyPresent };
}
/**
@@ -131,15 +163,29 @@ export function validateOwnershipParity(
);
mismatches++;
}
} else if (def.type === 'Property' || def.type === 'Variable') {
const found = model.fields.lookupFieldByOwner(ownerId, simple);
if (found === undefined || found.nodeId !== def.nodeId) {
} else if (
def.type === 'Property' ||
def.type === 'Variable' ||
def.type === 'Const' ||
def.type === 'Static'
) {
const found = model.fields.lookupAllByOwner(ownerId, simple);
if (!found.some((d) => d.nodeId === def.nodeId)) {
onWarn(
`semantic-model parity: ${def.type} ${def.nodeId} (${parsed.filePath}) ` +
`owned by ${ownerId} as "${simple}" not in FieldRegistry`,
);
mismatches++;
}
} else if (NESTED_TYPE_KINDS.has(def.type)) {
const found = model.types.lookupAllByOwner(ownerId, simple);
if (!found.some((d) => d.nodeId === def.nodeId)) {
onWarn(
`semantic-model parity: ${def.type} ${def.nodeId} (${parsed.filePath}) ` +
`owned by ${ownerId} as "${simple}" not in TypeRegistry owner index`,
);
mismatches++;
}
}
}
}
@@ -19,6 +19,7 @@ import { javaScopeResolver } from '../../languages/java/scope-resolver.js';
import { cScopeResolver } from '../../languages/c/scope-resolver.js';
import { cppScopeResolver } from '../../languages/cpp/scope-resolver.js';
import { phpScopeResolver } from '../../languages/php/scope-resolver.js';
import { javascriptScopeResolver } from '../../languages/javascript/scope-resolver.js';
/** Map of `SupportedLanguages` → `ScopeResolver`. The phase iterates
* this map intersected with `MIGRATED_LANGUAGES` (the per-language
@@ -36,4 +37,5 @@ export const SCOPE_RESOLVERS: ReadonlyMap<SupportedLanguages, ScopeResolver> = n
[SupportedLanguages.C, cScopeResolver],
[SupportedLanguages.CPlusPlus, cppScopeResolver],
[SupportedLanguages.PHP, phpScopeResolver],
[SupportedLanguages.JavaScript, javascriptScopeResolver],
]);
@@ -25,6 +25,7 @@
import type { ParsedFile, RegistryProviders } from 'gitnexus-shared';
import type { KnowledgeGraph } from '../../../graph/types.js';
import { lookupOwnedMembersByOwner } from '../../model/owned-members-lookup.js';
import type { MutableSemanticModel, SemanticModel } from '../../model/semantic-model.js';
import { reconcileOwnership, validateOwnershipParity } from './reconcile-ownership.js';
import { validateBindingsImmutability } from './validate-bindings-immutability.js';
@@ -342,6 +343,8 @@ export function runScopeResolution(
const { referenceIndex, stats: resolveStats } = resolveReferenceSites({
scopes: indexes,
providers: registryProviders,
ownedMembersByOwner: (ownerDefId, memberName) =>
lookupOwnedMembersByOwner(readonlyModel, ownerDefId, memberName),
});
const tResolve = PROF ? process.hrtime.bigint() : 0n;
+24 -21
View File
@@ -154,6 +154,7 @@ export const splitRelCsvByLabelPair = async (
let db: lbug.Database | null = null;
let conn: lbug.Connection | null = null;
let currentDbPath: string | null = null;
let currentDbReadOnly = false;
let ftsLoaded = false;
let vectorExtensionLoaded = false;
@@ -448,12 +449,17 @@ export const initLbug = async (dbPath: string) => {
* database is busy (e.g. `gitnexus analyze` holds the write lock).
* Each retry waits DB_LOCK_RETRY_DELAY_MS * attempt milliseconds.
*/
export const withLbugDb = async <T>(dbPath: string, operation: () => Promise<T>): Promise<T> => {
export const withLbugDb = async <T>(
dbPath: string,
operation: () => Promise<T>,
options: { readOnly?: boolean } = {},
): Promise<T> => {
let lastError: unknown;
const readOnly = options.readOnly === true;
for (let attempt = 1; attempt <= DB_LOCK_RETRY_ATTEMPTS; attempt++) {
try {
return await runWithSessionLock(async () => {
await ensureLbugInitialized(dbPath);
await ensureLbugInitialized(dbPath, readOnly);
return operation();
});
} catch (err) {
@@ -483,15 +489,15 @@ export const withLbugDb = async <T>(dbPath: string, operation: () => Promise<T>)
throw lastError;
};
const ensureLbugInitialized = async (dbPath: string) => {
if (conn && currentDbPath === dbPath) {
const ensureLbugInitialized = async (dbPath: string, readOnly: boolean = false) => {
if (conn && currentDbPath === dbPath && currentDbReadOnly === readOnly) {
return { db, conn };
}
await doInitLbug(dbPath);
await doInitLbug(dbPath, readOnly);
return { db, conn };
};
const doInitLbug = async (dbPath: string) => {
const doInitLbug = async (dbPath: string, readOnly: boolean = false) => {
// Different database requested — close the old one first
if (conn || db) {
await safeClose();
@@ -575,9 +581,12 @@ const doInitLbug = async (dbPath: string) => {
const parentDir = path.dirname(dbPath);
await fs.mkdir(parentDir, { recursive: true });
const opened = await openLbugConnection(lbug, dbPath);
const opened = readOnly
? await openLbugConnection(lbug, dbPath, { readOnly: true })
: await openLbugConnection(lbug, dbPath);
db = opened.db;
conn = opened.conn;
currentDbReadOnly = readOnly;
} finally {
await releaseInitLock();
}
@@ -614,7 +623,7 @@ const doInitLbug = async (dbPath: string) => {
` Original error: ${msg.slice(0, 200)}`,
);
}
if (!msg.includes('already exists') && !isDbBusyError(err)) {
if (!msg.includes('already exists') && !isDbBusyError(err) && !isReadOnlyDbError(err)) {
logger.warn(`⚠️ Schema creation warning: ${msg.slice(0, 120)}`);
}
}
@@ -1058,12 +1067,7 @@ export const batchInsertNodesToLbug = async (
};
export const executeQuery = async (cypher: string): Promise<any[]> => {
if (!conn) {
throw new Error('LadybugDB not initialized. Call initLbug first.');
}
const queryResult = await conn.query(cypher);
return await readQueryRows(queryResult);
return await executePrepared(cypher, {});
};
export const streamQuery = async (
@@ -1647,7 +1651,10 @@ export const createFTSIndex = async (
if (ensuredFTSIndexes.has(key)) return;
if (!(await loadFTSExtension())) {
return;
throw new Error(
`FTS extension unavailable - cannot create FTS index ${tableName}.${indexName}. ` +
'Run `gitnexus doctor` and ensure the LadybugDB FTS extension is installed and loadable on this machine.',
);
}
const propList = properties.map((p) => `'${p}'`).join(', ');
@@ -1726,19 +1733,15 @@ export const queryFTS = async (
throw new Error('LadybugDB not initialized. Call initLbug first.');
}
// Escape backslashes and single quotes to prevent Cypher injection
const escapedQuery = query.replace(/\\/g, '\\\\').replace(/'/g, "''");
const cypher = `
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', '${escapedQuery}', conjunctive := ${conjunctive})
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', $query, conjunctive := ${conjunctive})
RETURN node, score
ORDER BY score DESC
LIMIT ${limit}
`;
try {
const queryResult = await conn.query(cypher);
const rows = await readQueryRows(queryResult);
const rows = await executePrepared(cypher, { query });
return rows.map((row: any) => {
const node = row.node || row[0] || {};
+66 -46
View File
@@ -16,14 +16,52 @@
*/
import fs from 'fs/promises';
import os from 'os';
import path from 'path';
import lbug from '@ladybugdb/core';
import { loadFTSExtension } from './lbug-adapter.js';
import { isReadOnlyDbError, loadFTSExtension } from './lbug-adapter.js';
import {
createLbugDatabase,
isWalCorruptionError,
WAL_RECOVERY_SUGGESTION,
} from './lbug-config.js';
/**
* Probe whether a Windows FTS extension binary is locally installed under
* ~/.lbdb/extension/<any-version>/win_amd64/fts/. Returns true on the first
* version dir whose libfts.lbug_extension exists on disk; false if the
* extension root is missing or contains no FTS binary.
*
* Gates the Windows skip-FTS-load guard below so we only skip the load
* when no extension binary is present. When at least one binary exists,
* loadFTSExtension is called with policy: 'load-only' — LadybugDB resolves
* LOAD EXTENSION fts to its version-specific path internally, and the
* ExtensionManager's tryLoad try/catch handles version-mismatch errors
* cleanly without ever attempting dlopen of a stale binary. The install
* path that the #1199/#1217 SIGSEGV documented is never exercised at
* query time.
*
* Exported so unit tests can exercise the probe directly against a
* temp-dir plus spied `os.homedir()` — see lbug-pool-win-fts-probe.test.ts.
*/
export async function hasLocalWinFtsExtension(): Promise<boolean> {
try {
const extRoot = path.join(os.homedir(), '.lbdb', 'extension');
const versions = await fs.readdir(extRoot);
for (const v of versions) {
try {
await fs.stat(path.join(extRoot, v, 'win_amd64', 'fts', 'libfts.lbug_extension'));
return true;
} catch {
/* missing for this version, keep looking */
}
}
} catch {
/* no .lbdb/extension dir */
}
return false;
}
/** Per-repo pool: one Database, many Connections */
interface PoolEntry {
db: lbug.Database;
@@ -423,14 +461,24 @@ async function doInitLbug(repoId: string, dbPath: string): Promise<void> {
// install; analyze owns extension installation. If LOAD fails, search
// features degrade gracefully and the user-facing query path proceeds.
if (!shared.ftsLoaded) {
// Windows guard: LOAD EXTENSION fts crashes with SIGSEGV on Windows when
// the FTS extension binary is not installed locally (@ladybugdb/core native
// bug — the extension loader hits an unhandled error path that signals SIGSEGV
// rather than throwing a JS exception, so try/catch cannot protect here).
// Skip the load on Windows; bm25-index.js catches the resulting Kuzu catalog
// errors and returns empty BM25 results gracefully. Graph queries are unaffected.
// Windows guard: LOAD EXTENSION fts crashes with SIGSEGV on Windows during
// *install* — the @ladybugdb/core out-of-process installer hits an unhandled
// error path that signals SIGSEGV instead of throwing (see #1199, #1217).
// The previous unconditional skip was over-broad: it also disabled FTS on
// hosts where the binary was already on disk and only needed LOAD, leaving
// BM25 silently degraded with no error path (see #1690).
//
// Probe ~/.lbdb/extension/*/win_amd64/fts/ first. If any binary is on disk
// we run loadFTSExtension(..., 'load-only'); the install path is never
// exercised, and LadybugDB's version-specific resolution + ExtensionManager
// try/catch handle stale/zero-byte siblings cleanly (verified empirically
// on Win10 + Node 22.19 + gitnexus 1.6.5 + @ladybugdb/core 0.16.1). With
// no binary at all, we fall back to the upstream skip so install-time
// SIGSEGV continues to be avoided.
if (process.platform === 'win32') {
shared.ftsLoaded = true;
shared.ftsLoaded = (await hasLocalWinFtsExtension())
? await loadFTSExtension(available[0], { policy: 'load-only' })
: true;
} else {
shared.ftsLoaded = await loadFTSExtension(available[0], { policy: 'load-only' });
}
@@ -497,10 +545,12 @@ export async function initLbugWithDb(
// Load FTS extension if not already loaded on this Database.
// policy: 'load-only' — same contract as initLbug above; the read pool
// must not block on a network install during query execution.
// Windows guard: same SIGSEGV risk as doInitLbug above — skip on Windows.
// Windows guard: same probe-then-load policy as doInitLbug above.
if (!shared.ftsLoaded) {
if (process.platform === 'win32') {
shared.ftsLoaded = true;
shared.ftsLoaded = (await hasLocalWinFtsExtension())
? await loadFTSExtension(available[0], { policy: 'load-only' })
: true;
} else {
shared.ftsLoaded = await loadFTSExtension(available[0], { policy: 'load-only' });
}
@@ -598,30 +648,7 @@ function withTimeout<T>(promise: Promise<T>, ms: number, label: string): Promise
}
export const executeQuery = async (repoId: string, cypher: string): Promise<any[]> => {
const entry = pool.get(repoId);
if (!entry) {
throw new Error(`LadybugDB not initialized for repo "${repoId}". Call initLbug first.`);
}
if (isWriteQuery(cypher)) {
throw new Error('Write operations are not allowed. The pool adapter is read-only.');
}
entry.lastUsed = Date.now();
const conn = await checkout(entry);
silenceStdout();
activeQueryCount++;
try {
const queryResult = await withTimeout(conn.query(cypher), QUERY_TIMEOUT_MS, 'Query');
const result = Array.isArray(queryResult) ? queryResult[0] : queryResult;
const rows = await result.getAll();
return rows;
} finally {
activeQueryCount--;
restoreStdout();
checkin(entry, conn);
}
return await executeParameterized(repoId, cypher, {});
};
/**
@@ -653,6 +680,11 @@ export const executeParameterized = async (
const result = Array.isArray(queryResult) ? queryResult[0] : queryResult;
const rows = await result.getAll();
return rows;
} catch (err) {
if (isReadOnlyDbError(err)) {
throw new Error('Write operations are not allowed. The pool adapter is read-only.');
}
throw err;
} finally {
activeQueryCount--;
restoreStdout();
@@ -685,15 +717,3 @@ export const closeLbug = async (repoId?: string): Promise<void> => {
* Check if a specific repo's pool is active
*/
export const isLbugReady = (repoId: string): boolean => pool.has(repoId);
/** Regex to detect write operations in user-supplied Cypher queries.
* Note: CALL is NOT blocked — it's used for read-only FTS (CALL QUERY_FTS_INDEX)
* and vector search (CALL QUERY_VECTOR_INDEX). The database is opened in
* read-only mode as defense-in-depth against write procedures. */
export const CYPHER_WRITE_RE =
/(?<!:)\b(CREATE|DELETE|SET|MERGE|REMOVE|DROP|ALTER|COPY|DETACH|FOREACH|INSTALL|LOAD)\b/i;
/** Check if a Cypher query contains write operations */
export function isWriteQuery(query: string): boolean {
return CYPHER_WRITE_RE.test(query);
}
+24
View File
@@ -0,0 +1,24 @@
/**
* Return true only for plain-object payloads that can be safely used as
* named parameter maps in prepared Cypher execution.
*
* Validation criteria:
* - must be a JavaScript object (`typeof value === 'object'`)
* - must not be `null`
* - must not be an array
* - must have a plain-object prototype
* - values must be scalar bindable values (string | number | boolean | null)
*
* Rationale: prepared-statement params are key/value maps; rejecting null/array
* and non-plain objects keeps binding behavior predictable and avoids passing
* complex host objects to Ladybug parameter binding.
*/
const isBindableScalar = (value: unknown): value is string | number | boolean | null =>
value === null || ['string', 'number', 'boolean'].includes(typeof value);
export const isValidQueryParams = (value: unknown): value is Record<string, unknown> =>
value !== null &&
typeof value === 'object' &&
!Array.isArray(value) &&
(Object.getPrototypeOf(value) === Object.prototype || Object.getPrototypeOf(value) === null) &&
Object.values(value).every(isBindableScalar);
+94 -2
View File
@@ -25,7 +25,7 @@ import {
deleteAllCommunitiesAndProcesses,
queryImporters,
} from './lbug/lbug-adapter.js';
import { createSearchFTSIndexes } from './search/fts-indexes.js';
import { createSearchFTSIndexes, verifySearchFTSIndexes } from './search/fts-indexes.js';
import {
getStoragePaths,
saveMeta,
@@ -71,6 +71,10 @@ export interface AnalyzeOptions {
* bypass. See `allowDuplicateName` below.
*/
force?: boolean;
/** Repair only search indexes without re-running full parsing/indexing. */
repairFts?: boolean;
/** Emit per-index FTS create logs. */
verbose?: boolean;
embeddings?: boolean;
/**
* Override the auto-skip node-count cap for embedding generation.
@@ -126,6 +130,8 @@ export interface AnalyzeResult {
alreadyUpToDate?: boolean;
/** The raw pipeline result — only populated when needed by callers (e.g. skill generation). */
pipelineResult?: any;
/** True when analyze only repaired FTS indexes and skipped pipeline re-analysis. */
ftsRepairedOnly?: boolean;
}
// Re-export the pure flag-derivation helper so external callers (and tests)
@@ -190,6 +196,78 @@ export async function runFullAnalysis(
const currentCommit = repoHasGit ? getCurrentCommit(repoPath) : '';
const existingMeta = await loadMeta(storagePath);
// ── FTS-only repair path ────────────────────────────────────────────
if (options.repairFts) {
if (!existingMeta) {
throw new Error(
'Cannot repair FTS indexes because this repository has not been analyzed yet. ' +
'Run `gitnexus analyze` first to create the initial index, then retry `--repair-fts`.',
);
}
let lbugStat;
try {
lbugStat = await fs.lstat(lbugPath);
} catch {
throw new Error(
`Cannot repair FTS indexes: graph store at ${lbugPath} is missing. ` +
'Run `gitnexus analyze` (full) to rebuild from scratch.',
);
}
if (!lbugStat.isFile()) {
const foundType = lbugStat.isDirectory()
? 'a directory'
: lbugStat.isSymbolicLink()
? 'a symbolic link'
: lbugStat.isSocket()
? 'a socket'
: lbugStat.isBlockDevice()
? 'a block device'
: lbugStat.isCharacterDevice()
? 'a character device'
: lbugStat.isFIFO()
? 'a FIFO'
: 'not a regular file';
throw new Error(
`Cannot repair FTS indexes: graph store at ${lbugPath} is ${foundType} (expected a file). ` +
'Run `gitnexus analyze` (full) to rebuild from scratch.',
);
}
try {
await initLbug(lbugPath);
progress('fts', 85, 'Repairing search indexes...');
await createSearchFTSIndexes({
onIndexStart: options.verbose
? (table, indexName) => log(`FTS: creating ${table}.${indexName}`)
: undefined,
onIndexReady: options.verbose
? (table, indexName) => log(`FTS: ready ${table}.${indexName}`)
: undefined,
});
const missing = await verifySearchFTSIndexes(executeQuery);
if (missing.length > 0) {
throw new Error(
`FTS repair failed - missing indexes after rebuild: ${missing.join(', ')}. ` +
'Run `gitnexus analyze --force` to perform a full graph+FTS rebuild; ' +
'if that also fails, verify FTS extension availability via `gitnexus doctor`.',
);
}
await ensureGitNexusIgnored(repoPath);
progress('fts', 90, 'Search indexes ready');
progress('done', 100, 'Done');
return {
repoName:
options.registryName ??
getInferredRepoName(repoPath) ??
path.basename(resolveRepoIdentityRoot(repoPath)),
repoPath,
stats: existingMeta.stats ?? {},
ftsRepairedOnly: true,
};
} finally {
await closeLbug().catch(() => {});
}
}
// ── Crash recovery: dirty flag forces full rebuild ────────────────
// If the previous incremental run set incrementalInProgress and didn't
// clear it, the on-disk index may be in a half-state. Cheapest path
@@ -583,7 +661,21 @@ export async function runFullAnalysis(
// ── Phase 3: FTS (85–90%) ─────────────────────────────────────────
progress('fts', 85, 'Creating search indexes...');
await createSearchFTSIndexes();
await createSearchFTSIndexes({
onIndexStart: options.verbose
? (table, indexName) => log(`FTS: creating ${table}.${indexName}`)
: undefined,
onIndexReady: options.verbose
? (table, indexName) => log(`FTS: ready ${table}.${indexName}`)
: undefined,
});
const missingIndexNames = await verifySearchFTSIndexes(executeQuery);
if (missingIndexNames.length > 0) {
throw new Error(
`FTS verification failed - missing indexes after analyze: ${missingIndexNames.join(', ')}. ` +
'Check FTS extension availability, then retry `gitnexus analyze --force` for a full rebuild.',
);
}
progress('fts', 90, 'Search indexes ready');
// ── Phase 3.5: Re-insert cached embeddings ────────────────────────
+6 -7
View File
@@ -27,22 +27,20 @@ export interface FTSSearchResponse {
* caller can distinguish "zero matches" from "index missing".
*/
async function queryFTSViaExecutor(
executor: (cypher: string) => Promise<any[]>,
executor: (cypher: string, params: Record<string, any>) => Promise<any[]>,
tableName: string,
indexName: string,
query: string,
limit: number,
): Promise<Array<{ filePath: string; score: number; nodeId: string }> | null> {
// Escape single quotes and backslashes to prevent Cypher injection
const escapedQuery = query.replace(/\\/g, '\\\\').replace(/'/g, "''");
const cypher = `
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', '${escapedQuery}', conjunctive := false)
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', $query, conjunctive := false)
RETURN node, score
ORDER BY score DESC
LIMIT ${limit}
`;
try {
const rows = await executor(cypher);
const rows = await executor(cypher, { query });
return rows.map((row: any) => {
const node = row.node || row[0] || {};
const score = row.score ?? row[1] ?? 0;
@@ -81,8 +79,9 @@ export const searchFTSFromLbug = async (
// IMPORTANT: FTS queries run sequentially to avoid connection contention.
// The MCP pool supports multiple connections, but FTS is best run serially.
const poolMod = await import('../lbug/pool-adapter.js');
const { executeQuery } = poolMod;
const executor = (cypher: string) => executeQuery(repoId, cypher);
const { executeParameterized } = poolMod;
const executor = (cypher: string, params: Record<string, any>) =>
executeParameterized(repoId, cypher, params);
for (const { table, indexName } of FTS_INDEXES) {
const result = await queryFTSViaExecutor(executor, table, indexName, query, limit);
+38 -1
View File
@@ -1,8 +1,45 @@
import { createFTSIndex } from '../lbug/lbug-adapter.js';
import { FTS_INDEXES } from './fts-schema.js';
export async function createSearchFTSIndexes(): Promise<void> {
export interface CreateSearchFTSIndexesOptions {
onIndexStart?: (table: string, indexName: string) => void;
onIndexReady?: (table: string, indexName: string) => void;
}
export async function createSearchFTSIndexes(
options?: CreateSearchFTSIndexesOptions,
): Promise<void> {
for (const { table, indexName, properties } of FTS_INDEXES) {
options?.onIndexStart?.(table, indexName);
await createFTSIndex(table, indexName, [...properties]);
options?.onIndexReady?.(table, indexName);
}
}
export async function verifySearchFTSIndexes(
executeQuery: (cypher: string) => Promise<unknown[]>,
): Promise<string[]> {
const safeIdentifier = (value: string): string => {
if (!/^[A-Za-z_][A-Za-z0-9_]*$/.test(value)) {
throw new Error(`Invalid FTS identifier: ${value}`);
}
return value;
};
const missing: string[] = [];
for (const { table, indexName } of FTS_INDEXES) {
const safeTable = safeIdentifier(table);
const safeIndex = safeIdentifier(indexName);
const probe = `
CALL QUERY_FTS_INDEX('${safeTable}', '${safeIndex}', '__gitnexus_fts_probe__', conjunctive := false)
RETURN score
LIMIT 1
`;
try {
await executeQuery(probe);
} catch {
missing.push(`${table}.${indexName}`);
}
}
return missing;
}
+61 -4
View File
@@ -66,12 +66,15 @@ export interface WikiOptions {
concurrency?: number;
/** If true, stop after building module tree for user review */
reviewOnly?: boolean;
/** Output language for generated documentation (e.g. 'english', 'chinese', 'spanish') */
lang?: string;
}
export interface WikiMeta {
fromCommit: string;
generatedAt: string;
model: string;
lang: string;
moduleFiles: Record<string, string[]>;
moduleTree: ModuleTreeNode[];
}
@@ -177,6 +180,28 @@ export class WikiGenerator {
};
}
/**
* Return the effective lang string: strip control characters, trim, cap at 50 chars,
* then validate against a character allowlist. Returns '' if the value is absent or invalid.
* Used for both prompt construction and meta storage/comparison so they are always in sync.
*/
private effectiveLang(): string {
const lang = (this.options.lang ?? '')
.replace(/[\x00-\x1F\x7F]/g, '')
.trim()
.slice(0, 50);
return /^[a-zA-Z -]+$/.test(lang) ? lang : '';
}
/**
* Append an output-language instruction to a system prompt when --lang is set.
*/
private buildSystemPrompt(base: string): string {
const lang = this.effectiveLang();
if (!lang) return base;
return `${base}\n\nIMPORTANT: Write ALL documentation content in ${lang}. This includes prose, code comments in examples, and diagram labels. Note: page titles (H1 headings) are generated separately and will remain in English.`;
}
/**
* Route LLM call to the appropriate provider (OpenAI-compatible or Cursor CLI).
*/
@@ -207,6 +232,15 @@ export class WikiGenerator {
// Up-to-date check (skip if --force)
if (!forceMode && existingMeta && existingMeta.fromCommit === currentCommit) {
const currentLang = this.effectiveLang();
const metaLang = existingMeta.lang ?? '';
if (currentLang !== metaLang) {
const prevDisplay = metaLang || 'english (default)';
const nextDisplay = currentLang || 'english (default)';
throw new Error(
`Wiki was generated in ${prevDisplay}; use --force to regenerate in ${nextDisplay}.`,
);
}
// Still regenerate the HTML viewer in case it's missing
await this.ensureHTMLViewer();
return { pagesGenerated: 0, mode: 'up-to-date', failedModules: [] };
@@ -235,6 +269,15 @@ export class WikiGenerator {
let result: WikiRunResult;
try {
if (!forceMode && existingMeta && existingMeta.fromCommit) {
const currentLang = this.effectiveLang();
const metaLang = existingMeta.lang ?? '';
if (currentLang !== metaLang) {
const prevDisplay = metaLang || 'english (default)';
const nextDisplay = currentLang || 'english (default)';
throw new Error(
`Wiki was generated in ${prevDisplay}; use --force to regenerate in ${nextDisplay}.`,
);
}
result = await this.incrementalUpdate(existingMeta, currentCommit);
} else {
result = await this.fullGeneration(currentCommit);
@@ -368,6 +411,7 @@ export class WikiGenerator {
fromCommit: currentCommit,
generatedAt: new Date().toISOString(),
model: this.llmConfig.model,
lang: this.effectiveLang(),
moduleFiles,
moduleTree,
});
@@ -415,6 +459,9 @@ export class WikiGenerator {
DIRECTORY_TREE: dirTree,
});
// Grouping is a structured-data phase (JSON output), not documentation.
// Do NOT apply buildSystemPrompt here — a language instruction would risk
// translating module-name keys, breaking slug stability and JSON parsing.
const response = await this.invokeLLM(
prompt,
GROUPING_SYSTEM_PROMPT,
@@ -589,9 +636,13 @@ export class WikiGenerator {
PROCESSES: formatProcesses(processes),
});
const response = await this.invokeLLM(prompt, MODULE_SYSTEM_PROMPT, this.streamOpts(node.name));
const response = await this.invokeLLM(
prompt,
this.buildSystemPrompt(MODULE_SYSTEM_PROMPT),
this.streamOpts(node.name),
);
// Write page with front matter
// H1 uses the English module name (stable slug source); body is LLM-translated.
const pageContent = sanitizeMermaidMarkdown(`# ${node.name}\n\n${response.content}`);
await fs.writeFile(path.join(this.wikiDir, `${node.slug}.md`), pageContent, 'utf-8');
}
@@ -630,7 +681,11 @@ export class WikiGenerator {
CROSS_PROCESSES: formatProcesses(processes),
});
const response = await this.invokeLLM(prompt, PARENT_SYSTEM_PROMPT, this.streamOpts(node.name));
const response = await this.invokeLLM(
prompt,
this.buildSystemPrompt(PARENT_SYSTEM_PROMPT),
this.streamOpts(node.name),
);
const pageContent = sanitizeMermaidMarkdown(`# ${node.name}\n\n${response.content}`);
await fs.writeFile(path.join(this.wikiDir, `${node.slug}.md`), pageContent, 'utf-8');
@@ -678,7 +733,7 @@ export class WikiGenerator {
const response = await this.invokeLLM(
prompt,
OVERVIEW_SYSTEM_PROMPT,
this.buildSystemPrompt(OVERVIEW_SYSTEM_PROMPT),
this.streamOpts('Generating overview', 88),
);
@@ -713,6 +768,7 @@ export class WikiGenerator {
...existingMeta,
fromCommit: currentCommit,
generatedAt: new Date().toISOString(),
lang: this.effectiveLang(),
});
return { pagesGenerated: 0, mode: 'incremental', failedModules: [] };
}
@@ -817,6 +873,7 @@ export class WikiGenerator {
fromCommit: currentCommit,
generatedAt: new Date().toISOString(),
model: this.llmConfig.model,
lang: this.effectiveLang(),
});
this.onProgress('done', 100, 'Incremental update complete');
+149 -14
View File
@@ -14,15 +14,20 @@ import {
executeParameterized,
closeLbug,
isLbugReady,
isWriteQuery,
} from '../../core/lbug/pool-adapter.js';
import { isValidQueryParams } from '../../core/lbug/query-params.js';
import { isWalCorruptionError, WAL_RECOVERY_SUGGESTION } from '../../core/lbug/lbug-config.js';
export { isWriteQuery };
// Embedding imports are lazy (dynamic import) to avoid loading onnxruntime-node
// at MCP server startup — crashes on unsupported Node ABI versions (#89)
// git utilities available if needed
// import { isGitRepo, getCurrentCommit, getGitRoot } from '../../storage/git.js';
import { parseDiffHunks, type FileDiff } from '../../storage/git.js';
import {
parseDiffHunks,
getCanonicalRepoRoot,
getGitRoot,
type FileDiff,
} from '../../storage/git.js';
import { realpathSync } from 'fs';
import {
listRegisteredRepos,
cleanupOldKuzuFiles,
@@ -169,6 +174,9 @@ function logQueryError(context: string, err: unknown): void {
logger.error({ context, err: msg }, 'GitNexus query failed');
}
const isReadOnlyDbError = (err: unknown): boolean =>
/read-only database/i.test(err instanceof Error ? err.message : String(err));
/**
* Per-query latency telemetry for production aggregation (#553).
*
@@ -211,6 +219,82 @@ interface RepoHandle {
stats?: RegistryEntry['stats'];
}
/** Resolve symlinks for path comparison; falls back to path.resolve on error.
* Uses `realpathSync.native` (not the pure-JS `realpathSync`) so that Windows
* 8.3 short names (e.g. RUNNER~1 → runneradmin) are expanded to long form,
* matching the output of `git rev-parse --show-toplevel`. */
function tryRealpath(p: string): string {
try {
return realpathSync.native(p);
} catch {
return path.resolve(p);
}
}
/**
* Resolve the git diff cwd for detect_changes, auto-detecting linked worktrees.
*
* When `launchCwd` is a linked worktree of the same canonical repository as
* `repoPath` (i.e. `getGitRoot(launchCwd)` differs from `repoPath` but both
* share the same `getCanonicalRepoRoot`), returns the worktree's git root so
* that `git diff` sees the correct working directory and index.
*
* Returns `repoPath` unchanged in all other cases (non-worktree, git
* unavailable, unrelated repo).
*
* Extracted as a module-level export so tests can pass any `launchCwd` instead
* of relying on `process.cwd()`, which is fixed to the server launch directory
* and cannot be changed mid-process.
*/
export function resolveWorktreeCwd(repoPath: string, launchCwd: string): string {
try {
// Verify repoPath is a git root before comparing against its canonical
// root. If getGitRoot returns a different path, repoPath is an arbitrary
// subdirectory — skip both the linked-worktree guard and auto-detection
// and fall through to the repoPath fallback.
const repoGitRoot = getGitRoot(repoPath);
const repoCanonical =
repoGitRoot && tryRealpath(repoGitRoot) === tryRealpath(repoPath)
? getCanonicalRepoRoot(repoPath)
: null;
// Early exit: if repoPath is a linked worktree (differs from its canonical
// main-checkout root), return it unchanged. Do NOT override it with the
// server's launch directory — that would silently replace the explicitly-
// resolved worktree index with the main checkout.
//
// getCanonicalRepoRoot returns the main-checkout path for both the checkout
// and all linked worktrees:
// repoPath === canonical → main checkout (auto-detect may fire below)
// repoPath !== canonical → linked worktree (return as-is)
if (repoCanonical && tryRealpath(repoPath) !== tryRealpath(repoCanonical)) {
return repoPath;
}
const launchGitRoot = getGitRoot(launchCwd);
if (launchGitRoot) {
// Normalise via realpathSync before comparing so macOS /var → /private/var
// symlinks (and Windows 8.3 short names) don't create false mismatches.
const realLaunch = tryRealpath(launchGitRoot);
const realRepo = tryRealpath(repoPath);
if (realLaunch !== realRepo) {
const launchCanonical = getCanonicalRepoRoot(launchCwd);
// Use tryRealpath on both canonical values for cross-platform safety.
if (
launchCanonical &&
repoCanonical &&
tryRealpath(launchCanonical) === tryRealpath(repoCanonical)
) {
return launchGitRoot;
}
}
}
} catch {
// Best-effort; fall through to repoPath.
}
return repoPath;
}
export class LocalBackend {
private repos: Map<string, RepoHandle> = new Map();
private contextCache: Map<string, CodebaseContext> = new Map();
@@ -982,7 +1066,7 @@ export class LocalBackend {
timing,
...(!ftsUsed && {
warning:
'FTS indexes missing — keyword search degraded. Run: gitnexus analyze --force to rebuild indexes.',
'FTS indexes missing — keyword search degraded. Run: gitnexus analyze --repair-fts (or gitnexus analyze --force) to rebuild indexes.',
}),
};
}
@@ -1218,31 +1302,41 @@ export class LocalBackend {
}
}
async executeCypher(repoName: string, query: string): Promise<any> {
async executeCypher(
repoName: string,
query: string,
params: Record<string, unknown> = {},
): Promise<any> {
const repo = await this.resolveRepo(repoName);
return this.cypher(repo, { query });
return this.cypher(repo, { query, params });
}
private async cypher(repo: RepoHandle, params: { query: string }): Promise<any> {
private async cypher(
repo: RepoHandle,
request: { query: string; params?: Record<string, unknown> },
): Promise<any> {
await this.ensureInitialized(repo.id);
if (!isLbugReady(repo.id)) {
return { error: 'LadybugDB not ready. Index may be corrupted.' };
}
// Block write operations (defense-in-depth — DB is already read-only)
if (isWriteQuery(params.query)) {
if (request.params !== undefined && !isValidQueryParams(request.params)) {
return {
error:
'Write operations (CREATE, DELETE, SET, MERGE, REMOVE, DROP, ALTER, COPY, DETACH) are not allowed. The knowledge graph is read-only.',
error: '"params" must be a plain object with scalar values (string/number/boolean/null).',
};
}
try {
const result = await executeQuery(repo.id, params.query);
const result = await executeParameterized(repo.id, request.query, request.params ?? {});
return result;
} catch (err: any) {
const msg = err.message || 'Query failed';
if (isReadOnlyDbError(err)) {
return {
error:
'Write operations (CREATE, DELETE, SET, MERGE, REMOVE, DROP, ALTER, COPY, DETACH) are not allowed. The knowledge graph is read-only.',
};
}
if (isWalCorruptionError(err)) {
return {
error: msg,
@@ -2133,6 +2227,7 @@ export class LocalBackend {
params: {
scope?: string;
base_ref?: string;
worktree?: string;
},
): Promise<any> {
await this.ensureInitialized(repo.id);
@@ -2161,11 +2256,51 @@ export class LocalBackend {
let diffOutput: string;
try {
// Resolve the cwd for git diff.
//
// In a linked worktree (e.g. /repo/wt-feature/), the user's staged and
// unstaged changes live in that worktree's separate working directory and
// index. Running `git diff` from the canonical repo root sees a different
// working tree and returns empty output.
//
// Resolution order (see resolveWorktreeCwd for details):
// 1. params.worktree — explicit override, validated against the
// registered repo's canonical root.
// 2. Auto-detect — if the server's launch cwd (process.cwd()) is a
// linked worktree of the same canonical repo, use its git root.
// 3. repo.repoPath — fallback (original behaviour, handled inside
// resolveWorktreeCwd when no worktree is detected).
//
// Start with the auto-detected value; override with the validated
// explicit param when provided. This avoids a dead initial assignment.
let diffCwd = resolveWorktreeCwd(repo.repoPath, process.cwd());
if (params.worktree) {
if (!path.isAbsolute(params.worktree)) {
return {
error: `worktree must be an absolute path, got: "${params.worktree}"`,
};
}
const providedResolved = path.resolve(params.worktree);
const repoCanonical = getCanonicalRepoRoot(repo.repoPath);
if (!repoCanonical) {
return {
error: `Could not determine canonical root for repo "${repo.repoPath}". Is git available?`,
};
}
const worktreeCanonical = getCanonicalRepoRoot(providedResolved);
if (!worktreeCanonical || tryRealpath(worktreeCanonical) !== tryRealpath(repoCanonical)) {
return {
error: `worktree "${params.worktree}" is not a worktree of repo "${repo.repoPath}". Ensure the path is inside the same git repository.`,
};
}
diffCwd = providedResolved;
}
// maxBuffer raised from Node's 1MB default to 256MB to avoid ENOBUFS on
// repos with large unstaged/untracked diffs (e.g. unignored build folders).
// See issue: spawnSync git ENOBUFS in detect_changes(scope="unstaged").
diffOutput = execFileSync('git', diffArgs, {
cwd: repo.repoPath,
cwd: diffCwd,
encoding: 'utf-8',
maxBuffer: 256 * 1024 * 1024,
});
+12
View File
@@ -187,6 +187,11 @@ TIPS:
type: 'object',
properties: {
query: { type: 'string', description: 'Cypher query to execute' },
params: {
type: 'object',
description:
'Optional query parameters for placeholders (e.g. $name) to execute via prepared statement binding.',
},
repo: {
type: 'string',
description: 'Repository name or path. Omit if only one repo is indexed.',
@@ -253,6 +258,8 @@ Maps git diff hunks to indexed symbols, then traces which processes are impacted
WHEN TO USE: Before committing — to understand what your changes affect. Pre-commit review, PR preparation.
AFTER THIS: Review affected processes. Use context() on high-risk symbols. READ gitnexus://repo/{name}/process/{name} for full traces.
GIT WORKTREE SUPPORT: GitNexus automatically detects when the MCP server was launched from inside a linked git worktree and runs git diff against that worktree — no extra parameters needed in the common case. Pass "worktree" explicitly only when the server was started from a different directory than the worktree you are editing (e.g., the server runs from the canonical root but your changes are in a linked worktree at a different path).
Returns: changed symbols, affected processes, and a risk summary.`,
annotations: READ_ONLY_TOOL_ANNOTATIONS,
inputSchema: {
@@ -268,6 +275,11 @@ Returns: changed symbols, affected processes, and a risk summary.`,
type: 'string',
description: 'Branch/commit for "compare" scope (e.g., "main")',
},
worktree: {
type: 'string',
description:
'Absolute path to a linked git worktree. Pass this when your changes are in a worktree (the .git entry at that path is a file, not a directory). GitNexus will run git diff from that worktree so staged/unstaged changes are correctly detected.',
},
repo: {
type: 'string',
description: 'Repository name or path. Omit if only one repo is indexed.',
+164 -123
View File
@@ -22,8 +22,9 @@ import {
flushWAL,
closeLbug,
withLbugDb,
isReadOnlyDbError,
} from '../core/lbug/lbug-adapter.js';
import { isWriteQuery } from '../core/lbug/pool-adapter.js';
import { isValidQueryParams } from '../core/lbug/query-params.js';
import { NODE_TABLES, type GraphNode, type GraphRelationship } from 'gitnexus-shared';
import { searchFTSFromLbug } from '../core/search/bm25-index.js';
import { hybridSearch } from '../core/search/hybrid-search.js';
@@ -447,7 +448,14 @@ export const streamGraphNdjson = async (
*/
const mountSSEProgress = (app: express.Express, routePath: string, jm: JobManager) => {
app.get(routePath, (req, res) => {
const job = jm.getJob(req.params.jobId);
let jobId: string;
try {
jobId = assertString(req.params.jobId, 'jobId');
} catch (err: any) {
res.status(err.status ?? 400).json({ error: err.message });
return;
}
const job = jm.getJob(jobId);
if (!job) {
res.status(404).json({ error: 'Job not found' });
return;
@@ -493,7 +501,7 @@ const mountSSEProgress = (app: express.Express, routePath: string, jm: JobManage
try {
eventId++;
if (progress.phase === 'complete' || progress.phase === 'failed') {
const eventJob = jm.getJob(req.params.jobId);
const eventJob = jm.getJob(jobId);
res.write(
`id: ${eventId}\nevent: ${progress.phase}\ndata: ${JSON.stringify({
repoName: eventJob?.repoName,
@@ -621,6 +629,44 @@ export const handleFileRequest = async (
}
};
export const handleQueryRequest = async (
req: express.Request,
res: express.Response,
resolveRepo: (repoName?: string) => Promise<{ storagePath: string } | undefined>,
): Promise<void> => {
try {
const cypher = req.body.cypher as string;
if (!cypher) {
res.status(400).json({ error: 'Missing "cypher" in request body' });
return;
}
const queryParams = req.body.params;
if (queryParams !== undefined && !isValidQueryParams(queryParams)) {
res.status(400).json({
error: '"params" must be a plain object with scalar values (string/number/boolean/null)',
});
return;
}
const entry = await resolveRepo(requestedRepo(req));
if (!entry) {
res.status(404).json({ error: 'Repository not found' });
return;
}
const lbugPath = path.join(entry.storagePath, 'lbug');
const result = await withLbugDb(lbugPath, () => executePrepared(cypher, queryParams ?? {}), {
readOnly: true,
});
res.json({ result });
} catch (err: any) {
if (isReadOnlyDbError(err)) {
res.status(403).json({ error: 'Write queries are not allowed via the HTTP API' });
return;
}
res.status(500).json({ error: err.message || 'Query failed' });
}
};
export const createServer = async (port: number, host: string = '127.0.0.1') => {
const app = express();
app.disable('x-powered-by');
@@ -984,8 +1030,16 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
res.once('close', abortStreaming);
try {
await withLbugDb(lbugPath, async () =>
streamGraphNdjson(res, includeContent, abortController.signal),
// Read-only open: /api/graph never writes. Write-mode opens engage
// LadybugDB's checkpoint machinery (`.shadow` sidecar), which on
// Windows races with the OS file handle release and trips
// "Cannot open file ... lbug.shadow - Error 2". See pool-adapter.ts
// which already opens read-only for the same reason, and the
// /api/query precedent in PR #1655.
await withLbugDb(
lbugPath,
async () => streamGraphNdjson(res, includeContent, abortController.signal),
{ readOnly: true },
);
if (!abortController.signal.aborted && !res.writableEnded) {
res.end();
@@ -998,7 +1052,9 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
return;
}
const graph = await withLbugDb(lbugPath, async () => buildGraph(includeContent));
const graph = await withLbugDb(lbugPath, async () => buildGraph(includeContent), {
readOnly: true,
});
res.json(graph);
} catch (err: any) {
if (err instanceof ClientDisconnectedError) {
@@ -1020,29 +1076,7 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
// Execute Cypher query
app.post('/api/query', async (req, res) => {
try {
const cypher = req.body.cypher as string;
if (!cypher) {
res.status(400).json({ error: 'Missing "cypher" in request body' });
return;
}
if (isWriteQuery(cypher)) {
res.status(403).json({ error: 'Write queries are not allowed via the HTTP API' });
return;
}
const entry = await resolveRepo(requestedRepo(req));
if (!entry) {
res.status(404).json({ error: 'Repository not found' });
return;
}
const lbugPath = path.join(entry.storagePath, 'lbug');
const result = await withLbugDb(lbugPath, () => executeQuery(cypher));
res.json({ result });
} catch (err: any) {
res.status(500).json({ error: err.message || 'Query failed' });
}
await handleQueryRequest(req, res, resolveRepo);
});
// Search (supports mode: 'hybrid' | 'semantic' | 'bm25', and optional enrichment)
@@ -1067,68 +1101,70 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
const mode: string = req.body.mode ?? 'hybrid';
const enrich: boolean = req.body.enrich !== false; // default true
const results = await withLbugDb(lbugPath, async () => {
let searchResults: any[];
let ftsAvailable: boolean | undefined;
const results = await withLbugDb(
lbugPath,
async () => {
let searchResults: any[];
let ftsAvailable: boolean | undefined;
if (mode === 'semantic') {
const { isEmbedderReady } = await import('../core/embeddings/embedder.js');
if (!isEmbedderReady()) {
return { searchResults: [] as any[], ftsAvailable: undefined };
}
const { semanticSearch: semSearch } =
await import('../core/embeddings/embedding-pipeline.js');
searchResults = await semSearch(executeQuery, query, limit);
// Normalize semantic results to HybridSearchResult shape
searchResults = searchResults.map((r: any, i: number) => ({
...r,
score: r.score ?? 1 - (r.distance ?? 0),
rank: i + 1,
sources: ['semantic'],
}));
} else if (mode === 'bm25') {
const ftsResponse = await searchFTSFromLbug(query, limit);
ftsAvailable = ftsResponse.ftsAvailable;
searchResults = ftsResponse.results.map((r: any, i: number) => ({
...r,
rank: i + 1,
sources: ['bm25'],
}));
} else {
// hybrid (default)
const { isEmbedderReady } = await import('../core/embeddings/embedder.js');
if (isEmbedderReady()) {
if (mode === 'semantic') {
const { isEmbedderReady } = await import('../core/embeddings/embedder.js');
if (!isEmbedderReady()) {
return { searchResults: [] as any[], ftsAvailable: undefined };
}
const { semanticSearch: semSearch } =
await import('../core/embeddings/embedding-pipeline.js');
searchResults = await hybridSearch(query, limit, executeQuery, semSearch);
} else {
searchResults = await semSearch(executeQuery, query, limit);
// Normalize semantic results to HybridSearchResult shape
searchResults = searchResults.map((r: any, i: number) => ({
...r,
score: r.score ?? 1 - (r.distance ?? 0),
rank: i + 1,
sources: ['semantic'],
}));
} else if (mode === 'bm25') {
const ftsResponse = await searchFTSFromLbug(query, limit);
ftsAvailable = ftsResponse.ftsAvailable;
searchResults = ftsResponse.results;
searchResults = ftsResponse.results.map((r: any, i: number) => ({
...r,
rank: i + 1,
sources: ['bm25'],
}));
} else {
// hybrid (default)
const { isEmbedderReady } = await import('../core/embeddings/embedder.js');
if (isEmbedderReady()) {
const { semanticSearch: semSearch } =
await import('../core/embeddings/embedding-pipeline.js');
searchResults = await hybridSearch(query, limit, executeQuery, semSearch);
} else {
const ftsResponse = await searchFTSFromLbug(query, limit);
ftsAvailable = ftsResponse.ftsAvailable;
searchResults = ftsResponse.results;
}
}
}
if (!enrich) return { searchResults, ftsAvailable };
if (!enrich) return { searchResults, ftsAvailable };
// Server-side enrichment: add connections, cluster, processes per result
// Uses parameterized queries to prevent Cypher injection via nodeId
const validLabel = (label: string): boolean =>
(NODE_TABLES as readonly string[]).includes(label);
// Server-side enrichment: add connections, cluster, processes per result
// Uses parameterized queries to prevent Cypher injection via nodeId
const validLabel = (label: string): boolean =>
(NODE_TABLES as readonly string[]).includes(label);
const enriched = await Promise.all(
searchResults.slice(0, limit).map(async (r: any) => {
const nodeId: string = r.nodeId || r.id || '';
const nodeLabel = nodeId.split(':')[0];
const enrichment: { connections?: any; cluster?: string; processes?: any[] } = {};
const enriched = await Promise.all(
searchResults.slice(0, limit).map(async (r: any) => {
const nodeId: string = r.nodeId || r.id || '';
const nodeLabel = nodeId.split(':')[0];
const enrichment: { connections?: any; cluster?: string; processes?: any[] } = {};
if (!nodeId || !validLabel(nodeLabel)) return { ...r, ...enrichment };
if (!nodeId || !validLabel(nodeLabel)) return { ...r, ...enrichment };
// Run connections, cluster, and process queries in parallel
// Label is validated against NODE_TABLES (compile-time safe identifiers);
// nodeId uses $nid parameter binding to prevent injection
const [connRes, clusterRes, procRes] = await Promise.all([
executePrepared(
`
// Run connections, cluster, and process queries in parallel
// Label is validated against NODE_TABLES (compile-time safe identifiers);
// nodeId uses $nid parameter binding to prevent injection
const [connRes, clusterRes, procRes] = await Promise.all([
executePrepared(
`
MATCH (n:${nodeLabel} {id: $nid})
OPTIONAL MATCH (n)-[r1:CodeRelation]->(dst)
OPTIONAL MATCH (src)-[r2:CodeRelation]->(n)
@@ -1137,65 +1173,67 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
collect(DISTINCT {name: src.name, type: r2.type, confidence: r2.confidence}) AS incoming
LIMIT 1
`,
{ nid: nodeId },
).catch(() => []),
executePrepared(
`
{ nid: nodeId },
).catch(() => []),
executePrepared(
`
MATCH (n:${nodeLabel} {id: $nid})
MATCH (n)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
RETURN c.label AS label, c.description AS description
LIMIT 1
`,
{ nid: nodeId },
).catch(() => []),
executePrepared(
`
{ nid: nodeId },
).catch(() => []),
executePrepared(
`
MATCH (n:${nodeLabel} {id: $nid})
MATCH (n)-[rel:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
RETURN p.id AS id, p.label AS label, rel.step AS step, p.stepCount AS stepCount
ORDER BY rel.step
`,
{ nid: nodeId },
).catch(() => []),
]);
{ nid: nodeId },
).catch(() => []),
]);
if (connRes.length > 0) {
const row = connRes[0];
const outgoing = (Array.isArray(row) ? row[0] : row.outgoing || [])
.filter((c: any) => c?.name)
.slice(0, 5);
const incoming = (Array.isArray(row) ? row[1] : row.incoming || [])
.filter((c: any) => c?.name)
.slice(0, 5);
enrichment.connections = { outgoing, incoming };
}
if (connRes.length > 0) {
const row = connRes[0];
const outgoing = (Array.isArray(row) ? row[0] : row.outgoing || [])
.filter((c: any) => c?.name)
.slice(0, 5);
const incoming = (Array.isArray(row) ? row[1] : row.incoming || [])
.filter((c: any) => c?.name)
.slice(0, 5);
enrichment.connections = { outgoing, incoming };
}
if (clusterRes.length > 0) {
const row = clusterRes[0];
enrichment.cluster = Array.isArray(row) ? row[0] : row.label;
}
if (clusterRes.length > 0) {
const row = clusterRes[0];
enrichment.cluster = Array.isArray(row) ? row[0] : row.label;
}
if (procRes.length > 0) {
enrichment.processes = procRes
.map((row: any) => ({
id: Array.isArray(row) ? row[0] : row.id,
label: Array.isArray(row) ? row[1] : row.label,
step: Array.isArray(row) ? row[2] : row.step,
stepCount: Array.isArray(row) ? row[3] : row.stepCount,
}))
.filter((p: any) => p.id && p.label);
}
if (procRes.length > 0) {
enrichment.processes = procRes
.map((row: any) => ({
id: Array.isArray(row) ? row[0] : row.id,
label: Array.isArray(row) ? row[1] : row.label,
step: Array.isArray(row) ? row[2] : row.step,
stepCount: Array.isArray(row) ? row[3] : row.stepCount,
}))
.filter((p: any) => p.id && p.label);
}
return { ...r, ...enrichment };
}),
);
return { ...r, ...enrichment };
}),
);
return { searchResults: enriched, ftsAvailable };
});
return { searchResults: enriched, ftsAvailable };
},
{ readOnly: true },
);
const response: any = { results: results.searchResults ?? results };
if (results.ftsAvailable === false) {
response.warning =
'FTS indexes missing — keyword search degraded. Run: gitnexus analyze --force to rebuild indexes.';
'FTS indexes missing — keyword search degraded. Run: gitnexus analyze --repair-fts (or gitnexus analyze --force) to rebuild indexes.';
}
res.json(response);
} catch (err: any) {
@@ -1271,8 +1309,11 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
// Get file paths from the graph (lightweight — no content loaded)
const lbugPath = path.join(entry.storagePath, 'lbug');
const fileRows = await withLbugDb(lbugPath, () =>
executeQuery(`MATCH (n:File) WHERE n.content IS NOT NULL RETURN n.filePath AS filePath`),
const fileRows = await withLbugDb(
lbugPath,
() =>
executeQuery(`MATCH (n:File) WHERE n.content IS NOT NULL RETURN n.filePath AS filePath`),
{ readOnly: true },
);
// Search files on disk one at a time (constant memory)
@@ -2,6 +2,7 @@ export class User {
save() {}
}
/** @returns {User} */
export function getUser() {
return new User();
}
@@ -3,6 +3,7 @@ export class User {
getName() { return ''; }
}
/** @returns {User} */
export function getUser() {
return new User();
}
@@ -0,0 +1,12 @@
#include "lib.h"
void Service::f(int* p) {}
void Service::f(bool flag) {}
void Service::g(int a, int b) {}
void Service::g(int a, ...) {}
void Service::h(int a, double b) {}
void Service::h(int a, ...) {}
void Service::k(int a, ...) {}
@@ -0,0 +1,38 @@
#pragma once
class Service {
public:
void f(int* p);
void f(bool flag);
void g(int a, int b);
void g(int a, ...);
void h(int a, double b);
void h(int a, ...);
void k(int a, ...);
void runNullptr() {
f(nullptr);
}
void runPointer() {
int* p = nullptr;
f(p);
}
void runBoolConversion() {
f(42);
}
void run() {
int* p = nullptr;
f(nullptr);
f(p);
f(42);
g(1, 2);
h(1, 'a');
k(1, 2, 3);
}
};
@@ -0,0 +1,16 @@
#include <type_traits>
struct S {};
template <class T, std::enable_if_t<std::is_class_v<T>, int> = 0>
void pick(T value) {}
template <class T, std::enable_if_t<std::is_integral_v<T>, int> = 0>
void pick(T value) {}
void run() {
S s;
int n = 0;
pick(s);
pick(n);
}
@@ -0,0 +1,14 @@
#include <type_traits>
template <class T, std::enable_if_t<std::is_const_v<T>, int> = 0>
void pick(T value) {}
template <class T, std::enable_if_t<std::is_volatile_v<T>, int> = 0>
void pick(T value) {}
void run() {
const int c = 0;
volatile int v = 0;
pick(c);
pick(v);
}
@@ -0,0 +1,16 @@
#include <type_traits>
enum Color { Red };
template <class T, std::enable_if_t<std::is_enum_v<T>, int> = 0>
void pick(T value) {}
template <class T, std::enable_if_t<std::is_integral_v<T>, int> = 0>
void pick(T value) {}
void run() {
Color color = Red;
int n = 0;
pick(color);
pick(n);
}
@@ -0,0 +1,14 @@
#include <type_traits>
struct S {};
template <class T, std::enable_if_t<std::is_pointer_v<T>, int> = 0>
void pick(T value) {}
template <class T, std::enable_if_t<std::is_class_v<T>, int> = 0>
void pick(T value) {}
void run(S* p, S s) {
pick(p);
pick(s);
}
@@ -0,0 +1,14 @@
#include <type_traits>
template <class T, std::enable_if_t<std::is_reference_v<T>, int> = 0>
void pick(T value) {}
template <class T, std::enable_if_t<std::is_integral_v<T>, int> = 0>
void pick(T value) {}
void run() {
int n = 0;
int& r = n;
pick(r);
pick(n);
}
@@ -0,0 +1,12 @@
#include <type_traits>
template <class T, std::enable_if_t<std::is_void_v<T>, int> = 0>
void pick(T value) {}
template <class T, std::enable_if_t<std::is_pointer_v<T>, int> = 0>
void pick(T value) {}
void run() {
void* p;
pick(p);
}
@@ -0,0 +1,8 @@
package com.example;
public class Module1App {
public void run() {
UserService service = new UserService();
service.ping();
}
}
@@ -0,0 +1,6 @@
package com.example;
public class UserService {
public void ping() {
}
}
@@ -0,0 +1,8 @@
package com.example;
public class Module2App {
public void run() {
UserService service = new UserService();
service.ping();
}
}
@@ -0,0 +1,6 @@
package com.example;
public class UserService {
public void ping() {
}
}
@@ -0,0 +1,3 @@
import { AmbientBase } from './ambient';
export class Derived extends AmbientBase {}
@@ -0,0 +1,7 @@
// Ambient base class — simulates a .d.ts-declared external/library type
// whose body is never seen by the analyzer. Probes whether Step 2 MRO
// lookup can still resolve inherited members on owners that reconcile-
// ownership skipped because they have no parsed body.
export declare class AmbientBase {
ambientMethod(): string;
}
@@ -0,0 +1,6 @@
import { Derived } from './Derived';
export function run(): void {
const d = new Derived();
d.ambientMethod();
}
+5
View File
@@ -0,0 +1,5 @@
import fs from 'node:fs';
import path from 'node:path';
export const hasLadybugNative = (): boolean =>
fs.existsSync(path.join(process.cwd(), 'node_modules', '@ladybugdb', 'core', 'lbugjs.node'));
+43 -8
View File
@@ -4,7 +4,8 @@
* Creates temporary directories for tests and provides cleanup that tolerates
* LadybugDB's known Windows handle-release lag after retries.
*/
import fs from 'fs/promises';
import fs from 'fs';
import fsp from 'fs/promises';
import os from 'os';
import path from 'path';
@@ -13,22 +14,56 @@ export interface TestDBHandle {
cleanup: () => Promise<void>;
}
const CLEANUP_MAX_ATTEMPTS = 5;
const WINDOWS_NATIVE_LOCK_CODES = new Set(['EBUSY', 'EPERM', 'EACCES', 'ENOTEMPTY']);
export async function cleanupTempDir(tmpDir: string): Promise<void> {
const cleanupBackoffMs = (attempt: number): number => 100 * (attempt + 1);
const shouldSwallowCleanupError = (err: unknown): boolean => {
const code = (err as NodeJS.ErrnoException | undefined)?.code;
return process.platform === 'win32' && WINDOWS_NATIVE_LOCK_CODES.has(code ?? '');
};
const sleepSync = (ms: number): void => {
const view = new Int32Array(new SharedArrayBuffer(4));
Atomics.wait(view, 0, 0, ms);
};
export function cleanupTempDirSync(tmpDir: string): void {
let lastError: unknown;
for (let attempt = 0; attempt < 5; attempt++) {
for (let attempt = 0; attempt < CLEANUP_MAX_ATTEMPTS; attempt++) {
try {
await fs.rm(tmpDir, { recursive: true, force: true });
fs.rmSync(tmpDir, { recursive: true, force: true });
return;
} catch (err) {
lastError = err;
await new Promise((resolve) => setTimeout(resolve, 100 * (attempt + 1)));
if (attempt < CLEANUP_MAX_ATTEMPTS - 1) {
sleepSync(cleanupBackoffMs(attempt));
}
}
}
const code = (lastError as NodeJS.ErrnoException | undefined)?.code;
if (process.platform === 'win32' && WINDOWS_NATIVE_LOCK_CODES.has(code ?? '')) {
if (shouldSwallowCleanupError(lastError)) {
return;
}
throw lastError;
}
export async function cleanupTempDir(tmpDir: string): Promise<void> {
let lastError: unknown;
for (let attempt = 0; attempt < CLEANUP_MAX_ATTEMPTS; attempt++) {
try {
await fsp.rm(tmpDir, { recursive: true, force: true });
return;
} catch (err) {
lastError = err;
if (attempt < CLEANUP_MAX_ATTEMPTS - 1) {
await new Promise((resolve) => setTimeout(resolve, cleanupBackoffMs(attempt)));
}
}
}
if (shouldSwallowCleanupError(lastError)) {
return;
}
throw lastError;
@@ -46,7 +81,7 @@ export async function cleanupTempDir(tmpDir: string): Promise<void> {
* return.
*/
export async function createTempDir(prefix: string = 'gitnexus-test-'): Promise<TestDBHandle> {
const tmpDir = await fs.mkdtemp(path.join(os.tmpdir(), prefix));
const tmpDir = await fsp.mkdtemp(path.join(os.tmpdir(), prefix));
return {
dbPath: tmpDir,
cleanup: async () => {
+3 -2
View File
@@ -125,8 +125,9 @@ export function withTestLbugDB(
// LadybugDB enforces file locks — writable + read-only can't coexist
// on the same path, and db.close() segfaults on macOS due to N-API
// destructor issues. Reusing the writable Database avoids both problems.
// Write protection is enforced at the query validation layer (isWriteQuery)
// rather than at the native DB level.
// NOTE: This injected DB is writable by design for test setup.
// Read-only enforcement tests must initialize a separate pool entry
// via initLbug(...) so Ladybug native read-only mode is exercised.
if (options?.poolAdapter) {
const coreDb = adapter.getDatabase();
if (!coreDb) throw new Error('withTestLbugDB: core adapter has no open Database');
@@ -0,0 +1,74 @@
import { describe, it, expect } from 'vitest';
import { spawnSync } from 'node:child_process';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { fileURLToPath } from 'node:url';
const testDir = path.dirname(fileURLToPath(import.meta.url));
const repoRoot = path.resolve(testDir, '../..');
const distCli = path.join(repoRoot, 'dist', 'cli', 'index.js');
const fixtureSource = path.resolve(testDir, '..', 'fixtures', 'mini-repo');
const runAnalyzeWithForcedOom = (cwd: string, gitnexusHome: string) =>
spawnSync(process.execPath, [distCli, 'analyze'], {
cwd,
encoding: 'utf8',
timeout: process.env.CI ? 40_000 : 20_000,
stdio: ['pipe', 'pipe', 'pipe'],
env: {
...process.env,
GITNEXUS_HOME: gitnexusHome,
NODE_OPTIONS: '',
GITNEXUS_TEST_RESPAWN_HEAP_MB: '32',
GITNEXUS_TEST_FORCE_HEAP_OOM: '1',
CI: '1',
},
});
describe('analyze OOM guidance (real child-process OOM)', () => {
it('prints OOM guidance with Unix and Windows commands when respawned child truly OOMs', () => {
if (!fs.existsSync(distCli)) {
throw new Error(
'dist/cli/index.js missing — run `npm run build` first (or use `npm run test:integration`, which builds via pretest:integration).',
);
}
const oomTestRepoParent = fs.mkdtempSync(path.join(os.tmpdir(), 'gn-oom-e2e-repo-'));
const oomTestGitnexusHome = fs.mkdtempSync(path.join(os.tmpdir(), 'gn-oom-e2e-home-'));
const repoPath = path.join(oomTestRepoParent, 'mini-repo');
fs.cpSync(fixtureSource, repoPath, { recursive: true });
spawnSync('git', ['init'], { cwd: repoPath, stdio: 'pipe' });
spawnSync('git', ['add', '-A'], { cwd: repoPath, stdio: 'pipe' });
spawnSync('git', ['commit', '-m', 'initial commit'], {
cwd: repoPath,
stdio: 'pipe',
env: {
...process.env,
GIT_AUTHOR_NAME: 'test',
GIT_AUTHOR_EMAIL: 'test@test',
GIT_COMMITTER_NAME: 'test',
GIT_COMMITTER_EMAIL: 'test@test',
},
});
try {
const result = runAnalyzeWithForcedOom(repoPath, oomTestGitnexusHome);
const combinedOutput = `${result.stderr}\n${result.stdout}`;
expect(result.status).not.toBeNull();
expect(result.status).not.toBe(0);
expect(combinedOutput).toContain('Analysis likely ran out of memory.');
expect(combinedOutput).toContain(
'NODE_OPTIONS="--max-old-space-size=24576" gitnexus analyze [your-args]',
);
expect(combinedOutput).toContain(
'(Windows: set NODE_OPTIONS=--max-old-space-size=24576 && gitnexus analyze [your-args])',
);
} finally {
fs.rmSync(oomTestRepoParent, { recursive: true, force: true });
fs.rmSync(oomTestGitnexusHome, { recursive: true, force: true });
}
}, 60_000);
});
@@ -0,0 +1,97 @@
import express from 'express';
import http from 'node:http';
import { describe, expect, it, beforeAll, afterAll } from 'vitest';
import { withTestLbugDB } from '../helpers/test-indexed-db.js';
import { hasLadybugNative } from '../helpers/ladybug-native.js';
const WRITE_QUERY_TEST_CYPHER =
"CREATE (n:Function {id: 'api-write-test', name: 'api-write-test', filePath: '', startLine: 0, endLine: 0, isExported: false, content: '', description: ''})";
const startServer = (app: express.Express): Promise<{ server: http.Server; baseUrl: string }> =>
new Promise((resolve) => {
const server = app.listen(0, '127.0.0.1', () => {
const addr = server.address();
if (!addr || typeof addr === 'string') throw new Error('Failed to start test server');
resolve({ server, baseUrl: `http://127.0.0.1:${addr.port}` });
});
});
const stopServer = (server: http.Server): Promise<void> =>
new Promise((resolve, reject) => server.close((err) => (err ? reject(err) : resolve())));
withTestLbugDB(
'api-query-http',
(handle) => {
describe.skipIf(!hasLadybugNative())('/api/query runtime contract', () => {
let server: http.Server;
let baseUrl = '';
let handleQueryRequest: typeof import('../../src/server/api.js').handleQueryRequest;
beforeAll(async () => {
({ handleQueryRequest } = await import('../../src/server/api.js'));
const app = express();
app.use(express.json());
app.post('/api/query', async (req, res) => {
await handleQueryRequest(req, res, async () => ({
storagePath: handle.tmpHandle.dbPath,
}));
});
({ server, baseUrl } = await startServer(app));
});
afterAll(async () => {
await stopServer(server);
});
it('returns 200 for a valid read query', async () => {
const response = await fetch(`${baseUrl}/api/query`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ cypher: 'RETURN 1 AS one' }),
});
expect(response.status).toBe(200);
const body = await response.json();
expect(Array.isArray(body.result)).toBe(true);
expect(body.result[0].one).toBe(1);
});
it('returns 403 for a write query on read-only HTTP path', async () => {
const response = await fetch(`${baseUrl}/api/query`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({
cypher: WRITE_QUERY_TEST_CYPHER,
}),
});
expect(response.status).toBe(403);
const body = await response.json();
expect(body.error).toContain('Write queries are not allowed');
});
it('returns 400 for invalid params payload', async () => {
const response = await fetch(`${baseUrl}/api/query`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ cypher: 'RETURN 1 AS one', params: [1, 2, 3] }),
});
expect(response.status).toBe(400);
const body = await response.json();
expect(body.error).toContain('"params"');
});
it('returns 400 when cypher is missing', async () => {
const response = await fetch(`${baseUrl}/api/query`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({}),
});
expect(response.status).toBe(400);
const body = await response.json();
expect(body.error).toContain('Missing "cypher"');
});
});
},
{
poolAdapter: false,
},
);
+277 -26
View File
@@ -16,6 +16,7 @@ import os from 'os';
import { fileURLToPath, pathToFileURL } from 'url';
import { createRequire } from 'module';
import { cleanupTempDirSync } from '../helpers/test-db.js';
const testDir = path.dirname(fileURLToPath(import.meta.url));
const repoRoot = path.resolve(testDir, '../..');
@@ -75,10 +76,10 @@ afterAll(() => {
// Entire tmp copy goes away — no selective cleanup needed. The shared
// `test/fixtures/mini-repo/` source was never touched.
if (tmpParent) {
fs.rmSync(tmpParent, { recursive: true, force: true });
cleanupTempDirSync(tmpParent);
}
if (suiteGitnexusHome) {
fs.rmSync(suiteGitnexusHome, { recursive: true, force: true });
cleanupTempDirSync(suiteGitnexusHome);
}
});
@@ -268,8 +269,8 @@ describe('CLI end-to-end', () => {
`registry has no entry for ${repo}; entries: ${JSON.stringify(entries.map((e) => e.path))}`,
).toBe(true);
} finally {
fs.rmSync(gnHome, { recursive: true, force: true });
fs.rmSync(repoParent, { recursive: true, force: true });
cleanupTempDirSync(gnHome);
cleanupTempDirSync(repoParent);
}
}, 60_000);
@@ -310,8 +311,8 @@ describe('CLI end-to-end', () => {
expect(`${second.stdout}${second.stderr}`).toMatch(/registry entry/i);
expect(second.status).toBe(1);
} finally {
fs.rmSync(gnHome, { recursive: true, force: true });
fs.rmSync(repoParent, { recursive: true, force: true });
cleanupTempDirSync(gnHome);
cleanupTempDirSync(repoParent);
}
}, 60_000);
@@ -457,12 +458,12 @@ describe('CLI end-to-end', () => {
const afterStep4 = JSON.parse(fs.readFileSync(registryPath, 'utf-8'));
expect(afterStep4).toHaveLength(2);
} finally {
fs.rmSync(parentC, { recursive: true, force: true });
cleanupTempDirSync(parentC);
}
} finally {
fs.rmSync(gnHome, { recursive: true, force: true });
fs.rmSync(parentA, { recursive: true, force: true });
fs.rmSync(parentB, { recursive: true, force: true });
cleanupTempDirSync(gnHome);
cleanupTempDirSync(parentA);
cleanupTempDirSync(parentB);
}
}, 360000); // 6-min outer budget (4 × ~60s analyze calls + fixture setup)
});
@@ -571,8 +572,8 @@ describe('CLI end-to-end', () => {
expect(r4.status).toBe(0);
expect(`${r4.stdout}${r4.stderr}`).toMatch(/Nothing to remove/i);
} finally {
fs.rmSync(gnHome, { recursive: true, force: true });
fs.rmSync(parentA, { recursive: true, force: true });
cleanupTempDirSync(gnHome);
cleanupTempDirSync(parentA);
}
}, 180000); // 3-min outer budget (1 × ~60s analyze + 3 × fast remove calls)
@@ -675,9 +676,9 @@ describe('CLI end-to-end', () => {
// And it's NOT the one we just removed.
expect(finalEntries[0].path).not.toBe(repoAEntry.path);
} finally {
fs.rmSync(gnHome, { recursive: true, force: true });
fs.rmSync(parentA, { recursive: true, force: true });
fs.rmSync(parentB, { recursive: true, force: true });
cleanupTempDirSync(gnHome);
cleanupTempDirSync(parentA);
cleanupTempDirSync(parentB);
}
}, 240000); // 4-min outer budget (2 × ~60s analyze + 2 × fast remove)
@@ -759,8 +760,8 @@ describe('CLI end-to-end', () => {
expect(afterRegistry).toHaveLength(1);
expect(afterRegistry[0].storagePath).toBe(repo); // still poisoned (we did that)
} finally {
fs.rmSync(gnHome, { recursive: true, force: true });
fs.rmSync(parent, { recursive: true, force: true });
cleanupTempDirSync(gnHome);
cleanupTempDirSync(parent);
}
}, 120000); // 2-min budget (1 × ~60s analyze + 1 × fast remove-refused)
});
@@ -864,9 +865,9 @@ describe('CLI end-to-end', () => {
expect(afterRegistry).toHaveLength(1);
expect(afterRegistry[0].name).toBe('bad-alias');
} finally {
fs.rmSync(gnHome, { recursive: true, force: true });
fs.rmSync(parentBad, { recursive: true, force: true });
fs.rmSync(parentGood, { recursive: true, force: true });
cleanupTempDirSync(gnHome);
cleanupTempDirSync(parentBad);
cleanupTempDirSync(parentGood);
}
}, 240000); // 4-min budget (2 × ~60s analyze + 1 × fast clean --all)
});
@@ -954,7 +955,7 @@ describe('CLI end-to-end', () => {
expect(result.status).toBe(0);
expect(result.stdout).toMatch(/Repository not indexed/);
} finally {
fs.rmSync(tmpDir, { recursive: true, force: true });
cleanupTempDirSync(tmpDir);
}
});
@@ -968,7 +969,7 @@ describe('CLI end-to-end', () => {
expect(result.status).toBe(0);
expect(result.stdout).toMatch(/Not a git repository/);
} finally {
fs.rmSync(tmpDir, { recursive: true, force: true });
cleanupTempDirSync(tmpDir);
}
});
@@ -984,7 +985,7 @@ describe('CLI end-to-end', () => {
expect(result.status).toBe(1);
expect(result.stdout).toMatch(/not.*git repository/i);
} finally {
fs.rmSync(tmpDir, { recursive: true, force: true });
cleanupTempDirSync(tmpDir);
}
});
});
@@ -1014,7 +1015,7 @@ describe('CLI end-to-end', () => {
expect(result.status).toBe(1);
expect(result.stdout).toMatch(/not.*git repository/i);
} finally {
fs.rmSync(tmpDir, { recursive: true, force: true });
cleanupTempDirSync(tmpDir);
}
});
@@ -1051,7 +1052,7 @@ describe('CLI end-to-end', () => {
expect(result.status).toBe(1);
expect(result.stdout).toMatch(/No GitNexus index found/);
} finally {
fs.rmSync(tmpDir, { recursive: true, force: true });
cleanupTempDirSync(tmpDir);
}
});
@@ -1217,7 +1218,7 @@ describe('CLI end-to-end', () => {
child.stdout.on('data', (chunk: Buffer) => {
stdoutBuffer += chunk.toString();
if (stdoutBuffer.includes('GITNEXUS_EVAL_SERVER_READY:')) {
if (stdoutBuffer.includes('GITNEXUS_EVAL_SERVER_READY:127.0.0.1:')) {
foundOnStdout = true;
child.kill('SIGTERM');
}
@@ -1255,4 +1256,254 @@ describe('CLI end-to-end', () => {
});
}, 35000);
});
// ─── eval-server --host flag tests ───────────────────────────────────
// Verifies --host is wired to the actual bind address, not just accepted.
// Original flag registration test by Val Vladescu (PR #1602).
describe('eval-server --host flag', () => {
it('emits READY signal containing the bound host 127.0.0.1', () => {
return new Promise<void>((resolve, reject) => {
const child = spawn(
process.execPath,
[
'--import',
tsxImportUrl,
cliEntry,
'eval-server',
'--port',
'0',
'--host',
'127.0.0.1',
'--idle-timeout',
'3',
],
{
cwd: MINI_REPO,
stdio: ['ignore', 'pipe', 'pipe'],
env: cliEnv(),
},
);
let stdoutBuffer = '';
let stderrBuffer = '';
let settled = false;
const settle = (fn: () => void) => {
if (settled) return;
settled = true;
clearTimeout(timer);
child.kill('SIGTERM');
fn();
};
child.stdout.on('data', (chunk: Buffer) => {
stdoutBuffer += chunk.toString();
if (stdoutBuffer.includes('GITNEXUS_EVAL_SERVER_READY:')) {
if (stdoutBuffer.includes('GITNEXUS_EVAL_SERVER_READY:127.0.0.1:')) {
settle(resolve);
} else {
settle(() =>
reject(
new Error(
`READY signal did not contain expected host 127.0.0.1:\n${stdoutBuffer}`,
),
),
);
}
}
});
child.stderr.on('data', (chunk: Buffer) => {
stderrBuffer += chunk.toString();
if (stderrBuffer.includes('unknown option') || stderrBuffer.includes('error: unknown')) {
settle(() => reject(new Error(`eval-server rejected --host flag:\n${stderrBuffer}`)));
}
});
const timer = setTimeout(() => {
settle(() => reject(new Error('eval-server did not emit READY signal within 30s')));
}, 30000);
});
}, 35000);
it('binds to 0.0.0.0 and serves /health on 127.0.0.1 (cross-container use case)', () => {
return new Promise<void>((resolve, reject) => {
const child = spawn(
process.execPath,
[
'--import',
tsxImportUrl,
cliEntry,
'eval-server',
'--port',
'0',
'--host',
'0.0.0.0',
'--idle-timeout',
'3',
],
{
cwd: MINI_REPO,
stdio: ['ignore', 'pipe', 'pipe'],
env: cliEnv(),
},
);
let stdoutBuffer = '';
let settled = false;
const settle = (fn: () => void) => {
if (settled) return;
settled = true;
clearTimeout(timer);
child.kill('SIGTERM');
fn();
};
child.stdout.on('data', async (chunk: Buffer) => {
stdoutBuffer += chunk.toString();
const readyLine = stdoutBuffer
.split('\n')
.find((l) => l.startsWith('GITNEXUS_EVAL_SERVER_READY:0.0.0.0:'));
if (!readyLine || settled) return;
// Parse the actual OS-assigned port from the READY signal
const boundPort = readyLine.split(':').pop()?.trim();
if (!boundPort || isNaN(Number(boundPort))) {
settle(() => reject(new Error(`Could not parse port from READY signal: ${readyLine}`)));
return;
}
// A server bound to 0.0.0.0 must be reachable on 127.0.0.1 from the same host
try {
const res = await fetch(`http://127.0.0.1:${boundPort}/health`);
if (res.status === 200) {
settle(resolve);
} else {
settle(() => reject(new Error(`/health returned ${res.status}, expected 200`)));
}
} catch (err) {
settle(() =>
reject(
new Error(
`eval-server bound to 0.0.0.0 but /health unreachable on 127.0.0.1:${boundPort}: ${err}`,
),
),
);
}
});
child.stderr.on('data', (chunk: Buffer) => {
const text = chunk.toString();
if (text.includes('unknown option') || text.includes('error: unknown')) {
settle(() => reject(new Error(`eval-server rejected --host flag:\n${text}`)));
}
});
const timer = setTimeout(() => {
settle(() =>
reject(new Error('eval-server --host 0.0.0.0 did not emit READY signal within 30s')),
);
}, 30000);
});
}, 35000);
it('emits READY signal with bound IP (not literal "localhost") when --host localhost is used', () => {
return new Promise<void>((resolve, reject) => {
const child = spawn(
process.execPath,
[
'--import',
tsxImportUrl,
cliEntry,
'eval-server',
'--port',
'0',
'--host',
'localhost',
'--idle-timeout',
'3',
],
{
cwd: MINI_REPO,
stdio: ['ignore', 'pipe', 'pipe'],
env: cliEnv(),
},
);
let stdoutBuffer = '';
let settled = false;
const settle = (fn: () => void) => {
if (settled) return;
settled = true;
clearTimeout(timer);
child.kill('SIGTERM');
fn();
};
child.stdout.on('data', async (chunk: Buffer) => {
stdoutBuffer += chunk.toString();
const readyLine = stdoutBuffer
.split('\n')
.find((l) => l.startsWith('GITNEXUS_EVAL_SERVER_READY:'));
if (!readyLine || settled) return;
// The signal must contain a real bound IP, not the literal input string
if (readyLine.includes(':localhost:')) {
settle(() =>
reject(
new Error(
`READY signal contained literal "localhost" instead of a bound IP:\n${readyLine}`,
),
),
);
return;
}
// Parse host and port: everything after the prefix up to the last colon
const withoutPrefix = readyLine.slice('GITNEXUS_EVAL_SERVER_READY:'.length);
const lastColon = withoutPrefix.lastIndexOf(':');
const signalHost = withoutPrefix.slice(0, lastColon); // "127.0.0.1" or "[::1]"
const boundPort = withoutPrefix.slice(lastColon + 1).trim();
if (!boundPort || isNaN(Number(boundPort))) {
settle(() => reject(new Error(`Could not parse port from READY signal: ${readyLine}`)));
return;
}
// Probe /health at the bound address to confirm the server is reachable
try {
const res = await fetch(`http://${signalHost}:${boundPort}/health`);
if (res.status === 200) {
settle(resolve);
} else {
settle(() => reject(new Error(`/health returned ${res.status}, expected 200`)));
}
} catch (err) {
settle(() =>
reject(
new Error(
`eval-server bound to localhost but /health unreachable at ${signalHost}:${boundPort}: ${err}`,
),
),
);
}
});
child.stderr.on('data', (chunk: Buffer) => {
const text = chunk.toString();
if (text.includes('unknown option') || text.includes('error: unknown')) {
settle(() => reject(new Error(`eval-server rejected --host flag:\n${text}`)));
}
});
const timer = setTimeout(() => {
settle(() =>
reject(new Error('eval-server --host localhost did not emit READY signal within 30s')),
);
}, 30000);
});
}, 35000);
});
});
@@ -398,5 +398,172 @@ describe('filesystem-walker', () => {
expect(skipWarnings.length).toBeGreaterThan(0);
expect(String(skipWarnings[0].msg ?? '')).toContain('generated/vendored');
});
// Regression: issue #1659. The skipped-paths list and the
// GITNEXUS_MAX_FILE_SIZE hint must appear by default, otherwise users
// see "Skipped N large files" with no actionable detail and misdiagnose
// missing IMPORTS/CALLS edges as a resolver bug.
it('lists the skipped path by default (not gated behind GITNEXUS_VERBOSE)', async () => {
await walkRepositoryPaths(sizeDir);
const pathWarnings = cap.records().filter((r) => String(r.msg ?? '').includes(BIG_FILE));
expect(pathWarnings.length).toBeGreaterThan(0);
});
it('emits a GITNEXUS_MAX_FILE_SIZE hint when running with the default cap', async () => {
await walkRepositoryPaths(sizeDir);
const hint = cap
.records()
.filter((r) => String(r.msg ?? '').includes('GITNEXUS_MAX_FILE_SIZE=<KB>'));
expect(hint.length).toBe(1);
});
it('omits the GITNEXUS_MAX_FILE_SIZE hint when an override is active', async () => {
process.env.GITNEXUS_MAX_FILE_SIZE = '1';
await walkRepositoryPaths(sizeDir);
const hint = cap
.records()
.filter((r) => String(r.msg ?? '').includes('GITNEXUS_MAX_FILE_SIZE=<KB>'));
expect(hint.length).toBe(0);
});
// Edge case from the #1661 adversarial review: setting GITNEXUS_MAX_FILE_SIZE
// to the same value as the default (512KB) used to still print the hint
// because the byte comparison resolved to equal. The hint should care
// about whether the operator set the env var, not what value they chose.
it('omits the GITNEXUS_MAX_FILE_SIZE hint when the override equals the default value', async () => {
process.env.GITNEXUS_MAX_FILE_SIZE = '512';
await walkRepositoryPaths(sizeDir);
const hint = cap
.records()
.filter((r) => String(r.msg ?? '').includes('GITNEXUS_MAX_FILE_SIZE=<KB>'));
expect(hint.length).toBe(0);
});
});
describe('large file skip preview cap (#1659)', () => {
let manyDir: string;
const ORIGINAL_ENV = process.env.GITNEXUS_MAX_FILE_SIZE;
const ORIGINAL_VERBOSE = process.env.GITNEXUS_VERBOSE;
let cap: ReturnType<typeof _captureLogger>;
beforeAll(async () => {
manyDir = await fs.mkdtemp(path.join(os.tmpdir(), 'gn-walker-size-many-'));
await fs.mkdir(path.join(manyDir, 'src'), { recursive: true });
// 8 files >512KB so the preview-cap path (5) is exercised.
for (let i = 0; i < 8; i++) {
await fs.writeFile(path.join(manyDir, 'src', `big${i}.ts`), 'x'.repeat(600 * 1024));
}
});
afterAll(async () => {
await fs.rm(manyDir, { recursive: true, force: true });
});
beforeEach(() => {
delete process.env.GITNEXUS_MAX_FILE_SIZE;
delete process.env.GITNEXUS_VERBOSE;
_resetMaxFileSizeWarnings();
cap = _captureLogger();
});
afterEach(() => {
if (ORIGINAL_ENV === undefined) {
delete process.env.GITNEXUS_MAX_FILE_SIZE;
} else {
process.env.GITNEXUS_MAX_FILE_SIZE = ORIGINAL_ENV;
}
if (ORIGINAL_VERBOSE === undefined) {
delete process.env.GITNEXUS_VERBOSE;
} else {
process.env.GITNEXUS_VERBOSE = ORIGINAL_VERBOSE;
}
cap.restore();
});
it('truncates the path list to 5 and mentions GITNEXUS_VERBOSE when over the cap', async () => {
await walkRepositoryPaths(manyDir);
const pathLines = cap.records().filter((r) => /^\s*-\s/.test(String(r.msg ?? '')));
expect(pathLines.length).toBe(5);
const more = cap
.records()
.filter((r) => String(r.msg ?? '').includes('and 3 more (set GITNEXUS_VERBOSE=1'));
expect(more.length).toBe(1);
});
// Boundary check from the #1661 adversarial review: the SKIPPED_PREVIEW_CAP
// comparison is `<=`, so 5 paths should list all five without a truncation
// line and 6 paths should list exactly five plus "...and 1 more". Tested
// explicitly so a future off-by-one refactor (`<=` → `<`) fails fast.
it('lists all paths and omits the truncation line at exactly 5 skipped files', async () => {
const fiveDir = await fs.mkdtemp(path.join(os.tmpdir(), 'gn-walker-size-five-'));
try {
await fs.mkdir(path.join(fiveDir, 'src'), { recursive: true });
for (let i = 0; i < 5; i++) {
await fs.writeFile(path.join(fiveDir, 'src', `big${i}.ts`), 'x'.repeat(600 * 1024));
}
await walkRepositoryPaths(fiveDir);
const pathLines = cap.records().filter((r) => /^\s*-\s/.test(String(r.msg ?? '')));
expect(pathLines.length).toBe(5);
const more = cap.records().filter((r) => String(r.msg ?? '').includes('...and '));
expect(more.length).toBe(0);
} finally {
await fs.rm(fiveDir, { recursive: true, force: true });
}
});
it('lists exactly 5 paths plus "...and 1 more" at exactly 6 skipped files', async () => {
const sixDir = await fs.mkdtemp(path.join(os.tmpdir(), 'gn-walker-size-six-'));
try {
await fs.mkdir(path.join(sixDir, 'src'), { recursive: true });
for (let i = 0; i < 6; i++) {
await fs.writeFile(path.join(sixDir, 'src', `big${i}.ts`), 'x'.repeat(600 * 1024));
}
await walkRepositoryPaths(sixDir);
const pathLines = cap.records().filter((r) => /^\s*-\s/.test(String(r.msg ?? '')));
expect(pathLines.length).toBe(5);
const more = cap
.records()
.filter((r) => String(r.msg ?? '').includes('and 1 more (set GITNEXUS_VERBOSE=1'));
expect(more.length).toBe(1);
} finally {
await fs.rm(sixDir, { recursive: true, force: true });
}
});
it('lists every skipped path when GITNEXUS_VERBOSE=1', async () => {
process.env.GITNEXUS_VERBOSE = '1';
await walkRepositoryPaths(manyDir);
const pathLines = cap.records().filter((r) => /^\s*-\s/.test(String(r.msg ?? '')));
expect(pathLines.length).toBe(8);
const more = cap.records().filter((r) => String(r.msg ?? '').includes('and '));
expect(more.length).toBe(0);
});
// Issue #1659 follow-up (PR #1661 review): paths were pushed in fs.stat
// completion order, so the default preview could vary between runs on
// the same repo. The implementation sorts skippedLargePaths before
// slicing, so the listed paths come out in sorted order, which is the
// stable contract operators can rely on.
it('lists skipped paths in sorted order (deterministic preview)', async () => {
process.env.GITNEXUS_VERBOSE = '1';
await walkRepositoryPaths(manyDir);
const pathLines = cap
.records()
.map((r) => String(r.msg ?? ''))
.filter((m) => /^\s*-\s/.test(m))
.map((m) => m.replace(/^\s*-\s*/, ''));
expect(pathLines).toEqual([...pathLines].sort());
// sanity-check we actually saw all 8 of the manyDir fixture
expect(pathLines).toEqual([
'src/big0.ts',
'src/big1.ts',
'src/big2.ts',
'src/big3.ts',
'src/big4.ts',
'src/big5.ts',
'src/big6.ts',
'src/big7.ts',
]);
});
});
});
+32 -6
View File
@@ -118,6 +118,25 @@ withTestLbugDB(
// Should return 0 rows, not all rows
expect(rows).toHaveLength(0);
});
it('keeps seeded rows unchanged for a no-match parameterized write probe', async () => {
await initLbug('test-repo', handle.dbPath);
try {
const rows = await executeParameterized(
'test-repo',
'MATCH (n:Function) WHERE n.name = $target SET n.name = $name RETURN n.name AS name',
{ target: '__missing__', name: 'x' },
);
expect(rows).toEqual([]);
} catch (err) {
expect(String(err)).toMatch(/read-only database|write operations/i);
}
const rows = await executeQuery(
'test-repo',
'MATCH (n:Function) RETURN n.name AS name ORDER BY n.name',
);
expect(rows.map((r: any) => r.name)).toContain('main');
});
});
// ─── Error handling ──────────────────────────────────────────────────
@@ -133,14 +152,21 @@ withTestLbugDB(
await expect(initLbug('bad-repo', '/nonexistent/path/lbug')).rejects.toThrow();
});
it('read-only mode: write query throws', async () => {
it('keeps seeded data unchanged for a no-match write probe', async () => {
await initLbug('test-repo', handle.dbPath);
await expect(
executeQuery(
try {
await executeQuery(
'test-repo',
"CREATE (n:Function {id: 'new', name: 'new', filePath: '', startLine: 0, endLine: 0, isExported: false, content: '', description: ''})",
),
).rejects.toThrow();
"MATCH (n:Function) WHERE n.name = '__missing__' SET n.name = 'new' RETURN n",
);
} catch (err) {
expect(String(err)).toMatch(/read-only database|write operations/i);
}
const rows = await executeQuery(
'test-repo',
'MATCH (n:Function) RETURN n.name AS name ORDER BY n.name',
);
expect(rows.map((r: any) => r.name)).toContain('main');
});
});
@@ -52,13 +52,16 @@ withTestLbugDB(
expect(result.markdown).toContain('hash');
});
it('cypher tool blocks write queries', async () => {
it('cypher no-match write probe returns read-only error or empty rows', async () => {
const result = await backend.callTool('cypher', {
query:
"CREATE (n:Function {id: 'x', name: 'x', filePath: '', startLine: 0, endLine: 0, isExported: false, content: '', description: ''})",
"MATCH (n:Function) WHERE n.name = '__missing__' SET n.name = 'x' RETURN n.name AS name",
});
expect(result).toHaveProperty('error');
expect(result.error).toMatch(/write operations/i);
if (result?.error) {
expect(result.error).toMatch(/write operations|read-only/i);
return;
}
expect(result).toEqual([]);
});
it('context tool returns symbol info with callers and callees', async () => {
+31 -92
View File
@@ -4,21 +4,19 @@
* Tests tool implementations via direct LadybugDB queries.
* The full LocalBackend.callTool() requires a global registry,
* so here we test the security-critical behaviors directly:
* - Write-operation blocking in cypher
* - Query execution via the pool
* - Parameterized queries preventing injection
* - Read-only enforcement
*
* Covers hardening fixes: #1 (parameterized queries), #2 (write blocking),
* #3 (path traversal), #4 (relation allowlist), #25 (regex lastIndex),
* #26 (rename first-occurrence-only)
* Covers hardening fixes: #1 (parameterized queries), #3 (path traversal),
* #4 (relation allowlist), #26 (rename first-occurrence-only)
*/
import { describe, it, expect } from 'vitest';
import {
CYPHER_WRITE_RE,
initLbug,
closeLbug,
executeQuery,
executeParameterized,
isWriteQuery,
} from '../../src/mcp/core/lbug-adapter.js';
import { VALID_RELATION_TYPES } from '../../src/mcp/local/local-backend.js';
import { withTestLbugDB } from '../helpers/test-indexed-db.js';
@@ -29,35 +27,12 @@ import { LOCAL_BACKEND_SEED_DATA } from '../fixtures/local-backend-seed.js';
withTestLbugDB(
'local-backend',
(handle) => {
// ─── Cypher write blocking ───────────────────────────────────────────
describe('cypher write blocking', () => {
const allWriteKeywords = [
'CREATE',
'DELETE',
'SET',
'MERGE',
'REMOVE',
'DROP',
'ALTER',
'COPY',
'DETACH',
];
for (const keyword of allWriteKeywords) {
it(`blocks ${keyword} query`, () => {
const blocked = isWriteQuery(`MATCH (n) ${keyword} n.name = "x"`);
expect(blocked).toBe(true);
});
}
it('allows valid read queries through the pool', async () => {
const rows = await executeQuery(
handle.repoId,
'MATCH (n:Function) RETURN n.name AS name ORDER BY n.name',
);
expect(rows.length).toBeGreaterThanOrEqual(3);
});
it('allows valid read queries through the pool', async () => {
const rows = await executeQuery(
handle.repoId,
'MATCH (n:Function) RETURN n.name AS name ORDER BY n.name',
);
expect(rows.length).toBeGreaterThanOrEqual(3);
});
// ─── Parameterized queries ───────────────────────────────────────────
@@ -171,34 +146,27 @@ withTestLbugDB(
// ─── Read-only enforcement ───────────────────────────────────────────
describe('read-only database', () => {
it('rejects write operations at DB level', async () => {
await expect(
executeQuery(
handle.repoId,
`CREATE (n:Function {id: 'new', name: 'new', filePath: '', startLine: 0, endLine: 0, isExported: false, content: '', description: ''})`,
),
).rejects.toThrow();
});
});
// ─── Regex lastIndex hardening (#25) ─────────────────────────────────
describe('regex lastIndex (hardening #25)', () => {
it('CYPHER_WRITE_RE is non-global (no sticky lastIndex)', () => {
expect(CYPHER_WRITE_RE.global).toBe(false);
expect(CYPHER_WRITE_RE.sticky).toBe(false);
});
it('works correctly across multiple consecutive calls', () => {
// If the regex were global, lastIndex could cause false results
const results = [
isWriteQuery('CREATE (n)'), // true
isWriteQuery('MATCH (n) RETURN n'), // false
isWriteQuery('DELETE n'), // true
isWriteQuery('MATCH (n) RETURN n'), // false
isWriteQuery('SET n.x = 1'), // true
];
expect(results).toEqual([true, false, true, false, true]);
it('keeps seeded rows unchanged for a no-match write probe', async () => {
const readOnlyRepo = 'local-backend-read-only';
await initLbug(readOnlyRepo, handle.dbPath);
try {
const rows = await executeParameterized(
readOnlyRepo,
`MATCH (n:Function) WHERE n.name = $target SET n.name = $name RETURN n.name AS name`,
{ target: '__missing__', name: 'changed' },
);
expect(rows).toEqual([]);
} catch (err) {
expect(String(err)).toMatch(/Write operations are not allowed|read-only database/i);
}
const rows = await executeParameterized(
readOnlyRepo,
'MATCH (n:Function) WHERE n.name = $name RETURN n.name AS name',
{ name: 'login' },
);
expect(rows).toHaveLength(1);
expect(rows[0].name).toBe('login');
await closeLbug(readOnlyRepo);
});
});
@@ -215,35 +183,6 @@ withTestLbugDB(
});
});
// ─── Write blocking edge cases ──────────────────────────────────────
describe('write blocking edge cases', () => {
it('blocks lowercase write keywords (case-insensitive)', () => {
expect(isWriteQuery('create (n:Function {id: "x"})')).toBe(true);
expect(isWriteQuery('delete n')).toBe(true);
expect(isWriteQuery('set n.name = "x"')).toBe(true);
});
it('blocks write keyword in CREATED-like words (regex is keyword-boundary unaware)', () => {
// CYPHER_WRITE_RE uses \b word boundaries — "CREATED" does NOT match "CREATE"
const result = isWriteQuery("MATCH (n) WHERE n.name = 'CREATED' RETURN n");
// The regex uses word boundaries so substring "CREATE" inside "CREATED" is NOT matched
expect(result).toBe(false);
});
it('blocks multi-line queries with write keywords', () => {
expect(isWriteQuery('MATCH (n)\nDELETE n')).toBe(true);
});
it('returns false for empty string', () => {
expect(isWriteQuery('')).toBe(false);
});
it('returns false for whitespace-only query', () => {
expect(isWriteQuery(' ')).toBe(false);
});
});
// ─── Query error handling via pool ──────────────────────────────────
describe('query error handling via pool', () => {
@@ -1836,6 +1836,64 @@ describe('C++ overload resolution — conversion-rank disambiguation (#1578)', (
});
});
// C++ overload resolution: pointer/nullptr/ellipsis conversion ranks (#1637)
describe('C++ overload resolution — pointer/nullptr/ellipsis ranks (#1637)', () => {
let result: PipelineResult;
beforeAll(async () => {
result = await runPipelineFromRepo(
path.join(FIXTURES, 'cpp-overload-pointer-null-ellipsis'),
() => {},
);
}, 60000);
it('f(nullptr) and f(p) resolve to f(int*) while f(42) resolves to f(bool)', () => {
const calls = getRelationships(result, 'CALLS');
const nullptrCall = calls.find((c) => c.source === 'runNullptr' && c.target === 'f');
const pointerCall = calls.find((c) => c.source === 'runPointer' && c.target === 'f');
const boolCall = calls.find((c) => c.source === 'runBoolConversion' && c.target === 'f');
expect(
result.graph.getNode(nullptrCall?.rel.targetId ?? '')?.properties.parameterTypes,
).toEqual(['int']);
expect(
result.graph.getNode(pointerCall?.rel.targetId ?? '')?.properties.parameterTypes,
).toEqual(['int']);
expect(result.graph.getNode(boolCall?.rel.targetId ?? '')?.properties.parameterTypes).toEqual([
'bool',
]);
});
it('g(1, 2) resolves to fixed-arity g(int, int), not g(int, ...)', () => {
const calls = getRelationships(result, 'CALLS');
const gCalls = calls.filter((c) => c.source === 'run' && c.target === 'g');
expect(gCalls.length).toBe(1);
const tgt = result.graph.getNode(gCalls[0].rel.targetId);
expect(tgt?.properties.parameterTypes).toEqual(['int', 'int']);
});
it("h(1, 'a') resolves to h(int, double), not h(int, ...)", () => {
const calls = getRelationships(result, 'CALLS');
const hCalls = calls.filter((c) => c.source === 'run' && c.target === 'h');
expect(hCalls.length).toBe(1);
const tgt = result.graph.getNode(hCalls[0].rel.targetId);
expect(tgt?.properties.parameterTypes).toEqual(['int', 'double']);
});
it('k(1, 2, 3) keeps the ellipsis overload viable when it is the only match', () => {
const calls = getRelationships(result, 'CALLS');
const kCalls = calls.filter((c) => c.source === 'run' && c.target === 'k');
expect(kCalls.length).toBe(1);
const tgt = result.graph.getNode(kCalls[0].rel.targetId);
expect(tgt?.properties.parameterCount).toBeUndefined();
expect(tgt?.properties.parameterTypes).toEqual(['int']);
});
});
// ---------------------------------------------------------------------------
// U3: anonymous-namespace symbols MUST NOT leak across translation units
// (full-pipeline integration test; unit-level coverage exists separately)
@@ -3156,6 +3214,59 @@ describe('C++ SFINAE filter — C++20 requires-clause shape', () => {
});
});
describe('C++ SFINAE filter — Tier-A type_traits predicates', () => {
async function runFixture(name: string): Promise<PipelineResult> {
return runPipelineFromRepo(path.join(FIXTURES, name), () => {});
}
function callsFromRunToPick(result: PipelineResult) {
return getRelationships(result, 'CALLS').filter(
(c) => c.source === 'run' && c.target === 'pick',
);
}
it('is_pointer_v and is_class_v disambiguate pointer vs class arguments', async () => {
const result = await runFixture('cpp-sfinae-is-pointer');
const calls = callsFromRunToPick(result);
expect(calls.length).toBe(2);
expect(new Set(calls.map((c) => c.rel.targetId)).size).toBe(2);
}, 60000);
it('is_reference_v keeps reference-shaped arguments distinct from values', async () => {
const result = await runFixture('cpp-sfinae-is-reference');
const calls = callsFromRunToPick(result);
expect(calls.length).toBe(2);
expect(new Set(calls.map((c) => c.rel.targetId)).size).toBe(2);
}, 60000);
it('is_class_v rejects primitive arguments while keeping class arguments', async () => {
const result = await runFixture('cpp-sfinae-is-class');
const calls = callsFromRunToPick(result);
expect(calls.length).toBe(2);
expect(new Set(calls.map((c) => c.rel.targetId)).size).toBe(2);
}, 60000);
it('is_enum_v distinguishes known enum declarations from primitives', async () => {
const result = await runFixture('cpp-sfinae-is-enum');
const calls = callsFromRunToPick(result);
expect(calls.length).toBe(2);
expect(new Set(calls.map((c) => c.rel.targetId)).size).toBe(2);
}, 60000);
it('is_const_v and is_volatile_v disambiguate cv-qualified locals', async () => {
const result = await runFixture('cpp-sfinae-is-const-volatile');
const calls = callsFromRunToPick(result);
expect(calls.length).toBe(2);
expect(new Set(calls.map((c) => c.rel.targetId)).size).toBe(2);
}, 60000);
it('is_void_v does not misclassify void pointers as void values', async () => {
const result = await runFixture('cpp-sfinae-is-void');
const calls = callsFromRunToPick(result);
expect(calls.length).toBe(1);
}, 60000);
});
describe('C++ SFINAE filter — unknown predicate keeps both candidates (monotonicity contract)', () => {
let result: PipelineResult;
@@ -34,6 +34,13 @@ const LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES: Readonly<Record<string, Readonly
// which is only available in the registry-primary path.
'resolves user.Save() to the method whose receiver type is declared in another package file',
]),
java: new Set([
// Duplicate-FQN same-module path-affinity ordering is implemented in the
// Java provider hook for the scope-resolution path. Legacy DAG parity runs
// still use legacy owner/type resolution behavior and can bind cross-module.
'resolves Module1App.run calls to module1 UserService, not module2',
'resolves Module2App.run calls to module2 UserService, not module1',
]),
php: new Set([
// Arity-narrowing in `pickUniqueGlobalCallable` rejects free-call
// candidates that are definitively below required-parameter-count. The
@@ -189,6 +196,12 @@ const LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES: Readonly<Record<string, Readonly
// Multi-arg incomparable overloads: pairwise dominance check finds
// neither h(int,int) nor h(double,double) dominates. Scope-resolver-only.
'h(42, 2.5) emits zero CALLS edges — incomparable multi-arg overloads, ambiguous',
// Pointer/nullptr/ellipsis conversion ranks (#1637) need C++ type-class
// sidecars plus conversion-rank scoring. The legacy DAG has neither.
'f(nullptr) and f(p) resolve to f(int*) while f(42) resolves to f(bool)',
'g(1, 2) resolves to fixed-arity g(int, int), not g(int, ...)',
"h(1, 'a') resolves to h(int, double), not h(int, ...)",
'k(1, 2, 3) keeps the ellipsis overload viable when it is the only match',
// The legacy DAG path lacks the SFINAE / `requires`-clause aware
// overload filter (issue #1579). The two `process<T>` overloads
// guarded by mutually-exclusive `enable_if_t` predicates collapse
@@ -200,6 +213,12 @@ const LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES: Readonly<Record<string, Readonly
'enable_if_t<is_integral_v<T>> overload binds only on integral call sites',
'enable_if_t<is_floating_point_v<T>> overload binds only on floating call sites',
'requires-clause overloads disambiguate same as enable_if_t (F4 AST shape)',
'is_pointer_v and is_class_v disambiguate pointer vs class arguments',
'is_reference_v keeps reference-shaped arguments distinct from values',
'is_class_v rejects primitive arguments while keeping class arguments',
'is_enum_v distinguishes known enum declarations from primitives',
'is_const_v and is_volatile_v disambiguate cv-qualified locals',
'is_void_v does not misclassify void pointers as void values',
// The legacy DAG path has no inline-namespace same-name ambiguity
// detection. When two inline children declare the same name, the
// legacy path picks an arbitrary match. The scope-resolver returns
@@ -174,6 +174,72 @@ describe('Java call resolution with arity filtering', () => {
});
});
describe('Java same-module priority for duplicate FQNs', () => {
let result: PipelineResult;
beforeAll(async () => {
result = await runPipelineFromRepo(path.join(FIXTURES, 'java-duplicate-fqn-modules'), () => {});
}, 60000);
it('resolves Module1App.run calls to module1 UserService, not module2', () => {
const calls = getRelationships(result, 'CALLS');
const module1ToModule1 = calls.filter(
(c) =>
c.source === 'run' &&
c.target === 'UserService' &&
c.sourceFilePath === 'module1/src/main/java/com/example/Module1App.java' &&
c.targetFilePath === 'module1/src/main/java/com/example/UserService.java',
);
const module1ToModule2 = calls.filter(
(c) =>
c.source === 'run' &&
c.target === 'UserService' &&
c.sourceFilePath === 'module1/src/main/java/com/example/Module1App.java' &&
c.targetFilePath === 'module2/src/main/java/com/example/UserService.java',
);
const module1ToAnyUserService = calls.filter(
(c) =>
c.source === 'run' &&
c.target === 'UserService' &&
c.sourceFilePath === 'module1/src/main/java/com/example/Module1App.java' &&
/module[12]\/src\/main\/java\/com\/example\/UserService\.java/.test(c.targetFilePath),
);
expect(module1ToModule1.length).toBe(1);
expect(module1ToModule2.length).toBe(0);
expect(module1ToAnyUserService.length).toBe(1);
});
it('resolves Module2App.run calls to module2 UserService, not module1', () => {
const calls = getRelationships(result, 'CALLS');
const module2ToModule2 = calls.filter(
(c) =>
c.source === 'run' &&
c.target === 'UserService' &&
c.sourceFilePath === 'module2/src/main/java/com/example/Module2App.java' &&
c.targetFilePath === 'module2/src/main/java/com/example/UserService.java',
);
const module2ToModule1 = calls.filter(
(c) =>
c.source === 'run' &&
c.target === 'UserService' &&
c.sourceFilePath === 'module2/src/main/java/com/example/Module2App.java' &&
c.targetFilePath === 'module1/src/main/java/com/example/UserService.java',
);
const module2ToAnyUserService = calls.filter(
(c) =>
c.source === 'run' &&
c.target === 'UserService' &&
c.sourceFilePath === 'module2/src/main/java/com/example/Module2App.java' &&
/module[12]\/src\/main\/java\/com\/example\/UserService\.java/.test(c.targetFilePath),
);
expect(module2ToModule2.length).toBe(1);
expect(module2ToModule1.length).toBe(0);
expect(module2ToAnyUserService.length).toBe(1);
});
});
// ---------------------------------------------------------------------------
// Member-call resolution: obj.method() resolves through pipeline
// ---------------------------------------------------------------------------
@@ -2683,6 +2683,44 @@ describe('TypeScript Child extends Parent — inherited method resolution (SM-9)
});
});
// ---------------------------------------------------------------------------
// PR #1657 finding #6: ambient base class — Step 2 MRO ancestor whose body
// is never parsed (declare class). Probes whether the owner-keyed lookup
// can still resolve inherited members on owners that reconcile-ownership
// skipped because they have no parsed body.
// ---------------------------------------------------------------------------
describe('TypeScript Derived extends declare class AmbientBase — ambient MRO ancestor', () => {
let result: PipelineResult;
beforeAll(async () => {
result = await runPipelineFromRepo(
path.join(FIXTURES, 'typescript-ambient-base-class'),
() => {},
);
}, 60000);
it('detects AmbientBase and Derived classes', () => {
const classes = getNodesByLabel(result, 'Class');
expect(classes).toContain('AmbientBase');
expect(classes).toContain('Derived');
});
it('emits EXTENDS edge: Derived → AmbientBase', () => {
const extends_ = getRelationships(result, 'EXTENDS');
expect(edgeSet(extends_)).toContain('Derived → AmbientBase');
});
it('resolves d.ambientMethod() to AmbientBase.ambientMethod via MRO walk', () => {
const calls = getRelationships(result, 'CALLS');
const ambientCall = calls.find(
(c) => c.target === 'ambientMethod' && c.targetFilePath.includes('ambient.ts'),
);
expect(ambientCall).toBeDefined();
expect(ambientCall!.source).toBe('run');
});
});
// ---------------------------------------------------------------------------
// PR #1050: tsconfig path alias resolution under registry-primary path
// (Adversarial review Finding 1 — `@/services/user` must resolve via tsconfig
@@ -101,6 +101,15 @@ withTestLbugDB(
expect(Array.isArray(results)).toBe(true);
});
it('does not treat write-like words inside search text as write operations (#1608)', async () => {
const { results, ftsAvailable } = await searchFTSFromLbug(
'create user authentication delete',
10,
);
expect(ftsAvailable).toBe(true);
expect(results.length).toBeGreaterThan(0);
});
it('handles limit of 0', async () => {
const { results } = await searchFTSFromLbug('user authentication', 0);
expect(results).toEqual([]);
@@ -0,0 +1,200 @@
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';
const execFileSyncMock = vi.fn();
const getHeapStatisticsMock = vi.fn();
vi.mock('child_process', async () => {
const actual = await vi.importActual<typeof import('child_process')>('child_process');
return { ...actual, execFileSync: execFileSyncMock };
});
vi.mock('v8', () => ({
default: {
getHeapStatistics: getHeapStatisticsMock,
},
}));
vi.mock('../../src/core/lbug/lbug-adapter.js', () => ({
closeLbug: vi.fn(async () => undefined),
}));
describe('analyzeCommand heap respawn', () => {
let initialNodeOptions: string | undefined;
beforeEach(() => {
initialNodeOptions = process.env.NODE_OPTIONS;
vi.resetModules();
execFileSyncMock.mockReset();
getHeapStatisticsMock.mockReset();
process.exitCode = undefined;
});
afterEach(() => {
if (initialNodeOptions === undefined) delete process.env.NODE_OPTIONS;
else process.env.NODE_OPTIONS = initialNodeOptions;
});
it('re-execs analyze with 16GB heap when no max-old-space-size is present', async () => {
delete process.env.NODE_OPTIONS;
getHeapStatisticsMock.mockReturnValue({ heap_size_limit: 512 * 1024 * 1024 });
const { analyzeCommand } = await import('../../src/cli/analyze.js');
await analyzeCommand(undefined, {});
expect(execFileSyncMock).toHaveBeenCalledTimes(1);
const [, args, opts] = execFileSyncMock.mock.calls[0];
expect(args).toContain('--max-old-space-size=16384');
expect(opts.env.NODE_OPTIONS).toContain('--max-old-space-size=16384');
});
it('does not re-exec when NODE_OPTIONS already defines max-old-space-size', async () => {
process.env.NODE_OPTIONS = '--max-old-space-size=32768';
getHeapStatisticsMock.mockReturnValue({ heap_size_limit: 512 * 1024 * 1024 });
const { analyzeCommand } = await import('../../src/cli/analyze.js');
await analyzeCommand('/__gitnexus_nonexistent__', {});
expect(execFileSyncMock).not.toHaveBeenCalled();
});
it('prints heap guidance when respawned analyze exits with likely OOM', async () => {
delete process.env.NODE_OPTIONS;
getHeapStatisticsMock.mockReturnValue({ heap_size_limit: 512 * 1024 * 1024 });
execFileSyncMock.mockImplementationOnce(() => {
const err = new Error('child failed') as Error & { status?: number; signal?: string };
err.status = undefined;
err.signal = 'SIGABRT';
throw err;
});
const { _captureLogger } = await import('../../src/core/logger.js');
const cap = _captureLogger();
const { analyzeCommand } = await import('../../src/cli/analyze.js');
await analyzeCommand(undefined, {});
// Signal-only child failures do not carry a numeric status, so the CLI
// falls back to exit code 1.
expect(process.exitCode).toBe(1);
const oomGuidance = cap
.records()
.find((r) => r.msg.includes('Analysis likely ran out of memory.'));
expect(oomGuidance).toBeDefined();
const msg = oomGuidance?.msg ?? '';
expect(msg).toContain('NODE_OPTIONS="--max-old-space-size=24576"');
expect(msg).toContain('[your-args]');
expect(msg).toContain('native crash unrelated to heap size');
cap.restore();
});
it('prints heap guidance when child stderr contains heap OOM signature', async () => {
delete process.env.NODE_OPTIONS;
getHeapStatisticsMock.mockReturnValue({ heap_size_limit: 512 * 1024 * 1024 });
execFileSyncMock.mockImplementationOnce(() => {
const err = new Error('Command failed') as Error & {
status?: number;
signal?: string;
stderr?: Buffer;
};
err.status = 1;
err.signal = undefined;
err.stderr = Buffer.from(
'FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory',
);
throw err;
});
const { _captureLogger } = await import('../../src/core/logger.js');
const cap = _captureLogger();
const { analyzeCommand } = await import('../../src/cli/analyze.js');
await analyzeCommand(undefined, {});
expect(process.exitCode).toBe(1);
expect(cap.records().some((r) => r.msg.includes('Analysis likely ran out of memory.'))).toBe(
true,
);
cap.restore();
});
it('prints heap guidance when child stdout contains heap OOM signature', async () => {
delete process.env.NODE_OPTIONS;
getHeapStatisticsMock.mockReturnValue({ heap_size_limit: 512 * 1024 * 1024 });
execFileSyncMock.mockImplementationOnce(() => {
const err = new Error('Command failed') as Error & {
status?: number;
signal?: string;
stdout?: string;
};
err.status = 1;
err.signal = undefined;
err.stdout = 'FATAL ERROR: JavaScript heap out of memory';
throw err;
});
const { _captureLogger } = await import('../../src/core/logger.js');
const cap = _captureLogger();
const { analyzeCommand } = await import('../../src/cli/analyze.js');
await analyzeCommand(undefined, {});
expect(process.exitCode).toBe(1);
expect(cap.records().some((r) => r.msg.includes('Analysis likely ran out of memory.'))).toBe(
true,
);
cap.restore();
});
it('prints heap guidance when child exits 134 without output', async () => {
delete process.env.NODE_OPTIONS;
getHeapStatisticsMock.mockReturnValue({ heap_size_limit: 512 * 1024 * 1024 });
execFileSyncMock.mockImplementationOnce(() => {
const err = new Error('Command failed') as Error & {
status?: number;
signal?: string;
stderr?: string;
stdout?: string;
};
err.status = 134;
err.signal = undefined;
err.stderr = '';
err.stdout = '';
throw err;
});
const { _captureLogger } = await import('../../src/core/logger.js');
const cap = _captureLogger();
const { analyzeCommand } = await import('../../src/cli/analyze.js');
await analyzeCommand(undefined, {});
expect(process.exitCode).toBe(134);
expect(cap.records().some((r) => r.msg.includes('Analysis likely ran out of memory.'))).toBe(
true,
);
cap.restore();
});
it('does not print heap guidance for non-OOM child failures with output', async () => {
delete process.env.NODE_OPTIONS;
getHeapStatisticsMock.mockReturnValue({ heap_size_limit: 512 * 1024 * 1024 });
execFileSyncMock.mockImplementationOnce(() => {
const err = new Error('Command failed') as Error & {
status?: number;
signal?: string;
stderr?: Buffer;
};
err.status = 2;
err.signal = undefined;
err.stderr = Buffer.from('parser failed: invalid token');
throw err;
});
const { _captureLogger } = await import('../../src/core/logger.js');
const cap = _captureLogger();
const { analyzeCommand } = await import('../../src/cli/analyze.js');
await analyzeCommand(undefined, {});
expect(process.exitCode).toBe(2);
expect(cap.records().some((r) => r.msg.includes('Analysis likely ran out of memory.'))).toBe(
false,
);
cap.restore();
});
});
@@ -1,16 +1,21 @@
import { beforeEach, describe, expect, it, vi } from 'vitest';
const { runFullAnalysisMock, generateAIContextFilesMock, generateSkillFilesMock } = vi.hoisted(
() => {
const { runFullAnalysisMock, generateAIContextFilesMock, generateSkillFilesMock, cliErrorMock } =
vi.hoisted(() => {
const runFullAnalysisMock = vi.fn();
const generateAIContextFilesMock = vi.fn(async () => ({ files: [] as string[] }));
const generateSkillFilesMock = vi.fn(async () => ({
skills: [{ name: 'c', label: 'Community', symbolCount: 1, fileCount: 1 }],
outputPath: '/repo/.claude/skills/generated',
}));
return { runFullAnalysisMock, generateAIContextFilesMock, generateSkillFilesMock };
},
);
const cliErrorMock = vi.fn();
return {
runFullAnalysisMock,
generateAIContextFilesMock,
generateSkillFilesMock,
cliErrorMock,
};
});
vi.mock('../../src/core/run-analyze.js', () => ({
runFullAnalysis: runFullAnalysisMock,
@@ -24,6 +29,10 @@ vi.mock('../../src/cli/skill-gen.js', () => ({
generateSkillFiles: generateSkillFilesMock,
}));
vi.mock('../../src/cli/cli-message.js', () => ({
cliError: cliErrorMock,
}));
vi.mock('../../src/core/lbug/lbug-adapter.js', () => ({
closeLbug: vi.fn(async () => undefined),
}));
@@ -62,6 +71,7 @@ describe('analyzeCommand commander → runFullAnalysis noStats bridge (#1477)',
skills: [{ name: 'c', label: 'Community', symbolCount: 1, fileCount: 1 }],
outputPath: '/repo/.claude/skills/generated',
});
cliErrorMock.mockReset();
process.exitCode = undefined;
process.env.NODE_OPTIONS = `${process.env.NODE_OPTIONS ?? ''} --max-old-space-size=8192`.trim();
});
@@ -104,6 +114,27 @@ describe('analyzeCommand commander → runFullAnalysis noStats bridge (#1477)',
expect(opts.skipAgentsMd).toBe(true);
});
it('passes --repair-fts through to runFullAnalysis', async () => {
const { analyzeCommand } = await import('../../src/cli/analyze.js');
await analyzeCommand(undefined, { repairFts: true });
const opts = runFullAnalysisMock.mock.calls[0][1];
expect(opts.repairFts).toBe(true);
});
it('rejects combining --repair-fts with --force', async () => {
const { analyzeCommand } = await import('../../src/cli/analyze.js');
await analyzeCommand(undefined, { repairFts: true, force: true });
expect(process.exitCode).toBe(1);
expect(cliErrorMock).toHaveBeenCalledWith(
expect.stringMatching(/cannot combine `--repair-fts` with `--force`/i),
);
expect(runFullAnalysisMock).not.toHaveBeenCalled();
});
it('passes stats:false as noStats to generateAIContextFiles on the --skills regeneration path (#1477)', async () => {
runFullAnalysisMock.mockResolvedValueOnce({
repoName: 'repo',
@@ -0,0 +1,30 @@
import { describe, expect, it } from 'vitest';
import fs from 'node:fs/promises';
import path from 'node:path';
describe('api query read-only wiring', () => {
it('uses withLbugDb readOnly mode inside handleQueryRequest', async () => {
const source = await fs.readFile(
path.join(__dirname, '..', '..', 'src', 'server', 'api.ts'),
'utf-8',
);
expect(source).toMatch(/handleQueryRequest[\s\S]*withLbugDb\([\s\S]*readOnly:\s*true/);
});
it('routes /api/query through handleQueryRequest', async () => {
const source = await fs.readFile(
path.join(__dirname, '..', '..', 'src', 'server', 'api.ts'),
'utf-8',
);
expect(source).toContain("app.post('/api/query', async (req, res) => {");
expect(source).toContain('await handleQueryRequest(req, res, resolveRepo);');
});
it('opens Ladybug connection with readOnly option when requested', async () => {
const source = await fs.readFile(
path.join(__dirname, '..', '..', 'src', 'core', 'lbug', 'lbug-adapter.ts'),
'utf-8',
);
expect(source).toMatch(/openLbugConnection\(lbug,\s*dbPath,\s*\{\s*readOnly:\s*true\s*\}\)/);
});
});
@@ -0,0 +1,60 @@
import { describe, expect, it } from 'vitest';
import fs from 'node:fs/promises';
import path from 'node:path';
/**
* Regression guard for issue: "Cannot open file ... lbug.shadow - Error 2"
*
* Read-only HTTP endpoints (graph, search, grep) must open the LadybugDB with
* `{ readOnly: true }` so the engine never engages the checkpoint machinery
* (`.shadow` sidecar). Write-mode opens for read-only operations were the
* trigger for the Windows-only "Cannot open file ... lbug.shadow" failures
* observed in E2E runs.
*
* If you add another read-only endpoint and forget the option, this file
* fails — keeping the contract explicit at the static-analysis layer.
*
* Companion: api-query-readonly-wiring.test.ts (covers /api/query).
* Precedent: PR #1655 set the pattern for /api/query.
*/
describe('api read-only endpoint wiring', () => {
const readSource = () =>
fs.readFile(path.join(__dirname, '..', '..', 'src', 'server', 'api.ts'), 'utf-8');
it('/api/graph stream path opens read-only', async () => {
const source = await readSource();
expect(source).toMatch(
/streamGraphNdjson\(res, includeContent, abortController\.signal\)[\s\S]{0,200}readOnly:\s*true/,
);
});
it('/api/graph non-stream path opens read-only', async () => {
const source = await readSource();
expect(source).toMatch(/buildGraph\(includeContent\)[\s\S]{0,80}readOnly:\s*true/);
});
it('/api/search opens read-only', async () => {
const source = await readSource();
// The /api/search handler ends its withLbugDb callback with
// `return { searchResults: enriched, ftsAvailable };` immediately before
// the closing brace + options object. Match that suffix to confirm the
// search call site, not /api/query.
expect(source).toMatch(/searchResults: enriched, ftsAvailable[\s\S]{0,80}readOnly:\s*true/);
});
it('/api/grep opens read-only', async () => {
const source = await readSource();
expect(source).toMatch(/MATCH \(n:File\)[\s\S]{0,300}readOnly:\s*true/);
});
it('/api/embed remains write-mode (writes embeddings — must not be flipped to readOnly)', async () => {
const source = await readSource();
// Negative assertion: no `readOnly: true` between the embed job's
// `runEmbeddingPipeline` call site and its withLbugDb open. Embed writes
// back vector rows; flipping this to readOnly would silently break it.
const embedSection = source.match(/runEmbeddingPipeline[\s\S]{0,400}\}\s*\)\s*;[\s\S]{0,200}/);
if (embedSection) {
expect(embedSection[0]).not.toMatch(/readOnly:\s*true/);
}
});
});
+62 -15
View File
@@ -13,9 +13,10 @@ vi.mock('../../src/core/lbug/lbug-adapter.js', async (importOriginal) => {
// Pool adapter is dynamically imported by the MCP-pool path of
// `searchFTSFromLbug`. We mock it so we can drive the executor without
// spinning up a real LadybugDB pool.
const mockExecuteQuery = vi.fn();
const mockExecuteParameterized = vi.fn();
vi.mock('../../src/core/lbug/pool-adapter.js', () => ({
executeQuery: (repoId: string, cypher: string) => mockExecuteQuery(repoId, cypher),
executeParameterized: (repoId: string, cypher: string, params: Record<string, any>) =>
mockExecuteParameterized(repoId, cypher, params),
addPoolCloseListener: vi.fn(),
}));
@@ -39,6 +40,31 @@ describe('BM25 search', () => {
['Interface', 'interface_fts', ['name', 'content']],
]);
});
it('verifies all configured FTS indexes are queryable', async () => {
const executeQuery = vi.fn().mockResolvedValue([]);
const { verifySearchFTSIndexes } = await import('../../src/core/search/fts-indexes.js');
const missing = await verifySearchFTSIndexes(executeQuery);
expect(missing).toEqual([]);
expect(executeQuery).toHaveBeenCalledTimes(5);
});
it('reports missing indexes when an FTS probe fails', async () => {
const executeQuery = vi
.fn()
.mockResolvedValueOnce([])
.mockRejectedValueOnce(new Error('index does not exist'))
.mockResolvedValueOnce([])
.mockResolvedValueOnce([])
.mockResolvedValueOnce([]);
const { verifySearchFTSIndexes } = await import('../../src/core/search/fts-indexes.js');
const missing = await verifySearchFTSIndexes(executeQuery);
expect(missing).toEqual(['Function.function_fts']);
});
});
describe('searchFTSFromLbug', () => {
@@ -209,20 +235,22 @@ describe('BM25 search', () => {
const REPO = 'test-repo-readonly-fts';
beforeEach(() => {
mockExecuteQuery.mockReset();
mockExecuteParameterized.mockReset();
});
it('queries existing FTS indexes without issuing CREATE_FTS_INDEX', async () => {
mockExecuteQuery.mockImplementation(async (_repo: string, cypher: string) => {
if (cypher.includes('CREATE_FTS_INDEX')) {
throw new Error('query path must stay read-only');
}
mockExecuteParameterized.mockImplementation(
async (_repo: string, cypher: string, params: Record<string, any>) => {
if (cypher.includes('CREATE_FTS_INDEX')) {
throw new Error('query path must stay read-only');
}
if (cypher.includes("QUERY_FTS_INDEX('Function'")) {
return [{ node: { filePath: 'src/auth.ts', id: 'func:login' }, score: 8 }];
}
return [];
});
if (params.query === 'login' && cypher.includes("QUERY_FTS_INDEX('Function'")) {
return [{ node: { filePath: 'src/auth.ts', id: 'func:login' }, score: 8 }];
}
return [];
},
);
const { results } = await searchFTSFromLbug('login', 5, REPO);
@@ -230,16 +258,35 @@ describe('BM25 search', () => {
{ filePath: 'src/auth.ts', score: 8, rank: 1, nodeIds: ['func:login'] },
]);
expect(
mockExecuteQuery.mock.calls.some((c) => String(c[1]).includes('CREATE_FTS_INDEX')),
mockExecuteParameterized.mock.calls.some((c) => String(c[1]).includes('CREATE_FTS_INDEX')),
).toBe(false);
});
it('binds FTS user query text as a parameter in pool mode', async () => {
mockExecuteParameterized.mockResolvedValue([]);
const userQuery = "BrowserWindow create delete set remove 'main' window";
await searchFTSFromLbug(userQuery, 5, REPO);
expect(mockExecuteParameterized).toHaveBeenCalled();
for (const call of mockExecuteParameterized.mock.calls) {
const cypher = String(call[1]);
expect(cypher).toContain('$query');
expect(cypher).not.toContain(userQuery);
expect(cypher.toUpperCase()).not.toMatch(/\bCREATE\b/);
expect(cypher.toUpperCase()).not.toMatch(/\bDELETE\b/);
expect(cypher.toUpperCase()).not.toMatch(/\bSET\b/);
expect(cypher.toUpperCase()).not.toMatch(/\bREMOVE\b/);
expect(call[2]).toEqual({ query: userQuery });
}
});
it('uses the configured FTS query set on every call', async () => {
mockExecuteQuery.mockResolvedValue([]);
mockExecuteParameterized.mockResolvedValue([]);
await searchFTSFromLbug('anything', 5, REPO);
const queryCalls = mockExecuteQuery.mock.calls.filter((c) =>
const queryCalls = mockExecuteParameterized.mock.calls.filter((c) =>
String(c[1]).includes('QUERY_FTS_INDEX'),
);
expect(queryCalls.map((c) => String(c[1]).match(/QUERY_FTS_INDEX\('([^']+)'/)?.[1])).toEqual([
+8 -6
View File
@@ -203,7 +203,7 @@ describe('LocalBackend.callTool', () => {
const result = await backend.callTool('query', { query: 'ProcessActivity' });
expect(result).toHaveProperty('warning');
expect((result as any).warning).toMatch(/gitnexus analyze --force/);
expect((result as any).warning).toMatch(/gitnexus analyze --repair-fts/);
});
it('does not include warning when ftsAvailable is true with zero results', async () => {
@@ -292,13 +292,14 @@ describe('LocalBackend.callTool', () => {
});
it('dispatches cypher tool and blocks write queries', async () => {
(executeParameterized as any).mockRejectedValueOnce(new Error('read-only database'));
const result = await backend.callTool('cypher', { query: 'CREATE (n:Test)' });
expect(result).toHaveProperty('error');
expect(result.error).toContain('Write operations');
});
it('dispatches cypher tool with valid read query', async () => {
(executeQuery as any).mockResolvedValue([{ name: 'test', filePath: 'src/test.ts' }]);
(executeParameterized as any).mockResolvedValue([{ name: 'test', filePath: 'src/test.ts' }]);
const result = await backend.callTool('cypher', {
query: 'MATCH (n:Function) RETURN n.name AS name, n.filePath AS filePath LIMIT 5',
});
@@ -999,6 +1000,7 @@ describe('callTool cypher write blocking', () => {
for (const query of writeQueries) {
it(`blocks write query: ${query.slice(0, 30)}...`, async () => {
(executeParameterized as any).mockRejectedValueOnce(new Error('read-only database'));
const result = await backend.callTool('cypher', { query });
expect(result).toHaveProperty('error');
expect(result.error).toContain('Write operations');
@@ -1006,7 +1008,7 @@ describe('callTool cypher write blocking', () => {
}
it('allows read query through callTool', async () => {
(executeQuery as any).mockResolvedValue([]);
(executeParameterized as any).mockResolvedValue([]);
const result = await backend.callTool('cypher', {
query: 'MATCH (n:Function) RETURN n.name LIMIT 5',
});
@@ -1105,7 +1107,7 @@ describe('cypher result formatting', () => {
});
it('formats tabular results as markdown table', async () => {
(executeQuery as any).mockResolvedValue([
(executeParameterized as any).mockResolvedValue([
{ name: 'main', filePath: 'src/index.ts' },
{ name: 'helper', filePath: 'src/utils.ts' },
]);
@@ -1119,7 +1121,7 @@ describe('cypher result formatting', () => {
});
it('returns empty array as-is', async () => {
(executeQuery as any).mockResolvedValue([]);
(executeParameterized as any).mockResolvedValue([]);
const result = await backend.callTool('cypher', {
query: 'MATCH (n:Function) RETURN n.name LIMIT 0',
});
@@ -1127,7 +1129,7 @@ describe('cypher result formatting', () => {
});
it('returns error object when cypher fails', async () => {
(executeQuery as any).mockRejectedValue(new Error('Syntax error'));
(executeParameterized as any).mockRejectedValue(new Error('Syntax error'));
const result = await backend.callTool('cypher', {
query: 'INVALID CYPHER SYNTAX',
});

Some files were not shown because too many files have changed in this diff Show More