Compare commits
32
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
636c4046cb | ||
|
|
1b5c6e5b6a | ||
|
|
c34c36036f | ||
|
|
aa8f4d6efe | ||
|
|
4d2ed0e525 | ||
|
|
df1882d36b | ||
|
|
f350ae278a | ||
|
|
dae70a26ea | ||
|
|
b4a2a4b91e | ||
|
|
d7e1815aa3 | ||
|
|
92ad0f5491 | ||
|
|
803f0bed5f | ||
|
|
55f8d442f6 | ||
|
|
18167400c4 | ||
|
|
dad1ca7ab5 | ||
|
|
c746f30c90 | ||
|
|
15a667ae5e | ||
|
|
6210d80f1e | ||
|
|
637cfca39c | ||
|
|
73543a4714 | ||
|
|
b37974fdac | ||
|
|
ade2069633 | ||
|
|
5f0c0eba0e | ||
|
|
2632bcccc0 | ||
|
|
c9199b654f | ||
|
|
33f18ceaa2 | ||
|
|
c30833fad3 | ||
|
|
7d500390b9 | ||
|
|
bdc0439a10 | ||
|
|
105efd0f7c | ||
|
|
493827222d | ||
|
|
ed50a6729f |
@@ -17,11 +17,11 @@ npx gitnexus analyze
|
||||
|
||||
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates CLAUDE.md / AGENTS.md context files.
|
||||
|
||||
| Flag | Effect |
|
||||
| ------------------- | ------------------------------------------------------------------------------------------------------- |
|
||||
| `--force` | Force full re-index even if up to date |
|
||||
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
|
||||
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
|
||||
| Flag | Effect |
|
||||
| -------------- | ---------------------------------------------------------------- |
|
||||
| `--force` | Force full re-index even if up to date |
|
||||
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
|
||||
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
|
||||
|
||||
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
|
||||
|
||||
|
||||
@@ -62,131 +62,64 @@ Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.
|
||||
<!-- gitnexus:start -->
|
||||
# GitNexus — Code Intelligence
|
||||
|
||||
Indexed as **GitNexus** (4325 symbols, 10556 relationships, 300 execution flows). Use MCP tools to understand code, assess impact, and navigate safely.
|
||||
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||
|
||||
> If any tool warns the index is stale, run `npx gitnexus analyze` first.
|
||||
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
||||
|
||||
## Always Do
|
||||
|
||||
- **MUST run impact analysis before editing any symbol.** `gitnexus_impact({target: "symbolName", direction: "upstream"})` — report blast radius to the user.
|
||||
- **MUST run `gitnexus_detect_changes()` before committing** — verify only expected symbols and flows are affected.
|
||||
- **MUST warn the user** if impact returns HIGH or CRITICAL risk.
|
||||
- Explore unfamiliar code with `gitnexus_query({query: "concept"})` (process-grouped, ranked) instead of grepping.
|
||||
- Full context on a symbol: `gitnexus_context({name: "symbolName"})`.
|
||||
|
||||
## When Debugging
|
||||
|
||||
1. `gitnexus_query({query: "<error or symptom>"})` — find related execution flows
|
||||
2. `gitnexus_context({name: "<suspect function>"})` — callers, callees, process participation
|
||||
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace flow step by step
|
||||
4. Regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})`
|
||||
|
||||
## When Refactoring
|
||||
|
||||
- **Rename:** `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Graph edits are safe; text_search edits need manual review.
|
||||
- **Extract/Split:** `gitnexus_context` (incoming/outgoing refs) then `gitnexus_impact` (upstream callers) before moving code.
|
||||
- **After any refactor:** `gitnexus_detect_changes({scope: "all"})` to verify scope.
|
||||
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
|
||||
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
|
||||
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
|
||||
|
||||
## Never Do
|
||||
|
||||
- Edit a symbol without running `gitnexus_impact` first.
|
||||
- Ignore HIGH/CRITICAL risk warnings.
|
||||
- Rename with find-and-replace — use `gitnexus_rename`.
|
||||
- Commit without `gitnexus_detect_changes()`.
|
||||
- Add language-specific behavior to shared ingestion code (`gitnexus/src/core/ingestion/`) — use a `LanguageProvider` hook. Seeing `provider.mroStrategy === 'xxx'` or an import from `languages/xxx.ts` in shared code means stop and add a hook.
|
||||
|
||||
## Tools Quick Reference
|
||||
|
||||
| Tool | When to use | Example |
|
||||
|------|-------------|---------|
|
||||
| `list_repos` | Discover indexed repos | `gitnexus_list_repos({})` |
|
||||
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
|
||||
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
|
||||
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
|
||||
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
|
||||
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
|
||||
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
|
||||
| `api_impact` | Pre-change API route impact | `gitnexus_api_impact({route: "/api/users", method: "GET"})` |
|
||||
| `route_map` | Route → handler → consumer map | `gitnexus_route_map({})` |
|
||||
| `tool_map` | MCP/RPC tool definitions | `gitnexus_tool_map({})` |
|
||||
| `shape_check` | Response shape vs consumer access | `gitnexus_shape_check({route: "/api/users"})` |
|
||||
| `group_list` | List repo groups | `gitnexus_group_list({})` |
|
||||
| `group_sync` | Rebuild group Contract Registry | `gitnexus_group_sync({name: "myGroup"})` |
|
||||
| `query` (group mode) | Cross-repo search in a group (RRF-merged) | `gitnexus_query({repo: "@myGroup", query: "auth"})` |
|
||||
| `context` (group mode) | 360° view across all member repos | `gitnexus_context({repo: "@myGroup", name: "validateUser"})` |
|
||||
| `impact` (group mode) | Cross-repo blast radius via Contract Bridge | `gitnexus_impact({repo: "@myGroup", target: "X", direction: "upstream"})` |
|
||||
|
||||
> Group mode: pass `repo: "@<groupName>"` to fan out across all member repos, or `repo: "@<groupName>/<memberPath>"` to target a single member (path keys from `group.yaml`). Optional `service: "<monorepo/path>"` filters by service root. Group-level state (contracts, staleness) lives in the resources table below — there are **no** `group_query` / `group_context` / `group_impact` / `group_contracts` / `group_status` MCP tools.
|
||||
>
|
||||
> For a full walkthrough of setting up a group across multiple repos that communicate over gRPC, see [docs/guides/microservices-grpc.md](docs/guides/microservices-grpc.md).
|
||||
|
||||
## Impact Risk Levels
|
||||
|
||||
| Depth | Meaning | Action |
|
||||
|-------|---------|--------|
|
||||
| d=1 | WILL BREAK — direct callers/importers | MUST update |
|
||||
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
|
||||
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
|
||||
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
|
||||
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
|
||||
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
|
||||
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
|
||||
|
||||
## Resources
|
||||
|
||||
| Resource | Use for |
|
||||
|----------|---------|
|
||||
| `gitnexus://repo/GitNexus/context` | Codebase overview, index freshness |
|
||||
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
|
||||
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
|
||||
| `gitnexus://repo/GitNexus/processes` | All execution flows |
|
||||
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
|
||||
| `gitnexus://group/{name}/contracts` | Group Contract Registry (provider/consumer rows + cross-links) |
|
||||
| `gitnexus://group/{name}/status` | Per-member index + Contract Registry staleness report |
|
||||
|
||||
## Self-Check Before Finishing
|
||||
## CLI
|
||||
|
||||
1. `gitnexus_impact` was run for all modified symbols
|
||||
2. No HIGH/CRITICAL warnings were ignored
|
||||
3. `gitnexus_detect_changes()` confirms expected scope
|
||||
4. All d=1 dependents were updated
|
||||
|
||||
## Keeping the Index Fresh
|
||||
|
||||
```bash
|
||||
npx gitnexus analyze # incremental by default; preserves embeddings
|
||||
npx gitnexus analyze --force # full rebuild from scratch (opt out of incremental)
|
||||
npx gitnexus analyze --embeddings # also generate embeddings for new/changed nodes
|
||||
npx gitnexus analyze --drop-embeddings # explicit opt-in to wipe existing embeddings
|
||||
```
|
||||
|
||||
`analyze` runs **incrementally by default**. The pipeline still parses every file every run (cross-file resolution requires it), but tree-sitter parsing is **served from a content-addressed cache** under `.gitnexus/parse-cache/` (per-chunk JSON shards plus `index.json`) for chunks whose file contents haven't changed since the last run. Older installs may still have a legacy single file `.gitnexus/parse-cache.json`, which is read for backward compatibility but no longer written. Only changed-file rows (and their importers) are rewritten in LadybugDB; unchanged-file rows are preserved. Output is byte-equivalent to a full rebuild. Pass `--force` to wipe and re-index from scratch (e.g., to recover from a corrupt index, or after upgrading GitNexus).
|
||||
|
||||
The parse cache key is **content-addressed and version-tagged**: it survives `--force` runs, and is automatically invalidated by a `gitnexus` package upgrade (so a new tree-sitter grammar doesn't silently replay stale parse output). Safe to delete the whole `.gitnexus/parse-cache/` directory (and remove any legacy `.gitnexus/parse-cache.json` if present) at any time — it'll be rebuilt on the next analyze.
|
||||
|
||||
Check `.gitnexus/meta.json` `stats.embeddings` (0 = none). A plain `analyze` no longer drops existing vectors — pass `--drop-embeddings` to wipe.
|
||||
|
||||
> Claude Code: PostToolUse hook detects a stale index after `git commit` and `git merge` and prompts the agent to run `analyze`. The hook does not invoke `analyze` itself.
|
||||
|
||||
## CLI Skills
|
||||
|
||||
| Task | Skill file |
|
||||
|------|-----------|
|
||||
| Architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
|
||||
| Blast radius / "What breaks?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
|
||||
| Debugging / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
|
||||
| Refactoring | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
|
||||
| Tools/resources/schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
|
||||
| CLI commands (index, status, clean, wiki) | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
|
||||
|
||||
## Hook env knobs
|
||||
|
||||
The Claude Code hook (`gitnexus/hooks/claude/gitnexus-hook.cjs` and the mirrored plugin copy under `gitnexus-claude-plugin/hooks/`) honours these env vars. Defaults work for normal installations; set them only to override resolution. All path overrides ignore values that do not exist on disk and fall through to the standard resolution chain.
|
||||
|
||||
| Env var | Type | Default | Purpose |
|
||||
|---------|------|---------|---------|
|
||||
| `GITNEXUS_HOOK_CLI_PATH` | path | resolved via package layout / `require.resolve` | Override path to the `gitnexus` CLI entry the hook spawns for `augment`. |
|
||||
| `GITNEXUS_HOOK_LSOF_PATH` | path | `lsof` on `PATH` (with `/usr/bin/lsof`, `/usr/sbin/lsof`, `/sbin/lsof` fallbacks) | Override POSIX `lsof` location for the DB-lock probe. |
|
||||
| `GITNEXUS_HOOK_PS_PATH` | path | `ps` on `PATH` (with `/bin/ps`, `/usr/bin/ps` fallbacks) | Override POSIX `ps` location. |
|
||||
| `GITNEXUS_HOOK_POWERSHELL_PATH` | path | `%SystemRoot%\System32\WindowsPowerShell\v1.0\powershell.exe` (then `SysWOW64`, then `powershell.exe` on `PATH`) | Override Windows PowerShell location used by the Restart-Manager probe. |
|
||||
| `GITNEXUS_HOOK_LINUX_PROC_BUDGET_MS` | integer ms | `1200` | Max wall-clock for the Linux `/proc` fd scan before bailing out to the `lsof` fallback. |
|
||||
| `GITNEXUS_HOOK_RM_TARGET` | path | derived | Restart-Manager target file (the LadybugDB path under `.gitnexus/`). Set internally by the hook; rarely overridden manually. |
|
||||
| `GITNEXUS_DEBUG` | boolean (`1`/`true`) | unset | Verbose stderr from the hook: prints discarded augment-stderr prefixes and one-shot `.ps1` load-failure warnings. |
|
||||
| Task | Read this skill file |
|
||||
|------|---------------------|
|
||||
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
|
||||
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
|
||||
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
|
||||
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
|
||||
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
|
||||
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
|
||||
| Work in the Ingestion area (239 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
|
||||
| Work in the Extractors area (135 symbols) | `.claude/skills/generated/extractors/SKILL.md` |
|
||||
| Work in the Components area (112 symbols) | `.claude/skills/generated/components/SKILL.md` |
|
||||
| Work in the Lbug area (96 symbols) | `.claude/skills/generated/lbug/SKILL.md` |
|
||||
| Work in the Group area (94 symbols) | `.claude/skills/generated/group/SKILL.md` |
|
||||
| Work in the Cli area (92 symbols) | `.claude/skills/generated/cli/SKILL.md` |
|
||||
| Work in the Configs area (92 symbols) | `.claude/skills/generated/configs/SKILL.md` |
|
||||
| Work in the Type-extractors area (90 symbols) | `.claude/skills/generated/type-extractors/SKILL.md` |
|
||||
| Work in the Hooks area (88 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
|
||||
| Work in the Unit area (80 symbols) | `.claude/skills/generated/unit/SKILL.md` |
|
||||
| Work in the Cpp area (73 symbols) | `.claude/skills/generated/cpp/SKILL.md` |
|
||||
| Work in the Scope-resolution area (72 symbols) | `.claude/skills/generated/scope-resolution/SKILL.md` |
|
||||
| Work in the Server area (66 symbols) | `.claude/skills/generated/server/SKILL.md` |
|
||||
| Work in the Local area (61 symbols) | `.claude/skills/generated/local/SKILL.md` |
|
||||
| Work in the Wiki area (60 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
|
||||
| Work in the Workers area (57 symbols) | `.claude/skills/generated/workers/SKILL.md` |
|
||||
| Work in the Embeddings area (56 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
|
||||
| Work in the Typescript area (53 symbols) | `.claude/skills/generated/typescript/SKILL.md` |
|
||||
| Work in the Storage area (51 symbols) | `.claude/skills/generated/storage/SKILL.md` |
|
||||
| Work in the Php area (48 symbols) | `.claude/skills/generated/php/SKILL.md` |
|
||||
|
||||
<!-- gitnexus:end -->
|
||||
|
||||
|
||||
@@ -52,3 +52,67 @@ If always-on instructions grow, load deep conventions via conditional reads (e.g
|
||||
## GitNexus rules
|
||||
|
||||
See the `<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** for the canonical MCP tools, impact analysis rules, and index instructions.
|
||||
|
||||
<!-- gitnexus:start -->
|
||||
# GitNexus — Code Intelligence
|
||||
|
||||
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||
|
||||
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
||||
|
||||
## Always Do
|
||||
|
||||
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
|
||||
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
|
||||
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
|
||||
|
||||
## Never Do
|
||||
|
||||
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
|
||||
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
|
||||
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
|
||||
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
|
||||
|
||||
## Resources
|
||||
|
||||
| Resource | Use for |
|
||||
|----------|---------|
|
||||
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
|
||||
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
|
||||
| `gitnexus://repo/GitNexus/processes` | All execution flows |
|
||||
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
|
||||
|
||||
## CLI
|
||||
|
||||
| Task | Read this skill file |
|
||||
|------|---------------------|
|
||||
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
|
||||
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
|
||||
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
|
||||
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
|
||||
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
|
||||
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
|
||||
| Work in the Ingestion area (239 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
|
||||
| Work in the Extractors area (135 symbols) | `.claude/skills/generated/extractors/SKILL.md` |
|
||||
| Work in the Components area (112 symbols) | `.claude/skills/generated/components/SKILL.md` |
|
||||
| Work in the Lbug area (96 symbols) | `.claude/skills/generated/lbug/SKILL.md` |
|
||||
| Work in the Group area (94 symbols) | `.claude/skills/generated/group/SKILL.md` |
|
||||
| Work in the Cli area (92 symbols) | `.claude/skills/generated/cli/SKILL.md` |
|
||||
| Work in the Configs area (92 symbols) | `.claude/skills/generated/configs/SKILL.md` |
|
||||
| Work in the Type-extractors area (90 symbols) | `.claude/skills/generated/type-extractors/SKILL.md` |
|
||||
| Work in the Hooks area (88 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
|
||||
| Work in the Unit area (80 symbols) | `.claude/skills/generated/unit/SKILL.md` |
|
||||
| Work in the Cpp area (73 symbols) | `.claude/skills/generated/cpp/SKILL.md` |
|
||||
| Work in the Scope-resolution area (72 symbols) | `.claude/skills/generated/scope-resolution/SKILL.md` |
|
||||
| Work in the Server area (66 symbols) | `.claude/skills/generated/server/SKILL.md` |
|
||||
| Work in the Local area (61 symbols) | `.claude/skills/generated/local/SKILL.md` |
|
||||
| Work in the Wiki area (60 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
|
||||
| Work in the Workers area (57 symbols) | `.claude/skills/generated/workers/SKILL.md` |
|
||||
| Work in the Embeddings area (56 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
|
||||
| Work in the Typescript area (53 symbols) | `.claude/skills/generated/typescript/SKILL.md` |
|
||||
| Work in the Storage area (51 symbols) | `.claude/skills/generated/storage/SKILL.md` |
|
||||
| Work in the Php area (48 symbols) | `.claude/skills/generated/php/SKILL.md` |
|
||||
|
||||
<!-- gitnexus:end -->
|
||||
|
||||
@@ -1,4 +1,5 @@
|
||||
# GitNexus
|
||||
|
||||
**⚠️ Important Notice:** GitNexus has NO official cryptocurrency, token, or coin. Any token/coin using the GitNexus name on Pump.fun or any other platform is **not affiliated with, endorsed by, or created by** this project or its maintainers. Do not purchase any cryptocurrency claiming association with GitNexus.
|
||||
|
||||
<div align="center">
|
||||
@@ -30,14 +31,9 @@
|
||||
|
||||
Indexes any codebase into a knowledge graph — every dependency, call chain, cluster, and execution flow — then exposes it through smart tools so AI agents never miss code.
|
||||
|
||||
|
||||
|
||||
|
||||
https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
|
||||
|
||||
|
||||
|
||||
> *Like DeepWiki, but deeper.* DeepWiki helps you *understand* code. GitNexus lets you *analyze* it — because a knowledge graph tracks every relationship, not just descriptions.
|
||||
> _Like DeepWiki, but deeper._ DeepWiki helps you _understand_ code. GitNexus lets you _analyze_ it — because a knowledge graph tracks every relationship, not just descriptions.
|
||||
|
||||
**TL;DR:** The **Web UI** is a quick way to chat with any repo. The **CLI + MCP** is how you make your AI agent actually reliable — it gives Cursor, Claude Code, Codex, and friends a deep architectural view of your codebase so they stop missing dependencies, breaking call chains, and shipping blind edits. Even smaller models get full architectural clarity, making it compete with Goliath models.
|
||||
|
||||
@@ -47,18 +43,17 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
|
||||
|
||||
[](https://www.star-history.com/#abhigyanpatwari/GitNexus&type=date&legend=top-left)
|
||||
|
||||
|
||||
## Two Ways to Use GitNexus
|
||||
|
||||
| | **CLI + MCP** | **Web UI** |
|
||||
| ----------------- | -------------------------------------------------------------- | ------------------------------------------------------------ |
|
||||
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
|
||||
| **For** | Daily development with Cursor, Claude Code, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
|
||||
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
|
||||
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
|
||||
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
|
||||
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
|
||||
| **Privacy** | Everything local, no network | Everything in-browser, no server |
|
||||
| | **CLI + MCP** | **Web UI** |
|
||||
| ----------- | --------------------------------------------------------------------- | -------------------------------------------------------------------- |
|
||||
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
|
||||
| **For** | Daily development with Cursor, Claude Code, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
|
||||
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
|
||||
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
|
||||
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
|
||||
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
|
||||
| **Privacy** | Everything local, no network | Everything in-browser, no server |
|
||||
|
||||
> **Bridge mode:** `gitnexus serve` connects the two — the web UI auto-detects the local server and can browse all your CLI-indexed repos without re-uploading or re-indexing.
|
||||
|
||||
@@ -69,6 +64,7 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
|
||||
GitNexus is available as an **enterprise offering** - either as a fully managed **SaaS** or a **self-hosted** deployment. Also available for **commercial use** of the OSS version with proper licensing.
|
||||
|
||||
Enterprise includes:
|
||||
|
||||
- **PR Review** - automated blast radius analysis on pull requests
|
||||
- **Auto-updating Code Wiki** - always up-to-date documentation (Code Wiki is also available in OSS)
|
||||
- **Auto-reindexing** - knowledge graph stays fresh automatically
|
||||
@@ -77,6 +73,7 @@ Enterprise includes:
|
||||
- **Priority feature/language support** - request new languages or features
|
||||
|
||||
**Upcoming:**
|
||||
|
||||
- Auto regression forensics
|
||||
- End-to-end test generation
|
||||
|
||||
@@ -117,13 +114,13 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
|
||||
|
||||
### Editor Support
|
||||
|
||||
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
|
||||
| --------------------- | --- | ------ | -------------------- | -------------- |
|
||||
| **Claude Code** | Yes | Yes | Yes (PreToolUse + PostToolUse) | **Full** |
|
||||
| **Cursor** | Yes | Yes | Yes (postToolUse, [manual install](gitnexus-cursor-integration/README.md#hook-install)) | **Full** |
|
||||
| **Codex** | Yes | Yes | — | MCP + Skills |
|
||||
| **Windsurf** | Yes | — | — | MCP |
|
||||
| **OpenCode** | Yes | Yes | — | MCP + Skills |
|
||||
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
|
||||
| --------------- | --- | ------ | --------------------------------------------------------------------------------------- | ------------ |
|
||||
| **Claude Code** | Yes | Yes | Yes (PreToolUse + PostToolUse) | **Full** |
|
||||
| **Cursor** | Yes | Yes | Yes (postToolUse, [manual install](gitnexus-cursor-integration/README.md#hook-install)) | **Full** |
|
||||
| **Codex** | Yes | Yes | — | MCP + Skills |
|
||||
| **Windsurf** | Yes | — | — | MCP |
|
||||
| **OpenCode** | Yes | Yes | — | MCP + Skills |
|
||||
|
||||
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that enrich searches with graph context + PostToolUse hooks that detect a stale index after commits and prompt the agent to reindex.
|
||||
|
||||
@@ -131,10 +128,10 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
|
||||
|
||||
Built by the community — not officially maintained, but worth checking out.
|
||||
|
||||
| Project | Author | Description |
|
||||
|---------|--------|-------------|
|
||||
| [pi-gitnexus](https://github.com/tintinweb/pi-gitnexus) | [@tintinweb](https://github.com/tintinweb) | GitNexus plugin for [pi](https://pi.dev) — `pi install npm:pi-gitnexus` |
|
||||
| [gitnexus-stable-ops](https://github.com/ShunsukeHayashi/gitnexus-stable-ops) | [@ShunsukeHayashi](https://github.com/ShunsukeHayashi) | Stable ops & deployment workflows (Miyabi ecosystem) |
|
||||
| Project | Author | Description |
|
||||
| ----------------------------------------------------------------------------- | ------------------------------------------------------ | ----------------------------------------------------------------------- |
|
||||
| [pi-gitnexus](https://github.com/tintinweb/pi-gitnexus) | [@tintinweb](https://github.com/tintinweb) | GitNexus plugin for [pi](https://pi.dev) — `pi install npm:pi-gitnexus` |
|
||||
| [gitnexus-stable-ops](https://github.com/ShunsukeHayashi/gitnexus-stable-ops) | [@ShunsukeHayashi](https://github.com/ShunsukeHayashi) | Stable ops & deployment workflows (Miyabi ecosystem) |
|
||||
|
||||
> Have a project built on GitNexus? Open a PR to add it here!
|
||||
|
||||
@@ -197,7 +194,8 @@ args = ["-y", "gitnexus@latest", "mcp"]
|
||||
```bash
|
||||
gitnexus setup # Configure MCP for your editors (one-time)
|
||||
gitnexus analyze [path] # Index a repository (or update stale index)
|
||||
gitnexus analyze --force # Force full re-index
|
||||
gitnexus analyze --repair-fts # Fast path: rebuild/verify only FTS indexes on existing index data
|
||||
gitnexus analyze --force # Full rebuild: re-parse + graph rebuild + FTS rebuild
|
||||
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
|
||||
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
|
||||
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
|
||||
@@ -205,6 +203,7 @@ gitnexus analyze --skip-git # Index folders that are not Git repositories
|
||||
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
|
||||
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
|
||||
gitnexus analyze --worker-timeout 60 # Increase worker idle timeout for slow parses
|
||||
gitnexus analyze --workers <n> # Parse worker pool size (default: cores-1, capped at 16; 0 = sequential)
|
||||
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
|
||||
gitnexus serve # Start local HTTP server (multi-repo) for web UI connection
|
||||
gitnexus list # List all indexed repositories
|
||||
@@ -229,6 +228,25 @@ gitnexus group status <name> # Check staleness of repos in a group
|
||||
|
||||
If `analyze` reports a worker parse timeout on a large or unusual repository, it keeps running and falls back safely. To give slow worker jobs more time, use `gitnexus analyze --worker-timeout 60` or set `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=60000`. For very large files, `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` controls the worker job byte budget.
|
||||
|
||||
#### Environment variables
|
||||
|
||||
Most `analyze` knobs are also CLI flags (`--workers`, `--worker-timeout`, `--max-file-size`, `--verbose`). Use the env-var form when you'd otherwise repeat the same flag every run, or when invoking GitNexus from a long-running host (MCP server, eval-server, CI shell) that already manages its own environment. CLI flags take precedence over env vars; env vars take precedence over built-in defaults.
|
||||
|
||||
| Variable | Default | Effect | Tune when… |
|
||||
| -------------------------------------- | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITNEXUS_WORKER_POOL_SIZE` | `cores - 1`, capped at 16 | Parse worker pool size. `0` disables the pool (sequential fallback). Equivalent to `--workers <n>`. | Constrained containers (cgroup CPU limits), CI runners with explicit quotas, or debugging a worker-only crash via `0`. |
|
||||
| `GITNEXUS_PARSE_CHUNK_CONCURRENCY` | `2` | Number of chunks whose file contents may be read into memory in parallel while the pool dispatches the current chunk. Worker dispatch itself stays serial. | Repos large enough to chunk (multi-MB total source) where disk I/O is a measurable fraction of analyze wall-clock. |
|
||||
| `GITNEXUS_VERBOSE` | unset | When `1`, enables verbose ingestion logs (skipped-file warnings, per-chunk throughput, parse-cache stats). Equivalent to `--verbose`. | Debugging an analyze that "completed" but seems to have missed files; tuning `--workers` / chunk concurrency against observable throughput. |
|
||||
| `GITNEXUS_MAX_FILE_SIZE` | `512` (KB) | Walker skip threshold in KB. Hard cap is `32768` (tree-sitter buffer ceiling). Equivalent to `--max-file-size <kb>`. | Indexing repos with intentionally-large source files (generated parsers, vendored bundles) that should still be parsed. |
|
||||
| `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS` | `30000` | Worker idle timeout in milliseconds before retry/fallback. Equivalent to `--worker-timeout <seconds>` × 1000. | Slow-parsing files (large minified JS, deeply-nested TS types) that legitimately need more than 30s. |
|
||||
| `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` | `8388608` (8 MB) | Per-job byte budget the pool will send to a worker in one `postMessage`. | Very large individual files; mostly diagnostic — bumping past 8 MB risks structured-clone memory pressure. |
|
||||
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per worker slot before the slot is dropped from the active rotation. Bounds respawn loops on a chronically-crashing slot. | Hosts where a flaky worker should retry more (raise) or fail-fast (lower) before the slot is dropped. |
|
||||
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Combined with `timeoutBackoffFactor`, prevents exponentially-growing retries from stalling for hours. | Slow files that legitimately need long total retry windows; lower to fail-fast on stalls. |
|
||||
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD`| `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, every subsequent dispatch rejects until a fresh pool is created. | Hosts where a SIGSEGV-prone native grammar should trip the breaker sooner; CI runners that should fail loudly. |
|
||||
| `GITNEXUS_CHUNK_BYTE_BUDGET` | `2097152` (2 MB) | Chunk boundary used for cache-key composition and dispatch. Smaller = finer-grained cache hits but more dispatch overhead. | Tuning incremental-analyze cache behavior on monorepos. |
|
||||
| `GITNEXUS_NO_GITIGNORE` | unset | When set, skips `.gitignore` parsing. `.gitnexusignore` is still honored. | Indexing a repo whose `.gitignore` excludes files you actually want indexed (e.g., generated code committed for cross-repo lookup). |
|
||||
| `GITNEXUS_SKIP_OPTIONAL_GRAMMARS` | unset | When `=1` strictly, skips native builds for `tree-sitter-dart` / `tree-sitter-proto` at install time. | Installing on a host without a C++ toolchain; you're willing to skip Dart/Proto parsing. |
|
||||
|
||||
#### Publishing to understand-quickly (opt-in)
|
||||
|
||||
[`looptech-ai/understand-quickly`](https://github.com/looptech-ai/understand-quickly) is a public registry of code-knowledge graphs that lists `gitnexus@1` as a first-class format. After registering your repo once (`npx @understand-quickly/cli add` or the [wizard](https://looptech-ai.github.io/understand-quickly/add.html)), `gitnexus publish` fires a single `repository_dispatch` event so the registry resyncs your entry on demand instead of waiting for the nightly job.
|
||||
@@ -239,27 +257,27 @@ It is opt-in and a no-op without `UNDERSTAND_QUICKLY_TOKEN` — a fine-grained G
|
||||
|
||||
**16 tools** exposed via MCP (11 per-repo + 5 group):
|
||||
|
||||
| Tool | What It Does | `repo` Param |
|
||||
| ------------------ | ----------------------------------------------------------------- | -------------- |
|
||||
| `list_repos` | Discover all indexed repositories | — |
|
||||
| `query` | Process-grouped hybrid search (BM25 + semantic + RRF) | Optional |
|
||||
| `context` | 360-degree symbol view — categorized refs, process participation | Optional |
|
||||
| `impact` | Blast radius analysis with depth grouping and confidence | Optional |
|
||||
| `detect_changes` | Git-diff impact — maps changed lines to affected processes | Optional |
|
||||
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
|
||||
| `cypher` | Raw Cypher graph queries | Optional |
|
||||
| `group_list` | List configured repository groups | — |
|
||||
| `group_sync` | Extract contracts and match across repos/services | — |
|
||||
| `group_contracts`| Inspect extracted contracts and cross-links | — |
|
||||
| `group_query` | Search execution flows across all repos in a group | — |
|
||||
| `group_status` | Check staleness of repos in a group | — |
|
||||
| Tool | What It Does | `repo` Param |
|
||||
| ----------------- | ---------------------------------------------------------------- | ------------ |
|
||||
| `list_repos` | Discover all indexed repositories | — |
|
||||
| `query` | Process-grouped hybrid search (BM25 + semantic + RRF) | Optional |
|
||||
| `context` | 360-degree symbol view — categorized refs, process participation | Optional |
|
||||
| `impact` | Blast radius analysis with depth grouping and confidence | Optional |
|
||||
| `detect_changes` | Git-diff impact — maps changed lines to affected processes | Optional |
|
||||
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
|
||||
| `cypher` | Raw Cypher graph queries | Optional |
|
||||
| `group_list` | List configured repository groups | — |
|
||||
| `group_sync` | Extract contracts and match across repos/services | — |
|
||||
| `group_contracts` | Inspect extracted contracts and cross-links | — |
|
||||
| `group_query` | Search execution flows across all repos in a group | — |
|
||||
| `group_status` | Check staleness of repos in a group | — |
|
||||
|
||||
> When only one repo is indexed, the `repo` parameter is optional. With multiple repos, specify which one: `query({query: "auth", repo: "my-app"})`.
|
||||
|
||||
**Resources** for instant context:
|
||||
|
||||
| Resource | Purpose |
|
||||
| ----------------------------------------- | ---------------------------------------------------- |
|
||||
| Resource | Purpose |
|
||||
| --------------------------------------- | ---------------------------------------------------- |
|
||||
| `gitnexus://repos` | List all indexed repositories (read this first) |
|
||||
| `gitnexus://repo/{name}/context` | Codebase stats, staleness check, and available tools |
|
||||
| `gitnexus://repo/{name}/clusters` | All functional clusters with cohesion scores |
|
||||
@@ -270,9 +288,9 @@ It is opt-in and a no-op without `UNDERSTAND_QUICKLY_TOKEN` — a fine-grained G
|
||||
|
||||
**2 MCP prompts** for guided workflows:
|
||||
|
||||
| Prompt | What It Does |
|
||||
| ----------------- | ------------------------------------------------------------------------- |
|
||||
| `detect_impact` | Pre-commit change analysis — scope, affected processes, risk level |
|
||||
| Prompt | What It Does |
|
||||
| --------------- | ------------------------------------------------------------------------- |
|
||||
| `detect_impact` | Pre-commit change analysis — scope, affected processes, risk level |
|
||||
| `generate_map` | Architecture documentation from the knowledge graph with mermaid diagrams |
|
||||
|
||||
**4 agent skills** installed to `.claude/skills/` automatically:
|
||||
@@ -359,10 +377,10 @@ npx gitnexus@latest serve
|
||||
|
||||
The official Docker setup ships **two signed images** orchestrated by `docker-compose.yaml`. Each image is published to both **GitHub Container Registry** (GHCR) and **Docker Hub** — same build, same digest, same Cosign signature — so pick whichever registry you prefer:
|
||||
|
||||
| Purpose | GHCR (default in `docker-compose.yaml`) | Docker Hub mirror |
|
||||
| ---------------------------------------------------------------------- | --------------------------------------------- | ------------------------------------------- |
|
||||
| CLI / `gitnexus serve` backend (HTTP API on port `4747`, MCP, indexer) | `ghcr.io/abhigyanpatwari/gitnexus:latest` | `akonlabs/gitnexus:latest` |
|
||||
| Static web UI (port `4173`) | `ghcr.io/abhigyanpatwari/gitnexus-web:latest` | `akonlabs/gitnexus-web:latest` |
|
||||
| Purpose | GHCR (default in `docker-compose.yaml`) | Docker Hub mirror |
|
||||
| ---------------------------------------------------------------------- | --------------------------------------------- | ------------------------------ |
|
||||
| CLI / `gitnexus serve` backend (HTTP API on port `4747`, MCP, indexer) | `ghcr.io/abhigyanpatwari/gitnexus:latest` | `akonlabs/gitnexus:latest` |
|
||||
| Static web UI (port `4173`) | `ghcr.io/abhigyanpatwari/gitnexus-web:latest` | `akonlabs/gitnexus-web:latest` |
|
||||
|
||||
> **Heads-up — image rename.** Earlier releases published the web UI under
|
||||
> `ghcr.io/abhigyanpatwari/gitnexus`. Starting with the introduction of the
|
||||
@@ -578,22 +596,22 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
|
||||
|
||||
### Supported Languages
|
||||
|
||||
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
|
||||
|----------|---------|----------------|---------|----------|-----------------|---------------------|--------|------------|-------------|
|
||||
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
|
||||
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
|
||||
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
|
||||
| ---------- | ------- | -------------- | ------- | -------- | ---------------- | --------------------- | ------ | ---------- | ------------ |
|
||||
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
|
||||
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
|
||||
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
|
||||
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
|
||||
|
||||
@@ -725,9 +743,11 @@ gitnexus wiki --force
|
||||
|
||||
|
||||
# Increase the timeout or retries for large codebase or slow LLM providers
|
||||
gitnexus wiki --timeout <seconds> # Per-attempt LLM request timeout in seconds (default: 60)
|
||||
gitnexus wiki --timeout <seconds> # LLM request timeout in seconds (default: disabled)
|
||||
gitnexus wiki --retries <n> # Max LLM retry attempts per request (default: 3)
|
||||
|
||||
# Change the language generation for wiki
|
||||
gitnexus wiki --lang <lang> # Output language for generated documentation (e.g. english, chinese, spanish, japanese)
|
||||
```
|
||||
|
||||
The wiki generator reads the indexed graph structure, groups files into modules via LLM, generates per-module documentation pages, and creates an overview page — all with cross-references to the knowledge graph.
|
||||
@@ -736,16 +756,16 @@ The wiki generator reads the indexed graph structure, groups files into modules
|
||||
|
||||
## Tech Stack
|
||||
|
||||
| Layer | CLI | Web |
|
||||
| ------------------------- | ------------------------------------- | --------------------------------------- |
|
||||
| Layer | CLI | Web |
|
||||
| ------------------- | ------------------------------------- | --------------------------------------- |
|
||||
| **Runtime** | Node.js (native) | Browser (WASM) |
|
||||
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
|
||||
| **Database** | LadybugDB native | LadybugDB WASM |
|
||||
| **Database** | LadybugDB native | LadybugDB WASM |
|
||||
| **Embeddings** | HuggingFace transformers.js (GPU/CPU) | transformers.js (WebGPU/WASM) |
|
||||
| **Search** | BM25 + semantic + RRF | BM25 + semantic + RRF |
|
||||
| **Agent Interface** | MCP (stdio) | LangChain ReAct agent |
|
||||
| **Visualization** | — | Sigma.js + Graphology (WebGL) |
|
||||
| **Frontend** | — | React 18, TypeScript, Vite, Tailwind v4 |
|
||||
| **Visualization** | — | Sigma.js + Graphology (WebGL) |
|
||||
| **Frontend** | — | React 18, TypeScript, Vite, Tailwind v4 |
|
||||
| **Clustering** | Graphology | Graphology |
|
||||
| **Concurrency** | Worker threads + async | Web Workers + Comlink |
|
||||
|
||||
@@ -761,12 +781,12 @@ The wiki generator reads the indexed graph structure, groups files into modules
|
||||
|
||||
### Recently Completed
|
||||
|
||||
- [X] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
|
||||
- [X] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
|
||||
- [X] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
|
||||
- [X] Multi-Repo MCP, Zero-Config Setup, 14 Language Support
|
||||
- [X] Community Detection, Process Detection, Confidence Scoring
|
||||
- [X] Hybrid Search, Vector Index
|
||||
- [x] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
|
||||
- [x] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
|
||||
- [x] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
|
||||
- [x] Multi-Repo MCP, Zero-Config Setup, 14 Language Support
|
||||
- [x] Community Detection, Process Detection, Confidence Scoring
|
||||
- [x] Hybrid Search, Vector Index
|
||||
|
||||
---
|
||||
|
||||
|
||||
+57
-2
@@ -162,8 +162,8 @@ Each mode has a `system_{mode}.jinja` + `instance_{mode}.jinja` pair. The agent
|
||||
|
||||
```
|
||||
Agent → bash command → /usr/local/bin/gitnexus-query
|
||||
→ curl localhost:4848/tool/query (fast path: eval-server, ~100ms)
|
||||
→ npx gitnexus query (fallback: cold CLI, ~5-10s)
|
||||
→ curl http://127.0.0.1:4848/tool/query (fast path: eval-server, ~100ms)
|
||||
→ npx gitnexus query (fallback: cold CLI, ~5-10s)
|
||||
```
|
||||
|
||||
Each tool script in `/usr/local/bin/` is standalone — no sourcing, no env inheritance needed. This is critical because mini-swe-agent runs every command via `subprocess.run` in a fresh subshell.
|
||||
@@ -176,6 +176,61 @@ The eval-server is a lightweight HTTP daemon that:
|
||||
- Includes next-step hints to guide tool chaining (query → context → impact → fix)
|
||||
- Auto-shuts down after idle timeout
|
||||
|
||||
**CLI flags:**
|
||||
|
||||
| Flag | Default | Purpose |
|
||||
|------|---------|---------|
|
||||
| `--port <port>` | `4848` | Port to listen on |
|
||||
| `--host <host>` | `127.0.0.1` | Bind address — use `0.0.0.0` for cross-container access |
|
||||
| `--idle-timeout <seconds>` | `0` (disabled) | Auto-shutdown after N seconds of inactivity |
|
||||
|
||||
**READY signal:**
|
||||
|
||||
When the server is ready, it writes to stdout:
|
||||
|
||||
```
|
||||
# IPv4
|
||||
GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848
|
||||
|
||||
# IPv6 (bracketed to avoid colon ambiguity)
|
||||
GITNEXUS_EVAL_SERVER_READY:[::1]:4848
|
||||
```
|
||||
|
||||
Parse the port as the last colon-segment (`split(':').pop()`) — not `split(':')[1]`, which breaks for IPv6 and for non-loopback IPv4 hosts added in this release.
|
||||
|
||||
### Custom port and host
|
||||
|
||||
`run_eval.py` does not expose `--port` or `--host` as CLI flags. Configure them in your mode YAML under the `environment:` key:
|
||||
|
||||
```yaml
|
||||
# configs/modes/native_augment.yaml (or whichever mode you're running)
|
||||
environment:
|
||||
eval_server_port: 4849 # change if 4848 is already in use on the host
|
||||
eval_server_host: "0.0.0.0" # bind all interfaces — needed for cross-container setups
|
||||
```
|
||||
|
||||
Defaults are `port: 4848` and `host: 127.0.0.1` (loopback only). Use `0.0.0.0` only when the agent container needs to reach the eval-server from a separate network namespace. The health probe and tool scripts connect via the configured bind host (defaulting to `127.0.0.1`), which is reachable for both loopback and all-interface binds.
|
||||
|
||||
`"localhost"` is also a valid `eval_server_host` value. The OS resolves it at bind time — typically `127.0.0.1` on dual-stack or IPv4-only systems, and `::1` on IPv6-only systems. The exact result depends on your `/etc/hosts` and `gai.conf`. The READY signal will reflect the actual bound address (e.g. `GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848` or `GITNEXUS_EVAL_SERVER_READY:[::1]:4848`), not the literal string `localhost`. Use this when you want the server to bind to whichever loopback address the OS prefers rather than forcing IPv4.
|
||||
|
||||
**Running eval-server directly in Docker / Docker Compose:**
|
||||
|
||||
```bash
|
||||
# Bind to all interfaces so sibling containers can reach it
|
||||
gitnexus eval-server --host 0.0.0.0 --port 4848
|
||||
|
||||
# Then probe from a sibling container via its service hostname
|
||||
curl http://eval-container:4848/health
|
||||
```
|
||||
|
||||
If you need a non-default port (e.g. to avoid conflicts), pass `--port <port>` alongside `--host`. The READY signal will reflect both:
|
||||
|
||||
```
|
||||
GITNEXUS_EVAL_SERVER_READY:0.0.0.0:5000
|
||||
```
|
||||
|
||||
Parse the port as the last colon-segment (`split(':').pop()`) — safe for both IPv4 and bracketed IPv6 forms.
|
||||
|
||||
### Index caching
|
||||
|
||||
SWE-bench repos repeat (Django has 200+ instances at different commits). The harness caches GitNexus indexes per `(repo, commit)` hash in `~/.gitnexus-eval-cache/` to avoid redundant re-indexing.
|
||||
|
||||
@@ -39,6 +39,7 @@ logger = logging.getLogger("gitnexus_docker")
|
||||
|
||||
DEFAULT_CACHE_DIR = Path.home() / ".gitnexus-eval-cache"
|
||||
EVAL_SERVER_PORT = 4848
|
||||
EVAL_SERVER_HOST = "127.0.0.1"
|
||||
|
||||
|
||||
class GitNexusDockerEnvironment(DockerEnvironment):
|
||||
@@ -62,6 +63,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
|
||||
skip_embeddings: bool = True,
|
||||
gitnexus_timeout: int = 120,
|
||||
eval_server_port: int = EVAL_SERVER_PORT,
|
||||
eval_server_host: str = EVAL_SERVER_HOST,
|
||||
**kwargs,
|
||||
):
|
||||
super().__init__(**kwargs)
|
||||
@@ -70,6 +72,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
|
||||
self.skip_embeddings = skip_embeddings
|
||||
self.gitnexus_timeout = gitnexus_timeout
|
||||
self.eval_server_port = eval_server_port
|
||||
self.eval_server_host = eval_server_host
|
||||
self.index_time: float = 0.0
|
||||
self._gitnexus_ready = False
|
||||
|
||||
@@ -165,22 +168,29 @@ class GitNexusDockerEnvironment(DockerEnvironment):
|
||||
|
||||
def _start_eval_server(self):
|
||||
"""Start the GitNexus eval-server daemon in the background."""
|
||||
logger.info(f"Starting eval-server on port {self.eval_server_port}...")
|
||||
logger.info(
|
||||
f"Starting eval-server on {self.eval_server_host}:{self.eval_server_port}..."
|
||||
)
|
||||
|
||||
self.execute({
|
||||
"command": (
|
||||
f"nohup npx gitnexus eval-server --port {self.eval_server_port} "
|
||||
f"--host {self.eval_server_host} "
|
||||
f"--idle-timeout 600 "
|
||||
f"> /tmp/gitnexus-eval-server.log 2>&1 &"
|
||||
),
|
||||
"timeout": 5,
|
||||
})
|
||||
|
||||
# Use 127.0.0.1 for the health probe — reachable whether server binds
|
||||
# loopback or all interfaces (0.0.0.0), avoiding DNS resolution issues.
|
||||
health_host = "127.0.0.1"
|
||||
|
||||
# Wait for the server to be ready (up to ~15s for KuzuDB init)
|
||||
for i in range(EVAL_SERVER_HEALTH_RETRIES):
|
||||
time.sleep(EVAL_SERVER_HEALTH_INTERVAL_SECONDS)
|
||||
health = self.execute({
|
||||
"command": f"curl -sf http://127.0.0.1:{self.eval_server_port}/health 2>/dev/null || echo 'NOT_READY'",
|
||||
"command": f"curl -sf http://{health_host}:{self.eval_server_port}/health 2>/dev/null || echo 'NOT_READY'",
|
||||
"timeout": EVAL_SERVER_HEALTH_TIMEOUT_SECONDS,
|
||||
})
|
||||
output = health.get("output", "").strip()
|
||||
@@ -201,7 +211,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
|
||||
)
|
||||
|
||||
@staticmethod
|
||||
def _render_tool_script(spec: ToolScriptSpec, port: str) -> str:
|
||||
def _render_tool_script(spec: ToolScriptSpec, port: str, host: str = EVAL_SERVER_HOST) -> str:
|
||||
"""
|
||||
Render a standalone bash script for a GitNexus tool.
|
||||
|
||||
@@ -212,6 +222,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
|
||||
|
||||
if spec.endpoint:
|
||||
lines.append(f'PORT="${{GITNEXUS_EVAL_PORT:-{port}}}"')
|
||||
lines.append(f'HOST="${{GITNEXUS_EVAL_HOST:-{host}}}"')
|
||||
|
||||
if spec.header:
|
||||
lines.append(spec.header.strip())
|
||||
@@ -221,7 +232,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
|
||||
|
||||
if spec.endpoint:
|
||||
lines.append(
|
||||
f'result=$(curl -sf -X POST "http://127.0.0.1:${{PORT}}{spec.endpoint}" '
|
||||
f'result=$(curl -sf -X POST "http://${{HOST}}:${{PORT}}{spec.endpoint}" '
|
||||
'-H "Content-Type: application/json" -d "$payload" 2>/dev/null)'
|
||||
)
|
||||
lines.append('if [ $? -eq 0 ] && [ -n "$result" ]; then echo "$result"; exit 0; fi')
|
||||
@@ -244,9 +255,10 @@ class GitNexusDockerEnvironment(DockerEnvironment):
|
||||
Uses heredocs with quoted delimiter to avoid all quoting/escaping issues.
|
||||
"""
|
||||
port = str(self.eval_server_port)
|
||||
host = self.eval_server_host
|
||||
|
||||
for spec in TOOL_SPECS.values():
|
||||
script_content = self._render_tool_script(spec, port).strip()
|
||||
script_content = self._render_tool_script(spec, port, host).strip()
|
||||
# Use heredoc with quoted delimiter — prevents all variable expansion and quoting issues
|
||||
self.execute({
|
||||
"command": (
|
||||
@@ -387,5 +399,6 @@ class GitNexusDockerEnvironment(DockerEnvironment):
|
||||
"index_time_seconds": round(self.index_time, 2),
|
||||
"skip_embeddings": self.skip_embeddings,
|
||||
"eval_server_port": self.eval_server_port,
|
||||
"eval_server_host": self.eval_server_host,
|
||||
}
|
||||
return base
|
||||
|
||||
Generated
+3
-3
@@ -760,11 +760,11 @@ wheels = [
|
||||
|
||||
[[package]]
|
||||
name = "idna"
|
||||
version = "3.11"
|
||||
version = "3.15"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/6f/6d/0703ccc57f3a7233505399edb88de3cbd678da106337b9fcde432b65ed60/idna-3.11.tar.gz", hash = "sha256:795dafcc9c04ed0c1fb032c2aa73654d8e8c5023a7df64a53f39190ada629902", size = 194582, upload-time = "2025-10-12T14:55:20.501Z" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/82/77/7b3966d0b9d1d31a36ddf1746926a11dface89a83409bf1483f0237aa758/idna-3.15.tar.gz", hash = "sha256:ca962446ea538f7092a95e057da437618e886f4d349216d2b1e294abfdb65fdc", size = 199245, upload-time = "2026-05-12T22:45:57.011Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/0e/61/66938bbb5fc52dbdf84594873d5b51fb1f7c7794e9c0f5bd885f30bc507b/idna-3.11-py3-none-any.whl", hash = "sha256:771a87f49d9defaf64091e6e6fe9c18d4833f140bd19464795bc32d966ca37ea", size = 71008, upload-time = "2025-10-12T14:55:18.883Z" },
|
||||
{ url = "https://files.pythonhosted.org/packages/d2/23/408243171aa9aaba178d3e2559159c24c1171a641aa83b67bdd3394ead8e/idna-3.15-py3-none-any.whl", hash = "sha256:048adeaf8c2d788c40fee287673ccaa74c24ffd8dcf09ffa555a2fbb59f10ac8", size = 72340, upload-time = "2026-05-12T22:45:55.733Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
|
||||
@@ -56,15 +56,15 @@ Generates repository documentation from the knowledge graph using an LLM. Requir
|
||||
|
||||
| Flag | Effect |
|
||||
|------|--------|
|
||||
| `--force` | Force full regeneration |
|
||||
| `--force` | Force full regeneration, also required to re-gerenate an existing wiki in a different language |
|
||||
| `--model <model>` | LLM model (default: minimax/minimax-m2.5) |
|
||||
| `--base-url <url>` | LLM API base URL |
|
||||
| `--api-key <key>` | LLM API key |
|
||||
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
|
||||
| `--gist` | Publish wiki as a public GitHub Gist |
|
||||
| `--timeout <seconds>` | Per-attempt LLM request timeout in seconds (default: 60) |
|
||||
| `--timeout <seconds>` | LLM request timeout in seconds (default: disabled) |
|
||||
| `--retries <n>` | Max LLM retry attempts per request (default: 3) |
|
||||
|
||||
| `--lang <lang>` | Output language for generated documentation (e.g. english, chinese, spanish, japanese)|
|
||||
### list — Show all indexed repos
|
||||
|
||||
```bash
|
||||
|
||||
@@ -127,6 +127,7 @@ export { CLASS_KINDS, METHOD_KINDS, FIELD_KINDS } from './scope-resolution/regis
|
||||
export type {
|
||||
RegistryContext,
|
||||
RegistryProviders,
|
||||
OwnedMembersByOwnerLookup,
|
||||
OwnerScopedContributor,
|
||||
ArityVerdict,
|
||||
ConstraintContext,
|
||||
|
||||
@@ -21,6 +21,7 @@
|
||||
* (defined in `./types.ts`).
|
||||
*/
|
||||
|
||||
import type { ParameterTypeClass } from './symbol-definition.js';
|
||||
import type { Range, ScopeId } from './types.js';
|
||||
|
||||
/**
|
||||
@@ -79,4 +80,11 @@ export interface ReferenceSite {
|
||||
* (C#: `42` → `'int'`, `"alice"` → `'string'`).
|
||||
*/
|
||||
readonly argumentTypes?: readonly string[];
|
||||
/**
|
||||
* Optional per-argument type-shape sidecar for languages that need
|
||||
* cv/ref/pointer distinctions during constraint filtering. This is
|
||||
* intentionally separate from `argumentTypes`, which stays normalized
|
||||
* for existing overload narrowing and conversion-rank logic.
|
||||
*/
|
||||
readonly argumentTypeClasses?: readonly ParameterTypeClass[];
|
||||
}
|
||||
|
||||
@@ -13,7 +13,7 @@
|
||||
*/
|
||||
|
||||
import type { NodeLabel } from '../../graph/types.js';
|
||||
import type { SymbolDefinition } from '../symbol-definition.js';
|
||||
import type { ParameterTypeClass, SymbolDefinition } from '../symbol-definition.js';
|
||||
import type { Callsite, DefId } from '../types.js';
|
||||
import type { DefIndex } from '../def-index.js';
|
||||
import type { QualifiedNameIndex } from '../qualified-name-index.js';
|
||||
@@ -65,6 +65,13 @@ export interface ConstraintContext {
|
||||
* `narrowOverloadCandidates`' `argTypes` parameter.
|
||||
*/
|
||||
readonly argumentTypes?: readonly string[];
|
||||
/**
|
||||
* Optional shape-preserving sidecar aligned with `argumentTypes`.
|
||||
* Unknown or unsupported slots should be omitted by producers or
|
||||
* marked with `indirection: 'unknown'`; consumers must preserve the
|
||||
* monotonic fallback and return 'unknown' instead of guessing.
|
||||
*/
|
||||
readonly argumentTypeClasses?: readonly ParameterTypeClass[];
|
||||
}
|
||||
|
||||
// ─── Owner-scoped contributor (concrete shape for `RegistryContributor`) ────
|
||||
@@ -93,6 +100,19 @@ export interface OwnerScopedContributor {
|
||||
byName(name: string): readonly SymbolDefinition[];
|
||||
}
|
||||
|
||||
/**
|
||||
* Required owner-keyed lookup hook for Step 2 receiver/MRO member walks.
|
||||
* Production callers wire this to the SemanticModel's authoritative
|
||||
* method/field/nested-type registries so each `(ownerDefId, memberName)`
|
||||
* probe is O(1). Implementations MUST return `[]` on an indexed miss —
|
||||
* Step 2 treats `[]` as authoritative and does not consult `defs` for a
|
||||
* fallback scan.
|
||||
*/
|
||||
export type OwnedMembersByOwnerLookup = (
|
||||
ownerDefId: DefId,
|
||||
memberName: string,
|
||||
) => readonly SymbolDefinition[];
|
||||
|
||||
// ─── Top-level context threaded through every lookup ───────────────────────
|
||||
|
||||
export interface RegistryContext {
|
||||
@@ -100,6 +120,7 @@ export interface RegistryContext {
|
||||
readonly defs: DefIndex;
|
||||
readonly qualifiedNames: QualifiedNameIndex;
|
||||
readonly moduleScopes: ModuleScopeIndex;
|
||||
readonly ownedMembersByOwner: OwnedMembersByOwnerLookup;
|
||||
/**
|
||||
* Method-dispatch index; required for method/field registries that
|
||||
* honor `useReceiverTypeBinding`. Omit for class-only lookups.
|
||||
|
||||
@@ -27,8 +27,10 @@
|
||||
* is true, resolve the receiver's type at `startScope` (from
|
||||
* `scope.typeBindings`), then walk the MRO via
|
||||
* `MethodDispatchIndex.mroFor(ownerDefId)`. Membership per owner comes
|
||||
* through `RegistryContext.methodDispatch` + owner lookups into
|
||||
* `scope.ownedDefs`; each hit records a raw signal with the owner's
|
||||
* through an optional `RegistryContext.ownedMembersByOwner` hook when
|
||||
* supplied (`undefined` → fall back to `defs.byId`; `[]` → indexed
|
||||
* miss), otherwise via the compatibility fallback scan over
|
||||
* `defs.byId`; each hit records a raw signal with the owner's
|
||||
* MRO depth.
|
||||
*
|
||||
* **Step 3 — Owner-scoped contributor.** When
|
||||
@@ -263,13 +265,14 @@ function walkReceiverTypeBinding(
|
||||
// Walk the owner itself at depth 0, then its MRO chain.
|
||||
const walk: DefId[] = [ownerDefId, ...ctx.methodDispatch.mroFor(ownerDefId)];
|
||||
|
||||
for (let mroDepth = 0; mroDepth < walk.length; mroDepth++) {
|
||||
const currentOwnerId = walk[mroDepth]!;
|
||||
let mroDepth = 0;
|
||||
for (const currentOwnerId of walk) {
|
||||
const members = collectOwnedMembers(currentOwnerId, name, ctx);
|
||||
for (const def of members) {
|
||||
if (!acceptedKinds.has(def.type)) continue;
|
||||
recordTypeBindingHit(perCandidate, def, mroDepth, ownerDefId);
|
||||
}
|
||||
mroDepth++;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -333,23 +336,7 @@ function collectOwnedMembers(
|
||||
memberName: string,
|
||||
ctx: RegistryContext,
|
||||
): readonly SymbolDefinition[] {
|
||||
// An owner's members are defs whose `ownerId === ownerDefId` and whose
|
||||
// simple name matches `memberName`. We iterate `defs.byId` — O(D) per
|
||||
// call today. A future by-owner index would make this O(K); tracked as
|
||||
// a follow-up optimization before Ring 3 flips go production.
|
||||
const out: SymbolDefinition[] = [];
|
||||
for (const def of ctx.defs.byId.values()) {
|
||||
if (def.ownerId !== ownerDefId) continue;
|
||||
if (simpleNameOf(def) !== memberName) continue;
|
||||
out.push(def);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
function simpleNameOf(def: SymbolDefinition): string | undefined {
|
||||
if (def.qualifiedName === undefined || def.qualifiedName.length === 0) return undefined;
|
||||
const dot = def.qualifiedName.lastIndexOf('.');
|
||||
return dot === -1 ? def.qualifiedName : def.qualifiedName.slice(dot + 1);
|
||||
return ctx.ownedMembersByOwner(ownerDefId, memberName);
|
||||
}
|
||||
|
||||
function recordTypeBindingHit(
|
||||
|
||||
Generated
+8
-22
@@ -41,7 +41,7 @@
|
||||
"sigma": "^3.0.2",
|
||||
"tailwindcss": "^4.2.4",
|
||||
"uuid": "^14.0.0",
|
||||
"zod": "^3.25.76"
|
||||
"zod": "^4.3.6"
|
||||
},
|
||||
"devDependencies": {
|
||||
"@babel/types": "^7.29.0",
|
||||
@@ -5599,13 +5599,12 @@
|
||||
}
|
||||
},
|
||||
"node_modules/langsmith": {
|
||||
"version": "0.5.23",
|
||||
"resolved": "https://registry.npmjs.org/langsmith/-/langsmith-0.5.23.tgz",
|
||||
"integrity": "sha512-dE/M/2Gg2S2R8ygDdkWGJVO3JstijvsNvPXsy9V8WGbpb88Zn8xF/aTjPx4mIy5gIoo02T6FssOgYyLf51Dv1Q==",
|
||||
"version": "0.6.3",
|
||||
"resolved": "https://registry.npmjs.org/langsmith/-/langsmith-0.6.3.tgz",
|
||||
"integrity": "sha512-pXrQ4/4myQvjFFOAUmt5pWRrLEZR20gzIJD7MNdUH+5/S5nLI4ZRBo/SYKC6coaYj9pYTfQdBIzcs+3kfJ5uDA==",
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
"p-queue": "6.6.2",
|
||||
"uuid": "10.0.0"
|
||||
"p-queue": "6.6.2"
|
||||
},
|
||||
"peerDependencies": {
|
||||
"@opentelemetry/api": "*",
|
||||
@@ -5632,19 +5631,6 @@
|
||||
}
|
||||
}
|
||||
},
|
||||
"node_modules/langsmith/node_modules/uuid": {
|
||||
"version": "10.0.0",
|
||||
"resolved": "https://registry.npmjs.org/uuid/-/uuid-10.0.0.tgz",
|
||||
"integrity": "sha512-8XkAphELsDnEGrDxUOHB3RGvXz6TeuYSGEZBOjtTtPm2lwhGBjLgOzLHB63IUWfBpNucQjND6d3AOudO+H3RWQ==",
|
||||
"funding": [
|
||||
"https://github.com/sponsors/broofa",
|
||||
"https://github.com/sponsors/ctavan"
|
||||
],
|
||||
"license": "MIT",
|
||||
"bin": {
|
||||
"uuid": "dist/bin/uuid"
|
||||
}
|
||||
},
|
||||
"node_modules/layout-base": {
|
||||
"version": "1.0.2",
|
||||
"resolved": "https://registry.npmjs.org/layout-base/-/layout-base-1.0.2.tgz",
|
||||
@@ -8903,9 +8889,9 @@
|
||||
}
|
||||
},
|
||||
"node_modules/zod": {
|
||||
"version": "3.25.76",
|
||||
"resolved": "https://registry.npmjs.org/zod/-/zod-3.25.76.tgz",
|
||||
"integrity": "sha512-gzUt/qt81nXsFGKIFcC3YnfEAx5NkunCfnDlvuBSSFS02bcXu4Lmea0AFIUwbLWxWPx3d9p8S5QoaujKcNQxcQ==",
|
||||
"version": "4.3.6",
|
||||
"resolved": "https://registry.npmjs.org/zod/-/zod-4.3.6.tgz",
|
||||
"integrity": "sha512-rftlrkhHZOcjDwkGlnUtZZkvaPHCsDATp4pGpuOOMDaTdDDXF91wuVDJoWoPsKX/3YPQ5fHuF3STjcYyKr+Qhg==",
|
||||
"license": "MIT",
|
||||
"funding": {
|
||||
"url": "https://github.com/sponsors/colinhacks"
|
||||
|
||||
@@ -51,7 +51,7 @@
|
||||
"sigma": "^3.0.2",
|
||||
"tailwindcss": "^4.2.4",
|
||||
"uuid": "^14.0.0",
|
||||
"zod": "^3.25.76"
|
||||
"zod": "^4.3.6"
|
||||
},
|
||||
"devDependencies": {
|
||||
"@babel/types": "^7.29.0",
|
||||
|
||||
+12
-1
@@ -151,7 +151,8 @@ Your AI agent gets these tools automatically:
|
||||
```bash
|
||||
gitnexus setup # Configure MCP for your editors (one-time)
|
||||
gitnexus analyze [path] # Index a repository (or update stale index)
|
||||
gitnexus analyze --force # Force full re-index
|
||||
gitnexus analyze --repair-fts # Fast path: rebuild/verify only FTS indexes on existing index data
|
||||
gitnexus analyze --force # Full rebuild: re-parse + graph rebuild + FTS rebuild
|
||||
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
|
||||
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
|
||||
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
|
||||
@@ -358,6 +359,16 @@ npx gitnexus analyze
|
||||
|
||||
For repositories with very large source files, `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` controls the worker job byte budget. The default is **8388608 bytes (8 MB)**.
|
||||
|
||||
### Worker pool resilience tuning
|
||||
|
||||
Three env vars expose the pool's resilience layers (respawn budget, cumulative-timeout cap, circuit breaker). Defaults are tuned for typical repos; bump them when an analyze legitimately needs more retries, or lower them to fail-fast on a known-bad shape.
|
||||
|
||||
| Variable | Default | Effect |
|
||||
| ------------------------------------------------- | ------------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per slot before the slot is dropped from the active rotation. |
|
||||
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Bounds exponentially-growing retry waits. |
|
||||
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` | `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, dispatches require a fresh pool. |
|
||||
|
||||
## Privacy
|
||||
|
||||
- All processing happens locally on your machine
|
||||
|
||||
@@ -0,0 +1,175 @@
|
||||
# Parse-throughput benchmark (scaffold)
|
||||
|
||||
> **Status: methodology + harness scaffold, no measurement data yet.**
|
||||
> The Latest measurement table below contains `_TBD_` placeholders.
|
||||
> This file ships intentionally without numbers — populating it
|
||||
> requires a dedicated bench-pass against the U6 fixture (and ideally
|
||||
> a real-world TS-root-scale repo) on consistent hardware, which is
|
||||
> tracked as future work rather than gated on PR #1693's merge.
|
||||
> Until the table is populated, the load-bearing perf-regression
|
||||
> protection lives in `gitnexus/test/integration/parse-impl-large-fixture.test.ts`
|
||||
> (U6, 30 s wall-clock budget via `Promise.race`).
|
||||
|
||||
Tracks `runChunkedParseAndResolve` wall-clock + peak heap on a synthetic
|
||||
fixture so PR #1693's "analyze no longer hangs on TS-root-shaped loads"
|
||||
claim is measurable, not just asserted by smoke tests. The harness
|
||||
recipe below is deliberately small enough to re-run in a few minutes
|
||||
when the bench-pass is undertaken.
|
||||
|
||||
---
|
||||
|
||||
## Methodology
|
||||
|
||||
### Fixture
|
||||
|
||||
Synthetic TypeScript repo, _not_ a clone of microsoft/TypeScript. CI cost
|
||||
of cloning real-world repos is prohibitive; the synthetic shape exercises
|
||||
the same pipeline paths (chunking, deferred extraction, cross-chunk
|
||||
imports + heritage) without the disk-I/O overhead. Larger numbers can be
|
||||
manually captured against real repos and cross-referenced here, but the
|
||||
authoritative regression-tracking shape is the synthetic fixture so runs
|
||||
are reproducible across hardware.
|
||||
|
||||
The fixture matches the structure pinned by
|
||||
`gitnexus/test/integration/parse-impl-large-fixture.test.ts` (U6):
|
||||
|
||||
- 15 small modules (`mod0.ts` … `mod14.ts`), one exported function each.
|
||||
- 1 dense `complex.ts` with 30 functions + 1 class + 1 interface.
|
||||
- 1 `index.ts` re-exporting every symbol from every module.
|
||||
|
||||
`GITNEXUS_CHUNK_BYTE_BUDGET=64` forces multi-chunk parsing on this small
|
||||
fixture — without that override the whole thing fits in one chunk and
|
||||
the deferred-extraction path is not exercised end-to-end.
|
||||
|
||||
### What to measure
|
||||
|
||||
| Metric | How |
|
||||
| --------------------------- | -------------------------------------------------------------------------------- |
|
||||
| Wall-clock total | `Date.now()` delta around `runChunkedParseAndResolve` |
|
||||
| Peak heap | Sample `process.memoryUsage().heapUsed` every 50 ms during the run; keep the max |
|
||||
| Chunks observed | Count distinct `Parsing chunk X/Y` progress messages |
|
||||
| `getStats()` final snapshot | Quarantined paths, dropped slots, breaker state |
|
||||
|
||||
### Hardware shape (record alongside each measurement)
|
||||
|
||||
- OS + version
|
||||
- CPU model + logical core count
|
||||
- RAM
|
||||
- Node version
|
||||
- gitnexus commit SHA (so the snapshot is anchored to a tree, not "main")
|
||||
|
||||
---
|
||||
|
||||
## Harness recipe
|
||||
|
||||
The U6 test (`test/integration/parse-impl-large-fixture.test.ts`) is the
|
||||
checked-in mini-benchmark — it exercises the same fixture and bounds the
|
||||
wall-clock at 30 s via `Promise.race`. To produce a richer snapshot for
|
||||
this doc, run it under instrumentation:
|
||||
|
||||
```bash
|
||||
# From the gitnexus/ subdir:
|
||||
cd gitnexus
|
||||
# Single-threaded baseline (sequential fallback):
|
||||
npx vitest run test/integration/parse-impl-large-fixture.test.ts --reporter=verbose
|
||||
|
||||
# Worker-pool path (requires built dist/ — pre-built by `npm run build`):
|
||||
npm run build && \
|
||||
GITNEXUS_WORKER_POOL_SIZE=4 \
|
||||
GITNEXUS_PARSE_CHUNK_CONCURRENCY=2 \
|
||||
GITNEXUS_VERBOSE=1 \
|
||||
npx vitest run test/integration/parse-impl-large-fixture.test.ts --reporter=verbose
|
||||
```
|
||||
|
||||
For peak-heap sampling, wrap the dispatch call in a Node script that
|
||||
polls `process.memoryUsage()`. A future helper at
|
||||
`gitnexus/bench/scripts/parse-throughput.ts` would automate this — the
|
||||
plan's stretch goal. Until that lands, capture peak heap manually via:
|
||||
|
||||
```bash
|
||||
node --inspect=0 \
|
||||
--require ./scripts/heap-sampler.js \
|
||||
./node_modules/.bin/vitest run test/integration/parse-impl-large-fixture.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Latest measurement
|
||||
|
||||
> _No measurement data has been collected yet — this file is the
|
||||
> methodology + harness scaffold. The single recorded data point is the
|
||||
> U6 wall-clock smoke baseline below; the worker-pool rows are
|
||||
> placeholders for future bench-pass output._
|
||||
|
||||
The U6 integration test (`gitnexus/test/integration/parse-impl-large-fixture.test.ts`)
|
||||
was observed completing the synthetic fixture in **~6 seconds** under
|
||||
the sequential path (`skipWorkers: true`) on the development machine,
|
||||
well under the 30 s `Promise.race` wall-clock budget. That number is a
|
||||
smoke baseline only — recorded here for reference, not as a regression
|
||||
target.
|
||||
|
||||
| Path | files/s | wall-clock | peak heap | chunks | quarantined |
|
||||
| ------------------------------------------ | ------- | -------------------- | --------- | ------ | ----------- |
|
||||
| Sequential (`skipWorkers: true`, U6 smoke) | _TBD_ | ~6 s _(observation)_ | _TBD_ | 17 | 0 |
|
||||
| Worker pool, `--workers 4`, concurrency 2 | _TBD_ | _TBD_ | _TBD_ | _TBD_ | 0 |
|
||||
| Worker pool, `--workers 1`, concurrency 1 | _TBD_ | _TBD_ | _TBD_ | _TBD_ | 0 |
|
||||
|
||||
**Hardware:** _TBD — record OS, CPU, RAM, Node version, gitnexus SHA at
|
||||
the time of the bench-pass that populates the table above._
|
||||
|
||||
---
|
||||
|
||||
## Operator-tuning quick reference
|
||||
|
||||
Cross-links to the env vars documented in the [README](../../README.md#environment-variables).
|
||||
Use this section as a starting point when the benchmark numbers above
|
||||
suggest a tuning opportunity for your hardware shape.
|
||||
|
||||
- **CPU-bound, big repo, lots of cores:** raise `GITNEXUS_WORKER_POOL_SIZE`
|
||||
past the default cap of 16. The 16-worker cap exists because past that
|
||||
point main-thread merge / extraction dominates; if you've measurably
|
||||
ruled that out, the env var lifts the cap explicitly. (See
|
||||
`worker-pool.ts` `DEFAULT_POOL_SIZE_CAP`.)
|
||||
- **Slow files (large minified JS, deep TS types):** raise
|
||||
`GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS` past 30 000 ms. The cumulative
|
||||
budget is 5× this value (U10 pins this) so a 60 s idle timeout permits
|
||||
300 s of total retry-and-split wall-clock before quarantining the file.
|
||||
- **Constrained container (cgroup CPU limit):** the pool now uses
|
||||
`os.availableParallelism()` (U3 H2), which honors cgroup limits — no
|
||||
manual `GITNEXUS_WORKER_POOL_SIZE` override needed unless the auto-
|
||||
resolved value is too aggressive for your I/O budget.
|
||||
- **Long-running host (eval-server, MCP daemon) running back-to-back
|
||||
analyzes:** `--workers` is now threaded through `AnalyzeOptions`
|
||||
(U2 B2), so per-invocation sizing is honored without `process.env`
|
||||
state leaking across calls. `GITNEXUS_VERBOSE` is similarly snapshot/
|
||||
restore-bracketed.
|
||||
|
||||
---
|
||||
|
||||
## What this benchmark does NOT measure
|
||||
|
||||
- **Real-repo performance.** The synthetic fixture is sized for CI; it
|
||||
doesn't exercise the cumulative-load shape (50k files, occasional
|
||||
pathological file) that drove the original PR #1693 hang report. Real-
|
||||
repo numbers should be captured ad-hoc against the user's target repo
|
||||
and cross-referenced here only as supplementary evidence.
|
||||
- **Worker-pool resilience under real crashes.** That's verified by the
|
||||
`worker-pool.test.ts` integration tests (real `process.exit`, real
|
||||
`error` events, real protocol violations) and the unit suite. The
|
||||
benchmark cares about throughput on the happy path.
|
||||
- **IPC repack throughput.** Phase 3 of the PR #1693 plan introduces a
|
||||
transferList + binary wire-format IPC repack (U16-U17). Once that
|
||||
lands, an `IPC repack` row should be added to the "Latest measurement"
|
||||
table above with before/after numbers on the same hardware.
|
||||
|
||||
---
|
||||
|
||||
## Related artifacts
|
||||
|
||||
- Plan: `docs/plans/2026-05-20-001-feat-pr1693-resilience-hardening-and-ipc-repack-plan.md`
|
||||
- Integration test (mini-benchmark with wall-clock guard): `gitnexus/test/integration/parse-impl-large-fixture.test.ts` (U6)
|
||||
- Operator env-var reference: `README.md` → Environment variables
|
||||
- Resilience layer tests: `gitnexus/test/unit/worker-pool-resilience.test.ts`,
|
||||
`worker-pool-cumulative-timeout.test.ts`,
|
||||
`worker-pool-windows-quarantine.test.ts`,
|
||||
`worker-pool-slot-generation.test.ts`
|
||||
Generated
+209
-639
File diff suppressed because it is too large
Load Diff
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "gitnexus",
|
||||
"version": "1.6.5",
|
||||
"version": "1.6.6-rc.28",
|
||||
"description": "Graph-powered code intelligence for AI agents. Index any codebase, query via MCP or CLI.",
|
||||
"author": "Abhigyan Patwari",
|
||||
"license": "PolyForm-Noncommercial-1.0.0",
|
||||
@@ -60,7 +60,7 @@
|
||||
"cli-progress": "^3.12.0",
|
||||
"commander": "^14.0.3",
|
||||
"cors": "^2.8.5",
|
||||
"express": "^4.19.2",
|
||||
"express": "^5.2.1",
|
||||
"express-rate-limit": "^8.4.1",
|
||||
"glob": "^13.0.6",
|
||||
"graphology": "^0.26.0",
|
||||
@@ -100,7 +100,7 @@
|
||||
"devDependencies": {
|
||||
"@types/cli-progress": "^3.11.6",
|
||||
"@types/cors": "^2.8.17",
|
||||
"@types/express": "^4.17.21",
|
||||
"@types/express": "^5.0.6",
|
||||
"@types/js-yaml": "^4.0.9",
|
||||
"@types/node": "^25.6.0",
|
||||
"@types/uuid": "^11.0.0",
|
||||
|
||||
+180
-7
@@ -68,13 +68,69 @@ const installFatalHandlers = (): void => {
|
||||
});
|
||||
};
|
||||
|
||||
const HEAP_MB = 8192;
|
||||
const HEAP_FLAG = `--max-old-space-size=${HEAP_MB}`;
|
||||
const HEAP_MB = 16384;
|
||||
const TEST_RESPAWN_HEAP_MB = Number(process.env.GITNEXUS_TEST_RESPAWN_HEAP_MB);
|
||||
const RESPAWN_HEAP_MB =
|
||||
Number.isFinite(TEST_RESPAWN_HEAP_MB) && TEST_RESPAWN_HEAP_MB > 0
|
||||
? Math.floor(TEST_RESPAWN_HEAP_MB)
|
||||
: HEAP_MB;
|
||||
const HEAP_FLAG = `--max-old-space-size=${RESPAWN_HEAP_MB}`;
|
||||
/** Increase default stack size (KB) to prevent stack overflow on deep class hierarchies. */
|
||||
const STACK_KB = 4096;
|
||||
const STACK_FLAG = `--stack-size=${STACK_KB}`;
|
||||
|
||||
/** Re-exec the process with an 8GB heap and larger stack if we're currently below that. */
|
||||
/**
|
||||
* Heuristic for "child re-exec likely died from V8 OOM".
|
||||
*
|
||||
* Platform-independent detection is best-effort: V8/Node usually emit
|
||||
* stable heap-exhaustion phrases in stderr/message across Linux/macOS/Windows
|
||||
* (for example "JavaScript heap out of memory" or "Reached heap limit"),
|
||||
* while some environments only expose status/signal (e.g. 134/SIGABRT).
|
||||
* We combine both text signatures and process-exit signatures.
|
||||
*/
|
||||
const childProcessLikelyOom = (err: unknown): boolean => {
|
||||
if (!err || typeof err !== 'object') return false;
|
||||
const e = err as {
|
||||
status?: unknown;
|
||||
signal?: unknown;
|
||||
stderr?: unknown;
|
||||
stdout?: unknown;
|
||||
message?: unknown;
|
||||
};
|
||||
|
||||
const hasHeapOomSignature = (v: unknown): boolean => {
|
||||
const text = (
|
||||
Buffer.isBuffer(v) ? v.toString('utf8') : typeof v === 'string' ? v : ''
|
||||
).toLowerCase();
|
||||
if (!text) return false;
|
||||
return (
|
||||
text.includes('javascript heap out of memory') ||
|
||||
text.includes('reached heap limit') ||
|
||||
text.includes('allocation failed - javascript heap out of memory') ||
|
||||
text.includes('fatalprocessoutofmemory')
|
||||
);
|
||||
};
|
||||
|
||||
const fields = [e.message, e.stderr, e.stdout];
|
||||
if (fields.some((v) => hasHeapOomSignature(v))) return true;
|
||||
|
||||
const hasAnyChildOutput = [e.stderr, e.stdout].some(
|
||||
(v) => (Buffer.isBuffer(v) && v.length > 0) || (typeof v === 'string' && v.length > 0),
|
||||
);
|
||||
if (hasAnyChildOutput) return false;
|
||||
|
||||
return e.status === 134 || e.signal === 'SIGABRT';
|
||||
};
|
||||
|
||||
const forceHeapOOMForTestIfEnabled = (): void => {
|
||||
if (process.env.GITNEXUS_TEST_FORCE_HEAP_OOM !== '1') return;
|
||||
// Allocate JS strings (not Buffers) so pressure lands on V8 heap itself.
|
||||
// Buffers can allocate off-heap, which makes OOM triggering less reliable.
|
||||
const chunks: string[] = [];
|
||||
for (;;) chunks.push('x'.repeat(1024 * 1024));
|
||||
};
|
||||
|
||||
/** Re-exec the process with a 16GB heap and larger stack if we're currently below that. */
|
||||
function ensureHeap(): boolean {
|
||||
const nodeOpts = process.env.NODE_OPTIONS || '';
|
||||
if (nodeOpts.includes('--max-old-space-size')) return false;
|
||||
@@ -92,14 +148,64 @@ function ensureHeap(): boolean {
|
||||
stdio: 'inherit',
|
||||
env: { ...process.env, NODE_OPTIONS: `${nodeOpts} ${HEAP_FLAG}`.trim() },
|
||||
});
|
||||
} catch (e: any) {
|
||||
process.exitCode = e.status ?? 1;
|
||||
} catch (e: unknown) {
|
||||
if (childProcessLikelyOom(e)) {
|
||||
cliError(
|
||||
` Analysis likely ran out of memory.\n` +
|
||||
` Retry with a larger heap if your machine allows it:\n` +
|
||||
` NODE_OPTIONS="--max-old-space-size=24576" gitnexus analyze [your-args]\n` +
|
||||
` (Windows: set NODE_OPTIONS=--max-old-space-size=24576 && gitnexus analyze [your-args])\n` +
|
||||
` If this persists, it may be a native crash unrelated to heap size.\n`,
|
||||
{ recoveryHint: 'heap-oom-respawn' },
|
||||
);
|
||||
}
|
||||
const status =
|
||||
typeof e === 'object' && e !== null && 'status' in e && typeof e.status === 'number'
|
||||
? e.status
|
||||
: 1;
|
||||
process.exitCode = status;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* GITNEXUS_* env vars that `analyzeCommand` writes for backward-compatible
|
||||
* downstream consumption. Snapshotted at function entry and restored in the
|
||||
* finally block so that programmatic callers (tests, long-running hosts)
|
||||
* don't see leaked state across invocations. `GITNEXUS_WORKER_POOL_SIZE` is
|
||||
* NOT in this list: that knob is threaded through `runFullAnalysis` options
|
||||
* (see `workerPoolSize` plumbing) so the CLI never has to mutate `process.env`
|
||||
* for it in the first place.
|
||||
*/
|
||||
const ANALYZE_CLI_ENV_KEYS = [
|
||||
'GITNEXUS_VERBOSE',
|
||||
'GITNEXUS_MAX_FILE_SIZE',
|
||||
'GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS',
|
||||
'GITNEXUS_EMBEDDING_THREADS',
|
||||
'GITNEXUS_EMBEDDING_BATCH_SIZE',
|
||||
'GITNEXUS_EMBEDDING_SUB_BATCH_SIZE',
|
||||
'GITNEXUS_EMBEDDING_DEVICE',
|
||||
] as const;
|
||||
|
||||
type AnalyzeEnvSnapshot = Record<(typeof ANALYZE_CLI_ENV_KEYS)[number], string | undefined>;
|
||||
|
||||
const snapshotAnalyzeEnv = (): AnalyzeEnvSnapshot => {
|
||||
const snap = {} as AnalyzeEnvSnapshot;
|
||||
for (const k of ANALYZE_CLI_ENV_KEYS) snap[k] = process.env[k];
|
||||
return snap;
|
||||
};
|
||||
|
||||
const restoreAnalyzeEnv = (snap: AnalyzeEnvSnapshot): void => {
|
||||
for (const k of ANALYZE_CLI_ENV_KEYS) {
|
||||
const v = snap[k];
|
||||
if (v === undefined) delete process.env[k];
|
||||
else process.env[k] = v;
|
||||
}
|
||||
};
|
||||
|
||||
export interface AnalyzeOptions {
|
||||
force?: boolean;
|
||||
repairFts?: boolean;
|
||||
/**
|
||||
* Embedding generation toggle. Commander parses `--embeddings [limit]` as:
|
||||
* - `undefined` when the flag is omitted
|
||||
@@ -159,6 +265,8 @@ export interface AnalyzeOptions {
|
||||
maxFileSize?: string;
|
||||
/** Override worker sub-batch idle timeout in seconds. */
|
||||
workerTimeout?: string;
|
||||
/** Parse worker pool size; 0 disables workers (sequential fallback). */
|
||||
workers?: string;
|
||||
embeddingThreads?: string;
|
||||
embeddingBatchSize?: string;
|
||||
embeddingSubBatchSize?: string;
|
||||
@@ -185,12 +293,29 @@ export const shouldGenerateCommunitySkillFiles = (
|
||||
|
||||
export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOptions) => {
|
||||
if (ensureHeap()) return;
|
||||
forceHeapOOMForTestIfEnabled();
|
||||
|
||||
// Install fatal handlers immediately after re-exec resolution so any
|
||||
// async error that escapes the try/catch below (#1169) surfaces with
|
||||
// a stack trace and a non-zero exit code instead of a silent exit 0.
|
||||
installFatalHandlers();
|
||||
|
||||
// Snapshot the GITNEXUS_* env vars that the impl writes for downstream
|
||||
// consumption, so they don't leak across `analyzeCommand` invocations in
|
||||
// programmatic callers (tests, long-running hosts). `process.exit(0)` on
|
||||
// the success path bypasses `finally` — intentional: when the process is
|
||||
// exiting, restoration is moot. For early-return paths (validation
|
||||
// errors) and the alreadyUpToDate fast path the finally restores the
|
||||
// pre-call values.
|
||||
const envSnap = snapshotAnalyzeEnv();
|
||||
try {
|
||||
await analyzeCommandImpl(inputPath, options);
|
||||
} finally {
|
||||
restoreAnalyzeEnv(envSnap);
|
||||
}
|
||||
};
|
||||
|
||||
const analyzeCommandImpl = async (inputPath?: string, options?: AnalyzeOptions): Promise<void> => {
|
||||
if (options?.verbose) {
|
||||
process.env.GITNEXUS_VERBOSE = '1';
|
||||
}
|
||||
@@ -211,6 +336,26 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
|
||||
);
|
||||
}
|
||||
|
||||
// `--workers` is threaded through `runFullAnalysis` options → PipelineOptions
|
||||
// → createWorkerPool, intentionally bypassing the GITNEXUS_WORKER_POOL_SIZE
|
||||
// env channel so this CLI surface never mutates `process.env` for pool size.
|
||||
// Tests can therefore re-invoke analyzeCommand with different --workers
|
||||
// values back-to-back and observe the value they passed, not whatever the
|
||||
// previous call leaked.
|
||||
let workerPoolSize: number | undefined;
|
||||
if (options?.workers !== undefined) {
|
||||
const parsedWorkers = Number(options.workers);
|
||||
if (!Number.isInteger(parsedWorkers) || parsedWorkers < 0) {
|
||||
cliError(
|
||||
' --workers must be a non-negative integer. ' +
|
||||
'Pass 0 to disable the worker pool (sequential fallback).\n',
|
||||
);
|
||||
process.exitCode = 1;
|
||||
return;
|
||||
}
|
||||
workerPoolSize = parsedWorkers;
|
||||
}
|
||||
|
||||
// Parse `--embeddings [limit]`: `true` → default cap, string → numeric cap
|
||||
// (0 disables the cap entirely). Validated up here so failures match the
|
||||
// sibling-validation pattern (exit before bar.start() — otherwise
|
||||
@@ -276,6 +421,15 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
|
||||
process.env.GITNEXUS_EMBEDDING_DEVICE = options.embeddingDevice;
|
||||
}
|
||||
|
||||
if (options?.repairFts && options?.force) {
|
||||
cliError(
|
||||
' Cannot combine `--repair-fts` with `--force`. ' +
|
||||
'Use `--repair-fts` for fast FTS-only repair, or `--force` for a full rebuild.\n',
|
||||
);
|
||||
process.exitCode = 1;
|
||||
return;
|
||||
}
|
||||
|
||||
console.log('\n GitNexus Analyzer\n');
|
||||
|
||||
// `--index-only` is the stronger contract — it suppresses every form of file
|
||||
@@ -454,9 +608,11 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
|
||||
// needs a fresh pipelineResult. Has no bearing on the registry
|
||||
// collision guard (see allowDuplicateName below).
|
||||
force: options?.force || options?.skills,
|
||||
repairFts: options?.repairFts,
|
||||
embeddings: embeddingsEnabled,
|
||||
embeddingsNodeLimit,
|
||||
dropEmbeddings: options?.dropEmbeddings,
|
||||
verbose: options?.verbose,
|
||||
skipGit: options?.skipGit,
|
||||
skipAgentsMd,
|
||||
skipSkills,
|
||||
@@ -472,6 +628,10 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
|
||||
// be able to accept the duplicate name without also paying the
|
||||
// cost of a full pipeline re-index. See #829 review round 2.
|
||||
allowDuplicateName: options?.allowDuplicateName,
|
||||
// Worker pool size threaded from --workers, replacing the previous
|
||||
// GITNEXUS_WORKER_POOL_SIZE env mutation. `undefined` defers to the
|
||||
// env / auto-formula fallback inside the pipeline.
|
||||
workerPoolSize,
|
||||
},
|
||||
{
|
||||
onProgress: (_phase, percent, message) => {
|
||||
@@ -501,6 +661,19 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
|
||||
return;
|
||||
}
|
||||
|
||||
if (result.ftsRepairedOnly) {
|
||||
clearInterval(elapsedTimer);
|
||||
process.removeListener('SIGINT', sigintHandler);
|
||||
console.log = origLog;
|
||||
// eslint-disable-next-line no-console -- restoring after intentional progress-bar routing
|
||||
console.warn = origWarn;
|
||||
// eslint-disable-next-line no-console -- restoring after intentional progress-bar routing
|
||||
console.error = origError;
|
||||
bar.stop();
|
||||
console.log(' FTS indexes repaired successfully\n');
|
||||
return;
|
||||
}
|
||||
|
||||
// Post-finalize invariant (#1169): runFullAnalysis nominally writes
|
||||
// meta.json and registers the repo, but on Windows it has been
|
||||
// observed to return successfully with neither artifact present
|
||||
@@ -596,7 +769,7 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
|
||||
}
|
||||
|
||||
console.log('');
|
||||
} catch (err: any) {
|
||||
} catch (err: unknown) {
|
||||
clearInterval(elapsedTimer);
|
||||
process.removeListener('SIGINT', sigintHandler);
|
||||
console.log = origLog;
|
||||
@@ -606,7 +779,7 @@ export const analyzeCommand = async (inputPath?: string, options?: AnalyzeOption
|
||||
console.error = origError;
|
||||
bar.stop();
|
||||
|
||||
const msg = err.message || String(err);
|
||||
const msg = err instanceof Error ? err.message : String(err);
|
||||
|
||||
// Registry name-collision from --name (#829) — surface as an
|
||||
// actionable error rather than a generic stack-trace.
|
||||
|
||||
@@ -14,9 +14,14 @@
|
||||
* Agent bash cmd → curl localhost:PORT/tool/query → eval-server → LocalBackend → format → text
|
||||
*
|
||||
* Usage:
|
||||
* gitnexus eval-server # default port 4848
|
||||
* gitnexus eval-server --port 4848 # explicit port
|
||||
* gitnexus eval-server --idle-timeout 300 # auto-shutdown after 300s idle
|
||||
* gitnexus eval-server # default port 4848, binds 127.0.0.1
|
||||
* gitnexus eval-server --port 4848 # explicit port
|
||||
* gitnexus eval-server --host 0.0.0.0 # reachable from other VMs / containers
|
||||
* gitnexus eval-server --idle-timeout 300 # auto-shutdown after 300s idle
|
||||
*
|
||||
* READY signal format: GITNEXUS_EVAL_SERVER_READY:<host>:<port>
|
||||
* IPv4: GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848
|
||||
* IPv6: GITNEXUS_EVAL_SERVER_READY:[::1]:4848
|
||||
*
|
||||
* API:
|
||||
* POST /tool/:name — Call a tool. Body is JSON arguments. Returns formatted text.
|
||||
@@ -25,16 +30,30 @@
|
||||
*/
|
||||
|
||||
import http from 'http';
|
||||
import { isIPv4, isIPv6 } from 'node:net';
|
||||
import { writeSync } from 'node:fs';
|
||||
import { LocalBackend } from '../mcp/local/local-backend.js';
|
||||
import { logger } from '../core/logger.js';
|
||||
import { cliInfo, cliWarn } from './cli-message.js';
|
||||
import { cliInfo, cliWarn, cliError } from './cli-message.js';
|
||||
|
||||
export interface EvalServerOptions {
|
||||
port?: string;
|
||||
host?: string;
|
||||
idleTimeout?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Validate the --host value. Accepts IPv4, IPv6, or "localhost".
|
||||
* Returns the host string unchanged, or null if invalid.
|
||||
* "localhost" is passed through so the OS resolves it to the correct loopback
|
||||
* address (127.0.0.1 or ::1) at bind time rather than forcing IPv4.
|
||||
*/
|
||||
export function validateHost(raw: string): string | null {
|
||||
if (raw === 'localhost') return raw;
|
||||
if (isIPv4(raw) || isIPv6(raw)) return raw;
|
||||
return null;
|
||||
}
|
||||
|
||||
// ─── Text Formatters ──────────────────────────────────────────────────
|
||||
// Convert structured JSON results into compact, LLM-friendly text.
|
||||
// Design: minimize tokens, maximize actionability.
|
||||
@@ -330,6 +349,22 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
|
||||
const port = parseInt(options?.port || '4848');
|
||||
const idleTimeoutSec = parseInt(options?.idleTimeout || '0');
|
||||
|
||||
const rawHost = options?.host ?? '127.0.0.1';
|
||||
const host = validateHost(rawHost);
|
||||
if (!host) {
|
||||
cliError(
|
||||
`Invalid --host value "${rawHost}":\n` +
|
||||
` Must be an IP address or "localhost".\n\n` +
|
||||
` Examples:\n` +
|
||||
` gitnexus eval-server --host 127.0.0.1 (loopback only, default)\n` +
|
||||
` gitnexus eval-server --host 0.0.0.0 (all network interfaces)\n` +
|
||||
` gitnexus eval-server --host 192.168.1.5 (specific interface)\n` +
|
||||
` gitnexus eval-server --host localhost (OS-resolved loopback)\n`,
|
||||
{ flag: '--host', value: rawHost },
|
||||
);
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
const backend = new LocalBackend();
|
||||
const ok = await backend.init();
|
||||
|
||||
@@ -426,12 +461,72 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
|
||||
}
|
||||
});
|
||||
|
||||
server.listen(port, '127.0.0.1', () => {
|
||||
server.on('error', (err: NodeJS.ErrnoException) => {
|
||||
if (err.code === 'EADDRINUSE') {
|
||||
cliError(
|
||||
`\nGitNexus eval-server failed to start:\n` +
|
||||
` Port ${port} is already in use.\n\n` +
|
||||
` Either:\n` +
|
||||
` 1. Stop the process already using port ${port}\n` +
|
||||
` 2. Use a different port: gitnexus eval-server --port 4849\n`,
|
||||
{ code: err.code, port, host },
|
||||
);
|
||||
} else if (err.code === 'EADDRNOTAVAIL') {
|
||||
// "localhost" may resolve to ::1 on IPv6-only systems; treat it as
|
||||
// potentially IPv6 so the user gets the right diagnostic hint.
|
||||
const isIPv6Host = isIPv6(host) || host === 'localhost';
|
||||
cliError(
|
||||
`\nGitNexus eval-server failed to start:\n` +
|
||||
` Address ${host} is not available on this machine.\n\n` +
|
||||
(isIPv6Host
|
||||
? ` Address ${host} resolved but is not reachable — IPv6 may be disabled, or the loopback interface may be unavailable.\n` +
|
||||
` Docker containers and many CI environments disable IPv6 by default.\n\n`
|
||||
: ` The --host value must be an IP assigned to a local network interface.\n` +
|
||||
` Run \`ip addr\` (Linux) or \`ipconfig\` (Windows) to list available addresses.\n\n`) +
|
||||
` Common fixes:\n` +
|
||||
` gitnexus eval-server --host 127.0.0.1 (loopback, this machine only)\n` +
|
||||
` gitnexus eval-server --host 0.0.0.0 (all interfaces, reachable from other VMs)\n`,
|
||||
{ code: err.code, port, host },
|
||||
);
|
||||
} else if (err.code === 'EACCES') {
|
||||
cliError(
|
||||
`\nGitNexus eval-server failed to start:\n` +
|
||||
` Permission denied binding to port ${port}.\n\n` +
|
||||
` Ports below 1024 require elevated privileges.\n` +
|
||||
` Use a port above 1024: gitnexus eval-server --port 4848\n`,
|
||||
{ code: err.code, port, host },
|
||||
);
|
||||
} else {
|
||||
cliError(`\nGitNexus eval-server failed to start:\n ${err.message}\n`, {
|
||||
code: err.code,
|
||||
port,
|
||||
host,
|
||||
});
|
||||
}
|
||||
process.exit(1);
|
||||
});
|
||||
|
||||
server.listen(port, host, () => {
|
||||
// Plain-text banner for the human watching stderr; structured record
|
||||
// for log aggregation (split into two so the user sees a real banner
|
||||
// not `{"level":30,"msg":"...","port":4747,"endpoints":[...]}`).
|
||||
// Use server.address() so the banner and READY signal reflect what the OS
|
||||
// actually bound to, not the input host string. This matters when "localhost"
|
||||
// is passed: the OS may resolve it to ::1 on some systems.
|
||||
const addr = server.address();
|
||||
// server.listen callback only fires after a successful TCP bind, so
|
||||
// server.address() is guaranteed to return an AddressInfo object here.
|
||||
if (typeof addr !== 'object' || addr === null) {
|
||||
cliError(
|
||||
`\nGitNexus eval-server: unexpected server.address() value after bind: ${JSON.stringify(addr)}\n`,
|
||||
);
|
||||
process.exit(1);
|
||||
}
|
||||
const boundPort = addr.port;
|
||||
const boundAddress = addr.address;
|
||||
const displayHost = boundAddress.includes(':') ? `[${boundAddress}]` : boundAddress;
|
||||
const bannerLines = [
|
||||
`GitNexus eval-server: listening on http://127.0.0.1:${port}`,
|
||||
`GitNexus eval-server: listening on http://${displayHost}:${boundPort}`,
|
||||
` POST /tool/query — search execution flows`,
|
||||
` POST /tool/context — 360-degree symbol view`,
|
||||
` POST /tool/impact — blast radius analysis`,
|
||||
@@ -443,8 +538,8 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
|
||||
bannerLines.push(` Auto-shutdown after ${idleTimeoutSec}s idle`);
|
||||
}
|
||||
cliInfo(bannerLines.join('\n'), {
|
||||
port,
|
||||
host: '127.0.0.1',
|
||||
port: boundPort,
|
||||
host,
|
||||
idleTimeoutSec: idleTimeoutSec > 0 ? idleTimeoutSec : undefined,
|
||||
endpoints: [
|
||||
'POST /tool/query',
|
||||
@@ -457,7 +552,7 @@ export async function evalServerCommand(options?: EvalServerOptions): Promise<vo
|
||||
});
|
||||
try {
|
||||
// Use fd 1 directly — LadybugDB captures process.stdout (#324)
|
||||
writeSync(1, `GITNEXUS_EVAL_SERVER_READY:${port}\n`);
|
||||
writeSync(1, `GITNEXUS_EVAL_SERVER_READY:${displayHost}:${boundPort}\n`);
|
||||
} catch {
|
||||
// stdout may not be available (e.g., broken pipe)
|
||||
}
|
||||
|
||||
@@ -23,6 +23,7 @@ program
|
||||
.command('analyze [path]')
|
||||
.description('Index a repository (full analysis)')
|
||||
.option('-f, --force', 'Force full re-index even if up to date')
|
||||
.option('--repair-fts', 'Repair/rebuild search FTS indexes without full re-analysis')
|
||||
.option(
|
||||
'--embeddings [limit]',
|
||||
'Enable embedding generation for semantic search (off by default). ' +
|
||||
@@ -70,6 +71,10 @@ program
|
||||
'--worker-timeout <seconds>',
|
||||
'Worker sub-batch idle timeout before retry/fallback. Default: 30.',
|
||||
)
|
||||
.option(
|
||||
'--workers <n>',
|
||||
'Parse worker pool size. Default: cores-1 capped at 16. Pass 0 to disable workers (sequential).',
|
||||
)
|
||||
.option('--embedding-threads <n>', 'Limit local ONNX embedding CPU threads')
|
||||
.option('--embedding-batch-size <n>', 'Number of nodes per embedding batch')
|
||||
.option('--embedding-sub-batch-size <n>', 'Number of chunks per embedding model call')
|
||||
@@ -81,6 +86,11 @@ program
|
||||
' GITNEXUS_MAX_FILE_SIZE=N Override large-file skip threshold (KB). Default 512, max 32768.\n' +
|
||||
' GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=N Worker idle timeout in milliseconds. Default 30000.\n' +
|
||||
' GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES=N Worker job byte budget. Default 8388608.\n' +
|
||||
' GITNEXUS_WORKER_POOL_SIZE=N Parse worker count override. Default cores-1 capped at 16.\n' +
|
||||
' GITNEXUS_PARSE_CHUNK_CONCURRENCY=N Concurrent in-flight parse chunks. Default 2.\n' +
|
||||
' GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT=N Max replacement spawns per slot before drop. Default 3.\n' +
|
||||
' GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS=N Total retry wall-time per job. Default 5x sub-batch timeout.\n' +
|
||||
' GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD=N Per-slot deaths to trip circuit breaker. Default max(3, poolSize).\n' +
|
||||
' GITNEXUS_EMBEDDING_THREADS=N Limit local ONNX CPU threads for --embeddings.\n' +
|
||||
' GITNEXUS_SEMANTIC_EXACT_SCAN_LIMIT=N Max embedding chunks for exact-scan fallback. Default 10000.\n' +
|
||||
'\nTip: `.gitnexusignore` supports `.gitignore`-style negation. Add e.g.\n' +
|
||||
@@ -161,11 +171,15 @@ program
|
||||
)
|
||||
.option('--no-reasoning-model', 'Disable reasoning model mode (overrides saved config)')
|
||||
.option('--concurrency <n>', 'Parallel LLM calls (default: 3)', '3')
|
||||
.option('--timeout <seconds>', 'Per-attempt LLM request timeout in seconds (default: 60)')
|
||||
.option('--timeout <seconds>', 'LLM request timeout in seconds (default: disabled)')
|
||||
.option('--retries <n>', 'Max LLM retry attempts per request (default: 3)')
|
||||
.option('--gist', 'Publish wiki as a public GitHub Gist after generation')
|
||||
.option('-v, --verbose', 'Enable verbose output (show LLM commands and responses)')
|
||||
.option('--review', 'Stop after grouping to review module structure before generating pages')
|
||||
.option(
|
||||
'--lang <lang>',
|
||||
'Output language for generated documentation (e.g. english, chinese, spanish, japanese)',
|
||||
)
|
||||
.action(createLazyAction(() => import('./wiki.js'), 'wikiCommand'));
|
||||
|
||||
program
|
||||
@@ -237,6 +251,10 @@ program
|
||||
.command('eval-server')
|
||||
.description('Start lightweight HTTP server for fast tool calls during evaluation')
|
||||
.option('-p, --port <port>', 'Port number', '4848')
|
||||
.option(
|
||||
'--host <host>',
|
||||
'Bind address (default: 127.0.0.1, use 0.0.0.0 to expose to all interfaces)',
|
||||
)
|
||||
.option('--idle-timeout <seconds>', 'Auto-shutdown after N seconds idle (0 = disabled)', '0')
|
||||
.action(createLazyAction(() => import('./eval-server.js'), 'evalServerCommand'));
|
||||
|
||||
|
||||
@@ -61,11 +61,13 @@ function resolveGitnexusBin(): string | null {
|
||||
.filter(Boolean);
|
||||
|
||||
if (isWin) {
|
||||
// On Windows, `where` returns multiple entries (e.g. the POSIX shell
|
||||
// script AND the .cmd/.bat wrapper). Prefer the wrapper because
|
||||
// child_process.spawn() cannot execute a shell script directly.
|
||||
// On Windows, npm global installs can surface multiple launchers for the
|
||||
// same package (e.g. a POSIX shell shim plus .cmd/.bat wrappers). Claude
|
||||
// and the other MCP hosts need a directly spawnable command path, so only
|
||||
// accept the Windows wrapper. If it is missing, fall back to the slower
|
||||
// npx entry instead of persisting a non-spawnable shim path.
|
||||
const cmdLine = lines.find((l) => /\.(cmd|bat)$/i.test(l));
|
||||
return cmdLine || lines[0] || null;
|
||||
return cmdLine || null;
|
||||
}
|
||||
|
||||
return lines[0] || null;
|
||||
|
||||
@@ -35,6 +35,24 @@ export interface WikiCommandOptions {
|
||||
review?: boolean;
|
||||
timeout?: string;
|
||||
retries?: string;
|
||||
lang?: string;
|
||||
}
|
||||
|
||||
function parsePositiveIntegerOption(
|
||||
value: string | undefined,
|
||||
flag: string,
|
||||
multiplier = 1,
|
||||
): number | undefined {
|
||||
if (value === undefined) return undefined;
|
||||
const trimmed = value.trim();
|
||||
if (!/^[1-9]\d*$/.test(trimmed)) {
|
||||
throw new Error(`${flag} must be a positive integer`);
|
||||
}
|
||||
const parsed = parseInt(trimmed, 10);
|
||||
if (parsed > Math.floor(Number.MAX_SAFE_INTEGER / multiplier)) {
|
||||
throw new Error(`${flag} is too large`);
|
||||
}
|
||||
return parsed;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -89,6 +107,24 @@ function prompt(question: string, hide = false): Promise<string> {
|
||||
}
|
||||
|
||||
export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptions) => {
|
||||
// Snapshot GITNEXUS_VERBOSE at entry — wikiCommand mutates it (the impl
|
||||
// below) so cursor-client (process.env-driven) sees the right value during
|
||||
// this run. Restored in finally so back-to-back wiki calls in long-running
|
||||
// hosts don't leak verbose state from one invocation to the next. Pairs
|
||||
// with the same snapshot/restore pattern in `analyzeCommand`.
|
||||
const originalVerbose = process.env.GITNEXUS_VERBOSE;
|
||||
try {
|
||||
await wikiCommandImpl(inputPath, options);
|
||||
} finally {
|
||||
if (originalVerbose === undefined) {
|
||||
delete process.env.GITNEXUS_VERBOSE;
|
||||
} else {
|
||||
process.env.GITNEXUS_VERBOSE = originalVerbose;
|
||||
}
|
||||
}
|
||||
};
|
||||
|
||||
const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions): Promise<void> => {
|
||||
// Set verbose mode globally for cursor-client to pick up
|
||||
if (options?.verbose) {
|
||||
process.env.GITNEXUS_VERBOSE = '1';
|
||||
@@ -127,6 +163,17 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
|
||||
return;
|
||||
}
|
||||
|
||||
let timeoutSeconds: number | undefined;
|
||||
let retries: number | undefined;
|
||||
try {
|
||||
timeoutSeconds = parsePositiveIntegerOption(options?.timeout, '--timeout', 1000);
|
||||
retries = parsePositiveIntegerOption(options?.retries, '--retries');
|
||||
} catch (error) {
|
||||
console.log(` Error: ${(error as Error).message}\n`);
|
||||
process.exitCode = 1;
|
||||
return;
|
||||
}
|
||||
|
||||
// ── Resolve LLM config (with interactive fallback) ─────────────────
|
||||
// Save any CLI overrides immediately
|
||||
if (
|
||||
@@ -350,13 +397,11 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
|
||||
}
|
||||
|
||||
// ── Apply per-run overrides not saved to config ────────────────────
|
||||
if (options?.timeout) {
|
||||
const secs = parseInt(options.timeout, 10);
|
||||
if (!isNaN(secs) && secs > 0) llmConfig.requestTimeoutMs = secs * 1000;
|
||||
if (timeoutSeconds !== undefined) {
|
||||
llmConfig.requestTimeoutMs = timeoutSeconds * 1000;
|
||||
}
|
||||
if (options?.retries) {
|
||||
const n = parseInt(options.retries, 10);
|
||||
if (!isNaN(n) && n > 0) llmConfig.maxAttempts = n;
|
||||
if (retries !== undefined) {
|
||||
llmConfig.maxAttempts = retries;
|
||||
}
|
||||
|
||||
// ── Setup progress bar with elapsed timer ──────────────────────────
|
||||
@@ -395,6 +440,7 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
|
||||
force: options?.force,
|
||||
concurrency: options?.concurrency ? parseInt(options.concurrency, 10) : undefined,
|
||||
reviewOnly: options?.review,
|
||||
lang: options?.lang,
|
||||
};
|
||||
|
||||
const generator = new WikiGenerator(
|
||||
@@ -563,6 +609,8 @@ export const wikiCommand = async (inputPath?: string, options?: WikiCommandOptio
|
||||
|
||||
if (err.message?.includes('No source files')) {
|
||||
console.log(`\n ${err.message}\n`);
|
||||
} else if (err.message?.includes('LLM request timed out after')) {
|
||||
console.log(`\n Timeout: ${err.message}\n`);
|
||||
} else if (err.message?.includes('content filter')) {
|
||||
// Content filter block — actionable message
|
||||
console.log(`\n Content Filter: ${err.message}\n`);
|
||||
|
||||
@@ -18,12 +18,14 @@ import { getPluginForFile, HTTP_SCAN_GLOB, type HttpDetection } from './http-pat
|
||||
* the preferred path because the graph has richer symbol metadata
|
||||
* (real uids, class/method structure, etc.).
|
||||
*
|
||||
* 2. **Source-scan fallback (Strategy B)** — parse files directly with
|
||||
* the per-language plugin registry in `./http-patterns/`. Used when
|
||||
* the graph has no routes/fetches for this repo (e.g. a repo that
|
||||
* hasn't been indexed yet, or whose indexer doesn't know the
|
||||
* framework). Each plugin owns its tree-sitter grammar and query
|
||||
* sources — this orchestrator imports NO grammars or query strings.
|
||||
* 2. **Source-scan supplement (Strategy B)** — parse files directly with
|
||||
* the per-language plugin registry in `./http-patterns/`. Used to
|
||||
* fill gaps when graph extraction only covers part of a polyglot repo
|
||||
* (e.g. Java graph routes plus Go source-scan routes). Graph entries
|
||||
* remain authoritative for duplicate contract IDs because they carry
|
||||
* richer symbol metadata. Each plugin owns its tree-sitter grammar
|
||||
* and query sources — this orchestrator imports NO grammars or query
|
||||
* strings.
|
||||
*
|
||||
* Adding a new language for Strategy B is a one-file edit in
|
||||
* `http-patterns/index.ts`: register a new `HttpLanguagePlugin` and
|
||||
@@ -194,17 +196,19 @@ export class HttpRouteExtractor implements ContractExtractor {
|
||||
|
||||
const graphProviders =
|
||||
dbExecutor != null ? await this.extractProvidersGraph(dbExecutor, getDetections) : [];
|
||||
const providers =
|
||||
graphProviders.length > 0
|
||||
? graphProviders
|
||||
: this.extractProvidersSourceScan(await getScannedFiles(), getDetections);
|
||||
// Source scan always runs to capture routes in languages/files not covered
|
||||
// by graph edges; the glob and per-file parse results are cached above.
|
||||
const providers = this.mergeGraphAndSourceContracts(
|
||||
graphProviders,
|
||||
this.extractProvidersSourceScan(await getScannedFiles(), getDetections),
|
||||
);
|
||||
|
||||
const graphConsumers =
|
||||
dbExecutor != null ? await this.extractConsumersGraph(dbExecutor, getDetections) : [];
|
||||
const consumers =
|
||||
graphConsumers.length > 0
|
||||
? graphConsumers
|
||||
: this.extractConsumersSourceScan(await getScannedFiles(), getDetections);
|
||||
const consumers = this.mergeGraphAndSourceContracts(
|
||||
graphConsumers,
|
||||
this.extractConsumersSourceScan(await getScannedFiles(), getDetections),
|
||||
);
|
||||
|
||||
return [...providers, ...consumers];
|
||||
}
|
||||
@@ -473,4 +477,18 @@ export class HttpRouteExtractor implements ContractExtractor {
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
private mergeGraphAndSourceContracts(
|
||||
graphContracts: ExtractedContract[],
|
||||
sourceContracts: ExtractedContract[],
|
||||
): ExtractedContract[] {
|
||||
const seenContractIds = new Set(graphContracts.map((c) => c.contractId));
|
||||
const out = [...graphContracts];
|
||||
for (const contract of sourceContracts) {
|
||||
if (seenContractIds.has(contract.contractId)) continue;
|
||||
seenContractIds.add(contract.contractId);
|
||||
out.push(contract);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -74,12 +74,31 @@ export const walkRepositoryPaths = async (
|
||||
|
||||
if (skippedLarge > 0) {
|
||||
const isDefault = maxFileSizeBytes === DEFAULT_MAX_FILE_SIZE_BYTES;
|
||||
const isOverrideUnset = !process.env.GITNEXUS_MAX_FILE_SIZE;
|
||||
const suffix = isDefault ? ', likely generated/vendored' : '';
|
||||
logger.warn(` Skipped ${skippedLarge} large files (>${maxFileSizeBytes / 1024}KB${suffix})`);
|
||||
if (isVerboseIngestionEnabled()) {
|
||||
for (const p of skippedLargePaths) {
|
||||
logger.warn(` - ${p}`);
|
||||
}
|
||||
|
||||
// Always show at least the first few paths so users can diagnose why
|
||||
// edges are missing from a specific file (issue #1659). The full list is
|
||||
// gated behind GITNEXUS_VERBOSE=1 to avoid flooding output on repos with
|
||||
// many generated/vendored blobs. Sort before slicing so the preview is
|
||||
// stable across runs (fs.stat callbacks race within each batch).
|
||||
skippedLargePaths.sort();
|
||||
const SKIPPED_PREVIEW_CAP = 5;
|
||||
const showAll = isVerboseIngestionEnabled() || skippedLargePaths.length <= SKIPPED_PREVIEW_CAP;
|
||||
const preview = showAll ? skippedLargePaths : skippedLargePaths.slice(0, SKIPPED_PREVIEW_CAP);
|
||||
for (const p of preview) {
|
||||
logger.warn(` - ${p}`);
|
||||
}
|
||||
if (!showAll) {
|
||||
const remaining = skippedLargePaths.length - SKIPPED_PREVIEW_CAP;
|
||||
logger.warn(` ...and ${remaining} more (set GITNEXUS_VERBOSE=1 to list them all)`);
|
||||
}
|
||||
// Only hint about the env var when the user has not set it at all. An
|
||||
// explicit GITNEXUS_MAX_FILE_SIZE=512 happens to resolve to the same
|
||||
// bytes as the default but the operator clearly already knows the knob.
|
||||
if (isDefault && isOverrideUnset) {
|
||||
logger.warn(` Set GITNEXUS_MAX_FILE_SIZE=<KB> to include files above the default cap.`);
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
import type { Capture, CaptureMatch } from 'gitnexus-shared';
|
||||
import type { Capture, CaptureMatch, ParameterTypeClass } from 'gitnexus-shared';
|
||||
import {
|
||||
findNodeAtRange,
|
||||
nodeToCapture,
|
||||
@@ -9,7 +9,11 @@ import { getCppParser, getCppScopeQuery } from './query.js';
|
||||
import { getTreeSitterBufferSize } from '../../constants.js';
|
||||
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
|
||||
import { splitCppInclude, splitCppUsingDecl } from './import-decomposer.js';
|
||||
import { computeCppDeclarationArity, computeCppCallArity } from './arity-metadata.js';
|
||||
import {
|
||||
classifyCppParameterType,
|
||||
computeCppDeclarationArity,
|
||||
computeCppCallArity,
|
||||
} from './arity-metadata.js';
|
||||
import { markCppAnonymousNamespaceRange, markFileLocal } from './file-local-linkage.js';
|
||||
import { markCppDependentBase } from './two-phase-lookup.js';
|
||||
import { markCppAdlSiteArgs, markCppAdlSiteNoAdl, type CppAdlArgInfo } from './adl.js';
|
||||
@@ -217,6 +221,14 @@ export function emitCppScopeCaptures(
|
||||
JSON.stringify(argTypes),
|
||||
);
|
||||
}
|
||||
const argTypeClasses = inferCppCallArgTypeClasses(cNode);
|
||||
if (argTypeClasses !== undefined && argTypeClasses.length > 0) {
|
||||
grouped['@reference.parameter-type-classes'] = syntheticCapture(
|
||||
'@reference.parameter-type-classes',
|
||||
cNode,
|
||||
JSON.stringify(argTypeClasses),
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -683,6 +695,35 @@ function inferCppCallArgTypes(node: SyntaxNode): string[] | undefined {
|
||||
return types.length > 0 ? types : undefined;
|
||||
}
|
||||
|
||||
function inferCppCallArgTypeClasses(node: SyntaxNode): ParameterTypeClass[] | undefined {
|
||||
const argList = node.childForFieldName('arguments');
|
||||
if (argList === null) return undefined;
|
||||
|
||||
const classes: ParameterTypeClass[] = [];
|
||||
for (let i = 0; i < argList.childCount; i++) {
|
||||
const child = argList.child(i);
|
||||
if (child === null) continue;
|
||||
if (child.type === ',' || child.type === '(' || child.type === ')') continue;
|
||||
const litType = inferCppLiteralType(child);
|
||||
if (litType !== '') {
|
||||
classes.push(valueTypeClass(litType));
|
||||
} else if (child.type === 'identifier') {
|
||||
classes.push(lookupDeclaredTypeClassForIdentifier(child));
|
||||
} else {
|
||||
classes.push(unknownTypeClass('unknown'));
|
||||
}
|
||||
}
|
||||
return classes.length > 0 ? classes : undefined;
|
||||
}
|
||||
|
||||
function valueTypeClass(base: string): ParameterTypeClass {
|
||||
return { base, cv: 'none', indirection: 'value', pointerDepth: 0 };
|
||||
}
|
||||
|
||||
function unknownTypeClass(base: string): ParameterTypeClass {
|
||||
return { base, cv: 'unknown', indirection: 'unknown', pointerDepth: 0 };
|
||||
}
|
||||
|
||||
/**
|
||||
* Infer the canonical type name of a C++ literal AST node.
|
||||
* Returns empty string for non-literal / unknown nodes.
|
||||
@@ -750,6 +791,9 @@ function lookupDeclaredTypeForIdentifier(identNode: SyntaxNode): string {
|
||||
}
|
||||
if (scope === null) return '';
|
||||
|
||||
const paramType = lookupFunctionParameterType(scope, varName);
|
||||
if (paramType !== '') return paramType;
|
||||
|
||||
// Scan declarations in the scope for a matching variable name
|
||||
for (let i = 0; i < scope.childCount; i++) {
|
||||
const stmt = scope.child(i);
|
||||
@@ -763,18 +807,118 @@ function lookupDeclaredTypeForIdentifier(identNode: SyntaxNode): string {
|
||||
// Check init_declarator children for the variable name
|
||||
const declarator = stmt.childForFieldName('declarator');
|
||||
if (declarator === null) continue;
|
||||
if (declarator.type === 'init_declarator') {
|
||||
const nameChild = declarator.childForFieldName('declarator');
|
||||
if (nameChild !== null && nameChild.text === varName) {
|
||||
return normalizeCppTypeText(typeNode.text);
|
||||
}
|
||||
} else if (declarator.text === varName) {
|
||||
const nameChild = declaredNameNode(declarator);
|
||||
if (nameChild !== null && extractDeclaratorLeafName(nameChild) === varName) {
|
||||
return normalizeCppTypeText(typeNode.text);
|
||||
}
|
||||
}
|
||||
return '';
|
||||
}
|
||||
|
||||
function lookupDeclaredTypeClassForIdentifier(identNode: SyntaxNode): ParameterTypeClass {
|
||||
const varName = identNode.text;
|
||||
let scope: SyntaxNode | null = identNode.parent;
|
||||
while (
|
||||
scope !== null &&
|
||||
scope.type !== 'compound_statement' &&
|
||||
scope.type !== 'translation_unit'
|
||||
) {
|
||||
scope = scope.parent;
|
||||
}
|
||||
if (scope === null) return unknownTypeClass('unknown');
|
||||
|
||||
const paramTypeClass = lookupFunctionParameterTypeClass(scope, varName, identNode);
|
||||
if (paramTypeClass !== undefined) return paramTypeClass;
|
||||
|
||||
for (let i = 0; i < scope.childCount; i++) {
|
||||
const stmt = scope.child(i);
|
||||
if (stmt === null || stmt.type !== 'declaration') continue;
|
||||
|
||||
const typeNode = stmt.childForFieldName('type');
|
||||
if (typeNode === null) continue;
|
||||
if (typeNode.type === 'placeholder_type_specifier') continue;
|
||||
|
||||
const declarator = stmt.childForFieldName('declarator');
|
||||
if (declarator === null) continue;
|
||||
const nameChild = declaredNameNode(declarator);
|
||||
if (nameChild === null || extractDeclaratorLeafName(nameChild) !== varName) continue;
|
||||
|
||||
const typeClass = classifyCppParameterType(
|
||||
typeNode.text,
|
||||
nameChild.text,
|
||||
stmt.text.replace(/;\s*$/, ''),
|
||||
);
|
||||
if (isKnownEnumName(identNode, typeClass.base)) {
|
||||
return { ...typeClass, base: `enum:${typeClass.base}` };
|
||||
}
|
||||
return typeClass;
|
||||
}
|
||||
return unknownTypeClass('unknown');
|
||||
}
|
||||
|
||||
function lookupFunctionParameterType(scope: SyntaxNode, varName: string): string {
|
||||
const param = findEnclosingFunctionParameter(scope, varName);
|
||||
if (param === null) return '';
|
||||
const typeNode = param.childForFieldName('type');
|
||||
if (typeNode === null) return '';
|
||||
return normalizeCppTypeText(typeNode.text);
|
||||
}
|
||||
|
||||
function lookupFunctionParameterTypeClass(
|
||||
scope: SyntaxNode,
|
||||
varName: string,
|
||||
identNode: SyntaxNode,
|
||||
): ParameterTypeClass | undefined {
|
||||
const param = findEnclosingFunctionParameter(scope, varName);
|
||||
if (param === null) return undefined;
|
||||
const typeNode = param.childForFieldName('type');
|
||||
if (typeNode === null) return undefined;
|
||||
const declarator = param.childForFieldName('declarator');
|
||||
if (declarator === null) return undefined;
|
||||
const typeClass = classifyCppParameterType(typeNode.text, declarator.text, param.text);
|
||||
if (isKnownEnumName(identNode, typeClass.base)) {
|
||||
return { ...typeClass, base: `enum:${typeClass.base}` };
|
||||
}
|
||||
return typeClass;
|
||||
}
|
||||
|
||||
function findEnclosingFunctionParameter(scope: SyntaxNode, varName: string): SyntaxNode | null {
|
||||
let node: SyntaxNode | null = scope.parent;
|
||||
while (node !== null) {
|
||||
if (node.type === 'function_definition' || node.type === 'function_declarator') {
|
||||
const fnDecl =
|
||||
node.type === 'function_declarator'
|
||||
? node
|
||||
: findFirstDescendantOfType(node, 'function_declarator');
|
||||
const params = fnDecl?.childForFieldName('parameters') ?? null;
|
||||
if (params !== null) {
|
||||
for (let i = 0; i < params.namedChildCount; i++) {
|
||||
const param = params.namedChild(i);
|
||||
if (param === null || param.type !== 'parameter_declaration') continue;
|
||||
const declarator = param.childForFieldName('declarator');
|
||||
if (declarator !== null && extractDeclaratorLeafName(declarator) === varName) {
|
||||
return param;
|
||||
}
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
node = node.parent;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
function declaredNameNode(declarator: SyntaxNode): SyntaxNode | null {
|
||||
if (declarator.type !== 'init_declarator') return declarator;
|
||||
for (let i = 0; i < declarator.namedChildCount; i++) {
|
||||
const child = declarator.namedChild(i);
|
||||
if (child === null) continue;
|
||||
if (child.type === 'identifier') return child;
|
||||
if (child.type.endsWith('_declarator')) return child;
|
||||
}
|
||||
return declarator.childForFieldName('declarator');
|
||||
}
|
||||
|
||||
/** Normalize a type-specifier text for argument type matching.
|
||||
* Strips qualifiers (const, volatile), namespace prefixes (std::),
|
||||
* and pointer/reference markers. */
|
||||
@@ -786,6 +930,25 @@ function normalizeCppTypeText(text: string): string {
|
||||
return t;
|
||||
}
|
||||
|
||||
function isKnownEnumName(node: SyntaxNode, typeName: string): boolean {
|
||||
if (typeName === '' || typeName === 'unknown') return false;
|
||||
let root: SyntaxNode = node;
|
||||
while (root.parent !== null) root = root.parent;
|
||||
const stack: SyntaxNode[] = [root];
|
||||
while (stack.length > 0) {
|
||||
const cur = stack.pop()!;
|
||||
if (cur.type === 'enum_specifier') {
|
||||
const name = cur.childForFieldName('name');
|
||||
if (name?.text === typeName) return true;
|
||||
}
|
||||
for (let i = 0; i < cur.childCount; i++) {
|
||||
const child = cur.child(i);
|
||||
if (child !== null) stack.push(child);
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/**
|
||||
* Detect whether a `namespace_definition` AST node is inline.
|
||||
* Tree-sitter-cpp exposes the `inline` keyword as an anonymous child
|
||||
@@ -1247,7 +1410,9 @@ function extractDeclaratorLeafName(node: SyntaxNode): string | null {
|
||||
const next =
|
||||
cur.childForFieldName('declarator') ??
|
||||
// parenthesized_declarator: single named child
|
||||
(cur.type === 'parenthesized_declarator' ? cur.namedChild(0) : null);
|
||||
(cur.type === 'parenthesized_declarator' || cur.type.endsWith('_declarator')
|
||||
? cur.namedChild(0)
|
||||
: null);
|
||||
if (next === null) return null;
|
||||
cur = next;
|
||||
}
|
||||
|
||||
@@ -20,20 +20,27 @@
|
||||
* NOT: flip compatible↔incompatible; pass through unknown.
|
||||
*/
|
||||
|
||||
import type { ArityVerdict, Callsite, ConstraintContext, SymbolDefinition } from 'gitnexus-shared';
|
||||
import type {
|
||||
ArityVerdict,
|
||||
Callsite,
|
||||
ConstraintContext,
|
||||
ParameterTypeClass,
|
||||
SymbolDefinition,
|
||||
} from 'gitnexus-shared';
|
||||
import { classifyType, type TypeClass } from './type-classifier.js';
|
||||
import type { ConstraintExpr, CppConstraintPayload } from './constraint-extractor.js';
|
||||
|
||||
type AtomicEvaluator = (argClasses: readonly TypeClass[]) => ArityVerdict;
|
||||
interface ConstraintArgClass {
|
||||
readonly typeClass: TypeClass;
|
||||
readonly shape?: ParameterTypeClass;
|
||||
}
|
||||
|
||||
type AtomicEvaluator = (args: readonly ConstraintArgClass[]) => ArityVerdict;
|
||||
|
||||
/**
|
||||
* Curated Tier-A predicate registry — the four canonical
|
||||
* `<type_traits>` variable templates whose truth tables are closed-form
|
||||
* over our coarse `TypeClass` enum.
|
||||
*
|
||||
* Deferred predicates that need a cv/ref/pointer sidecar on
|
||||
* `normalizeCppParamType` (today the normalizer strips those markers
|
||||
* before storage) live in #1579 as one-line follow-up adds.
|
||||
* Curated Tier-A predicate registry. Predicates that depend on pointer,
|
||||
* reference, or cv shape consult `ConstraintContext.argumentTypeClasses`.
|
||||
* Missing or unsupported shape returns 'unknown' to preserve monotonicity.
|
||||
*/
|
||||
// ISO `<type_traits>` treats `bool`, `char`, and the signed/unsigned char
|
||||
// variants as integral types (§21.3.4 Table 48), so `is_integral_v<bool>`
|
||||
@@ -46,30 +53,135 @@ function isIntegralClass(c: TypeClass | undefined): boolean {
|
||||
}
|
||||
|
||||
const REGISTRY = new Map<string, AtomicEvaluator>([
|
||||
['is_integral_v', (cls) => verdictFromBool(isIntegralClass(cls[0]), cls)],
|
||||
['is_floating_point_v', (cls) => verdictFromBool(cls[0] === 'floating', cls)],
|
||||
[
|
||||
'is_void_v',
|
||||
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && arg.typeClass === 'void'),
|
||||
],
|
||||
[
|
||||
'is_integral_v',
|
||||
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && isIntegralClass(arg.typeClass)),
|
||||
],
|
||||
[
|
||||
'is_floating_point_v',
|
||||
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && arg.typeClass === 'floating'),
|
||||
],
|
||||
[
|
||||
'is_arithmetic_v',
|
||||
(cls) => verdictFromBool(isIntegralClass(cls[0]) || cls[0] === 'floating', cls),
|
||||
(args) =>
|
||||
unaryVerdict(
|
||||
args,
|
||||
(arg) =>
|
||||
isPlainValue(arg) && (isIntegralClass(arg.typeClass) || arg.typeClass === 'floating'),
|
||||
),
|
||||
],
|
||||
[
|
||||
'is_enum_v',
|
||||
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && arg.typeClass === 'enum'),
|
||||
],
|
||||
[
|
||||
'is_class_v',
|
||||
(args) => unaryVerdict(args, (arg) => isPlainValue(arg) && arg.typeClass === 'class'),
|
||||
],
|
||||
[
|
||||
'is_pointer_v',
|
||||
(args) =>
|
||||
unaryShapeVerdict(args, (shape) => shape.indirection === 'pointer' && shape.pointerDepth > 0),
|
||||
],
|
||||
[
|
||||
'is_reference_v',
|
||||
(args) =>
|
||||
unaryShapeVerdict(
|
||||
args,
|
||||
(shape) => shape.indirection === 'lvalue-ref' || shape.indirection === 'rvalue-ref',
|
||||
),
|
||||
],
|
||||
[
|
||||
'is_const_v',
|
||||
(args) =>
|
||||
unaryShapeVerdict(args, (shape) => shape.cv === 'const' || shape.cv === 'const volatile', {
|
||||
requireTopLevelCv: true,
|
||||
}),
|
||||
],
|
||||
[
|
||||
'is_volatile_v',
|
||||
(args) =>
|
||||
unaryShapeVerdict(args, (shape) => shape.cv === 'volatile' || shape.cv === 'const volatile', {
|
||||
requireTopLevelCv: true,
|
||||
}),
|
||||
],
|
||||
// NOTE: cv-qualifiers are stripped by `normalizeCppParamType` before the
|
||||
// type token reaches `classifyType`, so `is_same_v<const T, T>` returns
|
||||
// `'compatible'` instead of the ISO-correct `false`. Tracked under the
|
||||
// cv-sidecar refactor in #1579's "Out of scope" list; until that lands
|
||||
// this approximation matches the common `is_same_v<T, ConcreteType>`
|
||||
// dispatch idiom and silently degrades on cv-distinct compares.
|
||||
[
|
||||
'is_same_v',
|
||||
(cls) => {
|
||||
if (cls.length < 2 || cls[0] === 'unknown' || cls[1] === 'unknown') return 'unknown';
|
||||
return cls[0] === cls[1] ? 'compatible' : 'incompatible';
|
||||
(args) => {
|
||||
if (args.length < 2 || args[0].typeClass === 'unknown' || args[1].typeClass === 'unknown') {
|
||||
return 'unknown';
|
||||
}
|
||||
return args[0].typeClass === args[1].typeClass ? 'compatible' : 'incompatible';
|
||||
},
|
||||
],
|
||||
]);
|
||||
|
||||
function verdictFromBool(predicate: boolean, cls: readonly TypeClass[]): ArityVerdict {
|
||||
if (cls[0] === 'unknown') return 'unknown';
|
||||
return predicate ? 'compatible' : 'incompatible';
|
||||
function unaryVerdict(
|
||||
args: readonly ConstraintArgClass[],
|
||||
predicate: (arg: ConstraintArgClass) => boolean,
|
||||
): ArityVerdict {
|
||||
const arg = args[0];
|
||||
if (arg === undefined || arg.typeClass === 'unknown') return 'unknown';
|
||||
return predicate(arg) ? 'compatible' : 'incompatible';
|
||||
}
|
||||
|
||||
function unaryShapeVerdict(
|
||||
args: readonly ConstraintArgClass[],
|
||||
predicate: (shape: ParameterTypeClass) => boolean,
|
||||
options: { readonly requireTopLevelCv?: boolean } = {},
|
||||
): ArityVerdict {
|
||||
const arg = args[0];
|
||||
if (arg === undefined || arg.typeClass === 'unknown') return 'unknown';
|
||||
const shape = arg.shape;
|
||||
if (shape === undefined || shape.indirection === 'unknown' || shape.cv === 'unknown') {
|
||||
return 'unknown';
|
||||
}
|
||||
if (options.requireTopLevelCv === true && shape.indirection === 'pointer') {
|
||||
return 'unknown';
|
||||
}
|
||||
return predicate(shape) ? 'compatible' : 'incompatible';
|
||||
}
|
||||
|
||||
function isPlainValue(arg: ConstraintArgClass): boolean {
|
||||
const shape = arg.shape;
|
||||
if (shape === undefined) return true;
|
||||
return shape.indirection === 'value';
|
||||
}
|
||||
|
||||
function classifyConstraintArg(
|
||||
token: string | undefined,
|
||||
shape?: ParameterTypeClass,
|
||||
): ConstraintArgClass {
|
||||
if (shape !== undefined && shape.base.startsWith('enum:')) {
|
||||
return { typeClass: 'enum', shape };
|
||||
}
|
||||
const typeClass = token === undefined || token === '' ? 'unknown' : classifyType(token);
|
||||
return { typeClass, ...(shape !== undefined ? { shape } : {}) };
|
||||
}
|
||||
|
||||
function tokenForArg(ctx: ConstraintContext, argIdx: number): string | undefined {
|
||||
const shape = ctx.argumentTypeClasses?.[argIdx];
|
||||
if (shape?.base.startsWith('enum:')) return shape.base;
|
||||
return ctx.argumentTypes?.[argIdx];
|
||||
}
|
||||
|
||||
function shapeForTemplateParam(
|
||||
ctx: ConstraintContext,
|
||||
paramName: string,
|
||||
argIdx: number,
|
||||
def?: SymbolDefinition,
|
||||
): ParameterTypeClass | undefined {
|
||||
const argShape = ctx.argumentTypeClasses?.[argIdx];
|
||||
if (argShape === undefined) return undefined;
|
||||
|
||||
const paramShape = def?.parameterTypeClasses?.[argIdx];
|
||||
if (paramShape === undefined) return argShape;
|
||||
if (paramShape.base === paramName && paramShape.indirection === 'value') return argShape;
|
||||
return undefined;
|
||||
}
|
||||
|
||||
/** Public surface — registered as `ScopeResolver.constraintCompatibility`. */
|
||||
@@ -80,13 +192,14 @@ export function cppConstraintCompatibility(
|
||||
): ArityVerdict {
|
||||
const payload = def.templateConstraints as CppConstraintPayload | undefined;
|
||||
if (payload === undefined) return 'unknown';
|
||||
return evaluate(payload.expr, payload, ctx);
|
||||
return evaluate(payload.expr, payload, ctx, def);
|
||||
}
|
||||
|
||||
function evaluate(
|
||||
expr: ConstraintExpr,
|
||||
payload: CppConstraintPayload,
|
||||
ctx: ConstraintContext,
|
||||
def?: SymbolDefinition,
|
||||
): ArityVerdict {
|
||||
switch (expr.kind) {
|
||||
case 'unknown':
|
||||
@@ -96,17 +209,18 @@ function evaluate(
|
||||
if (evaluator === undefined) return 'unknown';
|
||||
const classes = expr.args.map((paramName) => {
|
||||
const argIdx = payload.paramArgIndex[paramName];
|
||||
if (argIdx === undefined) return 'unknown' as TypeClass;
|
||||
const token = ctx.argumentTypes?.[argIdx];
|
||||
if (token === undefined || token === '') return 'unknown' as TypeClass;
|
||||
return classifyType(token);
|
||||
if (argIdx === undefined) return { typeClass: 'unknown' as TypeClass };
|
||||
return classifyConstraintArg(
|
||||
tokenForArg(ctx, argIdx),
|
||||
shapeForTemplateParam(ctx, paramName, argIdx, def),
|
||||
);
|
||||
});
|
||||
return evaluator(classes);
|
||||
}
|
||||
case 'and': {
|
||||
let result: ArityVerdict = 'compatible';
|
||||
for (const child of expr.children) {
|
||||
const v = evaluate(child, payload, ctx);
|
||||
const v = evaluate(child, payload, ctx, def);
|
||||
if (v === 'incompatible') return 'incompatible';
|
||||
if (v === 'unknown') result = 'unknown';
|
||||
}
|
||||
@@ -115,14 +229,14 @@ function evaluate(
|
||||
case 'or': {
|
||||
let result: ArityVerdict = 'incompatible';
|
||||
for (const child of expr.children) {
|
||||
const v = evaluate(child, payload, ctx);
|
||||
const v = evaluate(child, payload, ctx, def);
|
||||
if (v === 'compatible') return 'compatible';
|
||||
if (v === 'unknown') result = 'unknown';
|
||||
}
|
||||
return result;
|
||||
}
|
||||
case 'not': {
|
||||
const v = evaluate(expr.child, payload, ctx);
|
||||
const v = evaluate(expr.child, payload, ctx, def);
|
||||
if (v === 'compatible') return 'incompatible';
|
||||
if (v === 'incompatible') return 'compatible';
|
||||
return 'unknown';
|
||||
|
||||
@@ -1,32 +1,30 @@
|
||||
/**
|
||||
* C++ conversion-rank scoring for overload resolution (#1578).
|
||||
* C++ conversion-rank scoring for overload resolution (#1578, #1637).
|
||||
*
|
||||
* Operates on **normalized** type strings (output of
|
||||
* `normalizeCppParamType` in `arity-metadata.ts`). After normalization:
|
||||
* - int/long/short/unsigned → 'int'
|
||||
* - float/double → 'double'
|
||||
* - char → 'char', bool → 'bool'
|
||||
*
|
||||
* Because the normalizer collapses promotion pairs (int↔long,
|
||||
* float↔double) to the same string, those promotions are invisible at
|
||||
* this layer — they appear as exact matches (rank 0).
|
||||
* Operates on normalized type strings (output of `normalizeCppParamType`
|
||||
* in `arity-metadata.ts`) plus optional shape sidecars from #1630.
|
||||
* Normalization intentionally collapses cv/ref/pointer spelling for stable
|
||||
* graph IDs, so pointer/nullptr rules must consult `ParameterTypeClass`.
|
||||
*
|
||||
* Post-normalization ranking:
|
||||
* - rank 0 — exact (same normalized type)
|
||||
* - rank 1 — integral promotion (char→int, bool→int)
|
||||
* - rank 2 — standard arithmetic conversion (int↔double, char→double,
|
||||
* bool→double)
|
||||
* - Infinity — mismatch (string↔int, user types, pointers, etc.)
|
||||
* - rank 0: exact (same normalized type)
|
||||
* - rank 1: integral promotion (char -> int, bool -> int)
|
||||
* - rank 2: standard conversion (arithmetic, nullptr -> T*, T* -> bool,
|
||||
* T* -> void*)
|
||||
* - rank 3: nullptr -> bool (kept worse than nullptr -> T*)
|
||||
* - rank 4: ellipsis conversion (worst viable)
|
||||
* - Infinity: mismatch (string -> int, user types, unsupported shapes)
|
||||
*
|
||||
* This function is intentionally C++-specific (issue #1578 pitfall:
|
||||
* keep conversion-rank tables out of shared overload-narrowing). Other
|
||||
* languages may define their own `ConversionRankFn` in the future.
|
||||
* This function is intentionally C++-specific. Other languages may define
|
||||
* their own `ConversionRankFn` in the future.
|
||||
*/
|
||||
|
||||
import type { ParameterTypeClass } from 'gitnexus-shared';
|
||||
|
||||
/** Set of normalized arithmetic types that support implicit conversion. */
|
||||
const ARITHMETIC = new Set(['int', 'double', 'char', 'bool']);
|
||||
|
||||
/** Integral promotion targets: char→int and bool→int are rank 1. */
|
||||
/** Integral promotion targets: char -> int and bool -> int are rank 1. */
|
||||
const INTEGRAL_PROMOTION = new Map([
|
||||
['char', 'int'],
|
||||
['bool', 'int'],
|
||||
@@ -35,13 +33,40 @@ const INTEGRAL_PROMOTION = new Map([
|
||||
/**
|
||||
* Return the conversion rank from `argType` to `paramType`.
|
||||
*
|
||||
* @returns 0 for exact match, 1 for integral promotion (char/bool→int),
|
||||
* 2 for standard arithmetic conversion, Infinity for mismatch.
|
||||
* @returns 0 for exact match, 1 for integral promotion, 2 for standard
|
||||
* conversion, 3 for nullptr -> bool, 4 for ellipsis, Infinity
|
||||
* for mismatch.
|
||||
*/
|
||||
export function cppConversionRank(argType: string, paramType: string): number {
|
||||
if (argType === paramType) return 0;
|
||||
// Integral promotions: char→int, bool→int (ISO C++ [conv.prom])
|
||||
export function cppConversionRank(
|
||||
argType: string,
|
||||
paramType: string,
|
||||
argTypeClass?: ParameterTypeClass,
|
||||
paramTypeClass?: ParameterTypeClass,
|
||||
): number {
|
||||
if (argType === paramType) {
|
||||
return exactShapeCompatible(argTypeClass, paramTypeClass) ? 0 : Infinity;
|
||||
}
|
||||
if (paramType === '...') return 4;
|
||||
if (INTEGRAL_PROMOTION.get(argType) === paramType) return 1;
|
||||
if (ARITHMETIC.has(argType) && ARITHMETIC.has(paramType)) return 2;
|
||||
if (argType === 'null' && isPointer(paramTypeClass)) return 2;
|
||||
if (argType === 'null' && paramType === 'bool') return 3;
|
||||
if (isPointer(argTypeClass) && paramType === 'bool') return 2;
|
||||
if (isPointer(argTypeClass) && isPointer(paramTypeClass) && paramType === 'void') return 2;
|
||||
return Infinity;
|
||||
}
|
||||
|
||||
function isPointer(typeClass: ParameterTypeClass | undefined): boolean {
|
||||
return typeClass?.indirection === 'pointer' && typeClass.pointerDepth > 0;
|
||||
}
|
||||
|
||||
function exactShapeCompatible(
|
||||
argTypeClass: ParameterTypeClass | undefined,
|
||||
paramTypeClass: ParameterTypeClass | undefined,
|
||||
): boolean {
|
||||
if (argTypeClass === undefined || paramTypeClass === undefined) return true;
|
||||
if (argTypeClass.indirection === 'unknown' || paramTypeClass.indirection === 'unknown') {
|
||||
return true;
|
||||
}
|
||||
return isPointer(argTypeClass) === isPointer(paramTypeClass);
|
||||
}
|
||||
|
||||
@@ -7,11 +7,10 @@
|
||||
* the call-site inference in `captures.ts`) to one of the categories
|
||||
* the `<type_traits>` predicate registry uses for SFINAE filtering.
|
||||
*
|
||||
* Intentionally coarse: cv / pointer / reference qualifiers are stripped
|
||||
* upstream by `normalizeCppParamType`. Tier-A predicates
|
||||
* (`is_integral_v`, `is_floating_point_v`, `is_arithmetic_v`, `is_same_v`)
|
||||
* are insensitive to those modifiers per ISO `<type_traits>` semantics
|
||||
* ("including any cv-qualified variants").
|
||||
* `argumentTypes` remain normalized for overload narrowing, while
|
||||
* constraint predicates that need cv/ref/pointer shape read the parallel
|
||||
* `argumentTypeClasses` sidecar. Unknown shapes must stay unknown rather
|
||||
* than being guessed as incompatible.
|
||||
*/
|
||||
|
||||
export type TypeClass =
|
||||
@@ -21,7 +20,11 @@ export type TypeClass =
|
||||
| 'char'
|
||||
| 'string'
|
||||
| 'null'
|
||||
| 'void'
|
||||
| 'enum'
|
||||
| 'class'
|
||||
| 'pointer'
|
||||
| 'reference'
|
||||
| 'unknown';
|
||||
|
||||
/**
|
||||
@@ -29,13 +32,17 @@ export type TypeClass =
|
||||
* inference table in `captures.ts:inferCppLiteralType` plus the std::
|
||||
* normalization in `arity-metadata.ts:normalizeCppParamType`.
|
||||
*
|
||||
* Caller note: token must already be normalized (no `const`, no `&` / `*`,
|
||||
* no `std::` prefix). Tokens passed via `ConstraintContext.argumentTypes`
|
||||
* coming from `inferCppCallArgTypes` satisfy this.
|
||||
* Caller note: token should be normalized for overload matching. Enum
|
||||
* tokens produced by the C++ adapter use the internal `enum:<Name>`
|
||||
* prefix so `is_enum_v` does not have to guess that every user token is
|
||||
* class-like.
|
||||
*/
|
||||
export function classifyType(token: string): TypeClass {
|
||||
if (token.length === 0) return 'unknown';
|
||||
if (token.startsWith('enum:')) return 'enum';
|
||||
switch (token) {
|
||||
case 'void':
|
||||
return 'void';
|
||||
case 'int':
|
||||
return 'integral';
|
||||
case 'double':
|
||||
|
||||
@@ -27,6 +27,7 @@ import { javaMethodConfig } from '../method-extractors/configs/jvm.js';
|
||||
import { createVariableExtractor } from '../variable-extractors/generic.js';
|
||||
import { javaVariableConfig } from '../variable-extractors/configs/jvm.js';
|
||||
import { createHeritageExtractor } from '../heritage-extractors/generic.js';
|
||||
import type { SymbolDefinition } from 'gitnexus-shared';
|
||||
import {
|
||||
emitJavaScopeCaptures,
|
||||
interpretJavaImport,
|
||||
@@ -39,6 +40,48 @@ import {
|
||||
resolveJavaImportTarget,
|
||||
} from './java/index.js';
|
||||
|
||||
const orderJavaSameNameTypeCandidates = ({
|
||||
callSiteFilePath,
|
||||
candidates,
|
||||
}: {
|
||||
readonly typeName: string;
|
||||
readonly callSiteFilePath: string;
|
||||
readonly candidates: readonly SymbolDefinition[];
|
||||
}): readonly SymbolDefinition[] | null => {
|
||||
if (!callSiteFilePath.endsWith('.java')) return null;
|
||||
if (candidates.length <= 1) return null;
|
||||
const callerDir = splitDirectorySegments(callSiteFilePath);
|
||||
|
||||
const scored = candidates.map((candidate, index) => ({
|
||||
candidate,
|
||||
index,
|
||||
score: sharedPrefixLength(callerDir, splitDirectorySegments(candidate.filePath)),
|
||||
}));
|
||||
const bestScore = Math.max(...scored.map((entry) => entry.score));
|
||||
// When all candidates tie, we have no structural signal to prefer one path.
|
||||
// Returning null keeps downstream ambiguity handling conservative.
|
||||
if (scored.every((entry) => entry.score === bestScore)) return null;
|
||||
|
||||
const ordered = [...scored]
|
||||
.sort((a, b) => b.score - a.score || a.index - b.index)
|
||||
.map((entry) => entry.candidate);
|
||||
return ordered;
|
||||
};
|
||||
|
||||
const splitDirectorySegments = (filePath: string): string[] => {
|
||||
const normalized = filePath.replace(/\\/g, '/');
|
||||
// Remove empty segments from leading/trailing/multiple slashes, then drop filename.
|
||||
const segments = normalized.split('/').filter(Boolean);
|
||||
return segments.slice(0, -1);
|
||||
};
|
||||
|
||||
const sharedPrefixLength = (left: readonly string[], right: readonly string[]): number => {
|
||||
const max = Math.min(left.length, right.length);
|
||||
let idx = 0;
|
||||
while (idx < max && left[idx] === right[idx]) idx += 1;
|
||||
return idx;
|
||||
};
|
||||
|
||||
export const javaProvider = defineLanguage({
|
||||
id: SupportedLanguages.Java,
|
||||
extensions: ['.java'],
|
||||
@@ -87,4 +130,5 @@ export const javaProvider = defineLanguage({
|
||||
receiverBinding: javaReceiverBinding,
|
||||
arityCompatibility: javaArityCompatibility,
|
||||
resolveImportTarget: resolveJavaImportTarget,
|
||||
orderSameNameTypeCandidates: orderJavaSameNameTypeCandidates,
|
||||
});
|
||||
|
||||
@@ -0,0 +1,12 @@
|
||||
/**
|
||||
* Arity compatibility for JavaScript.
|
||||
*
|
||||
* Delegates to `typescriptArityCompatibility` unchanged — JavaScript
|
||||
* supports the same arity constructs (rest parameters `...args`, default
|
||||
* parameters `p = v`) and the metadata shape (`parameterCount`,
|
||||
* `requiredParameterCount`, `parameterTypes`) is synthesized by the same
|
||||
* `computeTsArityMetadata` function (which understands both TS and JS
|
||||
* parameter node types via `extractTsJsParameters`).
|
||||
*/
|
||||
|
||||
export { typescriptArityCompatibility as jsArityCompatibility } from '../typescript/arity.js';
|
||||
@@ -0,0 +1,722 @@
|
||||
/**
|
||||
* `emitScopeCaptures` for JavaScript.
|
||||
*
|
||||
* Adapts `emitTsScopeCaptures` for the JavaScript grammar:
|
||||
*
|
||||
* 1. **JS grammar** — uses `tree-sitter-javascript` instead of
|
||||
* `tree-sitter-typescript`. The JS scope query is a subset of the
|
||||
* TypeScript one (TypeScript-only node types dropped).
|
||||
*
|
||||
* 2. **CJS `require()` decomposition** — `const { X } = require('./m')`
|
||||
* and `const X = require('./m')` are walked in a post-query pass and
|
||||
* synthesized as `@import.kind/name/alias/source` markers so that
|
||||
* `interpretJsImport` can recover a `ParsedImport` using the same
|
||||
* shape as the TypeScript ESM decomposer.
|
||||
*
|
||||
* 3. **JSDoc type bindings** — JavaScript has no static type annotations
|
||||
* so `@type-binding.parameter` / `@type-binding.return` must be
|
||||
* inferred from leading JSDoc comments. A lightweight regex scanner
|
||||
* (`parseJsDocParams` / `parseJsDocReturn`) extracts `@param {T} n`
|
||||
* and `@returns {T}` tags and emits synthetic captures positioned on
|
||||
* the annotated function node.
|
||||
*
|
||||
* 4. **Shared synthesis passes** — destructuring, for-of map-tuple, and
|
||||
* instanceof narrowing passes are duplicated from `typescript/captures.ts`
|
||||
* (they are pure AST operations with no grammar-specific logic).
|
||||
*
|
||||
* Pure given the input source text. No I/O, no globals consulted.
|
||||
*/
|
||||
|
||||
import type { Capture, CaptureMatch } from 'gitnexus-shared';
|
||||
import {
|
||||
findNodeAtRange,
|
||||
nodeToCapture,
|
||||
syntheticCapture,
|
||||
type SyntaxNode,
|
||||
} from '../../utils/ast-helpers.js';
|
||||
import { splitImportStatement } from '../typescript/import-decomposer.js';
|
||||
import { getJsParser, getJsScopeQuery, jsCachedTreeMatchesGrammar } from './query.js';
|
||||
import { computeTsArityMetadata } from '../typescript/arity-metadata.js';
|
||||
import { synthesizeTsReceiverBinding } from '../typescript/receiver-binding.js';
|
||||
import { getTreeSitterBufferSize } from '../../constants.js';
|
||||
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
|
||||
|
||||
/** JS function-like node types that may carry a synthesized `this` binding.
|
||||
* Kept in sync with the `@scope.function` patterns in `query.ts`. */
|
||||
const FUNCTION_NODE_TYPES = [
|
||||
'method_definition',
|
||||
'arrow_function',
|
||||
'function_expression',
|
||||
'function_declaration',
|
||||
'generator_function_declaration',
|
||||
] as const;
|
||||
|
||||
/** Declaration anchors that carry function-like arity metadata. */
|
||||
const FUNCTION_DECL_TAGS = ['@declaration.method', '@declaration.function'] as const;
|
||||
|
||||
/** Callsite anchors that should carry `@reference.arity` + param types. */
|
||||
const CALL_TAGS = [
|
||||
'@reference.call.free',
|
||||
'@reference.call.member',
|
||||
'@reference.call.constructor',
|
||||
] as const;
|
||||
|
||||
function pickFirstDefined(grouped: CaptureMatch, tags: readonly string[]): Capture | undefined {
|
||||
for (const tag of tags) {
|
||||
const cap = grouped[tag];
|
||||
if (cap !== undefined) return cap;
|
||||
}
|
||||
return undefined;
|
||||
}
|
||||
|
||||
/** Filter `@reference.read.member` in non-read contexts (same logic as TS). */
|
||||
function shouldEmitReadMember(memberNode: SyntaxNode): boolean {
|
||||
const parent = memberNode.parent;
|
||||
if (parent === null) return true;
|
||||
switch (parent.type) {
|
||||
case 'call_expression':
|
||||
return parent.childForFieldName('function')?.id !== memberNode.id;
|
||||
case 'new_expression':
|
||||
return parent.childForFieldName('constructor')?.id !== memberNode.id;
|
||||
case 'assignment_expression':
|
||||
case 'augmented_assignment_expression':
|
||||
return parent.childForFieldName('left')?.id !== memberNode.id;
|
||||
case 'jsx_self_closing_element':
|
||||
case 'jsx_opening_element':
|
||||
return parent.childForFieldName('name')?.id !== memberNode.id;
|
||||
default:
|
||||
return true;
|
||||
}
|
||||
}
|
||||
|
||||
/** Find the first JS function-like node at the given range. */
|
||||
function findFunctionNode(rootNode: SyntaxNode, range: Capture['range']): SyntaxNode | null {
|
||||
for (const nodeType of FUNCTION_NODE_TYPES) {
|
||||
const n = findNodeAtRange(rootNode, range, nodeType);
|
||||
if (n !== null) return n;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/** Infer a callsite argument's static type from literal shapes. */
|
||||
function inferArgType(argNode: SyntaxNode): string {
|
||||
switch (argNode.type) {
|
||||
case 'number':
|
||||
return 'number';
|
||||
case 'string':
|
||||
case 'template_string':
|
||||
return 'string';
|
||||
case 'true':
|
||||
case 'false':
|
||||
return 'boolean';
|
||||
case 'null':
|
||||
return 'null';
|
||||
case 'undefined':
|
||||
return 'undefined';
|
||||
case 'array':
|
||||
return 'Array';
|
||||
case 'object':
|
||||
return 'object';
|
||||
case 'regex':
|
||||
return 'RegExp';
|
||||
case 'new_expression': {
|
||||
const ctor = argNode.childForFieldName('constructor');
|
||||
return ctor?.text ?? '';
|
||||
}
|
||||
default:
|
||||
return '';
|
||||
}
|
||||
}
|
||||
|
||||
// ─── CJS require() decomposition ─────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Walk the AST and synthesize `@import.*` captures for CJS `require()` calls:
|
||||
*
|
||||
* - `const { X, Y } = require('./m')` → one match per destructured name,
|
||||
* `@import.kind = 'named'`, `@import.name = X / Y`.
|
||||
* - `const X = require('./m')` → `@import.kind = 'namespace'`,
|
||||
* `@import.alias = X` (the whole module is bound to X).
|
||||
* - `require('./m')` as a bare expression-statement → side-effect.
|
||||
*
|
||||
* CJS named-alias form (`const { X: alias } = require('./m')`) emits
|
||||
* `@import.kind = 'named-alias'` with `@import.name = X` and
|
||||
* `@import.alias = alias`.
|
||||
*
|
||||
* The synthesized markers are identical to those produced by
|
||||
* `splitImportStatement` for ESM, so `interpretJsImport` can delegate
|
||||
* unchanged to `interpretTsImport` for all cases.
|
||||
*/
|
||||
function synthesizeCjsImports(root: SyntaxNode, out: CaptureMatch[]): void {
|
||||
const stack: SyntaxNode[] = [root];
|
||||
for (;;) {
|
||||
const node = stack.pop();
|
||||
if (node === undefined) break;
|
||||
for (const child of node.namedChildren) {
|
||||
if (child !== null) stack.push(child);
|
||||
}
|
||||
|
||||
if (node.type !== 'call_expression') continue;
|
||||
|
||||
// Require call: function must be bare identifier "require".
|
||||
const fn = node.childForFieldName('function');
|
||||
if (fn === null || fn.type !== 'identifier' || fn.text !== 'require') continue;
|
||||
|
||||
const argsNode = node.childForFieldName('arguments');
|
||||
if (argsNode === null) continue;
|
||||
|
||||
// Source must be a string literal.
|
||||
const firstArg = argsNode.namedChild(0);
|
||||
if (firstArg === null || firstArg.type !== 'string') continue;
|
||||
const rawSource = firstArg.text; // includes surrounding quotes
|
||||
const source = firstArg.namedChild(0)?.text ?? rawSource.slice(1, -1);
|
||||
|
||||
const parent = node.parent;
|
||||
|
||||
// Case 1: const { X } = require('./m') OR const X = require('./m')
|
||||
if (parent?.type === 'variable_declarator') {
|
||||
const nameNode = parent.childForFieldName('name');
|
||||
if (nameNode === null) continue;
|
||||
|
||||
if (nameNode.type === 'object_pattern') {
|
||||
// Destructured: emit one match per specifier.
|
||||
for (const field of nameNode.namedChildren) {
|
||||
if (field === null) continue;
|
||||
if (field.type === 'shorthand_property_identifier_pattern') {
|
||||
const name = field.text;
|
||||
out.push({
|
||||
'@import.statement': syntheticCapture('@import.statement', node, rawSource),
|
||||
'@import.kind': syntheticCapture('@import.kind', node, 'named'),
|
||||
'@import.name': syntheticCapture('@import.name', field, name),
|
||||
'@import.source': syntheticCapture('@import.source', firstArg, source),
|
||||
});
|
||||
} else if (field.type === 'pair_pattern') {
|
||||
const key = field.childForFieldName('key');
|
||||
const value = field.childForFieldName('value');
|
||||
if (key === null || value === null || value.type !== 'identifier') continue;
|
||||
out.push({
|
||||
'@import.statement': syntheticCapture('@import.statement', node, rawSource),
|
||||
'@import.kind': syntheticCapture('@import.kind', node, 'named-alias'),
|
||||
'@import.name': syntheticCapture('@import.name', key, key.text),
|
||||
'@import.alias': syntheticCapture('@import.alias', value, value.text),
|
||||
'@import.source': syntheticCapture('@import.source', firstArg, source),
|
||||
});
|
||||
}
|
||||
}
|
||||
} else if (nameNode.type === 'identifier') {
|
||||
// Namespace-style: const X = require('./m') → bind whole module to X.
|
||||
out.push({
|
||||
'@import.statement': syntheticCapture('@import.statement', node, rawSource),
|
||||
'@import.kind': syntheticCapture('@import.kind', node, 'namespace'),
|
||||
'@import.alias': syntheticCapture('@import.alias', nameNode, nameNode.text),
|
||||
'@import.source': syntheticCapture('@import.source', firstArg, source),
|
||||
});
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
// Case 2: bare require('./m') — side-effect import.
|
||||
if (parent?.type === 'expression_statement') {
|
||||
out.push({
|
||||
'@import.statement': syntheticCapture('@import.statement', node, rawSource),
|
||||
'@import.kind': syntheticCapture('@import.kind', node, 'side-effect'),
|
||||
'@import.source': syntheticCapture('@import.source', firstArg, source),
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ─── JSDoc type binding synthesis ────────────────────────────────────────
|
||||
|
||||
interface JsDocParam {
|
||||
readonly name: string;
|
||||
readonly type: string;
|
||||
}
|
||||
|
||||
/** Extract `@param {Type} name` entries from a JSDoc comment block. */
|
||||
function parseJsDocParams(text: string): readonly JsDocParam[] {
|
||||
const results: JsDocParam[] = [];
|
||||
// Match @param {Type} name or @param {Type} [name] (optional)
|
||||
const re = /@param\s+\{([^}]+)\}\s+\[?(\w+)\]?/g;
|
||||
let m: RegExpExecArray | null;
|
||||
while ((m = re.exec(text)) !== null) {
|
||||
results.push({ type: m[1].trim(), name: m[2].trim() });
|
||||
}
|
||||
return results;
|
||||
}
|
||||
|
||||
/** Extract `@returns {Type}` or `@return {Type}` from a JSDoc comment. */
|
||||
function parseJsDocReturn(text: string): string | null {
|
||||
const m = /@returns?\s+\{([^}]+)\}/.exec(text);
|
||||
return m ? m[1].trim() : null;
|
||||
}
|
||||
|
||||
/** Extract `@type {Type}` from a JSDoc comment (variable-level annotation). */
|
||||
function parseJsDocType(text: string): string | null {
|
||||
const m = /@type\s+\{([^}]+)\}/.exec(text);
|
||||
return m ? m[1].trim() : null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Walk the AST and synthesize `@type-binding.*` captures from JSDoc
|
||||
* comments immediately preceding function declarations / expressions.
|
||||
*
|
||||
* Only `/** … */` block comments are scanned. Line comments (`//`) are
|
||||
* intentionally excluded — JSDoc lives in block comments.
|
||||
*
|
||||
* Emits:
|
||||
* - `@type-binding.parameter` for each `@param {T} n` tag.
|
||||
* - `@type-binding.return` for `@returns {T}` / `@return {T}`.
|
||||
* - `@type-binding.annotation` for `@type {T}` on `let`/`const`/`var`
|
||||
* declarations — covers the common `/** @type {User} */ const u = …`
|
||||
* pattern (ECMA-262 §14.3.1/§14.3.2 variable declarations).
|
||||
*
|
||||
* The binding is anchored on the function node so `tsBindingScopeFor`
|
||||
* can hoist method return-type bindings to Module scope (matching the
|
||||
* TypeScript path where `hoistTypeBindingsToModule: true`).
|
||||
*/
|
||||
function synthesizeJsDocBindings(root: SyntaxNode, out: CaptureMatch[]): void {
|
||||
const stack: SyntaxNode[] = [root];
|
||||
for (;;) {
|
||||
const node = stack.pop();
|
||||
if (node === undefined) break;
|
||||
for (const child of node.namedChildren) {
|
||||
if (child !== null) stack.push(child);
|
||||
}
|
||||
|
||||
const isFnDecl =
|
||||
node.type === 'function_declaration' || node.type === 'generator_function_declaration';
|
||||
const isMethodDef = node.type === 'method_definition';
|
||||
// Also check lexical_declaration containing an arrow/fn-expression
|
||||
const isLexDecl = node.type === 'lexical_declaration' || node.type === 'variable_declaration';
|
||||
|
||||
if (!isFnDecl && !isMethodDef && !isLexDecl) continue;
|
||||
|
||||
// For `export function foo() { ... }`, the JSDoc comment precedes the
|
||||
// wrapping export_statement, not the inner function_declaration.
|
||||
// Walk up to the export_statement so the preceding-sibling search finds it.
|
||||
const lookupNode =
|
||||
(isFnDecl || isLexDecl) && node.parent?.type === 'export_statement' ? node.parent : node;
|
||||
|
||||
// Find the preceding sibling comment.
|
||||
let sibling = lookupNode.previousNamedSibling;
|
||||
while (sibling !== null && sibling.type === 'comment') {
|
||||
const text = sibling.text;
|
||||
if (text.startsWith('/**')) {
|
||||
// Found a JSDoc block.
|
||||
const params = parseJsDocParams(text);
|
||||
const retType = parseJsDocReturn(text);
|
||||
const varType = isLexDecl ? parseJsDocType(text) : null;
|
||||
|
||||
// Determine the anchor node (the function-like node, for hoisting).
|
||||
const anchor = node;
|
||||
|
||||
for (const p of params) {
|
||||
out.push({
|
||||
'@type-binding.name': syntheticCapture('@type-binding.name', anchor, p.name),
|
||||
'@type-binding.type': syntheticCapture('@type-binding.type', anchor, p.type),
|
||||
'@type-binding.parameter': syntheticCapture('@type-binding.parameter', anchor, '1'),
|
||||
});
|
||||
}
|
||||
|
||||
if (retType !== null) {
|
||||
// For named functions, use the function name as the binding name so
|
||||
// `hoistTypeBindingsToModule` knows which function's return type this is.
|
||||
let fnName: string | null = null;
|
||||
if (isFnDecl) {
|
||||
fnName = node.childForFieldName('name')?.text ?? null;
|
||||
} else if (isMethodDef) {
|
||||
// method_definition uses `name:` field for the method name
|
||||
const nameNode = node.childForFieldName('name');
|
||||
if (nameNode?.type === 'property_identifier') fnName = nameNode.text;
|
||||
} else if (isLexDecl) {
|
||||
const declarator = node.namedChild(0);
|
||||
const nameNode = declarator?.childForFieldName('name');
|
||||
if (nameNode?.type === 'identifier') fnName = nameNode.text;
|
||||
}
|
||||
if (fnName !== null) {
|
||||
out.push({
|
||||
'@type-binding.name': syntheticCapture('@type-binding.name', anchor, fnName),
|
||||
'@type-binding.type': syntheticCapture('@type-binding.type', anchor, retType),
|
||||
'@type-binding.return': syntheticCapture('@type-binding.return', anchor, '1'),
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
// @type {T} on let/const/var: `/** @type {User} */ const u = getUser()`.
|
||||
// Emits annotation-strength binding (source = 'annotation') so it
|
||||
// overrides any weaker constructor/alias inference on the same name.
|
||||
if (varType !== null) {
|
||||
for (const declarator of node.namedChildren) {
|
||||
if (declarator === null || declarator.type !== 'variable_declarator') continue;
|
||||
const nameNode = declarator.childForFieldName('name');
|
||||
if (nameNode === null || nameNode.type !== 'identifier') continue;
|
||||
out.push({
|
||||
'@type-binding.name': syntheticCapture('@type-binding.name', nameNode, nameNode.text),
|
||||
'@type-binding.type': syntheticCapture('@type-binding.type', nameNode, varType),
|
||||
'@type-binding.annotation': syntheticCapture(
|
||||
'@type-binding.annotation',
|
||||
nameNode,
|
||||
'1',
|
||||
),
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
break;
|
||||
}
|
||||
sibling = sibling.previousNamedSibling;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ─── Destructuring / for-of / instanceof (shared with TS captures) ───────
|
||||
|
||||
function synthesizeDestructuringBindings(root: SyntaxNode, out: CaptureMatch[]): void {
|
||||
const stack: SyntaxNode[] = [root];
|
||||
for (;;) {
|
||||
const node = stack.pop();
|
||||
if (node === undefined) break;
|
||||
for (const child of node.namedChildren) {
|
||||
if (child !== null) stack.push(child);
|
||||
}
|
||||
if (node.type !== 'variable_declarator') continue;
|
||||
const nameNode = node.childForFieldName('name');
|
||||
const valueNode = node.childForFieldName('value');
|
||||
if (nameNode === null || valueNode === null) continue;
|
||||
if (nameNode.type !== 'object_pattern') continue;
|
||||
if (valueNode.type !== 'identifier') continue;
|
||||
const rhsName = valueNode.text;
|
||||
for (const fieldNode of nameNode.namedChildren) {
|
||||
if (fieldNode === null) continue;
|
||||
if (fieldNode.type === 'shorthand_property_identifier_pattern') {
|
||||
const localName = fieldNode.text;
|
||||
out.push({
|
||||
'@type-binding.name': syntheticCapture('@type-binding.name', fieldNode, localName),
|
||||
'@type-binding.type': syntheticCapture(
|
||||
'@type-binding.type',
|
||||
fieldNode,
|
||||
`${rhsName}.${localName}`,
|
||||
),
|
||||
'@type-binding.destructured': syntheticCapture(
|
||||
'@type-binding.destructured',
|
||||
fieldNode,
|
||||
fieldNode.text,
|
||||
),
|
||||
});
|
||||
} else if (fieldNode.type === 'pair_pattern') {
|
||||
const key = fieldNode.childForFieldName('key');
|
||||
const value = fieldNode.childForFieldName('value');
|
||||
if (key === null || value === null || value.type !== 'identifier') continue;
|
||||
const fieldName = key.text;
|
||||
const localName = value.text;
|
||||
out.push({
|
||||
'@type-binding.name': syntheticCapture('@type-binding.name', value, localName),
|
||||
'@type-binding.type': syntheticCapture(
|
||||
'@type-binding.type',
|
||||
fieldNode,
|
||||
`${rhsName}.${fieldName}`,
|
||||
),
|
||||
'@type-binding.destructured': syntheticCapture(
|
||||
'@type-binding.destructured',
|
||||
fieldNode,
|
||||
fieldNode.text,
|
||||
),
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
function synthesizeForOfMapTupleBindings(root: SyntaxNode, out: CaptureMatch[]): void {
|
||||
const stack: SyntaxNode[] = [root];
|
||||
for (;;) {
|
||||
const node = stack.pop();
|
||||
if (node === undefined) break;
|
||||
for (const child of node.namedChildren) {
|
||||
if (child !== null) stack.push(child);
|
||||
}
|
||||
if (node.type !== 'for_in_statement') continue;
|
||||
const left = node.childForFieldName('left');
|
||||
const right = node.childForFieldName('right');
|
||||
if (left === null || right === null) continue;
|
||||
if (left.type !== 'array_pattern' || right.type !== 'identifier') continue;
|
||||
const rhs = right.text;
|
||||
let slot = 0;
|
||||
for (const child of left.namedChildren) {
|
||||
if (child === null || child.type !== 'identifier') continue;
|
||||
const localName = child.text;
|
||||
out.push({
|
||||
'@type-binding.name': syntheticCapture('@type-binding.name', child, localName),
|
||||
'@type-binding.type': syntheticCapture(
|
||||
'@type-binding.type',
|
||||
child,
|
||||
`__MAP_TUPLE_${slot}__:${rhs}`,
|
||||
),
|
||||
'@type-binding.map-tuple-entry': syntheticCapture(
|
||||
'@type-binding.map-tuple-entry',
|
||||
child,
|
||||
String(slot),
|
||||
),
|
||||
});
|
||||
slot++;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
function synthesizeInstanceofNarrowings(root: SyntaxNode, out: CaptureMatch[]): void {
|
||||
const stack: SyntaxNode[] = [root];
|
||||
for (;;) {
|
||||
const node = stack.pop();
|
||||
if (node === undefined) break;
|
||||
for (const child of node.namedChildren) {
|
||||
if (child !== null) stack.push(child);
|
||||
}
|
||||
if (node.type !== 'if_statement') continue;
|
||||
const cond = node.childForFieldName('condition');
|
||||
if (cond === null) continue;
|
||||
const inner = cond.type === 'parenthesized_expression' ? cond.namedChildren[0] : cond;
|
||||
if (inner === null || inner.type !== 'binary_expression') continue;
|
||||
const op = inner.childForFieldName('operator');
|
||||
const left = inner.childForFieldName('left');
|
||||
const right = inner.childForFieldName('right');
|
||||
if (op === null || left === null || right === null) continue;
|
||||
if (op.type !== 'instanceof') continue;
|
||||
if (left.type !== 'identifier') continue;
|
||||
if (right.type !== 'identifier') continue;
|
||||
const varName = left.text;
|
||||
const typeName = right.text;
|
||||
const cons = node.childForFieldName('consequence');
|
||||
if (cons === null) continue;
|
||||
out.push({
|
||||
'@type-binding.name': syntheticCapture('@type-binding.name', cons, varName),
|
||||
'@type-binding.type': syntheticCapture('@type-binding.type', right, typeName),
|
||||
'@type-binding.instanceof-narrow': syntheticCapture(
|
||||
'@type-binding.instanceof-narrow',
|
||||
cons,
|
||||
'1',
|
||||
),
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
// ─── Constructor field type bindings ─────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Synthesize class-scope type bindings from `this.X = new Y()` assignments
|
||||
* inside constructor method bodies. Covers the traditional ES5+ OOP pattern:
|
||||
*
|
||||
* class User {
|
||||
* constructor() {
|
||||
* /** @type {Address} *\/
|
||||
* this.address = new Address();
|
||||
* }
|
||||
* }
|
||||
*
|
||||
* The emitted `@type-binding.class-field` is hoisted to the Class scope by
|
||||
* `tsBindingScopeFor` so that compound-receiver resolution can look up
|
||||
* `User.address → Address` when resolving `user.address.save()`.
|
||||
*
|
||||
* Type source priority:
|
||||
* 1. JSDoc `@type {T}` comment immediately preceding the statement
|
||||
* 2. `new Y()` constructor inference
|
||||
*/
|
||||
function synthesizeConstructorFieldBindings(root: SyntaxNode, out: CaptureMatch[]): void {
|
||||
const stack: SyntaxNode[] = [root];
|
||||
for (;;) {
|
||||
const node = stack.pop();
|
||||
if (node === undefined) break;
|
||||
for (const child of node.namedChildren) {
|
||||
if (child !== null) stack.push(child);
|
||||
}
|
||||
// Only process constructor method definitions
|
||||
if (node.type !== 'method_definition') continue;
|
||||
const nameNode = node.childForFieldName('name');
|
||||
if (nameNode?.text !== 'constructor') continue;
|
||||
|
||||
const body = node.childForFieldName('body');
|
||||
if (body === null) continue;
|
||||
|
||||
for (const stmt of body.namedChildren) {
|
||||
if (stmt === null || stmt.type !== 'expression_statement') continue;
|
||||
const expr = stmt.namedChild(0);
|
||||
if (expr === null || expr.type !== 'assignment_expression') continue;
|
||||
|
||||
const left = expr.childForFieldName('left');
|
||||
const right = expr.childForFieldName('right');
|
||||
if (left === null || right === null) continue;
|
||||
if (left.type !== 'member_expression') continue;
|
||||
|
||||
const obj = left.childForFieldName('object');
|
||||
const prop = left.childForFieldName('property');
|
||||
if (obj === null || prop === null) continue;
|
||||
if (obj.text !== 'this' || prop.type !== 'property_identifier') continue;
|
||||
|
||||
const fieldName = prop.text;
|
||||
|
||||
// Prefer JSDoc @type annotation on the preceding sibling comment.
|
||||
let typeName: string | null = null;
|
||||
const prevSib: SyntaxNode | null = stmt.previousNamedSibling;
|
||||
if (prevSib !== null && prevSib.type === 'comment') {
|
||||
const m = /@type\s*\{([^}]+)\}/.exec(prevSib.text);
|
||||
if (m?.[1]) typeName = m[1].trim();
|
||||
}
|
||||
// Fall back to constructor inference from `new Y()`.
|
||||
if (typeName === null && right.type === 'new_expression') {
|
||||
const ctor = right.childForFieldName('constructor');
|
||||
if (ctor !== null && ctor.type === 'identifier') typeName = ctor.text;
|
||||
}
|
||||
if (typeName === null) continue;
|
||||
|
||||
out.push({
|
||||
'@type-binding.name': syntheticCapture('@type-binding.name', prop, fieldName),
|
||||
'@type-binding.type': syntheticCapture('@type-binding.type', prop, typeName),
|
||||
// Anchor: positioned inside the constructor body so tsBindingScopeFor
|
||||
// can walk up from the Function (constructor) scope to the Class scope.
|
||||
'@type-binding.class-field': syntheticCapture('@type-binding.class-field', stmt, '1'),
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ─── Main emitter ──────────────────────────────────────────────────────────
|
||||
|
||||
export function emitJsScopeCaptures(
|
||||
sourceText: string,
|
||||
filePath: string,
|
||||
cachedTree?: unknown,
|
||||
): readonly CaptureMatch[] {
|
||||
let tree = cachedTree as ReturnType<ReturnType<typeof getJsParser>['parse']> | undefined;
|
||||
if (tree !== undefined && !jsCachedTreeMatchesGrammar(tree)) {
|
||||
tree = undefined;
|
||||
}
|
||||
if (tree === undefined) {
|
||||
tree = parseSourceSafe(getJsParser(filePath), sourceText, undefined, {
|
||||
bufferSize: getTreeSitterBufferSize(sourceText),
|
||||
});
|
||||
}
|
||||
|
||||
const rawMatches = getJsScopeQuery(filePath).matches(tree.rootNode);
|
||||
const out: CaptureMatch[] = [];
|
||||
|
||||
for (const m of rawMatches) {
|
||||
const grouped: Record<string, Capture> = {};
|
||||
for (const c of m.captures) {
|
||||
const tag = '@' + c.name;
|
||||
grouped[tag] = nodeToCapture(tag, c.node);
|
||||
}
|
||||
if (Object.keys(grouped).length === 0) continue;
|
||||
|
||||
// Decompose ESM import_statement / re-export export_statement.
|
||||
if (grouped['@import.statement'] !== undefined) {
|
||||
const stmtCapture = grouped['@import.statement'];
|
||||
const stmtNode =
|
||||
findNodeAtRange(tree.rootNode, stmtCapture.range, 'import_statement') ??
|
||||
findNodeAtRange(tree.rootNode, stmtCapture.range, 'export_statement');
|
||||
if (stmtNode !== null) {
|
||||
const decomposed = splitImportStatement(stmtNode);
|
||||
for (const d of decomposed) out.push(d);
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
// Decompose dynamic import() calls.
|
||||
if (grouped['@import.dynamic'] !== undefined) {
|
||||
const dynCapture = grouped['@import.dynamic'];
|
||||
const callNode = findNodeAtRange(tree.rootNode, dynCapture.range, 'call_expression');
|
||||
if (callNode !== null) {
|
||||
const decomposed = splitImportStatement(callNode);
|
||||
for (const d of decomposed) out.push(d);
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
// Filter @reference.read.member false-positives.
|
||||
if (grouped['@reference.read.member'] !== undefined) {
|
||||
const anchor = grouped['@reference.read.member'];
|
||||
const memberNode = findNodeAtRange(tree.rootNode, anchor.range, 'member_expression');
|
||||
if (memberNode === null || !shouldEmitReadMember(memberNode)) {
|
||||
continue;
|
||||
}
|
||||
}
|
||||
|
||||
// Synthesize arity metadata on function-like declarations.
|
||||
const declAnchor = pickFirstDefined(grouped, FUNCTION_DECL_TAGS);
|
||||
if (declAnchor !== undefined) {
|
||||
const fnNode = findFunctionNode(tree.rootNode, declAnchor.range);
|
||||
if (fnNode !== null) {
|
||||
const arity = computeTsArityMetadata(fnNode);
|
||||
if (arity.parameterCount !== undefined) {
|
||||
grouped['@declaration.parameter-count'] = syntheticCapture(
|
||||
'@declaration.parameter-count',
|
||||
fnNode,
|
||||
String(arity.parameterCount),
|
||||
);
|
||||
}
|
||||
if (arity.requiredParameterCount !== undefined) {
|
||||
grouped['@declaration.required-parameter-count'] = syntheticCapture(
|
||||
'@declaration.required-parameter-count',
|
||||
fnNode,
|
||||
String(arity.requiredParameterCount),
|
||||
);
|
||||
}
|
||||
if (arity.parameterTypes !== undefined) {
|
||||
grouped['@declaration.parameter-types'] = syntheticCapture(
|
||||
'@declaration.parameter-types',
|
||||
fnNode,
|
||||
JSON.stringify(arity.parameterTypes),
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Synthesize @reference.arity on callsites.
|
||||
const callAnchor = pickFirstDefined(grouped, CALL_TAGS);
|
||||
if (callAnchor !== undefined && grouped['@reference.arity'] === undefined) {
|
||||
const callNode =
|
||||
findNodeAtRange(tree.rootNode, callAnchor.range, 'call_expression') ??
|
||||
findNodeAtRange(tree.rootNode, callAnchor.range, 'new_expression');
|
||||
if (callNode !== null) {
|
||||
const argList = callNode.childForFieldName('arguments');
|
||||
const args: SyntaxNode[] =
|
||||
argList === null
|
||||
? []
|
||||
: argList.namedChildren.filter(
|
||||
(c): c is SyntaxNode => c !== null && c.type !== 'comment',
|
||||
);
|
||||
grouped['@reference.arity'] = syntheticCapture(
|
||||
'@reference.arity',
|
||||
callNode,
|
||||
String(args.length),
|
||||
);
|
||||
grouped['@reference.parameter-types'] = syntheticCapture(
|
||||
'@reference.parameter-types',
|
||||
callNode,
|
||||
JSON.stringify(args.map(inferArgType)),
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
out.push(grouped);
|
||||
|
||||
// Synthesize `this` receiver type-bindings on class member functions.
|
||||
const scopeFnAnchor = grouped['@scope.function'];
|
||||
if (scopeFnAnchor !== undefined) {
|
||||
const fnNode = findFunctionNode(tree.rootNode, scopeFnAnchor.range);
|
||||
if (fnNode !== null) {
|
||||
const synth = synthesizeTsReceiverBinding(fnNode);
|
||||
if (synth !== null) out.push(synth);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Post-query synthesis passes.
|
||||
synthesizeCjsImports(tree.rootNode, out);
|
||||
synthesizeJsDocBindings(tree.rootNode, out);
|
||||
synthesizeConstructorFieldBindings(tree.rootNode, out);
|
||||
synthesizeDestructuringBindings(tree.rootNode, out);
|
||||
synthesizeForOfMapTupleBindings(tree.rootNode, out);
|
||||
synthesizeInstanceofNarrowings(tree.rootNode, out);
|
||||
|
||||
return out;
|
||||
}
|
||||
@@ -0,0 +1,72 @@
|
||||
/**
|
||||
* Import-target resolver for JavaScript.
|
||||
*
|
||||
* Delegates to the TypeScript `resolveTsTarget` standard-strategy resolver
|
||||
* with `language: SupportedLanguages.JavaScript` so the resolver tries
|
||||
* `.js` / `.jsx` extensions in addition to (or instead of) `.ts` / `.tsx`.
|
||||
*
|
||||
* The `TsResolveContext.language` flag already exists in `import-target.ts`
|
||||
* and the resolver (`resolveImportPath`) already branches on it — this
|
||||
* adapter just wires the right value in.
|
||||
*
|
||||
* CJS `require()` calls reference the same module-path strings as ESM
|
||||
* `import` statements, so the resolver handles them uniformly without any
|
||||
* CJS-specific logic here.
|
||||
*
|
||||
* No `tsconfig.json` path-alias support (JavaScript projects don't use
|
||||
* `tsconfig.json` compilerOptions.paths in general). Projects that DO use
|
||||
* tsconfig-based aliases alongside JavaScript can still resolve via the
|
||||
* standard extension-suffix fallback; the alias branch is a no-op when
|
||||
* `tsconfigPaths` is null.
|
||||
*/
|
||||
|
||||
import { SupportedLanguages } from 'gitnexus-shared';
|
||||
import { resolveTsTarget, type TsResolveContext } from '../typescript/import-target.js';
|
||||
|
||||
export type JsResolveContext = TsResolveContext;
|
||||
|
||||
type PassCache = {
|
||||
readonly key: ReadonlySet<string>;
|
||||
readonly allFilePaths: Set<string>;
|
||||
readonly allFileList: readonly string[];
|
||||
readonly normalizedFileList: readonly string[];
|
||||
readonly resolveCache: Map<string, string | null>;
|
||||
};
|
||||
|
||||
/**
|
||||
* Build a memoized `resolveImportTarget` adapter for JavaScript.
|
||||
* Caches the derived arrays and per-pass resolve cache across
|
||||
* `resolveImportTarget` calls within a single workspace pass.
|
||||
*/
|
||||
export function makeJsResolveImportTarget(): (
|
||||
targetRaw: string,
|
||||
fromFile: string,
|
||||
allFilePaths: ReadonlySet<string>,
|
||||
resolutionConfig?: unknown,
|
||||
) => string | readonly string[] | null {
|
||||
let cached: PassCache | null = null;
|
||||
|
||||
return (targetRaw, fromFile, allFilePaths) => {
|
||||
if (cached === null || cached.key !== allFilePaths) {
|
||||
const allFileList = Array.from(allFilePaths);
|
||||
cached = {
|
||||
key: allFilePaths,
|
||||
allFilePaths: new Set(allFilePaths),
|
||||
allFileList,
|
||||
normalizedFileList: allFileList.map((f) => f.toLowerCase()),
|
||||
resolveCache: new Map(),
|
||||
};
|
||||
}
|
||||
|
||||
const ws: JsResolveContext = {
|
||||
fromFile,
|
||||
language: SupportedLanguages.JavaScript,
|
||||
allFilePaths: cached.allFilePaths,
|
||||
allFileList: cached.allFileList,
|
||||
normalizedFileList: cached.normalizedFileList,
|
||||
resolveCache: cached.resolveCache,
|
||||
tsconfigPaths: null,
|
||||
};
|
||||
return resolveTsTarget(targetRaw, ws);
|
||||
};
|
||||
}
|
||||
@@ -0,0 +1,49 @@
|
||||
/**
|
||||
* JavaScript scope-resolution hooks (RFC #909 Ring 3, issue #928).
|
||||
*
|
||||
* Public API barrel. Consumers should import from this file rather
|
||||
* than the individual modules.
|
||||
*
|
||||
* Module layout (each file is a single concern):
|
||||
*
|
||||
* - `query.ts` — JS scope query string + lazy parser/query
|
||||
* singletons (`getJsParser`, `getJsScopeQuery`)
|
||||
* - `captures.ts` — `emitJsScopeCaptures` — runs the JS scope query,
|
||||
* synthesizes CJS require() imports and JSDoc-
|
||||
* derived type bindings, delegates arity synthesis
|
||||
* and destructuring/instanceof passes to shared
|
||||
* or TypeScript utilities
|
||||
* - `interpret.ts` — `interpretJsImport` / `interpretJsTypeBinding`
|
||||
* (delegate to TypeScript interpreters — same
|
||||
* capture-marker vocabulary)
|
||||
* - `simple-hooks.ts` — `jsBindingScopeFor` (var hoisting),
|
||||
* `jsImportOwningScope`, `jsReceiverBinding`
|
||||
* (all delegate to TypeScript counterparts)
|
||||
* - `merge-bindings.ts` — `jsMergeBindings` (LEGB via typescriptMergeBindings)
|
||||
* - `arity.ts` — `jsArityCompatibility` (delegates to TS function)
|
||||
* - `import-target.ts` — `makeJsResolveImportTarget` (memoized adapter)
|
||||
* - `scope-resolver.ts` — `javascriptScopeResolver` wiring object
|
||||
*
|
||||
* ## Known limitations
|
||||
*
|
||||
* 1. **JSDoc coverage** — `@param {T} name`, `@returns {T}` / `@return {T}`,
|
||||
* and `@type {T}` on variable declarations are synthesized. `@typedef`
|
||||
* is not yet synthesized (tracked in #1646).
|
||||
* 2. **CJS chained destructuring** — `const { X: { Y } } = require(...)`
|
||||
* (nested destructuring) emits only the outer `X` binding; `Y` is not
|
||||
* resolved.
|
||||
* 3. **Dynamic require** — `require(computedPath)` is skipped (non-literal
|
||||
* argument — cannot statically resolve the target).
|
||||
* 4. **`module.exports` / `exports.X`** — CJS export forms are not yet
|
||||
* modeled as re-exports. The finalize algorithm treats the exporting
|
||||
* module as a namespace; importers that do `const X = require('./m')`
|
||||
* bind the module namespace, and member-call resolution walks the
|
||||
* class graph from there.
|
||||
*/
|
||||
|
||||
export { emitJsScopeCaptures } from './captures.js';
|
||||
export { interpretJsImport, interpretJsTypeBinding } from './interpret.js';
|
||||
export { jsMergeBindings } from './merge-bindings.js';
|
||||
export { jsArityCompatibility } from './arity.js';
|
||||
export { makeJsResolveImportTarget } from './import-target.js';
|
||||
export { jsBindingScopeFor, jsImportOwningScope, jsReceiverBinding } from './simple-hooks.js';
|
||||
@@ -0,0 +1,45 @@
|
||||
/**
|
||||
* Capture-match → semantic-shape interpreters for JavaScript.
|
||||
*
|
||||
* `interpretJsImport` delegates to `interpretTsImport` for all cases
|
||||
* because `emitJsScopeCaptures` synthesizes the same
|
||||
* `@import.kind/name/alias/source` markers for both ESM and CJS imports.
|
||||
*
|
||||
* The `@import.kind` values emitted for CJS by `captures.ts`:
|
||||
*
|
||||
* - `'named'` : `const { X } = require('./m')` → named import
|
||||
* - `'named-alias'` : `const { X: Y } = require('./m')` → aliased import
|
||||
* - `'namespace'` : `const X = require('./m')` → namespace import
|
||||
* - `'side-effect'` : `require('./m')` bare expression → side-effect
|
||||
*
|
||||
* These match the kinds `interpretTsImport` already handles for ESM
|
||||
* (`import { X }`, `import { X as Y }`, `import * as X`, `import './m'`),
|
||||
* so no new branch is needed here.
|
||||
*
|
||||
* `interpretJsTypeBinding` handles the JS-only `@type-binding.class-field`
|
||||
* tag before delegating to `interpretTsTypeBinding`. The class-field tag
|
||||
* is emitted by `synthesizeConstructorFieldBindings` and should produce
|
||||
* `source = 'annotation'` — the same strength as an explicit type
|
||||
* annotation. Remapping it to `@type-binding.annotation` achieves this
|
||||
* without adding a JS-specific branch to the shared TS interpreter
|
||||
* (DoD.md §2.2).
|
||||
*/
|
||||
|
||||
import type { CaptureMatch, ParsedImport, ParsedTypeBinding } from 'gitnexus-shared';
|
||||
import { interpretTsImport, interpretTsTypeBinding } from '../typescript/interpret.js';
|
||||
|
||||
export function interpretJsImport(captures: CaptureMatch): ParsedImport | null {
|
||||
return interpretTsImport(captures);
|
||||
}
|
||||
|
||||
export function interpretJsTypeBinding(captures: CaptureMatch): ParsedTypeBinding | null {
|
||||
// @type-binding.class-field is a JS-only tag emitted by
|
||||
// synthesizeConstructorFieldBindings. Remap it to the standard
|
||||
// @type-binding.annotation tag so interpretTsTypeBinding assigns
|
||||
// source = 'annotation' without a JS-specific branch in shared code.
|
||||
if (captures['@type-binding.class-field'] !== undefined) {
|
||||
const { '@type-binding.class-field': classField, ...rest } = captures;
|
||||
return interpretTsTypeBinding({ ...rest, '@type-binding.annotation': classField });
|
||||
}
|
||||
return interpretTsTypeBinding(captures);
|
||||
}
|
||||
@@ -0,0 +1,21 @@
|
||||
/**
|
||||
* Binding-merge precedence for JavaScript.
|
||||
*
|
||||
* JavaScript has no TypeScript declaration-merging (no `interface + class`
|
||||
* coexisting in the same scope, no `namespace + class` dual-space declarations).
|
||||
* However, `typescriptMergeBindings` handles these by falling back to
|
||||
* `['value']` for any `NodeLabel` not explicitly mapped to multiple spaces —
|
||||
* which is what every JavaScript declaration produces. The result is pure
|
||||
* LEGB precedence without any cross-space logic, which is exactly what
|
||||
* JavaScript needs.
|
||||
*
|
||||
* Reuse rather than reimplementing to keep the single source of truth for
|
||||
* the tier (local 0 / import-namespace-reexport 1 / wildcard 2) ordering.
|
||||
*/
|
||||
|
||||
import type { BindingRef } from 'gitnexus-shared';
|
||||
import { typescriptMergeBindings } from '../typescript/merge-bindings.js';
|
||||
|
||||
export function jsMergeBindings(bindings: readonly BindingRef[]): readonly BindingRef[] {
|
||||
return typescriptMergeBindings(bindings);
|
||||
}
|
||||
@@ -0,0 +1,421 @@
|
||||
/**
|
||||
* Tree-sitter query for JavaScript scope captures (RFC §5.1, Ring 3).
|
||||
*
|
||||
* Subset of the TypeScript scope query (`languages/typescript/query.ts`)
|
||||
* compiled against `tree-sitter-javascript`. TypeScript-only node types
|
||||
* (`interface_declaration`, `type_alias_declaration`, `enum_declaration`,
|
||||
* `internal_module`, `abstract_class_declaration`, `function_signature`,
|
||||
* `method_signature`, `abstract_method_signature`, `type_annotation`,
|
||||
* `public_field_definition`) are dropped because:
|
||||
*
|
||||
* 1. The JS grammar doesn't define them — the query compiler would
|
||||
* throw `InvalidNodeType` if they were included.
|
||||
* 2. JavaScript has no static type annotations, so the `@type-binding.*`
|
||||
* patterns derived from TS annotation nodes don't apply.
|
||||
*
|
||||
* What IS shared with the TypeScript query:
|
||||
*
|
||||
* - Scope patterns: `program`, `class_declaration`, `(class)` (the JS
|
||||
* grammar node for class expressions — NOT `class_expression`, which
|
||||
* does not exist in `tree-sitter-javascript`), `function_declaration`,
|
||||
* `generator_function_declaration`, `function_expression`,
|
||||
* `arrow_function`, `method_definition`.
|
||||
* - Declaration patterns for functions, classes, const/let/var,
|
||||
* object-property arrows (Zustand, TanStack, etc.), and HOC-wrapped
|
||||
* variable declarations (forwardRef / memo / useCallback / useMemo).
|
||||
* - Import patterns: `import_statement`, `export_statement` re-exports,
|
||||
* and dynamic `import()` (represented as `call_expression(import)` in
|
||||
* both grammars — the `import` leaf node exists in tree-sitter-javascript
|
||||
* as well as tree-sitter-typescript).
|
||||
* - Type-binding patterns that work without static annotations:
|
||||
* constructor inference (`new User()`), call-result alias
|
||||
* (`const u = getUser()`), member-access alias (`const a = u.addr`),
|
||||
* identifier alias, assignment rebind, and for-of element bindings.
|
||||
* JSDoc-derived type bindings (`@param {User} u`, `@returns {User}`)
|
||||
* are handled separately in `captures.ts` via comment-node scanning.
|
||||
* - Reference patterns: free calls, member calls, constructor calls,
|
||||
* write-access, read-access, and dynamic import.
|
||||
*
|
||||
* CJS `require()` is NOT captured here; it is handled in `captures.ts`
|
||||
* by scanning parent context (destructured vs. namespace) of `call_expression`
|
||||
* nodes whose callee is the identifier `require`.
|
||||
*
|
||||
* Grammar version: `tree-sitter-javascript` pinned in gitnexus/package.json.
|
||||
*
|
||||
* Exposes lazy `Parser` and `Query` singletons so callers don't pay
|
||||
* tree-sitter init cost per file.
|
||||
*/
|
||||
|
||||
import Parser from 'tree-sitter';
|
||||
import JS from 'tree-sitter-javascript';
|
||||
|
||||
const JS_GRAMMAR = JS as Parameters<Parser['setLanguage']>[0];
|
||||
|
||||
/** True when the file should be parsed with the JSX-extended query. */
|
||||
function isJsxFile(filePath: string): boolean {
|
||||
return filePath.endsWith('.jsx');
|
||||
}
|
||||
|
||||
const JAVASCRIPT_SCOPE_QUERY = `
|
||||
;; Scopes — module / class-likes / function-likes
|
||||
(program) @scope.module
|
||||
|
||||
(class_declaration) @scope.class
|
||||
(class) @scope.class
|
||||
|
||||
(function_declaration) @scope.function
|
||||
(generator_function_declaration) @scope.function
|
||||
(function_expression) @scope.function
|
||||
(arrow_function) @scope.function
|
||||
(method_definition) @scope.function
|
||||
|
||||
;; Declarations — classes
|
||||
(class_declaration
|
||||
name: (identifier) @declaration.name) @declaration.class
|
||||
|
||||
;; Declarations — methods (inside class bodies)
|
||||
(method_definition
|
||||
name: (property_identifier) @declaration.name) @declaration.method
|
||||
|
||||
;; Declarations — class fields (JS uses field_definition, not public_field_definition)
|
||||
(field_definition
|
||||
property: (property_identifier) @declaration.name) @declaration.property
|
||||
|
||||
;; Declarations — free functions
|
||||
(function_declaration
|
||||
name: (identifier) @declaration.name) @declaration.function
|
||||
|
||||
(generator_function_declaration
|
||||
name: (identifier) @declaration.name) @declaration.function
|
||||
|
||||
;; Arrow / function-expression assigned to a const/let/var.
|
||||
;; Anchor discipline: @declaration.function sits on the INNER arrow or
|
||||
;; function_expression, NOT on the lexical_declaration wrapper. This
|
||||
;; aligns anchor.range with the @scope.function range so
|
||||
;; pass2AttachDeclarations resolves the innermost scope correctly and
|
||||
;; resolveCallerGraphId walks up to the right caller anchor.
|
||||
(lexical_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name
|
||||
value: (arrow_function) @declaration.function))
|
||||
|
||||
(lexical_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name
|
||||
value: (function_expression) @declaration.function))
|
||||
|
||||
(export_statement
|
||||
declaration: (lexical_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name
|
||||
value: (arrow_function) @declaration.function)))
|
||||
|
||||
(export_statement
|
||||
declaration: (lexical_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name
|
||||
value: (function_expression) @declaration.function)))
|
||||
|
||||
(variable_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name
|
||||
value: (arrow_function) @declaration.function))
|
||||
|
||||
(variable_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name
|
||||
value: (function_expression) @declaration.function))
|
||||
|
||||
;; Object-property arrows / function expressions named by their pair key.
|
||||
;; Same anchor discipline as the lexical_declaration block above: the
|
||||
;; @declaration.function capture must sit on the INNER arrow/fn-expression.
|
||||
(pair
|
||||
key: (property_identifier) @declaration.name
|
||||
value: (arrow_function) @declaration.function)
|
||||
|
||||
(pair
|
||||
key: (property_identifier) @declaration.name
|
||||
value: (function_expression) @declaration.function)
|
||||
|
||||
(pair
|
||||
key: (string (string_fragment) @declaration.name)
|
||||
value: (arrow_function) @declaration.function)
|
||||
|
||||
(pair
|
||||
key: (string (string_fragment) @declaration.name)
|
||||
value: (function_expression) @declaration.function)
|
||||
|
||||
;; HOC-wrapped variable declarations: const X = HOC((args) => { ... }).
|
||||
;; Covers React.forwardRef, memo, useCallback, useMemo, observer,
|
||||
;; debounce, and any user-defined HOC factory.
|
||||
(lexical_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name
|
||||
value: (call_expression
|
||||
arguments: (arguments
|
||||
(arrow_function) @declaration.function))))
|
||||
|
||||
(lexical_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name
|
||||
value: (call_expression
|
||||
arguments: (arguments
|
||||
(function_expression) @declaration.function))))
|
||||
|
||||
(export_statement
|
||||
declaration: (lexical_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name
|
||||
value: (call_expression
|
||||
arguments: (arguments
|
||||
(arrow_function) @declaration.function)))))
|
||||
|
||||
(export_statement
|
||||
declaration: (lexical_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name
|
||||
value: (call_expression
|
||||
arguments: (arguments
|
||||
(function_expression) @declaration.function)))))
|
||||
|
||||
(variable_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name
|
||||
value: (call_expression
|
||||
arguments: (arguments
|
||||
(arrow_function) @declaration.function))))
|
||||
|
||||
(variable_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name
|
||||
value: (call_expression
|
||||
arguments: (arguments
|
||||
(function_expression) @declaration.function))))
|
||||
|
||||
;; Variable / constant declarations (non-function values).
|
||||
(lexical_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name)) @declaration.const
|
||||
|
||||
(export_statement
|
||||
declaration: (lexical_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name))) @declaration.const
|
||||
|
||||
(variable_declaration
|
||||
(variable_declarator
|
||||
name: (identifier) @declaration.name)) @declaration.variable
|
||||
|
||||
;; Imports (ESM) — single anchor per statement; decomposer emits per-specifier markers.
|
||||
(import_statement) @import.statement
|
||||
|
||||
;; Re-exports with a source clause.
|
||||
(export_statement
|
||||
source: (string)) @import.statement
|
||||
|
||||
;; Dynamic imports: import('./m') — tree-sitter-javascript represents this
|
||||
;; as call_expression with a named import leaf as the function field,
|
||||
;; identical to tree-sitter-typescript.
|
||||
(call_expression
|
||||
function: (import)) @import.dynamic
|
||||
|
||||
;; ── Type bindings (no static annotations in JS; inferred from AST shape) ──
|
||||
|
||||
;; Constructor-inferred: const u = new User()
|
||||
(variable_declarator
|
||||
name: (identifier) @type-binding.name
|
||||
value: (new_expression
|
||||
constructor: (identifier) @type-binding.type)) @type-binding.constructor
|
||||
|
||||
;; Qualified constructor: const u = new models.User()
|
||||
(variable_declarator
|
||||
name: (identifier) @type-binding.name
|
||||
value: (new_expression
|
||||
constructor: (member_expression) @type-binding.type)) @type-binding.constructor
|
||||
|
||||
;; Call-result alias: const u = getUser()
|
||||
(variable_declarator
|
||||
name: (identifier) @type-binding.name
|
||||
value: (call_expression
|
||||
function: (identifier) @type-binding.type)) @type-binding.alias
|
||||
|
||||
;; Member-call alias: const u = svc.getUser()
|
||||
(variable_declarator
|
||||
name: (identifier) @type-binding.name
|
||||
value: (call_expression
|
||||
function: (member_expression) @type-binding.type)) @type-binding.alias
|
||||
|
||||
;; Await chain: const u = await getUser() / await svc.getUser()
|
||||
(variable_declarator
|
||||
name: (identifier) @type-binding.name
|
||||
value: (await_expression
|
||||
(call_expression
|
||||
function: (identifier) @type-binding.type))) @type-binding.alias
|
||||
|
||||
(variable_declarator
|
||||
name: (identifier) @type-binding.name
|
||||
value: (await_expression
|
||||
(call_expression
|
||||
function: (member_expression) @type-binding.type))) @type-binding.alias
|
||||
|
||||
;; Member-access alias: const addr = user.address
|
||||
(variable_declarator
|
||||
name: (identifier) @type-binding.name
|
||||
value: (member_expression) @type-binding.type) @type-binding.member-alias
|
||||
|
||||
;; Identifier alias: const alias = user
|
||||
(variable_declarator
|
||||
name: (identifier) @type-binding.name
|
||||
value: (identifier) @type-binding.type) @type-binding.alias
|
||||
|
||||
;; Assignment rebind: u = new User() / u = getUser()
|
||||
(assignment_expression
|
||||
left: (identifier) @type-binding.name
|
||||
right: (new_expression
|
||||
constructor: (identifier) @type-binding.type)) @type-binding.constructor
|
||||
|
||||
(assignment_expression
|
||||
left: (identifier) @type-binding.name
|
||||
right: (call_expression
|
||||
function: (identifier) @type-binding.type)) @type-binding.alias
|
||||
|
||||
(assignment_expression
|
||||
left: (identifier) @type-binding.name
|
||||
right: (identifier) @type-binding.type) @type-binding.alias
|
||||
|
||||
;; For-of element: for (const u of users) / for (const u of getUsers())
|
||||
(for_in_statement
|
||||
left: (identifier) @type-binding.name
|
||||
right: (identifier) @type-binding.type) @type-binding.alias
|
||||
|
||||
(for_in_statement
|
||||
left: (identifier) @type-binding.name
|
||||
right: (call_expression
|
||||
function: (identifier) @type-binding.type)) @type-binding.alias
|
||||
|
||||
(for_in_statement
|
||||
left: (identifier) @type-binding.name
|
||||
right: (call_expression
|
||||
function: (member_expression) @type-binding.type)) @type-binding.alias
|
||||
|
||||
(for_in_statement
|
||||
left: (identifier) @type-binding.name
|
||||
right: (member_expression
|
||||
property: (property_identifier) @type-binding.type)) @type-binding.alias
|
||||
|
||||
;; ── References ────────────────────────────────────────────────────────────
|
||||
|
||||
;; Free calls: fn(args). The dynamic-import filter runs in captures.ts.
|
||||
(call_expression
|
||||
function: (identifier) @reference.name) @reference.call.free
|
||||
|
||||
;; Awaited free call: await fn<T>(...) re-associated by tree-sitter.
|
||||
(call_expression
|
||||
function: (await_expression
|
||||
(identifier) @reference.name)) @reference.call.free
|
||||
|
||||
;; Member calls: obj.method() (includes optional chain).
|
||||
(call_expression
|
||||
function: (member_expression
|
||||
object: (_) @reference.receiver
|
||||
property: (property_identifier) @reference.name)) @reference.call.member
|
||||
|
||||
;; Awaited member call: await svc.m<T>(...)
|
||||
(call_expression
|
||||
function: (await_expression
|
||||
(member_expression
|
||||
object: (_) @reference.receiver
|
||||
property: (property_identifier) @reference.name))) @reference.call.member
|
||||
|
||||
;; Constructor calls: new User() / new ns.User()
|
||||
(new_expression
|
||||
constructor: (identifier) @reference.name) @reference.call.constructor
|
||||
|
||||
(new_expression
|
||||
constructor: (member_expression) @reference.call.constructor.qualified) @reference.call.constructor
|
||||
|
||||
;; Write access: obj.field = value
|
||||
(assignment_expression
|
||||
left: (member_expression
|
||||
object: (_) @reference.receiver
|
||||
property: (property_identifier) @reference.name)) @reference.write.member
|
||||
|
||||
(augmented_assignment_expression
|
||||
left: (member_expression
|
||||
object: (_) @reference.receiver
|
||||
property: (property_identifier) @reference.name)) @reference.write.member
|
||||
|
||||
;; Read access: obj.field (in read context; captures.ts filters non-reads).
|
||||
(member_expression
|
||||
object: (_) @reference.receiver
|
||||
property: (property_identifier) @reference.name) @reference.read.member
|
||||
`;
|
||||
|
||||
/** JSX-only suffix — appended when compiling against the JSX grammar for .jsx files. */
|
||||
const JSX_QUERY_SUFFIX = `
|
||||
;; <Foo />
|
||||
((jsx_self_closing_element
|
||||
name: (identifier) @reference.name) @reference.call.free
|
||||
(#match? @reference.name "^[A-Z]"))
|
||||
|
||||
;; <Foo> ... </Foo>
|
||||
((jsx_opening_element
|
||||
name: (identifier) @reference.name) @reference.call.free
|
||||
(#match? @reference.name "^[A-Z]"))
|
||||
|
||||
;; <Foo.Bar />
|
||||
(jsx_self_closing_element
|
||||
name: (member_expression
|
||||
object: (_) @reference.receiver
|
||||
property: (property_identifier) @reference.name)) @reference.call.member
|
||||
|
||||
(jsx_opening_element
|
||||
name: (member_expression
|
||||
object: (_) @reference.receiver
|
||||
property: (property_identifier) @reference.name)) @reference.call.member
|
||||
`;
|
||||
|
||||
let _jsParser: Parser | null = null;
|
||||
let _jsQuery: Parser.Query | null = null;
|
||||
let _jsxParser: Parser | null = null;
|
||||
let _jsxQuery: Parser.Query | null = null;
|
||||
|
||||
export function getJsParser(filePath?: string): Parser {
|
||||
// JSX files use the same JavaScript grammar in tree-sitter-javascript;
|
||||
// both .js and .jsx parse with the same grammar object. We keep separate
|
||||
// singletons only to mirror the TypeScript pattern and in case a future
|
||||
// version of the grammar diverges.
|
||||
if (filePath !== undefined && isJsxFile(filePath)) {
|
||||
if (_jsxParser === null) {
|
||||
_jsxParser = new Parser();
|
||||
_jsxParser.setLanguage(JS_GRAMMAR);
|
||||
}
|
||||
return _jsxParser;
|
||||
}
|
||||
if (_jsParser === null) {
|
||||
_jsParser = new Parser();
|
||||
_jsParser.setLanguage(JS_GRAMMAR);
|
||||
}
|
||||
return _jsParser;
|
||||
}
|
||||
|
||||
export function getJsScopeQuery(filePath?: string): Parser.Query {
|
||||
if (filePath !== undefined && isJsxFile(filePath)) {
|
||||
if (_jsxQuery === null) {
|
||||
_jsxQuery = new Parser.Query(JS_GRAMMAR, JAVASCRIPT_SCOPE_QUERY + JSX_QUERY_SUFFIX);
|
||||
}
|
||||
return _jsxQuery;
|
||||
}
|
||||
if (_jsQuery === null) {
|
||||
_jsQuery = new Parser.Query(JS_GRAMMAR, JAVASCRIPT_SCOPE_QUERY);
|
||||
}
|
||||
return _jsQuery;
|
||||
}
|
||||
|
||||
/** Validate that a cached Tree was produced by the JS grammar. */
|
||||
export function jsCachedTreeMatchesGrammar(tree: unknown): boolean {
|
||||
// eslint-disable-next-line @typescript-eslint/no-explicit-any
|
||||
const lang = (tree as any)?.getLanguage?.();
|
||||
if (lang === undefined || lang === null) return true;
|
||||
return lang === JS_GRAMMAR;
|
||||
}
|
||||
@@ -0,0 +1,89 @@
|
||||
/**
|
||||
* JavaScript `ScopeResolver` registered in `SCOPE_RESOLVERS` and
|
||||
* consumed by the generic `runScopeResolution` orchestrator
|
||||
* (RFC #909 Ring 3, issue #928).
|
||||
*
|
||||
* Follows the same minimal wiring-only pattern as TypeScript (the third
|
||||
* migration). Per-hook logic lives in sibling modules:
|
||||
*
|
||||
* - `query.ts` — JS scope query + parser/query singletons
|
||||
* - `captures.ts` — `emitJsScopeCaptures` (JS grammar, CJS, JSDoc)
|
||||
* - `interpret.ts` — `interpretJsImport` (delegates to TS interpreter)
|
||||
* - `simple-hooks.ts` — `jsBindingScopeFor`, `jsImportOwningScope`,
|
||||
* `jsReceiverBinding` (all delegate to TS hooks)
|
||||
* - `merge-bindings.ts` — `jsMergeBindings` (delegates to TS function)
|
||||
* - `arity.ts` — `jsArityCompatibility` (delegates to TS function)
|
||||
* - `import-target.ts` — `makeJsResolveImportTarget` (TS resolver, JS extensions)
|
||||
*
|
||||
* See `./index.ts` for the full per-module rationale.
|
||||
*
|
||||
* ## Key differences from TypeScript resolver
|
||||
*
|
||||
* - `fieldFallbackOnMethodLookup: true` — JavaScript is dynamically typed;
|
||||
* the field-fallback heuristic is ENABLED (unlike TypeScript, which
|
||||
* disables it because the type-binding layer is precise).
|
||||
* - `allowGlobalFreeCallFallback: true` — CJS `require` patterns and
|
||||
* global helpers (e.g. `process`, `console`) benefit from workspace-
|
||||
* wide unique-name fallback. TypeScript uses explicit imports.
|
||||
* - `loadResolutionConfig` is omitted — JavaScript projects don't use
|
||||
* `tsconfig.json` path aliases in general. `tsconfigPaths: null` is
|
||||
* threaded through the resolver adapter.
|
||||
* - `hoistTypeBindingsToModule: true` — JSDoc `@returns {T}` bindings are
|
||||
* synthesized on the function scope and hoisted, matching TypeScript's
|
||||
* method return-type hoisting strategy for cross-file chain resolution.
|
||||
*/
|
||||
|
||||
import type { ParsedFile } from 'gitnexus-shared';
|
||||
import { SupportedLanguages } from 'gitnexus-shared';
|
||||
import { buildMro, defaultLinearize } from '../../scope-resolution/passes/mro.js';
|
||||
import { populateClassOwnedMembers } from '../../scope-resolution/scope/walkers.js';
|
||||
import type { ScopeResolver } from '../../scope-resolution/contract/scope-resolver.js';
|
||||
import { javascriptProvider } from '../typescript.js';
|
||||
import { jsMergeBindings } from './merge-bindings.js';
|
||||
import { jsArityCompatibility } from './arity.js';
|
||||
import { makeJsResolveImportTarget } from './import-target.js';
|
||||
|
||||
const javascriptScopeResolver: ScopeResolver = {
|
||||
language: SupportedLanguages.JavaScript,
|
||||
languageProvider: javascriptProvider,
|
||||
importEdgeReason: 'javascript-scope: import',
|
||||
|
||||
resolveImportTarget: makeJsResolveImportTarget(),
|
||||
|
||||
// JavaScript LEGB — same tier ordering as TypeScript; no declaration-
|
||||
// merging across type/value/namespace spaces.
|
||||
mergeBindings: (existing, incoming) => [...jsMergeBindings([...existing, ...incoming])],
|
||||
|
||||
// Adapter: jsArityCompatibility uses (def, callsite); contract is (callsite, def).
|
||||
arityCompatibility: (callsite, def) => jsArityCompatibility(def, callsite),
|
||||
|
||||
buildMro: (graph, parsedFiles, nodeLookup) =>
|
||||
buildMro(graph, parsedFiles, nodeLookup, defaultLinearize),
|
||||
|
||||
populateOwners: (parsed: ParsedFile) => populateClassOwnedMembers(parsed),
|
||||
|
||||
// JavaScript `super` keyword: same pattern as TypeScript.
|
||||
isSuperReceiver: (text) => /^super(\s*\(|\s*\.|\s*\[|\s*$)/.test(text.trim()),
|
||||
|
||||
// JavaScript is dynamically typed — enable the field-fallback heuristic
|
||||
// so member-call receivers without type annotations can still resolve
|
||||
// through declared class fields (e.g. JSDoc-typed fields).
|
||||
fieldFallbackOnMethodLookup: true,
|
||||
|
||||
// Return-type propagation (across ESM imports) mirrors TypeScript's
|
||||
// default behavior. JSDoc @returns bindings are hoisted to Module scope
|
||||
// and propagated to importers via the standard mechanism.
|
||||
propagatesReturnTypesAcrossImports: true,
|
||||
|
||||
// JSDoc @returns bindings are synthesized on the function/method node
|
||||
// and hoisted to Module scope by `jsBindingScopeFor` (identical to the
|
||||
// TypeScript `tsBindingScopeFor` `@type-binding.return` branch).
|
||||
hoistTypeBindingsToModule: true,
|
||||
|
||||
// CJS-heavy codebases often have utility functions exported without
|
||||
// explicit imports at the call site. Workspace-wide unique-name fallback
|
||||
// recovers these edges.
|
||||
allowGlobalFreeCallFallback: true,
|
||||
};
|
||||
|
||||
export { javascriptScopeResolver };
|
||||
@@ -0,0 +1,48 @@
|
||||
/**
|
||||
* Simple hooks for the JavaScript scope-resolution provider.
|
||||
*
|
||||
* `jsBindingScopeFor` wraps `tsBindingScopeFor` and adds the JS-only
|
||||
* `@type-binding.class-field` hoisting rule. The other two hooks
|
||||
* (`jsImportOwningScope`, `jsReceiverBinding`) are identical to their
|
||||
* TypeScript counterparts and are re-exported directly.
|
||||
*
|
||||
* ## Why class-field hoisting lives here (not in `tsBindingScopeFor`)
|
||||
*
|
||||
* `@type-binding.class-field` is emitted exclusively by
|
||||
* `synthesizeConstructorFieldBindings` in `captures.ts`, which is a
|
||||
* JavaScript-only synthesis pass. TypeScript uses
|
||||
* `@type-binding.parameter-property` for constructor parameter
|
||||
* properties instead. Keeping the JS-only rule in the JS hook file
|
||||
* prevents language-specific logic from leaking into shared TypeScript
|
||||
* infrastructure (DoD.md §2.2).
|
||||
*/
|
||||
|
||||
import type { CaptureMatch, Scope, ScopeId, ScopeTree } from 'gitnexus-shared';
|
||||
import { tsBindingScopeFor, walkToScope } from '../typescript/simple-hooks.js';
|
||||
|
||||
export {
|
||||
tsImportOwningScope as jsImportOwningScope,
|
||||
tsReceiverBinding as jsReceiverBinding,
|
||||
} from '../typescript/simple-hooks.js';
|
||||
|
||||
/**
|
||||
* Like `tsBindingScopeFor` but additionally hoists
|
||||
* `@type-binding.class-field` captures to the enclosing Class scope.
|
||||
*
|
||||
* `@type-binding.class-field` is anchored inside the constructor body
|
||||
* (by `synthesizeConstructorFieldBindings`) so that `walkToScope` can
|
||||
* walk up from the Function (constructor) scope to the Class scope.
|
||||
* This puts `User.address → Address` in the class's typeBindings so
|
||||
* compound-receiver resolution finds it when resolving
|
||||
* `user.address.save()`.
|
||||
*/
|
||||
export function jsBindingScopeFor(
|
||||
decl: CaptureMatch,
|
||||
innermost: Scope,
|
||||
tree: ScopeTree,
|
||||
): ScopeId | null {
|
||||
if (decl['@type-binding.class-field'] !== undefined) {
|
||||
return walkToScope(innermost, tree, 'Class');
|
||||
}
|
||||
return tsBindingScopeFor(decl, innermost, tree);
|
||||
}
|
||||
@@ -29,6 +29,16 @@ import { kotlinMethodConfig } from '../method-extractors/configs/jvm.js';
|
||||
import { createVariableExtractor } from '../variable-extractors/generic.js';
|
||||
import { kotlinVariableConfig } from '../variable-extractors/configs/jvm.js';
|
||||
import { createHeritageExtractor } from '../heritage-extractors/generic.js';
|
||||
import {
|
||||
emitKotlinScopeCaptures,
|
||||
interpretKotlinImport,
|
||||
interpretKotlinTypeBinding,
|
||||
kotlinArityCompatibility,
|
||||
kotlinBindingScopeFor,
|
||||
kotlinImportOwningScope,
|
||||
kotlinMergeBindings,
|
||||
kotlinReceiverBinding,
|
||||
} from './kotlin/index.js';
|
||||
|
||||
/** Check if a Kotlin function_declaration capture is inside a class_body (i.e., a method).
|
||||
* Kotlin grammar uses function_declaration for both top-level functions and class methods.
|
||||
@@ -166,4 +176,14 @@ export const kotlinProvider = defineLanguage({
|
||||
if (isKotlinClassMethod(functionNode)) return 'Method';
|
||||
return defaultLabel;
|
||||
},
|
||||
|
||||
// ── RFC #909 Ring 3: scope-based resolution hooks ──
|
||||
emitScopeCaptures: emitKotlinScopeCaptures,
|
||||
interpretImport: interpretKotlinImport,
|
||||
interpretTypeBinding: interpretKotlinTypeBinding,
|
||||
bindingScopeFor: kotlinBindingScopeFor,
|
||||
importOwningScope: kotlinImportOwningScope,
|
||||
mergeBindings: (_scope, bindings) => kotlinMergeBindings(bindings),
|
||||
receiverBinding: kotlinReceiverBinding,
|
||||
arityCompatibility: kotlinArityCompatibility,
|
||||
});
|
||||
|
||||
@@ -0,0 +1,26 @@
|
||||
import type { SyntaxNode } from '../../utils/ast-helpers.js';
|
||||
import { kotlinMethodConfig } from '../../method-extractors/configs/jvm.js';
|
||||
|
||||
export interface KotlinArityMetadata {
|
||||
readonly parameterCount: number | undefined;
|
||||
readonly requiredParameterCount: number | undefined;
|
||||
readonly parameterTypes: readonly string[] | undefined;
|
||||
}
|
||||
|
||||
export function computeKotlinArityMetadata(fnNode: SyntaxNode): KotlinArityMetadata {
|
||||
const params = kotlinMethodConfig.extractParameters?.(fnNode) ?? [];
|
||||
let hasVararg = false;
|
||||
const parameterTypes: string[] = [];
|
||||
for (const param of params) {
|
||||
if (param.isVariadic) hasVararg = true;
|
||||
if (param.type !== null) parameterTypes.push(param.type);
|
||||
}
|
||||
if (hasVararg) parameterTypes.push('vararg');
|
||||
|
||||
const required = params.filter((p) => !p.isOptional && !p.isVariadic).length;
|
||||
return {
|
||||
parameterCount: hasVararg ? undefined : params.length,
|
||||
requiredParameterCount: required,
|
||||
parameterTypes: parameterTypes.length > 0 ? parameterTypes : undefined,
|
||||
};
|
||||
}
|
||||
@@ -0,0 +1,18 @@
|
||||
import type { Callsite, SymbolDefinition } from 'gitnexus-shared';
|
||||
|
||||
export function kotlinArityCompatibility(
|
||||
def: SymbolDefinition,
|
||||
callsite: Callsite,
|
||||
): 'compatible' | 'unknown' | 'incompatible' {
|
||||
const min = def.requiredParameterCount;
|
||||
const max = def.parameterCount;
|
||||
if (min === undefined && max === undefined) return 'unknown';
|
||||
|
||||
const argCount = callsite.arity;
|
||||
if (!Number.isFinite(argCount) || argCount < 0) return 'unknown';
|
||||
|
||||
const hasVararg = def.parameterTypes?.some((t) => t === 'vararg') ?? false;
|
||||
if (min !== undefined && argCount < min) return 'incompatible';
|
||||
if (max !== undefined && argCount > max && !hasVararg) return 'incompatible';
|
||||
return 'compatible';
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
let hits = 0;
|
||||
let misses = 0;
|
||||
|
||||
export function recordKotlinCacheHit(): void {
|
||||
hits += 1;
|
||||
}
|
||||
|
||||
export function recordKotlinCacheMiss(): void {
|
||||
misses += 1;
|
||||
}
|
||||
|
||||
export function getKotlinCaptureCacheStats(): { readonly hits: number; readonly misses: number } {
|
||||
return { hits, misses };
|
||||
}
|
||||
|
||||
export function resetKotlinCaptureCacheStats(): void {
|
||||
hits = 0;
|
||||
misses = 0;
|
||||
}
|
||||
@@ -0,0 +1,484 @@
|
||||
import type { Capture, CaptureMatch } from 'gitnexus-shared';
|
||||
import {
|
||||
findNodeAtRange,
|
||||
nodeToCapture,
|
||||
syntheticCapture,
|
||||
type SyntaxNode,
|
||||
} from '../../utils/ast-helpers.js';
|
||||
import { getTreeSitterBufferSize } from '../../constants.js';
|
||||
import { parseSourceSafe } from '../../../tree-sitter/safe-parse.js';
|
||||
import { computeKotlinArityMetadata } from './arity-metadata.js';
|
||||
import { splitKotlinImportHeader } from './import-decomposer.js';
|
||||
import { recordKotlinCacheHit, recordKotlinCacheMiss } from './cache-stats.js';
|
||||
import { normalizeKotlinType } from './interpret.js';
|
||||
import { synthesizeKotlinReceiverBinding } from './receiver-binding.js';
|
||||
import { getKotlinParser, getKotlinScopeQuery } from './query.js';
|
||||
|
||||
const FUNCTION_DECL_TAGS = ['@declaration.function'] as const;
|
||||
|
||||
export function emitKotlinScopeCaptures(
|
||||
sourceText: string,
|
||||
_filePath: string,
|
||||
cachedTree?: unknown,
|
||||
): readonly CaptureMatch[] {
|
||||
let tree = cachedTree as ReturnType<ReturnType<typeof getKotlinParser>['parse']> | undefined;
|
||||
if (tree === undefined) {
|
||||
tree = parseSourceSafe(getKotlinParser(), sourceText, undefined, {
|
||||
bufferSize: getTreeSitterBufferSize(sourceText),
|
||||
});
|
||||
recordKotlinCacheMiss();
|
||||
} else {
|
||||
recordKotlinCacheHit();
|
||||
}
|
||||
|
||||
const out: CaptureMatch[] = [];
|
||||
const returnTypes = collectKotlinReturnTypeTexts(tree.rootNode);
|
||||
out.push(...synthesizeKotlinLocalAssignmentBindings(tree.rootNode, returnTypes));
|
||||
out.push(...synthesizeKotlinLoopBindings(tree.rootNode, returnTypes));
|
||||
|
||||
for (const match of getKotlinScopeQuery().matches(tree.rootNode)) {
|
||||
const grouped: Record<string, Capture> = {};
|
||||
for (const capture of match.captures) {
|
||||
const tag = '@' + capture.name;
|
||||
grouped[tag] = nodeToCapture(tag, capture.node);
|
||||
}
|
||||
if (Object.keys(grouped).length === 0) continue;
|
||||
|
||||
if (grouped['@import.statement'] !== undefined) {
|
||||
const importNode = findNodeAtRange(
|
||||
tree.rootNode,
|
||||
grouped['@import.statement']!.range,
|
||||
'import_header',
|
||||
);
|
||||
if (importNode !== null) {
|
||||
const decomposed = splitKotlinImportHeader(importNode);
|
||||
if (decomposed !== null) {
|
||||
out.push(decomposed);
|
||||
continue;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (
|
||||
grouped['@reference.call.free'] !== undefined &&
|
||||
grouped['@reference.receiver'] !== undefined
|
||||
) {
|
||||
continue;
|
||||
}
|
||||
|
||||
if (grouped['@reference.read.member'] !== undefined) {
|
||||
const anchor = grouped['@reference.read.member']!;
|
||||
const navNode = findNodeAtRange(tree.rootNode, anchor.range, 'navigation_expression');
|
||||
if (navNode === null || !shouldEmitReadMember(navNode)) continue;
|
||||
}
|
||||
|
||||
if (grouped['@scope.function'] !== undefined) {
|
||||
out.push(grouped);
|
||||
const fnNode = findNodeAtRange(
|
||||
tree.rootNode,
|
||||
grouped['@scope.function']!.range,
|
||||
'function_declaration',
|
||||
);
|
||||
if (fnNode !== null) {
|
||||
out.push(...synthesizeKotlinReceiverBinding(fnNode));
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
const declTag = FUNCTION_DECL_TAGS.find((tag) => grouped[tag] !== undefined);
|
||||
if (declTag !== undefined) {
|
||||
const fnNode = findNodeAtRange(
|
||||
tree.rootNode,
|
||||
grouped[declTag]!.range,
|
||||
'function_declaration',
|
||||
);
|
||||
if (fnNode !== null) {
|
||||
const arity = computeKotlinArityMetadata(fnNode);
|
||||
if (arity.parameterCount !== undefined) {
|
||||
grouped['@declaration.parameter-count'] = syntheticCapture(
|
||||
'@declaration.parameter-count',
|
||||
fnNode,
|
||||
String(arity.parameterCount),
|
||||
);
|
||||
}
|
||||
if (arity.requiredParameterCount !== undefined) {
|
||||
grouped['@declaration.required-parameter-count'] = syntheticCapture(
|
||||
'@declaration.required-parameter-count',
|
||||
fnNode,
|
||||
String(arity.requiredParameterCount),
|
||||
);
|
||||
}
|
||||
if (arity.parameterTypes !== undefined) {
|
||||
grouped['@declaration.parameter-types'] = syntheticCapture(
|
||||
'@declaration.parameter-types',
|
||||
fnNode,
|
||||
JSON.stringify(arity.parameterTypes),
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
const callTag = (
|
||||
['@reference.call.free', '@reference.call.member', '@reference.call.constructor'] as const
|
||||
).find((tag) => grouped[tag] !== undefined);
|
||||
if (callTag !== undefined && grouped['@reference.arity'] === undefined) {
|
||||
const callNode = findNodeAtRange(tree.rootNode, grouped[callTag]!.range, 'call_expression');
|
||||
if (callNode !== null) {
|
||||
const args = callArguments(callNode);
|
||||
grouped['@reference.arity'] = syntheticCapture(
|
||||
'@reference.arity',
|
||||
callNode,
|
||||
String(args.length),
|
||||
);
|
||||
grouped['@reference.parameter-types'] = syntheticCapture(
|
||||
'@reference.parameter-types',
|
||||
callNode,
|
||||
JSON.stringify(args.map(inferArgType)),
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
out.push(grouped);
|
||||
|
||||
const extensionFallback = extensionFreeCallFallback(grouped, tree.rootNode);
|
||||
if (extensionFallback !== null) out.push(extensionFallback);
|
||||
}
|
||||
|
||||
return out;
|
||||
}
|
||||
|
||||
function synthesizeKotlinLoopBindings(
|
||||
rootNode: SyntaxNode,
|
||||
returnTypes: ReadonlyMap<string, string>,
|
||||
): CaptureMatch[] {
|
||||
const out: CaptureMatch[] = [];
|
||||
for (const fnNode of descendantsOfType(rootNode, 'function_declaration')) {
|
||||
const localTypes = collectKotlinLocalTypeTexts(fnNode, returnTypes);
|
||||
for (const forNode of descendantsOfType(fnNode, 'for_statement')) {
|
||||
const variable = forNode.namedChildren.find((child) => child.type === 'variable_declaration');
|
||||
const name = variable?.namedChildren.find((child) => child.type === 'simple_identifier');
|
||||
if (variable === undefined || name === undefined) continue;
|
||||
|
||||
const explicitType = variable.namedChildren.find((child) => isKotlinTypeNode(child));
|
||||
const iterable = forNode.namedChildren.find(
|
||||
(child) => child.id !== variable.id && child.type !== 'control_structure_body',
|
||||
);
|
||||
const rawType =
|
||||
explicitType?.text ??
|
||||
(iterable === undefined
|
||||
? null
|
||||
: inferKotlinIterableElementType(iterable, localTypes, returnTypes));
|
||||
if (rawType === null || rawType.trim() === '') continue;
|
||||
|
||||
const anchor =
|
||||
forNode.namedChildren.find((child) => child.type === 'control_structure_body') ?? forNode;
|
||||
out.push({
|
||||
'@type-binding.annotation': nodeToCapture('@type-binding.annotation', anchor),
|
||||
'@type-binding.name': syntheticCapture('@type-binding.name', name, name.text),
|
||||
'@type-binding.type': syntheticCapture(
|
||||
'@type-binding.type',
|
||||
explicitType ?? iterable ?? name,
|
||||
normalizeKotlinType(rawType),
|
||||
),
|
||||
});
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
function synthesizeKotlinLocalAssignmentBindings(
|
||||
rootNode: SyntaxNode,
|
||||
returnTypes: ReadonlyMap<string, string>,
|
||||
): CaptureMatch[] {
|
||||
const out: CaptureMatch[] = [];
|
||||
for (const fnNode of descendantsOfType(rootNode, 'function_declaration')) {
|
||||
const localTypes = new Map<string, string>();
|
||||
for (const prop of descendantsOfType(fnNode, 'property_declaration')) {
|
||||
const inferred = inferKotlinPropertyType(prop, localTypes, returnTypes);
|
||||
if (inferred === null) continue;
|
||||
localTypes.set(inferred.name.text, inferred.rawType);
|
||||
if (inferred.synthetic) {
|
||||
out.push({
|
||||
'@type-binding.annotation': nodeToCapture('@type-binding.annotation', prop),
|
||||
'@type-binding.name': syntheticCapture(
|
||||
'@type-binding.name',
|
||||
inferred.name,
|
||||
inferred.name.text,
|
||||
),
|
||||
'@type-binding.type': syntheticCapture(
|
||||
'@type-binding.type',
|
||||
inferred.source,
|
||||
normalizeKotlinType(inferred.rawType),
|
||||
),
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
function collectKotlinLocalTypeTexts(
|
||||
fnNode: SyntaxNode,
|
||||
returnTypes: ReadonlyMap<string, string>,
|
||||
): Map<string, string> {
|
||||
const out = new Map<string, string>();
|
||||
for (const node of descendants(fnNode)) {
|
||||
if (node.type === 'parameter') {
|
||||
const name = descendantsOfType(node, 'simple_identifier')[0];
|
||||
const type = node.namedChildren.find((child) => isKotlinTypeNode(child));
|
||||
if (name !== undefined && type !== undefined) out.set(name.text, type.text);
|
||||
continue;
|
||||
}
|
||||
|
||||
if (node.type === 'property_declaration') {
|
||||
const inferred = inferKotlinPropertyType(node, out, returnTypes);
|
||||
if (inferred !== null) out.set(inferred.name.text, inferred.rawType);
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
function collectKotlinReturnTypeTexts(rootNode: SyntaxNode): Map<string, string> {
|
||||
const out = new Map<string, string>();
|
||||
for (const fnNode of descendantsOfType(rootNode, 'function_declaration')) {
|
||||
const name = fnNode.namedChildren.find((child) => child.type === 'simple_identifier');
|
||||
const paramsIndex = fnNode.namedChildren.findIndex(
|
||||
(child) => child.type === 'function_value_parameters',
|
||||
);
|
||||
const type =
|
||||
paramsIndex < 0
|
||||
? undefined
|
||||
: fnNode.namedChildren.slice(paramsIndex + 1).find((child) => isKotlinTypeNode(child));
|
||||
if (name !== undefined && type !== undefined) out.set(name.text, type.text);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
function inferKotlinPropertyType(
|
||||
prop: SyntaxNode,
|
||||
localTypes: ReadonlyMap<string, string>,
|
||||
returnTypes: ReadonlyMap<string, string>,
|
||||
): { name: SyntaxNode; rawType: string; source: SyntaxNode; synthetic: boolean } | null {
|
||||
const variable = prop.namedChildren.find((child) => child.type === 'variable_declaration');
|
||||
const name = variable?.namedChildren.find((child) => child.type === 'simple_identifier');
|
||||
if (variable === undefined || name === undefined) return null;
|
||||
|
||||
const explicitType = variable.namedChildren.find((child) => isKotlinTypeNode(child));
|
||||
if (explicitType !== undefined) {
|
||||
return { name, rawType: explicitType.text, source: explicitType, synthetic: false };
|
||||
}
|
||||
|
||||
const value = prop.namedChildren.find(
|
||||
(child) => child.id !== variable.id && child.type !== 'binding_pattern_kind',
|
||||
);
|
||||
if (value?.type === 'simple_identifier') {
|
||||
const rawType = localTypes.get(value.text);
|
||||
return rawType === undefined ? null : { name, rawType, source: value, synthetic: true };
|
||||
}
|
||||
|
||||
if (value?.type === 'call_expression') {
|
||||
const callee = value.namedChildren.find((child) => child.type === 'simple_identifier');
|
||||
if (callee === undefined) return null;
|
||||
const rawType =
|
||||
returnTypes.get(callee.text) ?? (isUppercaseName(callee.text) ? callee.text : null);
|
||||
if (rawType === null) return null;
|
||||
return { name, rawType, source: callee, synthetic: true };
|
||||
}
|
||||
|
||||
return null;
|
||||
}
|
||||
|
||||
function inferKotlinIterableElementType(
|
||||
iterable: SyntaxNode,
|
||||
localTypes: ReadonlyMap<string, string>,
|
||||
returnTypes: ReadonlyMap<string, string>,
|
||||
): string | null {
|
||||
if (iterable.type === 'simple_identifier') {
|
||||
const raw = localTypes.get(iterable.text);
|
||||
return raw === undefined ? null : kotlinContainerElementType(raw, 'values');
|
||||
}
|
||||
|
||||
if (iterable.type === 'navigation_expression') {
|
||||
const receiver = iterable.namedChildren[0];
|
||||
const member = iterable.namedChildren
|
||||
.find((child) => child.type === 'navigation_suffix')
|
||||
?.namedChildren.find((child) => child.type === 'simple_identifier')?.text;
|
||||
if (receiver?.type !== 'simple_identifier') return null;
|
||||
const raw = localTypes.get(receiver.text);
|
||||
return raw === undefined ? null : kotlinContainerElementType(raw, member ?? 'values');
|
||||
}
|
||||
|
||||
if (iterable.type === 'call_expression') {
|
||||
const callee = iterable.namedChildren.find((child) => child.type === 'simple_identifier');
|
||||
if (callee === undefined) return null;
|
||||
const raw = returnTypes.get(callee.text);
|
||||
return raw === undefined ? null : kotlinContainerElementType(raw, 'values');
|
||||
}
|
||||
|
||||
return null;
|
||||
}
|
||||
|
||||
function isUppercaseName(text: string): boolean {
|
||||
return /^[A-Z]/.test(text);
|
||||
}
|
||||
|
||||
function kotlinContainerElementType(rawType: string, member: string): string | null {
|
||||
const parsed = parseKotlinGeneric(rawType);
|
||||
if (parsed === null) return normalizeKotlinType(rawType);
|
||||
|
||||
const base = parsed.base.split('.').pop() ?? parsed.base;
|
||||
if (isKotlinMapType(base)) {
|
||||
if (member === 'keys') return parsed.args[0] ?? null;
|
||||
return parsed.args[1] ?? null;
|
||||
}
|
||||
if (isKotlinIterableType(base)) return parsed.args[0] ?? null;
|
||||
return normalizeKotlinType(rawType);
|
||||
}
|
||||
|
||||
function parseKotlinGeneric(text: string): { base: string; args: string[] } | null {
|
||||
const trimmed = text.trim().replace(/\?$/, '');
|
||||
const open = trimmed.indexOf('<');
|
||||
const close = trimmed.lastIndexOf('>');
|
||||
if (open < 0 || close < open) return null;
|
||||
return {
|
||||
base: trimmed.slice(0, open).trim(),
|
||||
args: splitTopLevelKotlinArgs(trimmed.slice(open + 1, close)),
|
||||
};
|
||||
}
|
||||
|
||||
function splitTopLevelKotlinArgs(text: string): string[] {
|
||||
const out: string[] = [];
|
||||
let depth = 0;
|
||||
let start = 0;
|
||||
for (let i = 0; i < text.length; i++) {
|
||||
const ch = text[i];
|
||||
if (ch === '<') depth++;
|
||||
else if (ch === '>') depth--;
|
||||
else if (ch === ',' && depth === 0) {
|
||||
out.push(text.slice(start, i).trim());
|
||||
start = i + 1;
|
||||
}
|
||||
}
|
||||
out.push(text.slice(start).trim());
|
||||
return out.filter((arg) => arg.length > 0);
|
||||
}
|
||||
|
||||
function isKotlinMapType(base: string): boolean {
|
||||
return ['Map', 'MutableMap', 'HashMap', 'LinkedHashMap'].includes(base);
|
||||
}
|
||||
|
||||
function isKotlinIterableType(base: string): boolean {
|
||||
return [
|
||||
'List',
|
||||
'MutableList',
|
||||
'ArrayList',
|
||||
'Set',
|
||||
'MutableSet',
|
||||
'Collection',
|
||||
'Iterable',
|
||||
'Sequence',
|
||||
'Array',
|
||||
].includes(base);
|
||||
}
|
||||
|
||||
function isKotlinTypeNode(node: SyntaxNode): boolean {
|
||||
return (
|
||||
node.type === 'user_type' || node.type === 'nullable_type' || node.type === 'function_type'
|
||||
);
|
||||
}
|
||||
|
||||
function descendantsOfType(node: SyntaxNode, type: string): SyntaxNode[] {
|
||||
return descendants(node).filter((child) => child.type === type);
|
||||
}
|
||||
|
||||
function descendants(node: SyntaxNode): SyntaxNode[] {
|
||||
const out: SyntaxNode[] = [];
|
||||
for (let i = 0; i < node.namedChildCount; i++) {
|
||||
const child = node.namedChild(i);
|
||||
if (child === null) continue;
|
||||
out.push(child, ...descendants(child));
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
function shouldEmitReadMember(navNode: SyntaxNode): boolean {
|
||||
const parent = navNode.parent;
|
||||
if (parent === null) return true;
|
||||
if (parent.type === 'call_expression') return false;
|
||||
if (parent.type === 'directly_assignable_expression') return false;
|
||||
return true;
|
||||
}
|
||||
|
||||
function callArguments(callNode: SyntaxNode): SyntaxNode[] {
|
||||
const suffix = callNode.namedChildren.find((child) => child.type === 'call_suffix');
|
||||
if (suffix === undefined) return [];
|
||||
|
||||
const valueArgs = suffix?.namedChildren.find((child) => child.type === 'value_arguments');
|
||||
const args = valueArgs?.namedChildren.filter((child) => child.type === 'value_argument') ?? [];
|
||||
const trailingLambdas = suffix.namedChildren.filter((child) => child.type === 'annotated_lambda');
|
||||
return [...args, ...trailingLambdas];
|
||||
}
|
||||
|
||||
function inferArgType(argNode: SyntaxNode): string {
|
||||
const value = argNode.namedChild(0) ?? argNode;
|
||||
switch (value.type) {
|
||||
case 'integer_literal':
|
||||
case 'long_literal':
|
||||
return 'Int';
|
||||
case 'real_literal':
|
||||
return 'Double';
|
||||
case 'string_literal':
|
||||
case 'line_string_literal':
|
||||
case 'multi_line_string_literal':
|
||||
return 'String';
|
||||
case 'character_literal':
|
||||
return 'Char';
|
||||
case 'boolean_literal':
|
||||
return 'Boolean';
|
||||
case 'call_expression': {
|
||||
const first = value.namedChild(0);
|
||||
return first?.type === 'simple_identifier' ? first.text : '';
|
||||
}
|
||||
default:
|
||||
return '';
|
||||
}
|
||||
}
|
||||
|
||||
function extensionFreeCallFallback(
|
||||
grouped: Record<string, Capture>,
|
||||
rootNode: SyntaxNode,
|
||||
): CaptureMatch | null {
|
||||
const member = grouped['@reference.call.member'];
|
||||
const receiver = grouped['@reference.receiver'];
|
||||
const name = grouped['@reference.name'];
|
||||
if (member === undefined || receiver === undefined || name === undefined) return null;
|
||||
|
||||
const callNode = findNodeAtRange(rootNode, member.range, 'call_expression');
|
||||
if (callNode === null) return null;
|
||||
const receiverNode = findNodeAtRange(rootNode, receiver.range);
|
||||
if (receiverNode === null || !isLiteralReceiver(receiverNode)) return null;
|
||||
|
||||
const out: Record<string, Capture> = {
|
||||
'@reference.call.free': syntheticCapture('@reference.call.free', callNode, callNode.text),
|
||||
'@reference.name': syntheticCapture('@reference.name', callNode, name.text),
|
||||
};
|
||||
if (grouped['@reference.arity'] !== undefined)
|
||||
out['@reference.arity'] = grouped['@reference.arity'];
|
||||
if (grouped['@reference.parameter-types'] !== undefined) {
|
||||
out['@reference.parameter-types'] = grouped['@reference.parameter-types'];
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
function isLiteralReceiver(node: SyntaxNode): boolean {
|
||||
return [
|
||||
'integer_literal',
|
||||
'long_literal',
|
||||
'real_literal',
|
||||
'string_literal',
|
||||
'line_string_literal',
|
||||
'multi_line_string_literal',
|
||||
'character_literal',
|
||||
'boolean_literal',
|
||||
].includes(node.type);
|
||||
}
|
||||
@@ -0,0 +1,49 @@
|
||||
import type { Capture, CaptureMatch } from 'gitnexus-shared';
|
||||
import { nodeToCapture, syntheticCapture, type SyntaxNode } from '../../utils/ast-helpers.js';
|
||||
|
||||
type KotlinImportKind = 'named' | 'alias' | 'wildcard';
|
||||
|
||||
interface KotlinImportSpec {
|
||||
readonly kind: KotlinImportKind;
|
||||
readonly source: string;
|
||||
readonly name: string;
|
||||
readonly alias?: string;
|
||||
readonly atNode: SyntaxNode;
|
||||
}
|
||||
|
||||
export function splitKotlinImportHeader(importNode: SyntaxNode): CaptureMatch | null {
|
||||
if (importNode.type !== 'import_header') return null;
|
||||
const spec = parseKotlinImport(importNode);
|
||||
if (spec === null) return null;
|
||||
|
||||
const out: Record<string, Capture> = {
|
||||
'@import.statement': nodeToCapture('@import.statement', importNode),
|
||||
'@import.kind': syntheticCapture('@import.kind', spec.atNode, spec.kind),
|
||||
'@import.source': syntheticCapture('@import.source', spec.atNode, spec.source),
|
||||
'@import.name': syntheticCapture('@import.name', spec.atNode, spec.name),
|
||||
};
|
||||
if (spec.alias !== undefined) {
|
||||
out['@import.alias'] = syntheticCapture('@import.alias', spec.atNode, spec.alias);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
function parseKotlinImport(node: SyntaxNode): KotlinImportSpec | null {
|
||||
const identifier = node.namedChildren.find((child) => child.type === 'identifier');
|
||||
if (identifier === undefined) return null;
|
||||
const source = identifier.text.trim();
|
||||
if (source.length === 0) return null;
|
||||
|
||||
const hasWildcard = node.namedChildren.some((child) => child.type === 'wildcard_import');
|
||||
if (hasWildcard) {
|
||||
return { kind: 'wildcard', source, name: '*', atNode: node };
|
||||
}
|
||||
|
||||
const aliasNode = node.namedChildren.find((child) => child.type === 'import_alias');
|
||||
const alias = aliasNode?.namedChildren.find((child) => child.type === 'type_identifier')?.text;
|
||||
const importedName = source.split('.').pop() ?? source;
|
||||
if (alias !== undefined && alias.length > 0) {
|
||||
return { kind: 'alias', source, name: importedName, alias, atNode: node };
|
||||
}
|
||||
return { kind: 'named', source, name: importedName, atNode: node };
|
||||
}
|
||||
@@ -0,0 +1,76 @@
|
||||
import type { ParsedImport, WorkspaceIndex } from 'gitnexus-shared';
|
||||
|
||||
export interface KotlinResolveContext {
|
||||
readonly fromFile: string;
|
||||
readonly allFilePaths: ReadonlySet<string>;
|
||||
}
|
||||
|
||||
export function resolveKotlinImportTarget(
|
||||
parsedImport: ParsedImport,
|
||||
workspaceIndex: WorkspaceIndex,
|
||||
): string | null {
|
||||
const ctx = workspaceIndex as KotlinResolveContext | undefined;
|
||||
if (
|
||||
ctx === undefined ||
|
||||
typeof (ctx as { fromFile?: unknown }).fromFile !== 'string' ||
|
||||
!((ctx as { allFilePaths?: unknown }).allFilePaths instanceof Set)
|
||||
) {
|
||||
return null;
|
||||
}
|
||||
if (parsedImport.kind === 'dynamic-unresolved') return null;
|
||||
if (parsedImport.targetRaw === null || parsedImport.targetRaw === '') return null;
|
||||
|
||||
const target = parsedImport.targetRaw.endsWith('.*')
|
||||
? parsedImport.targetRaw.slice(0, -2)
|
||||
: parsedImport.targetRaw;
|
||||
const pathLike = target.replace(/\./g, '/');
|
||||
|
||||
return (
|
||||
findKotlinFile(ctx.allFilePaths, pathLike) ??
|
||||
findKotlinFile(ctx.allFilePaths, pathLike.split('/').slice(0, -1).join('/')) ??
|
||||
findByProgressivePrefixStrip(ctx.allFilePaths, pathLike)
|
||||
);
|
||||
}
|
||||
|
||||
function findKotlinFile(allFilePaths: ReadonlySet<string>, pathLike: string): string | null {
|
||||
if (pathLike === '') return null;
|
||||
const extensions = ['.kt', '.kts'];
|
||||
const suffix = `/${pathLike}`;
|
||||
const dirPrefix = `${pathLike}/`;
|
||||
const suffixDirPrefix = `/${dirPrefix}`;
|
||||
|
||||
let suffixFile: string | null = null;
|
||||
let directoryChild: string | null = null;
|
||||
|
||||
for (const raw of allFilePaths) {
|
||||
const file = raw.replace(/\\/g, '/');
|
||||
if (!extensions.some((ext) => file.endsWith(ext))) continue;
|
||||
for (const ext of extensions) {
|
||||
if (file === `${pathLike}${ext}`) return raw;
|
||||
if (suffixFile === null && file.endsWith(`${suffix}${ext}`)) suffixFile = raw;
|
||||
}
|
||||
if (directoryChild === null) {
|
||||
const atRoot = file.startsWith(dirPrefix);
|
||||
const atNested = file.includes(suffixDirPrefix);
|
||||
if (atRoot || atNested) {
|
||||
const idx = atRoot ? 0 : file.indexOf(suffixDirPrefix) + 1;
|
||||
const after = file.slice(idx + dirPrefix.length);
|
||||
if (after.length > 0 && !after.includes('/')) directoryChild = raw;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return suffixFile ?? directoryChild;
|
||||
}
|
||||
|
||||
function findByProgressivePrefixStrip(
|
||||
allFilePaths: ReadonlySet<string>,
|
||||
pathLike: string,
|
||||
): string | null {
|
||||
const segments = pathLike.split('/').filter(Boolean);
|
||||
for (let skip = 1; skip < segments.length; skip++) {
|
||||
const found = findKotlinFile(allFilePaths, segments.slice(skip).join('/'));
|
||||
if (found !== null) return found;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
@@ -0,0 +1,12 @@
|
||||
export { emitKotlinScopeCaptures } from './captures.js';
|
||||
export { getKotlinCaptureCacheStats, resetKotlinCaptureCacheStats } from './cache-stats.js';
|
||||
export { interpretKotlinImport, interpretKotlinTypeBinding } from './interpret.js';
|
||||
export { kotlinArityCompatibility } from './arity.js';
|
||||
export { resolveKotlinImportTarget, type KotlinResolveContext } from './import-target.js';
|
||||
export { kotlinMergeBindings } from './merge-bindings.js';
|
||||
export { populateKotlinOwners } from './owners.js';
|
||||
export {
|
||||
kotlinBindingScopeFor,
|
||||
kotlinImportOwningScope,
|
||||
kotlinReceiverBinding,
|
||||
} from './simple-hooks.js';
|
||||
@@ -0,0 +1,71 @@
|
||||
import type { CaptureMatch, ParsedImport, ParsedTypeBinding, TypeRef } from 'gitnexus-shared';
|
||||
|
||||
export function interpretKotlinImport(captures: CaptureMatch): ParsedImport | null {
|
||||
const kind = captures['@import.kind']?.text;
|
||||
const source = captures['@import.source']?.text;
|
||||
const name = captures['@import.name']?.text;
|
||||
if (kind === undefined || source === undefined) return null;
|
||||
|
||||
switch (kind) {
|
||||
case 'named':
|
||||
return {
|
||||
kind: 'named',
|
||||
localName: name ?? source.split('.').pop() ?? source,
|
||||
importedName: name ?? source.split('.').pop() ?? source,
|
||||
targetRaw: source,
|
||||
};
|
||||
case 'alias': {
|
||||
const alias = captures['@import.alias']?.text;
|
||||
if (alias === undefined || name === undefined) return null;
|
||||
return {
|
||||
kind: 'alias',
|
||||
localName: alias,
|
||||
importedName: name,
|
||||
alias,
|
||||
targetRaw: source,
|
||||
};
|
||||
}
|
||||
case 'wildcard':
|
||||
return { kind: 'wildcard', targetRaw: source.endsWith('.*') ? source : `${source}.*` };
|
||||
default:
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
export function interpretKotlinTypeBinding(captures: CaptureMatch): ParsedTypeBinding | null {
|
||||
const nameCap = captures['@type-binding.name'];
|
||||
const typeCap = captures['@type-binding.type'];
|
||||
if (nameCap === undefined || typeCap === undefined) return null;
|
||||
|
||||
let source: TypeRef['source'] = 'annotation';
|
||||
if (captures['@type-binding.self'] !== undefined) source = 'self';
|
||||
else if (captures['@type-binding.parameter'] !== undefined) source = 'parameter-annotation';
|
||||
else if (captures['@type-binding.return'] !== undefined) source = 'return-annotation';
|
||||
else if (captures['@type-binding.constructor'] !== undefined) source = 'constructor-inferred';
|
||||
|
||||
return {
|
||||
boundName: nameCap.text,
|
||||
rawTypeName: normalizeKotlinType(typeCap.text),
|
||||
source,
|
||||
};
|
||||
}
|
||||
|
||||
export function normalizeKotlinType(text: string): string {
|
||||
let out = text.trim();
|
||||
while (out.endsWith('?')) out = out.slice(0, -1).trim();
|
||||
const lastDot = out.lastIndexOf('.');
|
||||
if (lastDot >= 0) out = out.slice(lastDot + 1);
|
||||
|
||||
const collection = out.match(
|
||||
/^(?:List|MutableList|ArrayList|Set|MutableSet|Collection|Iterable|Sequence|Array)<([^,<>]+)>$/,
|
||||
);
|
||||
if (collection !== null) return normalizeKotlinType(collection[1]!);
|
||||
|
||||
const map = out.match(/^(?:Map|MutableMap|HashMap|LinkedHashMap)<[^,<>]+,\s*([^,<>]+)>$/);
|
||||
if (map !== null) return normalizeKotlinType(map[1]!);
|
||||
|
||||
const erased = out.match(/^([A-Za-z_][A-Za-z0-9_]*)<.+>$/s);
|
||||
if (erased !== null) return erased[1]!;
|
||||
|
||||
return out;
|
||||
}
|
||||
@@ -0,0 +1,26 @@
|
||||
import type { BindingRef } from 'gitnexus-shared';
|
||||
|
||||
function tierOf(binding: BindingRef): number {
|
||||
switch (binding.origin) {
|
||||
case 'local':
|
||||
return 0;
|
||||
case 'import':
|
||||
case 'namespace':
|
||||
case 'reexport':
|
||||
return 1;
|
||||
case 'wildcard':
|
||||
return 2;
|
||||
default:
|
||||
return 3;
|
||||
}
|
||||
}
|
||||
|
||||
export function kotlinMergeBindings(bindings: readonly BindingRef[]): readonly BindingRef[] {
|
||||
if (bindings.length === 0) return bindings;
|
||||
const best = Math.min(...bindings.map(tierOf));
|
||||
const seen = new Map<string, BindingRef>();
|
||||
for (const binding of bindings) {
|
||||
if (tierOf(binding) === best) seen.set(binding.def.nodeId, binding);
|
||||
}
|
||||
return [...seen.values()];
|
||||
}
|
||||
@@ -0,0 +1,50 @@
|
||||
import type { ParsedFile, ScopeId, SymbolDefinition } from 'gitnexus-shared';
|
||||
import { isClassLike, populateClassOwnedMembers } from '../../scope-resolution/scope/walkers.js';
|
||||
|
||||
export function populateKotlinOwners(parsed: ParsedFile): void {
|
||||
populateClassOwnedMembers(parsed);
|
||||
populateCompanionMembersOnEnclosingClass(parsed);
|
||||
}
|
||||
|
||||
function populateCompanionMembersOnEnclosingClass(parsed: ParsedFile): void {
|
||||
const scopesById = new Map<ScopeId, ParsedFile['scopes'][number]>();
|
||||
for (const scope of parsed.scopes) scopesById.set(scope.id, scope);
|
||||
|
||||
for (const scope of parsed.scopes) {
|
||||
if (scope.kind !== 'Function' || scope.parent === null) continue;
|
||||
const parent = scopesById.get(scope.parent);
|
||||
if (parent === undefined || parent.kind !== 'Class') continue;
|
||||
if (parent.ownedDefs.some((def) => isClassLike(def.type))) continue;
|
||||
|
||||
const enclosing = findEnclosingClassWithDef(parent.parent, scopesById);
|
||||
if (enclosing === undefined) continue;
|
||||
for (const def of scope.ownedDefs) {
|
||||
if (def.ownerId !== undefined) continue;
|
||||
(def as { ownerId?: string }).ownerId = enclosing.nodeId;
|
||||
qualify(def, enclosing);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
function findEnclosingClassWithDef(
|
||||
start: ScopeId | null,
|
||||
scopesById: ReadonlyMap<ScopeId, ParsedFile['scopes'][number]>,
|
||||
): SymbolDefinition | undefined {
|
||||
let current = start;
|
||||
while (current !== null) {
|
||||
const scope = scopesById.get(current);
|
||||
if (scope === undefined) return undefined;
|
||||
if (scope.kind === 'Class') {
|
||||
const classDef = scope.ownedDefs.find((def) => isClassLike(def.type));
|
||||
if (classDef !== undefined) return classDef;
|
||||
}
|
||||
current = scope.parent;
|
||||
}
|
||||
return undefined;
|
||||
}
|
||||
|
||||
function qualify(def: SymbolDefinition, owner: SymbolDefinition): void {
|
||||
if (def.qualifiedName === undefined || def.qualifiedName.includes('.')) return;
|
||||
if (owner.qualifiedName === undefined || owner.qualifiedName.length === 0) return;
|
||||
(def as { qualifiedName: string }).qualifiedName = `${owner.qualifiedName}.${def.qualifiedName}`;
|
||||
}
|
||||
@@ -0,0 +1,116 @@
|
||||
import Parser from 'tree-sitter';
|
||||
import Kotlin from 'tree-sitter-kotlin';
|
||||
|
||||
const KOTLIN_SCOPE_QUERY = `
|
||||
;; Scopes
|
||||
(source_file) @scope.module
|
||||
(class_declaration) @scope.class
|
||||
(object_declaration) @scope.class
|
||||
(companion_object) @scope.class
|
||||
(function_declaration) @scope.function
|
||||
|
||||
;; Declarations — types
|
||||
(class_declaration
|
||||
"interface"
|
||||
(type_identifier) @declaration.name) @declaration.interface
|
||||
|
||||
(class_declaration
|
||||
"class"
|
||||
(type_identifier) @declaration.name) @declaration.class
|
||||
|
||||
(object_declaration
|
||||
(type_identifier) @declaration.name) @declaration.class
|
||||
|
||||
(companion_object
|
||||
(type_identifier) @declaration.name) @declaration.class
|
||||
|
||||
(type_alias
|
||||
(type_identifier) @declaration.name) @declaration.type_alias
|
||||
|
||||
;; Declarations — functions / methods / properties
|
||||
(function_declaration
|
||||
(simple_identifier) @declaration.name) @declaration.function
|
||||
|
||||
(property_declaration
|
||||
(variable_declaration
|
||||
(simple_identifier) @declaration.name)) @declaration.property
|
||||
|
||||
(class_parameter
|
||||
(binding_pattern_kind)
|
||||
(simple_identifier) @declaration.name) @declaration.property
|
||||
|
||||
;; Imports
|
||||
(import_header) @import.statement
|
||||
|
||||
;; Type bindings — parameters
|
||||
(parameter
|
||||
(simple_identifier) @type-binding.name
|
||||
[(user_type) (nullable_type) (function_type)] @type-binding.type) @type-binding.parameter
|
||||
|
||||
;; Type bindings — property / local annotations
|
||||
(property_declaration
|
||||
(variable_declaration
|
||||
(simple_identifier) @type-binding.name
|
||||
[(user_type) (nullable_type) (function_type)] @type-binding.type)) @type-binding.annotation
|
||||
|
||||
(class_parameter
|
||||
(binding_pattern_kind)
|
||||
(simple_identifier) @type-binding.name
|
||||
[(user_type) (nullable_type) (function_type)] @type-binding.type) @type-binding.annotation
|
||||
|
||||
;; Type bindings — constructor-inferred val user = User(...)
|
||||
(property_declaration
|
||||
(variable_declaration
|
||||
(simple_identifier) @type-binding.name)
|
||||
(call_expression
|
||||
(simple_identifier) @type-binding.type)) @type-binding.constructor
|
||||
|
||||
;; Type bindings — return annotations after function parameters
|
||||
(function_declaration
|
||||
(simple_identifier) @type-binding.name
|
||||
(function_value_parameters)
|
||||
[(user_type) (nullable_type) (function_type)] @type-binding.type) @type-binding.return
|
||||
|
||||
;; References — direct calls / constructor syntax
|
||||
(call_expression
|
||||
(simple_identifier) @reference.name) @reference.call.free
|
||||
|
||||
;; References — member calls: obj.method()
|
||||
(call_expression
|
||||
(navigation_expression
|
||||
(_) @reference.receiver
|
||||
(navigation_suffix
|
||||
(simple_identifier) @reference.name))) @reference.call.member
|
||||
|
||||
;; References — property writes
|
||||
(assignment
|
||||
(directly_assignable_expression
|
||||
(_) @reference.receiver
|
||||
(navigation_suffix
|
||||
(simple_identifier) @reference.name))
|
||||
(_)) @reference.write.member
|
||||
|
||||
;; References — property reads
|
||||
(navigation_expression
|
||||
(_) @reference.receiver
|
||||
(navigation_suffix
|
||||
(simple_identifier) @reference.name)) @reference.read.member
|
||||
`;
|
||||
|
||||
let parser: Parser | null = null;
|
||||
let query: Parser.Query | null = null;
|
||||
|
||||
export function getKotlinParser(): Parser {
|
||||
if (parser === null) {
|
||||
parser = new Parser();
|
||||
parser.setLanguage(Kotlin as Parameters<Parser['setLanguage']>[0]);
|
||||
}
|
||||
return parser;
|
||||
}
|
||||
|
||||
export function getKotlinScopeQuery(): Parser.Query {
|
||||
if (query === null) {
|
||||
query = new Parser.Query(Kotlin as Parameters<Parser['setLanguage']>[0], KOTLIN_SCOPE_QUERY);
|
||||
}
|
||||
return query;
|
||||
}
|
||||
@@ -0,0 +1,106 @@
|
||||
import type { Capture, CaptureMatch } from 'gitnexus-shared';
|
||||
import { nodeToCapture, syntheticCapture, type SyntaxNode } from '../../utils/ast-helpers.js';
|
||||
import { normalizeKotlinType } from './interpret.js';
|
||||
|
||||
const TYPE_DECL_NODE_TYPES = new Set([
|
||||
'class_declaration',
|
||||
'object_declaration',
|
||||
'companion_object',
|
||||
]);
|
||||
|
||||
export function synthesizeKotlinReceiverBinding(fnNode: SyntaxNode): CaptureMatch[] {
|
||||
if (fnNode.type !== 'function_declaration') return [];
|
||||
|
||||
const anchorNode = findFunctionBody(fnNode);
|
||||
if (anchorNode === null) return [];
|
||||
|
||||
const extensionReceiver = extensionReceiverType(fnNode);
|
||||
if (extensionReceiver !== null) {
|
||||
return [buildReceiverMatch(anchorNode, 'this', extensionReceiver)];
|
||||
}
|
||||
|
||||
const enclosingType = findEnclosingTypeDeclaration(fnNode);
|
||||
if (enclosingType === null) return [];
|
||||
|
||||
const enclosingName = typeDeclarationName(enclosingType);
|
||||
if (enclosingName === null) return [];
|
||||
|
||||
const out = [buildReceiverMatch(anchorNode, 'this', enclosingName)];
|
||||
const superName = firstSuperclassText(enclosingType);
|
||||
if (superName !== null) out.push(buildReceiverMatch(anchorNode, 'super', superName));
|
||||
return out;
|
||||
}
|
||||
|
||||
function findFunctionBody(fnNode: SyntaxNode): SyntaxNode | null {
|
||||
for (let i = 0; i < fnNode.namedChildCount; i++) {
|
||||
const child = fnNode.namedChild(i);
|
||||
if (child?.type === 'function_body') return child;
|
||||
}
|
||||
return fnNode;
|
||||
}
|
||||
|
||||
function extensionReceiverType(fnNode: SyntaxNode): string | null {
|
||||
for (let i = 0; i < fnNode.namedChildCount; i++) {
|
||||
const child = fnNode.namedChild(i);
|
||||
if (child === null) continue;
|
||||
if (child.type === 'simple_identifier') return null;
|
||||
if (child.type === 'user_type' || child.type === 'nullable_type') {
|
||||
return normalizeKotlinType(child.text);
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
function findEnclosingTypeDeclaration(node: SyntaxNode): SyntaxNode | null {
|
||||
let current = node.parent;
|
||||
while (current !== null) {
|
||||
if (TYPE_DECL_NODE_TYPES.has(current.type)) return current;
|
||||
current = current.parent;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
function typeDeclarationName(typeNode: SyntaxNode): string | null {
|
||||
if (typeNode.type === 'companion_object') {
|
||||
return (
|
||||
typeNode.namedChildren.find((child) => child.type === 'type_identifier')?.text ??
|
||||
enclosingNonCompanionTypeName(typeNode) ??
|
||||
'Companion'
|
||||
);
|
||||
}
|
||||
return typeNode.namedChildren.find((child) => child.type === 'type_identifier')?.text ?? null;
|
||||
}
|
||||
|
||||
function enclosingNonCompanionTypeName(node: SyntaxNode): string | null {
|
||||
let current = node.parent;
|
||||
while (current !== null) {
|
||||
if (current.type === 'class_declaration' || current.type === 'object_declaration') {
|
||||
return current.namedChildren.find((child) => child.type === 'type_identifier')?.text ?? null;
|
||||
}
|
||||
current = current.parent;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
function firstSuperclassText(typeNode: SyntaxNode): string | null {
|
||||
if (typeNode.type !== 'class_declaration') return null;
|
||||
for (const child of typeNode.namedChildren) {
|
||||
if (child.type !== 'delegation_specifier') continue;
|
||||
const ctor = child.namedChildren.find((n) => n.type === 'constructor_invocation');
|
||||
const userType =
|
||||
ctor?.namedChildren.find((n) => n.type === 'user_type') ??
|
||||
child.namedChildren.find((n) => n.type === 'user_type');
|
||||
const name = userType?.namedChildren.find((n) => n.type === 'type_identifier')?.text;
|
||||
if (name !== undefined) return normalizeKotlinType(name);
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
function buildReceiverMatch(anchorNode: SyntaxNode, name: string, typeText: string): CaptureMatch {
|
||||
const out: Record<string, Capture> = {
|
||||
'@type-binding.self': nodeToCapture('@type-binding.self', anchorNode),
|
||||
'@type-binding.name': syntheticCapture('@type-binding.name', anchorNode, name),
|
||||
'@type-binding.type': syntheticCapture('@type-binding.type', anchorNode, typeText),
|
||||
};
|
||||
return out;
|
||||
}
|
||||
@@ -0,0 +1,56 @@
|
||||
import { SupportedLanguages, type ParsedFile } from 'gitnexus-shared';
|
||||
import { buildMro, defaultLinearize } from '../../scope-resolution/passes/mro.js';
|
||||
import type { ScopeResolver } from '../../scope-resolution/contract/scope-resolver.js';
|
||||
import { kotlinProvider } from '../kotlin.js';
|
||||
import {
|
||||
kotlinArityCompatibility,
|
||||
kotlinMergeBindings,
|
||||
populateKotlinOwners,
|
||||
resolveKotlinImportTarget,
|
||||
type KotlinResolveContext,
|
||||
} from './index.js';
|
||||
|
||||
/**
|
||||
* Kotlin scope resolver for RFC #909 Ring 3.
|
||||
*
|
||||
* Kotlin is intentionally registered but not yet listed in
|
||||
* `MIGRATED_LANGUAGES`, matching the Java migration pattern from #1482:
|
||||
* the resolver can run in shadow/forced mode, while production default
|
||||
* stays on the legacy DAG until registry-primary parity reaches the
|
||||
* RFC threshold. Forced mode currently passes 154/175 fixtures (88%),
|
||||
* including core import, receiver, companion, default-param, vararg,
|
||||
* constructor, local assignment-chain, and collection-iteration fixtures.
|
||||
* Remaining gaps are advanced TypeEnv behaviors such as smart casts,
|
||||
* cross-file iterable return propagation, method-chain fixpoint cases,
|
||||
* overload target-id selection, virtual dispatch, and interface default
|
||||
* method dispatch.
|
||||
*/
|
||||
export const kotlinScopeResolver: ScopeResolver = {
|
||||
language: SupportedLanguages.Kotlin,
|
||||
languageProvider: kotlinProvider,
|
||||
importEdgeReason: 'kotlin-scope: import',
|
||||
|
||||
resolveImportTarget: (targetRaw, fromFile, allFilePaths) => {
|
||||
const ws: KotlinResolveContext = { fromFile, allFilePaths };
|
||||
return resolveKotlinImportTarget(
|
||||
{ kind: 'named', localName: '_', importedName: '_', targetRaw },
|
||||
ws,
|
||||
);
|
||||
},
|
||||
|
||||
mergeBindings: (existing, incoming) => [...kotlinMergeBindings([...existing, ...incoming])],
|
||||
|
||||
arityCompatibility: (callsite, def) => kotlinArityCompatibility(def, callsite),
|
||||
|
||||
buildMro: (graph, parsedFiles, nodeLookup) =>
|
||||
buildMro(graph, parsedFiles, nodeLookup, defaultLinearize),
|
||||
|
||||
populateOwners: (parsed: ParsedFile) => populateKotlinOwners(parsed),
|
||||
|
||||
isSuperReceiver: (text) => text.trim() === 'super',
|
||||
|
||||
fieldFallbackOnMethodLookup: false,
|
||||
propagatesReturnTypesAcrossImports: true,
|
||||
collapseMemberCallsByCallerTarget: false,
|
||||
hoistTypeBindingsToModule: true,
|
||||
};
|
||||
@@ -0,0 +1,36 @@
|
||||
import type {
|
||||
CaptureMatch,
|
||||
ParsedImport,
|
||||
Scope,
|
||||
ScopeId,
|
||||
ScopeTree,
|
||||
TypeRef,
|
||||
} from 'gitnexus-shared';
|
||||
|
||||
export function kotlinBindingScopeFor(
|
||||
decl: CaptureMatch,
|
||||
innermost: Scope,
|
||||
tree: ScopeTree,
|
||||
): ScopeId | null {
|
||||
if (decl['@type-binding.return'] === undefined) return null;
|
||||
|
||||
let current: Scope | undefined = innermost;
|
||||
while (current !== undefined && current.kind !== 'Module') {
|
||||
if (current.parent === null) break;
|
||||
current = tree.getScope(current.parent);
|
||||
}
|
||||
return current?.kind === 'Module' ? current.id : null;
|
||||
}
|
||||
|
||||
export function kotlinImportOwningScope(
|
||||
_imp: ParsedImport,
|
||||
_innermost: Scope,
|
||||
_tree: ScopeTree,
|
||||
): ScopeId | null {
|
||||
return null;
|
||||
}
|
||||
|
||||
export function kotlinReceiverBinding(functionScope: Scope): TypeRef | null {
|
||||
if (functionScope.kind !== 'Function') return null;
|
||||
return functionScope.typeBindings.get('this') ?? functionScope.typeBindings.get('super') ?? null;
|
||||
}
|
||||
@@ -56,6 +56,16 @@ import {
|
||||
typescriptArityCompatibility,
|
||||
resolveTsImportTarget,
|
||||
} from './typescript/index.js';
|
||||
import {
|
||||
emitJsScopeCaptures,
|
||||
interpretJsImport,
|
||||
interpretJsTypeBinding,
|
||||
jsBindingScopeFor,
|
||||
jsImportOwningScope,
|
||||
jsReceiverBinding,
|
||||
jsMergeBindings,
|
||||
jsArityCompatibility,
|
||||
} from './javascript/index.js';
|
||||
|
||||
/**
|
||||
* TypeScript/JavaScript: arrow_function and function_expression are
|
||||
@@ -359,4 +369,19 @@ export const javascriptProvider = defineLanguage({
|
||||
classExtractor: createClassExtractor(javascriptClassConfig),
|
||||
heritageExtractor: createHeritageExtractor(SupportedLanguages.JavaScript),
|
||||
builtInNames: BUILT_INS,
|
||||
|
||||
// ── RFC #909 Ring 3: scope-based resolution hooks (RFC §5) ──────────
|
||||
// JavaScript is the fourth migration after Python, C#, and TypeScript.
|
||||
// Hooks are thin wrappers over the TypeScript implementations where
|
||||
// semantics are identical; JS-specific additions (CJS require(),
|
||||
// JSDoc type bindings) live in ./javascript/captures.ts.
|
||||
// See ./javascript/index.ts for the full per-module rationale.
|
||||
emitScopeCaptures: emitJsScopeCaptures,
|
||||
interpretImport: interpretJsImport,
|
||||
interpretTypeBinding: interpretJsTypeBinding,
|
||||
bindingScopeFor: jsBindingScopeFor,
|
||||
importOwningScope: jsImportOwningScope,
|
||||
mergeBindings: (_scope, bindings) => jsMergeBindings(bindings),
|
||||
receiverBinding: jsReceiverBinding,
|
||||
arityCompatibility: jsArityCompatibility,
|
||||
});
|
||||
|
||||
@@ -64,7 +64,7 @@ const CALL_TAGS = [
|
||||
'@reference.call.constructor',
|
||||
] as const;
|
||||
|
||||
function pickFirstDefined(grouped: CaptureMatch, tags: readonly string[]): Capture | undefined {
|
||||
function pickFirstCapture(grouped: CaptureMatch, tags: readonly string[]): Capture | undefined {
|
||||
for (const tag of tags) {
|
||||
const cap = grouped[tag];
|
||||
if (cap !== undefined) return cap;
|
||||
@@ -72,6 +72,17 @@ function pickFirstDefined(grouped: CaptureMatch, tags: readonly string[]): Captu
|
||||
return undefined;
|
||||
}
|
||||
|
||||
function pickFirstNode(
|
||||
grouped: Record<string, SyntaxNode | undefined>,
|
||||
tags: readonly string[],
|
||||
): SyntaxNode | undefined {
|
||||
for (const tag of tags) {
|
||||
const node = grouped[tag];
|
||||
if (node !== undefined) return node;
|
||||
}
|
||||
return undefined;
|
||||
}
|
||||
|
||||
/**
|
||||
* Drop `@reference.read.member` matches whose underlying `member_expression`
|
||||
* is NOT actually a read context:
|
||||
@@ -113,6 +124,34 @@ function shouldEmitReadMember(memberNode: SyntaxNode): boolean {
|
||||
}
|
||||
}
|
||||
|
||||
/** Walks the parent chain from `node` (inclusive), returning the first node
|
||||
* whose type matches, or null. Faster than `findNodeAtRange` when the caller
|
||||
* already holds the anchor node — avoids re-scanning the tree from the root. */
|
||||
function findSelfOrAncestorOfType(node: SyntaxNode | undefined, type: string): SyntaxNode | null {
|
||||
if (node === undefined) return null;
|
||||
let current: SyntaxNode | null = node;
|
||||
while (current !== null) {
|
||||
if (current.type === type) return current;
|
||||
current = current.parent;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/** Walks the parent chain from `node` (inclusive), returning the first node
|
||||
* whose type is in the set, or null. Plural form of {@link findSelfOrAncestorOfType}. */
|
||||
function findSelfOrAncestorOfTypes(
|
||||
node: SyntaxNode | undefined,
|
||||
types: readonly string[],
|
||||
): SyntaxNode | null {
|
||||
if (node === undefined) return null;
|
||||
let current: SyntaxNode | null = node;
|
||||
while (current !== null) {
|
||||
if (types.includes(current.type)) return current;
|
||||
current = current.parent;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
export function emitTsScopeCaptures(
|
||||
sourceText: string,
|
||||
filePath: string,
|
||||
@@ -151,9 +190,11 @@ export function emitTsScopeCaptures(
|
||||
// `@`; we put it back so the central extractor's prefix lookups
|
||||
// (`@scope.`, `@declaration.`, …) work.
|
||||
const grouped: Record<string, Capture> = {};
|
||||
const groupedNodes: Record<string, SyntaxNode> = {};
|
||||
for (const c of m.captures) {
|
||||
const tag = '@' + c.name;
|
||||
grouped[tag] = nodeToCapture(tag, c.node);
|
||||
groupedNodes[tag] = c.node;
|
||||
}
|
||||
if (Object.keys(grouped).length === 0) continue;
|
||||
|
||||
@@ -165,6 +206,10 @@ export function emitTsScopeCaptures(
|
||||
if (grouped['@import.statement'] !== undefined) {
|
||||
const stmtCapture = grouped['@import.statement'];
|
||||
const stmtNode =
|
||||
findSelfOrAncestorOfTypes(groupedNodes['@import.statement'], [
|
||||
'import_statement',
|
||||
'export_statement',
|
||||
]) ??
|
||||
findNodeAtRange(tree.rootNode, stmtCapture.range, 'import_statement') ??
|
||||
findNodeAtRange(tree.rootNode, stmtCapture.range, 'export_statement');
|
||||
if (stmtNode !== null) {
|
||||
@@ -183,7 +228,9 @@ export function emitTsScopeCaptures(
|
||||
// `splitDynamicImport` branch consumes.
|
||||
if (grouped['@import.dynamic'] !== undefined) {
|
||||
const dynCapture = grouped['@import.dynamic'];
|
||||
const callNode = findNodeAtRange(tree.rootNode, dynCapture.range, 'call_expression');
|
||||
const callNode =
|
||||
findSelfOrAncestorOfType(groupedNodes['@import.dynamic'], 'call_expression') ??
|
||||
findNodeAtRange(tree.rootNode, dynCapture.range, 'call_expression');
|
||||
if (callNode !== null) {
|
||||
const decomposed = splitImportStatement(callNode);
|
||||
for (const d of decomposed) out.push(d);
|
||||
@@ -197,7 +244,9 @@ export function emitTsScopeCaptures(
|
||||
// we rely on this emit-side filter so the query stays simple.
|
||||
if (grouped['@reference.read.member'] !== undefined) {
|
||||
const anchor = grouped['@reference.read.member'];
|
||||
const memberNode = findNodeAtRange(tree.rootNode, anchor.range, 'member_expression');
|
||||
const memberNode =
|
||||
findSelfOrAncestorOfType(groupedNodes['@reference.read.member'], 'member_expression') ??
|
||||
findNodeAtRange(tree.rootNode, anchor.range, 'member_expression');
|
||||
if (memberNode === null || !shouldEmitReadMember(memberNode)) {
|
||||
continue;
|
||||
}
|
||||
@@ -208,9 +257,10 @@ export function emitTsScopeCaptures(
|
||||
// overloads — TypeScript supports overload signatures via
|
||||
// function_signature, so `parameterTypes` is populated when
|
||||
// available.
|
||||
const declAnchor = pickFirstDefined(grouped, FUNCTION_DECL_TAGS);
|
||||
const declAnchor = pickFirstCapture(grouped, FUNCTION_DECL_TAGS);
|
||||
const declAnchorNode = pickFirstNode(groupedNodes, FUNCTION_DECL_TAGS);
|
||||
if (declAnchor !== undefined) {
|
||||
const fnNode = findFunctionNode(tree.rootNode, declAnchor.range);
|
||||
const fnNode = findFunctionNode(tree.rootNode, declAnchor.range, declAnchorNode);
|
||||
if (fnNode !== null) {
|
||||
const arity = computeTsArityMetadata(fnNode);
|
||||
if (arity.parameterCount !== undefined) {
|
||||
@@ -255,9 +305,11 @@ export function emitTsScopeCaptures(
|
||||
// calls to disambiguate by props-arity, a JSX-aware arity
|
||||
// synthesizer would need to count `jsx_attribute` children of the
|
||||
// opening tag instead of `arguments`.
|
||||
const callAnchor = pickFirstDefined(grouped, CALL_TAGS);
|
||||
const callAnchor = pickFirstCapture(grouped, CALL_TAGS);
|
||||
const callAnchorNode = pickFirstNode(groupedNodes, CALL_TAGS);
|
||||
if (callAnchor !== undefined && grouped['@reference.arity'] === undefined) {
|
||||
const callNode =
|
||||
findSelfOrAncestorOfTypes(callAnchorNode, ['call_expression', 'new_expression']) ??
|
||||
findNodeAtRange(tree.rootNode, callAnchor.range, 'call_expression') ??
|
||||
findNodeAtRange(tree.rootNode, callAnchor.range, 'new_expression');
|
||||
if (callNode !== null) {
|
||||
@@ -293,7 +345,11 @@ export function emitTsScopeCaptures(
|
||||
// lookup instead of synthesis — covered by `tsReceiverBinding`.
|
||||
const scopeFnAnchor = grouped['@scope.function'];
|
||||
if (scopeFnAnchor !== undefined) {
|
||||
const fnNode = findFunctionNode(tree.rootNode, scopeFnAnchor.range);
|
||||
const fnNode = findFunctionNode(
|
||||
tree.rootNode,
|
||||
scopeFnAnchor.range,
|
||||
groupedNodes['@scope.function'],
|
||||
);
|
||||
if (fnNode !== null) {
|
||||
const synth = synthesizeTsReceiverBinding(fnNode);
|
||||
if (synth !== null) out.push(synth);
|
||||
@@ -518,7 +574,13 @@ function inferArgType(argNode: SyntaxNode): string {
|
||||
* The `@scope.function` anchor range covers the whole node, but the
|
||||
* tag alone doesn't identify which node type among the many TS
|
||||
* function-likes. */
|
||||
function findFunctionNode(rootNode: SyntaxNode, range: Capture['range']): SyntaxNode | null {
|
||||
function findFunctionNode(
|
||||
rootNode: SyntaxNode,
|
||||
range: Capture['range'],
|
||||
anchorNode?: SyntaxNode,
|
||||
): SyntaxNode | null {
|
||||
const fromAnchor = findSelfOrAncestorOfTypes(anchorNode, FUNCTION_NODE_TYPES);
|
||||
if (fromAnchor !== null) return fromAnchor;
|
||||
for (const nodeType of FUNCTION_NODE_TYPES) {
|
||||
const n = findNodeAtRange(rootNode, range, nodeType);
|
||||
if (n !== null) return n;
|
||||
|
||||
@@ -75,8 +75,11 @@ export function tsBindingScopeFor(
|
||||
* any of `kinds`. Returns the matching scope's id or `null` when no
|
||||
* ancestor matches (e.g., a return type binding emitted outside any
|
||||
* Module scope — shouldn't happen in well-formed input).
|
||||
*
|
||||
* Exported so language-specific hook wrappers (e.g. `jsBindingScopeFor`)
|
||||
* can reuse it without duplicating the traversal logic.
|
||||
*/
|
||||
function walkToScope(
|
||||
export function walkToScope(
|
||||
from: Scope,
|
||||
tree: ScopeTree,
|
||||
...kinds: readonly Scope['kind'][]
|
||||
|
||||
@@ -2,18 +2,35 @@
|
||||
* Field Registry
|
||||
*
|
||||
* Owner-scoped field/property index extracted from SymbolTable.
|
||||
* Stores Property symbols keyed by `ownerNodeId\0fieldName` for O(1) lookup.
|
||||
* Stores Property / Variable / Const / Static symbols keyed by
|
||||
* `ownerNodeId\0fieldName` for O(1) lookup. Supports multiple defs
|
||||
* under the same (owner, name) — e.g. legacy Property plus a
|
||||
* scope-resolution Variable reconciliation entry.
|
||||
*/
|
||||
|
||||
import type { SymbolDefinition } from 'gitnexus-shared';
|
||||
|
||||
const EMPTY: readonly SymbolDefinition[] = Object.freeze([]);
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Public read-only interface
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
export interface FieldRegistry {
|
||||
/** Look up a field/property by its owning class nodeId and field name. */
|
||||
/**
|
||||
* First field registered under `(ownerNodeId, fieldName)`, if any.
|
||||
* Registration order is first-wins: when a Property and a Variable share
|
||||
* an `(owner, simpleName)` key, the earlier `register(...)` call's def is
|
||||
* returned. Prefer `lookupAllByOwner` when overloads or duplicate-kind
|
||||
* entries under the same name must all be visible.
|
||||
*/
|
||||
lookupFieldByOwner(ownerNodeId: string, fieldName: string): SymbolDefinition | undefined;
|
||||
|
||||
/**
|
||||
* Every field registered under `(ownerNodeId, fieldName)` in registration
|
||||
* order. Returns `[]` on miss.
|
||||
*/
|
||||
lookupAllByOwner(ownerNodeId: string, fieldName: string): readonly SymbolDefinition[];
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -21,7 +38,7 @@ export interface FieldRegistry {
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
export interface MutableFieldRegistry extends FieldRegistry {
|
||||
/** Register a field/property under its owner. */
|
||||
/** Register a field under its owner. Appends when the key already exists. */
|
||||
register(ownerNodeId: string, fieldName: string, def: SymbolDefinition): void;
|
||||
/** Clear all entries. */
|
||||
clear(): void;
|
||||
@@ -32,22 +49,36 @@ export interface MutableFieldRegistry extends FieldRegistry {
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
export const createFieldRegistry = (): MutableFieldRegistry => {
|
||||
const fieldByOwner = new Map<string, SymbolDefinition>();
|
||||
const fieldByOwner = new Map<string, SymbolDefinition[]>();
|
||||
|
||||
const lookupAllByOwner = (
|
||||
ownerNodeId: string,
|
||||
fieldName: string,
|
||||
): readonly SymbolDefinition[] => {
|
||||
return fieldByOwner.get(`${ownerNodeId}\0${fieldName}`) ?? EMPTY;
|
||||
};
|
||||
|
||||
const lookupFieldByOwner = (
|
||||
ownerNodeId: string,
|
||||
fieldName: string,
|
||||
): SymbolDefinition | undefined => {
|
||||
return fieldByOwner.get(`${ownerNodeId}\0${fieldName}`);
|
||||
const pool = lookupAllByOwner(ownerNodeId, fieldName);
|
||||
return pool.length === 0 ? undefined : pool[0];
|
||||
};
|
||||
|
||||
const register = (ownerNodeId: string, fieldName: string, def: SymbolDefinition): void => {
|
||||
fieldByOwner.set(`${ownerNodeId}\0${fieldName}`, def);
|
||||
const key = `${ownerNodeId}\0${fieldName}`;
|
||||
const existing = fieldByOwner.get(key);
|
||||
if (existing) {
|
||||
existing.push(def);
|
||||
} else {
|
||||
fieldByOwner.set(key, [def]);
|
||||
}
|
||||
};
|
||||
|
||||
const clear = (): void => {
|
||||
fieldByOwner.clear();
|
||||
};
|
||||
|
||||
return { lookupFieldByOwner, register, clear };
|
||||
return { lookupFieldByOwner, lookupAllByOwner, register, clear };
|
||||
};
|
||||
|
||||
@@ -0,0 +1,45 @@
|
||||
/**
|
||||
* Owner-keyed member lookup for Step 2 (RFC #909 / PR #1656).
|
||||
*
|
||||
* Merges MethodRegistry + FieldRegistry hits for `(ownerDefId, memberName)`
|
||||
* in O(1) map time per registry — no `defs.byId` scan. Callers that omit
|
||||
* this helper and leave `ownedMembersByOwner` unset fall back to an O(|defs|)
|
||||
* compatibility scan inside `lookupCore.collectOwnedMembers`.
|
||||
*/
|
||||
|
||||
import type { DefId, SymbolDefinition } from 'gitnexus-shared';
|
||||
import type { SemanticModel } from './semantic-model.js';
|
||||
|
||||
const EMPTY: readonly SymbolDefinition[] = Object.freeze([]);
|
||||
|
||||
/**
|
||||
* Production hook for `RegistryContext.ownedMembersByOwner`.
|
||||
* Returns `[]` on miss (authoritative indexed empty) — never `undefined`.
|
||||
*
|
||||
* Merges hits from all three owner-keyed registries (methods, fields,
|
||||
* nested types) under the same `(ownerDefId, memberName)` key. The
|
||||
* caller's `acceptedKinds` filter in `lookupCore` picks the right subset.
|
||||
*/
|
||||
export function lookupOwnedMembersByOwner(
|
||||
model: Pick<SemanticModel, 'methods' | 'fields' | 'types'>,
|
||||
ownerDefId: DefId,
|
||||
memberName: string,
|
||||
): readonly SymbolDefinition[] {
|
||||
const methods = model.methods.lookupAllByOwner(ownerDefId, memberName);
|
||||
const fields = model.fields.lookupAllByOwner(ownerDefId, memberName);
|
||||
const nestedTypes = model.types.lookupAllByOwner(ownerDefId, memberName);
|
||||
const methodCount = methods.length;
|
||||
const fieldCount = fields.length;
|
||||
const typeCount = nestedTypes.length;
|
||||
const total = methodCount + fieldCount + typeCount;
|
||||
if (total === 0) return EMPTY;
|
||||
if (methodCount === total) return methods;
|
||||
if (fieldCount === total) return fields;
|
||||
if (typeCount === total) return nestedTypes;
|
||||
const merged = new Array<SymbolDefinition>(total);
|
||||
let i = 0;
|
||||
for (let j = 0; j < methodCount; j++) merged[i++] = methods[j]!;
|
||||
for (let j = 0; j < fieldCount; j++) merged[i++] = fields[j]!;
|
||||
for (let j = 0; j < typeCount; j++) merged[i++] = nestedTypes[j]!;
|
||||
return merged;
|
||||
}
|
||||
@@ -8,6 +8,8 @@
|
||||
|
||||
import type { SymbolDefinition } from 'gitnexus-shared';
|
||||
|
||||
const EMPTY: readonly SymbolDefinition[] = Object.freeze([]);
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Public read-only interface
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -35,6 +37,14 @@ export interface TypeRegistry {
|
||||
* Returned array is a view into the live index — do not mutate.
|
||||
*/
|
||||
lookupImplByName(name: string): readonly SymbolDefinition[];
|
||||
|
||||
/**
|
||||
* Look up nested-type defs registered under `(ownerNodeId, simpleName)`
|
||||
* in registration order. Returns `[]` on miss. Used by Step 2 Receiver/MRO
|
||||
* resolution when the receiver's owner declares nested classes/structs/
|
||||
* enums/typedefs/etc. that the caller's `acceptedKinds` includes.
|
||||
*/
|
||||
lookupAllByOwner(ownerNodeId: string, simpleName: string): readonly SymbolDefinition[];
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -46,6 +56,8 @@ export interface MutableTypeRegistry extends TypeRegistry {
|
||||
registerClass(name: string, qualifiedName: string, def: SymbolDefinition): void;
|
||||
/** Register a Rust Impl block by name. */
|
||||
registerImpl(name: string, def: SymbolDefinition): void;
|
||||
/** Register a nested type under its owner. Appends when the key already exists. */
|
||||
registerByOwner(ownerNodeId: string, simpleName: string, def: SymbolDefinition): void;
|
||||
/** Clear all entries. */
|
||||
clear(): void;
|
||||
}
|
||||
@@ -58,6 +70,7 @@ export const createTypeRegistry = (): MutableTypeRegistry => {
|
||||
const classByName = new Map<string, SymbolDefinition[]>();
|
||||
const classByQualifiedName = new Map<string, SymbolDefinition[]>();
|
||||
const implByName = new Map<string, SymbolDefinition[]>();
|
||||
const nestedByOwner = new Map<string, SymbolDefinition[]>();
|
||||
|
||||
const lookupClassByName = (name: string): SymbolDefinition[] => {
|
||||
return classByName.get(name) ?? [];
|
||||
@@ -71,6 +84,13 @@ export const createTypeRegistry = (): MutableTypeRegistry => {
|
||||
return implByName.get(name) ?? [];
|
||||
};
|
||||
|
||||
const lookupAllByOwner = (
|
||||
ownerNodeId: string,
|
||||
simpleName: string,
|
||||
): readonly SymbolDefinition[] => {
|
||||
return nestedByOwner.get(`${ownerNodeId}\0${simpleName}`) ?? EMPTY;
|
||||
};
|
||||
|
||||
const registerClass = (name: string, qualifiedName: string, def: SymbolDefinition): void => {
|
||||
const existing = classByName.get(name);
|
||||
if (existing) {
|
||||
@@ -96,18 +116,35 @@ export const createTypeRegistry = (): MutableTypeRegistry => {
|
||||
}
|
||||
};
|
||||
|
||||
const registerByOwner = (
|
||||
ownerNodeId: string,
|
||||
simpleName: string,
|
||||
def: SymbolDefinition,
|
||||
): void => {
|
||||
const key = `${ownerNodeId}\0${simpleName}`;
|
||||
const existing = nestedByOwner.get(key);
|
||||
if (existing) {
|
||||
existing.push(def);
|
||||
} else {
|
||||
nestedByOwner.set(key, [def]);
|
||||
}
|
||||
};
|
||||
|
||||
const clear = (): void => {
|
||||
classByName.clear();
|
||||
classByQualifiedName.clear();
|
||||
implByName.clear();
|
||||
nestedByOwner.clear();
|
||||
};
|
||||
|
||||
return {
|
||||
lookupClassByName,
|
||||
lookupClassByQualifiedName,
|
||||
lookupImplByName,
|
||||
lookupAllByOwner,
|
||||
registerClass,
|
||||
registerImpl,
|
||||
registerByOwner,
|
||||
clear,
|
||||
};
|
||||
};
|
||||
|
||||
@@ -832,6 +832,14 @@ const processParsingSequential = async (
|
||||
// Public API
|
||||
// ============================================================================
|
||||
|
||||
/**
|
||||
* Per-`WorkerPool` log-dedup state for quarantine reporting. Keyed on the
|
||||
* pool instance so multiple concurrent pools (test fixtures, future
|
||||
* multi-pool callers) each get their own seen-set. WeakMap entries vanish
|
||||
* when the pool is garbage-collected.
|
||||
*/
|
||||
const loggedQuarantineByPool = new WeakMap<WorkerPool, Set<string>>();
|
||||
|
||||
export const processParsing = async (
|
||||
graph: KnowledgeGraph,
|
||||
files: { path: string; content: string }[],
|
||||
@@ -874,25 +882,75 @@ export const processParsing = async (
|
||||
`[scope-resolution prof] worker pool engaged for ${files.length} files — cross-phase tree cache will be empty; scope-resolution re-parses.`,
|
||||
);
|
||||
}
|
||||
try {
|
||||
return await processParsingWithWorkers(
|
||||
graph,
|
||||
files,
|
||||
symbolTable,
|
||||
astCache,
|
||||
workerPool,
|
||||
reportProgress,
|
||||
outRawResults,
|
||||
);
|
||||
} catch (err) {
|
||||
const message = err instanceof Error ? err.message : String(err);
|
||||
logger.warn({ message }, 'Worker pool parsing stopped; continuing with sequential parser:');
|
||||
reportProgress?.(
|
||||
lastProgress,
|
||||
files.length,
|
||||
`Sequential fallback after worker issue: ${message}`,
|
||||
);
|
||||
// U20 design pivot: the worker pool's resilience layers
|
||||
// (respawn budget, circuit breaker, quarantine, slot-attribution,
|
||||
// cumulative timeout) are the SOLE contract for handling worker
|
||||
// failures. There is no sequential-parser fallback for either
|
||||
// partial quarantine or full pool failure — the operator must see
|
||||
// a clear hard signal when workers can't recover, instead of a
|
||||
// silently-degraded graph from a possibly-crashing main-thread
|
||||
// sequential parser. A failing tree-sitter native binding that
|
||||
// quarantined a worker would, under the previous design, re-trigger
|
||||
// the same SIGSEGV on the main thread; we avoid that risk entirely.
|
||||
//
|
||||
// - Partial quarantine: the file is missing from this run's graph;
|
||||
// the per-chunk warn log below surfaces it; U2's chunk-cache
|
||||
// write-guard in parse-impl.ts keeps the chunk uncached so the
|
||||
// next analyze gets a cache miss and a fresh pool retries.
|
||||
// - Full pool failure: `WorkerPoolDispatchError` propagates from
|
||||
// `processParsingWithWorkers` up through this function. The
|
||||
// analyze run errors out instead of falling back to sequential.
|
||||
const data = await processParsingWithWorkers(
|
||||
graph,
|
||||
files,
|
||||
symbolTable,
|
||||
astCache,
|
||||
workerPool,
|
||||
reportProgress,
|
||||
outRawResults,
|
||||
);
|
||||
// Session-scoped quarantine (worker-pool resilience Layer 3): surface
|
||||
// any files this pool has decided are unsafe for workers so the
|
||||
// operator can see what was skipped. The pool already filtered them
|
||||
// out of dispatch; we only need to log + progress-report. Quarantine
|
||||
// is session-scoped per pool instance — a fresh `createWorkerPool`
|
||||
// call clears it.
|
||||
//
|
||||
// Dedup: log full path list only for entries newly quarantined since
|
||||
// the previous dispatch on the same pool. The per-chunk progress
|
||||
// message still surfaces the count for UX continuity, but the
|
||||
// structured `quarantinedFiles` payload is only emitted when there
|
||||
// is new signal — prevents O(quarantine × chunks) log spam.
|
||||
const quarantineSnapshot = workerPool.getQuarantinedPaths?.() ?? [];
|
||||
const quarantineSet = new Set(quarantineSnapshot);
|
||||
if (quarantineSet.size > 0) {
|
||||
const quarantinedInChunk = files.filter((file) => quarantineSet.has(file.path));
|
||||
if (quarantinedInChunk.length > 0) {
|
||||
const seenForPool = loggedQuarantineByPool.get(workerPool) ?? new Set<string>();
|
||||
const newlyQuarantined = quarantinedInChunk
|
||||
.map((file) => file.path)
|
||||
.filter((p) => !seenForPool.has(p));
|
||||
for (const p of newlyQuarantined) seenForPool.add(p);
|
||||
loggedQuarantineByPool.set(workerPool, seenForPool);
|
||||
if (newlyQuarantined.length > 0) {
|
||||
logger.warn(
|
||||
{
|
||||
newlyQuarantined,
|
||||
cumulativeQuarantine: quarantineSet.size,
|
||||
chunkSkipped: quarantinedInChunk.length,
|
||||
},
|
||||
`Worker quarantine: ${newlyQuarantined.length} new file(s) skipped this chunk ` +
|
||||
`(${quarantinedInChunk.length} skipped total, ${quarantineSet.size} cumulative).`,
|
||||
);
|
||||
}
|
||||
reportProgress?.(
|
||||
lastProgress,
|
||||
files.length,
|
||||
`${quarantinedInChunk.length} worker-quarantined file(s) skipped`,
|
||||
);
|
||||
}
|
||||
}
|
||||
return data;
|
||||
}
|
||||
|
||||
// Fallback: sequential parsing (no pre-extracted data)
|
||||
|
||||
@@ -55,6 +55,7 @@ import type {
|
||||
ExtractedCall,
|
||||
ExtractedDecoratorRoute,
|
||||
ExtractedFetchCall,
|
||||
ExtractedImport,
|
||||
ExtractedORMQuery,
|
||||
ExtractedRoute,
|
||||
ExtractedToolDef,
|
||||
@@ -69,6 +70,7 @@ import path from 'node:path';
|
||||
import { fileURLToPath, pathToFileURL } from 'node:url';
|
||||
|
||||
import { isDev } from '../utils/env.js';
|
||||
import { isVerboseIngestionEnabled } from '../utils/verbose.js';
|
||||
import { synthesizeWildcardImportBindings, needsSynthesis } from './wildcard-synthesis.js';
|
||||
import { extractORMQueriesInline } from './orm-extraction.js';
|
||||
|
||||
@@ -85,11 +87,24 @@ import { logger } from '../../logger.js';
|
||||
* gives a useful invalidation floor (~1/N chunks on a multi-MB repo)
|
||||
* while keeping worker dispatch overhead under 5% on cold runs.
|
||||
*/
|
||||
const CHUNK_BYTE_BUDGET = (() => {
|
||||
/**
|
||||
* Built-in chunk byte budget when neither `PipelineOptions.chunkByteBudget`
|
||||
* nor `GITNEXUS_CHUNK_BYTE_BUDGET` is set. Tuned to give a useful
|
||||
* cache-invalidation floor (~1/N chunks on a multi-MB repo) while keeping
|
||||
* worker dispatch overhead under 5% on cold runs. Resolution happens at
|
||||
* call time inside `runChunkedParseAndResolve` (U14 from PR #1693 review)
|
||||
* — previously this was a module-load IIFE, which froze the env value at
|
||||
* import time and meant per-call option threading silently no-op'd.
|
||||
*/
|
||||
const DEFAULT_CHUNK_BYTE_BUDGET = 2 * 1024 * 1024;
|
||||
|
||||
function resolveChunkByteBudget(options?: PipelineOptions): number {
|
||||
const opt = options?.chunkByteBudget;
|
||||
if (typeof opt === 'number' && Number.isFinite(opt) && opt > 0) return opt;
|
||||
const env = Number(process.env.GITNEXUS_CHUNK_BYTE_BUDGET);
|
||||
if (Number.isFinite(env) && env > 0) return env;
|
||||
return 2 * 1024 * 1024;
|
||||
})();
|
||||
return DEFAULT_CHUNK_BYTE_BUDGET;
|
||||
}
|
||||
|
||||
// ── Main parse + resolve function ──────────────────────────────────────────
|
||||
|
||||
@@ -177,18 +192,28 @@ export async function runChunkedParseAndResolve(
|
||||
if (totalParseable === 0) {
|
||||
onProgress({
|
||||
phase: 'parsing',
|
||||
percent: 82,
|
||||
// Skip directly to the end of the parse-phase progress band (M2 from PR
|
||||
// #1693 review). Parse 20-70%, deferred 70-95%; nothing in either runs
|
||||
// when there's no parseable file, so jump to 95.
|
||||
percent: 95,
|
||||
message: 'No parseable files found — skipping parsing phase',
|
||||
stats: { filesProcessed: 0, totalFiles: 0, nodesCreated: graph.nodeCount },
|
||||
});
|
||||
}
|
||||
|
||||
// Build byte-budget chunks
|
||||
// Build byte-budget chunks. The budget is resolved per-call (U14): options
|
||||
// first, then env, then the built-in default. Pre-U14 this was a
|
||||
// module-load IIFE constant, which froze the env value at import time
|
||||
// and made `PipelineOptions.chunkByteBudget` silently no-op on warm test
|
||||
// runs. Resolving in the function body restores per-call configurability
|
||||
// and matches the pattern used by resolveAutoPoolSize and the U1
|
||||
// parseChunkConcurrency resolver.
|
||||
const chunkByteBudget = resolveChunkByteBudget(options);
|
||||
const chunks: string[][] = [];
|
||||
let currentChunk: string[] = [];
|
||||
let currentBytes = 0;
|
||||
for (const file of parseableScanned) {
|
||||
if (currentChunk.length > 0 && currentBytes + file.size > CHUNK_BYTE_BUDGET) {
|
||||
if (currentChunk.length > 0 && currentBytes + file.size > chunkByteBudget) {
|
||||
chunks.push(currentChunk);
|
||||
currentChunk = [];
|
||||
currentBytes = 0;
|
||||
@@ -203,16 +228,22 @@ export async function runChunkedParseAndResolve(
|
||||
if (isDev) {
|
||||
const totalMB = parseableScanned.reduce((s, f) => s + f.size, 0) / (1024 * 1024);
|
||||
logger.info(
|
||||
`📂 Scan: ${totalFiles} paths, ${totalParseable} parseable (${totalMB.toFixed(0)}MB), ${numChunks} chunks @ ${CHUNK_BYTE_BUDGET / (1024 * 1024)}MB budget`,
|
||||
`📂 Scan: ${totalFiles} paths, ${totalParseable} parseable (${totalMB.toFixed(0)}MB), ${numChunks} chunks @ ${chunkByteBudget / (1024 * 1024)}MB budget`,
|
||||
);
|
||||
}
|
||||
|
||||
onProgress({
|
||||
phase: 'parsing',
|
||||
percent: 20,
|
||||
message: `Parsing ${totalParseable} files in ${numChunks} chunk${numChunks !== 1 ? 's' : ''}...`,
|
||||
stats: { filesProcessed: 0, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
|
||||
});
|
||||
// Skip the "Parsing N files..." announcement when there's nothing to parse
|
||||
// — the early-return branch above already emitted percent 95 ("skipping
|
||||
// parsing phase"), and emitting percent 20 here would regress the
|
||||
// progress stream non-monotonically (M2 from PR #1693 review).
|
||||
if (totalParseable > 0) {
|
||||
onProgress({
|
||||
phase: 'parsing',
|
||||
percent: 20,
|
||||
message: `Parsing ${totalParseable} files in ${numChunks} chunk${numChunks !== 1 ? 's' : ''}...`,
|
||||
stats: { filesProcessed: 0, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
|
||||
});
|
||||
}
|
||||
|
||||
// Don't spawn workers for tiny repos — overhead exceeds benefit.
|
||||
// Test suites may lower the thresholds via `options.workerThresholdsForTest`
|
||||
@@ -221,18 +252,33 @@ export async function runChunkedParseAndResolve(
|
||||
const MIN_BYTES_FOR_WORKERS = options?.workerThresholdsForTest?.minBytes ?? 512 * 1024;
|
||||
const totalBytes = parseableScanned.reduce((s, f) => s + f.size, 0);
|
||||
|
||||
// Create worker pool once, reuse across chunks
|
||||
// Create worker pool once, reuse across chunks.
|
||||
//
|
||||
// `workerPoolSize === 0` is a programmatic equivalent of `skipWorkers:
|
||||
// true` per the `PipelineOptions.workerPoolSize` contract. Short-
|
||||
// circuiting here avoids constructing a useless pool that rejects
|
||||
// every dispatch (with a `Worker pool parsing stopped` warn log per
|
||||
// chunk) just to fall back to the sequential path via the error
|
||||
// catch — the gate honors the docstring directly.
|
||||
let workerPool: WorkerPool | undefined;
|
||||
if (
|
||||
!options?.skipWorkers &&
|
||||
options?.workerPoolSize !== 0 &&
|
||||
(totalParseable >= MIN_FILES_FOR_WORKERS || totalBytes >= MIN_BYTES_FOR_WORKERS)
|
||||
) {
|
||||
try {
|
||||
let workerUrl = new URL('../workers/parse-worker.js', import.meta.url);
|
||||
// U20.U3 test-only injection: integration tests pass a custom
|
||||
// worker script URL via `workerUrlForTest` (mirrors the
|
||||
// `workerThresholdsForTest` precedent) so they can drive the
|
||||
// chunk-loop with deterministically-misbehaving workers without
|
||||
// mocking the module import graph. When unset, the normal src/
|
||||
// → dist/ resolution runs.
|
||||
let workerUrl =
|
||||
options?.workerUrlForTest ?? new URL('../workers/parse-worker.js', import.meta.url);
|
||||
// When running under vitest, import.meta.url points to src/ where no .js exists.
|
||||
// Fall back to the compiled dist/ worker so the pool can spawn real worker threads.
|
||||
const thisDir = fileURLToPath(new URL('.', import.meta.url));
|
||||
if (!fs.existsSync(fileURLToPath(workerUrl))) {
|
||||
if (!options?.workerUrlForTest && !fs.existsSync(fileURLToPath(workerUrl))) {
|
||||
const distWorker = path.resolve(
|
||||
thisDir,
|
||||
'..',
|
||||
@@ -249,7 +295,7 @@ export async function runChunkedParseAndResolve(
|
||||
workerUrl = pathToFileURL(distWorker);
|
||||
}
|
||||
}
|
||||
workerPool = createWorkerPool(workerUrl);
|
||||
workerPool = createWorkerPool(workerUrl, options?.workerPoolSize);
|
||||
} catch (err) {
|
||||
logger.warn(
|
||||
{ err: (err as Error).message },
|
||||
@@ -301,6 +347,16 @@ export async function runChunkedParseAndResolve(
|
||||
const deferredWorkerHeritage: ExtractedHeritage[] = [];
|
||||
const deferredConstructorBindings: FileConstructorBindings[] = [];
|
||||
const deferredAssignments: ExtractedAssignment[] = [];
|
||||
// Imports accumulated across chunks. Previously processed per-chunk
|
||||
// via `processImportsFromExtracted` inside the chunk loop, which
|
||||
// forced workers to sit idle on the main thread's extraction pass
|
||||
// between chunk dispatches (4-5% CPU utilization symptom). Deferring
|
||||
// to a single end-of-loop pass lets the worker pool start chunk N+1
|
||||
// immediately after chunk N's worker dispatch returns. Resolution is
|
||||
// strictly-more-information at end-of-loop because graph now has
|
||||
// every chunk's symbols — improves cross-chunk import targets.
|
||||
const deferredWorkerImports: ExtractedImport[] = [];
|
||||
let anyChunkNeedsWildcardSynth = false;
|
||||
// Aggregated per-file ParsedFile artifacts produced by workers' calls
|
||||
// to `extractParsedFile`. Threaded through to the scope-resolution
|
||||
// phase so it can SKIP its own re-extraction on cache hits — this is
|
||||
@@ -317,10 +373,54 @@ export async function runChunkedParseAndResolve(
|
||||
let chunkCacheMisses = 0;
|
||||
|
||||
try {
|
||||
// U1 — bounded chunk concurrency (B1 from PR #1693 review): pre-fetch
|
||||
// chunk file contents up to `parseChunkConcurrency` chunks ahead of the
|
||||
// dispatch cursor so file I/O overlaps with worker compute. Worker
|
||||
// dispatch itself stays serial because `WorkerPool.dispatch` is not
|
||||
// reentrant (concurrent calls would race on the shared per-slot
|
||||
// busy/in-flight state). With concurrency=1 behavior is identical to
|
||||
// the pure-serial loop. F4: deferred-state aggregation still happens
|
||||
// in chunkIdx order (the for-loop below iterates sequentially), so
|
||||
// cross-chunk processors see deterministic input regardless of
|
||||
// file-read completion order. Honors options.parseChunkConcurrency
|
||||
// (threaded from the CLI), then GITNEXUS_PARSE_CHUNK_CONCURRENCY env
|
||||
// (default 2 — matches the help text the CLI advertises).
|
||||
const parseChunkConcurrency = ((): number => {
|
||||
const opt = options?.parseChunkConcurrency;
|
||||
if (typeof opt === 'number' && Number.isInteger(opt) && opt >= 1) return opt;
|
||||
const env = Number(process.env.GITNEXUS_PARSE_CHUNK_CONCURRENCY);
|
||||
if (Number.isInteger(env) && env >= 1) return env;
|
||||
return 2;
|
||||
})();
|
||||
const chunkContentPromises = new Array<Promise<Map<string, string>> | undefined>(numChunks);
|
||||
const startChunkPrefetch = (i: number): void => {
|
||||
if (i >= numChunks || chunkContentPromises[i] !== undefined) return;
|
||||
chunkContentPromises[i] = readFileContents(repoPath, chunks[i]);
|
||||
};
|
||||
for (let i = 0; i < Math.min(parseChunkConcurrency, numChunks); i++) {
|
||||
startChunkPrefetch(i);
|
||||
}
|
||||
|
||||
// Hoisted loop-invariant: GITNEXUS_VERBOSE / NODE_ENV are read once
|
||||
// (not on every chunk). Previously evaluated at the top of the loop
|
||||
// body, which re-read process.env on every iteration even though
|
||||
// the env can't change mid-run.
|
||||
const verboseThroughputLog = isDev || isVerboseIngestionEnabled();
|
||||
|
||||
for (let chunkIdx = 0; chunkIdx < numChunks; chunkIdx++) {
|
||||
const chunkPaths = chunks[chunkIdx];
|
||||
// Start wall-clock for the per-chunk throughput log emitted at end
|
||||
// of this iteration. The gate is computed once above; here we just
|
||||
// sample the clock if the gate is on. Computed when either
|
||||
// NODE_ENV=development OR the operator passed `--verbose`
|
||||
// (GITNEXUS_VERBOSE) — the previous `isDev`-only gate meant
|
||||
// operators running `gitnexus analyze --verbose` in production
|
||||
// never saw the log (M3 from PR #1693 review).
|
||||
const chunkStartMs: number | null = verboseThroughputLog ? Date.now() : null;
|
||||
|
||||
const chunkContents = await readFileContents(repoPath, chunkPaths);
|
||||
const chunkContents = await chunkContentPromises[chunkIdx]!;
|
||||
chunkContentPromises[chunkIdx] = undefined; // release the in-memory copy
|
||||
startChunkPrefetch(chunkIdx + parseChunkConcurrency);
|
||||
const chunkFiles = chunkPaths
|
||||
.filter((p) => chunkContents.has(p))
|
||||
.map((p) => ({ path: p, content: chunkContents.get(p)! }));
|
||||
@@ -357,7 +457,11 @@ export async function runChunkedParseAndResolve(
|
||||
const cachedFiles = chunkFiles.length;
|
||||
onProgress({
|
||||
phase: 'parsing',
|
||||
percent: Math.round(20 + ((filesParsedSoFar + cachedFiles) / totalParseable) * 62),
|
||||
// Parse phase covers 20-70 (50 points). Deferred extraction below
|
||||
// takes 70-95 so the UI advances through the (potentially long)
|
||||
// resolution stages instead of holding at 82 (M2 from PR #1693
|
||||
// review).
|
||||
percent: Math.round(20 + ((filesParsedSoFar + cachedFiles) / totalParseable) * 50),
|
||||
message: `Parsing chunk ${chunkIdx + 1}/${numChunks} (cache)...`,
|
||||
stats: {
|
||||
filesProcessed: filesParsedSoFar + cachedFiles,
|
||||
@@ -378,7 +482,8 @@ export async function runChunkedParseAndResolve(
|
||||
scopeTreeCache,
|
||||
(current, _total, filePath) => {
|
||||
const globalCurrent = filesParsedSoFar + current;
|
||||
const parsingProgress = 20 + (globalCurrent / totalParseable) * 62;
|
||||
// Parse phase covers 20-70 (M2). Deferred extraction handles 70-95.
|
||||
const parsingProgress = 20 + (globalCurrent / totalParseable) * 50;
|
||||
onProgress({
|
||||
phase: 'parsing',
|
||||
percent: Math.round(parsingProgress),
|
||||
@@ -399,56 +504,63 @@ export async function runChunkedParseAndResolve(
|
||||
// Persist the raw results for this chunk hash. Sequential path
|
||||
// doesn't populate rawResults (it writes directly to graph), so
|
||||
// small repos without worker pool simply don't cache. That's fine.
|
||||
//
|
||||
// U20.U2: refuse the write when any chunk file is in the
|
||||
// worker pool's cumulative quarantine snapshot. The chunkHash
|
||||
// is computed from EVERY file in the chunk, but the pool's
|
||||
// Layer 3 quarantine filters quarantined files out of dispatch
|
||||
// — so `rawResults` is narrower than the chunkHash key implies.
|
||||
// Caching it would silently replay incomplete results on the
|
||||
// next run with unchanged content (the corruption class Codex's
|
||||
// adversarial review of PR #1693 flagged).
|
||||
//
|
||||
// Skipping the write means the next analyze gets a cache miss
|
||||
// for this chunk and re-dispatches against a fresh worker pool
|
||||
// (quarantine is session-scoped — `createQuarantine` is called
|
||||
// per-pool at worker-pool.ts), giving the quarantined file
|
||||
// another chance. If quarantine fires again, U20.U1's
|
||||
// sequential gap-fill still produces a complete graph for this
|
||||
// run; the cache just stays empty for this chunk until a fully-
|
||||
// clean dispatch lands.
|
||||
if (parseCache && chunkHash && rawResults.length > 0) {
|
||||
parseCache.entries.set(chunkHash, rawResults);
|
||||
if (isDev) {
|
||||
logger.info(
|
||||
`📦 parse-cache MISS+store: chunk ${chunkIdx + 1}/${numChunks} (${chunkFiles.length} files, ${chunkHash.slice(0, 8)})`,
|
||||
);
|
||||
const quarantineSnapshot = workerPool?.getQuarantinedPaths?.() ?? [];
|
||||
const quarantineSet = new Set(quarantineSnapshot);
|
||||
const chunkHadQuarantine = chunkFiles.some((f) => quarantineSet.has(f.path));
|
||||
if (chunkHadQuarantine) {
|
||||
if (isDev) {
|
||||
const quarantinedInChunk = chunkFiles.filter((f) => quarantineSet.has(f.path)).length;
|
||||
logger.info(
|
||||
`📦 parse-cache SKIP: chunk ${chunkIdx + 1}/${numChunks} ` +
|
||||
`had ${quarantinedInChunk} worker-quarantined file(s); ` +
|
||||
`next run will rediscover (${chunkHash.slice(0, 8)})`,
|
||||
);
|
||||
}
|
||||
} else {
|
||||
parseCache.entries.set(chunkHash, rawResults);
|
||||
if (isDev) {
|
||||
logger.info(
|
||||
`📦 parse-cache MISS+store: chunk ${chunkIdx + 1}/${numChunks} (${chunkFiles.length} files, ${chunkHash.slice(0, 8)})`,
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
const chunkBasePercent = 20 + (filesParsedSoFar / totalParseable) * 62;
|
||||
|
||||
// Per-chunk extraction passes (processImportsFromExtracted,
|
||||
// processHeritageFromExtracted, processRoutesFromExtracted,
|
||||
// synthesizeWildcardImportBindings, seedCrossFileReceiverTypes)
|
||||
// moved out of the chunk loop into a single end-of-loop pass below.
|
||||
// Reason: per-chunk extraction blocked the chunk loop on
|
||||
// main-thread work between worker dispatches — workers sat idle
|
||||
// and total CPU utilization plateaued at 4-5% on multi-core boxes.
|
||||
// Deferring keeps workers busy chunk-after-chunk; resolution sees
|
||||
// strictly-more-information (full repo graph) so cross-chunk import
|
||||
// and heritage targets resolve at least as well as before.
|
||||
if (chunkWorkerData) {
|
||||
await processImportsFromExtracted(
|
||||
graph,
|
||||
allPathObjects,
|
||||
chunkWorkerData.imports,
|
||||
ctx,
|
||||
(current, total) => {
|
||||
onProgress({
|
||||
phase: 'parsing',
|
||||
percent: Math.round(chunkBasePercent),
|
||||
message: `Resolving imports (chunk ${chunkIdx + 1}/${numChunks})...`,
|
||||
detail: `${current}/${total} files`,
|
||||
stats: {
|
||||
filesProcessed: filesParsedSoFar,
|
||||
totalFiles: totalParseable,
|
||||
nodesCreated: graph.nodeCount,
|
||||
},
|
||||
});
|
||||
},
|
||||
repoPath,
|
||||
importCtx,
|
||||
);
|
||||
if (chunkNeedsSynthesis[chunkIdx]) {
|
||||
synthesizeWildcardImportBindings(graph, ctx);
|
||||
hasSynthesized = true;
|
||||
}
|
||||
if (exportedTypeMap.size > 0 && ctx.namedImportMap.size > 0) {
|
||||
const { enrichedCount } = seedCrossFileReceiverTypes(
|
||||
chunkWorkerData.calls,
|
||||
ctx.namedImportMap,
|
||||
exportedTypeMap,
|
||||
);
|
||||
if (isDev && enrichedCount > 0) {
|
||||
logger.info(
|
||||
`🔗 E1: Seeded ${enrichedCount} cross-file receiver types (chunk ${chunkIdx + 1})`,
|
||||
);
|
||||
}
|
||||
anyChunkNeedsWildcardSynth = true;
|
||||
}
|
||||
for (const item of chunkWorkerData.imports) deferredWorkerImports.push(item);
|
||||
for (const item of chunkWorkerData.calls) deferredWorkerCalls.push(item);
|
||||
for (const item of chunkWorkerData.heritage) deferredWorkerHeritage.push(item);
|
||||
for (const item of chunkWorkerData.constructorBindings)
|
||||
@@ -463,35 +575,6 @@ export async function runChunkedParseAndResolve(
|
||||
for (const item of chunkWorkerData.assignments) deferredAssignments.push(item);
|
||||
}
|
||||
|
||||
await Promise.all([
|
||||
processHeritageFromExtracted(graph, chunkWorkerData.heritage, ctx, (current, total) => {
|
||||
onProgress({
|
||||
phase: 'parsing',
|
||||
percent: Math.round(chunkBasePercent),
|
||||
message: `Resolving heritage (chunk ${chunkIdx + 1}/${numChunks})...`,
|
||||
detail: `${current}/${total} records`,
|
||||
stats: {
|
||||
filesProcessed: filesParsedSoFar,
|
||||
totalFiles: totalParseable,
|
||||
nodesCreated: graph.nodeCount,
|
||||
},
|
||||
});
|
||||
}),
|
||||
processRoutesFromExtracted(graph, chunkWorkerData.routes ?? [], ctx, (current, total) => {
|
||||
onProgress({
|
||||
phase: 'parsing',
|
||||
percent: Math.round(chunkBasePercent),
|
||||
message: `Resolving routes (chunk ${chunkIdx + 1}/${numChunks})...`,
|
||||
detail: `${current}/${total} routes`,
|
||||
stats: {
|
||||
filesProcessed: filesParsedSoFar,
|
||||
totalFiles: totalParseable,
|
||||
nodesCreated: graph.nodeCount,
|
||||
},
|
||||
});
|
||||
}),
|
||||
]);
|
||||
|
||||
if (chunkWorkerData.fileScopeBindings?.length) {
|
||||
for (const { filePath, bindings } of chunkWorkerData.fileScopeBindings) {
|
||||
if (typeof filePath !== 'string' || filePath.length === 0) continue;
|
||||
@@ -530,6 +613,24 @@ export async function runChunkedParseAndResolve(
|
||||
|
||||
filesParsedSoFar += chunkFiles.length;
|
||||
astCache.clear();
|
||||
|
||||
// Throughput observability (U3): emit a per-chunk metrics line
|
||||
// under verbose ingestion mode so operators can verify CPU
|
||||
// utilization moved + tune `--workers` / batch sizes without
|
||||
// guessing. Cheap snapshot — just reads pool closure state.
|
||||
if (verboseThroughputLog && chunkStartMs !== null) {
|
||||
const elapsedMs = Date.now() - chunkStartMs;
|
||||
const filesPerSec = elapsedMs > 0 ? (chunkFiles.length * 1000) / elapsedMs : 0;
|
||||
const stats = workerPool?.getStats?.();
|
||||
const poolFrag = stats
|
||||
? ` pool: ${stats.activeSlots}/${stats.size} active, ` +
|
||||
`${stats.quarantined} quarantined${stats.poolBroken ? ', BROKEN' : ''}`
|
||||
: ' (sequential)';
|
||||
logger.info(
|
||||
`📊 chunk ${chunkIdx + 1}/${numChunks}: ${chunkFiles.length} files in ${elapsedMs}ms ` +
|
||||
`(${filesPerSec.toFixed(1)} files/s)${poolFrag}`,
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
if (isDev && parseCache && (chunkCacheHits > 0 || chunkCacheMisses > 0)) {
|
||||
@@ -538,10 +639,129 @@ export async function runChunkedParseAndResolve(
|
||||
);
|
||||
}
|
||||
|
||||
// Deferred end-of-loop extraction (moved out of the per-chunk block):
|
||||
// 1. processImportsFromExtracted on all chunks' imports
|
||||
// 2. synthesizeWildcardImportBindings (if any chunk had wildcards)
|
||||
// 3. seedCrossFileReceiverTypes on deferred calls (depends on
|
||||
// namedImportMap populated by step 1)
|
||||
// 4. processHeritageFromExtracted on all chunks' heritage
|
||||
// 5. processRoutesFromExtracted on all chunks' routes
|
||||
// Same logic as the prior per-chunk passes, just batched — resolution
|
||||
// sees the full repo graph instead of just current-and-earlier chunks.
|
||||
// Deferred extraction band (M2 from PR #1693 review): the 4 stages below
|
||||
// each get their own 5-10 point slice of the 70-95 range so percent
|
||||
// advances monotonically through the (potentially long) resolution work
|
||||
// instead of holding flat at 82. Stages that are skipped (zero-length
|
||||
// input) leave their band as a no-op jump — the next stage still starts
|
||||
// at its own band, preserving monotonicity.
|
||||
// imports: 70 -> 75 (5)
|
||||
// heritage: 75 -> 80 (5)
|
||||
// routes: 80 -> 85 (5)
|
||||
// calls: 85 -> 95 (10)
|
||||
if (deferredWorkerImports.length > 0) {
|
||||
await processImportsFromExtracted(
|
||||
graph,
|
||||
allPathObjects,
|
||||
deferredWorkerImports,
|
||||
ctx,
|
||||
(current, total) => {
|
||||
const ratio = total > 0 ? current / total : 1;
|
||||
onProgress({
|
||||
phase: 'parsing',
|
||||
percent: 70 + Math.round(ratio * 5),
|
||||
message: 'Resolving imports (all chunks)...',
|
||||
detail: `${current}/${total} files`,
|
||||
stats: {
|
||||
filesProcessed: filesParsedSoFar,
|
||||
totalFiles: totalParseable,
|
||||
nodesCreated: graph.nodeCount,
|
||||
},
|
||||
});
|
||||
},
|
||||
repoPath,
|
||||
importCtx,
|
||||
);
|
||||
// U15 (lightweight M1): processImportsFromExtracted is the sole
|
||||
// consumer of `deferredWorkerImports`. Free the array now so the
|
||||
// GC can reclaim the per-file ExtractedImport records before the
|
||||
// heavier downstream stages run (heritage, routes, calls). Peak
|
||||
// accumulator memory drops from O(repo) to O(repo - imports) for
|
||||
// the remainder of the deferred phase. The future per-chunk
|
||||
// streaming upgrade can rewrite this with the same correctness
|
||||
// contract once profile data shows it's warranted.
|
||||
deferredWorkerImports.length = 0;
|
||||
}
|
||||
if (anyChunkNeedsWildcardSynth) {
|
||||
synthesizeWildcardImportBindings(graph, ctx);
|
||||
hasSynthesized = true;
|
||||
}
|
||||
// L5 from PR #1693 review: populate `exportedTypeMap` from the in-progress
|
||||
// graph BEFORE `seedCrossFileReceiverTypes` runs. Previously the seeding
|
||||
// branch below was reached with `exportedTypeMap.size === 0` in the
|
||||
// worker path (the map was only built at the post-parse block far below,
|
||||
// AFTER the seeding branch), so the seed dead-coded itself silently and
|
||||
// call resolution never got the cross-file receiver-type enrichment.
|
||||
// The post-parse builder still runs as a defensive fallback on the
|
||||
// sequential path; its `size === 0` guard means we don't pay the cost
|
||||
// twice on the worker path.
|
||||
if (exportedTypeMap.size === 0 && graph.nodeCount > 0) {
|
||||
const graphExports = buildExportedTypeMapFromGraph(graph, ctx.model.symbols);
|
||||
for (const [fp, exports] of graphExports) exportedTypeMap.set(fp, exports);
|
||||
}
|
||||
if (exportedTypeMap.size > 0 && ctx.namedImportMap.size > 0 && deferredWorkerCalls.length > 0) {
|
||||
const { enrichedCount } = seedCrossFileReceiverTypes(
|
||||
deferredWorkerCalls,
|
||||
ctx.namedImportMap,
|
||||
exportedTypeMap,
|
||||
);
|
||||
if (isDev && enrichedCount > 0) {
|
||||
logger.info(`🔗 E1: Seeded ${enrichedCount} cross-file receiver types (all chunks)`);
|
||||
}
|
||||
}
|
||||
if (deferredWorkerHeritage.length > 0) {
|
||||
await processHeritageFromExtracted(graph, deferredWorkerHeritage, ctx, (current, total) => {
|
||||
const ratio = total > 0 ? current / total : 1;
|
||||
onProgress({
|
||||
phase: 'parsing',
|
||||
percent: 75 + Math.round(ratio * 5),
|
||||
message: 'Resolving heritage (all chunks)...',
|
||||
detail: `${current}/${total} records`,
|
||||
stats: {
|
||||
filesProcessed: filesParsedSoFar,
|
||||
totalFiles: totalParseable,
|
||||
nodesCreated: graph.nodeCount,
|
||||
},
|
||||
});
|
||||
});
|
||||
}
|
||||
if (allExtractedRoutes.length > 0) {
|
||||
await processRoutesFromExtracted(graph, allExtractedRoutes, ctx, (current, total) => {
|
||||
const ratio = total > 0 ? current / total : 1;
|
||||
onProgress({
|
||||
phase: 'parsing',
|
||||
percent: 80 + Math.round(ratio * 5),
|
||||
message: 'Resolving routes (all chunks)...',
|
||||
detail: `${current}/${total} routes`,
|
||||
stats: {
|
||||
filesProcessed: filesParsedSoFar,
|
||||
totalFiles: totalParseable,
|
||||
nodesCreated: graph.nodeCount,
|
||||
},
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
const fullWorkerHeritageMap =
|
||||
deferredWorkerHeritage.length > 0
|
||||
? buildHeritageMap(deferredWorkerHeritage, ctx, getHeritageStrategyForLanguage)
|
||||
: undefined;
|
||||
// U15 (lightweight M1): buildHeritageMap is the LAST consumer of the
|
||||
// raw `deferredWorkerHeritage` records — processCallsFromExtracted
|
||||
// below reads from the derived `fullWorkerHeritageMap` instead. Free
|
||||
// the raw heritage array now so the GC can reclaim it before the
|
||||
// (potentially long) call-resolution stage. processHeritageFromExtracted
|
||||
// earlier was a read-only consumer (pushed to graph, didn't drain).
|
||||
deferredWorkerHeritage.length = 0;
|
||||
|
||||
if (deferredWorkerCalls.length > 0) {
|
||||
await processCallsFromExtracted(
|
||||
@@ -549,9 +769,13 @@ export async function runChunkedParseAndResolve(
|
||||
deferredWorkerCalls,
|
||||
ctx,
|
||||
(current, total) => {
|
||||
const ratio = total > 0 ? current / total : 1;
|
||||
onProgress({
|
||||
phase: 'parsing',
|
||||
percent: 82,
|
||||
// Calls is the longest deferred stage on real repos — give it the
|
||||
// 10-point tail 85-95 so the progress bar visibly advances during
|
||||
// call resolution instead of holding at 82 (M2).
|
||||
percent: 85 + Math.round(ratio * 10),
|
||||
message: 'Resolving calls (all chunks)...',
|
||||
detail: `${current}/${total} files`,
|
||||
stats: {
|
||||
@@ -576,6 +800,20 @@ export async function runChunkedParseAndResolve(
|
||||
bindingAccumulator,
|
||||
);
|
||||
}
|
||||
// U15 (lightweight M1): all three arrays have had their last consumer
|
||||
// by the time we reach this point — processCallsFromExtracted drained
|
||||
// `deferredWorkerCalls` and read `deferredConstructorBindings`;
|
||||
// processAssignmentsFromExtracted drained `deferredAssignments` and
|
||||
// also read `deferredConstructorBindings`. Free them now so the
|
||||
// function-scope references die before downstream graph-build /
|
||||
// scope-resolution starts using its own working memory. Note: arrays
|
||||
// returned in the function result object (allFetchCalls,
|
||||
// allExtractedRoutes, allDecoratorRoutes, allToolDefs, allORMQueries,
|
||||
// allParsedFiles) intentionally stay live — downstream consumers
|
||||
// need them.
|
||||
deferredWorkerCalls.length = 0;
|
||||
deferredConstructorBindings.length = 0;
|
||||
deferredAssignments.length = 0;
|
||||
} finally {
|
||||
await workerPool?.terminate();
|
||||
}
|
||||
|
||||
@@ -55,6 +55,16 @@ export interface PipelineOptions {
|
||||
minFiles?: number;
|
||||
minBytes?: number;
|
||||
};
|
||||
/**
|
||||
* @internal Test-only override for the worker script URL the pool
|
||||
* spawns. When unset, parse-impl resolves `parse-worker.js` from the
|
||||
* adjacent `workers/` directory (or the compiled `dist/` fallback
|
||||
* under vitest). Integration tests use this to inject a custom
|
||||
* worker script that deterministically triggers worker-pool
|
||||
* resilience paths (e.g., crash-on-poison-file) — same precedent as
|
||||
* `workerThresholdsForTest`. Do not use from production call sites.
|
||||
*/
|
||||
workerUrlForTest?: URL;
|
||||
/**
|
||||
* Incremental-indexing parse cache. When provided:
|
||||
* - The parse phase looks up each chunk's content hash in
|
||||
@@ -68,6 +78,46 @@ export interface PipelineOptions {
|
||||
* See `gitnexus/src/storage/parse-cache.ts`.
|
||||
*/
|
||||
parseCache?: import('../../storage/parse-cache.js').ParseCache;
|
||||
/**
|
||||
* Worker pool size override, threaded from the CLI `--workers` flag
|
||||
* via `AnalyzeOptions`. When set, parse-impl passes this directly to
|
||||
* `createWorkerPool` so the pool sizing bypasses the env-var fallback
|
||||
* in `resolveAutoPoolSize`. The env-var channel
|
||||
* (`GITNEXUS_WORKER_POOL_SIZE`) remains as a back-compat fallback when
|
||||
* this field is undefined. Setting `workerPoolSize: 0` disables the
|
||||
* pool entirely (sequential fallback) — equivalent to `skipWorkers`
|
||||
* but expressed in the same units as `--workers <N>` so long-running
|
||||
* hosts (eval-server, MCP daemon) can size per-call without leaking
|
||||
* `process.env` state across analyze invocations.
|
||||
*/
|
||||
workerPoolSize?: number;
|
||||
/**
|
||||
* Number of chunks whose file contents may be read into memory in
|
||||
* parallel while the worker pool is busy dispatching the current
|
||||
* chunk. Pre-fetching overlaps disk I/O for chunk N+1..N+K with the
|
||||
* worker compute on chunk N — modest but real wall-clock win on
|
||||
* repos large enough to chunk. Worker dispatch itself remains serial
|
||||
* because `WorkerPool.dispatch` is not reentrant (concurrent calls
|
||||
* would race on the shared per-slot busy/in-flight state).
|
||||
*
|
||||
* `1` matches today's pure-serial behavior; `2` is the documented
|
||||
* default (`GITNEXUS_PARSE_CHUNK_CONCURRENCY`). Falls back to the
|
||||
* env var when undefined; defaults to 2 when neither is set.
|
||||
*/
|
||||
parseChunkConcurrency?: number;
|
||||
/**
|
||||
* Byte budget per parse chunk (in bytes). When set, parse-impl uses
|
||||
* this instead of the `GITNEXUS_CHUNK_BYTE_BUDGET` env var or the
|
||||
* built-in 2 MB default. Smaller values produce more chunks (finer
|
||||
* cache-hit granularity, more worker dispatches); larger values
|
||||
* batch more files per dispatch.
|
||||
*
|
||||
* Threading the value through options instead of the env var lets
|
||||
* tests vary the chunk layout per-call without `vi.resetModules` and
|
||||
* lets long-running hosts (eval-server, MCP daemon) size per-call
|
||||
* without leaking `process.env` state across invocations.
|
||||
*/
|
||||
chunkByteBudget?: number;
|
||||
}
|
||||
|
||||
// ── Phase registry ─────────────────────────────────────────────────────────
|
||||
|
||||
@@ -74,6 +74,7 @@ export const MIGRATED_LANGUAGES: ReadonlySet<SupportedLanguages> = new Set<Suppo
|
||||
SupportedLanguages.C,
|
||||
SupportedLanguages.CPlusPlus,
|
||||
SupportedLanguages.PHP,
|
||||
SupportedLanguages.JavaScript,
|
||||
]);
|
||||
|
||||
/**
|
||||
|
||||
@@ -64,6 +64,8 @@ export interface ResolveReferencesInput {
|
||||
readonly scopes: ScopeResolutionIndexes;
|
||||
/** Provider hooks consumed by the registries (e.g. `arityCompatibility`). */
|
||||
readonly providers?: RegistryProviders;
|
||||
/** Required owner-keyed member lookup used by Step 2 receiver/MRO walks. */
|
||||
readonly ownedMembersByOwner: RegistryContext['ownedMembersByOwner'];
|
||||
}
|
||||
|
||||
export interface ResolveStats {
|
||||
@@ -92,6 +94,7 @@ export function resolveReferenceSites(input: ResolveReferencesInput): ResolveRef
|
||||
defs: scopes.defs,
|
||||
qualifiedNames: scopes.qualifiedNames,
|
||||
moduleScopes: scopes.moduleScopes,
|
||||
ownedMembersByOwner: input.ownedMembersByOwner,
|
||||
methodDispatch: scopes.methodDispatch,
|
||||
providers,
|
||||
};
|
||||
@@ -191,7 +194,10 @@ function lookupForSite(
|
||||
case 'write': {
|
||||
// Try field first; fall through to method then class so bare-name
|
||||
// reads of a function (e.g. `cb = save`) still resolve.
|
||||
const fieldHits = fieldRegistry.lookup(site.name, site.inScope);
|
||||
const fieldOpts: Parameters<FieldRegistry['lookup']>[2] = {
|
||||
...(site.explicitReceiver !== undefined ? { explicitReceiver: site.explicitReceiver } : {}),
|
||||
};
|
||||
const fieldHits = fieldRegistry.lookup(site.name, site.inScope, fieldOpts);
|
||||
if (fieldHits.length > 0) return fieldHits;
|
||||
const methodHits = methodRegistry.lookup(site.name, site.inScope);
|
||||
if (methodHits.length > 0) return methodHits;
|
||||
|
||||
@@ -913,6 +913,9 @@ function pass5CollectReferences(
|
||||
const explicitReceiver = extractExplicitReceiver(match);
|
||||
const arity = extractArity(match);
|
||||
const argumentTypes = extractArgumentTypes(match);
|
||||
const argumentTypeClasses = parseJsonParameterTypeClassesCapture(
|
||||
match['@reference.parameter-type-classes'],
|
||||
);
|
||||
|
||||
const site: ReferenceSite = {
|
||||
name: nameCap.text,
|
||||
@@ -923,6 +926,7 @@ function pass5CollectReferences(
|
||||
...(explicitReceiver !== undefined ? { explicitReceiver } : {}),
|
||||
...(arity !== undefined ? { arity } : {}),
|
||||
...(argumentTypes !== undefined ? { argumentTypes } : {}),
|
||||
...(argumentTypeClasses !== undefined ? { argumentTypeClasses } : {}),
|
||||
};
|
||||
referenceSites.push(site);
|
||||
}
|
||||
@@ -1040,9 +1044,11 @@ const KNOWN_SUB_TAGS: ReadonlySet<string> = new Set<string>([
|
||||
'@reference.receiver',
|
||||
'@reference.arity',
|
||||
'@reference.parameter-types',
|
||||
'@reference.parameter-type-classes',
|
||||
'@declaration.parameter-count',
|
||||
'@declaration.required-parameter-count',
|
||||
'@declaration.parameter-types',
|
||||
'@declaration.parameter-type-classes',
|
||||
'@declaration.template-constraints',
|
||||
]);
|
||||
|
||||
|
||||
@@ -17,7 +17,13 @@
|
||||
* generalization plan.
|
||||
*/
|
||||
|
||||
import type { ParsedFile, Reference, ScopeId, SymbolDefinition } from 'gitnexus-shared';
|
||||
import type {
|
||||
ParameterTypeClass,
|
||||
ParsedFile,
|
||||
Reference,
|
||||
ScopeId,
|
||||
SymbolDefinition,
|
||||
} from 'gitnexus-shared';
|
||||
import type { KnowledgeGraph } from '../../../graph/types.js';
|
||||
import type { ScopeResolutionIndexes } from '../../model/scope-resolution-indexes.js';
|
||||
import type { SemanticModel } from '../../model/semantic-model.js';
|
||||
@@ -78,6 +84,12 @@ export function emitFreeCallFallback(
|
||||
let emitted = 0;
|
||||
const seen = new Set<string>();
|
||||
|
||||
// Build an O(1) simple-name -> callable defs index over scopes.defs once
|
||||
// per pass so pickUniqueGlobalCallable doesn't re-scan defs.byId.values()
|
||||
// per call site. Same name + callable-kind filter that the previous scan
|
||||
// applied (see pickUniqueGlobalCallable JSDoc). Cost: O(|defs|) once.
|
||||
const globalCallablesBySimpleName = buildGlobalCallableIndex(scopes);
|
||||
|
||||
for (const parsed of parsedFiles) {
|
||||
for (const site of parsed.referenceSites) {
|
||||
if (site.kind !== 'call') continue;
|
||||
@@ -126,6 +138,7 @@ export function emitFreeCallFallback(
|
||||
site.arity,
|
||||
site.argumentTypes,
|
||||
{
|
||||
argumentTypeClasses: site.argumentTypeClasses,
|
||||
conversionRankFn: options.conversionRankFn,
|
||||
constraintCompatibility: options.constraintCompatibility,
|
||||
},
|
||||
@@ -190,6 +203,7 @@ export function emitFreeCallFallback(
|
||||
fnDef = ordinary[0];
|
||||
} else {
|
||||
const narrowed = narrowOverloadCandidates(ordinary, site.arity, site.argumentTypes, {
|
||||
argumentTypeClasses: site.argumentTypeClasses,
|
||||
conversionRankFn: options.conversionRankFn,
|
||||
constraintCompatibility: options.constraintCompatibility,
|
||||
});
|
||||
@@ -225,6 +239,7 @@ export function emitFreeCallFallback(
|
||||
push(adl);
|
||||
|
||||
const narrowed = narrowOverloadCandidates(merged, site.arity, site.argumentTypes, {
|
||||
argumentTypeClasses: site.argumentTypeClasses,
|
||||
conversionRankFn: options.conversionRankFn,
|
||||
constraintCompatibility: options.constraintCompatibility,
|
||||
});
|
||||
@@ -254,7 +269,7 @@ export function emitFreeCallFallback(
|
||||
fnDef = pickUniqueGlobalCallable(
|
||||
site.name,
|
||||
model,
|
||||
scopes,
|
||||
globalCallablesBySimpleName,
|
||||
parsed.filePath,
|
||||
options.isFileLocalDef,
|
||||
site.arity,
|
||||
@@ -268,6 +283,7 @@ export function emitFreeCallFallback(
|
||||
})
|
||||
: undefined,
|
||||
site.argumentTypes,
|
||||
site.argumentTypeClasses,
|
||||
options.conversionRankFn,
|
||||
);
|
||||
}
|
||||
@@ -299,23 +315,46 @@ export function emitFreeCallFallback(
|
||||
return emitted;
|
||||
}
|
||||
|
||||
/**
|
||||
* Build a `simpleName -> callable defs` index from `scopes.defs` once per
|
||||
* pass. Mirrors the filter the old per-site scan applied: Function /
|
||||
* Method / Constructor, keyed by the last `.`-segment of `qualifiedName`
|
||||
* (falling back to the qualifiedName itself when undotted). Used by
|
||||
* `pickUniqueGlobalCallable` so every free-call fallback site is O(1)
|
||||
* instead of O(|defs|).
|
||||
*/
|
||||
function buildGlobalCallableIndex(
|
||||
scopes: ScopeResolutionIndexes,
|
||||
): ReadonlyMap<string, readonly SymbolDefinition[]> {
|
||||
const out = new Map<string, SymbolDefinition[]>();
|
||||
for (const def of scopes.defs.byId.values()) {
|
||||
if (def.type !== 'Function' && def.type !== 'Method' && def.type !== 'Constructor') continue;
|
||||
const qualified = def.qualifiedName;
|
||||
if (qualified === undefined || qualified.length === 0) continue;
|
||||
const dot = qualified.lastIndexOf('.');
|
||||
const simple = dot === -1 ? qualified : qualified.slice(dot + 1);
|
||||
const bucket = out.get(simple);
|
||||
if (bucket) bucket.push(def);
|
||||
else out.set(simple, [def]);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
function pickUniqueGlobalCallable(
|
||||
name: string,
|
||||
model: SemanticModel,
|
||||
scopes: ScopeResolutionIndexes,
|
||||
globalCallablesBySimpleName: ReadonlyMap<string, readonly SymbolDefinition[]>,
|
||||
callerFilePath: string,
|
||||
isFileLocalDef?: (def: SymbolDefinition) => boolean,
|
||||
callArity?: number,
|
||||
isCallerVisible?: (candidate: SymbolDefinition) => boolean,
|
||||
callArgTypes?: readonly string[],
|
||||
callArgTypeClasses?: readonly ParameterTypeClass[],
|
||||
conversionRankFn?: ConversionRankFn,
|
||||
): SymbolDefinition | undefined {
|
||||
const scopeDefs: SymbolDefinition[] = [];
|
||||
const scopeSeen = new Set<string>();
|
||||
for (const def of scopes.defs.byId.values()) {
|
||||
const simple = def.qualifiedName?.split('.').pop() ?? def.qualifiedName;
|
||||
if (simple !== name) continue;
|
||||
if (def.type !== 'Function' && def.type !== 'Method' && def.type !== 'Constructor') continue;
|
||||
for (const def of globalCallablesBySimpleName.get(name) ?? []) {
|
||||
// Skip file-local defs (e.g. C `static` functions) that live in a
|
||||
// different file from the caller — they are logically invisible.
|
||||
if (isFileLocalDef !== undefined && def.filePath !== callerFilePath && isFileLocalDef(def)) {
|
||||
@@ -349,6 +388,7 @@ function pickUniqueGlobalCallable(
|
||||
// disambiguate (e.g., `f(int)` vs `f(double)` called with `f(2.5)`).
|
||||
if (scopeDefs.length > 1) {
|
||||
const narrowed = narrowOverloadCandidates(scopeDefs, callArity, callArgTypes, {
|
||||
argumentTypeClasses: callArgTypeClasses,
|
||||
conversionRankFn,
|
||||
});
|
||||
if (narrowed.length === 1) return narrowed[0];
|
||||
@@ -389,6 +429,7 @@ function pickUniqueGlobalCallable(
|
||||
// Same argument-type + conversion-rank narrowing for the model pool.
|
||||
if (defs.length > 1) {
|
||||
const narrowed = narrowOverloadCandidates(defs, callArity, callArgTypes, {
|
||||
argumentTypeClasses: callArgTypeClasses,
|
||||
conversionRankFn,
|
||||
});
|
||||
if (narrowed.length === 1) return narrowed[0];
|
||||
@@ -462,6 +503,7 @@ export function pickImplicitThisOverload(
|
||||
readonly name: string;
|
||||
readonly arity?: number;
|
||||
readonly argumentTypes?: readonly string[];
|
||||
readonly argumentTypeClasses?: readonly import('gitnexus-shared').ParameterTypeClass[];
|
||||
},
|
||||
scopes: ScopeResolutionIndexes,
|
||||
workspaceIndex: WorkspaceResolutionIndex,
|
||||
@@ -498,6 +540,7 @@ export function pickImplicitThisOverload(
|
||||
// disambiguating signal) leaves the call unresolved rather than
|
||||
// routing to an arbitrary first overload by registration order.
|
||||
const candidates = narrowOverloadCandidates(overloads, site.arity, site.argumentTypes, {
|
||||
argumentTypeClasses: site.argumentTypeClasses,
|
||||
conversionRankFn: hookCtx?.conversionRankFn,
|
||||
constraintCompatibility: hookCtx?.constraintCompatibility,
|
||||
});
|
||||
|
||||
@@ -38,7 +38,13 @@
|
||||
* 5. Empty input returns empty output.
|
||||
*/
|
||||
|
||||
import type { ArityVerdict, Callsite, ConstraintContext, SymbolDefinition } from 'gitnexus-shared';
|
||||
import type {
|
||||
ArityVerdict,
|
||||
Callsite,
|
||||
ConstraintContext,
|
||||
ParameterTypeClass,
|
||||
SymbolDefinition,
|
||||
} from 'gitnexus-shared';
|
||||
|
||||
/**
|
||||
* Per-slot conversion-rank function. Returns a numeric cost for
|
||||
@@ -51,7 +57,12 @@ import type { ArityVerdict, Callsite, ConstraintContext, SymbolDefinition } from
|
||||
* Each language provides its own implementation. The function operates
|
||||
* on normalized type strings (output of the language's type normalizer).
|
||||
*/
|
||||
export type ConversionRankFn = (argType: string, paramType: string) => number;
|
||||
export type ConversionRankFn = (
|
||||
argType: string,
|
||||
paramType: string,
|
||||
argTypeClass?: ParameterTypeClass,
|
||||
paramTypeClass?: ParameterTypeClass,
|
||||
) => number;
|
||||
|
||||
/**
|
||||
* Optional hook bundle for narrowing extension points. Threaded in
|
||||
@@ -62,6 +73,8 @@ export type ConversionRankFn = (argType: string, paramType: string) => number;
|
||||
* undefined preserves the legacy arity + exact-type behavior.
|
||||
*/
|
||||
export interface OverloadNarrowingHookCtx {
|
||||
/** Shape-preserving per-argument sidecar aligned with `argTypes`. */
|
||||
readonly argumentTypeClasses?: ConstraintContext['argumentTypeClasses'];
|
||||
/** Conversion-rank scoring fallback (step 4b). Engages when the
|
||||
* exact-type filter rejects every candidate. */
|
||||
readonly conversionRankFn?: ConversionRankFn;
|
||||
@@ -128,7 +141,16 @@ export function narrowOverloadCandidates(
|
||||
if (params === undefined) return false;
|
||||
for (let i = 0; i < argTypes.length && i < params.length; i++) {
|
||||
if (argTypes[i] === '') continue;
|
||||
if (argTypes[i] !== params[i]) return false;
|
||||
if (
|
||||
!exactTypeSlotMatches(
|
||||
argTypes[i],
|
||||
params[i],
|
||||
hookCtx?.argumentTypeClasses?.[i],
|
||||
d.parameterTypeClasses?.[i],
|
||||
)
|
||||
) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
return true;
|
||||
});
|
||||
@@ -142,7 +164,12 @@ export function narrowOverloadCandidates(
|
||||
// are returned; multiple survivors are genuinely ambiguous. When
|
||||
// ranking also yields empty, fall through to the arity-filtered
|
||||
// `candidates` set — matches pre-#1606 behavior.
|
||||
const ranked = rankByConversion(candidates, argTypes, hookCtx.conversionRankFn);
|
||||
const ranked = rankByConversion(
|
||||
candidates,
|
||||
argTypes,
|
||||
hookCtx.conversionRankFn,
|
||||
hookCtx.argumentTypeClasses,
|
||||
);
|
||||
if (ranked.length > 0) result = ranked;
|
||||
}
|
||||
}
|
||||
@@ -163,7 +190,15 @@ export function narrowOverloadCandidates(
|
||||
// than emitting a wrong edge.
|
||||
if (hookCtx?.constraintCompatibility !== undefined && argCount !== undefined) {
|
||||
const callsite: Callsite = { arity: argCount };
|
||||
const ctx: ConstraintContext = argTypes !== undefined ? { argumentTypes: argTypes } : {};
|
||||
const ctx: ConstraintContext =
|
||||
argTypes !== undefined
|
||||
? {
|
||||
argumentTypes: argTypes,
|
||||
...(hookCtx.argumentTypeClasses !== undefined
|
||||
? { argumentTypeClasses: hookCtx.argumentTypeClasses }
|
||||
: {}),
|
||||
}
|
||||
: {};
|
||||
result = result.filter((def) => {
|
||||
if (def.templateConstraints === undefined) return true;
|
||||
return hookCtx.constraintCompatibility!(callsite, def, ctx) !== 'incompatible';
|
||||
@@ -173,6 +208,27 @@ export function narrowOverloadCandidates(
|
||||
return result;
|
||||
}
|
||||
|
||||
function exactTypeSlotMatches(
|
||||
argType: string,
|
||||
paramType: string,
|
||||
argTypeClass?: ParameterTypeClass,
|
||||
paramTypeClass?: ParameterTypeClass,
|
||||
): boolean {
|
||||
if (argType !== paramType) return false;
|
||||
// C++ normalizes away pointer markers (`int*` -> `int`). When both sides
|
||||
// provide shape sidecars, do not let that collapse make `int` exactly match
|
||||
// `int*`. Unknown sidecar evidence preserves the previous string-only path.
|
||||
if (argTypeClass === undefined || paramTypeClass === undefined) return true;
|
||||
if (argTypeClass.indirection === 'unknown' || paramTypeClass.indirection === 'unknown') {
|
||||
return true;
|
||||
}
|
||||
return isPointerShape(argTypeClass) === isPointerShape(paramTypeClass);
|
||||
}
|
||||
|
||||
function isPointerShape(typeClass: ParameterTypeClass): boolean {
|
||||
return typeClass.indirection === 'pointer' && typeClass.pointerDepth > 0;
|
||||
}
|
||||
|
||||
/**
|
||||
* Pairwise dominance comparison (ISO C++ [over.ics.rank]).
|
||||
*
|
||||
@@ -189,6 +245,7 @@ function rankByConversion(
|
||||
candidates: readonly SymbolDefinition[],
|
||||
argTypes: readonly string[],
|
||||
rankFn: ConversionRankFn,
|
||||
argTypeClasses?: readonly ParameterTypeClass[],
|
||||
): readonly SymbolDefinition[] {
|
||||
// Step 1: compute per-slot ranks and exclude non-viable candidates.
|
||||
const viable: Array<{ def: SymbolDefinition; ranks: number[] }> = [];
|
||||
@@ -197,12 +254,22 @@ function rankByConversion(
|
||||
if (params === undefined) continue;
|
||||
const ranks: number[] = [];
|
||||
let ok = true;
|
||||
for (let i = 0; i < argTypes.length && i < params.length; i++) {
|
||||
for (let i = 0; i < argTypes.length; i++) {
|
||||
const paramType = parameterTypeAt(params, i);
|
||||
if (paramType === undefined) {
|
||||
ok = false;
|
||||
break;
|
||||
}
|
||||
if (argTypes[i] === '') {
|
||||
ranks.push(0); // unknown arg → any-match (rank 0)
|
||||
continue;
|
||||
}
|
||||
const r = rankFn(argTypes[i], params[i]);
|
||||
const r = rankFn(
|
||||
argTypes[i],
|
||||
paramType,
|
||||
argTypeClasses?.[i],
|
||||
parameterTypeClassAt(d.parameterTypeClasses, i),
|
||||
);
|
||||
if (!isFinite(r)) {
|
||||
ok = false;
|
||||
break;
|
||||
@@ -229,6 +296,20 @@ function rankByConversion(
|
||||
return viable.filter((_, idx) => !dominated.has(idx)).map((v) => v.def);
|
||||
}
|
||||
|
||||
function parameterTypeAt(params: readonly string[], argIndex: number): string | undefined {
|
||||
if (argIndex < params.length) return params[argIndex];
|
||||
return params[params.length - 1] === '...' ? '...' : undefined;
|
||||
}
|
||||
|
||||
function parameterTypeClassAt(
|
||||
params: readonly ParameterTypeClass[] | undefined,
|
||||
argIndex: number,
|
||||
): ParameterTypeClass | undefined {
|
||||
if (params === undefined) return undefined;
|
||||
if (argIndex < params.length) return params[argIndex];
|
||||
return params[params.length - 1]?.base === '...' ? params[params.length - 1] : undefined;
|
||||
}
|
||||
|
||||
/**
|
||||
* Compare two per-slot rank vectors.
|
||||
* Returns -1 if `a` dominates `b` (not worse everywhere, better somewhere),
|
||||
|
||||
@@ -346,6 +346,7 @@ export function emitReceiverBoundCalls(
|
||||
site.arity,
|
||||
site.argumentTypes,
|
||||
{
|
||||
argumentTypeClasses: site.argumentTypeClasses,
|
||||
conversionRankFn: provider.conversionRankFn,
|
||||
constraintCompatibility: provider.constraintCompatibility,
|
||||
},
|
||||
@@ -732,6 +733,7 @@ function pickOverload(
|
||||
if (overloads.length === 1) return overloads[0];
|
||||
|
||||
const candidates = narrowOverloadCandidates(overloads, site.arity, site.argumentTypes, {
|
||||
argumentTypeClasses: site.argumentTypeClasses,
|
||||
conversionRankFn: provider.conversionRankFn,
|
||||
constraintCompatibility: provider.constraintCompatibility,
|
||||
});
|
||||
|
||||
@@ -18,8 +18,8 @@
|
||||
* undefined `ownerId` is reachable via either:
|
||||
* - `model.methods.lookupAllByOwner(ownerId, simpleName)` — if the
|
||||
* def is a Method / Function / Constructor, OR
|
||||
* - `model.fields.lookupFieldByOwner(ownerId, simpleName)` — if the
|
||||
* def is a Property / Variable.
|
||||
* - `model.fields.lookupAllByOwner(ownerId, simpleName)` — if the
|
||||
* def is a Property / Variable / Const / Static.
|
||||
*
|
||||
* This invariant is the foundation of Contract Invariant I9
|
||||
* (`contract/scope-resolver.ts`): scope-resolution passes MUST read
|
||||
@@ -45,11 +45,29 @@ import type { ParsedFile } from 'gitnexus-shared';
|
||||
import type { MutableSemanticModel, SemanticModel } from '../../model/semantic-model.js';
|
||||
import { simpleQualifiedName } from '../graph-bridge/ids.js';
|
||||
|
||||
const NESTED_TYPE_KINDS = new Set<string>([
|
||||
'Class',
|
||||
'Interface',
|
||||
'Enum',
|
||||
'Struct',
|
||||
'Union',
|
||||
'Trait',
|
||||
'TypeAlias',
|
||||
'Typedef',
|
||||
'Record',
|
||||
'Delegate',
|
||||
'Annotation',
|
||||
'Template',
|
||||
'Namespace',
|
||||
]);
|
||||
|
||||
export interface ReconcileStats {
|
||||
/** Method/Function/Constructor defs registered into MethodRegistry. */
|
||||
readonly methodsRegistered: number;
|
||||
/** Property/Variable defs registered into FieldRegistry. */
|
||||
readonly fieldsRegistered: number;
|
||||
/** Class-like nested type defs registered into TypeRegistry by owner. */
|
||||
readonly nestedTypesRegistered: number;
|
||||
/** Defs already present (idempotent skip). */
|
||||
readonly skippedAlreadyPresent: number;
|
||||
}
|
||||
@@ -60,6 +78,7 @@ export function reconcileOwnership(
|
||||
): ReconcileStats {
|
||||
let methodsRegistered = 0;
|
||||
let fieldsRegistered = 0;
|
||||
let nestedTypesRegistered = 0;
|
||||
let skippedAlreadyPresent = 0;
|
||||
|
||||
for (const parsed of parsedFiles) {
|
||||
@@ -77,19 +96,32 @@ export function reconcileOwnership(
|
||||
}
|
||||
model.methods.register(ownerId, simple, def);
|
||||
methodsRegistered++;
|
||||
} else if (def.type === 'Property' || def.type === 'Variable') {
|
||||
const existing = model.fields.lookupFieldByOwner(ownerId, simple);
|
||||
if (existing !== undefined && existing.nodeId === def.nodeId) {
|
||||
} else if (
|
||||
def.type === 'Property' ||
|
||||
def.type === 'Variable' ||
|
||||
def.type === 'Const' ||
|
||||
def.type === 'Static'
|
||||
) {
|
||||
const existing = model.fields.lookupAllByOwner(ownerId, simple);
|
||||
if (existing.some((e) => e.nodeId === def.nodeId)) {
|
||||
skippedAlreadyPresent++;
|
||||
continue;
|
||||
}
|
||||
model.fields.register(ownerId, simple, def);
|
||||
fieldsRegistered++;
|
||||
} else if (NESTED_TYPE_KINDS.has(def.type)) {
|
||||
const existing = model.types.lookupAllByOwner(ownerId, simple);
|
||||
if (existing.some((e) => e.nodeId === def.nodeId)) {
|
||||
skippedAlreadyPresent++;
|
||||
continue;
|
||||
}
|
||||
model.types.registerByOwner(ownerId, simple, def);
|
||||
nestedTypesRegistered++;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return { methodsRegistered, fieldsRegistered, skippedAlreadyPresent };
|
||||
return { methodsRegistered, fieldsRegistered, nestedTypesRegistered, skippedAlreadyPresent };
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -131,15 +163,29 @@ export function validateOwnershipParity(
|
||||
);
|
||||
mismatches++;
|
||||
}
|
||||
} else if (def.type === 'Property' || def.type === 'Variable') {
|
||||
const found = model.fields.lookupFieldByOwner(ownerId, simple);
|
||||
if (found === undefined || found.nodeId !== def.nodeId) {
|
||||
} else if (
|
||||
def.type === 'Property' ||
|
||||
def.type === 'Variable' ||
|
||||
def.type === 'Const' ||
|
||||
def.type === 'Static'
|
||||
) {
|
||||
const found = model.fields.lookupAllByOwner(ownerId, simple);
|
||||
if (!found.some((d) => d.nodeId === def.nodeId)) {
|
||||
onWarn(
|
||||
`semantic-model parity: ${def.type} ${def.nodeId} (${parsed.filePath}) ` +
|
||||
`owned by ${ownerId} as "${simple}" not in FieldRegistry`,
|
||||
);
|
||||
mismatches++;
|
||||
}
|
||||
} else if (NESTED_TYPE_KINDS.has(def.type)) {
|
||||
const found = model.types.lookupAllByOwner(ownerId, simple);
|
||||
if (!found.some((d) => d.nodeId === def.nodeId)) {
|
||||
onWarn(
|
||||
`semantic-model parity: ${def.type} ${def.nodeId} (${parsed.filePath}) ` +
|
||||
`owned by ${ownerId} as "${simple}" not in TypeRegistry owner index`,
|
||||
);
|
||||
mismatches++;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -19,6 +19,8 @@ import { javaScopeResolver } from '../../languages/java/scope-resolver.js';
|
||||
import { cScopeResolver } from '../../languages/c/scope-resolver.js';
|
||||
import { cppScopeResolver } from '../../languages/cpp/scope-resolver.js';
|
||||
import { phpScopeResolver } from '../../languages/php/scope-resolver.js';
|
||||
import { javascriptScopeResolver } from '../../languages/javascript/scope-resolver.js';
|
||||
import { kotlinScopeResolver } from '../../languages/kotlin/scope-resolver.js';
|
||||
|
||||
/** Map of `SupportedLanguages` → `ScopeResolver`. The phase iterates
|
||||
* this map intersected with `MIGRATED_LANGUAGES` (the per-language
|
||||
@@ -36,4 +38,6 @@ export const SCOPE_RESOLVERS: ReadonlyMap<SupportedLanguages, ScopeResolver> = n
|
||||
[SupportedLanguages.C, cScopeResolver],
|
||||
[SupportedLanguages.CPlusPlus, cppScopeResolver],
|
||||
[SupportedLanguages.PHP, phpScopeResolver],
|
||||
[SupportedLanguages.JavaScript, javascriptScopeResolver],
|
||||
[SupportedLanguages.Kotlin, kotlinScopeResolver],
|
||||
]);
|
||||
|
||||
@@ -25,6 +25,7 @@
|
||||
|
||||
import type { ParsedFile, RegistryProviders } from 'gitnexus-shared';
|
||||
import type { KnowledgeGraph } from '../../../graph/types.js';
|
||||
import { lookupOwnedMembersByOwner } from '../../model/owned-members-lookup.js';
|
||||
import type { MutableSemanticModel, SemanticModel } from '../../model/semantic-model.js';
|
||||
import { reconcileOwnership, validateOwnershipParity } from './reconcile-ownership.js';
|
||||
import { validateBindingsImmutability } from './validate-bindings-immutability.js';
|
||||
@@ -342,6 +343,8 @@ export function runScopeResolution(
|
||||
const { referenceIndex, stats: resolveStats } = resolveReferenceSites({
|
||||
scopes: indexes,
|
||||
providers: registryProviders,
|
||||
ownedMembersByOwner: (ownerDefId, memberName) =>
|
||||
lookupOwnedMembersByOwner(readonlyModel, ownerDefId, memberName),
|
||||
});
|
||||
const tResolve = PROF ? process.hrtime.bigint() : 0n;
|
||||
|
||||
|
||||
@@ -301,10 +301,7 @@ export interface ParseWorkerInput {
|
||||
content: string;
|
||||
}
|
||||
|
||||
type WorkerIncomingMessage =
|
||||
| { type: 'sub-batch'; files: ParseWorkerInput[] }
|
||||
| { type: 'flush' }
|
||||
| ParseWorkerInput[];
|
||||
type WorkerIncomingMessage = { type: 'sub-batch'; files: ParseWorkerInput[] } | { type: 'flush' };
|
||||
|
||||
// ============================================================================
|
||||
// Worker-local parser + language map
|
||||
@@ -1401,6 +1398,15 @@ const processFileGroup = (
|
||||
// Skip files larger than the max tree-sitter buffer (32 MB)
|
||||
if (getTreeSitterContentByteLength(file.content) > TREE_SITTER_MAX_BUFFER) continue;
|
||||
|
||||
// Authoritative in-flight signal for the pool: lets `WorkerPool` exclude
|
||||
// exactly this file if the worker dies during parse/extract, instead of
|
||||
// guessing from `items[lastProgress]` (which the language-grouped order
|
||||
// here would defeat). The pool gracefully ignores this when running an
|
||||
// older worker build that doesn't emit it.
|
||||
if (parentPort) {
|
||||
parentPort.postMessage({ type: 'starting-file', path: file.path });
|
||||
}
|
||||
|
||||
// Vue SFC preprocessing: extract <script> block content
|
||||
let parseContent = file.content;
|
||||
let lineOffset = 0;
|
||||
@@ -1458,8 +1464,11 @@ const processFileGroup = (
|
||||
parseContent,
|
||||
file.path,
|
||||
(message) => {
|
||||
if (parentPort) parentPort.postMessage({ type: 'warning', message });
|
||||
else logger.warn(message);
|
||||
if (parentPort) {
|
||||
parentPort.postMessage({ type: 'warning', message });
|
||||
} else {
|
||||
logger.warn(message);
|
||||
}
|
||||
},
|
||||
tree,
|
||||
);
|
||||
@@ -2438,20 +2447,59 @@ const mergeResult = (target: ParseWorkerResult, src: ParseWorkerResult) => {
|
||||
target.fileCount += src.fileCount;
|
||||
};
|
||||
|
||||
// Signal the pool that worker-side initialization (parser imports, language
|
||||
// grammars, type-env setup, all helper modules) is complete and the message
|
||||
// handler below is about to be attached. The pool's `waitForWorkerReady`
|
||||
// resolves on this handshake — without it, a worker that crashes during
|
||||
// top-of-script init slips past pool startup (Node's `online` event fires
|
||||
// before the script body runs) and the pool only notices via the first
|
||||
// dispatch's idle timeout (~30s). Emit once; the dispatch handler treats
|
||||
// any subsequent `ready` message as a benign no-op.
|
||||
//
|
||||
// Native postMessage carries the ready handshake — Node's structured
|
||||
// clone delivers `{type:'ready'}` to the pool's waitForWorkerReady
|
||||
// listener directly. The pool drops the slot if this isn't seen within
|
||||
// `WORKER_READY_TIMEOUT_MS` (5s), so emitting it AFTER all top-of-script
|
||||
// init (imports, native binding loads, type-env setup) completes is the
|
||||
// load-bearing signal that this worker is ready for dispatch.
|
||||
parentPort!.postMessage({ type: 'ready' });
|
||||
|
||||
// Module-scope `TextDecoder` for sub-batch content. The pool sends each
|
||||
// file's content as a `Uint8Array` (zero-copy ArrayBuffer transfer); we
|
||||
// decode to string lazily here, once per file, before handing to
|
||||
// tree-sitter. Hoisted to module scope so we don't allocate a new
|
||||
// ICU-backed decoder per sub-batch — `TextDecoder.decode()` is
|
||||
// stateless across calls and safe to share.
|
||||
const sharedContentDecoder = new TextDecoder('utf-8');
|
||||
|
||||
/**
|
||||
* Convert the pool's sub-batch `files` array (content as `Uint8Array`,
|
||||
* transferred zero-copy) into the `ParseWorkerInput[]` shape
|
||||
* `processBatch` expects (content as `string`). This is the one place
|
||||
* the UTF-8 decode happens — runs on the worker thread in parallel with
|
||||
* continued main-thread work.
|
||||
*/
|
||||
function decodeSubBatchFiles(
|
||||
files: Array<{ path: string; content: Uint8Array | string }>,
|
||||
): ParseWorkerInput[] {
|
||||
return files.map((f) => ({
|
||||
path: f.path,
|
||||
// Test scaffolding (the writeReadyWorker preamble that wraps
|
||||
// parentPort.on) may already convert content to string before
|
||||
// calling here; tolerate both shapes so the same worker code
|
||||
// exercises real and synthetic dispatches.
|
||||
content: typeof f.content === 'string' ? f.content : sharedContentDecoder.decode(f.content),
|
||||
}));
|
||||
}
|
||||
|
||||
parentPort!.on('message', (msg: WorkerIncomingMessage) => {
|
||||
try {
|
||||
// Legacy single-message mode (backward compat): array of files
|
||||
if (Array.isArray(msg)) {
|
||||
const result = processBatch(msg, (filesProcessed) => {
|
||||
parentPort!.postMessage({ type: 'progress', filesProcessed });
|
||||
});
|
||||
parentPort!.postMessage({ type: 'result', data: result });
|
||||
return;
|
||||
}
|
||||
|
||||
// Sub-batch mode: { type: 'sub-batch', files: [...] }
|
||||
if (msg.type === 'sub-batch') {
|
||||
const result = processBatch(msg.files, (filesProcessed) => {
|
||||
const files = decodeSubBatchFiles(
|
||||
msg.files as Array<{ path: string; content: Uint8Array | string }>,
|
||||
);
|
||||
const result = processBatch(files, (filesProcessed) => {
|
||||
parentPort!.postMessage({
|
||||
type: 'progress',
|
||||
filesProcessed: cumulativeProcessed + filesProcessed,
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
/**
|
||||
* Quarantine layer (Layer 3 of the worker-pool resilience model).
|
||||
*
|
||||
* Tracks paths that caused a worker death this pool lifetime and must
|
||||
* not be re-dispatched to a worker. Session-scoped — created once per
|
||||
* `createWorkerPool` invocation and discarded with the pool.
|
||||
*
|
||||
* This module is the first piece of the U13 layer-extraction work. The
|
||||
* doc-review's A10 finding flagged the full 5-module split as
|
||||
* abstraction-without-multi-consumer-demand, so the rest of the
|
||||
* extraction is deferred until a real second consumer emerges (e.g., a
|
||||
* non-parse worker pool that reuses the same resilience layers).
|
||||
* Extracting the smallest self-contained layer first validates the
|
||||
* factory + interface pattern with minimal risk: behavior is unchanged,
|
||||
* the worker-pool.ts public API is unchanged, and existing tests act as
|
||||
* the regression net.
|
||||
*/
|
||||
|
||||
/**
|
||||
* Operations a {@link createQuarantine} instance exposes to the worker
|
||||
* pool. Intentionally tiny — anything more would invite the abstraction
|
||||
* overhead doc-review A10 cautioned against. Snapshot returns a fresh
|
||||
* `string[]` (not a `Set` or iterator) so callers can pass it directly
|
||||
* to `WorkerPoolDispatchError` without an `Array.from` dance and so
|
||||
* mutations to the returned array can't accidentally leak back into the
|
||||
* internal set.
|
||||
*/
|
||||
export interface Quarantine {
|
||||
/** Mark `path` as known-bad for the remainder of this pool's life. */
|
||||
add(path: string): void;
|
||||
/** Whether `path` has been quarantined. */
|
||||
has(path: string): boolean;
|
||||
/** Defensive copy of every quarantined path. */
|
||||
snapshot(): string[];
|
||||
/** How many distinct paths are currently quarantined. */
|
||||
readonly size: number;
|
||||
}
|
||||
|
||||
/**
|
||||
* Construct a fresh quarantine. Each `createWorkerPool` invocation gets
|
||||
* its own instance — quarantines never outlive the pool that created
|
||||
* them. The implementation is a thin wrapper around `Set<string>`; the
|
||||
* named interface exists to make the resilience layer addressable as a
|
||||
* unit (named module, dedicated tests) instead of an inline Set field
|
||||
* tangled into 1100+ LOC of pool plumbing.
|
||||
*/
|
||||
export function createQuarantine(): Quarantine {
|
||||
const paths = new Set<string>();
|
||||
return {
|
||||
add: (path) => {
|
||||
paths.add(path);
|
||||
},
|
||||
has: (path) => paths.has(path),
|
||||
snapshot: () => Array.from(paths),
|
||||
get size() {
|
||||
return paths.size;
|
||||
},
|
||||
};
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -154,6 +154,7 @@ export const splitRelCsvByLabelPair = async (
|
||||
let db: lbug.Database | null = null;
|
||||
let conn: lbug.Connection | null = null;
|
||||
let currentDbPath: string | null = null;
|
||||
let currentDbReadOnly = false;
|
||||
let ftsLoaded = false;
|
||||
let vectorExtensionLoaded = false;
|
||||
|
||||
@@ -448,12 +449,17 @@ export const initLbug = async (dbPath: string) => {
|
||||
* database is busy (e.g. `gitnexus analyze` holds the write lock).
|
||||
* Each retry waits DB_LOCK_RETRY_DELAY_MS * attempt milliseconds.
|
||||
*/
|
||||
export const withLbugDb = async <T>(dbPath: string, operation: () => Promise<T>): Promise<T> => {
|
||||
export const withLbugDb = async <T>(
|
||||
dbPath: string,
|
||||
operation: () => Promise<T>,
|
||||
options: { readOnly?: boolean } = {},
|
||||
): Promise<T> => {
|
||||
let lastError: unknown;
|
||||
const readOnly = options.readOnly === true;
|
||||
for (let attempt = 1; attempt <= DB_LOCK_RETRY_ATTEMPTS; attempt++) {
|
||||
try {
|
||||
return await runWithSessionLock(async () => {
|
||||
await ensureLbugInitialized(dbPath);
|
||||
await ensureLbugInitialized(dbPath, readOnly);
|
||||
return operation();
|
||||
});
|
||||
} catch (err) {
|
||||
@@ -483,15 +489,15 @@ export const withLbugDb = async <T>(dbPath: string, operation: () => Promise<T>)
|
||||
throw lastError;
|
||||
};
|
||||
|
||||
const ensureLbugInitialized = async (dbPath: string) => {
|
||||
if (conn && currentDbPath === dbPath) {
|
||||
const ensureLbugInitialized = async (dbPath: string, readOnly: boolean = false) => {
|
||||
if (conn && currentDbPath === dbPath && currentDbReadOnly === readOnly) {
|
||||
return { db, conn };
|
||||
}
|
||||
await doInitLbug(dbPath);
|
||||
await doInitLbug(dbPath, readOnly);
|
||||
return { db, conn };
|
||||
};
|
||||
|
||||
const doInitLbug = async (dbPath: string) => {
|
||||
const doInitLbug = async (dbPath: string, readOnly: boolean = false) => {
|
||||
// Different database requested — close the old one first
|
||||
if (conn || db) {
|
||||
await safeClose();
|
||||
@@ -575,9 +581,12 @@ const doInitLbug = async (dbPath: string) => {
|
||||
const parentDir = path.dirname(dbPath);
|
||||
await fs.mkdir(parentDir, { recursive: true });
|
||||
|
||||
const opened = await openLbugConnection(lbug, dbPath);
|
||||
const opened = readOnly
|
||||
? await openLbugConnection(lbug, dbPath, { readOnly: true })
|
||||
: await openLbugConnection(lbug, dbPath);
|
||||
db = opened.db;
|
||||
conn = opened.conn;
|
||||
currentDbReadOnly = readOnly;
|
||||
} finally {
|
||||
await releaseInitLock();
|
||||
}
|
||||
@@ -614,7 +623,7 @@ const doInitLbug = async (dbPath: string) => {
|
||||
` Original error: ${msg.slice(0, 200)}`,
|
||||
);
|
||||
}
|
||||
if (!msg.includes('already exists') && !isDbBusyError(err)) {
|
||||
if (!msg.includes('already exists') && !isDbBusyError(err) && !isReadOnlyDbError(err)) {
|
||||
logger.warn(`⚠️ Schema creation warning: ${msg.slice(0, 120)}`);
|
||||
}
|
||||
}
|
||||
@@ -1058,12 +1067,7 @@ export const batchInsertNodesToLbug = async (
|
||||
};
|
||||
|
||||
export const executeQuery = async (cypher: string): Promise<any[]> => {
|
||||
if (!conn) {
|
||||
throw new Error('LadybugDB not initialized. Call initLbug first.');
|
||||
}
|
||||
|
||||
const queryResult = await conn.query(cypher);
|
||||
return await readQueryRows(queryResult);
|
||||
return await executePrepared(cypher, {});
|
||||
};
|
||||
|
||||
export const streamQuery = async (
|
||||
@@ -1647,7 +1651,10 @@ export const createFTSIndex = async (
|
||||
if (ensuredFTSIndexes.has(key)) return;
|
||||
|
||||
if (!(await loadFTSExtension())) {
|
||||
return;
|
||||
throw new Error(
|
||||
`FTS extension unavailable - cannot create FTS index ${tableName}.${indexName}. ` +
|
||||
'Run `gitnexus doctor` and ensure the LadybugDB FTS extension is installed and loadable on this machine.',
|
||||
);
|
||||
}
|
||||
|
||||
const propList = properties.map((p) => `'${p}'`).join(', ');
|
||||
@@ -1726,19 +1733,15 @@ export const queryFTS = async (
|
||||
throw new Error('LadybugDB not initialized. Call initLbug first.');
|
||||
}
|
||||
|
||||
// Escape backslashes and single quotes to prevent Cypher injection
|
||||
const escapedQuery = query.replace(/\\/g, '\\\\').replace(/'/g, "''");
|
||||
|
||||
const cypher = `
|
||||
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', '${escapedQuery}', conjunctive := ${conjunctive})
|
||||
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', $query, conjunctive := ${conjunctive})
|
||||
RETURN node, score
|
||||
ORDER BY score DESC
|
||||
LIMIT ${limit}
|
||||
`;
|
||||
|
||||
try {
|
||||
const queryResult = await conn.query(cypher);
|
||||
const rows = await readQueryRows(queryResult);
|
||||
const rows = await executePrepared(cypher, { query });
|
||||
|
||||
return rows.map((row: any) => {
|
||||
const node = row.node || row[0] || {};
|
||||
|
||||
@@ -16,14 +16,52 @@
|
||||
*/
|
||||
|
||||
import fs from 'fs/promises';
|
||||
import os from 'os';
|
||||
import path from 'path';
|
||||
import lbug from '@ladybugdb/core';
|
||||
import { loadFTSExtension } from './lbug-adapter.js';
|
||||
import { isReadOnlyDbError, loadFTSExtension } from './lbug-adapter.js';
|
||||
import {
|
||||
createLbugDatabase,
|
||||
isWalCorruptionError,
|
||||
WAL_RECOVERY_SUGGESTION,
|
||||
} from './lbug-config.js';
|
||||
|
||||
/**
|
||||
* Probe whether a Windows FTS extension binary is locally installed under
|
||||
* ~/.lbdb/extension/<any-version>/win_amd64/fts/. Returns true on the first
|
||||
* version dir whose libfts.lbug_extension exists on disk; false if the
|
||||
* extension root is missing or contains no FTS binary.
|
||||
*
|
||||
* Gates the Windows skip-FTS-load guard below so we only skip the load
|
||||
* when no extension binary is present. When at least one binary exists,
|
||||
* loadFTSExtension is called with policy: 'load-only' — LadybugDB resolves
|
||||
* LOAD EXTENSION fts to its version-specific path internally, and the
|
||||
* ExtensionManager's tryLoad try/catch handles version-mismatch errors
|
||||
* cleanly without ever attempting dlopen of a stale binary. The install
|
||||
* path that the #1199/#1217 SIGSEGV documented is never exercised at
|
||||
* query time.
|
||||
*
|
||||
* Exported so unit tests can exercise the probe directly against a
|
||||
* temp-dir plus spied `os.homedir()` — see lbug-pool-win-fts-probe.test.ts.
|
||||
*/
|
||||
export async function hasLocalWinFtsExtension(): Promise<boolean> {
|
||||
try {
|
||||
const extRoot = path.join(os.homedir(), '.lbdb', 'extension');
|
||||
const versions = await fs.readdir(extRoot);
|
||||
for (const v of versions) {
|
||||
try {
|
||||
await fs.stat(path.join(extRoot, v, 'win_amd64', 'fts', 'libfts.lbug_extension'));
|
||||
return true;
|
||||
} catch {
|
||||
/* missing for this version, keep looking */
|
||||
}
|
||||
}
|
||||
} catch {
|
||||
/* no .lbdb/extension dir */
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/** Per-repo pool: one Database, many Connections */
|
||||
interface PoolEntry {
|
||||
db: lbug.Database;
|
||||
@@ -423,14 +461,24 @@ async function doInitLbug(repoId: string, dbPath: string): Promise<void> {
|
||||
// install; analyze owns extension installation. If LOAD fails, search
|
||||
// features degrade gracefully and the user-facing query path proceeds.
|
||||
if (!shared.ftsLoaded) {
|
||||
// Windows guard: LOAD EXTENSION fts crashes with SIGSEGV on Windows when
|
||||
// the FTS extension binary is not installed locally (@ladybugdb/core native
|
||||
// bug — the extension loader hits an unhandled error path that signals SIGSEGV
|
||||
// rather than throwing a JS exception, so try/catch cannot protect here).
|
||||
// Skip the load on Windows; bm25-index.js catches the resulting Kuzu catalog
|
||||
// errors and returns empty BM25 results gracefully. Graph queries are unaffected.
|
||||
// Windows guard: LOAD EXTENSION fts crashes with SIGSEGV on Windows during
|
||||
// *install* — the @ladybugdb/core out-of-process installer hits an unhandled
|
||||
// error path that signals SIGSEGV instead of throwing (see #1199, #1217).
|
||||
// The previous unconditional skip was over-broad: it also disabled FTS on
|
||||
// hosts where the binary was already on disk and only needed LOAD, leaving
|
||||
// BM25 silently degraded with no error path (see #1690).
|
||||
//
|
||||
// Probe ~/.lbdb/extension/*/win_amd64/fts/ first. If any binary is on disk
|
||||
// we run loadFTSExtension(..., 'load-only'); the install path is never
|
||||
// exercised, and LadybugDB's version-specific resolution + ExtensionManager
|
||||
// try/catch handle stale/zero-byte siblings cleanly (verified empirically
|
||||
// on Win10 + Node 22.19 + gitnexus 1.6.5 + @ladybugdb/core 0.16.1). With
|
||||
// no binary at all, we fall back to the upstream skip so install-time
|
||||
// SIGSEGV continues to be avoided.
|
||||
if (process.platform === 'win32') {
|
||||
shared.ftsLoaded = true;
|
||||
shared.ftsLoaded = (await hasLocalWinFtsExtension())
|
||||
? await loadFTSExtension(available[0], { policy: 'load-only' })
|
||||
: true;
|
||||
} else {
|
||||
shared.ftsLoaded = await loadFTSExtension(available[0], { policy: 'load-only' });
|
||||
}
|
||||
@@ -497,10 +545,12 @@ export async function initLbugWithDb(
|
||||
// Load FTS extension if not already loaded on this Database.
|
||||
// policy: 'load-only' — same contract as initLbug above; the read pool
|
||||
// must not block on a network install during query execution.
|
||||
// Windows guard: same SIGSEGV risk as doInitLbug above — skip on Windows.
|
||||
// Windows guard: same probe-then-load policy as doInitLbug above.
|
||||
if (!shared.ftsLoaded) {
|
||||
if (process.platform === 'win32') {
|
||||
shared.ftsLoaded = true;
|
||||
shared.ftsLoaded = (await hasLocalWinFtsExtension())
|
||||
? await loadFTSExtension(available[0], { policy: 'load-only' })
|
||||
: true;
|
||||
} else {
|
||||
shared.ftsLoaded = await loadFTSExtension(available[0], { policy: 'load-only' });
|
||||
}
|
||||
@@ -598,30 +648,7 @@ function withTimeout<T>(promise: Promise<T>, ms: number, label: string): Promise
|
||||
}
|
||||
|
||||
export const executeQuery = async (repoId: string, cypher: string): Promise<any[]> => {
|
||||
const entry = pool.get(repoId);
|
||||
if (!entry) {
|
||||
throw new Error(`LadybugDB not initialized for repo "${repoId}". Call initLbug first.`);
|
||||
}
|
||||
|
||||
if (isWriteQuery(cypher)) {
|
||||
throw new Error('Write operations are not allowed. The pool adapter is read-only.');
|
||||
}
|
||||
|
||||
entry.lastUsed = Date.now();
|
||||
|
||||
const conn = await checkout(entry);
|
||||
silenceStdout();
|
||||
activeQueryCount++;
|
||||
try {
|
||||
const queryResult = await withTimeout(conn.query(cypher), QUERY_TIMEOUT_MS, 'Query');
|
||||
const result = Array.isArray(queryResult) ? queryResult[0] : queryResult;
|
||||
const rows = await result.getAll();
|
||||
return rows;
|
||||
} finally {
|
||||
activeQueryCount--;
|
||||
restoreStdout();
|
||||
checkin(entry, conn);
|
||||
}
|
||||
return await executeParameterized(repoId, cypher, {});
|
||||
};
|
||||
|
||||
/**
|
||||
@@ -653,6 +680,11 @@ export const executeParameterized = async (
|
||||
const result = Array.isArray(queryResult) ? queryResult[0] : queryResult;
|
||||
const rows = await result.getAll();
|
||||
return rows;
|
||||
} catch (err) {
|
||||
if (isReadOnlyDbError(err)) {
|
||||
throw new Error('Write operations are not allowed. The pool adapter is read-only.');
|
||||
}
|
||||
throw err;
|
||||
} finally {
|
||||
activeQueryCount--;
|
||||
restoreStdout();
|
||||
@@ -685,15 +717,3 @@ export const closeLbug = async (repoId?: string): Promise<void> => {
|
||||
* Check if a specific repo's pool is active
|
||||
*/
|
||||
export const isLbugReady = (repoId: string): boolean => pool.has(repoId);
|
||||
|
||||
/** Regex to detect write operations in user-supplied Cypher queries.
|
||||
* Note: CALL is NOT blocked — it's used for read-only FTS (CALL QUERY_FTS_INDEX)
|
||||
* and vector search (CALL QUERY_VECTOR_INDEX). The database is opened in
|
||||
* read-only mode as defense-in-depth against write procedures. */
|
||||
export const CYPHER_WRITE_RE =
|
||||
/(?<!:)\b(CREATE|DELETE|SET|MERGE|REMOVE|DROP|ALTER|COPY|DETACH|FOREACH|INSTALL|LOAD)\b/i;
|
||||
|
||||
/** Check if a Cypher query contains write operations */
|
||||
export function isWriteQuery(query: string): boolean {
|
||||
return CYPHER_WRITE_RE.test(query);
|
||||
}
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
/**
|
||||
* Return true only for plain-object payloads that can be safely used as
|
||||
* named parameter maps in prepared Cypher execution.
|
||||
*
|
||||
* Validation criteria:
|
||||
* - must be a JavaScript object (`typeof value === 'object'`)
|
||||
* - must not be `null`
|
||||
* - must not be an array
|
||||
* - must have a plain-object prototype
|
||||
* - values must be scalar bindable values (string | number | boolean | null)
|
||||
*
|
||||
* Rationale: prepared-statement params are key/value maps; rejecting null/array
|
||||
* and non-plain objects keeps binding behavior predictable and avoids passing
|
||||
* complex host objects to Ladybug parameter binding.
|
||||
*/
|
||||
const isBindableScalar = (value: unknown): value is string | number | boolean | null =>
|
||||
value === null || ['string', 'number', 'boolean'].includes(typeof value);
|
||||
|
||||
export const isValidQueryParams = (value: unknown): value is Record<string, unknown> =>
|
||||
value !== null &&
|
||||
typeof value === 'object' &&
|
||||
!Array.isArray(value) &&
|
||||
(Object.getPrototypeOf(value) === Object.prototype || Object.getPrototypeOf(value) === null) &&
|
||||
Object.values(value).every(isBindableScalar);
|
||||
@@ -25,7 +25,7 @@ import {
|
||||
deleteAllCommunitiesAndProcesses,
|
||||
queryImporters,
|
||||
} from './lbug/lbug-adapter.js';
|
||||
import { createSearchFTSIndexes } from './search/fts-indexes.js';
|
||||
import { createSearchFTSIndexes, verifySearchFTSIndexes } from './search/fts-indexes.js';
|
||||
import {
|
||||
getStoragePaths,
|
||||
saveMeta,
|
||||
@@ -71,6 +71,10 @@ export interface AnalyzeOptions {
|
||||
* bypass. See `allowDuplicateName` below.
|
||||
*/
|
||||
force?: boolean;
|
||||
/** Repair only search indexes without re-running full parsing/indexing. */
|
||||
repairFts?: boolean;
|
||||
/** Emit per-index FTS create logs. */
|
||||
verbose?: boolean;
|
||||
embeddings?: boolean;
|
||||
/**
|
||||
* Override the auto-skip node-count cap for embedding generation.
|
||||
@@ -110,6 +114,14 @@ export interface AnalyzeOptions {
|
||||
* of a pipeline re-index.
|
||||
*/
|
||||
allowDuplicateName?: boolean;
|
||||
/**
|
||||
* Worker pool size override, threaded from the CLI `--workers` flag.
|
||||
* Forwarded to `PipelineOptions.workerPoolSize` so the parse phase
|
||||
* sizes the pool without `analyzeCommand` mutating `process.env`.
|
||||
* `0` disables the pool (sequential fallback); positive integer sets
|
||||
* the count; `undefined` defers to the env / auto-formula fallback.
|
||||
*/
|
||||
workerPoolSize?: number;
|
||||
}
|
||||
|
||||
export interface AnalyzeResult {
|
||||
@@ -126,6 +138,8 @@ export interface AnalyzeResult {
|
||||
alreadyUpToDate?: boolean;
|
||||
/** The raw pipeline result — only populated when needed by callers (e.g. skill generation). */
|
||||
pipelineResult?: any;
|
||||
/** True when analyze only repaired FTS indexes and skipped pipeline re-analysis. */
|
||||
ftsRepairedOnly?: boolean;
|
||||
}
|
||||
|
||||
// Re-export the pure flag-derivation helper so external callers (and tests)
|
||||
@@ -190,6 +204,78 @@ export async function runFullAnalysis(
|
||||
const currentCommit = repoHasGit ? getCurrentCommit(repoPath) : '';
|
||||
const existingMeta = await loadMeta(storagePath);
|
||||
|
||||
// ── FTS-only repair path ────────────────────────────────────────────
|
||||
if (options.repairFts) {
|
||||
if (!existingMeta) {
|
||||
throw new Error(
|
||||
'Cannot repair FTS indexes because this repository has not been analyzed yet. ' +
|
||||
'Run `gitnexus analyze` first to create the initial index, then retry `--repair-fts`.',
|
||||
);
|
||||
}
|
||||
let lbugStat;
|
||||
try {
|
||||
lbugStat = await fs.lstat(lbugPath);
|
||||
} catch {
|
||||
throw new Error(
|
||||
`Cannot repair FTS indexes: graph store at ${lbugPath} is missing. ` +
|
||||
'Run `gitnexus analyze` (full) to rebuild from scratch.',
|
||||
);
|
||||
}
|
||||
if (!lbugStat.isFile()) {
|
||||
const foundType = lbugStat.isDirectory()
|
||||
? 'a directory'
|
||||
: lbugStat.isSymbolicLink()
|
||||
? 'a symbolic link'
|
||||
: lbugStat.isSocket()
|
||||
? 'a socket'
|
||||
: lbugStat.isBlockDevice()
|
||||
? 'a block device'
|
||||
: lbugStat.isCharacterDevice()
|
||||
? 'a character device'
|
||||
: lbugStat.isFIFO()
|
||||
? 'a FIFO'
|
||||
: 'not a regular file';
|
||||
throw new Error(
|
||||
`Cannot repair FTS indexes: graph store at ${lbugPath} is ${foundType} (expected a file). ` +
|
||||
'Run `gitnexus analyze` (full) to rebuild from scratch.',
|
||||
);
|
||||
}
|
||||
try {
|
||||
await initLbug(lbugPath);
|
||||
progress('fts', 85, 'Repairing search indexes...');
|
||||
await createSearchFTSIndexes({
|
||||
onIndexStart: options.verbose
|
||||
? (table, indexName) => log(`FTS: creating ${table}.${indexName}`)
|
||||
: undefined,
|
||||
onIndexReady: options.verbose
|
||||
? (table, indexName) => log(`FTS: ready ${table}.${indexName}`)
|
||||
: undefined,
|
||||
});
|
||||
const missing = await verifySearchFTSIndexes(executeQuery);
|
||||
if (missing.length > 0) {
|
||||
throw new Error(
|
||||
`FTS repair failed - missing indexes after rebuild: ${missing.join(', ')}. ` +
|
||||
'Run `gitnexus analyze --force` to perform a full graph+FTS rebuild; ' +
|
||||
'if that also fails, verify FTS extension availability via `gitnexus doctor`.',
|
||||
);
|
||||
}
|
||||
await ensureGitNexusIgnored(repoPath);
|
||||
progress('fts', 90, 'Search indexes ready');
|
||||
progress('done', 100, 'Done');
|
||||
return {
|
||||
repoName:
|
||||
options.registryName ??
|
||||
getInferredRepoName(repoPath) ??
|
||||
path.basename(resolveRepoIdentityRoot(repoPath)),
|
||||
repoPath,
|
||||
stats: existingMeta.stats ?? {},
|
||||
ftsRepairedOnly: true,
|
||||
};
|
||||
} finally {
|
||||
await closeLbug().catch(() => {});
|
||||
}
|
||||
}
|
||||
|
||||
// ── Crash recovery: dirty flag forces full rebuild ────────────────
|
||||
// If the previous incremental run set incrementalInProgress and didn't
|
||||
// clear it, the on-disk index may be in a half-state. Cheapest path
|
||||
@@ -366,7 +452,7 @@ export async function runFullAnalysis(
|
||||
: p.message || phaseLabel;
|
||||
progress(p.phase, scaled, message);
|
||||
},
|
||||
{ parseCache },
|
||||
{ parseCache, workerPoolSize: options.workerPoolSize },
|
||||
);
|
||||
|
||||
// ── Phase 2: LadybugDB (60–85%) ──────────────────────────────────
|
||||
@@ -583,7 +669,21 @@ export async function runFullAnalysis(
|
||||
|
||||
// ── Phase 3: FTS (85–90%) ─────────────────────────────────────────
|
||||
progress('fts', 85, 'Creating search indexes...');
|
||||
await createSearchFTSIndexes();
|
||||
await createSearchFTSIndexes({
|
||||
onIndexStart: options.verbose
|
||||
? (table, indexName) => log(`FTS: creating ${table}.${indexName}`)
|
||||
: undefined,
|
||||
onIndexReady: options.verbose
|
||||
? (table, indexName) => log(`FTS: ready ${table}.${indexName}`)
|
||||
: undefined,
|
||||
});
|
||||
const missingIndexNames = await verifySearchFTSIndexes(executeQuery);
|
||||
if (missingIndexNames.length > 0) {
|
||||
throw new Error(
|
||||
`FTS verification failed - missing indexes after analyze: ${missingIndexNames.join(', ')}. ` +
|
||||
'Check FTS extension availability, then retry `gitnexus analyze --force` for a full rebuild.',
|
||||
);
|
||||
}
|
||||
progress('fts', 90, 'Search indexes ready');
|
||||
|
||||
// ── Phase 3.5: Re-insert cached embeddings ────────────────────────
|
||||
|
||||
@@ -27,22 +27,20 @@ export interface FTSSearchResponse {
|
||||
* caller can distinguish "zero matches" from "index missing".
|
||||
*/
|
||||
async function queryFTSViaExecutor(
|
||||
executor: (cypher: string) => Promise<any[]>,
|
||||
executor: (cypher: string, params: Record<string, any>) => Promise<any[]>,
|
||||
tableName: string,
|
||||
indexName: string,
|
||||
query: string,
|
||||
limit: number,
|
||||
): Promise<Array<{ filePath: string; score: number; nodeId: string }> | null> {
|
||||
// Escape single quotes and backslashes to prevent Cypher injection
|
||||
const escapedQuery = query.replace(/\\/g, '\\\\').replace(/'/g, "''");
|
||||
const cypher = `
|
||||
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', '${escapedQuery}', conjunctive := false)
|
||||
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', $query, conjunctive := false)
|
||||
RETURN node, score
|
||||
ORDER BY score DESC
|
||||
LIMIT ${limit}
|
||||
`;
|
||||
try {
|
||||
const rows = await executor(cypher);
|
||||
const rows = await executor(cypher, { query });
|
||||
return rows.map((row: any) => {
|
||||
const node = row.node || row[0] || {};
|
||||
const score = row.score ?? row[1] ?? 0;
|
||||
@@ -81,8 +79,9 @@ export const searchFTSFromLbug = async (
|
||||
// IMPORTANT: FTS queries run sequentially to avoid connection contention.
|
||||
// The MCP pool supports multiple connections, but FTS is best run serially.
|
||||
const poolMod = await import('../lbug/pool-adapter.js');
|
||||
const { executeQuery } = poolMod;
|
||||
const executor = (cypher: string) => executeQuery(repoId, cypher);
|
||||
const { executeParameterized } = poolMod;
|
||||
const executor = (cypher: string, params: Record<string, any>) =>
|
||||
executeParameterized(repoId, cypher, params);
|
||||
|
||||
for (const { table, indexName } of FTS_INDEXES) {
|
||||
const result = await queryFTSViaExecutor(executor, table, indexName, query, limit);
|
||||
|
||||
@@ -1,8 +1,45 @@
|
||||
import { createFTSIndex } from '../lbug/lbug-adapter.js';
|
||||
import { FTS_INDEXES } from './fts-schema.js';
|
||||
|
||||
export async function createSearchFTSIndexes(): Promise<void> {
|
||||
export interface CreateSearchFTSIndexesOptions {
|
||||
onIndexStart?: (table: string, indexName: string) => void;
|
||||
onIndexReady?: (table: string, indexName: string) => void;
|
||||
}
|
||||
|
||||
export async function createSearchFTSIndexes(
|
||||
options?: CreateSearchFTSIndexesOptions,
|
||||
): Promise<void> {
|
||||
for (const { table, indexName, properties } of FTS_INDEXES) {
|
||||
options?.onIndexStart?.(table, indexName);
|
||||
await createFTSIndex(table, indexName, [...properties]);
|
||||
options?.onIndexReady?.(table, indexName);
|
||||
}
|
||||
}
|
||||
|
||||
export async function verifySearchFTSIndexes(
|
||||
executeQuery: (cypher: string) => Promise<unknown[]>,
|
||||
): Promise<string[]> {
|
||||
const safeIdentifier = (value: string): string => {
|
||||
if (!/^[A-Za-z_][A-Za-z0-9_]*$/.test(value)) {
|
||||
throw new Error(`Invalid FTS identifier: ${value}`);
|
||||
}
|
||||
return value;
|
||||
};
|
||||
|
||||
const missing: string[] = [];
|
||||
for (const { table, indexName } of FTS_INDEXES) {
|
||||
const safeTable = safeIdentifier(table);
|
||||
const safeIndex = safeIdentifier(indexName);
|
||||
const probe = `
|
||||
CALL QUERY_FTS_INDEX('${safeTable}', '${safeIndex}', '__gitnexus_fts_probe__', conjunctive := false)
|
||||
RETURN score
|
||||
LIMIT 1
|
||||
`;
|
||||
try {
|
||||
await executeQuery(probe);
|
||||
} catch {
|
||||
missing.push(`${table}.${indexName}`);
|
||||
}
|
||||
}
|
||||
return missing;
|
||||
}
|
||||
|
||||
@@ -66,12 +66,15 @@ export interface WikiOptions {
|
||||
concurrency?: number;
|
||||
/** If true, stop after building module tree for user review */
|
||||
reviewOnly?: boolean;
|
||||
/** Output language for generated documentation (e.g. 'english', 'chinese', 'spanish') */
|
||||
lang?: string;
|
||||
}
|
||||
|
||||
export interface WikiMeta {
|
||||
fromCommit: string;
|
||||
generatedAt: string;
|
||||
model: string;
|
||||
lang: string;
|
||||
moduleFiles: Record<string, string[]>;
|
||||
moduleTree: ModuleTreeNode[];
|
||||
}
|
||||
@@ -177,6 +180,28 @@ export class WikiGenerator {
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Return the effective lang string: strip control characters, trim, cap at 50 chars,
|
||||
* then validate against a character allowlist. Returns '' if the value is absent or invalid.
|
||||
* Used for both prompt construction and meta storage/comparison so they are always in sync.
|
||||
*/
|
||||
private effectiveLang(): string {
|
||||
const lang = (this.options.lang ?? '')
|
||||
.replace(/[\x00-\x1F\x7F]/g, '')
|
||||
.trim()
|
||||
.slice(0, 50);
|
||||
return /^[a-zA-Z -]+$/.test(lang) ? lang : '';
|
||||
}
|
||||
|
||||
/**
|
||||
* Append an output-language instruction to a system prompt when --lang is set.
|
||||
*/
|
||||
private buildSystemPrompt(base: string): string {
|
||||
const lang = this.effectiveLang();
|
||||
if (!lang) return base;
|
||||
return `${base}\n\nIMPORTANT: Write ALL documentation content in ${lang}. This includes prose, code comments in examples, and diagram labels. Note: page titles (H1 headings) are generated separately and will remain in English.`;
|
||||
}
|
||||
|
||||
/**
|
||||
* Route LLM call to the appropriate provider (OpenAI-compatible or Cursor CLI).
|
||||
*/
|
||||
@@ -207,6 +232,15 @@ export class WikiGenerator {
|
||||
|
||||
// Up-to-date check (skip if --force)
|
||||
if (!forceMode && existingMeta && existingMeta.fromCommit === currentCommit) {
|
||||
const currentLang = this.effectiveLang();
|
||||
const metaLang = existingMeta.lang ?? '';
|
||||
if (currentLang !== metaLang) {
|
||||
const prevDisplay = metaLang || 'english (default)';
|
||||
const nextDisplay = currentLang || 'english (default)';
|
||||
throw new Error(
|
||||
`Wiki was generated in ${prevDisplay}; use --force to regenerate in ${nextDisplay}.`,
|
||||
);
|
||||
}
|
||||
// Still regenerate the HTML viewer in case it's missing
|
||||
await this.ensureHTMLViewer();
|
||||
return { pagesGenerated: 0, mode: 'up-to-date', failedModules: [] };
|
||||
@@ -235,6 +269,15 @@ export class WikiGenerator {
|
||||
let result: WikiRunResult;
|
||||
try {
|
||||
if (!forceMode && existingMeta && existingMeta.fromCommit) {
|
||||
const currentLang = this.effectiveLang();
|
||||
const metaLang = existingMeta.lang ?? '';
|
||||
if (currentLang !== metaLang) {
|
||||
const prevDisplay = metaLang || 'english (default)';
|
||||
const nextDisplay = currentLang || 'english (default)';
|
||||
throw new Error(
|
||||
`Wiki was generated in ${prevDisplay}; use --force to regenerate in ${nextDisplay}.`,
|
||||
);
|
||||
}
|
||||
result = await this.incrementalUpdate(existingMeta, currentCommit);
|
||||
} else {
|
||||
result = await this.fullGeneration(currentCommit);
|
||||
@@ -368,6 +411,7 @@ export class WikiGenerator {
|
||||
fromCommit: currentCommit,
|
||||
generatedAt: new Date().toISOString(),
|
||||
model: this.llmConfig.model,
|
||||
lang: this.effectiveLang(),
|
||||
moduleFiles,
|
||||
moduleTree,
|
||||
});
|
||||
@@ -415,6 +459,9 @@ export class WikiGenerator {
|
||||
DIRECTORY_TREE: dirTree,
|
||||
});
|
||||
|
||||
// Grouping is a structured-data phase (JSON output), not documentation.
|
||||
// Do NOT apply buildSystemPrompt here — a language instruction would risk
|
||||
// translating module-name keys, breaking slug stability and JSON parsing.
|
||||
const response = await this.invokeLLM(
|
||||
prompt,
|
||||
GROUPING_SYSTEM_PROMPT,
|
||||
@@ -589,9 +636,13 @@ export class WikiGenerator {
|
||||
PROCESSES: formatProcesses(processes),
|
||||
});
|
||||
|
||||
const response = await this.invokeLLM(prompt, MODULE_SYSTEM_PROMPT, this.streamOpts(node.name));
|
||||
const response = await this.invokeLLM(
|
||||
prompt,
|
||||
this.buildSystemPrompt(MODULE_SYSTEM_PROMPT),
|
||||
this.streamOpts(node.name),
|
||||
);
|
||||
|
||||
// Write page with front matter
|
||||
// H1 uses the English module name (stable slug source); body is LLM-translated.
|
||||
const pageContent = sanitizeMermaidMarkdown(`# ${node.name}\n\n${response.content}`);
|
||||
await fs.writeFile(path.join(this.wikiDir, `${node.slug}.md`), pageContent, 'utf-8');
|
||||
}
|
||||
@@ -630,7 +681,11 @@ export class WikiGenerator {
|
||||
CROSS_PROCESSES: formatProcesses(processes),
|
||||
});
|
||||
|
||||
const response = await this.invokeLLM(prompt, PARENT_SYSTEM_PROMPT, this.streamOpts(node.name));
|
||||
const response = await this.invokeLLM(
|
||||
prompt,
|
||||
this.buildSystemPrompt(PARENT_SYSTEM_PROMPT),
|
||||
this.streamOpts(node.name),
|
||||
);
|
||||
|
||||
const pageContent = sanitizeMermaidMarkdown(`# ${node.name}\n\n${response.content}`);
|
||||
await fs.writeFile(path.join(this.wikiDir, `${node.slug}.md`), pageContent, 'utf-8');
|
||||
@@ -678,7 +733,7 @@ export class WikiGenerator {
|
||||
|
||||
const response = await this.invokeLLM(
|
||||
prompt,
|
||||
OVERVIEW_SYSTEM_PROMPT,
|
||||
this.buildSystemPrompt(OVERVIEW_SYSTEM_PROMPT),
|
||||
this.streamOpts('Generating overview', 88),
|
||||
);
|
||||
|
||||
@@ -713,6 +768,7 @@ export class WikiGenerator {
|
||||
...existingMeta,
|
||||
fromCommit: currentCommit,
|
||||
generatedAt: new Date().toISOString(),
|
||||
lang: this.effectiveLang(),
|
||||
});
|
||||
return { pagesGenerated: 0, mode: 'incremental', failedModules: [] };
|
||||
}
|
||||
@@ -817,6 +873,7 @@ export class WikiGenerator {
|
||||
fromCommit: currentCommit,
|
||||
generatedAt: new Date().toISOString(),
|
||||
model: this.llmConfig.model,
|
||||
lang: this.effectiveLang(),
|
||||
});
|
||||
|
||||
this.onProgress('done', 100, 'Incremental update complete');
|
||||
|
||||
@@ -23,7 +23,7 @@ export interface LLMConfig {
|
||||
apiVersion?: string;
|
||||
/** When true, strips sampling params and uses max_completion_tokens instead of max_tokens */
|
||||
isReasoningModel?: boolean;
|
||||
/** Per-attempt fetch timeout in ms (default: 60_000). */
|
||||
/** Per-attempt fetch timeout in ms. Omit to disable request timeouts. */
|
||||
requestTimeoutMs?: number;
|
||||
/** Max fetch attempts before giving up (default: 3). */
|
||||
maxAttempts?: number;
|
||||
@@ -81,6 +81,19 @@ export function estimateTokens(text: string): number {
|
||||
return Math.ceil(text.length / 4);
|
||||
}
|
||||
|
||||
function formatTimeoutDuration(timeoutMs: number): string {
|
||||
if (timeoutMs >= 1000 && timeoutMs % 1000 === 0) {
|
||||
return `${timeoutMs / 1000}s`;
|
||||
}
|
||||
return `${timeoutMs}ms`;
|
||||
}
|
||||
|
||||
function isTimeoutLikeError(err: unknown): boolean {
|
||||
if (!(err instanceof Error)) return false;
|
||||
if (err.name === 'TimeoutError' || err.name === 'AbortError') return true;
|
||||
return /time(d)?\s*out|timeout/i.test(err.message);
|
||||
}
|
||||
|
||||
/**
|
||||
* Validate that a base URL supplied for LLM API calls is a safe HTTP/HTTPS
|
||||
* endpoint (CWE-918 / CodeQL js/http-to-file-access).
|
||||
@@ -237,12 +250,13 @@ export async function callLLM(
|
||||
...authHeaders,
|
||||
},
|
||||
body: JSON.stringify(body),
|
||||
// Per-attempt timeout. Without this each retry can hang
|
||||
// indefinitely on a frozen TCP connection — the per-call
|
||||
// signal is the only timeout `resilientFetch` honors;
|
||||
// `capDelayMs` only bounds the *backoff* between attempts.
|
||||
// Default 60s; raise via --timeout for slow models or large pages.
|
||||
signal: AbortSignal.timeout(config.requestTimeoutMs ?? 60_000),
|
||||
// Request timeout is opt-in for wiki generation. Large local
|
||||
// model runs can legitimately take well over a minute, so the
|
||||
// default runtime path must not impose a hidden 60s ceiling.
|
||||
signal:
|
||||
config.requestTimeoutMs !== undefined
|
||||
? AbortSignal.timeout(config.requestTimeoutMs)
|
||||
: undefined,
|
||||
},
|
||||
{
|
||||
breakerKey: `wiki-llm-${new URL(url).host}`,
|
||||
@@ -261,6 +275,12 @@ export async function callLLM(
|
||||
`LLM API error (${err.response.status} after retries): ${errorText.slice(0, 500)}`,
|
||||
);
|
||||
}
|
||||
if (config.requestTimeoutMs !== undefined && isTimeoutLikeError(err)) {
|
||||
throw new Error(
|
||||
`LLM request timed out after ${formatTimeoutDuration(config.requestTimeoutMs)}. ` +
|
||||
'Increase --timeout or omit it to disable the request timeout.',
|
||||
);
|
||||
}
|
||||
throw err;
|
||||
}
|
||||
|
||||
|
||||
@@ -14,15 +14,20 @@ import {
|
||||
executeParameterized,
|
||||
closeLbug,
|
||||
isLbugReady,
|
||||
isWriteQuery,
|
||||
} from '../../core/lbug/pool-adapter.js';
|
||||
import { isValidQueryParams } from '../../core/lbug/query-params.js';
|
||||
import { isWalCorruptionError, WAL_RECOVERY_SUGGESTION } from '../../core/lbug/lbug-config.js';
|
||||
export { isWriteQuery };
|
||||
// Embedding imports are lazy (dynamic import) to avoid loading onnxruntime-node
|
||||
// at MCP server startup — crashes on unsupported Node ABI versions (#89)
|
||||
// git utilities available if needed
|
||||
// import { isGitRepo, getCurrentCommit, getGitRoot } from '../../storage/git.js';
|
||||
import { parseDiffHunks, type FileDiff } from '../../storage/git.js';
|
||||
import {
|
||||
parseDiffHunks,
|
||||
getCanonicalRepoRoot,
|
||||
getGitRoot,
|
||||
type FileDiff,
|
||||
} from '../../storage/git.js';
|
||||
import { realpathSync } from 'fs';
|
||||
import {
|
||||
listRegisteredRepos,
|
||||
cleanupOldKuzuFiles,
|
||||
@@ -169,6 +174,9 @@ function logQueryError(context: string, err: unknown): void {
|
||||
logger.error({ context, err: msg }, 'GitNexus query failed');
|
||||
}
|
||||
|
||||
const isReadOnlyDbError = (err: unknown): boolean =>
|
||||
/read-only database/i.test(err instanceof Error ? err.message : String(err));
|
||||
|
||||
/**
|
||||
* Per-query latency telemetry for production aggregation (#553).
|
||||
*
|
||||
@@ -211,6 +219,82 @@ interface RepoHandle {
|
||||
stats?: RegistryEntry['stats'];
|
||||
}
|
||||
|
||||
/** Resolve symlinks for path comparison; falls back to path.resolve on error.
|
||||
* Uses `realpathSync.native` (not the pure-JS `realpathSync`) so that Windows
|
||||
* 8.3 short names (e.g. RUNNER~1 → runneradmin) are expanded to long form,
|
||||
* matching the output of `git rev-parse --show-toplevel`. */
|
||||
function tryRealpath(p: string): string {
|
||||
try {
|
||||
return realpathSync.native(p);
|
||||
} catch {
|
||||
return path.resolve(p);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve the git diff cwd for detect_changes, auto-detecting linked worktrees.
|
||||
*
|
||||
* When `launchCwd` is a linked worktree of the same canonical repository as
|
||||
* `repoPath` (i.e. `getGitRoot(launchCwd)` differs from `repoPath` but both
|
||||
* share the same `getCanonicalRepoRoot`), returns the worktree's git root so
|
||||
* that `git diff` sees the correct working directory and index.
|
||||
*
|
||||
* Returns `repoPath` unchanged in all other cases (non-worktree, git
|
||||
* unavailable, unrelated repo).
|
||||
*
|
||||
* Extracted as a module-level export so tests can pass any `launchCwd` instead
|
||||
* of relying on `process.cwd()`, which is fixed to the server launch directory
|
||||
* and cannot be changed mid-process.
|
||||
*/
|
||||
export function resolveWorktreeCwd(repoPath: string, launchCwd: string): string {
|
||||
try {
|
||||
// Verify repoPath is a git root before comparing against its canonical
|
||||
// root. If getGitRoot returns a different path, repoPath is an arbitrary
|
||||
// subdirectory — skip both the linked-worktree guard and auto-detection
|
||||
// and fall through to the repoPath fallback.
|
||||
const repoGitRoot = getGitRoot(repoPath);
|
||||
const repoCanonical =
|
||||
repoGitRoot && tryRealpath(repoGitRoot) === tryRealpath(repoPath)
|
||||
? getCanonicalRepoRoot(repoPath)
|
||||
: null;
|
||||
|
||||
// Early exit: if repoPath is a linked worktree (differs from its canonical
|
||||
// main-checkout root), return it unchanged. Do NOT override it with the
|
||||
// server's launch directory — that would silently replace the explicitly-
|
||||
// resolved worktree index with the main checkout.
|
||||
//
|
||||
// getCanonicalRepoRoot returns the main-checkout path for both the checkout
|
||||
// and all linked worktrees:
|
||||
// repoPath === canonical → main checkout (auto-detect may fire below)
|
||||
// repoPath !== canonical → linked worktree (return as-is)
|
||||
if (repoCanonical && tryRealpath(repoPath) !== tryRealpath(repoCanonical)) {
|
||||
return repoPath;
|
||||
}
|
||||
|
||||
const launchGitRoot = getGitRoot(launchCwd);
|
||||
if (launchGitRoot) {
|
||||
// Normalise via realpathSync before comparing so macOS /var → /private/var
|
||||
// symlinks (and Windows 8.3 short names) don't create false mismatches.
|
||||
const realLaunch = tryRealpath(launchGitRoot);
|
||||
const realRepo = tryRealpath(repoPath);
|
||||
if (realLaunch !== realRepo) {
|
||||
const launchCanonical = getCanonicalRepoRoot(launchCwd);
|
||||
// Use tryRealpath on both canonical values for cross-platform safety.
|
||||
if (
|
||||
launchCanonical &&
|
||||
repoCanonical &&
|
||||
tryRealpath(launchCanonical) === tryRealpath(repoCanonical)
|
||||
) {
|
||||
return launchGitRoot;
|
||||
}
|
||||
}
|
||||
}
|
||||
} catch {
|
||||
// Best-effort; fall through to repoPath.
|
||||
}
|
||||
return repoPath;
|
||||
}
|
||||
|
||||
export class LocalBackend {
|
||||
private repos: Map<string, RepoHandle> = new Map();
|
||||
private contextCache: Map<string, CodebaseContext> = new Map();
|
||||
@@ -982,7 +1066,7 @@ export class LocalBackend {
|
||||
timing,
|
||||
...(!ftsUsed && {
|
||||
warning:
|
||||
'FTS indexes missing — keyword search degraded. Run: gitnexus analyze --force to rebuild indexes.',
|
||||
'FTS indexes missing — keyword search degraded. Run: gitnexus analyze --repair-fts (or gitnexus analyze --force) to rebuild indexes.',
|
||||
}),
|
||||
};
|
||||
}
|
||||
@@ -1218,31 +1302,41 @@ export class LocalBackend {
|
||||
}
|
||||
}
|
||||
|
||||
async executeCypher(repoName: string, query: string): Promise<any> {
|
||||
async executeCypher(
|
||||
repoName: string,
|
||||
query: string,
|
||||
params: Record<string, unknown> = {},
|
||||
): Promise<any> {
|
||||
const repo = await this.resolveRepo(repoName);
|
||||
return this.cypher(repo, { query });
|
||||
return this.cypher(repo, { query, params });
|
||||
}
|
||||
|
||||
private async cypher(repo: RepoHandle, params: { query: string }): Promise<any> {
|
||||
private async cypher(
|
||||
repo: RepoHandle,
|
||||
request: { query: string; params?: Record<string, unknown> },
|
||||
): Promise<any> {
|
||||
await this.ensureInitialized(repo.id);
|
||||
|
||||
if (!isLbugReady(repo.id)) {
|
||||
return { error: 'LadybugDB not ready. Index may be corrupted.' };
|
||||
}
|
||||
|
||||
// Block write operations (defense-in-depth — DB is already read-only)
|
||||
if (isWriteQuery(params.query)) {
|
||||
if (request.params !== undefined && !isValidQueryParams(request.params)) {
|
||||
return {
|
||||
error:
|
||||
'Write operations (CREATE, DELETE, SET, MERGE, REMOVE, DROP, ALTER, COPY, DETACH) are not allowed. The knowledge graph is read-only.',
|
||||
error: '"params" must be a plain object with scalar values (string/number/boolean/null).',
|
||||
};
|
||||
}
|
||||
|
||||
try {
|
||||
const result = await executeQuery(repo.id, params.query);
|
||||
const result = await executeParameterized(repo.id, request.query, request.params ?? {});
|
||||
return result;
|
||||
} catch (err: any) {
|
||||
const msg = err.message || 'Query failed';
|
||||
if (isReadOnlyDbError(err)) {
|
||||
return {
|
||||
error:
|
||||
'Write operations (CREATE, DELETE, SET, MERGE, REMOVE, DROP, ALTER, COPY, DETACH) are not allowed. The knowledge graph is read-only.',
|
||||
};
|
||||
}
|
||||
if (isWalCorruptionError(err)) {
|
||||
return {
|
||||
error: msg,
|
||||
@@ -2133,6 +2227,7 @@ export class LocalBackend {
|
||||
params: {
|
||||
scope?: string;
|
||||
base_ref?: string;
|
||||
worktree?: string;
|
||||
},
|
||||
): Promise<any> {
|
||||
await this.ensureInitialized(repo.id);
|
||||
@@ -2161,11 +2256,51 @@ export class LocalBackend {
|
||||
|
||||
let diffOutput: string;
|
||||
try {
|
||||
// Resolve the cwd for git diff.
|
||||
//
|
||||
// In a linked worktree (e.g. /repo/wt-feature/), the user's staged and
|
||||
// unstaged changes live in that worktree's separate working directory and
|
||||
// index. Running `git diff` from the canonical repo root sees a different
|
||||
// working tree and returns empty output.
|
||||
//
|
||||
// Resolution order (see resolveWorktreeCwd for details):
|
||||
// 1. params.worktree — explicit override, validated against the
|
||||
// registered repo's canonical root.
|
||||
// 2. Auto-detect — if the server's launch cwd (process.cwd()) is a
|
||||
// linked worktree of the same canonical repo, use its git root.
|
||||
// 3. repo.repoPath — fallback (original behaviour, handled inside
|
||||
// resolveWorktreeCwd when no worktree is detected).
|
||||
//
|
||||
// Start with the auto-detected value; override with the validated
|
||||
// explicit param when provided. This avoids a dead initial assignment.
|
||||
let diffCwd = resolveWorktreeCwd(repo.repoPath, process.cwd());
|
||||
if (params.worktree) {
|
||||
if (!path.isAbsolute(params.worktree)) {
|
||||
return {
|
||||
error: `worktree must be an absolute path, got: "${params.worktree}"`,
|
||||
};
|
||||
}
|
||||
const providedResolved = path.resolve(params.worktree);
|
||||
const repoCanonical = getCanonicalRepoRoot(repo.repoPath);
|
||||
if (!repoCanonical) {
|
||||
return {
|
||||
error: `Could not determine canonical root for repo "${repo.repoPath}". Is git available?`,
|
||||
};
|
||||
}
|
||||
const worktreeCanonical = getCanonicalRepoRoot(providedResolved);
|
||||
if (!worktreeCanonical || tryRealpath(worktreeCanonical) !== tryRealpath(repoCanonical)) {
|
||||
return {
|
||||
error: `worktree "${params.worktree}" is not a worktree of repo "${repo.repoPath}". Ensure the path is inside the same git repository.`,
|
||||
};
|
||||
}
|
||||
diffCwd = providedResolved;
|
||||
}
|
||||
|
||||
// maxBuffer raised from Node's 1MB default to 256MB to avoid ENOBUFS on
|
||||
// repos with large unstaged/untracked diffs (e.g. unignored build folders).
|
||||
// See issue: spawnSync git ENOBUFS in detect_changes(scope="unstaged").
|
||||
diffOutput = execFileSync('git', diffArgs, {
|
||||
cwd: repo.repoPath,
|
||||
cwd: diffCwd,
|
||||
encoding: 'utf-8',
|
||||
maxBuffer: 256 * 1024 * 1024,
|
||||
});
|
||||
|
||||
@@ -187,6 +187,11 @@ TIPS:
|
||||
type: 'object',
|
||||
properties: {
|
||||
query: { type: 'string', description: 'Cypher query to execute' },
|
||||
params: {
|
||||
type: 'object',
|
||||
description:
|
||||
'Optional query parameters for placeholders (e.g. $name) to execute via prepared statement binding.',
|
||||
},
|
||||
repo: {
|
||||
type: 'string',
|
||||
description: 'Repository name or path. Omit if only one repo is indexed.',
|
||||
@@ -253,6 +258,8 @@ Maps git diff hunks to indexed symbols, then traces which processes are impacted
|
||||
WHEN TO USE: Before committing — to understand what your changes affect. Pre-commit review, PR preparation.
|
||||
AFTER THIS: Review affected processes. Use context() on high-risk symbols. READ gitnexus://repo/{name}/process/{name} for full traces.
|
||||
|
||||
GIT WORKTREE SUPPORT: GitNexus automatically detects when the MCP server was launched from inside a linked git worktree and runs git diff against that worktree — no extra parameters needed in the common case. Pass "worktree" explicitly only when the server was started from a different directory than the worktree you are editing (e.g., the server runs from the canonical root but your changes are in a linked worktree at a different path).
|
||||
|
||||
Returns: changed symbols, affected processes, and a risk summary.`,
|
||||
annotations: READ_ONLY_TOOL_ANNOTATIONS,
|
||||
inputSchema: {
|
||||
@@ -268,6 +275,11 @@ Returns: changed symbols, affected processes, and a risk summary.`,
|
||||
type: 'string',
|
||||
description: 'Branch/commit for "compare" scope (e.g., "main")',
|
||||
},
|
||||
worktree: {
|
||||
type: 'string',
|
||||
description:
|
||||
'Absolute path to a linked git worktree. Pass this when your changes are in a worktree (the .git entry at that path is a file, not a directory). GitNexus will run git diff from that worktree so staged/unstaged changes are correctly detected.',
|
||||
},
|
||||
repo: {
|
||||
type: 'string',
|
||||
description: 'Repository name or path. Omit if only one repo is indexed.',
|
||||
|
||||
+164
-123
@@ -22,8 +22,9 @@ import {
|
||||
flushWAL,
|
||||
closeLbug,
|
||||
withLbugDb,
|
||||
isReadOnlyDbError,
|
||||
} from '../core/lbug/lbug-adapter.js';
|
||||
import { isWriteQuery } from '../core/lbug/pool-adapter.js';
|
||||
import { isValidQueryParams } from '../core/lbug/query-params.js';
|
||||
import { NODE_TABLES, type GraphNode, type GraphRelationship } from 'gitnexus-shared';
|
||||
import { searchFTSFromLbug } from '../core/search/bm25-index.js';
|
||||
import { hybridSearch } from '../core/search/hybrid-search.js';
|
||||
@@ -447,7 +448,14 @@ export const streamGraphNdjson = async (
|
||||
*/
|
||||
const mountSSEProgress = (app: express.Express, routePath: string, jm: JobManager) => {
|
||||
app.get(routePath, (req, res) => {
|
||||
const job = jm.getJob(req.params.jobId);
|
||||
let jobId: string;
|
||||
try {
|
||||
jobId = assertString(req.params.jobId, 'jobId');
|
||||
} catch (err: any) {
|
||||
res.status(err.status ?? 400).json({ error: err.message });
|
||||
return;
|
||||
}
|
||||
const job = jm.getJob(jobId);
|
||||
if (!job) {
|
||||
res.status(404).json({ error: 'Job not found' });
|
||||
return;
|
||||
@@ -493,7 +501,7 @@ const mountSSEProgress = (app: express.Express, routePath: string, jm: JobManage
|
||||
try {
|
||||
eventId++;
|
||||
if (progress.phase === 'complete' || progress.phase === 'failed') {
|
||||
const eventJob = jm.getJob(req.params.jobId);
|
||||
const eventJob = jm.getJob(jobId);
|
||||
res.write(
|
||||
`id: ${eventId}\nevent: ${progress.phase}\ndata: ${JSON.stringify({
|
||||
repoName: eventJob?.repoName,
|
||||
@@ -621,6 +629,44 @@ export const handleFileRequest = async (
|
||||
}
|
||||
};
|
||||
|
||||
export const handleQueryRequest = async (
|
||||
req: express.Request,
|
||||
res: express.Response,
|
||||
resolveRepo: (repoName?: string) => Promise<{ storagePath: string } | undefined>,
|
||||
): Promise<void> => {
|
||||
try {
|
||||
const cypher = req.body.cypher as string;
|
||||
if (!cypher) {
|
||||
res.status(400).json({ error: 'Missing "cypher" in request body' });
|
||||
return;
|
||||
}
|
||||
const queryParams = req.body.params;
|
||||
if (queryParams !== undefined && !isValidQueryParams(queryParams)) {
|
||||
res.status(400).json({
|
||||
error: '"params" must be a plain object with scalar values (string/number/boolean/null)',
|
||||
});
|
||||
return;
|
||||
}
|
||||
|
||||
const entry = await resolveRepo(requestedRepo(req));
|
||||
if (!entry) {
|
||||
res.status(404).json({ error: 'Repository not found' });
|
||||
return;
|
||||
}
|
||||
const lbugPath = path.join(entry.storagePath, 'lbug');
|
||||
const result = await withLbugDb(lbugPath, () => executePrepared(cypher, queryParams ?? {}), {
|
||||
readOnly: true,
|
||||
});
|
||||
res.json({ result });
|
||||
} catch (err: any) {
|
||||
if (isReadOnlyDbError(err)) {
|
||||
res.status(403).json({ error: 'Write queries are not allowed via the HTTP API' });
|
||||
return;
|
||||
}
|
||||
res.status(500).json({ error: err.message || 'Query failed' });
|
||||
}
|
||||
};
|
||||
|
||||
export const createServer = async (port: number, host: string = '127.0.0.1') => {
|
||||
const app = express();
|
||||
app.disable('x-powered-by');
|
||||
@@ -984,8 +1030,16 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
|
||||
res.once('close', abortStreaming);
|
||||
|
||||
try {
|
||||
await withLbugDb(lbugPath, async () =>
|
||||
streamGraphNdjson(res, includeContent, abortController.signal),
|
||||
// Read-only open: /api/graph never writes. Write-mode opens engage
|
||||
// LadybugDB's checkpoint machinery (`.shadow` sidecar), which on
|
||||
// Windows races with the OS file handle release and trips
|
||||
// "Cannot open file ... lbug.shadow - Error 2". See pool-adapter.ts
|
||||
// which already opens read-only for the same reason, and the
|
||||
// /api/query precedent in PR #1655.
|
||||
await withLbugDb(
|
||||
lbugPath,
|
||||
async () => streamGraphNdjson(res, includeContent, abortController.signal),
|
||||
{ readOnly: true },
|
||||
);
|
||||
if (!abortController.signal.aborted && !res.writableEnded) {
|
||||
res.end();
|
||||
@@ -998,7 +1052,9 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
|
||||
return;
|
||||
}
|
||||
|
||||
const graph = await withLbugDb(lbugPath, async () => buildGraph(includeContent));
|
||||
const graph = await withLbugDb(lbugPath, async () => buildGraph(includeContent), {
|
||||
readOnly: true,
|
||||
});
|
||||
res.json(graph);
|
||||
} catch (err: any) {
|
||||
if (err instanceof ClientDisconnectedError) {
|
||||
@@ -1020,29 +1076,7 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
|
||||
|
||||
// Execute Cypher query
|
||||
app.post('/api/query', async (req, res) => {
|
||||
try {
|
||||
const cypher = req.body.cypher as string;
|
||||
if (!cypher) {
|
||||
res.status(400).json({ error: 'Missing "cypher" in request body' });
|
||||
return;
|
||||
}
|
||||
|
||||
if (isWriteQuery(cypher)) {
|
||||
res.status(403).json({ error: 'Write queries are not allowed via the HTTP API' });
|
||||
return;
|
||||
}
|
||||
|
||||
const entry = await resolveRepo(requestedRepo(req));
|
||||
if (!entry) {
|
||||
res.status(404).json({ error: 'Repository not found' });
|
||||
return;
|
||||
}
|
||||
const lbugPath = path.join(entry.storagePath, 'lbug');
|
||||
const result = await withLbugDb(lbugPath, () => executeQuery(cypher));
|
||||
res.json({ result });
|
||||
} catch (err: any) {
|
||||
res.status(500).json({ error: err.message || 'Query failed' });
|
||||
}
|
||||
await handleQueryRequest(req, res, resolveRepo);
|
||||
});
|
||||
|
||||
// Search (supports mode: 'hybrid' | 'semantic' | 'bm25', and optional enrichment)
|
||||
@@ -1067,68 +1101,70 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
|
||||
const mode: string = req.body.mode ?? 'hybrid';
|
||||
const enrich: boolean = req.body.enrich !== false; // default true
|
||||
|
||||
const results = await withLbugDb(lbugPath, async () => {
|
||||
let searchResults: any[];
|
||||
let ftsAvailable: boolean | undefined;
|
||||
const results = await withLbugDb(
|
||||
lbugPath,
|
||||
async () => {
|
||||
let searchResults: any[];
|
||||
let ftsAvailable: boolean | undefined;
|
||||
|
||||
if (mode === 'semantic') {
|
||||
const { isEmbedderReady } = await import('../core/embeddings/embedder.js');
|
||||
if (!isEmbedderReady()) {
|
||||
return { searchResults: [] as any[], ftsAvailable: undefined };
|
||||
}
|
||||
const { semanticSearch: semSearch } =
|
||||
await import('../core/embeddings/embedding-pipeline.js');
|
||||
searchResults = await semSearch(executeQuery, query, limit);
|
||||
// Normalize semantic results to HybridSearchResult shape
|
||||
searchResults = searchResults.map((r: any, i: number) => ({
|
||||
...r,
|
||||
score: r.score ?? 1 - (r.distance ?? 0),
|
||||
rank: i + 1,
|
||||
sources: ['semantic'],
|
||||
}));
|
||||
} else if (mode === 'bm25') {
|
||||
const ftsResponse = await searchFTSFromLbug(query, limit);
|
||||
ftsAvailable = ftsResponse.ftsAvailable;
|
||||
searchResults = ftsResponse.results.map((r: any, i: number) => ({
|
||||
...r,
|
||||
rank: i + 1,
|
||||
sources: ['bm25'],
|
||||
}));
|
||||
} else {
|
||||
// hybrid (default)
|
||||
const { isEmbedderReady } = await import('../core/embeddings/embedder.js');
|
||||
if (isEmbedderReady()) {
|
||||
if (mode === 'semantic') {
|
||||
const { isEmbedderReady } = await import('../core/embeddings/embedder.js');
|
||||
if (!isEmbedderReady()) {
|
||||
return { searchResults: [] as any[], ftsAvailable: undefined };
|
||||
}
|
||||
const { semanticSearch: semSearch } =
|
||||
await import('../core/embeddings/embedding-pipeline.js');
|
||||
searchResults = await hybridSearch(query, limit, executeQuery, semSearch);
|
||||
} else {
|
||||
searchResults = await semSearch(executeQuery, query, limit);
|
||||
// Normalize semantic results to HybridSearchResult shape
|
||||
searchResults = searchResults.map((r: any, i: number) => ({
|
||||
...r,
|
||||
score: r.score ?? 1 - (r.distance ?? 0),
|
||||
rank: i + 1,
|
||||
sources: ['semantic'],
|
||||
}));
|
||||
} else if (mode === 'bm25') {
|
||||
const ftsResponse = await searchFTSFromLbug(query, limit);
|
||||
ftsAvailable = ftsResponse.ftsAvailable;
|
||||
searchResults = ftsResponse.results;
|
||||
searchResults = ftsResponse.results.map((r: any, i: number) => ({
|
||||
...r,
|
||||
rank: i + 1,
|
||||
sources: ['bm25'],
|
||||
}));
|
||||
} else {
|
||||
// hybrid (default)
|
||||
const { isEmbedderReady } = await import('../core/embeddings/embedder.js');
|
||||
if (isEmbedderReady()) {
|
||||
const { semanticSearch: semSearch } =
|
||||
await import('../core/embeddings/embedding-pipeline.js');
|
||||
searchResults = await hybridSearch(query, limit, executeQuery, semSearch);
|
||||
} else {
|
||||
const ftsResponse = await searchFTSFromLbug(query, limit);
|
||||
ftsAvailable = ftsResponse.ftsAvailable;
|
||||
searchResults = ftsResponse.results;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (!enrich) return { searchResults, ftsAvailable };
|
||||
if (!enrich) return { searchResults, ftsAvailable };
|
||||
|
||||
// Server-side enrichment: add connections, cluster, processes per result
|
||||
// Uses parameterized queries to prevent Cypher injection via nodeId
|
||||
const validLabel = (label: string): boolean =>
|
||||
(NODE_TABLES as readonly string[]).includes(label);
|
||||
// Server-side enrichment: add connections, cluster, processes per result
|
||||
// Uses parameterized queries to prevent Cypher injection via nodeId
|
||||
const validLabel = (label: string): boolean =>
|
||||
(NODE_TABLES as readonly string[]).includes(label);
|
||||
|
||||
const enriched = await Promise.all(
|
||||
searchResults.slice(0, limit).map(async (r: any) => {
|
||||
const nodeId: string = r.nodeId || r.id || '';
|
||||
const nodeLabel = nodeId.split(':')[0];
|
||||
const enrichment: { connections?: any; cluster?: string; processes?: any[] } = {};
|
||||
const enriched = await Promise.all(
|
||||
searchResults.slice(0, limit).map(async (r: any) => {
|
||||
const nodeId: string = r.nodeId || r.id || '';
|
||||
const nodeLabel = nodeId.split(':')[0];
|
||||
const enrichment: { connections?: any; cluster?: string; processes?: any[] } = {};
|
||||
|
||||
if (!nodeId || !validLabel(nodeLabel)) return { ...r, ...enrichment };
|
||||
if (!nodeId || !validLabel(nodeLabel)) return { ...r, ...enrichment };
|
||||
|
||||
// Run connections, cluster, and process queries in parallel
|
||||
// Label is validated against NODE_TABLES (compile-time safe identifiers);
|
||||
// nodeId uses $nid parameter binding to prevent injection
|
||||
const [connRes, clusterRes, procRes] = await Promise.all([
|
||||
executePrepared(
|
||||
`
|
||||
// Run connections, cluster, and process queries in parallel
|
||||
// Label is validated against NODE_TABLES (compile-time safe identifiers);
|
||||
// nodeId uses $nid parameter binding to prevent injection
|
||||
const [connRes, clusterRes, procRes] = await Promise.all([
|
||||
executePrepared(
|
||||
`
|
||||
MATCH (n:${nodeLabel} {id: $nid})
|
||||
OPTIONAL MATCH (n)-[r1:CodeRelation]->(dst)
|
||||
OPTIONAL MATCH (src)-[r2:CodeRelation]->(n)
|
||||
@@ -1137,65 +1173,67 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
|
||||
collect(DISTINCT {name: src.name, type: r2.type, confidence: r2.confidence}) AS incoming
|
||||
LIMIT 1
|
||||
`,
|
||||
{ nid: nodeId },
|
||||
).catch(() => []),
|
||||
executePrepared(
|
||||
`
|
||||
{ nid: nodeId },
|
||||
).catch(() => []),
|
||||
executePrepared(
|
||||
`
|
||||
MATCH (n:${nodeLabel} {id: $nid})
|
||||
MATCH (n)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
|
||||
RETURN c.label AS label, c.description AS description
|
||||
LIMIT 1
|
||||
`,
|
||||
{ nid: nodeId },
|
||||
).catch(() => []),
|
||||
executePrepared(
|
||||
`
|
||||
{ nid: nodeId },
|
||||
).catch(() => []),
|
||||
executePrepared(
|
||||
`
|
||||
MATCH (n:${nodeLabel} {id: $nid})
|
||||
MATCH (n)-[rel:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
|
||||
RETURN p.id AS id, p.label AS label, rel.step AS step, p.stepCount AS stepCount
|
||||
ORDER BY rel.step
|
||||
`,
|
||||
{ nid: nodeId },
|
||||
).catch(() => []),
|
||||
]);
|
||||
{ nid: nodeId },
|
||||
).catch(() => []),
|
||||
]);
|
||||
|
||||
if (connRes.length > 0) {
|
||||
const row = connRes[0];
|
||||
const outgoing = (Array.isArray(row) ? row[0] : row.outgoing || [])
|
||||
.filter((c: any) => c?.name)
|
||||
.slice(0, 5);
|
||||
const incoming = (Array.isArray(row) ? row[1] : row.incoming || [])
|
||||
.filter((c: any) => c?.name)
|
||||
.slice(0, 5);
|
||||
enrichment.connections = { outgoing, incoming };
|
||||
}
|
||||
if (connRes.length > 0) {
|
||||
const row = connRes[0];
|
||||
const outgoing = (Array.isArray(row) ? row[0] : row.outgoing || [])
|
||||
.filter((c: any) => c?.name)
|
||||
.slice(0, 5);
|
||||
const incoming = (Array.isArray(row) ? row[1] : row.incoming || [])
|
||||
.filter((c: any) => c?.name)
|
||||
.slice(0, 5);
|
||||
enrichment.connections = { outgoing, incoming };
|
||||
}
|
||||
|
||||
if (clusterRes.length > 0) {
|
||||
const row = clusterRes[0];
|
||||
enrichment.cluster = Array.isArray(row) ? row[0] : row.label;
|
||||
}
|
||||
if (clusterRes.length > 0) {
|
||||
const row = clusterRes[0];
|
||||
enrichment.cluster = Array.isArray(row) ? row[0] : row.label;
|
||||
}
|
||||
|
||||
if (procRes.length > 0) {
|
||||
enrichment.processes = procRes
|
||||
.map((row: any) => ({
|
||||
id: Array.isArray(row) ? row[0] : row.id,
|
||||
label: Array.isArray(row) ? row[1] : row.label,
|
||||
step: Array.isArray(row) ? row[2] : row.step,
|
||||
stepCount: Array.isArray(row) ? row[3] : row.stepCount,
|
||||
}))
|
||||
.filter((p: any) => p.id && p.label);
|
||||
}
|
||||
if (procRes.length > 0) {
|
||||
enrichment.processes = procRes
|
||||
.map((row: any) => ({
|
||||
id: Array.isArray(row) ? row[0] : row.id,
|
||||
label: Array.isArray(row) ? row[1] : row.label,
|
||||
step: Array.isArray(row) ? row[2] : row.step,
|
||||
stepCount: Array.isArray(row) ? row[3] : row.stepCount,
|
||||
}))
|
||||
.filter((p: any) => p.id && p.label);
|
||||
}
|
||||
|
||||
return { ...r, ...enrichment };
|
||||
}),
|
||||
);
|
||||
return { ...r, ...enrichment };
|
||||
}),
|
||||
);
|
||||
|
||||
return { searchResults: enriched, ftsAvailable };
|
||||
});
|
||||
return { searchResults: enriched, ftsAvailable };
|
||||
},
|
||||
{ readOnly: true },
|
||||
);
|
||||
const response: any = { results: results.searchResults ?? results };
|
||||
if (results.ftsAvailable === false) {
|
||||
response.warning =
|
||||
'FTS indexes missing — keyword search degraded. Run: gitnexus analyze --force to rebuild indexes.';
|
||||
'FTS indexes missing — keyword search degraded. Run: gitnexus analyze --repair-fts (or gitnexus analyze --force) to rebuild indexes.';
|
||||
}
|
||||
res.json(response);
|
||||
} catch (err: any) {
|
||||
@@ -1271,8 +1309,11 @@ export const createServer = async (port: number, host: string = '127.0.0.1') =>
|
||||
|
||||
// Get file paths from the graph (lightweight — no content loaded)
|
||||
const lbugPath = path.join(entry.storagePath, 'lbug');
|
||||
const fileRows = await withLbugDb(lbugPath, () =>
|
||||
executeQuery(`MATCH (n:File) WHERE n.content IS NOT NULL RETURN n.filePath AS filePath`),
|
||||
const fileRows = await withLbugDb(
|
||||
lbugPath,
|
||||
() =>
|
||||
executeQuery(`MATCH (n:File) WHERE n.content IS NOT NULL RETURN n.filePath AS filePath`),
|
||||
{ readOnly: true },
|
||||
);
|
||||
|
||||
// Search files on disk one at a time (constant memory)
|
||||
|
||||
+1
@@ -2,6 +2,7 @@ export class User {
|
||||
save() {}
|
||||
}
|
||||
|
||||
/** @returns {User} */
|
||||
export function getUser() {
|
||||
return new User();
|
||||
}
|
||||
|
||||
@@ -3,6 +3,7 @@ export class User {
|
||||
getName() { return ''; }
|
||||
}
|
||||
|
||||
/** @returns {User} */
|
||||
export function getUser() {
|
||||
return new User();
|
||||
}
|
||||
|
||||
+12
@@ -0,0 +1,12 @@
|
||||
#include "lib.h"
|
||||
|
||||
void Service::f(int* p) {}
|
||||
void Service::f(bool flag) {}
|
||||
|
||||
void Service::g(int a, int b) {}
|
||||
void Service::g(int a, ...) {}
|
||||
|
||||
void Service::h(int a, double b) {}
|
||||
void Service::h(int a, ...) {}
|
||||
|
||||
void Service::k(int a, ...) {}
|
||||
+38
@@ -0,0 +1,38 @@
|
||||
#pragma once
|
||||
|
||||
class Service {
|
||||
public:
|
||||
void f(int* p);
|
||||
void f(bool flag);
|
||||
|
||||
void g(int a, int b);
|
||||
void g(int a, ...);
|
||||
|
||||
void h(int a, double b);
|
||||
void h(int a, ...);
|
||||
|
||||
void k(int a, ...);
|
||||
|
||||
void runNullptr() {
|
||||
f(nullptr);
|
||||
}
|
||||
|
||||
void runPointer() {
|
||||
int* p = nullptr;
|
||||
f(p);
|
||||
}
|
||||
|
||||
void runBoolConversion() {
|
||||
f(42);
|
||||
}
|
||||
|
||||
void run() {
|
||||
int* p = nullptr;
|
||||
f(nullptr);
|
||||
f(p);
|
||||
f(42);
|
||||
g(1, 2);
|
||||
h(1, 'a');
|
||||
k(1, 2, 3);
|
||||
}
|
||||
};
|
||||
@@ -0,0 +1,16 @@
|
||||
#include <type_traits>
|
||||
|
||||
struct S {};
|
||||
|
||||
template <class T, std::enable_if_t<std::is_class_v<T>, int> = 0>
|
||||
void pick(T value) {}
|
||||
|
||||
template <class T, std::enable_if_t<std::is_integral_v<T>, int> = 0>
|
||||
void pick(T value) {}
|
||||
|
||||
void run() {
|
||||
S s;
|
||||
int n = 0;
|
||||
pick(s);
|
||||
pick(n);
|
||||
}
|
||||
@@ -0,0 +1,14 @@
|
||||
#include <type_traits>
|
||||
|
||||
template <class T, std::enable_if_t<std::is_const_v<T>, int> = 0>
|
||||
void pick(T value) {}
|
||||
|
||||
template <class T, std::enable_if_t<std::is_volatile_v<T>, int> = 0>
|
||||
void pick(T value) {}
|
||||
|
||||
void run() {
|
||||
const int c = 0;
|
||||
volatile int v = 0;
|
||||
pick(c);
|
||||
pick(v);
|
||||
}
|
||||
@@ -0,0 +1,16 @@
|
||||
#include <type_traits>
|
||||
|
||||
enum Color { Red };
|
||||
|
||||
template <class T, std::enable_if_t<std::is_enum_v<T>, int> = 0>
|
||||
void pick(T value) {}
|
||||
|
||||
template <class T, std::enable_if_t<std::is_integral_v<T>, int> = 0>
|
||||
void pick(T value) {}
|
||||
|
||||
void run() {
|
||||
Color color = Red;
|
||||
int n = 0;
|
||||
pick(color);
|
||||
pick(n);
|
||||
}
|
||||
@@ -0,0 +1,14 @@
|
||||
#include <type_traits>
|
||||
|
||||
struct S {};
|
||||
|
||||
template <class T, std::enable_if_t<std::is_pointer_v<T>, int> = 0>
|
||||
void pick(T value) {}
|
||||
|
||||
template <class T, std::enable_if_t<std::is_class_v<T>, int> = 0>
|
||||
void pick(T value) {}
|
||||
|
||||
void run(S* p, S s) {
|
||||
pick(p);
|
||||
pick(s);
|
||||
}
|
||||
@@ -0,0 +1,14 @@
|
||||
#include <type_traits>
|
||||
|
||||
template <class T, std::enable_if_t<std::is_reference_v<T>, int> = 0>
|
||||
void pick(T value) {}
|
||||
|
||||
template <class T, std::enable_if_t<std::is_integral_v<T>, int> = 0>
|
||||
void pick(T value) {}
|
||||
|
||||
void run() {
|
||||
int n = 0;
|
||||
int& r = n;
|
||||
pick(r);
|
||||
pick(n);
|
||||
}
|
||||
@@ -0,0 +1,12 @@
|
||||
#include <type_traits>
|
||||
|
||||
template <class T, std::enable_if_t<std::is_void_v<T>, int> = 0>
|
||||
void pick(T value) {}
|
||||
|
||||
template <class T, std::enable_if_t<std::is_pointer_v<T>, int> = 0>
|
||||
void pick(T value) {}
|
||||
|
||||
void run() {
|
||||
void* p;
|
||||
pick(p);
|
||||
}
|
||||
+8
@@ -0,0 +1,8 @@
|
||||
package com.example;
|
||||
|
||||
public class Module1App {
|
||||
public void run() {
|
||||
UserService service = new UserService();
|
||||
service.ping();
|
||||
}
|
||||
}
|
||||
+6
@@ -0,0 +1,6 @@
|
||||
package com.example;
|
||||
|
||||
public class UserService {
|
||||
public void ping() {
|
||||
}
|
||||
}
|
||||
+8
@@ -0,0 +1,8 @@
|
||||
package com.example;
|
||||
|
||||
public class Module2App {
|
||||
public void run() {
|
||||
UserService service = new UserService();
|
||||
service.ping();
|
||||
}
|
||||
}
|
||||
+6
@@ -0,0 +1,6 @@
|
||||
package com.example;
|
||||
|
||||
public class UserService {
|
||||
public void ping() {
|
||||
}
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user