Compare commits

..
Author SHA1 Message Date
abhigyantrumioandClaude Opus 4.6 3ea8e9e88a fix(ingestion): align CALLS edge sourceId with node ID format
findEnclosingFunctionId generated IDs without :startLine suffix,
but node creation includes it. This caused every CALLS edge to
reference a non-existent source node, making the process detector
find 0 entry points and produce 0 execution flows.

Bumps to 1.3.9.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 13:55:42 +05:30
426 changed files with 3139 additions and 26296 deletions
@@ -22,7 +22,7 @@ Run from the project root. This parses all source files, builds the knowledge gr
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook runs `analyze` automatically after `git commit` and `git merge`, preserving embeddings if previously generated.
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale.
### status — Check index freshness
@@ -1,89 +1,89 @@
---
name: gitnexus-debugging
description: "Use when the user is debugging a bug, tracing an error, or asking why something fails. Examples: \"Why is X failing?\", \"Where does this error come from?\", \"Trace this bug\""
---
# Debugging with GitNexus
## When to Use
- "Why is this function failing?"
- "Trace where this error comes from"
- "Who calls this method?"
- "This endpoint returns 500"
- Investigating bugs, errors, or unexpected behavior
## Workflow
```
1. gitnexus_query({query: "<error or symptom>"}) → Find related execution flows
2. gitnexus_context({name: "<suspect>"}) → See callers/callees/processes
3. READ gitnexus://repo/{name}/process/{name} → Trace execution flow
4. gitnexus_cypher({query: "MATCH path..."}) → Custom traces if needed
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] Understand the symptom (error message, unexpected behavior)
- [ ] gitnexus_query for error text or related code
- [ ] Identify the suspect function from returned processes
- [ ] gitnexus_context to see callers and callees
- [ ] Trace execution flow via process resource if applicable
- [ ] gitnexus_cypher for custom call chain traces if needed
- [ ] Read source files to confirm root cause
```
## Debugging Patterns
| Symptom | GitNexus Approach |
| -------------------- | ---------------------------------------------------------- |
| Error message | `gitnexus_query` for error text → `context` on throw sites |
| Wrong return value | `context` on the function → trace callees for data flow |
| Intermittent failure | `context` → look for external calls, async deps |
| Performance issue | `context` → find symbols with many callers (hot paths) |
| Recent regression | `detect_changes` to see what your changes affect |
## Tools
**gitnexus_query** — find code related to error:
```
gitnexus_query({query: "payment validation error"})
→ Processes: CheckoutFlow, ErrorHandling
→ Symbols: validatePayment, handlePaymentError, PaymentException
```
**gitnexus_context** — full context for a suspect:
```
gitnexus_context({name: "validatePayment"})
→ Incoming calls: processCheckout, webhookHandler
→ Outgoing calls: verifyCard, fetchRates (external API!)
→ Processes: CheckoutFlow (step 3/7)
```
**gitnexus_cypher** — custom call chain traces:
```cypher
MATCH path = (a)-[:CodeRelation {type: 'CALLS'}*1..2]->(b:Function {name: "validatePayment"})
RETURN [n IN nodes(path) | n.name] AS chain
```
## Example: "Payment endpoint returns 500 intermittently"
```
1. gitnexus_query({query: "payment error handling"})
→ Processes: CheckoutFlow, ErrorHandling
→ Symbols: validatePayment, handlePaymentError
2. gitnexus_context({name: "validatePayment"})
→ Outgoing calls: verifyCard, fetchRates (external API!)
3. READ gitnexus://repo/my-app/process/CheckoutFlow
→ Step 3: validatePayment → calls fetchRates (external)
4. Root cause: fetchRates calls external API without proper timeout
```
---
name: gitnexus-debugging
description: "Use when the user is debugging a bug, tracing an error, or asking why something fails. Examples: \"Why is X failing?\", \"Where does this error come from?\", \"Trace this bug\""
---
# Debugging with GitNexus
## When to Use
- "Why is this function failing?"
- "Trace where this error comes from"
- "Who calls this method?"
- "This endpoint returns 500"
- Investigating bugs, errors, or unexpected behavior
## Workflow
```
1. gitnexus_query({query: "<error or symptom>"}) → Find related execution flows
2. gitnexus_context({name: "<suspect>"}) → See callers/callees/processes
3. READ gitnexus://repo/{name}/process/{name} → Trace execution flow
4. gitnexus_cypher({query: "MATCH path..."}) → Custom traces if needed
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] Understand the symptom (error message, unexpected behavior)
- [ ] gitnexus_query for error text or related code
- [ ] Identify the suspect function from returned processes
- [ ] gitnexus_context to see callers and callees
- [ ] Trace execution flow via process resource if applicable
- [ ] gitnexus_cypher for custom call chain traces if needed
- [ ] Read source files to confirm root cause
```
## Debugging Patterns
| Symptom | GitNexus Approach |
| -------------------- | ---------------------------------------------------------- |
| Error message | `gitnexus_query` for error text → `context` on throw sites |
| Wrong return value | `context` on the function → trace callees for data flow |
| Intermittent failure | `context` → look for external calls, async deps |
| Performance issue | `context` → find symbols with many callers (hot paths) |
| Recent regression | `detect_changes` to see what your changes affect |
## Tools
**gitnexus_query** — find code related to error:
```
gitnexus_query({query: "payment validation error"})
→ Processes: CheckoutFlow, ErrorHandling
→ Symbols: validatePayment, handlePaymentError, PaymentException
```
**gitnexus_context** — full context for a suspect:
```
gitnexus_context({name: "validatePayment"})
→ Incoming calls: processCheckout, webhookHandler
→ Outgoing calls: verifyCard, fetchRates (external API!)
→ Processes: CheckoutFlow (step 3/7)
```
**gitnexus_cypher** — custom call chain traces:
```cypher
MATCH path = (a)-[:CodeRelation {type: 'CALLS'}*1..2]->(b:Function {name: "validatePayment"})
RETURN [n IN nodes(path) | n.name] AS chain
```
## Example: "Payment endpoint returns 500 intermittently"
```
1. gitnexus_query({query: "payment error handling"})
→ Processes: CheckoutFlow, ErrorHandling
→ Symbols: validatePayment, handlePaymentError
2. gitnexus_context({name: "validatePayment"})
→ Outgoing calls: verifyCard, fetchRates (external API!)
3. READ gitnexus://repo/my-app/process/CheckoutFlow
→ Step 3: validatePayment → calls fetchRates (external)
4. Root cause: fetchRates calls external API without proper timeout
```
@@ -1,78 +1,78 @@
---
name: gitnexus-exploring
description: "Use when the user asks how code works, wants to understand architecture, trace execution flows, or explore unfamiliar parts of the codebase. Examples: \"How does X work?\", \"What calls this function?\", \"Show me the auth flow\""
---
# Exploring Codebases with GitNexus
## When to Use
- "How does authentication work?"
- "What's the project structure?"
- "Show me the main components"
- "Where is the database logic?"
- Understanding code you haven't seen before
## Workflow
```
1. READ gitnexus://repos → Discover indexed repos
2. READ gitnexus://repo/{name}/context → Codebase overview, check staleness
3. gitnexus_query({query: "<what you want to understand>"}) → Find related execution flows
4. gitnexus_context({name: "<symbol>"}) → Deep dive on specific symbol
5. READ gitnexus://repo/{name}/process/{name} → Trace full execution flow
```
> If step 2 says "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] READ gitnexus://repo/{name}/context
- [ ] gitnexus_query for the concept you want to understand
- [ ] Review returned processes (execution flows)
- [ ] gitnexus_context on key symbols for callers/callees
- [ ] READ process resource for full execution traces
- [ ] Read source files for implementation details
```
## Resources
| Resource | What you get |
| --------------------------------------- | ------------------------------------------------------- |
| `gitnexus://repo/{name}/context` | Stats, staleness warning (~150 tokens) |
| `gitnexus://repo/{name}/clusters` | All functional areas with cohesion scores (~300 tokens) |
| `gitnexus://repo/{name}/cluster/{name}` | Area members with file paths (~500 tokens) |
| `gitnexus://repo/{name}/process/{name}` | Step-by-step execution trace (~200 tokens) |
## Tools
**gitnexus_query** — find execution flows related to a concept:
```
gitnexus_query({query: "payment processing"})
→ Processes: CheckoutFlow, RefundFlow, WebhookHandler
→ Symbols grouped by flow with file locations
```
**gitnexus_context** — 360-degree view of a symbol:
```
gitnexus_context({name: "validateUser"})
→ Incoming calls: loginHandler, apiMiddleware
→ Outgoing calls: checkToken, getUserById
→ Processes: LoginFlow (step 2/5), TokenRefresh (step 1/3)
```
## Example: "How does payment processing work?"
```
1. READ gitnexus://repo/my-app/context → 918 symbols, 45 processes
2. gitnexus_query({query: "payment processing"})
→ CheckoutFlow: processPayment → validateCard → chargeStripe
→ RefundFlow: initiateRefund → calculateRefund → processRefund
3. gitnexus_context({name: "processPayment"})
→ Incoming: checkoutHandler, webhookHandler
→ Outgoing: validateCard, chargeStripe, saveTransaction
4. Read src/payments/processor.ts for implementation details
```
---
name: gitnexus-exploring
description: "Use when the user asks how code works, wants to understand architecture, trace execution flows, or explore unfamiliar parts of the codebase. Examples: \"How does X work?\", \"What calls this function?\", \"Show me the auth flow\""
---
# Exploring Codebases with GitNexus
## When to Use
- "How does authentication work?"
- "What's the project structure?"
- "Show me the main components"
- "Where is the database logic?"
- Understanding code you haven't seen before
## Workflow
```
1. READ gitnexus://repos → Discover indexed repos
2. READ gitnexus://repo/{name}/context → Codebase overview, check staleness
3. gitnexus_query({query: "<what you want to understand>"}) → Find related execution flows
4. gitnexus_context({name: "<symbol>"}) → Deep dive on specific symbol
5. READ gitnexus://repo/{name}/process/{name} → Trace full execution flow
```
> If step 2 says "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] READ gitnexus://repo/{name}/context
- [ ] gitnexus_query for the concept you want to understand
- [ ] Review returned processes (execution flows)
- [ ] gitnexus_context on key symbols for callers/callees
- [ ] READ process resource for full execution traces
- [ ] Read source files for implementation details
```
## Resources
| Resource | What you get |
| --------------------------------------- | ------------------------------------------------------- |
| `gitnexus://repo/{name}/context` | Stats, staleness warning (~150 tokens) |
| `gitnexus://repo/{name}/clusters` | All functional areas with cohesion scores (~300 tokens) |
| `gitnexus://repo/{name}/cluster/{name}` | Area members with file paths (~500 tokens) |
| `gitnexus://repo/{name}/process/{name}` | Step-by-step execution trace (~200 tokens) |
## Tools
**gitnexus_query** — find execution flows related to a concept:
```
gitnexus_query({query: "payment processing"})
→ Processes: CheckoutFlow, RefundFlow, WebhookHandler
→ Symbols grouped by flow with file locations
```
**gitnexus_context** — 360-degree view of a symbol:
```
gitnexus_context({name: "validateUser"})
→ Incoming calls: loginHandler, apiMiddleware
→ Outgoing calls: checkToken, getUserById
→ Processes: LoginFlow (step 2/5), TokenRefresh (step 1/3)
```
## Example: "How does payment processing work?"
```
1. READ gitnexus://repo/my-app/context → 918 symbols, 45 processes
2. gitnexus_query({query: "payment processing"})
→ CheckoutFlow: processPayment → validateCard → chargeStripe
→ RefundFlow: initiateRefund → calculateRefund → processRefund
3. gitnexus_context({name: "processPayment"})
→ Incoming: checkoutHandler, webhookHandler
→ Outgoing: validateCard, chargeStripe, saveTransaction
4. Read src/payments/processor.ts for implementation details
```
@@ -1,97 +1,97 @@
---
name: gitnexus-impact-analysis
description: "Use when the user wants to know what will break if they change something, or needs safety analysis before editing code. Examples: \"Is it safe to change X?\", \"What depends on this?\", \"What will break?\""
---
# Impact Analysis with GitNexus
## When to Use
- "Is it safe to change this function?"
- "What will break if I modify X?"
- "Show me the blast radius"
- "Who uses this code?"
- Before making non-trivial code changes
- Before committing — to understand what your changes affect
## Workflow
```
1. gitnexus_impact({target: "X", direction: "upstream"}) → What depends on this
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
3. gitnexus_detect_changes() → Map current git changes to affected flows
4. Assess risk and report to user
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] gitnexus_impact({target, direction: "upstream"}) to find dependents
- [ ] Review d=1 items first (these WILL BREAK)
- [ ] Check high-confidence (>0.8) dependencies
- [ ] READ processes to check affected execution flows
- [ ] gitnexus_detect_changes() for pre-commit check
- [ ] Assess risk level and report to user
```
## Understanding Output
| Depth | Risk Level | Meaning |
| ----- | ---------------- | ------------------------ |
| d=1 | **WILL BREAK** | Direct callers/importers |
| d=2 | LIKELY AFFECTED | Indirect dependencies |
| d=3 | MAY NEED TESTING | Transitive effects |
## Risk Assessment
| Affected | Risk |
| ------------------------------ | -------- |
| <5 symbols, few processes | LOW |
| 5-15 symbols, 2-5 processes | MEDIUM |
| >15 symbols or many processes | HIGH |
| Critical path (auth, payments) | CRITICAL |
## Tools
**gitnexus_impact** — the primary tool for symbol blast radius:
```
gitnexus_impact({
target: "validateUser",
direction: "upstream",
minConfidence: 0.8,
maxDepth: 3
})
→ d=1 (WILL BREAK):
- loginHandler (src/auth/login.ts:42) [CALLS, 100%]
- apiMiddleware (src/api/middleware.ts:15) [CALLS, 100%]
→ d=2 (LIKELY AFFECTED):
- authRouter (src/routes/auth.ts:22) [CALLS, 95%]
```
**gitnexus_detect_changes** — git-diff based impact analysis:
```
gitnexus_detect_changes({scope: "staged"})
→ Changed: 5 symbols in 3 files
→ Affected: LoginFlow, TokenRefresh, APIMiddlewarePipeline
→ Risk: MEDIUM
```
## Example: "What breaks if I change validateUser?"
```
1. gitnexus_impact({target: "validateUser", direction: "upstream"})
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
2. READ gitnexus://repo/my-app/processes
→ LoginFlow and TokenRefresh touch validateUser
3. Risk: 2 direct callers, 2 processes = MEDIUM
```
---
name: gitnexus-impact-analysis
description: "Use when the user wants to know what will break if they change something, or needs safety analysis before editing code. Examples: \"Is it safe to change X?\", \"What depends on this?\", \"What will break?\""
---
# Impact Analysis with GitNexus
## When to Use
- "Is it safe to change this function?"
- "What will break if I modify X?"
- "Show me the blast radius"
- "Who uses this code?"
- Before making non-trivial code changes
- Before committing — to understand what your changes affect
## Workflow
```
1. gitnexus_impact({target: "X", direction: "upstream"}) → What depends on this
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
3. gitnexus_detect_changes() → Map current git changes to affected flows
4. Assess risk and report to user
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] gitnexus_impact({target, direction: "upstream"}) to find dependents
- [ ] Review d=1 items first (these WILL BREAK)
- [ ] Check high-confidence (>0.8) dependencies
- [ ] READ processes to check affected execution flows
- [ ] gitnexus_detect_changes() for pre-commit check
- [ ] Assess risk level and report to user
```
## Understanding Output
| Depth | Risk Level | Meaning |
| ----- | ---------------- | ------------------------ |
| d=1 | **WILL BREAK** | Direct callers/importers |
| d=2 | LIKELY AFFECTED | Indirect dependencies |
| d=3 | MAY NEED TESTING | Transitive effects |
## Risk Assessment
| Affected | Risk |
| ------------------------------ | -------- |
| <5 symbols, few processes | LOW |
| 5-15 symbols, 2-5 processes | MEDIUM |
| >15 symbols or many processes | HIGH |
| Critical path (auth, payments) | CRITICAL |
## Tools
**gitnexus_impact** — the primary tool for symbol blast radius:
```
gitnexus_impact({
target: "validateUser",
direction: "upstream",
minConfidence: 0.8,
maxDepth: 3
})
→ d=1 (WILL BREAK):
- loginHandler (src/auth/login.ts:42) [CALLS, 100%]
- apiMiddleware (src/api/middleware.ts:15) [CALLS, 100%]
→ d=2 (LIKELY AFFECTED):
- authRouter (src/routes/auth.ts:22) [CALLS, 95%]
```
**gitnexus_detect_changes** — git-diff based impact analysis:
```
gitnexus_detect_changes({scope: "staged"})
→ Changed: 5 symbols in 3 files
→ Affected: LoginFlow, TokenRefresh, APIMiddlewarePipeline
→ Risk: MEDIUM
```
## Example: "What breaks if I change validateUser?"
```
1. gitnexus_impact({target: "validateUser", direction: "upstream"})
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
2. READ gitnexus://repo/my-app/processes
→ LoginFlow and TokenRefresh touch validateUser
3. Risk: 2 direct callers, 2 processes = MEDIUM
```
@@ -1,121 +1,121 @@
---
name: gitnexus-refactoring
description: "Use when the user wants to rename, extract, split, move, or restructure code safely. Examples: \"Rename this function\", \"Extract this into a module\", \"Refactor this class\", \"Move this to a separate file\""
---
# Refactoring with GitNexus
## When to Use
- "Rename this function safely"
- "Extract this into a module"
- "Split this service"
- "Move this to a new file"
- Any task involving renaming, extracting, splitting, or restructuring code
## Workflow
```
1. gitnexus_impact({target: "X", direction: "upstream"}) → Map all dependents
2. gitnexus_query({query: "X"}) → Find execution flows involving X
3. gitnexus_context({name: "X"}) → See all incoming/outgoing refs
4. Plan update order: interfaces → implementations → callers → tests
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklists
### Rename Symbol
```
- [ ] gitnexus_rename({symbol_name: "oldName", new_name: "newName", dry_run: true}) — preview all edits
- [ ] Review graph edits (high confidence) and ast_search edits (review carefully)
- [ ] If satisfied: gitnexus_rename({..., dry_run: false}) — apply edits
- [ ] gitnexus_detect_changes() — verify only expected files changed
- [ ] Run tests for affected processes
```
### Extract Module
```
- [ ] gitnexus_context({name: target}) — see all incoming/outgoing refs
- [ ] gitnexus_impact({target, direction: "upstream"}) — find all external callers
- [ ] Define new module interface
- [ ] Extract code, update imports
- [ ] gitnexus_detect_changes() — verify affected scope
- [ ] Run tests for affected processes
```
### Split Function/Service
```
- [ ] gitnexus_context({name: target}) — understand all callees
- [ ] Group callees by responsibility
- [ ] gitnexus_impact({target, direction: "upstream"}) — map callers to update
- [ ] Create new functions/services
- [ ] Update callers
- [ ] gitnexus_detect_changes() — verify affected scope
- [ ] Run tests for affected processes
```
## Tools
**gitnexus_rename** — automated multi-file rename:
```
gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
→ 12 edits across 8 files
→ 10 graph edits (high confidence), 2 ast_search edits (review)
→ Changes: [{file_path, edits: [{line, old_text, new_text, confidence}]}]
```
**gitnexus_impact** — map all dependents first:
```
gitnexus_impact({target: "validateUser", direction: "upstream"})
→ d=1: loginHandler, apiMiddleware, testUtils
→ Affected Processes: LoginFlow, TokenRefresh
```
**gitnexus_detect_changes** — verify your changes after refactoring:
```
gitnexus_detect_changes({scope: "all"})
→ Changed: 8 files, 12 symbols
→ Affected processes: LoginFlow, TokenRefresh
→ Risk: MEDIUM
```
**gitnexus_cypher** — custom reference queries:
```cypher
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(f:Function {name: "validateUser"})
RETURN caller.name, caller.filePath ORDER BY caller.filePath
```
## Risk Rules
| Risk Factor | Mitigation |
| ------------------- | ----------------------------------------- |
| Many callers (>5) | Use gitnexus_rename for automated updates |
| Cross-area refs | Use detect_changes after to verify scope |
| String/dynamic refs | gitnexus_query to find them |
| External/public API | Version and deprecate properly |
## Example: Rename `validateUser` to `authenticateUser`
```
1. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
→ 12 edits: 10 graph (safe), 2 ast_search (review)
→ Files: validator.ts, login.ts, middleware.ts, config.json...
2. Review ast_search edits (config.json: dynamic reference!)
3. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: false})
→ Applied 12 edits across 8 files
4. gitnexus_detect_changes({scope: "all"})
→ Affected: LoginFlow, TokenRefresh
→ Risk: MEDIUM — run tests for these flows
```
---
name: gitnexus-refactoring
description: "Use when the user wants to rename, extract, split, move, or restructure code safely. Examples: \"Rename this function\", \"Extract this into a module\", \"Refactor this class\", \"Move this to a separate file\""
---
# Refactoring with GitNexus
## When to Use
- "Rename this function safely"
- "Extract this into a module"
- "Split this service"
- "Move this to a new file"
- Any task involving renaming, extracting, splitting, or restructuring code
## Workflow
```
1. gitnexus_impact({target: "X", direction: "upstream"}) → Map all dependents
2. gitnexus_query({query: "X"}) → Find execution flows involving X
3. gitnexus_context({name: "X"}) → See all incoming/outgoing refs
4. Plan update order: interfaces → implementations → callers → tests
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklists
### Rename Symbol
```
- [ ] gitnexus_rename({symbol_name: "oldName", new_name: "newName", dry_run: true}) — preview all edits
- [ ] Review graph edits (high confidence) and ast_search edits (review carefully)
- [ ] If satisfied: gitnexus_rename({..., dry_run: false}) — apply edits
- [ ] gitnexus_detect_changes() — verify only expected files changed
- [ ] Run tests for affected processes
```
### Extract Module
```
- [ ] gitnexus_context({name: target}) — see all incoming/outgoing refs
- [ ] gitnexus_impact({target, direction: "upstream"}) — find all external callers
- [ ] Define new module interface
- [ ] Extract code, update imports
- [ ] gitnexus_detect_changes() — verify affected scope
- [ ] Run tests for affected processes
```
### Split Function/Service
```
- [ ] gitnexus_context({name: target}) — understand all callees
- [ ] Group callees by responsibility
- [ ] gitnexus_impact({target, direction: "upstream"}) — map callers to update
- [ ] Create new functions/services
- [ ] Update callers
- [ ] gitnexus_detect_changes() — verify affected scope
- [ ] Run tests for affected processes
```
## Tools
**gitnexus_rename** — automated multi-file rename:
```
gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
→ 12 edits across 8 files
→ 10 graph edits (high confidence), 2 ast_search edits (review)
→ Changes: [{file_path, edits: [{line, old_text, new_text, confidence}]}]
```
**gitnexus_impact** — map all dependents first:
```
gitnexus_impact({target: "validateUser", direction: "upstream"})
→ d=1: loginHandler, apiMiddleware, testUtils
→ Affected Processes: LoginFlow, TokenRefresh
```
**gitnexus_detect_changes** — verify your changes after refactoring:
```
gitnexus_detect_changes({scope: "all"})
→ Changed: 8 files, 12 symbols
→ Affected processes: LoginFlow, TokenRefresh
→ Risk: MEDIUM
```
**gitnexus_cypher** — custom reference queries:
```cypher
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(f:Function {name: "validateUser"})
RETURN caller.name, caller.filePath ORDER BY caller.filePath
```
## Risk Rules
| Risk Factor | Mitigation |
| ------------------- | ----------------------------------------- |
| Many callers (>5) | Use gitnexus_rename for automated updates |
| Cross-area refs | Use detect_changes after to verify scope |
| String/dynamic refs | gitnexus_query to find them |
| External/public API | Version and deprecate properly |
## Example: Rename `validateUser` to `authenticateUser`
```
1. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
→ 12 edits: 10 graph (safe), 2 ast_search (review)
→ Files: validator.ts, login.ts, middleware.ts, config.json...
2. Review ast_search edits (config.json: dynamic reference!)
3. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: false})
→ Applied 12 edits across 8 files
4. gitnexus_detect_changes({scope: "all"})
→ Affected: LoginFlow, TokenRefresh
→ Risk: MEDIUM — run tests for these flows
```
-28
View File
@@ -1,28 +0,0 @@
name: Setup GitNexus
description: Setup Node.js 20, install dependencies, and optionally build
inputs:
build:
description: Whether to run npm run build after install
required: false
default: 'false'
runs:
using: composite
steps:
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus/package-lock.json
- name: Install dependencies
run: npm ci
shell: bash
working-directory: gitnexus
- name: Build
if: ${{ inputs.build == 'true' }}
run: npm run build
shell: bash
working-directory: gitnexus
-45
View File
@@ -1,45 +0,0 @@
changelog:
exclude:
labels:
- chore
authors:
- dependabot
- dependabot[bot]
categories:
- title: "\U0001F6A8 Security"
labels:
- security
- title: "\U0001F4A5 Breaking Changes"
labels:
- breaking
- title: "\U0001F680 Features"
labels:
- enhancement
- title: "\U0001F41B Bug Fixes"
labels:
- bug
- title: "\U0001F3CE\uFE0F Performance"
labels:
- performance
- title: "\U0001F9EA Tests"
labels:
- test
- title: "\U0001F504 Refactoring"
labels:
- refactor
- title: "\U0001F477 CI/CD"
labels:
- ci
- title: "\U0001F4E6 Dependencies"
labels:
- dependencies
- title: "\U0001F4DD Other Changes"
labels:
- "*"
exclude:
labels:
- dependencies
- ci
- test
- refactor
- chore
-192
View File
@@ -1,192 +0,0 @@
name: Integration Tests
on:
workflow_call:
inputs:
collect-coverage:
description: 'Whether to run the coverage collection job (only needed for PR reports)'
required: false
default: true
type: boolean
jobs:
# ── Integration test matrix ─────────────────────────────────────────
# Each test-group runs on a SEPARATE runner per OS, giving full process
# isolation for the KuzuDB native C++ addon.
# 3 OS x 4 groups = 12 parallel jobs.
#
# Groups:
# kuzu-db — 7 files using withTestKuzuDB / kuzu-adapter (native addon)
# Each file runs as its own `vitest run` invocation for full
# process isolation. KuzuDB's native N-API addon registers
# persistent handles that prevent fork workers from exiting
# on Linux, and its C++ destructors segfault during
# process.exit(). Running each file in its own process lets
# the OS reclaim all resources cleanly.
# pipeline — 12 files: ingestion pipeline + csv + 9 resolver tests
# e2e — 2 files: child-process only (spawnSync), no in-process kuzu
# standalone — 4 files: pure logic, no kuzu, no child processes
test-matrix:
name: integration (${{ matrix.os }} / ${{ matrix.test-group }})
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, windows-latest, macos-latest]
test-group: [kuzu-db, pipeline, e2e, standalone]
include:
- test-group: kuzu-db
# Marker — actual files are listed in the run step below
test-glob: ''
- test-group: pipeline
test-glob: >-
test/integration/pipeline.test.ts
test/integration/csv-pipeline.test.ts
test/integration/parsing.test.ts
test/integration/resolvers/typescript.test.ts
test/integration/resolvers/csharp.test.ts
test/integration/resolvers/cpp.test.ts
test/integration/resolvers/java.test.ts
test/integration/resolvers/python.test.ts
test/integration/resolvers/rust.test.ts
test/integration/resolvers/go.test.ts
test/integration/resolvers/kotlin.test.ts
test/integration/resolvers/php.test.ts
- test-group: e2e
test-glob: >-
test/integration/cli-e2e.test.ts
test/integration/hooks-e2e.test.ts
test/integration/skills-e2e.test.ts
- test-group: standalone
test-glob: >-
test/integration/filesystem-walker.test.ts
test/integration/enrichment.test.ts
test/integration/tree-sitter-languages.test.ts
test/integration/worker-pool.test.ts
runs-on: ${{ matrix.os }}
timeout-minutes: 25
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
with:
build: 'true'
# kuzu-db: run each file in its own vitest process for full isolation.
# KuzuDB's native addon hangs fork workers on Linux — process isolation
# is the only reliable fix boundary.
- name: Run integration tests — kuzu-db (process-isolated)
if: matrix.test-group == 'kuzu-db'
working-directory: gitnexus
shell: bash
run: |
set -e
files=(
test/integration/kuzu-core-adapter.test.ts
test/integration/kuzu-pool.test.ts
test/integration/local-backend.test.ts
test/integration/local-backend-calltool.test.ts
test/integration/search-core.test.ts
test/integration/search-pool.test.ts
test/integration/augmentation.test.ts
)
exit_code=0
for f in "${files[@]}"; do
echo "::group::$f"
if ! npx vitest run --reporter=verbose --pool=forks "$f"; then
exit_code=1
echo "::error::Test file failed: $f"
fi
echo "::endgroup::"
done
exit $exit_code
# Non-kuzu groups: run all files in a single vitest invocation
- name: Run integration tests — ${{ matrix.test-group }}
if: matrix.test-group != 'kuzu-db'
shell: bash
env:
TEST_GLOB: ${{ matrix.test-glob }}
run: npx vitest run --reporter=verbose $TEST_GLOB
working-directory: gitnexus
# ── Coverage collection (ubuntu only) ─────────────────────────────────
# Runs non-kuzu integration tests with coverage enabled so the PR report
# can merge integration + unit coverage for a combined view.
# kuzu-db tests are excluded because each file must run in its own vitest
# process (native addon isolation) which prevents single-run coverage merge.
coverage:
name: integration (ubuntu / coverage)
if: inputs.collect-coverage
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
with:
build: 'true'
- name: Run integration tests with coverage
working-directory: gitnexus
run: >-
npx vitest run
--reporter=default
--reporter=json
--outputFile=integration-results.json
--coverage
--coverage.reporter=json-summary
--coverage.reporter=json
--coverage.reporter=text
--coverage.thresholdAutoUpdate=false
--coverage.reportOnFailure=true
--coverage.thresholds.statements=0
--coverage.thresholds.branches=0
--coverage.thresholds.functions=0
--coverage.thresholds.lines=0
test/integration/pipeline.test.ts
test/integration/csv-pipeline.test.ts
test/integration/parsing.test.ts
test/integration/cli-e2e.test.ts
test/integration/hooks-e2e.test.ts
test/integration/filesystem-walker.test.ts
test/integration/enrichment.test.ts
test/integration/tree-sitter-languages.test.ts
test/integration/worker-pool.test.ts
test/integration/resolvers/typescript.test.ts
test/integration/resolvers/csharp.test.ts
test/integration/resolvers/cpp.test.ts
test/integration/resolvers/java.test.ts
test/integration/resolvers/python.test.ts
test/integration/resolvers/rust.test.ts
test/integration/resolvers/go.test.ts
test/integration/resolvers/kotlin.test.ts
test/integration/resolvers/php.test.ts
- name: Upload integration coverage
if: always()
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
name: integration-reports
path: |
gitnexus/coverage/coverage-summary.json
gitnexus/coverage/coverage-final.json
gitnexus/integration-results.json
retention-days: 5
# ── Unified status gate ──────────────────────────────────────────────
# Branch protection should require THIS job, not the matrix jobs directly.
# ci.yml's needs.integration.result aggregates through this gate.
status:
name: integration (all groups)
needs: test-matrix
if: always()
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- name: Check all matrix jobs passed
shell: bash
env:
RESULT: ${{ needs.test-matrix.result }}
run: |
if [[ "$RESULT" != "success" ]]; then
echo "::error::Integration matrix failed or cancelled: $RESULT"
exit 1
fi
-14
View File
@@ -1,14 +0,0 @@
name: Quality Checks
on:
workflow_call:
jobs:
typecheck:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
- run: npx tsc --noEmit
working-directory: gitnexus
-432
View File
@@ -1,432 +0,0 @@
name: CI Report
# Triggered after the CI workflow completes. Because workflow_run
# always runs code from the *default branch*, it receives a read/write
# GITHUB_TOKEN — even when the triggering PR comes from a fork.
on:
workflow_run:
workflows: ["CI"]
types: [completed]
permissions:
actions: read # needed to list/download workflow run artifacts
contents: read # needed for sparse checkout of vitest.config.ts
pull-requests: write # needed to post sticky PR comment
jobs:
pr-report:
name: PR Report
# Only run for pull-request CI runs
if: >-
github.event.workflow_run.event == 'pull_request' &&
github.event.workflow_run.conclusion != 'cancelled'
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
# ── Download artifacts from the CI run ────────────────────────
- name: Download artifacts
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
with:
script: |
const fs = require('fs');
const path = require('path');
const runId = context.payload.workflow_run.id;
const allArtifacts = await github.rest.actions.listWorkflowRunArtifacts({
owner: context.repo.owner,
repo: context.repo.repo,
run_id: runId,
});
async function downloadArtifact(name, dest) {
const match = allArtifacts.data.artifacts.find(a => a.name === name);
if (!match) {
core.warning(`Artifact "${name}" not found`);
return false;
}
const zip = await github.rest.actions.downloadArtifact({
owner: context.repo.owner,
repo: context.repo.repo,
artifact_id: match.id,
archive_format: 'zip',
});
fs.mkdirSync(dest, { recursive: true });
fs.writeFileSync(path.join(dest, `${name}.zip`), Buffer.from(zip.data));
return true;
}
const temp = process.env.RUNNER_TEMP;
await downloadArtifact('pr-meta', path.join(temp, 'dl'));
await downloadArtifact('test-reports', path.join(temp, 'dl'));
await downloadArtifact('integration-reports', path.join(temp, 'dl'));
- name: Extract artifacts
shell: bash
run: |
cd "$RUNNER_TEMP/dl"
# Extract each artifact into its own directory to avoid filename collisions
for z in *.zip; do
[ -f "$z" ] || continue
name="${z%.zip}"
mkdir -p "$RUNNER_TEMP/artifacts/$name"
unzip -o "$z" -d "$RUNNER_TEMP/artifacts/$name"
done
- name: Read PR metadata
id: meta
shell: bash
run: |
DIR="$RUNNER_TEMP/artifacts/pr-meta"
if [ ! -f "$DIR/pr_number" ]; then
echo "skip=true" >> "$GITHUB_OUTPUT"
echo "::warning::pr_number artifact missing — skipping report"
exit 0
fi
# Validate PR number is a positive integer (artifact comes from
# untrusted fork code, so treat contents defensively).
PR_NUM=$(cat "$DIR/pr_number" | tr -d '[:space:]')
if ! [[ "$PR_NUM" =~ ^[0-9]+$ ]]; then
echo "skip=true" >> "$GITHUB_OUTPUT"
echo "::error::Invalid PR number in artifact: '$PR_NUM'"
exit 0
fi
echo "skip=false" >> "$GITHUB_OUTPUT"
echo "pr_number=$PR_NUM" >> "$GITHUB_OUTPUT"
# Validate job-result strings against known GitHub Actions values.
# Artifact contents come from the PR workflow (potentially untrusted
# fork code), so we whitelist to prevent newline injection into
# GITHUB_OUTPUT.
validate_result() {
local val
val=$(cat "$1" | tr -d '[:space:]')
case "$val" in
success|failure|cancelled|skipped) echo "$val" ;;
*) echo "unknown" ;;
esac
}
echo "quality=$(validate_result "$DIR/quality_result")" >> "$GITHUB_OUTPUT"
echo "unit=$(validate_result "$DIR/unit_result")" >> "$GITHUB_OUTPUT"
echo "integration=$(validate_result "$DIR/integration_result")" >> "$GITHUB_OUTPUT"
- name: Checkout (for vitest config)
if: steps.meta.outputs.skip != 'true'
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
with:
sparse-checkout: gitnexus/vitest.config.ts
sparse-checkout-cone-mode: false
# ── Merge coverage from unit + integration ─────────────────────
- name: Setup Node.js
if: steps.meta.outputs.skip != 'true'
uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
- name: Install coverage merge tools
if: steps.meta.outputs.skip != 'true'
run: npm install --no-save istanbul-lib-coverage istanbul-lib-report istanbul-reports
- name: Merge coverage reports
if: steps.meta.outputs.skip != 'true'
id: coverage
shell: bash
run: |
DIR="$RUNNER_TEMP/artifacts"
UNIT_COV=$(find "$DIR/test-reports" -name "coverage-final.json" -type f 2>/dev/null | head -1)
INTEG_COV=$(find "$DIR/integration-reports" -name "coverage-final.json" -type f 2>/dev/null | head -1)
MERGED_DIR="$RUNNER_TEMP/merged-coverage"
mkdir -p "$MERGED_DIR"
if [ -n "$UNIT_COV" ] && [ -n "$INTEG_COV" ]; then
echo "has_merged=true" >> "$GITHUB_OUTPUT"
# Merge using Node.js + istanbul-lib-coverage.
# Paths are passed via env vars to avoid shell interpolation
# inside the script string.
UNIT_COV_PATH="$UNIT_COV" \
INTEG_COV_PATH="$INTEG_COV" \
MERGED_OUT_DIR="$MERGED_DIR" \
node -e "
const libCoverage = require('istanbul-lib-coverage');
const libReport = require('istanbul-lib-report');
const reports = require('istanbul-reports');
const fs = require('fs');
const map = libCoverage.createCoverageMap({});
map.merge(JSON.parse(fs.readFileSync(process.env.UNIT_COV_PATH, 'utf8')));
map.merge(JSON.parse(fs.readFileSync(process.env.INTEG_COV_PATH, 'utf8')));
const context = libReport.createContext({
coverageMap: map,
dir: process.env.MERGED_OUT_DIR,
});
reports.create('json-summary').execute(context);
console.log('Merged coverage written to ' + process.env.MERGED_OUT_DIR + '/coverage-summary.json');
"
elif [ -n "$UNIT_COV" ]; then
echo "has_merged=false" >> "$GITHUB_OUTPUT"
echo "::warning::Integration coverage not found — using unit coverage only"
else
echo "has_merged=false" >> "$GITHUB_OUTPUT"
echo "::warning::No coverage data found"
fi
- name: Build report
if: steps.meta.outputs.skip != 'true'
id: report
shell: bash
env:
QUALITY: ${{ steps.meta.outputs.quality }}
UNIT: ${{ steps.meta.outputs.unit }}
INTEG: ${{ steps.meta.outputs.integration }}
HAS_MERGED: ${{ steps.coverage.outputs.has_merged }}
RUN_URL: ${{ github.event.workflow_run.html_url }}
run: |
DIR="$RUNNER_TEMP/artifacts"
MERGED_DIR="$RUNNER_TEMP/merged-coverage"
# ── Helper: read coverage summary into prefixed vars ──
# Uses printf -v for safe variable assignment (no eval).
read_cov() {
local prefix=$1 file=$2
if [ -n "$file" ] && [ -f "$file" ]; then
local val
val=$(jq -r '.total.statements.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_STMTS" '%s' "$val"
val=$(jq -r '.total.branches.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_BRANCH" '%s' "$val"
val=$(jq -r '.total.functions.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_FUNCS" '%s' "$val"
val=$(jq -r '.total.lines.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_LINES" '%s' "$val"
val=$(jq -r '"\(.total.statements.covered)/\(.total.statements.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_STMTS_COV" '%s' "$val"
val=$(jq -r '"\(.total.branches.covered)/\(.total.branches.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_BRANCH_COV" '%s' "$val"
val=$(jq -r '"\(.total.functions.covered)/\(.total.functions.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_FUNCS_COV" '%s' "$val"
val=$(jq -r '"\(.total.lines.covered)/\(.total.lines.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_LINES_COV" '%s' "$val"
return 0
else
printf -v "${prefix}_STMTS" '%s' "N/A"
printf -v "${prefix}_BRANCH" '%s' "N/A"
printf -v "${prefix}_FUNCS" '%s' "N/A"
printf -v "${prefix}_LINES" '%s' "N/A"
printf -v "${prefix}_STMTS_COV" '%s' ""
printf -v "${prefix}_BRANCH_COV" '%s' ""
printf -v "${prefix}_FUNCS_COV" '%s' ""
printf -v "${prefix}_LINES_COV" '%s' ""
return 1
fi
}
# ── Read all three coverage reports ──
UNIT_SUMMARY=$(find "$DIR/test-reports" -name "coverage-summary.json" -type f 2>/dev/null | head -1)
INTEG_SUMMARY=$(find "$DIR/integration-reports" -name "coverage-summary.json" -type f 2>/dev/null | head -1)
MERGED_SUMMARY="$MERGED_DIR/coverage-summary.json"
read_cov "U" "$UNIT_SUMMARY"
HAS_UNIT=$?
read_cov "I" "$INTEG_SUMMARY"
HAS_INTEG=$?
read_cov "M" "$MERGED_SUMMARY"
# ── Locate test results (unit) ──
RESULTS_FILE=$(find "$DIR/test-reports" -name "test-results.json" -type f 2>/dev/null | head -1)
INTEG_RESULTS=$(find "$DIR/integration-reports" -name "integration-results.json" -type f 2>/dev/null | head -1)
if [ -n "$RESULTS_FILE" ]; then
U_TOTAL=$(jq -r '.numTotalTests' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_PASSED=$(jq -r '.numPassedTests' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_FAILED=$(jq -r '.numFailedTests' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_SKIPPED=$(jq -r '.numPendingTests' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_SUITES=$(jq -r '.numTotalTestSuites' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_DURATION=$(jq -r '((.testResults | map(.endTime) | max) - (.startTime)) / 1000 | floor' "$RESULTS_FILE" 2>/dev/null || echo 0)
else
U_TOTAL=0; U_PASSED=0; U_FAILED=0; U_SKIPPED=0; U_SUITES=0; U_DURATION=0
fi
if [ -n "$INTEG_RESULTS" ]; then
I_TOTAL=$(jq -r '.numTotalTests' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_PASSED=$(jq -r '.numPassedTests' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_FAILED=$(jq -r '.numFailedTests' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_SKIPPED=$(jq -r '.numPendingTests' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_SUITES=$(jq -r '.numTotalTestSuites' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_DURATION=$(jq -r '((.testResults | map(.endTime) | max) - (.startTime)) / 1000 | floor' "$INTEG_RESULTS" 2>/dev/null || echo 0)
else
I_TOTAL=0; I_PASSED=0; I_FAILED=0; I_SKIPPED=0; I_SUITES=0; I_DURATION=0
fi
# ── Sum test results ──
TOTAL=$((U_TOTAL + I_TOTAL))
PASSED=$((U_PASSED + I_PASSED))
FAILED=$((U_FAILED + I_FAILED))
SKIPPED=$((U_SKIPPED + I_SKIPPED))
SUITES=$((U_SUITES + I_SUITES))
DURATION=$((U_DURATION + I_DURATION))
# ── Coverage thresholds (read from vitest.config.ts) ──
if [ -f gitnexus/vitest.config.ts ]; then
THRESH_STMTS=$(grep -oP 'statements:\s*\K[0-9]+' gitnexus/vitest.config.ts || echo 0)
THRESH_BRANCH=$(grep -oP 'branches:\s*\K[0-9]+' gitnexus/vitest.config.ts || echo 0)
THRESH_FUNCS=$(grep -oP 'functions:\s*\K[0-9]+' gitnexus/vitest.config.ts || echo 0)
THRESH_LINES=$(grep -oP 'lines:\s*\K[0-9]+' gitnexus/vitest.config.ts || echo 0)
else
THRESH_STMTS=0; THRESH_BRANCH=0; THRESH_FUNCS=0; THRESH_LINES=0
fi
# ── Status helpers ──
status_icon() {
case "$1" in
success) echo "✅" ;;
failure) echo "❌" ;;
cancelled) echo "⏭️" ;;
*) echo "❓" ;;
esac
}
cov_bar() {
local pct=$1 thresh=$2
if [ "$pct" = "N/A" ]; then echo "—"; return; fi
local filled
filled=$(awk "BEGIN { printf \"%d\", $pct / 5 }")
(( filled < 0 )) && filled=0
(( filled > 20 )) && filled=20
local empty=$((20 - filled))
local bar=""
for ((i=0; i<filled; i++)); do bar+="█"; done
for ((i=0; i<empty; i++)); do bar+="░"; done
if [ "$(awk "BEGIN { print ($pct >= $thresh) ? 1 : 0 }")" = "1" ]; then
echo "🟢 ${bar}"
else
echo "🔴 ${bar}"
fi
}
# ── Overall status ──
if [[ "$QUALITY" == "success" && "$UNIT" == "success" && "$INTEG" == "success" ]]; then
OVERALL="✅ **All checks passed**"
else
OVERALL="❌ **Some checks failed**"
fi
# ── Build markdown ──
{
echo "body<<GITNEXUS_CI_REPORT_EOF_7f3a"
echo "## CI Report"
echo ""
echo "${OVERALL}"
echo ""
echo "### Pipeline Status"
echo ""
echo "| Stage | Status | Details |"
echo "|-------|--------|---------|"
echo "| $(status_icon "$QUALITY") Typecheck | \`${QUALITY}\` | tsc --noEmit |"
echo "| $(status_icon "$UNIT") Unit Tests | \`${UNIT}\` | 3 platforms |"
echo "| $(status_icon "$INTEG") Integration | \`${INTEG}\` | 3 OS x 4 groups = 12 jobs |"
echo ""
if [ "$TOTAL" -gt 0 ] 2>/dev/null; then
echo "### Test Results"
echo ""
if [ "$FAILED" = "0" ]; then
echo "✅ **${PASSED}** passed"
else
echo "❌ **${FAILED}** failed / **${PASSED}** passed"
fi
if [ "$SKIPPED" != "0" ]; then
echo " · ${SKIPPED} skipped"
fi
echo " · ${SUITES} suites · ${TOTAL} total"
echo " · ⏱️ ${DURATION}s"
if [ "$I_TOTAL" -gt 0 ] 2>/dev/null; then
echo " · 📊 ${U_TOTAL} unit + ${I_TOTAL} integration"
fi
echo ""
fi
# ── Coverage table helper ──
cov_table() {
local label=$1 s=$2 b=$3 f=$4 l=$5 sc=$6 bc=$7 fc=$8 lc=$9
shift 9
local ts=$1 tb=$2 tf=$3 tl=$4
echo "#### ${label}"
echo ""
echo "| Metric | Coverage | Covered | Threshold | Status |"
echo "|--------|----------|---------|-----------|--------|"
echo "| Statements | **${s}%** | ${sc} | ${ts}% | $(cov_bar "$s" "$ts") |"
echo "| Branches | **${b}%** | ${bc} | ${tb}% | $(cov_bar "$b" "$tb") |"
echo "| Functions | **${f}%** | ${fc} | ${tf}% | $(cov_bar "$f" "$tf") |"
echo "| Lines | **${l}%** | ${lc} | ${tl}% | $(cov_bar "$l" "$tl") |"
echo ""
}
if [ "$M_STMTS" != "N/A" ]; then
echo "### Code Coverage"
echo ""
cov_table "Combined (Unit + Integration)" \
"$M_STMTS" "$M_BRANCH" "$M_FUNCS" "$M_LINES" \
"$M_STMTS_COV" "$M_BRANCH_COV" "$M_FUNCS_COV" "$M_LINES_COV" \
"$THRESH_STMTS" "$THRESH_BRANCH" "$THRESH_FUNCS" "$THRESH_LINES"
echo "<details>"
echo "<summary>Coverage breakdown by test suite</summary>"
echo ""
if [ "$U_STMTS" != "N/A" ]; then
cov_table "Unit Tests" \
"$U_STMTS" "$U_BRANCH" "$U_FUNCS" "$U_LINES" \
"$U_STMTS_COV" "$U_BRANCH_COV" "$U_FUNCS_COV" "$U_LINES_COV" \
"$THRESH_STMTS" "$THRESH_BRANCH" "$THRESH_FUNCS" "$THRESH_LINES"
fi
if [ "$I_STMTS" != "N/A" ]; then
cov_table "Integration Tests" \
"$I_STMTS" "$I_BRANCH" "$I_FUNCS" "$I_LINES" \
"$I_STMTS_COV" "$I_BRANCH_COV" "$I_FUNCS_COV" "$I_LINES_COV" \
"$THRESH_STMTS" "$THRESH_BRANCH" "$THRESH_FUNCS" "$THRESH_LINES"
fi
echo "</details>"
echo ""
echo "<details>"
echo "<summary>Coverage thresholds are auto-ratcheted — they only go up</summary>"
echo ""
echo "Vitest \`thresholds.autoUpdate\` bumps the floor whenever local coverage exceeds it."
echo "CI enforces the current thresholds; developers commit the ratcheted values."
echo "</details>"
echo ""
elif [ "$U_STMTS" != "N/A" ]; then
echo "### Code Coverage (Unit only)"
echo ""
cov_table "Unit Tests" \
"$U_STMTS" "$U_BRANCH" "$U_FUNCS" "$U_LINES" \
"$U_STMTS_COV" "$U_BRANCH_COV" "$U_FUNCS_COV" "$U_LINES_COV" \
"$THRESH_STMTS" "$THRESH_BRANCH" "$THRESH_FUNCS" "$THRESH_LINES"
echo "<details>"
echo "<summary>Coverage thresholds are auto-ratcheted — they only go up</summary>"
echo ""
echo "Vitest \`thresholds.autoUpdate\` bumps the floor whenever local coverage exceeds it."
echo "CI enforces the current thresholds; developers commit the ratcheted values."
echo "</details>"
echo ""
else
echo "### Code Coverage"
echo ""
echo "⚠️ Coverage data unavailable - check the [unit test job](${RUN_URL}) for details."
echo ""
fi
echo "---"
echo "<sub>📋 [View full run](${RUN_URL}) · Generated by CI</sub>"
echo "GITNEXUS_CI_REPORT_EOF_7f3a"
} >> "$GITHUB_OUTPUT"
- name: Comment on PR
if: steps.meta.outputs.skip != 'true'
uses: marocchino/sticky-pull-request-comment@773744901bac0e8cbb5a0dc842800d45e9b2b405 # v2
with:
header: ci-report
number: ${{ steps.meta.outputs.pr_number }}
message: ${{ steps.report.outputs.body }}
-53
View File
@@ -1,53 +0,0 @@
name: Unit Tests
on:
workflow_call:
jobs:
unit-tests:
name: unit (ubuntu / coverage)
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
- name: Run unit tests with coverage
run: >-
npx vitest run test/unit
--reporter=default
--reporter=json
--outputFile=test-results.json
--coverage
--coverage.reporter=json-summary
--coverage.reporter=json
--coverage.reporter=text
--coverage.thresholdAutoUpdate=false
--coverage.reportOnFailure=true
working-directory: gitnexus
- name: Upload test reports
if: always()
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
name: test-reports
path: |
gitnexus/coverage/coverage-summary.json
gitnexus/coverage/coverage-final.json
gitnexus/test-results.json
retention-days: 5
cross-platform:
name: unit (${{ matrix.os }})
strategy:
fail-fast: false
matrix:
# Ubuntu already covered by the coverage job above
os: [windows-latest, macos-latest]
runs-on: ${{ matrix.os }}
timeout-minutes: 15
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
- run: npx vitest run test/unit
working-directory: gitnexus
+51 -81
View File
@@ -3,97 +3,67 @@ name: CI
on:
push:
branches: [main]
paths-ignore: ['**.md', 'docs/**', 'LICENSE']
pull_request:
branches: [main]
paths-ignore: ['**.md', 'docs/**', 'LICENSE']
workflow_call:
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
# ── Reusable workflow orchestration ─────────────────────────────────
# Each concern lives in its own workflow file for maintainability:
# ci-quality.yml — typecheck (tsc --noEmit)
# ci-unit-tests.yml — unit tests with coverage + cross-platform
# ci-integration.yml — integration test matrix (3 OS x 4 groups)
#
# Shared setup is DRY via .github/actions/setup-gitnexus composite action.
jobs:
quality:
uses: ./.github/workflows/ci-quality.yml
permissions:
contents: read
typecheck:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus/package-lock.json
- run: npm ci
working-directory: gitnexus
- run: npx tsc --noEmit
working-directory: gitnexus
unit-tests:
uses: ./.github/workflows/ci-unit-tests.yml
permissions:
contents: read
integration:
uses: ./.github/workflows/ci-integration.yml
with:
collect-coverage: ${{ github.event_name == 'pull_request' }}
permissions:
contents: read
# ── Save PR metadata for the reporting workflow ─────────────────
# The ci-report.yml workflow (triggered by workflow_run) needs the
# PR number and job results to post a comment. We save them as an
# artifact because workflow_run context doesn't reliably carry PR
# info for fork PRs.
save-pr-meta:
name: Save PR Metadata
if: always() && github.event_name == 'pull_request'
needs: [quality, unit-tests, integration]
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- name: Write metadata
shell: bash
env:
PR_NUMBER: ${{ github.event.number }}
QUALITY: ${{ needs.quality.result }}
UNIT: ${{ needs.unit-tests.result }}
INTEG: ${{ needs.integration.result }}
run: |
mkdir -p pr-meta
echo "$PR_NUMBER" > pr-meta/pr_number
echo "$QUALITY" > pr-meta/quality_result
echo "$UNIT" > pr-meta/unit_result
echo "$INTEG" > pr-meta/integration_result
- name: Upload PR metadata
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
name: pr-meta
path: pr-meta/
retention-days: 1
node-version: 20
cache: npm
cache-dependency-path: gitnexus/package-lock.json
- run: npm ci
working-directory: gitnexus
- run: npx vitest run test/unit --coverage --coverage.thresholdAutoUpdate=false
working-directory: gitnexus
# ── Unified CI gate ──────────────────────────────────────────────
# Single required check for branch protection.
ci-status:
name: CI Gate
needs: [quality, unit-tests, integration]
if: always()
integration-tests:
runs-on: ubuntu-latest
timeout-minutes: 5
timeout-minutes: 10
steps:
- name: Check all jobs passed
shell: bash
env:
QUALITY: ${{ needs.quality.result }}
UNIT: ${{ needs.unit-tests.result }}
INTEG: ${{ needs.integration.result }}
run: |
echo "Quality: $QUALITY"
echo "Unit Tests: $UNIT"
echo "Integration: $INTEG"
if [[ "$QUALITY" != "success" ]] ||
[[ "$UNIT" != "success" ]] ||
[[ "$INTEG" != "success" ]]; then
echo "::error::One or more CI jobs failed"
exit 1
fi
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus/package-lock.json
- run: npm ci
working-directory: gitnexus
- run: npx vitest run test/integration
working-directory: gitnexus
cross-platform:
strategy:
matrix:
os: [ubuntu-latest, windows-latest]
runs-on: ${{ matrix.os }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus/package-lock.json
- run: npm ci
working-directory: gitnexus
- run: npx vitest run test/unit
working-directory: gitnexus
+22 -75
View File
@@ -1,97 +1,44 @@
name: Claude Code Review
# Uses pull_request_target so the workflow runs as defined on the default branch,
# which allows access to secrets for posting review comments on fork PRs.
# SECURITY: The checkout below uses the PR head SHA to review the correct code.
# The claude-code-action sandboxes execution — it does NOT run arbitrary code
# from the checked-out source.
on:
# Trigger only when explicitly requested:
# - Add the "claude-review" label to a PR, OR
# - Comment "@claude" or "/review" on a PR
pull_request_target:
types: [labeled]
issue_comment:
types: [created]
pull_request:
types: [opened, synchronize, ready_for_review, reopened]
# Optional: Only run on specific file changes
# paths:
# - "src/**/*.ts"
# - "src/**/*.tsx"
# - "src/**/*.js"
# - "src/**/*.jsx"
jobs:
claude-review:
# Run only when:
# 1. The "claude-review" label is added to a non-draft PR by a trusted contributor, OR
# 2. A trusted contributor comments "@claude" or "/review" on a PR
if: |
(
github.event_name == 'pull_request_target' &&
github.event.label.name == 'claude-review' &&
github.event.pull_request.draft == false &&
(github.event.pull_request.author_association == 'OWNER' ||
github.event.pull_request.author_association == 'MEMBER' ||
github.event.pull_request.author_association == 'COLLABORATOR')
) ||
(
github.event_name == 'issue_comment' &&
github.event.issue.pull_request &&
(contains(github.event.comment.body, '@claude') ||
contains(github.event.comment.body, '/review')) &&
(github.event.comment.author_association == 'OWNER' ||
github.event.comment.author_association == 'MEMBER' ||
github.event.comment.author_association == 'COLLABORATOR')
)
# Optional: Filter by PR author
# if: |
# github.event.pull_request.user.login == 'external-contributor' ||
# github.event.pull_request.user.login == 'new-developer' ||
# github.event.pull_request.author_association == 'FIRST_TIME_CONTRIBUTOR'
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
contents: write # needed to push fork branch to origin
pull-requests: write
contents: read
pull-requests: read
issues: read
id-token: write
steps:
# For issue_comment triggers, resolve the PR number, head SHA, and branch name
- name: Resolve PR context
id: pr
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
- name: Checkout repository
uses: actions/checkout@v4
with:
script: |
let pr;
if (context.eventName === 'issue_comment') {
const resp = await github.rest.pulls.get({
owner: context.repo.owner,
repo: context.repo.repo,
pull_number: context.payload.issue.number,
});
pr = resp.data;
} else {
pr = context.payload.pull_request;
}
core.setOutput('number', pr.number);
core.setOutput('sha', pr.head.sha);
core.setOutput('branch', pr.head.ref);
core.setOutput('is_fork', String(pr.head.repo.full_name !== pr.base.repo.full_name));
- name: Checkout PR head
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
with:
ref: ${{ steps.pr.outputs.sha }}
fetch-depth: 1
# claude-code-action fetches branches by name from origin, which fails
# for fork PRs. Work around by pushing the fork branch to origin so
# the action can find it. Cleaned up in the post step below.
- name: Push fork branch to origin
if: steps.pr.outputs.is_fork == 'true'
run: git push origin HEAD:refs/heads/${{ steps.pr.outputs.branch }}
- name: Run Claude Code Review
id: claude-review
uses: anthropics/claude-code-action@9469d113c6afd29550c402740f22d1a97dd1209b # v1
uses: anthropics/claude-code-action@v1
with:
claude_code_oauth_token: ${{ secrets.CLAUDE_CODE_OAUTH_TOKEN }}
plugin_marketplaces: 'https://github.com/anthropics/claude-code.git'
plugins: 'code-review@claude-code-plugins'
prompt: '/code-review:code-review ${{ github.repository }}/pull/${{ steps.pr.outputs.number }}'
prompt: '/code-review:code-review ${{ github.repository }}/pull/${{ github.event.pull_request.number }}'
# See https://github.com/anthropics/claude-code-action/blob/main/docs/usage.md
# or https://code.claude.com/docs/en/cli-reference for available options
# Clean up the temporary branch we pushed for fork PRs
- name: Delete fork branch from origin
if: always() && steps.pr.outputs.is_fork == 'true'
run: git push origin --delete refs/heads/${{ steps.pr.outputs.branch }} || true
+13 -5
View File
@@ -18,25 +18,33 @@ jobs:
(github.event_name == 'pull_request_review' && contains(github.event.review.body, '@claude')) ||
(github.event_name == 'issues' && (contains(github.event.issue.body, '@claude') || contains(github.event.issue.title, '@claude')))
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
contents: read
pull-requests: write
issues: write
pull-requests: read
issues: read
id-token: write
actions: read # Required for Claude to read CI results on PRs
steps:
- name: Checkout repository
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
uses: actions/checkout@v4
with:
fetch-depth: 1
- name: Run Claude Code
id: claude
uses: anthropics/claude-code-action@9469d113c6afd29550c402740f22d1a97dd1209b # v1
uses: anthropics/claude-code-action@v1
with:
claude_code_oauth_token: ${{ secrets.CLAUDE_CODE_OAUTH_TOKEN }}
# This is an optional setting that allows Claude to read CI results on PRs
additional_permissions: |
actions: read
# Optional: Give a custom prompt to Claude. If this is not specified, Claude will perform the instructions specified in the comment that tagged it.
# prompt: 'Update the pull request description to include a summary of changes.'
# Optional: Add claude_args to customize behavior and configuration
# See https://github.com/anthropics/claude-code-action/blob/main/docs/usage.md
# or https://code.claude.com/docs/en/cli-reference for available options
# claude_args: '--allowed-tools Bash(gh pr:*)'
+3 -14
View File
@@ -5,25 +5,19 @@ on:
tags:
- 'v*'
# No workflow-level permissions — scoped per job below.
jobs:
ci:
uses: ./.github/workflows/ci.yml
permissions:
contents: read
pull-requests: write
publish:
needs: ci
runs-on: ubuntu-latest
timeout-minutes: 15
permissions:
contents: write
id-token: write
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
registry-url: https://registry.npmjs.org
@@ -33,13 +27,8 @@ jobs:
working-directory: gitnexus
- name: Verify version consistency
shell: bash
run: |
TAG_VERSION="${GITHUB_REF#refs/tags/v}"
if ! [[ "$TAG_VERSION" =~ ^[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$ ]]; then
echo "::error::Tag does not follow semver: v$TAG_VERSION"
exit 1
fi
PKG_VERSION=$(node -p "require('./package.json').version")
if [ "$TAG_VERSION" != "$PKG_VERSION" ]; then
echo "::error::Tag version (v$TAG_VERSION) does not match package.json version ($PKG_VERSION)"
@@ -63,6 +52,6 @@ jobs:
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Create GitHub Release
uses: softprops/action-gh-release@a06a81a03ee405af7f2048a818ed3f03bbf83c7b # v2
uses: softprops/action-gh-release@v2
with:
generate_release_notes: true
-7
View File
@@ -48,9 +48,6 @@ coverage/
# Claude Code worktrees
.claude/worktrees/
# Claude code skills
.claude/skills/generated/
# Assets (screenshots, images)
assets/
@@ -59,7 +56,3 @@ repomix-output*
# Design docs (local only)
docs/plans/
gitnexus/test/fixtures/mini-repo/*.md
gitnexus/test/fixtures/mini-repo/.claude
gitnexus/test/fixtures/mini-repo/.gitignore
+8 -86
View File
@@ -1,75 +1,17 @@
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
# GitNexus MCP
This project is indexed by GitNexus as **GitNexus** (1747 symbols, 4569 relationships, 130 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitnexusV2** (1444 symbols, 3700 relationships, 111 execution flows).
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Start Here
## Always Do
1. **Read `gitnexus://repo/{name}/context`** — codebase overview + check index freshness
2. **Match your task to a skill below** and **read that skill file**
3. **Follow the skill's workflow and checklist**
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
> If step 1 warns the index is stale, run `npx gitnexus analyze` in the terminal first.
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find execution flows related to the issue
2. `gitnexus_context({name: "<suspect function>"})` — see all callers, callees, and process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace the full execution flow step by step
4. For regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})` — see what your branch changed
## When Refactoring
- **Renaming**: MUST use `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Review the preview — graph edits are safe, text_search edits need manual review. Then run with `dry_run: false`.
- **Extracting/Splitting**: MUST run `gitnexus_context({name: "target"})` to see all incoming/outgoing refs, then `gitnexus_impact({target: "target", direction: "upstream"})` to find all external callers before moving code.
- After any refactor: run `gitnexus_detect_changes({scope: "all"})` to verify only expected files changed.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Tools Quick Reference
| Tool | When to use | Command |
|------|-------------|---------|
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update these |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## Self-Check Before Finishing
Before completing any code modification task, verify:
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL risk warnings were ignored
3. `gitnexus_detect_changes()` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## CLI
## Skills
| Task | Read this skill file |
|------|---------------------|
@@ -79,25 +21,5 @@ Before completing any code modification task, verify:
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (135 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Workers area (70 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Cli area (63 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Kuzu area (52 symbols) | `.claude/skills/generated/kuzu/SKILL.md` |
| Work in the Wiki area (52 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Embeddings area (48 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Components area (42 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Local area (36 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Storage area (36 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Services area (35 symbols) | `.claude/skills/generated/services/SKILL.md` |
| Work in the Mcp area (32 symbols) | `.claude/skills/generated/mcp/SKILL.md` |
| Work in the Llm area (30 symbols) | `.claude/skills/generated/llm/SKILL.md` |
| Work in the Eval area (18 symbols) | `.claude/skills/generated/eval/SKILL.md` |
| Work in the Bridge area (15 symbols) | `.claude/skills/generated/bridge/SKILL.md` |
| Work in the Hooks area (14 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Search area (11 symbols) | `.claude/skills/generated/search/SKILL.md` |
| Work in the Environments area (11 symbols) | `.claude/skills/generated/environments/SKILL.md` |
| Work in the Analysis area (10 symbols) | `.claude/skills/generated/analysis/SKILL.md` |
| Work in the Agents area (9 symbols) | `.claude/skills/generated/agents/SKILL.md` |
| Work in the Graph area (6 symbols) | `.claude/skills/generated/graph/SKILL.md` |
<!-- gitnexus:end -->
-101
View File
@@ -1,101 +0,0 @@
# Changelog
All notable changes to GitNexus will be documented in this file.
## [1.4.0] - 2026-03-13
### Added
- **Language-aware symbol resolution engine** with 3-tier resolver: exact FQN → scope-walk → guarded fuzzy fallback that refuses ambiguous matches (#238) — @magyargergo
- **Method Resolution Order (MRO)** with 5 language-specific strategies: C++ leftmost-base, C#/Java class-over-interface, Python C3 linearization, Rust qualified syntax, default BFS (#238) — @magyargergo
- **Constructor & struct literal resolution** across all languages — `new Foo()`, `User{...}`, C# primary constructors, target-typed new (#238) — @magyargergo
- **Receiver-constrained resolution** using per-file TypeEnv — disambiguates `user.save()` vs `repo.save()` via `ownerId` matching (#238) — @magyargergo
- **Heritage & ownership edges** — HAS_METHOD, OVERRIDES, Go struct embedding, Swift extension heritage, method signatures (`parameterCount`, `returnType`) (#238) — @magyargergo
- **Language-specific resolver directory** (`resolvers/`) — extracted JVM, Go, C#, PHP, Rust resolvers from monolithic import-processor (#238) — @magyargergo
- **Type extractor directory** (`type-extractors/`) — per-language type binding extraction with `Record<SupportedLanguages, Handler>` + `satisfies` dispatch (#238) — @magyargergo
- **Export detection dispatch table** — compile-time exhaustive `Record` + `satisfies` pattern replacing switch/if chains (#238) — @magyargergo
- **Language config module** (`language-config.ts`) — centralized tsconfig, go.mod, composer.json, .csproj, Swift package config loaders (#238) — @magyargergo
- **Optional skill generation** via `npx gitnexus analyze --skills` — generates AI agent skills from KuzuDB knowledge graph (#171) — @zander-raycraft
- **First-class C# support** — sibling-based modifier scanning, record/delegate/property/field/event declaration types (#163, #170, #178 via #237) — @Alice523, @benny-yamagata, @jnMetaCode
- **C/C++ support fixes** — `.h` → C++ mapping, static-linkage export detection, qualified/parenthesized declarators, 48 entry point patterns (#163, #227 via #237) — @Alice523, @bitgineer
- **Rust support fixes** — sibling-based `visibility_modifier` scanning for `pub` detection (#227 via #237) — @bitgineer
- **Adaptive tree-sitter buffer sizing** — `Math.min(Math.max(contentLength * 2, 512KB), 32MB)` (#216 via #237) — @JasonOA888
- **Call expression matching** in tree-sitter queries (#234 via #237) — @ex-nihilo-jg
- **DeepSeek model configurations** (#217) — @JasonOA888
- 282+ new unit tests, 178 integration resolver tests across 9 languages, 53 test files, 1146 total tests passing
### Fixed
- Skip unavailable native Swift parsers in sequential ingestion (#188) — @Gujiassh
- Heritage heuristic language-gated — no longer applies class/interface rules to wrong languages (#238) — @magyargergo
- C# `base_list` distinguishes EXTENDS vs IMPLEMENTS via symbol table + `I[A-Z]` heuristic (#238) — @magyargergo
- Go `qualified_type` (`models.User`) correctly unwrapped in TypeEnv (#238) — @magyargergo
- Global tier no longer blocks resolution when kind/arity filtering can narrow to 1 candidate (#238) — @magyargergo
### Changed
- `import-processor.ts` reduced from 1412 → 711 lines (50% reduction) via resolver and config extraction (#238) — @magyargergo
- `type-env.ts` reduced from 635 → ~125 lines via type-extractor extraction (#238) — @magyargergo
- CI/CD workflows hardened with security fixes and fork PR support (#222, #225) — @magyargergo
## [1.3.11] - 2026-03-08
### Security
- Fix FTS Cypher injection by escaping backslashes in search queries (#209) — @magyargergo
### Added
- Auto-reindex hook that runs `gitnexus analyze` after commits and merges, with automatic embeddings preservation (#205) — @L1nusB
- 968 integration tests (up from ~840) covering unhappy paths across search, enrichment, CLI, pipeline, worker pool, and KuzuDB (#209) — @magyargergo
- Coverage auto-ratcheting so thresholds bump automatically on CI (#209) — @magyargergo
- Rich CI PR report with coverage bars, test counts, and threshold tracking (#209) — @magyargergo
- Modular CI workflow architecture with separate unit-test, integration-test, and orchestrator jobs (#209) — @magyargergo
### Fixed
- KuzuDB native addon crashes on Linux/macOS by running integration tests in isolated vitest processes with `--pool=forks` (#209) — @magyargergo
- Worker pool `MODULE_NOT_FOUND` crash when script path is invalid (#209) — @magyargergo
### Changed
- Added macOS to the cross-platform CI test matrix (#208) — @magyargergo
## [1.3.10] - 2026-03-07
### Security
- **MCP transport buffer cap**: Added 10 MB `MAX_BUFFER_SIZE` limit to prevent out-of-memory attacks via oversized `Content-Length` headers or unbounded newline-delimited input
- **Content-Length validation**: Reject `Content-Length` values exceeding the buffer cap before allocating memory
- **Stack overflow prevention**: Replaced recursive `readNewlineMessage` with iterative loop to prevent stack overflow from consecutive empty lines
- **Ambiguous prefix hardening**: Tightened `looksLikeContentLength` to require 14+ bytes before matching, preventing false framing detection on short input
- **Closed transport guard**: `send()` now rejects with a clear error when called after `close()`, with proper write-error propagation
### Added
- **Dual-framing MCP transport** (`CompatibleStdioServerTransport`): Auto-detects Content-Length (Codex/OpenCode) and newline-delimited JSON (Cursor/Claude Code) framing on the first message, responds in the same format (#207)
- **Lazy CLI module loading**: All CLI subcommands now use `createLazyAction()` to defer heavy imports (tree-sitter, ONNX, KuzuDB) until invocation, significantly improving `gitnexus mcp` startup time (#207)
- **Type-safe lazy actions**: `createLazyAction` uses constrained generics to validate export names against module types at compile time
- **Regression test suite**: 13 unit tests covering transport framing, security hardening, buffer limits, and lazy action loading
### Fixed
- **CALLS edge sourceId alignment**: `findEnclosingFunctionId` now generates IDs with `:startLine` suffix matching node creation format, fixing process detector finding 0 entry points (#194)
- **LRU cache zero maxSize crash**: Guard `createASTCache` against `maxSize=0` when repos have no parseable files (#144)
### Changed
- Transport constructor accepts `NodeJS.ReadableStream` / `NodeJS.WritableStream` (widened from concrete `ReadStream`/`WriteStream`)
- `processReadBuffer` simplified to break on first error instead of stale-buffer retry loop
## [1.3.9] - 2026-03-06
### Fixed
- Aligned CALLS edge sourceId with node ID format in parse worker (#194)
## [1.3.8] - 2026-03-05
### Fixed
- Force-exit after analyze to prevent KuzuDB native cleanup hang (#192)
+8 -86
View File
@@ -1,75 +1,17 @@
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
# GitNexus MCP
This project is indexed by GitNexus as **GitNexus** (1747 symbols, 4569 relationships, 130 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitnexusV2** (1444 symbols, 3700 relationships, 111 execution flows).
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Start Here
## Always Do
1. **Read `gitnexus://repo/{name}/context`** — codebase overview + check index freshness
2. **Match your task to a skill below** and **read that skill file**
3. **Follow the skill's workflow and checklist**
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
> If step 1 warns the index is stale, run `npx gitnexus analyze` in the terminal first.
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find execution flows related to the issue
2. `gitnexus_context({name: "<suspect function>"})` — see all callers, callees, and process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace the full execution flow step by step
4. For regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})` — see what your branch changed
## When Refactoring
- **Renaming**: MUST use `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Review the preview — graph edits are safe, text_search edits need manual review. Then run with `dry_run: false`.
- **Extracting/Splitting**: MUST run `gitnexus_context({name: "target"})` to see all incoming/outgoing refs, then `gitnexus_impact({target: "target", direction: "upstream"})` to find all external callers before moving code.
- After any refactor: run `gitnexus_detect_changes({scope: "all"})` to verify only expected files changed.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Tools Quick Reference
| Tool | When to use | Command |
|------|-------------|---------|
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update these |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## Self-Check Before Finishing
Before completing any code modification task, verify:
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL risk warnings were ignored
3. `gitnexus_detect_changes()` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## CLI
## Skills
| Task | Read this skill file |
|------|---------------------|
@@ -79,25 +21,5 @@ Before completing any code modification task, verify:
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (135 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Workers area (70 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Cli area (63 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Kuzu area (52 symbols) | `.claude/skills/generated/kuzu/SKILL.md` |
| Work in the Wiki area (52 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Embeddings area (48 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Components area (42 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Local area (36 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Storage area (36 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Services area (35 symbols) | `.claude/skills/generated/services/SKILL.md` |
| Work in the Mcp area (32 symbols) | `.claude/skills/generated/mcp/SKILL.md` |
| Work in the Llm area (30 symbols) | `.claude/skills/generated/llm/SKILL.md` |
| Work in the Eval area (18 symbols) | `.claude/skills/generated/eval/SKILL.md` |
| Work in the Bridge area (15 symbols) | `.claude/skills/generated/bridge/SKILL.md` |
| Work in the Hooks area (14 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Search area (11 symbols) | `.claude/skills/generated/search/SKILL.md` |
| Work in the Environments area (11 symbols) | `.claude/skills/generated/environments/SKILL.md` |
| Work in the Analysis area (10 symbols) | `.claude/skills/generated/analysis/SKILL.md` |
| Work in the Agents area (9 symbols) | `.claude/skills/generated/agents/SKILL.md` |
| Work in the Graph area (6 symbols) | `.claude/skills/generated/graph/SKILL.md` |
<!-- gitnexus:end -->
+2 -9
View File
@@ -82,12 +82,12 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
| --------------------- | --- | ------ | -------------------- | -------------- |
| **Claude Code** | Yes | Yes | Yes (PreToolUse + PostToolUse) | **Full** |
| **Claude Code** | Yes | Yes | Yes (PreToolUse) | **Full** |
| **Cursor** | Yes | Yes | — | MCP + Skills |
| **Windsurf** | Yes | — | — | MCP |
| **OpenCode** | Yes | Yes | — | MCP + Skills |
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that enrich searches with graph context + PostToolUse hooks that auto-reindex after commits.
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that automatically enrich grep/glob/bash calls with knowledge graph context.
### Community Integrations
@@ -135,10 +135,7 @@ claude mcp add gitnexus -- npx -y gitnexus@latest mcp
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
gitnexus serve # Start local HTTP server (multi-repo) for web UI connection
gitnexus list # List all indexed repositories
@@ -192,10 +189,6 @@ gitnexus wiki --base-url <url> # Wiki with custom LLM API base URL
- **Impact Analysis** — Analyze blast radius before changes
- **Refactoring** — Plan safe refactors using dependency mapping
**Repo-specific skills** generated with `--skills`:
When you run `gitnexus analyze --skills`, GitNexus detects the functional areas of your codebase (via Leiden community detection) and generates a `SKILL.md` file for each one under `.claude/skills/generated/`. Each skill describes a module's key files, entry points, execution flows, and cross-area connections — so your AI agent gets targeted context for the exact area of code you're working in. Skills are regenerated on each `--skills` run to stay current with the codebase.
---
## Multi-Repo MCP Architecture
-13
View File
@@ -1,13 +0,0 @@
model: deepseek-ai/deepseek-chat
provider: openrouter
cost:
input: 0.14 # per 1M tokens
output: 0.28 # per 1M tokens
# Native DeepSeek API (direct)
api_key: null
base_url: null
# For OpenRouter, uncomment below and comment out direct config above
# api_key: \${OPENROUTER_API_KEY}
# base_url: https://openrouter.ai/api/v1
-15
View File
@@ -1,15 +0,0 @@
model: deepseek-ai/DeepSeek-V3
provider: openrouter
cost:
input: 0.27 # per 1M tokens
output: 1.10 # per 1M tokens
# Native DeepSeek API (direct)
# Get your API key at: https://platform.deepseek.com/
# Or use OpenRouter with: OPENROUTER_API_KEY
api_key: null
base_url: null
# For OpenRouter, uncomment below and comment out direct config above
# api_key: \${OPENROUTER_API_KEY}
# base_url: https://openrouter.ai/api/v1
+64 -146
View File
@@ -2,10 +2,8 @@
/**
* GitNexus Claude Code Plugin Hook
*
* PreToolUse — intercepts Grep/Glob/Bash searches and augments
* with graph context from the GitNexus index.
* PostToolUse — detects stale index after git mutations and notifies
* the agent to reindex.
* PreToolUse handler — intercepts Grep/Glob/Bash searches
* and augments with graph context from the GitNexus index.
*
* NOTE: SessionStart hooks are broken on Windows (Claude Code bug #23576).
* Session context is injected via CLAUDE.md / skills instead.
@@ -28,19 +26,19 @@ function readInput() {
}
/**
* Find the .gitnexus directory by walking up from startDir.
* Returns the path to .gitnexus/ or null if not found.
* Check if a directory (or ancestor) has a .gitnexus index.
*/
function findGitNexusDir(startDir) {
function findGitNexusIndex(startDir) {
let dir = startDir || process.cwd();
for (let i = 0; i < 5; i++) {
const candidate = path.join(dir, '.gitnexus');
if (fs.existsSync(candidate)) return candidate;
if (fs.existsSync(path.join(dir, '.gitnexus'))) {
return true;
}
const parent = path.dirname(dir);
if (parent === dir) break;
dir = parent;
}
return null;
return false;
}
/**
@@ -85,146 +83,66 @@ function extractPattern(toolName, toolInput) {
return null;
}
/**
* Spawn a gitnexus CLI command synchronously.
* Detects binary on PATH once, then runs exactly once.
*
* SECURITY: Never use shell: true with user-controlled arguments.
* On Windows, invoke gitnexus.cmd directly (no shell needed).
*/
function runGitNexusCli(args, cwd, timeout) {
const isWin = process.platform === 'win32';
// Detect whether 'gitnexus' is on PATH (cheap check, no execution)
let useDirectBinary = false;
try {
const which = spawnSync(
isWin ? 'where' : 'which', ['gitnexus'],
{ encoding: 'utf-8', timeout: 3000, stdio: ['pipe', 'pipe', 'pipe'] }
);
useDirectBinary = which.status === 0;
} catch { /* not on PATH */ }
if (useDirectBinary) {
return spawnSync(
isWin ? 'gitnexus.cmd' : 'gitnexus', args,
{ encoding: 'utf-8', timeout, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
}
// npx fallback needs shell on Windows since npx is a .cmd script
return spawnSync(
isWin ? 'npx.cmd' : 'npx', ['-y', 'gitnexus', ...args],
{ encoding: 'utf-8', timeout: timeout + 5000, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
}
/**
* Emit a hook response with additional context for the agent.
*/
function sendHookResponse(hookEventName, message) {
console.log(JSON.stringify({
hookSpecificOutput: { hookEventName, additionalContext: message }
}));
}
/**
* PreToolUse handler — augment searches with graph context.
*/
function handlePreToolUse(input) {
const cwd = input.cwd || process.cwd();
if (!path.isAbsolute(cwd)) return;
if (!findGitNexusDir(cwd)) return;
const toolName = input.tool_name || '';
const toolInput = input.tool_input || {};
if (toolName !== 'Grep' && toolName !== 'Glob' && toolName !== 'Bash') return;
const pattern = extractPattern(toolName, toolInput);
if (!pattern || pattern.length < 3) return;
let result = '';
try {
const child = runGitNexusCli(['augment', '--', pattern], cwd, 7000);
if (!child.error && child.status === 0) {
result = child.stderr || '';
}
} catch { /* graceful failure */ }
if (result && result.trim()) {
sendHookResponse('PreToolUse', result.trim());
}
}
/**
* PostToolUse handler — detect index staleness after git mutations.
*
* Instead of spawning a full `gitnexus analyze` synchronously (which blocks
* the agent for up to 120s and risks KuzuDB corruption on timeout), we do a
* lightweight staleness check: compare `git rev-parse HEAD` against the
* lastCommit stored in `.gitnexus/meta.json`. If they differ, notify the
* agent so it can decide when to reindex.
*/
function handlePostToolUse(input) {
const toolName = input.tool_name || '';
if (toolName !== 'Bash') return;
const command = (input.tool_input || {}).command || '';
if (!/\bgit\s+(commit|merge|rebase|cherry-pick|pull)(\s|$)/.test(command)) return;
// Only proceed if the command succeeded
const toolOutput = input.tool_output || {};
if (toolOutput.exit_code !== undefined && toolOutput.exit_code !== 0) return;
const cwd = input.cwd || process.cwd();
if (!path.isAbsolute(cwd)) return;
const gitNexusDir = findGitNexusDir(cwd);
if (!gitNexusDir) return;
// Compare HEAD against last indexed commit — skip if unchanged
let currentHead = '';
try {
const headResult = spawnSync('git', ['rev-parse', 'HEAD'], {
encoding: 'utf-8', timeout: 3000, cwd, stdio: ['pipe', 'pipe', 'pipe'],
});
currentHead = (headResult.stdout || '').trim();
} catch { return; }
if (!currentHead) return;
let lastCommit = '';
let hadEmbeddings = false;
try {
const meta = JSON.parse(fs.readFileSync(path.join(gitNexusDir, 'meta.json'), 'utf-8'));
lastCommit = meta.lastCommit || '';
hadEmbeddings = (meta.stats && meta.stats.embeddings > 0);
} catch { /* no meta — treat as stale */ }
// If HEAD matches last indexed commit, no reindex needed
if (currentHead && currentHead === lastCommit) return;
const analyzeCmd = `npx gitnexus analyze${hadEmbeddings ? ' --embeddings' : ''}`;
sendHookResponse('PostToolUse',
`GitNexus index is stale (last indexed: ${lastCommit ? lastCommit.slice(0, 7) : 'never'}). ` +
`Run \`${analyzeCmd}\` to update the knowledge graph.`
);
}
// Dispatch map for hook events
const handlers = {
PreToolUse: handlePreToolUse,
PostToolUse: handlePostToolUse,
};
function main() {
try {
const input = readInput();
const handler = handlers[input.hook_event_name || ''];
if (handler) handler(input);
} catch (err) {
if (process.env.GITNEXUS_DEBUG) {
console.error('GitNexus hook error:', (err.message || '').slice(0, 200));
const hookEvent = input.hook_event_name || '';
if (hookEvent !== 'PreToolUse') return;
const cwd = input.cwd || process.cwd();
if (!findGitNexusIndex(cwd)) return;
const toolName = input.tool_name || '';
const toolInput = input.tool_input || {};
if (toolName !== 'Grep' && toolName !== 'Glob' && toolName !== 'Bash') return;
const pattern = extractPattern(toolName, toolInput);
if (!pattern || pattern.length < 3) return;
// augment CLI writes result to stderr (KuzuDB's native module captures
// stdout fd at OS level, making it unusable in subprocess contexts).
let result = '';
const isWin = process.platform === 'win32';
// Try direct gitnexus binary first (faster if globally installed)
try {
const child = spawnSync(
'gitnexus',
['augment', pattern],
{ encoding: 'utf-8', timeout: 8000, cwd, stdio: ['pipe', 'pipe', 'pipe'], shell: isWin }
);
if (child.status === 0 && child.stderr && child.stderr.trim()) {
result = child.stderr;
}
} catch { /* not on PATH */ }
// Fallback to npx if direct binary didn't produce output
if (!result || !result.trim()) {
try {
const child = spawnSync(
'npx',
['-y', 'gitnexus', 'augment', pattern],
{ encoding: 'utf-8', timeout: 15000, cwd, stdio: ['pipe', 'pipe', 'pipe'], shell: isWin }
);
if (child.status === 0 && child.stderr && child.stderr.trim()) {
result = child.stderr;
}
} catch { /* graceful failure */ }
}
if (result && result.trim()) {
console.log(JSON.stringify({
hookSpecificOutput: {
hookEventName: 'PreToolUse',
additionalContext: result.trim()
}
}));
}
} catch {
// Graceful failure
}
}
-13
View File
@@ -12,19 +12,6 @@
}
]
}
],
"PostToolUse": [
{
"matcher": "Bash",
"hooks": [
{
"type": "command",
"command": "node ${CLAUDE_PLUGIN_ROOT}/hooks/gitnexus-hook.js",
"timeout": 10,
"statusMessage": "Checking GitNexus index freshness..."
}
]
}
]
}
}
@@ -330,20 +330,25 @@ const calculateCohesion = (memberIds: string[], graph: Graph): number => {
const memberSet = new Set(memberIds);
let internalEdges = 0;
let totalEdges = 0;
// Count internal vs total edges for community members
// Count edges within the community
memberIds.forEach(nodeId => {
if (graph.hasNode(nodeId)) {
graph.forEachNeighbor(nodeId, neighbor => {
totalEdges++;
if (memberSet.has(neighbor)) {
internalEdges++;
}
});
}
});
if (totalEdges === 0) return 1.0;
return Math.min(1.0, internalEdges / totalEdges);
// Each edge is counted twice (once from each end), so divide by 2
internalEdges = internalEdges / 2;
// Maximum possible internal edges for n nodes: n*(n-1)/2
const maxPossibleEdges = (memberIds.length * (memberIds.length - 1)) / 2;
if (maxPossibleEdges === 0) return 1.0;
return Math.min(1.0, internalEdges / maxPossibleEdges);
};
+2 -22
View File
@@ -139,8 +139,7 @@ Your AI agent gets these tools automatically:
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
gitnexus serve # Start local HTTP server (multi-repo) for web UI
gitnexus list # List all indexed repositories
@@ -157,26 +156,7 @@ GitNexus supports indexing multiple repositories. Each `gitnexus analyze` regist
## Supported Languages
TypeScript, JavaScript, Python, Java, C, C++, C#, Go, Rust, PHP, Kotlin, Swift
### Language Feature Matrix
| Language | Imports | Types | Exports | Named Bindings | Config | Frameworks | Entry Points | Heritage |
|----------|---------|-------|---------|----------------|--------|------------|-------------|----------|
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ |
| Go | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — |
| Swift | — | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| C | — | ✓ | ✓ | — | — | ✓ | ✓ | ✓ |
| C++ | — | ✓ | ✓ | — | — | ✓ | ✓ | ✓ |
**Imports** — cross-file import resolution · **Types** — type annotation extraction · **Exports** — public/exported symbol detection · **Named Bindings** — `import { X }` tracking · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics · **Heritage** — class inheritance / interface implementation
TypeScript, JavaScript, Python, Java, C, C++, C#, Go, Rust, PHP, Swift
## Agent Skills
+71 -154
View File
@@ -2,10 +2,8 @@
/**
* GitNexus Claude Code Hook
*
* PreToolUse — intercepts Grep/Glob/Bash searches and augments
* with graph context from the GitNexus index.
* PostToolUse — detects stale index after git mutations and notifies
* the agent to reindex.
* PreToolUse handler — intercepts Grep/Glob/Bash searches
* and augments with graph context from the GitNexus index.
*
* NOTE: SessionStart hooks are broken on Windows (Claude Code bug).
* Session context is injected via CLAUDE.md / skills instead.
@@ -13,7 +11,7 @@
const fs = require('fs');
const path = require('path');
const { spawnSync } = require('child_process');
const { execFileSync } = require('child_process');
/**
* Read JSON input from stdin synchronously.
@@ -28,19 +26,19 @@ function readInput() {
}
/**
* Find the .gitnexus directory by walking up from startDir.
* Returns the path to .gitnexus/ or null if not found.
* Check if a directory (or ancestor) has a .gitnexus index.
*/
function findGitNexusDir(startDir) {
function findGitNexusIndex(startDir) {
let dir = startDir || process.cwd();
for (let i = 0; i < 5; i++) {
const candidate = path.join(dir, '.gitnexus');
if (fs.existsSync(candidate)) return candidate;
if (fs.existsSync(path.join(dir, '.gitnexus'))) {
return true;
}
const parent = path.dirname(dir);
if (parent === dir) break;
dir = parent;
}
return null;
return false;
}
/**
@@ -85,153 +83,72 @@ function extractPattern(toolName, toolInput) {
return null;
}
/**
* Resolve the gitnexus CLI path.
* 1. Relative path (works when script is inside npm package)
* 2. require.resolve (works when gitnexus is globally installed)
* 3. Fall back to npx (returns empty string)
*/
function resolveCliPath() {
let cliPath = path.resolve(__dirname, '..', '..', 'dist', 'cli', 'index.js');
if (!fs.existsSync(cliPath)) {
try {
cliPath = require.resolve('gitnexus/dist/cli/index.js');
} catch {
cliPath = '';
}
}
return cliPath;
}
/**
* Spawn a gitnexus CLI command synchronously.
* Returns the stderr output (KuzuDB captures stdout at OS level).
*/
function runGitNexusCli(cliPath, args, cwd, timeout) {
const isWin = process.platform === 'win32';
if (cliPath) {
return spawnSync(
process.execPath,
[cliPath, ...args],
{ encoding: 'utf-8', timeout, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
}
// On Windows, invoke npx.cmd directly (no shell needed)
return spawnSync(
isWin ? 'npx.cmd' : 'npx',
['-y', 'gitnexus', ...args],
{ encoding: 'utf-8', timeout: timeout + 5000, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
}
/**
* PreToolUse handler — augment searches with graph context.
*/
function handlePreToolUse(input) {
const cwd = input.cwd || process.cwd();
if (!path.isAbsolute(cwd)) return;
if (!findGitNexusDir(cwd)) return;
const toolName = input.tool_name || '';
const toolInput = input.tool_input || {};
if (toolName !== 'Grep' && toolName !== 'Glob' && toolName !== 'Bash') return;
const pattern = extractPattern(toolName, toolInput);
if (!pattern || pattern.length < 3) return;
const cliPath = resolveCliPath();
let result = '';
try {
const child = runGitNexusCli(cliPath, ['augment', '--', pattern], cwd, 7000);
if (!child.error && child.status === 0) {
result = child.stderr || '';
}
} catch { /* graceful failure */ }
if (result && result.trim()) {
sendHookResponse('PreToolUse', result.trim());
}
}
/**
* Emit a PostToolUse hook response with additional context for the agent.
*/
function sendHookResponse(hookEventName, message) {
console.log(JSON.stringify({
hookSpecificOutput: { hookEventName, additionalContext: message }
}));
}
/**
* PostToolUse handler — detect index staleness after git mutations.
*
* Instead of spawning a full `gitnexus analyze` synchronously (which blocks
* the agent for up to 120s and risks KuzuDB corruption on timeout), we do a
* lightweight staleness check: compare `git rev-parse HEAD` against the
* lastCommit stored in `.gitnexus/meta.json`. If they differ, notify the
* agent so it can decide when to reindex.
*/
function handlePostToolUse(input) {
const toolName = input.tool_name || '';
if (toolName !== 'Bash') return;
const command = (input.tool_input || {}).command || '';
if (!/\bgit\s+(commit|merge|rebase|cherry-pick|pull)(\s|$)/.test(command)) return;
// Only proceed if the command succeeded
const toolOutput = input.tool_output || {};
if (toolOutput.exit_code !== undefined && toolOutput.exit_code !== 0) return;
const cwd = input.cwd || process.cwd();
if (!path.isAbsolute(cwd)) return;
const gitNexusDir = findGitNexusDir(cwd);
if (!gitNexusDir) return;
// Compare HEAD against last indexed commit — skip if unchanged
let currentHead = '';
try {
const headResult = spawnSync('git', ['rev-parse', 'HEAD'], {
encoding: 'utf-8', timeout: 3000, cwd, stdio: ['pipe', 'pipe', 'pipe'],
});
currentHead = (headResult.stdout || '').trim();
} catch { return; }
if (!currentHead) return;
let lastCommit = '';
let hadEmbeddings = false;
try {
const meta = JSON.parse(fs.readFileSync(path.join(gitNexusDir, 'meta.json'), 'utf-8'));
lastCommit = meta.lastCommit || '';
hadEmbeddings = (meta.stats && meta.stats.embeddings > 0);
} catch { /* no meta — treat as stale */ }
// If HEAD matches last indexed commit, no reindex needed
if (currentHead && currentHead === lastCommit) return;
const analyzeCmd = `npx gitnexus analyze${hadEmbeddings ? ' --embeddings' : ''}`;
sendHookResponse('PostToolUse',
`GitNexus index is stale (last indexed: ${lastCommit ? lastCommit.slice(0, 7) : 'never'}). ` +
`Run \`${analyzeCmd}\` to update the knowledge graph.`
);
}
// Dispatch map for hook events
const handlers = {
PreToolUse: handlePreToolUse,
PostToolUse: handlePostToolUse,
};
function main() {
try {
const input = readInput();
const handler = handlers[input.hook_event_name || ''];
if (handler) handler(input);
} catch (err) {
if (process.env.GITNEXUS_DEBUG) {
console.error('GitNexus hook error:', (err.message || '').slice(0, 200));
const hookEvent = input.hook_event_name || '';
if (hookEvent !== 'PreToolUse') return;
const cwd = input.cwd || process.cwd();
if (!findGitNexusIndex(cwd)) return;
const toolName = input.tool_name || '';
const toolInput = input.tool_input || {};
if (toolName !== 'Grep' && toolName !== 'Glob' && toolName !== 'Bash') return;
const pattern = extractPattern(toolName, toolInput);
if (!pattern || pattern.length < 3) return;
// Resolve CLI path — try multiple strategies:
// 1. Relative path (works when script is inside npm package)
// 2. require.resolve (works when gitnexus is globally installed)
// 3. Fall back to npx (works when neither is available)
let cliPath = path.resolve(__dirname, '..', '..', 'dist', 'cli', 'index.js');
if (!fs.existsSync(cliPath)) {
try {
cliPath = require.resolve('gitnexus/dist/cli/index.js');
} catch {
cliPath = ''; // will use npx fallback
}
}
// augment CLI writes result to stderr (KuzuDB's native module captures
// stdout fd at OS level, making it unusable in subprocess contexts).
const { spawnSync } = require('child_process');
let result = '';
try {
let child;
if (cliPath) {
child = spawnSync(
process.execPath,
[cliPath, 'augment', pattern],
{ encoding: 'utf-8', timeout: 8000, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
} else {
// npx fallback
const cmd = process.platform === 'win32' ? 'npx.cmd' : 'npx';
child = spawnSync(
cmd,
['-y', 'gitnexus', 'augment', pattern],
{ encoding: 'utf-8', timeout: 15000, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
}
result = child.stderr || '';
} catch { /* graceful failure */ }
if (result && result.trim()) {
console.log(JSON.stringify({
hookSpecificOutput: {
hookEventName: 'PreToolUse',
additionalContext: result.trim()
}
}));
}
} catch (err) {
// Graceful failure — log to stderr for debugging
console.error('GitNexus hook error:', err.message);
}
}
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "gitnexus",
"version": "1.4.0",
"version": "1.3.9",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "gitnexus",
"version": "1.4.0",
"version": "1.3.9",
"hasInstallScript": true,
"license": "PolyForm-Noncommercial-1.0.0",
"dependencies": {
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "gitnexus",
"version": "1.4.0",
"version": "1.3.9",
"description": "Graph-powered code intelligence for AI agents. Index any codebase, query via MCP or CLI.",
"author": "Abhigyan Patwari",
"license": "PolyForm-Noncommercial-1.0.0",
+1 -1
View File
@@ -22,7 +22,7 @@ Run from the project root. This parses all source files, builds the knowledge gr
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook runs `analyze` automatically after `git commit` and `git merge`, preserving embeddings if previously generated.
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale.
### status — Check index freshness
+6 -39
View File
@@ -9,7 +9,6 @@
import fs from 'fs/promises';
import path from 'path';
import { fileURLToPath } from 'url';
import { type GeneratedSkillInfo } from './skill-gen.js';
// ESM equivalent of __dirname
const __filename = fileURLToPath(import.meta.url);
@@ -38,22 +37,7 @@ const GITNEXUS_END_MARKER = '<!-- gitnexus:end -->';
* - Exact tool commands with parameters — vague directives get ignored
* - Self-review checklist — forces model to verify its own work
*/
function generateGitNexusContent(projectName: string, stats: RepoStats, generatedSkills?: GeneratedSkillInfo[]): string {
const generatedRows = (generatedSkills && generatedSkills.length > 0)
? generatedSkills.map(s =>
`| Work in the ${s.label} area (${s.symbolCount} symbols) | \`.claude/skills/generated/${s.name}/SKILL.md\` |`
).join('\n')
: '';
const skillsTable = `| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | \`.claude/skills/gitnexus/gitnexus-exploring/SKILL.md\` |
| Blast radius / "What breaks if I change X?" | \`.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md\` |
| Trace bugs / "Why is X failing?" | \`.claude/skills/gitnexus/gitnexus-debugging/SKILL.md\` |
| Rename / extract / split / refactor | \`.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md\` |
| Tools, resources, schema reference | \`.claude/skills/gitnexus/gitnexus-guide/SKILL.md\` |
| Index, status, clean, wiki CLI commands | \`.claude/skills/gitnexus/gitnexus-cli/SKILL.md\` |${generatedRows ? '\n' + generatedRows : ''}`;
function generateGitNexusContent(projectName: string, stats: RepoStats): string {
return `${GITNEXUS_START_MARKER}
# GitNexus — Code Intelligence
@@ -125,27 +109,11 @@ Before completing any code modification task, verify:
3. \`gitnexus_detect_changes()\` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## Keeping the Index Fresh
After committing code changes, the GitNexus index becomes stale. Re-run analyze to update it:
\`\`\`bash
npx gitnexus analyze
\`\`\`
If the index previously included embeddings, preserve them by adding \`--embeddings\`:
\`\`\`bash
npx gitnexus analyze --embeddings
\`\`\`
To check whether embeddings exist, inspect \`.gitnexus/meta.json\` — the \`stats.embeddings\` field shows the count (0 means no embeddings). **Running analyze without \`--embeddings\` will delete any previously generated embeddings.**
> Claude Code users: A PostToolUse hook handles this automatically after \`git commit\` and \`git merge\`.
## CLI
${skillsTable}
- Re-index: \`npx gitnexus analyze\`
- Check freshness: \`npx gitnexus status\`
- Generate docs: \`npx gitnexus wiki\`
${GITNEXUS_END_MARKER}`;
}
@@ -284,10 +252,9 @@ export async function generateAIContextFiles(
repoPath: string,
_storagePath: string,
projectName: string,
stats: RepoStats,
generatedSkills?: GeneratedSkillInfo[]
stats: RepoStats
): Promise<{ files: string[] }> {
const content = generateGitNexusContent(projectName, stats, generatedSkills);
const content = generateGitNexusContent(projectName, stats);
const createdFiles: string[] = [];
// Create AGENTS.md (standard for Cursor, Windsurf, OpenCode, Cline, etc.)
+2 -24
View File
@@ -17,7 +17,6 @@ import { initKuzu, loadGraphToKuzu, getKuzuStats, executeQuery, executeWithReuse
import { getStoragePaths, saveMeta, loadMeta, addToGitignore, registerRepo, getGlobalRegistryPath } from '../storage/repo-manager.js';
import { getCurrentCommit, isGitRepo, getGitRoot } from '../storage/git.js';
import { generateAIContextFiles } from './ai-context.js';
import { generateSkillFiles, type GeneratedSkillInfo } from './skill-gen.js';
import fs from 'fs/promises';
@@ -46,8 +45,6 @@ function ensureHeap(): boolean {
export interface AnalyzeOptions {
force?: boolean;
embeddings?: boolean;
skills?: boolean;
verbose?: boolean;
}
/** Threshold: auto-skip embeddings for repos with more nodes than this */
@@ -75,10 +72,6 @@ export const analyzeCommand = async (
) => {
if (ensureHeap()) return;
if (options?.verbose) {
process.env.GITNEXUS_VERBOSE = '1';
}
console.log('\n GitNexus Analyzer\n');
let repoPath: string;
@@ -104,7 +97,7 @@ export const analyzeCommand = async (
const currentCommit = getCurrentCommit(repoPath);
const existingMeta = await loadMeta(storagePath);
if (existingMeta && !options?.force && !options?.skills && existingMeta.lastCommit === currentCommit) {
if (existingMeta && !options?.force && existingMeta.lastCommit === currentCommit) {
console.log(' Already up to date\n');
return;
}
@@ -283,13 +276,6 @@ export const analyzeCommand = async (
// ── Phase 5: Finalize (98–100%) ───────────────────────────────────
updateBar(98, 'Saving metadata...');
// Count embeddings in the index (cached + newly generated)
let embeddingCount = 0;
try {
const embResult = await executeQuery(`MATCH (e:CodeEmbedding) RETURN count(e) AS cnt`);
embeddingCount = embResult?.[0]?.cnt ?? 0;
} catch { /* table may not exist if embeddings never ran */ }
const meta = {
repoPath,
lastCommit: currentCommit,
@@ -300,7 +286,6 @@ export const analyzeCommand = async (
edges: stats.edges,
communities: pipelineResult.communityResult?.stats.totalCommunities,
processes: pipelineResult.processResult?.stats.totalProcesses,
embeddings: embeddingCount,
},
};
await saveMeta(storagePath, meta);
@@ -318,13 +303,6 @@ export const analyzeCommand = async (
aggregatedClusterCount = Array.from(groups.values()).filter(count => count >= 5).length;
}
let generatedSkills: GeneratedSkillInfo[] = [];
if (options?.skills && pipelineResult.communityResult) {
updateBar(99, 'Generating skill files...');
const skillResult = await generateSkillFiles(repoPath, projectName, pipelineResult);
generatedSkills = skillResult.skills;
}
const aiContext = await generateAIContextFiles(repoPath, storagePath, projectName, {
files: pipelineResult.totalFileCount,
nodes: stats.nodes,
@@ -332,7 +310,7 @@ export const analyzeCommand = async (
communities: pipelineResult.communityResult?.stats.totalCommunities,
clusters: aggregatedClusterCount,
processes: pipelineResult.processResult?.stats.totalProcesses,
}, generatedSkills);
});
await closeKuzu();
// Note: we intentionally do NOT call disposeEmbedder() here.
+25 -18
View File
@@ -4,9 +4,18 @@
// Removing it from here improves MCP server startup time significantly.
import { Command } from 'commander';
import { analyzeCommand } from './analyze.js';
import { serveCommand } from './serve.js';
import { listCommand } from './list.js';
import { statusCommand } from './status.js';
import { mcpCommand } from './mcp.js';
import { cleanCommand } from './clean.js';
import { setupCommand } from './setup.js';
import { augmentCommand } from './augment.js';
import { wikiCommand } from './wiki.js';
import { queryCommand, contextCommand, impactCommand, cypherCommand } from './tool.js';
import { evalServerCommand } from './eval-server.js';
import { createRequire } from 'node:module';
import { createLazyAction } from './lazy-action.js';
const _require = createRequire(import.meta.url);
const pkg = _require('../../package.json');
const program = new Command();
@@ -19,45 +28,43 @@ program
program
.command('setup')
.description('One-time setup: configure MCP for Cursor, Claude Code, OpenCode')
.action(createLazyAction(() => import('./setup.js'), 'setupCommand'));
.action(setupCommand);
program
.command('analyze [path]')
.description('Index a repository (full analysis)')
.option('-f, --force', 'Force full re-index even if up to date')
.option('--embeddings', 'Enable embedding generation for semantic search (off by default)')
.option('--skills', 'Generate repo-specific skill files from detected communities')
.option('-v, --verbose', 'Enable verbose ingestion warnings (default: false)')
.action(createLazyAction(() => import('./analyze.js'), 'analyzeCommand'));
.action(analyzeCommand);
program
.command('serve')
.description('Start local HTTP server for web UI connection')
.option('-p, --port <port>', 'Port number', '4747')
.option('--host <host>', 'Bind address (default: 127.0.0.1, use 0.0.0.0 for remote access)')
.action(createLazyAction(() => import('./serve.js'), 'serveCommand'));
.action(serveCommand);
program
.command('mcp')
.description('Start MCP server (stdio) — serves all indexed repos')
.action(createLazyAction(() => import('./mcp.js'), 'mcpCommand'));
.action(mcpCommand);
program
.command('list')
.description('List all indexed repositories')
.action(createLazyAction(() => import('./list.js'), 'listCommand'));
.action(listCommand);
program
.command('status')
.description('Show index status for current repo')
.action(createLazyAction(() => import('./status.js'), 'statusCommand'));
.action(statusCommand);
program
.command('clean')
.description('Delete GitNexus index for current repo')
.option('-f, --force', 'Skip confirmation prompt')
.option('--all', 'Clean all indexed repos')
.action(createLazyAction(() => import('./clean.js'), 'cleanCommand'));
.action(cleanCommand);
program
.command('wiki [path]')
@@ -68,12 +75,12 @@ program
.option('--api-key <key>', 'LLM API key (saved to ~/.gitnexus/config.json)')
.option('--concurrency <n>', 'Parallel LLM calls (default: 3)', '3')
.option('--gist', 'Publish wiki as a public GitHub Gist after generation')
.action(createLazyAction(() => import('./wiki.js'), 'wikiCommand'));
.action(wikiCommand);
program
.command('augment <pattern>')
.description('Augment a search pattern with knowledge graph context (used by hooks)')
.action(createLazyAction(() => import('./augment.js'), 'augmentCommand'));
.action(augmentCommand);
// ─── Direct Tool Commands (no MCP overhead) ────────────────────────
// These invoke LocalBackend directly for use in eval, scripts, and CI.
@@ -86,7 +93,7 @@ program
.option('-g, --goal <text>', 'What you want to find')
.option('-l, --limit <n>', 'Max processes to return (default: 5)')
.option('--content', 'Include full symbol source code')
.action(createLazyAction(() => import('./tool.js'), 'queryCommand'));
.action(queryCommand);
program
.command('context [name]')
@@ -95,7 +102,7 @@ program
.option('-u, --uid <uid>', 'Direct symbol UID (zero-ambiguity lookup)')
.option('-f, --file <path>', 'File path to disambiguate common names')
.option('--content', 'Include full symbol source code')
.action(createLazyAction(() => import('./tool.js'), 'contextCommand'));
.action(contextCommand);
program
.command('impact <target>')
@@ -104,13 +111,13 @@ program
.option('-r, --repo <name>', 'Target repository')
.option('--depth <n>', 'Max relationship depth (default: 3)')
.option('--include-tests', 'Include test files in results')
.action(createLazyAction(() => import('./tool.js'), 'impactCommand'));
.action(impactCommand);
program
.command('cypher <query>')
.description('Execute raw Cypher query against the knowledge graph')
.option('-r, --repo <name>', 'Target repository')
.action(createLazyAction(() => import('./tool.js'), 'cypherCommand'));
.action(cypherCommand);
// ─── Eval Server (persistent daemon for SWE-bench) ─────────────────
@@ -119,6 +126,6 @@ program
.description('Start lightweight HTTP server for fast tool calls during evaluation')
.option('-p, --port <port>', 'Port number', '4848')
.option('--idle-timeout <seconds>', 'Auto-shutdown after N seconds idle (0 = disabled)', '0')
.action(createLazyAction(() => import('./eval-server.js'), 'evalServerCommand'));
.action(evalServerCommand);
program.parse(process.argv);
-26
View File
@@ -1,26 +0,0 @@
/**
* Creates a lazy-loaded CLI action that defers module import until invocation.
* The generic constraints ensure the export name is a valid key of the module
* at compile time — catching typos when used with concrete module imports.
*/
function isCallable(value: unknown): value is (...args: unknown[]) => unknown {
return typeof value === 'function';
}
export function createLazyAction<
TModule extends Record<string, unknown>,
TKey extends string & keyof TModule,
>(
loader: () => Promise<TModule>,
exportName: TKey,
): (...args: unknown[]) => Promise<void> {
return async (...args: unknown[]): Promise<void> => {
const module = await loader();
const action = module[exportName];
if (!isCallable(action)) {
throw new Error(`Lazy action export not found: ${exportName}`);
}
await action(...args);
};
}
+33 -53
View File
@@ -10,7 +10,6 @@ import fs from 'fs/promises';
import path from 'path';
import os from 'os';
import { fileURLToPath } from 'url';
import { glob } from 'glob';
import { getGlobalDir } from '../storage/repo-manager.js';
const __filename = fileURLToPath(import.meta.url);
@@ -169,18 +168,16 @@ async function installClaudeCodeHooks(result: SetupResult): Promise<void> {
// even when it's no longer inside the npm package tree
const resolvedCli = path.join(__dirname, '..', 'cli', 'index.js');
const normalizedCli = path.resolve(resolvedCli).replace(/\\/g, '/');
const jsonCli = JSON.stringify(normalizedCli);
content = content.replace(
"let cliPath = path.resolve(__dirname, '..', '..', 'dist', 'cli', 'index.js');",
`let cliPath = ${jsonCli};`
`let cliPath = '${normalizedCli}';`
);
await fs.writeFile(dest, content, 'utf-8');
} catch {
// Script not found in source — skip
}
const hookPath = path.join(destHooksDir, 'gitnexus-hook.cjs').replace(/\\/g, '/');
const hookCmd = `node "${hookPath.replace(/"/g, '\\"')}"`;
const hookCmd = `node "${path.join(destHooksDir, 'gitnexus-hook.cjs').replace(/\\/g, '/')}"`;
// Merge hook config into ~/.claude/settings.json
const existing = await readJsonFile(settingsPath) || {};
@@ -189,31 +186,25 @@ async function installClaudeCodeHooks(result: SetupResult): Promise<void> {
// NOTE: SessionStart hooks are broken on Windows (Claude Code bug #23576).
// Session context is delivered via CLAUDE.md / skills instead.
// Helper: add a hook entry if one with 'gitnexus-hook' isn't already registered
interface HookEntry { hooks?: Array<{ command?: string }> }
function ensureHookEntry(
eventName: string,
matcher: string,
timeout: number,
statusMessage: string,
) {
if (!existing.hooks[eventName]) existing.hooks[eventName] = [];
const hasHook = existing.hooks[eventName].some(
(h: HookEntry) => h.hooks?.some(hh => hh.command?.includes('gitnexus-hook'))
);
if (!hasHook) {
existing.hooks[eventName].push({
matcher,
hooks: [{ type: 'command', command: hookCmd, timeout, statusMessage }],
});
}
// Add PreToolUse hook if not already present
if (!existing.hooks.PreToolUse) existing.hooks.PreToolUse = [];
const hasPreToolHook = existing.hooks.PreToolUse.some(
(h: any) => h.hooks?.some((hh: any) => hh.command?.includes('gitnexus'))
);
if (!hasPreToolHook) {
existing.hooks.PreToolUse.push({
matcher: 'Grep|Glob|Bash',
hooks: [{
type: 'command',
command: hookCmd,
timeout: 8000,
statusMessage: 'Enriching with GitNexus graph context...',
}],
});
}
ensureHookEntry('PreToolUse', 'Grep|Glob|Bash', 10, 'Enriching with GitNexus graph context...');
ensureHookEntry('PostToolUse', 'Bash', 10, 'Checking GitNexus index freshness...');
await writeJsonFile(settingsPath, existing);
result.configured.push('Claude Code hooks (PreToolUse, PostToolUse)');
result.configured.push('Claude Code hooks (PreToolUse)');
} catch (err: any) {
result.errors.push(`Claude Code hooks: ${err.message}`);
}
@@ -241,6 +232,8 @@ async function setupOpenCode(result: SetupResult): Promise<void> {
// ─── Skill Installation ───────────────────────────────────────────
const SKILL_NAMES = ['gitnexus-exploring', 'gitnexus-debugging', 'gitnexus-impact-analysis', 'gitnexus-refactoring', 'gitnexus-guide', 'gitnexus-cli'];
/**
* Install GitNexus skills to a target directory.
* Each skill is installed as {targetDir}/gitnexus-{skillName}/SKILL.md
@@ -254,38 +247,25 @@ async function installSkillsTo(targetDir: string): Promise<string[]> {
const installed: string[] = [];
const skillsRoot = path.join(__dirname, '..', '..', 'skills');
let flatFiles: string[] = [];
let dirSkillFiles: string[] = [];
try {
[flatFiles, dirSkillFiles] = await Promise.all([
glob('*.md', { cwd: skillsRoot }),
glob('*/SKILL.md', { cwd: skillsRoot }),
]);
} catch {
return [];
}
const skillSources = new Map<string, { isDirectory: boolean }>();
for (const relPath of dirSkillFiles) {
skillSources.set(path.dirname(relPath), { isDirectory: true });
}
for (const relPath of flatFiles) {
const skillName = path.basename(relPath, '.md');
if (!skillSources.has(skillName)) {
skillSources.set(skillName, { isDirectory: false });
}
}
for (const [skillName, source] of skillSources) {
for (const skillName of SKILL_NAMES) {
const skillDir = path.join(targetDir, skillName);
try {
if (source.isDirectory) {
const dirSource = path.join(skillsRoot, skillName);
// Try directory-based skill first (skills/{name}/SKILL.md)
const dirSource = path.join(skillsRoot, skillName);
const dirSkillFile = path.join(dirSource, 'SKILL.md');
let isDirectory = false;
try {
const stat = await fs.stat(dirSource);
isDirectory = stat.isDirectory();
} catch { /* not a directory */ }
if (isDirectory) {
await copyDirRecursive(dirSource, skillDir);
installed.push(skillName);
} else {
// Fall back to flat file (skills/{name}.md)
const flatSource = path.join(skillsRoot, `${skillName}.md`);
const content = await fs.readFile(flatSource, 'utf-8');
await fs.mkdir(skillDir, { recursive: true });
-712
View File
@@ -1,712 +0,0 @@
/**
* Skill File Generator
*
* Generates repo-specific SKILL.md files from detected Leiden communities.
* Each significant community becomes a skill that describes a functional area
* of the codebase, including key files, entry points, execution flows, and
* cross-community connections.
*/
import fs from 'fs/promises';
import path from 'path';
import { PipelineResult } from '../types/pipeline.js';
import { CommunityNode, CommunityMembership } from '../core/ingestion/community-processor.js';
import { ProcessNode } from '../core/ingestion/process-processor.js';
import { GraphNode, KnowledgeGraph } from '../core/graph/types.js';
// ============================================================================
// TYPES
// ============================================================================
export interface GeneratedSkillInfo {
name: string;
label: string;
symbolCount: number;
fileCount: number;
}
interface AggregatedCommunity {
label: string;
rawIds: string[];
symbolCount: number;
cohesion: number;
}
interface MemberSymbol {
id: string;
name: string;
label: string;
filePath: string;
startLine: number;
isExported: boolean;
}
interface FileInfo {
relativePath: string;
symbols: string[];
}
interface CrossConnection {
targetLabel: string;
count: number;
}
// ============================================================================
// MAIN EXPORT
// ============================================================================
/**
* @brief Generate repo-specific skill files from detected communities
* @param {string} repoPath - Absolute path to the repository root
* @param {string} projectName - Human-readable project name
* @param {PipelineResult} pipelineResult - In-memory pipeline data with communities, processes, graph
* @returns {Promise<{ skills: GeneratedSkillInfo[], outputPath: string }>} Generated skill metadata
*/
export const generateSkillFiles = async (
repoPath: string,
projectName: string,
pipelineResult: PipelineResult
): Promise<{ skills: GeneratedSkillInfo[]; outputPath: string }> => {
const { communityResult, processResult, graph } = pipelineResult;
const outputDir = path.join(repoPath, '.claude', 'skills', 'generated');
if (!communityResult || !communityResult.memberships.length) {
console.log('\n Skills: no communities detected, skipping skill generation');
return { skills: [], outputPath: outputDir };
}
console.log('\n Generating repo-specific skills...');
// Step 1: Build communities from memberships (not the filtered communities array).
// The community processor skips singletons from its communities array but memberships
// include ALL assignments. For repos with sparse CALLS edges, the communities array
// can be empty while memberships still has useful groupings.
const communities = communityResult.communities.length > 0
? communityResult.communities
: buildCommunitiesFromMemberships(communityResult.memberships, graph, repoPath);
const aggregated = aggregateCommunities(communities);
// Step 2: Filter to significant communities
// Keep communities with >= 3 symbols after aggregation.
const significant = aggregated
.filter(c => c.symbolCount >= 3)
.sort((a, b) => b.symbolCount - a.symbolCount)
.slice(0, 20);
if (significant.length === 0) {
console.log('\n Skills: no significant communities found (all below 3-symbol threshold)');
return { skills: [], outputPath: outputDir };
}
// Step 3: Build lookup maps
const membershipsByComm = buildMembershipMap(communityResult.memberships);
const nodeIdToCommunityLabel = buildNodeCommunityLabelMap(
communityResult.memberships,
communities
);
// Step 4: Clear and recreate output directory
try {
await fs.rm(outputDir, { recursive: true, force: true });
} catch { /* may not exist */ }
await fs.mkdir(outputDir, { recursive: true });
// Step 5: Generate skill files
const skills: GeneratedSkillInfo[] = [];
const usedNames = new Set<string>();
for (const community of significant) {
// Gather member symbols
const members = gatherMembers(community.rawIds, membershipsByComm, graph);
if (members.length === 0) continue;
// Gather file info
const files = gatherFiles(members, repoPath);
// Gather entry points
const entryPoints = gatherEntryPoints(members);
// Gather execution flows
const flows = gatherFlows(community.rawIds, processResult?.processes || []);
// Gather cross-community connections
const connections = gatherCrossConnections(
community.rawIds,
community.label,
membershipsByComm,
nodeIdToCommunityLabel,
graph
);
// Generate kebab name
const kebabName = toKebabName(community.label, usedNames);
usedNames.add(kebabName);
// Generate SKILL.md content
const content = renderSkillMarkdown(
community,
projectName,
members,
files,
entryPoints,
flows,
connections,
kebabName
);
// Write file
const skillDir = path.join(outputDir, kebabName);
await fs.mkdir(skillDir, { recursive: true });
await fs.writeFile(path.join(skillDir, 'SKILL.md'), content, 'utf-8');
const info: GeneratedSkillInfo = {
name: kebabName,
label: community.label,
symbolCount: community.symbolCount,
fileCount: files.length,
};
skills.push(info);
console.log(` \u2713 ${community.label} (${community.symbolCount} symbols, ${files.length} files)`);
}
console.log(`\n ${skills.length} skills generated \u2192 .claude/skills/generated/`);
return { skills, outputPath: outputDir };
};
// ============================================================================
// FALLBACK COMMUNITY BUILDER
// ============================================================================
/**
* @brief Build CommunityNode-like objects from raw memberships when the community
* processor's communities array is empty (all singletons were filtered out)
* @param {CommunityMembership[]} memberships - All node-to-community assignments
* @param {KnowledgeGraph} graph - The knowledge graph for resolving node metadata
* @param {string} repoPath - Repository root for path normalization
* @returns {CommunityNode[]} Synthetic community nodes built from membership data
*/
const buildCommunitiesFromMemberships = (
memberships: CommunityMembership[],
graph: KnowledgeGraph,
repoPath: string
): CommunityNode[] => {
// Group memberships by communityId
const groups = new Map<string, string[]>();
for (const m of memberships) {
const arr = groups.get(m.communityId);
if (arr) {
arr.push(m.nodeId);
} else {
groups.set(m.communityId, [m.nodeId]);
}
}
const communities: CommunityNode[] = [];
for (const [commId, nodeIds] of groups) {
// Derive a heuristic label from the most common parent directory
const folderCounts = new Map<string, number>();
for (const nodeId of nodeIds) {
const node = graph.getNode(nodeId);
if (!node?.properties.filePath) continue;
const normalized = node.properties.filePath.replace(/\\/g, '/');
const parts = normalized.split('/').filter(Boolean);
if (parts.length >= 2) {
const folder = parts[parts.length - 2];
if (!['src', 'lib', 'core', 'utils', 'common', 'shared', 'helpers'].includes(folder.toLowerCase())) {
folderCounts.set(folder, (folderCounts.get(folder) || 0) + 1);
}
}
}
let bestFolder = '';
let bestCount = 0;
for (const [folder, count] of folderCounts) {
if (count > bestCount) {
bestCount = count;
bestFolder = folder;
}
}
const label = bestFolder
? bestFolder.charAt(0).toUpperCase() + bestFolder.slice(1)
: `Cluster_${commId.replace('comm_', '')}`;
// Compute cohesion as internal-edge ratio (matches backend calculateCohesion).
// For each member node, count edges that stay inside the community vs total.
const nodeSet = new Set(nodeIds);
let internalEdges = 0;
let totalEdges = 0;
graph.forEachRelationship(rel => {
if (nodeSet.has(rel.sourceId)) {
totalEdges++;
if (nodeSet.has(rel.targetId)) internalEdges++;
}
});
const cohesion = totalEdges > 0 ? Math.min(1.0, internalEdges / totalEdges) : 1.0;
communities.push({
id: commId,
label,
heuristicLabel: label,
cohesion,
symbolCount: nodeIds.length,
});
}
return communities.sort((a, b) => b.symbolCount - a.symbolCount);
};
// ============================================================================
// AGGREGATION
// ============================================================================
/**
* @brief Aggregate raw Leiden communities by heuristicLabel
* @param {CommunityNode[]} communities - Raw community nodes from Leiden detection
* @returns {AggregatedCommunity[]} Aggregated communities grouped by label
*/
const aggregateCommunities = (communities: CommunityNode[]): AggregatedCommunity[] => {
const groups = new Map<string, {
rawIds: string[];
totalSymbols: number;
weightedCohesion: number;
}>();
for (const c of communities) {
const label = c.heuristicLabel || c.label || 'Unknown';
const symbols = c.symbolCount || 0;
const cohesion = c.cohesion || 0;
const existing = groups.get(label);
if (!existing) {
groups.set(label, {
rawIds: [c.id],
totalSymbols: symbols,
weightedCohesion: cohesion * symbols,
});
} else {
existing.rawIds.push(c.id);
existing.totalSymbols += symbols;
existing.weightedCohesion += cohesion * symbols;
}
}
return Array.from(groups.entries()).map(([label, g]) => ({
label,
rawIds: g.rawIds,
symbolCount: g.totalSymbols,
cohesion: g.totalSymbols > 0 ? g.weightedCohesion / g.totalSymbols : 0,
}));
};
// ============================================================================
// LOOKUP MAP BUILDERS
// ============================================================================
/**
* @brief Build a map from communityId to member nodeIds
* @param {CommunityMembership[]} memberships - All membership records
* @returns {Map<string, string[]>} Map of communityId -> nodeId[]
*/
const buildMembershipMap = (memberships: CommunityMembership[]): Map<string, string[]> => {
const map = new Map<string, string[]>();
for (const m of memberships) {
const arr = map.get(m.communityId);
if (arr) {
arr.push(m.nodeId);
} else {
map.set(m.communityId, [m.nodeId]);
}
}
return map;
};
/**
* @brief Build a map from nodeId to aggregated community label
* @param {CommunityMembership[]} memberships - All membership records
* @param {CommunityNode[]} communities - Community nodes with labels
* @returns {Map<string, string>} Map of nodeId -> community label
*/
const buildNodeCommunityLabelMap = (
memberships: CommunityMembership[],
communities: CommunityNode[]
): Map<string, string> => {
const commIdToLabel = new Map<string, string>();
for (const c of communities) {
commIdToLabel.set(c.id, c.heuristicLabel || c.label || 'Unknown');
}
const map = new Map<string, string>();
for (const m of memberships) {
const label = commIdToLabel.get(m.communityId);
if (label) {
map.set(m.nodeId, label);
}
}
return map;
};
// ============================================================================
// DATA GATHERING
// ============================================================================
/**
* @brief Gather member symbols for an aggregated community
* @param {string[]} rawIds - Raw community IDs belonging to this aggregated community
* @param {Map<string, string[]>} membershipsByComm - communityId -> nodeIds
* @param {KnowledgeGraph} graph - The knowledge graph
* @returns {MemberSymbol[]} Array of member symbol information
*/
const gatherMembers = (
rawIds: string[],
membershipsByComm: Map<string, string[]>,
graph: KnowledgeGraph
): MemberSymbol[] => {
const seen = new Set<string>();
const members: MemberSymbol[] = [];
for (const commId of rawIds) {
const nodeIds = membershipsByComm.get(commId) || [];
for (const nodeId of nodeIds) {
if (seen.has(nodeId)) continue;
seen.add(nodeId);
const node = graph.getNode(nodeId);
if (!node) continue;
members.push({
id: node.id,
name: node.properties.name,
label: node.label,
filePath: node.properties.filePath || '',
startLine: node.properties.startLine || 0,
isExported: node.properties.isExported === true,
});
}
}
return members;
};
/**
* @brief Gather deduplicated file info with per-file symbol names
* @param {MemberSymbol[]} members - Member symbols
* @param {string} repoPath - Repository root for relative path computation
* @returns {FileInfo[]} Sorted by symbol count descending
*/
const gatherFiles = (members: MemberSymbol[], repoPath: string): FileInfo[] => {
const fileMap = new Map<string, string[]>();
for (const m of members) {
if (!m.filePath) continue;
const rel = toRelativePath(m.filePath, repoPath);
const arr = fileMap.get(rel);
if (arr) {
arr.push(m.name);
} else {
fileMap.set(rel, [m.name]);
}
}
return Array.from(fileMap.entries())
.map(([relativePath, symbols]) => ({ relativePath, symbols }))
.sort((a, b) => b.symbols.length - a.symbols.length);
};
/**
* @brief Gather exported entry points prioritized by type
* @param {MemberSymbol[]} members - Member symbols
* @returns {MemberSymbol[]} Exported symbols sorted by type priority
*/
const gatherEntryPoints = (members: MemberSymbol[]): MemberSymbol[] => {
const typePriority: Record<string, number> = {
Function: 0,
Class: 1,
Method: 2,
Interface: 3,
};
return members
.filter(m => m.isExported)
.sort((a, b) => {
const pa = typePriority[a.label] ?? 99;
const pb = typePriority[b.label] ?? 99;
return pa - pb;
});
};
/**
* @brief Gather execution flows touching this community
* @param {string[]} rawIds - Raw community IDs for this aggregated community
* @param {ProcessNode[]} processes - All detected processes
* @returns {ProcessNode[]} Processes whose communities intersect rawIds, sorted by stepCount
*/
const gatherFlows = (rawIds: string[], processes: ProcessNode[]): ProcessNode[] => {
const rawIdSet = new Set(rawIds);
return processes
.filter(proc => proc.communities.some(cid => rawIdSet.has(cid)))
.sort((a, b) => b.stepCount - a.stepCount);
};
/**
* @brief Gather cross-community call connections
* @param {string[]} rawIds - Raw community IDs for this aggregated community
* @param {string} ownLabel - This community's aggregated label
* @param {Map<string, string[]>} membershipsByComm - communityId -> nodeIds
* @param {Map<string, string>} nodeIdToCommunityLabel - nodeId -> community label
* @param {KnowledgeGraph} graph - The knowledge graph
* @returns {CrossConnection[]} Aggregated cross-community connections sorted by count
*/
const gatherCrossConnections = (
rawIds: string[],
ownLabel: string,
membershipsByComm: Map<string, string[]>,
nodeIdToCommunityLabel: Map<string, string>,
graph: KnowledgeGraph
): CrossConnection[] => {
// Collect all node IDs in this aggregated community
const ownNodeIds = new Set<string>();
for (const commId of rawIds) {
const nodeIds = membershipsByComm.get(commId) || [];
for (const nid of nodeIds) {
ownNodeIds.add(nid);
}
}
// Count outgoing CALLS to nodes in different communities
const targetCounts = new Map<string, number>();
graph.forEachRelationship(rel => {
if (rel.type !== 'CALLS') return;
if (!ownNodeIds.has(rel.sourceId)) return;
if (ownNodeIds.has(rel.targetId)) return; // same community
const targetLabel = nodeIdToCommunityLabel.get(rel.targetId);
if (!targetLabel || targetLabel === ownLabel) return;
targetCounts.set(targetLabel, (targetCounts.get(targetLabel) || 0) + 1);
});
return Array.from(targetCounts.entries())
.map(([targetLabel, count]) => ({ targetLabel, count }))
.sort((a, b) => b.count - a.count);
};
// ============================================================================
// MARKDOWN RENDERING
// ============================================================================
/**
* @brief Render SKILL.md content for a single community
* @param {AggregatedCommunity} community - The aggregated community data
* @param {string} projectName - Project name for the description
* @param {MemberSymbol[]} members - All member symbols
* @param {FileInfo[]} files - File info with symbol names
* @param {MemberSymbol[]} entryPoints - Exported entry point symbols
* @param {ProcessNode[]} flows - Execution flows touching this community
* @param {CrossConnection[]} connections - Cross-community connections
* @param {string} kebabName - Kebab-case name for the skill
* @returns {string} Full SKILL.md content
*/
const renderSkillMarkdown = (
community: AggregatedCommunity,
projectName: string,
members: MemberSymbol[],
files: FileInfo[],
entryPoints: MemberSymbol[],
flows: ProcessNode[],
connections: CrossConnection[],
kebabName: string
): string => {
const cohesionPct = Math.round(community.cohesion * 100);
// Dominant directory: most common top-level directory
const dominantDir = getDominantDirectory(files);
// Top symbol names for "When to Use"
const topNames = entryPoints.slice(0, 3).map(e => e.name);
if (topNames.length === 0) {
// Fallback to any members
topNames.push(...members.slice(0, 3).map(m => m.name));
}
const lines: string[] = [];
// Frontmatter
lines.push('---');
lines.push(`name: ${kebabName}`);
lines.push(`description: "Skill for the ${community.label} area of ${projectName}. ${community.symbolCount} symbols across ${files.length} files."`);
lines.push('---');
lines.push('');
// Title
lines.push(`# ${community.label}`);
lines.push('');
lines.push(`${community.symbolCount} symbols | ${files.length} files | Cohesion: ${cohesionPct}%`);
lines.push('');
// When to Use
lines.push('## When to Use');
lines.push('');
if (dominantDir) {
lines.push(`- Working with code in \`${dominantDir}/\``);
}
if (topNames.length > 0) {
lines.push(`- Understanding how ${topNames.join(', ')} work`);
}
lines.push(`- Modifying ${community.label.toLowerCase()}-related functionality`);
lines.push('');
// Key Files (top 10)
lines.push('## Key Files');
lines.push('');
lines.push('| File | Symbols |');
lines.push('|------|---------|');
for (const f of files.slice(0, 10)) {
const symbolList = f.symbols.slice(0, 5).join(', ');
const suffix = f.symbols.length > 5 ? ` (+${f.symbols.length - 5})` : '';
lines.push(`| \`${f.relativePath}\` | ${symbolList}${suffix} |`);
}
lines.push('');
// Entry Points (top 5)
if (entryPoints.length > 0) {
lines.push('## Entry Points');
lines.push('');
lines.push('Start here when exploring this area:');
lines.push('');
for (const ep of entryPoints.slice(0, 5)) {
lines.push(`- **\`${ep.name}\`** (${ep.label}) \u2014 \`${ep.filePath}:${ep.startLine}\``);
}
lines.push('');
}
// Key Symbols (top 20, exported first, then by type)
lines.push('## Key Symbols');
lines.push('');
lines.push('| Symbol | Type | File | Line |');
lines.push('|--------|------|------|------|');
const sortedMembers = [...members].sort((a, b) => {
if (a.isExported !== b.isExported) return a.isExported ? -1 : 1;
return a.label.localeCompare(b.label);
});
for (const m of sortedMembers.slice(0, 20)) {
lines.push(`| \`${m.name}\` | ${m.label} | \`${m.filePath}\` | ${m.startLine} |`);
}
lines.push('');
// Execution Flows
if (flows.length > 0) {
lines.push('## Execution Flows');
lines.push('');
lines.push('| Flow | Type | Steps |');
lines.push('|------|------|-------|');
for (const f of flows.slice(0, 10)) {
lines.push(`| \`${f.heuristicLabel}\` | ${f.processType} | ${f.stepCount} |`);
}
lines.push('');
}
// Connected Areas
if (connections.length > 0) {
lines.push('## Connected Areas');
lines.push('');
lines.push('| Area | Connections |');
lines.push('|------|-------------|');
for (const c of connections.slice(0, 8)) {
lines.push(`| ${c.targetLabel} | ${c.count} calls |`);
}
lines.push('');
}
// How to Explore
const firstEntry = entryPoints.length > 0 ? entryPoints[0].name : (members.length > 0 ? members[0].name : community.label);
lines.push('## How to Explore');
lines.push('');
lines.push(`1. \`gitnexus_context({name: "${firstEntry}"})\` \u2014 see callers and callees`);
lines.push(`2. \`gitnexus_query({query: "${community.label.toLowerCase()}"})\` \u2014 find related execution flows`);
lines.push('3. Read key files listed above for implementation details');
lines.push('');
return lines.join('\n');
};
// ============================================================================
// UTILITY HELPERS
// ============================================================================
/**
* @brief Convert a community label to a kebab-case directory name
* @param {string} label - The community label
* @param {Set<string>} usedNames - Already-used names for collision detection
* @returns {string} Unique kebab-case name capped at 50 characters
*/
const toKebabName = (label: string, usedNames: Set<string>): string => {
let name = label
.toLowerCase()
.replace(/[^a-z0-9]+/g, '-')
.replace(/^-+|-+$/g, '')
.slice(0, 50);
if (!name) name = 'skill';
let candidate = name;
let counter = 2;
while (usedNames.has(candidate)) {
candidate = `${name}-${counter}`;
counter++;
}
return candidate;
};
/**
* @brief Convert an absolute or repo-relative file path to a clean relative path
* @param {string} filePath - The file path from the graph node
* @param {string} repoPath - Repository root path
* @returns {string} Relative path using forward slashes
*/
const toRelativePath = (filePath: string, repoPath: string): string => {
// Normalize to forward slashes for cross-platform consistency
const normalizedFile = filePath.replace(/\\/g, '/');
const normalizedRepo = repoPath.replace(/\\/g, '/');
if (normalizedFile.startsWith(normalizedRepo)) {
return normalizedFile.slice(normalizedRepo.length).replace(/^\//, '');
}
// Already relative or different root
return normalizedFile.replace(/^\//, '');
};
/**
* @brief Find the dominant (most common) top-level directory across files
* @param {FileInfo[]} files - File info entries
* @returns {string | null} Most common directory or null
*/
const getDominantDirectory = (files: FileInfo[]): string | null => {
const dirCounts = new Map<string, number>();
for (const f of files) {
const parts = f.relativePath.split('/');
if (parts.length >= 2) {
const dir = parts[0];
dirCounts.set(dir, (dirCounts.get(dir) || 0) + f.symbols.length);
}
}
let best: string | null = null;
let bestCount = 0;
for (const [dir, count] of dirCounts) {
if (count > bestCount) {
bestCount = count;
best = dir;
}
}
return best;
};
+6 -12
View File
@@ -35,14 +35,12 @@ export type NodeLabel =
| 'Template';
import { SupportedLanguages } from '../../config/supported-languages.js';
export type NodeProperties = {
name: string,
filePath: string,
startLine?: number,
endLine?: number,
language?: SupportedLanguages,
language?: string,
isExported?: boolean,
// Optional AST-derived framework hint (e.g. @Controller, @GetMapping)
astFrameworkMultiplier?: number,
@@ -63,23 +61,19 @@ export type NodeProperties = {
// Entry point scoring (computed by process detection)
entryPointScore?: number,
entryPointReason?: string,
// Method signature (for MRO disambiguation)
parameterCount?: number,
returnType?: string,
}
export type RelationshipType =
| 'CONTAINS'
| 'CALLS'
| 'INHERITS'
| 'OVERRIDES'
export type RelationshipType =
| 'CONTAINS'
| 'CALLS'
| 'INHERITS'
| 'OVERRIDES'
| 'IMPORTS'
| 'USES'
| 'DEFINES'
| 'DECORATES'
| 'IMPLEMENTS'
| 'EXTENDS'
| 'HAS_METHOD'
| 'MEMBER_OF'
| 'STEP_IN_PROCESS'
+2 -3
View File
@@ -10,11 +10,10 @@ export interface ASTCache {
}
export const createASTCache = (maxSize: number = 50): ASTCache => {
const effectiveMax = Math.max(maxSize, 1);
// Initialize the cache with a 'dispose' handler
// This is the magic: When an item is evicted (dropped), this runs automatically.
const cache = new LRUCache<string, Parser.Tree>({
max: effectiveMax,
max: maxSize,
dispose: (tree) => {
try {
// NOTE: web-tree-sitter has tree.delete(); native tree-sitter trees are GC-managed.
@@ -42,7 +41,7 @@ export const createASTCache = (maxSize: number = 50): ASTCache => {
stats: () => ({
size: cache.size,
maxSize: effectiveMax
maxSize: maxSize
})
};
};
+278 -243
View File
@@ -1,28 +1,51 @@
import { KnowledgeGraph } from '../graph/types.js';
import { ASTCache } from './ast-cache.js';
import type { SymbolDefinition, SymbolTable } from './symbol-table.js';
import { ImportMap, PackageMap, NamedImportMap, isFileInPackageDir } from './import-processor.js';
import { resolveSymbol, resolveSymbolInternal } from './symbol-resolver.js';
import { walkBindingChain } from './named-binding-extraction.js';
import { SymbolTable } from './symbol-table.js';
import { ImportMap } from './import-processor.js';
import Parser from 'tree-sitter';
import { isLanguageAvailable, loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { generateId } from '../../lib/utils.js';
import {
getLanguageFromFilename,
isVerboseIngestionEnabled,
yieldToEventLoop,
FUNCTION_NODE_TYPES,
extractFunctionName,
isBuiltInOrNoise,
countCallArguments,
inferCallForm,
extractReceiverName,
} from './utils.js';
import { buildTypeEnv, lookupTypeEnv } from './type-env.js';
import { getTreeSitterBufferSize } from './constants.js';
import { getLanguageFromFilename, yieldToEventLoop } from './utils.js';
import type { ExtractedCall, ExtractedRoute } from './workers/parse-worker.js';
/**
* Node types that represent function/method definitions across languages.
* Used to find the enclosing function for a call site.
*/
const FUNCTION_NODE_TYPES = new Set([
// TypeScript/JavaScript
'function_declaration',
'arrow_function',
'function_expression',
'method_definition',
'generator_function_declaration',
// Python
'function_definition',
// Common async variants
'async_function_declaration',
'async_arrow_function',
// Java
'method_declaration',
'constructor_declaration',
// C/C++
// 'function_definition' already included above
// Go
// 'method_declaration' already included from Java
// C#
'local_function_statement',
// Rust
'function_item',
'impl_item', // Methods inside impl blocks
// Kotlin (function_declaration already included above via JS/TS)
'anonymous_function',
'lambda_literal',
// PHP — no additional node types needed
// Swift
'init_declaration',
'deinit_declaration',
]);
/**
* Walk up the AST from a node to find the enclosing function/method.
* Returns null if the call is at module/file level (top-level code).
@@ -33,22 +56,89 @@ const findEnclosingFunction = (
symbolTable: SymbolTable
): string | null => {
let current = node.parent;
while (current) {
if (FUNCTION_NODE_TYPES.has(current.type)) {
const { funcName, label } = extractFunctionName(current);
// Found enclosing function - try to get its name
let funcName: string | null = null;
let label = 'Function';
// Different node types have different name locations
// Swift init/deinit — handle before generic cases (more specific)
if (current.type === 'init_declaration' || current.type === 'deinit_declaration') {
const funcName = current.type === 'init_declaration' ? 'init' : 'deinit';
return generateId('Constructor', `${filePath}:${funcName}`);
}
if (current.type === 'function_declaration' ||
current.type === 'function_definition' ||
current.type === 'async_function_declaration' ||
current.type === 'generator_function_declaration' ||
current.type === 'function_item') { // Rust function
// Named function: function foo() {}
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier' || c.type === 'property_identifier');
funcName = nameNode?.text;
} else if (current.type === 'impl_item') {
// Rust method inside impl block: wrapper around function_item or const_item
// We need to look inside for the function_item
const funcItem = current.children?.find((c: any) => c.type === 'function_item');
if (funcItem) {
const nameNode = funcItem.childForFieldName?.('name') ||
funcItem.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
}
} else if (current.type === 'method_definition') {
// Method: foo() {} inside class (JS/TS)
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'property_identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'method_declaration') {
// Java method: public void foo() {}
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'constructor_declaration') {
// Java constructor: public ClassName() {}
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method'; // Treat constructors as methods for process detection
} else if (current.type === 'arrow_function' || current.type === 'function_expression') {
// Arrow/expression: const foo = () => {} - check parent variable declarator
const parent = current.parent;
if (parent?.type === 'variable_declarator') {
const nameNode = parent.childForFieldName?.('name') ||
parent.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
}
}
if (funcName) {
// Look up the function in symbol table to get its node ID
// Try exact match first
const nodeId = symbolTable.lookupExact(filePath, funcName);
if (nodeId) return nodeId;
return generateId(label, `${filePath}:${funcName}`);
// Try construct ID manually if lookup fails (common for non-exported internal functions)
// Format should match what parsing-processor generates: "Function:path/to/file:funcName"
// Check if we already have a node with this ID in the symbol table to be safe
const generatedId = generateId(label, `${filePath}:${funcName}`);
// Ideally we should verify this ID exists, but strictly speaking if we are inside it,
// it SHOULD exist. Returning it is better than falling back to File.
return generatedId;
}
// Couldn't determine function name - try parent (might be nested)
}
current = current.parent;
}
return null;
return null; // Top-level call (not inside any function)
};
export const processCalls = async (
@@ -57,13 +147,9 @@ export const processCalls = async (
astCache: ASTCache,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void,
namedImportMap?: NamedImportMap,
onProgress?: (current: number, total: number) => void
) => {
const parser = await loadParser();
const logSkipped = isVerboseIngestionEnabled();
const skippedByLang = logSkipped ? new Map<string, number>() : null;
for (let i = 0; i < files.length; i++) {
const file = files[i];
@@ -73,12 +159,6 @@ export const processCalls = async (
// 1. Check language support first
const language = getLanguageFromFilename(file.path);
if (!language) continue;
if (!isLanguageAvailable(language)) {
if (skippedByLang) {
skippedByLang.set(language, (skippedByLang.get(language) ?? 0) + 1);
}
continue;
}
const queryStr = LANGUAGE_QUERIES[language];
if (!queryStr) continue;
@@ -94,7 +174,7 @@ export const processCalls = async (
// Cache Miss: Re-parse
// Use larger bufferSize for files > 32KB
try {
tree = parser.parse(file.content, undefined, { bufferSize: getTreeSitterBufferSize(file.content.length) });
tree = parser.parse(file.content, undefined, { bufferSize: 1024 * 256 });
} catch (parseError) {
// Skip files that can't be parsed
continue;
@@ -115,10 +195,6 @@ export const processCalls = async (
continue;
}
// Build per-file TypeEnv for receiver resolution
const lang = getLanguageFromFilename(file.path);
const typeEnv = lang ? buildTypeEnv(tree, lang) : new Map();
// 3. Process each call match
matches.forEach(match => {
const captureMap: Record<string, any> = {};
@@ -135,22 +211,18 @@ export const processCalls = async (
// Skip common built-ins and noise
if (isBuiltInOrNoise(calledName)) return;
const callNode = captureMap['call'];
const callForm = inferCallForm(callNode, nameNode);
const receiverName = callForm === 'member' ? extractReceiverName(nameNode) : undefined;
const receiverTypeName = receiverName ? lookupTypeEnv(typeEnv, receiverName, callNode) : undefined;
// 4. Resolve the target using priority strategy (returns confidence)
const resolved = resolveCallTarget({
const resolved = resolveCallTarget(
calledName,
argCount: countCallArguments(callNode),
callForm,
receiverTypeName,
}, file.path, symbolTable, importMap, packageMap, namedImportMap);
file.path,
symbolTable,
importMap
);
if (!resolved) return;
// 5. Find the enclosing function (caller)
const callNode = captureMap['call'];
const enclosingFuncId = findEnclosingFunction(callNode, file.path, symbolTable);
// Use enclosing function as source, fallback to file for top-level calls
@@ -170,14 +242,6 @@ export const processCalls = async (
// Tree is now owned by the LRU cache — no manual delete needed
}
if (skippedByLang && skippedByLang.size > 0) {
for (const [lang, count] of skippedByLang.entries()) {
console.warn(
`[ingestion] Skipped ${count} ${lang} file(s) in call processing — ${lang} parser not available.`
);
}
}
};
/**
@@ -186,169 +250,156 @@ export const processCalls = async (
interface ResolveResult {
nodeId: string;
confidence: number; // 0-1: how sure are we?
reason: string; // 'import-resolved' | 'same-file' | 'unique-global'
reason: string; // 'import-resolved' | 'same-file' | 'fuzzy-global'
}
type ResolutionTier = 'same-file' | 'import-scoped' | 'unique-global';
interface TieredCandidates {
candidates: SymbolDefinition[];
tier: ResolutionTier;
}
const CALLABLE_SYMBOL_TYPES = new Set([
'Function',
'Method',
'Constructor',
'Macro',
'Delegate',
]);
const collectTieredCandidates = (
calledName: string,
currentFile: string,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
): TieredCandidates | null => {
const allDefs = symbolTable.lookupFuzzy(calledName);
// Tier 1: Same-file — highest priority, prevents imports from shadowing local defs
// (matches resolveSymbolInternal which checks lookupExactFull before named bindings)
const localDefs = allDefs.filter(def => def.filePath === currentFile);
if (localDefs.length > 0) {
return { candidates: localDefs, tier: 'same-file' };
}
// Tier 2a-named: Check named bindings with re-export chain following.
// Aliased imports (import { User as U }) mean lookupFuzzy('U') returns
// empty but we can resolve via the exported name.
// Re-exports (export { User } from './base') are followed up to 5 hops.
if (namedImportMap) {
const chainResult = resolveNamedBindingChainForCandidates(
calledName, currentFile, symbolTable, namedImportMap, allDefs,
);
if (chainResult) return chainResult;
}
if (allDefs.length === 0) return null;
const importedFiles = importMap.get(currentFile);
if (importedFiles) {
const importedDefs = allDefs.filter(def => importedFiles.has(def.filePath));
if (importedDefs.length > 0) {
return { candidates: importedDefs, tier: 'import-scoped' };
}
}
const importedPackages = packageMap?.get(currentFile);
if (importedPackages) {
const packageDefs = allDefs.filter(def => {
for (const dirSuffix of importedPackages) {
if (isFileInPackageDir(def.filePath, dirSuffix)) return true;
}
return false;
});
if (packageDefs.length > 0) {
return { candidates: packageDefs, tier: 'import-scoped' };
}
}
// Tier 3: Global — pass all candidates through; filterCallableCandidates
// will narrow by kind/arity and resolveCallTarget only emits when exactly 1 remains.
return { candidates: allDefs, tier: 'unique-global' };
};
const CONSTRUCTOR_TARGET_TYPES = new Set(['Constructor', 'Class', 'Struct', 'Record']);
const filterCallableCandidates = (
candidates: SymbolDefinition[],
argCount?: number,
callForm?: 'free' | 'member' | 'constructor',
): SymbolDefinition[] => {
let kindFiltered: SymbolDefinition[];
if (callForm === 'constructor') {
// For constructor calls, prefer Constructor > Class/Struct/Record > callable fallback
const constructors = candidates.filter(c => c.type === 'Constructor');
if (constructors.length > 0) {
kindFiltered = constructors;
} else {
const types = candidates.filter(c => CONSTRUCTOR_TARGET_TYPES.has(c.type));
kindFiltered = types.length > 0 ? types : candidates.filter(c => CALLABLE_SYMBOL_TYPES.has(c.type));
}
} else {
kindFiltered = candidates.filter(c => CALLABLE_SYMBOL_TYPES.has(c.type));
}
if (kindFiltered.length === 0) return [];
if (argCount === undefined) return kindFiltered;
const hasParameterMetadata = kindFiltered.some(candidate => candidate.parameterCount !== undefined);
if (!hasParameterMetadata) return kindFiltered;
return kindFiltered.filter(candidate =>
candidate.parameterCount === undefined || candidate.parameterCount === argCount
);
};
const toResolveResult = (
definition: SymbolDefinition,
tier: ResolutionTier,
): ResolveResult => {
if (tier === 'same-file') {
return { nodeId: definition.nodeId, confidence: 0.95, reason: 'same-file' };
}
if (tier === 'import-scoped') {
return { nodeId: definition.nodeId, confidence: 0.9, reason: 'import-resolved' };
}
return { nodeId: definition.nodeId, confidence: 0.5, reason: 'unique-global' };
};
/**
* Resolve a function call to its target node ID using priority strategy:
* A. Narrow candidates by scope tier (same-file, import-scoped, unique-global)
* B. Filter to callable symbol kinds (constructor-aware when callForm is set)
* C. Apply arity filtering when parameter metadata is available
* D. Apply receiver-type filtering for member calls with typed receivers
*
* If filtering still leaves multiple candidates, refuse to emit a CALLS edge.
* A. Check imported files first (highest confidence)
* B. Check local file definitions
* C. Fuzzy global search (lowest confidence)
*
* Returns confidence score so agents know what to trust.
*/
const resolveCallTarget = (
call: Pick<ExtractedCall, 'calledName' | 'argCount' | 'callForm' | 'receiverTypeName'>,
calledName: string,
currentFile: string,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
importMap: ImportMap
): ResolveResult | null => {
const tiered = collectTieredCandidates(call.calledName, currentFile, symbolTable, importMap, packageMap, namedImportMap);
if (!tiered) return null;
const filteredCandidates = filterCallableCandidates(tiered.candidates, call.argCount, call.callForm);
// D. Receiver-type filtering: for member calls with a known receiver type,
// filter candidates by ownerId matching the resolved type's nodeId
if (call.callForm === 'member' && call.receiverTypeName && filteredCandidates.length > 1) {
const typeDefs = symbolTable.lookupFuzzy(call.receiverTypeName);
if (typeDefs.length > 0) {
const typeNodeIds = new Set(typeDefs.map(d => d.nodeId));
const ownerFiltered = filteredCandidates.filter(c => c.ownerId && typeNodeIds.has(c.ownerId));
if (ownerFiltered.length === 1) {
return toResolveResult(ownerFiltered[0], tiered.tier);
}
// If receiver filtering narrows to 0, fall through to name-only resolution
// If still 2+, refuse (don't guess)
if (ownerFiltered.length > 1) return null;
}
// Strategy B first (cheapest — single map lookup): Check local file
const localNodeId = symbolTable.lookupExact(currentFile, calledName);
if (localNodeId) {
return { nodeId: localNodeId, confidence: 0.85, reason: 'same-file' };
}
if (filteredCandidates.length !== 1) return null;
// Strategy A: Check if any definition of calledName is in an imported file
// Reversed: instead of iterating all imports and checking each, get all definitions
// and check if any is imported. O(definitions) instead of O(imports).
const allDefs = symbolTable.lookupFuzzy(calledName);
if (allDefs.length > 0) {
const importedFiles = importMap.get(currentFile);
if (importedFiles) {
for (const def of allDefs) {
if (importedFiles.has(def.filePath)) {
return { nodeId: def.nodeId, confidence: 0.9, reason: 'import-resolved' };
}
}
}
return toResolveResult(filteredCandidates[0], tiered.tier);
// Strategy C: Fuzzy global (no import match found)
const confidence = allDefs.length === 1 ? 0.5 : 0.3;
return { nodeId: allDefs[0].nodeId, confidence, reason: 'fuzzy-global' };
}
return null;
};
/**
* Filter out common built-in functions and noise
* that shouldn't be tracked as calls
*/
/** Pre-built set (module-level singleton) to avoid re-creating per call */
const BUILT_IN_NAMES = new Set([
// JavaScript/TypeScript built-ins
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export',
'fetch', 'Response', 'Request',
// React hooks and common functions
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
// Common array/object methods
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python built-ins
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'open', 'read', 'write', 'close', 'append', 'extend', 'update',
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
// Kotlin stdlib (IMPORTANT: keep in sync with parse-worker.ts BUILT_IN_NAMES)
'println', 'print', 'readLine', 'require', 'requireNotNull', 'check', 'assert', 'lazy', 'error',
'listOf', 'mapOf', 'setOf', 'mutableListOf', 'mutableMapOf', 'mutableSetOf',
'arrayOf', 'sequenceOf', 'also', 'apply', 'run', 'with', 'takeIf', 'takeUnless',
'TODO', 'buildString', 'buildList', 'buildMap', 'buildSet',
'repeat', 'synchronized',
// Kotlin coroutine builders & scope functions
'launch', 'async', 'runBlocking', 'withContext', 'coroutineScope',
'supervisorScope', 'delay',
// Kotlin Flow operators
'flow', 'flowOf', 'collect', 'emit', 'onEach', 'catch',
'buffer', 'conflate', 'distinctUntilChanged',
'flatMapLatest', 'flatMapMerge', 'combine',
'stateIn', 'shareIn', 'launchIn',
// Kotlin infix stdlib functions
'to', 'until', 'downTo', 'step',
// C/C++ standard library and common kernel helpers
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
'scanf', 'fscanf', 'sscanf',
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
'sizeof', 'offsetof', 'typeof',
'assert', 'abort', 'exit', '_exit',
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
// Linux kernel common macros/helpers (not real call targets)
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
'min', 'max', 'clamp', 'abs', 'swap',
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
'GFP_KERNEL', 'GFP_ATOMIC',
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
'mutex_lock', 'mutex_unlock', 'mutex_init',
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
'get', 'put',
// Swift/iOS built-ins and standard library
'print', 'debugPrint', 'dump', 'fatalError', 'precondition', 'preconditionFailure',
'assert', 'assertionFailure', 'NSLog',
'abs', 'min', 'max', 'zip', 'stride', 'sequence', 'repeatElement',
'swap', 'withUnsafePointer', 'withUnsafeMutablePointer', 'withUnsafeBytes',
'autoreleasepool', 'unsafeBitCast', 'unsafeDowncast', 'numericCast',
'type', 'MemoryLayout',
// Swift collection/string methods (common noise)
'map', 'flatMap', 'compactMap', 'filter', 'reduce', 'forEach', 'contains',
'first', 'last', 'prefix', 'suffix', 'dropFirst', 'dropLast',
'sorted', 'reversed', 'enumerated', 'joined', 'split',
'append', 'insert', 'remove', 'removeAll', 'removeFirst', 'removeLast',
'isEmpty', 'count', 'index', 'startIndex', 'endIndex',
// UIKit/Foundation common methods (noise in call graph)
'addSubview', 'removeFromSuperview', 'layoutSubviews', 'setNeedsLayout',
'layoutIfNeeded', 'setNeedsDisplay', 'invalidateIntrinsicContentSize',
'addTarget', 'removeTarget', 'addGestureRecognizer',
'addConstraint', 'addConstraints', 'removeConstraint', 'removeConstraints',
'NSLocalizedString', 'Bundle',
'reloadData', 'reloadSections', 'reloadRows', 'performBatchUpdates',
'register', 'dequeueReusableCell', 'dequeueReusableSupplementaryView',
'beginUpdates', 'endUpdates', 'insertRows', 'deleteRows', 'insertSections', 'deleteSections',
'present', 'dismiss', 'pushViewController', 'popViewController', 'popToRootViewController',
'performSegue', 'prepare',
// GCD / async
'DispatchQueue', 'async', 'sync', 'asyncAfter',
'Task', 'withCheckedContinuation', 'withCheckedThrowingContinuation',
// Combine
'sink', 'store', 'assign', 'receive', 'subscribe',
// Notification / KVO
'addObserver', 'removeObserver', 'post', 'NotificationCenter',
]);
const isBuiltInOrNoise = (name: string): boolean => BUILT_IN_NAMES.has(name);
/**
* Fast path: resolve pre-extracted call sites from workers.
* No AST parsing — workers already extracted calledName + sourceId.
@@ -359,9 +410,7 @@ export const processCallsFromExtracted = async (
extractedCalls: ExtractedCall[],
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void,
namedImportMap?: NamedImportMap,
onProgress?: (current: number, total: number) => void
) => {
// Group by file for progress reporting
const byFile = new Map<string, ExtractedCall[]>();
@@ -386,12 +435,10 @@ export const processCallsFromExtracted = async (
for (const call of calls) {
const resolved = resolveCallTarget(
call,
call.calledName,
call.filePath,
symbolTable,
importMap,
packageMap,
namedImportMap,
importMap
);
if (!resolved) continue;
@@ -418,7 +465,6 @@ export const processRoutesFromExtracted = async (
extractedRoutes: ExtractedRoute[],
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void
) => {
for (let i = 0; i < extractedRoutes.length; i++) {
@@ -430,16 +476,24 @@ export const processRoutesFromExtracted = async (
if (!route.controllerName || !route.methodName) continue;
// Resolve controller class using shared resolver (Tier 1: same file,
// Tier 2: import-scoped, Tier 3: unique global).
const resolution = resolveSymbolInternal(route.controllerName, route.filePath, symbolTable, importMap, packageMap);
if (!resolution) continue;
// Resolve controller class in symbol table
const controllerDefs = symbolTable.lookupFuzzy(route.controllerName);
if (controllerDefs.length === 0) continue;
const controllerDef = resolution.definition;
// Derive confidence from the resolution tier
const confidence = resolution.tier === 'same-file' ? 0.95
: resolution.tier === 'import-scoped' ? 0.9
: 0.7;
// Prefer import-resolved match
const importedFiles = importMap.get(route.filePath);
let controllerDef = controllerDefs[0];
let confidence = controllerDefs.length === 1 ? 0.7 : 0.5;
if (importedFiles) {
for (const def of controllerDefs) {
if (importedFiles.has(def.filePath)) {
controllerDef = def;
confidence = 0.9;
break;
}
}
}
// Find the method on the controller
const methodId = symbolTable.lookupExact(controllerDef.filePath, route.methodName);
@@ -473,22 +527,3 @@ export const processRoutesFromExtracted = async (
onProgress?.(extractedRoutes.length, extractedRoutes.length);
};
/**
* Follow re-export chains through NamedImportMap for call candidate collection.
* Delegates chain-walking to the shared walkBindingChain utility, then
* applies call-processor semantics: any number of matches accepted.
*/
const resolveNamedBindingChainForCandidates = (
calledName: string,
currentFile: string,
symbolTable: SymbolTable,
namedImportMap: NamedImportMap,
allDefs: SymbolDefinition[],
): TieredCandidates | null => {
const defs = walkBindingChain(calledName, currentFile, symbolTable, namedImportMap, allDefs);
if (defs && defs.length > 0) {
return { candidates: defs, tier: 'import-scoped' };
}
return null;
};
-19
View File
@@ -1,19 +0,0 @@
/**
* Default minimum buffer size for tree-sitter parsing (512 KB).
* tree-sitter requires bufferSize >= file size in bytes.
*/
export const TREE_SITTER_BUFFER_SIZE = 512 * 1024;
/**
* Maximum buffer size cap (32 MB) to prevent OOM on huge files.
* Also used as the file-size skip threshold — files larger than this are not parsed.
*/
export const TREE_SITTER_MAX_BUFFER = 32 * 1024 * 1024;
/**
* Compute adaptive buffer size for tree-sitter parsing.
* Uses 2× file size, clamped between 512 KB and 32 MB.
* Previous 256 KB fixed limit silently skipped files > ~200 KB (e.g., imgui.h at 411 KB).
*/
export const getTreeSitterBufferSize = (contentLength: number): number =>
Math.min(Math.max(contentLength * 2, TREE_SITTER_BUFFER_SIZE), TREE_SITTER_MAX_BUFFER);
@@ -11,7 +11,6 @@
*/
import { detectFrameworkFromPath } from './framework-detection.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
// ============================================================================
// NAME PATTERNS - All 9 supported languages
@@ -39,47 +38,39 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
],
// JavaScript/TypeScript
[SupportedLanguages.JavaScript]: [
'javascript': [
/^use[A-Z]/, // React hooks (useEffect, etc.)
],
[SupportedLanguages.TypeScript]: [
'typescript': [
/^use[A-Z]/, // React hooks
],
// Python
[SupportedLanguages.Python]: [
'python': [
/^app$/, // Flask/FastAPI app
/^(get|post|put|delete|patch)_/i, // REST conventions
/^api_/, // API functions
/^view_/, // Django views
],
// Java
[SupportedLanguages.Java]: [
'java': [
/^do[A-Z]/, // doGet, doPost (Servlets)
/^create[A-Z]/, // Factory patterns
/^build[A-Z]/, // Builder patterns
/Service$/, // UserService
],
// C#
[SupportedLanguages.CSharp]: [
/^(Get|Post|Put|Delete|Patch)/, // ASP.NET action methods
/Action$/, // MVC actions
/^On[A-Z]/, // Event handlers / Blazor lifecycle
/Async$/, // Async entry points
/^Configure$/, // Startup.Configure
/^ConfigureServices$/, // Startup.ConfigureServices
/^Handle$/, // MediatR / generic handler
/^Execute$/, // Command pattern
/^Invoke$/, // Middleware Invoke
/^Map[A-Z]/, // Minimal API MapGet, MapPost
/Service$/, // Service classes
/^Seed/, // Database seeding
'csharp': [
/^(Get|Post|Put|Delete)/, // ASP.NET conventions
/Action$/, // MVC actions
/^On[A-Z]/, // Event handlers
/Async$/, // Async entry points
],
// Go
[SupportedLanguages.Go]: [
'go': [
/Handler$/, // http.Handler pattern
/^Serve/, // ServeHTTP
/^New[A-Z]/, // Constructor pattern (returns new instance)
@@ -87,7 +78,7 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
],
// Rust
[SupportedLanguages.Rust]: [
'rust': [
/^(get|post|put|delete)_handler$/i,
/^handle_/, // handle_request
/^new$/, // Constructor pattern
@@ -95,64 +86,25 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
/^spawn/, // Async spawn
],
// C - explicit main() boost plus common C entry point conventions
[SupportedLanguages.C]: [
// C - explicit main() boost (critical for C programs)
'c': [
/^main$/, // THE entry point
/^init_/, // init_server, init_client
/_init$/, // module_init, server_init
/^start_/, // start_server
/_start$/, // thread_start
/^run_/, // run_loop
/_run$/, // event_run
/^stop_/, // stop_server
/_stop$/, // service_stop
/^open_/, // open_connection
/_open$/, // file_open
/^close_/, // close_connection
/_close$/, // socket_close
/^create_/, // create_session
/_create$/, // object_create
/^destroy_/, // destroy_session
/_destroy$/, // object_destroy
/^handle_/, // handle_request
/_handler$/, // signal_handler
/_callback$/, // event_callback
/^cmd_/, // tmux: cmd_new_window, cmd_attach_session
/^server_/, // server_start, server_loop
/^client_/, // client_connect
/^session_/, // session_create
/^window_/, // window_resize (tmux)
/^key_/, // key_press
/^input_/, // input_parse
/^output_/, // output_write
/^notify_/, // notify_client
/^control_/, // control_start
/^init_/, // Initialization functions
/^start_/, // Start functions
/^run_/, // Run functions
],
// C++ - same as C plus OOP/template patterns
[SupportedLanguages.CPlusPlus]: [
// C++ - same as C plus class patterns
'cpp': [
/^main$/, // THE entry point
/^init_/,
/_init$/,
/^Create[A-Z]/, // Factory patterns
/^create_/,
/^Run$/, // Run methods
/^run$/,
/^Start$/, // Start methods
/^start$/,
/^handle_/,
/_handler$/,
/_callback$/,
/^OnEvent/, // Event callbacks
/^on_/,
/::Run$/, // Class::Run
/::Start$/, // Class::Start
/::Init$/, // Class::Init
/::Execute$/, // Class::Execute
],
// Swift / iOS
[SupportedLanguages.Swift]: [
'swift': [
/^viewDidLoad$/, // UIKit lifecycle
/^viewWillAppear$/, // UIKit lifecycle
/^viewDidAppear$/, // UIKit lifecycle
@@ -172,7 +124,7 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
],
// PHP / Laravel
[SupportedLanguages.PHP]: [
'php': [
/Controller$/, // UserController (class name convention)
/^handle$/, // Job::handle(), Listener::handle()
/^execute$/, // Command::execute()
@@ -193,14 +145,6 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
],
};
/** Pre-computed merged patterns (universal + language-specific) to avoid per-call array allocation. */
const MERGED_ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {};
const UNIVERSAL_PATTERNS = ENTRY_POINT_PATTERNS['*'] || [];
for (const [lang, patterns] of Object.entries(ENTRY_POINT_PATTERNS)) {
if (lang === '*') continue;
MERGED_ENTRY_POINT_PATTERNS[lang] = [...UNIVERSAL_PATTERNS, ...patterns];
}
// ============================================================================
// UTILITY PATTERNS - Functions that should be penalized
// ============================================================================
@@ -255,7 +199,7 @@ export interface EntryPointScoreResult {
*/
export function calculateEntryPointScore(
name: string,
language: SupportedLanguages,
language: string,
isExported: boolean,
callerCount: number,
calleeCount: number,
@@ -288,7 +232,9 @@ export function calculateEntryPointScore(
reasons.push('utility-pattern');
} else {
// Check positive patterns
const allPatterns = MERGED_ENTRY_POINT_PATTERNS[language] || UNIVERSAL_PATTERNS;
const universalPatterns = ENTRY_POINT_PATTERNS['*'] || [];
const langPatterns = ENTRY_POINT_PATTERNS[language] || [];
const allPatterns = [...universalPatterns, ...langPatterns];
if (allPatterns.some(p => p.test(name))) {
nameMultiplier = 1.5; // Bonus for matching entry point pattern
@@ -350,13 +296,8 @@ export function isTestFile(filePath: string): boolean {
p.endsWith('test.swift') ||
p.includes('uitests/') ||
// C# test patterns
p.endsWith('tests.cs') ||
p.endsWith('test.cs') ||
p.includes('.tests/') ||
p.includes('.test/') ||
p.includes('.integrationtests/') ||
p.includes('.unittests/') ||
p.includes('/testproject/') ||
p.includes('tests.cs') ||
// PHP/Laravel test patterns
p.endsWith('test.php') ||
p.endsWith('spec.php') ||
@@ -1,242 +0,0 @@
/**
* Export Detection
*
* Determines whether a symbol (function, class, etc.) is exported/public
* in its language. This is a pure function — safe for use in worker threads.
*
* Shared between parse-worker.ts (worker pool) and parsing-processor.ts (sequential fallback).
*/
import { findSiblingChild, SyntaxNode } from './utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
/** Handler type: given a node and symbol name, return true if the symbol is exported/public. */
type ExportChecker = (node: SyntaxNode, name: string) => boolean;
// ============================================================================
// Per-language export checkers
// ============================================================================
/** JS/TS: walk ancestors looking for export_statement or export_specifier. */
const tsExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
const type = current.type;
if (type === 'export_statement' ||
type === 'export_specifier' ||
(type === 'lexical_declaration' && current.parent?.type === 'export_statement')) {
return true;
}
// Fallback: check if node text starts with 'export ' for edge cases
if (current.text?.startsWith('export ')) {
return true;
}
current = current.parent;
}
return false;
};
/** Python: public if no leading underscore (convention). */
const pythonExportChecker: ExportChecker = (_node, name) => !name.startsWith('_');
/** Java: check for 'public' modifier — modifiers are siblings of the name node, not parents. */
const javaExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.parent) {
const parent = current.parent;
for (let i = 0; i < parent.childCount; i++) {
const child = parent.child(i);
if (child?.type === 'modifiers' && child.text?.includes('public')) {
return true;
}
}
if (parent.type === 'method_declaration' || parent.type === 'constructor_declaration') {
if (parent.text?.trimStart().startsWith('public')) {
return true;
}
}
}
current = current.parent;
}
return false;
};
/** C# declaration node types for sibling modifier scanning. */
const CSHARP_DECL_TYPES = new Set([
'method_declaration', 'local_function_statement', 'constructor_declaration',
'class_declaration', 'interface_declaration', 'struct_declaration',
'enum_declaration', 'record_declaration', 'record_struct_declaration',
'record_class_declaration', 'delegate_declaration',
'property_declaration', 'field_declaration', 'event_declaration',
'namespace_declaration', 'file_scoped_namespace_declaration',
]);
/**
* C#: modifier nodes are SIBLINGS of the name node inside the declaration.
* Walk up to the declaration node, then scan its direct children.
*/
const csharpExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (CSHARP_DECL_TYPES.has(current.type)) {
for (let i = 0; i < current.childCount; i++) {
const child = current.child(i);
if (child?.type === 'modifier' && child.text === 'public') return true;
}
return false;
}
current = current.parent;
}
return false;
};
/** Go: uppercase first letter = exported. */
const goExportChecker: ExportChecker = (_node, name) => {
if (name.length === 0) return false;
const first = name[0];
return first === first.toUpperCase() && first !== first.toLowerCase();
};
/** Rust declaration node types for sibling visibility_modifier scanning. */
const RUST_DECL_TYPES = new Set([
'function_item', 'struct_item', 'enum_item', 'trait_item', 'impl_item',
'union_item', 'type_item', 'const_item', 'static_item', 'mod_item',
'use_declaration', 'associated_type', 'function_signature_item',
]);
/**
* Rust: visibility_modifier is a SIBLING of the name node within the declaration node
* (function_item, struct_item, etc.), not a parent. Walk up to the declaration node,
* then scan its direct children.
*/
const rustExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (RUST_DECL_TYPES.has(current.type)) {
for (let i = 0; i < current.childCount; i++) {
const child = current.child(i);
if (child?.type === 'visibility_modifier' && child.text?.startsWith('pub')) return true;
}
return false;
}
current = current.parent;
}
return false;
};
/**
* Kotlin: default visibility is public (unlike Java).
* visibility_modifier is inside modifiers, a sibling of the name node within the declaration.
*/
const kotlinExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.parent) {
const visMod = findSiblingChild(current.parent, 'modifiers', 'visibility_modifier');
if (visMod) {
const text = visMod.text;
if (text === 'private' || text === 'internal' || text === 'protected') return false;
if (text === 'public') return true;
}
}
current = current.parent;
}
// No visibility modifier = public (Kotlin default)
return true;
};
/**
* C/C++: functions without 'static' storage class have external linkage by default,
* making them globally accessible (equivalent to exported). Only functions explicitly
* marked 'static' are file-scoped (not exported). C++ anonymous namespaces
* (namespace { ... }) also give internal linkage.
*/
const cCppExportChecker: ExportChecker = (node, _name) => {
let cur: SyntaxNode | null = node;
while (cur) {
if (cur.type === 'function_definition' || cur.type === 'declaration') {
// Check for 'static' storage class specifier as a direct child node.
// This avoids reading the full function text (which can be very large).
for (let i = 0; i < cur.childCount; i++) {
const child = cur.child(i);
if (child?.type === 'storage_class_specifier' && child.text === 'static') return false;
}
}
// C++ anonymous namespace: namespace_definition with no name child = internal linkage
if (cur.type === 'namespace_definition') {
const hasName = cur.childForFieldName?.('name');
if (!hasName) return false;
}
cur = cur.parent;
}
return true; // Top-level C/C++ functions default to external linkage
};
/** PHP: check for visibility modifier or top-level scope. */
const phpExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.type === 'class_declaration' ||
current.type === 'interface_declaration' ||
current.type === 'trait_declaration' ||
current.type === 'enum_declaration') {
return true;
}
if (current.type === 'visibility_modifier') {
return current.text === 'public';
}
current = current.parent;
}
// Top-level functions are globally accessible
return true;
};
/** Swift: check for 'public' or 'open' access modifiers. */
const swiftExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.type === 'modifiers' || current.type === 'visibility_modifier') {
const text = current.text || '';
if (text.includes('public') || text.includes('open')) return true;
}
current = current.parent;
}
return false;
};
// ============================================================================
// Exhaustive dispatch table — satisfies enforces all SupportedLanguages are covered
// ============================================================================
const exportCheckers = {
[SupportedLanguages.JavaScript]: tsExportChecker,
[SupportedLanguages.TypeScript]: tsExportChecker,
[SupportedLanguages.Python]: pythonExportChecker,
[SupportedLanguages.Java]: javaExportChecker,
[SupportedLanguages.CSharp]: csharpExportChecker,
[SupportedLanguages.Go]: goExportChecker,
[SupportedLanguages.Rust]: rustExportChecker,
[SupportedLanguages.Kotlin]: kotlinExportChecker,
[SupportedLanguages.C]: cCppExportChecker,
[SupportedLanguages.CPlusPlus]: cCppExportChecker,
[SupportedLanguages.PHP]: phpExportChecker,
[SupportedLanguages.Swift]: swiftExportChecker,
} satisfies Record<SupportedLanguages, ExportChecker>;
// ============================================================================
// Public API
// ============================================================================
/**
* Check if a tree-sitter node is exported/public in its language.
* @param node - The tree-sitter AST node
* @param name - The symbol name
* @param language - The programming language
* @returns true if the symbol is exported/public
*/
export const isNodeExported = (node: SyntaxNode, name: string, language: SupportedLanguages): boolean => {
const checker = exportCheckers[language];
if (!checker) return false;
return checker(node, name);
};
@@ -183,35 +183,7 @@ export function detectFrameworkFromPath(filePath: string): FrameworkHint | null
if (p.endsWith('controller.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 3.0, reason: 'aspnet-controller-file' };
}
// ASP.NET Services
if ((p.includes('/services/') || p.includes('/service/')) && p.endsWith('.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 1.8, reason: 'aspnet-service' };
}
// ASP.NET Middleware
if (p.includes('/middleware/') && p.endsWith('.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 2.5, reason: 'aspnet-middleware' };
}
// SignalR Hubs
if (p.includes('/hubs/') && p.endsWith('.cs')) {
return { framework: 'signalr', entryPointMultiplier: 2.5, reason: 'signalr-hub' };
}
if (p.endsWith('hub.cs')) {
return { framework: 'signalr', entryPointMultiplier: 2.5, reason: 'signalr-hub-file' };
}
// Minimal API / Program.cs / Startup.cs
if (p.endsWith('/program.cs') || p.endsWith('/startup.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 3.0, reason: 'aspnet-entry' };
}
// Background services / Hosted services
if ((p.includes('/backgroundservices/') || p.includes('/hostedservices/')) && p.endsWith('.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 2.0, reason: 'aspnet-background-service' };
}
// Blazor pages
if (p.includes('/pages/') && p.endsWith('.razor')) {
return { framework: 'blazor', entryPointMultiplier: 2.5, reason: 'blazor-page' };
@@ -413,11 +385,7 @@ export const FRAMEWORK_AST_PATTERNS = {
'jaxrs': ['@Path', '@GET', '@POST', '@PUT', '@DELETE'],
// C# attributes
'aspnet': ['[ApiController]', '[HttpGet]', '[HttpPost]', '[HttpPut]', '[HttpDelete]',
'[Route]', '[Authorize]', '[AllowAnonymous]'],
'signalr': ['[HubMethodName]', ': Hub', ': Hub<'],
'blazor': ['@page', '[Parameter]', '@inject'],
'efcore': ['DbContext', 'DbSet<', 'OnModelCreating'],
'aspnet': ['[ApiController]', '[HttpGet]', '[HttpPost]', '[Route]'],
// Go patterns (function signatures)
'go-http': ['http.Handler', 'http.HandlerFunc', 'ServeHTTP'],
@@ -437,8 +405,6 @@ export const FRAMEWORK_AST_PATTERNS = {
'combine': ['sink', 'assign', 'Publisher', 'Subscriber'],
};
import { SupportedLanguages } from '../../config/supported-languages.js';
interface AstFrameworkPatternConfig {
framework: string;
entryPointMultiplier: number;
@@ -447,33 +413,30 @@ interface AstFrameworkPatternConfig {
}
const AST_FRAMEWORK_PATTERNS_BY_LANGUAGE: Record<string, AstFrameworkPatternConfig[]> = {
[SupportedLanguages.JavaScript]: [
javascript: [
{ framework: 'nestjs', entryPointMultiplier: 3.2, reason: 'nestjs-decorator', patterns: FRAMEWORK_AST_PATTERNS.nestjs },
],
[SupportedLanguages.TypeScript]: [
typescript: [
{ framework: 'nestjs', entryPointMultiplier: 3.2, reason: 'nestjs-decorator', patterns: FRAMEWORK_AST_PATTERNS.nestjs },
],
[SupportedLanguages.Python]: [
python: [
{ framework: 'fastapi', entryPointMultiplier: 3.0, reason: 'fastapi-decorator', patterns: FRAMEWORK_AST_PATTERNS.fastapi },
{ framework: 'flask', entryPointMultiplier: 2.8, reason: 'flask-decorator', patterns: FRAMEWORK_AST_PATTERNS.flask },
],
[SupportedLanguages.Java]: [
java: [
{ framework: 'spring', entryPointMultiplier: 3.2, reason: 'spring-annotation', patterns: FRAMEWORK_AST_PATTERNS.spring },
{ framework: 'jaxrs', entryPointMultiplier: 3.0, reason: 'jaxrs-annotation', patterns: FRAMEWORK_AST_PATTERNS.jaxrs },
],
[SupportedLanguages.Kotlin]: [
kotlin: [
{ framework: 'spring-kotlin', entryPointMultiplier: 3.2, reason: 'spring-kotlin-annotation', patterns: FRAMEWORK_AST_PATTERNS.spring },
{ framework: 'jaxrs', entryPointMultiplier: 3.0, reason: 'jaxrs-annotation', patterns: FRAMEWORK_AST_PATTERNS.jaxrs },
{ framework: 'ktor', entryPointMultiplier: 2.8, reason: 'ktor-routing', patterns: ['routing', 'embeddedServer', 'Application.module'] },
{ framework: 'android-kotlin', entryPointMultiplier: 2.5, reason: 'android-annotation', patterns: ['@AndroidEntryPoint', 'AppCompatActivity', 'Fragment('] },
],
[SupportedLanguages.CSharp]: [
csharp: [
{ framework: 'aspnet', entryPointMultiplier: 3.2, reason: 'aspnet-attribute', patterns: FRAMEWORK_AST_PATTERNS.aspnet },
{ framework: 'signalr', entryPointMultiplier: 2.8, reason: 'signalr-attribute', patterns: FRAMEWORK_AST_PATTERNS.signalr },
{ framework: 'blazor', entryPointMultiplier: 2.5, reason: 'blazor-attribute', patterns: FRAMEWORK_AST_PATTERNS.blazor },
{ framework: 'efcore', entryPointMultiplier: 2.0, reason: 'efcore-pattern', patterns: FRAMEWORK_AST_PATTERNS.efcore },
],
[SupportedLanguages.PHP]: [
php: [
{ framework: 'laravel', entryPointMultiplier: 3.0, reason: 'php-route-attribute', patterns: FRAMEWORK_AST_PATTERNS.laravel },
],
};
@@ -493,7 +456,7 @@ const AST_PATTERNS_LOWERED: Record<string, Array<{ framework: string; entryPoint
* Note: callers should slice definitionText to ~300 chars since annotations appear at the start.
*/
export function detectFrameworkFromAST(
language: SupportedLanguages,
language: string,
definitionText: string
): FrameworkHint | null {
if (!language || !definitionText) return null;
+32 -113
View File
@@ -1,83 +1,29 @@
/**
* Heritage Processor
*
*
* Extracts class inheritance relationships:
* - EXTENDS: Class extends another Class (TS, JS, Python, C#, C++)
* - IMPLEMENTS: Class implements an Interface (TS, C#, Java, Kotlin, PHP)
*
* Languages like C# use a single `base_list` for both class and interface parents.
* We resolve the correct edge type by checking the symbol table: if the parent is
* registered as an Interface, we emit IMPLEMENTS; otherwise EXTENDS. For unresolved
* external symbols, the fallback heuristic is language-gated:
* - C# / Java: apply the `I[A-Z]` naming convention (e.g. IDisposable → IMPLEMENTS)
* - Swift: default to IMPLEMENTS (protocol conformance is more common than class inheritance)
* - All other languages: default to EXTENDS
* - EXTENDS: Class extends another Class (TS, JS, Python)
* - IMPLEMENTS: Class implements an Interface (TS only)
*/
import { KnowledgeGraph } from '../graph/types.js';
import { ASTCache } from './ast-cache.js';
import { SymbolTable, SymbolDefinition } from './symbol-table.js';
import { SymbolTable } from './symbol-table.js';
import Parser from 'tree-sitter';
import { isLanguageAvailable, loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { generateId } from '../../lib/utils.js';
import { getLanguageFromFilename, isVerboseIngestionEnabled, yieldToEventLoop } from './utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
import { getTreeSitterBufferSize } from './constants.js';
import { getLanguageFromFilename, yieldToEventLoop } from './utils.js';
import type { ExtractedHeritage } from './workers/parse-worker.js';
import { resolveSymbol } from './symbol-resolver.js';
import type { ImportMap, PackageMap } from './import-processor.js';
/** C#/Java convention: interfaces start with I followed by an uppercase letter */
const INTERFACE_NAME_RE = /^I[A-Z]/;
/**
* Determine whether a heritage.extends capture is actually an IMPLEMENTS relationship.
* Uses the symbol table first (authoritative — Tier 1); falls back to a language-gated
* heuristic for external symbols not present in the graph:
* - C# / Java: `I[A-Z]` naming convention
* - Swift: default IMPLEMENTS (protocol conformance is the norm)
* - All others: default EXTENDS
*/
const resolveExtendsType = (
parentName: string,
currentFilePath: string,
symbolTable: SymbolTable,
importMap: ImportMap,
language: SupportedLanguages,
packageMap?: PackageMap,
): { type: 'EXTENDS' | 'IMPLEMENTS'; idPrefix: string } => {
const resolved = resolveSymbol(parentName, currentFilePath, symbolTable, importMap, packageMap);
if (resolved) {
const isInterface = resolved.type === 'Interface';
return isInterface
? { type: 'IMPLEMENTS', idPrefix: 'Interface' }
: { type: 'EXTENDS', idPrefix: 'Class' };
}
// Unresolved symbol — fall back to language-specific heuristic
if (language === SupportedLanguages.CSharp || language === SupportedLanguages.Java) {
if (INTERFACE_NAME_RE.test(parentName)) {
return { type: 'IMPLEMENTS', idPrefix: 'Interface' };
}
} else if (language === SupportedLanguages.Swift) {
// Protocol conformance is far more common than class inheritance in Swift
return { type: 'IMPLEMENTS', idPrefix: 'Interface' };
}
return { type: 'EXTENDS', idPrefix: 'Class' };
};
export const processHeritage = async (
graph: KnowledgeGraph,
files: { path: string; content: string }[],
astCache: ASTCache,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void
) => {
const parser = await loadParser();
const logSkipped = isVerboseIngestionEnabled();
const skippedByLang = logSkipped ? new Map<string, number>() : null;
for (let i = 0; i < files.length; i++) {
const file = files[i];
@@ -87,12 +33,6 @@ export const processHeritage = async (
// 1. Check language support
const language = getLanguageFromFilename(file.path);
if (!language) continue;
if (!isLanguageAvailable(language)) {
if (skippedByLang) {
skippedByLang.set(language, (skippedByLang.get(language) ?? 0) + 1);
}
continue;
}
const queryStr = LANGUAGE_QUERIES[language];
if (!queryStr) continue;
@@ -107,7 +47,7 @@ export const processHeritage = async (
if (!tree) {
// Use larger bufferSize for files > 32KB
try {
tree = parser.parse(file.content, undefined, { bufferSize: getTreeSitterBufferSize(file.content.length) });
tree = parser.parse(file.content, undefined, { bufferSize: 1024 * 256 });
} catch (parseError) {
// Skip files that can't be parsed
continue;
@@ -135,34 +75,27 @@ export const processHeritage = async (
captureMap[c.name] = c.node;
});
// EXTENDS or IMPLEMENTS: resolve via symbol table for languages where
// the tree-sitter query can't distinguish classes from interfaces (C#, Java)
// EXTENDS: Class extends another Class
if (captureMap['heritage.class'] && captureMap['heritage.extends']) {
// Go struct embedding: skip named fields (only anonymous fields are embedded)
const extendsNode = captureMap['heritage.extends'];
const fieldDecl = extendsNode.parent;
if (fieldDecl?.type === 'field_declaration' && fieldDecl.childForFieldName('name')) {
return; // Named field, not struct embedding
}
const className = captureMap['heritage.class'].text;
const parentClassName = captureMap['heritage.extends'].text;
const { type: relType, idPrefix } = resolveExtendsType(parentClassName, file.path, symbolTable, importMap, language, packageMap);
// Resolve both class IDs
const childId = symbolTable.lookupExact(file.path, className) ||
resolveSymbol(className, file.path, symbolTable, importMap, packageMap)?.nodeId ||
symbolTable.lookupFuzzy(className)[0]?.nodeId ||
generateId('Class', `${file.path}:${className}`);
const parentId = resolveSymbol(parentClassName, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId(idPrefix, `${parentClassName}`);
const parentId = symbolTable.lookupFuzzy(parentClassName)[0]?.nodeId ||
generateId('Class', `${parentClassName}`);
if (childId && parentId && childId !== parentId) {
const relId = generateId('EXTENDS', `${childId}->${parentId}`);
graph.addRelationship({
id: generateId(relType, `${childId}->${parentId}`),
id: relId,
sourceId: childId,
targetId: parentId,
type: relType,
type: 'EXTENDS',
confidence: 1.0,
reason: '',
});
@@ -176,10 +109,10 @@ export const processHeritage = async (
// Resolve class and interface IDs
const classId = symbolTable.lookupExact(file.path, className) ||
resolveSymbol(className, file.path, symbolTable, importMap, packageMap)?.nodeId ||
symbolTable.lookupFuzzy(className)[0]?.nodeId ||
generateId('Class', `${file.path}:${className}`);
const interfaceId = resolveSymbol(interfaceName, file.path, symbolTable, importMap, packageMap)?.nodeId ||
const interfaceId = symbolTable.lookupFuzzy(interfaceName)[0]?.nodeId ||
generateId('Interface', `${interfaceName}`);
if (classId && interfaceId) {
@@ -203,10 +136,10 @@ export const processHeritage = async (
// Resolve struct and trait IDs
const structId = symbolTable.lookupExact(file.path, structName) ||
resolveSymbol(structName, file.path, symbolTable, importMap, packageMap)?.nodeId ||
symbolTable.lookupFuzzy(structName)[0]?.nodeId ||
generateId('Struct', `${file.path}:${structName}`);
const traitId = resolveSymbol(traitName, file.path, symbolTable, importMap, packageMap)?.nodeId ||
const traitId = symbolTable.lookupFuzzy(traitName)[0]?.nodeId ||
generateId('Trait', `${traitName}`);
if (structId && traitId) {
@@ -226,14 +159,6 @@ export const processHeritage = async (
// Tree is now owned by the LRU cache — no manual delete needed
}
if (skippedByLang && skippedByLang.size > 0) {
for (const [lang, count] of skippedByLang.entries()) {
console.warn(
`[ingestion] Skipped ${count} ${lang} file(s) in heritage processing — ${lang} parser not available.`
);
}
}
};
/**
@@ -244,8 +169,6 @@ export const processHeritageFromExtracted = async (
graph: KnowledgeGraph,
extractedHeritage: ExtractedHeritage[],
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void
) => {
const total = extractedHeritage.length;
@@ -259,33 +182,29 @@ export const processHeritageFromExtracted = async (
const h = extractedHeritage[i];
if (h.kind === 'extends') {
const fileLanguage = getLanguageFromFilename(h.filePath);
if (!fileLanguage) continue;
const { type: relType, idPrefix } = resolveExtendsType(h.parentName, h.filePath, symbolTable, importMap, fileLanguage, packageMap);
const childId = symbolTable.lookupExact(h.filePath, h.className) ||
resolveSymbol(h.className, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
symbolTable.lookupFuzzy(h.className)[0]?.nodeId ||
generateId('Class', `${h.filePath}:${h.className}`);
const parentId = resolveSymbol(h.parentName, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId(idPrefix, `${h.parentName}`);
const parentId = symbolTable.lookupFuzzy(h.parentName)[0]?.nodeId ||
generateId('Class', `${h.parentName}`);
if (childId && parentId && childId !== parentId) {
graph.addRelationship({
id: generateId(relType, `${childId}->${parentId}`),
id: generateId('EXTENDS', `${childId}->${parentId}`),
sourceId: childId,
targetId: parentId,
type: relType,
type: 'EXTENDS',
confidence: 1.0,
reason: '',
});
}
} else if (h.kind === 'implements') {
const classId = symbolTable.lookupExact(h.filePath, h.className) ||
resolveSymbol(h.className, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
symbolTable.lookupFuzzy(h.className)[0]?.nodeId ||
generateId('Class', `${h.filePath}:${h.className}`);
const interfaceId = resolveSymbol(h.parentName, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
const interfaceId = symbolTable.lookupFuzzy(h.parentName)[0]?.nodeId ||
generateId('Interface', `${h.parentName}`);
if (classId && interfaceId) {
@@ -300,10 +219,10 @@ export const processHeritageFromExtracted = async (
}
} else if (h.kind === 'trait-impl') {
const structId = symbolTable.lookupExact(h.filePath, h.className) ||
resolveSymbol(h.className, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
symbolTable.lookupFuzzy(h.className)[0]?.nodeId ||
generateId('Struct', `${h.filePath}:${h.className}`);
const traitId = resolveSymbol(h.parentName, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
const traitId = symbolTable.lookupFuzzy(h.parentName)[0]?.nodeId ||
generateId('Trait', `${h.parentName}`);
if (structId && traitId) {
File diff suppressed because it is too large Load Diff
@@ -1,215 +0,0 @@
import fs from 'fs/promises';
import path from 'path';
const isDev = process.env.NODE_ENV === 'development';
// ============================================================================
// LANGUAGE-SPECIFIC CONFIG TYPES
// ============================================================================
/** TypeScript path alias config parsed from tsconfig.json */
export interface TsconfigPaths {
/** Map of alias prefix -> target prefix (e.g., "@/" -> "src/") */
aliases: Map<string, string>;
/** Base URL for path resolution (relative to repo root) */
baseUrl: string;
}
/** Go module config parsed from go.mod */
export interface GoModuleConfig {
/** Module path (e.g., "github.com/user/repo") */
modulePath: string;
}
/** PHP Composer PSR-4 autoload config */
export interface ComposerConfig {
/** Map of namespace prefix -> directory (e.g., "App\\" -> "app/") */
psr4: Map<string, string>;
}
/** C# project config parsed from .csproj files */
export interface CSharpProjectConfig {
/** Root namespace from <RootNamespace> or assembly name (default: project directory name) */
rootNamespace: string;
/** Directory containing the .csproj file */
projectDir: string;
}
/** Swift Package Manager module config */
export interface SwiftPackageConfig {
/** Map of target name -> source directory path (e.g., "SiuperModel" -> "Package/Sources/SiuperModel") */
targets: Map<string, string>;
}
// ============================================================================
// LANGUAGE-SPECIFIC CONFIG LOADERS
// ============================================================================
/**
* Parse tsconfig.json to extract path aliases.
* Tries tsconfig.json, tsconfig.app.json, tsconfig.base.json in order.
*/
export async function loadTsconfigPaths(repoRoot: string): Promise<TsconfigPaths | null> {
const candidates = ['tsconfig.json', 'tsconfig.app.json', 'tsconfig.base.json'];
for (const filename of candidates) {
try {
const tsconfigPath = path.join(repoRoot, filename);
const raw = await fs.readFile(tsconfigPath, 'utf-8');
// Strip JSON comments (// and /* */ style) for robustness
const stripped = raw.replace(/\/\/.*$/gm, '').replace(/\/\*[\s\S]*?\*\//g, '');
const tsconfig = JSON.parse(stripped);
const compilerOptions = tsconfig.compilerOptions;
if (!compilerOptions?.paths) continue;
const baseUrl = compilerOptions.baseUrl || '.';
const aliases = new Map<string, string>();
for (const [pattern, targets] of Object.entries(compilerOptions.paths)) {
if (!Array.isArray(targets) || targets.length === 0) continue;
const target = targets[0] as string;
// Convert glob patterns: "@/*" -> "@/", "src/*" -> "src/"
const aliasPrefix = pattern.endsWith('/*') ? pattern.slice(0, -1) : pattern;
const targetPrefix = target.endsWith('/*') ? target.slice(0, -1) : target;
aliases.set(aliasPrefix, targetPrefix);
}
if (aliases.size > 0) {
if (isDev) {
console.log(`📦 Loaded ${aliases.size} path aliases from ${filename}`);
}
return { aliases, baseUrl };
}
} catch {
// File doesn't exist or isn't valid JSON - try next
}
}
return null;
}
/**
* Parse go.mod to extract module path.
*/
export async function loadGoModulePath(repoRoot: string): Promise<GoModuleConfig | null> {
try {
const goModPath = path.join(repoRoot, 'go.mod');
const content = await fs.readFile(goModPath, 'utf-8');
const match = content.match(/^module\s+(\S+)/m);
if (match) {
if (isDev) {
console.log(`📦 Loaded Go module path: ${match[1]}`);
}
return { modulePath: match[1] };
}
} catch {
// No go.mod
}
return null;
}
/** Parse composer.json to extract PSR-4 autoload mappings (including autoload-dev). */
export async function loadComposerConfig(repoRoot: string): Promise<ComposerConfig | null> {
try {
const composerPath = path.join(repoRoot, 'composer.json');
const raw = await fs.readFile(composerPath, 'utf-8');
const composer = JSON.parse(raw);
const psr4Raw = composer.autoload?.['psr-4'] ?? {};
const psr4Dev = composer['autoload-dev']?.['psr-4'] ?? {};
const merged = { ...psr4Raw, ...psr4Dev };
const psr4 = new Map<string, string>();
for (const [ns, dir] of Object.entries(merged)) {
const nsNorm = (ns as string).replace(/\\+$/, '');
const dirNorm = (dir as string).replace(/\\/g, '/').replace(/\/+$/, '');
psr4.set(nsNorm, dirNorm);
}
if (isDev) {
console.log(`📦 Loaded ${psr4.size} PSR-4 mappings from composer.json`);
}
return { psr4 };
} catch {
return null;
}
}
/**
* Parse .csproj files to extract RootNamespace.
* Scans the repo root for .csproj files and returns configs for each.
*/
export async function loadCSharpProjectConfig(repoRoot: string): Promise<CSharpProjectConfig[]> {
const configs: CSharpProjectConfig[] = [];
// BFS scan for .csproj files up to 5 levels deep, cap at 100 dirs to avoid runaway scanning
const scanQueue: { dir: string; depth: number }[] = [{ dir: repoRoot, depth: 0 }];
const maxDepth = 5;
const maxDirs = 100;
let dirsScanned = 0;
while (scanQueue.length > 0 && dirsScanned < maxDirs) {
const { dir, depth } = scanQueue.shift()!;
dirsScanned++;
try {
const entries = await fs.readdir(dir, { withFileTypes: true });
for (const entry of entries) {
if (entry.isDirectory() && depth < maxDepth) {
// Skip common non-project directories
if (entry.name === 'node_modules' || entry.name === '.git' || entry.name === 'bin' || entry.name === 'obj') continue;
scanQueue.push({ dir: path.join(dir, entry.name), depth: depth + 1 });
}
if (entry.isFile() && entry.name.endsWith('.csproj')) {
try {
const csprojPath = path.join(dir, entry.name);
const content = await fs.readFile(csprojPath, 'utf-8');
const nsMatch = content.match(/<RootNamespace>\s*([^<]+)\s*<\/RootNamespace>/);
const rootNamespace = nsMatch
? nsMatch[1].trim()
: entry.name.replace(/\.csproj$/, '');
const projectDir = path.relative(repoRoot, dir).replace(/\\/g, '/');
configs.push({ rootNamespace, projectDir });
if (isDev) {
console.log(`📦 Loaded C# project: ${entry.name} (namespace: ${rootNamespace}, dir: ${projectDir})`);
}
} catch {
// Can't read .csproj
}
}
}
} catch {
// Can't read directory
}
}
return configs;
}
export async function loadSwiftPackageConfig(repoRoot: string): Promise<SwiftPackageConfig | null> {
// Swift imports are module-name based (e.g., `import SiuperModel`)
// SPM convention: Sources/<TargetName>/ or Package/Sources/<TargetName>/
// We scan for these directories to build a target map
const targets = new Map<string, string>();
const sourceDirs = ['Sources', 'Package/Sources', 'src'];
for (const sourceDir of sourceDirs) {
try {
const fullPath = path.join(repoRoot, sourceDir);
const entries = await fs.readdir(fullPath, { withFileTypes: true });
for (const entry of entries) {
if (entry.isDirectory()) {
targets.set(entry.name, sourceDir + '/' + entry.name);
}
}
} catch {
// Directory doesn't exist
}
}
if (targets.size > 0) {
if (isDev) {
console.log(`📦 Loaded ${targets.size} Swift package targets`);
}
return { targets };
}
return null;
}
@@ -1,465 +0,0 @@
/**
* MRO (Method Resolution Order) Processor
*
* Walks the inheritance DAG (EXTENDS/IMPLEMENTS edges), collects methods from
* each ancestor via HAS_METHOD edges, detects method-name collisions across
* parents, and applies language-specific resolution rules to emit OVERRIDES edges.
*
* Language-specific rules:
* - C++: leftmost base class in declaration order wins
* - C#/Java: class method wins over interface default; multiple interface
* methods with same name are ambiguous (null resolution)
* - Python: C3 linearization determines MRO; first in linearized order wins
* - Rust: no auto-resolution — requires qualified syntax, resolvedTo = null
* - Default: single inheritance — first definition wins
*
* OVERRIDES edge direction: Class → Method (not Method → Method).
* The source is the child class that inherits conflicting methods,
* the target is the winning ancestor method node.
* Cypher: MATCH (c:Class)-[r:CodeRelation {type: 'OVERRIDES'}]->(m:Method)
*/
import { KnowledgeGraph, GraphRelationship } from '../graph/types.js';
import { generateId } from '../../lib/utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
// ---------------------------------------------------------------------------
// Public types
// ---------------------------------------------------------------------------
export interface MROEntry {
classId: string;
className: string;
language: SupportedLanguages;
mro: string[]; // linearized parent names
ambiguities: MethodAmbiguity[];
}
export interface MethodAmbiguity {
methodName: string;
definedIn: Array<{ classId: string; className: string; methodId: string }>;
resolvedTo: string | null; // winning methodId or null if truly ambiguous
reason: string;
}
export interface MROResult {
entries: MROEntry[];
overrideEdges: number;
ambiguityCount: number;
}
// ---------------------------------------------------------------------------
// Internal helpers
// ---------------------------------------------------------------------------
/** Collect EXTENDS, IMPLEMENTS, and HAS_METHOD adjacency from the graph. */
function buildAdjacency(graph: KnowledgeGraph) {
// parentMap: childId → parentIds[] (in insertion / declaration order)
const parentMap = new Map<string, string[]>();
// methodMap: classId → methodIds[]
const methodMap = new Map<string, string[]>();
// Track which edge type each parent link came from
const parentEdgeType = new Map<string, Map<string, 'EXTENDS' | 'IMPLEMENTS'>>();
graph.forEachRelationship((rel) => {
if (rel.type === 'EXTENDS' || rel.type === 'IMPLEMENTS') {
let parents = parentMap.get(rel.sourceId);
if (!parents) {
parents = [];
parentMap.set(rel.sourceId, parents);
}
parents.push(rel.targetId);
let edgeTypes = parentEdgeType.get(rel.sourceId);
if (!edgeTypes) {
edgeTypes = new Map();
parentEdgeType.set(rel.sourceId, edgeTypes);
}
edgeTypes.set(rel.targetId, rel.type);
}
if (rel.type === 'HAS_METHOD') {
let methods = methodMap.get(rel.sourceId);
if (!methods) {
methods = [];
methodMap.set(rel.sourceId, methods);
}
methods.push(rel.targetId);
}
});
return { parentMap, methodMap, parentEdgeType };
}
/**
* Gather all ancestor IDs in BFS / topological order.
* Returns the linearized list of ancestor IDs (excluding the class itself).
*/
function gatherAncestors(
classId: string,
parentMap: Map<string, string[]>,
): string[] {
const visited = new Set<string>();
const order: string[] = [];
const queue: string[] = [...(parentMap.get(classId) ?? [])];
while (queue.length > 0) {
const id = queue.shift()!;
if (visited.has(id)) continue;
visited.add(id);
order.push(id);
const grandparents = parentMap.get(id);
if (grandparents) {
for (const gp of grandparents) {
if (!visited.has(gp)) queue.push(gp);
}
}
}
return order;
}
// ---------------------------------------------------------------------------
// C3 linearization (Python MRO)
// ---------------------------------------------------------------------------
/**
* Compute C3 linearization for a class given a parentMap.
* Returns an array of ancestor IDs in C3 order (excluding the class itself),
* or null if linearization fails (inconsistent or cyclic hierarchy).
*/
function c3Linearize(
classId: string,
parentMap: Map<string, string[]>,
cache: Map<string, string[] | null>,
inProgress?: Set<string>,
): string[] | null {
if (cache.has(classId)) return cache.get(classId)!;
// Cycle detection: if we're already computing this class, the hierarchy is cyclic
const visiting = inProgress ?? new Set<string>();
if (visiting.has(classId)) {
cache.set(classId, null);
return null;
}
visiting.add(classId);
const directParents = parentMap.get(classId);
if (!directParents || directParents.length === 0) {
visiting.delete(classId);
cache.set(classId, []);
return [];
}
// Compute linearization for each parent first
const parentLinearizations: string[][] = [];
for (const pid of directParents) {
const pLin = c3Linearize(pid, parentMap, cache, visiting);
if (pLin === null) {
visiting.delete(classId);
cache.set(classId, null);
return null;
}
parentLinearizations.push([pid, ...pLin]);
}
// Add the direct parents list as the final sequence
const sequences = [...parentLinearizations, [...directParents]];
const result: string[] = [];
while (sequences.some(s => s.length > 0)) {
// Find a good head: one that doesn't appear in the tail of any other sequence
let head: string | null = null;
for (const seq of sequences) {
if (seq.length === 0) continue;
const candidate = seq[0];
const inTail = sequences.some(
other => other.length > 1 && other.indexOf(candidate, 1) !== -1
);
if (!inTail) {
head = candidate;
break;
}
}
if (head === null) {
// Inconsistent hierarchy
visiting.delete(classId);
cache.set(classId, null);
return null;
}
result.push(head);
// Remove the chosen head from all sequences
for (const seq of sequences) {
if (seq.length > 0 && seq[0] === head) {
seq.shift();
}
}
}
visiting.delete(classId);
cache.set(classId, result);
return result;
}
// ---------------------------------------------------------------------------
// Language-specific resolution
// ---------------------------------------------------------------------------
type MethodDef = { classId: string; className: string; methodId: string };
type Resolution = { resolvedTo: string | null; reason: string };
/** Resolve by MRO order — first ancestor in linearized order wins. */
function resolveByMroOrder(
methodName: string,
defs: MethodDef[],
mroOrder: string[],
reasonPrefix: string,
): Resolution {
for (const ancestorId of mroOrder) {
const match = defs.find(d => d.classId === ancestorId);
if (match) {
return {
resolvedTo: match.methodId,
reason: `${reasonPrefix}: ${match.className}::${methodName}`,
};
}
}
return { resolvedTo: defs[0].methodId, reason: `${reasonPrefix} fallback: first definition` };
}
function resolveCsharpJava(
methodName: string,
defs: MethodDef[],
parentEdgeTypes: Map<string, 'EXTENDS' | 'IMPLEMENTS'> | undefined,
): Resolution {
const classDefs: MethodDef[] = [];
const interfaceDefs: MethodDef[] = [];
for (const def of defs) {
const edgeType = parentEdgeTypes?.get(def.classId);
if (edgeType === 'IMPLEMENTS') {
interfaceDefs.push(def);
} else {
classDefs.push(def);
}
}
if (classDefs.length > 0) {
return {
resolvedTo: classDefs[0].methodId,
reason: `class method wins: ${classDefs[0].className}::${methodName}`,
};
}
if (interfaceDefs.length > 1) {
return {
resolvedTo: null,
reason: `ambiguous: ${methodName} defined in multiple interfaces: ${interfaceDefs.map(d => d.className).join(', ')}`,
};
}
if (interfaceDefs.length === 1) {
return {
resolvedTo: interfaceDefs[0].methodId,
reason: `single interface default: ${interfaceDefs[0].className}::${methodName}`,
};
}
return { resolvedTo: null, reason: 'no resolution found' };
}
// ---------------------------------------------------------------------------
// Main entry point
// ---------------------------------------------------------------------------
export function computeMRO(graph: KnowledgeGraph): MROResult {
const { parentMap, methodMap, parentEdgeType } = buildAdjacency(graph);
const c3Cache = new Map<string, string[] | null>();
const entries: MROEntry[] = [];
let overrideEdges = 0;
let ambiguityCount = 0;
// Process every class that has at least one parent
for (const [classId, directParents] of parentMap) {
if (directParents.length === 0) continue;
const classNode = graph.getNode(classId);
if (!classNode) continue;
const language = classNode.properties.language;
if (!language) continue;
const className = classNode.properties.name;
// Compute linearized MRO depending on language
let mroOrder: string[];
if (language === SupportedLanguages.Python) {
const c3Result = c3Linearize(classId, parentMap, c3Cache);
mroOrder = c3Result ?? gatherAncestors(classId, parentMap);
} else {
mroOrder = gatherAncestors(classId, parentMap);
}
// Get the parent names for the MRO entry
const mroNames: string[] = mroOrder
.map(id => graph.getNode(id)?.properties.name)
.filter((n): n is string => n !== undefined);
// Collect methods from all ancestors, grouped by method name
const methodsByName = new Map<string, MethodDef[]>();
for (const ancestorId of mroOrder) {
const ancestorNode = graph.getNode(ancestorId);
if (!ancestorNode) continue;
const methods = methodMap.get(ancestorId) ?? [];
for (const methodId of methods) {
const methodNode = graph.getNode(methodId);
if (!methodNode) continue;
// Properties don't participate in method resolution order
if (methodNode.label === 'Property') continue;
const methodName = methodNode.properties.name;
let defs = methodsByName.get(methodName);
if (!defs) {
defs = [];
methodsByName.set(methodName, defs);
}
// Avoid duplicates (same method seen via multiple paths)
if (!defs.some(d => d.methodId === methodId)) {
defs.push({
classId: ancestorId,
className: ancestorNode.properties.name,
methodId,
});
}
}
}
// Detect collisions: methods defined in 2+ different ancestors
const ambiguities: MethodAmbiguity[] = [];
// Compute transitive edge types once per class (only needed for C#/Java)
const needsEdgeTypes = language === SupportedLanguages.CSharp || language === SupportedLanguages.Java || language === SupportedLanguages.Kotlin;
const classEdgeTypes = needsEdgeTypes
? buildTransitiveEdgeTypes(classId, parentMap, parentEdgeType)
: undefined;
for (const [methodName, defs] of methodsByName) {
if (defs.length < 2) continue;
// Own method shadows inherited — no ambiguity
const ownMethods = methodMap.get(classId) ?? [];
const ownDefinesIt = ownMethods.some(mid => {
const mn = graph.getNode(mid);
return mn?.properties.name === methodName;
});
if (ownDefinesIt) continue;
let resolution: Resolution;
switch (language) {
case SupportedLanguages.CPlusPlus:
resolution = resolveByMroOrder(methodName, defs, mroOrder, 'C++ leftmost base');
break;
case SupportedLanguages.CSharp:
case SupportedLanguages.Java:
case SupportedLanguages.Kotlin:
resolution = resolveCsharpJava(methodName, defs, classEdgeTypes);
break;
case SupportedLanguages.Python:
resolution = resolveByMroOrder(methodName, defs, mroOrder, 'Python C3 MRO');
break;
case SupportedLanguages.Rust:
resolution = {
resolvedTo: null,
reason: `Rust requires qualified syntax: <Type as Trait>::${methodName}()`,
};
break;
default:
resolution = resolveByMroOrder(methodName, defs, mroOrder, 'first definition');
break;
}
const ambiguity: MethodAmbiguity = {
methodName,
definedIn: defs,
resolvedTo: resolution.resolvedTo,
reason: resolution.reason,
};
ambiguities.push(ambiguity);
if (resolution.resolvedTo === null) {
ambiguityCount++;
}
// Emit OVERRIDES edge if resolution found
if (resolution.resolvedTo !== null) {
graph.addRelationship({
id: generateId('OVERRIDES', `${classId}->${resolution.resolvedTo}`),
sourceId: classId,
targetId: resolution.resolvedTo,
type: 'OVERRIDES',
confidence: 1.0,
reason: resolution.reason,
});
overrideEdges++;
}
}
entries.push({
classId,
className,
language,
mro: mroNames,
ambiguities,
});
}
return { entries, overrideEdges, ambiguityCount };
}
/**
* Build transitive edge types for a class using BFS from the class to all ancestors.
*
* Known limitation: BFS first-reach heuristic can misclassify an interface as
* EXTENDS if it's reachable via a class chain before being seen via IMPLEMENTS.
* E.g. if BaseClass also implements IFoo, IFoo may be classified as EXTENDS.
* This affects C#/Java/Kotlin conflict resolution in rare diamond hierarchies.
*/
function buildTransitiveEdgeTypes(
classId: string,
parentMap: Map<string, string[]>,
parentEdgeType: Map<string, Map<string, 'EXTENDS' | 'IMPLEMENTS'>>,
): Map<string, 'EXTENDS' | 'IMPLEMENTS'> {
const result = new Map<string, 'EXTENDS' | 'IMPLEMENTS'>();
const directEdges = parentEdgeType.get(classId);
if (!directEdges) return result;
// BFS: propagate edge type from direct parents
const queue: Array<{ id: string; edgeType: 'EXTENDS' | 'IMPLEMENTS' }> = [];
const directParents = parentMap.get(classId) ?? [];
for (const pid of directParents) {
const et = directEdges.get(pid) ?? 'EXTENDS';
if (!result.has(pid)) {
result.set(pid, et);
queue.push({ id: pid, edgeType: et });
}
}
while (queue.length > 0) {
const { id, edgeType } = queue.shift()!;
const grandparents = parentMap.get(id) ?? [];
for (const gp of grandparents) {
if (!result.has(gp)) {
result.set(gp, edgeType);
queue.push({ id: gp, edgeType });
}
}
}
return result;
}
@@ -1,384 +0,0 @@
import { SupportedLanguages } from '../../config/supported-languages.js';
import type { SymbolTable, SymbolDefinition } from './symbol-table.js';
import type { NamedImportMap } from './import-processor.js';
/**
* Walk a named-binding re-export chain through NamedImportMap.
*
* When file A imports { User } from B, and B re-exports { User } from C,
* the NamedImportMap for A points to B, but B has no User definition.
* This function follows the chain: A→B→C until a definition is found.
*
* Returns the definitions found at the end of the chain, or null if the
* chain breaks (missing binding, circular reference, or depth exceeded).
* Max depth 5 to prevent infinite loops.
*
* @param allDefs Pre-computed `symbolTable.lookupFuzzy(name)` result — must be the
* complete unfiltered result. Passing a file-filtered subset will cause
* silent misses at depth=0 for non-aliased bindings.
*/
export function walkBindingChain(
name: string,
currentFilePath: string,
symbolTable: SymbolTable,
namedImportMap: NamedImportMap,
allDefs: SymbolDefinition[],
): SymbolDefinition[] | null {
let lookupFile = currentFilePath;
let lookupName = name;
const visited = new Set<string>();
for (let depth = 0; depth < 5; depth++) {
const bindings = namedImportMap.get(lookupFile);
if (!bindings) return null;
const binding = bindings.get(lookupName);
if (!binding) return null;
const key = `${binding.sourcePath}:${binding.exportedName}`;
if (visited.has(key)) return null; // circular
visited.add(key);
const targetName = binding.exportedName;
const resolvedDefs = targetName !== lookupName || depth > 0
? symbolTable.lookupFuzzy(targetName).filter(def => def.filePath === binding.sourcePath)
: allDefs.filter(def => def.filePath === binding.sourcePath);
if (resolvedDefs.length > 0) return resolvedDefs;
// No definition in source file → follow re-export chain
lookupFile = binding.sourcePath;
lookupName = targetName;
}
return null;
}
/**
* Extract named bindings from an import AST node.
* Returns undefined if the import is not a named import (e.g., import * or default).
*
* TS: import { User, Repo as R } from './models'
* → [{local:'User', exported:'User'}, {local:'R', exported:'Repo'}]
*
* Python: from models import User, Repo as R
* → [{local:'User', exported:'User'}, {local:'R', exported:'Repo'}]
*/
export function extractNamedBindings(
importNode: any,
language: SupportedLanguages,
): { local: string; exported: string }[] | undefined {
if (language === SupportedLanguages.TypeScript || language === SupportedLanguages.JavaScript) {
return extractTsNamedBindings(importNode);
}
if (language === SupportedLanguages.Python) {
return extractPythonNamedBindings(importNode);
}
if (language === SupportedLanguages.Kotlin) {
return extractKotlinNamedBindings(importNode);
}
if (language === SupportedLanguages.Rust) {
return extractRustNamedBindings(importNode);
}
if (language === SupportedLanguages.PHP) {
return extractPhpNamedBindings(importNode);
}
if (language === SupportedLanguages.CSharp) {
return extractCsharpNamedBindings(importNode);
}
if (language === SupportedLanguages.Java) {
return extractJavaNamedBindings(importNode);
}
return undefined;
}
export function extractTsNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// import_statement > import_clause > named_imports > import_specifier*
const importClause = findChild(importNode, 'import_clause');
if (importClause) {
const namedImports = findChild(importClause, 'named_imports');
if (!namedImports) return undefined; // default import, namespace import, or side-effect
const bindings: { local: string; exported: string }[] = [];
for (let i = 0; i < namedImports.namedChildCount; i++) {
const specifier = namedImports.namedChild(i);
if (specifier?.type !== 'import_specifier') continue;
const identifiers: string[] = [];
for (let j = 0; j < specifier.namedChildCount; j++) {
const child = specifier.namedChild(j);
if (child?.type === 'identifier') identifiers.push(child.text);
}
if (identifiers.length === 1) {
bindings.push({ local: identifiers[0], exported: identifiers[0] });
} else if (identifiers.length === 2) {
// import { Foo as Bar } → exported='Foo', local='Bar'
bindings.push({ local: identifiers[1], exported: identifiers[0] });
}
}
return bindings.length > 0 ? bindings : undefined;
}
// Re-export: export { X } from './y' → export_statement > export_clause > export_specifier
const exportClause = findChild(importNode, 'export_clause');
if (exportClause) {
const bindings: { local: string; exported: string }[] = [];
for (let i = 0; i < exportClause.namedChildCount; i++) {
const specifier = exportClause.namedChild(i);
if (specifier?.type !== 'export_specifier') continue;
const identifiers: string[] = [];
for (let j = 0; j < specifier.namedChildCount; j++) {
const child = specifier.namedChild(j);
if (child?.type === 'identifier') identifiers.push(child.text);
}
if (identifiers.length === 1) {
// export { User } from './base' → re-exports User as User
bindings.push({ local: identifiers[0], exported: identifiers[0] });
} else if (identifiers.length === 2) {
// export { Repo as Repository } from './models' → name=Repo, alias=Repository
// For re-exports, the first id is the source name, second is what's exported
// When another file imports { Repository }, they get Repo from the source
bindings.push({ local: identifiers[1], exported: identifiers[0] });
}
}
return bindings.length > 0 ? bindings : undefined;
}
return undefined;
}
export function extractPythonNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// Only from import_from_statement, not plain import_statement
if (importNode.type !== 'import_from_statement') return undefined;
const bindings: { local: string; exported: string }[] = [];
for (let i = 0; i < importNode.namedChildCount; i++) {
const child = importNode.namedChild(i);
if (!child) continue;
if (child.type === 'dotted_name') {
// Skip the module_name (first dotted_name is the source module)
const fieldName = importNode.childForFieldName?.('module_name');
if (fieldName && child.startIndex === fieldName.startIndex) continue;
// This is an imported name: from x import User
const name = child.text;
if (name) bindings.push({ local: name, exported: name });
}
if (child.type === 'aliased_import') {
// from x import Repo as R
const dottedName = findChild(child, 'dotted_name');
const aliasIdent = findChild(child, 'identifier');
if (dottedName && aliasIdent) {
bindings.push({ local: aliasIdent.text, exported: dottedName.text });
}
}
}
return bindings.length > 0 ? bindings : undefined;
}
export function extractKotlinNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// import_header > identifier + import_alias > simple_identifier
if (importNode.type !== 'import_header') return undefined;
const fullIdent = findChild(importNode, 'identifier');
if (!fullIdent) return undefined;
const fullText = fullIdent.text;
const exportedName = fullText.includes('.') ? fullText.split('.').pop()! : fullText;
const importAlias = findChild(importNode, 'import_alias');
if (importAlias) {
// Aliased: import com.example.User as U
const aliasIdent = findChild(importAlias, 'simple_identifier');
if (!aliasIdent) return undefined;
return [{ local: aliasIdent.text, exported: exportedName }];
}
// Non-aliased: import com.example.User → local="User", exported="User"
// Skip wildcard imports (ending in *)
if (fullText.endsWith('.*') || fullText.endsWith('*')) return undefined;
// Skip lowercase last segments — those are member/function imports (e.g.,
// import util.OneArg.writeAudit), not class imports. Multiple member imports
// with the same function name would collide in NamedImportMap, breaking
// arity-based disambiguation.
if (exportedName[0] && exportedName[0] === exportedName[0].toLowerCase()) return undefined;
return [{ local: exportedName, exported: exportedName }];
}
export function extractRustNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// use_declaration may contain use_as_clause at any depth
if (importNode.type !== 'use_declaration') return undefined;
const bindings: { local: string; exported: string }[] = [];
collectRustBindings(importNode, bindings);
return bindings.length > 0 ? bindings : undefined;
}
function collectRustBindings(node: any, bindings: { local: string; exported: string }[]): void {
if (node.type === 'use_as_clause') {
// First identifier = exported name, second identifier = local alias
const idents: string[] = [];
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'identifier') idents.push(child.text);
// For scoped_identifier, extract the last segment
if (child?.type === 'scoped_identifier') {
const nameNode = child.childForFieldName?.('name');
if (nameNode) idents.push(nameNode.text);
}
}
if (idents.length === 2) {
bindings.push({ local: idents[1], exported: idents[0] });
}
return;
}
// Terminal identifier in a use_list: use crate::models::{User, Repo}
if (node.type === 'identifier' && node.parent?.type === 'use_list') {
bindings.push({ local: node.text, exported: node.text });
return;
}
// Skip scoped_identifier that serves as path prefix in scoped_use_list
// e.g. use crate::models::{User, Repo} — the path node "crate::models" is not an importable symbol
if (node.type === 'scoped_identifier' && node.parent?.type === 'scoped_use_list') {
return; // path prefix — the use_list sibling handles the actual symbols
}
// Terminal scoped_identifier: use crate::models::User;
// Only extract if this is a leaf (no deeper use_list/use_as_clause/scoped_use_list)
if (node.type === 'scoped_identifier') {
let hasDeeper = false;
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'use_list' || child?.type === 'use_as_clause' || child?.type === 'scoped_use_list') {
hasDeeper = true;
break;
}
}
if (!hasDeeper) {
const nameNode = node.childForFieldName?.('name');
if (nameNode) {
bindings.push({ local: nameNode.text, exported: nameNode.text });
}
return;
}
}
// Recurse into children
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child) collectRustBindings(child, bindings);
}
}
export function extractPhpNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// namespace_use_declaration > namespace_use_clause* (flat)
// namespace_use_declaration > namespace_use_group > namespace_use_clause* (grouped)
if (importNode.type !== 'namespace_use_declaration') return undefined;
const bindings: { local: string; exported: string }[] = [];
// Collect all clauses — from direct children AND from namespace_use_group
const clauses: any[] = [];
for (let i = 0; i < importNode.namedChildCount; i++) {
const child = importNode.namedChild(i);
if (child?.type === 'namespace_use_clause') {
clauses.push(child);
} else if (child?.type === 'namespace_use_group') {
for (let j = 0; j < child.namedChildCount; j++) {
const groupChild = child.namedChild(j);
if (groupChild?.type === 'namespace_use_clause') clauses.push(groupChild);
}
}
}
for (const clause of clauses) {
// Flat imports: qualified_name + name (alias)
let qualifiedName: any = null;
const names: any[] = [];
for (let j = 0; j < clause.namedChildCount; j++) {
const child = clause.namedChild(j);
if (child?.type === 'qualified_name') qualifiedName = child;
else if (child?.type === 'name') names.push(child);
}
if (qualifiedName && names.length > 0) {
// Flat aliased import: use App\Models\Repo as R;
const fullText = qualifiedName.text;
const exportedName = fullText.includes('\\') ? fullText.split('\\').pop()! : fullText;
bindings.push({ local: names[0].text, exported: exportedName });
} else if (qualifiedName && names.length === 0) {
// Flat non-aliased import: use App\Models\User;
const fullText = qualifiedName.text;
const lastSegment = fullText.includes('\\') ? fullText.split('\\').pop()! : fullText;
bindings.push({ local: lastSegment, exported: lastSegment });
} else if (!qualifiedName && names.length >= 2) {
// Grouped aliased import: {Repo as R} — first name = exported, second = alias
bindings.push({ local: names[1].text, exported: names[0].text });
} else if (!qualifiedName && names.length === 1) {
// Grouped non-aliased import: {User} in use App\Models\{User, Repo as R}
bindings.push({ local: names[0].text, exported: names[0].text });
}
}
return bindings.length > 0 ? bindings : undefined;
}
export function extractCsharpNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// using_directive with identifier (alias) + qualified_name (target)
if (importNode.type !== 'using_directive') return undefined;
let aliasIdent: any = null;
let qualifiedName: any = null;
for (let i = 0; i < importNode.namedChildCount; i++) {
const child = importNode.namedChild(i);
if (child?.type === 'identifier' && !aliasIdent) aliasIdent = child;
else if (child?.type === 'qualified_name') qualifiedName = child;
}
if (!aliasIdent || !qualifiedName) return undefined;
const fullText = qualifiedName.text;
const exportedName = fullText.includes('.') ? fullText.split('.').pop()! : fullText;
return [{ local: aliasIdent.text, exported: exportedName }];
}
export function extractJavaNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// import_declaration > scoped_identifier "com.example.models.User"
// Wildcard imports (.*) don't produce named bindings
if (importNode.type !== 'import_declaration') return undefined;
// Check for asterisk (wildcard import) — skip those
for (let i = 0; i < importNode.childCount; i++) {
const child = importNode.child(i);
if (child?.type === 'asterisk') return undefined;
}
const scopedId = findChild(importNode, 'scoped_identifier');
if (!scopedId) return undefined;
const fullText = scopedId.text;
const lastDot = fullText.lastIndexOf('.');
if (lastDot === -1) return undefined;
const className = fullText.slice(lastDot + 1);
// Skip lowercase names — those are package imports, not class imports
if (className[0] && className[0] === className[0].toLowerCase()) return undefined;
return [{ local: className, exported: className }];
}
function findChild(node: any, type: string): any {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === type) return child;
}
return null;
}
+184 -44
View File
@@ -5,12 +5,10 @@ import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { generateId } from '../../lib/utils.js';
import { SymbolTable } from './symbol-table.js';
import { ASTCache } from './ast-cache.js';
import { getLanguageFromFilename, yieldToEventLoop, DEFINITION_CAPTURE_KEYS, getDefinitionNodeFromCaptures, findEnclosingClassId, extractMethodSignature } from './utils.js';
import { isNodeExported } from './export-detection.js';
import { findSiblingChild, getLanguageFromFilename, yieldToEventLoop } from './utils.js';
import { detectFrameworkFromAST } from './framework-detection.js';
import { WorkerPool } from './workers/worker-pool.js';
import type { ParseWorkerResult, ParseWorkerInput, ExtractedImport, ExtractedCall, ExtractedHeritage, ExtractedRoute } from './workers/parse-worker.js';
import { getTreeSitterBufferSize, TREE_SITTER_MAX_BUFFER } from './constants.js';
export type FileProgressCallback = (current: number, total: number, filePath: string) => void;
@@ -21,9 +19,183 @@ export interface WorkerExtractedData {
routes: ExtractedRoute[];
}
// isNodeExported imported from ./export-detection.js (shared module)
// Re-export for backward compatibility with any external consumers
export { isNodeExported } from './export-detection.js';
const DEFINITION_CAPTURE_KEYS = [
'definition.function',
'definition.class',
'definition.interface',
'definition.method',
'definition.struct',
'definition.enum',
'definition.namespace',
'definition.module',
'definition.trait',
'definition.impl',
'definition.type',
'definition.const',
'definition.static',
'definition.typedef',
'definition.macro',
'definition.union',
'definition.property',
'definition.record',
'definition.delegate',
'definition.annotation',
'definition.constructor',
'definition.template',
] as const;
const getDefinitionNodeFromCaptures = (captureMap: Record<string, any>): any | null => {
for (const key of DEFINITION_CAPTURE_KEYS) {
if (captureMap[key]) return captureMap[key];
}
return null;
};
// ============================================================================
// EXPORT DETECTION - Language-specific visibility detection
// ============================================================================
/**
* Check if a symbol (function, class, etc.) is exported/public
* Handles all 9 supported languages with explicit logic
*
* @param node - The AST node for the symbol name
* @param name - The symbol name
* @param language - The programming language
* @returns true if the symbol is exported/public
*/
export const isNodeExported = (node: any, name: string, language: string): boolean => {
let current = node;
switch (language) {
// JavaScript/TypeScript: Check for export keyword in ancestors
case 'javascript':
case 'typescript':
while (current) {
const type = current.type;
if (type === 'export_statement' ||
type === 'export_specifier' ||
type === 'lexical_declaration' && current.parent?.type === 'export_statement') {
return true;
}
// Also check if text starts with 'export '
if (current.text?.startsWith('export ')) {
return true;
}
current = current.parent;
}
return false;
// Python: Public if no leading underscore (convention)
case 'python':
return !name.startsWith('_');
// Java: Check for 'public' modifier
// In tree-sitter Java, modifiers are siblings of the name node, not parents
case 'java':
while (current) {
// Check if this node or any sibling is a 'modifiers' node containing 'public'
if (current.parent) {
const parent = current.parent;
// Check all children of the parent for modifiers
for (let i = 0; i < parent.childCount; i++) {
const child = parent.child(i);
if (child?.type === 'modifiers' && child.text?.includes('public')) {
return true;
}
}
// Also check if the parent's text starts with 'public' (fallback)
if (parent.type === 'method_declaration' || parent.type === 'constructor_declaration') {
if (parent.text?.trimStart().startsWith('public')) {
return true;
}
}
}
current = current.parent;
}
return false;
// C#: Check for 'public' modifier in ancestors
case 'csharp':
while (current) {
if (current.type === 'modifier' || current.type === 'modifiers') {
if (current.text?.includes('public')) return true;
}
current = current.parent;
}
return false;
// Go: Uppercase first letter = exported
case 'go':
if (name.length === 0) return false;
const first = name[0];
// Must be uppercase letter (not a number or symbol)
return first === first.toUpperCase() && first !== first.toLowerCase();
// Rust: Check for 'pub' visibility modifier
case 'rust':
while (current) {
if (current.type === 'visibility_modifier') {
if (current.text?.includes('pub')) return true;
}
current = current.parent;
}
return false;
// Kotlin: Default visibility is public (unlike Java)
// visibility_modifier is inside modifiers, a sibling of the name node within the declaration
case 'kotlin':
while (current) {
if (current.parent) {
const visMod = findSiblingChild(current.parent, 'modifiers', 'visibility_modifier');
if (visMod) {
const text = visMod.text;
if (text === 'private' || text === 'internal' || text === 'protected') return false;
if (text === 'public') return true;
}
}
current = current.parent;
}
// No visibility modifier = public (Kotlin default)
return true;
// C/C++: No native export concept at language level
// Entry points will be detected via name patterns (main, etc.)
case 'c':
case 'cpp':
return false;
// Swift: Check for 'public' or 'open' access modifiers
case 'swift':
while (current) {
if (current.type === 'modifiers' || current.type === 'visibility_modifier') {
const text = current.text || '';
if (text.includes('public') || text.includes('open')) return true;
}
current = current.parent;
}
return false;
// PHP: Check for visibility modifier or top-level scope
case 'php':
while (current) {
if (current.type === 'class_declaration' ||
current.type === 'interface_declaration' ||
current.type === 'trait_declaration' ||
current.type === 'enum_declaration') {
return true;
}
if (current.type === 'visibility_modifier') {
return current.text === 'public';
}
current = current.parent;
}
return true; // Top-level functions are globally accessible
default:
return false;
}
};
// ============================================================================
// Worker-based parallel parsing
@@ -75,10 +247,7 @@ const processParsingWithWorkers = async (
}
for (const sym of result.symbols) {
symbolTable.add(sym.filePath, sym.name, sym.nodeId, sym.type, {
parameterCount: sym.parameterCount,
ownerId: sym.ownerId,
});
symbolTable.add(sym.filePath, sym.name, sym.nodeId, sym.type);
}
allImports.push(...result.imports);
@@ -117,8 +286,8 @@ const processParsingSequential = async (
if (!language) continue;
// Skip files larger than the max tree-sitter buffer (32 MB)
if (file.content.length > TREE_SITTER_MAX_BUFFER) continue;
// Skip very large files — they can crash tree-sitter or cause OOM
if (file.content.length > 512 * 1024) continue;
try {
await loadLanguage(language, file.path);
@@ -128,7 +297,7 @@ const processParsingSequential = async (
let tree;
try {
tree = parser.parse(file.content, undefined, { bufferSize: getTreeSitterBufferSize(file.content.length) });
tree = parser.parse(file.content, undefined, { bufferSize: 1024 * 256 });
} catch (parseError) {
console.warn(`Skipping unparseable file: ${file.path}`);
continue;
@@ -199,18 +368,13 @@ const processParsingSequential = async (
const definitionNodeForRange = getDefinitionNodeFromCaptures(captureMap);
const startLine = definitionNodeForRange ? definitionNodeForRange.startPosition.row : (nameNode ? nameNode.startPosition.row : 0);
const nodeId = generateId(nodeLabel, `${file.path}:${nodeName}`);
const nodeId = generateId(nodeLabel, `${file.path}:${nodeName}:${startLine}`);
const definitionNode = getDefinitionNodeFromCaptures(captureMap);
const frameworkHint = definitionNode
? detectFrameworkFromAST(language, (definitionNode.text || '').slice(0, 300))
: null;
// Extract method signature for Method/Constructor nodes
const methodSig = (nodeLabel === 'Function' || nodeLabel === 'Method' || nodeLabel === 'Constructor')
? extractMethodSignature(definitionNode)
: undefined;
const node: GraphNode = {
id: nodeId,
label: nodeLabel as any,
@@ -225,24 +389,12 @@ const processParsingSequential = async (
astFrameworkMultiplier: frameworkHint.entryPointMultiplier,
astFrameworkReason: frameworkHint.reason,
} : {}),
...(methodSig ? {
parameterCount: methodSig.parameterCount,
returnType: methodSig.returnType,
} : {}),
},
};
graph.addNode(node);
// Compute enclosing class for Method/Constructor/Property/Function — used for both ownerId and HAS_METHOD
// Function is included because Kotlin/Rust/Python capture class methods as Function nodes
const needsOwner = nodeLabel === 'Method' || nodeLabel === 'Constructor' || nodeLabel === 'Property' || nodeLabel === 'Function';
const enclosingClassId = needsOwner ? findEnclosingClassId(nameNode || definitionNodeForRange, file.path) : null;
symbolTable.add(file.path, nodeName, nodeId, nodeLabel, {
parameterCount: methodSig?.parameterCount,
ownerId: enclosingClassId ?? undefined,
});
symbolTable.add(file.path, nodeName, nodeId, nodeLabel);
const fileId = generateId('File', file.path);
@@ -258,18 +410,6 @@ const processParsingSequential = async (
};
graph.addRelationship(relationship);
// ── HAS_METHOD: link method/constructor/property to enclosing class ──
if (enclosingClassId) {
graph.addRelationship({
id: generateId('HAS_METHOD', `${enclosingClassId}->${nodeId}`),
sourceId: enclosingClassId,
targetId: nodeId,
type: 'HAS_METHOD',
confidence: 1.0,
reason: '',
});
}
});
}
};
+25 -76
View File
@@ -1,17 +1,9 @@
import { createKnowledgeGraph } from '../graph/graph.js';
import { processStructure } from './structure-processor.js';
import { processParsing } from './parsing-processor.js';
import {
processImports,
processImportsFromExtracted,
createImportMap,
createPackageMap,
createNamedImportMap,
buildImportResolutionContext
} from './import-processor.js';
import { processImports, processImportsFromExtracted, createImportMap, buildImportResolutionContext } from './import-processor.js';
import { processCalls, processCallsFromExtracted, processRoutesFromExtracted } from './call-processor.js';
import { processHeritage, processHeritageFromExtracted } from './heritage-processor.js';
import { computeMRO } from './mro-processor.js';
import { processCommunities } from './community-processor.js';
import { processProcesses } from './process-processor.js';
import { createSymbolTable } from './symbol-table.js';
@@ -21,9 +13,6 @@ import { walkRepositoryPaths, readFileContents } from './filesystem-walker.js';
import { getLanguageFromFilename } from './utils.js';
import { isLanguageAvailable } from '../tree-sitter/parser-loader.js';
import { createWorkerPool, WorkerPool } from './workers/worker-pool.js';
import fs from 'node:fs';
import path from 'node:path';
import { fileURLToPath, pathToFileURL } from 'node:url';
const isDev = process.env.NODE_ENV === 'development';
@@ -44,8 +33,6 @@ export const runPipelineFromRepo = async (
const symbolTable = createSymbolTable();
let astCache = createASTCache(AST_CACHE_CAP);
const importMap = createImportMap();
const packageMap = createPackageMap();
const namedImportMap = createNamedImportMap();
const cleanup = () => {
astCache.clear();
@@ -121,15 +108,6 @@ export const runPipelineFromRepo = async (
const totalParseable = parseableScanned.length;
if (totalParseable === 0) {
onProgress({
phase: 'parsing',
percent: 82,
message: 'No parseable files found — skipping parsing phase',
stats: { filesProcessed: 0, totalFiles: 0, nodesCreated: graph.nodeCount },
});
}
// Build byte-budget chunks
const chunks: string[][] = [];
let currentChunk: string[] = [];
@@ -162,19 +140,10 @@ export const runPipelineFromRepo = async (
// Create worker pool once, reuse across chunks
let workerPool: WorkerPool | undefined;
try {
let workerUrl = new URL('./workers/parse-worker.js', import.meta.url);
// When running under vitest, import.meta.url points to src/ where no .js exists.
// Fall back to the compiled dist/ worker so the pool can spawn real worker threads.
const thisDir = fileURLToPath(new URL('.', import.meta.url));
if (!fs.existsSync(fileURLToPath(workerUrl))) {
const distWorker = path.resolve(thisDir, '..', '..', '..', 'dist', 'core', 'ingestion', 'workers', 'parse-worker.js');
if (fs.existsSync(distWorker)) {
workerUrl = pathToFileURL(distWorker) as URL;
}
}
const workerUrl = new URL('./workers/parse-worker.js', import.meta.url);
workerPool = createWorkerPool(workerUrl);
} catch (err) {
if (isDev) console.warn('Worker pool creation failed, using sequential fallback:', (err as Error).message);
// Worker pool creation failed — sequential fallback
}
let filesParsedSoFar = 0;
@@ -223,36 +192,21 @@ export const runPipelineFromRepo = async (
if (chunkWorkerData) {
// Imports
await processImportsFromExtracted(graph, allPathObjects, chunkWorkerData.imports, importMap, undefined, repoPath, importCtx, packageMap, namedImportMap);
// Calls + Heritage + Routes — resolve in parallel (no shared mutable state between them)
// This is safe because each writes disjoint relationship types into idempotent id-keyed Maps,
// and the single-threaded event loop prevents races between synchronous addRelationship calls.
await Promise.all([
processCallsFromExtracted(
graph,
chunkWorkerData.calls,
symbolTable, importMap,
packageMap,
undefined,
namedImportMap
),
processHeritageFromExtracted(
graph,
chunkWorkerData.heritage,
symbolTable,
importMap,
packageMap
),
processRoutesFromExtracted(
graph,
chunkWorkerData.routes ?? [],
symbolTable,
importMap,
packageMap
),
]);
await processImportsFromExtracted(graph, allPathObjects, chunkWorkerData.imports, importMap, undefined, repoPath, importCtx);
// Calls — resolve immediately, then free the array
if (chunkWorkerData.calls.length > 0) {
await processCallsFromExtracted(graph, chunkWorkerData.calls, symbolTable, importMap);
}
// Heritage — resolve immediately, then free
if (chunkWorkerData.heritage.length > 0) {
await processHeritageFromExtracted(graph, chunkWorkerData.heritage, symbolTable);
}
// Routes — resolve immediately (Laravel route→controller CALLS edges)
if (chunkWorkerData.routes && chunkWorkerData.routes.length > 0) {
await processRoutesFromExtracted(graph, chunkWorkerData.routes, symbolTable, importMap);
}
} else {
await processImports(graph, chunkFiles, astCache, importMap, undefined, repoPath, allPaths, packageMap, namedImportMap);
await processImports(graph, chunkFiles, astCache, importMap, undefined, repoPath, allPaths);
sequentialChunkPaths.push(chunkPaths);
}
@@ -273,8 +227,8 @@ export const runPipelineFromRepo = async (
.filter(p => chunkContents.has(p))
.map(p => ({ path: p, content: chunkContents.get(p)! }));
astCache = createASTCache(chunkFiles.length);
await processCalls(graph, chunkFiles, astCache, symbolTable, importMap, packageMap, undefined, namedImportMap);
await processHeritage(graph, chunkFiles, astCache, symbolTable, importMap, packageMap);
await processCalls(graph, chunkFiles, astCache, symbolTable, importMap);
await processHeritage(graph, chunkFiles, astCache, symbolTable);
astCache.clear();
}
@@ -285,17 +239,12 @@ export const runPipelineFromRepo = async (
(importCtx as any).suffixIndex = null;
(importCtx as any).normalizedFileList = null;
// ── Phase 4.5: Method Resolution Order ──────────────────────────────
onProgress({
phase: 'parsing',
percent: 81,
message: 'Computing method resolution order...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
const mroResult = computeMRO(graph);
if (isDev && mroResult.entries.length > 0) {
console.log(`🔀 MRO: ${mroResult.entries.length} classes analyzed, ${mroResult.ambiguityCount} ambiguities found, ${mroResult.overrideEdges} OVERRIDES edges`);
if (isDev) {
let importsCount = 0;
for (const r of graph.iterRelationships()) {
if (r.type === 'IMPORTS') importsCount++;
}
console.log(`📊 Pipeline: graph has ${importsCount} IMPORTS, ${graph.relationshipCount} total relationships`);
}
// ── Phase 5: Communities ───────────────────────────────────────────
@@ -13,7 +13,6 @@
import { KnowledgeGraph, GraphNode, GraphRelationship, NodeLabel } from '../graph/types.js';
import { CommunityMembership } from './community-processor.js';
import { calculateEntryPointScore, isTestFile } from './entry-point-scoring.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
const isDev = process.env.NODE_ENV === 'development';
@@ -288,7 +287,7 @@ const findEntryPoints = (
// Calculate entry point score using new scoring system
const { score: baseScore, reasons } = calculateEntryPointScore(
node.properties.name,
node.properties.language ?? SupportedLanguages.JavaScript,
node.properties.language || 'javascript',
node.properties.isExported ?? false,
callers.length,
callees.length,
@@ -1,128 +0,0 @@
/**
* C# namespace import resolution.
* Handles using-directive resolution via .csproj root namespace stripping.
*/
import type { SuffixIndex } from './utils.js';
import { suffixResolve } from './utils.js';
/** C# project config parsed from .csproj files */
export interface CSharpProjectConfig {
/** Root namespace from <RootNamespace> or assembly name (default: project directory name) */
rootNamespace: string;
/** Directory containing the .csproj file */
projectDir: string;
}
/**
* Resolve a C# using-directive import path to matching .cs files.
* Tries single-file match first, then directory match for namespace imports.
*/
export function resolveCSharpImport(
importPath: string,
csharpConfigs: CSharpProjectConfig[],
normalizedFileList: string[],
allFileList: string[],
index?: SuffixIndex,
): string[] {
const namespacePath = importPath.replace(/\./g, '/');
const results: string[] = [];
for (const config of csharpConfigs) {
const nsPath = config.rootNamespace.replace(/\./g, '/');
let relative: string;
if (namespacePath.startsWith(nsPath + '/')) {
relative = namespacePath.slice(nsPath.length + 1);
} else if (namespacePath === nsPath) {
// The import IS the root namespace — resolve to all .cs files in project root
relative = '';
} else {
continue;
}
const dirPrefix = config.projectDir
? (relative ? config.projectDir + '/' + relative : config.projectDir)
: relative;
// 1. Try as single file: relative.cs (e.g., "Models/DlqMessage.cs")
if (relative) {
const candidate = dirPrefix + '.cs';
if (index) {
const result = index.get(candidate) || index.getInsensitive(candidate);
if (result) return [result];
}
// Also try suffix match
const suffixResult = index?.get(relative + '.cs') || index?.getInsensitive(relative + '.cs');
if (suffixResult) return [suffixResult];
}
// 2. Try as directory: all .cs files directly inside (namespace import)
if (index) {
const dirFiles = index.getFilesInDir(dirPrefix, '.cs');
for (const f of dirFiles) {
const normalized = f.replace(/\\/g, '/');
// Check it's a direct child by finding the dirPrefix and ensuring no deeper slashes
const prefixIdx = normalized.indexOf(dirPrefix + '/');
if (prefixIdx < 0) continue;
const afterDir = normalized.substring(prefixIdx + dirPrefix.length + 1);
if (!afterDir.includes('/')) {
results.push(f);
}
}
if (results.length > 0) return results;
}
// 3. Linear scan fallback for directory matching
if (results.length === 0) {
const dirTrail = dirPrefix + '/';
for (let i = 0; i < normalizedFileList.length; i++) {
const normalized = normalizedFileList[i];
if (!normalized.endsWith('.cs')) continue;
const prefixIdx = normalized.indexOf(dirTrail);
if (prefixIdx < 0) continue;
const afterDir = normalized.substring(prefixIdx + dirTrail.length);
if (!afterDir.includes('/')) {
results.push(allFileList[i]);
}
}
if (results.length > 0) return results;
}
}
// Fallback: suffix matching without namespace stripping (single file)
const pathParts = namespacePath.split('/').filter(Boolean);
const fallback = suffixResolve(pathParts, normalizedFileList, allFileList, index);
return fallback ? [fallback] : [];
}
/**
* Compute the directory suffix for a C# namespace import (for PackageMap).
* Returns a suffix like "/ProjectDir/Models/" or null if no config matches.
*/
export function resolveCSharpNamespaceDir(
importPath: string,
csharpConfigs: CSharpProjectConfig[],
): string | null {
const namespacePath = importPath.replace(/\./g, '/');
for (const config of csharpConfigs) {
const nsPath = config.rootNamespace.replace(/\./g, '/');
let relative: string;
if (namespacePath.startsWith(nsPath + '/')) {
relative = namespacePath.slice(nsPath.length + 1);
} else if (namespacePath === nsPath) {
relative = '';
} else {
continue;
}
const dirPrefix = config.projectDir
? (relative ? config.projectDir + '/' + relative : config.projectDir)
: relative;
if (!dirPrefix) continue;
return '/' + dirPrefix + '/';
}
return null;
}
@@ -1,58 +0,0 @@
/**
* Go package import resolution.
* Handles Go module path-based package imports.
*/
/** Go module config parsed from go.mod */
export interface GoModuleConfig {
/** Module path (e.g., "github.com/user/repo") */
modulePath: string;
}
/**
* Extract the package directory suffix from a Go import path.
* Returns the suffix string (e.g., "/internal/auth/") or null if invalid.
*/
export function resolveGoPackageDir(
importPath: string,
goModule: GoModuleConfig,
): string | null {
if (!importPath.startsWith(goModule.modulePath)) return null;
const relativePkg = importPath.slice(goModule.modulePath.length + 1);
if (!relativePkg) return null;
return '/' + relativePkg + '/';
}
/**
* Resolve a Go internal package import to all .go files in the package directory.
* Returns an array of file paths.
*/
export function resolveGoPackage(
importPath: string,
goModule: GoModuleConfig,
normalizedFileList: string[],
allFileList: string[],
): string[] {
if (!importPath.startsWith(goModule.modulePath)) return [];
// Strip module path to get relative package path
const relativePkg = importPath.slice(goModule.modulePath.length + 1); // e.g., "internal/auth"
if (!relativePkg) return [];
const pkgSuffix = '/' + relativePkg + '/';
const matches: string[] = [];
for (let i = 0; i < normalizedFileList.length; i++) {
// Prepend '/' so paths like "internal/auth/service.go" match suffix "/internal/auth/"
const normalized = '/' + normalizedFileList[i];
// File must be directly in the package directory (not a subdirectory)
if (normalized.includes(pkgSuffix) && normalized.endsWith('.go') && !normalized.endsWith('_test.go')) {
const afterPkg = normalized.substring(normalized.indexOf(pkgSuffix) + pkgSuffix.length);
if (!afterPkg.includes('/')) {
matches.push(allFileList[i]);
}
}
}
return matches;
}
@@ -1,23 +0,0 @@
/**
* Language-specific import resolvers.
* Extracted from import-processor.ts for maintainability.
*/
export { EXTENSIONS, tryResolveWithExtensions, buildSuffixIndex, suffixResolve } from './utils.js';
export type { SuffixIndex } from './utils.js';
export { KOTLIN_EXTENSIONS, appendKotlinWildcard, resolveJvmWildcard, resolveJvmMemberImport } from './jvm.js';
export { resolveGoPackageDir, resolveGoPackage } from './go.js';
export type { GoModuleConfig } from './go.js';
export { resolveCSharpImport, resolveCSharpNamespaceDir } from './csharp.js';
export type { CSharpProjectConfig } from './csharp.js';
export { resolvePhpImport } from './php.js';
export type { ComposerConfig } from './php.js';
export { resolveRustImport, tryRustModulePath } from './rust.js';
export { resolveImportPath, RESOLVE_CACHE_CAP } from './standard.js';
export type { TsconfigPaths } from './standard.js';
@@ -1,106 +0,0 @@
/**
* JVM import resolution (Java + Kotlin).
* Handles wildcard imports, member/static imports, and Kotlin-specific patterns.
*/
import type { SuffixIndex } from './utils.js';
/** Kotlin file extensions for JVM resolver reuse */
export const KOTLIN_EXTENSIONS: readonly string[] = ['.kt', '.kts'];
/**
* Append .* to a Kotlin import path if the AST has a wildcard_import sibling node.
* Pure function — returns a new string without mutating the input.
*/
export const appendKotlinWildcard = (importPath: string, importNode: any): string => {
for (let i = 0; i < importNode.childCount; i++) {
if (importNode.child(i)?.type === 'wildcard_import') {
return importPath.endsWith('.*') ? importPath : `${importPath}.*`;
}
}
return importPath;
};
/**
* Resolve a JVM wildcard import (com.example.*) to all matching files.
* Works for both Java (.java) and Kotlin (.kt, .kts).
*/
export function resolveJvmWildcard(
importPath: string,
normalizedFileList: string[],
allFileList: string[],
extensions: readonly string[],
index?: SuffixIndex,
): string[] {
// "com.example.util.*" -> "com/example/util"
const packagePath = importPath.slice(0, -2).replace(/\./g, '/');
if (index) {
const candidates = extensions.flatMap(ext => index.getFilesInDir(packagePath, ext));
// Filter to only direct children (no subdirectories)
const packageSuffix = '/' + packagePath + '/';
return candidates.filter(f => {
const normalized = f.replace(/\\/g, '/');
const idx = normalized.indexOf(packageSuffix);
if (idx < 0) return false;
const afterPkg = normalized.substring(idx + packageSuffix.length);
return !afterPkg.includes('/');
});
}
// Fallback: linear scan
const packageSuffix = '/' + packagePath + '/';
const matches: string[] = [];
for (let i = 0; i < normalizedFileList.length; i++) {
const normalized = normalizedFileList[i];
if (normalized.includes(packageSuffix) &&
extensions.some(ext => normalized.endsWith(ext))) {
const afterPackage = normalized.substring(normalized.indexOf(packageSuffix) + packageSuffix.length);
if (!afterPackage.includes('/')) {
matches.push(allFileList[i]);
}
}
}
return matches;
}
/**
* Try to resolve a JVM member/static import by stripping the member name.
* Java: "com.example.Constants.VALUE" -> resolve "com.example.Constants"
* Kotlin: "com.example.Constants.VALUE" -> resolve "com.example.Constants"
*/
export function resolveJvmMemberImport(
importPath: string,
normalizedFileList: string[],
allFileList: string[],
extensions: readonly string[],
index?: SuffixIndex,
): string | null {
// Member imports: com.example.Constants.VALUE or com.example.Constants.*
// The last segment is a member name if it starts with lowercase, is ALL_CAPS, or is a wildcard
const segments = importPath.split('.');
if (segments.length < 3) return null;
const lastSeg = segments[segments.length - 1];
if (lastSeg === '*' || /^[a-z]/.test(lastSeg) || /^[A-Z_]+$/.test(lastSeg)) {
const classPath = segments.slice(0, -1).join('/');
for (const ext of extensions) {
const classSuffix = classPath + ext;
if (index) {
const result = index.get(classSuffix) || index.getInsensitive(classSuffix);
if (result) return result;
} else {
const fullSuffix = '/' + classSuffix;
for (let i = 0; i < normalizedFileList.length; i++) {
if (normalizedFileList[i].endsWith(fullSuffix) ||
normalizedFileList[i].toLowerCase().endsWith(fullSuffix.toLowerCase())) {
return allFileList[i];
}
}
}
}
}
return null;
}
@@ -1,51 +0,0 @@
/**
* PHP PSR-4 import resolution.
* Handles use-statement resolution via composer.json autoload mappings.
*/
import type { SuffixIndex } from './utils.js';
import { suffixResolve } from './utils.js';
/** PHP Composer PSR-4 autoload config */
export interface ComposerConfig {
/** Map of namespace prefix -> directory (e.g., "App\\" -> "app/") */
psr4: Map<string, string>;
}
/**
* Resolve a PHP use-statement import path using PSR-4 mappings.
* e.g. "App\Http\Controllers\UserController" -> "app/Http/Controllers/UserController.php"
*/
export function resolvePhpImport(
importPath: string,
composerConfig: ComposerConfig | null,
allFiles: Set<string>,
normalizedFileList: string[],
allFileList: string[],
index?: SuffixIndex,
): string | null {
// Normalize: replace backslashes with forward slashes
const normalized = importPath.replace(/\\/g, '/');
// Try PSR-4 resolution if composer.json was found
if (composerConfig) {
// Sort namespaces by length descending (longest match wins)
const sorted = [...composerConfig.psr4.entries()].sort((a, b) => b[0].length - a[0].length);
for (const [nsPrefix, dirPrefix] of sorted) {
const nsPrefixSlash = nsPrefix.replace(/\\/g, '/');
if (normalized.startsWith(nsPrefixSlash + '/') || normalized === nsPrefixSlash) {
const remainder = normalized.slice(nsPrefixSlash.length).replace(/^\//, '');
const filePath = dirPrefix + (remainder ? '/' + remainder : '') + '.php';
if (allFiles.has(filePath)) return filePath;
if (index) {
const result = index.getInsensitive(filePath);
if (result) return result;
}
}
}
}
// Fallback: suffix matching (works without composer.json)
const pathParts = normalized.split('/').filter(Boolean);
return suffixResolve(pathParts, normalizedFileList, allFileList, index);
}
@@ -1,82 +0,0 @@
/**
* Rust module import resolution.
* Handles crate::, super::, self:: prefix paths and :: separators.
*/
/**
* Resolve Rust use-path to a file.
* Handles crate::, super::, self:: prefixes and :: path separators.
*/
export function resolveRustImport(
currentFile: string,
importPath: string,
allFiles: Set<string>,
): string | null {
let rustPath: string;
if (importPath.startsWith('crate::')) {
// crate:: resolves from src/ directory (standard Rust layout)
rustPath = importPath.slice(7).replace(/::/g, '/');
// Try from src/ (standard layout)
const fromSrc = tryRustModulePath('src/' + rustPath, allFiles);
if (fromSrc) return fromSrc;
// Try from repo root (non-standard)
const fromRoot = tryRustModulePath(rustPath, allFiles);
if (fromRoot) return fromRoot;
return null;
}
if (importPath.startsWith('super::')) {
// super:: = parent directory of current file's module
const currentDir = currentFile.split('/').slice(0, -1);
currentDir.pop(); // Go up one level for super::
rustPath = importPath.slice(7).replace(/::/g, '/');
const fullPath = [...currentDir, rustPath].join('/');
return tryRustModulePath(fullPath, allFiles);
}
if (importPath.startsWith('self::')) {
// self:: = current module's directory
const currentDir = currentFile.split('/').slice(0, -1);
rustPath = importPath.slice(6).replace(/::/g, '/');
const fullPath = [...currentDir, rustPath].join('/');
return tryRustModulePath(fullPath, allFiles);
}
// Bare path without prefix (e.g., from a use in a nested module)
// Convert :: to / and try suffix matching
if (importPath.includes('::')) {
rustPath = importPath.replace(/::/g, '/');
return tryRustModulePath(rustPath, allFiles);
}
return null;
}
/**
* Try to resolve a Rust module path to a file.
* Tries: path.rs, path/mod.rs, and with the last segment stripped
* (last segment might be a symbol name, not a module).
*/
export function tryRustModulePath(modulePath: string, allFiles: Set<string>): string | null {
// Try direct: path.rs
if (allFiles.has(modulePath + '.rs')) return modulePath + '.rs';
// Try directory: path/mod.rs
if (allFiles.has(modulePath + '/mod.rs')) return modulePath + '/mod.rs';
// Try path/lib.rs (for crate root)
if (allFiles.has(modulePath + '/lib.rs')) return modulePath + '/lib.rs';
// The last segment might be a symbol (function, struct, etc.), not a module.
// Strip it and try again.
const lastSlash = modulePath.lastIndexOf('/');
if (lastSlash > 0) {
const parentPath = modulePath.substring(0, lastSlash);
if (allFiles.has(parentPath + '.rs')) return parentPath + '.rs';
if (allFiles.has(parentPath + '/mod.rs')) return parentPath + '/mod.rs';
}
return null;
}
@@ -1,177 +0,0 @@
/**
* Standard import path resolution.
* Handles relative imports, path alias rewriting, and generic suffix matching.
* Used as the fallback when language-specific resolvers don't match.
*/
import type { SuffixIndex } from './utils.js';
import { tryResolveWithExtensions, suffixResolve } from './utils.js';
import { resolveRustImport } from './rust.js';
import { SupportedLanguages } from '../../../config/supported-languages.js';
/** TypeScript path alias config parsed from tsconfig.json */
export interface TsconfigPaths {
/** Map of alias prefix -> target prefix (e.g., "@/" -> "src/") */
aliases: Map<string, string>;
/** Base URL for path resolution (relative to repo root) */
baseUrl: string;
}
/** Max entries in the resolve cache. Beyond this, entries are evicted.
* 100K entries ≈ 15MB — covers the most common import patterns. */
export const RESOLVE_CACHE_CAP = 100_000;
/**
* Resolve an import path to a file path in the repository.
*
* Language-specific preprocessing is applied before the generic resolution:
* - TypeScript/JavaScript: rewrites tsconfig path aliases
* - Rust: converts crate::/super::/self:: to relative paths
*
* Java wildcards and Go package imports are handled separately in processImports
* because they resolve to multiple files.
*/
export const resolveImportPath = (
currentFile: string,
importPath: string,
allFiles: Set<string>,
allFileList: string[],
normalizedFileList: string[],
resolveCache: Map<string, string | null>,
language: SupportedLanguages,
tsconfigPaths: TsconfigPaths | null,
index?: SuffixIndex,
): string | null => {
const cacheKey = `${currentFile}::${importPath}`;
if (resolveCache.has(cacheKey)) return resolveCache.get(cacheKey) ?? null;
const cache = (result: string | null): string | null => {
// Evict oldest 20% when cap is reached instead of clearing all
if (resolveCache.size >= RESOLVE_CACHE_CAP) {
const evictCount = Math.floor(RESOLVE_CACHE_CAP * 0.2);
const iter = resolveCache.keys();
for (let i = 0; i < evictCount; i++) {
const key = iter.next().value;
if (key !== undefined) resolveCache.delete(key);
}
}
resolveCache.set(cacheKey, result);
return result;
};
// ---- TypeScript/JavaScript: rewrite path aliases ----
if (
(language === SupportedLanguages.TypeScript || language === SupportedLanguages.JavaScript) &&
tsconfigPaths &&
!importPath.startsWith('.')
) {
for (const [aliasPrefix, targetPrefix] of tsconfigPaths.aliases) {
if (importPath.startsWith(aliasPrefix)) {
const remainder = importPath.slice(aliasPrefix.length);
// Build the rewritten path relative to baseUrl
const rewritten = tsconfigPaths.baseUrl === '.'
? targetPrefix + remainder
: tsconfigPaths.baseUrl + '/' + targetPrefix + remainder;
// Try direct resolution from repo root
const resolved = tryResolveWithExtensions(rewritten, allFiles);
if (resolved) return cache(resolved);
// Try suffix matching as fallback
const parts = rewritten.split('/').filter(Boolean);
const suffixResult = suffixResolve(parts, normalizedFileList, allFileList, index);
if (suffixResult) return cache(suffixResult);
}
}
}
// ---- Rust: convert module path syntax to file paths ----
if (language === SupportedLanguages.Rust) {
// Handle grouped imports: use crate::module::{Foo, Bar, Baz}
// Extract the prefix path before ::{...} and resolve the module, not the symbols
let rustImportPath = importPath;
const braceIdx = importPath.indexOf('::{');
if (braceIdx !== -1) {
rustImportPath = importPath.substring(0, braceIdx);
} else if (importPath.startsWith('{') && importPath.endsWith('}')) {
// Top-level grouped imports: use {crate::a, crate::b}
// Iterate each part and return the first that resolves. This function returns a single
// string, so callers that need ALL edges must intercept before reaching here (see the
// Rust grouped-import blocks in processImports / processImportsBatch). This fallback
// handles any path that reaches resolveImportPath directly.
const inner = importPath.slice(1, -1);
const parts = inner.split(',').map(p => p.trim()).filter(Boolean);
for (const part of parts) {
const partResult = resolveRustImport(currentFile, part, allFiles);
if (partResult) return cache(partResult);
}
return cache(null);
}
const rustResult = resolveRustImport(currentFile, rustImportPath, allFiles);
if (rustResult) return cache(rustResult);
// Fall through to generic resolution if Rust-specific didn't match
}
// ---- Python relative imports (PEP 328): .module, ..module, ... ----
if (language === SupportedLanguages.Python && importPath.startsWith('.')) {
const dotMatch = importPath.match(/^(\.+)(.*)/);
if (dotMatch) {
const dotCount = dotMatch[1].length;
const modulePart = dotMatch[2]; // e.g., "models" from ".models"
const dirParts = currentFile.split('/').slice(0, -1); // remove filename
// Navigate up: 1 dot = same package, 2 dots = parent package, etc.
// First dot means "current package", each additional dot goes up one level
for (let i = 1; i < dotCount; i++) {
dirParts.pop();
}
if (modulePart) {
// from .models import User → resolve "models" relative to current package
const modulePath = modulePart.replace(/\./g, '/');
dirParts.push(...modulePath.split('/'));
}
const basePath = dirParts.join('/');
const resolved = tryResolveWithExtensions(basePath, allFiles);
return cache(resolved);
}
}
// ---- Generic relative import resolution (./ and ../) ----
const currentDir = currentFile.split('/').slice(0, -1);
const parts = importPath.split('/');
for (const part of parts) {
if (part === '.') continue;
if (part === '..') {
currentDir.pop();
} else {
currentDir.push(part);
}
}
const basePath = currentDir.join('/');
if (importPath.startsWith('.')) {
const resolved = tryResolveWithExtensions(basePath, allFiles);
return cache(resolved);
}
// ---- Generic package/absolute import resolution (suffix matching) ----
// Java wildcards are handled in processImports, not here
if (importPath.endsWith('.*')) {
return cache(null);
}
// C/C++ includes use actual file paths (e.g. "animal.h") — don't convert dots to slashes
const isCpp = language === SupportedLanguages.C || language === SupportedLanguages.CPlusPlus;
const pathLike = importPath.includes('/') || isCpp
? importPath
: importPath.replace(/\./g, '/');
const pathParts = pathLike.split('/').filter(Boolean);
const resolved = suffixResolve(pathParts, normalizedFileList, allFileList, index);
return cache(resolved);
};
@@ -1,156 +0,0 @@
/**
* Shared utilities for import resolution.
* Extracted from import-processor.ts to reduce file size.
*/
/** All file extensions to try during resolution */
export const EXTENSIONS = [
'',
// TypeScript/JavaScript
'.tsx', '.ts', '.jsx', '.js', '/index.tsx', '/index.ts', '/index.jsx', '/index.js',
// Python
'.py', '/__init__.py',
// Java
'.java',
// Kotlin
'.kt', '.kts',
// C/C++
'.c', '.h', '.cpp', '.hpp', '.cc', '.cxx', '.hxx', '.hh',
// C#
'.cs',
// Go
'.go',
// Rust
'.rs', '/mod.rs',
// PHP
'.php', '.phtml',
// Swift
'.swift',
];
/**
* Try to match a path (with extensions) against the known file set.
* Returns the matched file path or null.
*/
export function tryResolveWithExtensions(
basePath: string,
allFiles: Set<string>,
): string | null {
for (const ext of EXTENSIONS) {
const candidate = basePath + ext;
if (allFiles.has(candidate)) return candidate;
}
return null;
}
/**
* Build a suffix index for O(1) endsWith lookups.
* Maps every possible path suffix to its original file path.
* e.g. for "src/com/example/Foo.java":
* "Foo.java" -> "src/com/example/Foo.java"
* "example/Foo.java" -> "src/com/example/Foo.java"
* "com/example/Foo.java" -> "src/com/example/Foo.java"
* etc.
*/
export interface SuffixIndex {
/** Exact suffix lookup (case-sensitive) */
get(suffix: string): string | undefined;
/** Case-insensitive suffix lookup */
getInsensitive(suffix: string): string | undefined;
/** Get all files in a directory suffix */
getFilesInDir(dirSuffix: string, extension: string): string[];
}
export function buildSuffixIndex(normalizedFileList: string[], allFileList: string[]): SuffixIndex {
// Map: normalized suffix -> original file path
const exactMap = new Map<string, string>();
// Map: lowercase suffix -> original file path
const lowerMap = new Map<string, string>();
// Map: directory suffix -> list of file paths in that directory
const dirMap = new Map<string, string[]>();
for (let i = 0; i < normalizedFileList.length; i++) {
const normalized = normalizedFileList[i];
const original = allFileList[i];
const parts = normalized.split('/');
// Index all suffixes: "a/b/c.java" -> ["c.java", "b/c.java", "a/b/c.java"]
for (let j = parts.length - 1; j >= 0; j--) {
const suffix = parts.slice(j).join('/');
// Only store first match (longest path wins for ambiguous suffixes)
if (!exactMap.has(suffix)) {
exactMap.set(suffix, original);
}
const lower = suffix.toLowerCase();
if (!lowerMap.has(lower)) {
lowerMap.set(lower, original);
}
}
// Index directory membership
const lastSlash = normalized.lastIndexOf('/');
if (lastSlash >= 0) {
// Build all directory suffixes
const dirParts = parts.slice(0, -1);
const fileName = parts[parts.length - 1];
const ext = fileName.substring(fileName.lastIndexOf('.'));
for (let j = dirParts.length - 1; j >= 0; j--) {
const dirSuffix = dirParts.slice(j).join('/');
const key = `${dirSuffix}:${ext}`;
let list = dirMap.get(key);
if (!list) {
list = [];
dirMap.set(key, list);
}
list.push(original);
}
}
}
return {
get: (suffix: string) => exactMap.get(suffix),
getInsensitive: (suffix: string) => lowerMap.get(suffix.toLowerCase()),
getFilesInDir: (dirSuffix: string, extension: string) => {
return dirMap.get(`${dirSuffix}:${extension}`) || [];
},
};
}
/**
* Suffix-based resolution using index. O(1) per lookup instead of O(files).
*/
export function suffixResolve(
pathParts: string[],
normalizedFileList: string[],
allFileList: string[],
index?: SuffixIndex,
): string | null {
if (index) {
for (let i = 0; i < pathParts.length; i++) {
const suffix = pathParts.slice(i).join('/');
for (const ext of EXTENSIONS) {
const suffixWithExt = suffix + ext;
const result = index.get(suffixWithExt) || index.getInsensitive(suffixWithExt);
if (result) return result;
}
}
return null;
}
// Fallback: linear scan (for backward compatibility)
for (let i = 0; i < pathParts.length; i++) {
const suffix = pathParts.slice(i).join('/');
for (const ext of EXTENSIONS) {
const suffixWithExt = suffix + ext;
const suffixPattern = '/' + suffixWithExt;
const matchIdx = normalizedFileList.findIndex(filePath =>
filePath.endsWith(suffixPattern) || filePath.toLowerCase().endsWith(suffixPattern.toLowerCase())
);
if (matchIdx !== -1) {
return allFileList[matchIdx];
}
}
}
return null;
}
@@ -1,123 +0,0 @@
/**
* Symbol Resolver
*
* Import-filtered candidate narrowing for bare identifier resolution.
* NOT FQN resolution — does not parse qualifiers (ns::Bar, com.foo.Bar).
*
* Shared between heritage-processor.ts and call-processor.ts.
*/
import type { SymbolTable, SymbolDefinition } from './symbol-table.js';
import type { ImportMap, PackageMap, NamedImportMap } from './import-processor.js';
import { isFileInPackageDir } from './import-processor.js';
import { walkBindingChain } from './named-binding-extraction.js';
/** Resolution tier for internal tracking, logging, and test assertions. */
export type ResolutionTier = 'same-file' | 'import-scoped' | 'unique-global';
/** Internal resolution result preserving tier metadata. */
export interface InternalResolution {
definition: SymbolDefinition;
tier: ResolutionTier;
candidateCount: number;
}
/**
* Resolve a bare identifier to its best-matching definition using import context.
*
* Resolution tiers (highest confidence first):
* 1. Same file (lookupExactFull — authoritative)
* 2. Import-scoped (lookupFuzzy filtered by importMap — acceptable)
* 3. Unique global (lookupFuzzy with exactly 1 match — acceptable fallback)
*
* If multiple global candidates remain after filtering, returns null.
* A wrong edge is worse than no edge.
*/
export const resolveSymbol = (
name: string,
currentFilePath: string,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
): SymbolDefinition | null => {
return resolveSymbolInternal(name, currentFilePath, symbolTable, importMap, packageMap, namedImportMap)?.definition ?? null;
};
/** Internal resolver preserving tier metadata for logging and test assertions. */
export const resolveSymbolInternal = (
name: string,
currentFilePath: string,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
): InternalResolution | null => {
// Tier 1: Same file — authoritative match
const localDef = symbolTable.lookupExactFull(currentFilePath, name);
if (localDef) return { definition: localDef, tier: 'same-file', candidateCount: 1 };
// Get all global definitions for subsequent tiers
const allDefs = symbolTable.lookupFuzzy(name);
// Tier 2a-named: Check named bindings BEFORE the empty-allDefs early return,
// because aliased imports (import { User as U }) mean lookupFuzzy('U') returns
// empty but we can resolve via the exported name.
if (namedImportMap) {
const result = resolveNamedBindingChain(name, currentFilePath, symbolTable, namedImportMap, allDefs);
if (result) return result;
}
if (allDefs.length === 0) return null;
// Tier 2a: Import-scoped — check if any definition is in a file imported by currentFile
const importedFiles = importMap.get(currentFilePath);
if (importedFiles) {
for (const def of allDefs) {
if (importedFiles.has(def.filePath)) {
return { definition: def, tier: 'import-scoped', candidateCount: allDefs.length };
}
}
}
// Tier 2b: Package-scoped — check if any definition is in a package/namespace dir imported by currentFile
// Used for Go packages and C# namespace imports to avoid ImportMap expansion bloat
const importedPackages = packageMap?.get(currentFilePath);
if (importedPackages) {
for (const def of allDefs) {
for (const dirSuffix of importedPackages) {
if (isFileInPackageDir(def.filePath, dirSuffix)) {
return { definition: def, tier: 'import-scoped', candidateCount: allDefs.length };
}
}
}
}
// Tier 3: Unique global — ONLY if exactly one candidate exists
// Ambiguous global matches are refused. A wrong edge is worse than no edge.
if (allDefs.length === 1) {
return { definition: allDefs[0], tier: 'unique-global', candidateCount: 1 };
}
// Ambiguous: multiple global candidates, no import or same-file match → refuse
return null;
};
/**
* Follow re-export chains through NamedImportMap.
* Delegates chain-walking to the shared walkBindingChain utility, then
* applies symbol-resolver semantics: exactly one match required.
*/
const resolveNamedBindingChain = (
name: string,
currentFilePath: string,
symbolTable: SymbolTable,
namedImportMap: NamedImportMap,
allDefs: SymbolDefinition[],
): InternalResolution | null => {
const defs = walkBindingChain(name, currentFilePath, symbolTable, namedImportMap, allDefs);
if (defs?.length === 1) {
return { definition: defs[0], tier: 'import-scoped', candidateCount: defs.length };
}
return null;
};
+14 -45
View File
@@ -2,22 +2,13 @@ export interface SymbolDefinition {
nodeId: string;
filePath: string;
type: string; // 'Function', 'Class', etc.
parameterCount?: number;
/** Links Method/Constructor to owning Class/Struct/Trait nodeId */
ownerId?: string;
}
export interface SymbolTable {
/**
* Register a new symbol definition
*/
add: (
filePath: string,
name: string,
nodeId: string,
type: string,
metadata?: { parameterCount?: number; ownerId?: string }
) => void;
add: (filePath: string, name: string, nodeId: string, type: string) => void;
/**
* High Confidence: Look for a symbol specifically inside a file
@@ -25,12 +16,6 @@ export interface SymbolTable {
*/
lookupExact: (filePath: string, name: string) => string | undefined;
/**
* High Confidence: Look for a symbol in a specific file, returning full definition.
* Includes type information needed for heritage resolution (Class vs Interface).
*/
lookupExactFull: (filePath: string, name: string) => SymbolDefinition | undefined;
/**
* Low Confidence: Look for a symbol anywhere in the project
* Used when imports are missing or for framework magic
@@ -49,48 +34,32 @@ export interface SymbolTable {
}
export const createSymbolTable = (): SymbolTable => {
// 1. File-Specific Index — stores full SymbolDefinition for O(1) lookupExactFull
// Structure: FilePath -> (SymbolName -> SymbolDefinition)
const fileIndex = new Map<string, Map<string, SymbolDefinition>>();
// 1. File-Specific Index (The "Good" one)
// Structure: FilePath -> (SymbolName -> NodeID)
const fileIndex = new Map<string, Map<string, string>>();
// 2. Global Reverse Index (The "Backup")
// Structure: SymbolName -> [List of Definitions]
const globalIndex = new Map<string, SymbolDefinition[]>();
const add = (
filePath: string,
name: string,
nodeId: string,
type: string,
metadata?: { parameterCount?: number; ownerId?: string }
) => {
const def: SymbolDefinition = {
nodeId,
filePath,
type,
...(metadata?.parameterCount !== undefined ? { parameterCount: metadata.parameterCount } : {}),
...(metadata?.ownerId !== undefined ? { ownerId: metadata.ownerId } : {}),
};
// A. Add to File Index (shared reference — zero additional memory)
const add = (filePath: string, name: string, nodeId: string, type: string) => {
// A. Add to File Index
if (!fileIndex.has(filePath)) {
fileIndex.set(filePath, new Map());
}
fileIndex.get(filePath)!.set(name, def);
fileIndex.get(filePath)!.set(name, nodeId);
// B. Add to Global Index (same object reference)
// B. Add to Global Index
if (!globalIndex.has(name)) {
globalIndex.set(name, []);
}
globalIndex.get(name)!.push(def);
globalIndex.get(name)!.push({ nodeId, filePath, type });
};
const lookupExact = (filePath: string, name: string): string | undefined => {
return fileIndex.get(filePath)?.get(name)?.nodeId;
};
const lookupExactFull = (filePath: string, name: string): SymbolDefinition | undefined => {
return fileIndex.get(filePath)?.get(name);
const fileSymbols = fileIndex.get(filePath);
if (!fileSymbols) return undefined;
return fileSymbols.get(name);
};
const lookupFuzzy = (name: string): SymbolDefinition[] => {
@@ -107,5 +76,5 @@ export const createSymbolTable = (): SymbolTable => {
globalIndex.clear();
};
return { add, lookupExact, lookupExactFull, lookupFuzzy, getStats, clear };
};
return { add, lookupExact, lookupFuzzy, getStats, clear };
};
@@ -47,10 +47,6 @@ export const TYPESCRIPT_QUERIES = `
(import_statement
source: (string) @import.source) @import
; Re-export statements: export { X } from './y'
(export_statement
source: (string) @import.source) @import
(call_expression
function: (identifier) @call.name) @call
@@ -58,10 +54,6 @@ export const TYPESCRIPT_QUERIES = `
function: (member_expression
property: (property_identifier) @call.name)) @call
; Constructor calls: new Foo()
(new_expression
constructor: (identifier) @call.name) @call
; Heritage queries - class extends
(class_declaration
name: (type_identifier) @heritage.class
@@ -77,7 +69,7 @@ export const TYPESCRIPT_QUERIES = `
(type_identifier) @heritage.implements))) @heritage.impl
`;
// JavaScript queries - works with tree-sitter-javascript
// JavaScript queries - works with tree-sitter-javascript
export const JAVASCRIPT_QUERIES = `
(class_declaration
name: (identifier) @name) @definition.class
@@ -113,10 +105,6 @@ export const JAVASCRIPT_QUERIES = `
(import_statement
source: (string) @import.source) @import
; Re-export statements: export { X } from './y'
(export_statement
source: (string) @import.source) @import
(call_expression
function: (identifier) @call.name) @call
@@ -124,10 +112,6 @@ export const JAVASCRIPT_QUERIES = `
function: (member_expression
property: (property_identifier) @call.name)) @call
; Constructor calls: new Foo()
(new_expression
constructor: (identifier) @call.name) @call
; Heritage queries - class extends (JavaScript uses different AST than TypeScript)
; In tree-sitter-javascript, class_heritage directly contains the parent identifier
(class_declaration
@@ -150,9 +134,6 @@ export const PYTHON_QUERIES = `
(import_from_statement
module_name: (dotted_name) @import.source) @import
(import_from_statement
module_name: (relative_import) @import.source) @import
(call
function: (identifier) @call.name) @call
@@ -186,9 +167,6 @@ export const JAVA_QUERIES = `
(method_invocation name: (identifier) @call.name) @call
(method_invocation object: (_) name: (identifier) @call.name) @call
; Constructor calls: new Foo()
(object_creation_expression type: (type_identifier) @call.name) @call
; Heritage - extends class
(class_declaration name: (identifier) @heritage.class
(superclass (type_identifier) @heritage.extends)) @heritage
@@ -200,17 +178,10 @@ export const JAVA_QUERIES = `
// C queries - works with tree-sitter-c
export const C_QUERIES = `
; Functions (direct declarator)
; Functions
(function_definition declarator: (function_declarator declarator: (identifier) @name)) @definition.function
(declaration declarator: (function_declarator declarator: (identifier) @name)) @definition.function
; Functions returning pointers (pointer_declarator wraps function_declarator)
(function_definition declarator: (pointer_declarator declarator: (function_declarator declarator: (identifier) @name))) @definition.function
(declaration declarator: (pointer_declarator declarator: (function_declarator declarator: (identifier) @name))) @definition.function
; Functions returning double pointers (nested pointer_declarator)
(function_definition declarator: (pointer_declarator declarator: (pointer_declarator declarator: (function_declarator declarator: (identifier) @name)))) @definition.function
; Structs, Unions, Enums, Typedefs
(struct_specifier name: (type_identifier) @name) @definition.struct
(union_specifier name: (type_identifier) @name) @definition.union
@@ -238,26 +209,15 @@ export const GO_QUERIES = `
; Types
(type_declaration (type_spec name: (type_identifier) @name type: (struct_type))) @definition.struct
(type_declaration (type_spec name: (type_identifier) @name type: (interface_type))) @definition.interface
(type_declaration (type_spec name: (type_identifier) @name)) @definition.type
; Imports
(import_declaration (import_spec path: (interpreted_string_literal) @import.source)) @import
(import_declaration (import_spec_list (import_spec path: (interpreted_string_literal) @import.source))) @import
; Struct embedding (anonymous fields = inheritance)
(type_declaration
(type_spec
name: (type_identifier) @heritage.class
type: (struct_type
(field_declaration_list
(field_declaration
type: (type_identifier) @heritage.extends))))) @definition.struct
; Calls
(call_expression function: (identifier) @call.name) @call
(call_expression function: (selector_expression field: (field_identifier) @call.name)) @call
; Struct literal construction: User{Name: "Alice"}
(composite_literal type: (type_identifier) @call.name) @call
`;
// C++ queries - works with tree-sitter-cpp
@@ -268,46 +228,10 @@ export const CPP_QUERIES = `
(namespace_definition name: (namespace_identifier) @name) @definition.namespace
(enum_specifier name: (type_identifier) @name) @definition.enum
; Typedefs and unions (common in C-style headers and mixed C/C++ code)
(type_definition declarator: (type_identifier) @name) @definition.typedef
(union_specifier name: (type_identifier) @name) @definition.union
; Macros
(preproc_function_def name: (identifier) @name) @definition.macro
(preproc_def name: (identifier) @name) @definition.macro
; Functions & Methods (direct declarator)
; Functions & Methods
(function_definition declarator: (function_declarator declarator: (identifier) @name)) @definition.function
(function_definition declarator: (function_declarator declarator: (qualified_identifier name: (identifier) @name))) @definition.method
; Functions/methods returning pointers (pointer_declarator wraps function_declarator)
(function_definition declarator: (pointer_declarator declarator: (function_declarator declarator: (identifier) @name))) @definition.function
(function_definition declarator: (pointer_declarator declarator: (function_declarator declarator: (qualified_identifier name: (identifier) @name)))) @definition.method
; Functions/methods returning double pointers (nested pointer_declarator)
(function_definition declarator: (pointer_declarator declarator: (pointer_declarator declarator: (function_declarator declarator: (identifier) @name)))) @definition.function
(function_definition declarator: (pointer_declarator declarator: (pointer_declarator declarator: (function_declarator declarator: (qualified_identifier name: (identifier) @name))))) @definition.method
; Functions/methods returning references (reference_declarator wraps function_declarator)
(function_definition declarator: (reference_declarator (function_declarator declarator: (identifier) @name))) @definition.function
(function_definition declarator: (reference_declarator (function_declarator declarator: (qualified_identifier name: (identifier) @name)))) @definition.method
; Destructors (destructor_name is distinct from identifier in tree-sitter-cpp)
(function_definition declarator: (function_declarator declarator: (qualified_identifier name: (destructor_name) @name))) @definition.method
; Function declarations / prototypes (common in headers)
(declaration declarator: (function_declarator declarator: (identifier) @name)) @definition.function
(declaration declarator: (pointer_declarator declarator: (function_declarator declarator: (identifier) @name))) @definition.function
; Inline class method declarations (inside class body, no body: void Foo();)
(field_declaration declarator: (function_declarator declarator: (identifier) @name)) @definition.method
; Inline class method definitions (inside class body, with body: void Foo() { ... })
(field_declaration_list
(function_definition
declarator: (function_declarator
declarator: [(field_identifier) (identifier) (operator_name) (destructor_name)] @name))) @definition.method
; Templates
(template_declaration (class_specifier name: (type_identifier) @name)) @definition.template
(template_declaration (function_definition declarator: (function_declarator declarator: (identifier) @name))) @definition.template
@@ -321,9 +245,6 @@ export const CPP_QUERIES = `
(call_expression function: (qualified_identifier name: (identifier) @call.name)) @call
(call_expression function: (template_function name: (identifier) @call.name)) @call
; Constructor calls: new User()
(new_expression type: (type_identifier) @call.name) @call
; Heritage
(class_specifier name: (type_identifier) @heritage.class
(base_class_clause (type_identifier) @heritage.extends)) @heritage
@@ -341,11 +262,9 @@ export const CSHARP_QUERIES = `
(record_declaration name: (identifier) @name) @definition.record
(delegate_declaration name: (identifier) @name) @definition.delegate
; Namespaces (block form and C# 10+ file-scoped form)
; Namespaces
(namespace_declaration name: (identifier) @name) @definition.namespace
(namespace_declaration name: (qualified_name) @name) @definition.namespace
(file_scoped_namespace_declaration name: (identifier) @name) @definition.namespace
(file_scoped_namespace_declaration name: (qualified_name) @name) @definition.namespace
; Methods & Properties
(method_declaration name: (identifier) @name) @definition.method
@@ -353,10 +272,6 @@ export const CSHARP_QUERIES = `
(constructor_declaration name: (identifier) @name) @definition.constructor
(property_declaration name: (identifier) @name) @definition.property
; Primary constructors (C# 12): class User(string name, int age) { }
(class_declaration name: (identifier) @name (parameter_list) @definition.constructor)
(record_declaration name: (identifier) @name (parameter_list) @definition.constructor)
; Using
(using_directive (qualified_name) @import.source) @import
(using_directive (identifier) @import.source) @import
@@ -365,17 +280,11 @@ export const CSHARP_QUERIES = `
(invocation_expression function: (identifier) @call.name) @call
(invocation_expression function: (member_access_expression name: (identifier) @call.name)) @call
; Constructor calls: new Foo() and new Foo { Props }
(object_creation_expression type: (identifier) @call.name) @call
; Target-typed new (C# 9): User u = new("x", 5)
(variable_declaration type: (identifier) @call.name (variable_declarator (implicit_object_creation_expression) @call))
; Heritage
(class_declaration name: (identifier) @heritage.class
(base_list (identifier) @heritage.extends)) @heritage
(base_list (simple_base_type (identifier) @heritage.extends))) @heritage
(class_declaration name: (identifier) @heritage.class
(base_list (generic_name (identifier) @heritage.extends))) @heritage
(base_list (simple_base_type (generic_name (identifier) @heritage.extends)))) @heritage
`;
// Rust queries - works with tree-sitter-rust
@@ -385,8 +294,7 @@ export const RUST_QUERIES = `
(struct_item name: (type_identifier) @name) @definition.struct
(enum_item name: (type_identifier) @name) @definition.enum
(trait_item name: (type_identifier) @name) @definition.trait
(impl_item type: (type_identifier) @name !trait) @definition.impl
(impl_item type: (generic_type type: (type_identifier) @name) !trait) @definition.impl
(impl_item type: (type_identifier) @name) @definition.impl
(mod_item name: (identifier) @name) @definition.module
; Type aliases, const, static, macros
@@ -404,14 +312,9 @@ export const RUST_QUERIES = `
(call_expression function: (scoped_identifier name: (identifier) @call.name)) @call
(call_expression function: (generic_function function: (identifier) @call.name)) @call
; Struct literal construction: User { name: value }
(struct_expression name: (type_identifier) @call.name) @call
; Heritage (trait implementation) — all combinations of concrete/generic trait × concrete/generic type
; Heritage (trait implementation)
(impl_item trait: (type_identifier) @heritage.trait type: (type_identifier) @heritage.class) @heritage
(impl_item trait: (generic_type type: (type_identifier) @heritage.trait) type: (type_identifier) @heritage.class) @heritage
(impl_item trait: (type_identifier) @heritage.trait type: (generic_type type: (type_identifier) @heritage.class)) @heritage
(impl_item trait: (generic_type type: (type_identifier) @heritage.trait) type: (generic_type type: (type_identifier) @heritage.class)) @heritage
`;
// PHP queries - works with tree-sitter-php (php_only grammar)
@@ -473,9 +376,6 @@ export const PHP_QUERIES = `
(scoped_call_expression
name: (name) @call.name) @call
; Constructor call: new User()
(object_creation_expression (name) @call.name) @call
; ── Heritage: extends ────────────────────────────────────────────────────────
(class_declaration
name: (name) @heritage.class
@@ -627,11 +527,6 @@ export const SWIFT_QUERIES = `
; Heritage - protocol inheritance
(protocol_declaration name: (type_identifier) @heritage.class
(inheritance_specifier inherits_from: (user_type (type_identifier) @heritage.extends))) @heritage
; Heritage - extension protocol conformance (e.g. extension Foo: SomeProtocol)
; Extensions wrap the name in user_type unlike class/struct/enum declarations
(class_declaration "extension" name: (user_type (type_identifier) @heritage.class)
(inheritance_specifier inherits_from: (user_type (type_identifier) @heritage.extends))) @heritage
`;
export const LANGUAGE_QUERIES: Record<SupportedLanguages, string> = {
-124
View File
@@ -1,124 +0,0 @@
import type { SyntaxNode } from './utils.js';
import { FUNCTION_NODE_TYPES, extractFunctionName } from './utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
import { typeConfigs, TYPED_PARAMETER_TYPES } from './type-extractors/index.js';
/**
* Per-file scoped type environment: maps (scope, variableName) → typeName.
* Scope-aware: variables inside functions are keyed by function name,
* file-level variables use the '' (empty string) scope.
*
* Design constraints:
* - Explicit-only: only type annotations, never inferred types
* - Scope-aware: function-local variables don't collide across functions
* - Conservative: complex/generic types extract the base name only
* - Per-file: built once, used for receiver resolution, then discarded
*/
export type TypeEnv = Map<string, Map<string, string>>;
/** File-level scope key */
const FILE_SCOPE = '';
/**
* Look up a variable's type in the TypeEnv, trying the call's enclosing
* function scope first, then falling back to file-level scope.
*/
export const lookupTypeEnv = (
env: TypeEnv,
varName: string,
callNode: SyntaxNode,
): string | undefined => {
// Determine the enclosing function scope for the call
const scopeKey = findEnclosingScopeKey(callNode);
// Try function-local scope first
if (scopeKey) {
const scopeEnv = env.get(scopeKey);
if (scopeEnv) {
const result = scopeEnv.get(varName);
if (result) return result;
}
}
// Fall back to file-level scope
const fileEnv = env.get(FILE_SCOPE);
return fileEnv?.get(varName);
};
/** Find the enclosing function name for scope lookup. */
const findEnclosingScopeKey = (node: SyntaxNode): string | undefined => {
let current = node.parent;
while (current) {
if (FUNCTION_NODE_TYPES.has(current.type)) {
const { funcName } = extractFunctionName(current);
if (funcName) return `${funcName}@${current.startIndex}`;
}
current = current.parent;
}
return undefined;
};
/**
* Build a scoped TypeEnv from a tree-sitter AST for a given language.
* Walks the tree tracking enclosing function scopes, so that variables
* inside different functions don't collide.
*/
export const buildTypeEnv = (
tree: { rootNode: SyntaxNode },
language: SupportedLanguages,
): TypeEnv => {
const env: TypeEnv = new Map();
walkForTypes(tree.rootNode, language, env, FILE_SCOPE);
return env;
};
const walkForTypes = (
node: SyntaxNode,
language: SupportedLanguages,
env: TypeEnv,
currentScope: string,
): void => {
// Detect scope boundaries (function/method definitions)
let scope = currentScope;
if (FUNCTION_NODE_TYPES.has(node.type)) {
const { funcName } = extractFunctionName(node);
if (funcName) scope = `${funcName}@${node.startIndex}`;
}
// Get or create the sub-map for this scope
if (!env.has(scope)) env.set(scope, new Map());
const scopeEnv = env.get(scope)!;
// Check if this node provides type information
extractTypeBinding(node, language, scopeEnv);
// Recurse into children
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (child) walkForTypes(child, language, env, scope);
}
};
/**
* Try to extract a (variableName → typeName) binding from a single AST node.
* Delegates to per-language type configurations.
*/
const extractTypeBinding = (
node: SyntaxNode,
language: SupportedLanguages,
env: Map<string, string>,
): void => {
// === PARAMETERS (most languages) ===
// This guard eliminates 90%+ of calls before any language dispatch.
if (TYPED_PARAMETER_TYPES.has(node.type)) {
const config = typeConfigs[language];
config.extractParameter(node, env);
return;
}
// === Per-language declaration extraction ===
const config = typeConfigs[language];
if (config.declarationNodeTypes.has(node.type)) {
config.extractDeclaration(node, env);
}
};
@@ -1,63 +0,0 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'declaration',
]);
/** C++: Type x = ...; Type* x; Type& x; */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
const typeNode = node.childForFieldName('type');
if (!typeNode) return;
const typeName = extractSimpleTypeName(typeNode);
if (!typeName) return;
const declarator = node.childForFieldName('declarator');
if (!declarator) return;
// init_declarator: Type x = value
const nameNode = declarator.type === 'init_declarator'
? declarator.childForFieldName('declarator')
: declarator;
if (!nameNode) return;
// Handle pointer/reference declarators
const finalName = nameNode.type === 'pointer_declarator' || nameNode.type === 'reference_declarator'
? nameNode.firstNamedChild
: nameNode;
if (!finalName) return;
const varName = extractVarName(finalName);
if (varName) env.set(varName, typeName);
};
/** C/C++: parameter_declaration → type declarator */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'parameter_declaration') {
typeNode = node.childForFieldName('type');
const declarator = node.childForFieldName('declarator');
if (declarator) {
nameNode = declarator.type === 'pointer_declarator' || declarator.type === 'reference_declarator'
? declarator.firstNamedChild
: declarator;
}
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -1,93 +0,0 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, findChildByType } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'local_declaration_statement',
'variable_declaration',
'field_declaration',
]);
/** C#: Type x = ...; var x = new Type(); */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
// C# tree-sitter: local_declaration_statement > variable_declaration > ...
// Recursively descend through wrapper nodes
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child) continue;
if (child.type === 'variable_declaration' || child.type === 'local_declaration_statement') {
extractDeclaration(child, env);
return;
}
}
// At variable_declaration level: first child is type, rest are variable_declarators
let typeNode: SyntaxNode | null = null;
const declarators: SyntaxNode[] = [];
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child) continue;
if (!typeNode && child.type !== 'variable_declarator' && child.type !== 'equals_value_clause') {
// First non-declarator child is the type (identifier, implicit_type, generic_name, etc.)
typeNode = child;
}
if (child.type === 'variable_declarator') {
declarators.push(child);
}
}
if (!typeNode || declarators.length === 0) return;
// Handle 'var x = new Foo()' — infer from object_creation_expression
let typeName: string | undefined;
if (typeNode.type === 'implicit_type' && typeNode.text === 'var') {
// Try to infer from initializer: var x = new Foo()
// C# tree-sitter puts object_creation_expression as direct child of variable_declarator
if (declarators.length === 1) {
const initializer = findChildByType(declarators[0], 'object_creation_expression')
?? findChildByType(declarators[0], 'equals_value_clause')?.firstNamedChild;
if (initializer?.type === 'object_creation_expression') {
const ctorType = initializer.childForFieldName('type');
if (ctorType) typeName = extractSimpleTypeName(ctorType);
}
}
} else {
typeName = extractSimpleTypeName(typeNode);
}
if (!typeName) return;
for (const decl of declarators) {
const nameNode = decl.childForFieldName('name') ?? decl.firstNamedChild;
if (nameNode) {
const varName = extractVarName(nameNode);
if (varName) env.set(varName, typeName);
}
}
};
/** C#: parameter → type name */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'parameter') {
typeNode = node.childForFieldName('type');
nameNode = node.childForFieldName('name');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -1,104 +0,0 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'var_declaration',
'var_spec',
'short_var_declaration',
]);
/** Go: var x Foo */
const extractGoVarDeclaration = (node: SyntaxNode, env: Map<string, string>): void => {
// Go var_declaration contains var_spec children
if (node.type === 'var_declaration') {
for (let i = 0; i < node.namedChildCount; i++) {
const spec = node.namedChild(i);
if (spec?.type === 'var_spec') extractGoVarDeclaration(spec, env);
}
return;
}
// var_spec: name type [= value]
const nameNode = node.childForFieldName('name');
const typeNode = node.childForFieldName('type');
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
/** Go: x := Foo{...} — infer type from composite literal (handles multi-assignment) */
const extractGoShortVarDeclaration = (node: SyntaxNode, env: Map<string, string>): void => {
const left = node.childForFieldName('left');
const right = node.childForFieldName('right');
if (!left || !right) return;
// Collect LHS names and RHS values (may be expression_lists for multi-assignment)
const lhsNodes: SyntaxNode[] = [];
const rhsNodes: SyntaxNode[] = [];
if (left.type === 'expression_list') {
for (let i = 0; i < left.namedChildCount; i++) {
const c = left.namedChild(i);
if (c) lhsNodes.push(c);
}
} else {
lhsNodes.push(left);
}
if (right.type === 'expression_list') {
for (let i = 0; i < right.namedChildCount; i++) {
const c = right.namedChild(i);
if (c) rhsNodes.push(c);
}
} else {
rhsNodes.push(right);
}
// Pair each LHS name with its corresponding RHS value
const count = Math.min(lhsNodes.length, rhsNodes.length);
for (let i = 0; i < count; i++) {
const valueNode = rhsNodes[i];
if (valueNode.type !== 'composite_literal') continue;
const typeNode = valueNode.childForFieldName('type');
if (!typeNode) continue;
const typeName = extractSimpleTypeName(typeNode);
if (!typeName) continue;
const varName = extractVarName(lhsNodes[i]);
if (varName) env.set(varName, typeName);
}
};
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
if (node.type === 'var_declaration' || node.type === 'var_spec') {
extractGoVarDeclaration(node, env);
} else if (node.type === 'short_var_declaration') {
extractGoShortVarDeclaration(node, env);
}
};
/** Go: parameter → name type */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'parameter') {
nameNode = node.childForFieldName('name');
typeNode = node.childForFieldName('type');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -1,35 +0,0 @@
/**
* Per-language type extraction configurations.
* Assembled here into a dispatch map keyed by SupportedLanguages.
*/
import { SupportedLanguages } from '../../../config/supported-languages.js';
import type { LanguageTypeConfig } from './types.js';
import { typeConfig as typescriptConfig } from './typescript.js';
import { javaTypeConfig, kotlinTypeConfig } from './jvm.js';
import { typeConfig as csharpConfig } from './csharp.js';
import { typeConfig as goConfig } from './go.js';
import { typeConfig as rustConfig } from './rust.js';
import { typeConfig as pythonConfig } from './python.js';
import { typeConfig as swiftConfig } from './swift.js';
import { typeConfig as cCppConfig } from './c-cpp.js';
import { typeConfig as phpConfig } from './php.js';
export const typeConfigs = {
[SupportedLanguages.JavaScript]: typescriptConfig,
[SupportedLanguages.TypeScript]: typescriptConfig,
[SupportedLanguages.Java]: javaTypeConfig,
[SupportedLanguages.Kotlin]: kotlinTypeConfig,
[SupportedLanguages.CSharp]: csharpConfig,
[SupportedLanguages.Go]: goConfig,
[SupportedLanguages.Rust]: rustConfig,
[SupportedLanguages.Python]: pythonConfig,
[SupportedLanguages.Swift]: swiftConfig,
[SupportedLanguages.C]: cCppConfig,
[SupportedLanguages.CPlusPlus]: cCppConfig,
[SupportedLanguages.PHP]: phpConfig,
} satisfies Record<SupportedLanguages, LanguageTypeConfig>;
export type { LanguageTypeConfig, TypeBindingExtractor, ParameterExtractor } from './types.js';
export { TYPED_PARAMETER_TYPES, extractSimpleTypeName, extractVarName, findChildByType } from './shared.js';
@@ -1,122 +0,0 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, findChildByType } from './shared.js';
// ── Java ──────────────────────────────────────────────────────────────────
const JAVA_DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'local_variable_declaration',
'field_declaration',
]);
/** Java: Type x = ...; Type x; */
const extractJavaDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
const typeNode = node.childForFieldName('type');
if (!typeNode) return;
const typeName = extractSimpleTypeName(typeNode);
if (!typeName) return;
// Find variable_declarator children
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type !== 'variable_declarator') continue;
const nameNode = child.childForFieldName('name');
if (nameNode) {
const varName = extractVarName(nameNode);
if (varName) env.set(varName, typeName);
}
}
};
/** Java: formal_parameter → type name */
const extractJavaParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'formal_parameter') {
typeNode = node.childForFieldName('type');
nameNode = node.childForFieldName('name');
} else {
// Generic fallback
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const javaTypeConfig: LanguageTypeConfig = {
declarationNodeTypes: JAVA_DECLARATION_NODE_TYPES,
extractDeclaration: extractJavaDeclaration,
extractParameter: extractJavaParameter,
};
// ── Kotlin ────────────────────────────────────────────────────────────────
const KOTLIN_DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'property_declaration',
'variable_declaration',
]);
/** Kotlin: val x: Foo = ... */
const extractKotlinDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
if (node.type === 'property_declaration') {
// Kotlin property_declaration: name/type are inside a variable_declaration child
const varDecl = findChildByType(node, 'variable_declaration');
if (varDecl) {
const nameNode = findChildByType(varDecl, 'simple_identifier');
const typeNode = findChildByType(varDecl, 'user_type');
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
return;
}
// Fallback: try direct fields
const nameNode = node.childForFieldName('name')
?? findChildByType(node, 'simple_identifier');
const typeNode = node.childForFieldName('type')
?? findChildByType(node, 'user_type');
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
} else if (node.type === 'variable_declaration') {
// variable_declaration directly inside functions
const nameNode = findChildByType(node, 'simple_identifier');
const typeNode = findChildByType(node, 'user_type');
if (nameNode && typeNode) {
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
}
}
};
/** Kotlin: formal_parameter → type name */
const extractKotlinParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'formal_parameter') {
typeNode = node.childForFieldName('type');
nameNode = node.childForFieldName('name');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const kotlinTypeConfig: LanguageTypeConfig = {
declarationNodeTypes: KOTLIN_DECLARATION_NODE_TYPES,
extractDeclaration: extractKotlinDeclaration,
extractParameter: extractKotlinParameter,
};
@@ -1,36 +0,0 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
// PHP has no local variable type annotations; only params carry types
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set<string>();
/** PHP: no typed local variable declarations */
const extractDeclaration: TypeBindingExtractor = (_node: SyntaxNode, _env: Map<string, string>): void => {
// PHP has no local variable type annotations
};
/** PHP: simple_parameter → type $name */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'simple_parameter') {
typeNode = node.childForFieldName('type');
nameNode = node.childForFieldName('name');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -1,44 +0,0 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'assignment',
]);
/** Python: x: Foo = ... (PEP 484 annotations) */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
// Python annotated assignment: left : type = value
// tree-sitter represents this differently based on grammar version
const left = node.childForFieldName('left');
const typeNode = node.childForFieldName('type');
if (!left || !typeNode) return;
const varName = extractVarName(left);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
/** Python: parameter with type annotation */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'parameter') {
nameNode = node.childForFieldName('name');
typeNode = node.childForFieldName('type');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -1,42 +0,0 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'let_declaration',
]);
/** Rust: let x: Foo = ... */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
const pattern = node.childForFieldName('pattern');
const typeNode = node.childForFieldName('type');
if (!pattern || !typeNode) return;
const varName = extractVarName(pattern);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
/** Rust: parameter → pattern: type */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'parameter') {
nameNode = node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -1,103 +0,0 @@
import type { SyntaxNode } from '../utils.js';
/**
* Extract the simple type name from a type AST node.
* Handles generic types (e.g., List<User> → List), qualified names
* (e.g., models.User → User), and nullable types (e.g., User? → User).
* Returns undefined for complex types (unions, intersections, function types).
*/
export const extractSimpleTypeName = (typeNode: SyntaxNode): string | undefined => {
// Direct type identifier
if (typeNode.type === 'type_identifier' || typeNode.type === 'identifier'
|| typeNode.type === 'simple_identifier') {
return typeNode.text;
}
// Qualified/scoped names: take the last segment (e.g., models.User → User)
if (typeNode.type === 'scoped_identifier' || typeNode.type === 'qualified_identifier'
|| typeNode.type === 'scoped_type_identifier' || typeNode.type === 'qualified_name'
|| typeNode.type === 'qualified_type'
|| typeNode.type === 'member_expression' || typeNode.type === 'attribute') {
const last = typeNode.lastNamedChild;
if (last && (last.type === 'type_identifier' || last.type === 'identifier'
|| last.type === 'simple_identifier' || last.type === 'name')) {
return last.text;
}
}
// Generic types: extract the base type (e.g., List<User> → List)
if (typeNode.type === 'generic_type' || typeNode.type === 'parameterized_type') {
const base = typeNode.childForFieldName('name')
?? typeNode.childForFieldName('type')
?? typeNode.firstNamedChild;
if (base) return extractSimpleTypeName(base);
}
// Nullable types (Kotlin User?, C# User?)
if (typeNode.type === 'nullable_type') {
const inner = typeNode.firstNamedChild;
if (inner) return extractSimpleTypeName(inner);
}
// Type annotations that wrap the actual type (TS/Python: `: Foo`, Kotlin: user_type)
if (typeNode.type === 'type_annotation' || typeNode.type === 'type'
|| typeNode.type === 'user_type') {
const inner = typeNode.firstNamedChild;
if (inner) return extractSimpleTypeName(inner);
}
// Pointer/reference types (C++, Rust): User*, &User, &mut User
if (typeNode.type === 'pointer_type' || typeNode.type === 'reference_type') {
const inner = typeNode.firstNamedChild;
if (inner) return extractSimpleTypeName(inner);
}
// PHP named_type / optional_type
if (typeNode.type === 'named_type' || typeNode.type === 'optional_type') {
const inner = typeNode.childForFieldName('name') ?? typeNode.firstNamedChild;
if (inner) return extractSimpleTypeName(inner);
}
// Name node (PHP)
if (typeNode.type === 'name') {
return typeNode.text;
}
return undefined;
};
/**
* Extract variable name from a declarator or pattern node.
* Returns the simple identifier text, or undefined for destructuring/complex patterns.
*/
export const extractVarName = (node: SyntaxNode): string | undefined => {
if (node.type === 'identifier' || node.type === 'simple_identifier'
|| node.type === 'variable_name' || node.type === 'name') {
return node.text;
}
// variable_declarator (Java/C#): has a 'name' field
if (node.type === 'variable_declarator') {
const nameChild = node.childForFieldName('name');
if (nameChild) return extractVarName(nameChild);
}
return undefined;
};
/** Node types for function/method parameters with type annotations */
export const TYPED_PARAMETER_TYPES = new Set([
'required_parameter', // TS: (x: Foo)
'optional_parameter', // TS: (x?: Foo)
'formal_parameter', // Java/Kotlin
'parameter', // C#/Rust/Go/Python/Swift
'parameter_declaration', // C/C++ void f(Type name)
'simple_parameter', // PHP function(Foo $x)
]);
/** Find the first named child with the given node type */
export const findChildByType = (node: SyntaxNode, type: string): SyntaxNode | null => {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === type) return child;
}
return null;
};
@@ -1,46 +0,0 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, findChildByType } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'property_declaration',
]);
/** Swift: let x: Foo = ... */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
// Swift property_declaration has pattern and type_annotation
const pattern = node.childForFieldName('pattern')
?? findChildByType(node, 'pattern');
const typeAnnotation = node.childForFieldName('type')
?? findChildByType(node, 'type_annotation');
if (!pattern || !typeAnnotation) return;
const varName = extractVarName(pattern) ?? pattern.text;
const typeName = extractSimpleTypeName(typeAnnotation);
if (varName && typeName) env.set(varName, typeName);
};
/** Swift: parameter → name: type */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'parameter') {
nameNode = node.childForFieldName('name')
?? node.childForFieldName('internal_name');
typeNode = node.childForFieldName('type');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -1,17 +0,0 @@
import type { SyntaxNode } from '../utils.js';
/** Extracts type bindings from a declaration node into the env map */
export type TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>) => void;
/** Extracts type bindings from a parameter node into the env map */
export type ParameterExtractor = (node: SyntaxNode, env: Map<string, string>) => void;
/** Per-language type extraction configuration */
export interface LanguageTypeConfig {
/** Node types that represent typed declarations for this language */
declarationNodeTypes: ReadonlySet<string>;
/** Extract a (varName → typeName) binding from a declaration node */
extractDeclaration: TypeBindingExtractor;
/** Extract a (varName → typeName) binding from a parameter node */
extractParameter: ParameterExtractor;
}
@@ -1,48 +0,0 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'lexical_declaration',
'variable_declaration',
]);
/** TypeScript: const x: Foo = ..., let x: Foo */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
for (let i = 0; i < node.namedChildCount; i++) {
const declarator = node.namedChild(i);
if (declarator?.type !== 'variable_declarator') continue;
const nameNode = declarator.childForFieldName('name');
const typeAnnotation = declarator.childForFieldName('type');
if (!nameNode || !typeAnnotation) continue;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeAnnotation);
if (varName && typeName) env.set(varName, typeName);
}
};
/** TypeScript: required_parameter / optional_parameter → name: type */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'required_parameter' || node.type === 'optional_parameter') {
nameNode = node.childForFieldName('pattern') ?? node.childForFieldName('name');
typeNode = node.childForFieldName('type');
} else {
// Generic fallback
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
+4 -751
View File
@@ -1,415 +1,4 @@
import type Parser from 'tree-sitter';
import { SupportedLanguages } from '../../config/supported-languages.js';
import { generateId } from '../../lib/utils.js';
/** Tree-sitter AST node. Re-exported for use across ingestion modules. */
export type SyntaxNode = Parser.SyntaxNode;
/**
* Ordered list of definition capture keys for tree-sitter query matches.
* Used to extract the definition node from a capture map.
*/
export const DEFINITION_CAPTURE_KEYS = [
'definition.function',
'definition.class',
'definition.interface',
'definition.method',
'definition.struct',
'definition.enum',
'definition.namespace',
'definition.module',
'definition.trait',
'definition.impl',
'definition.type',
'definition.const',
'definition.static',
'definition.typedef',
'definition.macro',
'definition.union',
'definition.property',
'definition.record',
'definition.delegate',
'definition.annotation',
'definition.constructor',
'definition.template',
] as const;
/** Extract the definition node from a tree-sitter query capture map. */
export const getDefinitionNodeFromCaptures = (captureMap: Record<string, any>): any | null => {
for (const key of DEFINITION_CAPTURE_KEYS) {
if (captureMap[key]) return captureMap[key];
}
return null;
};
/**
* Node types that represent function/method definitions across languages.
* Used to find the enclosing function for a call site.
*/
export const FUNCTION_NODE_TYPES = new Set([
// TypeScript/JavaScript
'function_declaration',
'arrow_function',
'function_expression',
'method_definition',
'generator_function_declaration',
// Python
'function_definition',
// Common async variants
'async_function_declaration',
'async_arrow_function',
// Java
'method_declaration',
'constructor_declaration',
// C/C++
// 'function_definition' already included above
// Go
// 'method_declaration' already included from Java
// C#
'local_function_statement',
// Rust
'function_item',
'impl_item', // Methods inside impl blocks
// PHP
'anonymous_function',
// Kotlin
'lambda_literal',
// Swift
'init_declaration',
'deinit_declaration',
]);
/**
* Node types for standard function declarations that need C/C++ declarator handling.
* Used by extractFunctionName to determine how to extract the function name.
*/
export const FUNCTION_DECLARATION_TYPES = new Set([
'function_declaration',
'function_definition',
'async_function_declaration',
'generator_function_declaration',
'function_item',
]);
/**
* Built-in function/method names that should not be tracked as call targets.
* Covers JS/TS, Python, Kotlin, C/C++, PHP, Swift standard library functions.
*/
export const BUILT_IN_NAMES = new Set([
// JavaScript/TypeScript
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export', 'fetch', 'Response', 'Request',
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'append', 'extend', 'update',
// NOTE: 'open', 'read', 'write', 'close' removed — these are real C POSIX syscalls
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
// Kotlin stdlib
'println', 'print', 'readLine', 'require', 'requireNotNull', 'check', 'assert', 'lazy', 'error',
'listOf', 'mapOf', 'setOf', 'mutableListOf', 'mutableMapOf', 'mutableSetOf',
'arrayOf', 'sequenceOf', 'also', 'apply', 'run', 'with', 'takeIf', 'takeUnless',
'TODO', 'buildString', 'buildList', 'buildMap', 'buildSet',
'repeat', 'synchronized',
// Kotlin coroutine builders & scope functions
'launch', 'async', 'runBlocking', 'withContext', 'coroutineScope',
'supervisorScope', 'delay',
// Kotlin Flow operators
'flow', 'flowOf', 'collect', 'emit', 'onEach', 'catch',
'buffer', 'conflate', 'distinctUntilChanged',
'flatMapLatest', 'flatMapMerge', 'combine',
'stateIn', 'shareIn', 'launchIn',
// Kotlin infix stdlib functions
'to', 'until', 'downTo', 'step',
// C/C++ standard library
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
'scanf', 'fscanf', 'sscanf',
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
'sizeof', 'offsetof', 'typeof',
'assert', 'abort', 'exit', '_exit',
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
// Linux kernel common macros/helpers (not real call targets)
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
'min', 'max', 'clamp', 'abs', 'swap',
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
'GFP_KERNEL', 'GFP_ATOMIC',
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
'mutex_lock', 'mutex_unlock', 'mutex_init',
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
'get', 'put',
// C# / .NET built-ins
'Console', 'WriteLine', 'ReadLine', 'Write',
'Task', 'Run', 'Wait', 'WhenAll', 'WhenAny', 'FromResult', 'Delay', 'ContinueWith',
'ConfigureAwait', 'GetAwaiter', 'GetResult',
'ToString', 'GetType', 'Equals', 'GetHashCode', 'ReferenceEquals',
'Add', 'Remove', 'Contains', 'Clear', 'Count', 'Any', 'All',
'Where', 'Select', 'SelectMany', 'OrderBy', 'OrderByDescending', 'GroupBy',
'First', 'FirstOrDefault', 'Single', 'SingleOrDefault', 'Last', 'LastOrDefault',
'ToList', 'ToArray', 'ToDictionary', 'AsEnumerable', 'AsQueryable',
'Aggregate', 'Sum', 'Average', 'Min', 'Max', 'Distinct', 'Skip', 'Take',
'String', 'Format', 'IsNullOrEmpty', 'IsNullOrWhiteSpace', 'Concat', 'Join',
'Trim', 'TrimStart', 'TrimEnd', 'Split', 'Replace', 'StartsWith', 'EndsWith',
'Convert', 'ToInt32', 'ToDouble', 'ToBoolean', 'ToByte',
'Math', 'Abs', 'Ceiling', 'Floor', 'Round', 'Pow', 'Sqrt',
'Dispose', 'Close',
'TryParse', 'Parse',
'AddRange', 'RemoveAt', 'RemoveAll', 'FindAll', 'Exists', 'TrueForAll',
'ContainsKey', 'TryGetValue', 'AddOrUpdate',
'Throw', 'ThrowIfNull',
// PHP built-ins
'echo', 'isset', 'empty', 'unset', 'list', 'array', 'compact', 'extract',
'count', 'strlen', 'strpos', 'strrpos', 'substr', 'strtolower', 'strtoupper', 'trim',
'ltrim', 'rtrim', 'str_replace', 'str_contains', 'str_starts_with', 'str_ends_with',
'sprintf', 'vsprintf', 'printf', 'number_format',
'array_map', 'array_filter', 'array_reduce', 'array_push', 'array_pop', 'array_shift',
'array_unshift', 'array_slice', 'array_splice', 'array_merge', 'array_keys', 'array_values',
'array_key_exists', 'in_array', 'array_search', 'array_unique', 'usort', 'rsort',
'json_encode', 'json_decode', 'serialize', 'unserialize',
'intval', 'floatval', 'strval', 'boolval', 'is_null', 'is_string', 'is_int', 'is_array',
'is_object', 'is_numeric', 'is_bool', 'is_float',
'var_dump', 'print_r', 'var_export',
'date', 'time', 'strtotime', 'mktime', 'microtime',
'file_exists', 'file_get_contents', 'file_put_contents', 'is_file', 'is_dir',
'preg_match', 'preg_match_all', 'preg_replace', 'preg_split',
'header', 'session_start', 'session_destroy', 'ob_start', 'ob_end_clean', 'ob_get_clean',
'dd', 'dump',
// Swift/iOS built-ins and standard library
'print', 'debugPrint', 'dump', 'fatalError', 'precondition', 'preconditionFailure',
'assert', 'assertionFailure', 'NSLog',
'abs', 'min', 'max', 'zip', 'stride', 'sequence', 'repeatElement',
'swap', 'withUnsafePointer', 'withUnsafeMutablePointer', 'withUnsafeBytes',
'autoreleasepool', 'unsafeBitCast', 'unsafeDowncast', 'numericCast',
'type', 'MemoryLayout',
// Swift collection/string methods (common noise)
'map', 'flatMap', 'compactMap', 'filter', 'reduce', 'forEach', 'contains',
'first', 'last', 'prefix', 'suffix', 'dropFirst', 'dropLast',
'sorted', 'reversed', 'enumerated', 'joined', 'split',
'append', 'insert', 'remove', 'removeAll', 'removeFirst', 'removeLast',
'isEmpty', 'count', 'index', 'startIndex', 'endIndex',
// UIKit/Foundation common methods (noise in call graph)
'addSubview', 'removeFromSuperview', 'layoutSubviews', 'setNeedsLayout',
'layoutIfNeeded', 'setNeedsDisplay', 'invalidateIntrinsicContentSize',
'addTarget', 'removeTarget', 'addGestureRecognizer',
'addConstraint', 'addConstraints', 'removeConstraint', 'removeConstraints',
'NSLocalizedString', 'Bundle',
'reloadData', 'reloadSections', 'reloadRows', 'performBatchUpdates',
'register', 'dequeueReusableCell', 'dequeueReusableSupplementaryView',
'beginUpdates', 'endUpdates', 'insertRows', 'deleteRows', 'insertSections', 'deleteSections',
'present', 'dismiss', 'pushViewController', 'popViewController', 'popToRootViewController',
'performSegue', 'prepare',
// GCD / async
'DispatchQueue', 'async', 'sync', 'asyncAfter',
'Task', 'withCheckedContinuation', 'withCheckedThrowingContinuation',
// Combine
'sink', 'store', 'assign', 'receive', 'subscribe',
// Notification / KVO
'addObserver', 'removeObserver', 'post', 'NotificationCenter',
// Rust standard library (common noise in call graphs)
'unwrap', 'expect', 'unwrap_or', 'unwrap_or_else', 'unwrap_or_default',
'ok', 'err', 'is_ok', 'is_err', 'map', 'map_err', 'and_then', 'or_else',
'clone', 'to_string', 'to_owned', 'into', 'from', 'as_ref', 'as_mut',
'iter', 'into_iter', 'collect', 'map', 'filter', 'fold', 'for_each',
'len', 'is_empty', 'push', 'pop', 'insert', 'remove', 'contains',
'format', 'write', 'writeln', 'panic', 'unreachable', 'todo', 'unimplemented',
'vec', 'println', 'eprintln', 'dbg',
'lock', 'read', 'write', 'try_lock',
'spawn', 'join', 'sleep',
'Some', 'None', 'Ok', 'Err',
]);
/** Check if a name is a built-in function or common noise that should be filtered out */
export const isBuiltInOrNoise = (name: string): boolean => BUILT_IN_NAMES.has(name);
/** AST node types that represent a class-like container (for HAS_METHOD edge extraction) */
export const CLASS_CONTAINER_TYPES = new Set([
'class_declaration', 'abstract_class_declaration',
'interface_declaration', 'struct_declaration', 'record_declaration',
'class_specifier', 'struct_specifier',
'impl_item', 'trait_item',
'class_definition',
'trait_declaration',
'protocol_declaration',
]);
export const CONTAINER_TYPE_TO_LABEL: Record<string, string> = {
class_declaration: 'Class',
abstract_class_declaration: 'Class',
interface_declaration: 'Interface',
struct_declaration: 'Struct',
struct_specifier: 'Struct',
class_specifier: 'Class',
class_definition: 'Class',
impl_item: 'Impl',
trait_item: 'Trait',
trait_declaration: 'Trait',
record_declaration: 'Record',
protocol_declaration: 'Interface',
};
/** Walk up AST to find enclosing class/struct/interface/impl, return its generateId or null.
* For Go method_declaration nodes, extracts receiver type (e.g. `func (u *User) Save()` → User struct). */
export const findEnclosingClassId = (node: any, filePath: string): string | null => {
let current = node.parent;
while (current) {
// Go: method_declaration has a receiver parameter with the struct type
if (current.type === 'method_declaration') {
const receiver = current.childForFieldName?.('receiver');
if (receiver) {
// receiver is a parameter_list: (u *User) or (u User)
const paramDecl = receiver.namedChildren?.find?.((c: any) => c.type === 'parameter_declaration');
if (paramDecl) {
const typeNode = paramDecl.childForFieldName?.('type');
if (typeNode) {
// Unwrap pointer_type (*User → User)
const inner = typeNode.type === 'pointer_type' ? typeNode.firstNamedChild : typeNode;
if (inner && (inner.type === 'type_identifier' || inner.type === 'identifier')) {
return generateId('Struct', `${filePath}:${inner.text}`);
}
}
}
}
}
if (CLASS_CONTAINER_TYPES.has(current.type)) {
// Rust impl_item: for `impl Trait for Struct {}`, pick the type after `for`
if (current.type === 'impl_item') {
const children = current.children ?? [];
const forIdx = children.findIndex((c: any) => c.text === 'for');
if (forIdx !== -1) {
const nameNode = children.slice(forIdx + 1).find((c: any) =>
c.type === 'type_identifier' || c.type === 'identifier'
);
if (nameNode) {
return generateId('Impl', `${filePath}:${nameNode.text}`);
}
}
// Fall through: plain `impl Struct {}` — use first type_identifier below
}
const nameNode = current.childForFieldName?.('name')
?? current.children?.find((c: any) =>
c.type === 'type_identifier' || c.type === 'identifier' || c.type === 'name'
);
if (nameNode) {
const label = CONTAINER_TYPE_TO_LABEL[current.type] || 'Class';
return generateId(label, `${filePath}:${nameNode.text}`);
}
}
current = current.parent;
}
return null;
};
/**
* Extract function name and label from a function_definition or similar AST node.
* Handles C/C++ qualified_identifier (ClassName::MethodName) and other language patterns.
*/
export const extractFunctionName = (node: any): { funcName: string | null; label: string } => {
let funcName: string | null = null;
let label = 'Function';
// Swift init/deinit
if (node.type === 'init_declaration' || node.type === 'deinit_declaration') {
return {
funcName: node.type === 'init_declaration' ? 'init' : 'deinit',
label: 'Constructor',
};
}
if (FUNCTION_DECLARATION_TYPES.has(node.type)) {
// C/C++: function_definition -> [pointer_declarator ->] function_declarator -> qualified_identifier/identifier
// Unwrap pointer_declarator / reference_declarator wrappers to reach function_declarator
let declarator = node.childForFieldName?.('declarator') ||
node.children?.find((c: any) => c.type === 'function_declarator');
while (declarator && (declarator.type === 'pointer_declarator' || declarator.type === 'reference_declarator')) {
declarator = declarator.childForFieldName?.('declarator') ||
declarator.children?.find((c: any) =>
c.type === 'function_declarator' || c.type === 'pointer_declarator' || c.type === 'reference_declarator');
}
if (declarator) {
const innerDeclarator = declarator.childForFieldName?.('declarator') ||
declarator.children?.find((c: any) =>
c.type === 'qualified_identifier' || c.type === 'identifier' || c.type === 'parenthesized_declarator');
if (innerDeclarator?.type === 'qualified_identifier') {
const nameNode = innerDeclarator.childForFieldName?.('name') ||
innerDeclarator.children?.find((c: any) => c.type === 'identifier');
if (nameNode?.text) {
funcName = nameNode.text;
label = 'Method';
}
} else if (innerDeclarator?.type === 'identifier') {
funcName = innerDeclarator.text;
} else if (innerDeclarator?.type === 'parenthesized_declarator') {
const nestedId = innerDeclarator.children?.find((c: any) =>
c.type === 'qualified_identifier' || c.type === 'identifier');
if (nestedId?.type === 'qualified_identifier') {
const nameNode = nestedId.childForFieldName?.('name') ||
nestedId.children?.find((c: any) => c.type === 'identifier');
if (nameNode?.text) {
funcName = nameNode.text;
label = 'Method';
}
} else if (nestedId?.type === 'identifier') {
funcName = nestedId.text;
}
}
}
// Fallback for other languages (Kotlin uses simple_identifier, Swift uses simple_identifier)
if (!funcName) {
const nameNode = node.childForFieldName?.('name') ||
node.children?.find((c: any) => c.type === 'identifier' || c.type === 'property_identifier' || c.type === 'simple_identifier');
funcName = nameNode?.text;
}
} else if (node.type === 'impl_item') {
const funcItem = node.children?.find((c: any) => c.type === 'function_item');
if (funcItem) {
const nameNode = funcItem.childForFieldName?.('name') ||
funcItem.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
}
} else if (node.type === 'method_definition') {
const nameNode = node.childForFieldName?.('name') ||
node.children?.find((c: any) => c.type === 'property_identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (node.type === 'method_declaration' || node.type === 'constructor_declaration') {
const nameNode = node.childForFieldName?.('name') ||
node.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (node.type === 'arrow_function' || node.type === 'function_expression') {
const parent = node.parent;
if (parent?.type === 'variable_declarator') {
const nameNode = parent.childForFieldName?.('name') ||
parent.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
}
}
return { funcName, label };
};
/**
* Yield control to the event loop so spinners/progress can render.
@@ -448,13 +37,11 @@ export const getLanguageFromFilename = (filename: string): SupportedLanguages |
if (filename.endsWith('.py')) return SupportedLanguages.Python;
// Java
if (filename.endsWith('.java')) return SupportedLanguages.Java;
// C source files
if (filename.endsWith('.c')) return SupportedLanguages.C;
// C++ (all common extensions, including .h)
// .h is parsed as C++ because tree-sitter-cpp is a strict superset of C, so pure-C
// headers parse correctly, and C++ headers (classes, templates) are handled properly.
// C (source and headers)
if (filename.endsWith('.c') || filename.endsWith('.h')) return SupportedLanguages.C;
// C++ (all common extensions)
if (filename.endsWith('.cpp') || filename.endsWith('.cc') || filename.endsWith('.cxx') ||
filename.endsWith('.h') || filename.endsWith('.hpp') || filename.endsWith('.hxx') || filename.endsWith('.hh')) return SupportedLanguages.CPlusPlus;
filename.endsWith('.hpp') || filename.endsWith('.hxx') || filename.endsWith('.hh')) return SupportedLanguages.CPlusPlus;
// C#
if (filename.endsWith('.cs')) return SupportedLanguages.CSharp;
// Go
@@ -473,337 +60,3 @@ export const getLanguageFromFilename = (filename: string): SupportedLanguages |
return null;
};
export interface MethodSignature {
parameterCount: number | undefined;
returnType: string | undefined;
}
const CALL_ARGUMENT_LIST_TYPES = new Set([
'arguments',
'argument_list',
'value_arguments',
]);
/**
* Extract parameter count and return type text from an AST method/function node.
* Works across languages by looking for common AST patterns.
*/
export const extractMethodSignature = (node: SyntaxNode | null | undefined): MethodSignature => {
let parameterCount: number | undefined = 0;
let returnType: string | undefined;
let isVariadic = false;
if (!node) return { parameterCount, returnType };
const paramListTypes = new Set([
'formal_parameters', 'parameters', 'parameter_list',
'function_parameters', 'method_parameters', 'function_value_parameters',
]);
// Node types that indicate variadic/rest parameters
const VARIADIC_PARAM_TYPES = new Set([
'variadic_parameter_declaration', // Go: ...string
'variadic_parameter', // Rust: extern "C" fn(...)
'spread_parameter', // Java: Object... args
'list_splat_pattern', // Python: *args
'dictionary_splat_pattern', // Python: **kwargs
]);
const findParameterList = (current: SyntaxNode): SyntaxNode | null => {
for (const child of current.children) {
if (paramListTypes.has(child.type)) return child;
}
for (const child of current.children) {
const nested = findParameterList(child);
if (nested) return nested;
}
return null;
};
const parameterList = (
paramListTypes.has(node.type) ? node // node itself IS the parameter list (e.g. C# primary constructors)
: node.childForFieldName?.('parameters')
?? findParameterList(node)
);
if (parameterList && paramListTypes.has(parameterList.type)) {
for (const param of parameterList.namedChildren) {
if (param.type === 'comment') continue;
if (param.text === 'self' || param.text === '&self' || param.text === '&mut self' ||
param.type === 'self_parameter') {
continue;
}
// Check for variadic parameter types
if (VARIADIC_PARAM_TYPES.has(param.type)) {
isVariadic = true;
continue;
}
// TypeScript/JavaScript: rest parameter — required_parameter containing rest_pattern
if (param.type === 'required_parameter' || param.type === 'optional_parameter') {
for (const child of param.children) {
if (child.type === 'rest_pattern') {
isVariadic = true;
break;
}
}
if (isVariadic) continue;
}
// Kotlin: vararg modifier on a regular parameter
if (param.type === 'parameter' || param.type === 'formal_parameter') {
const prev = param.previousSibling;
if (prev?.type === 'parameter_modifiers' && prev.text.includes('vararg')) {
isVariadic = true;
}
}
parameterCount++;
}
// C/C++: bare `...` token in parameter list (not a named child — check all children)
if (!isVariadic) {
for (const child of parameterList.children) {
if (!child.isNamed && child.text === '...') {
isVariadic = true;
break;
}
}
}
}
// Return type extraction — language-specific field names
// Go: 'result' field is either a type_identifier or parameter_list (multi-return)
const goResult = node.childForFieldName?.('result');
if (goResult) {
returnType = goResult.type === 'parameter_list'
? goResult.text // multi-return: "(string, error)"
: goResult.text; // single return: "int"
}
// Rust: 'return_type' field — the value IS the type node (e.g. primitive_type, type_identifier).
// Skip if the node is a type_annotation (TS/Python), which is handled by the generic loop below.
if (!returnType) {
const rustReturn = node.childForFieldName?.('return_type');
if (rustReturn && rustReturn.type !== 'type_annotation') {
returnType = rustReturn.text;
}
}
// C/C++: 'type' field on function_definition
if (!returnType) {
const cppType = node.childForFieldName?.('type');
if (cppType && cppType.text !== 'void') {
returnType = cppType.text;
}
}
// TS/Rust/Python/C#/Kotlin: type_annotation or return_type child
if (!returnType) {
for (const child of node.children) {
if (child.type === 'type_annotation' || child.type === 'return_type') {
const typeNode = child.children.find((c) => c.isNamed);
if (typeNode) returnType = typeNode.text;
}
}
}
if (isVariadic) parameterCount = undefined;
return { parameterCount, returnType };
};
/**
* Count direct arguments for a call expression across common tree-sitter grammars.
* Returns undefined when the argument container cannot be located cheaply.
*/
export const countCallArguments = (callNode: SyntaxNode | null | undefined): number | undefined => {
if (!callNode) return undefined;
// Direct field or direct child (most languages)
let argsNode: SyntaxNode | null | undefined = callNode.childForFieldName('arguments')
?? callNode.children.find((child) => CALL_ARGUMENT_LIST_TYPES.has(child.type));
// Kotlin/Swift: call_expression → call_suffix → value_arguments
// Search one level deeper for languages that wrap arguments in a suffix node
if (!argsNode) {
for (const child of callNode.children) {
if (!child.isNamed) continue;
const nested = child.children.find((gc) => CALL_ARGUMENT_LIST_TYPES.has(gc.type));
if (nested) { argsNode = nested; break; }
}
}
if (!argsNode) return undefined;
let count = 0;
for (const child of argsNode.children) {
if (!child.isNamed) continue;
if (child.type === 'comment') continue;
count++;
}
return count;
};
// ── Call-form discrimination (Phase 1, Step D) ─────────────────────────
/**
* AST node types that indicate a member-access wrapper around the callee name.
* When nameNode.parent.type is one of these, the call is a member call.
*/
const MEMBER_ACCESS_NODE_TYPES = new Set([
'member_expression', // TS/JS: obj.method()
'attribute', // Python: obj.method()
'member_access_expression', // C#: obj.Method()
'field_expression', // Rust/C++: obj.method() / ptr->method()
'selector_expression', // Go: obj.Method()
'navigation_suffix', // Kotlin/Swift: obj.method() — nameNode sits inside navigation_suffix
]);
/**
* Call node types that are inherently constructor invocations.
* Only includes patterns that the tree-sitter queries already capture as @call.
*/
const CONSTRUCTOR_CALL_NODE_TYPES = new Set([
'constructor_invocation', // Kotlin: Foo()
'new_expression', // TS/JS/C++: new Foo()
'object_creation_expression', // Java/C#/PHP: new Foo()
'implicit_object_creation_expression', // C# 9: User u = new(...)
'composite_literal', // Go: User{...}
'struct_expression', // Rust: User { ... }
]);
/**
* AST node types for scoped/qualified calls (e.g., Foo::new() in Rust, Foo::bar() in C++).
*/
const SCOPED_CALL_NODE_TYPES = new Set([
'scoped_identifier', // Rust: Foo::new()
'qualified_identifier', // C++: ns::func()
]);
type CallForm = 'free' | 'member' | 'constructor';
/**
* Infer whether a captured call site is a free call, member call, or constructor.
* Returns undefined if the form cannot be determined.
*
* Works by inspecting the AST structure between callNode (@call) and nameNode (@call.name).
* No tree-sitter query changes needed — the distinction is in the node types.
*/
export const inferCallForm = (
callNode: SyntaxNode,
nameNode: SyntaxNode,
): CallForm | undefined => {
// 1. Constructor: callNode itself is a constructor invocation (Kotlin)
if (CONSTRUCTOR_CALL_NODE_TYPES.has(callNode.type)) {
return 'constructor';
}
// 2. Member call: nameNode's parent is a member-access wrapper
const nameParent = nameNode.parent;
if (nameParent && MEMBER_ACCESS_NODE_TYPES.has(nameParent.type)) {
return 'member';
}
// 3. PHP: the callNode itself distinguishes member vs free calls
if (callNode.type === 'member_call_expression' || callNode.type === 'nullsafe_member_call_expression') {
return 'member';
}
if (callNode.type === 'scoped_call_expression') {
return 'member'; // static call Foo::bar()
}
// 4. Java method_invocation: member if it has an 'object' field
if (callNode.type === 'method_invocation' && callNode.childForFieldName('object')) {
return 'member';
}
// 5. Scoped calls (Rust Foo::new(), C++ ns::func()): treat as free
// The receiver is a type, not an instance — handled differently in Phase 3
if (nameParent && SCOPED_CALL_NODE_TYPES.has(nameParent.type)) {
return 'free';
}
// 6. Default: if nameNode is a direct child of callNode, it's a free call
if (nameNode.parent === callNode || nameParent?.parent === callNode) {
return 'free';
}
return undefined;
};
/**
* Extract the receiver identifier for member calls.
* Only captures simple identifiers — returns undefined for complex expressions
* like getUser().save() or arr[0].method().
*/
const SIMPLE_RECEIVER_TYPES = new Set([
'identifier',
'simple_identifier',
'variable_name', // PHP $variable (tree-sitter-php)
'name', // PHP name node
'this', // TS/JS/Java/C# this.method()
'self', // Rust/Python self.method()
]);
export const extractReceiverName = (
nameNode: SyntaxNode,
): string | undefined => {
const parent = nameNode.parent;
if (!parent) return undefined;
// PHP: member_call_expression / nullsafe_member_call_expression — receiver is on the callNode
// Java: method_invocation — receiver is the 'object' field on callNode
// For these, parent of nameNode is the call itself, so check the call's object field
const callNode = parent.parent ?? parent;
let receiver: SyntaxNode | null = null;
// Try standard field names used across grammars
receiver = parent.childForFieldName('object') // TS/JS member_expression, Python attribute, PHP, Java
?? parent.childForFieldName('value') // Rust field_expression
?? parent.childForFieldName('operand') // Go selector_expression
?? parent.childForFieldName('expression') // C# member_access_expression
?? parent.childForFieldName('argument'); // C++ field_expression
// Java method_invocation: 'object' field is on the callNode, not on nameNode's parent
if (!receiver && callNode.type === 'method_invocation') {
receiver = callNode.childForFieldName('object');
}
// PHP: member_call_expression has 'object' on the call node
if (!receiver && (callNode.type === 'member_call_expression' || callNode.type === 'nullsafe_member_call_expression')) {
receiver = callNode.childForFieldName('object');
}
// Kotlin/Swift: navigation_expression target is the first child
if (!receiver && parent.type === 'navigation_suffix') {
const navExpr = parent.parent;
if (navExpr?.type === 'navigation_expression') {
// First named child is the target (receiver)
for (const child of navExpr.children) {
if (child.isNamed && child !== parent) {
receiver = child;
break;
}
}
}
}
if (!receiver) return undefined;
// Only capture simple identifiers — refuse complex expressions
if (SIMPLE_RECEIVER_TYPES.has(receiver.type)) {
return receiver.text;
}
return undefined;
};
export const isVerboseIngestionEnabled = (): boolean => {
const raw = process.env.GITNEXUS_VERBOSE;
if (!raw) return false;
const value = raw.toLowerCase();
return value === '1' || value === 'true' || value === 'yes';
};
@@ -14,30 +14,14 @@ import PHP from 'tree-sitter-php';
import { createRequire } from 'node:module';
import { SupportedLanguages } from '../../../config/supported-languages.js';
import { LANGUAGE_QUERIES } from '../tree-sitter-queries.js';
import { getTreeSitterBufferSize, TREE_SITTER_MAX_BUFFER } from '../constants.js';
// tree-sitter-swift is an optionalDependency — may not be installed
const _require = createRequire(import.meta.url);
let Swift: any = null;
try { Swift = _require('tree-sitter-swift'); } catch {}
import {
getLanguageFromFilename,
FUNCTION_NODE_TYPES,
extractFunctionName,
isBuiltInOrNoise,
getDefinitionNodeFromCaptures,
findEnclosingClassId,
extractMethodSignature,
countCallArguments,
inferCallForm,
extractReceiverName
} from '../utils.js';
import { buildTypeEnv, lookupTypeEnv } from '../type-env.js';
import { isNodeExported } from '../export-detection.js';
import { findSiblingChild, getLanguageFromFilename } from '../utils.js';
import { detectFrameworkFromAST } from '../framework-detection.js';
import { generateId } from '../../../lib/utils.js';
import { extractNamedBindings } from '../named-binding-extraction.js';
import { appendKotlinWildcard } from '../resolvers/index.js';
// ============================================================================
// Types for serializable results
@@ -51,13 +35,11 @@ interface ParsedNode {
filePath: string;
startLine: number;
endLine: number;
language: SupportedLanguages;
language: string;
isExported: boolean;
astFrameworkMultiplier?: number;
astFrameworkReason?: string;
description?: string;
parameterCount?: number;
returnType?: string;
};
}
@@ -65,7 +47,7 @@ interface ParsedRelationship {
id: string;
sourceId: string;
targetId: string;
type: 'DEFINES' | 'HAS_METHOD';
type: 'DEFINES';
confidence: number;
reason: string;
}
@@ -75,16 +57,12 @@ interface ParsedSymbol {
name: string;
nodeId: string;
type: string;
parameterCount?: number;
ownerId?: string;
}
export interface ExtractedImport {
filePath: string;
rawImportPath: string;
language: SupportedLanguages;
/** Named bindings from the import (e.g., import {User as U} → [{local:'U', exported:'User'}]) */
namedBindings?: { local: string; exported: string }[];
language: string;
}
export interface ExtractedCall {
@@ -92,13 +70,6 @@ export interface ExtractedCall {
calledName: string;
/** generateId of enclosing function, or generateId('File', filePath) for top-level */
sourceId: string;
argCount?: number;
/** Discriminates free function calls from member/constructor calls */
callForm?: 'free' | 'member' | 'constructor';
/** Simple identifier of the receiver for member calls (e.g., 'user' in user.save()) */
receiverName?: string;
/** Resolved type name of the receiver (e.g., 'User' for user.save() when user: User) */
receiverTypeName?: string;
}
export interface ExtractedHeritage {
@@ -167,20 +138,198 @@ const setLanguage = (language: SupportedLanguages, filePath: string): void => {
parser.setLanguage(lang);
};
// isNodeExported imported from ../export-detection.js (shared module)
// ============================================================================
// Export detection (copied — needs AST parent traversal, can't cross threads)
// ============================================================================
const isNodeExported = (node: any, name: string, language: string): boolean => {
let current = node;
switch (language) {
case 'javascript':
case 'typescript':
while (current) {
const type = current.type;
if (type === 'export_statement' ||
type === 'export_specifier' ||
type === 'lexical_declaration' && current.parent?.type === 'export_statement') {
return true;
}
if (current.text?.startsWith('export ')) {
return true;
}
current = current.parent;
}
return false;
case 'python':
return !name.startsWith('_');
case 'java':
while (current) {
if (current.parent) {
const parent = current.parent;
for (let i = 0; i < parent.childCount; i++) {
const child = parent.child(i);
if (child?.type === 'modifiers' && child.text?.includes('public')) {
return true;
}
}
if (parent.type === 'method_declaration' || parent.type === 'constructor_declaration') {
if (parent.text?.trimStart().startsWith('public')) {
return true;
}
}
}
current = current.parent;
}
return false;
case 'csharp':
while (current) {
if (current.type === 'modifier' || current.type === 'modifiers') {
if (current.text?.includes('public')) return true;
}
current = current.parent;
}
return false;
case 'go':
if (name.length === 0) return false;
const first = name[0];
return first === first.toUpperCase() && first !== first.toLowerCase();
case 'rust':
while (current) {
if (current.type === 'visibility_modifier') {
if (current.text?.includes('pub')) return true;
}
current = current.parent;
}
return false;
// Kotlin: Default visibility is public (unlike Java)
// visibility_modifier is inside modifiers, a sibling of the name node within the declaration
case 'kotlin':
while (current) {
if (current.parent) {
const visMod = findSiblingChild(current.parent, 'modifiers', 'visibility_modifier');
if (visMod) {
const text = visMod.text;
if (text === 'private' || text === 'internal' || text === 'protected') return false;
if (text === 'public') return true;
}
}
current = current.parent;
}
// No visibility modifier = public (Kotlin default)
return true;
case 'c':
case 'cpp':
return false;
case 'php':
// Top-level classes/interfaces/traits are always accessible
// Methods/properties are exported only if they have 'public' modifier
while (current) {
if (current.type === 'class_declaration' ||
current.type === 'interface_declaration' ||
current.type === 'trait_declaration' ||
current.type === 'enum_declaration') {
return true;
}
if (current.type === 'visibility_modifier') {
return current.text === 'public';
}
current = current.parent;
}
// Top-level functions (no parent class) are globally accessible
return true;
case 'swift':
while (current) {
if (current.type === 'modifiers' || current.type === 'visibility_modifier') {
const text = current.text || '';
if (text.includes('public') || text.includes('open')) return true;
}
current = current.parent;
}
return false;
default:
return false;
}
};
// ============================================================================
// Enclosing function detection (for call extraction)
// ============================================================================
const FUNCTION_NODE_TYPES = new Set([
'function_declaration', 'arrow_function', 'function_expression',
'method_definition', 'generator_function_declaration',
'function_definition', 'async_function_declaration', 'async_arrow_function',
'method_declaration', 'constructor_declaration',
'local_function_statement', 'function_item', 'impl_item',
// Kotlin
'lambda_literal',
// PHP
'anonymous_function',
// Swift initializers/deinitializers
'init_declaration', 'deinit_declaration',
]);
/** Walk up AST to find enclosing function, return its generateId or null for top-level */
const findEnclosingFunctionId = (node: any, filePath: string): string | null => {
let current = node.parent;
while (current) {
if (FUNCTION_NODE_TYPES.has(current.type)) {
const { funcName, label } = extractFunctionName(current);
let funcName: string | null = null;
let label = 'Function';
if (current.type === 'init_declaration' || current.type === 'deinit_declaration') {
const funcName = current.type === 'init_declaration' ? 'init' : 'deinit';
const label = 'Constructor';
const startLine = current.startPosition?.row ?? 0;
return generateId(label, `${filePath}:${funcName}:${startLine}`);
}
if (['function_declaration', 'function_definition', 'async_function_declaration',
'generator_function_declaration', 'function_item'].includes(current.type)) {
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier' || c.type === 'property_identifier');
funcName = nameNode?.text;
} else if (current.type === 'impl_item') {
const funcItem = current.children?.find((c: any) => c.type === 'function_item');
if (funcItem) {
const nameNode = funcItem.childForFieldName?.('name') ||
funcItem.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
}
} else if (current.type === 'method_definition') {
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'property_identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'method_declaration' || current.type === 'constructor_declaration') {
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'arrow_function' || current.type === 'function_expression') {
const parent = current.parent;
if (parent?.type === 'variable_declarator') {
const nameNode = parent.childForFieldName?.('name') ||
parent.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
}
}
if (funcName) {
return generateId(label, `${filePath}:${funcName}`);
const startLine = current.startPosition?.row ?? 0;
return generateId(label, `${filePath}:${funcName}:${startLine}`);
}
}
current = current.parent;
@@ -188,6 +337,118 @@ const findEnclosingFunctionId = (node: any, filePath: string): string | null =>
return null;
};
const BUILT_INS = new Set([
// JavaScript/TypeScript
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export', 'fetch', 'Response', 'Request',
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'open', 'read', 'write', 'close', 'append', 'extend', 'update',
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
// Kotlin stdlib (IMPORTANT: keep in sync with call-processor.ts BUILT_IN_NAMES)
'println', 'print', 'readLine', 'require', 'requireNotNull', 'check', 'assert', 'lazy', 'error',
'listOf', 'mapOf', 'setOf', 'mutableListOf', 'mutableMapOf', 'mutableSetOf',
'arrayOf', 'sequenceOf', 'also', 'apply', 'run', 'with', 'takeIf', 'takeUnless',
'TODO', 'buildString', 'buildList', 'buildMap', 'buildSet',
'repeat', 'synchronized',
// Kotlin coroutine builders & scope functions
'launch', 'async', 'runBlocking', 'withContext', 'coroutineScope',
'supervisorScope', 'delay',
// Kotlin Flow operators
'flow', 'flowOf', 'collect', 'emit', 'onEach', 'catch',
'buffer', 'conflate', 'distinctUntilChanged',
'flatMapLatest', 'flatMapMerge', 'combine',
'stateIn', 'shareIn', 'launchIn',
// Kotlin infix stdlib functions
'to', 'until', 'downTo', 'step',
// C/C++ standard library
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
'scanf', 'fscanf', 'sscanf',
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
'sizeof', 'offsetof', 'typeof',
'assert', 'abort', 'exit', '_exit',
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
// Linux kernel common macros/helpers (not real call targets)
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
'min', 'max', 'clamp', 'abs', 'swap',
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
'GFP_KERNEL', 'GFP_ATOMIC',
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
'mutex_lock', 'mutex_unlock', 'mutex_init',
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
'get', 'put',
// PHP built-ins
'echo', 'isset', 'empty', 'unset', 'list', 'array', 'compact', 'extract',
'count', 'strlen', 'strpos', 'strrpos', 'substr', 'strtolower', 'strtoupper', 'trim',
'ltrim', 'rtrim', 'str_replace', 'str_contains', 'str_starts_with', 'str_ends_with',
'sprintf', 'vsprintf', 'printf', 'number_format',
'array_map', 'array_filter', 'array_reduce', 'array_push', 'array_pop', 'array_shift',
'array_unshift', 'array_slice', 'array_splice', 'array_merge', 'array_keys', 'array_values',
'array_key_exists', 'in_array', 'array_search', 'array_unique', 'usort', 'rsort',
'json_encode', 'json_decode', 'serialize', 'unserialize',
'intval', 'floatval', 'strval', 'boolval', 'is_null', 'is_string', 'is_int', 'is_array',
'is_object', 'is_numeric', 'is_bool', 'is_float',
'var_dump', 'print_r', 'var_export',
'date', 'time', 'strtotime', 'mktime', 'microtime',
'file_exists', 'file_get_contents', 'file_put_contents', 'is_file', 'is_dir',
'preg_match', 'preg_match_all', 'preg_replace', 'preg_split',
'header', 'session_start', 'session_destroy', 'ob_start', 'ob_end_clean', 'ob_get_clean',
'dd', 'dump',
// Swift/iOS built-ins and standard library
'print', 'debugPrint', 'dump', 'fatalError', 'precondition', 'preconditionFailure',
'assert', 'assertionFailure', 'NSLog',
'abs', 'min', 'max', 'zip', 'stride', 'sequence', 'repeatElement',
'swap', 'withUnsafePointer', 'withUnsafeMutablePointer', 'withUnsafeBytes',
'autoreleasepool', 'unsafeBitCast', 'unsafeDowncast', 'numericCast',
'type', 'MemoryLayout',
// Swift collection/string methods (common noise)
'map', 'flatMap', 'compactMap', 'filter', 'reduce', 'forEach', 'contains',
'first', 'last', 'prefix', 'suffix', 'dropFirst', 'dropLast',
'sorted', 'reversed', 'enumerated', 'joined', 'split',
'append', 'insert', 'remove', 'removeAll', 'removeFirst', 'removeLast',
'isEmpty', 'count', 'index', 'startIndex', 'endIndex',
// UIKit/Foundation common methods (noise in call graph)
'addSubview', 'removeFromSuperview', 'layoutSubviews', 'setNeedsLayout',
'layoutIfNeeded', 'setNeedsDisplay', 'invalidateIntrinsicContentSize',
'addTarget', 'removeTarget', 'addGestureRecognizer',
'addConstraint', 'addConstraints', 'removeConstraint', 'removeConstraints',
'NSLocalizedString', 'Bundle',
'reloadData', 'reloadSections', 'reloadRows', 'performBatchUpdates',
'register', 'dequeueReusableCell', 'dequeueReusableSupplementaryView',
'beginUpdates', 'endUpdates', 'insertRows', 'deleteRows', 'insertSections', 'deleteSections',
'present', 'dismiss', 'pushViewController', 'popViewController', 'popToRootViewController',
'performSegue', 'prepare',
// GCD / async
'DispatchQueue', 'async', 'sync', 'asyncAfter',
'Task', 'withCheckedContinuation', 'withCheckedThrowingContinuation',
// Combine
'sink', 'store', 'assign', 'receive', 'subscribe',
// Notification / KVO
'addObserver', 'removeObserver', 'post', 'NotificationCenter',
]);
// ============================================================================
// Label detection from capture map
// ============================================================================
@@ -222,8 +483,50 @@ const getLabelFromCaptures = (captureMap: Record<string, any>): string | null =>
return 'CodeElement';
};
// DEFINITION_CAPTURE_KEYS and getDefinitionNodeFromCaptures imported from ../utils.js
const DEFINITION_CAPTURE_KEYS = [
'definition.function',
'definition.class',
'definition.interface',
'definition.method',
'definition.struct',
'definition.enum',
'definition.namespace',
'definition.module',
'definition.trait',
'definition.impl',
'definition.type',
'definition.const',
'definition.static',
'definition.typedef',
'definition.macro',
'definition.union',
'definition.property',
'definition.record',
'definition.delegate',
'definition.annotation',
'definition.constructor',
'definition.template',
] as const;
const getDefinitionNodeFromCaptures = (captureMap: Record<string, any>): any | null => {
for (const key of DEFINITION_CAPTURE_KEYS) {
if (captureMap[key]) return captureMap[key];
}
return null;
};
/**
* Append .* to a Kotlin import path if the AST has a wildcard_import sibling node.
* Pure function — returns a new string without mutating the input.
*/
const appendKotlinWildcard = (importPath: string, importNode: any): string => {
for (let i = 0; i < importNode.childCount; i++) {
if (importNode.child(i)?.type === 'wildcard_import') {
return importPath.endsWith('.*') ? importPath : `${importPath}.*`;
}
}
return importPath;
};
// ============================================================================
// Process a batch of files
@@ -796,39 +1099,28 @@ const processFileGroup = (
try {
const lang = parser.getLanguage();
query = new Parser.Query(lang, queryString);
} catch (err) {
const message = `Query compilation failed for ${language}: ${err instanceof Error ? err.message : String(err)}`;
if (parentPort) {
parentPort.postMessage({ type: 'warning', message });
} else {
console.warn(message);
}
} catch {
return;
}
for (const file of files) {
// Skip files larger than the max tree-sitter buffer (32 MB)
if (file.content.length > TREE_SITTER_MAX_BUFFER) continue;
// Skip very large files — they can crash tree-sitter or cause OOM
if (file.content.length > 512 * 1024) continue;
let tree;
try {
tree = parser.parse(file.content, undefined, { bufferSize: getTreeSitterBufferSize(file.content.length) });
} catch (err) {
console.warn(`Failed to parse file ${file.path}: ${err instanceof Error ? err.message : String(err)}`);
tree = parser.parse(file.content, undefined, { bufferSize: 1024 * 256 });
} catch {
continue;
}
result.fileCount++;
onFileProcessed?.();
// Build per-file TypeEnv from explicit type annotations (for receiver resolution)
const typeEnv = buildTypeEnv(tree, language);
let matches;
try {
matches = query.matches(tree.rootNode);
} catch (err) {
console.warn(`Query execution failed for ${file.path}: ${err instanceof Error ? err.message : String(err)}`);
} catch {
continue;
}
@@ -843,12 +1135,10 @@ const processFileGroup = (
const rawImportPath = language === SupportedLanguages.Kotlin
? appendKotlinWildcard(captureMap['import.source'].text.replace(/['"<>]/g, ''), captureMap['import'])
: captureMap['import.source'].text.replace(/['"<>]/g, '');
const namedBindings = extractNamedBindings(captureMap['import'], language);
result.imports.push({
filePath: file.path,
rawImportPath,
language: language,
...(namedBindings ? { namedBindings } : {}),
});
continue;
}
@@ -858,22 +1148,11 @@ const processFileGroup = (
const callNameNode = captureMap['call.name'];
if (callNameNode) {
const calledName = callNameNode.text;
if (!isBuiltInOrNoise(calledName)) {
if (!BUILT_INS.has(calledName)) {
const callNode = captureMap['call'];
const sourceId = findEnclosingFunctionId(callNode, file.path)
|| generateId('File', file.path);
const callForm = inferCallForm(callNode, callNameNode);
const receiverName = callForm === 'member' ? extractReceiverName(callNameNode) : undefined;
const receiverTypeName = receiverName ? lookupTypeEnv(typeEnv, receiverName, callNode) : undefined;
result.calls.push({
filePath: file.path,
calledName,
sourceId,
argCount: countCallArguments(callNode),
...(callForm !== undefined ? { callForm } : {}),
...(receiverName !== undefined ? { receiverName } : {}),
...(receiverTypeName !== undefined ? { receiverTypeName } : {}),
});
result.calls.push({ filePath: file.path, calledName, sourceId });
}
}
continue;
@@ -882,21 +1161,12 @@ const processFileGroup = (
// Extract heritage (extends/implements)
if (captureMap['heritage.class']) {
if (captureMap['heritage.extends']) {
// Go struct embedding: the query matches ALL field_declarations with
// type_identifier, but only anonymous fields (no name) are embedded.
// Named fields like `Breed string` also match — skip them.
const extendsNode = captureMap['heritage.extends'];
const fieldDecl = extendsNode.parent;
const isNamedField = fieldDecl?.type === 'field_declaration'
&& fieldDecl.childForFieldName('name');
if (!isNamedField) {
result.heritage.push({
filePath: file.path,
className: captureMap['heritage.class'].text,
parentName: captureMap['heritage.extends'].text,
kind: 'extends',
});
}
result.heritage.push({
filePath: file.path,
className: captureMap['heritage.class'].text,
parentName: captureMap['heritage.extends'].text,
kind: 'extends',
});
}
if (captureMap['heritage.implements']) {
result.heritage.push({
@@ -928,7 +1198,7 @@ const processFileGroup = (
const nodeName = nameNode ? nameNode.text : 'init';
const definitionNode = getDefinitionNodeFromCaptures(captureMap);
const startLine = definitionNode ? definitionNode.startPosition.row : (nameNode ? nameNode.startPosition.row : 0);
const nodeId = generateId(nodeLabel, `${file.path}:${nodeName}`);
const nodeId = generateId(nodeLabel, `${file.path}:${nodeName}:${startLine}`);
let description: string | undefined;
if (language === SupportedLanguages.PHP) {
@@ -943,14 +1213,6 @@ const processFileGroup = (
? detectFrameworkFromAST(language, (definitionNode.text || '').slice(0, 300))
: null;
let parameterCount: number | undefined;
let returnType: string | undefined;
if (nodeLabel === 'Function' || nodeLabel === 'Method' || nodeLabel === 'Constructor') {
const sig = extractMethodSignature(definitionNode);
parameterCount = sig.parameterCount;
returnType = sig.returnType;
}
result.nodes.push({
id: nodeId,
label: nodeLabel,
@@ -966,23 +1228,14 @@ const processFileGroup = (
astFrameworkReason: frameworkHint.reason,
} : {}),
...(description !== undefined ? { description } : {}),
...(parameterCount !== undefined ? { parameterCount } : {}),
...(returnType !== undefined ? { returnType } : {}),
},
});
// Compute enclosing class for Method/Constructor/Property/Function — used for both ownerId and HAS_METHOD
// Function is included because Kotlin/Rust/Python capture class methods as Function nodes
const needsOwner = nodeLabel === 'Method' || nodeLabel === 'Constructor' || nodeLabel === 'Property' || nodeLabel === 'Function';
const enclosingClassId = needsOwner ? findEnclosingClassId(nameNode || definitionNode, file.path) : null;
result.symbols.push({
filePath: file.path,
name: nodeName,
nodeId,
type: nodeLabel,
...(parameterCount !== undefined ? { parameterCount } : {}),
...(enclosingClassId ? { ownerId: enclosingClassId } : {}),
});
const fileId = generateId('File', file.path);
@@ -995,18 +1248,6 @@ const processFileGroup = (
confidence: 1.0,
reason: '',
});
// ── HAS_METHOD: link method/constructor/property to enclosing class ──
if (enclosingClassId) {
result.relationships.push({
id: generateId('HAS_METHOD', `${enclosingClassId}->${nodeId}`),
sourceId: enclosingClassId,
targetId: nodeId,
type: 'HAS_METHOD',
confidence: 1.0,
reason: '',
});
}
}
// Extract Laravel routes from route files via procedural AST walk
@@ -1,7 +1,5 @@
import { Worker } from 'node:worker_threads';
import os from 'node:os';
import fs from 'node:fs';
import { fileURLToPath } from 'node:url';
export interface WorkerPool {
/**
@@ -32,13 +30,6 @@ const SUB_BATCH_TIMEOUT_MS = 30_000;
* Create a pool of worker threads.
*/
export const createWorkerPool = (workerUrl: URL, poolSize?: number): WorkerPool => {
// Validate worker script exists before spawning to prevent uncaught
// MODULE_NOT_FOUND crashes in worker threads (e.g. when running from src/ via vitest)
const workerPath = fileURLToPath(workerUrl);
if (!fs.existsSync(workerPath)) {
throw new Error(`Worker script not found: ${workerPath}`);
}
const size = poolSize ?? Math.min(8, Math.max(1, os.cpus().length - 1));
const workers: Worker[] = [];
+3 -19
View File
@@ -232,8 +232,7 @@ export const streamAllCSVsToDisk = async (
const functionWriter = new BufferedCSVWriter(path.join(csvDir, 'function.csv'), codeElementHeader);
const classWriter = new BufferedCSVWriter(path.join(csvDir, 'class.csv'), codeElementHeader);
const interfaceWriter = new BufferedCSVWriter(path.join(csvDir, 'interface.csv'), codeElementHeader);
const methodHeader = 'id,name,filePath,startLine,endLine,isExported,content,description,parameterCount,returnType';
const methodWriter = new BufferedCSVWriter(path.join(csvDir, 'method.csv'), methodHeader);
const methodWriter = new BufferedCSVWriter(path.join(csvDir, 'method.csv'), codeElementHeader);
const codeElemWriter = new BufferedCSVWriter(path.join(csvDir, 'codeelement.csv'), codeElementHeader);
const communityWriter = new BufferedCSVWriter(path.join(csvDir, 'community.csv'), 'id,label,heuristicLabel,keywords,description,enrichedBy,cohesion,symbolCount');
const processWriter = new BufferedCSVWriter(path.join(csvDir, 'process.csv'), 'id,label,heuristicLabel,processType,stepCount,communities,entryPointId,terminalId');
@@ -251,6 +250,7 @@ export const streamAllCSVsToDisk = async (
'Function': functionWriter,
'Class': classWriter,
'Interface': interfaceWriter,
'Method': methodWriter,
'CodeElement': codeElemWriter,
};
@@ -308,24 +308,8 @@ export const streamAllCSVsToDisk = async (
].join(','));
break;
}
case 'Method': {
const content = await extractContent(node, contentCache);
await methodWriter.addRow([
escapeCSVField(node.id),
escapeCSVField(node.properties.name || ''),
escapeCSVField(node.properties.filePath || ''),
escapeCSVNumber(node.properties.startLine, -1),
escapeCSVNumber(node.properties.endLine, -1),
node.properties.isExported ? 'true' : 'false',
escapeCSVField(content),
escapeCSVField((node.properties as any).description || ''),
escapeCSVNumber(node.properties.parameterCount, 0),
escapeCSVField(node.properties.returnType || ''),
].join(','));
break;
}
default: {
// Code element nodes (Function, Class, Interface, CodeElement)
// Code element nodes (Function, Class, Interface, Method, CodeElement)
const writer = codeWriterMap[node.label];
if (writer) {
const content = await extractContent(node, contentCache);
+2 -6
View File
@@ -328,9 +328,6 @@ const getCopyQuery = (table: NodeTableName, filePath: string): string => {
if (table === 'Process') {
return `COPY ${t}(id, label, heuristicLabel, processType, stepCount, communities, entryPointId, terminalId) FROM "${filePath}" ${COPY_CSV_OPTS}`;
}
if (table === 'Method') {
return `COPY ${t}(id, name, filePath, startLine, endLine, isExported, content, description, parameterCount, returnType) FROM "${filePath}" ${COPY_CSV_OPTS}`;
}
// TypeScript/JS code element tables have isExported; multi-language tables do not
if (TABLES_WITH_EXPORTED.has(table)) {
return `COPY ${t}(id, name, filePath, startLine, endLine, isExported, content, description) FROM "${filePath}" ${COPY_CSV_OPTS}`;
@@ -594,7 +591,6 @@ export const closeKuzu = async (): Promise<void> => {
export const isKuzuReady = (): boolean => conn !== null && db !== null;
/**
* Delete all nodes (and their relationships) for a specific file from KuzuDB
* @param filePath - The file path to delete nodes for
@@ -750,8 +746,8 @@ export const queryFTS = async (
throw new Error('KuzuDB not initialized. Call initKuzu first.');
}
// Escape backslashes and single quotes to prevent Cypher injection
const escapedQuery = query.replace(/\\/g, '\\\\').replace(/'/g, "''");
// Escape single quotes in query
const escapedQuery = query.replace(/'/g, "''");
const cypher = `
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', '${escapedQuery}', conjunctive := ${conjunctive})
+1 -16
View File
@@ -26,7 +26,7 @@ export type NodeTableName = typeof NODE_TABLES[number];
export const REL_TABLE_NAME = 'CodeRelation';
// Valid relation types
export const REL_TYPES = ['CONTAINS', 'DEFINES', 'IMPORTS', 'CALLS', 'EXTENDS', 'IMPLEMENTS', 'HAS_METHOD', 'OVERRIDES', 'MEMBER_OF', 'STEP_IN_PROCESS'] as const;
export const REL_TYPES = ['CONTAINS', 'DEFINES', 'IMPORTS', 'CALLS', 'EXTENDS', 'IMPLEMENTS', 'MEMBER_OF', 'STEP_IN_PROCESS'] as const;
export type RelType = typeof REL_TYPES[number];
// ============================================================================
@@ -104,8 +104,6 @@ CREATE NODE TABLE Method (
isExported BOOLEAN,
content STRING,
description STRING,
parameterCount INT32,
returnType STRING,
PRIMARY KEY (id)
)`;
@@ -262,7 +260,6 @@ CREATE REL TABLE ${REL_TABLE_NAME} (
FROM Class TO \`Union\`,
FROM Class TO \`Namespace\`,
FROM Class TO \`Typedef\`,
FROM Class TO \`Property\`,
FROM Method TO Function,
FROM Method TO Method,
FROM Method TO Class,
@@ -298,7 +295,6 @@ CREATE REL TABLE ${REL_TABLE_NAME} (
FROM Interface TO \`TypeAlias\`,
FROM Interface TO \`Struct\`,
FROM Interface TO \`Constructor\`,
FROM Interface TO \`Property\`,
FROM \`Struct\` TO Community,
FROM \`Struct\` TO \`Trait\`,
FROM \`Struct\` TO \`Struct\`,
@@ -307,8 +303,6 @@ CREATE REL TABLE ${REL_TABLE_NAME} (
FROM \`Struct\` TO Function,
FROM \`Struct\` TO Method,
FROM \`Struct\` TO Interface,
FROM \`Struct\` TO \`Constructor\`,
FROM \`Struct\` TO \`Property\`,
FROM \`Enum\` TO \`Enum\`,
FROM \`Enum\` TO Community,
FROM \`Enum\` TO Class,
@@ -322,13 +316,7 @@ CREATE REL TABLE ${REL_TABLE_NAME} (
FROM \`Union\` TO Community,
FROM \`Namespace\` TO Community,
FROM \`Namespace\` TO \`Struct\`,
FROM \`Trait\` TO Method,
FROM \`Trait\` TO \`Constructor\`,
FROM \`Trait\` TO \`Property\`,
FROM \`Trait\` TO Community,
FROM \`Impl\` TO Method,
FROM \`Impl\` TO \`Constructor\`,
FROM \`Impl\` TO \`Property\`,
FROM \`Impl\` TO Community,
FROM \`Impl\` TO \`Trait\`,
FROM \`Impl\` TO \`Struct\`,
@@ -339,9 +327,6 @@ CREATE REL TABLE ${REL_TABLE_NAME} (
FROM \`Const\` TO Community,
FROM \`Static\` TO Community,
FROM \`Property\` TO Community,
FROM \`Record\` TO Method,
FROM \`Record\` TO \`Constructor\`,
FROM \`Record\` TO \`Property\`,
FROM \`Record\` TO Community,
FROM \`Delegate\` TO Community,
FROM \`Annotation\` TO Community,
+1 -2
View File
@@ -24,8 +24,7 @@ async function queryFTSViaExecutor(
query: string,
limit: number,
): Promise<Array<{ filePath: string; score: number }>> {
// Escape single quotes and backslashes to prevent Cypher injection
const escapedQuery = query.replace(/\\/g, '\\\\').replace(/'/g, "''");
const escapedQuery = query.replace(/'/g, "''");
const cypher = `
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', '${escapedQuery}', conjunctive := false)
RETURN node, score
@@ -1,240 +0,0 @@
import process from 'node:process';
import type { Transport, TransportSendOptions } from '@modelcontextprotocol/sdk/shared/transport.js';
import { JSONRPCMessageSchema, type JSONRPCMessage } from '@modelcontextprotocol/sdk/types.js';
export type StdioFraming = 'content-length' | 'newline';
function deserializeMessage(raw: string): JSONRPCMessage {
return JSONRPCMessageSchema.parse(JSON.parse(raw));
}
function serializeNewlineMessage(message: JSONRPCMessage): string {
return `${JSON.stringify(message)}\n`;
}
function serializeContentLengthMessage(message: JSONRPCMessage): string {
const body = JSON.stringify(message);
return `Content-Length: ${Buffer.byteLength(body, 'utf8')}\r\n\r\n${body}`;
}
function findHeaderEnd(buffer: Buffer): { index: number; separatorLength: number } | null {
const crlfEnd = buffer.indexOf('\r\n\r\n');
if (crlfEnd !== -1) {
return { index: crlfEnd, separatorLength: 4 };
}
const lfEnd = buffer.indexOf('\n\n');
if (lfEnd !== -1) {
return { index: lfEnd, separatorLength: 2 };
}
return null;
}
function looksLikeContentLength(buffer: Buffer): boolean {
if (buffer.length < 14) {
return false;
}
const probe = buffer.toString('utf8', 0, Math.min(buffer.length, 32));
return /^content-length\s*:/i.test(probe);
}
const MAX_BUFFER_SIZE = 10 * 1024 * 1024; // 10 MB — generous for JSON-RPC
export class CompatibleStdioServerTransport implements Transport {
private _readBuffer: Buffer | undefined;
private _started = false;
private _framing: StdioFraming | null = null;
onmessage?: (message: JSONRPCMessage) => void;
onerror?: (error: Error) => void;
onclose?: () => void;
constructor(
private readonly _stdin: NodeJS.ReadableStream = process.stdin,
private readonly _stdout: NodeJS.WritableStream = process.stdout,
) {}
private readonly _ondata = (chunk: Buffer) => {
this._readBuffer = this._readBuffer ? Buffer.concat([this._readBuffer, chunk]) : chunk;
if (this._readBuffer.length > MAX_BUFFER_SIZE) {
this.onerror?.(new Error(`Read buffer exceeded maximum size (${MAX_BUFFER_SIZE} bytes)`));
this.discardBufferedInput();
return;
}
this.processReadBuffer();
};
private readonly _onerror = (error: Error) => {
this.onerror?.(error);
};
async start() {
if (this._started) {
throw new Error('CompatibleStdioServerTransport already started!');
}
this._started = true;
this._stdin.on('data', this._ondata);
this._stdin.on('error', this._onerror);
}
private detectFraming(): StdioFraming | null {
if (!this._readBuffer || this._readBuffer.length === 0) {
return null;
}
const firstByte = this._readBuffer[0];
if (firstByte === 0x7b || firstByte === 0x5b) {
return 'newline';
}
if (looksLikeContentLength(this._readBuffer)) {
return 'content-length';
}
return null;
}
private discardBufferedInput() {
this._readBuffer = undefined;
this._framing = null;
}
private readContentLengthMessage(): JSONRPCMessage | null {
if (!this._readBuffer) {
return null;
}
const header = findHeaderEnd(this._readBuffer);
if (header === null) {
return null;
}
const headerText = this._readBuffer
.toString('utf8', 0, header.index)
.replace(/\r\n/g, '\n')
.replace(/\r/g, '\n');
const match = headerText.match(/(?:^|\n)content-length\s*:\s*(\d+)/i);
if (!match) {
this.discardBufferedInput();
throw new Error('Missing Content-Length header from MCP client');
}
const contentLength = Number.parseInt(match[1], 10);
if (!Number.isFinite(contentLength) || contentLength < 0) {
this.discardBufferedInput();
throw new Error('Invalid Content-Length header from MCP client');
}
if (contentLength > MAX_BUFFER_SIZE) {
this.discardBufferedInput();
throw new Error(`Content-Length ${contentLength} exceeds maximum allowed size (${MAX_BUFFER_SIZE} bytes)`);
}
const bodyStart = header.index + header.separatorLength;
const bodyEnd = bodyStart + contentLength;
if (this._readBuffer.length < bodyEnd) {
return null;
}
const body = this._readBuffer.toString('utf8', bodyStart, bodyEnd);
this._readBuffer = this._readBuffer.subarray(bodyEnd);
return deserializeMessage(body);
}
private readNewlineMessage(): JSONRPCMessage | null {
if (!this._readBuffer) {
return null;
}
while (true) {
const newlineIndex = this._readBuffer.indexOf('\n');
if (newlineIndex === -1) {
return null;
}
const line = this._readBuffer.toString('utf8', 0, newlineIndex).replace(/\r$/, '');
this._readBuffer = this._readBuffer.subarray(newlineIndex + 1);
if (line.trim().length === 0) {
continue;
}
return deserializeMessage(line);
}
}
private readMessage(): JSONRPCMessage | null {
if (!this._readBuffer || this._readBuffer.length === 0) {
return null;
}
if (this._framing === null) {
this._framing = this.detectFraming();
if (this._framing === null) {
return null;
}
}
return this._framing === 'content-length'
? this.readContentLengthMessage()
: this.readNewlineMessage();
}
private processReadBuffer() {
while (true) {
try {
const message = this.readMessage();
if (message === null) {
break;
}
this.onmessage?.(message);
} catch (error) {
this.onerror?.(error as Error);
break;
}
}
}
async close() {
this._stdin.off('data', this._ondata);
this._stdin.off('error', this._onerror);
const remainingDataListeners = this._stdin.listenerCount('data');
if (remainingDataListeners === 0) {
this._stdin.pause();
}
this._started = false;
this._readBuffer = undefined;
this.onclose?.();
}
send(message: JSONRPCMessage, _options?: TransportSendOptions) {
return new Promise<void>((resolve, reject) => {
if (!this._started) {
reject(new Error('Transport is closed'));
return;
}
const payload = this._framing === 'newline'
? serializeNewlineMessage(message)
: serializeContentLengthMessage(message);
const onError = (error: Error) => {
this._stdout.removeListener('error', onError);
reject(error);
};
this._stdout.on('error', onError);
if (this._stdout.write(payload)) {
this._stdout.removeListener('error', onError);
resolve();
} else {
this._stdout.once('drain', () => {
this._stdout.removeListener('error', onError);
resolve();
});
}
});
}
}
+7 -7
View File
@@ -84,14 +84,15 @@ function evictLRU(): void {
}
/**
* Remove a repo from the pool without calling native close methods.
*
* KuzuDB's native .closeSync() triggers N-API destructor hooks that
* segfault on Linux/macOS. Pool databases are opened read-only, so
* there is no WAL to flush — just deleting the pool entry and letting
* the GC (or process exit) reclaim native resources is safe.
* Close all connections for a repo and remove it from the pool
*/
function closeOne(repoId: string): void {
const entry = pool.get(repoId);
if (!entry) return;
for (const conn of entry.available) {
try { conn.close(); } catch (e) { console.error('GitNexus [pool:close-conn]:', e instanceof Error ? e.message : e); }
}
try { entry.db.close(); } catch (e) { console.error('GitNexus [pool:close-db]:', e instanceof Error ? e.message : e); }
pool.delete(repoId);
}
@@ -324,7 +325,6 @@ export const closeKuzu = async (repoId?: string): Promise<void> => {
}
};
/**
* Check if a specific repo's pool is active
*/
+2 -2
View File
@@ -13,7 +13,7 @@
import { createRequire } from 'module';
import { Server } from '@modelcontextprotocol/sdk/server/index.js';
import { CompatibleStdioServerTransport } from './compatible-stdio-transport.js';
import { StdioServerTransport } from '@modelcontextprotocol/sdk/server/stdio.js';
import {
CallToolRequestSchema,
ListToolsRequestSchema,
@@ -277,7 +277,7 @@ export async function startMCPServer(backend: LocalBackend): Promise<void> {
const server = createMCPServer(backend);
// Connect to stdio transport
const transport = new CompatibleStdioServerTransport();
const transport = new StdioServerTransport();
await server.connect(transport);
// Graceful shutdown helper
+3 -12
View File
@@ -78,7 +78,7 @@ SCHEMA:
- Nodes: File, Folder, Function, Class, Interface, Method, CodeElement, Community, Process
- Multi-language nodes (use backticks): \`Struct\`, \`Enum\`, \`Trait\`, \`Impl\`, etc.
- All edges via single CodeRelation table with 'type' property
- Edge types: CONTAINS, DEFINES, CALLS, IMPORTS, EXTENDS, IMPLEMENTS, HAS_METHOD, OVERRIDES, MEMBER_OF, STEP_IN_PROCESS
- Edge types: CONTAINS, DEFINES, CALLS, IMPORTS, EXTENDS, IMPLEMENTS, MEMBER_OF, STEP_IN_PROCESS
- Edge properties: type (STRING), confidence (DOUBLE), reason (STRING), step (INT32)
EXAMPLES:
@@ -91,15 +91,6 @@ EXAMPLES:
• Trace a process:
MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process) WHERE p.heuristicLabel = "UserLogin" RETURN s.name, r.step ORDER BY r.step
• Find all methods of a class:
MATCH (c:Class {name: "UserService"})-[r:CodeRelation {type: 'HAS_METHOD'}]->(m:Method) RETURN m.name, m.parameterCount, m.returnType
• Find method overrides (MRO resolution):
MATCH (winner:Method)-[r:CodeRelation {type: 'OVERRIDES'}]->(loser:Method) RETURN winner.name, winner.filePath, loser.filePath, r.reason
• Detect diamond inheritance:
MATCH (d:Class)-[:CodeRelation {type: 'EXTENDS'}]->(b1), (d)-[:CodeRelation {type: 'EXTENDS'}]->(b2), (b1)-[:CodeRelation {type: 'EXTENDS'}]->(a), (b2)-[:CodeRelation {type: 'EXTENDS'}]->(a) WHERE b1 <> b2 RETURN d.name, b1.name, b2.name, a.name
OUTPUT: Returns { markdown, row_count } — results formatted as a Markdown table for easy reading.
TIPS:
@@ -200,7 +191,7 @@ Depth groups:
- d=2: LIKELY AFFECTED (indirect)
- d=3: MAY NEED TESTING (transitive)
EdgeType: CALLS, IMPORTS, EXTENDS, IMPLEMENTS, HAS_METHOD, OVERRIDES
EdgeType: CALLS, IMPORTS, EXTENDS, IMPLEMENTS
Confidence: 1.0 = certain, <0.8 = fuzzy match`,
inputSchema: {
type: 'object',
@@ -208,7 +199,7 @@ Confidence: 1.0 = certain, <0.8 = fuzzy match`,
target: { type: 'string', description: 'Name of function, class, or file to analyze' },
direction: { type: 'string', description: 'upstream (what depends on this) or downstream (what this depends on)' },
maxDepth: { type: 'number', description: 'Max relationship depth (default: 3)', default: 3 },
relationTypes: { type: 'array', items: { type: 'string' }, description: 'Filter: CALLS, IMPORTS, EXTENDS, IMPLEMENTS, HAS_METHOD, OVERRIDES (default: usage-based)' },
relationTypes: { type: 'array', items: { type: 'string' }, description: 'Filter: CALLS, IMPORTS, EXTENDS, IMPLEMENTS (default: usage-based)' },
includeTests: { type: 'boolean', description: 'Include test files (default: false)' },
minConfidence: { type: 'number', description: 'Minimum confidence 0-1 (default: 0.7)' },
repo: { type: 'string', description: 'Repository name or path. Omit if only one repo is indexed.' },
@@ -1,6 +0,0 @@
#pragma once
class Handler {
public:
void handle();
};
@@ -1,6 +0,0 @@
#pragma once
class Handler {
public:
void process();
};
@@ -1,8 +0,0 @@
#pragma once
#include "handler_a.h"
class Processor : public Handler {
public:
void run();
};
@@ -1,6 +0,0 @@
#include "one.h"
#include "zero.h"
void run() {
write_audit("hello");
}
@@ -1,3 +0,0 @@
inline const char* write_audit(const char* message) {
return message;
}
@@ -1,3 +0,0 @@
inline const char* write_audit() {
return "zero";
}
@@ -1,6 +0,0 @@
#include "user.h"
void processUser(const std::string& name) {
auto user = new User(name);
user->save();
}
@@ -1,10 +0,0 @@
#pragma once
#include <string>
class User {
public:
User(const std::string& name) : name_(name) {}
bool save() { return true; }
private:
std::string name_;
};
@@ -1,7 +0,0 @@
#pragma once
class Animal {
public:
virtual void speak();
virtual void move();
};
@@ -1,5 +0,0 @@
#include "duck.h"
void Duck::speak() {
// quack
}
@@ -1,8 +0,0 @@
#pragma once
#include "flyer.h"
#include "swimmer.h"
class Duck : public Flyer, public Swimmer {
public:
void speak() override;
};
@@ -1,8 +0,0 @@
#pragma once
#include "animal.h"
class Flyer : public Animal {
public:
void move() override;
void fly();
};
@@ -1,8 +0,0 @@
#pragma once
#include "animal.h"
class Swimmer : public Animal {
public:
void move() override;
void swim();
};
@@ -1,3 +0,0 @@
cmake_minimum_required(VERSION 3.10)
project(cpp-local-shadow)
add_executable(main src/main.cpp src/utils.cpp)

Some files were not shown because too many files have changed in this diff Show More