Compare commits

...
Author SHA1 Message Date
claude[bot]andGergő Magyar 8d39532185 test(ruby): add ingestion integration tests for Ruby language support
Adds 8 fixture-driven integration test scenarios modelled after the
existing C/C++ and Python resolver test suites:

- ruby-pkg: require_relative imports, class inheritance (EXTENDS), module
  include (IMPLEMENTS), method detection across 5 files
- ruby-ambiguous: two Handler classes in a/ and b/, require_relative to a/
  disambiguates the EXTENDS target
- ruby-calls: arity-filtered call resolution — write_audit("hello") resolves
  to one.rb (1-param) not zero.rb (0-params) via import-resolved
- ruby-member-calls: obj.method() member call resolves User.save via the
  import map; verifies HAS_METHOD edge
- ruby-mixin-heritage: include / extend / prepend all produce IMPLEMENTS
  edges; no EXTENDS emitted for any mixin form
- ruby-attr-properties: attr_accessor / attr_reader / attr_writer create
  Property nodes with correct accessorType in description
- ruby-receiver-resolution: both User and Repo classes exist with save
  methods; HAS_METHOD verified for each; at least one save call resolved
- ruby-local-shadow: local def save in main.rb shadows the require'd
  utils.rb#save; run → save resolves to same-file definition

Co-authored-by: Gergő Magyar <magyargergo@users.noreply.github.com>
2026-03-13 19:47:23 +00:00
Candido Gomes fbd1a6f0d8 feat: add Ruby language support in parsing tests and loader functionality 2026-03-13 14:58:08 -04:00
Candido Gomes c2c2937978 feat: implement Ruby call routing for imports, heritage, and properties 2026-03-13 14:47:32 -04:00
Candido Gomes 055fa96c0d feat: add Ruby import resolution support with require and require_relative handling 2026-03-13 14:36:29 -04:00
Candido Gomes 527d0bf775 revert 2026-03-13 12:30:37 -04:00
Candido Gomes c1bd27a7e0 revert 2026-03-13 12:28:48 -04:00
Candido Gomes eb73f0f600 feat: enhance syntax highlighting support for Ruby and other file types 2026-03-13 12:06:09 -04:00
Candido Gomes 743245a620 feat: add Kotlin support to tree-sitter integration and update related configurations 2026-03-13 11:56:07 -04:00
Candido Gomes ec1815cf0b feat: add Ruby framework detection for Rakefile and .rake extensions 2026-03-13 11:48:26 -04:00
Candido Gomes fc025cfd22 fix: correct Ruby queries syntax by closing string literal 2026-03-13 11:26:19 -04:00
Candido Gomes 27415265c7 feat: enhance Ruby framework detection and update tree-sitter queries 2026-03-13 11:24:17 -04:00
Candido Gomes a2814dc26e Merge upstream/main: resolve conflicts keeping Ruby support with shared modules
Conflicts resolved by taking upstream refactored shared module approach
(extractFunctionName, isBuiltInOrNoise, isNodeExported from utils.js and
export-detection.js) while preserving Ruby-specific call handling (require,
include/extend/prepend, attr_accessor) and adding Ruby to the language matrix.
2026-03-13 09:54:03 -04:00
Gergo Magyar 3dbe08fab6 chore: release v1.4.0 2026-03-13 13:18:14 +00:00
Gergő Magyar 1afe9166aa feat: language-aware code intelligence — symbol resolution, MRO, constructor discrimination (#238)
* feat: add Method Resolution Order (MRO) with language-specific rules

Implement full MRO computation for multi-language inheritance hierarchies:

- HAS_METHOD edges: Class→Method ownership edges emitted during parsing
  (both worker pool and sequential fallback paths)
- Method signatures: extract parameterCount and returnType from AST nodes
- C# heritage fix: distinguish EXTENDS vs IMPLEMENTS for base_list captures
  using symbol table lookup + I[A-Z] naming heuristic fallback
- MRO processor (Phase 4.5): walks inheritance DAG, detects method-name
  collisions across parents, applies language-specific resolution:
  - C++: leftmost base class in declaration order wins
  - C#/Java: class method wins over interface default
  - Python: C3 linearization with cycle detection
  - Rust: no auto-resolution (requires qualified syntax)
  - Default: first definition in BFS order wins
- OVERRIDES edges emitted for resolved method collisions
- KuzuDB schema: Method table extended with parameterCount/returnType;
  dedicated CSV writer and COPY query for 10-column Method rows
- MCP tools: updated Cypher examples for HAS_METHOD, OVERRIDES, diamond

72 tests across 5 test files covering MRO resolution, HAS_METHOD edges,
method signature extraction, C# heritage resolution, and integration
tests across C#/Rust/Python/TS/Java/C++.

* feat: add scope-based symbol resolution replacing raw lookupFuzzy

Introduces a shared 3-tier resolveSymbol function used by both
heritage-processor and call-processor:
1. Same-file (lookupExactFull — authoritative)
2. Import-scoped (filtered by ImportMap — high confidence)
3. Global fuzzy (first match — low confidence fallback)

Adds lookupExactFull to SymbolTable returning full SymbolDefinition
with type info needed for heritage Class/Interface disambiguation.

* refactor: tighten symbol resolution — Tier 3 refuses ambiguous matches

- lookupExactFull now O(1) via direct SymbolDefinition storage in fileIndex
  (shared object references with globalIndex — zero additional memory)
- Added resolveSymbolInternal() preserving { definition, tier, candidateCount }
  for test assertions and logging
- Tier 3 now returns null when multiple global candidates exist instead of
  arbitrary allDefs[0] — a wrong edge is worse than no edge
- call-processor: renamed fuzzy-global → unique-global, removed dead branch
- 12 new tests: tier assertions, ambiguous refusal per language family,
  heritage false-positive guard, O(1) shared reference verification

* fix: critical language support bugs in import resolution and MRO

Phase 5 critical fixes from all-language analysis:
- Python: add relative_import query capture (PEP 328) — `.models`, `..utils`
  were silently dropped, producing zero ImportMap entries
- Rust: extract prefix from grouped imports `crate::module::{A, B}` — brace
  groups previously failed resolution entirely
- Swift: use normalizedFileList for Windows path compatibility in module
  import resolution (matches Go's resolveGoPackage pattern)
- MRO: fix c_sharp → csharp language name mismatch (enum is 'csharp'),
  add Kotlin to C#/Java resolution rules (class method wins over interface)

* feat: add strict multi-language integration tests + fix C/C++ import resolution

Add 32 integration tests across 6 language fixtures (TypeScript, C#, C++,
Java, Python, Rust) with exact toBe/toEqual assertions validating heritage
edges, import resolution, and trait implementations.

Fix C/C++ import resolution bug where dot-to-slash conversion mangled
include paths (e.g. "animal.h" became "animal/h"). Now skips conversion
for C/C++ languages which use actual file paths in #include directives.

* fix: language-gate heritage heuristic, add Swift extension heritage, handle Rust grouped imports

- Gate I[A-Z] naming heuristic to C#/Java only (was firing for all languages)
- Swift unresolved types default to IMPLEMENTS (protocol conformance is the norm)
- Add tree-sitter query for Swift extension protocol conformance (extension Foo: Protocol)
- Handle Rust top-level grouped imports (use {crate::a, crate::b}) in both import loops
- Add 4 new heritage-processor tests (TypeScript refusal, Swift default, Swift Tier 1)

* feat: add Go struct embedding heritage + PackageMap optimization

Add Go struct embedding detection (anonymous fields → EXTENDS edges) via
new tree-sitter heritage query with named-field filtering in both
parse-worker and heritage-processor paths.

Implement PackageMap optimization for Go cross-package resolution:
replace O(N) file-level ImportMap expansion with directory-level suffix
matching (Tier 2b in symbol resolver). Graph IMPORTS edges are preserved
via addImportGraphEdge split.

Remove overly broad @definition.type from GO_QUERIES that was
double-matching structs/interfaces as TypeAlias nodes, breaking Tier 3
unique-global resolution.

Add Go fixture (go-pkg) with Admin→User embedding, cross-package calls,
and 7 integration tests covering structs, functions, imports, calls,
and heritage edges.

* test: add Kotlin heritage integration tests

Adds a kotlin-heritage fixture and 7 integration tests validating
class inheritance, interface implementation, JVM-style import
resolution, and symbol-table-driven EXTENDS/IMPLEMENTS disambiguation
via Kotlin delegation specifiers.

* feat: extract resolvers, add PHP tests, ambiguous tests for all languages

- Extract language-specific resolvers from import-processor.ts into
  resolvers/ directory (P7): jvm, go, csharp, php, rust, standard, utils
- import-processor.ts reduced from 1412 to 711 lines (50% reduction)
- Add comprehensive PHP integration tests: PSR-4 imports, traits, enums,
  heritage edges, method calls, MRO overrides
- Add ambiguous symbol resolution tests for all 9 languages verifying
  correct disambiguation via import chains
- Split monolithic lang-resolution.test.ts (1080 lines) into 9 per-language
  files under test/integration/resolvers/ with shared helpers

* feat: update integration tests to include resolver tests for multiple languages

* fix: address code review — schema gap, Rust impl name, Property OVERRIDES

Bugs fixed:
- Add 13 missing FROM/TO pairs in RELATION_SCHEMA for HAS_METHOD edges
  (Class/Interface/Struct/Trait/Impl/Record to Method/Constructor/Property)
- Fix findEnclosingClassId to pick implementing type for Rust
  impl Trait for Struct blocks (was picking trait name)
- Exclude Property nodes from MRO OVERRIDES collision detection
- Change MRO language fallback from typescript to unknown

Tests added:
- Unit: Property OVERRIDES exclusion (2 tests), Rust impl Trait for
  Struct name resolution (2 tests), schema HAS_METHOD pair coverage
- Integration: no OVERRIDES targets Property nodes across all 9 languages
- PHP fixture: added shared $status property to both traits to create
  real collision scenario for Property OVERRIDES exclusion test

Documentation:
- OVERRIDES edge direction (Class to Method), Go return type gap,
  BFS first-reach heuristic limitation

* feat: harden CALLS-edge resolution — Phase 0 validation

- Fix same-file confidence (0.85 → 0.95) to correctly outrank import-scoped (0.9)
- Fix Tier 1 overload preservation: use globalIndex filter instead of fileIndex lookup
- Add callable-kind guard: refuse CALLS edges to Interface and Enum symbols
- Fix Kotlin countCallArguments: handle call_suffix → value_arguments nesting
- Fix Kotlin extractFunctionName: add simple_identifier to fallback search
- Strictly type findParameterList and countCallArguments (remove all `any`)
- Add arity-based call resolution integration tests for 9 languages
- Add unit regression tests for Interface/Enum CALLS refusal

* chore: remove C# build artifacts from fixtures

* feat: add call-form discrimination and ownerId to symbol table (Phase 1)

Add inferCallForm() and extractReceiverName() to distinguish free/member/constructor
calls at the AST level across all 9 languages. Add ownerId field to SymbolDefinition
linking Method/Constructor/Property to their owning class. Includes 36 unit tests
and member-call integration tests for all 9 languages (132 tests, 0 failures).

* feat: constructor/struct-literal resolution across all languages (Phase 2)

Add constructor discrimination to CALLS-edge resolution: new Foo(),
User{...} struct literals, and C# primary constructors now resolve to
Constructor/Class/Struct/Record nodes instead of being filtered out.

Queries: new_expression (C++), object_creation_expression (PHP),
composite_literal (Go), struct_expression (Rust), primary constructor
and implicit_object_creation_expression (C#).

Relaxes global tier in collectTieredCandidates to pass all candidates
through filterCallableCandidates, allowing kind/arity narrowing to
disambiguate at lower confidence.

* feat: receiver-constrained resolution with integration tests for all 9 languages

Add receiver-type filtering (Phase 3): when a member call like `user.save()`
has a known receiver type from TypeEnv, filter candidates by ownerId to
disambiguate methods with the same name across different classes.

Key changes:
- call-processor: build per-file TypeEnv, pass receiverTypeName to resolveCallTarget
- parse-worker: extract receiverTypeName from TypeEnv in worker thread
- resolveCallTarget: new step D filters by ownerId matching receiver type
- utils: extractReceiverName supports C++ field_expression (argument field)
- utils: findEnclosingClassId extracts Go method receiver types
- type-env: handle Go qualified_type, Kotlin user_type/variable_declaration
- parse-worker + parsing-processor: Function added to needsOwner for
  Kotlin/Rust/Python class methods captured as Function nodes

Integration tests added for receiver-constrained resolution across all 9
languages: TypeScript, Java, Python, Go, Rust, C++, C#, Kotlin, PHP.

* feat: NamedImportMap, scoped TypeEnv, broadened signatures + TS rest-param variadic fix

Address all 4 PR #238 review items:
1. Remove redundant lookupFuzzy in processRoutesFromExtracted
2. Add NamedImportMap for TS/Python symbol-level import tracking (Tier 2a)
3. Make TypeEnv scope-aware (Map<scopeKey, Map<varName, type>>) to fix
   non-deterministic receiver resolution across functions
4. Broaden extractMethodSignature: Go/Rust/C++ return types, variadic
   detection for Go/Java/Python/C++/Kotlin/TypeScript rest params

Discovered and fixed: TS rest params (...args) were not detected as
variadic — added rest_pattern detection inside required_parameter nodes.

Integration tests added: scoped receiver, named import disambiguation,
and variadic call resolution for both TypeScript and Python.

* fix: alias import resolution, Go multi-assign TypeEnv, dead code removal

- NamedImportMap now stores {sourcePath, exportedName} so aliased imports
  (import { User as U }) resolve U → User in the source file
- Named binding check moved before empty-allDefs early return in both
  call-processor and symbol-resolver, fixing constructor calls via aliases
- Go extractFromGoShortVarDeclaration iterates all LHS/RHS pairs for
  multi-assignment (user, repo := User{}, Repo{}) instead of only first
- Remove unused TYPED_DECLARATION_TYPES set (TYPED_PARAMETER_TYPES kept)
- Integration tests for both fixes (go-multi-assign, typescript-alias-imports)

* feat: alias import extraction for Kotlin, Rust, PHP, C# + integration tests

Add named import alias extraction to both pipeline paths
(import-processor.ts and parse-worker.ts) for Kotlin, Rust, PHP,
and C#. Add integration test fixtures and tests for all 5 languages
(Python alias extraction already worked, just needed the test).

Each test verifies: class detection, member call resolution through
aliases to correct target files, and IMPORTS edge emission.

* refactor: use SupportedLanguages enum everywhere instead of raw strings

Replace all raw language string literals and `language: string` types
with the SupportedLanguages enum across 10 files. This ensures
compile-time safety for language dispatch and eliminates dead
`language === 'tsx'` checks (tsx maps to TypeScript in the enum).

* fix: tier-ordering bug, re-export chains, PHP grouped imports, Java named imports

- Fix collectTieredCandidates tier-ordering: same-file now checked before
  named bindings, preventing imports from shadowing local definitions
  (matches resolveSymbolInternal priority order)
- Add re-export chain resolution for TypeScript/JavaScript barrel files:
  export { X } from './base' and export type { X } from './base' now
  followed up to 5 hops through NamedImportMap
- Fix PHP grouped import alias extraction: use App\Models\{User, Repo as R}
  now correctly handled in both parse-worker and import-processor
- Add Java NamedImportMap support: import com.example.models.User now
  records User as a named binding for precise disambiguation
- Add 16 new integration tests across TypeScript, PHP, and Java resolvers
  (220 total resolver tests, all passing)

* refactor: consolidate alias extraction + add variadic/constructor/shadow integration tests

- Extract shared named-binding-extraction.ts from duplicate logic in
  import-processor.ts and parse-worker.ts (net -200 lines)
- Deduplicate appendKotlinWildcard (now imported from resolvers/index.ts)
- Add integration tests: constructor calls (Kotlin, Python), variadic
  resolution (Go, Java, C#, C++, Kotlin), re-export chains (Python),
  local definition shadowing (Python, Go)
- Add TODO(stack-graph) for TypeEnv scope key collision
- 225 integration tests passing (was 223)

* fix: PHP non-aliased imports, Python node identity, re-export chain dedup + local-shadow tests

- PHP flat non-aliased imports (use App\Models\User) now stored in NamedImportMap
- PHP grouped non-aliased imports ({User} in {User, Repo as R}) now stored in NamedImportMap
- Python: replace non-public child.id with child.startIndex for node identity
- Extract shared walkBindingChain() from symbol-resolver and call-processor
- Add PHP variadic resolution fixture + test (variadic_parameter already covers PHP)
- Add local-shadow integration tests for Java, C#, Kotlin, Rust, PHP, C++ (6 languages)

* feat: Rust non-aliased use bindings, Kotlin non-aliased imports, re-export chain resolution

Extend NamedImportMap coverage for Rust and Kotlin non-aliased imports:

- Rust: rename collectUseAsClauses → collectRustBindings, extract terminal
  scoped_identifier (use crate::models::User) and identifier in use_list
  (use crate::models::{User, Repo}) into NamedImportMap. This also enables
  pub use re-export chain following via walkBindingChain.
- Kotlin: extend extractKotlinNamedBindings to handle non-aliased imports
  (import com.example.User), skipping wildcard imports.
- Add rust-reexport-chain fixture + 3 integration tests verifying Handler{}
  resolves through mod.rs pub use to handler.rs.
- Add Kotlin heritage + constructor-calls reason assertions for non-aliased
  import-resolved resolution.
- Add C# heritage test documenting namespace import tier behavior.

* fix: skip Kotlin lowercase member imports in NamedImportMap

Member imports like `import util.OneArg.writeAudit` (lowercase last
segment) must not populate NamedImportMap — same-named function imports
from different classes collide, breaking arity-based disambiguation.
Apply the same guard Java already uses: skip lowercase last segments.

* fix: skip spurious path-prefix bindings in Rust grouped imports

collectRustBindings was extracting the path segment (e.g. "models") from
`use crate::models::{User, Repo}` as a spurious NamedImportMap entry.
Skip scoped_identifier nodes that are direct children of scoped_use_list
since they are path prefixes, not importable symbols.

Adds rust-grouped-imports fixture and 4 integration tests verifying both
symbols resolve correctly and no spurious binding leaks through.

* fix: use startIndex in TypeEnv scope key to prevent same-name method collision

Two methods named identically in different classes within the same file
previously shared a scope key, causing non-deterministic type resolution.
Now keys use funcName@startIndex for uniqueness.

Also adds tests documenting destructuring assignment extraction gap.

* test: document C# namespace-level import limitation in named binding extraction

* test: document same-arity overload discrimination limitation in call processor

* perf: parallelize calls/heritage/routes processing in worker path

Worker path now runs processCallsFromExtracted, processHeritageFromExtracted,
and processRoutesFromExtracted via Promise.all instead of sequentially.
Safe because all three only read shared state and write via addRelationship's
dedup guard. Sequential fallback path stays sequential (shared LRU astCache).

Also fixes Rust collectRustBindings spurious path-prefix bindings for 3+ level
grouped imports, and adds @param JSDoc for walkBindingChain's allDefs invariant.

* docs: improve Promise.all safety comment and walkBindingChain JSDoc

Clarify that the parallelization safety comes from disjoint relationship
types + idempotent id-keyed Maps, not from lack of shared state (the
graph is shared). Strengthen allDefs JSDoc to describe silent-miss
consequence of passing pre-filtered results.

* refactor: extract language-specific processing into modular dispatch tables

Phase 1: Extract type binding logic from type-env.ts (635→125 LOC) into
type-extractors/ directory with per-language files and Record<SupportedLanguages,
LanguageTypeConfig> + satisfies dispatch.

Phase 2: Extract 5 config loaders from import-processor.ts into
language-config.ts (removed ~196 LOC of inline loaders).

Phase 3: Convert export-detection.ts switch/case to exhaustive
Record<SupportedLanguages, ExportChecker> + satisfies dispatch table,
fix node: any → SyntaxNode.

Also adds language feature matrix to README.

All 1146 unit tests and 433 integration tests pass.

* refactor: extract type binding logic into type-extractors/ directory (Phase 1)

Extract per-language type extraction from type-env.ts (635→125 LOC) into
type-extractors/ with Record<SupportedLanguages, LanguageTypeConfig> + satisfies
dispatch. 9 per-language files, shared helpers, and barrel index.

* refactor: extract config loaders to language-config.ts (Phase 2)

Move 5 language-specific config loaders and their type interfaces from
import-processor.ts into standalone language-config.ts module.
2026-03-13 13:12:23 +00:00
Zander Raycraft 03bfa3c4d9 FEAT: Added support for optional skill generation based on KuzuDB after initial repo analysis (npx gitnexus analyze --skills) (#171)
* calm fix 4 adding skills to repo [ISSUE #140]

* inspect

* unit and integration tests

* fixed hardcoded cohesion miss

* e2e tests for --skills flag for langauge/repo support

* Cohesion test e2e tests
2026-03-13 08:29:13 +00:00
Subham Kundu 74c0e462c3 Merge pull request #217 from JasonOA888/feat/issue-215-deepseek-model
feat(models): add DeepSeek model configurations
2026-03-13 00:31:02 +05:30
Gergő Magyar 7376e92063 fix: consolidate C/C++/C#/Rust language support from 6 overlapping PRs (#237)
* fix: consolidate C/C++/C#/Rust language support from 6 overlapping PRs

Merges fixes from PRs #163, #170, #178, #216, #227, #234 into a single
coherent changeset with shared modules and deduplication.

Phase 0 — Pre-merge consolidation:
- Extract isNodeExported to shared export-detection.ts module
- Extract TREE_SITTER_BUFFER_SIZE to shared constants.ts with adaptive sizing
- Consolidate FUNCTION_NODE_TYPES, extractFunctionName, isBuiltInOrNoise
  from duplicated call-processor.ts and parse-worker.ts into shared utils.ts
- Add query compilation smoke tests for all 12 languages

Language fixes:
- fix(c/cpp): isExported checks static linkage instead of returning false
- fix(c/cpp): .h files parsed as C++ (tree-sitter-cpp is superset of C)
- fix(c/cpp): expanded entry point patterns (~30 new for C, ~18 for C++)
- fix(cpp): add typedef, union, macro, prototype, inline method queries
- fix(c#): isExported scans sibling modifiers instead of parent walk
- fix(c#): heritage queries use correct base_list AST structure
- fix(c#): add framework detection, import resolution, entry point scoring
- fix(rust): isExported scans sibling visibility_modifier in declaration
- fix(builtins): remove open/read/write/close (real C POSIX syscalls)
- fix(buffer): adaptive bufferSize (2x fileSize, 512KB-32MB range)
- feat(ts/js): add call_expression query patterns for const assignments

Deduplication:
- call-processor.ts: -226 lines (uses shared utils)
- parse-worker.ts: -320 lines (uses shared utils)
- parsing-processor.ts: -156 lines (uses shared export-detection)

* perf: fix review findings — hoist Sets, deduplicate DEFINITION_CAPTURE_KEYS

- Hoist CSHARP_DECL_TYPES and RUST_DECL_TYPES to module-level constants
  in export-detection.ts (was allocating new Set on every isNodeExported call)
- Extract DEFINITION_CAPTURE_KEYS and getDefinitionNodeFromCaptures to
  shared utils.ts (was duplicated in parsing-processor.ts and parse-worker.ts)
- Pre-compute merged entry point patterns to avoid per-call array spread
  in calculateEntryPointScore

* test: add C, C++, and Tree-sitter buffer size tests

* fix: C/C++/Rust review findings + comprehensive test coverage (+72 tests)

Source fixes:
- Add Rust built-in noise (unwrap, clone, into, collect, panic, etc.)
- C++ anonymous namespace → internal linkage (not exported)
- Replace .text regex with storage_class_specifier child scan (perf)
- Raise file skip threshold from 512KB to 32MB (TREE_SITTER_MAX_BUFFER)
- Export TREE_SITTER_MAX_BUFFER from constants.ts
- Add C++ double pointer query patterns to CPP_QUERIES
- Add C#: record_struct, record_class, file_scoped_namespace to decl types
- Add Rust: union_item to visibility scanning set

Tests (214 → 286):
- ingestion-utils: +24 (Rust/C# noise, pointer/ref/destructor extraction, buffer)
- parsing: +36 (real AST C/C++ static/namespace, Rust/C#/Java/PHP/Swift edge cases)
- tree-sitter-languages: +12 (query accuracy for C/C++/C#/Rust captures)
2026-03-10 23:03:32 +00:00
RyanbaandGergo Magyar 1be910f54a fix: skip unavailable native Swift parsers in sequential ingestion (#188)
* fix: skip unavailable native Swift parsers in sequential ingestion

* fix: warn when ingestion skips languages in verbose mode

* test: cover verbose skip warnings

* docs: update analyze flags

* docs: clarify verbose default

---------

Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-03-10 14:26:36 +00:00
Gergo Magyar fa9ba8925c fix(ci): support fork PRs in Claude Code Review workflow
claude-code-action fetches branches by name from origin, which fails
for fork PRs since the branch only exists on the fork remote. Work
around by detecting fork PRs and temporarily pushing the branch to
origin before the action runs, then cleaning up afterwards.

Also changed trigger from automatic (every push) to on-demand only
(label "claude-review" or comment "@claude" / "/review").
2026-03-09 15:33:41 +00:00
8efc272609 fix(ci): move PR report to workflow_run for fork PR support (#225)
* Initial plan

* fix: add pull-requests write permissions to GitHub Actions workflows

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(ci): remove ineffective job-level permissions from reusable workflow

* fix(ci): pass PR write permission from caller to reusable unit-tests workflow

* fix(ci): harden CI/CD workflows with security fixes and reliability improvements

- Pin all actions to commit SHAs to prevent supply-chain attacks
- Fix shell injection in ci-integration.yml by using env vars instead of direct interpolation
- Scope permissions per-job in publish.yml (was granting pull-requests:write to publish job)
- Restrict claude-code-review to trusted contributors only (OWNER/MEMBER/COLLABORATOR)
- Switch claude-code-review to pull_request_target for fork PR support
- Fix fail-fast: false in ci-unit-tests.yml cross-platform matrix
- Remove duplicate ubuntu-latest from unit test matrix
- Add timeouts to all workflow jobs
- Improve kuzu-db test loop to continue on failure and report per-file errors

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ci): read thresholds from `vitest.config.ts`

* fix(ci): move PR report to workflow_run for fork PR support

The sticky-pull-request-comment and vitest-coverage-report-action
both fail on fork PRs because pull_request events receive a read-only
GITHUB_TOKEN. This extracts PR reporting into a separate ci-report.yml
workflow triggered by workflow_run, which always gets read/write tokens.

Changes:
- ci.yml: replace pr-report job with save-pr-meta artifact upload
- ci-unit-tests.yml: remove davelosert/vitest-coverage-report-action,
  add coverage-final.json to artifact for merging
- ci-integration.yml: add ubuntu coverage job for non-kuzu groups
- ci-report.yml (new): workflow_run handler that downloads artifacts,
  merges unit + integration coverage via Istanbul, and posts combined
  PR comment with sticky-pull-request-comment

* feat(ci): show unit, integration, and merged coverage in PR report

- Disable coverage thresholds for integration-only run (partial coverage)
- Display combined coverage as the primary metric
- Show per-suite breakdown (unit / integration) in expandable details
- Thresholds applied against combined coverage, not individual suites

* fix(ci): add coverage collection input for PR reports and validate job results

* fix(ci): refine Claude Code Review workflow to support issue comments and enhance trusted contributor checks

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 15:07:46 +00:00
c990d7e6c6 fix(ci): harden CI/CD workflows with security fixes and reliability improvements (#222)
* Initial plan

* fix: add pull-requests write permissions to GitHub Actions workflows

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(ci): remove ineffective job-level permissions from reusable workflow

* fix(ci): pass PR write permission from caller to reusable unit-tests workflow

* fix(ci): harden CI/CD workflows with security fixes and reliability improvements

- Pin all actions to commit SHAs to prevent supply-chain attacks
- Fix shell injection in ci-integration.yml by using env vars instead of direct interpolation
- Scope permissions per-job in publish.yml (was granting pull-requests:write to publish job)
- Restrict claude-code-review to trusted contributors only (OWNER/MEMBER/COLLABORATOR)
- Switch claude-code-review to pull_request_target for fork PR support
- Fix fail-fast: false in ci-unit-tests.yml cross-platform matrix
- Remove duplicate ubuntu-latest from unit test matrix
- Add timeouts to all workflow jobs
- Improve kuzu-db test loop to continue on failure and report per-file errors

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ci): read thresholds from `vitest.config.ts`

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 08:25:03 +00:00
Jason c2bd8667a3 feat(models): add DeepSeek model configurations
Add configurations for DeepSeek-V3 and DeepSeek-Chat models
via OpenRouter integration.

- DeepSeek-V3: Reasoning model (input: /usr/bin/bash.27, output: .10)
- DeepSeek-Chat: Chat model (input: /usr/bin/bash.14, output: /usr/bin/bash.28)

Fixes #215
2026-03-08 15:50:25 +08:00
Candido Sales Gomes 631fec371f Merge branch 'main' into add-support-ruby-rails 2026-03-02 17:52:50 -05:00
Candido Sales Gomes d9dafdf792 Merge branch 'main' into add-support-ruby-rails 2026-03-01 14:33:09 -05:00
Candido Sales Gomes 6d0a761910 Merge branch 'main' into add-support-ruby-rails 2026-02-28 09:36:47 -05:00
Candido Gomes edd3aec044 refactor(ruby): remove outdated Ruby call routing documentation 2026-02-27 20:43:56 -05:00
Candido Gomes d51fb35634 feat(ruby): implement Ruby support with enhanced call processing and import handling 2026-02-27 20:40:52 -05:00
Candido GomesandClaude Opus 4.6 520e3ecf24 feat(ruby): add Ruby language support for CLI and web
Add Ruby as the 11th supported language in GitNexus, enabling code
intelligence for Ruby codebases. This includes:

- Tree-sitter parsing for classes, modules, methods, and singleton methods
- Import extraction for require/require_relative via call post-processing
- Mixin heritage detection for include/extend/prepend
- attr_accessor/attr_reader/attr_writer property extraction
- Ruby-specific built-in filtering (Kernel methods + enumerables)
- Framework detection for lib/, bin/, exe/, and Rake patterns
- Entry point scoring for service objects, jobs, and CLI commands
- Test file detection for _spec.rb, _test.rb, and spec/ directories
- WASM grammar for web version

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 17:53:41 -05:00
420 changed files with 22395 additions and 2185 deletions
@@ -22,7 +22,7 @@ Run from the project root. This parses all source files, builds the knowledge gr
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale.
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook runs `analyze` automatically after `git commit` and `git merge`, preserving embeddings if previously generated.
### status — Check index freshness
+1 -1
View File
@@ -10,7 +10,7 @@ inputs:
runs:
using: composite
steps:
- uses: actions/setup-node@v4
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
cache: npm
+99 -8
View File
@@ -2,6 +2,12 @@ name: Integration Tests
on:
workflow_call:
inputs:
collect-coverage:
description: 'Whether to run the coverage collection job (only needed for PR reports)'
required: false
default: true
type: boolean
jobs:
# ── Integration test matrix ─────────────────────────────────────────
@@ -17,7 +23,7 @@ jobs:
# on Linux, and its C++ destructors segfault during
# process.exit(). Running each file in its own process lets
# the OS reclaim all resources cleanly.
# pipeline — 3 files: ingestion pipeline + csv, each creates own temp DB
# pipeline — 12 files: ingestion pipeline + csv + 9 resolver tests
# e2e — 2 files: child-process only (spawnSync), no in-process kuzu
# standalone — 4 files: pure logic, no kuzu, no child processes
test-matrix:
@@ -36,10 +42,20 @@ jobs:
test/integration/pipeline.test.ts
test/integration/csv-pipeline.test.ts
test/integration/parsing.test.ts
test/integration/resolvers/typescript.test.ts
test/integration/resolvers/csharp.test.ts
test/integration/resolvers/cpp.test.ts
test/integration/resolvers/java.test.ts
test/integration/resolvers/python.test.ts
test/integration/resolvers/rust.test.ts
test/integration/resolvers/go.test.ts
test/integration/resolvers/kotlin.test.ts
test/integration/resolvers/php.test.ts
- test-group: e2e
test-glob: >-
test/integration/cli-e2e.test.ts
test/integration/hooks-e2e.test.ts
test/integration/skills-e2e.test.ts
- test-group: standalone
test-glob: >-
test/integration/filesystem-walker.test.ts
@@ -47,9 +63,9 @@ jobs:
test/integration/tree-sitter-languages.test.ts
test/integration/worker-pool.test.ts
runs-on: ${{ matrix.os }}
timeout-minutes: 15
timeout-minutes: 25
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
with:
build: 'true'
@@ -72,30 +88,105 @@ jobs:
test/integration/search-pool.test.ts
test/integration/augmentation.test.ts
)
exit_code=0
for f in "${files[@]}"; do
echo "::group::$f"
npx vitest run --reporter=verbose --pool=forks "$f"
if ! npx vitest run --reporter=verbose --pool=forks "$f"; then
exit_code=1
echo "::error::Test file failed: $f"
fi
echo "::endgroup::"
done
exit $exit_code
# Non-kuzu groups: run all files in a single vitest invocation
- name: Run integration tests — ${{ matrix.test-group }}
if: matrix.test-group != 'kuzu-db'
run: npx vitest run --reporter=verbose ${{ matrix.test-glob }}
shell: bash
env:
TEST_GLOB: ${{ matrix.test-glob }}
run: npx vitest run --reporter=verbose $TEST_GLOB
working-directory: gitnexus
# ── Coverage collection (ubuntu only) ─────────────────────────────────
# Runs non-kuzu integration tests with coverage enabled so the PR report
# can merge integration + unit coverage for a combined view.
# kuzu-db tests are excluded because each file must run in its own vitest
# process (native addon isolation) which prevents single-run coverage merge.
coverage:
name: integration (ubuntu / coverage)
if: inputs.collect-coverage
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
with:
build: 'true'
- name: Run integration tests with coverage
working-directory: gitnexus
run: >-
npx vitest run
--reporter=default
--reporter=json
--outputFile=integration-results.json
--coverage
--coverage.reporter=json-summary
--coverage.reporter=json
--coverage.reporter=text
--coverage.thresholdAutoUpdate=false
--coverage.reportOnFailure=true
--coverage.thresholds.statements=0
--coverage.thresholds.branches=0
--coverage.thresholds.functions=0
--coverage.thresholds.lines=0
test/integration/pipeline.test.ts
test/integration/csv-pipeline.test.ts
test/integration/parsing.test.ts
test/integration/cli-e2e.test.ts
test/integration/hooks-e2e.test.ts
test/integration/filesystem-walker.test.ts
test/integration/enrichment.test.ts
test/integration/tree-sitter-languages.test.ts
test/integration/worker-pool.test.ts
test/integration/resolvers/typescript.test.ts
test/integration/resolvers/csharp.test.ts
test/integration/resolvers/cpp.test.ts
test/integration/resolvers/java.test.ts
test/integration/resolvers/python.test.ts
test/integration/resolvers/rust.test.ts
test/integration/resolvers/go.test.ts
test/integration/resolvers/kotlin.test.ts
test/integration/resolvers/php.test.ts
- name: Upload integration coverage
if: always()
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
name: integration-reports
path: |
gitnexus/coverage/coverage-summary.json
gitnexus/coverage/coverage-final.json
gitnexus/integration-results.json
retention-days: 5
# ── Unified status gate ──────────────────────────────────────────────
# Branch protection should require THIS job, not the matrix jobs directly.
# ci.yml's needs.integration.result aggregates through this gate.
status:
name: integration (all groups)
needs: test-matrix
if: always()
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- name: Check all matrix jobs passed
shell: bash
env:
RESULT: ${{ needs.test-matrix.result }}
run: |
result="${{ needs.test-matrix.result }}"
if [[ "$result" != "success" ]]; then
echo "::error::Integration matrix failed or cancelled: $result"
if [[ "$RESULT" != "success" ]]; then
echo "::error::Integration matrix failed or cancelled: $RESULT"
exit 1
fi
+2 -1
View File
@@ -6,8 +6,9 @@ on:
jobs:
typecheck:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
- run: npx tsc --noEmit
working-directory: gitnexus
+432
View File
@@ -0,0 +1,432 @@
name: CI Report
# Triggered after the CI workflow completes. Because workflow_run
# always runs code from the *default branch*, it receives a read/write
# GITHUB_TOKEN — even when the triggering PR comes from a fork.
on:
workflow_run:
workflows: ["CI"]
types: [completed]
permissions:
actions: read # needed to list/download workflow run artifacts
contents: read # needed for sparse checkout of vitest.config.ts
pull-requests: write # needed to post sticky PR comment
jobs:
pr-report:
name: PR Report
# Only run for pull-request CI runs
if: >-
github.event.workflow_run.event == 'pull_request' &&
github.event.workflow_run.conclusion != 'cancelled'
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
# ── Download artifacts from the CI run ────────────────────────
- name: Download artifacts
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
with:
script: |
const fs = require('fs');
const path = require('path');
const runId = context.payload.workflow_run.id;
const allArtifacts = await github.rest.actions.listWorkflowRunArtifacts({
owner: context.repo.owner,
repo: context.repo.repo,
run_id: runId,
});
async function downloadArtifact(name, dest) {
const match = allArtifacts.data.artifacts.find(a => a.name === name);
if (!match) {
core.warning(`Artifact "${name}" not found`);
return false;
}
const zip = await github.rest.actions.downloadArtifact({
owner: context.repo.owner,
repo: context.repo.repo,
artifact_id: match.id,
archive_format: 'zip',
});
fs.mkdirSync(dest, { recursive: true });
fs.writeFileSync(path.join(dest, `${name}.zip`), Buffer.from(zip.data));
return true;
}
const temp = process.env.RUNNER_TEMP;
await downloadArtifact('pr-meta', path.join(temp, 'dl'));
await downloadArtifact('test-reports', path.join(temp, 'dl'));
await downloadArtifact('integration-reports', path.join(temp, 'dl'));
- name: Extract artifacts
shell: bash
run: |
cd "$RUNNER_TEMP/dl"
# Extract each artifact into its own directory to avoid filename collisions
for z in *.zip; do
[ -f "$z" ] || continue
name="${z%.zip}"
mkdir -p "$RUNNER_TEMP/artifacts/$name"
unzip -o "$z" -d "$RUNNER_TEMP/artifacts/$name"
done
- name: Read PR metadata
id: meta
shell: bash
run: |
DIR="$RUNNER_TEMP/artifacts/pr-meta"
if [ ! -f "$DIR/pr_number" ]; then
echo "skip=true" >> "$GITHUB_OUTPUT"
echo "::warning::pr_number artifact missing — skipping report"
exit 0
fi
# Validate PR number is a positive integer (artifact comes from
# untrusted fork code, so treat contents defensively).
PR_NUM=$(cat "$DIR/pr_number" | tr -d '[:space:]')
if ! [[ "$PR_NUM" =~ ^[0-9]+$ ]]; then
echo "skip=true" >> "$GITHUB_OUTPUT"
echo "::error::Invalid PR number in artifact: '$PR_NUM'"
exit 0
fi
echo "skip=false" >> "$GITHUB_OUTPUT"
echo "pr_number=$PR_NUM" >> "$GITHUB_OUTPUT"
# Validate job-result strings against known GitHub Actions values.
# Artifact contents come from the PR workflow (potentially untrusted
# fork code), so we whitelist to prevent newline injection into
# GITHUB_OUTPUT.
validate_result() {
local val
val=$(cat "$1" | tr -d '[:space:]')
case "$val" in
success|failure|cancelled|skipped) echo "$val" ;;
*) echo "unknown" ;;
esac
}
echo "quality=$(validate_result "$DIR/quality_result")" >> "$GITHUB_OUTPUT"
echo "unit=$(validate_result "$DIR/unit_result")" >> "$GITHUB_OUTPUT"
echo "integration=$(validate_result "$DIR/integration_result")" >> "$GITHUB_OUTPUT"
- name: Checkout (for vitest config)
if: steps.meta.outputs.skip != 'true'
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
with:
sparse-checkout: gitnexus/vitest.config.ts
sparse-checkout-cone-mode: false
# ── Merge coverage from unit + integration ─────────────────────
- name: Setup Node.js
if: steps.meta.outputs.skip != 'true'
uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
- name: Install coverage merge tools
if: steps.meta.outputs.skip != 'true'
run: npm install --no-save istanbul-lib-coverage istanbul-lib-report istanbul-reports
- name: Merge coverage reports
if: steps.meta.outputs.skip != 'true'
id: coverage
shell: bash
run: |
DIR="$RUNNER_TEMP/artifacts"
UNIT_COV=$(find "$DIR/test-reports" -name "coverage-final.json" -type f 2>/dev/null | head -1)
INTEG_COV=$(find "$DIR/integration-reports" -name "coverage-final.json" -type f 2>/dev/null | head -1)
MERGED_DIR="$RUNNER_TEMP/merged-coverage"
mkdir -p "$MERGED_DIR"
if [ -n "$UNIT_COV" ] && [ -n "$INTEG_COV" ]; then
echo "has_merged=true" >> "$GITHUB_OUTPUT"
# Merge using Node.js + istanbul-lib-coverage.
# Paths are passed via env vars to avoid shell interpolation
# inside the script string.
UNIT_COV_PATH="$UNIT_COV" \
INTEG_COV_PATH="$INTEG_COV" \
MERGED_OUT_DIR="$MERGED_DIR" \
node -e "
const libCoverage = require('istanbul-lib-coverage');
const libReport = require('istanbul-lib-report');
const reports = require('istanbul-reports');
const fs = require('fs');
const map = libCoverage.createCoverageMap({});
map.merge(JSON.parse(fs.readFileSync(process.env.UNIT_COV_PATH, 'utf8')));
map.merge(JSON.parse(fs.readFileSync(process.env.INTEG_COV_PATH, 'utf8')));
const context = libReport.createContext({
coverageMap: map,
dir: process.env.MERGED_OUT_DIR,
});
reports.create('json-summary').execute(context);
console.log('Merged coverage written to ' + process.env.MERGED_OUT_DIR + '/coverage-summary.json');
"
elif [ -n "$UNIT_COV" ]; then
echo "has_merged=false" >> "$GITHUB_OUTPUT"
echo "::warning::Integration coverage not found — using unit coverage only"
else
echo "has_merged=false" >> "$GITHUB_OUTPUT"
echo "::warning::No coverage data found"
fi
- name: Build report
if: steps.meta.outputs.skip != 'true'
id: report
shell: bash
env:
QUALITY: ${{ steps.meta.outputs.quality }}
UNIT: ${{ steps.meta.outputs.unit }}
INTEG: ${{ steps.meta.outputs.integration }}
HAS_MERGED: ${{ steps.coverage.outputs.has_merged }}
RUN_URL: ${{ github.event.workflow_run.html_url }}
run: |
DIR="$RUNNER_TEMP/artifacts"
MERGED_DIR="$RUNNER_TEMP/merged-coverage"
# ── Helper: read coverage summary into prefixed vars ──
# Uses printf -v for safe variable assignment (no eval).
read_cov() {
local prefix=$1 file=$2
if [ -n "$file" ] && [ -f "$file" ]; then
local val
val=$(jq -r '.total.statements.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_STMTS" '%s' "$val"
val=$(jq -r '.total.branches.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_BRANCH" '%s' "$val"
val=$(jq -r '.total.functions.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_FUNCS" '%s' "$val"
val=$(jq -r '.total.lines.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_LINES" '%s' "$val"
val=$(jq -r '"\(.total.statements.covered)/\(.total.statements.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_STMTS_COV" '%s' "$val"
val=$(jq -r '"\(.total.branches.covered)/\(.total.branches.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_BRANCH_COV" '%s' "$val"
val=$(jq -r '"\(.total.functions.covered)/\(.total.functions.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_FUNCS_COV" '%s' "$val"
val=$(jq -r '"\(.total.lines.covered)/\(.total.lines.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_LINES_COV" '%s' "$val"
return 0
else
printf -v "${prefix}_STMTS" '%s' "N/A"
printf -v "${prefix}_BRANCH" '%s' "N/A"
printf -v "${prefix}_FUNCS" '%s' "N/A"
printf -v "${prefix}_LINES" '%s' "N/A"
printf -v "${prefix}_STMTS_COV" '%s' ""
printf -v "${prefix}_BRANCH_COV" '%s' ""
printf -v "${prefix}_FUNCS_COV" '%s' ""
printf -v "${prefix}_LINES_COV" '%s' ""
return 1
fi
}
# ── Read all three coverage reports ──
UNIT_SUMMARY=$(find "$DIR/test-reports" -name "coverage-summary.json" -type f 2>/dev/null | head -1)
INTEG_SUMMARY=$(find "$DIR/integration-reports" -name "coverage-summary.json" -type f 2>/dev/null | head -1)
MERGED_SUMMARY="$MERGED_DIR/coverage-summary.json"
read_cov "U" "$UNIT_SUMMARY"
HAS_UNIT=$?
read_cov "I" "$INTEG_SUMMARY"
HAS_INTEG=$?
read_cov "M" "$MERGED_SUMMARY"
# ── Locate test results (unit) ──
RESULTS_FILE=$(find "$DIR/test-reports" -name "test-results.json" -type f 2>/dev/null | head -1)
INTEG_RESULTS=$(find "$DIR/integration-reports" -name "integration-results.json" -type f 2>/dev/null | head -1)
if [ -n "$RESULTS_FILE" ]; then
U_TOTAL=$(jq -r '.numTotalTests' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_PASSED=$(jq -r '.numPassedTests' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_FAILED=$(jq -r '.numFailedTests' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_SKIPPED=$(jq -r '.numPendingTests' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_SUITES=$(jq -r '.numTotalTestSuites' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_DURATION=$(jq -r '((.testResults | map(.endTime) | max) - (.startTime)) / 1000 | floor' "$RESULTS_FILE" 2>/dev/null || echo 0)
else
U_TOTAL=0; U_PASSED=0; U_FAILED=0; U_SKIPPED=0; U_SUITES=0; U_DURATION=0
fi
if [ -n "$INTEG_RESULTS" ]; then
I_TOTAL=$(jq -r '.numTotalTests' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_PASSED=$(jq -r '.numPassedTests' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_FAILED=$(jq -r '.numFailedTests' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_SKIPPED=$(jq -r '.numPendingTests' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_SUITES=$(jq -r '.numTotalTestSuites' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_DURATION=$(jq -r '((.testResults | map(.endTime) | max) - (.startTime)) / 1000 | floor' "$INTEG_RESULTS" 2>/dev/null || echo 0)
else
I_TOTAL=0; I_PASSED=0; I_FAILED=0; I_SKIPPED=0; I_SUITES=0; I_DURATION=0
fi
# ── Sum test results ──
TOTAL=$((U_TOTAL + I_TOTAL))
PASSED=$((U_PASSED + I_PASSED))
FAILED=$((U_FAILED + I_FAILED))
SKIPPED=$((U_SKIPPED + I_SKIPPED))
SUITES=$((U_SUITES + I_SUITES))
DURATION=$((U_DURATION + I_DURATION))
# ── Coverage thresholds (read from vitest.config.ts) ──
if [ -f gitnexus/vitest.config.ts ]; then
THRESH_STMTS=$(grep -oP 'statements:\s*\K[0-9]+' gitnexus/vitest.config.ts || echo 0)
THRESH_BRANCH=$(grep -oP 'branches:\s*\K[0-9]+' gitnexus/vitest.config.ts || echo 0)
THRESH_FUNCS=$(grep -oP 'functions:\s*\K[0-9]+' gitnexus/vitest.config.ts || echo 0)
THRESH_LINES=$(grep -oP 'lines:\s*\K[0-9]+' gitnexus/vitest.config.ts || echo 0)
else
THRESH_STMTS=0; THRESH_BRANCH=0; THRESH_FUNCS=0; THRESH_LINES=0
fi
# ── Status helpers ──
status_icon() {
case "$1" in
success) echo "✅" ;;
failure) echo "❌" ;;
cancelled) echo "⏭️" ;;
*) echo "❓" ;;
esac
}
cov_bar() {
local pct=$1 thresh=$2
if [ "$pct" = "N/A" ]; then echo "—"; return; fi
local filled
filled=$(awk "BEGIN { printf \"%d\", $pct / 5 }")
(( filled < 0 )) && filled=0
(( filled > 20 )) && filled=20
local empty=$((20 - filled))
local bar=""
for ((i=0; i<filled; i++)); do bar+="█"; done
for ((i=0; i<empty; i++)); do bar+="░"; done
if [ "$(awk "BEGIN { print ($pct >= $thresh) ? 1 : 0 }")" = "1" ]; then
echo "🟢 ${bar}"
else
echo "🔴 ${bar}"
fi
}
# ── Overall status ──
if [[ "$QUALITY" == "success" && "$UNIT" == "success" && "$INTEG" == "success" ]]; then
OVERALL="✅ **All checks passed**"
else
OVERALL="❌ **Some checks failed**"
fi
# ── Build markdown ──
{
echo "body<<GITNEXUS_CI_REPORT_EOF_7f3a"
echo "## CI Report"
echo ""
echo "${OVERALL}"
echo ""
echo "### Pipeline Status"
echo ""
echo "| Stage | Status | Details |"
echo "|-------|--------|---------|"
echo "| $(status_icon "$QUALITY") Typecheck | \`${QUALITY}\` | tsc --noEmit |"
echo "| $(status_icon "$UNIT") Unit Tests | \`${UNIT}\` | 3 platforms |"
echo "| $(status_icon "$INTEG") Integration | \`${INTEG}\` | 3 OS x 4 groups = 12 jobs |"
echo ""
if [ "$TOTAL" -gt 0 ] 2>/dev/null; then
echo "### Test Results"
echo ""
if [ "$FAILED" = "0" ]; then
echo "✅ **${PASSED}** passed"
else
echo "❌ **${FAILED}** failed / **${PASSED}** passed"
fi
if [ "$SKIPPED" != "0" ]; then
echo " · ${SKIPPED} skipped"
fi
echo " · ${SUITES} suites · ${TOTAL} total"
echo " · ⏱️ ${DURATION}s"
if [ "$I_TOTAL" -gt 0 ] 2>/dev/null; then
echo " · 📊 ${U_TOTAL} unit + ${I_TOTAL} integration"
fi
echo ""
fi
# ── Coverage table helper ──
cov_table() {
local label=$1 s=$2 b=$3 f=$4 l=$5 sc=$6 bc=$7 fc=$8 lc=$9
shift 9
local ts=$1 tb=$2 tf=$3 tl=$4
echo "#### ${label}"
echo ""
echo "| Metric | Coverage | Covered | Threshold | Status |"
echo "|--------|----------|---------|-----------|--------|"
echo "| Statements | **${s}%** | ${sc} | ${ts}% | $(cov_bar "$s" "$ts") |"
echo "| Branches | **${b}%** | ${bc} | ${tb}% | $(cov_bar "$b" "$tb") |"
echo "| Functions | **${f}%** | ${fc} | ${tf}% | $(cov_bar "$f" "$tf") |"
echo "| Lines | **${l}%** | ${lc} | ${tl}% | $(cov_bar "$l" "$tl") |"
echo ""
}
if [ "$M_STMTS" != "N/A" ]; then
echo "### Code Coverage"
echo ""
cov_table "Combined (Unit + Integration)" \
"$M_STMTS" "$M_BRANCH" "$M_FUNCS" "$M_LINES" \
"$M_STMTS_COV" "$M_BRANCH_COV" "$M_FUNCS_COV" "$M_LINES_COV" \
"$THRESH_STMTS" "$THRESH_BRANCH" "$THRESH_FUNCS" "$THRESH_LINES"
echo "<details>"
echo "<summary>Coverage breakdown by test suite</summary>"
echo ""
if [ "$U_STMTS" != "N/A" ]; then
cov_table "Unit Tests" \
"$U_STMTS" "$U_BRANCH" "$U_FUNCS" "$U_LINES" \
"$U_STMTS_COV" "$U_BRANCH_COV" "$U_FUNCS_COV" "$U_LINES_COV" \
"$THRESH_STMTS" "$THRESH_BRANCH" "$THRESH_FUNCS" "$THRESH_LINES"
fi
if [ "$I_STMTS" != "N/A" ]; then
cov_table "Integration Tests" \
"$I_STMTS" "$I_BRANCH" "$I_FUNCS" "$I_LINES" \
"$I_STMTS_COV" "$I_BRANCH_COV" "$I_FUNCS_COV" "$I_LINES_COV" \
"$THRESH_STMTS" "$THRESH_BRANCH" "$THRESH_FUNCS" "$THRESH_LINES"
fi
echo "</details>"
echo ""
echo "<details>"
echo "<summary>Coverage thresholds are auto-ratcheted — they only go up</summary>"
echo ""
echo "Vitest \`thresholds.autoUpdate\` bumps the floor whenever local coverage exceeds it."
echo "CI enforces the current thresholds; developers commit the ratcheted values."
echo "</details>"
echo ""
elif [ "$U_STMTS" != "N/A" ]; then
echo "### Code Coverage (Unit only)"
echo ""
cov_table "Unit Tests" \
"$U_STMTS" "$U_BRANCH" "$U_FUNCS" "$U_LINES" \
"$U_STMTS_COV" "$U_BRANCH_COV" "$U_FUNCS_COV" "$U_LINES_COV" \
"$THRESH_STMTS" "$THRESH_BRANCH" "$THRESH_FUNCS" "$THRESH_LINES"
echo "<details>"
echo "<summary>Coverage thresholds are auto-ratcheted — they only go up</summary>"
echo ""
echo "Vitest \`thresholds.autoUpdate\` bumps the floor whenever local coverage exceeds it."
echo "CI enforces the current thresholds; developers commit the ratcheted values."
echo "</details>"
echo ""
else
echo "### Code Coverage"
echo ""
echo "⚠️ Coverage data unavailable - check the [unit test job](${RUN_URL}) for details."
echo ""
fi
echo "---"
echo "<sub>📋 [View full run](${RUN_URL}) · Generated by CI</sub>"
echo "GITNEXUS_CI_REPORT_EOF_7f3a"
} >> "$GITHUB_OUTPUT"
- name: Comment on PR
if: steps.meta.outputs.skip != 'true'
uses: marocchino/sticky-pull-request-comment@773744901bac0e8cbb5a0dc842800d45e9b2b405 # v2
with:
header: ci-report
number: ${{ steps.meta.outputs.pr_number }}
message: ${{ steps.report.outputs.body }}
+9 -11
View File
@@ -7,8 +7,9 @@ jobs:
unit-tests:
name: unit (ubuntu / coverage)
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
- name: Run unit tests with coverage
@@ -25,31 +26,28 @@ jobs:
--coverage.reportOnFailure=true
working-directory: gitnexus
- name: Coverage report
if: always()
uses: davelosert/vitest-coverage-report-action@v2
with:
working-directory: gitnexus
- name: Upload test reports
if: always()
uses: actions/upload-artifact@v4
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
name: test-reports
path: |
gitnexus/coverage/coverage-summary.json
gitnexus/coverage/coverage-final.json
gitnexus/test-results.json
retention-days: 5
cross-platform:
name: unit (${{ matrix.os }})
strategy:
fail-fast: true
fail-fast: false
matrix:
os: [ubuntu-latest, windows-latest, macos-latest]
# Ubuntu already covered by the coverage job above
os: [windows-latest, macos-latest]
runs-on: ${{ matrix.os }}
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
- run: npx vitest run test/unit
working-directory: gitnexus
+51 -161
View File
@@ -3,10 +3,16 @@ name: CI
on:
push:
branches: [main]
paths-ignore: ['**.md', 'docs/**', 'LICENSE']
pull_request:
branches: [main]
paths-ignore: ['**.md', 'docs/**', 'LICENSE']
workflow_call:
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
# ── Reusable workflow orchestration ─────────────────────────────────
# Each concern lives in its own workflow file for maintainability:
# ci-quality.yml — typecheck (tsc --noEmit)
@@ -18,175 +24,53 @@ on:
jobs:
quality:
uses: ./.github/workflows/ci-quality.yml
permissions:
contents: read
unit-tests:
uses: ./.github/workflows/ci-unit-tests.yml
permissions:
contents: read
integration:
uses: ./.github/workflows/ci-integration.yml
with:
collect-coverage: ${{ github.event_name == 'pull_request' }}
permissions:
contents: read
# ── PR test & coverage report ────────────────────────────────────
# Downloads coverage artifacts from unit tests and posts a summary
# comment on the PR with test results and coverage metrics.
pr-report:
name: PR Report
# ── Save PR metadata for the reporting workflow ─────────────────
# The ci-report.yml workflow (triggered by workflow_run) needs the
# PR number and job results to post a comment. We save them as an
# artifact because workflow_run context doesn't reliably carry PR
# info for fork PRs.
save-pr-meta:
name: Save PR Metadata
if: always() && github.event_name == 'pull_request'
needs: [quality, unit-tests, integration]
runs-on: ubuntu-latest
permissions:
pull-requests: write
timeout-minutes: 5
steps:
- name: Download test reports
uses: actions/download-artifact@v4
with:
name: test-reports
path: reports
continue-on-error: true
- name: Debug artifact contents
run: find reports -type f 2>/dev/null || echo "No reports directory"
continue-on-error: true
- name: Build report
id: report
- name: Write metadata
shell: bash
env:
PR_NUMBER: ${{ github.event.number }}
QUALITY: ${{ needs.quality.result }}
UNIT: ${{ needs.unit-tests.result }}
INTEG: ${{ needs.integration.result }}
run: |
# ── Locate coverage file (artifact path may vary) ──
COV_FILE=$(find reports -name "coverage-summary.json" -type f 2>/dev/null | head -1)
if [ -n "$COV_FILE" ]; then
STMTS=$(jq -r '.total.statements.pct' "$COV_FILE")
BRANCH=$(jq -r '.total.branches.pct' "$COV_FILE")
FUNCS=$(jq -r '.total.functions.pct' "$COV_FILE")
LINES=$(jq -r '.total.lines.pct' "$COV_FILE")
STMTS_COV=$(jq -r '"\(.total.statements.covered)/\(.total.statements.total)"' "$COV_FILE")
BRANCH_COV=$(jq -r '"\(.total.branches.covered)/\(.total.branches.total)"' "$COV_FILE")
FUNCS_COV=$(jq -r '"\(.total.functions.covered)/\(.total.functions.total)"' "$COV_FILE")
LINES_COV=$(jq -r '"\(.total.lines.covered)/\(.total.lines.total)"' "$COV_FILE")
else
STMTS="N/A"; BRANCH="N/A"; FUNCS="N/A"; LINES="N/A"
STMTS_COV=""; BRANCH_COV=""; FUNCS_COV=""; LINES_COV=""
fi
mkdir -p pr-meta
echo "$PR_NUMBER" > pr-meta/pr_number
echo "$QUALITY" > pr-meta/quality_result
echo "$UNIT" > pr-meta/unit_result
echo "$INTEG" > pr-meta/integration_result
# ── Locate test results ──
RESULTS_FILE=$(find reports -name "test-results.json" -type f 2>/dev/null | head -1)
if [ -n "$RESULTS_FILE" ]; then
TOTAL=$(jq -r '.numTotalTests' "$RESULTS_FILE")
PASSED=$(jq -r '.numPassedTests' "$RESULTS_FILE")
FAILED=$(jq -r '.numFailedTests' "$RESULTS_FILE")
SKIPPED=$(jq -r '.numPendingTests' "$RESULTS_FILE")
SUITES=$(jq -r '.numTotalTestSuites' "$RESULTS_FILE")
DURATION=$(jq -r '((.testResults | map(.endTime) | max) - (.startTime)) / 1000 | floor' "$RESULTS_FILE" 2>/dev/null || echo "N/A")
else
TOTAL="N/A"; PASSED="N/A"; FAILED="N/A"; SKIPPED="N/A"
SUITES="N/A"; DURATION="N/A"
fi
# ── Coverage thresholds (from vitest.config.ts P0 settings) ──
THRESH_STMTS=26; THRESH_BRANCH=23; THRESH_FUNCS=28; THRESH_LINES=27
# ── Status helpers ──
status_icon() {
case "$1" in
success) echo "✅" ;;
failure) echo "❌" ;;
cancelled) echo "⏭️" ;;
*) echo "❓" ;;
esac
}
cov_bar() {
local pct=$1 thresh=$2
if [ "$pct" = "N/A" ]; then echo "—"; return; fi
local filled=$(echo "$pct / 5" | bc 2>/dev/null || echo 0)
local empty=$((20 - filled))
local bar=""
for ((i=0; i<filled; i++)); do bar+="█"; done
for ((i=0; i<empty; i++)); do bar+="░"; done
if [ "$(echo "$pct >= $thresh" | bc 2>/dev/null)" = "1" ]; then
echo "🟢 ${bar}"
else
echo "🔴 ${bar}"
fi
}
QUALITY="${{ needs.quality.result }}"
UNIT="${{ needs.unit-tests.result }}"
INTEG="${{ needs.integration.result }}"
# ── Overall status ──
if [[ "$QUALITY" == "success" && "$UNIT" == "success" && "$INTEG" == "success" ]]; then
OVERALL="✅ **All checks passed**"
else
OVERALL="❌ **Some checks failed**"
fi
# ── Build markdown ──
{
echo "body<<REPORT_EOF"
echo "## CI Report"
echo ""
echo "${OVERALL}"
echo ""
echo "### Pipeline Status"
echo ""
echo "| Stage | Status | Details |"
echo "|-------|--------|---------|"
echo "| $(status_icon "$QUALITY") Typecheck | \`${QUALITY}\` | tsc --noEmit |"
echo "| $(status_icon "$UNIT") Unit Tests | \`${UNIT}\` | 3 platforms |"
echo "| $(status_icon "$INTEG") Integration | \`${INTEG}\` | 3 OS × 4 groups = 12 jobs |"
echo ""
if [ "$TOTAL" != "N/A" ]; then
echo "### Test Results"
echo ""
if [ "$FAILED" = "0" ]; then
echo "✅ **${PASSED}** passed"
else
echo "❌ **${FAILED}** failed / **${PASSED}** passed"
fi
if [ "$SKIPPED" != "0" ]; then
echo " · ${SKIPPED} skipped"
fi
echo " · ${SUITES} suites · ${TOTAL} total"
if [ "$DURATION" != "N/A" ]; then
echo " · ⏱️ ${DURATION}s"
fi
echo ""
fi
if [ "$STMTS" != "N/A" ]; then
echo "### Code Coverage"
echo ""
echo "| Metric | Coverage | Covered | Threshold | Status |"
echo "|--------|----------|---------|-----------|--------|"
echo "| Statements | **${STMTS}%** | ${STMTS_COV} | ${THRESH_STMTS}% | $(cov_bar "$STMTS" "$THRESH_STMTS") |"
echo "| Branches | **${BRANCH}%** | ${BRANCH_COV} | ${THRESH_BRANCH}% | $(cov_bar "$BRANCH" "$THRESH_BRANCH") |"
echo "| Functions | **${FUNCS}%** | ${FUNCS_COV} | ${THRESH_FUNCS}% | $(cov_bar "$FUNCS" "$THRESH_FUNCS") |"
echo "| Lines | **${LINES}%** | ${LINES_COV} | ${THRESH_LINES}% | $(cov_bar "$LINES" "$THRESH_LINES") |"
echo ""
echo "<details>"
echo "<summary>Coverage thresholds are auto-ratcheted — they only go up</summary>"
echo ""
echo "Vitest \`thresholds.autoUpdate\` bumps the floor whenever local coverage exceeds it."
echo "CI enforces the current thresholds; developers commit the ratcheted values."
echo "</details>"
echo ""
else
echo "### Code Coverage"
echo ""
echo "⚠️ Coverage data unavailable — check the [unit test job](${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}) for details."
echo ""
fi
echo "---"
echo "<sub>📋 [View full run](${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}) · Generated by CI</sub>"
echo "REPORT_EOF"
} >> "$GITHUB_OUTPUT"
- name: Comment on PR
uses: marocchino/sticky-pull-request-comment@v2
- name: Upload PR metadata
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
header: ci-report
message: ${{ steps.report.outputs.body }}
name: pr-meta
path: pr-meta/
retention-days: 1
# ── Unified CI gate ──────────────────────────────────────────────
# Single required check for branch protection.
@@ -195,15 +79,21 @@ jobs:
needs: [quality, unit-tests, integration]
if: always()
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- name: Check all jobs passed
shell: bash
env:
QUALITY: ${{ needs.quality.result }}
UNIT: ${{ needs.unit-tests.result }}
INTEG: ${{ needs.integration.result }}
run: |
echo "Quality: ${{ needs.quality.result }}"
echo "Unit Tests: ${{ needs.unit-tests.result }}"
echo "Integration: ${{ needs.integration.result }}"
if [[ "${{ needs.quality.result }}" != "success" ]] ||
[[ "${{ needs.unit-tests.result }}" != "success" ]] ||
[[ "${{ needs.integration.result }}" != "success" ]]; then
echo "Quality: $QUALITY"
echo "Unit Tests: $UNIT"
echo "Integration: $INTEG"
if [[ "$QUALITY" != "success" ]] ||
[[ "$UNIT" != "success" ]] ||
[[ "$INTEG" != "success" ]]; then
echo "::error::One or more CI jobs failed"
exit 1
fi
+75 -22
View File
@@ -1,44 +1,97 @@
name: Claude Code Review
# Uses pull_request_target so the workflow runs as defined on the default branch,
# which allows access to secrets for posting review comments on fork PRs.
# SECURITY: The checkout below uses the PR head SHA to review the correct code.
# The claude-code-action sandboxes execution — it does NOT run arbitrary code
# from the checked-out source.
on:
pull_request:
types: [opened, synchronize, ready_for_review, reopened]
# Optional: Only run on specific file changes
# paths:
# - "src/**/*.ts"
# - "src/**/*.tsx"
# - "src/**/*.js"
# - "src/**/*.jsx"
# Trigger only when explicitly requested:
# - Add the "claude-review" label to a PR, OR
# - Comment "@claude" or "/review" on a PR
pull_request_target:
types: [labeled]
issue_comment:
types: [created]
jobs:
claude-review:
# Optional: Filter by PR author
# if: |
# github.event.pull_request.user.login == 'external-contributor' ||
# github.event.pull_request.user.login == 'new-developer' ||
# github.event.pull_request.author_association == 'FIRST_TIME_CONTRIBUTOR'
# Run only when:
# 1. The "claude-review" label is added to a non-draft PR by a trusted contributor, OR
# 2. A trusted contributor comments "@claude" or "/review" on a PR
if: |
(
github.event_name == 'pull_request_target' &&
github.event.label.name == 'claude-review' &&
github.event.pull_request.draft == false &&
(github.event.pull_request.author_association == 'OWNER' ||
github.event.pull_request.author_association == 'MEMBER' ||
github.event.pull_request.author_association == 'COLLABORATOR')
) ||
(
github.event_name == 'issue_comment' &&
github.event.issue.pull_request &&
(contains(github.event.comment.body, '@claude') ||
contains(github.event.comment.body, '/review')) &&
(github.event.comment.author_association == 'OWNER' ||
github.event.comment.author_association == 'MEMBER' ||
github.event.comment.author_association == 'COLLABORATOR')
)
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
contents: read
pull-requests: read
contents: write # needed to push fork branch to origin
pull-requests: write
issues: read
id-token: write
steps:
- name: Checkout repository
uses: actions/checkout@v4
# For issue_comment triggers, resolve the PR number, head SHA, and branch name
- name: Resolve PR context
id: pr
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
with:
script: |
let pr;
if (context.eventName === 'issue_comment') {
const resp = await github.rest.pulls.get({
owner: context.repo.owner,
repo: context.repo.repo,
pull_number: context.payload.issue.number,
});
pr = resp.data;
} else {
pr = context.payload.pull_request;
}
core.setOutput('number', pr.number);
core.setOutput('sha', pr.head.sha);
core.setOutput('branch', pr.head.ref);
core.setOutput('is_fork', String(pr.head.repo.full_name !== pr.base.repo.full_name));
- name: Checkout PR head
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
with:
ref: ${{ steps.pr.outputs.sha }}
fetch-depth: 1
# claude-code-action fetches branches by name from origin, which fails
# for fork PRs. Work around by pushing the fork branch to origin so
# the action can find it. Cleaned up in the post step below.
- name: Push fork branch to origin
if: steps.pr.outputs.is_fork == 'true'
run: git push origin HEAD:refs/heads/${{ steps.pr.outputs.branch }}
- name: Run Claude Code Review
id: claude-review
uses: anthropics/claude-code-action@v1
uses: anthropics/claude-code-action@9469d113c6afd29550c402740f22d1a97dd1209b # v1
with:
claude_code_oauth_token: ${{ secrets.CLAUDE_CODE_OAUTH_TOKEN }}
plugin_marketplaces: 'https://github.com/anthropics/claude-code.git'
plugins: 'code-review@claude-code-plugins'
prompt: '/code-review:code-review ${{ github.repository }}/pull/${{ github.event.pull_request.number }}'
# See https://github.com/anthropics/claude-code-action/blob/main/docs/usage.md
# or https://code.claude.com/docs/en/cli-reference for available options
prompt: '/code-review:code-review ${{ github.repository }}/pull/${{ steps.pr.outputs.number }}'
# Clean up the temporary branch we pushed for fork PRs
- name: Delete fork branch from origin
if: always() && steps.pr.outputs.is_fork == 'true'
run: git push origin --delete refs/heads/${{ steps.pr.outputs.branch }} || true
+5 -13
View File
@@ -18,33 +18,25 @@ jobs:
(github.event_name == 'pull_request_review' && contains(github.event.review.body, '@claude')) ||
(github.event_name == 'issues' && (contains(github.event.issue.body, '@claude') || contains(github.event.issue.title, '@claude')))
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
contents: read
pull-requests: read
issues: read
pull-requests: write
issues: write
id-token: write
actions: read # Required for Claude to read CI results on PRs
steps:
- name: Checkout repository
uses: actions/checkout@v4
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
with:
fetch-depth: 1
- name: Run Claude Code
id: claude
uses: anthropics/claude-code-action@v1
uses: anthropics/claude-code-action@9469d113c6afd29550c402740f22d1a97dd1209b # v1
with:
claude_code_oauth_token: ${{ secrets.CLAUDE_CODE_OAUTH_TOKEN }}
# This is an optional setting that allows Claude to read CI results on PRs
additional_permissions: |
actions: read
# Optional: Give a custom prompt to Claude. If this is not specified, Claude will perform the instructions specified in the comment that tagged it.
# prompt: 'Update the pull request description to include a summary of changes.'
# Optional: Add claude_args to customize behavior and configuration
# See https://github.com/anthropics/claude-code-action/blob/main/docs/usage.md
# or https://code.claude.com/docs/en/cli-reference for available options
# claude_args: '--allowed-tools Bash(gh pr:*)'
+13 -7
View File
@@ -5,24 +5,25 @@ on:
tags:
- 'v*'
permissions:
contents: write
id-token: write
pull-requests: write
# No workflow-level permissions — scoped per job below.
jobs:
ci:
uses: ./.github/workflows/ci.yml
permissions:
contents: read
pull-requests: write
publish:
needs: ci
runs-on: ubuntu-latest
timeout-minutes: 15
permissions:
contents: write
id-token: write
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
registry-url: https://registry.npmjs.org
@@ -32,8 +33,13 @@ jobs:
working-directory: gitnexus
- name: Verify version consistency
shell: bash
run: |
TAG_VERSION="${GITHUB_REF#refs/tags/v}"
if ! [[ "$TAG_VERSION" =~ ^[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$ ]]; then
echo "::error::Tag does not follow semver: v$TAG_VERSION"
exit 1
fi
PKG_VERSION=$(node -p "require('./package.json').version")
if [ "$TAG_VERSION" != "$PKG_VERSION" ]; then
echo "::error::Tag version (v$TAG_VERSION) does not match package.json version ($PKG_VERSION)"
@@ -57,6 +63,6 @@ jobs:
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Create GitHub Release
uses: softprops/action-gh-release@v2
uses: softprops/action-gh-release@a06a81a03ee405af7f2048a818ed3f03bbf83c7b # v2
with:
generate_release_notes: true
+3
View File
@@ -48,6 +48,9 @@ coverage/
# Claude Code worktrees
.claude/worktrees/
# Claude code skills
.claude/skills/generated/
# Assets (screenshots, images)
assets/
+29 -4
View File
@@ -1,7 +1,7 @@
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (1650 symbols, 4291 relationships, 125 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (1747 symbols, 4569 relationships, 130 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
@@ -71,8 +71,33 @@ Before completing any code modification task, verify:
## CLI
- Re-index: `npx gitnexus analyze`
- Check freshness: `npx gitnexus status`
- Generate docs: `npx gitnexus wiki`
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (135 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Workers area (70 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Cli area (63 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Kuzu area (52 symbols) | `.claude/skills/generated/kuzu/SKILL.md` |
| Work in the Wiki area (52 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Embeddings area (48 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Components area (42 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Local area (36 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Storage area (36 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Services area (35 symbols) | `.claude/skills/generated/services/SKILL.md` |
| Work in the Mcp area (32 symbols) | `.claude/skills/generated/mcp/SKILL.md` |
| Work in the Llm area (30 symbols) | `.claude/skills/generated/llm/SKILL.md` |
| Work in the Eval area (18 symbols) | `.claude/skills/generated/eval/SKILL.md` |
| Work in the Bridge area (15 symbols) | `.claude/skills/generated/bridge/SKILL.md` |
| Work in the Hooks area (14 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Search area (11 symbols) | `.claude/skills/generated/search/SKILL.md` |
| Work in the Environments area (11 symbols) | `.claude/skills/generated/environments/SKILL.md` |
| Work in the Analysis area (10 symbols) | `.claude/skills/generated/analysis/SKILL.md` |
| Work in the Agents area (9 symbols) | `.claude/skills/generated/agents/SKILL.md` |
| Work in the Graph area (6 symbols) | `.claude/skills/generated/graph/SKILL.md` |
<!-- gitnexus:end -->
+36
View File
@@ -2,6 +2,42 @@
All notable changes to GitNexus will be documented in this file.
## [1.4.0] - 2026-03-13
### Added
- **Language-aware symbol resolution engine** with 3-tier resolver: exact FQN → scope-walk → guarded fuzzy fallback that refuses ambiguous matches (#238) — @magyargergo
- **Method Resolution Order (MRO)** with 5 language-specific strategies: C++ leftmost-base, C#/Java class-over-interface, Python C3 linearization, Rust qualified syntax, default BFS (#238) — @magyargergo
- **Constructor & struct literal resolution** across all languages — `new Foo()`, `User{...}`, C# primary constructors, target-typed new (#238) — @magyargergo
- **Receiver-constrained resolution** using per-file TypeEnv — disambiguates `user.save()` vs `repo.save()` via `ownerId` matching (#238) — @magyargergo
- **Heritage & ownership edges** — HAS_METHOD, OVERRIDES, Go struct embedding, Swift extension heritage, method signatures (`parameterCount`, `returnType`) (#238) — @magyargergo
- **Language-specific resolver directory** (`resolvers/`) — extracted JVM, Go, C#, PHP, Rust resolvers from monolithic import-processor (#238) — @magyargergo
- **Type extractor directory** (`type-extractors/`) — per-language type binding extraction with `Record<SupportedLanguages, Handler>` + `satisfies` dispatch (#238) — @magyargergo
- **Export detection dispatch table** — compile-time exhaustive `Record` + `satisfies` pattern replacing switch/if chains (#238) — @magyargergo
- **Language config module** (`language-config.ts`) — centralized tsconfig, go.mod, composer.json, .csproj, Swift package config loaders (#238) — @magyargergo
- **Optional skill generation** via `npx gitnexus analyze --skills` — generates AI agent skills from KuzuDB knowledge graph (#171) — @zander-raycraft
- **First-class C# support** — sibling-based modifier scanning, record/delegate/property/field/event declaration types (#163, #170, #178 via #237) — @Alice523, @benny-yamagata, @jnMetaCode
- **C/C++ support fixes** — `.h` → C++ mapping, static-linkage export detection, qualified/parenthesized declarators, 48 entry point patterns (#163, #227 via #237) — @Alice523, @bitgineer
- **Rust support fixes** — sibling-based `visibility_modifier` scanning for `pub` detection (#227 via #237) — @bitgineer
- **Adaptive tree-sitter buffer sizing** — `Math.min(Math.max(contentLength * 2, 512KB), 32MB)` (#216 via #237) — @JasonOA888
- **Call expression matching** in tree-sitter queries (#234 via #237) — @ex-nihilo-jg
- **DeepSeek model configurations** (#217) — @JasonOA888
- 282+ new unit tests, 178 integration resolver tests across 9 languages, 53 test files, 1146 total tests passing
### Fixed
- Skip unavailable native Swift parsers in sequential ingestion (#188) — @Gujiassh
- Heritage heuristic language-gated — no longer applies class/interface rules to wrong languages (#238) — @magyargergo
- C# `base_list` distinguishes EXTENDS vs IMPLEMENTS via symbol table + `I[A-Z]` heuristic (#238) — @magyargergo
- Go `qualified_type` (`models.User`) correctly unwrapped in TypeEnv (#238) — @magyargergo
- Global tier no longer blocks resolution when kind/arity filtering can narrow to 1 candidate (#238) — @magyargergo
### Changed
- `import-processor.ts` reduced from 1412 → 711 lines (50% reduction) via resolver and config extraction (#238) — @magyargergo
- `type-env.ts` reduced from 635 → ~125 lines via type-extractor extraction (#238) — @magyargergo
- CI/CD workflows hardened with security fixes and fork PR support (#222, #225) — @magyargergo
## [1.3.11] - 2026-03-08
### Security
+29 -4
View File
@@ -1,7 +1,7 @@
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (1650 symbols, 4291 relationships, 125 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (1747 symbols, 4569 relationships, 130 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
@@ -71,8 +71,33 @@ Before completing any code modification task, verify:
## CLI
- Re-index: `npx gitnexus analyze`
- Check freshness: `npx gitnexus status`
- Generate docs: `npx gitnexus wiki`
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (135 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Workers area (70 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Cli area (63 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Kuzu area (52 symbols) | `.claude/skills/generated/kuzu/SKILL.md` |
| Work in the Wiki area (52 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Embeddings area (48 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Components area (42 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Local area (36 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Storage area (36 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Services area (35 symbols) | `.claude/skills/generated/services/SKILL.md` |
| Work in the Mcp area (32 symbols) | `.claude/skills/generated/mcp/SKILL.md` |
| Work in the Llm area (30 symbols) | `.claude/skills/generated/llm/SKILL.md` |
| Work in the Eval area (18 symbols) | `.claude/skills/generated/eval/SKILL.md` |
| Work in the Bridge area (15 symbols) | `.claude/skills/generated/bridge/SKILL.md` |
| Work in the Hooks area (14 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Search area (11 symbols) | `.claude/skills/generated/search/SKILL.md` |
| Work in the Environments area (11 symbols) | `.claude/skills/generated/environments/SKILL.md` |
| Work in the Analysis area (10 symbols) | `.claude/skills/generated/analysis/SKILL.md` |
| Work in the Agents area (9 symbols) | `.claude/skills/generated/agents/SKILL.md` |
| Work in the Graph area (6 symbols) | `.claude/skills/generated/graph/SKILL.md` |
<!-- gitnexus:end -->
+8 -1
View File
@@ -135,7 +135,10 @@ claude mcp add gitnexus -- npx -y gitnexus@latest mcp
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
gitnexus serve # Start local HTTP server (multi-repo) for web UI connection
gitnexus list # List all indexed repositories
@@ -189,6 +192,10 @@ gitnexus wiki --base-url <url> # Wiki with custom LLM API base URL
- **Impact Analysis** — Analyze blast radius before changes
- **Refactoring** — Plan safe refactors using dependency mapping
**Repo-specific skills** generated with `--skills`:
When you run `gitnexus analyze --skills`, GitNexus detects the functional areas of your codebase (via Leiden community detection) and generates a `SKILL.md` file for each one under `.claude/skills/generated/`. Each skill describes a module's key files, entry points, execution flows, and cross-area connections — so your AI agent gets targeted context for the exact area of code you're working in. Skills are regenerated on each `--skills` run to stay current with the codebase.
---
## Multi-Repo MCP Architecture
@@ -320,7 +327,7 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
### Supported Languages
TypeScript, JavaScript, Python, Java, Kotlin, C, C++, C#, Go, Rust, PHP, Swift
TypeScript, JavaScript, Python, Java, Kotlin, C, C++, C#, Go, Ruby, Rust, PHP, Swift
---
+13
View File
@@ -0,0 +1,13 @@
model: deepseek-ai/deepseek-chat
provider: openrouter
cost:
input: 0.14 # per 1M tokens
output: 0.28 # per 1M tokens
# Native DeepSeek API (direct)
api_key: null
base_url: null
# For OpenRouter, uncomment below and comment out direct config above
# api_key: \${OPENROUTER_API_KEY}
# base_url: https://openrouter.ai/api/v1
+15
View File
@@ -0,0 +1,15 @@
model: deepseek-ai/DeepSeek-V3
provider: openrouter
cost:
input: 0.27 # per 1M tokens
output: 1.10 # per 1M tokens
# Native DeepSeek API (direct)
# Get your API key at: https://platform.deepseek.com/
# Or use OpenRouter with: OPENROUTER_API_KEY
api_key: null
base_url: null
# For OpenRouter, uncomment below and comment out direct config above
# api_key: \${OPENROUTER_API_KEY}
# base_url: https://openrouter.ai/api/v1
Binary file not shown.
@@ -5,6 +5,42 @@ import { vscDarkPlus } from 'react-syntax-highlighter/dist/esm/styles/prism';
import { useAppState } from '../hooks/useAppState';
import { NODE_COLORS } from '../lib/constants';
/** Map file extension to Prism syntax highlighter language identifier */
const getSyntaxLanguage = (filePath: string | undefined): string => {
if (!filePath) return 'text';
const ext = filePath.split('.').pop()?.toLowerCase();
switch (ext) {
case 'js': case 'jsx': case 'mjs': case 'cjs': return 'javascript';
case 'ts': case 'tsx': case 'mts': case 'cts': return 'typescript';
case 'py': case 'pyw': return 'python';
case 'rb': case 'rake': case 'gemspec': return 'ruby';
case 'java': return 'java';
case 'go': return 'go';
case 'rs': return 'rust';
case 'c': case 'h': return 'c';
case 'cpp': case 'cc': case 'cxx': case 'hpp': case 'hxx': case 'hh': return 'cpp';
case 'cs': return 'csharp';
case 'php': return 'php';
case 'kt': case 'kts': return 'kotlin';
case 'swift': return 'swift';
case 'json': return 'json';
case 'yaml': case 'yml': return 'yaml';
case 'md': case 'mdx': return 'markdown';
case 'html': case 'htm': case 'erb': return 'markup';
case 'css': case 'scss': case 'sass': return 'css';
case 'sh': case 'bash': case 'zsh': return 'bash';
case 'sql': return 'sql';
case 'xml': return 'xml';
default: break;
}
// Handle extensionless Ruby files
const basename = filePath.split('/').pop() || '';
if (['Rakefile', 'Gemfile', 'Guardfile', 'Vagrantfile', 'Brewfile'].includes(basename)) return 'ruby';
if (['Makefile'].includes(basename)) return 'makefile';
if (['Dockerfile'].includes(basename)) return 'docker';
return 'text';
};
// Match the code theme used elsewhere in the app
const customTheme = {
...vscDarkPlus,
@@ -267,12 +303,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
<div className="flex-1 min-h-0 overflow-auto scrollbar-thin">
{selectedFileContent ? (
<SyntaxHighlighter
language={
selectedFilePath?.endsWith('.py') ? 'python' :
selectedFilePath?.endsWith('.js') || selectedFilePath?.endsWith('.jsx') ? 'javascript' :
selectedFilePath?.endsWith('.ts') || selectedFilePath?.endsWith('.tsx') ? 'typescript' :
'text'
}
language={getSyntaxLanguage(selectedFilePath)}
style={customTheme as any}
showLineNumbers
startingLineNumber={1}
@@ -339,11 +370,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
const hasRange = typeof ref.startLine === 'number';
const startDisplay = hasRange ? (ref.startLine ?? 0) + 1 : undefined;
const endDisplay = hasRange ? (ref.endLine ?? ref.startLine ?? 0) + 1 : undefined;
const language =
ref.filePath.endsWith('.py') ? 'python' :
ref.filePath.endsWith('.js') || ref.filePath.endsWith('.jsx') ? 'javascript' :
ref.filePath.endsWith('.ts') || ref.filePath.endsWith('.tsx') ? 'typescript' :
'text';
const language = getSyntaxLanguage(ref.filePath);
const isGlowing = glowRefId === ref.id;
@@ -9,6 +9,6 @@ export enum SupportedLanguages {
Go = 'go',
Rust = 'rust',
PHP = 'php',
// Ruby = 'ruby',
Ruby = 'ruby',
Swift = 'swift',
}
+141 -34
View File
@@ -6,6 +6,8 @@ import { loadParser, loadLanguage } from '../tree-sitter/parser-loader';
import { LANGUAGE_QUERIES } from './tree-sitter-queries';
import { generateId } from '../../lib/utils';
import { getLanguageFromFilename } from './utils';
import { SupportedLanguages } from '../../config/supported-languages';
import { routeRubyCall } from './ruby-call-routing';
/**
* Node types that represent function/method definitions across languages.
@@ -35,6 +37,9 @@ const FUNCTION_NODE_TYPES = new Set([
// Rust
'function_item',
'impl_item', // Methods inside impl blocks
// Ruby
'method', // def foo
'singleton_method', // def self.foo
]);
/**
@@ -92,6 +97,18 @@ const findEnclosingFunction = (
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method'; // Treat constructors as methods for process detection
} else if (current.type === 'method') {
// Ruby instance method: def foo
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'singleton_method') {
// Ruby class method: def self.foo
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Function';
} else if (current.type === 'arrow_function' || current.type === 'function_expression') {
// Arrow/expression: const foo = () => {} - check parent variable declarator
const parent = current.parent;
@@ -184,6 +201,61 @@ export const processCalls = async (
const calledName = nameNode.text;
// Ruby: route special calls to heritage or properties (imports handled by import-processor)
if (language === SupportedLanguages.Ruby) {
const callNode = captureMap['call'];
const routed = routeRubyCall(calledName, callNode);
switch (routed.kind) {
case 'skip':
case 'import': // handled by import-processor
return;
case 'heritage':
for (const item of routed.items) {
const childId = symbolTable.lookupExact(file.path, item.enclosingClass) ||
symbolTable.lookupFuzzy(item.enclosingClass)[0]?.nodeId ||
generateId('Class', `${file.path}:${item.enclosingClass}`);
const parentId = symbolTable.lookupFuzzy(item.mixinName)[0]?.nodeId ||
generateId('Module', `${item.mixinName}`);
if (childId && parentId) {
const relId = generateId('IMPLEMENTS', `${childId}->${parentId}`);
graph.addRelationship({
id: relId, sourceId: childId, targetId: parentId,
type: 'IMPLEMENTS', confidence: 1.0, reason: 'trait-impl',
});
}
}
return;
case 'properties':
for (const item of routed.items) {
const nodeId = generateId('Property', `${file.path}:${item.propName}`);
graph.addNode({
id: nodeId,
label: 'Property' as any,
properties: {
name: item.propName, filePath: file.path,
startLine: item.startLine, endLine: item.endLine,
language: SupportedLanguages.Ruby, isExported: true,
description: item.accessorType,
},
});
symbolTable.add(file.path, item.propName, nodeId, 'Property');
const fileId = generateId('File', file.path);
const relId = generateId('DEFINES', `${fileId}->${nodeId}`);
graph.addRelationship({
id: relId, sourceId: fileId, targetId: nodeId,
type: 'DEFINES', confidence: 1.0, reason: '',
});
}
return;
case 'call':
break; // fall through to normal call processing below
}
}
// Skip common built-ins and noise
if (isBuiltInOrNoise(calledName)) return;
@@ -200,10 +272,10 @@ export const processCalls = async (
// 5. Find the enclosing function (caller)
const callNode = captureMap['call'];
const enclosingFuncId = findEnclosingFunction(callNode, file.path, symbolTable);
// Use enclosing function as source, fallback to file for top-level calls
const sourceId = enclosingFuncId || generateId('File', file.path);
const relId = generateId('CALLS', `${sourceId}:${calledName}->${resolved.nodeId}`);
graph.addRelationship({
@@ -711,37 +783,72 @@ const resolveCallTarget = (
* Filter out common built-in functions and noise
* that shouldn't be tracked as calls
*/
const isBuiltInOrNoise = (name: string): boolean => {
const builtIns = new Set([
// JavaScript/TypeScript built-ins
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export',
'fetch', 'Response', 'Request',
// React hooks and common functions
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
// Common array/object methods
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python built-ins
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'open', 'read', 'write', 'close', 'append', 'extend', 'update',
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
]);
/** Pre-built set (module-level singleton) to avoid re-creating per call */
const BUILT_IN_NAMES = new Set([
// JavaScript/TypeScript built-ins
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export',
'fetch', 'Response', 'Request',
// React hooks and common functions
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
// Common array/object methods
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python built-ins
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'open', 'read', 'write', 'close', 'append', 'extend', 'update',
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
// C/C++ standard library and common kernel helpers
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
'scanf', 'fscanf', 'sscanf',
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
'sizeof', 'offsetof', 'typeof',
'assert', 'abort', 'exit', '_exit',
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
// Linux kernel common macros/helpers (not real call targets)
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
'min', 'max', 'clamp', 'abs', 'swap',
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
'GFP_KERNEL', 'GFP_ATOMIC',
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
'mutex_lock', 'mutex_unlock', 'mutex_init',
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
'get', 'put',
// Ruby built-ins and Kernel methods
'puts', 'print', 'p', 'pp', 'warn', 'raise', 'fail',
'require', 'require_relative', 'load', 'autoload',
'include', 'extend', 'prepend',
'attr_accessor', 'attr_reader', 'attr_writer',
'public', 'private', 'protected', 'module_function',
'lambda', 'proc', 'block_given?',
'nil?', 'is_a?', 'kind_of?', 'instance_of?', 'respond_to?',
'freeze', 'frozen?', 'dup', 'clone', 'tap', 'then', 'yield_self',
// Ruby enumerables
'each', 'map', 'select', 'reject', 'find', 'detect', 'collect',
'inject', 'reduce', 'flat_map', 'each_with_object', 'each_with_index',
'any?', 'all?', 'none?', 'count', 'first', 'last',
'sort', 'sort_by', 'min', 'max', 'min_by', 'max_by',
'group_by', 'partition', 'zip', 'compact', 'flatten', 'uniq',
]);
return builtIns.has(name);
};
const isBuiltInOrNoise = (name: string): boolean => BUILT_IN_NAMES.has(name);
@@ -330,25 +330,20 @@ const calculateCohesion = (memberIds: string[], graph: Graph): number => {
const memberSet = new Set(memberIds);
let internalEdges = 0;
// Count edges within the community
let totalEdges = 0;
// Count internal vs total edges for community members
memberIds.forEach(nodeId => {
if (graph.hasNode(nodeId)) {
graph.forEachNeighbor(nodeId, neighbor => {
totalEdges++;
if (memberSet.has(neighbor)) {
internalEdges++;
}
});
}
});
// Each edge is counted twice (once from each end), so divide by 2
internalEdges = internalEdges / 2;
// Maximum possible internal edges for n nodes: n*(n-1)/2
const maxPossibleEdges = (memberIds.length * (memberIds.length - 1)) / 2;
if (maxPossibleEdges === 0) return 1.0;
return Math.min(1.0, internalEdges / maxPossibleEdges);
if (totalEdges === 0) return 1.0;
return Math.min(1.0, internalEdges / totalEdges);
};
@@ -13,7 +13,7 @@
import { detectFrameworkFromPath } from './framework-detection';
// ============================================================================
// NAME PATTERNS - All 9 supported languages
// NAME PATTERNS - All 11 supported languages
// ============================================================================
/**
@@ -143,6 +143,13 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
/^save$/, // Repository::save()
/^delete$/, // Repository::delete()
],
// Ruby
'ruby': [
/^call$/, // Service objects (MyService.call)
/^perform$/, // Background jobs (Sidekiq, ActiveJob)
/^execute$/, // Command pattern
],
};
// ============================================================================
@@ -302,7 +309,12 @@ export function isTestFile(filePath: string): boolean {
p.endsWith('test.php') ||
p.endsWith('spec.php') ||
p.includes('/tests/feature/') ||
p.includes('/tests/unit/')
p.includes('/tests/unit/') ||
// Ruby test patterns
p.endsWith('_spec.rb') ||
p.endsWith('_test.rb') ||
p.includes('/spec/') ||
p.includes('/test/fixtures/')
);
}
@@ -257,6 +257,17 @@ export function detectFrameworkFromPath(filePath: string): FrameworkHint | null
return { framework: 'laravel', entryPointMultiplier: 1.5, reason: 'laravel-repository' };
}
// ========== RUBY ==========
// Ruby: bin/ or exe/ (CLI entry points)
if ((p.includes('/bin/') || p.includes('/exe/')) && p.endsWith('.rb')) {
return { framework: 'ruby', entryPointMultiplier: 2.5, reason: 'ruby-executable' };
}
// Ruby: Rakefile or *.rake (task definitions)
if (p.endsWith('/rakefile') || p.endsWith('.rake')) {
return { framework: 'ruby', entryPointMultiplier: 1.5, reason: 'ruby-rake' };
}
// ========== SWIFT / iOS ==========
// iOS App entry points (highest priority)
@@ -53,7 +53,9 @@ const resolveImportPath = (
// Go
'.go',
// Rust
'.rs', '/mod.rs'
'.rs', '/mod.rs',
// Ruby
'.rb', '.rake',
];
if (importPath.startsWith('.')) {
@@ -220,6 +222,45 @@ export const processImports = async (
importMap.get(file.path)!.add(resolvedPath);
}
}
// ---- Ruby: require/require_relative come through @call, not @import ----
if (language === 'ruby' && captureMap['call']) {
const callNameNode = captureMap['call.name'];
if (callNameNode) {
const calledName = callNameNode.text;
if (calledName === 'require' || calledName === 'require_relative') {
const callNode = captureMap['call'];
const argList = callNode.childForFieldName?.('arguments');
const stringNode = argList?.children?.find((c: any) => c.type === 'string');
const contentNode = stringNode?.children?.find((c: any) => c.type === 'string_content');
if (contentNode) {
let importPath = contentNode.text;
// require_relative always resolves relative to current file
if (calledName === 'require_relative' && !importPath.startsWith('.')) {
importPath = './' + importPath;
}
totalImportsFound++;
const resolvedPath = resolveImportPath(
file.path, importPath, allFilePaths, allFileList, resolveCache
);
if (resolvedPath) {
const sourceId = generateId('File', file.path);
const targetId = generateId('File', resolvedPath);
const relId = generateId('IMPORTS', `${file.path}->${resolvedPath}`);
totalImportsResolved++;
graph.addRelationship({
id: relId, sourceId, targetId,
type: 'IMPORTS', confidence: 1.0, reason: '',
});
if (!importMap.has(file.path)) {
importMap.set(file.path, new Set());
}
importMap.get(file.path)!.add(resolvedPath);
}
}
}
}
}
});
// If re-parsed just for this, delete the tree to save memory
@@ -14,7 +14,7 @@ export type FileProgressCallback = (current: number, total: number, filePath: st
/**
* Check if a symbol (function, class, etc.) is exported/public
* Handles all 9 supported languages with explicit logic
* Handles all 11 supported languages with explicit logic
*
* @param node - The AST node for the symbol name
* @param name - The symbol name
@@ -104,7 +104,11 @@ const isNodeExported = (node: any, name: string, language: string): boolean => {
case 'c':
case 'cpp':
return false;
// Ruby: All top-level definitions are public by default
case 'ruby':
return true;
default:
return false;
}
@@ -0,0 +1,99 @@
/**
* Shared Ruby call routing logic.
*
* Ruby expresses imports, heritage (mixins), and property definitions as
* method calls rather than syntax-level constructs. This module provides a
* single routing function used by the CLI call-processor, CLI parse-worker,
* and the web call-processor so that the classification logic lives in one
* place.
*/
// ── Result types ────────────────────────────────────────────────────────────
export type RubyCallRouting =
| { kind: 'import'; importPath: string; isRelative: boolean }
| { kind: 'heritage'; items: RubyHeritageItem[] }
| { kind: 'properties'; items: RubyPropertyItem[] }
| { kind: 'call' }
| { kind: 'skip' };
export interface RubyHeritageItem {
enclosingClass: string;
mixinName: string;
}
export interface RubyPropertyItem {
propName: string;
accessorType: string;
startLine: number;
endLine: number;
}
// ── Routing function ────────────────────────────────────────────────────────
/**
* Classify a Ruby call node and extract its semantic payload.
*
* @param calledName - The method name (e.g. 'require', 'include', 'attr_accessor')
* @param callNode - The tree-sitter `call` AST node
* @returns A discriminated union describing the call's semantic role
*/
export function routeRubyCall(calledName: string, callNode: any): RubyCallRouting {
// ── require / require_relative → import ─────────────────────────────────
if (calledName === 'require' || calledName === 'require_relative') {
const argList = callNode.childForFieldName?.('arguments');
const stringNode = argList?.children?.find((c: any) => c.type === 'string');
const contentNode = stringNode?.children?.find((c: any) => c.type === 'string_content');
if (!contentNode) return { kind: 'skip' };
let importPath: string = contentNode.text;
const isRelative = calledName === 'require_relative';
if (isRelative && !importPath.startsWith('.')) {
importPath = './' + importPath;
}
return { kind: 'import', importPath, isRelative };
}
// ── include / extend / prepend → heritage (mixin) ──────────────────────
if (calledName === 'include' || calledName === 'extend' || calledName === 'prepend') {
let enclosingClass: string | null = null;
let current = callNode.parent;
while (current) {
if (current.type === 'class' || current.type === 'module') {
const nameNode = current.childForFieldName?.('name');
if (nameNode) { enclosingClass = nameNode.text; break; }
}
current = current.parent;
}
if (!enclosingClass) return { kind: 'skip' };
const items: RubyHeritageItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of (argList?.children ?? [])) {
if (arg.type === 'constant' || arg.type === 'scope_resolution') {
items.push({ enclosingClass, mixinName: arg.text });
}
}
return items.length > 0 ? { kind: 'heritage', items } : { kind: 'skip' };
}
// ── attr_accessor / attr_reader / attr_writer → property definitions ───
if (calledName === 'attr_accessor' || calledName === 'attr_reader' || calledName === 'attr_writer') {
const items: RubyPropertyItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of (argList?.children ?? [])) {
if (arg.type === 'simple_symbol') {
items.push({
propName: arg.text.replace(/^:/, ''),
accessorType: calledName,
startLine: arg.startPosition.row,
endLine: arg.endPosition.row,
});
}
}
return items.length > 0 ? { kind: 'properties', items } : { kind: 'skip' };
}
// ── Everything else → regular call ─────────────────────────────────────
return { kind: 'call' };
}
@@ -396,6 +396,40 @@ export const PHP_QUERIES = `
[(name) (qualified_name)] @heritage.trait))) @heritage
`;
// Ruby queries - works with tree-sitter-ruby
// NOTE: Ruby uses `call` for require, include, extend, prepend, attr_* etc.
// These are all captured as @call and routed in JS post-processing:
// - require/require_relative → import extraction
// - include/extend/prepend → heritage (mixin) extraction
// - attr_accessor/attr_reader/attr_writer → property definition extraction
// - everything else → regular call extraction
export const RUBY_QUERIES = `
; ── Modules ──────────────────────────────────────────────────────────────────
(module
name: (constant) @name) @definition.module
; ── Classes ──────────────────────────────────────────────────────────────────
(class
name: (constant) @name) @definition.class
; ── Instance methods ─────────────────────────────────────────────────────────
(method
name: (identifier) @name) @definition.method
; ── Singleton (class-level) methods ──────────────────────────────────────────
(singleton_method
name: (identifier) @name) @definition.function
; ── All calls (require, include, attr_*, and regular calls routed in JS) ─────
(call
method: (identifier) @call.name) @call
; ── Heritage: class < SuperClass ─────────────────────────────────────────────
(class
name: (constant) @heritage.class
superclass: (superclass
(constant) @heritage.extends)) @heritage`;
// Swift queries - works with tree-sitter-swift
export const SWIFT_QUERIES = `
; Classes
@@ -460,6 +494,7 @@ export const LANGUAGE_QUERIES: Record<SupportedLanguages, string> = {
[SupportedLanguages.CSharp]: CSHARP_QUERIES,
[SupportedLanguages.Rust]: RUST_QUERIES,
[SupportedLanguages.PHP]: PHP_QUERIES,
[SupportedLanguages.Ruby]: RUBY_QUERIES,
[SupportedLanguages.Swift]: SWIFT_QUERIES,
};
+12
View File
@@ -1,5 +1,8 @@
import { SupportedLanguages } from '../../config/supported-languages';
/** Ruby extensionless filenames recognised as Ruby source */
const RUBY_EXTENSIONLESS_FILES = new Set(['Rakefile', 'Gemfile', 'Guardfile', 'Vagrantfile', 'Brewfile']);
/**
* Map file extension to SupportedLanguage enum
*/
@@ -31,6 +34,15 @@ export const getLanguageFromFilename = (filename: string): SupportedLanguages |
filename.endsWith('.php5') || filename.endsWith('.php8')) {
return SupportedLanguages.PHP;
}
// Ruby (extensions)
if (filename.endsWith('.rb') || filename.endsWith('.rake') || filename.endsWith('.gemspec')) {
return SupportedLanguages.Ruby;
}
// Ruby (extensionless files)
const basename = filename.split('/').pop() || filename;
if (RUBY_EXTENSIONLESS_FILES.has(basename)) {
return SupportedLanguages.Ruby;
}
// Swift
if (filename.endsWith('.swift')) return SupportedLanguages.Swift;
return null;
@@ -40,6 +40,7 @@ const getWasmPath = (language: SupportedLanguages, filePath?: string): string =>
[SupportedLanguages.Go]: '/wasm/go/tree-sitter-go.wasm',
[SupportedLanguages.Rust]: '/wasm/rust/tree-sitter-rust.wasm',
[SupportedLanguages.PHP]: '/wasm/php/tree-sitter-php.wasm',
[SupportedLanguages.Ruby]: '/wasm/ruby/tree-sitter-ruby.wasm',
[SupportedLanguages.Swift]: '/wasm/swift/tree-sitter-swift.wasm',
};
+23 -2
View File
@@ -139,7 +139,8 @@ Your AI agent gets these tools automatically:
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
gitnexus serve # Start local HTTP server (multi-repo) for web UI
gitnexus list # List all indexed repositories
@@ -156,7 +157,27 @@ GitNexus supports indexing multiple repositories. Each `gitnexus analyze` regist
## Supported Languages
TypeScript, JavaScript, Python, Java, C, C++, C#, Go, Rust, PHP, Swift
TypeScript, JavaScript, Python, Java, C, C++, C#, Go, Rust, PHP, Kotlin, Swift, Ruby
### Language Feature Matrix
| Language | Imports | Types | Exports | Named Bindings | Config | Frameworks | Entry Points | Heritage |
|----------|---------|-------|---------|----------------|--------|------------|-------------|----------|
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ |
| Go | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — |
| Swift | — | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Ruby | ✓ | ✓ | ✓ | — | — | ✓ | ✓ | ✓ |
| C | — | ✓ | ✓ | — | — | ✓ | ✓ | ✓ |
| C++ | — | ✓ | ✓ | — | — | ✓ | ✓ | ✓ |
**Imports** — cross-file import resolution · **Types** — type annotation extraction · **Exports** — public/exported symbol detection · **Named Bindings** — `import { X }` tracking · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics · **Heritage** — class inheritance / interface implementation
## Agent Skills
+35 -4
View File
@@ -1,12 +1,12 @@
{
"name": "gitnexus",
"version": "1.3.11",
"version": "1.4.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "gitnexus",
"version": "1.3.11",
"version": "1.4.0",
"hasInstallScript": true,
"license": "PolyForm-Noncommercial-1.0.0",
"dependencies": {
@@ -31,9 +31,9 @@
"tree-sitter-go": "^0.21.0",
"tree-sitter-java": "^0.21.0",
"tree-sitter-javascript": "^0.21.0",
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-php": "^0.23.12",
"tree-sitter-python": "^0.21.0",
"tree-sitter-ruby": "^0.23.1",
"tree-sitter-rust": "^0.21.0",
"tree-sitter-typescript": "^0.21.0",
"uuid": "^13.0.0"
@@ -56,6 +56,7 @@
"node": ">=18.0.0"
},
"optionalDependencies": {
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-swift": "^0.6.0"
}
},
@@ -5415,6 +5416,7 @@
"integrity": "sha512-A4obq6bjzmYrA+F0JLLoheFPcofFkctNaZSpnDd+GPn1SfVZLY4/GG4C0cYVBTOShuPBGGAOPLM1JWLZQV4m1g==",
"hasInstallScript": true,
"license": "MIT",
"optional": true,
"dependencies": {
"node-addon-api": "^7.1.0",
"node-gyp-build": "^4.8.0"
@@ -5432,7 +5434,8 @@
"version": "7.1.1",
"resolved": "https://registry.npmjs.org/node-addon-api/-/node-addon-api-7.1.1.tgz",
"integrity": "sha512-5m3bsyrjFWE1xf7nz7YXdN4udnVtXK6/Yfgn5qnahL6bCkf2yKt4k3nuTKAtT4r3IG8JNR2ncsIMdZuAzJjHQQ==",
"license": "MIT"
"license": "MIT",
"optional": true
},
"node_modules/tree-sitter-php": {
"version": "0.23.12",
@@ -5487,6 +5490,34 @@
"integrity": "sha512-5m3bsyrjFWE1xf7nz7YXdN4udnVtXK6/Yfgn5qnahL6bCkf2yKt4k3nuTKAtT4r3IG8JNR2ncsIMdZuAzJjHQQ==",
"license": "MIT"
},
"node_modules/tree-sitter-ruby": {
"version": "0.23.1",
"resolved": "https://registry.npmjs.org/tree-sitter-ruby/-/tree-sitter-ruby-0.23.1.tgz",
"integrity": "sha512-d9/RXgWjR6HanN7wTYhS5bpBQLz1VkH048Vm3CodPGyJVnamXMGb8oEhDypVCBq4QnHui9sTXuJBBP3WtCw5RA==",
"hasInstallScript": true,
"license": "MIT",
"dependencies": {
"node-addon-api": "^8.2.2",
"node-gyp-build": "^4.8.2"
},
"peerDependencies": {
"tree-sitter": "^0.21.1"
},
"peerDependenciesMeta": {
"tree-sitter": {
"optional": true
}
}
},
"node_modules/tree-sitter-ruby/node_modules/node-addon-api": {
"version": "8.6.0",
"resolved": "https://registry.npmjs.org/node-addon-api/-/node-addon-api-8.6.0.tgz",
"integrity": "sha512-gBVjCaqDlRUk0EwoPNKzIr9KkS9041G/q31IBShPs1Xz6UTA+EXdZADbzqAJQrpDRq71CIMnOP5VMut3SL0z5Q==",
"license": "MIT",
"engines": {
"node": "^18 || ^20 || >= 21"
}
},
"node_modules/tree-sitter-rust": {
"version": "0.21.0",
"resolved": "https://registry.npmjs.org/tree-sitter-rust/-/tree-sitter-rust-0.21.0.tgz",
+3 -1
View File
@@ -1,6 +1,6 @@
{
"name": "gitnexus",
"version": "1.3.11",
"version": "1.4.0",
"description": "Graph-powered code intelligence for AI agents. Index any codebase, query via MCP or CLI.",
"author": "Abhigyan Patwari",
"license": "PolyForm-Noncommercial-1.0.0",
@@ -72,11 +72,13 @@
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-php": "^0.23.12",
"tree-sitter-python": "^0.21.0",
"tree-sitter-ruby": "^0.23.1",
"tree-sitter-rust": "^0.21.0",
"tree-sitter-typescript": "^0.21.0",
"uuid": "^13.0.0"
},
"optionalDependencies": {
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-swift": "^0.6.0"
},
"devDependencies": {
+21 -6
View File
@@ -9,6 +9,7 @@
import fs from 'fs/promises';
import path from 'path';
import { fileURLToPath } from 'url';
import { type GeneratedSkillInfo } from './skill-gen.js';
// ESM equivalent of __dirname
const __filename = fileURLToPath(import.meta.url);
@@ -37,7 +38,22 @@ const GITNEXUS_END_MARKER = '<!-- gitnexus:end -->';
* - Exact tool commands with parameters — vague directives get ignored
* - Self-review checklist — forces model to verify its own work
*/
function generateGitNexusContent(projectName: string, stats: RepoStats): string {
function generateGitNexusContent(projectName: string, stats: RepoStats, generatedSkills?: GeneratedSkillInfo[]): string {
const generatedRows = (generatedSkills && generatedSkills.length > 0)
? generatedSkills.map(s =>
`| Work in the ${s.label} area (${s.symbolCount} symbols) | \`.claude/skills/generated/${s.name}/SKILL.md\` |`
).join('\n')
: '';
const skillsTable = `| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | \`.claude/skills/gitnexus/gitnexus-exploring/SKILL.md\` |
| Blast radius / "What breaks if I change X?" | \`.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md\` |
| Trace bugs / "Why is X failing?" | \`.claude/skills/gitnexus/gitnexus-debugging/SKILL.md\` |
| Rename / extract / split / refactor | \`.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md\` |
| Tools, resources, schema reference | \`.claude/skills/gitnexus/gitnexus-guide/SKILL.md\` |
| Index, status, clean, wiki CLI commands | \`.claude/skills/gitnexus/gitnexus-cli/SKILL.md\` |${generatedRows ? '\n' + generatedRows : ''}`;
return `${GITNEXUS_START_MARKER}
# GitNexus — Code Intelligence
@@ -129,9 +145,7 @@ To check whether embeddings exist, inspect \`.gitnexus/meta.json\` — the \`sta
## CLI
- Re-index: \`npx gitnexus analyze\`
- Check freshness: \`npx gitnexus status\`
- Generate docs: \`npx gitnexus wiki\`
${skillsTable}
${GITNEXUS_END_MARKER}`;
}
@@ -270,9 +284,10 @@ export async function generateAIContextFiles(
repoPath: string,
_storagePath: string,
projectName: string,
stats: RepoStats
stats: RepoStats,
generatedSkills?: GeneratedSkillInfo[]
): Promise<{ files: string[] }> {
const content = generateGitNexusContent(projectName, stats);
const content = generateGitNexusContent(projectName, stats, generatedSkills);
const createdFiles: string[] = [];
// Create AGENTS.md (standard for Cursor, Windsurf, OpenCode, Cline, etc.)
+16 -2
View File
@@ -17,6 +17,7 @@ import { initKuzu, loadGraphToKuzu, getKuzuStats, executeQuery, executeWithReuse
import { getStoragePaths, saveMeta, loadMeta, addToGitignore, registerRepo, getGlobalRegistryPath } from '../storage/repo-manager.js';
import { getCurrentCommit, isGitRepo, getGitRoot } from '../storage/git.js';
import { generateAIContextFiles } from './ai-context.js';
import { generateSkillFiles, type GeneratedSkillInfo } from './skill-gen.js';
import fs from 'fs/promises';
@@ -45,6 +46,8 @@ function ensureHeap(): boolean {
export interface AnalyzeOptions {
force?: boolean;
embeddings?: boolean;
skills?: boolean;
verbose?: boolean;
}
/** Threshold: auto-skip embeddings for repos with more nodes than this */
@@ -72,6 +75,10 @@ export const analyzeCommand = async (
) => {
if (ensureHeap()) return;
if (options?.verbose) {
process.env.GITNEXUS_VERBOSE = '1';
}
console.log('\n GitNexus Analyzer\n');
let repoPath: string;
@@ -97,7 +104,7 @@ export const analyzeCommand = async (
const currentCommit = getCurrentCommit(repoPath);
const existingMeta = await loadMeta(storagePath);
if (existingMeta && !options?.force && existingMeta.lastCommit === currentCommit) {
if (existingMeta && !options?.force && !options?.skills && existingMeta.lastCommit === currentCommit) {
console.log(' Already up to date\n');
return;
}
@@ -311,6 +318,13 @@ export const analyzeCommand = async (
aggregatedClusterCount = Array.from(groups.values()).filter(count => count >= 5).length;
}
let generatedSkills: GeneratedSkillInfo[] = [];
if (options?.skills && pipelineResult.communityResult) {
updateBar(99, 'Generating skill files...');
const skillResult = await generateSkillFiles(repoPath, projectName, pipelineResult);
generatedSkills = skillResult.skills;
}
const aiContext = await generateAIContextFiles(repoPath, storagePath, projectName, {
files: pipelineResult.totalFileCount,
nodes: stats.nodes,
@@ -318,7 +332,7 @@ export const analyzeCommand = async (
communities: pipelineResult.communityResult?.stats.totalCommunities,
clusters: aggregatedClusterCount,
processes: pipelineResult.processResult?.stats.totalProcesses,
});
}, generatedSkills);
await closeKuzu();
// Note: we intentionally do NOT call disposeEmbedder() here.
+3 -1
View File
@@ -26,7 +26,9 @@ program
.description('Index a repository (full analysis)')
.option('-f, --force', 'Force full re-index even if up to date')
.option('--embeddings', 'Enable embedding generation for semantic search (off by default)')
.action(createLazyAction(() => import('./analyze.js'), 'analyzeCommand'));
.option('--skills', 'Generate repo-specific skill files from detected communities')
.option('-v, --verbose', 'Enable verbose ingestion warnings (default: false)')
.action(createLazyAction(() => import('./analyze.js'), 'analyzeCommand'));
program
.command('serve')
+712
View File
@@ -0,0 +1,712 @@
/**
* Skill File Generator
*
* Generates repo-specific SKILL.md files from detected Leiden communities.
* Each significant community becomes a skill that describes a functional area
* of the codebase, including key files, entry points, execution flows, and
* cross-community connections.
*/
import fs from 'fs/promises';
import path from 'path';
import { PipelineResult } from '../types/pipeline.js';
import { CommunityNode, CommunityMembership } from '../core/ingestion/community-processor.js';
import { ProcessNode } from '../core/ingestion/process-processor.js';
import { GraphNode, KnowledgeGraph } from '../core/graph/types.js';
// ============================================================================
// TYPES
// ============================================================================
export interface GeneratedSkillInfo {
name: string;
label: string;
symbolCount: number;
fileCount: number;
}
interface AggregatedCommunity {
label: string;
rawIds: string[];
symbolCount: number;
cohesion: number;
}
interface MemberSymbol {
id: string;
name: string;
label: string;
filePath: string;
startLine: number;
isExported: boolean;
}
interface FileInfo {
relativePath: string;
symbols: string[];
}
interface CrossConnection {
targetLabel: string;
count: number;
}
// ============================================================================
// MAIN EXPORT
// ============================================================================
/**
* @brief Generate repo-specific skill files from detected communities
* @param {string} repoPath - Absolute path to the repository root
* @param {string} projectName - Human-readable project name
* @param {PipelineResult} pipelineResult - In-memory pipeline data with communities, processes, graph
* @returns {Promise<{ skills: GeneratedSkillInfo[], outputPath: string }>} Generated skill metadata
*/
export const generateSkillFiles = async (
repoPath: string,
projectName: string,
pipelineResult: PipelineResult
): Promise<{ skills: GeneratedSkillInfo[]; outputPath: string }> => {
const { communityResult, processResult, graph } = pipelineResult;
const outputDir = path.join(repoPath, '.claude', 'skills', 'generated');
if (!communityResult || !communityResult.memberships.length) {
console.log('\n Skills: no communities detected, skipping skill generation');
return { skills: [], outputPath: outputDir };
}
console.log('\n Generating repo-specific skills...');
// Step 1: Build communities from memberships (not the filtered communities array).
// The community processor skips singletons from its communities array but memberships
// include ALL assignments. For repos with sparse CALLS edges, the communities array
// can be empty while memberships still has useful groupings.
const communities = communityResult.communities.length > 0
? communityResult.communities
: buildCommunitiesFromMemberships(communityResult.memberships, graph, repoPath);
const aggregated = aggregateCommunities(communities);
// Step 2: Filter to significant communities
// Keep communities with >= 3 symbols after aggregation.
const significant = aggregated
.filter(c => c.symbolCount >= 3)
.sort((a, b) => b.symbolCount - a.symbolCount)
.slice(0, 20);
if (significant.length === 0) {
console.log('\n Skills: no significant communities found (all below 3-symbol threshold)');
return { skills: [], outputPath: outputDir };
}
// Step 3: Build lookup maps
const membershipsByComm = buildMembershipMap(communityResult.memberships);
const nodeIdToCommunityLabel = buildNodeCommunityLabelMap(
communityResult.memberships,
communities
);
// Step 4: Clear and recreate output directory
try {
await fs.rm(outputDir, { recursive: true, force: true });
} catch { /* may not exist */ }
await fs.mkdir(outputDir, { recursive: true });
// Step 5: Generate skill files
const skills: GeneratedSkillInfo[] = [];
const usedNames = new Set<string>();
for (const community of significant) {
// Gather member symbols
const members = gatherMembers(community.rawIds, membershipsByComm, graph);
if (members.length === 0) continue;
// Gather file info
const files = gatherFiles(members, repoPath);
// Gather entry points
const entryPoints = gatherEntryPoints(members);
// Gather execution flows
const flows = gatherFlows(community.rawIds, processResult?.processes || []);
// Gather cross-community connections
const connections = gatherCrossConnections(
community.rawIds,
community.label,
membershipsByComm,
nodeIdToCommunityLabel,
graph
);
// Generate kebab name
const kebabName = toKebabName(community.label, usedNames);
usedNames.add(kebabName);
// Generate SKILL.md content
const content = renderSkillMarkdown(
community,
projectName,
members,
files,
entryPoints,
flows,
connections,
kebabName
);
// Write file
const skillDir = path.join(outputDir, kebabName);
await fs.mkdir(skillDir, { recursive: true });
await fs.writeFile(path.join(skillDir, 'SKILL.md'), content, 'utf-8');
const info: GeneratedSkillInfo = {
name: kebabName,
label: community.label,
symbolCount: community.symbolCount,
fileCount: files.length,
};
skills.push(info);
console.log(` \u2713 ${community.label} (${community.symbolCount} symbols, ${files.length} files)`);
}
console.log(`\n ${skills.length} skills generated \u2192 .claude/skills/generated/`);
return { skills, outputPath: outputDir };
};
// ============================================================================
// FALLBACK COMMUNITY BUILDER
// ============================================================================
/**
* @brief Build CommunityNode-like objects from raw memberships when the community
* processor's communities array is empty (all singletons were filtered out)
* @param {CommunityMembership[]} memberships - All node-to-community assignments
* @param {KnowledgeGraph} graph - The knowledge graph for resolving node metadata
* @param {string} repoPath - Repository root for path normalization
* @returns {CommunityNode[]} Synthetic community nodes built from membership data
*/
const buildCommunitiesFromMemberships = (
memberships: CommunityMembership[],
graph: KnowledgeGraph,
repoPath: string
): CommunityNode[] => {
// Group memberships by communityId
const groups = new Map<string, string[]>();
for (const m of memberships) {
const arr = groups.get(m.communityId);
if (arr) {
arr.push(m.nodeId);
} else {
groups.set(m.communityId, [m.nodeId]);
}
}
const communities: CommunityNode[] = [];
for (const [commId, nodeIds] of groups) {
// Derive a heuristic label from the most common parent directory
const folderCounts = new Map<string, number>();
for (const nodeId of nodeIds) {
const node = graph.getNode(nodeId);
if (!node?.properties.filePath) continue;
const normalized = node.properties.filePath.replace(/\\/g, '/');
const parts = normalized.split('/').filter(Boolean);
if (parts.length >= 2) {
const folder = parts[parts.length - 2];
if (!['src', 'lib', 'core', 'utils', 'common', 'shared', 'helpers'].includes(folder.toLowerCase())) {
folderCounts.set(folder, (folderCounts.get(folder) || 0) + 1);
}
}
}
let bestFolder = '';
let bestCount = 0;
for (const [folder, count] of folderCounts) {
if (count > bestCount) {
bestCount = count;
bestFolder = folder;
}
}
const label = bestFolder
? bestFolder.charAt(0).toUpperCase() + bestFolder.slice(1)
: `Cluster_${commId.replace('comm_', '')}`;
// Compute cohesion as internal-edge ratio (matches backend calculateCohesion).
// For each member node, count edges that stay inside the community vs total.
const nodeSet = new Set(nodeIds);
let internalEdges = 0;
let totalEdges = 0;
graph.forEachRelationship(rel => {
if (nodeSet.has(rel.sourceId)) {
totalEdges++;
if (nodeSet.has(rel.targetId)) internalEdges++;
}
});
const cohesion = totalEdges > 0 ? Math.min(1.0, internalEdges / totalEdges) : 1.0;
communities.push({
id: commId,
label,
heuristicLabel: label,
cohesion,
symbolCount: nodeIds.length,
});
}
return communities.sort((a, b) => b.symbolCount - a.symbolCount);
};
// ============================================================================
// AGGREGATION
// ============================================================================
/**
* @brief Aggregate raw Leiden communities by heuristicLabel
* @param {CommunityNode[]} communities - Raw community nodes from Leiden detection
* @returns {AggregatedCommunity[]} Aggregated communities grouped by label
*/
const aggregateCommunities = (communities: CommunityNode[]): AggregatedCommunity[] => {
const groups = new Map<string, {
rawIds: string[];
totalSymbols: number;
weightedCohesion: number;
}>();
for (const c of communities) {
const label = c.heuristicLabel || c.label || 'Unknown';
const symbols = c.symbolCount || 0;
const cohesion = c.cohesion || 0;
const existing = groups.get(label);
if (!existing) {
groups.set(label, {
rawIds: [c.id],
totalSymbols: symbols,
weightedCohesion: cohesion * symbols,
});
} else {
existing.rawIds.push(c.id);
existing.totalSymbols += symbols;
existing.weightedCohesion += cohesion * symbols;
}
}
return Array.from(groups.entries()).map(([label, g]) => ({
label,
rawIds: g.rawIds,
symbolCount: g.totalSymbols,
cohesion: g.totalSymbols > 0 ? g.weightedCohesion / g.totalSymbols : 0,
}));
};
// ============================================================================
// LOOKUP MAP BUILDERS
// ============================================================================
/**
* @brief Build a map from communityId to member nodeIds
* @param {CommunityMembership[]} memberships - All membership records
* @returns {Map<string, string[]>} Map of communityId -> nodeId[]
*/
const buildMembershipMap = (memberships: CommunityMembership[]): Map<string, string[]> => {
const map = new Map<string, string[]>();
for (const m of memberships) {
const arr = map.get(m.communityId);
if (arr) {
arr.push(m.nodeId);
} else {
map.set(m.communityId, [m.nodeId]);
}
}
return map;
};
/**
* @brief Build a map from nodeId to aggregated community label
* @param {CommunityMembership[]} memberships - All membership records
* @param {CommunityNode[]} communities - Community nodes with labels
* @returns {Map<string, string>} Map of nodeId -> community label
*/
const buildNodeCommunityLabelMap = (
memberships: CommunityMembership[],
communities: CommunityNode[]
): Map<string, string> => {
const commIdToLabel = new Map<string, string>();
for (const c of communities) {
commIdToLabel.set(c.id, c.heuristicLabel || c.label || 'Unknown');
}
const map = new Map<string, string>();
for (const m of memberships) {
const label = commIdToLabel.get(m.communityId);
if (label) {
map.set(m.nodeId, label);
}
}
return map;
};
// ============================================================================
// DATA GATHERING
// ============================================================================
/**
* @brief Gather member symbols for an aggregated community
* @param {string[]} rawIds - Raw community IDs belonging to this aggregated community
* @param {Map<string, string[]>} membershipsByComm - communityId -> nodeIds
* @param {KnowledgeGraph} graph - The knowledge graph
* @returns {MemberSymbol[]} Array of member symbol information
*/
const gatherMembers = (
rawIds: string[],
membershipsByComm: Map<string, string[]>,
graph: KnowledgeGraph
): MemberSymbol[] => {
const seen = new Set<string>();
const members: MemberSymbol[] = [];
for (const commId of rawIds) {
const nodeIds = membershipsByComm.get(commId) || [];
for (const nodeId of nodeIds) {
if (seen.has(nodeId)) continue;
seen.add(nodeId);
const node = graph.getNode(nodeId);
if (!node) continue;
members.push({
id: node.id,
name: node.properties.name,
label: node.label,
filePath: node.properties.filePath || '',
startLine: node.properties.startLine || 0,
isExported: node.properties.isExported === true,
});
}
}
return members;
};
/**
* @brief Gather deduplicated file info with per-file symbol names
* @param {MemberSymbol[]} members - Member symbols
* @param {string} repoPath - Repository root for relative path computation
* @returns {FileInfo[]} Sorted by symbol count descending
*/
const gatherFiles = (members: MemberSymbol[], repoPath: string): FileInfo[] => {
const fileMap = new Map<string, string[]>();
for (const m of members) {
if (!m.filePath) continue;
const rel = toRelativePath(m.filePath, repoPath);
const arr = fileMap.get(rel);
if (arr) {
arr.push(m.name);
} else {
fileMap.set(rel, [m.name]);
}
}
return Array.from(fileMap.entries())
.map(([relativePath, symbols]) => ({ relativePath, symbols }))
.sort((a, b) => b.symbols.length - a.symbols.length);
};
/**
* @brief Gather exported entry points prioritized by type
* @param {MemberSymbol[]} members - Member symbols
* @returns {MemberSymbol[]} Exported symbols sorted by type priority
*/
const gatherEntryPoints = (members: MemberSymbol[]): MemberSymbol[] => {
const typePriority: Record<string, number> = {
Function: 0,
Class: 1,
Method: 2,
Interface: 3,
};
return members
.filter(m => m.isExported)
.sort((a, b) => {
const pa = typePriority[a.label] ?? 99;
const pb = typePriority[b.label] ?? 99;
return pa - pb;
});
};
/**
* @brief Gather execution flows touching this community
* @param {string[]} rawIds - Raw community IDs for this aggregated community
* @param {ProcessNode[]} processes - All detected processes
* @returns {ProcessNode[]} Processes whose communities intersect rawIds, sorted by stepCount
*/
const gatherFlows = (rawIds: string[], processes: ProcessNode[]): ProcessNode[] => {
const rawIdSet = new Set(rawIds);
return processes
.filter(proc => proc.communities.some(cid => rawIdSet.has(cid)))
.sort((a, b) => b.stepCount - a.stepCount);
};
/**
* @brief Gather cross-community call connections
* @param {string[]} rawIds - Raw community IDs for this aggregated community
* @param {string} ownLabel - This community's aggregated label
* @param {Map<string, string[]>} membershipsByComm - communityId -> nodeIds
* @param {Map<string, string>} nodeIdToCommunityLabel - nodeId -> community label
* @param {KnowledgeGraph} graph - The knowledge graph
* @returns {CrossConnection[]} Aggregated cross-community connections sorted by count
*/
const gatherCrossConnections = (
rawIds: string[],
ownLabel: string,
membershipsByComm: Map<string, string[]>,
nodeIdToCommunityLabel: Map<string, string>,
graph: KnowledgeGraph
): CrossConnection[] => {
// Collect all node IDs in this aggregated community
const ownNodeIds = new Set<string>();
for (const commId of rawIds) {
const nodeIds = membershipsByComm.get(commId) || [];
for (const nid of nodeIds) {
ownNodeIds.add(nid);
}
}
// Count outgoing CALLS to nodes in different communities
const targetCounts = new Map<string, number>();
graph.forEachRelationship(rel => {
if (rel.type !== 'CALLS') return;
if (!ownNodeIds.has(rel.sourceId)) return;
if (ownNodeIds.has(rel.targetId)) return; // same community
const targetLabel = nodeIdToCommunityLabel.get(rel.targetId);
if (!targetLabel || targetLabel === ownLabel) return;
targetCounts.set(targetLabel, (targetCounts.get(targetLabel) || 0) + 1);
});
return Array.from(targetCounts.entries())
.map(([targetLabel, count]) => ({ targetLabel, count }))
.sort((a, b) => b.count - a.count);
};
// ============================================================================
// MARKDOWN RENDERING
// ============================================================================
/**
* @brief Render SKILL.md content for a single community
* @param {AggregatedCommunity} community - The aggregated community data
* @param {string} projectName - Project name for the description
* @param {MemberSymbol[]} members - All member symbols
* @param {FileInfo[]} files - File info with symbol names
* @param {MemberSymbol[]} entryPoints - Exported entry point symbols
* @param {ProcessNode[]} flows - Execution flows touching this community
* @param {CrossConnection[]} connections - Cross-community connections
* @param {string} kebabName - Kebab-case name for the skill
* @returns {string} Full SKILL.md content
*/
const renderSkillMarkdown = (
community: AggregatedCommunity,
projectName: string,
members: MemberSymbol[],
files: FileInfo[],
entryPoints: MemberSymbol[],
flows: ProcessNode[],
connections: CrossConnection[],
kebabName: string
): string => {
const cohesionPct = Math.round(community.cohesion * 100);
// Dominant directory: most common top-level directory
const dominantDir = getDominantDirectory(files);
// Top symbol names for "When to Use"
const topNames = entryPoints.slice(0, 3).map(e => e.name);
if (topNames.length === 0) {
// Fallback to any members
topNames.push(...members.slice(0, 3).map(m => m.name));
}
const lines: string[] = [];
// Frontmatter
lines.push('---');
lines.push(`name: ${kebabName}`);
lines.push(`description: "Skill for the ${community.label} area of ${projectName}. ${community.symbolCount} symbols across ${files.length} files."`);
lines.push('---');
lines.push('');
// Title
lines.push(`# ${community.label}`);
lines.push('');
lines.push(`${community.symbolCount} symbols | ${files.length} files | Cohesion: ${cohesionPct}%`);
lines.push('');
// When to Use
lines.push('## When to Use');
lines.push('');
if (dominantDir) {
lines.push(`- Working with code in \`${dominantDir}/\``);
}
if (topNames.length > 0) {
lines.push(`- Understanding how ${topNames.join(', ')} work`);
}
lines.push(`- Modifying ${community.label.toLowerCase()}-related functionality`);
lines.push('');
// Key Files (top 10)
lines.push('## Key Files');
lines.push('');
lines.push('| File | Symbols |');
lines.push('|------|---------|');
for (const f of files.slice(0, 10)) {
const symbolList = f.symbols.slice(0, 5).join(', ');
const suffix = f.symbols.length > 5 ? ` (+${f.symbols.length - 5})` : '';
lines.push(`| \`${f.relativePath}\` | ${symbolList}${suffix} |`);
}
lines.push('');
// Entry Points (top 5)
if (entryPoints.length > 0) {
lines.push('## Entry Points');
lines.push('');
lines.push('Start here when exploring this area:');
lines.push('');
for (const ep of entryPoints.slice(0, 5)) {
lines.push(`- **\`${ep.name}\`** (${ep.label}) \u2014 \`${ep.filePath}:${ep.startLine}\``);
}
lines.push('');
}
// Key Symbols (top 20, exported first, then by type)
lines.push('## Key Symbols');
lines.push('');
lines.push('| Symbol | Type | File | Line |');
lines.push('|--------|------|------|------|');
const sortedMembers = [...members].sort((a, b) => {
if (a.isExported !== b.isExported) return a.isExported ? -1 : 1;
return a.label.localeCompare(b.label);
});
for (const m of sortedMembers.slice(0, 20)) {
lines.push(`| \`${m.name}\` | ${m.label} | \`${m.filePath}\` | ${m.startLine} |`);
}
lines.push('');
// Execution Flows
if (flows.length > 0) {
lines.push('## Execution Flows');
lines.push('');
lines.push('| Flow | Type | Steps |');
lines.push('|------|------|-------|');
for (const f of flows.slice(0, 10)) {
lines.push(`| \`${f.heuristicLabel}\` | ${f.processType} | ${f.stepCount} |`);
}
lines.push('');
}
// Connected Areas
if (connections.length > 0) {
lines.push('## Connected Areas');
lines.push('');
lines.push('| Area | Connections |');
lines.push('|------|-------------|');
for (const c of connections.slice(0, 8)) {
lines.push(`| ${c.targetLabel} | ${c.count} calls |`);
}
lines.push('');
}
// How to Explore
const firstEntry = entryPoints.length > 0 ? entryPoints[0].name : (members.length > 0 ? members[0].name : community.label);
lines.push('## How to Explore');
lines.push('');
lines.push(`1. \`gitnexus_context({name: "${firstEntry}"})\` \u2014 see callers and callees`);
lines.push(`2. \`gitnexus_query({query: "${community.label.toLowerCase()}"})\` \u2014 find related execution flows`);
lines.push('3. Read key files listed above for implementation details');
lines.push('');
return lines.join('\n');
};
// ============================================================================
// UTILITY HELPERS
// ============================================================================
/**
* @brief Convert a community label to a kebab-case directory name
* @param {string} label - The community label
* @param {Set<string>} usedNames - Already-used names for collision detection
* @returns {string} Unique kebab-case name capped at 50 characters
*/
const toKebabName = (label: string, usedNames: Set<string>): string => {
let name = label
.toLowerCase()
.replace(/[^a-z0-9]+/g, '-')
.replace(/^-+|-+$/g, '')
.slice(0, 50);
if (!name) name = 'skill';
let candidate = name;
let counter = 2;
while (usedNames.has(candidate)) {
candidate = `${name}-${counter}`;
counter++;
}
return candidate;
};
/**
* @brief Convert an absolute or repo-relative file path to a clean relative path
* @param {string} filePath - The file path from the graph node
* @param {string} repoPath - Repository root path
* @returns {string} Relative path using forward slashes
*/
const toRelativePath = (filePath: string, repoPath: string): string => {
// Normalize to forward slashes for cross-platform consistency
const normalizedFile = filePath.replace(/\\/g, '/');
const normalizedRepo = repoPath.replace(/\\/g, '/');
if (normalizedFile.startsWith(normalizedRepo)) {
return normalizedFile.slice(normalizedRepo.length).replace(/^\//, '');
}
// Already relative or different root
return normalizedFile.replace(/^\//, '');
};
/**
* @brief Find the dominant (most common) top-level directory across files
* @param {FileInfo[]} files - File info entries
* @returns {string | null} Most common directory or null
*/
const getDominantDirectory = (files: FileInfo[]): string | null => {
const dirCounts = new Map<string, number>();
for (const f of files) {
const parts = f.relativePath.split('/');
if (parts.length >= 2) {
const dir = parts[0];
dirCounts.set(dir, (dirCounts.get(dir) || 0) + f.symbols.length);
}
}
let best: string | null = null;
let bestCount = 0;
for (const [dir, count] of dirCounts) {
if (count > bestCount) {
bestCount = count;
best = dir;
}
}
return best;
};
+1 -1
View File
@@ -7,9 +7,9 @@ export enum SupportedLanguages {
CPlusPlus = 'cpp',
CSharp = 'csharp',
Go = 'go',
Ruby = 'ruby',
Rust = 'rust',
PHP = 'php',
Kotlin = 'kotlin',
// Ruby = 'ruby',
Swift = 'swift',
}
+12 -6
View File
@@ -35,12 +35,14 @@ export type NodeLabel =
| 'Template';
import { SupportedLanguages } from '../../config/supported-languages.js';
export type NodeProperties = {
name: string,
filePath: string,
startLine?: number,
endLine?: number,
language?: string,
language?: SupportedLanguages,
isExported?: boolean,
// Optional AST-derived framework hint (e.g. @Controller, @GetMapping)
astFrameworkMultiplier?: number,
@@ -61,19 +63,23 @@ export type NodeProperties = {
// Entry point scoring (computed by process detection)
entryPointScore?: number,
entryPointReason?: string,
// Method signature (for MRO disambiguation)
parameterCount?: number,
returnType?: string,
}
export type RelationshipType =
| 'CONTAINS'
| 'CALLS'
| 'INHERITS'
| 'OVERRIDES'
export type RelationshipType =
| 'CONTAINS'
| 'CALLS'
| 'INHERITS'
| 'OVERRIDES'
| 'IMPORTS'
| 'USES'
| 'DEFINES'
| 'DECORATES'
| 'IMPLEMENTS'
| 'EXTENDS'
| 'HAS_METHOD'
| 'MEMBER_OF'
| 'STEP_IN_PROCESS'
+300 -281
View File
@@ -1,50 +1,29 @@
import { KnowledgeGraph } from '../graph/types.js';
import { ASTCache } from './ast-cache.js';
import { SymbolTable } from './symbol-table.js';
import { ImportMap } from './import-processor.js';
import type { SymbolDefinition, SymbolTable } from './symbol-table.js';
import { ImportMap, PackageMap, NamedImportMap, isFileInPackageDir } from './import-processor.js';
import { resolveSymbolInternal } from './symbol-resolver.js';
import { walkBindingChain } from './named-binding-extraction.js';
import Parser from 'tree-sitter';
import { loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { isLanguageAvailable, loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { generateId } from '../../lib/utils.js';
import { getLanguageFromFilename, yieldToEventLoop } from './utils.js';
import type { ExtractedCall, ExtractedRoute } from './workers/parse-worker.js';
/**
* Node types that represent function/method definitions across languages.
* Used to find the enclosing function for a call site.
*/
const FUNCTION_NODE_TYPES = new Set([
// TypeScript/JavaScript
'function_declaration',
'arrow_function',
'function_expression',
'method_definition',
'generator_function_declaration',
// Python
'function_definition',
// Common async variants
'async_function_declaration',
'async_arrow_function',
// Java
'method_declaration',
'constructor_declaration',
// C/C++
// 'function_definition' already included above
// Go
// 'method_declaration' already included from Java
// C#
'local_function_statement',
// Rust
'function_item',
'impl_item', // Methods inside impl blocks
// Kotlin (function_declaration already included above via JS/TS)
'anonymous_function',
'lambda_literal',
// PHP — no additional node types needed
// Swift
'init_declaration',
'deinit_declaration',
]);
import { SupportedLanguages } from '../../config/supported-languages.js';
import {
getLanguageFromFilename,
isVerboseIngestionEnabled,
yieldToEventLoop,
FUNCTION_NODE_TYPES,
extractFunctionName,
isBuiltInOrNoise,
countCallArguments,
inferCallForm,
extractReceiverName,
} from './utils.js';
import { buildTypeEnv, lookupTypeEnv } from './type-env.js';
import { getTreeSitterBufferSize } from './constants.js';
import type { ExtractedCall, ExtractedHeritage, ExtractedRoute } from './workers/parse-worker.js';
import { routeRubyCall } from './ruby-call-routing.js';
/**
* Walk up the AST from a node to find the enclosing function/method.
@@ -56,89 +35,22 @@ const findEnclosingFunction = (
symbolTable: SymbolTable
): string | null => {
let current = node.parent;
while (current) {
if (FUNCTION_NODE_TYPES.has(current.type)) {
// Found enclosing function - try to get its name
let funcName: string | null = null;
let label = 'Function';
// Different node types have different name locations
// Swift init/deinit — handle before generic cases (more specific)
if (current.type === 'init_declaration' || current.type === 'deinit_declaration') {
const funcName = current.type === 'init_declaration' ? 'init' : 'deinit';
return generateId('Constructor', `${filePath}:${funcName}`);
}
const { funcName, label } = extractFunctionName(current);
if (current.type === 'function_declaration' ||
current.type === 'function_definition' ||
current.type === 'async_function_declaration' ||
current.type === 'generator_function_declaration' ||
current.type === 'function_item') { // Rust function
// Named function: function foo() {}
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier' || c.type === 'property_identifier');
funcName = nameNode?.text;
} else if (current.type === 'impl_item') {
// Rust method inside impl block: wrapper around function_item or const_item
// We need to look inside for the function_item
const funcItem = current.children?.find((c: any) => c.type === 'function_item');
if (funcItem) {
const nameNode = funcItem.childForFieldName?.('name') ||
funcItem.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
}
} else if (current.type === 'method_definition') {
// Method: foo() {} inside class (JS/TS)
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'property_identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'method_declaration') {
// Java method: public void foo() {}
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'constructor_declaration') {
// Java constructor: public ClassName() {}
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method'; // Treat constructors as methods for process detection
} else if (current.type === 'arrow_function' || current.type === 'function_expression') {
// Arrow/expression: const foo = () => {} - check parent variable declarator
const parent = current.parent;
if (parent?.type === 'variable_declarator') {
const nameNode = parent.childForFieldName?.('name') ||
parent.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
}
}
if (funcName) {
// Look up the function in symbol table to get its node ID
// Try exact match first
const nodeId = symbolTable.lookupExact(filePath, funcName);
if (nodeId) return nodeId;
// Try construct ID manually if lookup fails (common for non-exported internal functions)
// Format should match what parsing-processor generates: "Function:path/to/file:funcName"
// Check if we already have a node with this ID in the symbol table to be safe
const generatedId = generateId(label, `${filePath}:${funcName}`);
// Ideally we should verify this ID exists, but strictly speaking if we are inside it,
// it SHOULD exist. Returning it is better than falling back to File.
return generatedId;
return generateId(label, `${filePath}:${funcName}`);
}
// Couldn't determine function name - try parent (might be nested)
}
current = current.parent;
}
return null; // Top-level call (not inside any function)
return null;
};
export const processCalls = async (
@@ -147,9 +59,14 @@ export const processCalls = async (
astCache: ASTCache,
symbolTable: SymbolTable,
importMap: ImportMap,
onProgress?: (current: number, total: number) => void
) => {
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void,
namedImportMap?: NamedImportMap,
): Promise<ExtractedHeritage[]> => {
const parser = await loadParser();
const collectedHeritage: ExtractedHeritage[] = [];
const logSkipped = isVerboseIngestionEnabled();
const skippedByLang = logSkipped ? new Map<string, number>() : null;
for (let i = 0; i < files.length; i++) {
const file = files[i];
@@ -159,6 +76,12 @@ export const processCalls = async (
// 1. Check language support first
const language = getLanguageFromFilename(file.path);
if (!language) continue;
if (!isLanguageAvailable(language)) {
if (skippedByLang) {
skippedByLang.set(language, (skippedByLang.get(language) ?? 0) + 1);
}
continue;
}
const queryStr = LANGUAGE_QUERIES[language];
if (!queryStr) continue;
@@ -174,7 +97,7 @@ export const processCalls = async (
// Cache Miss: Re-parse
// Use larger bufferSize for files > 32KB
try {
tree = parser.parse(file.content, undefined, { bufferSize: 1024 * 256 });
tree = parser.parse(file.content, undefined, { bufferSize: getTreeSitterBufferSize(file.content.length) });
} catch (parseError) {
// Skip files that can't be parsed
continue;
@@ -195,6 +118,10 @@ export const processCalls = async (
continue;
}
// Build per-file TypeEnv for receiver resolution
const lang = getLanguageFromFilename(file.path);
const typeEnv = lang ? buildTypeEnv(tree, lang) : new Map();
// 3. Process each call match
matches.forEach(match => {
const captureMap: Record<string, any> = {};
@@ -208,26 +135,79 @@ export const processCalls = async (
const calledName = nameNode.text;
// Ruby: route special calls to heritage or properties (imports handled by import-processor)
if (language === SupportedLanguages.Ruby) {
const callNode = captureMap['call'];
const routed = routeRubyCall(calledName, callNode);
switch (routed.kind) {
case 'skip':
case 'import': // handled by import-processor
return;
case 'heritage':
for (const item of routed.items) {
collectedHeritage.push({
filePath: file.path,
className: item.enclosingClass,
parentName: item.mixinName,
kind: 'trait-impl',
});
}
return;
case 'properties':
for (const item of routed.items) {
const nodeId = generateId('Property', `${file.path}:${item.propName}`);
graph.addNode({
id: nodeId,
label: 'Property' as any,
properties: {
name: item.propName, filePath: file.path,
startLine: item.startLine, endLine: item.endLine,
language: SupportedLanguages.Ruby, isExported: true,
description: item.accessorType,
},
});
symbolTable.add(file.path, item.propName, nodeId, 'Property');
const fileId = generateId('File', file.path);
const relId = generateId('DEFINES', `${fileId}->${nodeId}`);
graph.addRelationship({
id: relId, sourceId: fileId, targetId: nodeId,
type: 'DEFINES', confidence: 1.0, reason: '',
});
}
return;
case 'call':
break; // fall through to normal call processing below
}
}
// Skip common built-ins and noise
if (isBuiltInOrNoise(calledName)) return;
const callNode = captureMap['call'];
const callForm = inferCallForm(callNode, nameNode);
const receiverName = callForm === 'member' ? extractReceiverName(nameNode) : undefined;
const receiverTypeName = receiverName ? lookupTypeEnv(typeEnv, receiverName, callNode) : undefined;
// 4. Resolve the target using priority strategy (returns confidence)
const resolved = resolveCallTarget(
const resolved = resolveCallTarget({
calledName,
file.path,
symbolTable,
importMap
);
argCount: countCallArguments(callNode),
callForm,
receiverTypeName,
}, file.path, symbolTable, importMap, packageMap, namedImportMap);
if (!resolved) return;
// 5. Find the enclosing function (caller)
const callNode = captureMap['call'];
const enclosingFuncId = findEnclosingFunction(callNode, file.path, symbolTable);
// Use enclosing function as source, fallback to file for top-level calls
const sourceId = enclosingFuncId || generateId('File', file.path);
const relId = generateId('CALLS', `${sourceId}:${calledName}->${resolved.nodeId}`);
graph.addRelationship({
@@ -242,6 +222,16 @@ export const processCalls = async (
// Tree is now owned by the LRU cache — no manual delete needed
}
if (skippedByLang && skippedByLang.size > 0) {
for (const [lang, count] of skippedByLang.entries()) {
console.warn(
`[ingestion] Skipped ${count} ${lang} file(s) in call processing — ${lang} parser not available.`
);
}
}
return collectedHeritage;
};
/**
@@ -250,155 +240,168 @@ export const processCalls = async (
interface ResolveResult {
nodeId: string;
confidence: number; // 0-1: how sure are we?
reason: string; // 'import-resolved' | 'same-file' | 'fuzzy-global'
reason: string; // 'import-resolved' | 'same-file' | 'unique-global'
}
/**
* Resolve a function call to its target node ID using priority strategy:
* A. Check imported files first (highest confidence)
* B. Check local file definitions
* C. Fuzzy global search (lowest confidence)
*
* Returns confidence score so agents know what to trust.
*/
const resolveCallTarget = (
type ResolutionTier = 'same-file' | 'import-scoped' | 'unique-global';
interface TieredCandidates {
candidates: SymbolDefinition[];
tier: ResolutionTier;
}
const CALLABLE_SYMBOL_TYPES = new Set([
'Function',
'Method',
'Constructor',
'Macro',
'Delegate',
]);
const collectTieredCandidates = (
calledName: string,
currentFile: string,
symbolTable: SymbolTable,
importMap: ImportMap
): ResolveResult | null => {
// Strategy B first (cheapest — single map lookup): Check local file
const localNodeId = symbolTable.lookupExact(currentFile, calledName);
if (localNodeId) {
return { nodeId: localNodeId, confidence: 0.85, reason: 'same-file' };
}
// Strategy A: Check if any definition of calledName is in an imported file
// Reversed: instead of iterating all imports and checking each, get all definitions
// and check if any is imported. O(definitions) instead of O(imports).
importMap: ImportMap,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
): TieredCandidates | null => {
const allDefs = symbolTable.lookupFuzzy(calledName);
if (allDefs.length > 0) {
const importedFiles = importMap.get(currentFile);
if (importedFiles) {
for (const def of allDefs) {
if (importedFiles.has(def.filePath)) {
return { nodeId: def.nodeId, confidence: 0.9, reason: 'import-resolved' };
}
}
}
// Strategy C: Fuzzy global (no import match found)
const confidence = allDefs.length === 1 ? 0.5 : 0.3;
return { nodeId: allDefs[0].nodeId, confidence, reason: 'fuzzy-global' };
// Tier 1: Same-file — highest priority, prevents imports from shadowing local defs
// (matches resolveSymbolInternal which checks lookupExactFull before named bindings)
const localDefs = allDefs.filter(def => def.filePath === currentFile);
if (localDefs.length > 0) {
return { candidates: localDefs, tier: 'same-file' };
}
return null;
// Tier 2a-named: Check named bindings with re-export chain following.
// Aliased imports (import { User as U }) mean lookupFuzzy('U') returns
// empty but we can resolve via the exported name.
// Re-exports (export { User } from './base') are followed up to 5 hops.
if (namedImportMap) {
const chainResult = resolveNamedBindingChainForCandidates(
calledName, currentFile, symbolTable, namedImportMap, allDefs,
);
if (chainResult) return chainResult;
}
if (allDefs.length === 0) return null;
const importedFiles = importMap.get(currentFile);
if (importedFiles) {
const importedDefs = allDefs.filter(def => importedFiles.has(def.filePath));
if (importedDefs.length > 0) {
return { candidates: importedDefs, tier: 'import-scoped' };
}
}
const importedPackages = packageMap?.get(currentFile);
if (importedPackages) {
const packageDefs = allDefs.filter(def => {
for (const dirSuffix of importedPackages) {
if (isFileInPackageDir(def.filePath, dirSuffix)) return true;
}
return false;
});
if (packageDefs.length > 0) {
return { candidates: packageDefs, tier: 'import-scoped' };
}
}
// Tier 3: Global — pass all candidates through; filterCallableCandidates
// will narrow by kind/arity and resolveCallTarget only emits when exactly 1 remains.
return { candidates: allDefs, tier: 'unique-global' };
};
const CONSTRUCTOR_TARGET_TYPES = new Set(['Constructor', 'Class', 'Struct', 'Record']);
const filterCallableCandidates = (
candidates: SymbolDefinition[],
argCount?: number,
callForm?: 'free' | 'member' | 'constructor',
): SymbolDefinition[] => {
let kindFiltered: SymbolDefinition[];
if (callForm === 'constructor') {
// For constructor calls, prefer Constructor > Class/Struct/Record > callable fallback
const constructors = candidates.filter(c => c.type === 'Constructor');
if (constructors.length > 0) {
kindFiltered = constructors;
} else {
const types = candidates.filter(c => CONSTRUCTOR_TARGET_TYPES.has(c.type));
kindFiltered = types.length > 0 ? types : candidates.filter(c => CALLABLE_SYMBOL_TYPES.has(c.type));
}
} else {
kindFiltered = candidates.filter(c => CALLABLE_SYMBOL_TYPES.has(c.type));
}
if (kindFiltered.length === 0) return [];
if (argCount === undefined) return kindFiltered;
const hasParameterMetadata = kindFiltered.some(candidate => candidate.parameterCount !== undefined);
if (!hasParameterMetadata) return kindFiltered;
return kindFiltered.filter(candidate =>
candidate.parameterCount === undefined || candidate.parameterCount === argCount
);
};
const toResolveResult = (
definition: SymbolDefinition,
tier: ResolutionTier,
): ResolveResult => {
if (tier === 'same-file') {
return { nodeId: definition.nodeId, confidence: 0.95, reason: 'same-file' };
}
if (tier === 'import-scoped') {
return { nodeId: definition.nodeId, confidence: 0.9, reason: 'import-resolved' };
}
return { nodeId: definition.nodeId, confidence: 0.5, reason: 'unique-global' };
};
/**
* Filter out common built-in functions and noise
* that shouldn't be tracked as calls
* Resolve a function call to its target node ID using priority strategy:
* A. Narrow candidates by scope tier (same-file, import-scoped, unique-global)
* B. Filter to callable symbol kinds (constructor-aware when callForm is set)
* C. Apply arity filtering when parameter metadata is available
* D. Apply receiver-type filtering for member calls with typed receivers
*
* If filtering still leaves multiple candidates, refuse to emit a CALLS edge.
*/
/** Pre-built set (module-level singleton) to avoid re-creating per call */
const BUILT_IN_NAMES = new Set([
// JavaScript/TypeScript built-ins
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export',
'fetch', 'Response', 'Request',
// React hooks and common functions
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
// Common array/object methods
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python built-ins
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'open', 'read', 'write', 'close', 'append', 'extend', 'update',
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
// Kotlin stdlib (IMPORTANT: keep in sync with parse-worker.ts BUILT_IN_NAMES)
'println', 'print', 'readLine', 'require', 'requireNotNull', 'check', 'assert', 'lazy', 'error',
'listOf', 'mapOf', 'setOf', 'mutableListOf', 'mutableMapOf', 'mutableSetOf',
'arrayOf', 'sequenceOf', 'also', 'apply', 'run', 'with', 'takeIf', 'takeUnless',
'TODO', 'buildString', 'buildList', 'buildMap', 'buildSet',
'repeat', 'synchronized',
// Kotlin coroutine builders & scope functions
'launch', 'async', 'runBlocking', 'withContext', 'coroutineScope',
'supervisorScope', 'delay',
// Kotlin Flow operators
'flow', 'flowOf', 'collect', 'emit', 'onEach', 'catch',
'buffer', 'conflate', 'distinctUntilChanged',
'flatMapLatest', 'flatMapMerge', 'combine',
'stateIn', 'shareIn', 'launchIn',
// Kotlin infix stdlib functions
'to', 'until', 'downTo', 'step',
// C/C++ standard library and common kernel helpers
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
'scanf', 'fscanf', 'sscanf',
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
'sizeof', 'offsetof', 'typeof',
'assert', 'abort', 'exit', '_exit',
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
// Linux kernel common macros/helpers (not real call targets)
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
'min', 'max', 'clamp', 'abs', 'swap',
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
'GFP_KERNEL', 'GFP_ATOMIC',
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
'mutex_lock', 'mutex_unlock', 'mutex_init',
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
'get', 'put',
// Swift/iOS built-ins and standard library
'print', 'debugPrint', 'dump', 'fatalError', 'precondition', 'preconditionFailure',
'assert', 'assertionFailure', 'NSLog',
'abs', 'min', 'max', 'zip', 'stride', 'sequence', 'repeatElement',
'swap', 'withUnsafePointer', 'withUnsafeMutablePointer', 'withUnsafeBytes',
'autoreleasepool', 'unsafeBitCast', 'unsafeDowncast', 'numericCast',
'type', 'MemoryLayout',
// Swift collection/string methods (common noise)
'map', 'flatMap', 'compactMap', 'filter', 'reduce', 'forEach', 'contains',
'first', 'last', 'prefix', 'suffix', 'dropFirst', 'dropLast',
'sorted', 'reversed', 'enumerated', 'joined', 'split',
'append', 'insert', 'remove', 'removeAll', 'removeFirst', 'removeLast',
'isEmpty', 'count', 'index', 'startIndex', 'endIndex',
// UIKit/Foundation common methods (noise in call graph)
'addSubview', 'removeFromSuperview', 'layoutSubviews', 'setNeedsLayout',
'layoutIfNeeded', 'setNeedsDisplay', 'invalidateIntrinsicContentSize',
'addTarget', 'removeTarget', 'addGestureRecognizer',
'addConstraint', 'addConstraints', 'removeConstraint', 'removeConstraints',
'NSLocalizedString', 'Bundle',
'reloadData', 'reloadSections', 'reloadRows', 'performBatchUpdates',
'register', 'dequeueReusableCell', 'dequeueReusableSupplementaryView',
'beginUpdates', 'endUpdates', 'insertRows', 'deleteRows', 'insertSections', 'deleteSections',
'present', 'dismiss', 'pushViewController', 'popViewController', 'popToRootViewController',
'performSegue', 'prepare',
// GCD / async
'DispatchQueue', 'async', 'sync', 'asyncAfter',
'Task', 'withCheckedContinuation', 'withCheckedThrowingContinuation',
// Combine
'sink', 'store', 'assign', 'receive', 'subscribe',
// Notification / KVO
'addObserver', 'removeObserver', 'post', 'NotificationCenter',
]);
const resolveCallTarget = (
call: Pick<ExtractedCall, 'calledName' | 'argCount' | 'callForm' | 'receiverTypeName'>,
currentFile: string,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
): ResolveResult | null => {
const tiered = collectTieredCandidates(call.calledName, currentFile, symbolTable, importMap, packageMap, namedImportMap);
if (!tiered) return null;
const isBuiltInOrNoise = (name: string): boolean => BUILT_IN_NAMES.has(name);
const filteredCandidates = filterCallableCandidates(tiered.candidates, call.argCount, call.callForm);
// D. Receiver-type filtering: for member calls with a known receiver type,
// filter candidates by ownerId matching the resolved type's nodeId
if (call.callForm === 'member' && call.receiverTypeName && filteredCandidates.length > 1) {
const typeDefs = symbolTable.lookupFuzzy(call.receiverTypeName);
if (typeDefs.length > 0) {
const typeNodeIds = new Set(typeDefs.map(d => d.nodeId));
const ownerFiltered = filteredCandidates.filter(c => c.ownerId && typeNodeIds.has(c.ownerId));
if (ownerFiltered.length === 1) {
return toResolveResult(ownerFiltered[0], tiered.tier);
}
// If receiver filtering narrows to 0, fall through to name-only resolution
// If still 2+, refuse (don't guess)
if (ownerFiltered.length > 1) return null;
}
}
if (filteredCandidates.length !== 1) return null;
return toResolveResult(filteredCandidates[0], tiered.tier);
};
/**
* Fast path: resolve pre-extracted call sites from workers.
@@ -410,7 +413,9 @@ export const processCallsFromExtracted = async (
extractedCalls: ExtractedCall[],
symbolTable: SymbolTable,
importMap: ImportMap,
onProgress?: (current: number, total: number) => void
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void,
namedImportMap?: NamedImportMap,
) => {
// Group by file for progress reporting
const byFile = new Map<string, ExtractedCall[]>();
@@ -435,10 +440,12 @@ export const processCallsFromExtracted = async (
for (const call of calls) {
const resolved = resolveCallTarget(
call.calledName,
call,
call.filePath,
symbolTable,
importMap
importMap,
packageMap,
namedImportMap,
);
if (!resolved) continue;
@@ -465,6 +472,7 @@ export const processRoutesFromExtracted = async (
extractedRoutes: ExtractedRoute[],
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void
) => {
for (let i = 0; i < extractedRoutes.length; i++) {
@@ -476,24 +484,16 @@ export const processRoutesFromExtracted = async (
if (!route.controllerName || !route.methodName) continue;
// Resolve controller class in symbol table
const controllerDefs = symbolTable.lookupFuzzy(route.controllerName);
if (controllerDefs.length === 0) continue;
// Resolve controller class using shared resolver (Tier 1: same file,
// Tier 2: import-scoped, Tier 3: unique global).
const resolution = resolveSymbolInternal(route.controllerName, route.filePath, symbolTable, importMap, packageMap);
if (!resolution) continue;
// Prefer import-resolved match
const importedFiles = importMap.get(route.filePath);
let controllerDef = controllerDefs[0];
let confidence = controllerDefs.length === 1 ? 0.7 : 0.5;
if (importedFiles) {
for (const def of controllerDefs) {
if (importedFiles.has(def.filePath)) {
controllerDef = def;
confidence = 0.9;
break;
}
}
}
const controllerDef = resolution.definition;
// Derive confidence from the resolution tier
const confidence = resolution.tier === 'same-file' ? 0.95
: resolution.tier === 'import-scoped' ? 0.9
: 0.7;
// Find the method on the controller
const methodId = symbolTable.lookupExact(controllerDef.filePath, route.methodName);
@@ -527,3 +527,22 @@ export const processRoutesFromExtracted = async (
onProgress?.(extractedRoutes.length, extractedRoutes.length);
};
/**
* Follow re-export chains through NamedImportMap for call candidate collection.
* Delegates chain-walking to the shared walkBindingChain utility, then
* applies call-processor semantics: any number of matches accepted.
*/
const resolveNamedBindingChainForCandidates = (
calledName: string,
currentFile: string,
symbolTable: SymbolTable,
namedImportMap: NamedImportMap,
allDefs: SymbolDefinition[],
): TieredCandidates | null => {
const defs = walkBindingChain(calledName, currentFile, symbolTable, namedImportMap, allDefs);
if (defs && defs.length > 0) {
return { candidates: defs, tier: 'import-scoped' };
}
return null;
};
+19
View File
@@ -0,0 +1,19 @@
/**
* Default minimum buffer size for tree-sitter parsing (512 KB).
* tree-sitter requires bufferSize >= file size in bytes.
*/
export const TREE_SITTER_BUFFER_SIZE = 512 * 1024;
/**
* Maximum buffer size cap (32 MB) to prevent OOM on huge files.
* Also used as the file-size skip threshold — files larger than this are not parsed.
*/
export const TREE_SITTER_MAX_BUFFER = 32 * 1024 * 1024;
/**
* Compute adaptive buffer size for tree-sitter parsing.
* Uses 2× file size, clamped between 512 KB and 32 MB.
* Previous 256 KB fixed limit silently skipped files > ~200 KB (e.g., imgui.h at 411 KB).
*/
export const getTreeSitterBufferSize = (contentLength: number): number =>
Math.min(Math.max(contentLength * 2, TREE_SITTER_BUFFER_SIZE), TREE_SITTER_MAX_BUFFER);
@@ -11,9 +11,10 @@
*/
import { detectFrameworkFromPath } from './framework-detection.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
// ============================================================================
// NAME PATTERNS - All 9 supported languages
// NAME PATTERNS - All 11 supported languages
// ============================================================================
/**
@@ -38,39 +39,47 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
],
// JavaScript/TypeScript
'javascript': [
[SupportedLanguages.JavaScript]: [
/^use[A-Z]/, // React hooks (useEffect, etc.)
],
'typescript': [
[SupportedLanguages.TypeScript]: [
/^use[A-Z]/, // React hooks
],
// Python
'python': [
[SupportedLanguages.Python]: [
/^app$/, // Flask/FastAPI app
/^(get|post|put|delete|patch)_/i, // REST conventions
/^api_/, // API functions
/^view_/, // Django views
],
// Java
'java': [
[SupportedLanguages.Java]: [
/^do[A-Z]/, // doGet, doPost (Servlets)
/^create[A-Z]/, // Factory patterns
/^build[A-Z]/, // Builder patterns
/Service$/, // UserService
],
// C#
'csharp': [
/^(Get|Post|Put|Delete)/, // ASP.NET conventions
/Action$/, // MVC actions
/^On[A-Z]/, // Event handlers
/Async$/, // Async entry points
[SupportedLanguages.CSharp]: [
/^(Get|Post|Put|Delete|Patch)/, // ASP.NET action methods
/Action$/, // MVC actions
/^On[A-Z]/, // Event handlers / Blazor lifecycle
/Async$/, // Async entry points
/^Configure$/, // Startup.Configure
/^ConfigureServices$/, // Startup.ConfigureServices
/^Handle$/, // MediatR / generic handler
/^Execute$/, // Command pattern
/^Invoke$/, // Middleware Invoke
/^Map[A-Z]/, // Minimal API MapGet, MapPost
/Service$/, // Service classes
/^Seed/, // Database seeding
],
// Go
'go': [
[SupportedLanguages.Go]: [
/Handler$/, // http.Handler pattern
/^Serve/, // ServeHTTP
/^New[A-Z]/, // Constructor pattern (returns new instance)
@@ -78,7 +87,7 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
],
// Rust
'rust': [
[SupportedLanguages.Rust]: [
/^(get|post|put|delete)_handler$/i,
/^handle_/, // handle_request
/^new$/, // Constructor pattern
@@ -86,25 +95,64 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
/^spawn/, // Async spawn
],
// C - explicit main() boost (critical for C programs)
'c': [
// C - explicit main() boost plus common C entry point conventions
[SupportedLanguages.C]: [
/^main$/, // THE entry point
/^init_/, // Initialization functions
/^start_/, // Start functions
/^run_/, // Run functions
/^init_/, // init_server, init_client
/_init$/, // module_init, server_init
/^start_/, // start_server
/_start$/, // thread_start
/^run_/, // run_loop
/_run$/, // event_run
/^stop_/, // stop_server
/_stop$/, // service_stop
/^open_/, // open_connection
/_open$/, // file_open
/^close_/, // close_connection
/_close$/, // socket_close
/^create_/, // create_session
/_create$/, // object_create
/^destroy_/, // destroy_session
/_destroy$/, // object_destroy
/^handle_/, // handle_request
/_handler$/, // signal_handler
/_callback$/, // event_callback
/^cmd_/, // tmux: cmd_new_window, cmd_attach_session
/^server_/, // server_start, server_loop
/^client_/, // client_connect
/^session_/, // session_create
/^window_/, // window_resize (tmux)
/^key_/, // key_press
/^input_/, // input_parse
/^output_/, // output_write
/^notify_/, // notify_client
/^control_/, // control_start
],
// C++ - same as C plus class patterns
'cpp': [
// C++ - same as C plus OOP/template patterns
[SupportedLanguages.CPlusPlus]: [
/^main$/, // THE entry point
/^init_/,
/_init$/,
/^Create[A-Z]/, // Factory patterns
/^create_/,
/^Run$/, // Run methods
/^run$/,
/^Start$/, // Start methods
/^start$/,
/^handle_/,
/_handler$/,
/_callback$/,
/^OnEvent/, // Event callbacks
/^on_/,
/::Run$/, // Class::Run
/::Start$/, // Class::Start
/::Init$/, // Class::Init
/::Execute$/, // Class::Execute
],
// Swift / iOS
'swift': [
[SupportedLanguages.Swift]: [
/^viewDidLoad$/, // UIKit lifecycle
/^viewWillAppear$/, // UIKit lifecycle
/^viewDidAppear$/, // UIKit lifecycle
@@ -124,7 +172,7 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
],
// PHP / Laravel
'php': [
[SupportedLanguages.PHP]: [
/Controller$/, // UserController (class name convention)
/^handle$/, // Job::handle(), Listener::handle()
/^execute$/, // Command::execute()
@@ -143,8 +191,23 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
/^save$/, // Repository::save()
/^delete$/, // Repository::delete()
],
// Ruby
'ruby': [
/^call$/, // Service objects (MyService.call)
/^perform$/, // Background jobs (Sidekiq, ActiveJob)
/^execute$/, // Command pattern
],
};
/** Pre-computed merged patterns (universal + language-specific) to avoid per-call array allocation. */
const MERGED_ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {};
const UNIVERSAL_PATTERNS = ENTRY_POINT_PATTERNS['*'] || [];
for (const [lang, patterns] of Object.entries(ENTRY_POINT_PATTERNS)) {
if (lang === '*') continue;
MERGED_ENTRY_POINT_PATTERNS[lang] = [...UNIVERSAL_PATTERNS, ...patterns];
}
// ============================================================================
// UTILITY PATTERNS - Functions that should be penalized
// ============================================================================
@@ -199,7 +262,7 @@ export interface EntryPointScoreResult {
*/
export function calculateEntryPointScore(
name: string,
language: string,
language: SupportedLanguages,
isExported: boolean,
callerCount: number,
calleeCount: number,
@@ -232,9 +295,7 @@ export function calculateEntryPointScore(
reasons.push('utility-pattern');
} else {
// Check positive patterns
const universalPatterns = ENTRY_POINT_PATTERNS['*'] || [];
const langPatterns = ENTRY_POINT_PATTERNS[language] || [];
const allPatterns = [...universalPatterns, ...langPatterns];
const allPatterns = MERGED_ENTRY_POINT_PATTERNS[language] || UNIVERSAL_PATTERNS;
if (allPatterns.some(p => p.test(name))) {
nameMultiplier = 1.5; // Bonus for matching entry point pattern
@@ -296,13 +357,23 @@ export function isTestFile(filePath: string): boolean {
p.endsWith('test.swift') ||
p.includes('uitests/') ||
// C# test patterns
p.endsWith('tests.cs') ||
p.endsWith('test.cs') ||
p.includes('.tests/') ||
p.includes('tests.cs') ||
p.includes('.test/') ||
p.includes('.integrationtests/') ||
p.includes('.unittests/') ||
p.includes('/testproject/') ||
// PHP/Laravel test patterns
p.endsWith('test.php') ||
p.endsWith('spec.php') ||
p.includes('/tests/feature/') ||
p.includes('/tests/unit/')
p.includes('/tests/unit/') ||
// Ruby test patterns
p.endsWith('_spec.rb') ||
p.endsWith('_test.rb') ||
p.includes('/spec/') ||
p.includes('/test/fixtures/')
);
}
@@ -0,0 +1,243 @@
/**
* Export Detection
*
* Determines whether a symbol (function, class, etc.) is exported/public
* in its language. This is a pure function — safe for use in worker threads.
*
* Shared between parse-worker.ts (worker pool) and parsing-processor.ts (sequential fallback).
*/
import { findSiblingChild, SyntaxNode } from './utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
/** Handler type: given a node and symbol name, return true if the symbol is exported/public. */
type ExportChecker = (node: SyntaxNode, name: string) => boolean;
// ============================================================================
// Per-language export checkers
// ============================================================================
/** JS/TS: walk ancestors looking for export_statement or export_specifier. */
const tsExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
const type = current.type;
if (type === 'export_statement' ||
type === 'export_specifier' ||
(type === 'lexical_declaration' && current.parent?.type === 'export_statement')) {
return true;
}
// Fallback: check if node text starts with 'export ' for edge cases
if (current.text?.startsWith('export ')) {
return true;
}
current = current.parent;
}
return false;
};
/** Python: public if no leading underscore (convention). */
const pythonExportChecker: ExportChecker = (_node, name) => !name.startsWith('_');
/** Java: check for 'public' modifier — modifiers are siblings of the name node, not parents. */
const javaExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.parent) {
const parent = current.parent;
for (let i = 0; i < parent.childCount; i++) {
const child = parent.child(i);
if (child?.type === 'modifiers' && child.text?.includes('public')) {
return true;
}
}
if (parent.type === 'method_declaration' || parent.type === 'constructor_declaration') {
if (parent.text?.trimStart().startsWith('public')) {
return true;
}
}
}
current = current.parent;
}
return false;
};
/** C# declaration node types for sibling modifier scanning. */
const CSHARP_DECL_TYPES = new Set([
'method_declaration', 'local_function_statement', 'constructor_declaration',
'class_declaration', 'interface_declaration', 'struct_declaration',
'enum_declaration', 'record_declaration', 'record_struct_declaration',
'record_class_declaration', 'delegate_declaration',
'property_declaration', 'field_declaration', 'event_declaration',
'namespace_declaration', 'file_scoped_namespace_declaration',
]);
/**
* C#: modifier nodes are SIBLINGS of the name node inside the declaration.
* Walk up to the declaration node, then scan its direct children.
*/
const csharpExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (CSHARP_DECL_TYPES.has(current.type)) {
for (let i = 0; i < current.childCount; i++) {
const child = current.child(i);
if (child?.type === 'modifier' && child.text === 'public') return true;
}
return false;
}
current = current.parent;
}
return false;
};
/** Go: uppercase first letter = exported. */
const goExportChecker: ExportChecker = (_node, name) => {
if (name.length === 0) return false;
const first = name[0];
return first === first.toUpperCase() && first !== first.toLowerCase();
};
/** Rust declaration node types for sibling visibility_modifier scanning. */
const RUST_DECL_TYPES = new Set([
'function_item', 'struct_item', 'enum_item', 'trait_item', 'impl_item',
'union_item', 'type_item', 'const_item', 'static_item', 'mod_item',
'use_declaration', 'associated_type', 'function_signature_item',
]);
/**
* Rust: visibility_modifier is a SIBLING of the name node within the declaration node
* (function_item, struct_item, etc.), not a parent. Walk up to the declaration node,
* then scan its direct children.
*/
const rustExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (RUST_DECL_TYPES.has(current.type)) {
for (let i = 0; i < current.childCount; i++) {
const child = current.child(i);
if (child?.type === 'visibility_modifier' && child.text?.startsWith('pub')) return true;
}
return false;
}
current = current.parent;
}
return false;
};
/**
* Kotlin: default visibility is public (unlike Java).
* visibility_modifier is inside modifiers, a sibling of the name node within the declaration.
*/
const kotlinExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.parent) {
const visMod = findSiblingChild(current.parent, 'modifiers', 'visibility_modifier');
if (visMod) {
const text = visMod.text;
if (text === 'private' || text === 'internal' || text === 'protected') return false;
if (text === 'public') return true;
}
}
current = current.parent;
}
// No visibility modifier = public (Kotlin default)
return true;
};
/**
* C/C++: functions without 'static' storage class have external linkage by default,
* making them globally accessible (equivalent to exported). Only functions explicitly
* marked 'static' are file-scoped (not exported). C++ anonymous namespaces
* (namespace { ... }) also give internal linkage.
*/
const cCppExportChecker: ExportChecker = (node, _name) => {
let cur: SyntaxNode | null = node;
while (cur) {
if (cur.type === 'function_definition' || cur.type === 'declaration') {
// Check for 'static' storage class specifier as a direct child node.
// This avoids reading the full function text (which can be very large).
for (let i = 0; i < cur.childCount; i++) {
const child = cur.child(i);
if (child?.type === 'storage_class_specifier' && child.text === 'static') return false;
}
}
// C++ anonymous namespace: namespace_definition with no name child = internal linkage
if (cur.type === 'namespace_definition') {
const hasName = cur.childForFieldName?.('name');
if (!hasName) return false;
}
cur = cur.parent;
}
return true; // Top-level C/C++ functions default to external linkage
};
/** PHP: check for visibility modifier or top-level scope. */
const phpExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.type === 'class_declaration' ||
current.type === 'interface_declaration' ||
current.type === 'trait_declaration' ||
current.type === 'enum_declaration') {
return true;
}
if (current.type === 'visibility_modifier') {
return current.text === 'public';
}
current = current.parent;
}
// Top-level functions are globally accessible
return true;
};
/** Swift: check for 'public' or 'open' access modifiers. */
const swiftExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.type === 'modifiers' || current.type === 'visibility_modifier') {
const text = current.text || '';
if (text.includes('public') || text.includes('open')) return true;
}
current = current.parent;
}
return false;
};
// ============================================================================
// Exhaustive dispatch table — satisfies enforces all SupportedLanguages are covered
// ============================================================================
const exportCheckers = {
[SupportedLanguages.JavaScript]: tsExportChecker,
[SupportedLanguages.TypeScript]: tsExportChecker,
[SupportedLanguages.Python]: pythonExportChecker,
[SupportedLanguages.Java]: javaExportChecker,
[SupportedLanguages.CSharp]: csharpExportChecker,
[SupportedLanguages.Go]: goExportChecker,
[SupportedLanguages.Rust]: rustExportChecker,
[SupportedLanguages.Kotlin]: kotlinExportChecker,
[SupportedLanguages.C]: cCppExportChecker,
[SupportedLanguages.CPlusPlus]: cCppExportChecker,
[SupportedLanguages.PHP]: phpExportChecker,
[SupportedLanguages.Swift]: swiftExportChecker,
[SupportedLanguages.Ruby]: (_node, _name) => true,
} satisfies Record<SupportedLanguages, ExportChecker>;
// ============================================================================
// Public API
// ============================================================================
/**
* Check if a tree-sitter node is exported/public in its language.
* @param node - The tree-sitter AST node
* @param name - The symbol name
* @param language - The programming language
* @returns true if the symbol is exported/public
*/
export const isNodeExported = (node: SyntaxNode, name: string, language: SupportedLanguages): boolean => {
const checker = exportCheckers[language];
if (!checker) return false;
return checker(node, name);
};
@@ -183,7 +183,35 @@ export function detectFrameworkFromPath(filePath: string): FrameworkHint | null
if (p.endsWith('controller.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 3.0, reason: 'aspnet-controller-file' };
}
// ASP.NET Services
if ((p.includes('/services/') || p.includes('/service/')) && p.endsWith('.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 1.8, reason: 'aspnet-service' };
}
// ASP.NET Middleware
if (p.includes('/middleware/') && p.endsWith('.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 2.5, reason: 'aspnet-middleware' };
}
// SignalR Hubs
if (p.includes('/hubs/') && p.endsWith('.cs')) {
return { framework: 'signalr', entryPointMultiplier: 2.5, reason: 'signalr-hub' };
}
if (p.endsWith('hub.cs')) {
return { framework: 'signalr', entryPointMultiplier: 2.5, reason: 'signalr-hub-file' };
}
// Minimal API / Program.cs / Startup.cs
if (p.endsWith('/program.cs') || p.endsWith('/startup.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 3.0, reason: 'aspnet-entry' };
}
// Background services / Hosted services
if ((p.includes('/backgroundservices/') || p.includes('/hostedservices/')) && p.endsWith('.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 2.0, reason: 'aspnet-background-service' };
}
// Blazor pages
if (p.includes('/pages/') && p.endsWith('.razor')) {
return { framework: 'blazor', entryPointMultiplier: 2.5, reason: 'blazor-page' };
@@ -302,6 +330,18 @@ export function detectFrameworkFromPath(filePath: string): FrameworkHint | null
return { framework: 'laravel', entryPointMultiplier: 1.5, reason: 'laravel-repository' };
}
// ========== RUBY ==========
// Ruby: bin/ or exe/ (CLI entry points)
if ((p.includes('/bin/') || p.includes('/exe/')) && p.endsWith('.rb')) {
return { framework: 'ruby', entryPointMultiplier: 2.5, reason: 'ruby-executable' };
}
// Ruby: Rakefile or *.rake (task definitions)
if (p.endsWith('/rakefile') || p.endsWith('.rake')) {
return { framework: 'ruby', entryPointMultiplier: 1.5, reason: 'ruby-rake' };
}
// ========== SWIFT / iOS ==========
// iOS App entry points (highest priority)
@@ -385,7 +425,11 @@ export const FRAMEWORK_AST_PATTERNS = {
'jaxrs': ['@Path', '@GET', '@POST', '@PUT', '@DELETE'],
// C# attributes
'aspnet': ['[ApiController]', '[HttpGet]', '[HttpPost]', '[Route]'],
'aspnet': ['[ApiController]', '[HttpGet]', '[HttpPost]', '[HttpPut]', '[HttpDelete]',
'[Route]', '[Authorize]', '[AllowAnonymous]'],
'signalr': ['[HubMethodName]', ': Hub', ': Hub<'],
'blazor': ['@page', '[Parameter]', '@inject'],
'efcore': ['DbContext', 'DbSet<', 'OnModelCreating'],
// Go patterns (function signatures)
'go-http': ['http.Handler', 'http.HandlerFunc', 'ServeHTTP'],
@@ -405,6 +449,8 @@ export const FRAMEWORK_AST_PATTERNS = {
'combine': ['sink', 'assign', 'Publisher', 'Subscriber'],
};
import { SupportedLanguages } from '../../config/supported-languages.js';
interface AstFrameworkPatternConfig {
framework: string;
entryPointMultiplier: number;
@@ -413,30 +459,33 @@ interface AstFrameworkPatternConfig {
}
const AST_FRAMEWORK_PATTERNS_BY_LANGUAGE: Record<string, AstFrameworkPatternConfig[]> = {
javascript: [
[SupportedLanguages.JavaScript]: [
{ framework: 'nestjs', entryPointMultiplier: 3.2, reason: 'nestjs-decorator', patterns: FRAMEWORK_AST_PATTERNS.nestjs },
],
typescript: [
[SupportedLanguages.TypeScript]: [
{ framework: 'nestjs', entryPointMultiplier: 3.2, reason: 'nestjs-decorator', patterns: FRAMEWORK_AST_PATTERNS.nestjs },
],
python: [
[SupportedLanguages.Python]: [
{ framework: 'fastapi', entryPointMultiplier: 3.0, reason: 'fastapi-decorator', patterns: FRAMEWORK_AST_PATTERNS.fastapi },
{ framework: 'flask', entryPointMultiplier: 2.8, reason: 'flask-decorator', patterns: FRAMEWORK_AST_PATTERNS.flask },
],
java: [
[SupportedLanguages.Java]: [
{ framework: 'spring', entryPointMultiplier: 3.2, reason: 'spring-annotation', patterns: FRAMEWORK_AST_PATTERNS.spring },
{ framework: 'jaxrs', entryPointMultiplier: 3.0, reason: 'jaxrs-annotation', patterns: FRAMEWORK_AST_PATTERNS.jaxrs },
],
kotlin: [
[SupportedLanguages.Kotlin]: [
{ framework: 'spring-kotlin', entryPointMultiplier: 3.2, reason: 'spring-kotlin-annotation', patterns: FRAMEWORK_AST_PATTERNS.spring },
{ framework: 'jaxrs', entryPointMultiplier: 3.0, reason: 'jaxrs-annotation', patterns: FRAMEWORK_AST_PATTERNS.jaxrs },
{ framework: 'ktor', entryPointMultiplier: 2.8, reason: 'ktor-routing', patterns: ['routing', 'embeddedServer', 'Application.module'] },
{ framework: 'android-kotlin', entryPointMultiplier: 2.5, reason: 'android-annotation', patterns: ['@AndroidEntryPoint', 'AppCompatActivity', 'Fragment('] },
],
csharp: [
[SupportedLanguages.CSharp]: [
{ framework: 'aspnet', entryPointMultiplier: 3.2, reason: 'aspnet-attribute', patterns: FRAMEWORK_AST_PATTERNS.aspnet },
{ framework: 'signalr', entryPointMultiplier: 2.8, reason: 'signalr-attribute', patterns: FRAMEWORK_AST_PATTERNS.signalr },
{ framework: 'blazor', entryPointMultiplier: 2.5, reason: 'blazor-attribute', patterns: FRAMEWORK_AST_PATTERNS.blazor },
{ framework: 'efcore', entryPointMultiplier: 2.0, reason: 'efcore-pattern', patterns: FRAMEWORK_AST_PATTERNS.efcore },
],
php: [
[SupportedLanguages.PHP]: [
{ framework: 'laravel', entryPointMultiplier: 3.0, reason: 'php-route-attribute', patterns: FRAMEWORK_AST_PATTERNS.laravel },
],
};
@@ -456,7 +505,7 @@ const AST_PATTERNS_LOWERED: Record<string, Array<{ framework: string; entryPoint
* Note: callers should slice definitionText to ~300 chars since annotations appear at the start.
*/
export function detectFrameworkFromAST(
language: string,
language: SupportedLanguages,
definitionText: string
): FrameworkHint | null {
if (!language || !definitionText) return null;
+113 -32
View File
@@ -1,29 +1,83 @@
/**
* Heritage Processor
*
*
* Extracts class inheritance relationships:
* - EXTENDS: Class extends another Class (TS, JS, Python)
* - IMPLEMENTS: Class implements an Interface (TS only)
* - EXTENDS: Class extends another Class (TS, JS, Python, C#, C++)
* - IMPLEMENTS: Class implements an Interface (TS, C#, Java, Kotlin, PHP)
*
* Languages like C# use a single `base_list` for both class and interface parents.
* We resolve the correct edge type by checking the symbol table: if the parent is
* registered as an Interface, we emit IMPLEMENTS; otherwise EXTENDS. For unresolved
* external symbols, the fallback heuristic is language-gated:
* - C# / Java: apply the `I[A-Z]` naming convention (e.g. IDisposable → IMPLEMENTS)
* - Swift: default to IMPLEMENTS (protocol conformance is more common than class inheritance)
* - All other languages: default to EXTENDS
*/
import { KnowledgeGraph } from '../graph/types.js';
import { ASTCache } from './ast-cache.js';
import { SymbolTable } from './symbol-table.js';
import { SymbolTable, SymbolDefinition } from './symbol-table.js';
import Parser from 'tree-sitter';
import { loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { isLanguageAvailable, loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { generateId } from '../../lib/utils.js';
import { getLanguageFromFilename, yieldToEventLoop } from './utils.js';
import { getLanguageFromFilename, isVerboseIngestionEnabled, yieldToEventLoop } from './utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
import { getTreeSitterBufferSize } from './constants.js';
import type { ExtractedHeritage } from './workers/parse-worker.js';
import { resolveSymbol } from './symbol-resolver.js';
import type { ImportMap, PackageMap } from './import-processor.js';
/** C#/Java convention: interfaces start with I followed by an uppercase letter */
const INTERFACE_NAME_RE = /^I[A-Z]/;
/**
* Determine whether a heritage.extends capture is actually an IMPLEMENTS relationship.
* Uses the symbol table first (authoritative — Tier 1); falls back to a language-gated
* heuristic for external symbols not present in the graph:
* - C# / Java: `I[A-Z]` naming convention
* - Swift: default IMPLEMENTS (protocol conformance is the norm)
* - All others: default EXTENDS
*/
const resolveExtendsType = (
parentName: string,
currentFilePath: string,
symbolTable: SymbolTable,
importMap: ImportMap,
language: SupportedLanguages,
packageMap?: PackageMap,
): { type: 'EXTENDS' | 'IMPLEMENTS'; idPrefix: string } => {
const resolved = resolveSymbol(parentName, currentFilePath, symbolTable, importMap, packageMap);
if (resolved) {
const isInterface = resolved.type === 'Interface';
return isInterface
? { type: 'IMPLEMENTS', idPrefix: 'Interface' }
: { type: 'EXTENDS', idPrefix: 'Class' };
}
// Unresolved symbol — fall back to language-specific heuristic
if (language === SupportedLanguages.CSharp || language === SupportedLanguages.Java) {
if (INTERFACE_NAME_RE.test(parentName)) {
return { type: 'IMPLEMENTS', idPrefix: 'Interface' };
}
} else if (language === SupportedLanguages.Swift) {
// Protocol conformance is far more common than class inheritance in Swift
return { type: 'IMPLEMENTS', idPrefix: 'Interface' };
}
return { type: 'EXTENDS', idPrefix: 'Class' };
};
export const processHeritage = async (
graph: KnowledgeGraph,
files: { path: string; content: string }[],
astCache: ASTCache,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void
) => {
const parser = await loadParser();
const logSkipped = isVerboseIngestionEnabled();
const skippedByLang = logSkipped ? new Map<string, number>() : null;
for (let i = 0; i < files.length; i++) {
const file = files[i];
@@ -33,6 +87,12 @@ export const processHeritage = async (
// 1. Check language support
const language = getLanguageFromFilename(file.path);
if (!language) continue;
if (!isLanguageAvailable(language)) {
if (skippedByLang) {
skippedByLang.set(language, (skippedByLang.get(language) ?? 0) + 1);
}
continue;
}
const queryStr = LANGUAGE_QUERIES[language];
if (!queryStr) continue;
@@ -47,7 +107,7 @@ export const processHeritage = async (
if (!tree) {
// Use larger bufferSize for files > 32KB
try {
tree = parser.parse(file.content, undefined, { bufferSize: 1024 * 256 });
tree = parser.parse(file.content, undefined, { bufferSize: getTreeSitterBufferSize(file.content.length) });
} catch (parseError) {
// Skip files that can't be parsed
continue;
@@ -75,27 +135,34 @@ export const processHeritage = async (
captureMap[c.name] = c.node;
});
// EXTENDS: Class extends another Class
// EXTENDS or IMPLEMENTS: resolve via symbol table for languages where
// the tree-sitter query can't distinguish classes from interfaces (C#, Java)
if (captureMap['heritage.class'] && captureMap['heritage.extends']) {
// Go struct embedding: skip named fields (only anonymous fields are embedded)
const extendsNode = captureMap['heritage.extends'];
const fieldDecl = extendsNode.parent;
if (fieldDecl?.type === 'field_declaration' && fieldDecl.childForFieldName('name')) {
return; // Named field, not struct embedding
}
const className = captureMap['heritage.class'].text;
const parentClassName = captureMap['heritage.extends'].text;
// Resolve both class IDs
const { type: relType, idPrefix } = resolveExtendsType(parentClassName, file.path, symbolTable, importMap, language, packageMap);
const childId = symbolTable.lookupExact(file.path, className) ||
symbolTable.lookupFuzzy(className)[0]?.nodeId ||
resolveSymbol(className, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Class', `${file.path}:${className}`);
const parentId = symbolTable.lookupFuzzy(parentClassName)[0]?.nodeId ||
generateId('Class', `${parentClassName}`);
const parentId = resolveSymbol(parentClassName, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId(idPrefix, `${parentClassName}`);
if (childId && parentId && childId !== parentId) {
const relId = generateId('EXTENDS', `${childId}->${parentId}`);
graph.addRelationship({
id: relId,
id: generateId(relType, `${childId}->${parentId}`),
sourceId: childId,
targetId: parentId,
type: 'EXTENDS',
type: relType,
confidence: 1.0,
reason: '',
});
@@ -109,10 +176,10 @@ export const processHeritage = async (
// Resolve class and interface IDs
const classId = symbolTable.lookupExact(file.path, className) ||
symbolTable.lookupFuzzy(className)[0]?.nodeId ||
resolveSymbol(className, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Class', `${file.path}:${className}`);
const interfaceId = symbolTable.lookupFuzzy(interfaceName)[0]?.nodeId ||
const interfaceId = resolveSymbol(interfaceName, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Interface', `${interfaceName}`);
if (classId && interfaceId) {
@@ -136,10 +203,10 @@ export const processHeritage = async (
// Resolve struct and trait IDs
const structId = symbolTable.lookupExact(file.path, structName) ||
symbolTable.lookupFuzzy(structName)[0]?.nodeId ||
resolveSymbol(structName, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Struct', `${file.path}:${structName}`);
const traitId = symbolTable.lookupFuzzy(traitName)[0]?.nodeId ||
const traitId = resolveSymbol(traitName, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Trait', `${traitName}`);
if (structId && traitId) {
@@ -159,6 +226,14 @@ export const processHeritage = async (
// Tree is now owned by the LRU cache — no manual delete needed
}
if (skippedByLang && skippedByLang.size > 0) {
for (const [lang, count] of skippedByLang.entries()) {
console.warn(
`[ingestion] Skipped ${count} ${lang} file(s) in heritage processing — ${lang} parser not available.`
);
}
}
};
/**
@@ -169,6 +244,8 @@ export const processHeritageFromExtracted = async (
graph: KnowledgeGraph,
extractedHeritage: ExtractedHeritage[],
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void
) => {
const total = extractedHeritage.length;
@@ -182,29 +259,33 @@ export const processHeritageFromExtracted = async (
const h = extractedHeritage[i];
if (h.kind === 'extends') {
const fileLanguage = getLanguageFromFilename(h.filePath);
if (!fileLanguage) continue;
const { type: relType, idPrefix } = resolveExtendsType(h.parentName, h.filePath, symbolTable, importMap, fileLanguage, packageMap);
const childId = symbolTable.lookupExact(h.filePath, h.className) ||
symbolTable.lookupFuzzy(h.className)[0]?.nodeId ||
resolveSymbol(h.className, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Class', `${h.filePath}:${h.className}`);
const parentId = symbolTable.lookupFuzzy(h.parentName)[0]?.nodeId ||
generateId('Class', `${h.parentName}`);
const parentId = resolveSymbol(h.parentName, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId(idPrefix, `${h.parentName}`);
if (childId && parentId && childId !== parentId) {
graph.addRelationship({
id: generateId('EXTENDS', `${childId}->${parentId}`),
id: generateId(relType, `${childId}->${parentId}`),
sourceId: childId,
targetId: parentId,
type: 'EXTENDS',
type: relType,
confidence: 1.0,
reason: '',
});
}
} else if (h.kind === 'implements') {
const classId = symbolTable.lookupExact(h.filePath, h.className) ||
symbolTable.lookupFuzzy(h.className)[0]?.nodeId ||
resolveSymbol(h.className, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Class', `${h.filePath}:${h.className}`);
const interfaceId = symbolTable.lookupFuzzy(h.parentName)[0]?.nodeId ||
const interfaceId = resolveSymbol(h.parentName, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Interface', `${h.parentName}`);
if (classId && interfaceId) {
@@ -219,10 +300,10 @@ export const processHeritageFromExtracted = async (
}
} else if (h.kind === 'trait-impl') {
const structId = symbolTable.lookupExact(h.filePath, h.className) ||
symbolTable.lookupFuzzy(h.className)[0]?.nodeId ||
resolveSymbol(h.className, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Struct', `${h.filePath}:${h.className}`);
const traitId = symbolTable.lookupFuzzy(h.parentName)[0]?.nodeId ||
const traitId = resolveSymbol(h.parentName, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Trait', `${h.parentName}`);
if (structId && traitId) {
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,215 @@
import fs from 'fs/promises';
import path from 'path';
const isDev = process.env.NODE_ENV === 'development';
// ============================================================================
// LANGUAGE-SPECIFIC CONFIG TYPES
// ============================================================================
/** TypeScript path alias config parsed from tsconfig.json */
export interface TsconfigPaths {
/** Map of alias prefix -> target prefix (e.g., "@/" -> "src/") */
aliases: Map<string, string>;
/** Base URL for path resolution (relative to repo root) */
baseUrl: string;
}
/** Go module config parsed from go.mod */
export interface GoModuleConfig {
/** Module path (e.g., "github.com/user/repo") */
modulePath: string;
}
/** PHP Composer PSR-4 autoload config */
export interface ComposerConfig {
/** Map of namespace prefix -> directory (e.g., "App\\" -> "app/") */
psr4: Map<string, string>;
}
/** C# project config parsed from .csproj files */
export interface CSharpProjectConfig {
/** Root namespace from <RootNamespace> or assembly name (default: project directory name) */
rootNamespace: string;
/** Directory containing the .csproj file */
projectDir: string;
}
/** Swift Package Manager module config */
export interface SwiftPackageConfig {
/** Map of target name -> source directory path (e.g., "SiuperModel" -> "Package/Sources/SiuperModel") */
targets: Map<string, string>;
}
// ============================================================================
// LANGUAGE-SPECIFIC CONFIG LOADERS
// ============================================================================
/**
* Parse tsconfig.json to extract path aliases.
* Tries tsconfig.json, tsconfig.app.json, tsconfig.base.json in order.
*/
export async function loadTsconfigPaths(repoRoot: string): Promise<TsconfigPaths | null> {
const candidates = ['tsconfig.json', 'tsconfig.app.json', 'tsconfig.base.json'];
for (const filename of candidates) {
try {
const tsconfigPath = path.join(repoRoot, filename);
const raw = await fs.readFile(tsconfigPath, 'utf-8');
// Strip JSON comments (// and /* */ style) for robustness
const stripped = raw.replace(/\/\/.*$/gm, '').replace(/\/\*[\s\S]*?\*\//g, '');
const tsconfig = JSON.parse(stripped);
const compilerOptions = tsconfig.compilerOptions;
if (!compilerOptions?.paths) continue;
const baseUrl = compilerOptions.baseUrl || '.';
const aliases = new Map<string, string>();
for (const [pattern, targets] of Object.entries(compilerOptions.paths)) {
if (!Array.isArray(targets) || targets.length === 0) continue;
const target = targets[0] as string;
// Convert glob patterns: "@/*" -> "@/", "src/*" -> "src/"
const aliasPrefix = pattern.endsWith('/*') ? pattern.slice(0, -1) : pattern;
const targetPrefix = target.endsWith('/*') ? target.slice(0, -1) : target;
aliases.set(aliasPrefix, targetPrefix);
}
if (aliases.size > 0) {
if (isDev) {
console.log(`📦 Loaded ${aliases.size} path aliases from ${filename}`);
}
return { aliases, baseUrl };
}
} catch {
// File doesn't exist or isn't valid JSON - try next
}
}
return null;
}
/**
* Parse go.mod to extract module path.
*/
export async function loadGoModulePath(repoRoot: string): Promise<GoModuleConfig | null> {
try {
const goModPath = path.join(repoRoot, 'go.mod');
const content = await fs.readFile(goModPath, 'utf-8');
const match = content.match(/^module\s+(\S+)/m);
if (match) {
if (isDev) {
console.log(`📦 Loaded Go module path: ${match[1]}`);
}
return { modulePath: match[1] };
}
} catch {
// No go.mod
}
return null;
}
/** Parse composer.json to extract PSR-4 autoload mappings (including autoload-dev). */
export async function loadComposerConfig(repoRoot: string): Promise<ComposerConfig | null> {
try {
const composerPath = path.join(repoRoot, 'composer.json');
const raw = await fs.readFile(composerPath, 'utf-8');
const composer = JSON.parse(raw);
const psr4Raw = composer.autoload?.['psr-4'] ?? {};
const psr4Dev = composer['autoload-dev']?.['psr-4'] ?? {};
const merged = { ...psr4Raw, ...psr4Dev };
const psr4 = new Map<string, string>();
for (const [ns, dir] of Object.entries(merged)) {
const nsNorm = (ns as string).replace(/\\+$/, '');
const dirNorm = (dir as string).replace(/\\/g, '/').replace(/\/+$/, '');
psr4.set(nsNorm, dirNorm);
}
if (isDev) {
console.log(`📦 Loaded ${psr4.size} PSR-4 mappings from composer.json`);
}
return { psr4 };
} catch {
return null;
}
}
/**
* Parse .csproj files to extract RootNamespace.
* Scans the repo root for .csproj files and returns configs for each.
*/
export async function loadCSharpProjectConfig(repoRoot: string): Promise<CSharpProjectConfig[]> {
const configs: CSharpProjectConfig[] = [];
// BFS scan for .csproj files up to 5 levels deep, cap at 100 dirs to avoid runaway scanning
const scanQueue: { dir: string; depth: number }[] = [{ dir: repoRoot, depth: 0 }];
const maxDepth = 5;
const maxDirs = 100;
let dirsScanned = 0;
while (scanQueue.length > 0 && dirsScanned < maxDirs) {
const { dir, depth } = scanQueue.shift()!;
dirsScanned++;
try {
const entries = await fs.readdir(dir, { withFileTypes: true });
for (const entry of entries) {
if (entry.isDirectory() && depth < maxDepth) {
// Skip common non-project directories
if (entry.name === 'node_modules' || entry.name === '.git' || entry.name === 'bin' || entry.name === 'obj') continue;
scanQueue.push({ dir: path.join(dir, entry.name), depth: depth + 1 });
}
if (entry.isFile() && entry.name.endsWith('.csproj')) {
try {
const csprojPath = path.join(dir, entry.name);
const content = await fs.readFile(csprojPath, 'utf-8');
const nsMatch = content.match(/<RootNamespace>\s*([^<]+)\s*<\/RootNamespace>/);
const rootNamespace = nsMatch
? nsMatch[1].trim()
: entry.name.replace(/\.csproj$/, '');
const projectDir = path.relative(repoRoot, dir).replace(/\\/g, '/');
configs.push({ rootNamespace, projectDir });
if (isDev) {
console.log(`📦 Loaded C# project: ${entry.name} (namespace: ${rootNamespace}, dir: ${projectDir})`);
}
} catch {
// Can't read .csproj
}
}
}
} catch {
// Can't read directory
}
}
return configs;
}
export async function loadSwiftPackageConfig(repoRoot: string): Promise<SwiftPackageConfig | null> {
// Swift imports are module-name based (e.g., `import SiuperModel`)
// SPM convention: Sources/<TargetName>/ or Package/Sources/<TargetName>/
// We scan for these directories to build a target map
const targets = new Map<string, string>();
const sourceDirs = ['Sources', 'Package/Sources', 'src'];
for (const sourceDir of sourceDirs) {
try {
const fullPath = path.join(repoRoot, sourceDir);
const entries = await fs.readdir(fullPath, { withFileTypes: true });
for (const entry of entries) {
if (entry.isDirectory()) {
targets.set(entry.name, sourceDir + '/' + entry.name);
}
}
} catch {
// Directory doesn't exist
}
}
if (targets.size > 0) {
if (isDev) {
console.log(`📦 Loaded ${targets.size} Swift package targets`);
}
return { targets };
}
return null;
}
@@ -0,0 +1,465 @@
/**
* MRO (Method Resolution Order) Processor
*
* Walks the inheritance DAG (EXTENDS/IMPLEMENTS edges), collects methods from
* each ancestor via HAS_METHOD edges, detects method-name collisions across
* parents, and applies language-specific resolution rules to emit OVERRIDES edges.
*
* Language-specific rules:
* - C++: leftmost base class in declaration order wins
* - C#/Java: class method wins over interface default; multiple interface
* methods with same name are ambiguous (null resolution)
* - Python: C3 linearization determines MRO; first in linearized order wins
* - Rust: no auto-resolution — requires qualified syntax, resolvedTo = null
* - Default: single inheritance — first definition wins
*
* OVERRIDES edge direction: Class → Method (not Method → Method).
* The source is the child class that inherits conflicting methods,
* the target is the winning ancestor method node.
* Cypher: MATCH (c:Class)-[r:CodeRelation {type: 'OVERRIDES'}]->(m:Method)
*/
import { KnowledgeGraph, GraphRelationship } from '../graph/types.js';
import { generateId } from '../../lib/utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
// ---------------------------------------------------------------------------
// Public types
// ---------------------------------------------------------------------------
export interface MROEntry {
classId: string;
className: string;
language: SupportedLanguages;
mro: string[]; // linearized parent names
ambiguities: MethodAmbiguity[];
}
export interface MethodAmbiguity {
methodName: string;
definedIn: Array<{ classId: string; className: string; methodId: string }>;
resolvedTo: string | null; // winning methodId or null if truly ambiguous
reason: string;
}
export interface MROResult {
entries: MROEntry[];
overrideEdges: number;
ambiguityCount: number;
}
// ---------------------------------------------------------------------------
// Internal helpers
// ---------------------------------------------------------------------------
/** Collect EXTENDS, IMPLEMENTS, and HAS_METHOD adjacency from the graph. */
function buildAdjacency(graph: KnowledgeGraph) {
// parentMap: childId → parentIds[] (in insertion / declaration order)
const parentMap = new Map<string, string[]>();
// methodMap: classId → methodIds[]
const methodMap = new Map<string, string[]>();
// Track which edge type each parent link came from
const parentEdgeType = new Map<string, Map<string, 'EXTENDS' | 'IMPLEMENTS'>>();
graph.forEachRelationship((rel) => {
if (rel.type === 'EXTENDS' || rel.type === 'IMPLEMENTS') {
let parents = parentMap.get(rel.sourceId);
if (!parents) {
parents = [];
parentMap.set(rel.sourceId, parents);
}
parents.push(rel.targetId);
let edgeTypes = parentEdgeType.get(rel.sourceId);
if (!edgeTypes) {
edgeTypes = new Map();
parentEdgeType.set(rel.sourceId, edgeTypes);
}
edgeTypes.set(rel.targetId, rel.type);
}
if (rel.type === 'HAS_METHOD') {
let methods = methodMap.get(rel.sourceId);
if (!methods) {
methods = [];
methodMap.set(rel.sourceId, methods);
}
methods.push(rel.targetId);
}
});
return { parentMap, methodMap, parentEdgeType };
}
/**
* Gather all ancestor IDs in BFS / topological order.
* Returns the linearized list of ancestor IDs (excluding the class itself).
*/
function gatherAncestors(
classId: string,
parentMap: Map<string, string[]>,
): string[] {
const visited = new Set<string>();
const order: string[] = [];
const queue: string[] = [...(parentMap.get(classId) ?? [])];
while (queue.length > 0) {
const id = queue.shift()!;
if (visited.has(id)) continue;
visited.add(id);
order.push(id);
const grandparents = parentMap.get(id);
if (grandparents) {
for (const gp of grandparents) {
if (!visited.has(gp)) queue.push(gp);
}
}
}
return order;
}
// ---------------------------------------------------------------------------
// C3 linearization (Python MRO)
// ---------------------------------------------------------------------------
/**
* Compute C3 linearization for a class given a parentMap.
* Returns an array of ancestor IDs in C3 order (excluding the class itself),
* or null if linearization fails (inconsistent or cyclic hierarchy).
*/
function c3Linearize(
classId: string,
parentMap: Map<string, string[]>,
cache: Map<string, string[] | null>,
inProgress?: Set<string>,
): string[] | null {
if (cache.has(classId)) return cache.get(classId)!;
// Cycle detection: if we're already computing this class, the hierarchy is cyclic
const visiting = inProgress ?? new Set<string>();
if (visiting.has(classId)) {
cache.set(classId, null);
return null;
}
visiting.add(classId);
const directParents = parentMap.get(classId);
if (!directParents || directParents.length === 0) {
visiting.delete(classId);
cache.set(classId, []);
return [];
}
// Compute linearization for each parent first
const parentLinearizations: string[][] = [];
for (const pid of directParents) {
const pLin = c3Linearize(pid, parentMap, cache, visiting);
if (pLin === null) {
visiting.delete(classId);
cache.set(classId, null);
return null;
}
parentLinearizations.push([pid, ...pLin]);
}
// Add the direct parents list as the final sequence
const sequences = [...parentLinearizations, [...directParents]];
const result: string[] = [];
while (sequences.some(s => s.length > 0)) {
// Find a good head: one that doesn't appear in the tail of any other sequence
let head: string | null = null;
for (const seq of sequences) {
if (seq.length === 0) continue;
const candidate = seq[0];
const inTail = sequences.some(
other => other.length > 1 && other.indexOf(candidate, 1) !== -1
);
if (!inTail) {
head = candidate;
break;
}
}
if (head === null) {
// Inconsistent hierarchy
visiting.delete(classId);
cache.set(classId, null);
return null;
}
result.push(head);
// Remove the chosen head from all sequences
for (const seq of sequences) {
if (seq.length > 0 && seq[0] === head) {
seq.shift();
}
}
}
visiting.delete(classId);
cache.set(classId, result);
return result;
}
// ---------------------------------------------------------------------------
// Language-specific resolution
// ---------------------------------------------------------------------------
type MethodDef = { classId: string; className: string; methodId: string };
type Resolution = { resolvedTo: string | null; reason: string };
/** Resolve by MRO order — first ancestor in linearized order wins. */
function resolveByMroOrder(
methodName: string,
defs: MethodDef[],
mroOrder: string[],
reasonPrefix: string,
): Resolution {
for (const ancestorId of mroOrder) {
const match = defs.find(d => d.classId === ancestorId);
if (match) {
return {
resolvedTo: match.methodId,
reason: `${reasonPrefix}: ${match.className}::${methodName}`,
};
}
}
return { resolvedTo: defs[0].methodId, reason: `${reasonPrefix} fallback: first definition` };
}
function resolveCsharpJava(
methodName: string,
defs: MethodDef[],
parentEdgeTypes: Map<string, 'EXTENDS' | 'IMPLEMENTS'> | undefined,
): Resolution {
const classDefs: MethodDef[] = [];
const interfaceDefs: MethodDef[] = [];
for (const def of defs) {
const edgeType = parentEdgeTypes?.get(def.classId);
if (edgeType === 'IMPLEMENTS') {
interfaceDefs.push(def);
} else {
classDefs.push(def);
}
}
if (classDefs.length > 0) {
return {
resolvedTo: classDefs[0].methodId,
reason: `class method wins: ${classDefs[0].className}::${methodName}`,
};
}
if (interfaceDefs.length > 1) {
return {
resolvedTo: null,
reason: `ambiguous: ${methodName} defined in multiple interfaces: ${interfaceDefs.map(d => d.className).join(', ')}`,
};
}
if (interfaceDefs.length === 1) {
return {
resolvedTo: interfaceDefs[0].methodId,
reason: `single interface default: ${interfaceDefs[0].className}::${methodName}`,
};
}
return { resolvedTo: null, reason: 'no resolution found' };
}
// ---------------------------------------------------------------------------
// Main entry point
// ---------------------------------------------------------------------------
export function computeMRO(graph: KnowledgeGraph): MROResult {
const { parentMap, methodMap, parentEdgeType } = buildAdjacency(graph);
const c3Cache = new Map<string, string[] | null>();
const entries: MROEntry[] = [];
let overrideEdges = 0;
let ambiguityCount = 0;
// Process every class that has at least one parent
for (const [classId, directParents] of parentMap) {
if (directParents.length === 0) continue;
const classNode = graph.getNode(classId);
if (!classNode) continue;
const language = classNode.properties.language;
if (!language) continue;
const className = classNode.properties.name;
// Compute linearized MRO depending on language
let mroOrder: string[];
if (language === SupportedLanguages.Python) {
const c3Result = c3Linearize(classId, parentMap, c3Cache);
mroOrder = c3Result ?? gatherAncestors(classId, parentMap);
} else {
mroOrder = gatherAncestors(classId, parentMap);
}
// Get the parent names for the MRO entry
const mroNames: string[] = mroOrder
.map(id => graph.getNode(id)?.properties.name)
.filter((n): n is string => n !== undefined);
// Collect methods from all ancestors, grouped by method name
const methodsByName = new Map<string, MethodDef[]>();
for (const ancestorId of mroOrder) {
const ancestorNode = graph.getNode(ancestorId);
if (!ancestorNode) continue;
const methods = methodMap.get(ancestorId) ?? [];
for (const methodId of methods) {
const methodNode = graph.getNode(methodId);
if (!methodNode) continue;
// Properties don't participate in method resolution order
if (methodNode.label === 'Property') continue;
const methodName = methodNode.properties.name;
let defs = methodsByName.get(methodName);
if (!defs) {
defs = [];
methodsByName.set(methodName, defs);
}
// Avoid duplicates (same method seen via multiple paths)
if (!defs.some(d => d.methodId === methodId)) {
defs.push({
classId: ancestorId,
className: ancestorNode.properties.name,
methodId,
});
}
}
}
// Detect collisions: methods defined in 2+ different ancestors
const ambiguities: MethodAmbiguity[] = [];
// Compute transitive edge types once per class (only needed for C#/Java)
const needsEdgeTypes = language === SupportedLanguages.CSharp || language === SupportedLanguages.Java || language === SupportedLanguages.Kotlin;
const classEdgeTypes = needsEdgeTypes
? buildTransitiveEdgeTypes(classId, parentMap, parentEdgeType)
: undefined;
for (const [methodName, defs] of methodsByName) {
if (defs.length < 2) continue;
// Own method shadows inherited — no ambiguity
const ownMethods = methodMap.get(classId) ?? [];
const ownDefinesIt = ownMethods.some(mid => {
const mn = graph.getNode(mid);
return mn?.properties.name === methodName;
});
if (ownDefinesIt) continue;
let resolution: Resolution;
switch (language) {
case SupportedLanguages.CPlusPlus:
resolution = resolveByMroOrder(methodName, defs, mroOrder, 'C++ leftmost base');
break;
case SupportedLanguages.CSharp:
case SupportedLanguages.Java:
case SupportedLanguages.Kotlin:
resolution = resolveCsharpJava(methodName, defs, classEdgeTypes);
break;
case SupportedLanguages.Python:
resolution = resolveByMroOrder(methodName, defs, mroOrder, 'Python C3 MRO');
break;
case SupportedLanguages.Rust:
resolution = {
resolvedTo: null,
reason: `Rust requires qualified syntax: <Type as Trait>::${methodName}()`,
};
break;
default:
resolution = resolveByMroOrder(methodName, defs, mroOrder, 'first definition');
break;
}
const ambiguity: MethodAmbiguity = {
methodName,
definedIn: defs,
resolvedTo: resolution.resolvedTo,
reason: resolution.reason,
};
ambiguities.push(ambiguity);
if (resolution.resolvedTo === null) {
ambiguityCount++;
}
// Emit OVERRIDES edge if resolution found
if (resolution.resolvedTo !== null) {
graph.addRelationship({
id: generateId('OVERRIDES', `${classId}->${resolution.resolvedTo}`),
sourceId: classId,
targetId: resolution.resolvedTo,
type: 'OVERRIDES',
confidence: 1.0,
reason: resolution.reason,
});
overrideEdges++;
}
}
entries.push({
classId,
className,
language,
mro: mroNames,
ambiguities,
});
}
return { entries, overrideEdges, ambiguityCount };
}
/**
* Build transitive edge types for a class using BFS from the class to all ancestors.
*
* Known limitation: BFS first-reach heuristic can misclassify an interface as
* EXTENDS if it's reachable via a class chain before being seen via IMPLEMENTS.
* E.g. if BaseClass also implements IFoo, IFoo may be classified as EXTENDS.
* This affects C#/Java/Kotlin conflict resolution in rare diamond hierarchies.
*/
function buildTransitiveEdgeTypes(
classId: string,
parentMap: Map<string, string[]>,
parentEdgeType: Map<string, Map<string, 'EXTENDS' | 'IMPLEMENTS'>>,
): Map<string, 'EXTENDS' | 'IMPLEMENTS'> {
const result = new Map<string, 'EXTENDS' | 'IMPLEMENTS'>();
const directEdges = parentEdgeType.get(classId);
if (!directEdges) return result;
// BFS: propagate edge type from direct parents
const queue: Array<{ id: string; edgeType: 'EXTENDS' | 'IMPLEMENTS' }> = [];
const directParents = parentMap.get(classId) ?? [];
for (const pid of directParents) {
const et = directEdges.get(pid) ?? 'EXTENDS';
if (!result.has(pid)) {
result.set(pid, et);
queue.push({ id: pid, edgeType: et });
}
}
while (queue.length > 0) {
const { id, edgeType } = queue.shift()!;
const grandparents = parentMap.get(id) ?? [];
for (const gp of grandparents) {
if (!result.has(gp)) {
result.set(gp, edgeType);
queue.push({ id: gp, edgeType });
}
}
}
return result;
}
@@ -0,0 +1,384 @@
import { SupportedLanguages } from '../../config/supported-languages.js';
import type { SymbolTable, SymbolDefinition } from './symbol-table.js';
import type { NamedImportMap } from './import-processor.js';
/**
* Walk a named-binding re-export chain through NamedImportMap.
*
* When file A imports { User } from B, and B re-exports { User } from C,
* the NamedImportMap for A points to B, but B has no User definition.
* This function follows the chain: A→B→C until a definition is found.
*
* Returns the definitions found at the end of the chain, or null if the
* chain breaks (missing binding, circular reference, or depth exceeded).
* Max depth 5 to prevent infinite loops.
*
* @param allDefs Pre-computed `symbolTable.lookupFuzzy(name)` result — must be the
* complete unfiltered result. Passing a file-filtered subset will cause
* silent misses at depth=0 for non-aliased bindings.
*/
export function walkBindingChain(
name: string,
currentFilePath: string,
symbolTable: SymbolTable,
namedImportMap: NamedImportMap,
allDefs: SymbolDefinition[],
): SymbolDefinition[] | null {
let lookupFile = currentFilePath;
let lookupName = name;
const visited = new Set<string>();
for (let depth = 0; depth < 5; depth++) {
const bindings = namedImportMap.get(lookupFile);
if (!bindings) return null;
const binding = bindings.get(lookupName);
if (!binding) return null;
const key = `${binding.sourcePath}:${binding.exportedName}`;
if (visited.has(key)) return null; // circular
visited.add(key);
const targetName = binding.exportedName;
const resolvedDefs = targetName !== lookupName || depth > 0
? symbolTable.lookupFuzzy(targetName).filter(def => def.filePath === binding.sourcePath)
: allDefs.filter(def => def.filePath === binding.sourcePath);
if (resolvedDefs.length > 0) return resolvedDefs;
// No definition in source file → follow re-export chain
lookupFile = binding.sourcePath;
lookupName = targetName;
}
return null;
}
/**
* Extract named bindings from an import AST node.
* Returns undefined if the import is not a named import (e.g., import * or default).
*
* TS: import { User, Repo as R } from './models'
* → [{local:'User', exported:'User'}, {local:'R', exported:'Repo'}]
*
* Python: from models import User, Repo as R
* → [{local:'User', exported:'User'}, {local:'R', exported:'Repo'}]
*/
export function extractNamedBindings(
importNode: any,
language: SupportedLanguages,
): { local: string; exported: string }[] | undefined {
if (language === SupportedLanguages.TypeScript || language === SupportedLanguages.JavaScript) {
return extractTsNamedBindings(importNode);
}
if (language === SupportedLanguages.Python) {
return extractPythonNamedBindings(importNode);
}
if (language === SupportedLanguages.Kotlin) {
return extractKotlinNamedBindings(importNode);
}
if (language === SupportedLanguages.Rust) {
return extractRustNamedBindings(importNode);
}
if (language === SupportedLanguages.PHP) {
return extractPhpNamedBindings(importNode);
}
if (language === SupportedLanguages.CSharp) {
return extractCsharpNamedBindings(importNode);
}
if (language === SupportedLanguages.Java) {
return extractJavaNamedBindings(importNode);
}
return undefined;
}
export function extractTsNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// import_statement > import_clause > named_imports > import_specifier*
const importClause = findChild(importNode, 'import_clause');
if (importClause) {
const namedImports = findChild(importClause, 'named_imports');
if (!namedImports) return undefined; // default import, namespace import, or side-effect
const bindings: { local: string; exported: string }[] = [];
for (let i = 0; i < namedImports.namedChildCount; i++) {
const specifier = namedImports.namedChild(i);
if (specifier?.type !== 'import_specifier') continue;
const identifiers: string[] = [];
for (let j = 0; j < specifier.namedChildCount; j++) {
const child = specifier.namedChild(j);
if (child?.type === 'identifier') identifiers.push(child.text);
}
if (identifiers.length === 1) {
bindings.push({ local: identifiers[0], exported: identifiers[0] });
} else if (identifiers.length === 2) {
// import { Foo as Bar } → exported='Foo', local='Bar'
bindings.push({ local: identifiers[1], exported: identifiers[0] });
}
}
return bindings.length > 0 ? bindings : undefined;
}
// Re-export: export { X } from './y' → export_statement > export_clause > export_specifier
const exportClause = findChild(importNode, 'export_clause');
if (exportClause) {
const bindings: { local: string; exported: string }[] = [];
for (let i = 0; i < exportClause.namedChildCount; i++) {
const specifier = exportClause.namedChild(i);
if (specifier?.type !== 'export_specifier') continue;
const identifiers: string[] = [];
for (let j = 0; j < specifier.namedChildCount; j++) {
const child = specifier.namedChild(j);
if (child?.type === 'identifier') identifiers.push(child.text);
}
if (identifiers.length === 1) {
// export { User } from './base' → re-exports User as User
bindings.push({ local: identifiers[0], exported: identifiers[0] });
} else if (identifiers.length === 2) {
// export { Repo as Repository } from './models' → name=Repo, alias=Repository
// For re-exports, the first id is the source name, second is what's exported
// When another file imports { Repository }, they get Repo from the source
bindings.push({ local: identifiers[1], exported: identifiers[0] });
}
}
return bindings.length > 0 ? bindings : undefined;
}
return undefined;
}
export function extractPythonNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// Only from import_from_statement, not plain import_statement
if (importNode.type !== 'import_from_statement') return undefined;
const bindings: { local: string; exported: string }[] = [];
for (let i = 0; i < importNode.namedChildCount; i++) {
const child = importNode.namedChild(i);
if (!child) continue;
if (child.type === 'dotted_name') {
// Skip the module_name (first dotted_name is the source module)
const fieldName = importNode.childForFieldName?.('module_name');
if (fieldName && child.startIndex === fieldName.startIndex) continue;
// This is an imported name: from x import User
const name = child.text;
if (name) bindings.push({ local: name, exported: name });
}
if (child.type === 'aliased_import') {
// from x import Repo as R
const dottedName = findChild(child, 'dotted_name');
const aliasIdent = findChild(child, 'identifier');
if (dottedName && aliasIdent) {
bindings.push({ local: aliasIdent.text, exported: dottedName.text });
}
}
}
return bindings.length > 0 ? bindings : undefined;
}
export function extractKotlinNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// import_header > identifier + import_alias > simple_identifier
if (importNode.type !== 'import_header') return undefined;
const fullIdent = findChild(importNode, 'identifier');
if (!fullIdent) return undefined;
const fullText = fullIdent.text;
const exportedName = fullText.includes('.') ? fullText.split('.').pop()! : fullText;
const importAlias = findChild(importNode, 'import_alias');
if (importAlias) {
// Aliased: import com.example.User as U
const aliasIdent = findChild(importAlias, 'simple_identifier');
if (!aliasIdent) return undefined;
return [{ local: aliasIdent.text, exported: exportedName }];
}
// Non-aliased: import com.example.User → local="User", exported="User"
// Skip wildcard imports (ending in *)
if (fullText.endsWith('.*') || fullText.endsWith('*')) return undefined;
// Skip lowercase last segments — those are member/function imports (e.g.,
// import util.OneArg.writeAudit), not class imports. Multiple member imports
// with the same function name would collide in NamedImportMap, breaking
// arity-based disambiguation.
if (exportedName[0] && exportedName[0] === exportedName[0].toLowerCase()) return undefined;
return [{ local: exportedName, exported: exportedName }];
}
export function extractRustNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// use_declaration may contain use_as_clause at any depth
if (importNode.type !== 'use_declaration') return undefined;
const bindings: { local: string; exported: string }[] = [];
collectRustBindings(importNode, bindings);
return bindings.length > 0 ? bindings : undefined;
}
function collectRustBindings(node: any, bindings: { local: string; exported: string }[]): void {
if (node.type === 'use_as_clause') {
// First identifier = exported name, second identifier = local alias
const idents: string[] = [];
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'identifier') idents.push(child.text);
// For scoped_identifier, extract the last segment
if (child?.type === 'scoped_identifier') {
const nameNode = child.childForFieldName?.('name');
if (nameNode) idents.push(nameNode.text);
}
}
if (idents.length === 2) {
bindings.push({ local: idents[1], exported: idents[0] });
}
return;
}
// Terminal identifier in a use_list: use crate::models::{User, Repo}
if (node.type === 'identifier' && node.parent?.type === 'use_list') {
bindings.push({ local: node.text, exported: node.text });
return;
}
// Skip scoped_identifier that serves as path prefix in scoped_use_list
// e.g. use crate::models::{User, Repo} — the path node "crate::models" is not an importable symbol
if (node.type === 'scoped_identifier' && node.parent?.type === 'scoped_use_list') {
return; // path prefix — the use_list sibling handles the actual symbols
}
// Terminal scoped_identifier: use crate::models::User;
// Only extract if this is a leaf (no deeper use_list/use_as_clause/scoped_use_list)
if (node.type === 'scoped_identifier') {
let hasDeeper = false;
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'use_list' || child?.type === 'use_as_clause' || child?.type === 'scoped_use_list') {
hasDeeper = true;
break;
}
}
if (!hasDeeper) {
const nameNode = node.childForFieldName?.('name');
if (nameNode) {
bindings.push({ local: nameNode.text, exported: nameNode.text });
}
return;
}
}
// Recurse into children
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child) collectRustBindings(child, bindings);
}
}
export function extractPhpNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// namespace_use_declaration > namespace_use_clause* (flat)
// namespace_use_declaration > namespace_use_group > namespace_use_clause* (grouped)
if (importNode.type !== 'namespace_use_declaration') return undefined;
const bindings: { local: string; exported: string }[] = [];
// Collect all clauses — from direct children AND from namespace_use_group
const clauses: any[] = [];
for (let i = 0; i < importNode.namedChildCount; i++) {
const child = importNode.namedChild(i);
if (child?.type === 'namespace_use_clause') {
clauses.push(child);
} else if (child?.type === 'namespace_use_group') {
for (let j = 0; j < child.namedChildCount; j++) {
const groupChild = child.namedChild(j);
if (groupChild?.type === 'namespace_use_clause') clauses.push(groupChild);
}
}
}
for (const clause of clauses) {
// Flat imports: qualified_name + name (alias)
let qualifiedName: any = null;
const names: any[] = [];
for (let j = 0; j < clause.namedChildCount; j++) {
const child = clause.namedChild(j);
if (child?.type === 'qualified_name') qualifiedName = child;
else if (child?.type === 'name') names.push(child);
}
if (qualifiedName && names.length > 0) {
// Flat aliased import: use App\Models\Repo as R;
const fullText = qualifiedName.text;
const exportedName = fullText.includes('\\') ? fullText.split('\\').pop()! : fullText;
bindings.push({ local: names[0].text, exported: exportedName });
} else if (qualifiedName && names.length === 0) {
// Flat non-aliased import: use App\Models\User;
const fullText = qualifiedName.text;
const lastSegment = fullText.includes('\\') ? fullText.split('\\').pop()! : fullText;
bindings.push({ local: lastSegment, exported: lastSegment });
} else if (!qualifiedName && names.length >= 2) {
// Grouped aliased import: {Repo as R} — first name = exported, second = alias
bindings.push({ local: names[1].text, exported: names[0].text });
} else if (!qualifiedName && names.length === 1) {
// Grouped non-aliased import: {User} in use App\Models\{User, Repo as R}
bindings.push({ local: names[0].text, exported: names[0].text });
}
}
return bindings.length > 0 ? bindings : undefined;
}
export function extractCsharpNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// using_directive with identifier (alias) + qualified_name (target)
if (importNode.type !== 'using_directive') return undefined;
let aliasIdent: any = null;
let qualifiedName: any = null;
for (let i = 0; i < importNode.namedChildCount; i++) {
const child = importNode.namedChild(i);
if (child?.type === 'identifier' && !aliasIdent) aliasIdent = child;
else if (child?.type === 'qualified_name') qualifiedName = child;
}
if (!aliasIdent || !qualifiedName) return undefined;
const fullText = qualifiedName.text;
const exportedName = fullText.includes('.') ? fullText.split('.').pop()! : fullText;
return [{ local: aliasIdent.text, exported: exportedName }];
}
export function extractJavaNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// import_declaration > scoped_identifier "com.example.models.User"
// Wildcard imports (.*) don't produce named bindings
if (importNode.type !== 'import_declaration') return undefined;
// Check for asterisk (wildcard import) — skip those
for (let i = 0; i < importNode.childCount; i++) {
const child = importNode.child(i);
if (child?.type === 'asterisk') return undefined;
}
const scopedId = findChild(importNode, 'scoped_identifier');
if (!scopedId) return undefined;
const fullText = scopedId.text;
const lastDot = fullText.lastIndexOf('.');
if (lastDot === -1) return undefined;
const className = fullText.slice(lastDot + 1);
// Skip lowercase names — those are package imports, not class imports
if (className[0] && className[0] === className[0].toLowerCase()) return undefined;
return [{ local: className, exported: className }];
}
function findChild(node: any, type: string): any {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === type) return child;
}
return null;
}
+44 -184
View File
@@ -5,10 +5,12 @@ import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { generateId } from '../../lib/utils.js';
import { SymbolTable } from './symbol-table.js';
import { ASTCache } from './ast-cache.js';
import { findSiblingChild, getLanguageFromFilename, yieldToEventLoop } from './utils.js';
import { getLanguageFromFilename, yieldToEventLoop, DEFINITION_CAPTURE_KEYS, getDefinitionNodeFromCaptures, findEnclosingClassId, extractMethodSignature } from './utils.js';
import { isNodeExported } from './export-detection.js';
import { detectFrameworkFromAST } from './framework-detection.js';
import { WorkerPool } from './workers/worker-pool.js';
import type { ParseWorkerResult, ParseWorkerInput, ExtractedImport, ExtractedCall, ExtractedHeritage, ExtractedRoute } from './workers/parse-worker.js';
import { getTreeSitterBufferSize, TREE_SITTER_MAX_BUFFER } from './constants.js';
export type FileProgressCallback = (current: number, total: number, filePath: string) => void;
@@ -19,183 +21,9 @@ export interface WorkerExtractedData {
routes: ExtractedRoute[];
}
const DEFINITION_CAPTURE_KEYS = [
'definition.function',
'definition.class',
'definition.interface',
'definition.method',
'definition.struct',
'definition.enum',
'definition.namespace',
'definition.module',
'definition.trait',
'definition.impl',
'definition.type',
'definition.const',
'definition.static',
'definition.typedef',
'definition.macro',
'definition.union',
'definition.property',
'definition.record',
'definition.delegate',
'definition.annotation',
'definition.constructor',
'definition.template',
] as const;
const getDefinitionNodeFromCaptures = (captureMap: Record<string, any>): any | null => {
for (const key of DEFINITION_CAPTURE_KEYS) {
if (captureMap[key]) return captureMap[key];
}
return null;
};
// ============================================================================
// EXPORT DETECTION - Language-specific visibility detection
// ============================================================================
/**
* Check if a symbol (function, class, etc.) is exported/public
* Handles all 9 supported languages with explicit logic
*
* @param node - The AST node for the symbol name
* @param name - The symbol name
* @param language - The programming language
* @returns true if the symbol is exported/public
*/
export const isNodeExported = (node: any, name: string, language: string): boolean => {
let current = node;
switch (language) {
// JavaScript/TypeScript: Check for export keyword in ancestors
case 'javascript':
case 'typescript':
while (current) {
const type = current.type;
if (type === 'export_statement' ||
type === 'export_specifier' ||
type === 'lexical_declaration' && current.parent?.type === 'export_statement') {
return true;
}
// Also check if text starts with 'export '
if (current.text?.startsWith('export ')) {
return true;
}
current = current.parent;
}
return false;
// Python: Public if no leading underscore (convention)
case 'python':
return !name.startsWith('_');
// Java: Check for 'public' modifier
// In tree-sitter Java, modifiers are siblings of the name node, not parents
case 'java':
while (current) {
// Check if this node or any sibling is a 'modifiers' node containing 'public'
if (current.parent) {
const parent = current.parent;
// Check all children of the parent for modifiers
for (let i = 0; i < parent.childCount; i++) {
const child = parent.child(i);
if (child?.type === 'modifiers' && child.text?.includes('public')) {
return true;
}
}
// Also check if the parent's text starts with 'public' (fallback)
if (parent.type === 'method_declaration' || parent.type === 'constructor_declaration') {
if (parent.text?.trimStart().startsWith('public')) {
return true;
}
}
}
current = current.parent;
}
return false;
// C#: Check for 'public' modifier in ancestors
case 'csharp':
while (current) {
if (current.type === 'modifier' || current.type === 'modifiers') {
if (current.text?.includes('public')) return true;
}
current = current.parent;
}
return false;
// Go: Uppercase first letter = exported
case 'go':
if (name.length === 0) return false;
const first = name[0];
// Must be uppercase letter (not a number or symbol)
return first === first.toUpperCase() && first !== first.toLowerCase();
// Rust: Check for 'pub' visibility modifier
case 'rust':
while (current) {
if (current.type === 'visibility_modifier') {
if (current.text?.includes('pub')) return true;
}
current = current.parent;
}
return false;
// Kotlin: Default visibility is public (unlike Java)
// visibility_modifier is inside modifiers, a sibling of the name node within the declaration
case 'kotlin':
while (current) {
if (current.parent) {
const visMod = findSiblingChild(current.parent, 'modifiers', 'visibility_modifier');
if (visMod) {
const text = visMod.text;
if (text === 'private' || text === 'internal' || text === 'protected') return false;
if (text === 'public') return true;
}
}
current = current.parent;
}
// No visibility modifier = public (Kotlin default)
return true;
// C/C++: No native export concept at language level
// Entry points will be detected via name patterns (main, etc.)
case 'c':
case 'cpp':
return false;
// Swift: Check for 'public' or 'open' access modifiers
case 'swift':
while (current) {
if (current.type === 'modifiers' || current.type === 'visibility_modifier') {
const text = current.text || '';
if (text.includes('public') || text.includes('open')) return true;
}
current = current.parent;
}
return false;
// PHP: Check for visibility modifier or top-level scope
case 'php':
while (current) {
if (current.type === 'class_declaration' ||
current.type === 'interface_declaration' ||
current.type === 'trait_declaration' ||
current.type === 'enum_declaration') {
return true;
}
if (current.type === 'visibility_modifier') {
return current.text === 'public';
}
current = current.parent;
}
return true; // Top-level functions are globally accessible
default:
return false;
}
};
// isNodeExported imported from ./export-detection.js (shared module)
// Re-export for backward compatibility with any external consumers
export { isNodeExported } from './export-detection.js';
// ============================================================================
// Worker-based parallel parsing
@@ -247,7 +75,10 @@ const processParsingWithWorkers = async (
}
for (const sym of result.symbols) {
symbolTable.add(sym.filePath, sym.name, sym.nodeId, sym.type);
symbolTable.add(sym.filePath, sym.name, sym.nodeId, sym.type, {
parameterCount: sym.parameterCount,
ownerId: sym.ownerId,
});
}
allImports.push(...result.imports);
@@ -286,8 +117,8 @@ const processParsingSequential = async (
if (!language) continue;
// Skip very large files — they can crash tree-sitter or cause OOM
if (file.content.length > 512 * 1024) continue;
// Skip files larger than the max tree-sitter buffer (32 MB)
if (file.content.length > TREE_SITTER_MAX_BUFFER) continue;
try {
await loadLanguage(language, file.path);
@@ -297,7 +128,7 @@ const processParsingSequential = async (
let tree;
try {
tree = parser.parse(file.content, undefined, { bufferSize: 1024 * 256 });
tree = parser.parse(file.content, undefined, { bufferSize: getTreeSitterBufferSize(file.content.length) });
} catch (parseError) {
console.warn(`Skipping unparseable file: ${file.path}`);
continue;
@@ -368,13 +199,18 @@ const processParsingSequential = async (
const definitionNodeForRange = getDefinitionNodeFromCaptures(captureMap);
const startLine = definitionNodeForRange ? definitionNodeForRange.startPosition.row : (nameNode ? nameNode.startPosition.row : 0);
const nodeId = generateId(nodeLabel, `${file.path}:${nodeName}:${startLine}`);
const nodeId = generateId(nodeLabel, `${file.path}:${nodeName}`);
const definitionNode = getDefinitionNodeFromCaptures(captureMap);
const frameworkHint = definitionNode
? detectFrameworkFromAST(language, (definitionNode.text || '').slice(0, 300))
: null;
// Extract method signature for Method/Constructor nodes
const methodSig = (nodeLabel === 'Function' || nodeLabel === 'Method' || nodeLabel === 'Constructor')
? extractMethodSignature(definitionNode)
: undefined;
const node: GraphNode = {
id: nodeId,
label: nodeLabel as any,
@@ -389,12 +225,24 @@ const processParsingSequential = async (
astFrameworkMultiplier: frameworkHint.entryPointMultiplier,
astFrameworkReason: frameworkHint.reason,
} : {}),
...(methodSig ? {
parameterCount: methodSig.parameterCount,
returnType: methodSig.returnType,
} : {}),
},
};
graph.addNode(node);
symbolTable.add(file.path, nodeName, nodeId, nodeLabel);
// Compute enclosing class for Method/Constructor/Property/Function — used for both ownerId and HAS_METHOD
// Function is included because Kotlin/Rust/Python capture class methods as Function nodes
const needsOwner = nodeLabel === 'Method' || nodeLabel === 'Constructor' || nodeLabel === 'Property' || nodeLabel === 'Function';
const enclosingClassId = needsOwner ? findEnclosingClassId(nameNode || definitionNodeForRange, file.path) : null;
symbolTable.add(file.path, nodeName, nodeId, nodeLabel, {
parameterCount: methodSig?.parameterCount,
ownerId: enclosingClassId ?? undefined,
});
const fileId = generateId('File', file.path);
@@ -410,6 +258,18 @@ const processParsingSequential = async (
};
graph.addRelationship(relationship);
// ── HAS_METHOD: link method/constructor/property to enclosing class ──
if (enclosingClassId) {
graph.addRelationship({
id: generateId('HAS_METHOD', `${enclosingClassId}->${nodeId}`),
sourceId: enclosingClassId,
targetId: nodeId,
type: 'HAS_METHOD',
confidence: 1.0,
reason: '',
});
}
});
}
};
+56 -23
View File
@@ -1,9 +1,17 @@
import { createKnowledgeGraph } from '../graph/graph.js';
import { processStructure } from './structure-processor.js';
import { processParsing } from './parsing-processor.js';
import { processImports, processImportsFromExtracted, createImportMap, buildImportResolutionContext } from './import-processor.js';
import {
processImports,
processImportsFromExtracted,
createImportMap,
createPackageMap,
createNamedImportMap,
buildImportResolutionContext
} from './import-processor.js';
import { processCalls, processCallsFromExtracted, processRoutesFromExtracted } from './call-processor.js';
import { processHeritage, processHeritageFromExtracted } from './heritage-processor.js';
import { computeMRO } from './mro-processor.js';
import { processCommunities } from './community-processor.js';
import { processProcesses } from './process-processor.js';
import { createSymbolTable } from './symbol-table.js';
@@ -36,6 +44,8 @@ export const runPipelineFromRepo = async (
const symbolTable = createSymbolTable();
let astCache = createASTCache(AST_CACHE_CAP);
const importMap = createImportMap();
const packageMap = createPackageMap();
const namedImportMap = createNamedImportMap();
const cleanup = () => {
astCache.clear();
@@ -213,21 +223,36 @@ export const runPipelineFromRepo = async (
if (chunkWorkerData) {
// Imports
await processImportsFromExtracted(graph, allPathObjects, chunkWorkerData.imports, importMap, undefined, repoPath, importCtx);
// Calls — resolve immediately, then free the array
if (chunkWorkerData.calls.length > 0) {
await processCallsFromExtracted(graph, chunkWorkerData.calls, symbolTable, importMap);
}
// Heritage — resolve immediately, then free
if (chunkWorkerData.heritage.length > 0) {
await processHeritageFromExtracted(graph, chunkWorkerData.heritage, symbolTable);
}
// Routes — resolve immediately (Laravel route→controller CALLS edges)
if (chunkWorkerData.routes && chunkWorkerData.routes.length > 0) {
await processRoutesFromExtracted(graph, chunkWorkerData.routes, symbolTable, importMap);
}
await processImportsFromExtracted(graph, allPathObjects, chunkWorkerData.imports, importMap, undefined, repoPath, importCtx, packageMap, namedImportMap);
// Calls + Heritage + Routes — resolve in parallel (no shared mutable state between them)
// This is safe because each writes disjoint relationship types into idempotent id-keyed Maps,
// and the single-threaded event loop prevents races between synchronous addRelationship calls.
await Promise.all([
processCallsFromExtracted(
graph,
chunkWorkerData.calls,
symbolTable, importMap,
packageMap,
undefined,
namedImportMap
),
processHeritageFromExtracted(
graph,
chunkWorkerData.heritage,
symbolTable,
importMap,
packageMap
),
processRoutesFromExtracted(
graph,
chunkWorkerData.routes ?? [],
symbolTable,
importMap,
packageMap
),
]);
} else {
await processImports(graph, chunkFiles, astCache, importMap, undefined, repoPath, allPaths);
await processImports(graph, chunkFiles, astCache, importMap, undefined, repoPath, allPaths, packageMap, namedImportMap);
sequentialChunkPaths.push(chunkPaths);
}
@@ -248,8 +273,11 @@ export const runPipelineFromRepo = async (
.filter(p => chunkContents.has(p))
.map(p => ({ path: p, content: chunkContents.get(p)! }));
astCache = createASTCache(chunkFiles.length);
await processCalls(graph, chunkFiles, astCache, symbolTable, importMap);
await processHeritage(graph, chunkFiles, astCache, symbolTable);
const rubyHeritage = await processCalls(graph, chunkFiles, astCache, symbolTable, importMap, packageMap, undefined, namedImportMap);
await processHeritage(graph, chunkFiles, astCache, symbolTable, importMap, packageMap);
if (rubyHeritage.length > 0) {
await processHeritageFromExtracted(graph, rubyHeritage, symbolTable, importMap, packageMap);
}
astCache.clear();
}
@@ -260,12 +288,17 @@ export const runPipelineFromRepo = async (
(importCtx as any).suffixIndex = null;
(importCtx as any).normalizedFileList = null;
if (isDev) {
let importsCount = 0;
for (const r of graph.iterRelationships()) {
if (r.type === 'IMPORTS') importsCount++;
}
console.log(`📊 Pipeline: graph has ${importsCount} IMPORTS, ${graph.relationshipCount} total relationships`);
// ── Phase 4.5: Method Resolution Order ──────────────────────────────
onProgress({
phase: 'parsing',
percent: 81,
message: 'Computing method resolution order...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
const mroResult = computeMRO(graph);
if (isDev && mroResult.entries.length > 0) {
console.log(`🔀 MRO: ${mroResult.entries.length} classes analyzed, ${mroResult.ambiguityCount} ambiguities found, ${mroResult.overrideEdges} OVERRIDES edges`);
}
// ── Phase 5: Communities ───────────────────────────────────────────
@@ -13,6 +13,7 @@
import { KnowledgeGraph, GraphNode, GraphRelationship, NodeLabel } from '../graph/types.js';
import { CommunityMembership } from './community-processor.js';
import { calculateEntryPointScore, isTestFile } from './entry-point-scoring.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
const isDev = process.env.NODE_ENV === 'development';
@@ -287,7 +288,7 @@ const findEntryPoints = (
// Calculate entry point score using new scoring system
const { score: baseScore, reasons } = calculateEntryPointScore(
node.properties.name,
node.properties.language || 'javascript',
node.properties.language ?? SupportedLanguages.JavaScript,
node.properties.isExported ?? false,
callers.length,
callees.length,
@@ -0,0 +1,128 @@
/**
* C# namespace import resolution.
* Handles using-directive resolution via .csproj root namespace stripping.
*/
import type { SuffixIndex } from './utils.js';
import { suffixResolve } from './utils.js';
/** C# project config parsed from .csproj files */
export interface CSharpProjectConfig {
/** Root namespace from <RootNamespace> or assembly name (default: project directory name) */
rootNamespace: string;
/** Directory containing the .csproj file */
projectDir: string;
}
/**
* Resolve a C# using-directive import path to matching .cs files.
* Tries single-file match first, then directory match for namespace imports.
*/
export function resolveCSharpImport(
importPath: string,
csharpConfigs: CSharpProjectConfig[],
normalizedFileList: string[],
allFileList: string[],
index?: SuffixIndex,
): string[] {
const namespacePath = importPath.replace(/\./g, '/');
const results: string[] = [];
for (const config of csharpConfigs) {
const nsPath = config.rootNamespace.replace(/\./g, '/');
let relative: string;
if (namespacePath.startsWith(nsPath + '/')) {
relative = namespacePath.slice(nsPath.length + 1);
} else if (namespacePath === nsPath) {
// The import IS the root namespace — resolve to all .cs files in project root
relative = '';
} else {
continue;
}
const dirPrefix = config.projectDir
? (relative ? config.projectDir + '/' + relative : config.projectDir)
: relative;
// 1. Try as single file: relative.cs (e.g., "Models/DlqMessage.cs")
if (relative) {
const candidate = dirPrefix + '.cs';
if (index) {
const result = index.get(candidate) || index.getInsensitive(candidate);
if (result) return [result];
}
// Also try suffix match
const suffixResult = index?.get(relative + '.cs') || index?.getInsensitive(relative + '.cs');
if (suffixResult) return [suffixResult];
}
// 2. Try as directory: all .cs files directly inside (namespace import)
if (index) {
const dirFiles = index.getFilesInDir(dirPrefix, '.cs');
for (const f of dirFiles) {
const normalized = f.replace(/\\/g, '/');
// Check it's a direct child by finding the dirPrefix and ensuring no deeper slashes
const prefixIdx = normalized.indexOf(dirPrefix + '/');
if (prefixIdx < 0) continue;
const afterDir = normalized.substring(prefixIdx + dirPrefix.length + 1);
if (!afterDir.includes('/')) {
results.push(f);
}
}
if (results.length > 0) return results;
}
// 3. Linear scan fallback for directory matching
if (results.length === 0) {
const dirTrail = dirPrefix + '/';
for (let i = 0; i < normalizedFileList.length; i++) {
const normalized = normalizedFileList[i];
if (!normalized.endsWith('.cs')) continue;
const prefixIdx = normalized.indexOf(dirTrail);
if (prefixIdx < 0) continue;
const afterDir = normalized.substring(prefixIdx + dirTrail.length);
if (!afterDir.includes('/')) {
results.push(allFileList[i]);
}
}
if (results.length > 0) return results;
}
}
// Fallback: suffix matching without namespace stripping (single file)
const pathParts = namespacePath.split('/').filter(Boolean);
const fallback = suffixResolve(pathParts, normalizedFileList, allFileList, index);
return fallback ? [fallback] : [];
}
/**
* Compute the directory suffix for a C# namespace import (for PackageMap).
* Returns a suffix like "/ProjectDir/Models/" or null if no config matches.
*/
export function resolveCSharpNamespaceDir(
importPath: string,
csharpConfigs: CSharpProjectConfig[],
): string | null {
const namespacePath = importPath.replace(/\./g, '/');
for (const config of csharpConfigs) {
const nsPath = config.rootNamespace.replace(/\./g, '/');
let relative: string;
if (namespacePath.startsWith(nsPath + '/')) {
relative = namespacePath.slice(nsPath.length + 1);
} else if (namespacePath === nsPath) {
relative = '';
} else {
continue;
}
const dirPrefix = config.projectDir
? (relative ? config.projectDir + '/' + relative : config.projectDir)
: relative;
if (!dirPrefix) continue;
return '/' + dirPrefix + '/';
}
return null;
}
@@ -0,0 +1,58 @@
/**
* Go package import resolution.
* Handles Go module path-based package imports.
*/
/** Go module config parsed from go.mod */
export interface GoModuleConfig {
/** Module path (e.g., "github.com/user/repo") */
modulePath: string;
}
/**
* Extract the package directory suffix from a Go import path.
* Returns the suffix string (e.g., "/internal/auth/") or null if invalid.
*/
export function resolveGoPackageDir(
importPath: string,
goModule: GoModuleConfig,
): string | null {
if (!importPath.startsWith(goModule.modulePath)) return null;
const relativePkg = importPath.slice(goModule.modulePath.length + 1);
if (!relativePkg) return null;
return '/' + relativePkg + '/';
}
/**
* Resolve a Go internal package import to all .go files in the package directory.
* Returns an array of file paths.
*/
export function resolveGoPackage(
importPath: string,
goModule: GoModuleConfig,
normalizedFileList: string[],
allFileList: string[],
): string[] {
if (!importPath.startsWith(goModule.modulePath)) return [];
// Strip module path to get relative package path
const relativePkg = importPath.slice(goModule.modulePath.length + 1); // e.g., "internal/auth"
if (!relativePkg) return [];
const pkgSuffix = '/' + relativePkg + '/';
const matches: string[] = [];
for (let i = 0; i < normalizedFileList.length; i++) {
// Prepend '/' so paths like "internal/auth/service.go" match suffix "/internal/auth/"
const normalized = '/' + normalizedFileList[i];
// File must be directly in the package directory (not a subdirectory)
if (normalized.includes(pkgSuffix) && normalized.endsWith('.go') && !normalized.endsWith('_test.go')) {
const afterPkg = normalized.substring(normalized.indexOf(pkgSuffix) + pkgSuffix.length);
if (!afterPkg.includes('/')) {
matches.push(allFileList[i]);
}
}
}
return matches;
}
@@ -0,0 +1,26 @@
/**
* Language-specific import resolvers.
* Extracted from import-processor.ts for maintainability.
*/
export { EXTENSIONS, tryResolveWithExtensions, buildSuffixIndex, suffixResolve } from './utils.js';
export type { SuffixIndex } from './utils.js';
export { KOTLIN_EXTENSIONS, appendKotlinWildcard, resolveJvmWildcard, resolveJvmMemberImport } from './jvm.js';
export { resolveGoPackageDir, resolveGoPackage } from './go.js';
export type { GoModuleConfig } from './go.js';
export { resolveCSharpImport, resolveCSharpNamespaceDir } from './csharp.js';
export type { CSharpProjectConfig } from './csharp.js';
export { resolvePhpImport } from './php.js';
export type { ComposerConfig } from './php.js';
export { resolveRustImport, tryRustModulePath } from './rust.js';
export { resolveRubyImport, extractRubyImportPath } from './ruby.js';
export type { GemfileConfig } from './ruby.js';
export { resolveImportPath, RESOLVE_CACHE_CAP } from './standard.js';
export type { TsconfigPaths } from './standard.js';
@@ -0,0 +1,106 @@
/**
* JVM import resolution (Java + Kotlin).
* Handles wildcard imports, member/static imports, and Kotlin-specific patterns.
*/
import type { SuffixIndex } from './utils.js';
/** Kotlin file extensions for JVM resolver reuse */
export const KOTLIN_EXTENSIONS: readonly string[] = ['.kt', '.kts'];
/**
* Append .* to a Kotlin import path if the AST has a wildcard_import sibling node.
* Pure function — returns a new string without mutating the input.
*/
export const appendKotlinWildcard = (importPath: string, importNode: any): string => {
for (let i = 0; i < importNode.childCount; i++) {
if (importNode.child(i)?.type === 'wildcard_import') {
return importPath.endsWith('.*') ? importPath : `${importPath}.*`;
}
}
return importPath;
};
/**
* Resolve a JVM wildcard import (com.example.*) to all matching files.
* Works for both Java (.java) and Kotlin (.kt, .kts).
*/
export function resolveJvmWildcard(
importPath: string,
normalizedFileList: string[],
allFileList: string[],
extensions: readonly string[],
index?: SuffixIndex,
): string[] {
// "com.example.util.*" -> "com/example/util"
const packagePath = importPath.slice(0, -2).replace(/\./g, '/');
if (index) {
const candidates = extensions.flatMap(ext => index.getFilesInDir(packagePath, ext));
// Filter to only direct children (no subdirectories)
const packageSuffix = '/' + packagePath + '/';
return candidates.filter(f => {
const normalized = f.replace(/\\/g, '/');
const idx = normalized.indexOf(packageSuffix);
if (idx < 0) return false;
const afterPkg = normalized.substring(idx + packageSuffix.length);
return !afterPkg.includes('/');
});
}
// Fallback: linear scan
const packageSuffix = '/' + packagePath + '/';
const matches: string[] = [];
for (let i = 0; i < normalizedFileList.length; i++) {
const normalized = normalizedFileList[i];
if (normalized.includes(packageSuffix) &&
extensions.some(ext => normalized.endsWith(ext))) {
const afterPackage = normalized.substring(normalized.indexOf(packageSuffix) + packageSuffix.length);
if (!afterPackage.includes('/')) {
matches.push(allFileList[i]);
}
}
}
return matches;
}
/**
* Try to resolve a JVM member/static import by stripping the member name.
* Java: "com.example.Constants.VALUE" -> resolve "com.example.Constants"
* Kotlin: "com.example.Constants.VALUE" -> resolve "com.example.Constants"
*/
export function resolveJvmMemberImport(
importPath: string,
normalizedFileList: string[],
allFileList: string[],
extensions: readonly string[],
index?: SuffixIndex,
): string | null {
// Member imports: com.example.Constants.VALUE or com.example.Constants.*
// The last segment is a member name if it starts with lowercase, is ALL_CAPS, or is a wildcard
const segments = importPath.split('.');
if (segments.length < 3) return null;
const lastSeg = segments[segments.length - 1];
if (lastSeg === '*' || /^[a-z]/.test(lastSeg) || /^[A-Z_]+$/.test(lastSeg)) {
const classPath = segments.slice(0, -1).join('/');
for (const ext of extensions) {
const classSuffix = classPath + ext;
if (index) {
const result = index.get(classSuffix) || index.getInsensitive(classSuffix);
if (result) return result;
} else {
const fullSuffix = '/' + classSuffix;
for (let i = 0; i < normalizedFileList.length; i++) {
if (normalizedFileList[i].endsWith(fullSuffix) ||
normalizedFileList[i].toLowerCase().endsWith(fullSuffix.toLowerCase())) {
return allFileList[i];
}
}
}
}
}
return null;
}
@@ -0,0 +1,51 @@
/**
* PHP PSR-4 import resolution.
* Handles use-statement resolution via composer.json autoload mappings.
*/
import type { SuffixIndex } from './utils.js';
import { suffixResolve } from './utils.js';
/** PHP Composer PSR-4 autoload config */
export interface ComposerConfig {
/** Map of namespace prefix -> directory (e.g., "App\\" -> "app/") */
psr4: Map<string, string>;
}
/**
* Resolve a PHP use-statement import path using PSR-4 mappings.
* e.g. "App\Http\Controllers\UserController" -> "app/Http/Controllers/UserController.php"
*/
export function resolvePhpImport(
importPath: string,
composerConfig: ComposerConfig | null,
allFiles: Set<string>,
normalizedFileList: string[],
allFileList: string[],
index?: SuffixIndex,
): string | null {
// Normalize: replace backslashes with forward slashes
const normalized = importPath.replace(/\\/g, '/');
// Try PSR-4 resolution if composer.json was found
if (composerConfig) {
// Sort namespaces by length descending (longest match wins)
const sorted = [...composerConfig.psr4.entries()].sort((a, b) => b[0].length - a[0].length);
for (const [nsPrefix, dirPrefix] of sorted) {
const nsPrefixSlash = nsPrefix.replace(/\\/g, '/');
if (normalized.startsWith(nsPrefixSlash + '/') || normalized === nsPrefixSlash) {
const remainder = normalized.slice(nsPrefixSlash.length).replace(/^\//, '');
const filePath = dirPrefix + (remainder ? '/' + remainder : '') + '.php';
if (allFiles.has(filePath)) return filePath;
if (index) {
const result = index.getInsensitive(filePath);
if (result) return result;
}
}
}
}
// Fallback: suffix matching (works without composer.json)
const pathParts = normalized.split('/').filter(Boolean);
return suffixResolve(pathParts, normalizedFileList, allFileList, index);
}
@@ -0,0 +1,54 @@
/**
* Ruby require/require_relative import resolution.
* Handles path resolution for Ruby's require and require_relative calls.
*/
import type { SuffixIndex } from './utils.js';
import { suffixResolve } from './utils.js';
/** Bundler config parsed from Gemfile (future: gem path resolution) */
export interface GemfileConfig {
gemPaths: Map<string, string>;
}
/**
* Resolve a Ruby require/require_relative path to a matching .rb file.
*
* require_relative paths are pre-normalized to './' prefix by the caller.
* require paths use suffix matching (gem-style paths like 'json', 'net/http').
*/
export function resolveRubyImport(
sourceFile: string,
importPath: string,
isRelative: boolean,
normalizedFileList: string[],
allFileList: string[],
index?: SuffixIndex,
): string | null {
const pathParts = importPath.replace(/^\.\//, '').split('/').filter(Boolean);
return suffixResolve(pathParts, normalizedFileList, allFileList, index);
}
/**
* Extract the import path string from a Ruby call AST node.
* Returns null if the call is not a require/require_relative.
*/
export function extractRubyImportPath(
calledName: string,
callNode: any,
): { importPath: string; isRelative: boolean } | null {
if (calledName !== 'require' && calledName !== 'require_relative') return null;
const argList = callNode.childForFieldName?.('arguments');
const stringNode = argList?.children?.find((c: any) => c.type === 'string');
const contentNode = stringNode?.children?.find((c: any) => c.type === 'string_content');
if (!contentNode) return null;
let importPath = contentNode.text;
const isRelative = calledName === 'require_relative';
if (isRelative && !importPath.startsWith('.')) {
importPath = './' + importPath;
}
return { importPath, isRelative };
}
@@ -0,0 +1,82 @@
/**
* Rust module import resolution.
* Handles crate::, super::, self:: prefix paths and :: separators.
*/
/**
* Resolve Rust use-path to a file.
* Handles crate::, super::, self:: prefixes and :: path separators.
*/
export function resolveRustImport(
currentFile: string,
importPath: string,
allFiles: Set<string>,
): string | null {
let rustPath: string;
if (importPath.startsWith('crate::')) {
// crate:: resolves from src/ directory (standard Rust layout)
rustPath = importPath.slice(7).replace(/::/g, '/');
// Try from src/ (standard layout)
const fromSrc = tryRustModulePath('src/' + rustPath, allFiles);
if (fromSrc) return fromSrc;
// Try from repo root (non-standard)
const fromRoot = tryRustModulePath(rustPath, allFiles);
if (fromRoot) return fromRoot;
return null;
}
if (importPath.startsWith('super::')) {
// super:: = parent directory of current file's module
const currentDir = currentFile.split('/').slice(0, -1);
currentDir.pop(); // Go up one level for super::
rustPath = importPath.slice(7).replace(/::/g, '/');
const fullPath = [...currentDir, rustPath].join('/');
return tryRustModulePath(fullPath, allFiles);
}
if (importPath.startsWith('self::')) {
// self:: = current module's directory
const currentDir = currentFile.split('/').slice(0, -1);
rustPath = importPath.slice(6).replace(/::/g, '/');
const fullPath = [...currentDir, rustPath].join('/');
return tryRustModulePath(fullPath, allFiles);
}
// Bare path without prefix (e.g., from a use in a nested module)
// Convert :: to / and try suffix matching
if (importPath.includes('::')) {
rustPath = importPath.replace(/::/g, '/');
return tryRustModulePath(rustPath, allFiles);
}
return null;
}
/**
* Try to resolve a Rust module path to a file.
* Tries: path.rs, path/mod.rs, and with the last segment stripped
* (last segment might be a symbol name, not a module).
*/
export function tryRustModulePath(modulePath: string, allFiles: Set<string>): string | null {
// Try direct: path.rs
if (allFiles.has(modulePath + '.rs')) return modulePath + '.rs';
// Try directory: path/mod.rs
if (allFiles.has(modulePath + '/mod.rs')) return modulePath + '/mod.rs';
// Try path/lib.rs (for crate root)
if (allFiles.has(modulePath + '/lib.rs')) return modulePath + '/lib.rs';
// The last segment might be a symbol (function, struct, etc.), not a module.
// Strip it and try again.
const lastSlash = modulePath.lastIndexOf('/');
if (lastSlash > 0) {
const parentPath = modulePath.substring(0, lastSlash);
if (allFiles.has(parentPath + '.rs')) return parentPath + '.rs';
if (allFiles.has(parentPath + '/mod.rs')) return parentPath + '/mod.rs';
}
return null;
}
@@ -0,0 +1,177 @@
/**
* Standard import path resolution.
* Handles relative imports, path alias rewriting, and generic suffix matching.
* Used as the fallback when language-specific resolvers don't match.
*/
import type { SuffixIndex } from './utils.js';
import { tryResolveWithExtensions, suffixResolve } from './utils.js';
import { resolveRustImport } from './rust.js';
import { SupportedLanguages } from '../../../config/supported-languages.js';
/** TypeScript path alias config parsed from tsconfig.json */
export interface TsconfigPaths {
/** Map of alias prefix -> target prefix (e.g., "@/" -> "src/") */
aliases: Map<string, string>;
/** Base URL for path resolution (relative to repo root) */
baseUrl: string;
}
/** Max entries in the resolve cache. Beyond this, entries are evicted.
* 100K entries ≈ 15MB — covers the most common import patterns. */
export const RESOLVE_CACHE_CAP = 100_000;
/**
* Resolve an import path to a file path in the repository.
*
* Language-specific preprocessing is applied before the generic resolution:
* - TypeScript/JavaScript: rewrites tsconfig path aliases
* - Rust: converts crate::/super::/self:: to relative paths
*
* Java wildcards and Go package imports are handled separately in processImports
* because they resolve to multiple files.
*/
export const resolveImportPath = (
currentFile: string,
importPath: string,
allFiles: Set<string>,
allFileList: string[],
normalizedFileList: string[],
resolveCache: Map<string, string | null>,
language: SupportedLanguages,
tsconfigPaths: TsconfigPaths | null,
index?: SuffixIndex,
): string | null => {
const cacheKey = `${currentFile}::${importPath}`;
if (resolveCache.has(cacheKey)) return resolveCache.get(cacheKey) ?? null;
const cache = (result: string | null): string | null => {
// Evict oldest 20% when cap is reached instead of clearing all
if (resolveCache.size >= RESOLVE_CACHE_CAP) {
const evictCount = Math.floor(RESOLVE_CACHE_CAP * 0.2);
const iter = resolveCache.keys();
for (let i = 0; i < evictCount; i++) {
const key = iter.next().value;
if (key !== undefined) resolveCache.delete(key);
}
}
resolveCache.set(cacheKey, result);
return result;
};
// ---- TypeScript/JavaScript: rewrite path aliases ----
if (
(language === SupportedLanguages.TypeScript || language === SupportedLanguages.JavaScript) &&
tsconfigPaths &&
!importPath.startsWith('.')
) {
for (const [aliasPrefix, targetPrefix] of tsconfigPaths.aliases) {
if (importPath.startsWith(aliasPrefix)) {
const remainder = importPath.slice(aliasPrefix.length);
// Build the rewritten path relative to baseUrl
const rewritten = tsconfigPaths.baseUrl === '.'
? targetPrefix + remainder
: tsconfigPaths.baseUrl + '/' + targetPrefix + remainder;
// Try direct resolution from repo root
const resolved = tryResolveWithExtensions(rewritten, allFiles);
if (resolved) return cache(resolved);
// Try suffix matching as fallback
const parts = rewritten.split('/').filter(Boolean);
const suffixResult = suffixResolve(parts, normalizedFileList, allFileList, index);
if (suffixResult) return cache(suffixResult);
}
}
}
// ---- Rust: convert module path syntax to file paths ----
if (language === SupportedLanguages.Rust) {
// Handle grouped imports: use crate::module::{Foo, Bar, Baz}
// Extract the prefix path before ::{...} and resolve the module, not the symbols
let rustImportPath = importPath;
const braceIdx = importPath.indexOf('::{');
if (braceIdx !== -1) {
rustImportPath = importPath.substring(0, braceIdx);
} else if (importPath.startsWith('{') && importPath.endsWith('}')) {
// Top-level grouped imports: use {crate::a, crate::b}
// Iterate each part and return the first that resolves. This function returns a single
// string, so callers that need ALL edges must intercept before reaching here (see the
// Rust grouped-import blocks in processImports / processImportsBatch). This fallback
// handles any path that reaches resolveImportPath directly.
const inner = importPath.slice(1, -1);
const parts = inner.split(',').map(p => p.trim()).filter(Boolean);
for (const part of parts) {
const partResult = resolveRustImport(currentFile, part, allFiles);
if (partResult) return cache(partResult);
}
return cache(null);
}
const rustResult = resolveRustImport(currentFile, rustImportPath, allFiles);
if (rustResult) return cache(rustResult);
// Fall through to generic resolution if Rust-specific didn't match
}
// ---- Python relative imports (PEP 328): .module, ..module, ... ----
if (language === SupportedLanguages.Python && importPath.startsWith('.')) {
const dotMatch = importPath.match(/^(\.+)(.*)/);
if (dotMatch) {
const dotCount = dotMatch[1].length;
const modulePart = dotMatch[2]; // e.g., "models" from ".models"
const dirParts = currentFile.split('/').slice(0, -1); // remove filename
// Navigate up: 1 dot = same package, 2 dots = parent package, etc.
// First dot means "current package", each additional dot goes up one level
for (let i = 1; i < dotCount; i++) {
dirParts.pop();
}
if (modulePart) {
// from .models import User → resolve "models" relative to current package
const modulePath = modulePart.replace(/\./g, '/');
dirParts.push(...modulePath.split('/'));
}
const basePath = dirParts.join('/');
const resolved = tryResolveWithExtensions(basePath, allFiles);
return cache(resolved);
}
}
// ---- Generic relative import resolution (./ and ../) ----
const currentDir = currentFile.split('/').slice(0, -1);
const parts = importPath.split('/');
for (const part of parts) {
if (part === '.') continue;
if (part === '..') {
currentDir.pop();
} else {
currentDir.push(part);
}
}
const basePath = currentDir.join('/');
if (importPath.startsWith('.')) {
const resolved = tryResolveWithExtensions(basePath, allFiles);
return cache(resolved);
}
// ---- Generic package/absolute import resolution (suffix matching) ----
// Java wildcards are handled in processImports, not here
if (importPath.endsWith('.*')) {
return cache(null);
}
// C/C++ includes use actual file paths (e.g. "animal.h") — don't convert dots to slashes
const isCpp = language === SupportedLanguages.C || language === SupportedLanguages.CPlusPlus;
const pathLike = importPath.includes('/') || isCpp
? importPath
: importPath.replace(/\./g, '/');
const pathParts = pathLike.split('/').filter(Boolean);
const resolved = suffixResolve(pathParts, normalizedFileList, allFileList, index);
return cache(resolved);
};
@@ -0,0 +1,158 @@
/**
* Shared utilities for import resolution.
* Extracted from import-processor.ts to reduce file size.
*/
/** All file extensions to try during resolution */
export const EXTENSIONS = [
'',
// TypeScript/JavaScript
'.tsx', '.ts', '.jsx', '.js', '/index.tsx', '/index.ts', '/index.jsx', '/index.js',
// Python
'.py', '/__init__.py',
// Java
'.java',
// Kotlin
'.kt', '.kts',
// C/C++
'.c', '.h', '.cpp', '.hpp', '.cc', '.cxx', '.hxx', '.hh',
// C#
'.cs',
// Go
'.go',
// Rust
'.rs', '/mod.rs',
// PHP
'.php', '.phtml',
// Swift
'.swift',
// Ruby
'.rb',
];
/**
* Try to match a path (with extensions) against the known file set.
* Returns the matched file path or null.
*/
export function tryResolveWithExtensions(
basePath: string,
allFiles: Set<string>,
): string | null {
for (const ext of EXTENSIONS) {
const candidate = basePath + ext;
if (allFiles.has(candidate)) return candidate;
}
return null;
}
/**
* Build a suffix index for O(1) endsWith lookups.
* Maps every possible path suffix to its original file path.
* e.g. for "src/com/example/Foo.java":
* "Foo.java" -> "src/com/example/Foo.java"
* "example/Foo.java" -> "src/com/example/Foo.java"
* "com/example/Foo.java" -> "src/com/example/Foo.java"
* etc.
*/
export interface SuffixIndex {
/** Exact suffix lookup (case-sensitive) */
get(suffix: string): string | undefined;
/** Case-insensitive suffix lookup */
getInsensitive(suffix: string): string | undefined;
/** Get all files in a directory suffix */
getFilesInDir(dirSuffix: string, extension: string): string[];
}
export function buildSuffixIndex(normalizedFileList: string[], allFileList: string[]): SuffixIndex {
// Map: normalized suffix -> original file path
const exactMap = new Map<string, string>();
// Map: lowercase suffix -> original file path
const lowerMap = new Map<string, string>();
// Map: directory suffix -> list of file paths in that directory
const dirMap = new Map<string, string[]>();
for (let i = 0; i < normalizedFileList.length; i++) {
const normalized = normalizedFileList[i];
const original = allFileList[i];
const parts = normalized.split('/');
// Index all suffixes: "a/b/c.java" -> ["c.java", "b/c.java", "a/b/c.java"]
for (let j = parts.length - 1; j >= 0; j--) {
const suffix = parts.slice(j).join('/');
// Only store first match (longest path wins for ambiguous suffixes)
if (!exactMap.has(suffix)) {
exactMap.set(suffix, original);
}
const lower = suffix.toLowerCase();
if (!lowerMap.has(lower)) {
lowerMap.set(lower, original);
}
}
// Index directory membership
const lastSlash = normalized.lastIndexOf('/');
if (lastSlash >= 0) {
// Build all directory suffixes
const dirParts = parts.slice(0, -1);
const fileName = parts[parts.length - 1];
const ext = fileName.substring(fileName.lastIndexOf('.'));
for (let j = dirParts.length - 1; j >= 0; j--) {
const dirSuffix = dirParts.slice(j).join('/');
const key = `${dirSuffix}:${ext}`;
let list = dirMap.get(key);
if (!list) {
list = [];
dirMap.set(key, list);
}
list.push(original);
}
}
}
return {
get: (suffix: string) => exactMap.get(suffix),
getInsensitive: (suffix: string) => lowerMap.get(suffix.toLowerCase()),
getFilesInDir: (dirSuffix: string, extension: string) => {
return dirMap.get(`${dirSuffix}:${extension}`) || [];
},
};
}
/**
* Suffix-based resolution using index. O(1) per lookup instead of O(files).
*/
export function suffixResolve(
pathParts: string[],
normalizedFileList: string[],
allFileList: string[],
index?: SuffixIndex,
): string | null {
if (index) {
for (let i = 0; i < pathParts.length; i++) {
const suffix = pathParts.slice(i).join('/');
for (const ext of EXTENSIONS) {
const suffixWithExt = suffix + ext;
const result = index.get(suffixWithExt) || index.getInsensitive(suffixWithExt);
if (result) return result;
}
}
return null;
}
// Fallback: linear scan (for backward compatibility)
for (let i = 0; i < pathParts.length; i++) {
const suffix = pathParts.slice(i).join('/');
for (const ext of EXTENSIONS) {
const suffixWithExt = suffix + ext;
const suffixPattern = '/' + suffixWithExt;
const matchIdx = normalizedFileList.findIndex(filePath =>
filePath.endsWith(suffixPattern) || filePath.toLowerCase().endsWith(suffixPattern.toLowerCase())
);
if (matchIdx !== -1) {
return allFileList[matchIdx];
}
}
}
return null;
}
@@ -0,0 +1,99 @@
/**
* Shared Ruby call routing logic.
*
* Ruby expresses imports, heritage (mixins), and property definitions as
* method calls rather than syntax-level constructs. This module provides a
* single routing function used by the CLI call-processor, CLI parse-worker,
* and the web call-processor so that the classification logic lives in one
* place.
*/
// ── Result types ────────────────────────────────────────────────────────────
export type RubyCallRouting =
| { kind: 'import'; importPath: string; isRelative: boolean }
| { kind: 'heritage'; items: RubyHeritageItem[] }
| { kind: 'properties'; items: RubyPropertyItem[] }
| { kind: 'call' }
| { kind: 'skip' };
export interface RubyHeritageItem {
enclosingClass: string;
mixinName: string;
}
export interface RubyPropertyItem {
propName: string;
accessorType: string;
startLine: number;
endLine: number;
}
// ── Routing function ────────────────────────────────────────────────────────
/**
* Classify a Ruby call node and extract its semantic payload.
*
* @param calledName - The method name (e.g. 'require', 'include', 'attr_accessor')
* @param callNode - The tree-sitter `call` AST node
* @returns A discriminated union describing the call's semantic role
*/
export function routeRubyCall(calledName: string, callNode: any): RubyCallRouting {
// ── require / require_relative → import ─────────────────────────────────
if (calledName === 'require' || calledName === 'require_relative') {
const argList = callNode.childForFieldName?.('arguments');
const stringNode = argList?.children?.find((c: any) => c.type === 'string');
const contentNode = stringNode?.children?.find((c: any) => c.type === 'string_content');
if (!contentNode) return { kind: 'skip' };
let importPath: string = contentNode.text;
const isRelative = calledName === 'require_relative';
if (isRelative && !importPath.startsWith('.')) {
importPath = './' + importPath;
}
return { kind: 'import', importPath, isRelative };
}
// ── include / extend / prepend → heritage (mixin) ──────────────────────
if (calledName === 'include' || calledName === 'extend' || calledName === 'prepend') {
let enclosingClass: string | null = null;
let current = callNode.parent;
while (current) {
if (current.type === 'class' || current.type === 'module') {
const nameNode = current.childForFieldName?.('name');
if (nameNode) { enclosingClass = nameNode.text; break; }
}
current = current.parent;
}
if (!enclosingClass) return { kind: 'skip' };
const items: RubyHeritageItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of (argList?.children ?? [])) {
if (arg.type === 'constant' || arg.type === 'scope_resolution') {
items.push({ enclosingClass, mixinName: arg.text });
}
}
return items.length > 0 ? { kind: 'heritage', items } : { kind: 'skip' };
}
// ── attr_accessor / attr_reader / attr_writer → property definitions ───
if (calledName === 'attr_accessor' || calledName === 'attr_reader' || calledName === 'attr_writer') {
const items: RubyPropertyItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of (argList?.children ?? [])) {
if (arg.type === 'simple_symbol') {
items.push({
propName: arg.text.replace(/^:/, ''),
accessorType: calledName,
startLine: arg.startPosition.row,
endLine: arg.endPosition.row,
});
}
}
return items.length > 0 ? { kind: 'properties', items } : { kind: 'skip' };
}
// ── Everything else → regular call ─────────────────────────────────────
return { kind: 'call' };
}
@@ -0,0 +1,123 @@
/**
* Symbol Resolver
*
* Import-filtered candidate narrowing for bare identifier resolution.
* NOT FQN resolution — does not parse qualifiers (ns::Bar, com.foo.Bar).
*
* Shared between heritage-processor.ts and call-processor.ts.
*/
import type { SymbolTable, SymbolDefinition } from './symbol-table.js';
import type { ImportMap, PackageMap, NamedImportMap } from './import-processor.js';
import { isFileInPackageDir } from './import-processor.js';
import { walkBindingChain } from './named-binding-extraction.js';
/** Resolution tier for internal tracking, logging, and test assertions. */
export type ResolutionTier = 'same-file' | 'import-scoped' | 'unique-global';
/** Internal resolution result preserving tier metadata. */
export interface InternalResolution {
definition: SymbolDefinition;
tier: ResolutionTier;
candidateCount: number;
}
/**
* Resolve a bare identifier to its best-matching definition using import context.
*
* Resolution tiers (highest confidence first):
* 1. Same file (lookupExactFull — authoritative)
* 2. Import-scoped (lookupFuzzy filtered by importMap — acceptable)
* 3. Unique global (lookupFuzzy with exactly 1 match — acceptable fallback)
*
* If multiple global candidates remain after filtering, returns null.
* A wrong edge is worse than no edge.
*/
export const resolveSymbol = (
name: string,
currentFilePath: string,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
): SymbolDefinition | null => {
return resolveSymbolInternal(name, currentFilePath, symbolTable, importMap, packageMap, namedImportMap)?.definition ?? null;
};
/** Internal resolver preserving tier metadata for logging and test assertions. */
export const resolveSymbolInternal = (
name: string,
currentFilePath: string,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
): InternalResolution | null => {
// Tier 1: Same file — authoritative match
const localDef = symbolTable.lookupExactFull(currentFilePath, name);
if (localDef) return { definition: localDef, tier: 'same-file', candidateCount: 1 };
// Get all global definitions for subsequent tiers
const allDefs = symbolTable.lookupFuzzy(name);
// Tier 2a-named: Check named bindings BEFORE the empty-allDefs early return,
// because aliased imports (import { User as U }) mean lookupFuzzy('U') returns
// empty but we can resolve via the exported name.
if (namedImportMap) {
const result = resolveNamedBindingChain(name, currentFilePath, symbolTable, namedImportMap, allDefs);
if (result) return result;
}
if (allDefs.length === 0) return null;
// Tier 2a: Import-scoped — check if any definition is in a file imported by currentFile
const importedFiles = importMap.get(currentFilePath);
if (importedFiles) {
for (const def of allDefs) {
if (importedFiles.has(def.filePath)) {
return { definition: def, tier: 'import-scoped', candidateCount: allDefs.length };
}
}
}
// Tier 2b: Package-scoped — check if any definition is in a package/namespace dir imported by currentFile
// Used for Go packages and C# namespace imports to avoid ImportMap expansion bloat
const importedPackages = packageMap?.get(currentFilePath);
if (importedPackages) {
for (const def of allDefs) {
for (const dirSuffix of importedPackages) {
if (isFileInPackageDir(def.filePath, dirSuffix)) {
return { definition: def, tier: 'import-scoped', candidateCount: allDefs.length };
}
}
}
}
// Tier 3: Unique global — ONLY if exactly one candidate exists
// Ambiguous global matches are refused. A wrong edge is worse than no edge.
if (allDefs.length === 1) {
return { definition: allDefs[0], tier: 'unique-global', candidateCount: 1 };
}
// Ambiguous: multiple global candidates, no import or same-file match → refuse
return null;
};
/**
* Follow re-export chains through NamedImportMap.
* Delegates chain-walking to the shared walkBindingChain utility, then
* applies symbol-resolver semantics: exactly one match required.
*/
const resolveNamedBindingChain = (
name: string,
currentFilePath: string,
symbolTable: SymbolTable,
namedImportMap: NamedImportMap,
allDefs: SymbolDefinition[],
): InternalResolution | null => {
const defs = walkBindingChain(name, currentFilePath, symbolTable, namedImportMap, allDefs);
if (defs?.length === 1) {
return { definition: defs[0], tier: 'import-scoped', candidateCount: defs.length };
}
return null;
};
+45 -14
View File
@@ -2,13 +2,22 @@ export interface SymbolDefinition {
nodeId: string;
filePath: string;
type: string; // 'Function', 'Class', etc.
parameterCount?: number;
/** Links Method/Constructor to owning Class/Struct/Trait nodeId */
ownerId?: string;
}
export interface SymbolTable {
/**
* Register a new symbol definition
*/
add: (filePath: string, name: string, nodeId: string, type: string) => void;
add: (
filePath: string,
name: string,
nodeId: string,
type: string,
metadata?: { parameterCount?: number; ownerId?: string }
) => void;
/**
* High Confidence: Look for a symbol specifically inside a file
@@ -16,6 +25,12 @@ export interface SymbolTable {
*/
lookupExact: (filePath: string, name: string) => string | undefined;
/**
* High Confidence: Look for a symbol in a specific file, returning full definition.
* Includes type information needed for heritage resolution (Class vs Interface).
*/
lookupExactFull: (filePath: string, name: string) => SymbolDefinition | undefined;
/**
* Low Confidence: Look for a symbol anywhere in the project
* Used when imports are missing or for framework magic
@@ -34,32 +49,48 @@ export interface SymbolTable {
}
export const createSymbolTable = (): SymbolTable => {
// 1. File-Specific Index (The "Good" one)
// Structure: FilePath -> (SymbolName -> NodeID)
const fileIndex = new Map<string, Map<string, string>>();
// 1. File-Specific Index — stores full SymbolDefinition for O(1) lookupExactFull
// Structure: FilePath -> (SymbolName -> SymbolDefinition)
const fileIndex = new Map<string, Map<string, SymbolDefinition>>();
// 2. Global Reverse Index (The "Backup")
// Structure: SymbolName -> [List of Definitions]
const globalIndex = new Map<string, SymbolDefinition[]>();
const add = (filePath: string, name: string, nodeId: string, type: string) => {
// A. Add to File Index
const add = (
filePath: string,
name: string,
nodeId: string,
type: string,
metadata?: { parameterCount?: number; ownerId?: string }
) => {
const def: SymbolDefinition = {
nodeId,
filePath,
type,
...(metadata?.parameterCount !== undefined ? { parameterCount: metadata.parameterCount } : {}),
...(metadata?.ownerId !== undefined ? { ownerId: metadata.ownerId } : {}),
};
// A. Add to File Index (shared reference — zero additional memory)
if (!fileIndex.has(filePath)) {
fileIndex.set(filePath, new Map());
}
fileIndex.get(filePath)!.set(name, nodeId);
fileIndex.get(filePath)!.set(name, def);
// B. Add to Global Index
// B. Add to Global Index (same object reference)
if (!globalIndex.has(name)) {
globalIndex.set(name, []);
}
globalIndex.get(name)!.push({ nodeId, filePath, type });
globalIndex.get(name)!.push(def);
};
const lookupExact = (filePath: string, name: string): string | undefined => {
const fileSymbols = fileIndex.get(filePath);
if (!fileSymbols) return undefined;
return fileSymbols.get(name);
return fileIndex.get(filePath)?.get(name)?.nodeId;
};
const lookupExactFull = (filePath: string, name: string): SymbolDefinition | undefined => {
return fileIndex.get(filePath)?.get(name);
};
const lookupFuzzy = (name: string): SymbolDefinition[] => {
@@ -76,5 +107,5 @@ export const createSymbolTable = (): SymbolTable => {
globalIndex.clear();
};
return { add, lookupExact, lookupFuzzy, getStats, clear };
};
return { add, lookupExact, lookupExactFull, lookupFuzzy, getStats, clear };
};
@@ -47,6 +47,10 @@ export const TYPESCRIPT_QUERIES = `
(import_statement
source: (string) @import.source) @import
; Re-export statements: export { X } from './y'
(export_statement
source: (string) @import.source) @import
(call_expression
function: (identifier) @call.name) @call
@@ -54,6 +58,10 @@ export const TYPESCRIPT_QUERIES = `
function: (member_expression
property: (property_identifier) @call.name)) @call
; Constructor calls: new Foo()
(new_expression
constructor: (identifier) @call.name) @call
; Heritage queries - class extends
(class_declaration
name: (type_identifier) @heritage.class
@@ -69,7 +77,7 @@ export const TYPESCRIPT_QUERIES = `
(type_identifier) @heritage.implements))) @heritage.impl
`;
// JavaScript queries - works with tree-sitter-javascript
// JavaScript queries - works with tree-sitter-javascript
export const JAVASCRIPT_QUERIES = `
(class_declaration
name: (identifier) @name) @definition.class
@@ -105,6 +113,10 @@ export const JAVASCRIPT_QUERIES = `
(import_statement
source: (string) @import.source) @import
; Re-export statements: export { X } from './y'
(export_statement
source: (string) @import.source) @import
(call_expression
function: (identifier) @call.name) @call
@@ -112,6 +124,10 @@ export const JAVASCRIPT_QUERIES = `
function: (member_expression
property: (property_identifier) @call.name)) @call
; Constructor calls: new Foo()
(new_expression
constructor: (identifier) @call.name) @call
; Heritage queries - class extends (JavaScript uses different AST than TypeScript)
; In tree-sitter-javascript, class_heritage directly contains the parent identifier
(class_declaration
@@ -134,6 +150,9 @@ export const PYTHON_QUERIES = `
(import_from_statement
module_name: (dotted_name) @import.source) @import
(import_from_statement
module_name: (relative_import) @import.source) @import
(call
function: (identifier) @call.name) @call
@@ -167,6 +186,9 @@ export const JAVA_QUERIES = `
(method_invocation name: (identifier) @call.name) @call
(method_invocation object: (_) name: (identifier) @call.name) @call
; Constructor calls: new Foo()
(object_creation_expression type: (type_identifier) @call.name) @call
; Heritage - extends class
(class_declaration name: (identifier) @heritage.class
(superclass (type_identifier) @heritage.extends)) @heritage
@@ -178,10 +200,17 @@ export const JAVA_QUERIES = `
// C queries - works with tree-sitter-c
export const C_QUERIES = `
; Functions
; Functions (direct declarator)
(function_definition declarator: (function_declarator declarator: (identifier) @name)) @definition.function
(declaration declarator: (function_declarator declarator: (identifier) @name)) @definition.function
; Functions returning pointers (pointer_declarator wraps function_declarator)
(function_definition declarator: (pointer_declarator declarator: (function_declarator declarator: (identifier) @name))) @definition.function
(declaration declarator: (pointer_declarator declarator: (function_declarator declarator: (identifier) @name))) @definition.function
; Functions returning double pointers (nested pointer_declarator)
(function_definition declarator: (pointer_declarator declarator: (pointer_declarator declarator: (function_declarator declarator: (identifier) @name)))) @definition.function
; Structs, Unions, Enums, Typedefs
(struct_specifier name: (type_identifier) @name) @definition.struct
(union_specifier name: (type_identifier) @name) @definition.union
@@ -209,15 +238,26 @@ export const GO_QUERIES = `
; Types
(type_declaration (type_spec name: (type_identifier) @name type: (struct_type))) @definition.struct
(type_declaration (type_spec name: (type_identifier) @name type: (interface_type))) @definition.interface
(type_declaration (type_spec name: (type_identifier) @name)) @definition.type
; Imports
(import_declaration (import_spec path: (interpreted_string_literal) @import.source)) @import
(import_declaration (import_spec_list (import_spec path: (interpreted_string_literal) @import.source))) @import
; Struct embedding (anonymous fields = inheritance)
(type_declaration
(type_spec
name: (type_identifier) @heritage.class
type: (struct_type
(field_declaration_list
(field_declaration
type: (type_identifier) @heritage.extends))))) @definition.struct
; Calls
(call_expression function: (identifier) @call.name) @call
(call_expression function: (selector_expression field: (field_identifier) @call.name)) @call
; Struct literal construction: User{Name: "Alice"}
(composite_literal type: (type_identifier) @call.name) @call
`;
// C++ queries - works with tree-sitter-cpp
@@ -228,10 +268,46 @@ export const CPP_QUERIES = `
(namespace_definition name: (namespace_identifier) @name) @definition.namespace
(enum_specifier name: (type_identifier) @name) @definition.enum
; Functions & Methods
; Typedefs and unions (common in C-style headers and mixed C/C++ code)
(type_definition declarator: (type_identifier) @name) @definition.typedef
(union_specifier name: (type_identifier) @name) @definition.union
; Macros
(preproc_function_def name: (identifier) @name) @definition.macro
(preproc_def name: (identifier) @name) @definition.macro
; Functions & Methods (direct declarator)
(function_definition declarator: (function_declarator declarator: (identifier) @name)) @definition.function
(function_definition declarator: (function_declarator declarator: (qualified_identifier name: (identifier) @name))) @definition.method
; Functions/methods returning pointers (pointer_declarator wraps function_declarator)
(function_definition declarator: (pointer_declarator declarator: (function_declarator declarator: (identifier) @name))) @definition.function
(function_definition declarator: (pointer_declarator declarator: (function_declarator declarator: (qualified_identifier name: (identifier) @name)))) @definition.method
; Functions/methods returning double pointers (nested pointer_declarator)
(function_definition declarator: (pointer_declarator declarator: (pointer_declarator declarator: (function_declarator declarator: (identifier) @name)))) @definition.function
(function_definition declarator: (pointer_declarator declarator: (pointer_declarator declarator: (function_declarator declarator: (qualified_identifier name: (identifier) @name))))) @definition.method
; Functions/methods returning references (reference_declarator wraps function_declarator)
(function_definition declarator: (reference_declarator (function_declarator declarator: (identifier) @name))) @definition.function
(function_definition declarator: (reference_declarator (function_declarator declarator: (qualified_identifier name: (identifier) @name)))) @definition.method
; Destructors (destructor_name is distinct from identifier in tree-sitter-cpp)
(function_definition declarator: (function_declarator declarator: (qualified_identifier name: (destructor_name) @name))) @definition.method
; Function declarations / prototypes (common in headers)
(declaration declarator: (function_declarator declarator: (identifier) @name)) @definition.function
(declaration declarator: (pointer_declarator declarator: (function_declarator declarator: (identifier) @name))) @definition.function
; Inline class method declarations (inside class body, no body: void Foo();)
(field_declaration declarator: (function_declarator declarator: (identifier) @name)) @definition.method
; Inline class method definitions (inside class body, with body: void Foo() { ... })
(field_declaration_list
(function_definition
declarator: (function_declarator
declarator: [(field_identifier) (identifier) (operator_name) (destructor_name)] @name))) @definition.method
; Templates
(template_declaration (class_specifier name: (type_identifier) @name)) @definition.template
(template_declaration (function_definition declarator: (function_declarator declarator: (identifier) @name))) @definition.template
@@ -245,6 +321,9 @@ export const CPP_QUERIES = `
(call_expression function: (qualified_identifier name: (identifier) @call.name)) @call
(call_expression function: (template_function name: (identifier) @call.name)) @call
; Constructor calls: new User()
(new_expression type: (type_identifier) @call.name) @call
; Heritage
(class_specifier name: (type_identifier) @heritage.class
(base_class_clause (type_identifier) @heritage.extends)) @heritage
@@ -262,9 +341,11 @@ export const CSHARP_QUERIES = `
(record_declaration name: (identifier) @name) @definition.record
(delegate_declaration name: (identifier) @name) @definition.delegate
; Namespaces
; Namespaces (block form and C# 10+ file-scoped form)
(namespace_declaration name: (identifier) @name) @definition.namespace
(namespace_declaration name: (qualified_name) @name) @definition.namespace
(file_scoped_namespace_declaration name: (identifier) @name) @definition.namespace
(file_scoped_namespace_declaration name: (qualified_name) @name) @definition.namespace
; Methods & Properties
(method_declaration name: (identifier) @name) @definition.method
@@ -272,6 +353,10 @@ export const CSHARP_QUERIES = `
(constructor_declaration name: (identifier) @name) @definition.constructor
(property_declaration name: (identifier) @name) @definition.property
; Primary constructors (C# 12): class User(string name, int age) { }
(class_declaration name: (identifier) @name (parameter_list) @definition.constructor)
(record_declaration name: (identifier) @name (parameter_list) @definition.constructor)
; Using
(using_directive (qualified_name) @import.source) @import
(using_directive (identifier) @import.source) @import
@@ -280,11 +365,17 @@ export const CSHARP_QUERIES = `
(invocation_expression function: (identifier) @call.name) @call
(invocation_expression function: (member_access_expression name: (identifier) @call.name)) @call
; Constructor calls: new Foo() and new Foo { Props }
(object_creation_expression type: (identifier) @call.name) @call
; Target-typed new (C# 9): User u = new("x", 5)
(variable_declaration type: (identifier) @call.name (variable_declarator (implicit_object_creation_expression) @call))
; Heritage
(class_declaration name: (identifier) @heritage.class
(base_list (simple_base_type (identifier) @heritage.extends))) @heritage
(base_list (identifier) @heritage.extends)) @heritage
(class_declaration name: (identifier) @heritage.class
(base_list (simple_base_type (generic_name (identifier) @heritage.extends)))) @heritage
(base_list (generic_name (identifier) @heritage.extends))) @heritage
`;
// Rust queries - works with tree-sitter-rust
@@ -294,7 +385,8 @@ export const RUST_QUERIES = `
(struct_item name: (type_identifier) @name) @definition.struct
(enum_item name: (type_identifier) @name) @definition.enum
(trait_item name: (type_identifier) @name) @definition.trait
(impl_item type: (type_identifier) @name) @definition.impl
(impl_item type: (type_identifier) @name !trait) @definition.impl
(impl_item type: (generic_type type: (type_identifier) @name) !trait) @definition.impl
(mod_item name: (identifier) @name) @definition.module
; Type aliases, const, static, macros
@@ -312,9 +404,14 @@ export const RUST_QUERIES = `
(call_expression function: (scoped_identifier name: (identifier) @call.name)) @call
(call_expression function: (generic_function function: (identifier) @call.name)) @call
; Heritage (trait implementation)
; Struct literal construction: User { name: value }
(struct_expression name: (type_identifier) @call.name) @call
; Heritage (trait implementation) — all combinations of concrete/generic trait × concrete/generic type
(impl_item trait: (type_identifier) @heritage.trait type: (type_identifier) @heritage.class) @heritage
(impl_item trait: (generic_type type: (type_identifier) @heritage.trait) type: (type_identifier) @heritage.class) @heritage
(impl_item trait: (type_identifier) @heritage.trait type: (generic_type type: (type_identifier) @heritage.class)) @heritage
(impl_item trait: (generic_type type: (type_identifier) @heritage.trait) type: (generic_type type: (type_identifier) @heritage.class)) @heritage
`;
// PHP queries - works with tree-sitter-php (php_only grammar)
@@ -376,6 +473,9 @@ export const PHP_QUERIES = `
(scoped_call_expression
name: (name) @call.name) @call
; Constructor call: new User()
(object_creation_expression (name) @call.name) @call
; ── Heritage: extends ────────────────────────────────────────────────────────
(class_declaration
name: (name) @heritage.class
@@ -396,6 +496,41 @@ export const PHP_QUERIES = `
[(name) (qualified_name)] @heritage.trait))) @heritage
`;
// Ruby queries - works with tree-sitter-ruby
// NOTE: Ruby uses `call` for require, include, extend, prepend, attr_* etc.
// These are all captured as @call and routed in JS post-processing:
// - require/require_relative → import extraction
// - include/extend/prepend → heritage (mixin) extraction
// - attr_accessor/attr_reader/attr_writer → property definition extraction
// - everything else → regular call extraction
export const RUBY_QUERIES = `
; ── Modules ──────────────────────────────────────────────────────────────────
(module
name: (constant) @name) @definition.module
; ── Classes ──────────────────────────────────────────────────────────────────
(class
name: (constant) @name) @definition.class
; ── Instance methods ─────────────────────────────────────────────────────────
(method
name: (identifier) @name) @definition.method
; ── Singleton (class-level) methods ──────────────────────────────────────────
(singleton_method
name: (identifier) @name) @definition.function
; ── All calls (require, include, attr_*, and regular calls routed in JS) ─────
(call
method: (identifier) @call.name) @call
; ── Heritage: class < SuperClass ─────────────────────────────────────────────
(class
name: (constant) @heritage.class
superclass: (superclass
(constant) @heritage.extends)) @heritage
`;
// Kotlin queries - works with tree-sitter-kotlin (fwcd/tree-sitter-kotlin)
// Based on official tags.scm; functions use simple_identifier, classes use type_identifier
export const KOTLIN_QUERIES = `
@@ -527,6 +662,11 @@ export const SWIFT_QUERIES = `
; Heritage - protocol inheritance
(protocol_declaration name: (type_identifier) @heritage.class
(inheritance_specifier inherits_from: (user_type (type_identifier) @heritage.extends))) @heritage
; Heritage - extension protocol conformance (e.g. extension Foo: SomeProtocol)
; Extensions wrap the name in user_type unlike class/struct/enum declarations
(class_declaration "extension" name: (user_type (type_identifier) @heritage.class)
(inheritance_specifier inherits_from: (user_type (type_identifier) @heritage.extends))) @heritage
`;
export const LANGUAGE_QUERIES: Record<SupportedLanguages, string> = {
@@ -538,6 +678,7 @@ export const LANGUAGE_QUERIES: Record<SupportedLanguages, string> = {
[SupportedLanguages.Go]: GO_QUERIES,
[SupportedLanguages.CPlusPlus]: CPP_QUERIES,
[SupportedLanguages.CSharp]: CSHARP_QUERIES,
[SupportedLanguages.Ruby]: RUBY_QUERIES,
[SupportedLanguages.Rust]: RUST_QUERIES,
[SupportedLanguages.PHP]: PHP_QUERIES,
[SupportedLanguages.Kotlin]: KOTLIN_QUERIES,
+124
View File
@@ -0,0 +1,124 @@
import type { SyntaxNode } from './utils.js';
import { FUNCTION_NODE_TYPES, extractFunctionName } from './utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
import { typeConfigs, TYPED_PARAMETER_TYPES } from './type-extractors/index.js';
/**
* Per-file scoped type environment: maps (scope, variableName) → typeName.
* Scope-aware: variables inside functions are keyed by function name,
* file-level variables use the '' (empty string) scope.
*
* Design constraints:
* - Explicit-only: only type annotations, never inferred types
* - Scope-aware: function-local variables don't collide across functions
* - Conservative: complex/generic types extract the base name only
* - Per-file: built once, used for receiver resolution, then discarded
*/
export type TypeEnv = Map<string, Map<string, string>>;
/** File-level scope key */
const FILE_SCOPE = '';
/**
* Look up a variable's type in the TypeEnv, trying the call's enclosing
* function scope first, then falling back to file-level scope.
*/
export const lookupTypeEnv = (
env: TypeEnv,
varName: string,
callNode: SyntaxNode,
): string | undefined => {
// Determine the enclosing function scope for the call
const scopeKey = findEnclosingScopeKey(callNode);
// Try function-local scope first
if (scopeKey) {
const scopeEnv = env.get(scopeKey);
if (scopeEnv) {
const result = scopeEnv.get(varName);
if (result) return result;
}
}
// Fall back to file-level scope
const fileEnv = env.get(FILE_SCOPE);
return fileEnv?.get(varName);
};
/** Find the enclosing function name for scope lookup. */
const findEnclosingScopeKey = (node: SyntaxNode): string | undefined => {
let current = node.parent;
while (current) {
if (FUNCTION_NODE_TYPES.has(current.type)) {
const { funcName } = extractFunctionName(current);
if (funcName) return `${funcName}@${current.startIndex}`;
}
current = current.parent;
}
return undefined;
};
/**
* Build a scoped TypeEnv from a tree-sitter AST for a given language.
* Walks the tree tracking enclosing function scopes, so that variables
* inside different functions don't collide.
*/
export const buildTypeEnv = (
tree: { rootNode: SyntaxNode },
language: SupportedLanguages,
): TypeEnv => {
const env: TypeEnv = new Map();
walkForTypes(tree.rootNode, language, env, FILE_SCOPE);
return env;
};
const walkForTypes = (
node: SyntaxNode,
language: SupportedLanguages,
env: TypeEnv,
currentScope: string,
): void => {
// Detect scope boundaries (function/method definitions)
let scope = currentScope;
if (FUNCTION_NODE_TYPES.has(node.type)) {
const { funcName } = extractFunctionName(node);
if (funcName) scope = `${funcName}@${node.startIndex}`;
}
// Get or create the sub-map for this scope
if (!env.has(scope)) env.set(scope, new Map());
const scopeEnv = env.get(scope)!;
// Check if this node provides type information
extractTypeBinding(node, language, scopeEnv);
// Recurse into children
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (child) walkForTypes(child, language, env, scope);
}
};
/**
* Try to extract a (variableName → typeName) binding from a single AST node.
* Delegates to per-language type configurations.
*/
const extractTypeBinding = (
node: SyntaxNode,
language: SupportedLanguages,
env: Map<string, string>,
): void => {
// === PARAMETERS (most languages) ===
// This guard eliminates 90%+ of calls before any language dispatch.
if (TYPED_PARAMETER_TYPES.has(node.type)) {
const config = typeConfigs[language];
config.extractParameter(node, env);
return;
}
// === Per-language declaration extraction ===
const config = typeConfigs[language];
if (config.declarationNodeTypes.has(node.type)) {
config.extractDeclaration(node, env);
}
};
@@ -0,0 +1,63 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'declaration',
]);
/** C++: Type x = ...; Type* x; Type& x; */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
const typeNode = node.childForFieldName('type');
if (!typeNode) return;
const typeName = extractSimpleTypeName(typeNode);
if (!typeName) return;
const declarator = node.childForFieldName('declarator');
if (!declarator) return;
// init_declarator: Type x = value
const nameNode = declarator.type === 'init_declarator'
? declarator.childForFieldName('declarator')
: declarator;
if (!nameNode) return;
// Handle pointer/reference declarators
const finalName = nameNode.type === 'pointer_declarator' || nameNode.type === 'reference_declarator'
? nameNode.firstNamedChild
: nameNode;
if (!finalName) return;
const varName = extractVarName(finalName);
if (varName) env.set(varName, typeName);
};
/** C/C++: parameter_declaration → type declarator */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'parameter_declaration') {
typeNode = node.childForFieldName('type');
const declarator = node.childForFieldName('declarator');
if (declarator) {
nameNode = declarator.type === 'pointer_declarator' || declarator.type === 'reference_declarator'
? declarator.firstNamedChild
: declarator;
}
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -0,0 +1,93 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, findChildByType } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'local_declaration_statement',
'variable_declaration',
'field_declaration',
]);
/** C#: Type x = ...; var x = new Type(); */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
// C# tree-sitter: local_declaration_statement > variable_declaration > ...
// Recursively descend through wrapper nodes
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child) continue;
if (child.type === 'variable_declaration' || child.type === 'local_declaration_statement') {
extractDeclaration(child, env);
return;
}
}
// At variable_declaration level: first child is type, rest are variable_declarators
let typeNode: SyntaxNode | null = null;
const declarators: SyntaxNode[] = [];
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child) continue;
if (!typeNode && child.type !== 'variable_declarator' && child.type !== 'equals_value_clause') {
// First non-declarator child is the type (identifier, implicit_type, generic_name, etc.)
typeNode = child;
}
if (child.type === 'variable_declarator') {
declarators.push(child);
}
}
if (!typeNode || declarators.length === 0) return;
// Handle 'var x = new Foo()' — infer from object_creation_expression
let typeName: string | undefined;
if (typeNode.type === 'implicit_type' && typeNode.text === 'var') {
// Try to infer from initializer: var x = new Foo()
// C# tree-sitter puts object_creation_expression as direct child of variable_declarator
if (declarators.length === 1) {
const initializer = findChildByType(declarators[0], 'object_creation_expression')
?? findChildByType(declarators[0], 'equals_value_clause')?.firstNamedChild;
if (initializer?.type === 'object_creation_expression') {
const ctorType = initializer.childForFieldName('type');
if (ctorType) typeName = extractSimpleTypeName(ctorType);
}
}
} else {
typeName = extractSimpleTypeName(typeNode);
}
if (!typeName) return;
for (const decl of declarators) {
const nameNode = decl.childForFieldName('name') ?? decl.firstNamedChild;
if (nameNode) {
const varName = extractVarName(nameNode);
if (varName) env.set(varName, typeName);
}
}
};
/** C#: parameter → type name */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'parameter') {
typeNode = node.childForFieldName('type');
nameNode = node.childForFieldName('name');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -0,0 +1,104 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'var_declaration',
'var_spec',
'short_var_declaration',
]);
/** Go: var x Foo */
const extractGoVarDeclaration = (node: SyntaxNode, env: Map<string, string>): void => {
// Go var_declaration contains var_spec children
if (node.type === 'var_declaration') {
for (let i = 0; i < node.namedChildCount; i++) {
const spec = node.namedChild(i);
if (spec?.type === 'var_spec') extractGoVarDeclaration(spec, env);
}
return;
}
// var_spec: name type [= value]
const nameNode = node.childForFieldName('name');
const typeNode = node.childForFieldName('type');
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
/** Go: x := Foo{...} — infer type from composite literal (handles multi-assignment) */
const extractGoShortVarDeclaration = (node: SyntaxNode, env: Map<string, string>): void => {
const left = node.childForFieldName('left');
const right = node.childForFieldName('right');
if (!left || !right) return;
// Collect LHS names and RHS values (may be expression_lists for multi-assignment)
const lhsNodes: SyntaxNode[] = [];
const rhsNodes: SyntaxNode[] = [];
if (left.type === 'expression_list') {
for (let i = 0; i < left.namedChildCount; i++) {
const c = left.namedChild(i);
if (c) lhsNodes.push(c);
}
} else {
lhsNodes.push(left);
}
if (right.type === 'expression_list') {
for (let i = 0; i < right.namedChildCount; i++) {
const c = right.namedChild(i);
if (c) rhsNodes.push(c);
}
} else {
rhsNodes.push(right);
}
// Pair each LHS name with its corresponding RHS value
const count = Math.min(lhsNodes.length, rhsNodes.length);
for (let i = 0; i < count; i++) {
const valueNode = rhsNodes[i];
if (valueNode.type !== 'composite_literal') continue;
const typeNode = valueNode.childForFieldName('type');
if (!typeNode) continue;
const typeName = extractSimpleTypeName(typeNode);
if (!typeName) continue;
const varName = extractVarName(lhsNodes[i]);
if (varName) env.set(varName, typeName);
}
};
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
if (node.type === 'var_declaration' || node.type === 'var_spec') {
extractGoVarDeclaration(node, env);
} else if (node.type === 'short_var_declaration') {
extractGoShortVarDeclaration(node, env);
}
};
/** Go: parameter → name type */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'parameter') {
nameNode = node.childForFieldName('name');
typeNode = node.childForFieldName('type');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -0,0 +1,40 @@
/**
* Per-language type extraction configurations.
* Assembled here into a dispatch map keyed by SupportedLanguages.
*/
import { SupportedLanguages } from '../../../config/supported-languages.js';
import type { LanguageTypeConfig } from './types.js';
import { typeConfig as typescriptConfig } from './typescript.js';
import { javaTypeConfig, kotlinTypeConfig } from './jvm.js';
import { typeConfig as csharpConfig } from './csharp.js';
import { typeConfig as goConfig } from './go.js';
import { typeConfig as rustConfig } from './rust.js';
import { typeConfig as pythonConfig } from './python.js';
import { typeConfig as swiftConfig } from './swift.js';
import { typeConfig as cCppConfig } from './c-cpp.js';
import { typeConfig as phpConfig } from './php.js';
export const typeConfigs = {
[SupportedLanguages.JavaScript]: typescriptConfig,
[SupportedLanguages.TypeScript]: typescriptConfig,
[SupportedLanguages.Java]: javaTypeConfig,
[SupportedLanguages.Kotlin]: kotlinTypeConfig,
[SupportedLanguages.CSharp]: csharpConfig,
[SupportedLanguages.Go]: goConfig,
[SupportedLanguages.Rust]: rustConfig,
[SupportedLanguages.Python]: pythonConfig,
[SupportedLanguages.Swift]: swiftConfig,
[SupportedLanguages.C]: cCppConfig,
[SupportedLanguages.CPlusPlus]: cCppConfig,
[SupportedLanguages.PHP]: phpConfig,
[SupportedLanguages.Ruby]: {
declarationNodeTypes: new Set<string>(),
extractDeclaration: () => {},
extractParameter: () => {},
},
} satisfies Record<SupportedLanguages, LanguageTypeConfig>;
export type { LanguageTypeConfig, TypeBindingExtractor, ParameterExtractor } from './types.js';
export { TYPED_PARAMETER_TYPES, extractSimpleTypeName, extractVarName, findChildByType } from './shared.js';
@@ -0,0 +1,122 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, findChildByType } from './shared.js';
// ── Java ──────────────────────────────────────────────────────────────────
const JAVA_DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'local_variable_declaration',
'field_declaration',
]);
/** Java: Type x = ...; Type x; */
const extractJavaDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
const typeNode = node.childForFieldName('type');
if (!typeNode) return;
const typeName = extractSimpleTypeName(typeNode);
if (!typeName) return;
// Find variable_declarator children
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type !== 'variable_declarator') continue;
const nameNode = child.childForFieldName('name');
if (nameNode) {
const varName = extractVarName(nameNode);
if (varName) env.set(varName, typeName);
}
}
};
/** Java: formal_parameter → type name */
const extractJavaParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'formal_parameter') {
typeNode = node.childForFieldName('type');
nameNode = node.childForFieldName('name');
} else {
// Generic fallback
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const javaTypeConfig: LanguageTypeConfig = {
declarationNodeTypes: JAVA_DECLARATION_NODE_TYPES,
extractDeclaration: extractJavaDeclaration,
extractParameter: extractJavaParameter,
};
// ── Kotlin ────────────────────────────────────────────────────────────────
const KOTLIN_DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'property_declaration',
'variable_declaration',
]);
/** Kotlin: val x: Foo = ... */
const extractKotlinDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
if (node.type === 'property_declaration') {
// Kotlin property_declaration: name/type are inside a variable_declaration child
const varDecl = findChildByType(node, 'variable_declaration');
if (varDecl) {
const nameNode = findChildByType(varDecl, 'simple_identifier');
const typeNode = findChildByType(varDecl, 'user_type');
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
return;
}
// Fallback: try direct fields
const nameNode = node.childForFieldName('name')
?? findChildByType(node, 'simple_identifier');
const typeNode = node.childForFieldName('type')
?? findChildByType(node, 'user_type');
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
} else if (node.type === 'variable_declaration') {
// variable_declaration directly inside functions
const nameNode = findChildByType(node, 'simple_identifier');
const typeNode = findChildByType(node, 'user_type');
if (nameNode && typeNode) {
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
}
}
};
/** Kotlin: formal_parameter → type name */
const extractKotlinParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'formal_parameter') {
typeNode = node.childForFieldName('type');
nameNode = node.childForFieldName('name');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const kotlinTypeConfig: LanguageTypeConfig = {
declarationNodeTypes: KOTLIN_DECLARATION_NODE_TYPES,
extractDeclaration: extractKotlinDeclaration,
extractParameter: extractKotlinParameter,
};
@@ -0,0 +1,36 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
// PHP has no local variable type annotations; only params carry types
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set<string>();
/** PHP: no typed local variable declarations */
const extractDeclaration: TypeBindingExtractor = (_node: SyntaxNode, _env: Map<string, string>): void => {
// PHP has no local variable type annotations
};
/** PHP: simple_parameter → type $name */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'simple_parameter') {
typeNode = node.childForFieldName('type');
nameNode = node.childForFieldName('name');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -0,0 +1,44 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'assignment',
]);
/** Python: x: Foo = ... (PEP 484 annotations) */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
// Python annotated assignment: left : type = value
// tree-sitter represents this differently based on grammar version
const left = node.childForFieldName('left');
const typeNode = node.childForFieldName('type');
if (!left || !typeNode) return;
const varName = extractVarName(left);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
/** Python: parameter with type annotation */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'parameter') {
nameNode = node.childForFieldName('name');
typeNode = node.childForFieldName('type');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -0,0 +1,42 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'let_declaration',
]);
/** Rust: let x: Foo = ... */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
const pattern = node.childForFieldName('pattern');
const typeNode = node.childForFieldName('type');
if (!pattern || !typeNode) return;
const varName = extractVarName(pattern);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
/** Rust: parameter → pattern: type */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'parameter') {
nameNode = node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -0,0 +1,103 @@
import type { SyntaxNode } from '../utils.js';
/**
* Extract the simple type name from a type AST node.
* Handles generic types (e.g., List<User> → List), qualified names
* (e.g., models.User → User), and nullable types (e.g., User? → User).
* Returns undefined for complex types (unions, intersections, function types).
*/
export const extractSimpleTypeName = (typeNode: SyntaxNode): string | undefined => {
// Direct type identifier
if (typeNode.type === 'type_identifier' || typeNode.type === 'identifier'
|| typeNode.type === 'simple_identifier') {
return typeNode.text;
}
// Qualified/scoped names: take the last segment (e.g., models.User → User)
if (typeNode.type === 'scoped_identifier' || typeNode.type === 'qualified_identifier'
|| typeNode.type === 'scoped_type_identifier' || typeNode.type === 'qualified_name'
|| typeNode.type === 'qualified_type'
|| typeNode.type === 'member_expression' || typeNode.type === 'attribute') {
const last = typeNode.lastNamedChild;
if (last && (last.type === 'type_identifier' || last.type === 'identifier'
|| last.type === 'simple_identifier' || last.type === 'name')) {
return last.text;
}
}
// Generic types: extract the base type (e.g., List<User> → List)
if (typeNode.type === 'generic_type' || typeNode.type === 'parameterized_type') {
const base = typeNode.childForFieldName('name')
?? typeNode.childForFieldName('type')
?? typeNode.firstNamedChild;
if (base) return extractSimpleTypeName(base);
}
// Nullable types (Kotlin User?, C# User?)
if (typeNode.type === 'nullable_type') {
const inner = typeNode.firstNamedChild;
if (inner) return extractSimpleTypeName(inner);
}
// Type annotations that wrap the actual type (TS/Python: `: Foo`, Kotlin: user_type)
if (typeNode.type === 'type_annotation' || typeNode.type === 'type'
|| typeNode.type === 'user_type') {
const inner = typeNode.firstNamedChild;
if (inner) return extractSimpleTypeName(inner);
}
// Pointer/reference types (C++, Rust): User*, &User, &mut User
if (typeNode.type === 'pointer_type' || typeNode.type === 'reference_type') {
const inner = typeNode.firstNamedChild;
if (inner) return extractSimpleTypeName(inner);
}
// PHP named_type / optional_type
if (typeNode.type === 'named_type' || typeNode.type === 'optional_type') {
const inner = typeNode.childForFieldName('name') ?? typeNode.firstNamedChild;
if (inner) return extractSimpleTypeName(inner);
}
// Name node (PHP)
if (typeNode.type === 'name') {
return typeNode.text;
}
return undefined;
};
/**
* Extract variable name from a declarator or pattern node.
* Returns the simple identifier text, or undefined for destructuring/complex patterns.
*/
export const extractVarName = (node: SyntaxNode): string | undefined => {
if (node.type === 'identifier' || node.type === 'simple_identifier'
|| node.type === 'variable_name' || node.type === 'name') {
return node.text;
}
// variable_declarator (Java/C#): has a 'name' field
if (node.type === 'variable_declarator') {
const nameChild = node.childForFieldName('name');
if (nameChild) return extractVarName(nameChild);
}
return undefined;
};
/** Node types for function/method parameters with type annotations */
export const TYPED_PARAMETER_TYPES = new Set([
'required_parameter', // TS: (x: Foo)
'optional_parameter', // TS: (x?: Foo)
'formal_parameter', // Java/Kotlin
'parameter', // C#/Rust/Go/Python/Swift
'parameter_declaration', // C/C++ void f(Type name)
'simple_parameter', // PHP function(Foo $x)
]);
/** Find the first named child with the given node type */
export const findChildByType = (node: SyntaxNode, type: string): SyntaxNode | null => {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === type) return child;
}
return null;
};
@@ -0,0 +1,46 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, findChildByType } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'property_declaration',
]);
/** Swift: let x: Foo = ... */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
// Swift property_declaration has pattern and type_annotation
const pattern = node.childForFieldName('pattern')
?? findChildByType(node, 'pattern');
const typeAnnotation = node.childForFieldName('type')
?? findChildByType(node, 'type_annotation');
if (!pattern || !typeAnnotation) return;
const varName = extractVarName(pattern) ?? pattern.text;
const typeName = extractSimpleTypeName(typeAnnotation);
if (varName && typeName) env.set(varName, typeName);
};
/** Swift: parameter → name: type */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'parameter') {
nameNode = node.childForFieldName('name')
?? node.childForFieldName('internal_name');
typeNode = node.childForFieldName('type');
} else {
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
@@ -0,0 +1,17 @@
import type { SyntaxNode } from '../utils.js';
/** Extracts type bindings from a declaration node into the env map */
export type TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>) => void;
/** Extracts type bindings from a parameter node into the env map */
export type ParameterExtractor = (node: SyntaxNode, env: Map<string, string>) => void;
/** Per-language type extraction configuration */
export interface LanguageTypeConfig {
/** Node types that represent typed declarations for this language */
declarationNodeTypes: ReadonlySet<string>;
/** Extract a (varName → typeName) binding from a declaration node */
extractDeclaration: TypeBindingExtractor;
/** Extract a (varName → typeName) binding from a parameter node */
extractParameter: ParameterExtractor;
}
@@ -0,0 +1,48 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'lexical_declaration',
'variable_declaration',
]);
/** TypeScript: const x: Foo = ..., let x: Foo */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
for (let i = 0; i < node.namedChildCount; i++) {
const declarator = node.namedChild(i);
if (declarator?.type !== 'variable_declarator') continue;
const nameNode = declarator.childForFieldName('name');
const typeAnnotation = declarator.childForFieldName('type');
if (!nameNode || !typeAnnotation) continue;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeAnnotation);
if (varName && typeName) env.set(varName, typeName);
}
};
/** TypeScript: required_parameter / optional_parameter → name: type */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
if (node.type === 'required_parameter' || node.type === 'optional_parameter') {
nameNode = node.childForFieldName('pattern') ?? node.childForFieldName('name');
typeNode = node.childForFieldName('type');
} else {
// Generic fallback
nameNode = node.childForFieldName('name') ?? node.childForFieldName('pattern');
typeNode = node.childForFieldName('type');
}
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
};
+764 -4
View File
@@ -1,4 +1,415 @@
import type Parser from 'tree-sitter';
import { SupportedLanguages } from '../../config/supported-languages.js';
import { generateId } from '../../lib/utils.js';
/** Tree-sitter AST node. Re-exported for use across ingestion modules. */
export type SyntaxNode = Parser.SyntaxNode;
/**
* Ordered list of definition capture keys for tree-sitter query matches.
* Used to extract the definition node from a capture map.
*/
export const DEFINITION_CAPTURE_KEYS = [
'definition.function',
'definition.class',
'definition.interface',
'definition.method',
'definition.struct',
'definition.enum',
'definition.namespace',
'definition.module',
'definition.trait',
'definition.impl',
'definition.type',
'definition.const',
'definition.static',
'definition.typedef',
'definition.macro',
'definition.union',
'definition.property',
'definition.record',
'definition.delegate',
'definition.annotation',
'definition.constructor',
'definition.template',
] as const;
/** Extract the definition node from a tree-sitter query capture map. */
export const getDefinitionNodeFromCaptures = (captureMap: Record<string, any>): any | null => {
for (const key of DEFINITION_CAPTURE_KEYS) {
if (captureMap[key]) return captureMap[key];
}
return null;
};
/**
* Node types that represent function/method definitions across languages.
* Used to find the enclosing function for a call site.
*/
export const FUNCTION_NODE_TYPES = new Set([
// TypeScript/JavaScript
'function_declaration',
'arrow_function',
'function_expression',
'method_definition',
'generator_function_declaration',
// Python
'function_definition',
// Common async variants
'async_function_declaration',
'async_arrow_function',
// Java
'method_declaration',
'constructor_declaration',
// C/C++
// 'function_definition' already included above
// Go
// 'method_declaration' already included from Java
// C#
'local_function_statement',
// Rust
'function_item',
'impl_item', // Methods inside impl blocks
// PHP
'anonymous_function',
// Kotlin
'lambda_literal',
// Swift
'init_declaration',
'deinit_declaration',
]);
/**
* Node types for standard function declarations that need C/C++ declarator handling.
* Used by extractFunctionName to determine how to extract the function name.
*/
export const FUNCTION_DECLARATION_TYPES = new Set([
'function_declaration',
'function_definition',
'async_function_declaration',
'generator_function_declaration',
'function_item',
]);
/**
* Built-in function/method names that should not be tracked as call targets.
* Covers JS/TS, Python, Kotlin, C/C++, PHP, Swift standard library functions.
*/
export const BUILT_IN_NAMES = new Set([
// JavaScript/TypeScript
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export', 'fetch', 'Response', 'Request',
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'append', 'extend', 'update',
// NOTE: 'open', 'read', 'write', 'close' removed — these are real C POSIX syscalls
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
// Kotlin stdlib
'println', 'print', 'readLine', 'require', 'requireNotNull', 'check', 'assert', 'lazy', 'error',
'listOf', 'mapOf', 'setOf', 'mutableListOf', 'mutableMapOf', 'mutableSetOf',
'arrayOf', 'sequenceOf', 'also', 'apply', 'run', 'with', 'takeIf', 'takeUnless',
'TODO', 'buildString', 'buildList', 'buildMap', 'buildSet',
'repeat', 'synchronized',
// Kotlin coroutine builders & scope functions
'launch', 'async', 'runBlocking', 'withContext', 'coroutineScope',
'supervisorScope', 'delay',
// Kotlin Flow operators
'flow', 'flowOf', 'collect', 'emit', 'onEach', 'catch',
'buffer', 'conflate', 'distinctUntilChanged',
'flatMapLatest', 'flatMapMerge', 'combine',
'stateIn', 'shareIn', 'launchIn',
// Kotlin infix stdlib functions
'to', 'until', 'downTo', 'step',
// C/C++ standard library
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
'scanf', 'fscanf', 'sscanf',
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
'sizeof', 'offsetof', 'typeof',
'assert', 'abort', 'exit', '_exit',
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
// Linux kernel common macros/helpers (not real call targets)
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
'min', 'max', 'clamp', 'abs', 'swap',
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
'GFP_KERNEL', 'GFP_ATOMIC',
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
'mutex_lock', 'mutex_unlock', 'mutex_init',
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
'get', 'put',
// C# / .NET built-ins
'Console', 'WriteLine', 'ReadLine', 'Write',
'Task', 'Run', 'Wait', 'WhenAll', 'WhenAny', 'FromResult', 'Delay', 'ContinueWith',
'ConfigureAwait', 'GetAwaiter', 'GetResult',
'ToString', 'GetType', 'Equals', 'GetHashCode', 'ReferenceEquals',
'Add', 'Remove', 'Contains', 'Clear', 'Count', 'Any', 'All',
'Where', 'Select', 'SelectMany', 'OrderBy', 'OrderByDescending', 'GroupBy',
'First', 'FirstOrDefault', 'Single', 'SingleOrDefault', 'Last', 'LastOrDefault',
'ToList', 'ToArray', 'ToDictionary', 'AsEnumerable', 'AsQueryable',
'Aggregate', 'Sum', 'Average', 'Min', 'Max', 'Distinct', 'Skip', 'Take',
'String', 'Format', 'IsNullOrEmpty', 'IsNullOrWhiteSpace', 'Concat', 'Join',
'Trim', 'TrimStart', 'TrimEnd', 'Split', 'Replace', 'StartsWith', 'EndsWith',
'Convert', 'ToInt32', 'ToDouble', 'ToBoolean', 'ToByte',
'Math', 'Abs', 'Ceiling', 'Floor', 'Round', 'Pow', 'Sqrt',
'Dispose', 'Close',
'TryParse', 'Parse',
'AddRange', 'RemoveAt', 'RemoveAll', 'FindAll', 'Exists', 'TrueForAll',
'ContainsKey', 'TryGetValue', 'AddOrUpdate',
'Throw', 'ThrowIfNull',
// PHP built-ins
'echo', 'isset', 'empty', 'unset', 'list', 'array', 'compact', 'extract',
'count', 'strlen', 'strpos', 'strrpos', 'substr', 'strtolower', 'strtoupper', 'trim',
'ltrim', 'rtrim', 'str_replace', 'str_contains', 'str_starts_with', 'str_ends_with',
'sprintf', 'vsprintf', 'printf', 'number_format',
'array_map', 'array_filter', 'array_reduce', 'array_push', 'array_pop', 'array_shift',
'array_unshift', 'array_slice', 'array_splice', 'array_merge', 'array_keys', 'array_values',
'array_key_exists', 'in_array', 'array_search', 'array_unique', 'usort', 'rsort',
'json_encode', 'json_decode', 'serialize', 'unserialize',
'intval', 'floatval', 'strval', 'boolval', 'is_null', 'is_string', 'is_int', 'is_array',
'is_object', 'is_numeric', 'is_bool', 'is_float',
'var_dump', 'print_r', 'var_export',
'date', 'time', 'strtotime', 'mktime', 'microtime',
'file_exists', 'file_get_contents', 'file_put_contents', 'is_file', 'is_dir',
'preg_match', 'preg_match_all', 'preg_replace', 'preg_split',
'header', 'session_start', 'session_destroy', 'ob_start', 'ob_end_clean', 'ob_get_clean',
'dd', 'dump',
// Swift/iOS built-ins and standard library
'print', 'debugPrint', 'dump', 'fatalError', 'precondition', 'preconditionFailure',
'assert', 'assertionFailure', 'NSLog',
'abs', 'min', 'max', 'zip', 'stride', 'sequence', 'repeatElement',
'swap', 'withUnsafePointer', 'withUnsafeMutablePointer', 'withUnsafeBytes',
'autoreleasepool', 'unsafeBitCast', 'unsafeDowncast', 'numericCast',
'type', 'MemoryLayout',
// Swift collection/string methods (common noise)
'map', 'flatMap', 'compactMap', 'filter', 'reduce', 'forEach', 'contains',
'first', 'last', 'prefix', 'suffix', 'dropFirst', 'dropLast',
'sorted', 'reversed', 'enumerated', 'joined', 'split',
'append', 'insert', 'remove', 'removeAll', 'removeFirst', 'removeLast',
'isEmpty', 'count', 'index', 'startIndex', 'endIndex',
// UIKit/Foundation common methods (noise in call graph)
'addSubview', 'removeFromSuperview', 'layoutSubviews', 'setNeedsLayout',
'layoutIfNeeded', 'setNeedsDisplay', 'invalidateIntrinsicContentSize',
'addTarget', 'removeTarget', 'addGestureRecognizer',
'addConstraint', 'addConstraints', 'removeConstraint', 'removeConstraints',
'NSLocalizedString', 'Bundle',
'reloadData', 'reloadSections', 'reloadRows', 'performBatchUpdates',
'register', 'dequeueReusableCell', 'dequeueReusableSupplementaryView',
'beginUpdates', 'endUpdates', 'insertRows', 'deleteRows', 'insertSections', 'deleteSections',
'present', 'dismiss', 'pushViewController', 'popViewController', 'popToRootViewController',
'performSegue', 'prepare',
// GCD / async
'DispatchQueue', 'async', 'sync', 'asyncAfter',
'Task', 'withCheckedContinuation', 'withCheckedThrowingContinuation',
// Combine
'sink', 'store', 'assign', 'receive', 'subscribe',
// Notification / KVO
'addObserver', 'removeObserver', 'post', 'NotificationCenter',
// Rust standard library (common noise in call graphs)
'unwrap', 'expect', 'unwrap_or', 'unwrap_or_else', 'unwrap_or_default',
'ok', 'err', 'is_ok', 'is_err', 'map', 'map_err', 'and_then', 'or_else',
'clone', 'to_string', 'to_owned', 'into', 'from', 'as_ref', 'as_mut',
'iter', 'into_iter', 'collect', 'map', 'filter', 'fold', 'for_each',
'len', 'is_empty', 'push', 'pop', 'insert', 'remove', 'contains',
'format', 'write', 'writeln', 'panic', 'unreachable', 'todo', 'unimplemented',
'vec', 'println', 'eprintln', 'dbg',
'lock', 'read', 'write', 'try_lock',
'spawn', 'join', 'sleep',
'Some', 'None', 'Ok', 'Err',
]);
/** Check if a name is a built-in function or common noise that should be filtered out */
export const isBuiltInOrNoise = (name: string): boolean => BUILT_IN_NAMES.has(name);
/** AST node types that represent a class-like container (for HAS_METHOD edge extraction) */
export const CLASS_CONTAINER_TYPES = new Set([
'class_declaration', 'abstract_class_declaration',
'interface_declaration', 'struct_declaration', 'record_declaration',
'class_specifier', 'struct_specifier',
'impl_item', 'trait_item',
'class_definition',
'trait_declaration',
'protocol_declaration',
]);
export const CONTAINER_TYPE_TO_LABEL: Record<string, string> = {
class_declaration: 'Class',
abstract_class_declaration: 'Class',
interface_declaration: 'Interface',
struct_declaration: 'Struct',
struct_specifier: 'Struct',
class_specifier: 'Class',
class_definition: 'Class',
impl_item: 'Impl',
trait_item: 'Trait',
trait_declaration: 'Trait',
record_declaration: 'Record',
protocol_declaration: 'Interface',
};
/** Walk up AST to find enclosing class/struct/interface/impl, return its generateId or null.
* For Go method_declaration nodes, extracts receiver type (e.g. `func (u *User) Save()` → User struct). */
export const findEnclosingClassId = (node: any, filePath: string): string | null => {
let current = node.parent;
while (current) {
// Go: method_declaration has a receiver parameter with the struct type
if (current.type === 'method_declaration') {
const receiver = current.childForFieldName?.('receiver');
if (receiver) {
// receiver is a parameter_list: (u *User) or (u User)
const paramDecl = receiver.namedChildren?.find?.((c: any) => c.type === 'parameter_declaration');
if (paramDecl) {
const typeNode = paramDecl.childForFieldName?.('type');
if (typeNode) {
// Unwrap pointer_type (*User → User)
const inner = typeNode.type === 'pointer_type' ? typeNode.firstNamedChild : typeNode;
if (inner && (inner.type === 'type_identifier' || inner.type === 'identifier')) {
return generateId('Struct', `${filePath}:${inner.text}`);
}
}
}
}
}
if (CLASS_CONTAINER_TYPES.has(current.type)) {
// Rust impl_item: for `impl Trait for Struct {}`, pick the type after `for`
if (current.type === 'impl_item') {
const children = current.children ?? [];
const forIdx = children.findIndex((c: any) => c.text === 'for');
if (forIdx !== -1) {
const nameNode = children.slice(forIdx + 1).find((c: any) =>
c.type === 'type_identifier' || c.type === 'identifier'
);
if (nameNode) {
return generateId('Impl', `${filePath}:${nameNode.text}`);
}
}
// Fall through: plain `impl Struct {}` — use first type_identifier below
}
const nameNode = current.childForFieldName?.('name')
?? current.children?.find((c: any) =>
c.type === 'type_identifier' || c.type === 'identifier' || c.type === 'name'
);
if (nameNode) {
const label = CONTAINER_TYPE_TO_LABEL[current.type] || 'Class';
return generateId(label, `${filePath}:${nameNode.text}`);
}
}
current = current.parent;
}
return null;
};
/**
* Extract function name and label from a function_definition or similar AST node.
* Handles C/C++ qualified_identifier (ClassName::MethodName) and other language patterns.
*/
export const extractFunctionName = (node: any): { funcName: string | null; label: string } => {
let funcName: string | null = null;
let label = 'Function';
// Swift init/deinit
if (node.type === 'init_declaration' || node.type === 'deinit_declaration') {
return {
funcName: node.type === 'init_declaration' ? 'init' : 'deinit',
label: 'Constructor',
};
}
if (FUNCTION_DECLARATION_TYPES.has(node.type)) {
// C/C++: function_definition -> [pointer_declarator ->] function_declarator -> qualified_identifier/identifier
// Unwrap pointer_declarator / reference_declarator wrappers to reach function_declarator
let declarator = node.childForFieldName?.('declarator') ||
node.children?.find((c: any) => c.type === 'function_declarator');
while (declarator && (declarator.type === 'pointer_declarator' || declarator.type === 'reference_declarator')) {
declarator = declarator.childForFieldName?.('declarator') ||
declarator.children?.find((c: any) =>
c.type === 'function_declarator' || c.type === 'pointer_declarator' || c.type === 'reference_declarator');
}
if (declarator) {
const innerDeclarator = declarator.childForFieldName?.('declarator') ||
declarator.children?.find((c: any) =>
c.type === 'qualified_identifier' || c.type === 'identifier' || c.type === 'parenthesized_declarator');
if (innerDeclarator?.type === 'qualified_identifier') {
const nameNode = innerDeclarator.childForFieldName?.('name') ||
innerDeclarator.children?.find((c: any) => c.type === 'identifier');
if (nameNode?.text) {
funcName = nameNode.text;
label = 'Method';
}
} else if (innerDeclarator?.type === 'identifier') {
funcName = innerDeclarator.text;
} else if (innerDeclarator?.type === 'parenthesized_declarator') {
const nestedId = innerDeclarator.children?.find((c: any) =>
c.type === 'qualified_identifier' || c.type === 'identifier');
if (nestedId?.type === 'qualified_identifier') {
const nameNode = nestedId.childForFieldName?.('name') ||
nestedId.children?.find((c: any) => c.type === 'identifier');
if (nameNode?.text) {
funcName = nameNode.text;
label = 'Method';
}
} else if (nestedId?.type === 'identifier') {
funcName = nestedId.text;
}
}
}
// Fallback for other languages (Kotlin uses simple_identifier, Swift uses simple_identifier)
if (!funcName) {
const nameNode = node.childForFieldName?.('name') ||
node.children?.find((c: any) => c.type === 'identifier' || c.type === 'property_identifier' || c.type === 'simple_identifier');
funcName = nameNode?.text;
}
} else if (node.type === 'impl_item') {
const funcItem = node.children?.find((c: any) => c.type === 'function_item');
if (funcItem) {
const nameNode = funcItem.childForFieldName?.('name') ||
funcItem.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
}
} else if (node.type === 'method_definition') {
const nameNode = node.childForFieldName?.('name') ||
node.children?.find((c: any) => c.type === 'property_identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (node.type === 'method_declaration' || node.type === 'constructor_declaration') {
const nameNode = node.childForFieldName?.('name') ||
node.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (node.type === 'arrow_function' || node.type === 'function_expression') {
const parent = node.parent;
if (parent?.type === 'variable_declarator') {
const nameNode = parent.childForFieldName?.('name') ||
parent.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
}
}
return { funcName, label };
};
/**
* Yield control to the event loop so spinners/progress can render.
@@ -6,6 +417,9 @@ import { SupportedLanguages } from '../../config/supported-languages.js';
*/
export const yieldToEventLoop = (): Promise<void> => new Promise(resolve => setImmediate(resolve));
/** Ruby extensionless filenames recognised as Ruby source */
const RUBY_EXTENSIONLESS_FILES = new Set(['Rakefile', 'Gemfile', 'Guardfile', 'Vagrantfile', 'Brewfile']);
/**
* Find a child of `childType` within a sibling node of `siblingType`.
* Used for Kotlin AST traversal where visibility_modifier lives inside a modifiers sibling.
@@ -37,11 +451,13 @@ export const getLanguageFromFilename = (filename: string): SupportedLanguages |
if (filename.endsWith('.py')) return SupportedLanguages.Python;
// Java
if (filename.endsWith('.java')) return SupportedLanguages.Java;
// C (source and headers)
if (filename.endsWith('.c') || filename.endsWith('.h')) return SupportedLanguages.C;
// C++ (all common extensions)
// C source files
if (filename.endsWith('.c')) return SupportedLanguages.C;
// C++ (all common extensions, including .h)
// .h is parsed as C++ because tree-sitter-cpp is a strict superset of C, so pure-C
// headers parse correctly, and C++ headers (classes, templates) are handled properly.
if (filename.endsWith('.cpp') || filename.endsWith('.cc') || filename.endsWith('.cxx') ||
filename.endsWith('.hpp') || filename.endsWith('.hxx') || filename.endsWith('.hh')) return SupportedLanguages.CPlusPlus;
filename.endsWith('.h') || filename.endsWith('.hpp') || filename.endsWith('.hxx') || filename.endsWith('.hh')) return SupportedLanguages.CPlusPlus;
// C#
if (filename.endsWith('.cs')) return SupportedLanguages.CSharp;
// Go
@@ -56,7 +472,351 @@ export const getLanguageFromFilename = (filename: string): SupportedLanguages |
filename.endsWith('.php5') || filename.endsWith('.php8')) {
return SupportedLanguages.PHP;
}
// Ruby (extensions)
if (filename.endsWith('.rb') || filename.endsWith('.rake') || filename.endsWith('.gemspec')) {
return SupportedLanguages.Ruby;
}
// Ruby (extensionless files)
const basename = filename.split('/').pop() || filename;
if (RUBY_EXTENSIONLESS_FILES.has(basename)) {
return SupportedLanguages.Ruby;
}
// Swift (extensions)
if (filename.endsWith('.swift')) return SupportedLanguages.Swift;
return null;
};
export interface MethodSignature {
parameterCount: number | undefined;
returnType: string | undefined;
}
const CALL_ARGUMENT_LIST_TYPES = new Set([
'arguments',
'argument_list',
'value_arguments',
]);
/**
* Extract parameter count and return type text from an AST method/function node.
* Works across languages by looking for common AST patterns.
*/
export const extractMethodSignature = (node: SyntaxNode | null | undefined): MethodSignature => {
let parameterCount: number | undefined = 0;
let returnType: string | undefined;
let isVariadic = false;
if (!node) return { parameterCount, returnType };
const paramListTypes = new Set([
'formal_parameters', 'parameters', 'parameter_list',
'function_parameters', 'method_parameters', 'function_value_parameters',
]);
// Node types that indicate variadic/rest parameters
const VARIADIC_PARAM_TYPES = new Set([
'variadic_parameter_declaration', // Go: ...string
'variadic_parameter', // Rust: extern "C" fn(...)
'spread_parameter', // Java: Object... args
'list_splat_pattern', // Python: *args
'dictionary_splat_pattern', // Python: **kwargs
]);
const findParameterList = (current: SyntaxNode): SyntaxNode | null => {
for (const child of current.children) {
if (paramListTypes.has(child.type)) return child;
}
for (const child of current.children) {
const nested = findParameterList(child);
if (nested) return nested;
}
return null;
};
const parameterList = (
paramListTypes.has(node.type) ? node // node itself IS the parameter list (e.g. C# primary constructors)
: node.childForFieldName?.('parameters')
?? findParameterList(node)
);
if (parameterList && paramListTypes.has(parameterList.type)) {
for (const param of parameterList.namedChildren) {
if (param.type === 'comment') continue;
if (param.text === 'self' || param.text === '&self' || param.text === '&mut self' ||
param.type === 'self_parameter') {
continue;
}
// Check for variadic parameter types
if (VARIADIC_PARAM_TYPES.has(param.type)) {
isVariadic = true;
continue;
}
// TypeScript/JavaScript: rest parameter — required_parameter containing rest_pattern
if (param.type === 'required_parameter' || param.type === 'optional_parameter') {
for (const child of param.children) {
if (child.type === 'rest_pattern') {
isVariadic = true;
break;
}
}
if (isVariadic) continue;
}
// Kotlin: vararg modifier on a regular parameter
if (param.type === 'parameter' || param.type === 'formal_parameter') {
const prev = param.previousSibling;
if (prev?.type === 'parameter_modifiers' && prev.text.includes('vararg')) {
isVariadic = true;
}
}
parameterCount++;
}
// C/C++: bare `...` token in parameter list (not a named child — check all children)
if (!isVariadic) {
for (const child of parameterList.children) {
if (!child.isNamed && child.text === '...') {
isVariadic = true;
break;
}
}
}
}
// Return type extraction — language-specific field names
// Go: 'result' field is either a type_identifier or parameter_list (multi-return)
const goResult = node.childForFieldName?.('result');
if (goResult) {
returnType = goResult.type === 'parameter_list'
? goResult.text // multi-return: "(string, error)"
: goResult.text; // single return: "int"
}
// Rust: 'return_type' field — the value IS the type node (e.g. primitive_type, type_identifier).
// Skip if the node is a type_annotation (TS/Python), which is handled by the generic loop below.
if (!returnType) {
const rustReturn = node.childForFieldName?.('return_type');
if (rustReturn && rustReturn.type !== 'type_annotation') {
returnType = rustReturn.text;
}
}
// C/C++: 'type' field on function_definition
if (!returnType) {
const cppType = node.childForFieldName?.('type');
if (cppType && cppType.text !== 'void') {
returnType = cppType.text;
}
}
// TS/Rust/Python/C#/Kotlin: type_annotation or return_type child
if (!returnType) {
for (const child of node.children) {
if (child.type === 'type_annotation' || child.type === 'return_type') {
const typeNode = child.children.find((c) => c.isNamed);
if (typeNode) returnType = typeNode.text;
}
}
}
if (isVariadic) parameterCount = undefined;
return { parameterCount, returnType };
};
/**
* Count direct arguments for a call expression across common tree-sitter grammars.
* Returns undefined when the argument container cannot be located cheaply.
*/
export const countCallArguments = (callNode: SyntaxNode | null | undefined): number | undefined => {
if (!callNode) return undefined;
// Direct field or direct child (most languages)
let argsNode: SyntaxNode | null | undefined = callNode.childForFieldName('arguments')
?? callNode.children.find((child) => CALL_ARGUMENT_LIST_TYPES.has(child.type));
// Kotlin/Swift: call_expression → call_suffix → value_arguments
// Search one level deeper for languages that wrap arguments in a suffix node
if (!argsNode) {
for (const child of callNode.children) {
if (!child.isNamed) continue;
const nested = child.children.find((gc) => CALL_ARGUMENT_LIST_TYPES.has(gc.type));
if (nested) { argsNode = nested; break; }
}
}
if (!argsNode) return undefined;
let count = 0;
for (const child of argsNode.children) {
if (!child.isNamed) continue;
if (child.type === 'comment') continue;
count++;
}
return count;
};
// ── Call-form discrimination (Phase 1, Step D) ─────────────────────────
/**
* AST node types that indicate a member-access wrapper around the callee name.
* When nameNode.parent.type is one of these, the call is a member call.
*/
const MEMBER_ACCESS_NODE_TYPES = new Set([
'member_expression', // TS/JS: obj.method()
'attribute', // Python: obj.method()
'member_access_expression', // C#: obj.Method()
'field_expression', // Rust/C++: obj.method() / ptr->method()
'selector_expression', // Go: obj.Method()
'navigation_suffix', // Kotlin/Swift: obj.method() — nameNode sits inside navigation_suffix
]);
/**
* Call node types that are inherently constructor invocations.
* Only includes patterns that the tree-sitter queries already capture as @call.
*/
const CONSTRUCTOR_CALL_NODE_TYPES = new Set([
'constructor_invocation', // Kotlin: Foo()
'new_expression', // TS/JS/C++: new Foo()
'object_creation_expression', // Java/C#/PHP: new Foo()
'implicit_object_creation_expression', // C# 9: User u = new(...)
'composite_literal', // Go: User{...}
'struct_expression', // Rust: User { ... }
]);
/**
* AST node types for scoped/qualified calls (e.g., Foo::new() in Rust, Foo::bar() in C++).
*/
const SCOPED_CALL_NODE_TYPES = new Set([
'scoped_identifier', // Rust: Foo::new()
'qualified_identifier', // C++: ns::func()
]);
type CallForm = 'free' | 'member' | 'constructor';
/**
* Infer whether a captured call site is a free call, member call, or constructor.
* Returns undefined if the form cannot be determined.
*
* Works by inspecting the AST structure between callNode (@call) and nameNode (@call.name).
* No tree-sitter query changes needed — the distinction is in the node types.
*/
export const inferCallForm = (
callNode: SyntaxNode,
nameNode: SyntaxNode,
): CallForm | undefined => {
// 1. Constructor: callNode itself is a constructor invocation (Kotlin)
if (CONSTRUCTOR_CALL_NODE_TYPES.has(callNode.type)) {
return 'constructor';
}
// 2. Member call: nameNode's parent is a member-access wrapper
const nameParent = nameNode.parent;
if (nameParent && MEMBER_ACCESS_NODE_TYPES.has(nameParent.type)) {
return 'member';
}
// 3. PHP: the callNode itself distinguishes member vs free calls
if (callNode.type === 'member_call_expression' || callNode.type === 'nullsafe_member_call_expression') {
return 'member';
}
if (callNode.type === 'scoped_call_expression') {
return 'member'; // static call Foo::bar()
}
// 4. Java method_invocation: member if it has an 'object' field
if (callNode.type === 'method_invocation' && callNode.childForFieldName('object')) {
return 'member';
}
// 5. Scoped calls (Rust Foo::new(), C++ ns::func()): treat as free
// The receiver is a type, not an instance — handled differently in Phase 3
if (nameParent && SCOPED_CALL_NODE_TYPES.has(nameParent.type)) {
return 'free';
}
// 6. Default: if nameNode is a direct child of callNode, it's a free call
if (nameNode.parent === callNode || nameParent?.parent === callNode) {
return 'free';
}
return undefined;
};
/**
* Extract the receiver identifier for member calls.
* Only captures simple identifiers — returns undefined for complex expressions
* like getUser().save() or arr[0].method().
*/
const SIMPLE_RECEIVER_TYPES = new Set([
'identifier',
'simple_identifier',
'variable_name', // PHP $variable (tree-sitter-php)
'name', // PHP name node
'this', // TS/JS/Java/C# this.method()
'self', // Rust/Python self.method()
]);
export const extractReceiverName = (
nameNode: SyntaxNode,
): string | undefined => {
const parent = nameNode.parent;
if (!parent) return undefined;
// PHP: member_call_expression / nullsafe_member_call_expression — receiver is on the callNode
// Java: method_invocation — receiver is the 'object' field on callNode
// For these, parent of nameNode is the call itself, so check the call's object field
const callNode = parent.parent ?? parent;
let receiver: SyntaxNode | null = null;
// Try standard field names used across grammars
receiver = parent.childForFieldName('object') // TS/JS member_expression, Python attribute, PHP, Java
?? parent.childForFieldName('value') // Rust field_expression
?? parent.childForFieldName('operand') // Go selector_expression
?? parent.childForFieldName('expression') // C# member_access_expression
?? parent.childForFieldName('argument'); // C++ field_expression
// Java method_invocation: 'object' field is on the callNode, not on nameNode's parent
if (!receiver && callNode.type === 'method_invocation') {
receiver = callNode.childForFieldName('object');
}
// PHP: member_call_expression has 'object' on the call node
if (!receiver && (callNode.type === 'member_call_expression' || callNode.type === 'nullsafe_member_call_expression')) {
receiver = callNode.childForFieldName('object');
}
// Kotlin/Swift: navigation_expression target is the first child
if (!receiver && parent.type === 'navigation_suffix') {
const navExpr = parent.parent;
if (navExpr?.type === 'navigation_expression') {
// First named child is the target (receiver)
for (const child of navExpr.children) {
if (child.isNamed && child !== parent) {
receiver = child;
break;
}
}
}
}
if (!receiver) return undefined;
// Only capture simple identifiers — refuse complex expressions
if (SIMPLE_RECEIVER_TYPES.has(receiver.type)) {
return receiver.text;
}
return undefined;
};
export const isVerboseIngestionEnabled = (): boolean => {
const raw = process.env.GITNEXUS_VERBOSE;
if (!raw) return false;
const value = raw.toLowerCase();
return value === '1' || value === 'true' || value === 'yes';
};
@@ -11,17 +11,35 @@ import Go from 'tree-sitter-go';
import Rust from 'tree-sitter-rust';
import Kotlin from 'tree-sitter-kotlin';
import PHP from 'tree-sitter-php';
import Ruby from 'tree-sitter-ruby';
import { createRequire } from 'node:module';
import { SupportedLanguages } from '../../../config/supported-languages.js';
import { LANGUAGE_QUERIES } from '../tree-sitter-queries.js';
import { getTreeSitterBufferSize, TREE_SITTER_MAX_BUFFER } from '../constants.js';
// tree-sitter-swift is an optionalDependency — may not be installed
const _require = createRequire(import.meta.url);
let Swift: any = null;
try { Swift = _require('tree-sitter-swift'); } catch {}
import { findSiblingChild, getLanguageFromFilename } from '../utils.js';
import {
getLanguageFromFilename,
FUNCTION_NODE_TYPES,
extractFunctionName,
isBuiltInOrNoise,
getDefinitionNodeFromCaptures,
findEnclosingClassId,
extractMethodSignature,
countCallArguments,
inferCallForm,
extractReceiverName
} from '../utils.js';
import { buildTypeEnv, lookupTypeEnv } from '../type-env.js';
import { isNodeExported } from '../export-detection.js';
import { detectFrameworkFromAST } from '../framework-detection.js';
import { generateId } from '../../../lib/utils.js';
import { extractNamedBindings } from '../named-binding-extraction.js';
import { appendKotlinWildcard } from '../resolvers/index.js';
import { routeRubyCall } from '../ruby-call-routing.js';
// ============================================================================
// Types for serializable results
@@ -35,11 +53,13 @@ interface ParsedNode {
filePath: string;
startLine: number;
endLine: number;
language: string;
language: SupportedLanguages;
isExported: boolean;
astFrameworkMultiplier?: number;
astFrameworkReason?: string;
description?: string;
parameterCount?: number;
returnType?: string;
};
}
@@ -47,7 +67,7 @@ interface ParsedRelationship {
id: string;
sourceId: string;
targetId: string;
type: 'DEFINES';
type: 'DEFINES' | 'HAS_METHOD';
confidence: number;
reason: string;
}
@@ -57,12 +77,16 @@ interface ParsedSymbol {
name: string;
nodeId: string;
type: string;
parameterCount?: number;
ownerId?: string;
}
export interface ExtractedImport {
filePath: string;
rawImportPath: string;
language: string;
language: SupportedLanguages;
/** Named bindings from the import (e.g., import {User as U} → [{local:'U', exported:'User'}]) */
namedBindings?: { local: string; exported: string }[];
}
export interface ExtractedCall {
@@ -70,6 +94,13 @@ export interface ExtractedCall {
calledName: string;
/** generateId of enclosing function, or generateId('File', filePath) for top-level */
sourceId: string;
argCount?: number;
/** Discriminates free function calls from member/constructor calls */
callForm?: 'free' | 'member' | 'constructor';
/** Simple identifier of the receiver for member calls (e.g., 'user' in user.save()) */
receiverName?: string;
/** Resolved type name of the receiver (e.g., 'User' for user.save() when user: User) */
receiverTypeName?: string;
}
export interface ExtractedHeritage {
@@ -126,6 +157,7 @@ const languageMap: Record<string, any> = {
[SupportedLanguages.Rust]: Rust,
[SupportedLanguages.Kotlin]: Kotlin,
[SupportedLanguages.PHP]: PHP.php_only,
[SupportedLanguages.Ruby]: Ruby,
...(Swift ? { [SupportedLanguages.Swift]: Swift } : {}),
};
@@ -138,198 +170,20 @@ const setLanguage = (language: SupportedLanguages, filePath: string): void => {
parser.setLanguage(lang);
};
// ============================================================================
// Export detection (copied — needs AST parent traversal, can't cross threads)
// ============================================================================
const isNodeExported = (node: any, name: string, language: string): boolean => {
let current = node;
switch (language) {
case 'javascript':
case 'typescript':
while (current) {
const type = current.type;
if (type === 'export_statement' ||
type === 'export_specifier' ||
type === 'lexical_declaration' && current.parent?.type === 'export_statement') {
return true;
}
if (current.text?.startsWith('export ')) {
return true;
}
current = current.parent;
}
return false;
case 'python':
return !name.startsWith('_');
case 'java':
while (current) {
if (current.parent) {
const parent = current.parent;
for (let i = 0; i < parent.childCount; i++) {
const child = parent.child(i);
if (child?.type === 'modifiers' && child.text?.includes('public')) {
return true;
}
}
if (parent.type === 'method_declaration' || parent.type === 'constructor_declaration') {
if (parent.text?.trimStart().startsWith('public')) {
return true;
}
}
}
current = current.parent;
}
return false;
case 'csharp':
while (current) {
if (current.type === 'modifier' || current.type === 'modifiers') {
if (current.text?.includes('public')) return true;
}
current = current.parent;
}
return false;
case 'go':
if (name.length === 0) return false;
const first = name[0];
return first === first.toUpperCase() && first !== first.toLowerCase();
case 'rust':
while (current) {
if (current.type === 'visibility_modifier') {
if (current.text?.includes('pub')) return true;
}
current = current.parent;
}
return false;
// Kotlin: Default visibility is public (unlike Java)
// visibility_modifier is inside modifiers, a sibling of the name node within the declaration
case 'kotlin':
while (current) {
if (current.parent) {
const visMod = findSiblingChild(current.parent, 'modifiers', 'visibility_modifier');
if (visMod) {
const text = visMod.text;
if (text === 'private' || text === 'internal' || text === 'protected') return false;
if (text === 'public') return true;
}
}
current = current.parent;
}
// No visibility modifier = public (Kotlin default)
return true;
case 'c':
case 'cpp':
return false;
case 'php':
// Top-level classes/interfaces/traits are always accessible
// Methods/properties are exported only if they have 'public' modifier
while (current) {
if (current.type === 'class_declaration' ||
current.type === 'interface_declaration' ||
current.type === 'trait_declaration' ||
current.type === 'enum_declaration') {
return true;
}
if (current.type === 'visibility_modifier') {
return current.text === 'public';
}
current = current.parent;
}
// Top-level functions (no parent class) are globally accessible
return true;
case 'swift':
while (current) {
if (current.type === 'modifiers' || current.type === 'visibility_modifier') {
const text = current.text || '';
if (text.includes('public') || text.includes('open')) return true;
}
current = current.parent;
}
return false;
default:
return false;
}
};
// isNodeExported imported from ../export-detection.js (shared module)
// ============================================================================
// Enclosing function detection (for call extraction)
// ============================================================================
const FUNCTION_NODE_TYPES = new Set([
'function_declaration', 'arrow_function', 'function_expression',
'method_definition', 'generator_function_declaration',
'function_definition', 'async_function_declaration', 'async_arrow_function',
'method_declaration', 'constructor_declaration',
'local_function_statement', 'function_item', 'impl_item',
// Kotlin
'lambda_literal',
// PHP
'anonymous_function',
// Swift initializers/deinitializers
'init_declaration', 'deinit_declaration',
]);
/** Walk up AST to find enclosing function, return its generateId or null for top-level */
const findEnclosingFunctionId = (node: any, filePath: string): string | null => {
let current = node.parent;
while (current) {
if (FUNCTION_NODE_TYPES.has(current.type)) {
let funcName: string | null = null;
let label = 'Function';
if (current.type === 'init_declaration' || current.type === 'deinit_declaration') {
const funcName = current.type === 'init_declaration' ? 'init' : 'deinit';
const label = 'Constructor';
const startLine = current.startPosition?.row ?? 0;
return generateId(label, `${filePath}:${funcName}:${startLine}`);
}
if (['function_declaration', 'function_definition', 'async_function_declaration',
'generator_function_declaration', 'function_item'].includes(current.type)) {
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier' || c.type === 'property_identifier');
funcName = nameNode?.text;
} else if (current.type === 'impl_item') {
const funcItem = current.children?.find((c: any) => c.type === 'function_item');
if (funcItem) {
const nameNode = funcItem.childForFieldName?.('name') ||
funcItem.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
}
} else if (current.type === 'method_definition') {
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'property_identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'method_declaration' || current.type === 'constructor_declaration') {
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'arrow_function' || current.type === 'function_expression') {
const parent = current.parent;
if (parent?.type === 'variable_declarator') {
const nameNode = parent.childForFieldName?.('name') ||
parent.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
}
}
const { funcName, label } = extractFunctionName(current);
if (funcName) {
const startLine = current.startPosition?.row ?? 0;
return generateId(label, `${filePath}:${funcName}:${startLine}`);
return generateId(label, `${filePath}:${funcName}`);
}
}
current = current.parent;
@@ -337,118 +191,6 @@ const findEnclosingFunctionId = (node: any, filePath: string): string | null =>
return null;
};
const BUILT_INS = new Set([
// JavaScript/TypeScript
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export', 'fetch', 'Response', 'Request',
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'open', 'read', 'write', 'close', 'append', 'extend', 'update',
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
// Kotlin stdlib (IMPORTANT: keep in sync with call-processor.ts BUILT_IN_NAMES)
'println', 'print', 'readLine', 'require', 'requireNotNull', 'check', 'assert', 'lazy', 'error',
'listOf', 'mapOf', 'setOf', 'mutableListOf', 'mutableMapOf', 'mutableSetOf',
'arrayOf', 'sequenceOf', 'also', 'apply', 'run', 'with', 'takeIf', 'takeUnless',
'TODO', 'buildString', 'buildList', 'buildMap', 'buildSet',
'repeat', 'synchronized',
// Kotlin coroutine builders & scope functions
'launch', 'async', 'runBlocking', 'withContext', 'coroutineScope',
'supervisorScope', 'delay',
// Kotlin Flow operators
'flow', 'flowOf', 'collect', 'emit', 'onEach', 'catch',
'buffer', 'conflate', 'distinctUntilChanged',
'flatMapLatest', 'flatMapMerge', 'combine',
'stateIn', 'shareIn', 'launchIn',
// Kotlin infix stdlib functions
'to', 'until', 'downTo', 'step',
// C/C++ standard library
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
'scanf', 'fscanf', 'sscanf',
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
'sizeof', 'offsetof', 'typeof',
'assert', 'abort', 'exit', '_exit',
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
// Linux kernel common macros/helpers (not real call targets)
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
'min', 'max', 'clamp', 'abs', 'swap',
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
'GFP_KERNEL', 'GFP_ATOMIC',
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
'mutex_lock', 'mutex_unlock', 'mutex_init',
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
'get', 'put',
// PHP built-ins
'echo', 'isset', 'empty', 'unset', 'list', 'array', 'compact', 'extract',
'count', 'strlen', 'strpos', 'strrpos', 'substr', 'strtolower', 'strtoupper', 'trim',
'ltrim', 'rtrim', 'str_replace', 'str_contains', 'str_starts_with', 'str_ends_with',
'sprintf', 'vsprintf', 'printf', 'number_format',
'array_map', 'array_filter', 'array_reduce', 'array_push', 'array_pop', 'array_shift',
'array_unshift', 'array_slice', 'array_splice', 'array_merge', 'array_keys', 'array_values',
'array_key_exists', 'in_array', 'array_search', 'array_unique', 'usort', 'rsort',
'json_encode', 'json_decode', 'serialize', 'unserialize',
'intval', 'floatval', 'strval', 'boolval', 'is_null', 'is_string', 'is_int', 'is_array',
'is_object', 'is_numeric', 'is_bool', 'is_float',
'var_dump', 'print_r', 'var_export',
'date', 'time', 'strtotime', 'mktime', 'microtime',
'file_exists', 'file_get_contents', 'file_put_contents', 'is_file', 'is_dir',
'preg_match', 'preg_match_all', 'preg_replace', 'preg_split',
'header', 'session_start', 'session_destroy', 'ob_start', 'ob_end_clean', 'ob_get_clean',
'dd', 'dump',
// Swift/iOS built-ins and standard library
'print', 'debugPrint', 'dump', 'fatalError', 'precondition', 'preconditionFailure',
'assert', 'assertionFailure', 'NSLog',
'abs', 'min', 'max', 'zip', 'stride', 'sequence', 'repeatElement',
'swap', 'withUnsafePointer', 'withUnsafeMutablePointer', 'withUnsafeBytes',
'autoreleasepool', 'unsafeBitCast', 'unsafeDowncast', 'numericCast',
'type', 'MemoryLayout',
// Swift collection/string methods (common noise)
'map', 'flatMap', 'compactMap', 'filter', 'reduce', 'forEach', 'contains',
'first', 'last', 'prefix', 'suffix', 'dropFirst', 'dropLast',
'sorted', 'reversed', 'enumerated', 'joined', 'split',
'append', 'insert', 'remove', 'removeAll', 'removeFirst', 'removeLast',
'isEmpty', 'count', 'index', 'startIndex', 'endIndex',
// UIKit/Foundation common methods (noise in call graph)
'addSubview', 'removeFromSuperview', 'layoutSubviews', 'setNeedsLayout',
'layoutIfNeeded', 'setNeedsDisplay', 'invalidateIntrinsicContentSize',
'addTarget', 'removeTarget', 'addGestureRecognizer',
'addConstraint', 'addConstraints', 'removeConstraint', 'removeConstraints',
'NSLocalizedString', 'Bundle',
'reloadData', 'reloadSections', 'reloadRows', 'performBatchUpdates',
'register', 'dequeueReusableCell', 'dequeueReusableSupplementaryView',
'beginUpdates', 'endUpdates', 'insertRows', 'deleteRows', 'insertSections', 'deleteSections',
'present', 'dismiss', 'pushViewController', 'popViewController', 'popToRootViewController',
'performSegue', 'prepare',
// GCD / async
'DispatchQueue', 'async', 'sync', 'asyncAfter',
'Task', 'withCheckedContinuation', 'withCheckedThrowingContinuation',
// Combine
'sink', 'store', 'assign', 'receive', 'subscribe',
// Notification / KVO
'addObserver', 'removeObserver', 'post', 'NotificationCenter',
]);
// ============================================================================
// Label detection from capture map
// ============================================================================
@@ -483,50 +225,8 @@ const getLabelFromCaptures = (captureMap: Record<string, any>): string | null =>
return 'CodeElement';
};
const DEFINITION_CAPTURE_KEYS = [
'definition.function',
'definition.class',
'definition.interface',
'definition.method',
'definition.struct',
'definition.enum',
'definition.namespace',
'definition.module',
'definition.trait',
'definition.impl',
'definition.type',
'definition.const',
'definition.static',
'definition.typedef',
'definition.macro',
'definition.union',
'definition.property',
'definition.record',
'definition.delegate',
'definition.annotation',
'definition.constructor',
'definition.template',
] as const;
// DEFINITION_CAPTURE_KEYS and getDefinitionNodeFromCaptures imported from ../utils.js
const getDefinitionNodeFromCaptures = (captureMap: Record<string, any>): any | null => {
for (const key of DEFINITION_CAPTURE_KEYS) {
if (captureMap[key]) return captureMap[key];
}
return null;
};
/**
* Append .* to a Kotlin import path if the AST has a wildcard_import sibling node.
* Pure function — returns a new string without mutating the input.
*/
const appendKotlinWildcard = (importPath: string, importNode: any): string => {
for (let i = 0; i < importNode.childCount; i++) {
if (importNode.child(i)?.type === 'wildcard_import') {
return importPath.endsWith('.*') ? importPath : `${importPath}.*`;
}
}
return importPath;
};
// ============================================================================
// Process a batch of files
@@ -1099,28 +799,39 @@ const processFileGroup = (
try {
const lang = parser.getLanguage();
query = new Parser.Query(lang, queryString);
} catch {
} catch (err) {
const message = `Query compilation failed for ${language}: ${err instanceof Error ? err.message : String(err)}`;
if (parentPort) {
parentPort.postMessage({ type: 'warning', message });
} else {
console.warn(message);
}
return;
}
for (const file of files) {
// Skip very large files — they can crash tree-sitter or cause OOM
if (file.content.length > 512 * 1024) continue;
// Skip files larger than the max tree-sitter buffer (32 MB)
if (file.content.length > TREE_SITTER_MAX_BUFFER) continue;
let tree;
try {
tree = parser.parse(file.content, undefined, { bufferSize: 1024 * 256 });
} catch {
tree = parser.parse(file.content, undefined, { bufferSize: getTreeSitterBufferSize(file.content.length) });
} catch (err) {
console.warn(`Failed to parse file ${file.path}: ${err instanceof Error ? err.message : String(err)}`);
continue;
}
result.fileCount++;
onFileProcessed?.();
// Build per-file TypeEnv from explicit type annotations (for receiver resolution)
const typeEnv = buildTypeEnv(tree, language);
let matches;
try {
matches = query.matches(tree.rootNode);
} catch {
} catch (err) {
console.warn(`Query execution failed for ${file.path}: ${err instanceof Error ? err.message : String(err)}`);
continue;
}
@@ -1135,10 +846,12 @@ const processFileGroup = (
const rawImportPath = language === SupportedLanguages.Kotlin
? appendKotlinWildcard(captureMap['import.source'].text.replace(/['"<>]/g, ''), captureMap['import'])
: captureMap['import.source'].text.replace(/['"<>]/g, '');
const namedBindings = extractNamedBindings(captureMap['import'], language);
result.imports.push({
filePath: file.path,
rawImportPath,
language: language,
...(namedBindings ? { namedBindings } : {}),
});
continue;
}
@@ -1148,11 +861,107 @@ const processFileGroup = (
const callNameNode = captureMap['call.name'];
if (callNameNode) {
const calledName = callNameNode.text;
if (!BUILT_INS.has(calledName)) {
// Ruby: route special calls to imports, heritage, or properties
if (language === SupportedLanguages.Ruby) {
const callNode = captureMap['call'];
const routed = routeRubyCall(calledName, callNode);
switch (routed.kind) {
case 'skip':
continue;
case 'import':
result.imports.push({
filePath: file.path,
rawImportPath: routed.importPath,
language,
});
continue;
case 'heritage':
for (const item of routed.items) {
result.heritage.push({
filePath: file.path,
className: item.enclosingClass,
parentName: item.mixinName,
kind: 'trait-impl',
});
}
continue;
case 'properties':
for (const item of routed.items) {
const nodeId = generateId('Property', `${file.path}:${item.propName}`);
result.nodes.push({
id: nodeId,
label: 'Property',
properties: {
name: item.propName,
filePath: file.path,
startLine: item.startLine,
endLine: item.endLine,
language,
isExported: true,
description: item.accessorType,
},
});
result.symbols.push({
filePath: file.path,
name: item.propName,
nodeId,
type: 'Property',
});
const fileId = generateId('File', file.path);
const relId = generateId('DEFINES', `${fileId}->${nodeId}`);
result.relationships.push({
id: relId,
sourceId: fileId,
targetId: nodeId,
type: 'DEFINES',
confidence: 1.0,
reason: '',
});
}
continue;
case 'call':
if (!isBuiltInOrNoise(calledName)) {
const sourceId = findEnclosingFunctionId(callNode, file.path)
|| generateId('File', file.path);
const callForm = inferCallForm(callNode, callNameNode);
const receiverName = callForm === 'member' ? extractReceiverName(callNameNode) : undefined;
const receiverTypeName = receiverName ? lookupTypeEnv(typeEnv, receiverName, callNode) : undefined;
result.calls.push({
filePath: file.path,
calledName,
sourceId,
argCount: countCallArguments(callNode),
...(callForm !== undefined ? { callForm } : {}),
...(receiverName !== undefined ? { receiverName } : {}),
...(receiverTypeName !== undefined ? { receiverTypeName } : {}),
});
}
continue;
}
}
if (!isBuiltInOrNoise(calledName)) {
const callNode = captureMap['call'];
const sourceId = findEnclosingFunctionId(callNode, file.path)
|| generateId('File', file.path);
result.calls.push({ filePath: file.path, calledName, sourceId });
const callForm = inferCallForm(callNode, callNameNode);
const receiverName = callForm === 'member' ? extractReceiverName(callNameNode) : undefined;
const receiverTypeName = receiverName ? lookupTypeEnv(typeEnv, receiverName, callNode) : undefined;
result.calls.push({
filePath: file.path,
calledName,
sourceId,
argCount: countCallArguments(callNode),
...(callForm !== undefined ? { callForm } : {}),
...(receiverName !== undefined ? { receiverName } : {}),
...(receiverTypeName !== undefined ? { receiverTypeName } : {}),
});
}
}
continue;
@@ -1161,12 +970,21 @@ const processFileGroup = (
// Extract heritage (extends/implements)
if (captureMap['heritage.class']) {
if (captureMap['heritage.extends']) {
result.heritage.push({
filePath: file.path,
className: captureMap['heritage.class'].text,
parentName: captureMap['heritage.extends'].text,
kind: 'extends',
});
// Go struct embedding: the query matches ALL field_declarations with
// type_identifier, but only anonymous fields (no name) are embedded.
// Named fields like `Breed string` also match — skip them.
const extendsNode = captureMap['heritage.extends'];
const fieldDecl = extendsNode.parent;
const isNamedField = fieldDecl?.type === 'field_declaration'
&& fieldDecl.childForFieldName('name');
if (!isNamedField) {
result.heritage.push({
filePath: file.path,
className: captureMap['heritage.class'].text,
parentName: captureMap['heritage.extends'].text,
kind: 'extends',
});
}
}
if (captureMap['heritage.implements']) {
result.heritage.push({
@@ -1198,7 +1016,7 @@ const processFileGroup = (
const nodeName = nameNode ? nameNode.text : 'init';
const definitionNode = getDefinitionNodeFromCaptures(captureMap);
const startLine = definitionNode ? definitionNode.startPosition.row : (nameNode ? nameNode.startPosition.row : 0);
const nodeId = generateId(nodeLabel, `${file.path}:${nodeName}:${startLine}`);
const nodeId = generateId(nodeLabel, `${file.path}:${nodeName}`);
let description: string | undefined;
if (language === SupportedLanguages.PHP) {
@@ -1213,6 +1031,14 @@ const processFileGroup = (
? detectFrameworkFromAST(language, (definitionNode.text || '').slice(0, 300))
: null;
let parameterCount: number | undefined;
let returnType: string | undefined;
if (nodeLabel === 'Function' || nodeLabel === 'Method' || nodeLabel === 'Constructor') {
const sig = extractMethodSignature(definitionNode);
parameterCount = sig.parameterCount;
returnType = sig.returnType;
}
result.nodes.push({
id: nodeId,
label: nodeLabel,
@@ -1228,14 +1054,23 @@ const processFileGroup = (
astFrameworkReason: frameworkHint.reason,
} : {}),
...(description !== undefined ? { description } : {}),
...(parameterCount !== undefined ? { parameterCount } : {}),
...(returnType !== undefined ? { returnType } : {}),
},
});
// Compute enclosing class for Method/Constructor/Property/Function — used for both ownerId and HAS_METHOD
// Function is included because Kotlin/Rust/Python capture class methods as Function nodes
const needsOwner = nodeLabel === 'Method' || nodeLabel === 'Constructor' || nodeLabel === 'Property' || nodeLabel === 'Function';
const enclosingClassId = needsOwner ? findEnclosingClassId(nameNode || definitionNode, file.path) : null;
result.symbols.push({
filePath: file.path,
name: nodeName,
nodeId,
type: nodeLabel,
...(parameterCount !== undefined ? { parameterCount } : {}),
...(enclosingClassId ? { ownerId: enclosingClassId } : {}),
});
const fileId = generateId('File', file.path);
@@ -1248,6 +1083,18 @@ const processFileGroup = (
confidence: 1.0,
reason: '',
});
// ── HAS_METHOD: link method/constructor/property to enclosing class ──
if (enclosingClassId) {
result.relationships.push({
id: generateId('HAS_METHOD', `${enclosingClassId}->${nodeId}`),
sourceId: enclosingClassId,
targetId: nodeId,
type: 'HAS_METHOD',
confidence: 1.0,
reason: '',
});
}
}
// Extract Laravel routes from route files via procedural AST walk
+19 -3
View File
@@ -232,7 +232,8 @@ export const streamAllCSVsToDisk = async (
const functionWriter = new BufferedCSVWriter(path.join(csvDir, 'function.csv'), codeElementHeader);
const classWriter = new BufferedCSVWriter(path.join(csvDir, 'class.csv'), codeElementHeader);
const interfaceWriter = new BufferedCSVWriter(path.join(csvDir, 'interface.csv'), codeElementHeader);
const methodWriter = new BufferedCSVWriter(path.join(csvDir, 'method.csv'), codeElementHeader);
const methodHeader = 'id,name,filePath,startLine,endLine,isExported,content,description,parameterCount,returnType';
const methodWriter = new BufferedCSVWriter(path.join(csvDir, 'method.csv'), methodHeader);
const codeElemWriter = new BufferedCSVWriter(path.join(csvDir, 'codeelement.csv'), codeElementHeader);
const communityWriter = new BufferedCSVWriter(path.join(csvDir, 'community.csv'), 'id,label,heuristicLabel,keywords,description,enrichedBy,cohesion,symbolCount');
const processWriter = new BufferedCSVWriter(path.join(csvDir, 'process.csv'), 'id,label,heuristicLabel,processType,stepCount,communities,entryPointId,terminalId');
@@ -250,7 +251,6 @@ export const streamAllCSVsToDisk = async (
'Function': functionWriter,
'Class': classWriter,
'Interface': interfaceWriter,
'Method': methodWriter,
'CodeElement': codeElemWriter,
};
@@ -308,8 +308,24 @@ export const streamAllCSVsToDisk = async (
].join(','));
break;
}
case 'Method': {
const content = await extractContent(node, contentCache);
await methodWriter.addRow([
escapeCSVField(node.id),
escapeCSVField(node.properties.name || ''),
escapeCSVField(node.properties.filePath || ''),
escapeCSVNumber(node.properties.startLine, -1),
escapeCSVNumber(node.properties.endLine, -1),
node.properties.isExported ? 'true' : 'false',
escapeCSVField(content),
escapeCSVField((node.properties as any).description || ''),
escapeCSVNumber(node.properties.parameterCount, 0),
escapeCSVField(node.properties.returnType || ''),
].join(','));
break;
}
default: {
// Code element nodes (Function, Class, Interface, Method, CodeElement)
// Code element nodes (Function, Class, Interface, CodeElement)
const writer = codeWriterMap[node.label];
if (writer) {
const content = await extractContent(node, contentCache);
+3
View File
@@ -328,6 +328,9 @@ const getCopyQuery = (table: NodeTableName, filePath: string): string => {
if (table === 'Process') {
return `COPY ${t}(id, label, heuristicLabel, processType, stepCount, communities, entryPointId, terminalId) FROM "${filePath}" ${COPY_CSV_OPTS}`;
}
if (table === 'Method') {
return `COPY ${t}(id, name, filePath, startLine, endLine, isExported, content, description, parameterCount, returnType) FROM "${filePath}" ${COPY_CSV_OPTS}`;
}
// TypeScript/JS code element tables have isExported; multi-language tables do not
if (TABLES_WITH_EXPORTED.has(table)) {
return `COPY ${t}(id, name, filePath, startLine, endLine, isExported, content, description) FROM "${filePath}" ${COPY_CSV_OPTS}`;
+16 -1
View File
@@ -26,7 +26,7 @@ export type NodeTableName = typeof NODE_TABLES[number];
export const REL_TABLE_NAME = 'CodeRelation';
// Valid relation types
export const REL_TYPES = ['CONTAINS', 'DEFINES', 'IMPORTS', 'CALLS', 'EXTENDS', 'IMPLEMENTS', 'MEMBER_OF', 'STEP_IN_PROCESS'] as const;
export const REL_TYPES = ['CONTAINS', 'DEFINES', 'IMPORTS', 'CALLS', 'EXTENDS', 'IMPLEMENTS', 'HAS_METHOD', 'OVERRIDES', 'MEMBER_OF', 'STEP_IN_PROCESS'] as const;
export type RelType = typeof REL_TYPES[number];
// ============================================================================
@@ -104,6 +104,8 @@ CREATE NODE TABLE Method (
isExported BOOLEAN,
content STRING,
description STRING,
parameterCount INT32,
returnType STRING,
PRIMARY KEY (id)
)`;
@@ -260,6 +262,7 @@ CREATE REL TABLE ${REL_TABLE_NAME} (
FROM Class TO \`Union\`,
FROM Class TO \`Namespace\`,
FROM Class TO \`Typedef\`,
FROM Class TO \`Property\`,
FROM Method TO Function,
FROM Method TO Method,
FROM Method TO Class,
@@ -295,6 +298,7 @@ CREATE REL TABLE ${REL_TABLE_NAME} (
FROM Interface TO \`TypeAlias\`,
FROM Interface TO \`Struct\`,
FROM Interface TO \`Constructor\`,
FROM Interface TO \`Property\`,
FROM \`Struct\` TO Community,
FROM \`Struct\` TO \`Trait\`,
FROM \`Struct\` TO \`Struct\`,
@@ -303,6 +307,8 @@ CREATE REL TABLE ${REL_TABLE_NAME} (
FROM \`Struct\` TO Function,
FROM \`Struct\` TO Method,
FROM \`Struct\` TO Interface,
FROM \`Struct\` TO \`Constructor\`,
FROM \`Struct\` TO \`Property\`,
FROM \`Enum\` TO \`Enum\`,
FROM \`Enum\` TO Community,
FROM \`Enum\` TO Class,
@@ -316,7 +322,13 @@ CREATE REL TABLE ${REL_TABLE_NAME} (
FROM \`Union\` TO Community,
FROM \`Namespace\` TO Community,
FROM \`Namespace\` TO \`Struct\`,
FROM \`Trait\` TO Method,
FROM \`Trait\` TO \`Constructor\`,
FROM \`Trait\` TO \`Property\`,
FROM \`Trait\` TO Community,
FROM \`Impl\` TO Method,
FROM \`Impl\` TO \`Constructor\`,
FROM \`Impl\` TO \`Property\`,
FROM \`Impl\` TO Community,
FROM \`Impl\` TO \`Trait\`,
FROM \`Impl\` TO \`Struct\`,
@@ -327,6 +339,9 @@ CREATE REL TABLE ${REL_TABLE_NAME} (
FROM \`Const\` TO Community,
FROM \`Static\` TO Community,
FROM \`Property\` TO Community,
FROM \`Record\` TO Method,
FROM \`Record\` TO \`Constructor\`,
FROM \`Record\` TO \`Property\`,
FROM \`Record\` TO Community,
FROM \`Delegate\` TO Community,
FROM \`Annotation\` TO Community,
@@ -10,6 +10,7 @@ import Go from 'tree-sitter-go';
import Rust from 'tree-sitter-rust';
import Kotlin from 'tree-sitter-kotlin';
import PHP from 'tree-sitter-php';
import Ruby from 'tree-sitter-ruby';
import { createRequire } from 'node:module';
import { SupportedLanguages } from '../../config/supported-languages.js';
@@ -33,6 +34,7 @@ const languageMap: Record<string, any> = {
[SupportedLanguages.Rust]: Rust,
[SupportedLanguages.Kotlin]: Kotlin,
[SupportedLanguages.PHP]: PHP.php_only,
[SupportedLanguages.Ruby]: Ruby,
...(Swift ? { [SupportedLanguages.Swift]: Swift } : {}),
};
+1
View File
@@ -32,6 +32,7 @@ export function isTestFilePath(filePath: string): boolean {
p.includes('/test/') || p.includes('/tests/') ||
p.includes('/testing/') || p.includes('/fixtures/') ||
p.endsWith('_test.go') || p.endsWith('_test.py') ||
p.endsWith('_spec.rb') || p.endsWith('_test.rb') || p.includes('/spec/') ||
p.includes('/test_') || p.includes('/conftest.')
);
}
+12 -3
View File
@@ -78,7 +78,7 @@ SCHEMA:
- Nodes: File, Folder, Function, Class, Interface, Method, CodeElement, Community, Process
- Multi-language nodes (use backticks): \`Struct\`, \`Enum\`, \`Trait\`, \`Impl\`, etc.
- All edges via single CodeRelation table with 'type' property
- Edge types: CONTAINS, DEFINES, CALLS, IMPORTS, EXTENDS, IMPLEMENTS, MEMBER_OF, STEP_IN_PROCESS
- Edge types: CONTAINS, DEFINES, CALLS, IMPORTS, EXTENDS, IMPLEMENTS, HAS_METHOD, OVERRIDES, MEMBER_OF, STEP_IN_PROCESS
- Edge properties: type (STRING), confidence (DOUBLE), reason (STRING), step (INT32)
EXAMPLES:
@@ -91,6 +91,15 @@ EXAMPLES:
• Trace a process:
MATCH (s)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process) WHERE p.heuristicLabel = "UserLogin" RETURN s.name, r.step ORDER BY r.step
• Find all methods of a class:
MATCH (c:Class {name: "UserService"})-[r:CodeRelation {type: 'HAS_METHOD'}]->(m:Method) RETURN m.name, m.parameterCount, m.returnType
• Find method overrides (MRO resolution):
MATCH (winner:Method)-[r:CodeRelation {type: 'OVERRIDES'}]->(loser:Method) RETURN winner.name, winner.filePath, loser.filePath, r.reason
• Detect diamond inheritance:
MATCH (d:Class)-[:CodeRelation {type: 'EXTENDS'}]->(b1), (d)-[:CodeRelation {type: 'EXTENDS'}]->(b2), (b1)-[:CodeRelation {type: 'EXTENDS'}]->(a), (b2)-[:CodeRelation {type: 'EXTENDS'}]->(a) WHERE b1 <> b2 RETURN d.name, b1.name, b2.name, a.name
OUTPUT: Returns { markdown, row_count } — results formatted as a Markdown table for easy reading.
TIPS:
@@ -191,7 +200,7 @@ Depth groups:
- d=2: LIKELY AFFECTED (indirect)
- d=3: MAY NEED TESTING (transitive)
EdgeType: CALLS, IMPORTS, EXTENDS, IMPLEMENTS
EdgeType: CALLS, IMPORTS, EXTENDS, IMPLEMENTS, HAS_METHOD, OVERRIDES
Confidence: 1.0 = certain, <0.8 = fuzzy match`,
inputSchema: {
type: 'object',
@@ -199,7 +208,7 @@ Confidence: 1.0 = certain, <0.8 = fuzzy match`,
target: { type: 'string', description: 'Name of function, class, or file to analyze' },
direction: { type: 'string', description: 'upstream (what depends on this) or downstream (what this depends on)' },
maxDepth: { type: 'number', description: 'Max relationship depth (default: 3)', default: 3 },
relationTypes: { type: 'array', items: { type: 'string' }, description: 'Filter: CALLS, IMPORTS, EXTENDS, IMPLEMENTS (default: usage-based)' },
relationTypes: { type: 'array', items: { type: 'string' }, description: 'Filter: CALLS, IMPORTS, EXTENDS, IMPLEMENTS, HAS_METHOD, OVERRIDES (default: usage-based)' },
includeTests: { type: 'boolean', description: 'Include test files (default: false)' },
minConfidence: { type: 'number', description: 'Minimum confidence 0-1 (default: 0.7)' },
repo: { type: 'string', description: 'Repository name or path. Omit if only one repo is indexed.' },
@@ -0,0 +1,6 @@
#pragma once
class Handler {
public:
void handle();
};
@@ -0,0 +1,6 @@
#pragma once
class Handler {
public:
void process();
};
@@ -0,0 +1,8 @@
#pragma once
#include "handler_a.h"
class Processor : public Handler {
public:
void run();
};
@@ -0,0 +1,6 @@
#include "one.h"
#include "zero.h"
void run() {
write_audit("hello");
}
@@ -0,0 +1,3 @@
inline const char* write_audit(const char* message) {
return message;
}
@@ -0,0 +1,3 @@
inline const char* write_audit() {
return "zero";
}
@@ -0,0 +1,6 @@
#include "user.h"
void processUser(const std::string& name) {
auto user = new User(name);
user->save();
}
@@ -0,0 +1,10 @@
#pragma once
#include <string>
class User {
public:
User(const std::string& name) : name_(name) {}
bool save() { return true; }
private:
std::string name_;
};
@@ -0,0 +1,7 @@
#pragma once
class Animal {
public:
virtual void speak();
virtual void move();
};
@@ -0,0 +1,5 @@
#include "duck.h"
void Duck::speak() {
// quack
}
@@ -0,0 +1,8 @@
#pragma once
#include "flyer.h"
#include "swimmer.h"
class Duck : public Flyer, public Swimmer {
public:
void speak() override;
};
@@ -0,0 +1,8 @@
#pragma once
#include "animal.h"
class Flyer : public Animal {
public:
void move() override;
void fly();
};
@@ -0,0 +1,8 @@
#pragma once
#include "animal.h"
class Swimmer : public Animal {
public:
void move() override;
void swim();
};
@@ -0,0 +1,3 @@
cmake_minimum_required(VERSION 3.10)
project(cpp-local-shadow)
add_executable(main src/main.cpp src/utils.cpp)

Some files were not shown because too many files have changed in this diff Show More