Compare commits

...
Author SHA1 Message Date
Gergo Magyar 23b25c1da4 feat: add type resolution system and roadmap documentation 2026-03-17 17:08:59 +00:00
Gergo Magyar bf731ec058 fix: review findings — remove template_string from SKIP_SUBTREE_TYPES, handle bare nullable keywords
- Remove template_string and concatenated_string from SKIP_SUBTREE_TYPES
  (template literals contain interpolated expressions with typed code)
- Add FAST_NULLABLE_KEYWORDS check to fastStripNullable for behavioral
  parity with stripNullable on bare null/undefined/void/None/nil
- Add explanatory comment on extractPendingAssignment scopeEnv guard
2026-03-17 17:02:30 +00:00
Gergo Magyar be1b34ec3d perf: optimize type resolution pipeline — worker threshold, skip graph phases, AST pruning
- Skip worker pool creation for small repos (<15 files or <512KB) — saves 100-400ms
- Add skipGraphPhases option to runPipelineFromRepo to skip MRO/community/process phases
- Add conservative SKIP_SUBTREE_TYPES for leaf-only AST nodes (string, comment, number)
- Pre-compute interestingNodeTypes set — single Set.has() replaces 3 checks per node
- Add fastStripNullable — skip full stripNullable for simple identifiers (90%+ case)
- Replace .children?.find() with manual for loops in extractFunctionName (no array alloc)
- Add hookTimeout: 120000 to vitest.config.ts for CI beforeAll hooks
2026-03-17 16:57:59 +00:00
Gergo Magyar f847685bfe feat: Phase 6.2 review findings — C# nested member foreach, C++ deref range-for, Java field_access
Close two gaps found during fourth-pass review of PR #318:

- C# foreach (var user in this.data.Values): nested member_access_expression
  now extracts intermediate property name for scopeEnv lookup
- C++ for (auto& user : *ptr): pointer_expression dereference now recognized
  as range-for iterable

Root causes fixed in shared infrastructure:
- extractSimpleTypeName: add template_type (C++) and generic_name (C#)
- extractGenericTypeArgs: add generic_name for consistency
- type-env.ts: unwrap variable_declaration wrapper in field_declaration
  for declarationTypeNodes capture (zero-allocation manual loop)

Additional review findings addressed:
- Java: add field_access handler for this.data.values() in method_invocation
- C++ pointer_expression: document limitation (*identifier only)
- TypeScript: fix stale comment about property_identifier

All 525 tests pass (278 unit + 247 integration).
2026-03-17 16:21:47 +00:00
Gergo Magyar ae49c4cce7 docs: add type resolution system documentation with roadmap
Covers the full architecture, resolution tiers (0-2), scope model,
language feature matrix, container descriptors, pipeline integration,
and the Phase 7-9 roadmap for cross-scope propagation, field-type
resolution, and return-type-aware binding.
2026-03-17 12:35:36 +00:00
Gergo Magyar f092b33e40 feat: Phase 6.1 type resolution gap closure — container descriptors, recursive_pattern, class fields
Add 13 missing container type descriptors (Collection, MutableMap, Stream, SortedSet, etc.)
to CONTAINER_DESCRIPTORS for correct element type extraction across C#, Kotlin, and Java.

Extend C# pattern binding to handle recursive_pattern (obj is User { Name: "Alice" } u)
in both is-expression and switch expression contexts.

Add TypeScript class field declaration support (public_field_definition) so for-loop
iteration over this.fieldName resolves element types from class field type annotations.
Includes file-scope fallback in resolveIterableElementType and nested member_expression
handling for this.field.method() patterns.
2026-03-17 12:25:36 +00:00
Gergo Magyar 3c5a62982d feat: enhance PHP type resolution for generics and member access in foreach loops 2026-03-17 11:05:28 +00:00
Gergo Magyar 08902f8a18 fix: position-indexed when/is bindings, Kotlin param extraction, HashMap.values for-loop
Three root causes for failing Kotlin integration tests:

1. When/is multi-arm resolution: flat scopeEnv stored only the last arm's
   type (last-writer-wins). Added PatternOverrides with AST range indexing
   so each when arm resolves to its narrowed type independently.

2. HashMap.values for-loop: navigation_expression without call_suffix was
   classified as bare property access (iterableName='values' instead of
   'data'). Now tries object-as-iterable + property-as-method first, with
   fallback to property-as-iterable for this.users patterns.

3. Kotlin parameter extraction: tree-sitter-kotlin parameter nodes use
   positional children (simple_identifier, user_type) not named fields
   (name, type). Added fallback to findChildByType in both
   extractKotlinParameter and extractTypeBinding.

Integration tests added for .keys/.values/Set/MutableMap iteration,
3-arm when/is, multi-call within arms, and when+else branch.
2026-03-17 10:14:03 +00:00
Gergo Magyar 294bfddaf2 feat: PR #318 review findings — pattern bindings, member access iterables, structured bindings
Address all 7 genuine gaps identified in PR #318 deep code review:

- Kotlin: add extractKotlinPatternBinding for when/is (type_test AST node)
  with allowPatternBindingOverwrite for smart-cast semantics
- Java: add type_pattern branch for Java 17+ switch pattern variables
- TypeScript: explicit object_pattern skip in for-of (no false bindings)
- Cross-language: member access iterables (self.users, this.users, repo.users)
  across all 10 language extractors
- C++: structured_binding_declarator handling in range-for (last-child heuristic)
- Rust: closure_parameter added to TYPED_PARAMETER_TYPES
- PHP: normalizePhpType handles angle-bracket generics (Collection<User>)

Code review fixes applied:
- Remove 4 debug console.log statements (c-cpp.ts, call-processor.ts)
- Hoist KNOWN_CONTAINER_PROPS to module scope (csharp.ts)
- Guard keysBefore allocation behind typeNode check (type-env.ts)
- Add depth limits (50) to 7 recursive type extraction functions
- Add 2048-char length cap to extractSimpleTypeName
- Fix PHP/Ruby missing typeArgPos parameter in resolveIterableElementType

Integration test fixtures: kotlin-when-pattern, java-switch-pattern,
cpp-structured-binding, typescript-member-access-for-loop,
python-member-access-for-loop
2026-03-17 09:21:27 +00:00
Gergo Magyar 7a92dce6f4 fix: rename C++ fixture files to correct case for case-sensitive CI
On case-sensitive filesystems (Linux/macOS CI), git tracked both the old
lowercase files (app.cpp, user.h) and the new uppercase files (App.cpp,
User.h) as separate files. The pipeline processed both, causing the old
app.cpp (with explicit User& type) to interfere with the new auto& test.

Removes old lowercase entries and re-adds with uppercase casing to match
the #include directives in the fixture.
2026-03-17 07:34:21 +00:00
Gergo Magyar c82aa58fe1 fix: update extractElementTypeFromString tests for last-arg default
TypeArgPosition change (default 'last') broke 5 existing tests expecting
first arg from multi-arg generics. Updated expectations and added explicit
pos='first' tests for key type extraction.
2026-03-17 07:04:04 +00:00
Gergo Magyar d05aa9ef6f feat: method-aware for-loop extractors + integration tests for all languages
Upgrade 4 existing extractors + create 3 new ones for full cross-language
coverage of call_expression iterables and container descriptor resolution:

Upgraded (add call expr iterable + methodToTypeArgPosition):
- Java: method_invocation (data.keySet(), data.values())
- Kotlin: navigation_expression + call_expression (data.keys, data.values())
- C#: member_access_expression + invocation_expression (data.Keys, data.Values)
- Go: TypeArgPosition threading for Go 1.18+ generics

New for-loop extractors:
- C++: for_range_loop with auto& unwrapping, template_type + qualified_identifier
  (std::vector<User>) extraction, explicit vs auto type handling
- PHP: foreach_statement with simple/key-value/by-reference forms, PHPDoc
  @param priority over AST array type
- Ruby: for-in with YARD @param type resolution via comment parsing

Integration test fixtures + tests for all 6 languages:
- java-map-keys-values (Map.values() + List iteration)
- kotlin-map-keys-values (HashMap.values + List iteration)
- csharp-dictionary-keys-values (Dictionary.Values foreach)
- cpp-range-for (auto& + const auto& range-based for)
- php-foreach-loop (foreach with PHPDoc @param User[])
- ruby-for-in-loop (for-in with YARD @param Array<User>)

Bugs fixed during integration testing:
- C++: qualified_identifier (std::vector) not unwrapped to template_type
- PHP: extractParameter overwrote PHPDoc-derived types with bare 'array'

252 unit tests pass, 201 integration tests pass across 6 languages.
2026-03-17 06:52:26 +00:00
Gergo Magyar b4986fdaba feat: container descriptor table for generic type arg resolution
Replace simple KEY_METHODS heuristic with CONTAINER_DESCRIPTORS table
that maps 30+ container types across all languages to their type parameter
semantics per access method.

Key improvements:
- Container-aware resolution: HashMap.iter() correctly yields V (arity 2),
  while Vec.iter() yields T (arity 1) — same method, different semantics
- Cross-language coverage: Map/HashMap/BTreeMap/dict/Dict/Dictionary/
  ConcurrentHashMap + List/Vec/Set/HashSet/Queue/Deque/Stack etc.
- Method categorization: keyMethods (keys/keySet/Keys) vs valueMethods
  (values/get/pop/iter/first/last) per container type
- Fallback for unknown containers: still uses method name heuristic,
  so MyCache<K,V>.keys() correctly returns first arg
- Exported getContainerDescriptor() for future heritage-chain lookups

Each language extractor now passes containerTypeName from scopeEnv to
methodToTypeArgPosition for descriptor-aware resolution.

252 unit tests pass (4 new descriptor tests), 1 skip (Ruby).
2026-03-17 06:27:05 +00:00
Gergo Magyar 656af32e52 feat: resolve 4 known limitation skip tests + method-aware type arg selection
Unskip 4 of 5 type-env known limitations with full integration test coverage:

1. TS destructured for-of: handle array_pattern by binding last named child
   to element type. Fix Map<K,V> to return last generic arg (value type).
2. Python dict.items() loop: handle `call` iterables + `pattern_list` left
   side. Fix dict[K,V] extraction via type_parameter with last-arg heuristic.
   Unwrap `type` wrapper in extractPyElementTypeFromAnnotation.
3. TS instanceof narrowing: add extractPatternBinding for binary_expression
   with positional child access. First-writer-wins (not block-scoped).
4. Rust .iter() for-loops: handle call_expression in for_expression value
   node by extracting receiver from field_expression.

Method-aware type arg resolution:
- Add TypeArgPosition ('first'|'last') to resolveIterableElementType
- .keys()/.keySet()/.Keys → first type arg (key); all else → last (value)
- Thread position through all 3 strategy callbacks in TS/Rust/Python
- Add predefined_type to extractSimpleTypeName for TS primitives (string etc)

New fixtures: rust-iter-for-loop, typescript-destructured-for-of,
typescript-instanceof-narrowing, python-dict-items-loop.
248 unit tests pass (6 new), 1 skip (Ruby block params).
2026-03-17 06:24:09 +00:00
Gergo Magyar 5f12ffccf2 test: add assertion bodies to known limitation skip tests
Convert empty skip test stubs to proper tests with parse/buildTypeEnv/expect
assertions following the codebase convention (e.g., call-processor.test.ts:319).
Each skip test now documents the exact expected behavior, so removing .skip
will cause a meaningful failure when the limitation is eventually fixed.

Also clarify Python integration skip tests as call-extraction issues (not
type-env) and Swift integration skips as build-dep issues (self/super
resolution code already exists in type-env.ts).
2026-03-16 23:13:05 +00:00
Gergo Magyar f1df9a12c1 test: integration tests for all Phase 6 language gaps + fix Rust param pattern field
Integration test fixtures and tests (30 new tests, all with exact match + negative):

Rust for-loop (5 tests):
- for user in &users with Vec<User> → User#save, negative Repo#save
- for repo in &repos with Vec<Repo> → Repo#save, negative User#save

Rust match arm (5 tests):
- match opt { Some(user) => user.save() } → User#save, negative Repo#save
- if let Ok(repo) = res → Repo#save, negative User#save

C# var foreach (5 tests):
- foreach (var user in users) with List<User> → User#Save, negative Repo#Save
- foreach (var repo in repos) with List<Repo> → Repo#Save

C# switch pattern (4 tests):
- is User user → User#Save, case Repo repo → Repo#Save

Kotlin unannotated for (4 tests):
- for (user in users) with List<User> → user.save, negative repo.save

Go map range (3 tests):
- for _, user := range userMap with map[string]User → User#Save, negative

TypeScript readonly (4 tests):
- for (const user of users) with readonly User[] → user.save, negative

Bug fix: type-env.ts parameter branch now falls back to childForFieldName('pattern')
for Rust parameters (Rust uses 'pattern' not 'name' for parameter names)
2026-03-16 22:52:11 +00:00
Gergo Magyar caa3310714 feat: Phase 4 — known limitation tests, match arm fix, final verification
- Fix Rust match_arm pattern extraction: unwrap match_pattern to get
  tuple_struct_pattern inside (tree-sitter-rust wraps in match_pattern node)
- Add first-writer-wins regression test for match arm scope leakage
- Add 5 documented skip tests for known limitations:
  - TS destructured for-of (tuple destructuring)
  - Python tuple unpacking in for-loops
  - TS instanceof narrowing (block-level scoping)
  - Rust for with .iter() (method call iterable)
  - Ruby block parameters (closure param inference)

Final: 238 passed, 5 skipped (documented limitations), tsc clean
2026-03-16 22:06:27 +00:00
Gergo Magyar 104f9cd311 feat: Phase 3 complete — all language gaps + pattern matching
Kotlin Tier 1c:
- Unannotated for-loop resolves via shared helper
- extractKotlinElementTypeFromTypeNode handles type_projection unwrapping
- findKotlinParamElementType walks to function_declaration

Java Tier 1c:
- var foreach resolves via shared helper
- extractJavaElementTypeFromTypeNode handles generic_type, array_type
- findJavaParamElementType walks to method_declaration

TypeScript:
- readonly User[] unwrapped via readonly_type → array_type recursion

C# switch patterns:
- declaration_pattern added to patternBindingNodeTypes
- extractPatternBinding handles standalone declaration_pattern (switch case/expr)

Rust match arms:
- match_arm added to patternBindingNodeTypes
- extractPatternBinding extended with match_arm → match_expression parent traversal

Python:
- as_pattern tries childForFieldName('alias') before positional fallback

Tests: 237 pass (was 224), 13 new tests added
2026-03-16 22:03:42 +00:00
Gergo Magyar d526ee927c feat: Phase 3 partial — Rust for-loop + C# var foreach Tier 1c
- Rust: add extractForLoopBinding with for_expression support
  - Handles &users, &mut users via reference_expression unwrapping
  - extractRustElementTypeFromTypeNode: generic_type, reference_type, slice/array
  - findRustParamElementType: AST walk with reference/mut pattern unwrapping
  - 4 unit tests (Vec<User>, &[User], range expr negative, no-annotation negative)

- C#: upgrade foreach to handle var (implicit_type) via Tier 1c
  - extractCSharpElementTypeFromTypeNode: generic_name, array_type, nullable_type
  - findCSharpParamElementType: AST walk to method_declaration parameters
  - 3 unit tests (var foreach, explicit type regression, no-annotation negative)
2026-03-16 21:56:19 +00:00
Gergo Magyar 1a90112495 refactor: Phase 2 architecture — shared helper, required params, decoupled type nodes
- Extract resolveIterableElementType shared helper in shared.ts implementing
  3-strategy fallback (declarationTypeNodes → scopeEnv string → AST walk)
- Refactor TS, Python, Go extractors to use shared helper (eliminates 3x duplication)
- Make ForLoopExtractor params required (aligned with PatternBindingExtractor)
- Update Java, Kotlin, C# extractor signatures to accept required params
- Decouple declarationTypeNodes from scopeEnv — capture raw type annotation
  nodes BEFORE extractDeclaration for container types (User[], []User, List[User])
- Hybrid approach: direct name extraction + keysBefore fallback for multi-declarator
- Document declarationTypeNodes invariant change (superset of scopeEnv)
2026-03-16 21:47:46 +00:00
Gergo Magyar 96cbe8cdae fix: Phase 1 bug fixes — Go range semantics, typed_parameter, bracket depth
- Go single-var range correctly returns early for slices/maps (index, not element)
- Go single-var range on channels correctly resolves element type
- Added map_type and channel_type to extractGoElementTypeFromTypeNode
- Added isChannelType helper for channel detection before skip decision
- Added 'typed_parameter' to TYPED_PARAMETER_TYPES for Python annotated params
- Fixed bracket depth tracking in extractElementTypeFromString — only match
  selected closeChar at depth 0, return undefined for mismatched brackets
- Un-skipped 3 prematurely skipped tests (TS local const, Python List/Sequence)
- Added tests for map range, single-var range semantics, bracket edge cases
2026-03-16 21:42:35 +00:00
Gergo Magyar 00c90d476f reorganise 2026-03-16 21:00:16 +00:00
Gergo Magyar 0c98923729 fix: address code review findings for Phase 6
- Add missing patternBindingNodeTypes to C# typeConfig (perf gate)
- Add 2048-char input length guard to extractElementTypeFromString
- Skip Python match/case integration tests (call extraction needs query updates)
2026-03-16 20:55:47 +00:00
Gergo Magyar 186bd5cf34 feat: Phase 6 type resolution — pattern matching, for-loop Tier 1c, coverage completion
- Add patternBindingNodeTypes gate to LanguageTypeConfig for 50% perf improvement
- Expand ForLoopExtractor signature with optional declarationTypeNodes + scope
- Add extractElementTypeFromString shared utility for container type parsing
- Python match/case: extractPatternBinding for `case User() as u:` pattern
- C# refactor: move is_pattern_expression from extractDeclaration to extractPatternBinding
- Ruby: add extractPendingAssignment for assignment chain propagation
- TS/JS: add for-loop Tier 1c for `for (const user of users)` with User[] inference
- Python: add for-loop Tier 1c for `for user in users:` with type annotation inference
- Go: add for-loop Tier 1c for `for _, user := range users` with []User inference
- Fix 'Property' as any stale cast in call-processor.ts
- Add dual return-type string length cap (2048 pre-cap, 512 post-cap)
- Add chain call integration tests for C#, Go, Rust, Python, JS, C++
- Add Python match/case integration test fixtures
- 27 new extractElementTypeFromString unit tests
- 3 for-loop edge cases skipped (declarationTypeNodes scope key lookup)
2026-03-16 20:42:38 +00:00
Gergő Magyar f2d3df48f6 feat: Phase 5 type resolution — chained calls, pattern matching, class-as-receiver (#315)
* feat: Phase 5 type resolution — chained calls, pattern matching, class-as-receiver, code review fixes

Phase 5.1: Chained method call resolution (depth-capped at 3)
- resolveChainedReceiver() resolves a.getUser().save() by walking the chain
  and looking up intermediate return types from the SymbolTable
- extractReceiverNode() + extractCallChain() shared in utils.ts
- receiverCallChain on ExtractedCall for worker path parity
- MAX_CHAIN_DEPTH=3 enforced in both extraction and resolution

Phase 5.2: Pattern matching binding extractors
- PatternBindingExtractor type added to LanguageTypeConfig
- declarationTypeNodes map tracks original type AST nodes for generic unwrapping
- Rust: if let Some(x)/Ok(x) unwrapping with extractGenericTypeArgs
- Java: instanceof pattern variables (Java 16+)
- C#: is-pattern disambiguation fixture (already working via extractDeclaration)

Phase 5.5d: Python standalone type annotations (name: str)
- expression_statement with type child now captured in DECLARATION_NODE_TYPES

Phase 5.5e: ReceiverKey collision fix for overloaded methods
- receiverKey preserves @startIndex to prevent same-name method collisions
- lookupReceiverType does prefix scan with ambiguity refusal

Class-as-receiver for static method calls (#289)
- UserService.find_user() now resolves via ctx.resolve() tiered lookup
- Respects import scoping — no false positives from unrelated packages

Code review fixes:
- Extracted CALL_EXPRESSION_TYPES + extractCallChain to utils.ts (eliminated duplication)
- Converted resolveChainedReceiver from recursion to loop (no exposed depth param)
- Added depth cap to extractReturnTypeName (defense against nested wrapper types)
- Replaced lookupFuzzy with ctx.resolve for class-as-receiver (architecturally consistent)

Closes #289

Test coverage: 6 new fixtures, 12+ new unit tests, 7 new integration test suites

* fix: Ruby chain calls, Rust Err(x) unwrap, Enum class-as-receiver (#315)

Address three per-language gaps identified in Phase 5 code review:

- Ruby: add `method`/`receiver` field fallbacks to extractCallChain
  (tree-sitter-ruby uses different field names than other grammars)
- Rust: handle `Err(e)` pattern binding via typeArgs[1] from Result<T,E>
- Enum: include Enum type in class-as-receiver filter (both paths)

Integration tests added for all three fixes.

* fix: chain base type resolution parity between serial and worker paths (#315)

- Worker path: add typeEnv.lookup for chain base receiver after extraction
  (typed parameters like `fn process(svc: &UserService)` were silently lost)
- Serial path: add ctx.resolve class-as-receiver fallback for chain base
  (class-name chains like `UserService.find_user().save()` failed)
- Fix misleading comment in parse-worker.ts that described unimplemented logic
- Integration tests: typed-parameter chain, static class-name chain

* fix: Kotlin chain call extraction, createClassNameLookup Enum/Struct (#315)

- Kotlin: extractCallChain now handles navigation_expression → navigation_suffix
  AST structure (Kotlin's call_expression has no 'function' field)
- createClassNameLookup: include Enum and Struct alongside Class for consistent
  constructor recognition in extractInitializer
- Integration test: kotlin-chain-call fixture verifying svc.getUser().save()
2026-03-16 19:38:09 +00:00
Gergő Magyar 5fa73bafdf feat: Phase 4 type resolution — nullable unwrapping, for-loop typing, assignment chains, code review fixes (#310)
* feat: Phase 4 type resolution — nullable unwrapping, for-loop typing, assignment chains, Kotlin return types

Phase 4.1: Nullable/optional chain unwrapping
- Add stripNullable utility in shared.ts for stripping nullable wrappers
- Apply in lookupInEnv to unwrap User | null → User, User? → User before receiver lookup
- Handles TS union, Kotlin/C#/Swift nullable suffix, Python Union[T, None], Rust Option<T>
- Enables receiver-type disambiguation through ?. optional chaining

Phase 4.2: For-loop element typing (Tier 0 — Java/C#/Kotlin)
- Add ForLoopExtractor type and forLoopNodeTypes to LanguageTypeConfig
- Java enhanced_for_statement, C# foreach_statement, Kotlin for_statement extractors
- Only explicit element types in AST (Tier 0); inference-based languages deferred

Phase 4.3: Assignment chain propagation (single-pass, depth-1)
- Add PendingAssignmentExtractor to LanguageTypeConfig with per-language implementations
- Handles TS/JS variable_declarator, Rust let_declaration, Python assignment,
  Go short_var_declaration, C# equals_value_clause, Java/Kotlin variable_declarator
- Single post-walk propagation pass (no fixpoint iteration per Sorbet/Pyright design)
- Resolves const b = a; b.save() when a has known type from Tier 0/1/1b

Phase 4.5: Kotlin return type extraction (bug fix)
- Fix extractMethodSignature to handle Kotlin user_type after function_value_parameters
- Remove lenient test assertions, add strict disambiguation proof

Integration tests across 10+ languages with competing same-name methods
and negative assertions proving disambiguation.

* fix: per-language assignment chain gaps from code review

- Kotlin: new extractKotlinPendingAssignment for property_declaration →
  variable_declaration AST (Java's variable_declarator doesn't exist in Kotlin)
- Go: handle var_spec (var b = u) alongside short_var_declaration (:=)
- PHP: add extractPendingAssignment for $alias = $user with $ prefix preserved

Integration tests added for all three languages with competing
same-name methods and negative disambiguation assertions.

* fix: code review fixes — DRY nullable keywords, avoid array allocations, clarify depth comment

Addresses findings from 6-agent code review on PR #310:

- Move stripNullable JSDoc to correct position (was orphaned above NULLABLE_KEYWORDS)
- DRY: reuse NULLABLE_KEYWORDS set in pipe-split filter instead of inline strings
- Replace node.children.find() with findChildByType/manual loops in jvm.ts,
  go.ts, csharp.ts to avoid unnecessary array allocations per tree-sitter call
- Clarify "depth-1" comment in type-env.ts: single-pass resolves multi-hop
  chains when forward-declared; reverse-order is depth-1 only
- Annotate extractGenericTypeArgs as Phase 5 infrastructure (zero production callers)
- Re-export PendingAssignmentExtractor from index.ts for API consistency
- Add explicit return undefined in Go extractPendingAssignment
- Remove redundant child.text === '=' check in Kotlin extractor

Test coverage:
- 20 new unit tests: stripNullable edge cases, per-language assignment chains,
  reverse-order depth limitation, nullable lookup resolution
- 15 new integration tests: multi-hop chains (a→b→c), nullable+chain combined
  (User|null + alias), Python User|None through stripNullable path
- 3 new fixtures: ts-multi-hop-chain, ts-nullable-chain, python-nullable-chain

* fix: third-pass review — walrus chain, scanner allocations, Kotlin variable_declaration, C# type guard

Addresses 4 new findings from third-pass CI review:

1. Python walrus operator (:=) now handled by extractPendingAssignment —
   named_expression nodes propagate alias chains alongside regular assignment
2. Scanner .namedChildren.find()/.some() in jvm.ts replaced with
   findChildByType() — consistent with 98daed4 code review fixes
3. Kotlin extractPendingAssignment extended to handle variable_declaration
   nodes in addition to property_declaration (function-local val/var)
4. C# extractPendingAssignment early-returns for is_pattern_expression and
   field_declaration nodes (never contain variable_declarator children)

Integration tests:
- Python: walrus chain (alias := u) with disambiguation (5 tests, 1 fixture)
- Kotlin: assignment chain with typed declarations (5 tests, 1 fixture)
- C#: assignment chain + is-pattern coexistence (6 tests, 1 fixture)
- Unit: Python walrus propagation (1 test)

* feat: nullable wrapper unwrapping + C++ assignment chains

Gaps 1, 2, 4 from code review — architectural changes to type resolution:

1. extractSimpleTypeName now unwraps nullable wrapper generics:
   - Optional<User> → "User" (Java), Option<User> → "User" (Rust),
     Maybe<User> → "User" (Kotlin Arrow/Haskell-style)
   - Containers (List, Map) and async wrappers (Promise, Future) are NOT
     unwrapped — methods are called on the container, not the inner type
   - Uses existing extractGenericTypeArgs (now production-active, was dead code)
   - NULLABLE_WRAPPER_TYPES set: Optional, Option, Maybe

2. C++ extractPendingAssignment added for auto alias chains:
   - auto alias = user; alias.save() now propagates User type
   - Handles pointer/reference declarators, auto/decltype(auto)

3. Updated existing Rust test: Option<User> parameter now correctly
   stores "User" instead of "Option" in TypeEnv

Integration tests with fixtures for Java Optional, Rust Option, C++ auto
chain. Full pipeline resolution marked .todo — requires call-processor
enhancement (TypeEnv stores correct types but call-processor needs
additional work to produce CALLS edges for these patterns).

Unit tests: 196 passed (7 new). Integration: all 9 languages green.

* fix: resolve .todo tests — stale dist/ was the root cause

The Rust Option<User> and C++ auto assignment chain integration tests
were marked .todo because the pipeline didn't produce CALLS edges.
Root cause: dist/ was compiled from pre-Phase 4 source and lacked:
- NULLABLE_WRAPPER_TYPES unwrapping in extractSimpleTypeName
- C++ extractPendingAssignment

After npm run build, all tests pass as real assertions:
- Rust: alias.save() resolves to User#save via Option<User> unwrap + chain
- C++: alias.save() and rAlias.save() resolve via auto assignment chain
  with correct disambiguation (User vs Repo)

Only remaining .todo: Rust user.unwrap().save() (Phase 5 — chained
return type inference, not a TypeEnv issue).
2026-03-16 15:21:54 +00:00
fbff6d08c0 feat(ingestion): respect .gitignore and .gitnexusignore during file discovery (#231)
* feat(ingestion): respect .gitignore and .gitnexusignore during file discovery

Add support for excluding files from indexing based on .gitignore and
.gitnexusignore patterns. Previously, GitNexus used only a hardcoded
ignore list, causing significant index pollution in repositories with
git-ignored directories containing code (e.g., Docker-mounted volumes).

Changes:
- Add `ignore` package for gitignore-spec pattern matching
- Add `loadIgnoreRules()` to parse .gitignore + .gitnexusignore
- Add `createIgnoreFilter()` returning glob-compatible IgnoreLike object
- Integrate filter into glob's `ignore` option for directory-level pruning
- Remove post-glob `.filter()` call (now handled during traversal)

The hardcoded DEFAULT_IGNORE_LIST remains as fallback for non-git repos.

Closes #228

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ingestion): address review feedback on ignore filtering

- Distinguish ENOENT vs EACCES in loadIgnoreRules (warn on permission errors)
- Add GITNEXUS_NO_GITIGNORE env var to bypass .gitignore parsing
- Fix bare-name pattern matching in childrenIgnored (check both with/without trailing slash)
- Rename isIgnoredDirectory to isHardcodedIgnoredDirectory for clarity
- Add clarifying comments for design decisions (D2 negation, D3 dot:false redundancy)
- Add tests for bare-name patterns, file-glob patterns, EACCES handling, env var

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ingestion): address second round of review feedback

- G1: Document GITNEXUS_NO_GITIGNORE in `analyze --help` and log when active
- G2: Add comment clarifying path-scurry POSIX normalization contract
- G3: Add IgnoreOptions interface — env var now falls back, callers can
  pass `noGitignore` explicitly for testability and future CLI flag
- G4: Add integration test verifying walkRepositoryPaths respects the env var

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(ingestion): gracefully skip files with unavailable tree-sitter grammars

Port unsupported language resilience from PR #301 by @jecanore.
- Make Kotlin import optional (like Swift) in parser-loader and parse-worker
- Add worker-local isLanguageAvailable() with filePath param for tsx distinction
- Track and log skipped files per language in both sequential and worker paths
- Add skippedLanguages to ParseWorkerResult for worker→main aggregation
- Add isLanguageAvailable unit tests

Refs: #301, #155, #228

Co-Authored-By: jecanore <juan@housingbase.io>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test(e2e): add ignore + language-skip end-to-end test with fixture repo

Add a fixture repo (test/fixtures/ignore-and-skip-repo/) with .gitignore,
.gitnexusignore, TypeScript source files, and a Swift file to exercise
all three features end-to-end:

- File discovery: verifies .gitignore excludes data/ and *.log,
  .gitnexusignore excludes vendor/, source files are discovered
- Parsing: verifies TypeScript files produce Function nodes and DEFINES
  relationships, Swift files are skipped gracefully when grammar is
  unavailable

Add the test to the standalone group in ci-integration.yml and coverage job.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ci): move ignore-and-skip-e2e test to e2e group per review feedback

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(test): use temp directory instead of fixture for e2e ignore test

The fixture's .gitignore prevented data/seed.json and debug.log from
being committed — these files would be missing after checkout in CI.

Switch to creating the entire test structure in a temp directory via
beforeAll (matching filesystem-walker.test.ts pattern). This ensures
all files exist regardless of git ignore rules.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(test): correct graph API usage in e2e ignore test

Use graph.nodes property getter instead of graph.getNodes(), and check
Function node filePath instead of non-existent File nodes (File nodes
are created by processStructure, not processParsing).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* ci: add workflows permission to ci-integration.yml

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* ci: change workflows permission to write per review

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* ci: move workflows permission from ci-integration.yml to ci.yml caller

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ci): fix Claude workflows for fork PRs, remove misplaced workflows perm

Three issues prevented Claude from running on fork PRs:

1. claude-code-review.yml lacked workflows:write — push failed when
   fork PRs modify .github/workflows/ files
2. claude.yml had no fork PR support — checked out main and couldn't
   fetch the fork's branch from origin
3. Cleanup step unconditionally deleted branches even when push failed,
   breaking the concurrent claude.yml workflow

Also removes workflows:write from ci.yml's integration job — CI tests
don't need that permission. The permission belongs on the claude
workflows that push fork branches.

Changes:
- Add workflows:write to both claude workflow permissions blocks
- Add fork PR detection + branch push/cleanup to claude.yml
- Add step id to push-fork; cleanup only runs if push succeeded
- Pass branch names via env vars to prevent shell injection (security)
- Add concurrency groups to prevent race conditions between workflows
- Remove misplaced workflows:write from ci.yml integration job

* fix(ci): use GitHub API for fork branch refs instead of git push

GITHUB_TOKEN cannot have 'workflows' permission — it's only valid for
PATs and GitHub Apps. This means git push fails whenever a fork PR
modifies .github/workflows/ files.

Replace git push with the GitHub REST API (POST/PATCH /git/refs) to
create temporary branch refs. The API creates a pointer to the
already-existing PR head commit without triggering the workflow file
push protection. Similarly, cleanup uses DELETE /git/refs instead of
git push --delete.

Also removes the invalid 'workflows: write' from permissions blocks.

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: jecanore <juan@housingbase.io>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-03-16 13:26:20 +00:00
Gergő Magyar 6c18ae08f7 feat: return type inference, doc-comment parsing, and per-language type extractors (#284)
* feat: Phase 3 — return type inference, generic args extraction, Ruby YARD type extractor

Three architectural improvements to the type resolution system:

1. Return type inference — wire extractMethodSignature returnType through
   SymbolDefinition into call-processor. When var = callee() and callee
   has a known return type, bind var to that type. Handles Promise<T>
   unwrapping, nullable stripping, pointer/reference removal.

2. Generic type argument extraction — new extractGenericTypeArgs() utility
   that extracts type parameters from List<User> → ['User']. Handles
   TS/Java/Kotlin/C#/Rust generic syntax. Building block for for-loop
   variable typing.

3. Ruby dedicated type extractor — replaces the stub with YARD annotation
   parsing (@param name [Type]), handling qualified types, nullable types,
   and singleton methods. Ruby now has real type resolution.

Unit tests: 127 → 192+ (type-env) + 65 (symbol-table, call-processor) + 18 (generics)
Integration tests: 8+ new test cases with fixtures across TS/Python/Go/Java/Ruby

* fix: Phase 3 gaps — WRAPPER_GENERICS correctness, Ruby :: qualifier, namespaced constructors

- Remove collection types (List, Array, Vec, Set) from WRAPPER_GENERICS to prevent
  false CALLS edges (e.g. List<User> no longer unwraps to User)
- Add :: qualifier handling in extractReturnTypeName for Ruby/C++/Rust namespaced types
- Add Ruby `constant` and `scope_resolution` node types to shared extractors
- Extract shared extractRubyConstructorAssignment helper (dedup type-env.ts + ruby.ts)
- Add integration tests for return type inference: Python, TypeScript, Go, Java, Ruby
- Add Ruby namespaced constructor fixture (Models::UserService.new)
- Add unit tests for collection reclassification and :: qualifiers

* feat: Phase 4 — CONSTRUCTOR_BINDING_SCANNERS for all languages + return type inference tests

Add CONSTRUCTOR_BINDING_SCANNERS for 6 missing languages, completing
return type inference coverage across all 11 supported languages:

- TypeScript/JS: variable_declarator with call_expression, unwraps await
- Go: short_var_declaration single-assignment (skips multi-return, new/make)
- Java: local_variable_declaration with `var` type + method_invocation
- C#: variable_declaration with implicit_type (var) + invocation_expression
- Rust: let_declaration without type annotation, handles mut_pattern
- PHP: assignment_expression with function_call_expression

Also adds property_identifier to extractSimpleTypeName for qualified
member calls (repo.getUser → getUser), fixing namespaced constructor
inference that was previously a known limitation.

Integration tests added for all 11 languages with correct label
assertions (Function vs Method per language's tree-sitter queries).

* refactor: merge CONSTRUCTOR_BINDING_SCANNERS into per-language LanguageTypeConfig

Eliminates the parallel dispatch map in type-env.ts by moving all 11
constructor binding scanners into their respective type-extractors/*.ts
files as `scanConstructorBinding` on LanguageTypeConfig.

- Add ConstructorBindingScanner type to types.ts
- Add shared helpers: hasTypeAnnotation, unwrapAwait, extractCalleeName
- Move scanners to typescript.ts, jvm.ts, python.ts, php.ts, go.ts,
  rust.ts, swift.ts, c-cpp.ts, csharp.ts, ruby.ts
- Fix `any` types in C# scanner → SyntaxNode | null
- Delete ~300 lines from type-env.ts (CONSTRUCTOR_BINDING_SCANNERS map)
- Update buildTypeEnv to use config.scanConstructorBinding

All 143 type-env unit tests and all 10 language integration suites pass.

* fix: remove unused import, fix any type in Java scanner, update stale comment

- Remove unused extractCalleeName import from jvm.ts
- Fix (c: any) → (c: SyntaxNode) in Java scanner
- Update stale CONSTRUCTOR_BINDING_SCANNERS reference in ruby.ts comment

* fix: C# and PHP return type inference — scanner fixes, method signature extraction, and cross-file resolution

Addresses code review findings on PR #284:

C# scanner (csharp.ts):
- Fix type node lookup: iterate children instead of childForFieldName('type')
  which returns undefined in tree-sitter-c-sharp
- Fix initializer lookup: handle direct invocation_expression children
  (no equals_value_clause wrapper in tree-sitter-c-sharp)

C# return type extraction (utils.ts):
- Add 'returns' field check to extractMethodSignature — tree-sitter-c-sharp
  uses 'returns', not 'type', for method return types

C# cross-file resolution (call-processor.ts + fixture):
- Add constructor binding verification to sequential processCalls path
  (was only in the worker processCallsFromExtracted path)
- Add ReturnType.csproj to csharp-return-type fixture
- Update fixture namespaces to use ReturnType.Models/ReturnType.Services
  prefix (matches real C# project conventions)

PHP scanner (php.ts):
- Extend scanConstructorBinding to handle member_call_expression
  ($this->getUser() patterns), not just function_call_expression

Shared (shared.ts):
- Add member_access_expression to extractSimpleTypeName qualified-names
  block (C# method calls like svc.GetUser())

Tests:
- Add Repo.cs/Repo.php disambiguation fixtures (two Save methods)
- Strengthen C# and PHP return type tests with hard disambiguation assertions
- Add C# scanner unit tests and return type extraction test

* feat: per-language ReturnTypeExtractor + doc-comment @param parsing for PHP, JS, Ruby

Add ReturnTypeExtractor to LanguageTypeConfig interface with implementations
for Ruby (YARD @return), PHP (PHPDoc @return), and JS/TS (JSDoc @returns).
The fallback is wired in both parsing-processor and parse-worker paths,
activating only when extractMethodSignature finds no AST-based return type.

Also add doc-comment @param type extraction for PHP and JS/TS, following
Ruby's existing collectYardParams pattern. This enables parameter.method()
resolution in loosely-typed codebases using PHPDoc @param or JSDoc @param.

Additional fixes from PR #284 code review:
- Go: add selector_expression + field_identifier to extractSimpleTypeName
  (enables package-qualified factory calls like models.NewUser())
- Ruby: broaden scanConstructorBinding to capture plain call assignments
  (user = get_user()) in addition to Class.new patterns
- Ruby: harden return-type fixture with disambiguation (two save methods)

Test coverage: +14 new integration tests across Go, Ruby, PHP, JS/TS

* fix: JSDoc async return type, PHP attribute walkers, and $this receiver disambiguation

Three fixes from fourth-pass code review on PR #284:

1. JSDoc `@returns {Promise<User>}` no longer stripped to `Promise` — extractReturnType
   now uses sanitizeReturnType (preserves generics) instead of normalizeJsDocType
   (which stripped them before extractReturnTypeName could unwrap WRAPPER_GENERICS).

2. PHP 8+ `#[Attribute]` and JS `@decorator` nodes no longer break doc-comment walkers.
   Both extractReturnType and collect*Params functions now skip attribute_list/decorator
   nodes instead of breaking on them as named siblings.

3. PHP `$this->method()` now provides receiverClassName for disambiguation.
   When two classes define the same method, the enclosing class narrows candidates
   via ownerId matching in call-processor, preventing false no-binding results.

* fix: sanitizeReturnType dot corruption, JS test assertions, Ruby constant receiver

- Remove redundant dot-path stripping from sanitizeReturnType that corrupted
  qualified names inside generics (e.g. Promise<models.User> → User>)
- Split JS async fixture into separate files and add negative assertions
  to properly verify disambiguation (mirroring PHP test pattern)
- Accept 'constant' node type in Ruby scanConstructorBinding for factory
  call assignments (SERVICE = build_service())
- Add 'constant' to SIMPLE_RECEIVER_TYPES so extractReceiverName handles
  Ruby constant receivers (SERVICE.process)

* fix: nested generic arg splitting, JS/Ruby test false positives

- Replace naive comma split in extractReturnTypeName with bracket-balanced
  extractFirstGenericArg so nested types like Future<Result<User, Error>>
  unwrap correctly instead of producing malformed "Result<User"
- Add CompletableFuture to WRAPPER_GENERICS for Java async unwrapping
- Split js-jsdoc-return-type fixture models.js into user.js/repo.js and
  add negative assertions to prove disambiguation (not just file match)
- Split ruby-constant-factory-call fixture into separate service files
  and add negative assertions against AdminService resolution

* fix: review findings — receiverClassName parity, Rust wrappers, Go multi-return, Kotlin/Swift qualified calls

P1: Sequential path now includes receiverClassName narrowing for PHP
$this->method() disambiguation (was missing vs worker path).

P2: Added Rc/Arc/Weak/MutexGuard/Cow + 6 more Rust Deref types to
WRAPPER_GENERICS (Box excluded — Java Swing collision). Extended
Kotlin/Swift scanners to handle navigation_expression callees.
Added Go multi-return support (user, err := f()) with blank/_/err/ok
guard + AST-level first-return extraction in extractMethodSignature.

P3: Extracted shared verifyConstructorBindings() eliminating 60 lines
of duplication between sequential and worker paths. Added return-type
inference integration tests for C++, Rust, Swift with competing
methods and negative disambiguation assertions.

* fix: Swift navigation_suffix unwrapping, Rust lifetime skipping, Kotlin disambiguation tests

- Swift scanConstructorBinding: handle tree-sitter wrapping qualified
  identifiers in navigation_suffix nodes
- Add extractFirstTypeArg to skip Rust lifetime parameters ('a, '_)
  when unwrapping wrapper generics like Ref<'_, User>
- Kotlin tests: add Repo class fixture with competing save() methods
  to prove disambiguation; assert no spurious edges on known gap
- Remove tree-sitter-kotlin from optionalDependencies (now regular dep)

* fix: C# null-conditional calls, Ruby YARD bracket-balanced split, PHPDoc alternate order, escapeValue hardening

- Add C# null-conditional call support (user?.Save()): tree-sitter query for
  conditional_access_expression, member_binding_expression in MEMBER_ACCESS_NODE_TYPES,
  receiver extraction via conditional_access_expression parent walk
- Fix Ruby YARD type parsing for nested generics (Hash<Symbol, User>): replace
  naive split(',') with bracket-balanced splitter respecting <> depth
- Add alternate YARD format (@param [Type] name) alongside standard (@param name [Type])
- Add alternate PHPDoc format (@param $name Type) alongside standard (@param Type $name)
- Harden escapeValue in kuzu-adapter.ts: escape \n and \r to prevent Cypher injection
- Integration tests: C# null-conditional fixture (5 tests), Ruby YARD generics fixture (6 tests)
- Unit tests: PHPDoc alternate order (2 tests), C# null-conditional call-form (updated)

* test: add Python static/classmethod integration tests (issue #289)

Verifies that classes using only @staticmethod/@classmethod have HAS_METHOD
edges connecting them to their child methods. This was the root cause of
issue #289 where context() and impact() returned empty for such classes.

Tests cover: HAS_METHOD edge emission, unique static method resolution
(create_user, delete_user), and ambiguous same-named method handling
(find_user on both UserService and AdminService — safely refused).

* fix: lbug batch escapeValue newline hardening, Rust ::default() scanner exclusion

- Apply \n/\r escaping to batch upsert escapeValue in lbug-adapter.ts:429
  (missed instance of the CREATE-path fix from ec4dca4)
- Exclude Rust ::default() from scanConstructorBinding to match
  extractInitializer behavior — avoids wasted cross-file lookups on
  the broadly-implemented Default trait
- Unit tests: 2 new scanner exclusion tests (::default and ::new)
- Integration tests: 6 new Rust ::default() constructor resolution tests
  with disambiguation fixture (User::default vs Repo::default)

* fix: C#/Rust async await unwrap, PHP backslash namespace, fallback escaping

- C# scanConstructorBinding: unwrap await_expression to find invocation_expression
  (var user = await svc.GetUserAsync() now produces constructor binding)
- Rust scanConstructorBinding: unwrap .await postfix via shared unwrapAwait helper
  (let user = get_user().await now produces constructor binding)
- extractReturnTypeName: handle PHP backslash namespace separator (\App\Models\User → User)
- fallbackRelationshipInserts: match batch escapeValue hardening with \n/\r escaping

Tests: 2 unit (type-env), 3 unit (call-processor), 7 integration (csharp+rust), 7 fixtures

* fix: C#/Rust async-binding test false positives — add competing types and negative assertions

C# fixture: add Order.cs with Order.Save(), change OrderService to return
Task<Order> via GetOrderAsync, add negative assertion proving user.Save()
does not resolve to Order#Save.

Rust fixture: split models.rs into user.rs/repo.rs, make process_user and
process_repo async fn, add bidirectional negative assertions proving no
cross-contamination between User#save and Repo#save.

* fix: C# async-binding broken assertion, bare wrapper type leak, JSDoc optional params

- Split Program.cs Main into ProcessUser/ProcessOrder so negative
  assertions use strict toBeUndefined() (matching Rust pattern)
- Guard bare wrapper types (Task, Promise, Option…) in
  extractReturnTypeName — return undefined instead of the wrapper name
- Update JSDOC_PARAM_RE to capture @param {Type} [optionalName] syntax

* fix: update symbol and relationship counts in documentation
2026-03-15 18:49:40 +00:00
Candido Sales GomesandClaude Opus 4.6 5a5850832c refactor: migrate from KuzuDB to LadybugDB v0.15 (#275)
* refactor: migrate from KuzuDB to LadybugDB v0.15

KuzuDB was archived (Apple acquisition, Oct 2025). LadybugDB is the
community fork with full API compatibility.

- Package swap: kuzu → @ladybugdb/core, kuzu-wasm → @ladybugdb/wasm-core
- Rename all internal paths: kuzu → lbug (adapters, schema, storage)
- Storage path: .gitnexus/kuzu → .gitnexus/lbug (with auto-cleanup)
- Add explicit VECTOR extension loading (required in v0.15)
- Update CI workflow, documentation, and all tests
- 1151 unit + 27 integration tests passing

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address code review findings (P1-P3)

P1: Fix WASM adapter to use getAll() API, wire cleanupOldKuzuFiles
into analyze command, add symlink path traversal protection.
P2: Cache VECTOR extension load state, batch augmentation engine
queries (20→4), fix web getCopyQuery for multi-language tables,
fix stale KuzuDB references, correct brainstorm package names.
P3: Complete lbug-wasm.d.ts type declarations, batch semantic
search per-label, update stale BM25 comment.

* chore: remove outdated KuzuDB migration brainstorming document

* fix: load FTS extension in MCP pool adapter on init

The read-only pool adapter never loaded the FTS extension, so all
QUERY_FTS_INDEX calls failed silently. This broke search-pool and
augmentation integration tests, and caused empty results in the
web UI server mode.

* feat: implement shared Database caching and connection reference counting

* feat: enhance KuzuDB migration handling and status reporting

* fix: mock cleanupOldKuzuFiles in local backend callTool tests

* fix: update mock for cleanupOldKuzuFiles and adjust imports in callTool tests

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 15:53:01 +00:00
Gergő Magyar 62242d5f44 feat: TypeEnvironment API with constructor inference, self/this/super resolution (#274)
* feat(type-env): constructor-call type inference for TypeEnv (Phase 1)

Add extractInitializer as a Tier 1 fallback in buildTypeEnv: when a
declaration node has no explicit type annotation, infer the type from
constructor-call patterns (new X(), X::new(), X::default(), $x = new X()).

Languages covered: TypeScript/JS, Java (var), Rust, PHP, C++ (auto).
Python/Kotlin/Swift deferred — need symbol-table access to distinguish
class constructors from function calls.

Adds 20 new unit tests covering constructor inference, annotation
precedence, and known limitations across all supported languages.

* fix(type-env): class-aware constructor resolution, multi-declarator fix

- Add collectClassNames pre-scan: walks AST to build Set<string> of
  class/struct names defined in the file
- C++ extractInitializer uses classNames.has() to verify identifier is
  a known class before inferring (auto x = User() resolves, auto x =
  getUser() does not — no false positives)
- Add InitializerExtractor type that receives classNames parameter
- Fix env.size gating: always call extractInitializer when available,
  so mixed declarators like const a: A = x, b = new B() resolve both
- Add env.has() guard in Java extractInitializer to skip already-bound vars
- Document Rust new/default whitelist rationale
- Pin all test assertions, add mixed multi-declarator test case

* fix(type-env): resolve Self/self/static/parent to actual type names

- Rust: Self::new()/Self::default() resolves to enclosing impl type
- PHP: new self()/static() resolves to enclosing class, parent() to superclass
- Rust: Tier 0 annotation guard prevents overwrite by constructor inference
- Rust: mut_pattern handling in extractVarName for let mut bindings
- TS: fix misleading comment in extractInitializer
- 58 tests passing (3 new Self/self resolution tests)

* perf(type-env): single-pass AST walk with closure-scoped state

Refactors buildTypeEnv to use closures instead of passing mutable state
as parameters. classNames, env, and config are captured by the inner
walk and extractTypeBinding functions — no parameter mutation.

- Eliminates separate collectClassNames pre-scan (O(2n) → O(n))
- config looked up once per file instead of per-node
- 29 fewer lines

* feat(type-env): constructor-inferred type resolution for all languages

Add cross-file constructor type inference to the ingestion pipeline,
enabling receiver-type disambiguation for member calls like
`user.save()` when the variable is assigned from a constructor without
explicit type annotations.

Pipeline changes:
- Add extractInitializer to Python and Swift type extractors
- Add CONSTRUCTOR_BINDING_SCANNERS for Python, Swift, C/C++ in type-env
- Wire constructorBindings through parse-worker → parsing-processor →
  pipeline → processCallsFromExtracted
- Rewrite resolveCallTarget receiver-type filtering (step D) to use
  tiered import resolution (same-file → import-scoped → global) before
  falling back to fuzzy ownerId matching
- Use collectTieredCandidates for constructor binding verification
  instead of raw lookupFuzzy

Bug fixes:
- Fix C++ inline method query: @definition.method was captured on
  field_declaration_list instead of function_definition, causing wrong
  parameterCount for all inline class methods
- Fix parse-worker accumulated/flush results missing constructorBindings

CI changes:
- Add swift.test.ts to ci-integration pipeline group and coverage job
- Update ci-report to fetch base branch (main) coverage for delta
  reporting instead of showing config thresholds
- Add per-suite timing breakdown table (unit/integration/total)
- Add expandable skipped test details section

Tests: 288 passed, 4 skipped (swift — macOS only) across 10 languages
- 36 new constructor-inferred integration tests (4 per language)
- 10 fixture directories with cross-file constructor patterns
- TypeScript, JavaScript, Java, Kotlin, Python, PHP, Rust, Go, C++, Swift

* fix(type-extractors): add type assertion for LanguageTypeConfig

* feat(ruby): constructor-inferred type resolution and self-receiver mapping

Add Ruby User.new constructor binding scanner to type-env, enabling
receiver-type disambiguation for member calls like user.save vs repo.save.
Add self/this → enclosing class resolution in lookupTypeEnv so self.method()
calls resolve to the correct class even when the method name is ambiguous.

* docs: update README with constructor inference and self/this resolution details

* refactor(ingestion): unified ResolutionContext replaces fragmented map passing

Introduce createResolutionContext() as the single resolution API for all
processors. Eliminates duplicated tier-selection logic, fixes heritage
namedImportMap bug, and adds per-file resolution caching.

- NEW resolution-context.ts: closure-factory with resolve(), per-file cache,
  TIER_CONFIDENCE constant, and shared ResolutionTier type
- DELETE symbol-resolver.ts: zero production importers, logic now in
  resolution-context.ts
- call-processor: all functions take ctx instead of 6 separate maps,
  collectTieredCandidates removed (ctx.resolve replaces it),
  D4 redundant re-resolve eliminated
- heritage-processor: takes ctx, resolveHeritageId helper extracts
  repeated 14-line fallback pattern, namedImportMap now included
- import-processor: takes ctx, dead createImportMap/createPackageMap/
  createNamedImportMap factories removed
- pipeline: creates single ctx, wires onProgress to all processors,
  logs cache hit rate in dev mode
- Tier renamed: unique-global → global (honest about returning all candidates)
- Tests migrated: 1178 unit + 84 integration passing

* feat(type-env): self/this/super resolution, TypeEnvironment API, and review fixes

Add cross-language receiver keyword resolution:
- self/this/$this → enclosing class name via AST walk
- super/base/parent → parent class name via heritage AST extraction
  (8 grammar variants: TS/JS, Java, Python, Ruby, C#, PHP, Kotlin, C++, Swift)
- D-phase widening in resolveCallTarget for super→parent method dispatch

Introduce TypeEnvironment API replacing loose TypeEnvResult + lookupTypeEnv:
- buildTypeEnv() returns TypeEnvironment with .lookup() method
- Single-pass AST walk merges constructor binding scan (was separate traversal)
- ClassNameLookup type replaces over-broad ReadonlySet<string> facade
- Memoized class name lookups to avoid redundant SymbolTable scans

Code review fixes (6 agents, 11 findings):
- Replace ctx.resolve(name, '') hack with direct symbols.lookupFuzzy()
- Extract scope key helpers (extractFuncNameFromScope, receiverKey)
- Simplify D-phase from 5 steps to 4 with deduped typeNodeIds
- Remove C from CONSTRUCTOR_BINDING_SCANNERS (YAGNI — C has no constructors)
- Cache Map reuse in ResolutionContext to reduce GC pressure
- Remove unused TieredCandidates import

Integration tests for self/this, parent, and super resolution across all
12 supported languages with per-language fixture directories.

* fix(type-env): generic parent resolution, TS cast inference, C++ brace-init

Fix generic parent class breaking super resolution:
- extractParentClassFromNode now uses extractSimpleTypeName to strip
  generic params (Base<T> → Base) and qualified names (models.Model → Model)
- Affects TS, Java, Python, C# heritage extraction

Fix TypeScript new X() as T / new X()! missed inference:
- Unwrap as_expression and non_null_expression before checking for
  new_expression in extractInitializer

Fix C++ brace-init User{} missed inference:
- Handle compound_literal_expression with type_identifier child
  in extractInitializer

Clean up deprecated lookupTypeEnv:
- Remove standalone lookupTypeEnv export, migrate all callers to
  TypeEnvironment.lookup() method
- Update all 80+ test assertions to use the new API

Integration test fixtures added:
- typescript-cast-constructor-inference (new X() as T, new X()!)
- typescript/java/csharp/kotlin-generic-parent-resolution
- cpp-brace-init-inference (auto x = User{})

* fix(type-extractors): Go &User{}, TS double-cast, Swift .init inference

Fix Go pointer-to-struct literal not inferred:
- Unwrap unary_expression (address-of &) before composite_literal check
- user := &User{} now correctly infers type User

Fix TypeScript double-cast only unwrapping one level:
- Change if to while loop for nested as_expression/non_null_expression
- new User() as unknown as Admin now correctly infers type User

Fix Swift User.init(name:) explicit init call missed:
- Handle navigation_expression callee with .init suffix in extractInitializer

Integration test fixtures:
- go-pointer-constructor-inference (&User{}, &Repo{})
- typescript-double-cast-inference (as unknown as T)

* feat: Rust struct literal, Python qualified ctor, Go new(), Swift .init scanner

- Rust: handle struct_expression in extractInitializer (User { name: "alice" })
- Python: support attribute nodes in extractInitializer (models.User("alice"))
  and the cross-file scanner — extractSimpleTypeName handles qualified names
- Go: handle new(User) built-in in extractGoShortVarDeclaration
- Swift: extend CONSTRUCTOR_BINDING_SCANNERS to handle navigation_expression
  callee for User.init(name:) cross-file resolution

Unit tests: 87 → 96 (Rust struct literal, Go new(), Python qualified ctor,
Python scanner qualified, plus edge cases)
Integration tests: 4 new describe blocks with fixtures

* fix: Rust Self{} resolution, C++ scoped brace-init, PHP promotion params, Ruby constants

- Rust: resolve Self {} struct literal to enclosing impl type (was stored as "Self")
- C++: replace type_identifier guard with extractSimpleTypeName for compound_literal_expression,
  enabling ns::User{} scoped brace-init (closes previously deferred gap)
- PHP: add property_promotion_parameter to TYPED_PARAMETER_TYPES for PHP 8.0+
  constructor property promotion (__construct(private Foo $x))
- Ruby: extend extractRubyConstructorBinding to accept constant left-hand side
  (REPO = Repo.new)

Unit tests: 96 → 101 (+5: Rust Self{} ×2, C++ ns::User{} ×1, PHP promotion ×1,
Ruby constant ×1)
Integration tests: 4 new describe blocks with fixtures

* feat: Phase 1 type resolution gaps — walrus, PHP properties, nullable, Go make/assert

Phase 1 quick wins from the type resolution gap analysis:

1. Python walrus operator := (named_expression) — extractInitializer + scanner
2. PHP 7.4+ typed class properties — property_declaration in extractDeclaration
3. Nullable union unwrapping — User | null → User in extractSimpleTypeName
4. Go make() builtin — slice/map element type extraction
5. Go type assertions — iface.(User) type extraction

Also: PHP primitive_type handling in extractSimpleTypeName (string, int, etc.)

Unit tests: 101 → 114 (+13)
Integration tests: 8 new describe blocks with fixtures

* feat: Phase 2 type resolution gaps — C++ range-for, Rust if-let, C# pattern matching, Python class annotations

Phase 2 medium-effort improvements:

1. C++ range-for with explicit type — for (User& u : vec) binds u: User
2. Rust if-let/while-let captured_pattern — user @ User { .. } binds user: User
3. C# is-pattern matching — if (obj is User user) binds user: User
4. Python class-level annotations — confirmed already working, added tests

Unit tests: 114 → 127 (+13)
Integration tests: 11 new test cases with fixtures
2026-03-14 19:05:49 +00:00
Gergő Magyar 6e38db879e fix(ruby): method-level call resolution, HAS_METHOD edges, and dispatch table (#278)
* fix(ruby): method-level call resolution, HAS_METHOD edges, and dispatch table refactoring

- Replace all `if (language === Ruby)` checks in processors with a
  `callRouters` dispatch table in call-routing.ts (renamed from
  ruby-call-routing.ts to preserve git history)
- Add Ruby `method` and `singleton_method` to FUNCTION_NODE_TYPES so
  findEnclosingFunction produces Method-level CALLS sources
- Add Ruby `class` and `module` to CLASS_CONTAINER_TYPES for HAS_METHOD
  edge generation
- Add bare call capture via tree-sitter query `(body_statement (identifier))`
  for Ruby methods called without parentheses
- Add Ruby member call detection (`call` node with `receiver` field) to
  inferCallForm and extractReceiverName
- Wire resolveRubyImport into resolveLanguageImport
- Add 24 integration tests across 5 suites: heritage/properties, arity
  filtering, member calls, ambiguous disambiguation, local shadow
- Add ruby.test.ts to CI integration workflow

* fix(ruby): resolve 6 Ruby resolution gaps from PR review

- Fix singleton_method label mismatch: @definition.function → @definition.method
  so CALLS edges from `def self.foo` bodies get correct sourceId
- Add ownerId and HAS_METHOD edges to attr_* Property nodes by calling
  findEnclosingClassId in both parse-worker and call-processor property branches
- Distinguish include/extend/prepend heritage: add heritageKind to
  RubyHeritageItem, propagate through heritage pipeline as IMPLEMENTS reason
- Document bare call over-capture limitation in tree-sitter query comment
- Add bare `require` (non-relative) import test coverage
- Add prepend/extend test coverage with distinct Loggable/Cacheable modules

31 Ruby integration tests passing, no regressions in other language resolvers.

* fix(ruby): web package parity — heritage reasons, property HAS_METHOD, singleton_method label

- Web call-processor: use item.heritageKind as IMPLEMENTS reason instead of
  hardcoded 'trait-impl', add :${kind} suffix to edge ID for uniqueness
- Web call-processor: port findEnclosingClassId, add HAS_METHOD edges for
  attr_* Property nodes to match CLI fix
- Web call-processor: singleton_method label 'Function' → 'Method' to match
  CLI tree-sitter query fix
- CLI parse-worker: update stale ExtractedHeritage.kind JSDoc to include
  'include' | 'extend' | 'prepend'

* fix(web): add HAS_METHOD to RelationshipType union

Web package was missing HAS_METHOD in the RelationshipType union,
causing a type mismatch with the HAS_METHOD edges emitted by the
attr_* property fix in call-processor.ts.
2026-03-14 11:34:30 +00:00
Candido Sales Gomes 0999595444 feat(ruby): Add Ruby language support for CLI and web (#111) 2026-03-13 21:46:59 +00:00
Chirag Nighut 649ad80dbb fix(cli): dynamically discover and install agent skills (#270) 2026-03-13 19:09:03 +00:00
696 changed files with 27485 additions and 2601 deletions
+23 -17
View File
@@ -12,29 +12,29 @@ on:
jobs:
# ── Integration test matrix ─────────────────────────────────────────
# Each test-group runs on a SEPARATE runner per OS, giving full process
# isolation for the KuzuDB native C++ addon.
# isolation for the LadybugDB native C++ addon.
# 3 OS x 4 groups = 12 parallel jobs.
#
# Groups:
# kuzu-db — 7 files using withTestKuzuDB / kuzu-adapter (native addon)
# lbug-db — 7 files using withTestLbugDB / lbug-adapter (native addon)
# Each file runs as its own `vitest run` invocation for full
# process isolation. KuzuDB's native N-API addon registers
# process isolation. LadybugDB's native N-API addon registers
# persistent handles that prevent fork workers from exiting
# on Linux, and its C++ destructors segfault during
# process.exit(). Running each file in its own process lets
# the OS reclaim all resources cleanly.
# pipeline — 12 files: ingestion pipeline + csv + 9 resolver tests
# e2e — 2 files: child-process only (spawnSync), no in-process kuzu
# standalone — 4 files: pure logic, no kuzu, no child processes
# e2e — 4 files: child-process only (spawnSync), no in-process lbug
# standalone — 4 files: pure logic, no lbug, no child processes
test-matrix:
name: integration (${{ matrix.os }} / ${{ matrix.test-group }})
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, windows-latest, macos-latest]
test-group: [kuzu-db, pipeline, e2e, standalone]
test-group: [lbug-db, pipeline, e2e, standalone]
include:
- test-group: kuzu-db
- test-group: lbug-db
# Marker — actual files are listed in the run step below
test-glob: ''
- test-group: pipeline
@@ -51,11 +51,14 @@ jobs:
test/integration/resolvers/go.test.ts
test/integration/resolvers/kotlin.test.ts
test/integration/resolvers/php.test.ts
test/integration/resolvers/ruby.test.ts
test/integration/resolvers/swift.test.ts
- test-group: e2e
test-glob: >-
test/integration/cli-e2e.test.ts
test/integration/hooks-e2e.test.ts
test/integration/skills-e2e.test.ts
test/integration/ignore-and-skip-e2e.test.ts
- test-group: standalone
test-glob: >-
test/integration/filesystem-walker.test.ts
@@ -70,18 +73,18 @@ jobs:
with:
build: 'true'
# kuzu-db: run each file in its own vitest process for full isolation.
# KuzuDB's native addon hangs fork workers on Linux — process isolation
# lbug-db: run each file in its own vitest process for full isolation.
# LadybugDB's native addon hangs fork workers on Linux — process isolation
# is the only reliable fix boundary.
- name: Run integration tests — kuzu-db (process-isolated)
if: matrix.test-group == 'kuzu-db'
- name: Run integration tests — lbug-db (process-isolated)
if: matrix.test-group == 'lbug-db'
working-directory: gitnexus
shell: bash
run: |
set -e
files=(
test/integration/kuzu-core-adapter.test.ts
test/integration/kuzu-pool.test.ts
test/integration/lbug-core-adapter.test.ts
test/integration/lbug-pool.test.ts
test/integration/local-backend.test.ts
test/integration/local-backend-calltool.test.ts
test/integration/search-core.test.ts
@@ -99,9 +102,9 @@ jobs:
done
exit $exit_code
# Non-kuzu groups: run all files in a single vitest invocation
# Non-lbug groups: run all files in a single vitest invocation
- name: Run integration tests — ${{ matrix.test-group }}
if: matrix.test-group != 'kuzu-db'
if: matrix.test-group != 'lbug-db'
shell: bash
env:
TEST_GLOB: ${{ matrix.test-glob }}
@@ -109,9 +112,9 @@ jobs:
working-directory: gitnexus
# ── Coverage collection (ubuntu only) ─────────────────────────────────
# Runs non-kuzu integration tests with coverage enabled so the PR report
# Runs non-lbug integration tests with coverage enabled so the PR report
# can merge integration + unit coverage for a combined view.
# kuzu-db tests are excluded because each file must run in its own vitest
# lbug-db tests are excluded because each file must run in its own vitest
# process (native addon isolation) which prevents single-run coverage merge.
coverage:
name: integration (ubuntu / coverage)
@@ -150,6 +153,7 @@ jobs:
test/integration/enrichment.test.ts
test/integration/tree-sitter-languages.test.ts
test/integration/worker-pool.test.ts
test/integration/ignore-and-skip-e2e.test.ts
test/integration/resolvers/typescript.test.ts
test/integration/resolvers/csharp.test.ts
test/integration/resolvers/cpp.test.ts
@@ -159,6 +163,8 @@ jobs:
test/integration/resolvers/go.test.ts
test/integration/resolvers/kotlin.test.ts
test/integration/resolvers/php.test.ts
test/integration/resolvers/ruby.test.ts
test/integration/resolvers/swift.test.ts
- name: Upload integration coverage
if: always()
+138 -48
View File
@@ -119,6 +119,66 @@ jobs:
sparse-checkout: gitnexus/vitest.config.ts
sparse-checkout-cone-mode: false
# ── Fetch base branch coverage for delta reporting ───────────
- name: Fetch base branch coverage
if: steps.meta.outputs.skip != 'true'
id: base-coverage
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
with:
script: |
const fs = require('fs');
const path = require('path');
// Find the latest successful CI run on main
const runs = await github.rest.actions.listWorkflowRuns({
owner: context.repo.owner,
repo: context.repo.repo,
workflow_id: 'ci.yml',
branch: 'main',
status: 'success',
per_page: 1,
});
if (runs.data.workflow_runs.length === 0) {
core.setOutput('found', 'false');
core.info('No successful main branch CI runs found');
return;
}
const mainRunId = runs.data.workflow_runs[0].id;
const artifacts = await github.rest.actions.listWorkflowRunArtifacts({
owner: context.repo.owner,
repo: context.repo.repo,
run_id: mainRunId,
});
const testReports = artifacts.data.artifacts.find(a => a.name === 'test-reports');
if (!testReports) {
core.setOutput('found', 'false');
core.info('No test-reports artifact on main branch');
return;
}
const zip = await github.rest.actions.downloadArtifact({
owner: context.repo.owner,
repo: context.repo.repo,
artifact_id: testReports.id,
archive_format: 'zip',
});
const dest = path.join(process.env.RUNNER_TEMP, 'base-coverage');
fs.mkdirSync(dest, { recursive: true });
fs.writeFileSync(path.join(dest, 'base.zip'), Buffer.from(zip.data));
core.setOutput('found', 'true');
core.setOutput('dir', dest);
- name: Extract base coverage
if: steps.meta.outputs.skip != 'true' && steps.base-coverage.outputs.found == 'true'
shell: bash
run: |
cd "${{ steps.base-coverage.outputs.dir }}"
unzip -o base.zip -d base
# ── Merge coverage from unit + integration ─────────────────────
- name: Setup Node.js
if: steps.meta.outputs.skip != 'true'
@@ -183,13 +243,14 @@ jobs:
UNIT: ${{ steps.meta.outputs.unit }}
INTEG: ${{ steps.meta.outputs.integration }}
HAS_MERGED: ${{ steps.coverage.outputs.has_merged }}
BASE_FOUND: ${{ steps.base-coverage.outputs.found }}
BASE_DIR: ${{ steps.base-coverage.outputs.dir }}
RUN_URL: ${{ github.event.workflow_run.html_url }}
run: |
DIR="$RUNNER_TEMP/artifacts"
MERGED_DIR="$RUNNER_TEMP/merged-coverage"
# ── Helper: read coverage summary into prefixed vars ──
# Uses printf -v for safe variable assignment (no eval).
read_cov() {
local prefix=$1 file=$2
if [ -n "$file" ] && [ -f "$file" ]; then
@@ -224,7 +285,7 @@ jobs:
fi
}
# ── Read all three coverage reports ──
# ── Read all coverage reports ──
UNIT_SUMMARY=$(find "$DIR/test-reports" -name "coverage-summary.json" -type f 2>/dev/null | head -1)
INTEG_SUMMARY=$(find "$DIR/integration-reports" -name "coverage-summary.json" -type f 2>/dev/null | head -1)
MERGED_SUMMARY="$MERGED_DIR/coverage-summary.json"
@@ -235,7 +296,14 @@ jobs:
HAS_INTEG=$?
read_cov "M" "$MERGED_SUMMARY"
# ── Locate test results (unit) ──
# ── Read base branch coverage (main) ──
BASE_SUMMARY=""
if [ "$BASE_FOUND" = "true" ] && [ -n "$BASE_DIR" ]; then
BASE_SUMMARY=$(find "$BASE_DIR/base" -name "coverage-summary.json" -type f 2>/dev/null | head -1)
fi
read_cov "B" "$BASE_SUMMARY"
# ── Locate test results ──
RESULTS_FILE=$(find "$DIR/test-reports" -name "test-results.json" -type f 2>/dev/null | head -1)
INTEG_RESULTS=$(find "$DIR/integration-reports" -name "integration-results.json" -type f 2>/dev/null | head -1)
@@ -267,17 +335,6 @@ jobs:
FAILED=$((U_FAILED + I_FAILED))
SKIPPED=$((U_SKIPPED + I_SKIPPED))
SUITES=$((U_SUITES + I_SUITES))
DURATION=$((U_DURATION + I_DURATION))
# ── Coverage thresholds (read from vitest.config.ts) ──
if [ -f gitnexus/vitest.config.ts ]; then
THRESH_STMTS=$(grep -oP 'statements:\s*\K[0-9]+' gitnexus/vitest.config.ts || echo 0)
THRESH_BRANCH=$(grep -oP 'branches:\s*\K[0-9]+' gitnexus/vitest.config.ts || echo 0)
THRESH_FUNCS=$(grep -oP 'functions:\s*\K[0-9]+' gitnexus/vitest.config.ts || echo 0)
THRESH_LINES=$(grep -oP 'lines:\s*\K[0-9]+' gitnexus/vitest.config.ts || echo 0)
else
THRESH_STMTS=0; THRESH_BRANCH=0; THRESH_FUNCS=0; THRESH_LINES=0
fi
# ── Status helpers ──
status_icon() {
@@ -289,8 +346,22 @@ jobs:
esac
}
cov_delta() {
local pct=$1 base=$2
if [ "$pct" = "N/A" ] || [ "$base" = "N/A" ]; then echo "—"; return; fi
local diff
diff=$(awk "BEGIN { printf \"%.1f\", $pct - $base }")
if [ "$(awk "BEGIN { print ($pct > $base) ? 1 : 0 }")" = "1" ]; then
echo "📈 +${diff}"
elif [ "$(awk "BEGIN { print ($pct < $base) ? 1 : 0 }")" = "1" ]; then
echo "📉 ${diff}"
else
echo "= ${diff}"
fi
}
cov_bar() {
local pct=$1 thresh=$2
local pct=$1 base=$2
if [ "$pct" = "N/A" ]; then echo "—"; return; fi
local filled
filled=$(awk "BEGIN { printf \"%d\", $pct / 5 }")
@@ -300,7 +371,8 @@ jobs:
local bar=""
for ((i=0; i<filled; i++)); do bar+="█"; done
for ((i=0; i<empty; i++)); do bar+="░"; done
if [ "$(awk "BEGIN { print ($pct >= $thresh) ? 1 : 0 }")" = "1" ]; then
# Green if >= base (or base unavailable), red if dropped
if [ "$base" = "N/A" ] || [ "$(awk "BEGIN { print ($pct >= $base) ? 1 : 0 }")" = "1" ]; then
echo "🟢 ${bar}"
else
echo "🔴 ${bar}"
@@ -333,18 +405,50 @@ jobs:
if [ "$TOTAL" -gt 0 ] 2>/dev/null; then
echo "### Test Results"
echo ""
echo "| Suite | Tests | Passed | Failed | Skipped | Duration |"
echo "|-------|-------|--------|--------|---------|----------|"
if [ "$U_TOTAL" -gt 0 ] 2>/dev/null; then
echo "| Unit | ${U_TOTAL} | ${U_PASSED} | ${U_FAILED} | ${U_SKIPPED} | ${U_DURATION}s |"
fi
if [ "$I_TOTAL" -gt 0 ] 2>/dev/null; then
echo "| Integration | ${I_TOTAL} | ${I_PASSED} | ${I_FAILED} | ${I_SKIPPED} | ${I_DURATION}s |"
fi
echo "| **Total** | **${TOTAL}** | **${PASSED}** | **${FAILED}** | **${SKIPPED}** | **$((U_DURATION + I_DURATION))s** |"
echo ""
if [ "$FAILED" = "0" ]; then
echo "✅ **${PASSED}** passed"
echo "✅ All **${PASSED}** tests passed"
else
echo "❌ **${FAILED}** failed / **${PASSED}** passed"
fi
if [ "$SKIPPED" != "0" ]; then
echo " · ${SKIPPED} skipped"
fi
echo " · ${SUITES} suites · ${TOTAL} total"
echo " · ⏱️ ${DURATION}s"
if [ "$I_TOTAL" -gt 0 ] 2>/dev/null; then
echo " · 📊 ${U_TOTAL} unit + ${I_TOTAL} integration"
echo ""
echo "<details>"
echo "<summary>${SKIPPED} test(s) skipped — expand for details</summary>"
echo ""
# Extract skipped test names from integration results
if [ -n "$INTEG_RESULTS" ] && [ "$I_SKIPPED" -gt 0 ] 2>/dev/null; then
echo "**Integration:**"
jq -r '
.testResults[]
| .assertionResults[]?
| select(.status == "pending" or .status == "skipped")
| "- \(.ancestorTitles | join(" > ")) > \(.title)"
' "$INTEG_RESULTS" 2>/dev/null || echo "- _(unable to parse skipped test details)_"
fi
# Extract skipped test names from unit results
if [ -n "$RESULTS_FILE" ] && [ "$U_SKIPPED" -gt 0 ] 2>/dev/null; then
echo ""
echo "**Unit:**"
jq -r '
.testResults[]
| .assertionResults[]?
| select(.status == "pending" or .status == "skipped")
| "- \(.ancestorTitles | join(" > ")) > \(.title)"
' "$RESULTS_FILE" 2>/dev/null || echo "- _(unable to parse skipped test details)_"
fi
echo ""
echo "</details>"
fi
echo ""
fi
@@ -353,15 +457,15 @@ jobs:
cov_table() {
local label=$1 s=$2 b=$3 f=$4 l=$5 sc=$6 bc=$7 fc=$8 lc=$9
shift 9
local ts=$1 tb=$2 tf=$3 tl=$4
local bs=$1 bb=$2 bf=$3 bl=$4
echo "#### ${label}"
echo ""
echo "| Metric | Coverage | Covered | Threshold | Status |"
echo "|--------|----------|---------|-----------|--------|"
echo "| Statements | **${s}%** | ${sc} | ${ts}% | $(cov_bar "$s" "$ts") |"
echo "| Branches | **${b}%** | ${bc} | ${tb}% | $(cov_bar "$b" "$tb") |"
echo "| Functions | **${f}%** | ${fc} | ${tf}% | $(cov_bar "$f" "$tf") |"
echo "| Lines | **${l}%** | ${lc} | ${tl}% | $(cov_bar "$l" "$tl") |"
echo "| Metric | Coverage | Covered | Base | Delta | Status |"
echo "|--------|----------|---------|------|-------|--------|"
echo "| Statements | **${s}%** | ${sc} | ${bs}% | $(cov_delta "$s" "$bs") | $(cov_bar "$s" "$bs") |"
echo "| Branches | **${b}%** | ${bc} | ${bb}% | $(cov_delta "$b" "$bb") | $(cov_bar "$b" "$bb") |"
echo "| Functions | **${f}%** | ${fc} | ${bf}% | $(cov_delta "$f" "$bf") | $(cov_bar "$f" "$bf") |"
echo "| Lines | **${l}%** | ${lc} | ${bl}% | $(cov_delta "$l" "$bl") | $(cov_bar "$l" "$bl") |"
echo ""
}
@@ -371,7 +475,7 @@ jobs:
cov_table "Combined (Unit + Integration)" \
"$M_STMTS" "$M_BRANCH" "$M_FUNCS" "$M_LINES" \
"$M_STMTS_COV" "$M_BRANCH_COV" "$M_FUNCS_COV" "$M_LINES_COV" \
"$THRESH_STMTS" "$THRESH_BRANCH" "$THRESH_FUNCS" "$THRESH_LINES"
"$B_STMTS" "$B_BRANCH" "$B_FUNCS" "$B_LINES"
echo "<details>"
echo "<summary>Coverage breakdown by test suite</summary>"
@@ -380,37 +484,23 @@ jobs:
cov_table "Unit Tests" \
"$U_STMTS" "$U_BRANCH" "$U_FUNCS" "$U_LINES" \
"$U_STMTS_COV" "$U_BRANCH_COV" "$U_FUNCS_COV" "$U_LINES_COV" \
"$THRESH_STMTS" "$THRESH_BRANCH" "$THRESH_FUNCS" "$THRESH_LINES"
"$B_STMTS" "$B_BRANCH" "$B_FUNCS" "$B_LINES"
fi
if [ "$I_STMTS" != "N/A" ]; then
cov_table "Integration Tests" \
"$I_STMTS" "$I_BRANCH" "$I_FUNCS" "$I_LINES" \
"$I_STMTS_COV" "$I_BRANCH_COV" "$I_FUNCS_COV" "$I_LINES_COV" \
"$THRESH_STMTS" "$THRESH_BRANCH" "$THRESH_FUNCS" "$THRESH_LINES"
"$B_STMTS" "$B_BRANCH" "$B_FUNCS" "$B_LINES"
fi
echo "</details>"
echo ""
echo "<details>"
echo "<summary>Coverage thresholds are auto-ratcheted — they only go up</summary>"
echo ""
echo "Vitest \`thresholds.autoUpdate\` bumps the floor whenever local coverage exceeds it."
echo "CI enforces the current thresholds; developers commit the ratcheted values."
echo "</details>"
echo ""
elif [ "$U_STMTS" != "N/A" ]; then
echo "### Code Coverage (Unit only)"
echo ""
cov_table "Unit Tests" \
"$U_STMTS" "$U_BRANCH" "$U_FUNCS" "$U_LINES" \
"$U_STMTS_COV" "$U_BRANCH_COV" "$U_FUNCS_COV" "$U_LINES_COV" \
"$THRESH_STMTS" "$THRESH_BRANCH" "$THRESH_FUNCS" "$THRESH_LINES"
echo "<details>"
echo "<summary>Coverage thresholds are auto-ratcheted — they only go up</summary>"
echo ""
echo "Vitest \`thresholds.autoUpdate\` bumps the floor whenever local coverage exceeds it."
echo "CI enforces the current thresholds; developers commit the ratcheted values."
echo "</details>"
echo ""
"$B_STMTS" "$B_BRANCH" "$B_FUNCS" "$B_LINES"
else
echo "### Code Coverage"
echo ""
+33 -9
View File
@@ -15,6 +15,12 @@ on:
issue_comment:
types: [created]
# Serialize per-PR so concurrent @claude comments don't race on the
# temporary fork branch push/delete.
concurrency:
group: claude-review-${{ github.event.issue.number || github.event.pull_request.number }}
cancel-in-progress: false
jobs:
claude-review:
# Run only when:
@@ -41,7 +47,7 @@ jobs:
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
contents: write # needed to push fork branch to origin
contents: write # needed to create fork branch ref via API
pull-requests: write
issues: read
id-token: write
@@ -76,11 +82,25 @@ jobs:
fetch-depth: 1
# claude-code-action fetches branches by name from origin, which fails
# for fork PRs. Work around by pushing the fork branch to origin so
# the action can find it. Cleaned up in the post step below.
- name: Push fork branch to origin
# for fork PRs. Create a temporary branch ref via the API so the action
# can find it. Using the API (not git push) avoids the GITHUB_TOKEN
# restriction that blocks pushing commits containing workflow file changes.
- name: Create fork branch ref on origin
id: push-fork
if: steps.pr.outputs.is_fork == 'true'
run: git push origin HEAD:refs/heads/${{ steps.pr.outputs.branch }}
env:
FORK_BRANCH: ${{ steps.pr.outputs.branch }}
FORK_SHA: ${{ steps.pr.outputs.sha }}
GH_TOKEN: ${{ github.token }}
run: |
gh api "repos/${{ github.repository }}/git/refs" \
--method POST \
-f ref="refs/heads/$FORK_BRANCH" \
-f sha="$FORK_SHA" \
|| gh api "repos/${{ github.repository }}/git/refs/heads/$FORK_BRANCH" \
--method PATCH \
-f sha="$FORK_SHA" \
-F force=true
- name: Run Claude Code Review
id: claude-review
@@ -91,7 +111,11 @@ jobs:
plugins: 'code-review@claude-code-plugins'
prompt: '/code-review:code-review ${{ github.repository }}/pull/${{ steps.pr.outputs.number }}'
# Clean up the temporary branch we pushed for fork PRs
- name: Delete fork branch from origin
if: always() && steps.pr.outputs.is_fork == 'true'
run: git push origin --delete refs/heads/${{ steps.pr.outputs.branch }} || true
# Clean up the temporary branch ref we created for fork PRs.
# Only delete if the create step actually succeeded.
- name: Delete fork branch ref from origin
if: always() && steps.push-fork.outcome == 'success'
env:
FORK_BRANCH: ${{ steps.pr.outputs.branch }}
GH_TOKEN: ${{ github.token }}
run: gh api "repos/${{ github.repository }}/git/refs/heads/$FORK_BRANCH" --method DELETE || true
+75 -2
View File
@@ -10,6 +10,12 @@ on:
pull_request_review:
types: [submitted]
# Serialize per-PR so concurrent @claude comments don't race on the
# temporary fork branch push/delete.
concurrency:
group: claude-code-${{ github.event.issue.number || github.event.pull_request.number || github.event.issue.id }}
cancel-in-progress: false
jobs:
claude:
if: |
@@ -20,17 +26,75 @@ jobs:
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
contents: read
contents: write # needed to create fork branch ref via API
pull-requests: write
issues: write
id-token: write
actions: read # Required for Claude to read CI results on PRs
actions: read # required for Claude to read CI results on PRs
steps:
# For PR-related triggers, resolve fork context so we can create a
# temporary branch ref (claude-code-action fetches by branch name).
- name: Resolve PR context
id: pr
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
with:
script: |
// Determine if this event is PR-related
let prNumber = null;
if (context.eventName === 'issue_comment' && context.payload.issue.pull_request) {
prNumber = context.payload.issue.number;
} else if (context.eventName === 'pull_request_review_comment') {
prNumber = context.payload.pull_request.number;
} else if (context.eventName === 'pull_request_review') {
prNumber = context.payload.pull_request.number;
}
if (!prNumber) {
core.setOutput('is_pr', 'false');
core.setOutput('is_fork', 'false');
return;
}
const resp = await github.rest.pulls.get({
owner: context.repo.owner,
repo: context.repo.repo,
pull_number: prNumber,
});
const pr = resp.data;
const isFork = pr.head.repo.full_name !== pr.base.repo.full_name;
core.setOutput('is_pr', 'true');
core.setOutput('is_fork', String(isFork));
core.setOutput('branch', pr.head.ref);
core.setOutput('sha', pr.head.sha);
- name: Checkout repository
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
with:
ref: ${{ steps.pr.outputs.is_fork == 'true' && steps.pr.outputs.sha || '' }}
fetch-depth: 1
# claude-code-action fetches branches by name from origin, which fails
# for fork PRs. Create a temporary branch ref via the API so the action
# can find it. Using the API (not git push) avoids the GITHUB_TOKEN
# restriction that blocks pushing commits containing workflow file changes.
- name: Create fork branch ref on origin
id: push-fork
if: steps.pr.outputs.is_fork == 'true'
env:
FORK_BRANCH: ${{ steps.pr.outputs.branch }}
FORK_SHA: ${{ steps.pr.outputs.sha }}
GH_TOKEN: ${{ github.token }}
run: |
gh api "repos/${{ github.repository }}/git/refs" \
--method POST \
-f ref="refs/heads/$FORK_BRANCH" \
-f sha="$FORK_SHA" \
|| gh api "repos/${{ github.repository }}/git/refs/heads/$FORK_BRANCH" \
--method PATCH \
-f sha="$FORK_SHA" \
-F force=true
- name: Run Claude Code
id: claude
uses: anthropics/claude-code-action@9469d113c6afd29550c402740f22d1a97dd1209b # v1
@@ -40,3 +104,12 @@ jobs:
# This is an optional setting that allows Claude to read CI results on PRs
additional_permissions: |
actions: read
# Clean up the temporary branch ref we created for fork PRs.
# Only delete if the create step actually succeeded.
- name: Delete fork branch ref from origin
if: always() && steps.push-fork.outcome == 'success'
env:
FORK_BRANCH: ${{ steps.pr.outputs.branch }}
GH_TOKEN: ${{ github.token }}
run: gh api "repos/${{ github.repository }}/git/refs/heads/$FORK_BRANCH" --method DELETE || true
+6 -1
View File
@@ -62,4 +62,9 @@ docs/plans/
gitnexus/test/fixtures/mini-repo/*.md
gitnexus/test/fixtures/mini-repo/.claude
gitnexus/test/fixtures/mini-repo/.gitignore
gitnexus/test/fixtures/mini-repo/.gitignore
# Ignore csharp generated obj and bin folders
gitnexus/test/fixtures/lang-resolution/**/obj
gitnexus/test/fixtures/lang-resolution/**/bin
GitNexus.sln
+19 -21
View File
@@ -1,7 +1,7 @@
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (1747 symbols, 4569 relationships, 130 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (2071 symbols, 4727 relationships, 154 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
@@ -69,6 +69,24 @@ Before completing any code modification task, verify:
3. `gitnexus_detect_changes()` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## Keeping the Index Fresh
After committing code changes, the GitNexus index becomes stale. Re-run analyze to update it:
```bash
npx gitnexus analyze
```
If the index previously included embeddings, preserve them by adding `--embeddings`:
```bash
npx gitnexus analyze --embeddings
```
To check whether embeddings exist, inspect `.gitnexus/meta.json` — the `stats.embeddings` field shows the count (0 means no embeddings). **Running analyze without `--embeddings` will delete any previously generated embeddings.**
> Claude Code users: A PostToolUse hook handles this automatically after `git commit` and `git merge`.
## CLI
| Task | Read this skill file |
@@ -79,25 +97,5 @@ Before completing any code modification task, verify:
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (135 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Workers area (70 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Cli area (63 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Kuzu area (52 symbols) | `.claude/skills/generated/kuzu/SKILL.md` |
| Work in the Wiki area (52 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Embeddings area (48 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Components area (42 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Local area (36 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Storage area (36 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Services area (35 symbols) | `.claude/skills/generated/services/SKILL.md` |
| Work in the Mcp area (32 symbols) | `.claude/skills/generated/mcp/SKILL.md` |
| Work in the Llm area (30 symbols) | `.claude/skills/generated/llm/SKILL.md` |
| Work in the Eval area (18 symbols) | `.claude/skills/generated/eval/SKILL.md` |
| Work in the Bridge area (15 symbols) | `.claude/skills/generated/bridge/SKILL.md` |
| Work in the Hooks area (14 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Search area (11 symbols) | `.claude/skills/generated/search/SKILL.md` |
| Work in the Environments area (11 symbols) | `.claude/skills/generated/environments/SKILL.md` |
| Work in the Analysis area (10 symbols) | `.claude/skills/generated/analysis/SKILL.md` |
| Work in the Agents area (9 symbols) | `.claude/skills/generated/agents/SKILL.md` |
| Work in the Graph area (6 symbols) | `.claude/skills/generated/graph/SKILL.md` |
<!-- gitnexus:end -->
+8
View File
@@ -2,6 +2,14 @@
All notable changes to GitNexus will be documented in this file.
## [Unreleased]
### Changed
- Migrated from KuzuDB to LadybugDB v0.15 (`@ladybugdb/core`, `@ladybugdb/wasm-core`)
- Renamed all internal paths from `kuzu` to `lbug` (storage: `.gitnexus/kuzu` → `.gitnexus/lbug`)
- Added automatic cleanup of stale KuzuDB index files
- LadybugDB v0.15 requires explicit VECTOR extension loading for semantic search
## [1.4.0] - 2026-03-13
### Added
+19 -21
View File
@@ -1,7 +1,7 @@
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (1747 symbols, 4569 relationships, 130 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (2071 symbols, 4727 relationships, 154 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
@@ -69,6 +69,24 @@ Before completing any code modification task, verify:
3. `gitnexus_detect_changes()` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## Keeping the Index Fresh
After committing code changes, the GitNexus index becomes stale. Re-run analyze to update it:
```bash
npx gitnexus analyze
```
If the index previously included embeddings, preserve them by adding `--embeddings`:
```bash
npx gitnexus analyze --embeddings
```
To check whether embeddings exist, inspect `.gitnexus/meta.json` — the `stats.embeddings` field shows the count (0 means no embeddings). **Running analyze without `--embeddings` will delete any previously generated embeddings.**
> Claude Code users: A PostToolUse hook handles this automatically after `git commit` and `git merge`.
## CLI
| Task | Read this skill file |
@@ -79,25 +97,5 @@ Before completing any code modification task, verify:
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (135 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Workers area (70 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Cli area (63 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Kuzu area (52 symbols) | `.claude/skills/generated/kuzu/SKILL.md` |
| Work in the Wiki area (52 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Embeddings area (48 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Components area (42 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Local area (36 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Storage area (36 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Services area (35 symbols) | `.claude/skills/generated/services/SKILL.md` |
| Work in the Mcp area (32 symbols) | `.claude/skills/generated/mcp/SKILL.md` |
| Work in the Llm area (30 symbols) | `.claude/skills/generated/llm/SKILL.md` |
| Work in the Eval area (18 symbols) | `.claude/skills/generated/eval/SKILL.md` |
| Work in the Bridge area (15 symbols) | `.claude/skills/generated/bridge/SKILL.md` |
| Work in the Hooks area (14 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Search area (11 symbols) | `.claude/skills/generated/search/SKILL.md` |
| Work in the Environments area (11 symbols) | `.claude/skills/generated/environments/SKILL.md` |
| Work in the Analysis area (10 symbols) | `.claude/skills/generated/analysis/SKILL.md` |
| Work in the Agents area (9 symbols) | `.claude/skills/generated/agents/SKILL.md` |
| Work in the Graph area (6 symbols) | `.claude/skills/generated/graph/SKILL.md` |
<!-- gitnexus:end -->
+27 -10
View File
@@ -51,7 +51,7 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
| **For** | Daily development with Cursor, Claude Code, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
| **Install** | `npm install -g gitnexus` | No install —[gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Storage** | KuzuDB native (fast, persistent) | KuzuDB WASM (in-memory, per session) |
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Privacy** | Everything local, no network | Everything in-browser, no server |
@@ -224,8 +224,8 @@ flowchart TD
Server["server.ts"]
Backend["LocalBackend"]
Pool["Connection Pool"]
ConnA["KuzuDB conn A"]
ConnB["KuzuDB conn B"]
ConnA["LadybugDB conn A"]
ConnB["LadybugDB conn B"]
end
Setup -->|"writes global MCP config"| CursorConfig["~/.cursor/mcp.json"]
@@ -242,7 +242,7 @@ flowchart TD
ConnB -->|"queries"| RepoB
```
**How it works:** Each `gitnexus analyze` stores the index in `.gitnexus/` inside the repo (portable, gitignored) and registers a pointer in `~/.gitnexus/registry.json`. When an AI agent starts, the MCP server reads the registry and can serve any indexed repo. KuzuDB connections are opened lazily on first query and evicted after 5 minutes of inactivity (max 5 concurrent). If only one repo is indexed, the `repo` parameter is optional on all tools — agents don't need to change anything.
**How it works:** Each `gitnexus analyze` stores the index in `.gitnexus/` inside the repo (portable, gitignored) and registers a pointer in `~/.gitnexus/registry.json`. When an AI agent starts, the MCP server reads the registry and can serve any indexed repo. LadybugDB connections are opened lazily on first query and evicted after 5 minutes of inactivity (max 5 concurrent). If only one repo is indexed, the `repo` parameter is optional on all tools — agents don't need to change anything.
---
@@ -263,7 +263,7 @@ npm install
npm run dev
```
The web UI uses the same indexing pipeline as the CLI but runs entirely in WebAssembly (Tree-sitter WASM, KuzuDB WASM, in-browser embeddings). It's great for quick exploration but limited by browser memory for larger repos.
The web UI uses the same indexing pipeline as the CLI but runs entirely in WebAssembly (Tree-sitter WASM, LadybugDB WASM, in-browser embeddings). It's great for quick exploration but limited by browser memory for larger repos.
**Local Backend Mode:** Run `gitnexus serve` and open the web UI locally — it auto-detects the server and shows all your indexed repos, with full AI chat support. No need to re-upload or re-index. The agent's tools (Cypher queries, search, code navigation) route through the backend HTTP API automatically.
@@ -320,14 +320,30 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
1. **Structure** — Walks the file tree and maps folder/file relationships
2. **Parsing** — Extracts functions, classes, methods, and interfaces using Tree-sitter ASTs
3. **Resolution** — Resolves imports and function calls across files with language-aware logic
3. **Resolution** — Resolves imports, function calls, heritage, constructor inference, and `self`/`this` receiver types across files with language-aware logic
4. **Clustering** — Groups related symbols into functional communities
5. **Processes** — Traces execution flows from entry points through call chains
6. **Search** — Builds hybrid search indexes for fast retrieval
### Supported Languages
TypeScript, JavaScript, Python, Java, Kotlin, C, C++, C#, Go, Rust, PHP, Swift
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
|----------|---------|----------------|---------|----------|-----------------|---------------------|--------|------------|-------------|
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
---
@@ -466,7 +482,7 @@ The wiki generator reads the indexed graph structure, groups files into modules
| ------------------------- | ------------------------------------- | --------------------------------------- |
| **Runtime** | Node.js (native) | Browser (WASM) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Database** | KuzuDB native | KuzuDB WASM |
| **Database** | LadybugDB native | LadybugDB WASM |
| **Embeddings** | HuggingFace transformers.js (GPU/CPU) | transformers.js (WebGPU/WASM) |
| **Search** | BM25 + semantic + RRF | BM25 + semantic + RRF |
| **Agent Interface** | MCP (stdio) | LangChain ReAct agent |
@@ -487,9 +503,10 @@ The wiki generator reads the indexed graph structure, groups files into modules
### Recently Completed
- [X] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
- [X] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
- [X] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
- [X] Multi-Repo MCP, Zero-Config Setup, 11 Language Support
- [X] Multi-Repo MCP, Zero-Config Setup, 13 Language Support
- [X] Community Detection, Process Detection, Confidence Scoring
- [X] Hybrid Search, Vector Index
@@ -506,7 +523,7 @@ The wiki generator reads the indexed graph structure, groups files into modules
## Acknowledgments
- [Tree-sitter](https://tree-sitter.github.io/) — AST parsing
- [KuzuDB](https://kuzudb.com/) — Embedded graph database with vector support
- [LadybugDB](https://ladybugdb.com/) — Embedded graph database with vector support (formerly KuzuDB)
- [Sigma.js](https://www.sigmajs.org/) — WebGL graph rendering
- [transformers.js](https://huggingface.co/docs/transformers.js) — Browser ML
- [Graphology](https://graphology.github.io/) — Graph data structures
+2 -2
View File
@@ -148,7 +148,7 @@ Each mode has a `system_{mode}.jinja` + `instance_{mode}.jinja` pair. The agent
1. Docker container starts with SWE-bench instance (repo at specific commit)
2. **GitNexus setup**: Node.js + gitnexus installed, `gitnexus analyze` runs (or restores from cache)
3. **Eval-server starts**: `gitnexus eval-server` daemon (persistent HTTP server, keeps KuzuDB warm)
3. **Eval-server starts**: `gitnexus eval-server` daemon (persistent HTTP server, keeps LadybugDB warm)
4. **Standalone tool scripts installed** in `/usr/local/bin/` — works with `subprocess.run` (no `.bashrc` needed)
5. Agent runs with the configured model + system prompt + GitNexus tools
6. Agent's patch is extracted as a git diff
@@ -167,7 +167,7 @@ Each tool script in `/usr/local/bin/` is standalone — no sourcing, no env inhe
### Eval-server
The eval-server is a lightweight HTTP daemon that:
- Keeps KuzuDB warm in memory (no cold start per tool call)
- Keeps LadybugDB warm in memory (no cold start per tool call)
- Returns LLM-friendly text (not raw JSON — saves tokens)
- Includes next-step hints to guide tool chaining (query → context → impact → fix)
- Auto-shuts down after idle timeout
@@ -160,7 +160,7 @@ function handlePreToolUse(input) {
* PostToolUse handler — detect index staleness after git mutations.
*
* Instead of spawning a full `gitnexus analyze` synchronously (which blocks
* the agent for up to 120s and risks KuzuDB corruption on timeout), we do a
* the agent for up to 120s and risks LadybugDB corruption on timeout), we do a
* lightweight staleness check: compare `git rev-parse HEAD` against the
* lastCommit stored in `.gitnexus/meta.json`. If they differ, notify the
* agent so it can decide when to reindex.
+25 -26
View File
@@ -10,6 +10,7 @@
"dependencies": {
"@huggingface/transformers": "^3.0.0",
"@isomorphic-git/lightning-fs": "^4.6.2",
"@ladybugdb/wasm-core": "^0.15.1",
"@langchain/anthropic": "^1.3.10",
"@langchain/core": "^1.1.15",
"@langchain/google-genai": "^2.1.10",
@@ -30,7 +31,6 @@
"graphology-utils": "^2.3.0",
"isomorphic-git": "^1.36.1",
"jszip": "^3.10.1",
"kuzu-wasm": "^0.11.1",
"langchain": "^1.2.10",
"lru-cache": "^11.2.4",
"lucide-react": "^0.562.0",
@@ -1643,6 +1643,30 @@
"@jridgewell/sourcemap-codec": "^1.4.14"
}
},
"node_modules/@ladybugdb/wasm-core": {
"version": "0.15.1",
"resolved": "https://registry.npmjs.org/@ladybugdb/wasm-core/-/wasm-core-0.15.1.tgz",
"integrity": "sha512-dHEq8inJQBkHnJrqZMKGdltSfeSv9OHECkzWQixqDLApXXGlbJ5Ugq5rRfk2PLJuZ74LVHT0cZvcn4JLmsnAIA==",
"license": "MIT",
"dependencies": {
"threads": "^1.7.0",
"tiny-worker": "^2.3.0",
"uuid": "^11.0.3"
}
},
"node_modules/@ladybugdb/wasm-core/node_modules/uuid": {
"version": "11.1.0",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-11.1.0.tgz",
"integrity": "sha512-0/A9rDy9P7cJ+8w1c9WD9V//9Wj15Ce2MPz8Ri6032usz+NfePxx5AcN3bN+r6ZL6jEo066/yNYB3tn4pQEx+A==",
"funding": [
"https://github.com/sponsors/broofa",
"https://github.com/sponsors/ctavan"
],
"license": "MIT",
"bin": {
"uuid": "dist/esm/bin/uuid"
}
},
"node_modules/@langchain/anthropic": {
"version": "1.3.10",
"resolved": "https://registry.npmjs.org/@langchain/anthropic/-/anthropic-1.3.10.tgz",
@@ -6194,31 +6218,6 @@
"resolved": "https://registry.npmjs.org/khroma/-/khroma-2.1.0.tgz",
"integrity": "sha512-Ls993zuzfayK269Svk9hzpeGUKob/sIgZzyHYdjQoAdQetRKpOLj+k/QQQ/6Qi0Yz65mlROrfd+Ev+1+7dz9Kw=="
},
"node_modules/kuzu-wasm": {
"version": "0.11.3",
"resolved": "https://registry.npmjs.org/kuzu-wasm/-/kuzu-wasm-0.11.3.tgz",
"integrity": "sha512-+bLOqXgYZJJ2dHJG1y9LTLyb9ZB73eLxErRZahZz2rPokfIdyLaktTJFzJH7wX39hgyukKn8QxeRNobH6gl27g==",
"deprecated": "Package no longer supported. Contact Support at https://www.npmjs.com/support for more info.",
"license": "MIT",
"dependencies": {
"threads": "^1.7.0",
"tiny-worker": "^2.3.0",
"uuid": "^11.0.3"
}
},
"node_modules/kuzu-wasm/node_modules/uuid": {
"version": "11.1.0",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-11.1.0.tgz",
"integrity": "sha512-0/A9rDy9P7cJ+8w1c9WD9V//9Wj15Ce2MPz8Ri6032usz+NfePxx5AcN3bN+r6ZL6jEo066/yNYB3tn4pQEx+A==",
"funding": [
"https://github.com/sponsors/broofa",
"https://github.com/sponsors/ctavan"
],
"license": "MIT",
"bin": {
"uuid": "dist/esm/bin/uuid"
}
},
"node_modules/langchain": {
"version": "1.2.10",
"resolved": "https://registry.npmjs.org/langchain/-/langchain-1.2.10.tgz",
+1 -1
View File
@@ -33,7 +33,7 @@
"graphology-layout-noverlap": "^0.4.2",
"isomorphic-git": "^1.36.1",
"jszip": "^3.10.1",
"kuzu-wasm": "^0.11.1",
"@ladybugdb/wasm-core": "^0.15.1",
"langchain": "^1.2.10",
"lru-cache": "^11.2.4",
"lucide-react": "^0.562.0",
Binary file not shown.
Binary file not shown.
@@ -5,6 +5,42 @@ import { vscDarkPlus } from 'react-syntax-highlighter/dist/esm/styles/prism';
import { useAppState } from '../hooks/useAppState';
import { NODE_COLORS } from '../lib/constants';
/** Map file extension to Prism syntax highlighter language identifier */
const getSyntaxLanguage = (filePath: string | undefined): string => {
if (!filePath) return 'text';
const ext = filePath.split('.').pop()?.toLowerCase();
switch (ext) {
case 'js': case 'jsx': case 'mjs': case 'cjs': return 'javascript';
case 'ts': case 'tsx': case 'mts': case 'cts': return 'typescript';
case 'py': case 'pyw': return 'python';
case 'rb': case 'rake': case 'gemspec': return 'ruby';
case 'java': return 'java';
case 'go': return 'go';
case 'rs': return 'rust';
case 'c': case 'h': return 'c';
case 'cpp': case 'cc': case 'cxx': case 'hpp': case 'hxx': case 'hh': return 'cpp';
case 'cs': return 'csharp';
case 'php': return 'php';
case 'kt': case 'kts': return 'kotlin';
case 'swift': return 'swift';
case 'json': return 'json';
case 'yaml': case 'yml': return 'yaml';
case 'md': case 'mdx': return 'markdown';
case 'html': case 'htm': case 'erb': return 'markup';
case 'css': case 'scss': case 'sass': return 'css';
case 'sh': case 'bash': case 'zsh': return 'bash';
case 'sql': return 'sql';
case 'xml': return 'xml';
default: break;
}
// Handle extensionless Ruby files
const basename = filePath.split('/').pop() || '';
if (['Rakefile', 'Gemfile', 'Guardfile', 'Vagrantfile', 'Brewfile'].includes(basename)) return 'ruby';
if (['Makefile'].includes(basename)) return 'makefile';
if (['Dockerfile'].includes(basename)) return 'docker';
return 'text';
};
// Match the code theme used elsewhere in the app
const customTheme = {
...vscDarkPlus,
@@ -267,12 +303,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
<div className="flex-1 min-h-0 overflow-auto scrollbar-thin">
{selectedFileContent ? (
<SyntaxHighlighter
language={
selectedFilePath?.endsWith('.py') ? 'python' :
selectedFilePath?.endsWith('.js') || selectedFilePath?.endsWith('.jsx') ? 'javascript' :
selectedFilePath?.endsWith('.ts') || selectedFilePath?.endsWith('.tsx') ? 'typescript' :
'text'
}
language={getSyntaxLanguage(selectedFilePath)}
style={customTheme as any}
showLineNumbers
startingLineNumber={1}
@@ -339,11 +370,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
const hasRange = typeof ref.startLine === 'number';
const startDisplay = hasRange ? (ref.startLine ?? 0) + 1 : undefined;
const endDisplay = hasRange ? (ref.endLine ?? ref.startLine ?? 0) + 1 : undefined;
const language =
ref.filePath.endsWith('.py') ? 'python' :
ref.filePath.endsWith('.js') || ref.filePath.endsWith('.jsx') ? 'javascript' :
ref.filePath.endsWith('.ts') || ref.filePath.endsWith('.tsx') ? 'typescript' :
'text';
const language = getSyntaxLanguage(ref.filePath);
const isGlowing = glowRefId === ref.id;
@@ -83,7 +83,7 @@ export const EmbeddingStatus = () => {
<button
onClick={handleTestArrayParams}
className="flex items-center gap-1 px-2 py-1.5 bg-surface border border-border-subtle rounded-lg text-xs text-text-muted hover:bg-hover hover:text-text-secondary transition-all"
title="Test if KuzuDB supports array params"
title="Test if LadybugDB supports array params"
>
<FlaskConical className="w-3 h-3" />
{testResult || 'Test'}
@@ -9,6 +9,6 @@ export enum SupportedLanguages {
Go = 'go',
Rust = 'rust',
PHP = 'php',
// Ruby = 'ruby',
Ruby = 'ruby',
Swift = 'swift',
}
+1 -1
View File
@@ -275,7 +275,7 @@ export const embedBatch = async (texts: string[]): Promise<Float32Array[]> => {
};
/**
* Convert Float32Array to regular number array (for KuzuDB storage)
* Convert Float32Array to regular number array (for LadybugDB storage)
*/
export const embeddingToArray = (embedding: Float32Array): number[] => {
return Array.from(embedding);
@@ -2,10 +2,10 @@
* Embedding Pipeline Module
*
* Orchestrates the background embedding process:
* 1. Query embeddable nodes from KuzuDB
* 1. Query embeddable nodes from LadybugDB
* 2. Generate text representations
* 3. Batch embed using transformers.js
* 4. Update KuzuDB with embeddings
* 4. Update LadybugDB with embeddings
* 5. Create vector index for semantic search
*/
@@ -27,7 +27,7 @@ import {
export type EmbeddingProgressCallback = (progress: EmbeddingProgress) => void;
/**
* Query all embeddable nodes from KuzuDB
* Query all embeddable nodes from LadybugDB
* Uses table-specific queries (File has different schema than code elements)
*/
const queryEmbeddableNodes = async (
@@ -102,9 +102,23 @@ const batchInsertEmbeddings = async (
* Create the vector index for semantic search
* Now indexes the separate CodeEmbedding table
*/
let vectorExtensionLoaded = false;
const createVectorIndex = async (
executeQuery: (cypher: string) => Promise<any[]>
): Promise<void> => {
// LadybugDB v0.15+ requires explicit VECTOR extension loading (once per session)
if (!vectorExtensionLoaded) {
try {
await executeQuery('INSTALL VECTOR');
await executeQuery('LOAD EXTENSION VECTOR');
vectorExtensionLoaded = true;
} catch {
// Extension may already be loaded — CREATE_VECTOR_INDEX will fail clearly if not
vectorExtensionLoaded = true;
}
}
const cypher = `
CALL CREATE_VECTOR_INDEX('CodeEmbedding', 'code_embedding_idx', 'embedding', metric := 'cosine')
`;
@@ -122,7 +136,7 @@ const createVectorIndex = async (
/**
* Run the embedding pipeline
*
* @param executeQuery - Function to execute Cypher queries against KuzuDB
* @param executeQuery - Function to execute Cypher queries against LadybugDB
* @param executeWithReusedStatement - Function to execute with reused prepared statement
* @param onProgress - Callback for progress updates
* @param config - Optional configuration override
@@ -206,7 +220,7 @@ export const runEmbeddingPipeline = async (
// Embed the batch
const embeddings = await embedBatch(texts);
// Update KuzuDB with embeddings
// Update LadybugDB with embeddings
const updates = batch.map((node, i) => ({
id: node.id,
embedding: embeddingToArray(embeddings[i]),
@@ -313,51 +327,64 @@ export const semanticSearch = async (
return [];
}
// Get metadata for each result by querying each node table
const results: SemanticSearchResult[] = [];
// Group results by label for batched metadata queries
const byLabel = new Map<string, Array<{ nodeId: string; distance: number }>>();
for (const embRow of embResults) {
const nodeId = embRow.nodeId ?? embRow[0];
const distance = embRow.distance ?? embRow[1];
// Extract label from node ID (format: Label:path:name)
const labelEndIdx = nodeId.indexOf(':');
const label = labelEndIdx > 0 ? nodeId.substring(0, labelEndIdx) : 'Unknown';
// Query the specific table for this node
// File nodes don't have startLine/endLine
if (!byLabel.has(label)) byLabel.set(label, []);
byLabel.get(label)!.push({ nodeId, distance });
}
// Batch-fetch metadata per label
const results: SemanticSearchResult[] = [];
for (const [label, items] of byLabel) {
const idList = items.map(i => `'${i.nodeId.replace(/'/g, "''")}'`).join(', ');
try {
let nodeQuery: string;
if (label === 'File') {
nodeQuery = `
MATCH (n:File {id: '${nodeId.replace(/'/g, "''")}'})
RETURN n.name AS name, n.filePath AS filePath
MATCH (n:File) WHERE n.id IN [${idList}]
RETURN n.id AS id, n.name AS name, n.filePath AS filePath
`;
} else {
nodeQuery = `
MATCH (n:${label} {id: '${nodeId.replace(/'/g, "''")}'})
RETURN n.name AS name, n.filePath AS filePath,
MATCH (n:${label}) WHERE n.id IN [${idList}]
RETURN n.id AS id, n.name AS name, n.filePath AS filePath,
n.startLine AS startLine, n.endLine AS endLine
`;
}
const nodeRows = await executeQuery(nodeQuery);
if (nodeRows.length > 0) {
const nodeRow = nodeRows[0];
results.push({
nodeId,
name: nodeRow.name ?? nodeRow[0] ?? '',
label,
filePath: nodeRow.filePath ?? nodeRow[1] ?? '',
distance,
startLine: label !== 'File' ? (nodeRow.startLine ?? nodeRow[2]) : undefined,
endLine: label !== 'File' ? (nodeRow.endLine ?? nodeRow[3]) : undefined,
});
const rowMap = new Map<string, any>();
for (const row of nodeRows) {
const id = row.id ?? row[0];
rowMap.set(id, row);
}
for (const item of items) {
const nodeRow = rowMap.get(item.nodeId);
if (nodeRow) {
results.push({
nodeId: item.nodeId,
name: nodeRow.name ?? nodeRow[1] ?? '',
label,
filePath: nodeRow.filePath ?? nodeRow[2] ?? '',
distance: item.distance,
startLine: label !== 'File' ? (nodeRow.startLine ?? nodeRow[3]) : undefined,
endLine: label !== 'File' ? (nodeRow.endLine ?? nodeRow[4]) : undefined,
});
}
}
} catch {
// Table might not exist, skip
}
}
// Re-sort by distance since batch queries may have mixed order
results.sort((a, b) => a.distance - b.distance);
return results;
};
+1 -1
View File
@@ -92,7 +92,7 @@ export interface SemanticSearchResult {
}
/**
* Node data for embedding (minimal structure from KuzuDB query)
* Node data for embedding (minimal structure from LadybugDB query)
*/
export interface EmbeddableNode {
id: string;
+1
View File
@@ -54,6 +54,7 @@ export type RelationshipType =
| 'DECORATES'
| 'IMPLEMENTS'
| 'EXTENDS'
| 'HAS_METHOD'
| 'MEMBER_OF'
| 'STEP_IN_PROCESS'
+190 -34
View File
@@ -6,6 +6,7 @@ import { loadParser, loadLanguage } from '../tree-sitter/parser-loader';
import { LANGUAGE_QUERIES } from './tree-sitter-queries';
import { generateId } from '../../lib/utils';
import { getLanguageFromFilename } from './utils';
import { callRouters } from './call-routing';
/**
* Node types that represent function/method definitions across languages.
@@ -35,6 +36,9 @@ const FUNCTION_NODE_TYPES = new Set([
// Rust
'function_item',
'impl_item', // Methods inside impl blocks
// Ruby
'method', // def foo
'singleton_method', // def self.foo
]);
/**
@@ -92,6 +96,18 @@ const findEnclosingFunction = (
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method'; // Treat constructors as methods for process detection
} else if (current.type === 'method') {
// Ruby instance method: def foo
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'singleton_method') {
// Ruby class method: def self.foo
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'arrow_function' || current.type === 'function_expression') {
// Arrow/expression: const foo = () => {} - check parent variable declarator
const parent = current.parent;
@@ -126,6 +142,47 @@ const findEnclosingFunction = (
return null; // Top-level call (not inside any function)
};
/** AST node types that represent a class-like container */
const CLASS_CONTAINER_TYPES = new Set([
'class_declaration', 'abstract_class_declaration',
'interface_declaration', 'struct_declaration', 'record_declaration',
'class_specifier', 'struct_specifier',
'impl_item', 'trait_item',
'class_definition',
'trait_declaration',
'protocol_declaration',
'class', 'module', // Ruby
]);
const CONTAINER_TYPE_TO_LABEL: Record<string, string> = {
class_declaration: 'Class', abstract_class_declaration: 'Class',
interface_declaration: 'Interface',
struct_declaration: 'Struct', struct_specifier: 'Struct',
class_specifier: 'Class', class_definition: 'Class',
impl_item: 'Impl', trait_item: 'Trait', trait_declaration: 'Trait',
record_declaration: 'Record', protocol_declaration: 'Interface',
class: 'Class', module: 'Module',
};
/** Walk up AST to find enclosing class/struct/interface, return its generateId or null. */
const findEnclosingClassId = (node: any, filePath: string): string | null => {
let current = node.parent;
while (current) {
if (CLASS_CONTAINER_TYPES.has(current.type)) {
const nameNode = current.childForFieldName?.('name')
?? current.children?.find((c: any) =>
c.type === 'type_identifier' || c.type === 'identifier' || c.type === 'name' || c.type === 'constant'
);
if (nameNode) {
const label = CONTAINER_TYPE_TO_LABEL[current.type] || 'Class';
return generateId(label, `${filePath}:${nameNode.text}`);
}
}
current = current.parent;
}
return null;
};
export const processCalls = async (
graph: KnowledgeGraph,
files: { path: string; content: string }[],
@@ -171,6 +228,8 @@ export const processCalls = async (
continue;
}
const callRouter = callRouters[language];
// 3. Process each call match
matches.forEach(match => {
const captureMap: Record<string, any> = {};
@@ -184,6 +243,68 @@ export const processCalls = async (
const calledName = nameNode.text;
// Dispatch: route language-specific calls (heritage, properties, imports)
const routed = callRouter(calledName, captureMap['call']);
if (routed) {
switch (routed.kind) {
case 'skip':
case 'import': // handled by import-processor
return;
case 'heritage':
for (const item of routed.items) {
const childId = symbolTable.lookupExact(file.path, item.enclosingClass) ||
symbolTable.lookupFuzzy(item.enclosingClass)[0]?.nodeId ||
generateId('Class', `${file.path}:${item.enclosingClass}`);
const parentId = symbolTable.lookupFuzzy(item.mixinName)[0]?.nodeId ||
generateId('Module', `${item.mixinName}`);
if (childId && parentId) {
const relId = generateId('IMPLEMENTS', `${childId}->${parentId}:${item.heritageKind}`);
graph.addRelationship({
id: relId, sourceId: childId, targetId: parentId,
type: 'IMPLEMENTS', confidence: 1.0, reason: item.heritageKind,
});
}
}
return;
case 'properties': {
const fileId = generateId('File', file.path);
const propEnclosingClassId = findEnclosingClassId(captureMap['call'], file.path);
for (const item of routed.items) {
const nodeId = generateId('Property', `${file.path}:${item.propName}`);
graph.addNode({
id: nodeId,
label: 'Property' as any, // TODO: add 'Property' to graph node label union
properties: {
name: item.propName, filePath: file.path,
startLine: item.startLine, endLine: item.endLine,
language, isExported: true,
description: item.accessorType,
},
});
symbolTable.add(file.path, item.propName, nodeId, 'Property');
const relId = generateId('DEFINES', `${fileId}->${nodeId}`);
graph.addRelationship({
id: relId, sourceId: fileId, targetId: nodeId,
type: 'DEFINES', confidence: 1.0, reason: '',
});
if (propEnclosingClassId) {
graph.addRelationship({
id: generateId('HAS_METHOD', `${propEnclosingClassId}->${nodeId}`),
sourceId: propEnclosingClassId, targetId: nodeId,
type: 'HAS_METHOD', confidence: 1.0, reason: '',
});
}
}
return;
}
case 'call':
break; // fall through to normal call processing below
}
}
// Skip common built-ins and noise
if (isBuiltInOrNoise(calledName)) return;
@@ -200,10 +321,10 @@ export const processCalls = async (
// 5. Find the enclosing function (caller)
const callNode = captureMap['call'];
const enclosingFuncId = findEnclosingFunction(callNode, file.path, symbolTable);
// Use enclosing function as source, fallback to file for top-level calls
const sourceId = enclosingFuncId || generateId('File', file.path);
const relId = generateId('CALLS', `${sourceId}:${calledName}->${resolved.nodeId}`);
graph.addRelationship({
@@ -711,37 +832,72 @@ const resolveCallTarget = (
* Filter out common built-in functions and noise
* that shouldn't be tracked as calls
*/
const isBuiltInOrNoise = (name: string): boolean => {
const builtIns = new Set([
// JavaScript/TypeScript built-ins
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export',
'fetch', 'Response', 'Request',
// React hooks and common functions
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
// Common array/object methods
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python built-ins
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'open', 'read', 'write', 'close', 'append', 'extend', 'update',
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
]);
/** Pre-built set (module-level singleton) to avoid re-creating per call */
const BUILT_IN_NAMES = new Set([
// JavaScript/TypeScript built-ins
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export',
'fetch', 'Response', 'Request',
// React hooks and common functions
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
// Common array/object methods
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python built-ins
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'open', 'read', 'write', 'close', 'append', 'extend', 'update',
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
// C/C++ standard library and common kernel helpers
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
'scanf', 'fscanf', 'sscanf',
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
'sizeof', 'offsetof', 'typeof',
'assert', 'abort', 'exit', '_exit',
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
// Linux kernel common macros/helpers (not real call targets)
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
'min', 'max', 'clamp', 'abs', 'swap',
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
'GFP_KERNEL', 'GFP_ATOMIC',
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
'mutex_lock', 'mutex_unlock', 'mutex_init',
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
'get', 'put',
// Ruby built-ins and Kernel methods
'puts', 'print', 'p', 'pp', 'warn', 'raise', 'fail',
'require', 'require_relative', 'load', 'autoload',
'include', 'extend', 'prepend',
'attr_accessor', 'attr_reader', 'attr_writer',
'public', 'private', 'protected', 'module_function',
'lambda', 'proc', 'block_given?',
'nil?', 'is_a?', 'kind_of?', 'instance_of?', 'respond_to?',
'freeze', 'frozen?', 'dup', 'clone', 'tap', 'then', 'yield_self',
// Ruby enumerables
'each', 'map', 'select', 'reject', 'find', 'detect', 'collect',
'inject', 'reduce', 'flat_map', 'each_with_object', 'each_with_index',
'any?', 'all?', 'none?', 'count', 'first', 'last',
'sort', 'sort_by', 'min', 'max', 'min_by', 'max_by',
'group_by', 'partition', 'zip', 'compact', 'flatten', 'uniq',
]);
return builtIns.has(name);
};
const isBuiltInOrNoise = (name: string): boolean => BUILT_IN_NAMES.has(name);
@@ -0,0 +1,148 @@
/**
* Shared Ruby call routing logic.
*
* Ruby expresses imports, heritage (mixins), and property definitions as
* method calls rather than syntax-level constructs. This module provides a
* routing function used by the CLI call-processor, CLI parse-worker, and
* the web call-processor so that the classification logic lives in one place.
*
* NOTE: This file is intentionally duplicated in gitnexus-web/ because the
* two packages have separate build targets (Node native vs WASM/browser).
* Keep both copies in sync until a shared package is introduced.
*/
import { SupportedLanguages } from '../../config/supported-languages';
// ── Call routing dispatch table ─────────────────────────────────────────────
/** null = this call was not routed; fall through to default call handling */
export type CallRoutingResult = RubyCallRouting | null;
export type CallRouter = (
calledName: string,
callNode: any,
) => CallRoutingResult;
/** No-op router: returns null for every call (passthrough to normal processing) */
const noRouting: CallRouter = () => null;
/** Per-language call routing. noRouting = no special routing (normal call processing) */
export const callRouters: Record<SupportedLanguages, CallRouter> = {
[SupportedLanguages.JavaScript]: noRouting,
[SupportedLanguages.TypeScript]: noRouting,
[SupportedLanguages.Python]: noRouting,
[SupportedLanguages.Java]: noRouting,
[SupportedLanguages.Go]: noRouting,
[SupportedLanguages.Rust]: noRouting,
[SupportedLanguages.CSharp]: noRouting,
[SupportedLanguages.PHP]: noRouting,
[SupportedLanguages.Swift]: noRouting,
[SupportedLanguages.CPlusPlus]: noRouting,
[SupportedLanguages.C]: noRouting,
[SupportedLanguages.Ruby]: routeRubyCall,
};
// ── Result types ────────────────────────────────────────────────────────────
export type RubyCallRouting =
| { kind: 'import'; importPath: string; isRelative: boolean }
| { kind: 'heritage'; items: RubyHeritageItem[] }
| { kind: 'properties'; items: RubyPropertyItem[] }
| { kind: 'call' }
| { kind: 'skip' };
export interface RubyHeritageItem {
enclosingClass: string;
mixinName: string;
heritageKind: 'include' | 'extend' | 'prepend';
}
export type RubyAccessorType = 'attr_accessor' | 'attr_reader' | 'attr_writer';
export interface RubyPropertyItem {
propName: string;
accessorType: RubyAccessorType;
startLine: number;
endLine: number;
}
// ── Pre-allocated singletons for common return values ────────────────────────
const CALL_RESULT: RubyCallRouting = { kind: 'call' };
const SKIP_RESULT: RubyCallRouting = { kind: 'skip' };
/** Max depth for parent-walking loops to prevent pathological AST traversals */
const MAX_PARENT_DEPTH = 50;
// ── Routing function ────────────────────────────────────────────────────────
/**
* Classify a Ruby call node and extract its semantic payload.
*
* @param calledName - The method name (e.g. 'require', 'include', 'attr_accessor')
* @param callNode - The tree-sitter `call` AST node
* @returns A discriminated union describing the call's semantic role
*/
export function routeRubyCall(calledName: string, callNode: any): RubyCallRouting {
// ── require / require_relative → import ─────────────────────────────────
if (calledName === 'require' || calledName === 'require_relative') {
const argList = callNode.childForFieldName?.('arguments');
const stringNode = argList?.children?.find((c: any) => c.type === 'string');
const contentNode = stringNode?.children?.find((c: any) => c.type === 'string_content');
if (!contentNode) return SKIP_RESULT;
let importPath: string = contentNode.text;
// Validate: reject null bytes, control chars, excessively long paths
if (!importPath || importPath.length > 1024 || /[\x00-\x1f]/.test(importPath)) {
return SKIP_RESULT;
}
const isRelative = calledName === 'require_relative';
if (isRelative && !importPath.startsWith('.')) {
importPath = './' + importPath;
}
return { kind: 'import', importPath, isRelative };
}
// ── include / extend / prepend → heritage (mixin) ──────────────────────
if (calledName === 'include' || calledName === 'extend' || calledName === 'prepend') {
let enclosingClass: string | null = null;
let current = callNode.parent;
let depth = 0;
while (current && ++depth <= MAX_PARENT_DEPTH) {
if (current.type === 'class' || current.type === 'module') {
const nameNode = current.childForFieldName?.('name');
if (nameNode) { enclosingClass = nameNode.text; break; }
}
current = current.parent;
}
if (!enclosingClass) return SKIP_RESULT;
const items: RubyHeritageItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of (argList?.children ?? [])) {
if (arg.type === 'constant' || arg.type === 'scope_resolution') {
items.push({ enclosingClass, mixinName: arg.text, heritageKind: calledName as 'include' | 'extend' | 'prepend' });
}
}
return items.length > 0 ? { kind: 'heritage', items } : SKIP_RESULT;
}
// ── attr_accessor / attr_reader / attr_writer → property definitions ───
if (calledName === 'attr_accessor' || calledName === 'attr_reader' || calledName === 'attr_writer') {
const items: RubyPropertyItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of (argList?.children ?? [])) {
if (arg.type === 'simple_symbol') {
items.push({
propName: arg.text.startsWith(':') ? arg.text.slice(1) : arg.text,
accessorType: calledName as RubyAccessorType,
startLine: arg.startPosition.row,
endLine: arg.endPosition.row,
});
}
}
return items.length > 0 ? { kind: 'properties', items } : SKIP_RESULT;
}
// ── Everything else → regular call ─────────────────────────────────────
return CALL_RESULT;
}
@@ -13,7 +13,7 @@
import { detectFrameworkFromPath } from './framework-detection';
// ============================================================================
// NAME PATTERNS - All 9 supported languages
// NAME PATTERNS - All 11 supported languages
// ============================================================================
/**
@@ -143,6 +143,13 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
/^save$/, // Repository::save()
/^delete$/, // Repository::delete()
],
// Ruby
'ruby': [
/^call$/, // Service objects (MyService.call)
/^perform$/, // Background jobs (Sidekiq, ActiveJob)
/^execute$/, // Command pattern
],
};
// ============================================================================
@@ -302,7 +309,12 @@ export function isTestFile(filePath: string): boolean {
p.endsWith('test.php') ||
p.endsWith('spec.php') ||
p.includes('/tests/feature/') ||
p.includes('/tests/unit/')
p.includes('/tests/unit/') ||
// Ruby test patterns
p.endsWith('_spec.rb') ||
p.endsWith('_test.rb') ||
p.includes('/spec/') ||
p.includes('/test/fixtures/')
);
}
@@ -257,6 +257,17 @@ export function detectFrameworkFromPath(filePath: string): FrameworkHint | null
return { framework: 'laravel', entryPointMultiplier: 1.5, reason: 'laravel-repository' };
}
// ========== RUBY ==========
// Ruby: bin/ or exe/ (CLI entry points)
if ((p.includes('/bin/') || p.includes('/exe/')) && p.endsWith('.rb')) {
return { framework: 'ruby', entryPointMultiplier: 2.5, reason: 'ruby-executable' };
}
// Ruby: Rakefile or *.rake (task definitions)
if (p.endsWith('/rakefile') || p.endsWith('.rake')) {
return { framework: 'ruby', entryPointMultiplier: 1.5, reason: 'ruby-rake' };
}
// ========== SWIFT / iOS ==========
// iOS App entry points (highest priority)
@@ -4,6 +4,7 @@ import { loadParser, loadLanguage } from '../tree-sitter/parser-loader';
import { LANGUAGE_QUERIES } from './tree-sitter-queries';
import { generateId } from '../../lib/utils';
import { getLanguageFromFilename } from './utils';
import { callRouters } from './call-routing';
// Type: Map<FilePath, Set<ResolvedFilePath>>
// Stores all files that a given file imports from
@@ -53,7 +54,9 @@ const resolveImportPath = (
// Go
'.go',
// Rust
'.rs', '/mod.rs'
'.rs', '/mod.rs',
// Ruby
'.rb', '.rake',
];
if (importPath.startsWith('.')) {
@@ -220,6 +223,35 @@ export const processImports = async (
importMap.get(file.path)!.add(resolvedPath);
}
}
// ---- Language-specific call-as-import routing (Ruby require, etc.) ----
if (captureMap['call']) {
const callNameNode = captureMap['call.name'];
if (callNameNode) {
const callRouter = callRouters[language];
const routed = callRouter(callNameNode.text, captureMap['call']);
if (routed && routed.kind === 'import') {
totalImportsFound++;
const resolvedPath = resolveImportPath(
file.path, routed.importPath, allFilePaths, allFileList, resolveCache
);
if (resolvedPath) {
const sourceId = generateId('File', file.path);
const targetId = generateId('File', resolvedPath);
const relId = generateId('IMPORTS', `${file.path}->${resolvedPath}`);
totalImportsResolved++;
graph.addRelationship({
id: relId, sourceId, targetId,
type: 'IMPORTS', confidence: 1.0, reason: '',
});
if (!importMap.has(file.path)) {
importMap.set(file.path, new Set());
}
importMap.get(file.path)!.add(resolvedPath);
}
}
}
}
});
// If re-parsed just for this, delete the tree to save memory
@@ -14,7 +14,7 @@ export type FileProgressCallback = (current: number, total: number, filePath: st
/**
* Check if a symbol (function, class, etc.) is exported/public
* Handles all 9 supported languages with explicit logic
* Handles all 11 supported languages with explicit logic
*
* @param node - The AST node for the symbol name
* @param name - The symbol name
@@ -104,7 +104,11 @@ const isNodeExported = (node: any, name: string, language: string): boolean => {
case 'c':
case 'cpp':
return false;
// Ruby: All top-level definitions are public by default
case 'ruby':
return true;
default:
return false;
}
@@ -396,6 +396,40 @@ export const PHP_QUERIES = `
[(name) (qualified_name)] @heritage.trait))) @heritage
`;
// Ruby queries - works with tree-sitter-ruby
// NOTE: Ruby uses `call` for require, include, extend, prepend, attr_* etc.
// These are all captured as @call and routed in JS post-processing:
// - require/require_relative → import extraction
// - include/extend/prepend → heritage (mixin) extraction
// - attr_accessor/attr_reader/attr_writer → property definition extraction
// - everything else → regular call extraction
export const RUBY_QUERIES = `
; ── Modules ──────────────────────────────────────────────────────────────────
(module
name: (constant) @name) @definition.module
; ── Classes ──────────────────────────────────────────────────────────────────
(class
name: (constant) @name) @definition.class
; ── Instance methods ─────────────────────────────────────────────────────────
(method
name: (identifier) @name) @definition.method
; ── Singleton (class-level) methods ──────────────────────────────────────────
(singleton_method
name: (identifier) @name) @definition.function
; ── All calls (require, include, attr_*, and regular calls routed in JS) ─────
(call
method: (identifier) @call.name) @call
; ── Heritage: class < SuperClass ─────────────────────────────────────────────
(class
name: (constant) @heritage.class
superclass: (superclass
(constant) @heritage.extends)) @heritage`;
// Swift queries - works with tree-sitter-swift
export const SWIFT_QUERIES = `
; Classes
@@ -460,6 +494,7 @@ export const LANGUAGE_QUERIES: Record<SupportedLanguages, string> = {
[SupportedLanguages.CSharp]: CSHARP_QUERIES,
[SupportedLanguages.Rust]: RUST_QUERIES,
[SupportedLanguages.PHP]: PHP_QUERIES,
[SupportedLanguages.Ruby]: RUBY_QUERIES,
[SupportedLanguages.Swift]: SWIFT_QUERIES,
};
+12
View File
@@ -1,5 +1,8 @@
import { SupportedLanguages } from '../../config/supported-languages';
/** Ruby extensionless filenames recognised as Ruby source */
const RUBY_EXTENSIONLESS_FILES = new Set(['Rakefile', 'Gemfile', 'Guardfile', 'Vagrantfile', 'Brewfile']);
/**
* Map file extension to SupportedLanguage enum
*/
@@ -31,6 +34,15 @@ export const getLanguageFromFilename = (filename: string): SupportedLanguages |
filename.endsWith('.php5') || filename.endsWith('.php8')) {
return SupportedLanguages.PHP;
}
// Ruby (extensions)
if (filename.endsWith('.rb') || filename.endsWith('.rake') || filename.endsWith('.gemspec')) {
return SupportedLanguages.Ruby;
}
// Ruby (extensionless files)
const basename = filename.split('/').pop() || filename;
if (RUBY_EXTENSIONLESS_FILES.has(basename)) {
return SupportedLanguages.Ruby;
}
// Swift
if (filename.endsWith('.swift')) return SupportedLanguages.Swift;
return null;
@@ -1,5 +1,5 @@
/**
* CSV Generator for KuzuDB Hybrid Schema
* CSV Generator for LadybugDB Hybrid Schema
*
* Generates separate CSV files for each node table and one relation CSV.
* This enables efficient bulk loading via COPY FROM for hybrid schema.
@@ -18,10 +18,10 @@ import { NODE_TABLES, NodeTableName } from './schema';
// ============================================================================
/**
* Sanitize string to ensure valid UTF-8 and safe CSV content for KuzuDB
* Sanitize string to ensure valid UTF-8 and safe CSV content for LadybugDB
* Removes or replaces invalid characters that would break CSV parsing.
*
* Critical: KuzuDB's CSV parser can misinterpret \r\n inside quoted fields.
* Critical: LadybugDB's CSV parser can misinterpret \r\n inside quoted fields.
* We normalize all line endings to \n only.
*/
const sanitizeUTF8 = (str: string): string => {
@@ -213,7 +213,7 @@ const generateCommunityCSV = (nodes: GraphNode[]): string => {
for (const node of nodes) {
if (node.label !== 'Community') continue;
// Handle keywords array - convert to KuzuDB array format
// Handle keywords array - convert to LadybugDB array format
const keywords = (node.properties as any).keywords || [];
const keywordsStr = `[${keywords.map((k: string) => `'${k.replace(/'/g, "''")}'`).join(',')}]`;
@@ -221,7 +221,7 @@ const generateCommunityCSV = (nodes: GraphNode[]): string => {
escapeCSVField(node.id),
escapeCSVField(node.properties.name || ''), // label is stored in name
escapeCSVField(node.properties.heuristicLabel || ''),
keywordsStr, // Array format for KuzuDB
keywordsStr, // Array format for LadybugDB
escapeCSVField((node.properties as any).description || ''),
escapeCSVField((node.properties as any).enrichedBy || 'heuristic'),
escapeCSVNumber(node.properties.cohesion, 0),
@@ -1,51 +1,51 @@
/**
* KuzuDB Adapter
*
* Manages the KuzuDB WASM instance for client-side graph database operations.
* LadybugDB Adapter
*
* Manages the LadybugDB WASM instance for client-side graph database operations.
* Uses the "Snapshot / Bulk Load" pattern with COPY FROM for performance.
*
*
* Multi-table schema: separate tables for File, Function, Class, etc.
*/
import { KnowledgeGraph } from '../graph/types';
import {
NODE_TABLES,
import {
NODE_TABLES,
REL_TABLE_NAME,
SCHEMA_QUERIES,
SCHEMA_QUERIES,
EMBEDDING_TABLE_NAME,
NodeTableName,
} from './schema';
import { generateAllCSVs } from './csv-generator';
// Holds the reference to the dynamically loaded module
let kuzu: any = null;
let lbug: any = null;
let db: any = null;
let conn: any = null;
/**
* Initialize KuzuDB WASM module and create in-memory database
* Initialize LadybugDB WASM module and create in-memory database
*/
export const initKuzu = async () => {
if (conn) return { db, conn, kuzu };
export const initLbug = async () => {
if (conn) return { db, conn, lbug };
try {
if (import.meta.env.DEV) console.log('🚀 Initializing KuzuDB...');
if (import.meta.env.DEV) console.log('🚀 Initializing LadybugDB...');
// 1. Dynamic Import (Fixes the "not a function" bundler issue)
const kuzuModule = await import('kuzu-wasm');
const lbugModule = await import('@ladybugdb/wasm-core');
// 2. Handle Vite/Webpack "default" wrapping
kuzu = kuzuModule.default || kuzuModule;
lbug = lbugModule.default || lbugModule;
// 3. Initialize WASM
await kuzu.init();
// 4. Create Database with 512MB buffer pool
await lbug.init();
// 4. Create Database with 512MB buffer manager
const BUFFER_POOL_SIZE = 512 * 1024 * 1024; // 512MB
db = new kuzu.Database(':memory:', BUFFER_POOL_SIZE);
conn = new kuzu.Connection(db);
if (import.meta.env.DEV) console.log('✅ KuzuDB WASM Initialized');
db = new lbug.Database(':memory:', BUFFER_POOL_SIZE);
conn = new lbug.Connection(db);
if (import.meta.env.DEV) console.log('✅ LadybugDB WASM Initialized');
// 5. Initialize Schema (all node tables, then rel tables, then embedding table)
for (const schemaQuery of SCHEMA_QUERIES) {
@@ -58,60 +58,60 @@ export const initKuzu = async () => {
}
}
}
if (import.meta.env.DEV) console.log('✅ KuzuDB Multi-Table Schema Created');
return { db, conn, kuzu };
if (import.meta.env.DEV) console.log('✅ LadybugDB Multi-Table Schema Created');
return { db, conn, lbug };
} catch (error) {
if (import.meta.env.DEV) console.error('❌ KuzuDB Initialization Failed:', error);
if (import.meta.env.DEV) console.error('❌ LadybugDB Initialization Failed:', error);
throw error;
}
};
/**
* Load a KnowledgeGraph into KuzuDB using COPY FROM (bulk load)
* Load a KnowledgeGraph into LadybugDB using COPY FROM (bulk load)
* Uses batched CSV writes and COPY statements for optimal performance
*/
export const loadGraphToKuzu = async (
graph: KnowledgeGraph,
export const loadGraphToLbug = async (
graph: KnowledgeGraph,
fileContents: Map<string, string>
) => {
const { conn, kuzu } = await initKuzu();
const { conn, lbug } = await initLbug();
try {
if (import.meta.env.DEV) console.log(`KuzuDB: Generating CSVs for ${graph.nodeCount} nodes...`);
if (import.meta.env.DEV) console.log(`LadybugDB: Generating CSVs for ${graph.nodeCount} nodes...`);
// 1. Generate all CSVs (per-table)
const csvData = generateAllCSVs(graph, fileContents);
const fs = kuzu.FS;
const fs = lbug.FS;
// 2. Write all node CSVs to virtual filesystem
const nodeFiles: Array<{ table: NodeTableName; path: string }> = [];
for (const [tableName, csv] of csvData.nodes.entries()) {
// Skip empty CSVs (only header row)
if (csv.split('\n').length <= 1) continue;
const path = `/${tableName.toLowerCase()}.csv`;
try { await fs.unlink(path); } catch {}
await fs.writeFile(path, csv);
nodeFiles.push({ table: tableName, path });
}
// 3. Parse relation CSV and prepare for INSERT (COPY FROM doesn't work with multi-pair tables)
const relLines = csvData.relCSV.split('\n').slice(1).filter(line => line.trim());
const relCount = relLines.length;
if (import.meta.env.DEV) {
console.log(`KuzuDB: Wrote ${nodeFiles.length} node CSVs, ${relCount} relations to insert`);
console.log(`LadybugDB: Wrote ${nodeFiles.length} node CSVs, ${relCount} relations to insert`);
}
// 4. COPY all node tables (must complete before rels due to FK constraints)
for (const { table, path } of nodeFiles) {
const copyQuery = getCopyQuery(table, path);
await conn.query(copyQuery);
}
// 5. INSERT relations one by one (COPY doesn't work with multi-pair REL tables)
// Build a set of valid table names for fast lookup
const validTables = new Set<string>(NODE_TABLES as readonly string[]);
@@ -135,13 +135,13 @@ export const loadGraphToKuzu = async (
// Format: "from","to","type",confidence,"reason",step
const match = line.match(/"([^"]*)","([^"]*)","([^"]*)",([0-9.]+),"([^"]*)",([0-9-]+)/);
if (!match) continue;
const [, fromId, toId, relType, confidenceStr, reason, stepStr] = match;
const fromLabel = getNodeLabel(fromId);
const toLabel = getNodeLabel(toId);
// Skip relationships where either node's label doesn't have a table in KuzuDB
// Skip relationships where either node's label doesn't have a table in LadybugDB
// Querying a non-existent table causes a fatal native crash
if (!validTables.has(fromLabel) || !validTables.has(toLabel)) {
skippedRels++;
@@ -150,7 +150,7 @@ export const loadGraphToKuzu = async (
const confidence = parseFloat(confidenceStr) || 1.0;
const step = parseInt(stepStr) || 0;
const insertQuery = `
MATCH (a:${escapeLabel(fromLabel)} {id: '${fromId.replace(/'/g, "''")}'}),
(b:${escapeLabel(toLabel)} {id: '${toId.replace(/'/g, "''")}'})
@@ -167,38 +167,39 @@ export const loadGraphToKuzu = async (
const toLabel = getNodeLabel(toId);
const key = `${relType}:${fromLabel}->` + toLabel;
skippedRelStats.set(key, (skippedRelStats.get(key) || 0) + 1);
if (import.meta.env.DEV) {
console.warn(`⚠️ Skipped: ${key} | "${fromId}" → "${toId}" | ${err instanceof Error ? err.message : String(err)}`);
}
}
}
}
if (import.meta.env.DEV) {
console.log(`KuzuDB: Inserted ${insertedRels}/${relCount} relations`);
console.log(`LadybugDB: Inserted ${insertedRels}/${relCount} relations`);
if (skippedRels > 0) {
const topSkipped = Array.from(skippedRelStats.entries())
.sort((a, b) => b[1] - a[1])
.slice(0, 10);
console.warn(`KuzuDB: Skipped ${skippedRels}/${relCount} relations (top by kind/pair):`, topSkipped);
console.warn(`LadybugDB: Skipped ${skippedRels}/${relCount} relations (top by kind/pair):`, topSkipped);
}
}
// 6. Verify results
let totalNodes = 0;
for (const tableName of NODE_TABLES) {
try {
const countRes = await conn.query(`MATCH (n:${tableName}) RETURN count(n) AS cnt`);
const countRow = await countRes.getNext();
const countRows = await countRes.getAll();
const countRow = countRows[0];
const count = countRow ? (countRow.cnt ?? countRow[0] ?? 0) : 0;
totalNodes += Number(count);
} catch {
// Table might be empty, skip
}
}
if (import.meta.env.DEV) console.log(`✅ KuzuDB Bulk Load Complete. Total nodes: ${totalNodes}, edges: ${insertedRels}`);
if (import.meta.env.DEV) console.log(`✅ LadybugDB Bulk Load Complete. Total nodes: ${totalNodes}, edges: ${insertedRels}`);
// 7. Cleanup CSV files
for (const { path } of nodeFiles) {
@@ -208,12 +209,12 @@ export const loadGraphToKuzu = async (
return { success: true, count: totalNodes };
} catch (error) {
if (import.meta.env.DEV) console.error('❌ KuzuDB Bulk Load Failed:', error);
if (import.meta.env.DEV) console.error('❌ LadybugDB Bulk Load Failed:', error);
return { success: false, count: 0 };
}
};
// KuzuDB default ESCAPE is '\' (backslash), but our CSV uses RFC 4180 escaping ("" for literal quotes).
// LadybugDB default ESCAPE is '\' (backslash), but our CSV uses RFC 4180 escaping ("" for literal quotes).
// Source code content is full of backslashes which confuse the auto-detection.
// We MUST explicitly set ESCAPE='"' and disable auto_detect.
const COPY_CSV_OPTS = `(HEADER=true, ESCAPE='"', DELIM=',', QUOTE='"', PARALLEL=false, auto_detect=false)`;
@@ -229,6 +230,9 @@ const escapeTableName = (table: string): string => {
return BACKTICK_TABLES.has(table) ? `\`${table}\`` : table;
};
/** Tables with isExported column (TypeScript/JS-native types) */
const TABLES_WITH_EXPORTED = new Set<string>(['Function', 'Class', 'Interface', 'Method', 'CodeElement']);
/**
* Get the COPY query for a node table with correct column mapping
*/
@@ -246,8 +250,12 @@ const getCopyQuery = (table: NodeTableName, path: string): string => {
if (table === 'Process') {
return `COPY ${t}(id, label, heuristicLabel, processType, stepCount, communities, entryPointId, terminalId) FROM "${path}" ${COPY_CSV_OPTS}`;
}
// Code element tables (Function, Class, Interface, Method, CodeElement, and multi-language)
return `COPY ${t}(id, name, filePath, startLine, endLine, isExported, content) FROM "${path}" ${COPY_CSV_OPTS}`;
// TypeScript/JS code element tables have isExported; multi-language tables do not
if (TABLES_WITH_EXPORTED.has(table)) {
return `COPY ${t}(id, name, filePath, startLine, endLine, isExported, content) FROM "${path}" ${COPY_CSV_OPTS}`;
}
// Multi-language tables (Struct, Impl, Trait, Macro, etc.)
return `COPY ${t}(id, name, filePath, startLine, endLine, content) FROM "${path}" ${COPY_CSV_OPTS}`;
};
/**
@@ -256,12 +264,12 @@ const getCopyQuery = (table: NodeTableName, path: string): string => {
*/
export const executeQuery = async (cypher: string): Promise<any[]> => {
if (!conn) {
await initKuzu();
await initLbug();
}
try {
const result = await conn.query(cypher);
// Extract column names from RETURN clause
const returnMatch = cypher.match(/RETURN\s+(.+?)(?:\s+ORDER|\s+LIMIT|\s+SKIP|\s*$)/is);
let columnNames: string[] = [];
@@ -284,12 +292,11 @@ export const executeQuery = async (cypher: string): Promise<any[]> => {
return col.replace(/[^a-zA-Z0-9_]/g, '_');
});
}
// Collect all rows
const allRows = await result.getAll();
const rows: any[] = [];
while (await result.hasNext()) {
const row = await result.getNext();
for (const row of allRows) {
// Convert tuple to named object if we have column names and row is array
if (Array.isArray(row) && columnNames.length === row.length) {
const namedRow: Record<string, any> = {};
@@ -302,7 +309,7 @@ export const executeQuery = async (cypher: string): Promise<any[]> => {
rows.push(row);
}
}
return rows;
} catch (error) {
if (import.meta.env.DEV) console.error('Query execution failed:', error);
@@ -313,7 +320,7 @@ export const executeQuery = async (cypher: string): Promise<any[]> => {
/**
* Get database statistics
*/
export const getKuzuStats = async (): Promise<{ nodes: number; edges: number }> => {
export const getLbugStats = async (): Promise<{ nodes: number; edges: number }> => {
if (!conn) {
return { nodes: 0, edges: 0 };
}
@@ -324,43 +331,45 @@ export const getKuzuStats = async (): Promise<{ nodes: number; edges: number }>
for (const tableName of NODE_TABLES) {
try {
const nodeResult = await conn.query(`MATCH (n:${tableName}) RETURN count(n) AS cnt`);
const nodeRow = await nodeResult.getNext();
const nodeRows = await nodeResult.getAll();
const nodeRow = nodeRows[0];
totalNodes += Number(nodeRow?.cnt ?? nodeRow?.[0] ?? 0);
} catch {
// Table might not exist or be empty
}
}
// Count edges from single relation table
let totalEdges = 0;
try {
const edgeResult = await conn.query(`MATCH ()-[r:${REL_TABLE_NAME}]->() RETURN count(r) AS cnt`);
const edgeRow = await edgeResult.getNext();
const edgeRows = await edgeResult.getAll();
const edgeRow = edgeRows[0];
totalEdges = Number(edgeRow?.cnt ?? edgeRow?.[0] ?? 0);
} catch {
// Table might not exist or be empty
}
return { nodes: totalNodes, edges: totalEdges };
} catch (error) {
if (import.meta.env.DEV) {
console.warn('Failed to get Kuzu stats:', error);
console.warn('Failed to get LadybugDB stats:', error);
}
return { nodes: 0, edges: 0 };
}
};
/**
* Check if KuzuDB is initialized and has data
* Check if LadybugDB is initialized and has data
*/
export const isKuzuReady = (): boolean => {
export const isLbugReady = (): boolean => {
return conn !== null && db !== null;
};
/**
* Close the database connection (cleanup)
*/
export const closeKuzu = async (): Promise<void> => {
export const closeLbug = async (): Promise<void> => {
if (conn) {
try {
await conn.close();
@@ -373,7 +382,7 @@ export const closeKuzu = async (): Promise<void> => {
} catch {}
db = null;
}
kuzu = null;
lbug = null;
};
/**
@@ -387,24 +396,20 @@ export const executePrepared = async (
params: Record<string, any>
): Promise<any[]> => {
if (!conn) {
await initKuzu();
await initLbug();
}
try {
const stmt = await conn.prepare(cypher);
if (!stmt.isSuccess()) {
const errMsg = await stmt.getErrorMessage();
throw new Error(`Prepare failed: ${errMsg}`);
}
const result = await conn.execute(stmt, params);
const rows: any[] = [];
while (await result.hasNext()) {
const row = await result.getNext();
rows.push(row);
}
const rows = await result.getAll();
await stmt.close();
return rows;
} catch (error) {
@@ -421,22 +426,22 @@ export const executeWithReusedStatement = async (
paramsList: Array<Record<string, any>>
): Promise<void> => {
if (!conn) {
await initKuzu();
await initLbug();
}
if (paramsList.length === 0) return;
const SUB_BATCH_SIZE = 4;
for (let i = 0; i < paramsList.length; i += SUB_BATCH_SIZE) {
const subBatch = paramsList.slice(i, i + SUB_BATCH_SIZE);
const stmt = await conn.prepare(cypher);
if (!stmt.isSuccess()) {
const errMsg = await stmt.getErrorMessage();
throw new Error(`Prepare failed: ${errMsg}`);
}
try {
for (const params of subBatch) {
await conn.execute(stmt, params);
@@ -444,7 +449,7 @@ export const executeWithReusedStatement = async (
} finally {
await stmt.close();
}
if (i + SUB_BATCH_SIZE < paramsList.length) {
await new Promise(r => setTimeout(r, 0));
}
@@ -456,65 +461,67 @@ export const executeWithReusedStatement = async (
*/
export const testArrayParams = async (): Promise<{ success: boolean; error?: string }> => {
if (!conn) {
await initKuzu();
await initLbug();
}
try {
const testEmbedding = new Array(384).fill(0).map((_, i) => i / 384);
// Get any node ID to test with (try File first, then others)
let testNodeId: string | null = null;
for (const tableName of NODE_TABLES) {
try {
const nodeResult = await conn.query(`MATCH (n:${tableName}) RETURN n.id AS id LIMIT 1`);
const nodeRow = await nodeResult.getNext();
const nodeRows = await nodeResult.getAll();
const nodeRow = nodeRows[0];
if (nodeRow) {
testNodeId = nodeRow.id ?? nodeRow[0];
break;
}
} catch {}
}
if (!testNodeId) {
return { success: false, error: 'No nodes found to test with' };
}
if (import.meta.env.DEV) {
console.log('🧪 Testing array params with node:', testNodeId);
}
// First create an embedding entry
const createQuery = `CREATE (e:${EMBEDDING_TABLE_NAME} {nodeId: $nodeId, embedding: $embedding})`;
const stmt = await conn.prepare(createQuery);
if (!stmt.isSuccess()) {
const errMsg = await stmt.getErrorMessage();
return { success: false, error: `Prepare failed: ${errMsg}` };
}
await conn.execute(stmt, {
nodeId: testNodeId,
embedding: testEmbedding,
});
await stmt.close();
// Verify it was stored
const verifyResult = await conn.query(
`MATCH (e:${EMBEDDING_TABLE_NAME} {nodeId: '${testNodeId}'}) RETURN e.embedding AS emb`
);
const verifyRow = await verifyResult.getNext();
const verifyRows = await verifyResult.getAll();
const verifyRow = verifyRows[0];
const storedEmb = verifyRow?.emb ?? verifyRow?.[0];
if (storedEmb && Array.isArray(storedEmb) && storedEmb.length === 384) {
if (import.meta.env.DEV) {
console.log('✅ Array params WORK! Stored embedding length:', storedEmb.length);
}
return { success: true };
} else {
return {
success: false,
error: `Embedding not stored correctly. Got: ${typeof storedEmb}, length: ${storedEmb?.length}`
return {
success: false,
error: `Embedding not stored correctly. Got: ${typeof storedEmb}, length: ${storedEmb?.length}`
};
}
} catch (error) {
@@ -1,5 +1,5 @@
/**
* KuzuDB Schema Definitions
* LadybugDB Schema Definitions
*
* Hybrid Schema:
* - Separate node tables for each code element type (File, Function, Class, etc.)
+2 -2
View File
@@ -17,7 +17,7 @@ import { z } from 'zod';
import { WebGPUNotAvailableError, embedText, embeddingToArray, initEmbedder, isEmbedderReady } from '../embeddings/embedder';
/**
* Tool factory - creates tools bound to the KuzuDB query functions
* Tool factory - creates tools bound to the LadybugDB query functions
*/
export const createGraphRAGTools = (
executeQuery: (cypher: string) => Promise<any[]>,
@@ -975,7 +975,7 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`,
// For code elements (Function, Class, etc.), use the direct id
const isFileTarget = targetType === 'File';
// Query each depth level separately (KuzuDB doesn't support list comprehensions on paths)
// Query each depth level separately (LadybugDB doesn't support list comprehensions on paths)
// For depth 1: direct connections only
// For depth 2+: chain multiple single-hop queries
const depthQueries: Promise<any[]>[] = [];
+1 -1
View File
@@ -224,7 +224,7 @@ export interface AgentStep {
* Graph schema information for LLM context
*/
export const GRAPH_SCHEMA_DESCRIPTION = `
KUZU GRAPH DATABASE SCHEMA (Multi-Table):
LADYBUG GRAPH DATABASE SCHEMA (Multi-Table):
NODE TABLES:
1. File - Source files
@@ -40,6 +40,7 @@ const getWasmPath = (language: SupportedLanguages, filePath?: string): string =>
[SupportedLanguages.Go]: '/wasm/go/tree-sitter-go.wasm',
[SupportedLanguages.Rust]: '/wasm/rust/tree-sitter-rust.wasm',
[SupportedLanguages.PHP]: '/wasm/php/tree-sitter-php.wasm',
[SupportedLanguages.Ruby]: '/wasm/ruby/tree-sitter-ruby.wasm',
[SupportedLanguages.Swift]: '/wasm/swift/tree-sitter-swift.wasm',
};
@@ -1,28 +1,35 @@
declare module 'kuzu-wasm' {
declare module '@ladybugdb/wasm-core' {
export function init(): Promise<void>;
export class Database {
constructor(path: string);
constructor(path: string, bufferPoolSize?: number);
close(): Promise<void>;
}
export class Connection {
constructor(db: Database);
query(cypher: string): Promise<QueryResult>;
prepare(cypher: string): Promise<PreparedStatement>;
execute(stmt: PreparedStatement, params?: Record<string, any>): Promise<QueryResult>;
close(): Promise<void>;
}
export interface QueryResult {
getAll(): Promise<any[]>;
hasNext(): Promise<boolean>;
getNext(): Promise<any>;
}
export interface PreparedStatement {
isSuccess(): boolean;
getErrorMessage(): Promise<string>;
close(): Promise<void>;
}
export const FS: {
writeFile(path: string, data: string): Promise<void>;
unlink(path: string): Promise<void>;
};
const kuzu: {
const lbug: {
init: typeof init;
Database: typeof Database;
Connection: typeof Connection;
FS: typeof FS;
};
export default kuzu;
export default lbug;
}
+66 -66
View File
@@ -26,13 +26,13 @@ import {
type HybridSearchResult,
} from '../core/search';
// Lazy import for Kuzu to avoid breaking worker if SharedArrayBuffer unavailable
let kuzuAdapter: typeof import('../core/kuzu/kuzu-adapter') | null = null;
const getKuzuAdapter = async () => {
if (!kuzuAdapter) {
kuzuAdapter = await import('../core/kuzu/kuzu-adapter');
// Lazy import for LadybugDB to avoid breaking worker if SharedArrayBuffer unavailable
let lbugAdapter: typeof import('../core/lbug/lbug-adapter') | null = null;
const getLbugAdapter = async () => {
if (!lbugAdapter) {
lbugAdapter = await import('../core/lbug/lbug-adapter');
}
return kuzuAdapter;
return lbugAdapter;
};
// Embedding state
@@ -172,52 +172,52 @@ const workerApi = {
console.log(`🔍 BM25 index built: ${bm25DocCount} documents`);
}
// Load graph into KuzuDB for querying (optional - gracefully degrades)
// Load graph into LadybugDB for querying (optional - gracefully degrades)
try {
onProgress({
phase: 'complete',
percent: 98,
message: 'Loading into KuzuDB...',
message: 'Loading into LadybugDB...',
stats: {
filesProcessed: result.graph.nodeCount,
totalFiles: result.graph.nodeCount,
nodesCreated: result.graph.nodeCount,
},
});
const kuzu = await getKuzuAdapter();
await kuzu.loadGraphToKuzu(result.graph, result.fileContents);
const lbug = await getLbugAdapter();
await lbug.loadGraphToLbug(result.graph, result.fileContents);
if (import.meta.env.DEV) {
const stats = await kuzu.getKuzuStats();
console.log('KuzuDB loaded:', stats);
const stats = await lbug.getLbugStats();
console.log('LadybugDB loaded:', stats);
console.log('📁 Stored', storedFileContents.size, 'files for grep/read tools');
}
} catch {
// KuzuDB is optional - silently continue without it
// LadybugDB is optional - silently continue without it
}
// Store clustering config for background enrichment (runs after graph loads)
if (clusteringConfig) {
pendingEnrichmentConfig = clusteringConfig;
console.log('📋 Clustering config saved for background enrichment');
}
// Convert to serializable format for transfer back to main thread
return serializePipelineResult(result);
},
/**
* Execute a Cypher query against the KuzuDB database
* Execute a Cypher query against the LadybugDB database
* @param cypher - The Cypher query string
* @returns Query results as an array of objects
*/
async runQuery(cypher: string): Promise<any[]> {
const kuzu = await getKuzuAdapter();
if (!kuzu.isKuzuReady()) {
const lbug = await getLbugAdapter();
if (!lbug.isLbugReady()) {
throw new Error('Database not ready. Please load a repository first.');
}
return kuzu.executeQuery(cypher);
return lbug.executeQuery(cypher);
},
/**
@@ -225,8 +225,8 @@ const workerApi = {
*/
async isReady(): Promise<boolean> {
try {
const kuzu = await getKuzuAdapter();
return kuzu.isKuzuReady();
const lbug = await getLbugAdapter();
return lbug.isLbugReady();
} catch {
return false;
}
@@ -237,8 +237,8 @@ const workerApi = {
*/
async getStats(): Promise<{ nodes: number; edges: number }> {
try {
const kuzu = await getKuzuAdapter();
return kuzu.getKuzuStats();
const lbug = await getLbugAdapter();
return lbug.getLbugStats();
} catch {
return { nodes: 0, edges: 0 };
}
@@ -276,29 +276,29 @@ const workerApi = {
console.log(`🔍 BM25 index built: ${bm25DocCount} documents`);
}
// Load graph into KuzuDB for querying (optional - gracefully degrades)
// Load graph into LadybugDB for querying (optional - gracefully degrades)
try {
onProgress({
phase: 'complete',
percent: 98,
message: 'Loading into KuzuDB...',
message: 'Loading into LadybugDB...',
stats: {
filesProcessed: result.graph.nodeCount,
totalFiles: result.graph.nodeCount,
nodesCreated: result.graph.nodeCount,
},
});
const kuzu = await getKuzuAdapter();
await kuzu.loadGraphToKuzu(result.graph, result.fileContents);
const lbug = await getLbugAdapter();
await lbug.loadGraphToLbug(result.graph, result.fileContents);
if (import.meta.env.DEV) {
const stats = await kuzu.getKuzuStats();
console.log('KuzuDB loaded:', stats);
const stats = await lbug.getLbugStats();
console.log('LadybugDB loaded:', stats);
console.log('📁 Stored', storedFileContents.size, 'files for grep/read tools');
}
} catch {
// KuzuDB is optional - silently continue without it
// LadybugDB is optional - silently continue without it
}
// Store clustering config for background enrichment (runs after graph loads)
@@ -325,8 +325,8 @@ const workerApi = {
onProgress: (progress: EmbeddingProgress) => void,
forceDevice?: 'webgpu' | 'wasm'
): Promise<void> {
const kuzu = await getKuzuAdapter();
if (!kuzu.isKuzuReady()) {
const lbug = await getLbugAdapter();
if (!lbug.isLbugReady()) {
throw new Error('Database not ready. Please load a repository first.');
}
@@ -343,8 +343,8 @@ const workerApi = {
};
await runEmbeddingPipeline(
kuzu.executeQuery,
kuzu.executeWithReusedStatement,
lbug.executeQuery,
lbug.executeWithReusedStatement,
progressCallback,
forceDevice ? { device: forceDevice } : {}
);
@@ -400,15 +400,15 @@ const workerApi = {
k: number = 10,
maxDistance: number = 0.5
): Promise<SemanticSearchResult[]> {
const kuzu = await getKuzuAdapter();
if (!kuzu.isKuzuReady()) {
const lbug = await getLbugAdapter();
if (!lbug.isLbugReady()) {
throw new Error('Database not ready. Please load a repository first.');
}
if (!isEmbeddingComplete) {
throw new Error('Embeddings not ready. Please wait for embedding pipeline to complete.');
}
return doSemanticSearch(kuzu.executeQuery, query, k, maxDistance);
return doSemanticSearch(lbug.executeQuery, query, k, maxDistance);
},
/**
@@ -424,15 +424,15 @@ const workerApi = {
k: number = 5,
hops: number = 2
): Promise<any[]> {
const kuzu = await getKuzuAdapter();
if (!kuzu.isKuzuReady()) {
const lbug = await getLbugAdapter();
if (!lbug.isLbugReady()) {
throw new Error('Database not ready. Please load a repository first.');
}
if (!isEmbeddingComplete) {
throw new Error('Embeddings not ready. Please wait for embedding pipeline to complete.');
}
return doSemanticSearchWithContext(kuzu.executeQuery, query, k, hops);
return doSemanticSearchWithContext(lbug.executeQuery, query, k, hops);
},
/**
@@ -458,9 +458,9 @@ const workerApi = {
let semanticResults: SemanticSearchResult[] = [];
if (isEmbeddingComplete) {
try {
const kuzu = await getKuzuAdapter();
if (kuzu.isKuzuReady()) {
semanticResults = await doSemanticSearch(kuzu.executeQuery, query, k * 3, 0.5);
const lbug = await getLbugAdapter();
if (lbug.isLbugReady()) {
semanticResults = await doSemanticSearch(lbug.executeQuery, query, k * 3, 0.5);
}
} catch {
// Semantic search failed, continue with BM25 only
@@ -516,15 +516,15 @@ const workerApi = {
},
/**
* Test if KuzuDB supports array parameters in prepared statements
* Test if LadybugDB supports array parameters in prepared statements
* This is a diagnostic function
*/
async testArrayParams(): Promise<{ success: boolean; error?: string }> {
const kuzu = await getKuzuAdapter();
if (!kuzu.isKuzuReady()) {
const lbug = await getLbugAdapter();
if (!lbug.isLbugReady()) {
return { success: false, error: 'Database not ready' };
}
return kuzu.testArrayParams();
return lbug.testArrayParams();
},
// ============================================================
@@ -539,8 +539,8 @@ const workerApi = {
*/
async initializeAgent(config: ProviderConfig, projectName?: string): Promise<{ success: boolean; error?: string }> {
try {
const kuzu = await getKuzuAdapter();
if (!kuzu.isKuzuReady()) {
const lbug = await getLbugAdapter();
if (!lbug.isLbugReady()) {
return { success: false, error: 'Database not ready. Please load a repository first.' };
}
@@ -549,31 +549,31 @@ const workerApi = {
if (!isEmbeddingComplete) {
throw new Error('Embeddings not ready');
}
return doSemanticSearch(kuzu.executeQuery, query, k, maxDistance);
return doSemanticSearch(lbug.executeQuery, query, k, maxDistance);
};
const semanticSearchWithContextWrapper = async (query: string, k?: number, hops?: number) => {
if (!isEmbeddingComplete) {
throw new Error('Embeddings not ready');
}
return doSemanticSearchWithContext(kuzu.executeQuery, query, k, hops);
return doSemanticSearchWithContext(lbug.executeQuery, query, k, hops);
};
// Hybrid search wrapper - combines BM25 + semantic
const hybridSearchWrapper = async (query: string, k?: number) => {
// Get BM25 results (always available after ingestion)
const bm25Results = searchBM25(query, (k ?? 10) * 3);
// Get semantic results if embeddings are ready
let semanticResults: any[] = [];
if (isEmbeddingComplete) {
try {
semanticResults = await doSemanticSearch(kuzu.executeQuery, query, (k ?? 10) * 3, 0.5);
semanticResults = await doSemanticSearch(lbug.executeQuery, query, (k ?? 10) * 3, 0.5);
} catch {
// Semantic search failed, continue with BM25 only
}
}
// Merge with RRF
return mergeWithRRF(bm25Results, semanticResults, k ?? 10);
};
@@ -586,7 +586,7 @@ const workerApi = {
let codebaseContext;
try {
codebaseContext = await buildCodebaseContext(kuzu.executeQuery, resolvedProjectName);
codebaseContext = await buildCodebaseContext(lbug.executeQuery, resolvedProjectName);
if (import.meta.env.DEV) {
console.log('📊 Codebase context built:', {
files: codebaseContext.stats.fileCount,
@@ -600,7 +600,7 @@ const workerApi = {
currentAgent = createGraphRAGAgent(
config,
kuzu.executeQuery,
lbug.executeQuery,
semanticSearchWrapper,
semanticSearchWithContextWrapper,
hybridSearchWrapper,
@@ -627,7 +627,7 @@ const workerApi = {
/**
* Initialize the Graph RAG agent in backend mode (HTTP-backed tools).
* Uses HTTP wrappers instead of local KuzuDB for all tool queries.
* Uses HTTP wrappers instead of local LadybugDB for all tool queries.
* @param config - Provider configuration for the LLM
* @param backendUrl - Base URL of the gitnexus serve backend
* @param repoName - Repository name on the backend
@@ -848,9 +848,9 @@ const workerApi = {
}
});
// Update KuzuDB with new data
// Update LadybugDB with new data
try {
const kuzu = await getKuzuAdapter();
const lbug = await getLbugAdapter();
onProgress(enrichments.size, enrichments.size); // Done
@@ -872,11 +872,11 @@ const workerApi = {
c.enrichedBy = "llm"
`;
await kuzu.executeQuery(query);
await lbug.executeQuery(query);
}
} catch (err) {
console.error('Failed to update KuzuDB with enrichment:', err);
console.error('Failed to update LadybugDB with enrichment:', err);
}
// Convert Map to Record for serialization
+5 -5
View File
@@ -12,11 +12,11 @@ export default defineConfig({
tailwindcss(),
wasm(),
topLevelAwait(),
// Copy kuzu-wasm worker file to assets folder for production
// Copy lbug-wasm worker file to assets folder for production
viteStaticCopy({
targets: [
{
src: 'node_modules/kuzu-wasm/kuzu_wasm_worker.js',
src: 'node_modules/@ladybugdb/wasm-core/lbug_wasm_worker.js',
dest: 'assets'
}
]
@@ -35,12 +35,12 @@ export default defineConfig({
define: {
global: 'globalThis',
},
// Optimize deps - exclude kuzu-wasm from pre-bundling (it has WASM files)
// Optimize deps - exclude lbug-wasm from pre-bundling (it has WASM files)
optimizeDeps: {
exclude: ['kuzu-wasm'],
exclude: ['@ladybugdb/wasm-core'],
include: ['buffer'],
},
// Required for KuzuDB WASM (SharedArrayBuffer needs Cross-Origin Isolation)
// Required for LadybugDB WASM (SharedArrayBuffer needs Cross-Origin Isolation)
server: {
headers: {
'Cross-Origin-Opener-Policy': 'same-origin',
+7
View File
@@ -0,0 +1,7 @@
{
"permissions": {
"allow": [
"mcp__plugin_claude-mem_mcp-search__get_observations"
]
}
}
+9
View File
@@ -0,0 +1,9 @@
FROM node:22-bookworm
WORKDIR /app
RUN apt-get update && apt-get install -y python3 make g++ && rm -rf /var/lib/apt/lists/*
COPY . .
RUN npm ci --ignore-scripts \
&& node scripts/patch-tree-sitter-swift.cjs \
&& (npm rebuild 2>&1 || true) \
&& cd node_modules/tree-sitter-kotlin && npx --yes node-gyp rebuild 2>&1
CMD ["npx", "vitest", "run", "test/integration", "--reporter=verbose"]
+18 -17
View File
@@ -96,7 +96,7 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
5. **Processes** — Traces execution flows from entry points through call chains
6. **Search** — Builds hybrid search indexes for fast retrieval
The result is a **KuzuDB graph database** stored locally in `.gitnexus/` with full-text search and semantic embeddings.
The result is a **LadybugDB graph database** stored locally in `.gitnexus/` with full-text search and semantic embeddings.
## MCP Tools
@@ -157,26 +157,27 @@ GitNexus supports indexing multiple repositories. Each `gitnexus analyze` regist
## Supported Languages
TypeScript, JavaScript, Python, Java, C, C++, C#, Go, Rust, PHP, Kotlin, Swift
TypeScript, JavaScript, Python, Java, C, C++, C#, Go, Rust, PHP, Kotlin, Swift, Ruby
### Language Feature Matrix
| Language | Imports | Types | Exports | Named Bindings | Config | Frameworks | Entry Points | Heritage |
|----------|---------|-------|---------|----------------|--------|------------|-------------|----------|
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ |
| Go | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — |
| Swift | — | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| C | — | ✓ | ✓ | — | — | ✓ | ✓ | ✓ |
| C++ | — | ✓ | ✓ | — | — | ✓ | ✓ | ✓ |
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
|----------|---------|----------------|---------|----------|-----------------|---------------------|--------|------------|-------------|
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
**Imports** — cross-file import resolution · **Types** — type annotation extraction · **Exports** — public/exported symbol detection · **Named Bindings** — `import { X }` tracking · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics · **Heritage** — class inheritance / interface implementation
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
## Agent Skills
+172 -501
View File
@@ -11,6 +11,7 @@
"license": "PolyForm-Noncommercial-1.0.0",
"dependencies": {
"@huggingface/transformers": "^3.0.0",
"@ladybugdb/core": "^0.15.1",
"@modelcontextprotocol/sdk": "^1.0.0",
"cli-progress": "^3.12.0",
"commander": "^12.0.0",
@@ -20,7 +21,7 @@
"graphology": "^0.25.4",
"graphology-indices": "^0.17.0",
"graphology-utils": "^2.3.0",
"kuzu": "^0.11.3",
"ignore": "^7.0.5",
"lru-cache": "^11.0.0",
"mnemonist": "^0.39.0",
"pandemonium": "^2.4.0",
@@ -31,10 +32,11 @@
"tree-sitter-go": "^0.21.0",
"tree-sitter-java": "^0.21.0",
"tree-sitter-javascript": "^0.21.0",
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-php": "^0.23.12",
"tree-sitter-python": "^0.21.0",
"tree-sitter-ruby": "^0.23.1",
"tree-sitter-rust": "^0.21.0",
"tree-sitter-swift": "^0.6.0",
"tree-sitter-typescript": "^0.21.0",
"uuid": "^13.0.0"
},
@@ -56,6 +58,7 @@
"node": ">=18.0.0"
},
"optionalDependencies": {
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-swift": "^0.6.0"
}
},
@@ -1147,6 +1150,133 @@
"@jridgewell/sourcemap-codec": "^1.4.14"
}
},
"node_modules/@ladybugdb/core": {
"version": "0.15.1",
"resolved": "https://registry.npmjs.org/@ladybugdb/core/-/core-0.15.1.tgz",
"integrity": "sha512-a+jhzIlS2+57Y2YWXlta7Dq5A3577dQ8YO7DzPCFZxozeiGIZn0K9v0ROO+ws4PW9BwuQYI5BXQxTEtaa1Otlg==",
"hasInstallScript": true,
"license": "MIT",
"dependencies": {
"cmake-js": "^8.0.0",
"node-addon-api": "^6.0.0"
}
},
"node_modules/@ladybugdb/core/node_modules/chownr": {
"version": "3.0.0",
"resolved": "https://registry.npmjs.org/chownr/-/chownr-3.0.0.tgz",
"integrity": "sha512-+IxzY9BZOQd/XuYPRmrvEVjF/nqj5kgT4kEq7VofrDoM1MxoRjEWkrCC3EtLi59TVawxTAn+orJwFQcrqEN1+g==",
"license": "BlueOak-1.0.0",
"engines": {
"node": ">=18"
}
},
"node_modules/@ladybugdb/core/node_modules/cmake-js": {
"version": "8.0.0",
"resolved": "https://registry.npmjs.org/cmake-js/-/cmake-js-8.0.0.tgz",
"integrity": "sha512-YbUP88RDwCvoQkZhRtGURYm9RIpWdtvZuhT87fKNoLjk8kIFIFeARpKfuZQGdwfH99GZpUmqSfcDrK62X7lTgg==",
"license": "MIT",
"dependencies": {
"debug": "^4.4.3",
"fs-extra": "^11.3.3",
"node-api-headers": "^1.8.0",
"rc": "1.2.8",
"semver": "^7.7.3",
"tar": "^7.5.6",
"url-join": "^4.0.1",
"which": "^6.0.0",
"yargs": "^17.7.2"
},
"bin": {
"cmake-js": "bin/cmake-js"
},
"engines": {
"node": "^20.17.0 || >=22.9.0"
}
},
"node_modules/@ladybugdb/core/node_modules/debug": {
"version": "4.4.3",
"resolved": "https://registry.npmjs.org/debug/-/debug-4.4.3.tgz",
"integrity": "sha512-RGwwWnwQvkVfavKVt22FGLw+xYSdzARwm0ru6DhTVA3umU5hZc28V3kO4stgYryrTlLpuvgI9GiijltAjNbcqA==",
"license": "MIT",
"dependencies": {
"ms": "^2.1.3"
},
"engines": {
"node": ">=6.0"
},
"peerDependenciesMeta": {
"supports-color": {
"optional": true
}
}
},
"node_modules/@ladybugdb/core/node_modules/isexe": {
"version": "4.0.0",
"resolved": "https://registry.npmjs.org/isexe/-/isexe-4.0.0.tgz",
"integrity": "sha512-FFUtZMpoZ8RqHS3XeXEmHWLA4thH+ZxCv2lOiPIn1Xc7CxrqhWzNSDzD+/chS/zbYezmiwWLdQC09JdQKmthOw==",
"license": "BlueOak-1.0.0",
"engines": {
"node": ">=20"
}
},
"node_modules/@ladybugdb/core/node_modules/minizlib": {
"version": "3.1.0",
"resolved": "https://registry.npmjs.org/minizlib/-/minizlib-3.1.0.tgz",
"integrity": "sha512-KZxYo1BUkWD2TVFLr0MQoM8vUUigWD3LlD83a/75BqC+4qE0Hb1Vo5v1FgcfaNXvfXzr+5EhQ6ing/CaBijTlw==",
"license": "MIT",
"dependencies": {
"minipass": "^7.1.2"
},
"engines": {
"node": ">= 18"
}
},
"node_modules/@ladybugdb/core/node_modules/ms": {
"version": "2.1.3",
"resolved": "https://registry.npmjs.org/ms/-/ms-2.1.3.tgz",
"integrity": "sha512-6FlzubTLZG3J2a/NVCAleEhjzq5oxgHyaCU9yYXvcLsvoVaHJq/s5xXI6/XXP6tz7R9xAOtHnSO/tXtF3WRTlA==",
"license": "MIT"
},
"node_modules/@ladybugdb/core/node_modules/tar": {
"version": "7.5.11",
"resolved": "https://registry.npmjs.org/tar/-/tar-7.5.11.tgz",
"integrity": "sha512-ChjMH33/KetonMTAtpYdgUFr0tbz69Fp2v7zWxQfYZX4g5ZN2nOBXm1R2xyA+lMIKrLKIoKAwFj93jE/avX9cQ==",
"license": "BlueOak-1.0.0",
"dependencies": {
"@isaacs/fs-minipass": "^4.0.0",
"chownr": "^3.0.0",
"minipass": "^7.1.2",
"minizlib": "^3.1.0",
"yallist": "^5.0.0"
},
"engines": {
"node": ">=18"
}
},
"node_modules/@ladybugdb/core/node_modules/which": {
"version": "6.0.1",
"resolved": "https://registry.npmjs.org/which/-/which-6.0.1.tgz",
"integrity": "sha512-oGLe46MIrCRqX7ytPUf66EAYvdeMIZYn3WaocqqKZAxrBpkqHfL/qvTyJ/bTk5+AqHCjXmrv3CEWgy368zhRUg==",
"license": "ISC",
"dependencies": {
"isexe": "^4.0.0"
},
"bin": {
"node-which": "bin/which.js"
},
"engines": {
"node": "^20.17.0 || >=22.9.0"
}
},
"node_modules/@ladybugdb/core/node_modules/yallist": {
"version": "5.0.0",
"resolved": "https://registry.npmjs.org/yallist/-/yallist-5.0.0.tgz",
"integrity": "sha512-YgvUTfwqyc7UXVMrB+SImsVYSmTS8X/tSrtdNZMImM+n7+QTriRXyXim0mBrTXNeqzVF0KWGgHPeiyViFFrNDw==",
"license": "BlueOak-1.0.0",
"engines": {
"node": ">=18"
}
},
"node_modules/@modelcontextprotocol/sdk": {
"version": "1.25.3",
"resolved": "https://registry.npmjs.org/@modelcontextprotocol/sdk/-/sdk-1.25.3.tgz",
@@ -2273,26 +2403,6 @@
"url": "https://github.com/chalk/ansi-styles?sponsor=1"
}
},
"node_modules/aproba": {
"version": "2.1.0",
"resolved": "https://registry.npmjs.org/aproba/-/aproba-2.1.0.tgz",
"integrity": "sha512-tLIEcj5GuR2RSTnxNKdkK0dJ/GrC7P38sUkiDmDuHfsHmbagTFAxDVIBltoklXEVIQ/f14IL8IMJ5pn9Hez1Ew==",
"license": "ISC"
},
"node_modules/are-we-there-yet": {
"version": "3.0.1",
"resolved": "https://registry.npmjs.org/are-we-there-yet/-/are-we-there-yet-3.0.1.tgz",
"integrity": "sha512-QZW4EDmGwlYur0Yyf/b2uGucHQMa8aFUP7eu9ddR73vvhFyt4V0Vl3QHPcTNJ8l6qYOBdxgXdnBXQrHilfRQBg==",
"deprecated": "This package is no longer supported.",
"license": "ISC",
"dependencies": {
"delegates": "^1.0.0",
"readable-stream": "^3.6.0"
},
"engines": {
"node": "^12.13.0 || ^14.15.0 || >=16.0.0"
}
},
"node_modules/array-flatten": {
"version": "1.1.1",
"resolved": "https://registry.npmjs.org/array-flatten/-/array-flatten-1.1.1.tgz",
@@ -2321,23 +2431,6 @@
"js-tokens": "^10.0.0"
}
},
"node_modules/asynckit": {
"version": "0.4.0",
"resolved": "https://registry.npmjs.org/asynckit/-/asynckit-0.4.0.tgz",
"integrity": "sha512-Oei9OH4tRh0YqU3GxhX79dM/mwVgvbZJaSNaRk+bshkj0S5cfHcgYakreBjrHwatXKbz+IoIdYLxrKim2MjW0Q==",
"license": "MIT"
},
"node_modules/axios": {
"version": "1.13.4",
"resolved": "https://registry.npmjs.org/axios/-/axios-1.13.4.tgz",
"integrity": "sha512-1wVkUaAO6WyaYtCkcYCOx12ZgpGf9Zif+qXa4n+oYzK558YryKqiL6UWwd5DqiH3VRW0GYhTZQ/vlgJrCoNQlg==",
"license": "MIT",
"dependencies": {
"follow-redirects": "^1.15.6",
"form-data": "^4.0.4",
"proxy-from-env": "^1.1.0"
}
},
"node_modules/body-parser": {
"version": "1.20.4",
"resolved": "https://registry.npmjs.org/body-parser/-/body-parser-1.20.4.tgz",
@@ -2432,15 +2525,6 @@
"node": ">=18"
}
},
"node_modules/chownr": {
"version": "2.0.0",
"resolved": "https://registry.npmjs.org/chownr/-/chownr-2.0.0.tgz",
"integrity": "sha512-bIomtDF5KGpdogkLd9VspvFzk9KfpyyGlS8YFVZl7TGPBHL5snIOnxeshwVgPteQ9b4Eydl+pVbIyE1DcvCWgQ==",
"license": "ISC",
"engines": {
"node": ">=10"
}
},
"node_modules/cli-progress": {
"version": "3.12.0",
"resolved": "https://registry.npmjs.org/cli-progress/-/cli-progress-3.12.0.tgz",
@@ -2581,55 +2665,6 @@
"url": "https://github.com/chalk/wrap-ansi?sponsor=1"
}
},
"node_modules/cmake-js": {
"version": "7.4.0",
"resolved": "https://registry.npmjs.org/cmake-js/-/cmake-js-7.4.0.tgz",
"integrity": "sha512-Lw0JxEHrmk+qNj1n9W9d4IvkDdYTBn7l2BW6XmtLj7WPpIo2shvxUy+YokfjMxAAOELNonQwX3stkPhM5xSC2Q==",
"license": "MIT",
"dependencies": {
"axios": "^1.6.5",
"debug": "^4",
"fs-extra": "^11.2.0",
"memory-stream": "^1.0.0",
"node-api-headers": "^1.1.0",
"npmlog": "^6.0.2",
"rc": "^1.2.7",
"semver": "^7.5.4",
"tar": "^6.2.0",
"url-join": "^4.0.1",
"which": "^2.0.2",
"yargs": "^17.7.2"
},
"bin": {
"cmake-js": "bin/cmake-js"
},
"engines": {
"node": ">= 14.15.0"
}
},
"node_modules/cmake-js/node_modules/debug": {
"version": "4.4.3",
"resolved": "https://registry.npmjs.org/debug/-/debug-4.4.3.tgz",
"integrity": "sha512-RGwwWnwQvkVfavKVt22FGLw+xYSdzARwm0ru6DhTVA3umU5hZc28V3kO4stgYryrTlLpuvgI9GiijltAjNbcqA==",
"license": "MIT",
"dependencies": {
"ms": "^2.1.3"
},
"engines": {
"node": ">=6.0"
},
"peerDependenciesMeta": {
"supports-color": {
"optional": true
}
}
},
"node_modules/cmake-js/node_modules/ms": {
"version": "2.1.3",
"resolved": "https://registry.npmjs.org/ms/-/ms-2.1.3.tgz",
"integrity": "sha512-6FlzubTLZG3J2a/NVCAleEhjzq5oxgHyaCU9yYXvcLsvoVaHJq/s5xXI6/XXP6tz7R9xAOtHnSO/tXtF3WRTlA==",
"license": "MIT"
},
"node_modules/color-convert": {
"version": "2.0.1",
"resolved": "https://registry.npmjs.org/color-convert/-/color-convert-2.0.1.tgz",
@@ -2648,27 +2683,6 @@
"integrity": "sha512-dOy+3AuW3a2wNbZHIuMZpTcgjGuLU/uBL/ubcZF9OXbDo8ff4O8yVp5Bf0efS8uEoYo5q4Fx7dY9OgQGXgAsQA==",
"license": "MIT"
},
"node_modules/color-support": {
"version": "1.1.3",
"resolved": "https://registry.npmjs.org/color-support/-/color-support-1.1.3.tgz",
"integrity": "sha512-qiBjkpbMLO/HL68y+lh4q0/O1MZFj2RX6X/KmMa3+gJD3z+WwI1ZzDHysvqHGS3mP6mznPckpXmw1nI9cJjyRg==",
"license": "ISC",
"bin": {
"color-support": "bin.js"
}
},
"node_modules/combined-stream": {
"version": "1.0.8",
"resolved": "https://registry.npmjs.org/combined-stream/-/combined-stream-1.0.8.tgz",
"integrity": "sha512-FQN4MRfuJeHf7cBbBMJFXhKSDq+2kAArBlmRBvcvFE5BB1HZKXtSFASDhdlz9zOYwxh8lDdnvmMOe/+5cdoEdg==",
"license": "MIT",
"dependencies": {
"delayed-stream": "~1.0.0"
},
"engines": {
"node": ">= 0.8"
}
},
"node_modules/commander": {
"version": "12.1.0",
"resolved": "https://registry.npmjs.org/commander/-/commander-12.1.0.tgz",
@@ -2678,12 +2692,6 @@
"node": ">=18"
}
},
"node_modules/console-control-strings": {
"version": "1.1.0",
"resolved": "https://registry.npmjs.org/console-control-strings/-/console-control-strings-1.1.0.tgz",
"integrity": "sha512-ty/fTekppD2fIwRvnZAVdeOiGd1c7YXEixbgJTNzqcxJWKQnjJ/V1bNEEE6hygpM3WjwHFUVK6HTjWSzV4a8sQ==",
"license": "ISC"
},
"node_modules/content-disposition": {
"version": "0.5.4",
"resolved": "https://registry.npmjs.org/content-disposition/-/content-disposition-0.5.4.tgz",
@@ -2803,21 +2811,6 @@
"url": "https://github.com/sponsors/ljharb"
}
},
"node_modules/delayed-stream": {
"version": "1.0.0",
"resolved": "https://registry.npmjs.org/delayed-stream/-/delayed-stream-1.0.0.tgz",
"integrity": "sha512-ZySD7Nf91aLB0RxL4KGrKHBXl7Eds1DAmEdcoVawXnLD7SDhpNgtuII2aAkg7a7QS41jxPSZ17p4VdGnMHk3MQ==",
"license": "MIT",
"engines": {
"node": ">=0.4.0"
}
},
"node_modules/delegates": {
"version": "1.0.0",
"resolved": "https://registry.npmjs.org/delegates/-/delegates-1.0.0.tgz",
"integrity": "sha512-bd2L678uiWATM6m5Z1VzNCErI3jiGzt6HGY8OVICs40JQq/HALfbyNJmp0UDakEY4pMMaN0Ly5om/B1VI/+xfQ==",
"license": "MIT"
},
"node_modules/depd": {
"version": "2.0.0",
"resolved": "https://registry.npmjs.org/depd/-/depd-2.0.0.tgz",
@@ -2930,21 +2923,6 @@
"node": ">= 0.4"
}
},
"node_modules/es-set-tostringtag": {
"version": "2.1.0",
"resolved": "https://registry.npmjs.org/es-set-tostringtag/-/es-set-tostringtag-2.1.0.tgz",
"integrity": "sha512-j6vWzfrGVfyXxge+O0x5sh6cvxAog0a/4Rdd2K36zCMV5eJ+/+tOAngRO8cODMNWbVRdVlmGZQL2YS3yR8bIUA==",
"license": "MIT",
"dependencies": {
"es-errors": "^1.3.0",
"get-intrinsic": "^1.2.6",
"has-tostringtag": "^1.0.2",
"hasown": "^2.0.2"
},
"engines": {
"node": ">= 0.4"
}
},
"node_modules/es6-error": {
"version": "4.1.1",
"resolved": "https://registry.npmjs.org/es6-error/-/es6-error-4.1.1.tgz",
@@ -3204,26 +3182,6 @@
"integrity": "sha512-MI1qs7Lo4Syw0EOzUl0xjs2lsoeqFku44KpngfIduHBYvzm8h2+7K8YMQh1JtVVVrUvhLpNwqVi4DERegUJhPQ==",
"license": "Apache-2.0"
},
"node_modules/follow-redirects": {
"version": "1.15.11",
"resolved": "https://registry.npmjs.org/follow-redirects/-/follow-redirects-1.15.11.tgz",
"integrity": "sha512-deG2P0JfjrTxl50XGCDyfI97ZGVCxIpfKYmfyrQ54n5FO/0gfIES8C/Psl6kWVDolizcaaxZJnTS0QSMxvnsBQ==",
"funding": [
{
"type": "individual",
"url": "https://github.com/sponsors/RubenVerborgh"
}
],
"license": "MIT",
"engines": {
"node": ">=4.0"
},
"peerDependenciesMeta": {
"debug": {
"optional": true
}
}
},
"node_modules/foreground-child": {
"version": "3.3.1",
"resolved": "https://registry.npmjs.org/foreground-child/-/foreground-child-3.3.1.tgz",
@@ -3240,22 +3198,6 @@
"url": "https://github.com/sponsors/isaacs"
}
},
"node_modules/form-data": {
"version": "4.0.5",
"resolved": "https://registry.npmjs.org/form-data/-/form-data-4.0.5.tgz",
"integrity": "sha512-8RipRLol37bNs2bhoV67fiTEvdTrbMUYcFTiy3+wuuOnUog2QBHCZWXDRijWQfAkhBj2Uf5UnVaiWwA5vdd82w==",
"license": "MIT",
"dependencies": {
"asynckit": "^0.4.0",
"combined-stream": "^1.0.8",
"es-set-tostringtag": "^2.1.0",
"hasown": "^2.0.2",
"mime-types": "^2.1.12"
},
"engines": {
"node": ">= 6"
}
},
"node_modules/forwarded": {
"version": "0.2.0",
"resolved": "https://registry.npmjs.org/forwarded/-/forwarded-0.2.0.tgz",
@@ -3288,30 +3230,6 @@
"node": ">=14.14"
}
},
"node_modules/fs-minipass": {
"version": "2.1.0",
"resolved": "https://registry.npmjs.org/fs-minipass/-/fs-minipass-2.1.0.tgz",
"integrity": "sha512-V/JgOLFCS+R6Vcq0slCuaeWEdNC3ouDlJMNIsacH2VtALiu9mV4LPrHc5cDl8k5aw6J8jwgWWpiTo5RYhmIzvg==",
"license": "ISC",
"dependencies": {
"minipass": "^3.0.0"
},
"engines": {
"node": ">= 8"
}
},
"node_modules/fs-minipass/node_modules/minipass": {
"version": "3.3.6",
"resolved": "https://registry.npmjs.org/minipass/-/minipass-3.3.6.tgz",
"integrity": "sha512-DxiNidxSEK+tHG6zOIklvNOwm3hvCrbUrdtzY74U6HKTJxvIDfOUL5W5P2Ghd3DTkhhKPYGqeNUIh5qcM4YBfw==",
"license": "ISC",
"dependencies": {
"yallist": "^4.0.0"
},
"engines": {
"node": ">=8"
}
},
"node_modules/fsevents": {
"version": "2.3.3",
"resolved": "https://registry.npmjs.org/fsevents/-/fsevents-2.3.3.tgz",
@@ -3336,73 +3254,6 @@
"url": "https://github.com/sponsors/ljharb"
}
},
"node_modules/gauge": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/gauge/-/gauge-4.0.4.tgz",
"integrity": "sha512-f9m+BEN5jkg6a0fZjleidjN51VE1X+mPFQ2DJ0uv1V39oCLCbsGe6yjbBnp7eK7z/+GAon99a3nHuqbuuthyPg==",
"deprecated": "This package is no longer supported.",
"license": "ISC",
"dependencies": {
"aproba": "^1.0.3 || ^2.0.0",
"color-support": "^1.1.3",
"console-control-strings": "^1.1.0",
"has-unicode": "^2.0.1",
"signal-exit": "^3.0.7",
"string-width": "^4.2.3",
"strip-ansi": "^6.0.1",
"wide-align": "^1.1.5"
},
"engines": {
"node": "^12.13.0 || ^14.15.0 || >=16.0.0"
}
},
"node_modules/gauge/node_modules/ansi-regex": {
"version": "5.0.1",
"resolved": "https://registry.npmjs.org/ansi-regex/-/ansi-regex-5.0.1.tgz",
"integrity": "sha512-quJQXlTSUGL2LH9SUXo8VwsY4soanhgo6LNSm84E1LBcE8s3O0wpdiRzyR9z/ZZJMlMWv37qOOb9pdJlMUEKFQ==",
"license": "MIT",
"engines": {
"node": ">=8"
}
},
"node_modules/gauge/node_modules/emoji-regex": {
"version": "8.0.0",
"resolved": "https://registry.npmjs.org/emoji-regex/-/emoji-regex-8.0.0.tgz",
"integrity": "sha512-MSjYzcWNOA0ewAHpz0MxpYFvwg6yjy1NG3xteoqz644VCo/RPgnr1/GGt+ic3iJTzQ8Eu3TdM14SawnVUmGE6A==",
"license": "MIT"
},
"node_modules/gauge/node_modules/signal-exit": {
"version": "3.0.7",
"resolved": "https://registry.npmjs.org/signal-exit/-/signal-exit-3.0.7.tgz",
"integrity": "sha512-wnD2ZE+l+SPC/uoS0vXeE9L1+0wuaMqKlfz9AMUo38JsyLSBWSFcHR1Rri62LZc12vLr1gb3jl7iwQhgwpAbGQ==",
"license": "ISC"
},
"node_modules/gauge/node_modules/string-width": {
"version": "4.2.3",
"resolved": "https://registry.npmjs.org/string-width/-/string-width-4.2.3.tgz",
"integrity": "sha512-wKyQRQpjJ0sIp62ErSZdGsjMJWsap5oRNihHhu6G7JVO/9jIB6UyevL+tXuOqrng8j/cxKTWyWUwvSTriiZz/g==",
"license": "MIT",
"dependencies": {
"emoji-regex": "^8.0.0",
"is-fullwidth-code-point": "^3.0.0",
"strip-ansi": "^6.0.1"
},
"engines": {
"node": ">=8"
}
},
"node_modules/gauge/node_modules/strip-ansi": {
"version": "6.0.1",
"resolved": "https://registry.npmjs.org/strip-ansi/-/strip-ansi-6.0.1.tgz",
"integrity": "sha512-Y38VPSHcqkFrCpFnQ9vuSXmquuv5oXOKpGeT6aGrr3o3Gc9AlVa6JBfUSOCnbxGGZF+/0ooI7KrPuUSztUdU5A==",
"license": "MIT",
"dependencies": {
"ansi-regex": "^5.0.1"
},
"engines": {
"node": ">=8"
}
},
"node_modules/get-caller-file": {
"version": "2.0.5",
"resolved": "https://registry.npmjs.org/get-caller-file/-/get-caller-file-2.0.5.tgz",
@@ -3618,27 +3469,6 @@
"url": "https://github.com/sponsors/ljharb"
}
},
"node_modules/has-tostringtag": {
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/has-tostringtag/-/has-tostringtag-1.0.2.tgz",
"integrity": "sha512-NqADB8VjPFLM2V0VvHUewwwsw0ZWBaIdgo+ieHtK3hasLz4qeCRjYcqfB6AQrBggRKppKF8L52/VqdVsO47Dlw==",
"license": "MIT",
"dependencies": {
"has-symbols": "^1.0.3"
},
"engines": {
"node": ">= 0.4"
},
"funding": {
"url": "https://github.com/sponsors/ljharb"
}
},
"node_modules/has-unicode": {
"version": "2.0.1",
"resolved": "https://registry.npmjs.org/has-unicode/-/has-unicode-2.0.1.tgz",
"integrity": "sha512-8Rf9Y83NBReMnx0gFzA8JImQACstCYWUplepDa9xprwwtmgEZUF0h/i5xSA625zB/I37EtrswSST6OXxwaaIJQ==",
"license": "ISC"
},
"node_modules/hasown": {
"version": "2.0.2",
"resolved": "https://registry.npmjs.org/hasown/-/hasown-2.0.2.tgz",
@@ -3700,6 +3530,15 @@
"node": ">=0.10.0"
}
},
"node_modules/ignore": {
"version": "7.0.5",
"resolved": "https://registry.npmjs.org/ignore/-/ignore-7.0.5.tgz",
"integrity": "sha512-Hs59xBNfUIunMFgWAbGX5cq6893IbWg4KnrjbYwX3tx0ztorVgTDA6B2sxf8ejHJ4wz8BqGUMYlnzNBer5NvGg==",
"license": "MIT",
"engines": {
"node": ">= 4"
}
},
"node_modules/inherits": {
"version": "2.0.4",
"resolved": "https://registry.npmjs.org/inherits/-/inherits-2.0.4.tgz",
@@ -3842,18 +3681,6 @@
"graceful-fs": "^4.1.6"
}
},
"node_modules/kuzu": {
"version": "0.11.3",
"resolved": "https://registry.npmjs.org/kuzu/-/kuzu-0.11.3.tgz",
"integrity": "sha512-4+hD3Y+YMV3e0uiqTv1/GUal47D04l8qluw1WFWg8Nx3k7rLsHG1Pmq9WHIOlf1742svxQvTYQiuY6oS1qxAZA==",
"deprecated": "Package no longer supported. Contact Support at https://www.npmjs.com/support for more info.",
"hasInstallScript": true,
"license": "MIT",
"dependencies": {
"cmake-js": "^7.3.0",
"node-addon-api": "^6.0.0"
}
},
"node_modules/long": {
"version": "5.3.2",
"resolved": "https://registry.npmjs.org/long/-/long-5.3.2.tgz",
@@ -3937,15 +3764,6 @@
"node": ">= 0.6"
}
},
"node_modules/memory-stream": {
"version": "1.0.0",
"resolved": "https://registry.npmjs.org/memory-stream/-/memory-stream-1.0.0.tgz",
"integrity": "sha512-Wm13VcsPIMdG96dzILfij09PvuS3APtcKNh7M28FsCA/w6+1mjR7hhPmfFNoilX9xU7wTdhsH5lJAm6XNzdtww==",
"license": "MIT",
"dependencies": {
"readable-stream": "^3.4.0"
}
},
"node_modules/merge-descriptors": {
"version": "1.0.3",
"resolved": "https://registry.npmjs.org/merge-descriptors/-/merge-descriptors-1.0.3.tgz",
@@ -4030,43 +3848,6 @@
"node": ">=16 || 14 >=14.17"
}
},
"node_modules/minizlib": {
"version": "2.1.2",
"resolved": "https://registry.npmjs.org/minizlib/-/minizlib-2.1.2.tgz",
"integrity": "sha512-bAxsR8BVfj60DWXHE3u30oHzfl4G7khkSuPW+qvpd7jFRHm7dLxOjUk1EHACJ/hxLY8phGJ0YhYHZo7jil7Qdg==",
"license": "MIT",
"dependencies": {
"minipass": "^3.0.0",
"yallist": "^4.0.0"
},
"engines": {
"node": ">= 8"
}
},
"node_modules/minizlib/node_modules/minipass": {
"version": "3.3.6",
"resolved": "https://registry.npmjs.org/minipass/-/minipass-3.3.6.tgz",
"integrity": "sha512-DxiNidxSEK+tHG6zOIklvNOwm3hvCrbUrdtzY74U6HKTJxvIDfOUL5W5P2Ghd3DTkhhKPYGqeNUIh5qcM4YBfw==",
"license": "ISC",
"dependencies": {
"yallist": "^4.0.0"
},
"engines": {
"node": ">=8"
}
},
"node_modules/mkdirp": {
"version": "1.0.4",
"resolved": "https://registry.npmjs.org/mkdirp/-/mkdirp-1.0.4.tgz",
"integrity": "sha512-vVqVZQyf3WLx2Shd0qJ9xuvqgAyKPLAiqITEtqW0oIUjzo3PePDd6fW9iFz30ef7Ysp/oiWqbhszeGWW2T6Gzw==",
"license": "MIT",
"bin": {
"mkdirp": "bin/cmd.js"
},
"engines": {
"node": ">=10"
}
},
"node_modules/mnemonist": {
"version": "0.39.8",
"resolved": "https://registry.npmjs.org/mnemonist/-/mnemonist-0.39.8.tgz",
@@ -4133,22 +3914,6 @@
"node-gyp-build-test": "build-test.js"
}
},
"node_modules/npmlog": {
"version": "6.0.2",
"resolved": "https://registry.npmjs.org/npmlog/-/npmlog-6.0.2.tgz",
"integrity": "sha512-/vBvz5Jfr9dT/aFWd0FIRf+T/Q2WBsLENygUaFUqstqsycmZAP/t5BvFJTK0viFmSUxiUKTUplWy5vt+rvKIxg==",
"deprecated": "This package is no longer supported.",
"license": "ISC",
"dependencies": {
"are-we-there-yet": "^3.0.0",
"console-control-strings": "^1.1.0",
"gauge": "^4.0.3",
"set-blocking": "^2.0.0"
},
"engines": {
"node": "^12.13.0 || ^14.15.0 || >=16.0.0"
}
},
"node_modules/object-assign": {
"version": "4.1.1",
"resolved": "https://registry.npmjs.org/object-assign/-/object-assign-4.1.1.tgz",
@@ -4469,12 +4234,6 @@
"node": ">= 0.10"
}
},
"node_modules/proxy-from-env": {
"version": "1.1.0",
"resolved": "https://registry.npmjs.org/proxy-from-env/-/proxy-from-env-1.1.0.tgz",
"integrity": "sha512-D+zkORCbA9f1tdWRK0RaCR3GPv50cMxcrz4X8k5LTSUD1Dkw47mKJEZQNunItRTkWwgtaUSo1RVFRIG9ZXiFYg==",
"license": "MIT"
},
"node_modules/qs": {
"version": "6.14.1",
"resolved": "https://registry.npmjs.org/qs/-/qs-6.14.1.tgz",
@@ -4545,20 +4304,6 @@
"rc": "cli.js"
}
},
"node_modules/readable-stream": {
"version": "3.6.2",
"resolved": "https://registry.npmjs.org/readable-stream/-/readable-stream-3.6.2.tgz",
"integrity": "sha512-9u/sniCrY3D5WdsERHzHE4G2YCXqoG5FTHUiCC4SIbr6XcLZBY05ya9EKjYek9O5xOAwjGq+1JdGBAS7Q9ScoA==",
"license": "MIT",
"dependencies": {
"inherits": "^2.0.3",
"string_decoder": "^1.1.1",
"util-deprecate": "^1.0.1"
},
"engines": {
"node": ">= 6"
}
},
"node_modules/require-directory": {
"version": "2.1.1",
"resolved": "https://registry.npmjs.org/require-directory/-/require-directory-2.1.1.tgz",
@@ -4802,12 +4547,6 @@
"node": ">= 0.8.0"
}
},
"node_modules/set-blocking": {
"version": "2.0.0",
"resolved": "https://registry.npmjs.org/set-blocking/-/set-blocking-2.0.0.tgz",
"integrity": "sha512-KiKBS8AnWGEyLzofFfmvKwpdPzqiy16LvQfK3yv/fVH7Bj13/wl3JSR1J+rfgRE9q7xUJK4qvgS8raSOeLUehw==",
"license": "ISC"
},
"node_modules/setprototypeof": {
"version": "1.2.0",
"resolved": "https://registry.npmjs.org/setprototypeof/-/setprototypeof-1.2.0.tgz",
@@ -5009,15 +4748,6 @@
"dev": true,
"license": "MIT"
},
"node_modules/string_decoder": {
"version": "1.3.0",
"resolved": "https://registry.npmjs.org/string_decoder/-/string_decoder-1.3.0.tgz",
"integrity": "sha512-hkRX8U1WjJFd8LsDJ2yQ/wWWxaopEsABU1XfkM8A+j0+85JAGppt16cr1Whg6KIbb4okU6Mql6BOj+uup/wKeA==",
"license": "MIT",
"dependencies": {
"safe-buffer": "~5.2.0"
}
},
"node_modules/string-width": {
"version": "5.1.2",
"resolved": "https://registry.npmjs.org/string-width/-/string-width-5.1.2.tgz",
@@ -5136,33 +4866,6 @@
"node": ">=8"
}
},
"node_modules/tar": {
"version": "6.2.1",
"resolved": "https://registry.npmjs.org/tar/-/tar-6.2.1.tgz",
"integrity": "sha512-DZ4yORTwrbTj/7MZYq2w+/ZFdI6OZ/f9SFHR+71gIVUZhOQPHzVCLpvRnPgyaMpfWxxk/4ONva3GQSyNIKRv6A==",
"deprecated": "Old versions of tar are not supported, and contain widely publicized security vulnerabilities, which have been fixed in the current version. Please update. Support for old versions may be purchased (at exhorbitant rates) by contacting i@izs.me",
"license": "ISC",
"dependencies": {
"chownr": "^2.0.0",
"fs-minipass": "^2.0.0",
"minipass": "^5.0.0",
"minizlib": "^2.1.1",
"mkdirp": "^1.0.3",
"yallist": "^4.0.0"
},
"engines": {
"node": ">=10"
}
},
"node_modules/tar/node_modules/minipass": {
"version": "5.0.0",
"resolved": "https://registry.npmjs.org/minipass/-/minipass-5.0.0.tgz",
"integrity": "sha512-3FnjYuehv9k6ovOEbyOswadCDPX1piCfhV8ncmYtHOjuPwylVWsghTLo7rabjC3Rx5xD4HDx8Wm1xnMF7S5qFQ==",
"license": "ISC",
"engines": {
"node": ">=8"
}
},
"node_modules/tinybench": {
"version": "2.9.0",
"resolved": "https://registry.npmjs.org/tinybench/-/tinybench-2.9.0.tgz",
@@ -5415,6 +5118,7 @@
"integrity": "sha512-A4obq6bjzmYrA+F0JLLoheFPcofFkctNaZSpnDd+GPn1SfVZLY4/GG4C0cYVBTOShuPBGGAOPLM1JWLZQV4m1g==",
"hasInstallScript": true,
"license": "MIT",
"optional": true,
"dependencies": {
"node-addon-api": "^7.1.0",
"node-gyp-build": "^4.8.0"
@@ -5432,7 +5136,8 @@
"version": "7.1.1",
"resolved": "https://registry.npmjs.org/node-addon-api/-/node-addon-api-7.1.1.tgz",
"integrity": "sha512-5m3bsyrjFWE1xf7nz7YXdN4udnVtXK6/Yfgn5qnahL6bCkf2yKt4k3nuTKAtT4r3IG8JNR2ncsIMdZuAzJjHQQ==",
"license": "MIT"
"license": "MIT",
"optional": true
},
"node_modules/tree-sitter-php": {
"version": "0.23.12",
@@ -5487,6 +5192,34 @@
"integrity": "sha512-5m3bsyrjFWE1xf7nz7YXdN4udnVtXK6/Yfgn5qnahL6bCkf2yKt4k3nuTKAtT4r3IG8JNR2ncsIMdZuAzJjHQQ==",
"license": "MIT"
},
"node_modules/tree-sitter-ruby": {
"version": "0.23.1",
"resolved": "https://registry.npmjs.org/tree-sitter-ruby/-/tree-sitter-ruby-0.23.1.tgz",
"integrity": "sha512-d9/RXgWjR6HanN7wTYhS5bpBQLz1VkH048Vm3CodPGyJVnamXMGb8oEhDypVCBq4QnHui9sTXuJBBP3WtCw5RA==",
"hasInstallScript": true,
"license": "MIT",
"dependencies": {
"node-addon-api": "^8.2.2",
"node-gyp-build": "^4.8.2"
},
"peerDependencies": {
"tree-sitter": "^0.21.1"
},
"peerDependenciesMeta": {
"tree-sitter": {
"optional": true
}
}
},
"node_modules/tree-sitter-ruby/node_modules/node-addon-api": {
"version": "8.6.0",
"resolved": "https://registry.npmjs.org/node-addon-api/-/node-addon-api-8.6.0.tgz",
"integrity": "sha512-gBVjCaqDlRUk0EwoPNKzIr9KkS9041G/q31IBShPs1Xz6UTA+EXdZADbzqAJQrpDRq71CIMnOP5VMut3SL0z5Q==",
"license": "MIT",
"engines": {
"node": "^18 || ^20 || >= 21"
}
},
"node_modules/tree-sitter-rust": {
"version": "0.21.0",
"resolved": "https://registry.npmjs.org/tree-sitter-rust/-/tree-sitter-rust-0.21.0.tgz",
@@ -5677,12 +5410,6 @@
"integrity": "sha512-jk1+QP6ZJqyOiuEI9AEWQfju/nB2Pw466kbA0LEZljHwKeMgd9WrAEgEGxjPDD2+TNbbb37rTyhEfrCXfuKXnA==",
"license": "MIT"
},
"node_modules/util-deprecate": {
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/util-deprecate/-/util-deprecate-1.0.2.tgz",
"integrity": "sha512-EPD5q1uXyFxJpCrLnCc1nHnq3gOa6DZBocAIiI2TaSCA7VCJ1UJDMagCzIkXNsUYfD1daK//LTEQ8xiIbrHtcw==",
"license": "MIT"
},
"node_modules/utils-merge": {
"version": "1.0.1",
"resolved": "https://registry.npmjs.org/utils-merge/-/utils-merge-1.0.1.tgz",
@@ -5899,56 +5626,6 @@
"node": ">=8"
}
},
"node_modules/wide-align": {
"version": "1.1.5",
"resolved": "https://registry.npmjs.org/wide-align/-/wide-align-1.1.5.tgz",
"integrity": "sha512-eDMORYaPNZ4sQIuuYPDHdQvf4gyCF9rEEV/yPxGfwPkRodwEgiMUUXTx/dex+Me0wxx53S+NgUHaP7y3MGlDmg==",
"license": "ISC",
"dependencies": {
"string-width": "^1.0.2 || 2 || 3 || 4"
}
},
"node_modules/wide-align/node_modules/ansi-regex": {
"version": "5.0.1",
"resolved": "https://registry.npmjs.org/ansi-regex/-/ansi-regex-5.0.1.tgz",
"integrity": "sha512-quJQXlTSUGL2LH9SUXo8VwsY4soanhgo6LNSm84E1LBcE8s3O0wpdiRzyR9z/ZZJMlMWv37qOOb9pdJlMUEKFQ==",
"license": "MIT",
"engines": {
"node": ">=8"
}
},
"node_modules/wide-align/node_modules/emoji-regex": {
"version": "8.0.0",
"resolved": "https://registry.npmjs.org/emoji-regex/-/emoji-regex-8.0.0.tgz",
"integrity": "sha512-MSjYzcWNOA0ewAHpz0MxpYFvwg6yjy1NG3xteoqz644VCo/RPgnr1/GGt+ic3iJTzQ8Eu3TdM14SawnVUmGE6A==",
"license": "MIT"
},
"node_modules/wide-align/node_modules/string-width": {
"version": "4.2.3",
"resolved": "https://registry.npmjs.org/string-width/-/string-width-4.2.3.tgz",
"integrity": "sha512-wKyQRQpjJ0sIp62ErSZdGsjMJWsap5oRNihHhu6G7JVO/9jIB6UyevL+tXuOqrng8j/cxKTWyWUwvSTriiZz/g==",
"license": "MIT",
"dependencies": {
"emoji-regex": "^8.0.0",
"is-fullwidth-code-point": "^3.0.0",
"strip-ansi": "^6.0.1"
},
"engines": {
"node": ">=8"
}
},
"node_modules/wide-align/node_modules/strip-ansi": {
"version": "6.0.1",
"resolved": "https://registry.npmjs.org/strip-ansi/-/strip-ansi-6.0.1.tgz",
"integrity": "sha512-Y38VPSHcqkFrCpFnQ9vuSXmquuv5oXOKpGeT6aGrr3o3Gc9AlVa6JBfUSOCnbxGGZF+/0ooI7KrPuUSztUdU5A==",
"license": "MIT",
"dependencies": {
"ansi-regex": "^5.0.1"
},
"engines": {
"node": ">=8"
}
},
"node_modules/wrap-ansi": {
"version": "8.1.0",
"resolved": "https://registry.npmjs.org/wrap-ansi/-/wrap-ansi-8.1.0.tgz",
@@ -6055,12 +5732,6 @@
"node": ">=10"
}
},
"node_modules/yallist": {
"version": "4.0.0",
"resolved": "https://registry.npmjs.org/yallist/-/yallist-4.0.0.tgz",
"integrity": "sha512-3wdGidZyq5PB084XLES5TpOSRA3wjXAlIWMhum2kRcv/41Sn2emQ0dycQW4uZXLejwKvg6EsvbdlVL+FYEct7A==",
"license": "ISC"
},
"node_modules/yargs": {
"version": "17.7.2",
"resolved": "https://registry.npmjs.org/yargs/-/yargs-17.7.2.tgz",
+4 -2
View File
@@ -58,7 +58,8 @@
"graphology": "^0.25.4",
"graphology-indices": "^0.17.0",
"graphology-utils": "^2.3.0",
"kuzu": "^0.11.3",
"@ladybugdb/core": "^0.15.1",
"ignore": "^7.0.5",
"lru-cache": "^11.0.0",
"mnemonist": "^0.39.0",
"pandemonium": "^2.4.0",
@@ -69,14 +70,15 @@
"tree-sitter-go": "^0.21.0",
"tree-sitter-java": "^0.21.0",
"tree-sitter-javascript": "^0.21.0",
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-php": "^0.23.12",
"tree-sitter-python": "^0.21.0",
"tree-sitter-ruby": "^0.23.1",
"tree-sitter-rust": "^0.21.0",
"tree-sitter-typescript": "^0.21.0",
"uuid": "^13.0.0"
},
"optionalDependencies": {
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-swift": "^0.6.0"
},
"devDependencies": {
+40 -28
View File
@@ -9,12 +9,12 @@ import { execFileSync } from 'child_process';
import v8 from 'v8';
import cliProgress from 'cli-progress';
import { runPipelineFromRepo } from '../core/ingestion/pipeline.js';
import { initKuzu, loadGraphToKuzu, getKuzuStats, executeQuery, executeWithReusedStatement, closeKuzu, createFTSIndex, loadCachedEmbeddings } from '../core/kuzu/kuzu-adapter.js';
import { initLbug, loadGraphToLbug, getLbugStats, executeQuery, executeWithReusedStatement, closeLbug, createFTSIndex, loadCachedEmbeddings } from '../core/lbug/lbug-adapter.js';
// Embedding imports are lazy (dynamic import) so onnxruntime-node is never
// loaded when embeddings are not requested. This avoids crashes on Node
// versions whose ABI is not yet supported by the native binary (#89).
// disposeEmbedder intentionally not called — ONNX Runtime segfaults on cleanup (see #38)
import { getStoragePaths, saveMeta, loadMeta, addToGitignore, registerRepo, getGlobalRegistryPath } from '../storage/repo-manager.js';
import { getStoragePaths, saveMeta, loadMeta, addToGitignore, registerRepo, getGlobalRegistryPath, cleanupOldKuzuFiles } from '../storage/repo-manager.js';
import { getCurrentCommit, isGitRepo, getGitRoot } from '../storage/git.js';
import { generateAIContextFiles } from './ai-context.js';
import { generateSkillFiles, type GeneratedSkillInfo } from './skill-gen.js';
@@ -63,7 +63,7 @@ const PHASE_LABELS: Record<string, string> = {
communities: 'Detecting communities',
processes: 'Detecting processes',
complete: 'Pipeline complete',
kuzu: 'Loading into KuzuDB',
lbug: 'Loading into LadybugDB',
fts: 'Creating search indexes',
embeddings: 'Generating embeddings',
done: 'Done',
@@ -100,7 +100,15 @@ export const analyzeCommand = async (
return;
}
const { storagePath, kuzuPath } = getStoragePaths(repoPath);
const { storagePath, lbugPath } = getStoragePaths(repoPath);
// Clean up stale KuzuDB files from before the LadybugDB migration.
// If kuzu existed but lbug doesn't, we're doing a migration re-index — say so.
const kuzuResult = await cleanupOldKuzuFiles(storagePath);
if (kuzuResult.found && kuzuResult.needsReindex) {
console.log(' Migrating from KuzuDB to LadybugDB — rebuilding index...\n');
}
const currentCommit = getCurrentCommit(repoPath);
const existingMeta = await loadMeta(storagePath);
@@ -109,6 +117,10 @@ export const analyzeCommand = async (
return;
}
if (process.env.GITNEXUS_NO_GITIGNORE) {
console.log(' GITNEXUS_NO_GITIGNORE is set — skipping .gitignore (still reading .gitnexusignore)\n');
}
// Single progress bar for entire pipeline
const bar = new cliProgress.SingleBar({
format: ' {bar} {percentage}% | {phase}',
@@ -130,7 +142,7 @@ export const analyzeCommand = async (
aborted = true;
bar.stop();
console.log('\n Interrupted — cleaning up...');
closeKuzu().catch(() => {}).finally(() => process.exit(130));
closeLbug().catch(() => {}).finally(() => process.exit(130));
};
process.on('SIGINT', sigintHandler);
@@ -180,13 +192,13 @@ export const analyzeCommand = async (
if (options?.embeddings && existingMeta && !options?.force) {
try {
updateBar(0, 'Caching embeddings...');
await initKuzu(kuzuPath);
await initLbug(lbugPath);
const cached = await loadCachedEmbeddings();
cachedEmbeddingNodeIds = cached.embeddingNodeIds;
cachedEmbeddings = cached.embeddings;
await closeKuzu();
await closeLbug();
} catch {
try { await closeKuzu(); } catch {}
try { await closeLbug(); } catch {}
}
}
@@ -197,25 +209,25 @@ export const analyzeCommand = async (
updateBar(scaled, phaseLabel);
});
// ── Phase 2: KuzuDB (60–85%) ──────────────────────────────────────
updateBar(60, 'Loading into KuzuDB...');
// ── Phase 2: LadybugDB (60–85%) ──────────────────────────────────────
updateBar(60, 'Loading into LadybugDB...');
await closeKuzu();
const kuzuFiles = [kuzuPath, `${kuzuPath}.wal`, `${kuzuPath}.lock`];
for (const f of kuzuFiles) {
await closeLbug();
const lbugFiles = [lbugPath, `${lbugPath}.wal`, `${lbugPath}.lock`];
for (const f of lbugFiles) {
try { await fs.rm(f, { recursive: true, force: true }); } catch {}
}
const t0Kuzu = Date.now();
await initKuzu(kuzuPath);
let kuzuMsgCount = 0;
const kuzuResult = await loadGraphToKuzu(pipelineResult.graph, pipelineResult.repoPath, storagePath, (msg) => {
kuzuMsgCount++;
const progress = Math.min(84, 60 + Math.round((kuzuMsgCount / (kuzuMsgCount + 10)) * 24));
const t0Lbug = Date.now();
await initLbug(lbugPath);
let lbugMsgCount = 0;
const lbugResult = await loadGraphToLbug(pipelineResult.graph, pipelineResult.repoPath, storagePath, (msg) => {
lbugMsgCount++;
const progress = Math.min(84, 60 + Math.round((lbugMsgCount / (lbugMsgCount + 10)) * 24));
updateBar(progress, msg);
});
const kuzuTime = ((Date.now() - t0Kuzu) / 1000).toFixed(1);
const kuzuWarnings = kuzuResult.warnings;
const lbugTime = ((Date.now() - t0Lbug) / 1000).toFixed(1);
const lbugWarnings = lbugResult.warnings;
// ── Phase 3: FTS (85–90%) ─────────────────────────────────────────
updateBar(85, 'Creating search indexes...');
@@ -249,7 +261,7 @@ export const analyzeCommand = async (
}
// ── Phase 4: Embeddings (90–98%) ──────────────────────────────────
const stats = await getKuzuStats();
const stats = await getLbugStats();
let embeddingTime = '0.0';
let embeddingSkipped = true;
let embeddingSkipReason = 'off (use --embeddings to enable)';
@@ -334,7 +346,7 @@ export const analyzeCommand = async (
processes: pipelineResult.processResult?.stats.totalProcesses,
}, generatedSkills);
await closeKuzu();
await closeLbug();
// Note: we intentionally do NOT call disposeEmbedder() here.
// ONNX Runtime's native cleanup segfaults on macOS and some Linux configs.
// Since the process exits immediately after, Node.js reclaims everything.
@@ -355,7 +367,7 @@ export const analyzeCommand = async (
const embeddingsCached = cachedEmbeddings.length > 0;
console.log(`\n Repository indexed successfully (${totalTime}s)${embeddingsCached ? ` [${cachedEmbeddings.length} embeddings cached]` : ''}\n`);
console.log(` ${stats.nodes.toLocaleString()} nodes | ${stats.edges.toLocaleString()} edges | ${pipelineResult.communityResult?.stats.totalCommunities || 0} clusters | ${pipelineResult.processResult?.stats.totalProcesses || 0} flows`);
console.log(` KuzuDB ${kuzuTime}s | FTS ${ftsTime}s | Embeddings ${embeddingSkipped ? embeddingSkipReason : embeddingTime + 's'}`);
console.log(` LadybugDB ${lbugTime}s | FTS ${ftsTime}s | Embeddings ${embeddingSkipped ? embeddingSkipReason : embeddingTime + 's'}`);
console.log(` ${repoPath}`);
if (aiContext.files.length > 0) {
@@ -363,12 +375,12 @@ export const analyzeCommand = async (
}
// Show a quiet summary if some edge types needed fallback insertion
if (kuzuWarnings.length > 0) {
const totalFallback = kuzuWarnings.reduce((sum, w) => {
if (lbugWarnings.length > 0) {
const totalFallback = lbugWarnings.reduce((sum, w) => {
const m = w.match(/\((\d+) edges\)/);
return sum + (m ? parseInt(m[1]) : 0);
}, 0);
console.log(` Note: ${totalFallback} edges across ${kuzuWarnings.length} types inserted via fallback (schema will be updated in next release)`);
console.log(` Note: ${totalFallback} edges across ${lbugWarnings.length} types inserted via fallback (schema will be updated in next release)`);
}
try {
@@ -379,7 +391,7 @@ export const analyzeCommand = async (
console.log('');
// KuzuDB's native module holds open handles that prevent Node from exiting.
// LadybugDB's native module holds open handles that prevent Node from exiting.
// ONNX Runtime also registers native atexit hooks that segfault on some
// platforms (#38, #40). Force-exit to ensure clean termination.
process.exit(0);
+1 -1
View File
@@ -23,7 +23,7 @@ export async function augmentCommand(pattern: string): Promise<void> {
if (result) {
// IMPORTANT: Write to stderr, NOT stdout.
// KuzuDB's native module captures stdout fd at OS level during init,
// LadybugDB's native module captures stdout fd at OS level during init,
// which makes stdout permanently broken in subprocess contexts.
// stderr is never captured, so it works reliably everywhere.
// The hook reads from the subprocess's stderr.
+1 -1
View File
@@ -1,7 +1,7 @@
/**
* Eval Server — Lightweight HTTP server for SWE-bench evaluation
*
* Keeps KuzuDB warm in memory so tool calls from the agent are near-instant.
* Keeps LadybugDB warm in memory so tool calls from the agent are near-instant.
* Designed to run inside Docker containers during SWE-bench evaluation.
*
* KEY DESIGN: Returns LLM-friendly text, not raw JSON.
+1
View File
@@ -28,6 +28,7 @@ program
.option('--embeddings', 'Enable embedding generation for semantic search (off by default)')
.option('--skills', 'Generate repo-specific skill files from detected communities')
.option('-v, --verbose', 'Enable verbose ingestion warnings (default: false)')
.addHelpText('after', '\nEnvironment variables:\n GITNEXUS_NO_GITIGNORE=1 Skip .gitignore parsing (still reads .gitnexusignore)')
.action(createLazyAction(() => import('./analyze.js'), 'analyzeCommand'));
program
+1 -1
View File
@@ -11,7 +11,7 @@ import { LocalBackend } from '../mcp/local/local-backend.js';
export const mcpCommand = async () => {
// Prevent unhandled errors from crashing the MCP server process.
// KuzuDB lock conflicts and transient errors should degrade gracefully.
// LadybugDB lock conflicts and transient errors should degrade gracefully.
process.on('uncaughtException', (err) => {
console.error(`GitNexus MCP: uncaught exception — ${err.message}`);
// Process is in an undefined state after uncaughtException — exit after flushing
+27 -15
View File
@@ -10,6 +10,7 @@ import fs from 'fs/promises';
import path from 'path';
import os from 'os';
import { fileURLToPath } from 'url';
import { glob } from 'glob';
import { getGlobalDir } from '../storage/repo-manager.js';
const __filename = fileURLToPath(import.meta.url);
@@ -240,8 +241,6 @@ async function setupOpenCode(result: SetupResult): Promise<void> {
// ─── Skill Installation ───────────────────────────────────────────
const SKILL_NAMES = ['gitnexus-exploring', 'gitnexus-debugging', 'gitnexus-impact-analysis', 'gitnexus-refactoring', 'gitnexus-guide', 'gitnexus-cli'];
/**
* Install GitNexus skills to a target directory.
* Each skill is installed as {targetDir}/gitnexus-{skillName}/SKILL.md
@@ -255,25 +254,38 @@ async function installSkillsTo(targetDir: string): Promise<string[]> {
const installed: string[] = [];
const skillsRoot = path.join(__dirname, '..', '..', 'skills');
for (const skillName of SKILL_NAMES) {
let flatFiles: string[] = [];
let dirSkillFiles: string[] = [];
try {
[flatFiles, dirSkillFiles] = await Promise.all([
glob('*.md', { cwd: skillsRoot }),
glob('*/SKILL.md', { cwd: skillsRoot }),
]);
} catch {
return [];
}
const skillSources = new Map<string, { isDirectory: boolean }>();
for (const relPath of dirSkillFiles) {
skillSources.set(path.dirname(relPath), { isDirectory: true });
}
for (const relPath of flatFiles) {
const skillName = path.basename(relPath, '.md');
if (!skillSources.has(skillName)) {
skillSources.set(skillName, { isDirectory: false });
}
}
for (const [skillName, source] of skillSources) {
const skillDir = path.join(targetDir, skillName);
try {
// Try directory-based skill first (skills/{name}/SKILL.md)
const dirSource = path.join(skillsRoot, skillName);
const dirSkillFile = path.join(dirSource, 'SKILL.md');
let isDirectory = false;
try {
const stat = await fs.stat(dirSource);
isDirectory = stat.isDirectory();
} catch { /* not a directory */ }
if (isDirectory) {
if (source.isDirectory) {
const dirSource = path.join(skillsRoot, skillName);
await copyDirRecursive(dirSource, skillDir);
installed.push(skillName);
} else {
// Fall back to flat file (skills/{name}.md)
const flatSource = path.join(skillsRoot, `${skillName}.md`);
const content = await fs.readFile(flatSource, 'utf-8');
await fs.mkdir(skillDir, { recursive: true });
+13 -5
View File
@@ -4,12 +4,12 @@
* Shows the indexing status of the current repository.
*/
import { findRepo } from '../storage/repo-manager.js';
import { getCurrentCommit, isGitRepo } from '../storage/git.js';
import { findRepo, getStoragePaths, hasKuzuIndex } from '../storage/repo-manager.js';
import { getCurrentCommit, isGitRepo, getGitRoot } from '../storage/git.js';
export const statusCommand = async () => {
const cwd = process.cwd();
if (!isGitRepo(cwd)) {
console.log('Not a git repository.');
return;
@@ -17,8 +17,16 @@ export const statusCommand = async () => {
const repo = await findRepo(cwd);
if (!repo) {
console.log('Repository not indexed.');
console.log('Run: gitnexus analyze');
// Check if there's a stale KuzuDB index that needs migration
const repoRoot = getGitRoot(cwd) ?? cwd;
const { storagePath } = getStoragePaths(repoRoot);
if (await hasKuzuIndex(storagePath)) {
console.log('Repository has a stale KuzuDB index from a previous version.');
console.log('Run: gitnexus analyze (rebuilds the index with LadybugDB)');
} else {
console.log('Repository not indexed.');
console.log('Run: gitnexus analyze');
}
return;
}
+2 -2
View File
@@ -10,7 +10,7 @@
* gitnexus impact --target "AuthService" --direction upstream
* gitnexus cypher "MATCH (n:Function) RETURN n.name LIMIT 10"
*
* Note: Output goes to stderr because KuzuDB's native module captures stdout
* Note: Output goes to stderr because LadybugDB's native module captures stdout
* at the OS level during init. This is consistent with augment.ts.
*/
@@ -31,7 +31,7 @@ async function getBackend(): Promise<LocalBackend> {
function output(data: any): void {
const text = typeof data === 'string' ? data : JSON.stringify(data, null, 2);
// stderr because KuzuDB captures stdout at OS level
// stderr because LadybugDB captures stdout at OS level
process.stderr.write(text + '\n');
}
+2 -2
View File
@@ -101,7 +101,7 @@ export const wikiCommand = async (
}
// ── Check for existing index ────────────────────────────────────────
const { storagePath, kuzuPath } = getStoragePaths(repoPath);
const { storagePath, lbugPath } = getStoragePaths(repoPath);
const meta = await loadMeta(storagePath);
if (!meta) {
@@ -247,7 +247,7 @@ export const wikiCommand = async (
const generator = new WikiGenerator(
repoPath,
storagePath,
kuzuPath,
lbugPath,
llmConfig,
wikiOptions,
(phase, percent, detail) => {
+92
View File
@@ -1,3 +1,8 @@
import ignore, { type Ignore } from 'ignore';
import fs from 'fs/promises';
import nodePath from 'path';
import type { Path } from 'path-scurry';
const DEFAULT_IGNORE_LIST = new Set([
// Version Control
'.git',
@@ -186,6 +191,10 @@ const IGNORED_FILES = new Set([
// NOTE: Negation patterns in .gitnexusignore (e.g. `!vendor/`) cannot override
// entries in DEFAULT_IGNORE_LIST — this is intentional. The hardcoded list protects
// against indexing directories that are almost never source code (node_modules, .git, etc.).
// Users who need to include such directories should remove them from the hardcoded list.
export const shouldIgnorePath = (filePath: string): boolean => {
const normalizedPath = filePath.replace(/\\/g, '/');
const parts = normalizedPath.split('/');
@@ -237,3 +246,86 @@ export const shouldIgnorePath = (filePath: string): boolean => {
return false;
}
/** Check if a directory name is in the hardcoded ignore list */
export const isHardcodedIgnoredDirectory = (name: string): boolean => {
return DEFAULT_IGNORE_LIST.has(name);
};
/**
* Load .gitignore and .gitnexusignore rules from the repo root.
* Returns an `ignore` instance with all patterns, or null if no files found.
*/
export interface IgnoreOptions {
/** Skip .gitignore parsing, only read .gitnexusignore. Defaults to GITNEXUS_NO_GITIGNORE env var. */
noGitignore?: boolean;
}
export const loadIgnoreRules = async (
repoPath: string,
options?: IgnoreOptions
): Promise<Ignore | null> => {
const ig = ignore();
let hasRules = false;
// Allow users to bypass .gitignore parsing (e.g. when .gitignore accidentally excludes source files)
const skipGitignore = options?.noGitignore ?? !!process.env.GITNEXUS_NO_GITIGNORE;
const filenames = skipGitignore
? ['.gitnexusignore']
: ['.gitignore', '.gitnexusignore'];
for (const filename of filenames) {
try {
const content = await fs.readFile(nodePath.join(repoPath, filename), 'utf-8');
ig.add(content);
hasRules = true;
} catch (err: unknown) {
const code = (err as NodeJS.ErrnoException).code;
if (code !== 'ENOENT') {
console.warn(` Warning: could not read ${filename}: ${(err as Error).message}`);
}
}
}
return hasRules ? ig : null;
};
/**
* Create a glob-compatible ignore filter combining:
* - .gitignore / .gitnexusignore patterns (via `ignore` package)
* - Hardcoded DEFAULT_IGNORE_LIST, IGNORED_EXTENSIONS, IGNORED_FILES
*
* Returns an IgnoreLike object for glob's `ignore` option,
* enabling directory-level pruning during traversal.
*/
export const createIgnoreFilter = async (repoPath: string, options?: IgnoreOptions) => {
const ig = await loadIgnoreRules(repoPath, options);
return {
ignored(p: Path): boolean {
// path-scurry's Path.relative() returns POSIX paths on all platforms,
// which is what the `ignore` package expects. No explicit normalization needed.
const rel = p.relative();
if (!rel) return false;
// Check .gitignore / .gitnexusignore patterns
if (ig && ig.ignores(rel)) return true;
// Fall back to hardcoded rules
return shouldIgnorePath(rel);
},
childrenIgnored(p: Path): boolean {
// Fast path: check directory name against hardcoded list.
// Note: dot-directories (.git, .vscode, etc.) are primarily excluded by
// glob's `dot: false` option in filesystem-walker.ts. This check is
// defense-in-depth — do not remove `dot: false` assuming this covers it.
if (DEFAULT_IGNORE_LIST.has(p.name)) return true;
// Check against .gitignore / .gitnexusignore patterns.
// Test both bare path and path with trailing slash to handle
// bare-name patterns (e.g. `local`) and dir-only patterns (e.g. `local/`).
if (ig) {
const rel = p.relative();
if (rel && (ig.ignores(rel) || ig.ignores(rel + '/'))) return true;
}
return false;
},
};
};
+1 -1
View File
@@ -7,9 +7,9 @@ export enum SupportedLanguages {
CPlusPlus = 'cpp',
CSharp = 'csharp',
Go = 'go',
Ruby = 'ruby',
Rust = 'rust',
PHP = 'php',
Kotlin = 'kotlin',
// Ruby = 'ruby',
Swift = 'swift',
}
+102 -77
View File
@@ -24,7 +24,7 @@ import { listRegisteredRepos } from '../../storage/repo-manager.js';
async function findRepoForCwd(cwd: string): Promise<{
name: string;
storagePath: string;
kuzuPath: string;
lbugPath: string;
} | null> {
try {
const entries = await listRegisteredRepos({ validate: true });
@@ -66,7 +66,7 @@ async function findRepoForCwd(cwd: string): Promise<{
return {
name: bestMatch.name,
storagePath: bestMatch.storagePath,
kuzuPath: path.join(bestMatch.storagePath, 'kuzu'),
lbugPath: path.join(bestMatch.storagePath, 'lbug'),
};
} catch {
return null;
@@ -92,19 +92,19 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
const repo = await findRepoForCwd(workDir);
if (!repo) return '';
// Lazy-load kuzu adapter (skip unnecessary init)
const { initKuzu, executeQuery, isKuzuReady } = await import('../../mcp/core/kuzu-adapter.js');
const { searchFTSFromKuzu } = await import('../search/bm25-index.js');
// Lazy-load lbug adapter (skip unnecessary init)
const { initLbug, executeQuery, isLbugReady } = await import('../../mcp/core/lbug-adapter.js');
const { searchFTSFromLbug } = await import('../search/bm25-index.js');
const repoId = repo.name.toLowerCase();
// Init KuzuDB if not already
if (!isKuzuReady(repoId)) {
await initKuzu(repoId, repo.kuzuPath);
// Init LadybugDB if not already
if (!isLbugReady(repoId)) {
await initLbug(repoId, repo.lbugPath);
}
// Step 1: BM25 search (fast, no embeddings)
const bm25Results = await searchFTSFromKuzu(pattern, 10, repoId);
const bm25Results = await searchFTSFromLbug(pattern, 10, repoId);
if (bm25Results.length === 0) return '';
@@ -140,8 +140,90 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
if (symbolMatches.length === 0) return '';
// Step 3: For top matches, fetch callers/callees/processes
// Also get cluster cohesion internally for ranking
// Step 3: Batch-fetch callers/callees/processes/cohesion for top matches
// Uses batched WHERE n.id IN [...] queries instead of per-symbol queries
const uniqueSymbols = symbolMatches.slice(0, 5).filter((sym, i, arr) =>
arr.findIndex(s => s.nodeId === sym.nodeId) === i
);
if (uniqueSymbols.length === 0) return '';
const idList = uniqueSymbols.map(s => `'${s.nodeId.replace(/'/g, "''")}'`).join(', ');
// Batch fetch callers
const callersMap = new Map<string, string[]>();
try {
const rows = await executeQuery(repoId, `
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(n)
WHERE n.id IN [${idList}]
RETURN n.id AS targetId, caller.name AS name
LIMIT 15
`);
for (const r of rows) {
const tid = r.targetId || r[0];
const name = r.name || r[1];
if (tid && name) {
if (!callersMap.has(tid)) callersMap.set(tid, []);
callersMap.get(tid)!.push(name);
}
}
} catch { /* skip */ }
// Batch fetch callees
const calleesMap = new Map<string, string[]>();
try {
const rows = await executeQuery(repoId, `
MATCH (n)-[:CodeRelation {type: 'CALLS'}]->(callee)
WHERE n.id IN [${idList}]
RETURN n.id AS sourceId, callee.name AS name
LIMIT 15
`);
for (const r of rows) {
const sid = r.sourceId || r[0];
const name = r.name || r[1];
if (sid && name) {
if (!calleesMap.has(sid)) calleesMap.set(sid, []);
calleesMap.get(sid)!.push(name);
}
}
} catch { /* skip */ }
// Batch fetch processes
const processesMap = new Map<string, string[]>();
try {
const rows = await executeQuery(repoId, `
MATCH (n)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
WHERE n.id IN [${idList}]
RETURN n.id AS nodeId, p.heuristicLabel AS label, r.step AS step, p.stepCount AS stepCount
`);
for (const r of rows) {
const nid = r.nodeId || r[0];
const label = r.label || r[1];
const step = r.step || r[2];
const stepCount = r.stepCount || r[3];
if (nid && label) {
if (!processesMap.has(nid)) processesMap.set(nid, []);
processesMap.get(nid)!.push(`${label} (step ${step}/${stepCount})`);
}
}
} catch { /* skip */ }
// Batch fetch cohesion
const cohesionMap = new Map<string, number>();
try {
const rows = await executeQuery(repoId, `
MATCH (n)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
WHERE n.id IN [${idList}]
RETURN n.id AS nodeId, c.cohesion AS cohesion
`);
for (const r of rows) {
const nid = r.nodeId || r[0];
const coh = r.cohesion ?? r[1] ?? 0;
if (nid) cohesionMap.set(nid, coh);
}
} catch { /* skip */ }
// Assemble enriched results
const enriched: Array<{
name: string;
filePath: string;
@@ -150,72 +232,15 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
processes: string[];
cohesion: number;
}> = [];
const seen = new Set<string>();
for (const sym of symbolMatches.slice(0, 5)) {
if (seen.has(sym.nodeId)) continue;
seen.add(sym.nodeId);
const escaped = sym.nodeId.replace(/'/g, "''");
// Callers
let callers: string[] = [];
try {
const rows = await executeQuery(repoId, `
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(n {id: '${escaped}'})
RETURN caller.name AS name
LIMIT 3
`);
callers = rows.map((r: any) => r.name || r[0]).filter(Boolean);
} catch { /* skip */ }
// Callees
let callees: string[] = [];
try {
const rows = await executeQuery(repoId, `
MATCH (n {id: '${escaped}'})-[:CodeRelation {type: 'CALLS'}]->(callee)
RETURN callee.name AS name
LIMIT 3
`);
callees = rows.map((r: any) => r.name || r[0]).filter(Boolean);
} catch { /* skip */ }
// Processes
let processes: string[] = [];
try {
const rows = await executeQuery(repoId, `
MATCH (n {id: '${escaped}'})-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
RETURN p.heuristicLabel AS label, r.step AS step, p.stepCount AS stepCount
`);
processes = rows.map((r: any) => {
const label = r.label || r[0];
const step = r.step || r[1];
const stepCount = r.stepCount || r[2];
return `${label} (step ${step}/${stepCount})`;
}).filter(Boolean);
} catch { /* skip */ }
// Cluster cohesion (internal ranking signal)
let cohesion = 0;
try {
const rows = await executeQuery(repoId, `
MATCH (n {id: '${escaped}'})-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
RETURN c.cohesion AS cohesion
LIMIT 1
`);
if (rows.length > 0) {
cohesion = (rows[0].cohesion ?? rows[0][0]) || 0;
}
} catch { /* skip */ }
for (const sym of uniqueSymbols) {
enriched.push({
name: sym.name,
filePath: sym.filePath,
callers,
callees,
processes,
cohesion,
callers: (callersMap.get(sym.nodeId) || []).slice(0, 3),
callees: (calleesMap.get(sym.nodeId) || []).slice(0, 3),
processes: processesMap.get(sym.nodeId) || [],
cohesion: cohesionMap.get(sym.nodeId) || 0,
});
}
+1 -1
View File
@@ -262,7 +262,7 @@ export const embedBatch = async (texts: string[]): Promise<Float32Array[]> => {
};
/**
* Convert Float32Array to regular number array (for KuzuDB storage)
* Convert Float32Array to regular number array (for LadybugDB storage)
*/
export const embeddingToArray = (embedding: Float32Array): number[] => {
return Array.from(embedding);
@@ -2,10 +2,10 @@
* Embedding Pipeline Module
*
* Orchestrates the background embedding process:
* 1. Query embeddable nodes from KuzuDB
* 1. Query embeddable nodes from LadybugDB
* 2. Generate text representations
* 3. Batch embed using transformers.js
* 4. Update KuzuDB with embeddings
* 4. Update LadybugDB with embeddings
* 5. Create vector index for semantic search
*/
@@ -29,7 +29,7 @@ const isDev = process.env.NODE_ENV === 'development';
export type EmbeddingProgressCallback = (progress: EmbeddingProgress) => void;
/**
* Query all embeddable nodes from KuzuDB
* Query all embeddable nodes from LadybugDB
* Uses table-specific queries (File has different schema than code elements)
*/
const queryEmbeddableNodes = async (
@@ -104,9 +104,23 @@ const batchInsertEmbeddings = async (
* Create the vector index for semantic search
* Now indexes the separate CodeEmbedding table
*/
let vectorExtensionLoaded = false;
const createVectorIndex = async (
executeQuery: (cypher: string) => Promise<any[]>
): Promise<void> => {
// LadybugDB v0.15+ requires explicit VECTOR extension loading (once per session)
if (!vectorExtensionLoaded) {
try {
await executeQuery('INSTALL VECTOR');
await executeQuery('LOAD EXTENSION VECTOR');
vectorExtensionLoaded = true;
} catch {
// Extension may already be loaded — CREATE_VECTOR_INDEX will fail clearly if not
vectorExtensionLoaded = true;
}
}
const cypher = `
CALL CREATE_VECTOR_INDEX('CodeEmbedding', 'code_embedding_idx', 'embedding', metric := 'cosine')
`;
@@ -124,7 +138,7 @@ const createVectorIndex = async (
/**
* Run the embedding pipeline
*
* @param executeQuery - Function to execute Cypher queries against KuzuDB
* @param executeQuery - Function to execute Cypher queries against LadybugDB
* @param executeWithReusedStatement - Function to execute with reused prepared statement
* @param onProgress - Callback for progress updates
* @param config - Optional configuration override
@@ -219,7 +233,7 @@ export const runEmbeddingPipeline = async (
// Embed the batch
const embeddings = await embedBatch(texts);
// Update KuzuDB with embeddings
// Update LadybugDB with embeddings
const updates = batch.map((node, i) => ({
id: node.id,
embedding: embeddingToArray(embeddings[i]),
@@ -326,51 +340,64 @@ export const semanticSearch = async (
return [];
}
// Get metadata for each result by querying each node table
const results: SemanticSearchResult[] = [];
// Group results by label for batched metadata queries
const byLabel = new Map<string, Array<{ nodeId: string; distance: number }>>();
for (const embRow of embResults) {
const nodeId = embRow.nodeId ?? embRow[0];
const distance = embRow.distance ?? embRow[1];
// Extract label from node ID (format: Label:path:name)
const labelEndIdx = nodeId.indexOf(':');
const label = labelEndIdx > 0 ? nodeId.substring(0, labelEndIdx) : 'Unknown';
// Query the specific table for this node
// File nodes don't have startLine/endLine
if (!byLabel.has(label)) byLabel.set(label, []);
byLabel.get(label)!.push({ nodeId, distance });
}
// Batch-fetch metadata per label
const results: SemanticSearchResult[] = [];
for (const [label, items] of byLabel) {
const idList = items.map(i => `'${i.nodeId.replace(/'/g, "''")}'`).join(', ');
try {
let nodeQuery: string;
if (label === 'File') {
nodeQuery = `
MATCH (n:File {id: '${nodeId.replace(/'/g, "''")}'})
RETURN n.name AS name, n.filePath AS filePath
MATCH (n:File) WHERE n.id IN [${idList}]
RETURN n.id AS id, n.name AS name, n.filePath AS filePath
`;
} else {
nodeQuery = `
MATCH (n:${label} {id: '${nodeId.replace(/'/g, "''")}'})
RETURN n.name AS name, n.filePath AS filePath,
MATCH (n:${label}) WHERE n.id IN [${idList}]
RETURN n.id AS id, n.name AS name, n.filePath AS filePath,
n.startLine AS startLine, n.endLine AS endLine
`;
}
const nodeRows = await executeQuery(nodeQuery);
if (nodeRows.length > 0) {
const nodeRow = nodeRows[0];
results.push({
nodeId,
name: nodeRow.name ?? nodeRow[0] ?? '',
label,
filePath: nodeRow.filePath ?? nodeRow[1] ?? '',
distance,
startLine: label !== 'File' ? (nodeRow.startLine ?? nodeRow[2]) : undefined,
endLine: label !== 'File' ? (nodeRow.endLine ?? nodeRow[3]) : undefined,
});
const rowMap = new Map<string, any>();
for (const row of nodeRows) {
const id = row.id ?? row[0];
rowMap.set(id, row);
}
for (const item of items) {
const nodeRow = rowMap.get(item.nodeId);
if (nodeRow) {
results.push({
nodeId: item.nodeId,
name: nodeRow.name ?? nodeRow[1] ?? '',
label,
filePath: nodeRow.filePath ?? nodeRow[2] ?? '',
distance: item.distance,
startLine: label !== 'File' ? (nodeRow.startLine ?? nodeRow[3]) : undefined,
endLine: label !== 'File' ? (nodeRow.endLine ?? nodeRow[4]) : undefined,
});
}
}
} catch {
// Table might not exist, skip
}
}
// Re-sort by distance since batch queries may have mixed order
results.sort((a, b) => a.distance - b.distance);
return results;
};
+1 -1
View File
@@ -92,7 +92,7 @@ export interface SemanticSearchResult {
}
/**
* Node data for embedding (minimal structure from KuzuDB query)
* Node data for embedding (minimal structure from LadybugDB query)
*/
export interface EmbeddableNode {
id: string;
+569 -185
View File
@@ -1,10 +1,9 @@
import { KnowledgeGraph } from '../graph/types.js';
import { ASTCache } from './ast-cache.js';
import type { SymbolDefinition, SymbolTable } from './symbol-table.js';
import { ImportMap, PackageMap, NamedImportMap, isFileInPackageDir } from './import-processor.js';
import { resolveSymbol, resolveSymbolInternal } from './symbol-resolver.js';
import { walkBindingChain } from './named-binding-extraction.js';
import type { SymbolDefinition } from './symbol-table.js';
import Parser from 'tree-sitter';
import type { ResolutionContext } from './resolution-context.js';
import { TIER_CONFIDENCE, type ResolutionTier } from './resolution-context.js';
import { isLanguageAvailable, loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { generateId } from '../../lib/utils.js';
@@ -18,10 +17,17 @@ import {
countCallArguments,
inferCallForm,
extractReceiverName,
extractReceiverNode,
findEnclosingClassId,
CALL_EXPRESSION_TYPES,
MAX_CHAIN_DEPTH,
extractCallChain,
} from './utils.js';
import { buildTypeEnv, lookupTypeEnv } from './type-env.js';
import { buildTypeEnv } from './type-env.js';
import type { ConstructorBinding } from './type-env.js';
import { getTreeSitterBufferSize } from './constants.js';
import type { ExtractedCall, ExtractedRoute } from './workers/parse-worker.js';
import type { ExtractedCall, ExtractedHeritage, ExtractedRoute, FileConstructorBindings } from './workers/parse-worker.js';
import { callRouters } from './call-routing.js';
/**
* Walk up the AST from a node to find the enclosing function/method.
@@ -30,7 +36,7 @@ import type { ExtractedCall, ExtractedRoute } from './workers/parse-worker.js';
const findEnclosingFunction = (
node: any,
filePath: string,
symbolTable: SymbolTable
ctx: ResolutionContext
): string | null => {
let current = node.parent;
@@ -39,8 +45,10 @@ const findEnclosingFunction = (
const { funcName, label } = extractFunctionName(current);
if (funcName) {
const nodeId = symbolTable.lookupExact(filePath, funcName);
if (nodeId) return nodeId;
const resolved = ctx.resolve(funcName, filePath);
if (resolved?.tier === 'same-file' && resolved.candidates.length > 0) {
return resolved.candidates[0].nodeId;
}
return generateId(label, `${filePath}:${funcName}`);
}
@@ -51,17 +59,74 @@ const findEnclosingFunction = (
return null;
};
/**
* Verify constructor bindings against SymbolTable and infer receiver types.
* Shared between sequential (processCalls) and worker (processCallsFromExtracted) paths.
*/
const verifyConstructorBindings = (
bindings: readonly ConstructorBinding[],
filePath: string,
ctx: ResolutionContext,
graph?: KnowledgeGraph,
): Map<string, string> => {
const verified = new Map<string, string>();
for (const { scope, varName, calleeName, receiverClassName } of bindings) {
const tiered = ctx.resolve(calleeName, filePath);
const isClass = tiered?.candidates.some(def => def.type === 'Class') ?? false;
if (isClass) {
verified.set(receiverKey(scope, varName), calleeName);
} else {
let callableDefs = tiered?.candidates.filter(d =>
d.type === 'Function' || d.type === 'Method'
);
// When receiver class is known (e.g. $this->method() in PHP), narrow
// candidates to methods owned by that class to avoid false disambiguation failures.
if (callableDefs && callableDefs.length > 1 && receiverClassName) {
if (graph) {
// Worker path: use graph.getNode (fast, already in-memory)
const narrowed = callableDefs.filter(d => {
if (!d.ownerId) return false;
const owner = graph.getNode(d.ownerId);
return owner?.properties.name === receiverClassName;
});
if (narrowed.length > 0) callableDefs = narrowed;
} else {
// Sequential path: use ctx.resolve (no graph available)
const classResolved = ctx.resolve(receiverClassName, filePath);
if (classResolved && classResolved.candidates.length > 0) {
const classNodeIds = new Set(classResolved.candidates.map(c => c.nodeId));
const narrowed = callableDefs.filter(d =>
d.ownerId && classNodeIds.has(d.ownerId)
);
if (narrowed.length > 0) callableDefs = narrowed;
}
}
}
if (callableDefs && callableDefs.length === 1 && callableDefs[0].returnType) {
const typeName = extractReturnTypeName(callableDefs[0].returnType);
if (typeName) {
verified.set(receiverKey(scope, varName), typeName);
}
}
}
}
return verified;
};
export const processCalls = async (
graph: KnowledgeGraph,
files: { path: string; content: string }[],
astCache: ASTCache,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
ctx: ResolutionContext,
onProgress?: (current: number, total: number) => void,
namedImportMap?: NamedImportMap,
) => {
): Promise<ExtractedHeritage[]> => {
const parser = await loadParser();
const collectedHeritage: ExtractedHeritage[] = [];
const logSkipped = isVerboseIngestionEnabled();
const skippedByLang = logSkipped ? new Map<string, number>() : null;
@@ -70,7 +135,6 @@ export const processCalls = async (
onProgress?.(i + 1, files.length);
if (i % 20 === 0) await yieldToEventLoop();
// 1. Check language support first
const language = getLanguageFromFilename(file.path);
if (!language) continue;
if (!isLanguageAvailable(language)) {
@@ -83,24 +147,15 @@ export const processCalls = async (
const queryStr = LANGUAGE_QUERIES[language];
if (!queryStr) continue;
// 2. ALWAYS load the language before querying (parser is stateful)
await loadLanguage(language, file.path);
// 3. Get AST (Try Cache First)
let tree = astCache.get(file.path);
let wasReparsed = false;
if (!tree) {
// Cache Miss: Re-parse
// Use larger bufferSize for files > 32KB
try {
tree = parser.parse(file.content, undefined, { bufferSize: getTreeSitterBufferSize(file.content.length) });
} catch (parseError) {
// Skip files that can't be parsed
continue;
}
wasReparsed = true;
// Cache re-parsed tree so heritage phase gets hits
astCache.set(file.path, tree);
}
@@ -115,16 +170,20 @@ export const processCalls = async (
continue;
}
// Build per-file TypeEnv for receiver resolution
const lang = getLanguageFromFilename(file.path);
const typeEnv = lang ? buildTypeEnv(tree, lang) : new Map();
const typeEnv = lang ? buildTypeEnv(tree, lang, ctx.symbols) : null;
const callRouter = callRouters[language];
const verifiedReceivers = typeEnv && typeEnv.constructorBindings.length > 0
? verifyConstructorBindings(typeEnv.constructorBindings, file.path, ctx)
: new Map<string, string>();
ctx.enableCache(file.path);
// 3. Process each call match
matches.forEach(match => {
const captureMap: Record<string, any> = {};
match.captures.forEach(c => captureMap[c.name] = c.node);
// Only process @call captures
if (!captureMap['call']) return;
const nameNode = captureMap['call.name'];
@@ -132,30 +191,126 @@ export const processCalls = async (
const calledName = nameNode.text;
// Skip common built-ins and noise
const routed = callRouter(calledName, captureMap['call']);
if (routed) {
switch (routed.kind) {
case 'skip':
case 'import':
return;
case 'heritage':
for (const item of routed.items) {
collectedHeritage.push({
filePath: file.path,
className: item.enclosingClass,
parentName: item.mixinName,
kind: item.heritageKind,
});
}
return;
case 'properties': {
const fileId = generateId('File', file.path);
const propEnclosingClassId = findEnclosingClassId(captureMap['call'], file.path);
for (const item of routed.items) {
const nodeId = generateId('Property', `${file.path}:${item.propName}`);
graph.addNode({
id: nodeId,
label: 'Property',
properties: {
name: item.propName, filePath: file.path,
startLine: item.startLine, endLine: item.endLine,
language, isExported: true,
description: item.accessorType,
},
});
ctx.symbols.add(file.path, item.propName, nodeId, 'Property',
propEnclosingClassId ? { ownerId: propEnclosingClassId } : undefined);
const relId = generateId('DEFINES', `${fileId}->${nodeId}`);
graph.addRelationship({
id: relId, sourceId: fileId, targetId: nodeId,
type: 'DEFINES', confidence: 1.0, reason: '',
});
if (propEnclosingClassId) {
graph.addRelationship({
id: generateId('HAS_METHOD', `${propEnclosingClassId}->${nodeId}`),
sourceId: propEnclosingClassId, targetId: nodeId,
type: 'HAS_METHOD', confidence: 1.0, reason: '',
});
}
}
return;
}
case 'call':
break;
}
}
if (isBuiltInOrNoise(calledName)) return;
const callNode = captureMap['call'];
const callForm = inferCallForm(callNode, nameNode);
const receiverName = callForm === 'member' ? extractReceiverName(nameNode) : undefined;
const receiverTypeName = receiverName ? lookupTypeEnv(typeEnv, receiverName, callNode) : undefined;
let receiverTypeName = receiverName && typeEnv ? typeEnv.lookup(receiverName, callNode) : undefined;
// Fall back to verified constructor bindings for return type inference
if (!receiverTypeName && receiverName && verifiedReceivers.size > 0) {
const enclosingFunc = findEnclosingFunction(callNode, file.path, ctx);
const funcName = enclosingFunc ? extractFuncNameFromSourceId(enclosingFunc) : '';
receiverTypeName = lookupReceiverType(verifiedReceivers, funcName, receiverName);
}
// Fall back to class-as-receiver for static method calls (e.g. UserService.find_user()).
// When the receiver name is not a variable in TypeEnv but resolves to a Class/Struct/Interface
// through the standard tiered resolution, use it directly as the receiver type.
if (!receiverTypeName && receiverName && callForm === 'member') {
const typeResolved = ctx.resolve(receiverName, file.path);
if (typeResolved && typeResolved.candidates.some(
d => d.type === 'Class' || d.type === 'Interface' || d.type === 'Struct' || d.type === 'Enum',
)) {
receiverTypeName = receiverName;
}
}
// Fall back to chained call resolution when the receiver is a call expression
// (e.g. svc.getUser().save() — receiver of save() is getUser(), not a simple identifier).
if (callForm === 'member' && !receiverTypeName && !receiverName) {
const receiverNode = extractReceiverNode(nameNode);
if (receiverNode && CALL_EXPRESSION_TYPES.has(receiverNode.type)) {
const extracted = extractCallChain(receiverNode);
if (extracted) {
// Resolve the base receiver type if possible
let baseType = extracted.baseReceiverName && typeEnv
? typeEnv.lookup(extracted.baseReceiverName, callNode)
: undefined;
if (!baseType && extracted.baseReceiverName && verifiedReceivers.size > 0) {
const enclosingFunc = findEnclosingFunction(callNode, file.path, ctx);
const funcName = enclosingFunc ? extractFuncNameFromSourceId(enclosingFunc) : '';
baseType = lookupReceiverType(verifiedReceivers, funcName, extracted.baseReceiverName);
}
// Class-as-receiver for chain base (e.g. UserService.find_user().save())
if (!baseType && extracted.baseReceiverName) {
const cr = ctx.resolve(extracted.baseReceiverName, file.path);
if (cr?.candidates.some(d =>
d.type === 'Class' || d.type === 'Interface' || d.type === 'Struct' || d.type === 'Enum',
)) {
baseType = extracted.baseReceiverName;
}
}
receiverTypeName = resolveChainedReceiver(extracted.chain, baseType, file.path, ctx);
}
}
}
// 4. Resolve the target using priority strategy (returns confidence)
const resolved = resolveCallTarget({
calledName,
argCount: countCallArguments(callNode),
callForm,
receiverTypeName,
}, file.path, symbolTable, importMap, packageMap, namedImportMap);
}, file.path, ctx);
if (!resolved) return;
// 5. Find the enclosing function (caller)
const enclosingFuncId = findEnclosingFunction(callNode, file.path, symbolTable);
// Use enclosing function as source, fallback to file for top-level calls
const enclosingFuncId = findEnclosingFunction(callNode, file.path, ctx);
const sourceId = enclosingFuncId || generateId('File', file.path);
const relId = generateId('CALLS', `${sourceId}:${calledName}->${resolved.nodeId}`);
graph.addRelationship({
@@ -168,7 +323,7 @@ export const processCalls = async (
});
});
// Tree is now owned by the LRU cache — no manual delete needed
ctx.clearCache();
}
if (skippedByLang && skippedByLang.size > 0) {
@@ -178,6 +333,8 @@ export const processCalls = async (
);
}
}
return collectedHeritage;
};
/**
@@ -185,15 +342,8 @@ export const processCalls = async (
*/
interface ResolveResult {
nodeId: string;
confidence: number; // 0-1: how sure are we?
reason: string; // 'import-resolved' | 'same-file' | 'unique-global'
}
type ResolutionTier = 'same-file' | 'import-scoped' | 'unique-global';
interface TieredCandidates {
candidates: SymbolDefinition[];
tier: ResolutionTier;
confidence: number;
reason: string;
}
const CALLABLE_SYMBOL_TYPES = new Set([
@@ -204,73 +354,16 @@ const CALLABLE_SYMBOL_TYPES = new Set([
'Delegate',
]);
const collectTieredCandidates = (
calledName: string,
currentFile: string,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
): TieredCandidates | null => {
const allDefs = symbolTable.lookupFuzzy(calledName);
// Tier 1: Same-file — highest priority, prevents imports from shadowing local defs
// (matches resolveSymbolInternal which checks lookupExactFull before named bindings)
const localDefs = allDefs.filter(def => def.filePath === currentFile);
if (localDefs.length > 0) {
return { candidates: localDefs, tier: 'same-file' };
}
// Tier 2a-named: Check named bindings with re-export chain following.
// Aliased imports (import { User as U }) mean lookupFuzzy('U') returns
// empty but we can resolve via the exported name.
// Re-exports (export { User } from './base') are followed up to 5 hops.
if (namedImportMap) {
const chainResult = resolveNamedBindingChainForCandidates(
calledName, currentFile, symbolTable, namedImportMap, allDefs,
);
if (chainResult) return chainResult;
}
if (allDefs.length === 0) return null;
const importedFiles = importMap.get(currentFile);
if (importedFiles) {
const importedDefs = allDefs.filter(def => importedFiles.has(def.filePath));
if (importedDefs.length > 0) {
return { candidates: importedDefs, tier: 'import-scoped' };
}
}
const importedPackages = packageMap?.get(currentFile);
if (importedPackages) {
const packageDefs = allDefs.filter(def => {
for (const dirSuffix of importedPackages) {
if (isFileInPackageDir(def.filePath, dirSuffix)) return true;
}
return false;
});
if (packageDefs.length > 0) {
return { candidates: packageDefs, tier: 'import-scoped' };
}
}
// Tier 3: Global — pass all candidates through; filterCallableCandidates
// will narrow by kind/arity and resolveCallTarget only emits when exactly 1 remains.
return { candidates: allDefs, tier: 'unique-global' };
};
const CONSTRUCTOR_TARGET_TYPES = new Set(['Constructor', 'Class', 'Struct', 'Record']);
const filterCallableCandidates = (
candidates: SymbolDefinition[],
candidates: readonly SymbolDefinition[],
argCount?: number,
callForm?: 'free' | 'member' | 'constructor',
): SymbolDefinition[] => {
let kindFiltered: SymbolDefinition[];
if (callForm === 'constructor') {
// For constructor calls, prefer Constructor > Class/Struct/Record > callable fallback
const constructors = candidates.filter(c => c.type === 'Constructor');
if (constructors.length > 0) {
kindFiltered = constructors;
@@ -296,19 +389,56 @@ const filterCallableCandidates = (
const toResolveResult = (
definition: SymbolDefinition,
tier: ResolutionTier,
): ResolveResult => {
if (tier === 'same-file') {
return { nodeId: definition.nodeId, confidence: 0.95, reason: 'same-file' };
): ResolveResult => ({
nodeId: definition.nodeId,
confidence: TIER_CONFIDENCE[tier],
reason: tier === 'same-file' ? 'same-file' : tier === 'import-scoped' ? 'import-resolved' : 'global',
});
/**
* Resolve a chain of intermediate method calls to find the receiver type for a
* final member call. Called when the receiver of a call is itself a call
* expression (e.g. `svc.getUser().save()`).
*
* @param chainNames Ordered list of method names from outermost to innermost
* intermediate call (e.g. ['getUser'] for `svc.getUser().save()`).
* @param baseReceiverTypeName The already-resolved type of the base receiver
* (e.g. 'UserService' for `svc`), or undefined.
* @param currentFile The file path for resolution context.
* @param ctx The resolution context for symbol lookup.
* @returns The type name of the final intermediate call's return type, or undefined
* if resolution fails at any step.
*/
function resolveChainedReceiver(
chainNames: string[],
baseReceiverTypeName: string | undefined,
currentFile: string,
ctx: ResolutionContext,
): string | undefined {
let currentType = baseReceiverTypeName;
for (const name of chainNames) {
const resolved = resolveCallTarget(
{ calledName: name, callForm: 'member', receiverTypeName: currentType },
currentFile,
ctx,
);
if (!resolved) return undefined;
const candidates = ctx.symbols.lookupFuzzy(name);
const symDef = candidates.find(c => c.nodeId === resolved.nodeId);
if (!symDef?.returnType) return undefined;
const returnTypeName = extractReturnTypeName(symDef.returnType);
if (!returnTypeName) return undefined;
currentType = returnTypeName;
}
if (tier === 'import-scoped') {
return { nodeId: definition.nodeId, confidence: 0.9, reason: 'import-resolved' };
}
return { nodeId: definition.nodeId, confidence: 0.5, reason: 'unique-global' };
};
return currentType;
}
/**
* Resolve a function call to its target node ID using priority strategy:
* A. Narrow candidates by scope tier (same-file, import-scoped, unique-global)
* A. Narrow candidates by scope tier via ctx.resolve()
* B. Filter to callable symbol kinds (constructor-aware when callForm is set)
* C. Apply arity filtering when parameter metadata is available
* D. Apply receiver-type filtering for member calls with typed receivers
@@ -318,29 +448,48 @@ const toResolveResult = (
const resolveCallTarget = (
call: Pick<ExtractedCall, 'calledName' | 'argCount' | 'callForm' | 'receiverTypeName'>,
currentFile: string,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
ctx: ResolutionContext,
): ResolveResult | null => {
const tiered = collectTieredCandidates(call.calledName, currentFile, symbolTable, importMap, packageMap, namedImportMap);
const tiered = ctx.resolve(call.calledName, currentFile);
if (!tiered) return null;
const filteredCandidates = filterCallableCandidates(tiered.candidates, call.argCount, call.callForm);
// D. Receiver-type filtering: for member calls with a known receiver type,
// filter candidates by ownerId matching the resolved type's nodeId
if (call.callForm === 'member' && call.receiverTypeName && filteredCandidates.length > 1) {
const typeDefs = symbolTable.lookupFuzzy(call.receiverTypeName);
if (typeDefs.length > 0) {
const typeNodeIds = new Set(typeDefs.map(d => d.nodeId));
const ownerFiltered = filteredCandidates.filter(c => c.ownerId && typeNodeIds.has(c.ownerId));
// resolve the type through the same tiered import infrastructure, then
// filter method candidates to the type's defining file. Fall back to
// fuzzy ownerId matching only when file-based narrowing is inconclusive.
//
// Applied regardless of candidate count — the sole same-file candidate may
// belong to the wrong class (e.g. super.save() should hit the parent's save,
// not the child's own save method in the same file).
if (call.callForm === 'member' && call.receiverTypeName) {
// D1. Resolve the receiver type
const typeResolved = ctx.resolve(call.receiverTypeName, currentFile);
if (typeResolved && typeResolved.candidates.length > 0) {
const typeNodeIds = new Set(typeResolved.candidates.map(d => d.nodeId));
const typeFiles = new Set(typeResolved.candidates.map(d => d.filePath));
// D2. Widen candidates: same-file tier may miss the parent's method when
// it lives in another file. Query the symbol table directly for all
// global methods with this name, then apply arity/kind filtering.
const methodPool = filteredCandidates.length <= 1
? filterCallableCandidates(ctx.symbols.lookupFuzzy(call.calledName), call.argCount, call.callForm)
: filteredCandidates;
// D3. File-based: prefer candidates whose filePath matches the resolved type's file
const fileFiltered = methodPool.filter(c => typeFiles.has(c.filePath));
if (fileFiltered.length === 1) {
return toResolveResult(fileFiltered[0], tiered.tier);
}
// D4. ownerId fallback: narrow by ownerId matching the type's nodeId
const pool = fileFiltered.length > 0 ? fileFiltered : methodPool;
const ownerFiltered = pool.filter(c => c.ownerId && typeNodeIds.has(c.ownerId));
if (ownerFiltered.length === 1) {
return toResolveResult(ownerFiltered[0], tiered.tier);
}
// If receiver filtering narrows to 0, fall through to name-only resolution
// If still 2+, refuse (don't guess)
if (ownerFiltered.length > 1) return null;
if (fileFiltered.length > 1 || ownerFiltered.length > 1) return null;
}
}
@@ -349,62 +498,323 @@ const resolveCallTarget = (
return toResolveResult(filteredCandidates[0], tiered.tier);
};
// ── Return type text helpers ─────────────────────────────────────────────
// extractSimpleTypeName works on AST nodes; this operates on raw return-type
// text already stored in SymbolDefinition (e.g. "User", "Promise<User>",
// "User | null", "*User"). Extracts the base user-defined type name.
/** Primitive / built-in types that should NOT produce a receiver binding. */
const PRIMITIVE_TYPES = new Set([
'string', 'number', 'boolean', 'void', 'int', 'float', 'double', 'long',
'short', 'byte', 'char', 'bool', 'str', 'i8', 'i16', 'i32', 'i64',
'u8', 'u16', 'u32', 'u64', 'f32', 'f64', 'usize', 'isize',
'undefined', 'null', 'None', 'nil',
]);
/**
* Extract a simple type name from raw return-type text.
* Handles common patterns:
* "User" → "User"
* "Promise<User>" → "User" (unwrap wrapper generics)
* "Option<User>" → "User"
* "Result<User, Error>" → "User" (first type arg)
* "User | null" → "User" (strip nullable union)
* "User?" → "User" (strip nullable suffix)
* "*User" → "User" (Go pointer)
* "&User" → "User" (Rust reference)
* Returns undefined for complex types or primitives.
*/
const WRAPPER_GENERICS = new Set([
'Promise', 'Observable', 'Future', 'CompletableFuture', 'Task', 'ValueTask', // async wrappers
'Option', 'Some', 'Optional', 'Maybe', // nullable wrappers
'Result', 'Either', // result wrappers
// Rust smart pointers (Deref to inner type)
'Rc', 'Arc', 'Weak', // pointer types
'MutexGuard', 'RwLockReadGuard', 'RwLockWriteGuard', // guard types
'Ref', 'RefMut', // RefCell guards
'Cow', // copy-on-write
// Containers (List, Array, Vec, Set, etc.) are intentionally excluded —
// methods are called on the container, not the element type.
// Non-wrapper generics return the base type (e.g., List) via the else branch.
]);
/**
* Extracts the first type argument from a comma-separated generic argument string,
* respecting nested angle brackets. For example:
* "Result<User, Error>" → "Result<User, Error>" (no top-level comma)
* "User, Error" → "User"
* "Map<K, V>, string" → "Map<K, V>"
*/
function extractFirstGenericArg(args: string): string {
let depth = 0;
for (let i = 0; i < args.length; i++) {
if (args[i] === '<') depth++;
else if (args[i] === '>') depth--;
else if (args[i] === ',' && depth === 0) return args.slice(0, i).trim();
}
return args.trim();
}
/**
* Extract the first non-lifetime type argument from a generic argument string.
* Skips Rust lifetime parameters (e.g., `'a`, `'_`) to find the actual type.
* "'_, User" → "User"
* "'a, User" → "User"
* "User, Error" → "User" (no lifetime — delegates to extractFirstGenericArg)
*/
function extractFirstTypeArg(args: string): string {
let remaining = args;
while (remaining) {
const first = extractFirstGenericArg(remaining);
if (!first.startsWith("'")) return first;
// Skip past this lifetime arg + the comma separator
const commaIdx = remaining.indexOf(',', first.length);
if (commaIdx < 0) return first; // only lifetimes — fall through
remaining = remaining.slice(commaIdx + 1).trim();
}
return args.trim();
}
const MAX_RETURN_TYPE_INPUT_LENGTH = 2048;
const MAX_RETURN_TYPE_LENGTH = 512;
export const extractReturnTypeName = (raw: string, depth = 0): string | undefined => {
if (depth > 10) return undefined;
if (raw.length > MAX_RETURN_TYPE_INPUT_LENGTH) return undefined;
let text = raw.trim();
if (!text) return undefined;
// Strip pointer/reference prefixes: *User, &User, &mut User
text = text.replace(/^[&*]+\s*(mut\s+)?/, '');
// Strip nullable suffix: User?
text = text.replace(/\?$/, '');
// Handle union types: "User | null" → "User"
if (text.includes('|')) {
const parts = text.split('|').map(p => p.trim()).filter(p =>
p !== 'null' && p !== 'undefined' && p !== 'void' && p !== 'None' && p !== 'nil'
);
if (parts.length === 1) text = parts[0];
else return undefined; // genuine union — too complex
}
// Handle generics: Promise<User> → unwrap if wrapper, else take base
const genericMatch = text.match(/^(\w+)\s*<(.+)>$/);
if (genericMatch) {
const [, base, args] = genericMatch;
if (WRAPPER_GENERICS.has(base)) {
// Take the first non-lifetime type argument, using bracket-balanced splitting
// so that nested generics like Result<User, Error> are not split at the inner
// comma. Lifetime parameters (Rust 'a, '_) are skipped.
const firstArg = extractFirstTypeArg(args);
return extractReturnTypeName(firstArg, depth + 1);
}
// Non-wrapper generic: return the base type (e.g., Map<K,V> → Map)
return PRIMITIVE_TYPES.has(base.toLowerCase()) ? undefined : base;
}
// Bare wrapper type without generic argument (e.g. Task, Promise, Option)
// should not produce a binding — these are meaningless without a type parameter
if (WRAPPER_GENERICS.has(text)) return undefined;
// Handle qualified names: models.User → User, Models::User → User, \App\Models\User → User
if (text.includes('::') || text.includes('.') || text.includes('\\')) {
text = text.split(/::|[.\\]/).pop()!;
}
// Final check: skip primitives
if (PRIMITIVE_TYPES.has(text) || PRIMITIVE_TYPES.has(text.toLowerCase())) return undefined;
// Must start with uppercase (class/type convention) or be a valid identifier
if (!/^[A-Z_]\w*$/.test(text)) return undefined;
// If the final extracted type name is too long, reject it
if (text.length > MAX_RETURN_TYPE_LENGTH) return undefined;
return text;
};
// ── Scope key helpers ────────────────────────────────────────────────────
// Scope keys use the format "funcName@startIndex" (produced by type-env.ts).
// Source IDs use "Label:filepath:funcName" (produced by parse-worker.ts).
// NUL (\0) is used as a composite-key separator because it cannot appear
// in source-code identifiers, preventing ambiguous concatenation.
//
// receiverKey stores the FULL scope (funcName@startIndex) to prevent
// collisions between overloaded methods with the same name in different
// classes (e.g. User.save@100 and Repo.save@200 are distinct keys).
// Lookup uses a secondary funcName-only index built in lookupReceiverType.
/** Extract the function name from a scope key ("funcName@startIndex" → "funcName"). */
const extractFuncNameFromScope = (scope: string): string =>
scope.slice(0, scope.indexOf('@'));
/** Extract the trailing function name from a sourceId ("Function:filepath:funcName" → "funcName"). */
const extractFuncNameFromSourceId = (sourceId: string): string => {
const lastColon = sourceId.lastIndexOf(':');
return lastColon >= 0 ? sourceId.slice(lastColon + 1) : '';
};
/**
* Build a composite key for receiver type storage.
* Uses the full scope string (e.g. "save@100") to distinguish overloaded
* methods with the same name in different classes.
*/
const receiverKey = (scope: string, varName: string): string =>
`${scope}\0${varName}`;
/**
* Look up a receiver type from a verified receiver map.
* The map is keyed by `scope\0varName` (full scope with @startIndex).
* Since the lookup side only has `funcName` (no startIndex), we scan for
* all entries whose key starts with `funcName@` and has the matching varName.
* If exactly one unique type is found, return it. If multiple distinct types
* exist (true overload collision), return undefined (refuse to guess).
* Falls back to the file-level scope key `\0varName` (empty funcName).
*/
const lookupReceiverType = (
map: Map<string, string>,
funcName: string,
varName: string,
): string | undefined => {
// Fast path: file-level scope (empty funcName — used as fallback)
const fileLevelKey = receiverKey('', varName);
const prefix = `${funcName}@`;
const suffix = `\0${varName}`;
let found: string | undefined;
let ambiguous = false;
for (const [key, value] of map) {
if (key === fileLevelKey) continue; // handled separately below
if (key.startsWith(prefix) && key.endsWith(suffix)) {
// Verify the key is exactly "funcName@<digits>\0varName" with no extra chars.
// The part between prefix and suffix should be the startIndex (digits only),
// but we accept any non-empty segment to be forward-compatible.
const middle = key.slice(prefix.length, key.length - suffix.length);
if (middle.length === 0) continue; // malformed key — skip
if (found === undefined) {
found = value;
} else if (found !== value) {
ambiguous = true;
break;
}
}
}
if (!ambiguous && found !== undefined) return found;
// Fallback: file-level scope (bindings outside any function)
return map.get(fileLevelKey);
};
/**
* Fast path: resolve pre-extracted call sites from workers.
* No AST parsing — workers already extracted calledName + sourceId.
* This function only does symbol table lookups + graph mutations.
*/
export const processCallsFromExtracted = async (
graph: KnowledgeGraph,
extractedCalls: ExtractedCall[],
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
ctx: ResolutionContext,
onProgress?: (current: number, total: number) => void,
namedImportMap?: NamedImportMap,
constructorBindings?: FileConstructorBindings[],
) => {
// Group by file for progress reporting
// Scope-aware receiver types: keyed by filePath → "funcName\0varName" → typeName.
// The scope dimension prevents collisions when two functions in the same file
// have same-named locals pointing to different constructor types.
const fileReceiverTypes = new Map<string, Map<string, string>>();
if (constructorBindings) {
for (const { filePath, bindings } of constructorBindings) {
const verified = verifyConstructorBindings(bindings, filePath, ctx, graph);
if (verified.size > 0) {
fileReceiverTypes.set(filePath, verified);
}
}
}
const byFile = new Map<string, ExtractedCall[]>();
for (const call of extractedCalls) {
let list = byFile.get(call.filePath);
if (!list) {
list = [];
byFile.set(call.filePath, list);
}
if (!list) { list = []; byFile.set(call.filePath, list); }
list.push(call);
}
const totalFiles = byFile.size;
let filesProcessed = 0;
for (const [_filePath, calls] of byFile) {
for (const [filePath, calls] of byFile) {
filesProcessed++;
if (filesProcessed % 100 === 0) {
onProgress?.(filesProcessed, totalFiles);
await yieldToEventLoop();
}
ctx.enableCache(filePath);
const receiverMap = fileReceiverTypes.get(filePath);
for (const call of calls) {
const resolved = resolveCallTarget(
call,
call.filePath,
symbolTable,
importMap,
packageMap,
namedImportMap,
);
let effectiveCall = call;
// Step 1: resolve receiver type from constructor bindings
if (!call.receiverTypeName && call.receiverName && receiverMap) {
const callFuncName = extractFuncNameFromSourceId(call.sourceId);
const resolvedType = lookupReceiverType(receiverMap, callFuncName, call.receiverName);
if (resolvedType) {
effectiveCall = { ...call, receiverTypeName: resolvedType };
}
}
// Step 1b: class-as-receiver for static method calls (e.g. UserService.find_user())
if (!effectiveCall.receiverTypeName && effectiveCall.receiverName && effectiveCall.callForm === 'member') {
const typeResolved = ctx.resolve(effectiveCall.receiverName, effectiveCall.filePath);
if (typeResolved && typeResolved.candidates.some(
d => d.type === 'Class' || d.type === 'Interface' || d.type === 'Struct' || d.type === 'Enum',
)) {
effectiveCall = { ...effectiveCall, receiverTypeName: effectiveCall.receiverName };
}
}
// Step 2: if the call has a receiver call chain (e.g. svc.getUser().save()),
// resolve the chain to determine the final receiver type.
// This runs whenever receiverCallChain is present — even when Step 1 set a
// receiverTypeName, that type is the BASE receiver (e.g. UserService for svc),
// and the chain must be walked to produce the FINAL receiver (e.g. User from
// getUser() : User).
if (effectiveCall.receiverCallChain?.length) {
// Step 1 may have resolved the base receiver type (e.g. svc → UserService).
// Use it as the starting point for chain resolution.
let baseType = effectiveCall.receiverTypeName;
// If Step 1 didn't resolve it, try the receiver map directly.
if (!baseType && effectiveCall.receiverName && receiverMap) {
const callFuncName = extractFuncNameFromSourceId(effectiveCall.sourceId);
baseType = lookupReceiverType(receiverMap, callFuncName, effectiveCall.receiverName);
}
const chainedType = resolveChainedReceiver(
effectiveCall.receiverCallChain,
baseType,
effectiveCall.filePath,
ctx,
);
if (chainedType) {
effectiveCall = { ...effectiveCall, receiverTypeName: chainedType };
}
}
const resolved = resolveCallTarget(effectiveCall, effectiveCall.filePath, ctx);
if (!resolved) continue;
const relId = generateId('CALLS', `${call.sourceId}:${call.calledName}->${resolved.nodeId}`);
const relId = generateId('CALLS', `${effectiveCall.sourceId}:${effectiveCall.calledName}->${resolved.nodeId}`);
graph.addRelationship({
id: relId,
sourceId: call.sourceId,
sourceId: effectiveCall.sourceId,
targetId: resolved.nodeId,
type: 'CALLS',
confidence: resolved.confidence,
reason: resolved.reason,
});
}
ctx.clearCache();
}
onProgress?.(totalFiles, totalFiles);
@@ -416,10 +826,8 @@ export const processCallsFromExtracted = async (
export const processRoutesFromExtracted = async (
graph: KnowledgeGraph,
extractedRoutes: ExtractedRoute[],
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void
ctx: ResolutionContext,
onProgress?: (current: number, total: number) => void,
) => {
for (let i = 0; i < extractedRoutes.length; i++) {
const route = extractedRoutes[i];
@@ -430,23 +838,18 @@ export const processRoutesFromExtracted = async (
if (!route.controllerName || !route.methodName) continue;
// Resolve controller class using shared resolver (Tier 1: same file,
// Tier 2: import-scoped, Tier 3: unique global).
const resolution = resolveSymbolInternal(route.controllerName, route.filePath, symbolTable, importMap, packageMap);
if (!resolution) continue;
const controllerResolved = ctx.resolve(route.controllerName, route.filePath);
if (!controllerResolved || controllerResolved.candidates.length === 0) continue;
if (controllerResolved.tier === 'global' && controllerResolved.candidates.length > 1) continue;
const controllerDef = resolution.definition;
// Derive confidence from the resolution tier
const confidence = resolution.tier === 'same-file' ? 0.95
: resolution.tier === 'import-scoped' ? 0.9
: 0.7;
const controllerDef = controllerResolved.candidates[0];
const confidence = TIER_CONFIDENCE[controllerResolved.tier];
// Find the method on the controller
const methodId = symbolTable.lookupExact(controllerDef.filePath, route.methodName);
const methodResolved = ctx.resolve(route.methodName, controllerDef.filePath);
const methodId = methodResolved?.tier === 'same-file' ? methodResolved.candidates[0]?.nodeId : undefined;
const sourceId = generateId('File', route.filePath);
if (!methodId) {
// Construct method ID manually
const guessedId = generateId('Method', `${controllerDef.filePath}:${route.methodName}`);
const relId = generateId('CALLS', `${sourceId}:route->${guessedId}`);
graph.addRelationship({
@@ -473,22 +876,3 @@ export const processRoutesFromExtracted = async (
onProgress?.(extractedRoutes.length, extractedRoutes.length);
};
/**
* Follow re-export chains through NamedImportMap for call candidate collection.
* Delegates chain-walking to the shared walkBindingChain utility, then
* applies call-processor semantics: any number of matches accepted.
*/
const resolveNamedBindingChainForCandidates = (
calledName: string,
currentFile: string,
symbolTable: SymbolTable,
namedImportMap: NamedImportMap,
allDefs: SymbolDefinition[],
): TieredCandidates | null => {
const defs = walkBindingChain(calledName, currentFile, symbolTable, namedImportMap, allDefs);
if (defs && defs.length > 0) {
return { candidates: defs, tier: 'import-scoped' };
}
return null;
};
+149
View File
@@ -0,0 +1,149 @@
/**
* Shared Ruby call routing logic.
*
* Ruby expresses imports, heritage (mixins), and property definitions as
* method calls rather than syntax-level constructs. This module provides a
* routing function used by the CLI call-processor, CLI parse-worker, and
* the web call-processor so that the classification logic lives in one place.
*
* NOTE: This file is intentionally duplicated in gitnexus-web/ because the
* two packages have separate build targets (Node native vs WASM/browser).
* Keep both copies in sync until a shared package is introduced.
*/
import { SupportedLanguages } from '../../config/supported-languages.js';
// ── Call routing dispatch table ─────────────────────────────────────────────
/** null = this call was not routed; fall through to default call handling */
export type CallRoutingResult = RubyCallRouting | null;
export type CallRouter = (
calledName: string,
callNode: any,
) => CallRoutingResult;
/** No-op router: returns null for every call (passthrough to normal processing) */
const noRouting: CallRouter = () => null;
/** Per-language call routing. noRouting = no special routing (normal call processing) */
export const callRouters: Record<SupportedLanguages, CallRouter> = {
[SupportedLanguages.JavaScript]: noRouting,
[SupportedLanguages.TypeScript]: noRouting,
[SupportedLanguages.Python]: noRouting,
[SupportedLanguages.Java]: noRouting,
[SupportedLanguages.Kotlin]: noRouting,
[SupportedLanguages.Go]: noRouting,
[SupportedLanguages.Rust]: noRouting,
[SupportedLanguages.CSharp]: noRouting,
[SupportedLanguages.PHP]: noRouting,
[SupportedLanguages.Swift]: noRouting,
[SupportedLanguages.CPlusPlus]: noRouting,
[SupportedLanguages.C]: noRouting,
[SupportedLanguages.Ruby]: routeRubyCall,
};
// ── Result types ────────────────────────────────────────────────────────────
export type RubyCallRouting =
| { kind: 'import'; importPath: string; isRelative: boolean }
| { kind: 'heritage'; items: RubyHeritageItem[] }
| { kind: 'properties'; items: RubyPropertyItem[] }
| { kind: 'call' }
| { kind: 'skip' };
export interface RubyHeritageItem {
enclosingClass: string;
mixinName: string;
heritageKind: 'include' | 'extend' | 'prepend';
}
export type RubyAccessorType = 'attr_accessor' | 'attr_reader' | 'attr_writer';
export interface RubyPropertyItem {
propName: string;
accessorType: RubyAccessorType;
startLine: number;
endLine: number;
}
// ── Pre-allocated singletons for common return values ────────────────────────
const CALL_RESULT: RubyCallRouting = { kind: 'call' };
const SKIP_RESULT: RubyCallRouting = { kind: 'skip' };
/** Max depth for parent-walking loops to prevent pathological AST traversals */
const MAX_PARENT_DEPTH = 50;
// ── Routing function ────────────────────────────────────────────────────────
/**
* Classify a Ruby call node and extract its semantic payload.
*
* @param calledName - The method name (e.g. 'require', 'include', 'attr_accessor')
* @param callNode - The tree-sitter `call` AST node
* @returns A discriminated union describing the call's semantic role
*/
export function routeRubyCall(calledName: string, callNode: any): RubyCallRouting {
// ── require / require_relative → import ─────────────────────────────────
if (calledName === 'require' || calledName === 'require_relative') {
const argList = callNode.childForFieldName?.('arguments');
const stringNode = argList?.children?.find((c: any) => c.type === 'string');
const contentNode = stringNode?.children?.find((c: any) => c.type === 'string_content');
if (!contentNode) return SKIP_RESULT;
let importPath: string = contentNode.text;
// Validate: reject null bytes, control chars, excessively long paths
if (!importPath || importPath.length > 1024 || /[\x00-\x1f]/.test(importPath)) {
return SKIP_RESULT;
}
const isRelative = calledName === 'require_relative';
if (isRelative && !importPath.startsWith('.')) {
importPath = './' + importPath;
}
return { kind: 'import', importPath, isRelative };
}
// ── include / extend / prepend → heritage (mixin) ──────────────────────
if (calledName === 'include' || calledName === 'extend' || calledName === 'prepend') {
let enclosingClass: string | null = null;
let current = callNode.parent;
let depth = 0;
while (current && ++depth <= MAX_PARENT_DEPTH) {
if (current.type === 'class' || current.type === 'module') {
const nameNode = current.childForFieldName?.('name');
if (nameNode) { enclosingClass = nameNode.text; break; }
}
current = current.parent;
}
if (!enclosingClass) return SKIP_RESULT;
const items: RubyHeritageItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of (argList?.children ?? [])) {
if (arg.type === 'constant' || arg.type === 'scope_resolution') {
items.push({ enclosingClass, mixinName: arg.text, heritageKind: calledName as 'include' | 'extend' | 'prepend' });
}
}
return items.length > 0 ? { kind: 'heritage', items } : SKIP_RESULT;
}
// ── attr_accessor / attr_reader / attr_writer → property definitions ───
if (calledName === 'attr_accessor' || calledName === 'attr_reader' || calledName === 'attr_writer') {
const items: RubyPropertyItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of (argList?.children ?? [])) {
if (arg.type === 'simple_symbol') {
items.push({
propName: arg.text.startsWith(':') ? arg.text.slice(1) : arg.text,
accessorType: calledName as RubyAccessorType,
startLine: arg.startPosition.row,
endLine: arg.endPosition.row,
});
}
}
return items.length > 0 ? { kind: 'properties', items } : SKIP_RESULT;
}
// ── Everything else → regular call ─────────────────────────────────────
return CALL_RESULT;
}
@@ -14,7 +14,7 @@ import { detectFrameworkFromPath } from './framework-detection.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
// ============================================================================
// NAME PATTERNS - All 9 supported languages
// NAME PATTERNS - All 11 supported languages
// ============================================================================
/**
@@ -191,6 +191,13 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
/^save$/, // Repository::save()
/^delete$/, // Repository::delete()
],
// Ruby
[SupportedLanguages.Ruby]: [
/^call$/, // Service objects (MyService.call)
/^perform$/, // Background jobs (Sidekiq, ActiveJob)
/^execute$/, // Command pattern
],
};
/** Pre-computed merged patterns (universal + language-specific) to avoid per-call array allocation. */
@@ -361,7 +368,12 @@ export function isTestFile(filePath: string): boolean {
p.endsWith('test.php') ||
p.endsWith('spec.php') ||
p.includes('/tests/feature/') ||
p.includes('/tests/unit/')
p.includes('/tests/unit/') ||
// Ruby test patterns
p.endsWith('_spec.rb') ||
p.endsWith('_test.rb') ||
p.includes('/spec/') ||
p.includes('/test/fixtures/')
);
}
@@ -222,6 +222,7 @@ const exportCheckers = {
[SupportedLanguages.CPlusPlus]: cCppExportChecker,
[SupportedLanguages.PHP]: phpExportChecker,
[SupportedLanguages.Swift]: swiftExportChecker,
[SupportedLanguages.Ruby]: (_node, _name) => true,
} satisfies Record<SupportedLanguages, ExportChecker>;
// ============================================================================
@@ -1,7 +1,7 @@
import fs from 'fs/promises';
import path from 'path';
import { glob } from 'glob';
import { shouldIgnorePath } from '../../config/ignore-service.js';
import { createIgnoreFilter } from '../../config/ignore-service.js';
export interface FileEntry {
path: string;
@@ -32,13 +32,14 @@ export const walkRepositoryPaths = async (
repoPath: string,
onProgress?: (current: number, total: number, filePath: string) => void
): Promise<ScannedFile[]> => {
const files = await glob('**/*', {
const ignoreFilter = await createIgnoreFilter(repoPath);
const filtered = await glob('**/*', {
cwd: repoPath,
nodir: true,
dot: false,
ignore: ignoreFilter,
});
const filtered = files.filter(file => !shouldIgnorePath(file));
const entries: ScannedFile[] = [];
let processed = 0;
let skippedLarge = 0;
@@ -330,6 +330,18 @@ export function detectFrameworkFromPath(filePath: string): FrameworkHint | null
return { framework: 'laravel', entryPointMultiplier: 1.5, reason: 'laravel-repository' };
}
// ========== RUBY ==========
// Ruby: bin/ or exe/ (CLI entry points)
if ((p.includes('/bin/') || p.includes('/exe/')) && p.endsWith('.rb')) {
return { framework: 'ruby', entryPointMultiplier: 2.5, reason: 'ruby-executable' };
}
// Ruby: Rakefile or *.rake (task definitions)
if (p.endsWith('/rakefile') || p.endsWith('.rake')) {
return { framework: 'ruby', entryPointMultiplier: 1.5, reason: 'ruby-rake' };
}
// ========== SWIFT / iOS ==========
// iOS App entry points (highest priority)
@@ -16,7 +16,6 @@
import { KnowledgeGraph } from '../graph/types.js';
import { ASTCache } from './ast-cache.js';
import { SymbolTable, SymbolDefinition } from './symbol-table.js';
import Parser from 'tree-sitter';
import { isLanguageAvailable, loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
@@ -25,8 +24,7 @@ import { getLanguageFromFilename, isVerboseIngestionEnabled, yieldToEventLoop }
import { SupportedLanguages } from '../../config/supported-languages.js';
import { getTreeSitterBufferSize } from './constants.js';
import type { ExtractedHeritage } from './workers/parse-worker.js';
import { resolveSymbol } from './symbol-resolver.js';
import type { ImportMap, PackageMap } from './import-processor.js';
import type { ResolutionContext } from './resolution-context.js';
/** C#/Java convention: interfaces start with I followed by an uppercase letter */
const INTERFACE_NAME_RE = /^I[A-Z]/;
@@ -42,14 +40,12 @@ const INTERFACE_NAME_RE = /^I[A-Z]/;
const resolveExtendsType = (
parentName: string,
currentFilePath: string,
symbolTable: SymbolTable,
importMap: ImportMap,
ctx: ResolutionContext,
language: SupportedLanguages,
packageMap?: PackageMap,
): { type: 'EXTENDS' | 'IMPLEMENTS'; idPrefix: string } => {
const resolved = resolveSymbol(parentName, currentFilePath, symbolTable, importMap, packageMap);
if (resolved) {
const isInterface = resolved.type === 'Interface';
const resolved = ctx.resolve(parentName, currentFilePath);
if (resolved && resolved.candidates.length > 0) {
const isInterface = resolved.candidates[0].type === 'Interface';
return isInterface
? { type: 'IMPLEMENTS', idPrefix: 'Interface' }
: { type: 'EXTENDS', idPrefix: 'Class' };
@@ -66,14 +62,34 @@ const resolveExtendsType = (
return { type: 'EXTENDS', idPrefix: 'Class' };
};
/**
* Resolve a symbol ID for heritage, with fallback to generated ID.
* Uses ctx.resolve() → pick first candidate's nodeId → generate synthetic ID.
*/
const resolveHeritageId = (
name: string,
filePath: string,
ctx: ResolutionContext,
fallbackLabel: string,
fallbackKey?: string,
): string => {
const resolved = ctx.resolve(name, filePath);
if (resolved && resolved.candidates.length > 0) {
// For global with multiple candidates, refuse (a wrong edge is worse than no edge)
if (resolved.tier === 'global' && resolved.candidates.length > 1) {
return generateId(fallbackLabel, fallbackKey ?? name);
}
return resolved.candidates[0].nodeId;
}
return generateId(fallbackLabel, fallbackKey ?? name);
};
export const processHeritage = async (
graph: KnowledgeGraph,
files: { path: string; content: string }[],
astCache: ASTCache,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void
ctx: ResolutionContext,
onProgress?: (current: number, total: number) => void,
) => {
const parser = await loadParser();
const logSkipped = isVerboseIngestionEnabled();
@@ -102,8 +118,6 @@ export const processHeritage = async (
// 3. Get AST
let tree = astCache.get(file.path);
let wasReparsed = false;
if (!tree) {
// Use larger bufferSize for files > 32KB
try {
@@ -112,7 +126,6 @@ export const processHeritage = async (
// Skip files that can't be parsed
continue;
}
wasReparsed = true;
// Cache re-parsed tree for potential future use
astCache.set(file.path, tree);
}
@@ -148,14 +161,10 @@ export const processHeritage = async (
const className = captureMap['heritage.class'].text;
const parentClassName = captureMap['heritage.extends'].text;
const { type: relType, idPrefix } = resolveExtendsType(parentClassName, file.path, symbolTable, importMap, language, packageMap);
const { type: relType, idPrefix } = resolveExtendsType(parentClassName, file.path, ctx, language);
const childId = symbolTable.lookupExact(file.path, className) ||
resolveSymbol(className, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Class', `${file.path}:${className}`);
const parentId = resolveSymbol(parentClassName, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId(idPrefix, `${parentClassName}`);
const childId = resolveHeritageId(className, file.path, ctx, 'Class', `${file.path}:${className}`);
const parentId = resolveHeritageId(parentClassName, file.path, ctx, idPrefix);
if (childId && parentId && childId !== parentId) {
graph.addRelationship({
@@ -174,19 +183,12 @@ export const processHeritage = async (
const className = captureMap['heritage.class'].text;
const interfaceName = captureMap['heritage.implements'].text;
// Resolve class and interface IDs
const classId = symbolTable.lookupExact(file.path, className) ||
resolveSymbol(className, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Class', `${file.path}:${className}`);
const interfaceId = resolveSymbol(interfaceName, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Interface', `${interfaceName}`);
const classId = resolveHeritageId(className, file.path, ctx, 'Class', `${file.path}:${className}`);
const interfaceId = resolveHeritageId(interfaceName, file.path, ctx, 'Interface');
if (classId && interfaceId) {
const relId = generateId('IMPLEMENTS', `${classId}->${interfaceId}`);
graph.addRelationship({
id: relId,
id: generateId('IMPLEMENTS', `${classId}->${interfaceId}`),
sourceId: classId,
targetId: interfaceId,
type: 'IMPLEMENTS',
@@ -201,19 +203,12 @@ export const processHeritage = async (
const structName = captureMap['heritage.class'].text;
const traitName = captureMap['heritage.trait'].text;
// Resolve struct and trait IDs
const structId = symbolTable.lookupExact(file.path, structName) ||
resolveSymbol(structName, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Struct', `${file.path}:${structName}`);
const traitId = resolveSymbol(traitName, file.path, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Trait', `${traitName}`);
const structId = resolveHeritageId(structName, file.path, ctx, 'Struct', `${file.path}:${structName}`);
const traitId = resolveHeritageId(traitName, file.path, ctx, 'Trait');
if (structId && traitId) {
const relId = generateId('IMPLEMENTS', `${structId}->${traitId}`);
graph.addRelationship({
id: relId,
id: generateId('IMPLEMENTS', `${structId}->${traitId}`),
sourceId: structId,
targetId: traitId,
type: 'IMPLEMENTS',
@@ -243,10 +238,8 @@ export const processHeritage = async (
export const processHeritageFromExtracted = async (
graph: KnowledgeGraph,
extractedHeritage: ExtractedHeritage[],
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
onProgress?: (current: number, total: number) => void
ctx: ResolutionContext,
onProgress?: (current: number, total: number) => void,
) => {
const total = extractedHeritage.length;
@@ -261,14 +254,10 @@ export const processHeritageFromExtracted = async (
if (h.kind === 'extends') {
const fileLanguage = getLanguageFromFilename(h.filePath);
if (!fileLanguage) continue;
const { type: relType, idPrefix } = resolveExtendsType(h.parentName, h.filePath, symbolTable, importMap, fileLanguage, packageMap);
const { type: relType, idPrefix } = resolveExtendsType(h.parentName, h.filePath, ctx, fileLanguage);
const childId = symbolTable.lookupExact(h.filePath, h.className) ||
resolveSymbol(h.className, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Class', `${h.filePath}:${h.className}`);
const parentId = resolveSymbol(h.parentName, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId(idPrefix, `${h.parentName}`);
const childId = resolveHeritageId(h.className, h.filePath, ctx, 'Class', `${h.filePath}:${h.className}`);
const parentId = resolveHeritageId(h.parentName, h.filePath, ctx, idPrefix);
if (childId && parentId && childId !== parentId) {
graph.addRelationship({
@@ -281,12 +270,8 @@ export const processHeritageFromExtracted = async (
});
}
} else if (h.kind === 'implements') {
const classId = symbolTable.lookupExact(h.filePath, h.className) ||
resolveSymbol(h.className, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Class', `${h.filePath}:${h.className}`);
const interfaceId = resolveSymbol(h.parentName, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Interface', `${h.parentName}`);
const classId = resolveHeritageId(h.className, h.filePath, ctx, 'Class', `${h.filePath}:${h.className}`);
const interfaceId = resolveHeritageId(h.parentName, h.filePath, ctx, 'Interface');
if (classId && interfaceId) {
graph.addRelationship({
@@ -298,22 +283,18 @@ export const processHeritageFromExtracted = async (
reason: '',
});
}
} else if (h.kind === 'trait-impl') {
const structId = symbolTable.lookupExact(h.filePath, h.className) ||
resolveSymbol(h.className, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Struct', `${h.filePath}:${h.className}`);
const traitId = resolveSymbol(h.parentName, h.filePath, symbolTable, importMap, packageMap)?.nodeId ||
generateId('Trait', `${h.parentName}`);
} else if (h.kind === 'trait-impl' || h.kind === 'include' || h.kind === 'extend' || h.kind === 'prepend') {
const structId = resolveHeritageId(h.className, h.filePath, ctx, 'Struct', `${h.filePath}:${h.className}`);
const traitId = resolveHeritageId(h.parentName, h.filePath, ctx, 'Trait');
if (structId && traitId) {
graph.addRelationship({
id: generateId('IMPLEMENTS', `${structId}->${traitId}`),
id: generateId('IMPLEMENTS', `${structId}->${traitId}:${h.kind}`),
sourceId: structId,
targetId: traitId,
type: 'IMPLEMENTS',
confidence: 1.0,
reason: 'trait-impl',
reason: h.kind,
});
}
}
+35 -16
View File
@@ -30,7 +30,10 @@ import {
resolveCSharpNamespaceDir,
resolvePhpImport,
resolveRustImport,
resolveRubyImport,
} from './resolvers/index.js';
import { callRouters } from './call-routing.js';
import type { ResolutionContext } from './resolution-context.js';
import type {
SuffixIndex,
TsconfigPaths,
@@ -54,15 +57,11 @@ const isDev = process.env.NODE_ENV === 'development';
// Stores all files that a given file imports from
export type ImportMap = Map<string, Set<string>>;
export const createImportMap = (): ImportMap => new Map();
// Type: Map<FilePath, Set<PackageDirSuffix>>
// Stores Go package directory suffixes imported by a file (e.g., "/internal/auth/").
// Avoids expanding every Go package import into N individual ImportMap edges.
export type PackageMap = Map<string, Set<string>>;
export const createPackageMap = (): PackageMap => new Map();
// Type: Map<ImportingFilePath, Map<LocalName, {sourcePath, exportedName}>>
// Tracks which specific names a file imports from which sources (TS/Python only).
// Used to tighten Tier 2a resolution: `import { User } from './models'`
@@ -72,8 +71,6 @@ export const createPackageMap = (): PackageMap => new Map();
export interface NamedImportBinding { sourcePath: string; exportedName: string }
export type NamedImportMap = Map<string, Map<string, NamedImportBinding>>;
export const createNamedImportMap = (): NamedImportMap => new Map();
/**
* Check if a file path is directly inside a package directory identified by its suffix.
* Used by the symbol resolver for Go and C# directory-level import matching.
@@ -222,6 +219,12 @@ function resolveLanguageImport(
return null; // External framework (Foundation, UIKit, etc.)
}
// Ruby: require / require_relative
if (language === SupportedLanguages.Ruby) {
const resolved = resolveRubyImport(rawImportPath, normalizedFileList, allFileList, index);
return resolved ? { kind: 'files', files: [resolved] } : null;
}
// Rust: expand top-level grouped imports: use {crate::a, crate::b}
if (language === SupportedLanguages.Rust && rawImportPath.startsWith('{') && rawImportPath.endsWith('}')) {
const inner = rawImportPath.slice(1, -1);
@@ -301,13 +304,14 @@ export const processImports = async (
graph: KnowledgeGraph,
files: { path: string; content: string }[],
astCache: ASTCache,
importMap: ImportMap,
ctx: ResolutionContext,
onProgress?: (current: number, total: number) => void,
repoRoot?: string,
allPaths?: string[],
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
) => {
const importMap = ctx.importMap;
const packageMap = ctx.packageMap;
const namedImportMap = ctx.namedImportMap;
// Use allPaths (full repo) when available for cross-chunk resolution, else fall back to chunk files
const allFileList = allPaths ?? files.map(f => f.path);
const allFilePaths = new Set(allFileList);
@@ -333,7 +337,7 @@ export const processImports = async (
swiftPackageConfig: await loadSwiftPackageConfig(effectiveRoot),
csharpConfigs: await loadCSharpProjectConfig(effectiveRoot),
};
const ctx: ResolveCtx = { allFilePaths, allFileList, normalizedFileList, index, resolveCache };
const resolveCtx: ResolveCtx = { allFilePaths, allFileList, normalizedFileList, index, resolveCache };
// Helper: add an IMPORTS edge to the graph only (no ImportMap update)
const addImportGraphEdge = (filePath: string, resolvedPath: string) => {
@@ -440,10 +444,24 @@ export const processImports = async (
: sourceNode.text.replace(/['"<>]/g, '');
totalImportsFound++;
const result = resolveLanguageImport(file.path, rawImportPath, language, configs, ctx);
const result = resolveLanguageImport(file.path, rawImportPath, language, configs, resolveCtx);
const bindings = namedImportMap ? extractNamedBindings(captureMap['import'], language) : undefined;
applyImportResult(result, file.path, importMap, packageMap, addImportEdge, addImportGraphEdge, bindings, namedImportMap);
}
// ---- Language-specific call-as-import routing (Ruby require, etc.) ----
if (captureMap['call']) {
const callNameNode = captureMap['call.name'];
if (callNameNode) {
const callRouter = callRouters[language];
const routed = callRouter(callNameNode.text, captureMap['call']);
if (routed && routed.kind === 'import') {
totalImportsFound++;
const result = resolveLanguageImport(file.path, routed.importPath, language, configs, resolveCtx);
applyImportResult(result, file.path, importMap, packageMap, addImportEdge, addImportGraphEdge);
}
}
}
});
// Tree is now owned by the LRU cache — no manual delete needed
@@ -470,15 +488,16 @@ export const processImportsFromExtracted = async (
graph: KnowledgeGraph,
files: { path: string }[],
extractedImports: ExtractedImport[],
importMap: ImportMap,
ctx: ResolutionContext,
onProgress?: (current: number, total: number) => void,
repoRoot?: string,
prebuiltCtx?: ImportResolutionContext,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
) => {
const ctx = prebuiltCtx ?? buildImportResolutionContext(files.map(f => f.path));
const { allFilePaths, allFileList, normalizedFileList, suffixIndex: index, resolveCache } = ctx;
const importMap = ctx.importMap;
const packageMap = ctx.packageMap;
const namedImportMap = ctx.namedImportMap;
const importCtx = prebuiltCtx ?? buildImportResolutionContext(files.map(f => f.path));
const { allFilePaths, allFileList, normalizedFileList, suffixIndex: index, resolveCache } = importCtx;
let totalImportsFound = 0;
let totalImportsResolved = 0;
@@ -1,6 +1,6 @@
import { KnowledgeGraph, GraphNode, GraphRelationship } from '../graph/types.js';
import Parser from 'tree-sitter';
import { loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { loadParser, loadLanguage, isLanguageAvailable } from '../tree-sitter/parser-loader.js';
import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { generateId } from '../../lib/utils.js';
import { SymbolTable } from './symbol-table.js';
@@ -8,8 +8,9 @@ import { ASTCache } from './ast-cache.js';
import { getLanguageFromFilename, yieldToEventLoop, DEFINITION_CAPTURE_KEYS, getDefinitionNodeFromCaptures, findEnclosingClassId, extractMethodSignature } from './utils.js';
import { isNodeExported } from './export-detection.js';
import { detectFrameworkFromAST } from './framework-detection.js';
import { typeConfigs } from './type-extractors/index.js';
import { WorkerPool } from './workers/worker-pool.js';
import type { ParseWorkerResult, ParseWorkerInput, ExtractedImport, ExtractedCall, ExtractedHeritage, ExtractedRoute } from './workers/parse-worker.js';
import type { ParseWorkerResult, ParseWorkerInput, ExtractedImport, ExtractedCall, ExtractedHeritage, ExtractedRoute, FileConstructorBindings } from './workers/parse-worker.js';
import { getTreeSitterBufferSize, TREE_SITTER_MAX_BUFFER } from './constants.js';
export type FileProgressCallback = (current: number, total: number, filePath: string) => void;
@@ -19,6 +20,7 @@ export interface WorkerExtractedData {
calls: ExtractedCall[];
heritage: ExtractedHeritage[];
routes: ExtractedRoute[];
constructorBindings: FileConstructorBindings[];
}
// isNodeExported imported from ./export-detection.js (shared module)
@@ -44,7 +46,7 @@ const processParsingWithWorkers = async (
if (lang) parseableFiles.push({ path: file.path, content: file.content });
}
if (parseableFiles.length === 0) return { imports: [], calls: [], heritage: [], routes: [] };
if (parseableFiles.length === 0) return { imports: [], calls: [], heritage: [], routes: [], constructorBindings: [] };
const total = files.length;
@@ -61,6 +63,7 @@ const processParsingWithWorkers = async (
const allCalls: ExtractedCall[] = [];
const allHeritage: ExtractedHeritage[] = [];
const allRoutes: ExtractedRoute[] = [];
const allConstructorBindings: FileConstructorBindings[] = [];
for (const result of chunkResults) {
for (const node of result.nodes) {
graph.addNode({
@@ -77,6 +80,7 @@ const processParsingWithWorkers = async (
for (const sym of result.symbols) {
symbolTable.add(sym.filePath, sym.name, sym.nodeId, sym.type, {
parameterCount: sym.parameterCount,
returnType: sym.returnType,
ownerId: sym.ownerId,
});
}
@@ -85,11 +89,26 @@ const processParsingWithWorkers = async (
allCalls.push(...result.calls);
allHeritage.push(...result.heritage);
allRoutes.push(...result.routes);
allConstructorBindings.push(...result.constructorBindings);
}
// Merge and log skipped languages from workers
const skippedLanguages = new Map<string, number>();
for (const result of chunkResults) {
for (const [lang, count] of Object.entries(result.skippedLanguages)) {
skippedLanguages.set(lang, (skippedLanguages.get(lang) || 0) + count);
}
}
if (skippedLanguages.size > 0) {
const summary = Array.from(skippedLanguages.entries())
.map(([lang, count]) => `${lang}: ${count}`)
.join(', ');
console.warn(` Skipped unsupported languages: ${summary}`);
}
// Final progress
onFileProgress?.(total, total, 'done');
return { imports: allImports, calls: allCalls, heritage: allHeritage, routes: allRoutes };
return { imports: allImports, calls: allCalls, heritage: allHeritage, routes: allRoutes, constructorBindings: allConstructorBindings };
};
// ============================================================================
@@ -105,6 +124,7 @@ const processParsingSequential = async (
) => {
const parser = await loadParser();
const total = files.length;
const skippedLanguages = new Map<string, number>();
for (let i = 0; i < files.length; i++) {
const file = files[i];
@@ -117,13 +137,19 @@ const processParsingSequential = async (
if (!language) continue;
// Skip unsupported languages (e.g. Swift when tree-sitter-swift not installed)
if (!isLanguageAvailable(language)) {
skippedLanguages.set(language, (skippedLanguages.get(language) || 0) + 1);
continue;
}
// Skip files larger than the max tree-sitter buffer (32 MB)
if (file.content.length > TREE_SITTER_MAX_BUFFER) continue;
try {
await loadLanguage(language, file.path);
} catch {
continue; // parser unavailable — already warned in pipeline
continue; // parser unavailable — safety net
}
let tree;
@@ -211,6 +237,14 @@ const processParsingSequential = async (
? extractMethodSignature(definitionNode)
: undefined;
// Language-specific return type fallback (e.g. Ruby YARD @return [Type])
if (methodSig && !methodSig.returnType && definitionNode) {
const tc = typeConfigs[language as keyof typeof typeConfigs];
if (tc?.extractReturnType) {
methodSig.returnType = tc.extractReturnType(definitionNode);
}
}
const node: GraphNode = {
id: nodeId,
label: nodeLabel as any,
@@ -241,6 +275,7 @@ const processParsingSequential = async (
symbolTable.add(file.path, nodeName, nodeId, nodeLabel, {
parameterCount: methodSig?.parameterCount,
returnType: methodSig?.returnType,
ownerId: enclosingClassId ?? undefined,
});
@@ -272,6 +307,13 @@ const processParsingSequential = async (
}
});
}
if (skippedLanguages.size > 0) {
const summary = Array.from(skippedLanguages.entries())
.map(([lang, count]) => `${lang}: ${count}`)
.join(', ');
console.warn(` Skipped unsupported languages: ${summary}`);
}
};
// ============================================================================
+205 -148
View File
@@ -4,9 +4,6 @@ import { processParsing } from './parsing-processor.js';
import {
processImports,
processImportsFromExtracted,
createImportMap,
createPackageMap,
createNamedImportMap,
buildImportResolutionContext
} from './import-processor.js';
import { processCalls, processCallsFromExtracted, processRoutesFromExtracted } from './call-processor.js';
@@ -14,7 +11,7 @@ import { processHeritage, processHeritageFromExtracted } from './heritage-proces
import { computeMRO } from './mro-processor.js';
import { processCommunities } from './community-processor.js';
import { processProcesses } from './process-processor.js';
import { createSymbolTable } from './symbol-table.js';
import { createResolutionContext } from './resolution-context.js';
import { createASTCache } from './ast-cache.js';
import { PipelineProgress, PipelineResult } from '../../types/pipeline.js';
import { walkRepositoryPaths, readFileContents } from './filesystem-walker.js';
@@ -36,20 +33,24 @@ const CHUNK_BYTE_BUDGET = 20 * 1024 * 1024; // 20MB
/** Max AST trees to keep in LRU cache */
const AST_CACHE_CAP = 50;
export interface PipelineOptions {
/** Skip MRO, community detection, and process extraction for faster test runs. */
skipGraphPhases?: boolean;
}
export const runPipelineFromRepo = async (
repoPath: string,
onProgress: (progress: PipelineProgress) => void
onProgress: (progress: PipelineProgress) => void,
options?: PipelineOptions,
): Promise<PipelineResult> => {
const graph = createKnowledgeGraph();
const symbolTable = createSymbolTable();
const ctx = createResolutionContext();
const symbolTable = ctx.symbols;
let astCache = createASTCache(AST_CACHE_CAP);
const importMap = createImportMap();
const packageMap = createPackageMap();
const namedImportMap = createNamedImportMap();
const cleanup = () => {
astCache.clear();
symbolTable.clear();
ctx.clear();
};
try {
@@ -159,22 +160,29 @@ export const runPipelineFromRepo = async (
stats: { filesProcessed: 0, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
});
// Don't spawn workers for tiny repos — overhead exceeds benefit
const MIN_FILES_FOR_WORKERS = 15;
const MIN_BYTES_FOR_WORKERS = 512 * 1024;
const totalBytes = parseableScanned.reduce((s, f) => s + f.size, 0);
// Create worker pool once, reuse across chunks
let workerPool: WorkerPool | undefined;
try {
let workerUrl = new URL('./workers/parse-worker.js', import.meta.url);
// When running under vitest, import.meta.url points to src/ where no .js exists.
// Fall back to the compiled dist/ worker so the pool can spawn real worker threads.
const thisDir = fileURLToPath(new URL('.', import.meta.url));
if (!fs.existsSync(fileURLToPath(workerUrl))) {
const distWorker = path.resolve(thisDir, '..', '..', '..', 'dist', 'core', 'ingestion', 'workers', 'parse-worker.js');
if (fs.existsSync(distWorker)) {
workerUrl = pathToFileURL(distWorker) as URL;
if (totalParseable >= MIN_FILES_FOR_WORKERS || totalBytes >= MIN_BYTES_FOR_WORKERS) {
try {
let workerUrl = new URL('./workers/parse-worker.js', import.meta.url);
// When running under vitest, import.meta.url points to src/ where no .js exists.
// Fall back to the compiled dist/ worker so the pool can spawn real worker threads.
const thisDir = fileURLToPath(new URL('.', import.meta.url));
if (!fs.existsSync(fileURLToPath(workerUrl))) {
const distWorker = path.resolve(thisDir, '..', '..', '..', 'dist', 'core', 'ingestion', 'workers', 'parse-worker.js');
if (fs.existsSync(distWorker)) {
workerUrl = pathToFileURL(distWorker) as URL;
}
}
workerPool = createWorkerPool(workerUrl);
} catch (err) {
if (isDev) console.warn('Worker pool creation failed, using sequential fallback:', (err as Error).message);
}
workerPool = createWorkerPool(workerUrl);
} catch (err) {
if (isDev) console.warn('Worker pool creation failed, using sequential fallback:', (err as Error).message);
}
let filesParsedSoFar = 0;
@@ -221,38 +229,69 @@ export const runPipelineFromRepo = async (
workerPool,
);
const chunkBasePercent = 20 + ((filesParsedSoFar / totalParseable) * 62);
if (chunkWorkerData) {
// Imports
await processImportsFromExtracted(graph, allPathObjects, chunkWorkerData.imports, importMap, undefined, repoPath, importCtx, packageMap, namedImportMap);
await processImportsFromExtracted(graph, allPathObjects, chunkWorkerData.imports, ctx, (current, total) => {
onProgress({
phase: 'parsing',
percent: Math.round(chunkBasePercent),
message: `Resolving imports (chunk ${chunkIdx + 1}/${numChunks})...`,
detail: `${current}/${total} files`,
stats: { filesProcessed: filesParsedSoFar, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
});
}, repoPath, importCtx);
// Calls + Heritage + Routes — resolve in parallel (no shared mutable state between them)
// This is safe because each writes disjoint relationship types into idempotent id-keyed Maps,
// and the single-threaded event loop prevents races between synchronous addRelationship calls.
await Promise.all([
processCallsFromExtracted(
graph,
chunkWorkerData.calls,
symbolTable, importMap,
packageMap,
undefined,
namedImportMap
graph,
chunkWorkerData.calls,
ctx,
(current, total) => {
onProgress({
phase: 'parsing',
percent: Math.round(chunkBasePercent),
message: `Resolving calls (chunk ${chunkIdx + 1}/${numChunks})...`,
detail: `${current}/${total} files`,
stats: { filesProcessed: filesParsedSoFar, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
});
},
chunkWorkerData.constructorBindings,
),
processHeritageFromExtracted(
graph,
chunkWorkerData.heritage,
symbolTable,
importMap,
packageMap
graph,
chunkWorkerData.heritage,
ctx,
(current, total) => {
onProgress({
phase: 'parsing',
percent: Math.round(chunkBasePercent),
message: `Resolving heritage (chunk ${chunkIdx + 1}/${numChunks})...`,
detail: `${current}/${total} records`,
stats: { filesProcessed: filesParsedSoFar, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
});
},
),
processRoutesFromExtracted(
graph,
chunkWorkerData.routes ?? [],
symbolTable,
importMap,
packageMap
graph,
chunkWorkerData.routes ?? [],
ctx,
(current, total) => {
onProgress({
phase: 'parsing',
percent: Math.round(chunkBasePercent),
message: `Resolving routes (chunk ${chunkIdx + 1}/${numChunks})...`,
detail: `${current}/${total} routes`,
stats: { filesProcessed: filesParsedSoFar, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
});
},
),
]);
} else {
await processImports(graph, chunkFiles, astCache, importMap, undefined, repoPath, allPaths, packageMap, namedImportMap);
await processImports(graph, chunkFiles, astCache, ctx, undefined, repoPath, allPaths);
sequentialChunkPaths.push(chunkPaths);
}
@@ -273,11 +312,22 @@ export const runPipelineFromRepo = async (
.filter(p => chunkContents.has(p))
.map(p => ({ path: p, content: chunkContents.get(p)! }));
astCache = createASTCache(chunkFiles.length);
await processCalls(graph, chunkFiles, astCache, symbolTable, importMap, packageMap, undefined, namedImportMap);
await processHeritage(graph, chunkFiles, astCache, symbolTable, importMap, packageMap);
const rubyHeritage = await processCalls(graph, chunkFiles, astCache, ctx);
await processHeritage(graph, chunkFiles, astCache, ctx);
if (rubyHeritage.length > 0) {
await processHeritageFromExtracted(graph, rubyHeritage, ctx);
}
astCache.clear();
}
// Log resolution cache stats
if (isDev) {
const rcStats = ctx.getStats();
const total = rcStats.cacheHits + rcStats.cacheMisses;
const hitRate = total > 0 ? ((rcStats.cacheHits / total) * 100).toFixed(1) : '0';
console.log(`🔍 Resolution cache: ${rcStats.cacheHits} hits, ${rcStats.cacheMisses} misses (${hitRate}% hit rate)`);
}
// Free import resolution context — suffix index + resolve cache no longer needed
// (allPathObjects and importCtx hold ~94MB+ for large repos)
allPathObjects.length = 0;
@@ -285,130 +335,137 @@ export const runPipelineFromRepo = async (
(importCtx as any).suffixIndex = null;
(importCtx as any).normalizedFileList = null;
// ── Phase 4.5: Method Resolution Order ──────────────────────────────
onProgress({
phase: 'parsing',
percent: 81,
message: 'Computing method resolution order...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
let communityResult: Awaited<ReturnType<typeof processCommunities>> | undefined;
let processResult: Awaited<ReturnType<typeof processProcesses>> | undefined;
const mroResult = computeMRO(graph);
if (isDev && mroResult.entries.length > 0) {
console.log(`🔀 MRO: ${mroResult.entries.length} classes analyzed, ${mroResult.ambiguityCount} ambiguities found, ${mroResult.overrideEdges} OVERRIDES edges`);
}
// ── Phase 5: Communities ───────────────────────────────────────────
onProgress({
phase: 'communities',
percent: 82,
message: 'Detecting code communities...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
const communityResult = await processCommunities(graph, (message, progress) => {
const communityProgress = 82 + (progress * 0.10);
if (!options?.skipGraphPhases) {
// ── Phase 4.5: Method Resolution Order ──────────────────────────────
onProgress({
phase: 'communities',
percent: Math.round(communityProgress),
message,
phase: 'parsing',
percent: 81,
message: 'Computing method resolution order...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
});
if (isDev) {
console.log(`🏘️ Community detection: ${communityResult.stats.totalCommunities} communities found (modularity: ${communityResult.stats.modularity.toFixed(3)})`);
}
const mroResult = computeMRO(graph);
if (isDev && mroResult.entries.length > 0) {
console.log(`🔀 MRO: ${mroResult.entries.length} classes analyzed, ${mroResult.ambiguityCount} ambiguities found, ${mroResult.overrideEdges} OVERRIDES edges`);
}
communityResult.communities.forEach(comm => {
graph.addNode({
id: comm.id,
label: 'Community' as const,
properties: {
name: comm.label,
filePath: '',
heuristicLabel: comm.heuristicLabel,
cohesion: comm.cohesion,
symbolCount: comm.symbolCount,
}
// ── Phase 5: Communities ───────────────────────────────────────────
onProgress({
phase: 'communities',
percent: 82,
message: 'Detecting code communities...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
});
communityResult.memberships.forEach(membership => {
graph.addRelationship({
id: `${membership.nodeId}_member_of_${membership.communityId}`,
type: 'MEMBER_OF',
sourceId: membership.nodeId,
targetId: membership.communityId,
confidence: 1.0,
reason: 'leiden-algorithm',
});
});
// ── Phase 6: Processes ─────────────────────────────────────────────
onProgress({
phase: 'processes',
percent: 94,
message: 'Detecting execution flows...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
let symbolCount = 0;
graph.forEachNode(n => { if (n.label !== 'File') symbolCount++; });
const dynamicMaxProcesses = Math.max(20, Math.min(300, Math.round(symbolCount / 10)));
const processResult = await processProcesses(
graph,
communityResult.memberships,
(message, progress) => {
const processProgress = 94 + (progress * 0.05);
communityResult = await processCommunities(graph, (message, progress) => {
const communityProgress = 82 + (progress * 0.10);
onProgress({
phase: 'processes',
percent: Math.round(processProgress),
phase: 'communities',
percent: Math.round(communityProgress),
message,
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
},
{ maxProcesses: dynamicMaxProcesses, minSteps: 3 }
);
});
if (isDev) {
console.log(`🔄 Process detection: ${processResult.stats.totalProcesses} processes found (${processResult.stats.crossCommunityCount} cross-community)`);
if (isDev) {
console.log(`🏘️ Community detection: ${communityResult.stats.totalCommunities} communities found (modularity: ${communityResult.stats.modularity.toFixed(3)})`);
}
communityResult.communities.forEach(comm => {
graph.addNode({
id: comm.id,
label: 'Community' as const,
properties: {
name: comm.label,
filePath: '',
heuristicLabel: comm.heuristicLabel,
cohesion: comm.cohesion,
symbolCount: comm.symbolCount,
}
});
});
communityResult.memberships.forEach(membership => {
graph.addRelationship({
id: `${membership.nodeId}_member_of_${membership.communityId}`,
type: 'MEMBER_OF',
sourceId: membership.nodeId,
targetId: membership.communityId,
confidence: 1.0,
reason: 'leiden-algorithm',
});
});
// ── Phase 6: Processes ─────────────────────────────────────────────
onProgress({
phase: 'processes',
percent: 94,
message: 'Detecting execution flows...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
let symbolCount = 0;
graph.forEachNode(n => { if (n.label !== 'File') symbolCount++; });
const dynamicMaxProcesses = Math.max(20, Math.min(300, Math.round(symbolCount / 10)));
processResult = await processProcesses(
graph,
communityResult.memberships,
(message, progress) => {
const processProgress = 94 + (progress * 0.05);
onProgress({
phase: 'processes',
percent: Math.round(processProgress),
message,
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
},
{ maxProcesses: dynamicMaxProcesses, minSteps: 3 }
);
if (isDev) {
console.log(`🔄 Process detection: ${processResult.stats.totalProcesses} processes found (${processResult.stats.crossCommunityCount} cross-community)`);
}
processResult.processes.forEach(proc => {
graph.addNode({
id: proc.id,
label: 'Process' as const,
properties: {
name: proc.label,
filePath: '',
heuristicLabel: proc.heuristicLabel,
processType: proc.processType,
stepCount: proc.stepCount,
communities: proc.communities,
entryPointId: proc.entryPointId,
terminalId: proc.terminalId,
}
});
});
processResult.steps.forEach(step => {
graph.addRelationship({
id: `${step.nodeId}_step_${step.step}_${step.processId}`,
type: 'STEP_IN_PROCESS',
sourceId: step.nodeId,
targetId: step.processId,
confidence: 1.0,
reason: 'trace-detection',
step: step.step,
});
});
}
processResult.processes.forEach(proc => {
graph.addNode({
id: proc.id,
label: 'Process' as const,
properties: {
name: proc.label,
filePath: '',
heuristicLabel: proc.heuristicLabel,
processType: proc.processType,
stepCount: proc.stepCount,
communities: proc.communities,
entryPointId: proc.entryPointId,
terminalId: proc.terminalId,
}
});
});
processResult.steps.forEach(step => {
graph.addRelationship({
id: `${step.nodeId}_step_${step.step}_${step.processId}`,
type: 'STEP_IN_PROCESS',
sourceId: step.nodeId,
targetId: step.processId,
confidence: 1.0,
reason: 'trace-detection',
step: step.step,
});
});
onProgress({
phase: 'complete',
percent: 100,
message: `Graph complete! ${communityResult.stats.totalCommunities} communities, ${processResult.stats.totalProcesses} processes detected.`,
message: communityResult && processResult
? `Graph complete! ${communityResult.stats.totalCommunities} communities, ${processResult.stats.totalProcesses} processes detected.`
: 'Graph complete! (graph phases skipped)',
stats: {
filesProcessed: totalFiles,
totalFiles,
@@ -0,0 +1,192 @@
/**
* Resolution Context
*
* Single implementation of tiered name resolution. Replaces the duplicated
* tier-selection logic previously split between symbol-resolver.ts and
* call-processor.ts.
*
* Resolution tiers (highest confidence first):
* 1. Same file (lookupExactFull — authoritative)
* 2a-named. Named binding chain (walkBindingChain via NamedImportMap)
* 2a. Import-scoped (lookupFuzzy filtered by ImportMap)
* 2b. Package-scoped (lookupFuzzy filtered by PackageMap)
* 3. Global (all candidates — consumers must check candidate count)
*/
import type { SymbolTable, SymbolDefinition } from './symbol-table.js';
import { createSymbolTable } from './symbol-table.js';
import type { NamedImportBinding } from './import-processor.js';
import { isFileInPackageDir } from './import-processor.js';
import { walkBindingChain } from './named-binding-extraction.js';
/** Resolution tier for tracking, logging, and test assertions. */
export type ResolutionTier = 'same-file' | 'import-scoped' | 'global';
/** Tier-selected candidates with metadata. */
export interface TieredCandidates {
readonly candidates: readonly SymbolDefinition[];
readonly tier: ResolutionTier;
}
/** Confidence scores per resolution tier. */
export const TIER_CONFIDENCE: Record<ResolutionTier, number> = {
'same-file': 0.95,
'import-scoped': 0.9,
'global': 0.5,
};
// --- Map types ---
export type ImportMap = Map<string, Set<string>>;
export type PackageMap = Map<string, Set<string>>;
export type NamedImportMap = Map<string, Map<string, NamedImportBinding>>;
export interface ResolutionContext {
/**
* The only resolution API. Returns all candidates at the winning tier.
*
* Tier 3 ('global') returns ALL candidates regardless of count —
* consumers must check candidates.length and refuse ambiguous matches.
*/
resolve(name: string, fromFile: string): TieredCandidates | null;
// --- Data access (for pipeline wiring, not resolution) ---
/** Symbol table — used by parsing-processor to populate symbols. */
readonly symbols: SymbolTable;
/** Raw maps — used by import-processor to populate import data. */
readonly importMap: ImportMap;
readonly packageMap: PackageMap;
readonly namedImportMap: NamedImportMap;
// --- Per-file cache lifecycle ---
enableCache(filePath: string): void;
clearCache(): void;
// --- Operational ---
getStats(): { fileCount: number; globalSymbolCount: number; cacheHits: number; cacheMisses: number };
clear(): void;
}
export const createResolutionContext = (): ResolutionContext => {
const symbols = createSymbolTable();
const importMap: ImportMap = new Map();
const packageMap: PackageMap = new Map();
const namedImportMap: NamedImportMap = new Map();
// Per-file cache state
let cacheFile: string | null = null;
let cache: Map<string, TieredCandidates | null> | null = null;
let cacheHits = 0;
let cacheMisses = 0;
// --- Core resolution (single implementation of tier logic) ---
const resolveUncached = (name: string, fromFile: string): TieredCandidates | null => {
// Tier 1: Same file — authoritative match
const localDef = symbols.lookupExactFull(fromFile, name);
if (localDef) {
return { candidates: [localDef], tier: 'same-file' };
}
// Get all global definitions for subsequent tiers
const allDefs = symbols.lookupFuzzy(name);
// Tier 2a-named: Check named bindings BEFORE empty-allDefs early return
// because aliased imports mean lookupFuzzy('U') returns empty but we
// can resolve via the exported name.
const chainResult = walkBindingChain(name, fromFile, symbols, namedImportMap, allDefs);
if (chainResult && chainResult.length > 0) {
return { candidates: chainResult, tier: 'import-scoped' };
}
if (allDefs.length === 0) return null;
// Tier 2a: Import-scoped — definition in a file imported by fromFile
const importedFiles = importMap.get(fromFile);
if (importedFiles) {
const importedDefs = allDefs.filter(def => importedFiles.has(def.filePath));
if (importedDefs.length > 0) {
return { candidates: importedDefs, tier: 'import-scoped' };
}
}
// Tier 2b: Package-scoped — definition in a package dir imported by fromFile
const importedPackages = packageMap.get(fromFile);
if (importedPackages) {
const packageDefs = allDefs.filter(def => {
for (const dirSuffix of importedPackages) {
if (isFileInPackageDir(def.filePath, dirSuffix)) return true;
}
return false;
});
if (packageDefs.length > 0) {
return { candidates: packageDefs, tier: 'import-scoped' };
}
}
// Tier 3: Global — pass all candidates through.
// Consumers must check candidate count and refuse ambiguous matches.
return { candidates: allDefs, tier: 'global' };
};
const resolve = (name: string, fromFile: string): TieredCandidates | null => {
// Check cache (only when enabled AND fromFile matches cached file)
if (cache && cacheFile === fromFile) {
if (cache.has(name)) {
cacheHits++;
return cache.get(name)!;
}
cacheMisses++;
}
const result = resolveUncached(name, fromFile);
// Store in cache if active and file matches
if (cache && cacheFile === fromFile) {
cache.set(name, result);
}
return result;
};
// --- Cache lifecycle ---
const enableCache = (filePath: string): void => {
cacheFile = filePath;
if (!cache) cache = new Map();
else cache.clear();
};
const clearCache = (): void => {
cacheFile = null;
// Reuse the Map instance — just clear entries to reduce GC pressure at scale.
cache?.clear();
};
const getStats = () => ({
...symbols.getStats(),
cacheHits,
cacheMisses,
});
const clear = (): void => {
symbols.clear();
importMap.clear();
packageMap.clear();
namedImportMap.clear();
clearCache();
cacheHits = 0;
cacheMisses = 0;
};
return {
resolve,
symbols,
importMap,
packageMap,
namedImportMap,
enableCache,
clearCache,
getStats,
clear,
};
};
@@ -19,5 +19,7 @@ export type { ComposerConfig } from './php.js';
export { resolveRustImport, tryRustModulePath } from './rust.js';
export { resolveRubyImport } from './ruby.js';
export { resolveImportPath, RESOLVE_CACHE_CAP } from './standard.js';
export type { TsconfigPaths } from './standard.js';
@@ -0,0 +1,23 @@
/**
* Ruby require/require_relative import resolution.
* Handles path resolution for Ruby's require and require_relative calls.
*/
import type { SuffixIndex } from './utils.js';
import { suffixResolve } from './utils.js';
/**
* Resolve a Ruby require/require_relative path to a matching .rb file.
*
* require_relative paths are pre-normalized to './' prefix by the caller.
* require paths use suffix matching (gem-style paths like 'json', 'net/http').
*/
export function resolveRubyImport(
importPath: string,
normalizedFileList: string[],
allFileList: string[],
index?: SuffixIndex,
): string | null {
const pathParts = importPath.replace(/^\.\//, '').split('/').filter(Boolean);
return suffixResolve(pathParts, normalizedFileList, allFileList, index);
}
@@ -26,6 +26,8 @@ export const EXTENSIONS = [
'.php', '.phtml',
// Swift
'.swift',
// Ruby
'.rb',
];
/**
@@ -1,123 +0,0 @@
/**
* Symbol Resolver
*
* Import-filtered candidate narrowing for bare identifier resolution.
* NOT FQN resolution — does not parse qualifiers (ns::Bar, com.foo.Bar).
*
* Shared between heritage-processor.ts and call-processor.ts.
*/
import type { SymbolTable, SymbolDefinition } from './symbol-table.js';
import type { ImportMap, PackageMap, NamedImportMap } from './import-processor.js';
import { isFileInPackageDir } from './import-processor.js';
import { walkBindingChain } from './named-binding-extraction.js';
/** Resolution tier for internal tracking, logging, and test assertions. */
export type ResolutionTier = 'same-file' | 'import-scoped' | 'unique-global';
/** Internal resolution result preserving tier metadata. */
export interface InternalResolution {
definition: SymbolDefinition;
tier: ResolutionTier;
candidateCount: number;
}
/**
* Resolve a bare identifier to its best-matching definition using import context.
*
* Resolution tiers (highest confidence first):
* 1. Same file (lookupExactFull — authoritative)
* 2. Import-scoped (lookupFuzzy filtered by importMap — acceptable)
* 3. Unique global (lookupFuzzy with exactly 1 match — acceptable fallback)
*
* If multiple global candidates remain after filtering, returns null.
* A wrong edge is worse than no edge.
*/
export const resolveSymbol = (
name: string,
currentFilePath: string,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
): SymbolDefinition | null => {
return resolveSymbolInternal(name, currentFilePath, symbolTable, importMap, packageMap, namedImportMap)?.definition ?? null;
};
/** Internal resolver preserving tier metadata for logging and test assertions. */
export const resolveSymbolInternal = (
name: string,
currentFilePath: string,
symbolTable: SymbolTable,
importMap: ImportMap,
packageMap?: PackageMap,
namedImportMap?: NamedImportMap,
): InternalResolution | null => {
// Tier 1: Same file — authoritative match
const localDef = symbolTable.lookupExactFull(currentFilePath, name);
if (localDef) return { definition: localDef, tier: 'same-file', candidateCount: 1 };
// Get all global definitions for subsequent tiers
const allDefs = symbolTable.lookupFuzzy(name);
// Tier 2a-named: Check named bindings BEFORE the empty-allDefs early return,
// because aliased imports (import { User as U }) mean lookupFuzzy('U') returns
// empty but we can resolve via the exported name.
if (namedImportMap) {
const result = resolveNamedBindingChain(name, currentFilePath, symbolTable, namedImportMap, allDefs);
if (result) return result;
}
if (allDefs.length === 0) return null;
// Tier 2a: Import-scoped — check if any definition is in a file imported by currentFile
const importedFiles = importMap.get(currentFilePath);
if (importedFiles) {
for (const def of allDefs) {
if (importedFiles.has(def.filePath)) {
return { definition: def, tier: 'import-scoped', candidateCount: allDefs.length };
}
}
}
// Tier 2b: Package-scoped — check if any definition is in a package/namespace dir imported by currentFile
// Used for Go packages and C# namespace imports to avoid ImportMap expansion bloat
const importedPackages = packageMap?.get(currentFilePath);
if (importedPackages) {
for (const def of allDefs) {
for (const dirSuffix of importedPackages) {
if (isFileInPackageDir(def.filePath, dirSuffix)) {
return { definition: def, tier: 'import-scoped', candidateCount: allDefs.length };
}
}
}
}
// Tier 3: Unique global — ONLY if exactly one candidate exists
// Ambiguous global matches are refused. A wrong edge is worse than no edge.
if (allDefs.length === 1) {
return { definition: allDefs[0], tier: 'unique-global', candidateCount: 1 };
}
// Ambiguous: multiple global candidates, no import or same-file match → refuse
return null;
};
/**
* Follow re-export chains through NamedImportMap.
* Delegates chain-walking to the shared walkBindingChain utility, then
* applies symbol-resolver semantics: exactly one match required.
*/
const resolveNamedBindingChain = (
name: string,
currentFilePath: string,
symbolTable: SymbolTable,
namedImportMap: NamedImportMap,
allDefs: SymbolDefinition[],
): InternalResolution | null => {
const defs = walkBindingChain(name, currentFilePath, symbolTable, namedImportMap, allDefs);
if (defs?.length === 1) {
return { definition: defs[0], tier: 'import-scoped', candidateCount: defs.length };
}
return null;
};
+5 -2
View File
@@ -3,6 +3,8 @@ export interface SymbolDefinition {
filePath: string;
type: string; // 'Function', 'Class', etc.
parameterCount?: number;
/** Raw return type text extracted from AST (e.g. 'User', 'Promise<User>') */
returnType?: string;
/** Links Method/Constructor to owning Class/Struct/Trait nodeId */
ownerId?: string;
}
@@ -16,7 +18,7 @@ export interface SymbolTable {
name: string,
nodeId: string,
type: string,
metadata?: { parameterCount?: number; ownerId?: string }
metadata?: { parameterCount?: number; returnType?: string; ownerId?: string }
) => void;
/**
@@ -62,13 +64,14 @@ export const createSymbolTable = (): SymbolTable => {
name: string,
nodeId: string,
type: string,
metadata?: { parameterCount?: number; ownerId?: string }
metadata?: { parameterCount?: number; returnType?: string; ownerId?: string }
) => {
const def: SymbolDefinition = {
nodeId,
filePath,
type,
...(metadata?.parameterCount !== undefined ? { parameterCount: metadata.parameterCount } : {}),
...(metadata?.returnType !== undefined ? { returnType: metadata.returnType } : {}),
...(metadata?.ownerId !== undefined ? { ownerId: metadata.ownerId } : {}),
};
@@ -306,7 +306,7 @@ export const CPP_QUERIES = `
(field_declaration_list
(function_definition
declarator: (function_declarator
declarator: [(field_identifier) (identifier) (operator_name) (destructor_name)] @name))) @definition.method
declarator: [(field_identifier) (identifier) (operator_name) (destructor_name)] @name)) @definition.method)
; Templates
(template_declaration (class_specifier name: (type_identifier) @name)) @definition.template
@@ -365,6 +365,13 @@ export const CSHARP_QUERIES = `
(invocation_expression function: (identifier) @call.name) @call
(invocation_expression function: (member_access_expression name: (identifier) @call.name)) @call
; Null-conditional method calls: user?.Save()
; Parses as: invocation_expression → conditional_access_expression → member_binding_expression → identifier
(invocation_expression
function: (conditional_access_expression
(member_binding_expression
(identifier) @call.name))) @call
; Constructor calls: new Foo() and new Foo { Props }
(object_creation_expression type: (identifier) @call.name) @call
@@ -496,6 +503,51 @@ export const PHP_QUERIES = `
[(name) (qualified_name)] @heritage.trait))) @heritage
`;
// Ruby queries - works with tree-sitter-ruby
// NOTE: Ruby uses `call` for require, include, extend, prepend, attr_* etc.
// These are all captured as @call and routed in JS post-processing:
// - require/require_relative → import extraction
// - include/extend/prepend → heritage (mixin) extraction
// - attr_accessor/attr_reader/attr_writer → property definition extraction
// - everything else → regular call extraction
export const RUBY_QUERIES = `
; ── Modules ──────────────────────────────────────────────────────────────────
(module
name: (constant) @name) @definition.module
; ── Classes ──────────────────────────────────────────────────────────────────
(class
name: (constant) @name) @definition.class
; ── Instance methods ─────────────────────────────────────────────────────────
(method
name: (identifier) @name) @definition.method
; ── Singleton (class-level) methods ──────────────────────────────────────────
(singleton_method
name: (identifier) @name) @definition.method
; ── All calls (require, include, attr_*, and regular calls routed in JS) ─────
(call
method: (identifier) @call.name) @call
; ── Bare calls without parens (identifiers at statement level are method calls) ─
; NOTE: This may over-capture variable reads as calls (e.g. 'result' at
; statement level). Ruby's grammar makes bare identifiers ambiguous — they
; could be local variables or zero-arity method calls. Post-processing via
; isBuiltInOrNoise and symbol resolution filtering suppresses most false
; positives, but a variable name that coincidentally matches a method name
; elsewhere may produce a false CALLS edge.
(body_statement
(identifier) @call.name @call)
; ── Heritage: class < SuperClass ─────────────────────────────────────────────
(class
name: (constant) @heritage.class
superclass: (superclass
(constant) @heritage.extends)) @heritage
`;
// Kotlin queries - works with tree-sitter-kotlin (fwcd/tree-sitter-kotlin)
// Based on official tags.scm; functions use simple_identifier, classes use type_identifier
export const KOTLIN_QUERIES = `
@@ -643,6 +695,7 @@ export const LANGUAGE_QUERIES: Record<SupportedLanguages, string> = {
[SupportedLanguages.Go]: GO_QUERIES,
[SupportedLanguages.CPlusPlus]: CPP_QUERIES,
[SupportedLanguages.CSharp]: CSHARP_QUERIES,
[SupportedLanguages.Ruby]: RUBY_QUERIES,
[SupportedLanguages.Rust]: RUST_QUERIES,
[SupportedLanguages.PHP]: PHP_QUERIES,
[SupportedLanguages.Kotlin]: KOTLIN_QUERIES,
+582 -60
View File
@@ -1,7 +1,10 @@
import type { SyntaxNode } from './utils.js';
import { FUNCTION_NODE_TYPES, extractFunctionName } from './utils.js';
import { FUNCTION_NODE_TYPES, extractFunctionName, CLASS_CONTAINER_TYPES } from './utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
import { typeConfigs, TYPED_PARAMETER_TYPES } from './type-extractors/index.js';
import type { ClassNameLookup } from './type-extractors/types.js';
import { extractSimpleTypeName, extractVarName, stripNullable } from './type-extractors/shared.js';
import type { SymbolTable } from './symbol-table.js';
/**
* Per-file scoped type environment: maps (scope, variableName) → typeName.
@@ -9,7 +12,9 @@ import { typeConfigs, TYPED_PARAMETER_TYPES } from './type-extractors/index.js';
* file-level variables use the '' (empty string) scope.
*
* Design constraints:
* - Explicit-only: only type annotations, never inferred types
* - Explicit-only: Tier 0 uses type annotations; Tier 1 infers from constructors
* - Tier 2: single-pass assignment chain propagation in source order — resolves
* `const b = a` when `a` already has a type from Tier 0/1
* - Scope-aware: function-local variables don't collide across functions
* - Conservative: complex/generic types extract the base name only
* - Per-file: built once, used for receiver resolution, then discarded
@@ -19,30 +24,271 @@ export type TypeEnv = Map<string, Map<string, string>>;
/** File-level scope key */
const FILE_SCOPE = '';
/** Fallback for languages where class names aren't in a 'name' field (e.g. Kotlin uses type_identifier). */
const findTypeIdentifierChild = (node: SyntaxNode): SyntaxNode | null => {
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (child && child.type === 'type_identifier') return child;
}
return null;
};
/**
* Look up a variable's type in the TypeEnv, trying the call's enclosing
* function scope first, then falling back to file-level scope.
* Per-file type environment with receiver resolution.
* Built once per file via `buildTypeEnv`, used for receiver-type filtering,
* then discarded. Encapsulates scope-aware type lookup and self/this/super
* AST resolution behind a single `.lookup()` method.
*/
export const lookupTypeEnv = (
export interface TypeEnvironment {
/** Look up a variable's resolved type, with self/this/super AST resolution. */
lookup(varName: string, callNode: SyntaxNode): string | undefined;
/** Unverified cross-file constructor bindings for SymbolTable verification. */
readonly constructorBindings: readonly ConstructorBinding[];
/** Raw per-scope type bindings — for testing and debugging. */
readonly env: TypeEnv;
}
/**
* Position-indexed pattern binding: active only within a specific AST range.
* Used for smart-cast narrowing in mutually exclusive branches (e.g., Kotlin when arms).
*/
interface PatternOverride {
rangeStart: number;
rangeEnd: number;
typeName: string;
}
/** scope → varName → overrides (checked in order, first range match wins) */
type PatternOverrides = Map<string, Map<string, PatternOverride[]>>;
/** AST node types that represent mutually exclusive branch containers for pattern bindings. */
const PATTERN_BRANCH_TYPES = new Set([
'when_entry', // Kotlin when
'switch_block_label', // Java switch (enhanced)
]);
/** Walk up the AST from a pattern node to find the enclosing branch container. */
const findPatternBranchScope = (node: SyntaxNode): SyntaxNode | undefined => {
let current = node.parent;
while (current) {
if (PATTERN_BRANCH_TYPES.has(current.type)) return current;
if (FUNCTION_NODE_TYPES.has(current.type)) return undefined;
current = current.parent;
}
return undefined;
};
/** Bare nullable keywords that fastStripNullable must reject. */
const FAST_NULLABLE_KEYWORDS = new Set(['null', 'undefined', 'void', 'None', 'nil']);
/**
* Fast-path nullable check: 90%+ of type names are simple identifiers (e.g. "User")
* that don't need the full stripNullable parse. Only call stripNullable when the
* string contains nullable markers ('|' for union types, '?' for nullable suffix).
*/
const fastStripNullable = (typeName: string): string | undefined => {
if (FAST_NULLABLE_KEYWORDS.has(typeName)) return undefined;
return (typeName.indexOf('|') === -1 && typeName.indexOf('?') === -1)
? typeName
: stripNullable(typeName);
};
/** Implementation of the lookup logic — shared between TypeEnvironment and the legacy export. */
const lookupInEnv = (
env: TypeEnv,
varName: string,
callNode: SyntaxNode,
patternOverrides?: PatternOverrides,
): string | undefined => {
// Self/this receiver: resolve to enclosing class name via AST walk
if (varName === 'self' || varName === 'this' || varName === '$this') {
return findEnclosingClassName(callNode);
}
// Super/base/parent receiver: resolve to the parent class name via AST walk.
// Walks up to the enclosing class, then extracts the superclass from its heritage node.
if (varName === 'super' || varName === 'base' || varName === 'parent') {
return findEnclosingParentClassName(callNode);
}
// Determine the enclosing function scope for the call
const scopeKey = findEnclosingScopeKey(callNode);
// Check position-indexed pattern overrides first (e.g., Kotlin when/is smart casts).
// These take priority over flat scopeEnv because they represent per-branch narrowing.
if (scopeKey && patternOverrides) {
const varOverrides = patternOverrides.get(scopeKey)?.get(varName);
if (varOverrides) {
const pos = callNode.startIndex;
for (const override of varOverrides) {
if (pos >= override.rangeStart && pos <= override.rangeEnd) {
return fastStripNullable(override.typeName);
}
}
}
}
// Try function-local scope first
if (scopeKey) {
const scopeEnv = env.get(scopeKey);
if (scopeEnv) {
const result = scopeEnv.get(varName);
if (result) return result;
if (result) return fastStripNullable(result);
}
}
// Fall back to file-level scope
const fileEnv = env.get(FILE_SCOPE);
return fileEnv?.get(varName);
const raw = fileEnv?.get(varName);
return raw ? fastStripNullable(raw) : undefined;
};
/**
* Walk up the AST from a node to find the enclosing class/module name.
* Used to resolve `self`/`this` receivers to their containing type.
*/
const findEnclosingClassName = (node: SyntaxNode): string | undefined => {
let current = node.parent;
while (current) {
if (CLASS_CONTAINER_TYPES.has(current.type)) {
const nameNode = current.childForFieldName('name')
?? findTypeIdentifierChild(current);
if (nameNode) return nameNode.text;
}
current = current.parent;
}
return undefined;
};
/**
* Walk up the AST to find the enclosing class, then extract its parent class name
* from the heritage/superclass AST node. Used to resolve `super`/`base`/`parent`.
*
* Supported patterns per tree-sitter grammar:
* - Java/Ruby: `superclass` field → type_identifier/constant
* - Python: `superclasses` field → argument_list → first identifier
* - TypeScript/JS: unnamed `class_heritage` child → `extends_clause` → identifier
* - C#: unnamed `base_list` child → first identifier
* - PHP: unnamed `base_clause` child → name
* - Kotlin: unnamed `delegation_specifier` child → constructor_invocation → user_type → type_identifier
* - C++: unnamed `base_class_clause` child → type_identifier
* - Swift: unnamed `inheritance_specifier` child → user_type → type_identifier
*/
const findEnclosingParentClassName = (node: SyntaxNode): string | undefined => {
let current = node.parent;
while (current) {
if (CLASS_CONTAINER_TYPES.has(current.type)) {
return extractParentClassFromNode(current);
}
current = current.parent;
}
return undefined;
};
/** Extract the parent/superclass name from a class declaration AST node. */
const extractParentClassFromNode = (classNode: SyntaxNode): string | undefined => {
// 1. Named fields: Java (superclass), Ruby (superclass), Python (superclasses)
const superclassNode = classNode.childForFieldName('superclass');
if (superclassNode) {
// Java: superclass > type_identifier or generic_type, Ruby: superclass > constant
const inner = superclassNode.childForFieldName('type')
?? superclassNode.firstNamedChild
?? superclassNode;
return extractSimpleTypeName(inner) ?? inner.text;
}
const superclassesNode = classNode.childForFieldName('superclasses');
if (superclassesNode) {
// Python: argument_list with identifiers or attribute nodes (e.g. models.Model)
const first = superclassesNode.firstNamedChild;
if (first) return extractSimpleTypeName(first) ?? first.text;
}
// 2. Unnamed children: walk class node's children looking for heritage nodes
for (let i = 0; i < classNode.childCount; i++) {
const child = classNode.child(i);
if (!child) continue;
switch (child.type) {
// TypeScript: class_heritage > extends_clause > type_identifier
// JavaScript: class_heritage > identifier (no extends_clause wrapper)
case 'class_heritage': {
for (let j = 0; j < child.childCount; j++) {
const clause = child.child(j);
if (clause?.type === 'extends_clause') {
const typeNode = clause.firstNamedChild;
if (typeNode) return extractSimpleTypeName(typeNode) ?? typeNode.text;
}
// JS: direct identifier child (no extends_clause wrapper)
if (clause?.type === 'identifier' || clause?.type === 'type_identifier') {
return clause.text;
}
}
break;
}
// C#: base_list > identifier or generic_name > identifier
case 'base_list': {
const first = child.firstNamedChild;
if (first) {
// generic_name wraps the identifier: BaseClass<T>
if (first.type === 'generic_name') {
const inner = first.childForFieldName('name') ?? first.firstNamedChild;
if (inner) return inner.text;
}
return first.text;
}
break;
}
// PHP: base_clause > name
case 'base_clause': {
const name = child.firstNamedChild;
if (name) return name.text;
break;
}
// C++: base_class_clause > type_identifier (with optional access_specifier before it)
case 'base_class_clause': {
for (let j = 0; j < child.childCount; j++) {
const inner = child.child(j);
if (inner?.type === 'type_identifier') return inner.text;
}
break;
}
// Kotlin: delegation_specifier > constructor_invocation > user_type > type_identifier
case 'delegation_specifier': {
const delegate = child.firstNamedChild;
if (delegate?.type === 'constructor_invocation') {
const userType = delegate.firstNamedChild;
if (userType?.type === 'user_type') {
const typeId = userType.firstNamedChild;
if (typeId) return typeId.text;
}
}
// Also handle plain user_type (interface conformance without parentheses)
if (delegate?.type === 'user_type') {
const typeId = delegate.firstNamedChild;
if (typeId) return typeId.text;
}
break;
}
// Swift: inheritance_specifier > user_type > type_identifier
case 'inheritance_specifier': {
const userType = child.childForFieldName('inherits_from') ?? child.firstNamedChild;
if (userType?.type === 'user_type') {
const typeId = userType.firstNamedChild;
if (typeId) return typeId.text;
}
break;
}
}
}
return undefined;
};
/** Find the enclosing function name for scope lookup. */
@@ -59,66 +305,342 @@ const findEnclosingScopeKey = (node: SyntaxNode): string | undefined => {
};
/**
* Build a scoped TypeEnv from a tree-sitter AST for a given language.
* Walks the tree tracking enclosing function scopes, so that variables
* inside different functions don't collide.
* Create a lookup that checks both local AST class names AND the SymbolTable's
* global index. This allows extractInitializer functions to distinguish
* constructor calls from function calls (e.g. Kotlin `User()` vs `getUser()`)
* using cross-file type information when available.
*
* Only `.has()` is exposed — the SymbolTable doesn't support iteration.
* Results are memoized to avoid redundant lookupFuzzy scans across declarations.
*/
export const buildTypeEnv = (
tree: { rootNode: SyntaxNode },
language: SupportedLanguages,
): TypeEnv => {
const env: TypeEnv = new Map();
walkForTypes(tree.rootNode, language, env, FILE_SCOPE);
return env;
};
const createClassNameLookup = (
localNames: Set<string>,
symbolTable?: SymbolTable,
): ClassNameLookup => {
if (!symbolTable) return localNames;
const walkForTypes = (
node: SyntaxNode,
language: SupportedLanguages,
env: TypeEnv,
currentScope: string,
): void => {
// Detect scope boundaries (function/method definitions)
let scope = currentScope;
if (FUNCTION_NODE_TYPES.has(node.type)) {
const { funcName } = extractFunctionName(node);
if (funcName) scope = `${funcName}@${node.startIndex}`;
}
// Get or create the sub-map for this scope
if (!env.has(scope)) env.set(scope, new Map());
const scopeEnv = env.get(scope)!;
// Check if this node provides type information
extractTypeBinding(node, language, scopeEnv);
// Recurse into children
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (child) walkForTypes(child, language, env, scope);
}
const memo = new Map<string, boolean>();
return {
has(name: string): boolean {
if (localNames.has(name)) return true;
const cached = memo.get(name);
if (cached !== undefined) return cached;
const result = symbolTable.lookupFuzzy(name).some(def =>
def.type === 'Class' || def.type === 'Enum' || def.type === 'Struct',
);
memo.set(name, result);
return result;
},
};
};
/**
* Try to extract a (variableName → typeName) binding from a single AST node.
* Delegates to per-language type configurations.
* Build a TypeEnvironment from a tree-sitter AST for a given language.
* Single-pass: collects class/struct names, type bindings, AND constructor
* bindings that couldn't be resolved locally — all in one AST walk.
*
* When a symbolTable is provided (call-processor path), class names from across
* the project are available for constructor inference in languages like Kotlin
* where constructors are syntactically identical to function calls.
*/
const extractTypeBinding = (
node: SyntaxNode,
/**
* Node types whose subtrees can NEVER contain type-relevant descendants
* (declarations, parameters, for-loops, class definitions, pattern bindings).
* Conservative leaf-only set — verified safe across all 12 supported language grammars.
* IMPORTANT: Do NOT add expression containers (arguments, binary_expression, etc.) —
* they can contain arrow functions with typed parameters.
*/
const SKIP_SUBTREE_TYPES = new Set([
// Plain string literals (NOT template_string — it contains interpolated expressions
// that can hold arrow functions with typed parameters, e.g. `${(x: T) => x}`)
'string', 'string_literal',
'string_content', 'string_fragment', 'heredoc_body',
// Comments
'comment', 'line_comment', 'block_comment',
// Numeric/boolean/null literals
'number', 'integer_literal', 'float_literal',
'true', 'false', 'null',
// Regex
'regex', 'regex_pattern',
]);
export const buildTypeEnv = (
tree: { rootNode: SyntaxNode },
language: SupportedLanguages,
env: Map<string, string>,
): void => {
// === PARAMETERS (most languages) ===
// This guard eliminates 90%+ of calls before any language dispatch.
if (TYPED_PARAMETER_TYPES.has(node.type)) {
const config = typeConfigs[language];
config.extractParameter(node, env);
return;
symbolTable?: SymbolTable,
): TypeEnvironment => {
const env: TypeEnv = new Map();
const patternOverrides: PatternOverrides = new Map();
const localClassNames = new Set<string>();
const classNames = createClassNameLookup(localClassNames, symbolTable);
const config = typeConfigs[language];
const bindings: ConstructorBinding[] = [];
// Pre-compute combined set of node types that need extractTypeBinding.
// Single Set.has() replaces 3 separate checks per node in walk().
const interestingNodeTypes = new Set<string>();
TYPED_PARAMETER_TYPES.forEach(t => interestingNodeTypes.add(t));
config.declarationNodeTypes.forEach(t => interestingNodeTypes.add(t));
config.forLoopNodeTypes?.forEach(t => interestingNodeTypes.add(t));
const pendingAssignments: Array<{ scope: string; lhs: string; rhs: string }> = [];
// Maps `scope\0varName` → the type annotation AST node from the original declaration.
// Allows pattern extractors to navigate back to the declaration's generic type arguments
// (e.g., to extract T from Result<T, E> for `if let Ok(x) = res`).
// NOTE: This is a SUPERSET of scopeEnv — entries exist even when extractSimpleTypeName
// returns undefined for container types (User[], []User, List[User]). This is intentional:
// for-loop Strategy 1 needs the raw AST type node for exactly those container types.
const declarationTypeNodes = new Map<string, SyntaxNode>();
/**
* Try to extract a (variableName → typeName) binding from a single AST node.
*
* Resolution tiers (first match wins):
* - Tier 0: explicit type annotations via extractDeclaration / extractForLoopBinding
* - Tier 1: constructor-call inference via extractInitializer (fallback)
*
* Side effect: populates declarationTypeNodes for variables that have an explicit
* type annotation field on the declaration node. This allows pattern extractors to
* retrieve generic type arguments from the original declaration (e.g., extracting T
* from Result<T, E> for `if let Ok(x) = res`).
*/
const extractTypeBinding = (node: SyntaxNode, scopeEnv: Map<string, string>, scope: string): void => {
// This guard eliminates 90%+ of calls before any language dispatch.
if (TYPED_PARAMETER_TYPES.has(node.type)) {
// Capture the raw type annotation BEFORE extractParameter.
// Most languages use 'name' field; Rust uses 'pattern'; TS uses 'pattern' for some param types.
// Kotlin `parameter` nodes use positional children instead of named fields,
// so we fall back to scanning children by type when childForFieldName returns null.
let typeNode = node.childForFieldName('type');
if (typeNode) {
const nameNode = node.childForFieldName('name')
?? node.childForFieldName('pattern');
if (nameNode) {
const varName = extractVarName(nameNode);
if (varName && !declarationTypeNodes.has(`${scope}\0${varName}`)) {
declarationTypeNodes.set(`${scope}\0${varName}`, typeNode);
}
}
} else {
// Fallback: positional children (Kotlin `parameter` → simple_identifier + user_type)
let fallbackName: SyntaxNode | null = null;
let fallbackType: SyntaxNode | null = null;
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child) continue;
if (!fallbackName && (child.type === 'simple_identifier' || child.type === 'identifier')) {
fallbackName = child;
}
if (!fallbackType && (child.type === 'user_type' || child.type === 'type_identifier'
|| child.type === 'generic_type' || child.type === 'parameterized_type')) {
fallbackType = child;
}
}
if (fallbackName && fallbackType) {
const varName = extractVarName(fallbackName);
if (varName && !declarationTypeNodes.has(`${scope}\0${varName}`)) {
declarationTypeNodes.set(`${scope}\0${varName}`, fallbackType);
}
}
}
config.extractParameter(node, scopeEnv);
return;
}
// For-each loop variable bindings (Java/C#/Kotlin): explicit element types in the AST.
// Checked before declarationNodeTypes — loop variables are not declarations.
if (config.forLoopNodeTypes?.has(node.type)) {
config.extractForLoopBinding?.(node, scopeEnv, declarationTypeNodes, scope);
return;
}
if (config.declarationNodeTypes.has(node.type)) {
// Capture the raw type annotation AST node BEFORE extractDeclaration.
// This decouples type node capture from scopeEnv success — container types
// (User[], []User, List[User]) that fail extractSimpleTypeName still get
// their AST type node recorded for Strategy 1 for-loop resolution.
// Try direct extraction first (works for Go var_spec, Python assignment, Rust let_declaration).
// Try direct type field first, then unwrap wrapper nodes (C# field_declaration,
// local_declaration_statement wrap their type inside a variable_declaration child).
let typeNode = node.childForFieldName('type');
if (!typeNode) {
// C# field_declaration / local_declaration_statement wrap type inside variable_declaration.
// Use manual loop instead of namedChildren.find() to avoid array allocation on hot path.
let wrapped = node.childForFieldName('declaration');
if (!wrapped) {
for (let i = 0; i < node.namedChildCount; i++) {
const c = node.namedChild(i);
if (c?.type === 'variable_declaration') { wrapped = c; break; }
}
}
if (wrapped) typeNode = wrapped.childForFieldName('type');
}
if (typeNode) {
const nameNode = node.childForFieldName('name')
?? node.childForFieldName('left')
?? node.childForFieldName('pattern');
if (nameNode) {
const varName = extractVarName(nameNode);
if (varName && !declarationTypeNodes.has(`${scope}\0${varName}`)) {
declarationTypeNodes.set(`${scope}\0${varName}`, typeNode);
}
}
}
// Run the language-specific declaration extractor (may or may not add to scopeEnv).
const keysBefore = typeNode ? new Set(scopeEnv.keys()) : undefined;
config.extractDeclaration(node, scopeEnv);
// Fallback: for multi-declarator languages (TS, C#, Java) where the type field
// is on variable_declarator children, capture via keysBefore/keysAfter diff.
if (typeNode && keysBefore) {
for (const varName of scopeEnv.keys()) {
if (!keysBefore.has(varName) && !declarationTypeNodes.has(`${scope}\0${varName}`)) {
declarationTypeNodes.set(`${scope}\0${varName}`, typeNode);
}
}
}
// Tier 1: constructor-call inference as fallback.
// Always called when available — each language's extractInitializer
// internally skips declarators that already have explicit annotations,
// so this handles mixed cases like `const a: A = x, b = new B()`.
if (config.extractInitializer) {
config.extractInitializer(node, scopeEnv, classNames);
}
}
};
const walk = (node: SyntaxNode, currentScope: string): void => {
// Fast skip: subtrees that can never contain type-relevant nodes (leaf-like literals).
if (SKIP_SUBTREE_TYPES.has(node.type)) return;
// Collect class/struct names as we encounter them (used by extractInitializer
// to distinguish constructor calls from function calls, e.g. C++ `User()` vs `getUser()`)
// Currently only C++ uses this locally; other languages rely on the SymbolTable path.
if (CLASS_CONTAINER_TYPES.has(node.type)) {
// Most languages use 'name' field; Kotlin uses a type_identifier child instead
const nameNode = node.childForFieldName('name')
?? findTypeIdentifierChild(node);
if (nameNode) localClassNames.add(nameNode.text);
}
// Detect scope boundaries (function/method definitions)
let scope = currentScope;
if (FUNCTION_NODE_TYPES.has(node.type)) {
const { funcName } = extractFunctionName(node);
if (funcName) scope = `${funcName}@${node.startIndex}`;
}
// Only create scope map and call extractTypeBinding for interesting node types.
// Single Set.has() replaces 3 separate checks inside extractTypeBinding.
if (interestingNodeTypes.has(node.type)) {
if (!env.has(scope)) env.set(scope, new Map());
const scopeEnv = env.get(scope)!;
extractTypeBinding(node, scopeEnv, scope);
}
// Pattern binding extraction: handles constructs that introduce NEW typed variables
// via pattern matching (e.g. `if let Some(x) = opt`, `x instanceof T t`).
// Runs after Tier 0/1 so scopeEnv already contains the source variable's type.
// Conservative: extractor returns undefined when source type is unknown.
if (config.extractPatternBinding && (!config.patternBindingNodeTypes || config.patternBindingNodeTypes.has(node.type))) {
// Ensure scopeEnv exists for pattern binding reads/writes
if (!env.has(scope)) env.set(scope, new Map());
const scopeEnv = env.get(scope)!;
const patternBinding = config.extractPatternBinding(node, scopeEnv, declarationTypeNodes, scope);
if (patternBinding) {
if (config.allowPatternBindingOverwrite) {
// Position-indexed: store per-branch binding for smart-cast narrowing.
// Each when arm / switch case gets its own type for the variable,
// preventing cross-arm contamination (e.g., Kotlin when/is).
const branchNode = findPatternBranchScope(node);
if (branchNode) {
if (!patternOverrides.has(scope)) patternOverrides.set(scope, new Map());
const varMap = patternOverrides.get(scope)!;
if (!varMap.has(patternBinding.varName)) varMap.set(patternBinding.varName, []);
varMap.get(patternBinding.varName)!.push({
rangeStart: branchNode.startIndex,
rangeEnd: branchNode.endIndex,
typeName: patternBinding.typeName,
});
}
// Also store in flat scopeEnv as fallback (last arm wins — same as before
// for code that doesn't use position-indexed lookup).
scopeEnv.set(patternBinding.varName, patternBinding.typeName);
} else if (!scopeEnv.has(patternBinding.varName)) {
// First-writer-wins for languages without smart-cast overwrite (Java instanceof, etc.)
scopeEnv.set(patternBinding.varName, patternBinding.typeName);
}
}
}
// Tier 2: collect plain-identifier RHS assignments for post-walk propagation.
// Delegates to per-language extractPendingAssignment — AST shapes differ widely
// (JS uses variable_declarator/name/value, Rust uses let_declaration/pattern/value,
// Python uses assignment/left/right, Go uses short_var_declaration/expression_list).
if (config.extractPendingAssignment && config.declarationNodeTypes.has(node.type)) {
// scopeEnv is guaranteed to exist here because declarationNodeTypes is a subset
// of interestingNodeTypes, so extractTypeBinding already created the scope map above.
const scopeEnv = env.get(scope);
if (scopeEnv) {
const pending = config.extractPendingAssignment(node, scopeEnv);
if (pending) {
pendingAssignments.push({ scope, ...pending });
}
}
}
// Scan for constructor bindings that couldn't be resolved locally.
// Only collect if TypeEnv didn't already resolve this binding.
if (config.scanConstructorBinding) {
const result = config.scanConstructorBinding(node);
if (result) {
const scopeEnv = env.get(scope);
if (!scopeEnv?.has(result.varName)) {
bindings.push({ scope, ...result });
}
}
}
// Recurse into children
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (child) walk(child, scope);
}
};
walk(tree.rootNode, FILE_SCOPE);
// Tier 2: single-pass assignment chain propagation in source order.
// Resolves `const b = a` where `a` has a known type from Tier 0/1.
// Multi-hop chains resolve when forward-declared (a→b→c in source order);
// reverse-order assignments are depth-1 only. No fixpoint iteration —
// this covers 95%+ of real-world patterns.
for (const { scope, lhs, rhs } of pendingAssignments) {
const scopeEnv = env.get(scope);
if (!scopeEnv || scopeEnv.has(lhs)) continue;
const rhsType = scopeEnv.get(rhs) ?? env.get(FILE_SCOPE)?.get(rhs);
if (rhsType) {
scopeEnv.set(lhs, rhsType);
}
}
// === Per-language declaration extraction ===
const config = typeConfigs[language];
if (config.declarationNodeTypes.has(node.type)) {
config.extractDeclaration(node, env);
}
return {
lookup: (varName, callNode) => lookupInEnv(env, varName, callNode, patternOverrides),
constructorBindings: bindings,
env,
};
};
/**
* Unverified constructor binding: a `val x = Callee()` pattern where we
* couldn't confirm the callee is a class (because it's defined in another file).
* The caller must verify `calleeName` against the SymbolTable before trusting.
*/
export interface ConstructorBinding {
/** Function scope key (matches TypeEnv scope keys) */
scope: string;
/** Variable name that received the constructor result */
varName: string;
/** Name of the callee (potential class constructor) */
calleeName: string;
/** Enclosing class name when callee is a method on a known receiver (e.g. $this) */
receiverClassName?: string;
}
@@ -1,6 +1,6 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor, InitializerExtractor, ClassNameLookup, ConstructorBindingScanner, PendingAssignmentExtractor, ForLoopExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, resolveIterableElementType, methodToTypeArgPosition, type TypeArgPosition } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'declaration',
@@ -32,6 +32,75 @@ const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<str
if (varName) env.set(varName, typeName);
};
/** C++: auto x = new User(); auto x = User(); */
const extractInitializer: InitializerExtractor = (node: SyntaxNode, env: Map<string, string>, classNames: ClassNameLookup): void => {
const typeNode = node.childForFieldName('type');
if (!typeNode) return;
// Only handle auto/placeholder — typed declarations are handled by extractDeclaration
const typeText = typeNode.text;
if (
typeText !== 'auto' &&
typeText !== 'decltype(auto)' &&
typeNode.type !== 'placeholder_type_specifier'
) return;
const declarator = node.childForFieldName('declarator');
if (!declarator) return;
// Must be an init_declarator (i.e., has an initializer value)
if (declarator.type !== 'init_declarator') return;
const value = declarator.childForFieldName('value');
if (!value) return;
// Resolve the variable name, unwrapping pointer/reference declarators
const nameNode = declarator.childForFieldName('declarator');
if (!nameNode) return;
const finalName =
nameNode.type === 'pointer_declarator' || nameNode.type === 'reference_declarator'
? nameNode.firstNamedChild
: nameNode;
if (!finalName) return;
const varName = extractVarName(finalName);
if (!varName) return;
// auto x = new User() — new_expression
if (value.type === 'new_expression') {
const ctorType = value.childForFieldName('type');
if (ctorType) {
const typeName = extractSimpleTypeName(ctorType);
if (typeName) env.set(varName, typeName);
}
return;
}
// auto x = User() — call_expression where function is a type name
// tree-sitter-cpp may parse the constructor name as type_identifier or identifier.
// For plain identifiers, verify against known class names from the file's AST
// to distinguish constructor calls (User()) from function calls (getUser()).
if (value.type === 'call_expression') {
const func = value.childForFieldName('function');
if (!func) return;
if (func.type === 'type_identifier') {
const typeName = func.text;
if (typeName) env.set(varName, typeName);
} else if (func.type === 'identifier') {
const text = func.text;
if (text && classNames.has(text)) env.set(varName, text);
}
return;
}
// auto x = User{} — compound_literal_expression (brace initialization)
// AST: compound_literal_expression > type_identifier + initializer_list
if (value.type === 'compound_literal_expression') {
const typeId = value.firstNamedChild;
const typeName = typeId ? extractSimpleTypeName(typeId) : undefined;
if (typeName) env.set(varName, typeName);
}
};
/** C/C++: parameter_declaration → type declarator */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
@@ -56,8 +125,240 @@ const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string,
if (varName && typeName) env.set(varName, typeName);
};
/** C/C++: auto x = User() where function is an identifier (not type_identifier) */
const scanConstructorBinding: ConstructorBindingScanner = (node) => {
if (node.type !== 'declaration') return undefined;
const typeNode = node.childForFieldName('type');
if (!typeNode) return undefined;
const typeText = typeNode.text;
if (typeText !== 'auto' && typeText !== 'decltype(auto)' && typeNode.type !== 'placeholder_type_specifier') return undefined;
const declarator = node.childForFieldName('declarator');
if (!declarator || declarator.type !== 'init_declarator') return undefined;
const value = declarator.childForFieldName('value');
if (!value || value.type !== 'call_expression') return undefined;
const func = value.childForFieldName('function');
if (!func) return undefined;
if (func.type === 'qualified_identifier' || func.type === 'scoped_identifier') {
const last = func.lastNamedChild;
if (!last) return undefined;
const nameNode = declarator.childForFieldName('declarator');
if (!nameNode) return undefined;
const finalName = nameNode.type === 'pointer_declarator' || nameNode.type === 'reference_declarator'
? nameNode.firstNamedChild : nameNode;
if (!finalName) return undefined;
return { varName: finalName.text, calleeName: last.text };
}
if (func.type !== 'identifier') return undefined;
const nameNode = declarator.childForFieldName('declarator');
if (!nameNode) return undefined;
const finalName = nameNode.type === 'pointer_declarator' || nameNode.type === 'reference_declarator'
? nameNode.firstNamedChild : nameNode;
if (!finalName) return undefined;
const varName = finalName.text;
if (!varName) return undefined;
return { varName, calleeName: func.text };
};
/** C++: auto alias = user → declaration with auto type + init_declarator where value is identifier */
const extractPendingAssignment: PendingAssignmentExtractor = (node, scopeEnv) => {
if (node.type !== 'declaration') return undefined;
const typeNode = node.childForFieldName('type');
if (!typeNode) return undefined;
// Only handle auto — typed declarations already resolved by extractDeclaration
const typeText = typeNode.text;
if (typeText !== 'auto' && typeText !== 'decltype(auto)'
&& typeNode.type !== 'placeholder_type_specifier') return undefined;
const declarator = node.childForFieldName('declarator');
if (!declarator || declarator.type !== 'init_declarator') return undefined;
const value = declarator.childForFieldName('value');
if (!value || value.type !== 'identifier') return undefined;
const nameNode = declarator.childForFieldName('declarator');
if (!nameNode) return undefined;
const finalName = nameNode.type === 'pointer_declarator' || nameNode.type === 'reference_declarator'
? nameNode.firstNamedChild : nameNode;
if (!finalName) return undefined;
const lhs = extractVarName(finalName);
if (!lhs || scopeEnv.has(lhs)) return undefined;
return { lhs, rhs: value.text };
};
// --- For-loop Tier 1c ---
const FOR_LOOP_NODE_TYPES: ReadonlySet<string> = new Set(['for_range_loop']);
/** Extract template type arguments from a C++ template_type node.
* C++ template_type uses template_argument_list (not type_arguments), and each
* argument is a type_descriptor with a 'type' field containing the type_specifier. */
const extractCppTemplateTypeArgs = (templateTypeNode: SyntaxNode): string[] => {
const argsNode = templateTypeNode.childForFieldName('arguments');
if (!argsNode || argsNode.type !== 'template_argument_list') return [];
const result: string[] = [];
for (let i = 0; i < argsNode.namedChildCount; i++) {
let argNode = argsNode.namedChild(i);
if (!argNode) continue;
// type_descriptor wraps the actual type specifier in a 'type' field
if (argNode.type === 'type_descriptor') {
const inner = argNode.childForFieldName('type');
if (inner) argNode = inner;
}
const name = extractSimpleTypeName(argNode);
if (name) result.push(name);
}
return result;
};
/** Extract element type from a C++ type annotation AST node.
* Handles: template_type (vector<User>, map<string, User>),
* pointer/reference types (User*, User&). */
const extractCppElementTypeFromTypeNode = (typeNode: SyntaxNode, pos: TypeArgPosition = 'last', depth = 0): string | undefined => {
if (depth > 50) return undefined;
// template_type: vector<User>, map<string, User> — extract type arg based on position
if (typeNode.type === 'template_type') {
const args = extractCppTemplateTypeArgs(typeNode);
if (args.length >= 1) return pos === 'first' ? args[0] : args[args.length - 1];
}
// reference/pointer types: unwrap and recurse (vector<User>& → vector<User>)
if (typeNode.type === 'reference_type' || typeNode.type === 'pointer_type'
|| typeNode.type === 'type_descriptor') {
const inner = typeNode.lastNamedChild;
if (inner) return extractCppElementTypeFromTypeNode(inner, pos, depth + 1);
}
// qualified/scoped types: std::vector<User> → unwrap to template_type child
if (typeNode.type === 'qualified_identifier' || typeNode.type === 'scoped_type_identifier') {
const inner = typeNode.lastNamedChild;
if (inner) return extractCppElementTypeFromTypeNode(inner, pos, depth + 1);
}
return undefined;
};
/** Walk up from a for-range-loop to the enclosing function_definition and search parameters
* for one named `iterableName`. Returns the element type from its annotation. */
const findCppParamElementType = (iterableName: string, startNode: SyntaxNode, pos: TypeArgPosition = 'last'): string | undefined => {
let current: SyntaxNode | null = startNode.parent;
while (current) {
if (current.type === 'function_definition') {
const declarator = current.childForFieldName('declarator');
// function_definition > declarator (function_declarator) > parameters (parameter_list)
const paramsNode = declarator?.childForFieldName('parameters');
if (paramsNode) {
for (let i = 0; i < paramsNode.namedChildCount; i++) {
const param = paramsNode.namedChild(i);
if (!param || param.type !== 'parameter_declaration') continue;
const paramDeclarator = param.childForFieldName('declarator');
if (!paramDeclarator) continue;
// Unwrap reference/pointer declarators: vector<User>& users → &users
let identNode = paramDeclarator;
if (identNode.type === 'reference_declarator' || identNode.type === 'pointer_declarator') {
identNode = identNode.firstNamedChild ?? identNode;
}
if (identNode.text !== iterableName) continue;
const typeNode = param.childForFieldName('type');
if (typeNode) return extractCppElementTypeFromTypeNode(typeNode, pos);
}
}
break;
}
current = current.parent;
}
return undefined;
};
/** C++: for (auto& user : users) — extract loop variable binding.
* Handles explicit types (for (User& user : users)) and auto (for (auto& user : users)).
* For auto, resolves element type from the iterable's container type. */
const extractForLoopBinding: ForLoopExtractor = (
node: SyntaxNode,
scopeEnv: Map<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
): void => {
if (node.type !== 'for_range_loop') return;
const typeNode = node.childForFieldName('type');
const declaratorNode = node.childForFieldName('declarator');
const rightNode = node.childForFieldName('right');
if (!typeNode || !declaratorNode || !rightNode) return;
// Unwrap reference/pointer declarator to get the loop variable name
let nameNode = declaratorNode;
if (nameNode.type === 'reference_declarator' || nameNode.type === 'pointer_declarator') {
nameNode = nameNode.firstNamedChild ?? nameNode;
}
// Handle structured bindings: auto& [key, value] or auto [key, value]
// Bind the last identifier (value heuristic for [key, value] patterns)
let loopVarName: string | undefined;
if (nameNode.type === 'structured_binding_declarator') {
const lastChild = nameNode.lastNamedChild;
if (lastChild?.type === 'identifier') {
loopVarName = lastChild.text;
}
} else if (declaratorNode.type === 'structured_binding_declarator') {
const lastChild = declaratorNode.lastNamedChild;
if (lastChild?.type === 'identifier') {
loopVarName = lastChild.text;
}
}
const varName = loopVarName ?? extractVarName(nameNode);
if (!varName) return;
// Check if the type is auto/placeholder — if not, use the explicit type directly
const isAuto = typeNode.type === 'placeholder_type_specifier'
|| typeNode.text === 'auto'
|| typeNode.text === 'const auto'
|| typeNode.text === 'decltype(auto)';
if (!isAuto) {
// Explicit type: for (User& user : users) — extract directly
const typeName = extractSimpleTypeName(typeNode);
if (typeName) scopeEnv.set(varName, typeName);
return;
}
// auto/const auto/auto& — resolve from the iterable's container type
// Extract iterable name + optional method
let iterableName: string | undefined;
let methodName: string | undefined;
if (rightNode.type === 'identifier') {
iterableName = rightNode.text;
} else if (rightNode.type === 'field_expression') {
const prop = rightNode.lastNamedChild;
if (prop) iterableName = prop.text;
} else if (rightNode.type === 'call_expression') {
// users.begin() is NOT used in range-for, but container.items() etc. might be
const fieldExpr = rightNode.childForFieldName('function');
if (fieldExpr?.type === 'field_expression') {
const obj = fieldExpr.firstNamedChild;
if (obj?.type === 'identifier') iterableName = obj.text;
const field = fieldExpr.lastNamedChild;
if (field?.type === 'field_identifier') methodName = field.text;
}
} else if (rightNode.type === 'pointer_expression') {
// Dereference: for (auto& user : *ptr) → pointer_expression > identifier
// Only handles simple *identifier; *this->field and **ptr are not resolved.
const operand = rightNode.lastNamedChild;
if (operand?.type === 'identifier') iterableName = operand.text;
}
if (!iterableName) return;
const containerTypeName = scopeEnv.get(iterableName);
const typeArgPos = methodToTypeArgPosition(methodName, containerTypeName);
const elementType = resolveIterableElementType(
iterableName, node, scopeEnv, declarationTypeNodes, scope,
extractCppElementTypeFromTypeNode, findCppParamElementType,
typeArgPos,
);
if (elementType) scopeEnv.set(varName, elementType);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
forLoopNodeTypes: FOR_LOOP_NODE_TYPES,
extractDeclaration,
extractParameter,
extractInitializer,
scanConstructorBinding,
extractForLoopBinding,
extractPendingAssignment,
};
@@ -1,6 +1,9 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, findChildByType } from './shared.js';
import type { ConstructorBindingScanner, ForLoopExtractor, LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor, PendingAssignmentExtractor, PatternBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, findChildByType, unwrapAwait, extractGenericTypeArgs, resolveIterableElementType, methodToTypeArgPosition, type TypeArgPosition } from './shared.js';
/** Known container property accessors that operate on the container itself (e.g., dict.Keys, dict.Values) */
const KNOWN_CONTAINER_PROPS: ReadonlySet<string> = new Set(['Keys', 'Values']);
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'local_declaration_statement',
@@ -44,7 +47,8 @@ const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<str
let typeName: string | undefined;
if (typeNode.type === 'implicit_type' && typeNode.text === 'var') {
// Try to infer from initializer: var x = new Foo()
// C# tree-sitter puts object_creation_expression as direct child of variable_declarator
// tree-sitter-c-sharp may put object_creation_expression as direct child
// or inside equals_value_clause depending on grammar version
if (declarators.length === 1) {
const initializer = findChildByType(declarators[0], 'object_creation_expression')
?? findChildByType(declarators[0], 'equals_value_clause')?.firstNamedChild;
@@ -86,8 +90,249 @@ const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string,
if (varName && typeName) env.set(varName, typeName);
};
/** C#: var x = SomeFactory(...) → bind x to SomeFactory (constructor-like call) */
const scanConstructorBinding: ConstructorBindingScanner = (node) => {
if (node.type !== 'variable_declaration') return undefined;
// Find type and declarator children by iterating (C# grammar doesn't expose 'type' as a named field)
let typeNode: SyntaxNode | null = null;
let declarator: SyntaxNode | null = null;
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child) continue;
if (child.type === 'variable_declarator') { if (!declarator) declarator = child; }
else if (!typeNode) { typeNode = child; }
}
// Only handle implicit_type (var) — explicit types handled by extractDeclaration
if (!typeNode || typeNode.type !== 'implicit_type') return undefined;
if (!declarator) return undefined;
const nameNode = declarator.childForFieldName('name') ?? declarator.firstNamedChild;
if (!nameNode || nameNode.type !== 'identifier') return undefined;
// Find the initializer value: either inside equals_value_clause or as a direct child
// (tree-sitter-c-sharp puts invocation_expression directly inside variable_declarator)
let value: SyntaxNode | null = null;
for (let i = 0; i < declarator.namedChildCount; i++) {
const child = declarator.namedChild(i);
if (!child) continue;
if (child.type === 'equals_value_clause') { value = child.firstNamedChild; break; }
if (child.type === 'invocation_expression' || child.type === 'object_creation_expression' || child.type === 'await_expression') { value = child; break; }
}
if (!value) return undefined;
// Unwrap await: `var user = await svc.GetUserAsync()` → await_expression wraps invocation_expression
value = unwrapAwait(value);
if (!value) return undefined;
// Skip object_creation_expression (new User()) — handled by extractInitializer
if (value.type === 'object_creation_expression') return undefined;
if (value.type !== 'invocation_expression') return undefined;
const func = value.firstNamedChild;
if (!func) return undefined;
const calleeName = extractSimpleTypeName(func);
if (!calleeName) return undefined;
return { varName: nameNode.text, calleeName };
};
const FOR_LOOP_NODE_TYPES: ReadonlySet<string> = new Set([
'foreach_statement',
]);
/** Extract element type from a C# type annotation AST node.
* Handles generic_name (List<User>), array_type (User[]), nullable_type (?).
* `pos` selects which type arg: 'first' for keys, 'last' for values (default). */
const extractCSharpElementTypeFromTypeNode = (typeNode: SyntaxNode, pos: TypeArgPosition = 'last', depth = 0): string | undefined => {
if (depth > 50) return undefined;
// generic_name: List<User>, IEnumerable<User>, Dictionary<string, User>
// C# uses generic_name (not generic_type)
if (typeNode.type === 'generic_name') {
const argList = findChildByType(typeNode, 'type_argument_list');
if (argList && argList.namedChildCount >= 1) {
if (pos === 'first') {
const firstArg = argList.namedChild(0);
if (firstArg) return extractSimpleTypeName(firstArg);
} else {
const lastArg = argList.namedChild(argList.namedChildCount - 1);
if (lastArg) return extractSimpleTypeName(lastArg);
}
}
}
// array_type: User[]
if (typeNode.type === 'array_type') {
const elemNode = typeNode.firstNamedChild;
if (elemNode) return extractSimpleTypeName(elemNode);
}
// nullable_type: unwrap and recurse (List<User>? → List<User> → User)
if (typeNode.type === 'nullable_type') {
const inner = typeNode.firstNamedChild;
if (inner) return extractCSharpElementTypeFromTypeNode(inner, pos, depth + 1);
}
return undefined;
};
/** Walk up from a foreach to the enclosing method and search parameters. */
const findCSharpParamElementType = (iterableName: string, startNode: SyntaxNode, pos: TypeArgPosition = 'last'): string | undefined => {
let current: SyntaxNode | null = startNode.parent;
while (current) {
if (current.type === 'method_declaration' || current.type === 'local_function_statement') {
const paramsNode = current.childForFieldName('parameters');
if (paramsNode) {
for (let i = 0; i < paramsNode.namedChildCount; i++) {
const param = paramsNode.namedChild(i);
if (!param || param.type !== 'parameter') continue;
const nameNode = param.childForFieldName('name');
if (nameNode?.text !== iterableName) continue;
const typeNode = param.childForFieldName('type');
if (typeNode) return extractCSharpElementTypeFromTypeNode(typeNode, pos);
}
}
break;
}
current = current.parent;
}
return undefined;
};
/** C#: foreach (User user in users) — extract loop variable binding.
* Tier 1c: for `foreach (var user in users)`, resolves element type from iterable. */
const extractForLoopBinding: ForLoopExtractor = (
node: SyntaxNode,
scopeEnv: Map<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
): void => {
const typeNode = node.childForFieldName('type');
const nameNode = node.childForFieldName('left');
if (!typeNode || !nameNode) return;
const varName = extractVarName(nameNode);
if (!varName) return;
// Explicit type (existing behavior): foreach (User user in users)
if (!(typeNode.type === 'implicit_type' && typeNode.text === 'var')) {
const typeName = extractSimpleTypeName(typeNode);
if (typeName) scopeEnv.set(varName, typeName);
return;
}
// Tier 1c: implicit type (var) — resolve from iterable's container type
const rightNode = node.childForFieldName('right');
let iterableName: string | undefined;
let methodName: string | undefined;
if (rightNode?.type === 'identifier') {
iterableName = rightNode.text;
} else if (rightNode?.type === 'member_access_expression') {
// C# property access: data.Keys, data.Values → member_access_expression
// Also handles bare member access: this.users, repo.users → use property as iterableName
const obj = rightNode.childForFieldName('expression');
const prop = rightNode.childForFieldName('name');
const propText = prop?.type === 'identifier' ? prop.text : undefined;
if (propText && KNOWN_CONTAINER_PROPS.has(propText)) {
if (obj?.type === 'identifier') {
iterableName = obj.text;
} else if (obj?.type === 'member_access_expression') {
// Nested member access: this.data.Values → obj is "this.data", extract "data"
const innerProp = obj.childForFieldName('name');
if (innerProp) iterableName = innerProp.text;
}
methodName = propText;
} else if (propText) {
// Bare member access: this.users → use property name for scopeEnv lookup
iterableName = propText;
}
} else if (rightNode?.type === 'invocation_expression') {
// C# method call: data.Select(...) → invocation_expression > member_access_expression
const fn = rightNode.firstNamedChild;
if (fn?.type === 'member_access_expression') {
const obj = fn.childForFieldName('expression');
const prop = fn.childForFieldName('name');
if (obj?.type === 'identifier') iterableName = obj.text;
if (prop?.type === 'identifier') methodName = prop.text;
}
}
if (!iterableName) return;
const containerTypeName = scopeEnv.get(iterableName);
const typeArgPos = methodToTypeArgPosition(methodName, containerTypeName);
const elementType = resolveIterableElementType(
iterableName, node, scopeEnv, declarationTypeNodes, scope,
extractCSharpElementTypeFromTypeNode, findCSharpParamElementType,
typeArgPos,
);
if (elementType) scopeEnv.set(varName, elementType);
};
/**
* C# pattern binding extractor for `obj is Type variable` (type pattern).
*
* AST structure:
* is_pattern_expression
* expression: (the variable being tested)
* pattern: declaration_pattern
* type: (the declared type)
* name: single_variable_designation > identifier (the new variable name)
*
* Conservative: returns undefined when the pattern field is absent, is not a
* declaration_pattern, or when the type/name cannot be extracted.
* No scopeEnv lookup is needed — the pattern explicitly declares the new variable's type.
*/
const extractPatternBinding: PatternBindingExtractor = (node) => {
// is_pattern_expression: `obj is User user` — has a declaration_pattern child
if (node.type === 'is_pattern_expression') {
const pattern = node.childForFieldName('pattern');
if (pattern?.type !== 'declaration_pattern' && pattern?.type !== 'recursive_pattern') return undefined;
const typeNode = pattern.childForFieldName('type');
const nameNode = pattern.childForFieldName('name');
if (!typeNode || !nameNode) return undefined;
const typeName = extractSimpleTypeName(typeNode);
const varName = extractVarName(nameNode);
if (!typeName || !varName) return undefined;
return { varName, typeName };
}
// declaration_pattern / recursive_pattern: standalone in switch statements and switch expressions
// `case User u:` or `User u =>` or `User { Name: "Alice" } u =>`
// Both use the same 'type' and 'name' fields.
if (node.type === 'declaration_pattern' || node.type === 'recursive_pattern') {
const typeNode = node.childForFieldName('type');
const nameNode = node.childForFieldName('name');
if (!typeNode || !nameNode) return undefined;
const typeName = extractSimpleTypeName(typeNode);
const varName = extractVarName(nameNode);
if (!typeName || !varName) return undefined;
return { varName, typeName };
}
return undefined;
};
/** C#: var alias = u → variable_declarator with name + equals_value_clause.
* Only local_declaration_statement and variable_declaration contain variable_declarator children;
* is_pattern_expression and field_declaration never do — skip them early. */
const extractPendingAssignment: PendingAssignmentExtractor = (node, scopeEnv) => {
if (node.type === 'is_pattern_expression' || node.type === 'field_declaration') return undefined;
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child || child.type !== 'variable_declarator') continue;
const nameNode = child.childForFieldName('name');
if (!nameNode) continue;
const lhs = nameNode.text;
if (scopeEnv.has(lhs)) continue;
// C# wraps value in equals_value_clause; fall back to last named child
let evc: SyntaxNode | null = null;
for (let j = 0; j < child.childCount; j++) {
if (child.child(j)?.type === 'equals_value_clause') { evc = child.child(j); break; }
}
const valueNode = evc?.firstNamedChild ?? child.namedChild(child.namedChildCount - 1);
if (valueNode && valueNode !== nameNode && (valueNode.type === 'identifier' || valueNode.type === 'simple_identifier')) {
return { lhs, rhs: valueNode.text };
}
}
return undefined;
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
forLoopNodeTypes: FOR_LOOP_NODE_TYPES,
patternBindingNodeTypes: new Set(['is_pattern_expression', 'declaration_pattern', 'recursive_pattern']),
extractDeclaration,
extractParameter,
scanConstructorBinding,
extractForLoopBinding,
extractPendingAssignment,
extractPatternBinding,
};
@@ -1,6 +1,6 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
import type { ConstructorBindingScanner, ForLoopExtractor, LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor, PendingAssignmentExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, extractElementTypeFromString, extractGenericTypeArgs, findChildByType, resolveIterableElementType, methodToTypeArgPosition, type TypeArgPosition } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'var_declaration',
@@ -59,7 +59,51 @@ const extractGoShortVarDeclaration = (node: SyntaxNode, env: Map<string, string>
// Pair each LHS name with its corresponding RHS value
const count = Math.min(lhsNodes.length, rhsNodes.length);
for (let i = 0; i < count; i++) {
const valueNode = rhsNodes[i];
let valueNode = rhsNodes[i];
// Unwrap &User{} — unary_expression (address-of) wrapping composite_literal
if (valueNode.type === 'unary_expression' && valueNode.firstNamedChild?.type === 'composite_literal') {
valueNode = valueNode.firstNamedChild;
}
// Go built-in new(User) — call_expression with 'new' callee and type argument
// Go built-in make([]User, 0) / make(map[string]User) — extract element/value type
if (valueNode.type === 'call_expression') {
const funcNode = valueNode.childForFieldName('function');
if (funcNode?.text === 'new') {
const args = valueNode.childForFieldName('arguments');
if (args?.firstNamedChild) {
const typeName = extractSimpleTypeName(args.firstNamedChild);
const varName = extractVarName(lhsNodes[i]);
if (varName && typeName) env.set(varName, typeName);
}
} else if (funcNode?.text === 'make') {
const args = valueNode.childForFieldName('arguments');
const firstArg = args?.firstNamedChild;
if (firstArg) {
let innerType: SyntaxNode | null = null;
if (firstArg.type === 'slice_type') {
innerType = firstArg.childForFieldName('element');
} else if (firstArg.type === 'map_type') {
innerType = firstArg.childForFieldName('value');
}
if (innerType) {
const typeName = extractSimpleTypeName(innerType);
const varName = extractVarName(lhsNodes[i]);
if (varName && typeName) env.set(varName, typeName);
}
}
}
continue;
}
// Go type assertion: user := iface.(User) — type_assertion_expression with 'type' field
if (valueNode.type === 'type_assertion_expression') {
const typeNode = valueNode.childForFieldName('type');
if (typeNode) {
const typeName = extractSimpleTypeName(typeNode);
const varName = extractVarName(lhsNodes[i]);
if (varName && typeName) env.set(varName, typeName);
}
continue;
}
if (valueNode.type !== 'composite_literal') continue;
const typeNode = valueNode.childForFieldName('type');
if (!typeNode) continue;
@@ -97,8 +141,277 @@ const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string,
if (varName && typeName) env.set(varName, typeName);
};
/** Go: user := NewUser(...) — infer type from single-assignment call expression */
const scanConstructorBinding: ConstructorBindingScanner = (node) => {
if (node.type !== 'short_var_declaration') return undefined;
const left = node.childForFieldName('left');
const right = node.childForFieldName('right');
if (!left || !right) return undefined;
const leftIds = left.type === 'expression_list' ? left.namedChildren : [left];
const rightExprs = right.type === 'expression_list' ? right.namedChildren : [right];
// Multi-return: user, err := NewUser() — bind first var when second is err/ok/_
if (leftIds.length === 2 && rightExprs.length === 1) {
const secondVar = leftIds[1];
const isErrorOrDiscard =
secondVar.text === '_' ||
secondVar.text === 'err' ||
secondVar.text === 'ok' ||
secondVar.text === 'error';
if (isErrorOrDiscard && leftIds[0].type === 'identifier') {
if (rightExprs[0].type !== 'call_expression') return undefined;
const func = rightExprs[0].childForFieldName('function');
if (!func) return undefined;
if (func.text === 'new' || func.text === 'make') return undefined;
const calleeName = extractSimpleTypeName(func);
if (!calleeName) return undefined;
return { varName: leftIds[0].text, calleeName };
}
}
// Single assignment only
if (leftIds.length !== 1 || leftIds[0].type !== 'identifier') return undefined;
if (rightExprs.length !== 1 || rightExprs[0].type !== 'call_expression') return undefined;
const func = rightExprs[0].childForFieldName('function');
if (!func) return undefined;
// Skip new() and make() — already handled by extractDeclaration
if (func.text === 'new' || func.text === 'make') return undefined;
const calleeName = extractSimpleTypeName(func);
if (!calleeName) return undefined;
return { varName: leftIds[0].text, calleeName };
};
const FOR_LOOP_NODE_TYPES: ReadonlySet<string> = new Set([
'for_statement',
]);
/** Go function/method node types that carry a parameter list. */
const GO_FUNCTION_NODE_TYPES = new Set([
'function_declaration', 'method_declaration', 'func_literal',
]);
/**
* Extract element type from a Go type annotation AST node.
* Handles:
* slice_type "[]User" → element field → type_identifier "User"
* array_type "[10]User" → element field → type_identifier "User"
* Falls back to text-based extraction via extractElementTypeFromString.
*/
const extractGoElementTypeFromTypeNode = (typeNode: SyntaxNode, pos: TypeArgPosition = 'last'): string | undefined => {
// slice_type: []User — element field is the element type
if (typeNode.type === 'slice_type' || typeNode.type === 'array_type') {
const elemNode = typeNode.childForFieldName('element');
if (elemNode) return extractSimpleTypeName(elemNode);
}
// map_type: map[string]User — value field is the element type (for range, second var gets value)
if (typeNode.type === 'map_type') {
const valueNode = typeNode.childForFieldName('value');
if (valueNode) return extractSimpleTypeName(valueNode);
}
// channel_type: chan User — the type argument is the element type
if (typeNode.type === 'channel_type') {
const valueNode = typeNode.childForFieldName('value') ?? typeNode.lastNamedChild;
if (valueNode) return extractSimpleTypeName(valueNode);
}
// generic_type: Go 1.18+ generics (e.g., MySlice[User], Cache[string, User])
// Use position-aware arg selection: 'first' for keys, 'last' for values.
if (typeNode.type === 'generic_type') {
const args = extractGenericTypeArgs(typeNode);
if (args.length >= 1) return pos === 'first' ? args[0] : args[args.length - 1];
}
// Fallback: text-based extraction ([]User → User, User[] → User)
return extractElementTypeFromString(typeNode.text, pos);
};
/** Check if a Go type node represents a channel type. Used to determine
* whether single-var range yields the element (channels) vs index (slices/maps). */
const isChannelType = (
iterableName: string,
scopeEnv: ReadonlyMap<string, string>,
declarationTypeNodes?: ReadonlyMap<string, SyntaxNode>,
scope?: string,
): boolean => {
if (declarationTypeNodes && scope) {
const typeNode = declarationTypeNodes.get(`${scope}\0${iterableName}`);
if (typeNode) return typeNode.type === 'channel_type';
}
const t = scopeEnv.get(iterableName);
return !!t && t.startsWith('chan ');
};
/**
* Walk up the AST from a for-statement to find the enclosing function declaration,
* then search its parameters for one named `iterableName`.
* Returns the element type extracted from its type annotation, or undefined.
*
* Go parameter_declaration has:
* name field: identifier (the parameter name)
* type field: the type node (slice_type for []User)
*/
const findGoParamElementType = (iterableName: string, startNode: SyntaxNode, pos: TypeArgPosition = 'last'): string | undefined => {
let current: SyntaxNode | null = startNode.parent;
while (current) {
if (GO_FUNCTION_NODE_TYPES.has(current.type)) {
const paramsNode = current.childForFieldName('parameters');
if (paramsNode) {
for (let i = 0; i < paramsNode.namedChildCount; i++) {
const paramDecl = paramsNode.namedChild(i);
if (!paramDecl || paramDecl.type !== 'parameter_declaration') continue;
// parameter_declaration: name type — name field is the identifier
const nameNode = paramDecl.childForFieldName('name');
if (nameNode?.text === iterableName) {
const typeNode = paramDecl.childForFieldName('type');
if (typeNode) return extractGoElementTypeFromTypeNode(typeNode, pos);
}
}
}
break;
}
current = current.parent;
}
return undefined;
};
/**
* Go: for _, user := range users where users has a known slice type.
*
* Go uses a single `for_statement` node for all for-loop forms. We detect
* range-based loops by looking for a `range_clause` child node. C-style for
* loops (with `for_clause`) and infinite loops (no clause) are ignored.
*
* Tier 1c: resolves the element type via three strategies in priority order:
* 1. declarationTypeNodes — raw type annotation AST node
* 2. scopeEnv string — extractElementTypeFromString on the stored type
* 3. AST walk — walks up to the enclosing function's parameters to read []User directly
* For `_, user := range users`, the loop variable is the second identifier in
* the `left` expression_list (index is discarded, value is the element).
*/
const extractForLoopBinding: ForLoopExtractor = (
node: SyntaxNode,
scopeEnv: Map<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
): void => {
if (node.type !== 'for_statement') return;
// Find the range_clause child — this distinguishes range loops from other for forms.
let rangeClause: SyntaxNode | null = null;
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'range_clause') {
rangeClause = child;
break;
}
}
if (!rangeClause) return;
// The iterable is the `right` field of the range_clause.
const rightNode = rangeClause.childForFieldName('right');
let iterableName: string | undefined;
if (rightNode?.type === 'identifier') {
iterableName = rightNode.text;
} else if (rightNode?.type === 'selector_expression') {
const field = rightNode.childForFieldName('field');
if (field) iterableName = field.text;
}
if (!iterableName) return;
const containerTypeName = scopeEnv.get(iterableName);
const typeArgPos = methodToTypeArgPosition(undefined, containerTypeName);
const elementType = resolveIterableElementType(
iterableName, node, scopeEnv, declarationTypeNodes, scope,
extractGoElementTypeFromTypeNode, findGoParamElementType,
typeArgPos,
);
if (!elementType) return;
// The loop variable(s) are in the `left` field.
// Go range semantics:
// Slice/Array/String: single-var → INDEX (int); two-var → (index, element)
// Map: single-var → KEY; two-var → (key, value)
// Channel: single-var → ELEMENT (channels have no index)
const leftNode = rangeClause.childForFieldName('left');
if (!leftNode) return;
let loopVarNode: SyntaxNode | null = null;
if (leftNode.type === 'expression_list') {
if (leftNode.namedChildCount >= 2) {
// Two-var form: `_, user` or `i, user` — second variable gets element/value type
loopVarNode = leftNode.namedChild(1);
} else {
// Single-var in expression_list — yields INDEX for slices/maps, ELEMENT for channels
if (isChannelType(iterableName, scopeEnv, declarationTypeNodes, scope)) {
loopVarNode = leftNode.namedChild(0);
} else {
return; // index-only range on slice/map — skip
}
}
} else {
// Plain identifier (single-var form without expression_list)
if (isChannelType(iterableName, scopeEnv, declarationTypeNodes, scope)) {
loopVarNode = leftNode;
} else {
return; // index-only range on slice/map — skip
}
}
if (!loopVarNode) return;
// Skip the blank identifier `_`
if (loopVarNode.text === '_') return;
const loopVarName = extractVarName(loopVarNode);
if (loopVarName) scopeEnv.set(loopVarName, elementType);
};
/** Go: alias := u (short_var_declaration) or var b = u (var_spec) */
const extractPendingAssignment: PendingAssignmentExtractor = (node, scopeEnv) => {
if (node.type === 'short_var_declaration') {
const left = node.childForFieldName('left');
const right = node.childForFieldName('right');
if (!left || !right) return undefined;
const lhsNode = left.type === 'expression_list' ? left.firstNamedChild : left;
const rhsNode = right.type === 'expression_list' ? right.firstNamedChild : right;
if (!lhsNode || !rhsNode) return undefined;
if (lhsNode.type !== 'identifier') return undefined;
const lhs = lhsNode.text;
if (scopeEnv.has(lhs)) return undefined;
if (rhsNode.type === 'identifier') return { lhs, rhs: rhsNode.text };
return undefined;
}
if (node.type === 'var_spec' || node.type === 'var_declaration') {
// var_declaration contains var_spec children; var_spec has name + expression_list value
const specs: SyntaxNode[] = [];
if (node.type === 'var_declaration') {
for (let i = 0; i < node.namedChildCount; i++) {
const c = node.namedChild(i);
if (c?.type === 'var_spec') specs.push(c);
}
} else {
specs.push(node);
}
for (const spec of specs) {
const nameNode = spec.childForFieldName('name');
if (!nameNode || nameNode.type !== 'identifier') continue;
const lhs = nameNode.text;
if (scopeEnv.has(lhs)) continue;
// Check if the last named child is a bare identifier (no type annotation between name and value)
let exprList: SyntaxNode | null = null;
for (let i = 0; i < spec.childCount; i++) {
if (spec.child(i)?.type === 'expression_list') { exprList = spec.child(i); break; }
}
const rhsNode = exprList?.firstNamedChild;
if (rhsNode?.type === 'identifier') return { lhs, rhs: rhsNode.text };
}
}
return undefined;
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
forLoopNodeTypes: FOR_LOOP_NODE_TYPES,
extractDeclaration,
extractParameter,
scanConstructorBinding,
extractForLoopBinding,
extractPendingAssignment,
};
@@ -15,6 +15,7 @@ import { typeConfig as pythonConfig } from './python.js';
import { typeConfig as swiftConfig } from './swift.js';
import { typeConfig as cCppConfig } from './c-cpp.js';
import { typeConfig as phpConfig } from './php.js';
import { typeConfig as rubyConfig } from './ruby.js';
export const typeConfigs = {
[SupportedLanguages.JavaScript]: typescriptConfig,
@@ -29,7 +30,23 @@ export const typeConfigs = {
[SupportedLanguages.C]: cCppConfig,
[SupportedLanguages.CPlusPlus]: cCppConfig,
[SupportedLanguages.PHP]: phpConfig,
[SupportedLanguages.Ruby]: rubyConfig,
} satisfies Record<SupportedLanguages, LanguageTypeConfig>;
export type { LanguageTypeConfig, TypeBindingExtractor, ParameterExtractor } from './types.js';
export { TYPED_PARAMETER_TYPES, extractSimpleTypeName, extractVarName, findChildByType } from './shared.js';
export type {
LanguageTypeConfig,
TypeBindingExtractor,
ParameterExtractor,
ConstructorBindingScanner,
ForLoopExtractor,
PendingAssignmentExtractor,
PatternBindingExtractor,
} from './types.js';
export {
TYPED_PARAMETER_TYPES,
extractSimpleTypeName,
extractGenericTypeArgs,
extractVarName,
findChildByType,
extractRubyConstructorAssignment
} from './shared.js';
@@ -1,6 +1,6 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, findChildByType } from './shared.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor, InitializerExtractor, ClassNameLookup, ConstructorBindingScanner, ForLoopExtractor, PendingAssignmentExtractor, PatternBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, findChildByType, extractGenericTypeArgs, resolveIterableElementType, methodToTypeArgPosition, type TypeArgPosition } from './shared.js';
// ── Java ──────────────────────────────────────────────────────────────────
@@ -14,7 +14,7 @@ const extractJavaDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map
const typeNode = node.childForFieldName('type');
if (!typeNode) return;
const typeName = extractSimpleTypeName(typeNode);
if (!typeName) return;
if (!typeName || typeName === 'var') return; // skip Java 10 var — handled by extractInitializer
// Find variable_declarator children
for (let i = 0; i < node.namedChildCount; i++) {
@@ -28,6 +28,25 @@ const extractJavaDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map
}
};
/** Java 10+: var x = new User() — infer type from object_creation_expression */
const extractJavaInitializer: InitializerExtractor = (node: SyntaxNode, env: Map<string, string>, _classNames: ClassNameLookup): void => {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type !== 'variable_declarator') continue;
const nameNode = child.childForFieldName('name');
const valueNode = child.childForFieldName('value');
if (!nameNode || !valueNode) continue;
// Skip declarators that already have a binding from extractDeclaration
const varName = extractVarName(nameNode);
if (!varName || env.has(varName)) continue;
if (valueNode.type !== 'object_creation_expression') continue;
const ctorType = valueNode.childForFieldName('type');
if (!ctorType) continue;
const typeName = extractSimpleTypeName(ctorType);
if (typeName) env.set(varName, typeName);
}
};
/** Java: formal_parameter → type name */
const extractJavaParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
@@ -48,10 +67,188 @@ const extractJavaParameter: ParameterExtractor = (node: SyntaxNode, env: Map<str
if (varName && typeName) env.set(varName, typeName);
};
/** Java: var x = SomeFactory.create() — constructor binding for `var` with method_invocation */
const scanJavaConstructorBinding: ConstructorBindingScanner = (node) => {
if (node.type !== 'local_variable_declaration') return undefined;
const typeNode = node.childForFieldName('type');
if (!typeNode) return undefined;
if (typeNode.text !== 'var') return undefined;
const declarator = findChildByType(node, 'variable_declarator');
if (!declarator) return undefined;
const nameNode = declarator.childForFieldName('name');
const value = declarator.childForFieldName('value');
if (!nameNode || !value) return undefined;
if (value.type === 'object_creation_expression') return undefined;
if (value.type !== 'method_invocation') return undefined;
const methodName = value.childForFieldName('name');
if (!methodName) return undefined;
return { varName: nameNode.text, calleeName: methodName.text };
};
const JAVA_FOR_LOOP_NODE_TYPES: ReadonlySet<string> = new Set([
'enhanced_for_statement',
]);
/** Extract element type from a Java type annotation AST node.
* Handles generic_type (List<User>), array_type (User[]). */
const extractJavaElementTypeFromTypeNode = (typeNode: SyntaxNode, pos: TypeArgPosition = 'last'): string | undefined => {
if (typeNode.type === 'generic_type') {
const args = extractGenericTypeArgs(typeNode);
if (args.length >= 1) return pos === 'first' ? args[0] : args[args.length - 1];
}
if (typeNode.type === 'array_type') {
const elemNode = typeNode.firstNamedChild;
if (elemNode) return extractSimpleTypeName(elemNode);
}
return undefined;
};
/** Walk up from a for-each to the enclosing method_declaration and search parameters. */
const findJavaParamElementType = (iterableName: string, startNode: SyntaxNode, pos: TypeArgPosition = 'last'): string | undefined => {
let current: SyntaxNode | null = startNode.parent;
while (current) {
if (current.type === 'method_declaration' || current.type === 'constructor_declaration') {
const paramsNode = current.childForFieldName('parameters');
if (paramsNode) {
for (let i = 0; i < paramsNode.namedChildCount; i++) {
const param = paramsNode.namedChild(i);
if (!param || param.type !== 'formal_parameter') continue;
const nameNode = param.childForFieldName('name');
if (nameNode?.text !== iterableName) continue;
const typeNode = param.childForFieldName('type');
if (typeNode) return extractJavaElementTypeFromTypeNode(typeNode, pos);
}
}
break;
}
current = current.parent;
}
return undefined;
};
/** Java: for (User user : users) — extract loop variable binding.
* Tier 1c: for `for (var user : users)`, resolves element type from iterable. */
const extractJavaForLoopBinding: ForLoopExtractor = (
node: SyntaxNode,
scopeEnv: Map<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
): void => {
const typeNode = node.childForFieldName('type');
const nameNode = node.childForFieldName('name');
if (!typeNode || !nameNode) return;
const varName = extractVarName(nameNode);
if (!varName) return;
// Explicit type (existing behavior): for (User user : users)
const typeName = extractSimpleTypeName(typeNode);
if (typeName && typeName !== 'var') {
scopeEnv.set(varName, typeName);
return;
}
// Tier 1c: var — resolve from iterable's container type
const iterableNode = node.childForFieldName('value');
if (!iterableNode) return;
let iterableName: string | undefined;
let methodName: string | undefined;
if (iterableNode.type === 'identifier') {
iterableName = iterableNode.text;
} else if (iterableNode.type === 'field_access') {
const field = iterableNode.childForFieldName('field');
if (field) iterableName = field.text;
} else if (iterableNode.type === 'method_invocation') {
// data.keySet() → method_invocation > object: identifier + name: identifier
// Also handles this.data.values() → object is field_access, extract inner field name
const obj = iterableNode.childForFieldName('object');
const name = iterableNode.childForFieldName('name');
if (obj?.type === 'identifier') {
iterableName = obj.text;
} else if (obj?.type === 'field_access') {
const innerField = obj.childForFieldName('field');
if (innerField) iterableName = innerField.text;
}
if (name) methodName = name.text;
}
if (!iterableName) return;
const containerTypeName = scopeEnv.get(iterableName);
const typeArgPos = methodToTypeArgPosition(methodName, containerTypeName);
const elementType = resolveIterableElementType(
iterableName, node, scopeEnv, declarationTypeNodes, scope,
extractJavaElementTypeFromTypeNode, findJavaParamElementType,
typeArgPos,
);
if (elementType) scopeEnv.set(varName, elementType);
};
/** Java: var alias = u → local_variable_declaration > variable_declarator with name/value */
const extractJavaPendingAssignment: PendingAssignmentExtractor = (node, scopeEnv) => {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child || child.type !== 'variable_declarator') continue;
const nameNode = child.childForFieldName('name');
const valueNode = child.childForFieldName('value');
if (!nameNode || !valueNode) continue;
const lhs = nameNode.text;
if (scopeEnv.has(lhs)) continue;
if (valueNode.type === 'identifier' || valueNode.type === 'simple_identifier') return { lhs, rhs: valueNode.text };
}
return undefined;
};
/**
* Java 16+ `instanceof` pattern variable: `x instanceof User user`
*
* AST structure:
* instanceof_expression
* left: expression (the variable being tested)
* instanceof keyword
* right: type (the type to test against)
* name: identifier (the pattern variable — optional, Java 16+)
*
* Conservative: returns undefined when the `name` field is absent (plain instanceof
* without pattern variable, e.g. `x instanceof User`) or when the type cannot be
* extracted. The source variable's existing type is NOT used — the pattern explicitly
* declares the new type, so no scopeEnv lookup is needed.
*/
const extractJavaPatternBinding: PatternBindingExtractor = (node) => {
if (node.type === 'type_pattern') {
// Java 17+ switch pattern: case User u -> ...
// type_pattern has positional children (NO named fields):
// namedChild(0) = type (type_identifier, e.g., User)
// namedChild(1) = identifier (e.g., u)
const typeNode = node.namedChild(0);
const nameNode = node.namedChild(1);
if (!typeNode || !nameNode) return undefined;
const typeName = extractSimpleTypeName(typeNode);
const varName = extractVarName(nameNode);
if (!typeName || !varName) return undefined;
return { varName, typeName };
}
if (node.type !== 'instanceof_expression') return undefined;
const nameNode = node.childForFieldName('name');
if (!nameNode) return undefined;
const typeNode = node.childForFieldName('right');
if (!typeNode) return undefined;
const typeName = extractSimpleTypeName(typeNode);
const varName = extractVarName(nameNode);
if (!typeName || !varName) return undefined;
return { varName, typeName };
};
export const javaTypeConfig: LanguageTypeConfig = {
declarationNodeTypes: JAVA_DECLARATION_NODE_TYPES,
forLoopNodeTypes: JAVA_FOR_LOOP_NODE_TYPES,
patternBindingNodeTypes: new Set(['instanceof_expression', 'type_pattern']),
extractDeclaration: extractJavaDeclaration,
extractParameter: extractJavaParameter,
extractInitializer: extractJavaInitializer,
scanConstructorBinding: scanJavaConstructorBinding,
extractForLoopBinding: extractJavaForLoopBinding,
extractPendingAssignment: extractJavaPendingAssignment,
extractPatternBinding: extractJavaPatternBinding,
};
// ── Kotlin ────────────────────────────────────────────────────────────────
@@ -96,7 +293,10 @@ const extractKotlinDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: M
}
};
/** Kotlin: formal_parameter → type name */
/** Kotlin: parameter / formal_parameter → type name.
* Kotlin's tree-sitter grammar uses positional children (simple_identifier, user_type)
* rather than named fields (name, type) on `parameter` nodes, so we fall back to
* findChildByType when childForFieldName returns null. */
const extractKotlinParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
let typeNode: SyntaxNode | null = null;
@@ -109,14 +309,301 @@ const extractKotlinParameter: ParameterExtractor = (node: SyntaxNode, env: Map<s
typeNode = node.childForFieldName('type');
}
// Fallback: Kotlin `parameter` nodes use positional children, not named fields
if (!nameNode) nameNode = findChildByType(node, 'simple_identifier');
if (!typeNode) typeNode = findChildByType(node, 'user_type');
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
};
/** Kotlin: val user = User() — infer type from call_expression when callee is a known class.
* Kotlin constructors are syntactically identical to function calls, so we verify
* against classNames (which may include cross-file SymbolTable lookups). */
const extractKotlinInitializer: InitializerExtractor = (node: SyntaxNode, env: Map<string, string>, classNames: ClassNameLookup): void => {
if (node.type !== 'property_declaration') return;
// Skip if there's an explicit type annotation — Tier 0 already handled it
const varDecl = findChildByType(node, 'variable_declaration');
if (varDecl && findChildByType(varDecl, 'user_type')) return;
// Get the initializer value — the call_expression after '='
const value = node.childForFieldName('value')
?? findChildByType(node, 'call_expression');
if (!value || value.type !== 'call_expression') return;
// The callee is the first child of call_expression (simple_identifier for direct calls)
const callee = value.firstNamedChild;
if (!callee || callee.type !== 'simple_identifier') return;
const calleeName = callee.text;
if (!calleeName || !classNames.has(calleeName)) return;
// Extract the variable name from the variable_declaration inside property_declaration
const nameNode = varDecl
? findChildByType(varDecl, 'simple_identifier')
: findChildByType(node, 'simple_identifier');
if (!nameNode) return;
const varName = extractVarName(nameNode);
if (varName) env.set(varName, calleeName);
};
/** Kotlin: val x = User(...) — constructor binding for property_declaration with call_expression */
const scanKotlinConstructorBinding: ConstructorBindingScanner = (node) => {
if (node.type !== 'property_declaration') return undefined;
const varDecl = findChildByType(node, 'variable_declaration');
if (!varDecl) return undefined;
if (findChildByType(varDecl, 'user_type')) return undefined;
const callExpr = findChildByType(node, 'call_expression');
if (!callExpr) return undefined;
const callee = callExpr.firstNamedChild;
if (!callee) return undefined;
let calleeName: string | undefined;
if (callee.type === 'simple_identifier') {
calleeName = callee.text;
} else if (callee.type === 'navigation_expression') {
// Extract method name from qualified call: service.getUser() → getUser
const suffix = callee.lastNamedChild;
if (suffix?.type === 'navigation_suffix') {
const methodName = suffix.lastNamedChild;
if (methodName?.type === 'simple_identifier') {
calleeName = methodName.text;
}
}
}
if (!calleeName) return undefined;
const nameNode = findChildByType(varDecl, 'simple_identifier');
if (!nameNode) return undefined;
return { varName: nameNode.text, calleeName };
};
const KOTLIN_FOR_LOOP_NODE_TYPES: ReadonlySet<string> = new Set([
'for_statement',
]);
/** Extract element type from a Kotlin type annotation AST node (user_type wrapping generic).
* Kotlin: user_type → [type_identifier, type_arguments → [type_projection → user_type]]
* Handles the type_projection wrapper that Kotlin uses for generic type arguments. */
const extractKotlinElementTypeFromTypeNode = (typeNode: SyntaxNode, pos: TypeArgPosition = 'last'): string | undefined => {
if (typeNode.type === 'user_type') {
const argsNode = findChildByType(typeNode, 'type_arguments');
if (argsNode && argsNode.namedChildCount >= 1) {
const targetArg = pos === 'first'
? argsNode.namedChild(0)
: argsNode.namedChild(argsNode.namedChildCount - 1);
if (!targetArg) return undefined;
// Kotlin wraps type args in type_projection — unwrap to get the inner type
const inner = targetArg.type === 'type_projection'
? targetArg.firstNamedChild
: targetArg;
if (inner) return extractSimpleTypeName(inner);
}
}
return undefined;
};
/** Walk up from a for-loop to the enclosing function_declaration and search parameters.
* Kotlin parameters use positional children (simple_identifier, user_type), not named fields. */
const findKotlinParamElementType = (iterableName: string, startNode: SyntaxNode, pos: TypeArgPosition = 'last'): string | undefined => {
let current: SyntaxNode | null = startNode.parent;
while (current) {
if (current.type === 'function_declaration') {
const paramsNode = findChildByType(current, 'function_value_parameters');
if (paramsNode) {
for (let i = 0; i < paramsNode.namedChildCount; i++) {
const param = paramsNode.namedChild(i);
if (!param || param.type !== 'parameter') continue;
const nameNode = findChildByType(param, 'simple_identifier');
if (nameNode?.text !== iterableName) continue;
const typeNode = findChildByType(param, 'user_type');
if (typeNode) return extractKotlinElementTypeFromTypeNode(typeNode, pos);
}
}
break;
}
current = current.parent;
}
return undefined;
};
/** Kotlin: for (user: User in users) — extract loop variable binding.
* Tier 1c: for `for (user in users)` without annotation, resolves from iterable. */
const extractKotlinForLoopBinding: ForLoopExtractor = (
node: SyntaxNode,
scopeEnv: Map<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
): void => {
const varDecl = findChildByType(node, 'variable_declaration');
if (!varDecl) return;
const nameNode = findChildByType(varDecl, 'simple_identifier');
if (!nameNode) return;
const varName = extractVarName(nameNode);
if (!varName) return;
// Explicit type annotation (existing behavior): for (user: User in users)
const typeNode = findChildByType(varDecl, 'user_type');
if (typeNode) {
const typeName = extractSimpleTypeName(typeNode);
if (typeName) scopeEnv.set(varName, typeName);
return;
}
// Tier 1c: no annotation — resolve from iterable's container type
// Kotlin for-loop children: [variable_declaration, iterable_expr, control_structure_body]
// The iterable is the second named child of the for_statement (after variable_declaration)
let iterableName: string | undefined;
let methodName: string | undefined;
let fallbackIterableName: string | undefined;
let foundVarDecl = false;
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child === varDecl) { foundVarDecl = true; continue; }
if (!foundVarDecl || !child) continue;
if (child.type === 'simple_identifier') {
iterableName = child.text;
break;
}
if (child.type === 'navigation_expression') {
// data.keys → navigation_expression > simple_identifier(data) + navigation_suffix > simple_identifier(keys)
const obj = child.firstNamedChild;
const suffix = findChildByType(child, 'navigation_suffix');
const prop = suffix ? findChildByType(suffix, 'simple_identifier') : null;
const hasCallSuffix = suffix ? findChildByType(suffix, 'call_suffix') !== null : false;
// Always try object as iterable + property as method first (handles data.values, data.keys).
// For bare property access without call_suffix, also save property as fallback
// (handles this.users, repo.items where the property IS the iterable).
if (obj?.type === 'simple_identifier') iterableName = obj.text;
if (prop) methodName = prop.text;
if (!hasCallSuffix && prop) {
fallbackIterableName = prop.text;
}
break;
}
if (child.type === 'call_expression') {
// data.values() → call_expression > navigation_expression > simple_identifier + navigation_suffix
const callee = child.firstNamedChild;
if (callee?.type === 'navigation_expression') {
const obj = callee.firstNamedChild;
if (obj?.type === 'simple_identifier') iterableName = obj.text;
const suffix = findChildByType(callee, 'navigation_suffix');
if (suffix) {
const prop = findChildByType(suffix, 'simple_identifier');
if (prop) methodName = prop.text;
}
}
break;
}
}
if (!iterableName) return;
let containerTypeName = scopeEnv.get(iterableName);
// Fallback: if object has no type in scope, try the property as the iterable name.
// Handles patterns like this.users where the property itself is the iterable variable.
if (!containerTypeName && fallbackIterableName) {
iterableName = fallbackIterableName;
methodName = undefined;
containerTypeName = scopeEnv.get(iterableName);
}
const typeArgPos = methodToTypeArgPosition(methodName, containerTypeName);
const elementType = resolveIterableElementType(
iterableName, node, scopeEnv, declarationTypeNodes, scope,
extractKotlinElementTypeFromTypeNode, findKotlinParamElementType,
typeArgPos,
);
if (elementType) scopeEnv.set(varName, elementType);
};
/** Kotlin: val alias = u → property_declaration or variable_declaration.
* property_declaration has: binding_pattern_kind("val"), variable_declaration("alias"),
* "=", and the RHS value (simple_identifier "u").
* variable_declaration appears directly inside functions and has simple_identifier children. */
const extractKotlinPendingAssignment: PendingAssignmentExtractor = (node, scopeEnv) => {
if (node.type === 'property_declaration') {
// Find the variable name from variable_declaration child
const varDecl = findChildByType(node, 'variable_declaration');
if (!varDecl) return undefined;
const nameNode = varDecl.firstNamedChild;
if (!nameNode || nameNode.type !== 'simple_identifier') return undefined;
const lhs = nameNode.text;
if (scopeEnv.has(lhs)) return undefined;
// Find the RHS: a simple_identifier sibling after the "=" token
let foundEq = false;
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (!child) continue;
if (child.type === '=') { foundEq = true; continue; }
if (foundEq && child.type === 'simple_identifier') {
return { lhs, rhs: child.text };
}
}
return undefined;
}
if (node.type === 'variable_declaration') {
// variable_declaration directly inside functions: simple_identifier children
const nameNode = findChildByType(node, 'simple_identifier');
if (!nameNode) return undefined;
const lhs = nameNode.text;
if (scopeEnv.has(lhs)) return undefined;
// Look for RHS simple_identifier after "=" in the parent (property_declaration)
// variable_declaration itself doesn't contain "=" — it's in the parent
const parent = node.parent;
if (!parent) return undefined;
let foundEq = false;
for (let i = 0; i < parent.childCount; i++) {
const child = parent.child(i);
if (!child) continue;
if (child.type === '=') { foundEq = true; continue; }
if (foundEq && child.type === 'simple_identifier') {
return { lhs, rhs: child.text };
}
}
return undefined;
}
return undefined;
};
/** Walk up from a node to find an ancestor of a given type. */
const findAncestorByType = (node: SyntaxNode, type: string): SyntaxNode | undefined => {
let current = node.parent;
while (current) {
if (current.type === type) return current;
current = current.parent;
}
return undefined;
};
const extractKotlinPatternBinding: PatternBindingExtractor = (node) => {
if (node.type !== 'type_test') return undefined;
const typeNode = node.lastNamedChild;
if (!typeNode) return undefined;
const typeName = extractSimpleTypeName(typeNode);
if (!typeName) return undefined;
const whenExpr = findAncestorByType(node, 'when_expression');
if (!whenExpr) return undefined;
const whenSubject = whenExpr.namedChild(0);
const subject = whenSubject?.firstNamedChild ?? whenSubject;
if (!subject) return undefined;
const varName = extractVarName(subject);
if (!varName) return undefined;
return { varName, typeName };
};
export const kotlinTypeConfig: LanguageTypeConfig = {
allowPatternBindingOverwrite: true,
declarationNodeTypes: KOTLIN_DECLARATION_NODE_TYPES,
forLoopNodeTypes: KOTLIN_FOR_LOOP_NODE_TYPES,
patternBindingNodeTypes: new Set(['type_test']),
extractDeclaration: extractKotlinDeclaration,
extractParameter: extractKotlinParameter,
extractInitializer: extractKotlinInitializer,
scanConstructorBinding: scanKotlinConstructorBinding,
extractForLoopBinding: extractKotlinForLoopBinding,
extractPendingAssignment: extractKotlinPendingAssignment,
extractPatternBinding: extractKotlinPatternBinding,
};
@@ -1,13 +1,187 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor, InitializerExtractor, ClassNameLookup, ConstructorBindingScanner, ReturnTypeExtractor, PendingAssignmentExtractor, ForLoopExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, extractCalleeName, resolveIterableElementType, extractElementTypeFromString } from './shared.js';
// PHP has no local variable type annotations; only params carry types
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set<string>();
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'assignment_expression', // For constructor inference: $x = new User()
'property_declaration', // PHP 7.4+ typed properties: private UserRepo $repo;
'method_declaration', // PHPDoc @param on class methods
'function_definition', // PHPDoc @param on top-level functions
]);
/** PHP: no typed local variable declarations */
const extractDeclaration: TypeBindingExtractor = (_node: SyntaxNode, _env: Map<string, string>): void => {
// PHP has no local variable type annotations
/** Walk up the AST to find the enclosing class declaration. */
const findEnclosingClass = (node: SyntaxNode): SyntaxNode | null => {
let current = node.parent;
while (current) {
if (current.type === 'class_declaration') return current;
current = current.parent;
}
return null;
};
/**
* Resolve PHP self/static/parent to the actual class name.
* - self/static → enclosing class name
* - parent → superclass from base_clause
*/
const resolvePhpKeyword = (keyword: string, node: SyntaxNode): string | undefined => {
if (keyword === 'self' || keyword === 'static') {
const cls = findEnclosingClass(node);
if (!cls) return undefined;
const nameNode = cls.childForFieldName('name');
return nameNode?.text;
}
if (keyword === 'parent') {
const cls = findEnclosingClass(node);
if (!cls) return undefined;
// base_clause contains the parent class name
for (let i = 0; i < cls.namedChildCount; i++) {
const child = cls.namedChild(i);
if (child?.type === 'base_clause') {
const parentName = child.firstNamedChild;
if (parentName) return extractSimpleTypeName(parentName);
}
}
return undefined;
}
return undefined;
};
const normalizePhpType = (raw: string): string | undefined => {
// Strip nullable prefix: ?User → User
let type = raw.startsWith('?') ? raw.slice(1) : raw;
// Strip array suffix: User[] → User
type = type.replace(/\[\]$/, '');
// Strip union with null/false/void: User|null → User
const parts = type.split('|').filter(p => p !== 'null' && p !== 'false' && p !== 'void' && p !== 'mixed');
if (parts.length !== 1) return undefined;
type = parts[0];
// Strip namespace: \App\Models\User → User
const segments = type.split('\\');
type = segments[segments.length - 1];
// Skip uninformative types
if (type === 'mixed' || type === 'void' || type === 'self' || type === 'static' || type === 'object') return undefined;
// Extract element type from generic: Collection<User> → User
// PHPDoc generics encode the element type in angle brackets. Since PHP's Strategy B
// uses the scopeEnv value directly as the element type, we must store the inner type,
// not the container name. This mirrors how User[] → User is handled by the [] strip above.
const genericMatch = type.match(/^(\w+)\s*</);
if (genericMatch) {
const elementType = extractElementTypeFromString(type);
return elementType ?? undefined;
}
if (/^\w+$/.test(type)) return type;
return undefined;
};
/** Node types to skip when walking backwards to find doc-comments.
* PHP 8+ attributes (#[Route(...)]) appear as named siblings between PHPDoc and method. */
const SKIP_NODE_TYPES: ReadonlySet<string> = new Set(['attribute_list', 'attribute']);
/** Regex to extract PHPDoc @param annotations: `@param Type $name` (standard order) */
const PHPDOC_PARAM_RE = /@param\s+(\S+)\s+\$(\w+)/g;
/** Alternate PHPDoc order: `@param $name Type` (name first) */
const PHPDOC_PARAM_ALT_RE = /@param\s+\$(\w+)\s+(\S+)/g;
/**
* Collect PHPDoc @param type bindings from comment nodes preceding a method/function.
* Returns a map of paramName → typeName (without $ prefix).
*/
const collectPhpDocParams = (methodNode: SyntaxNode): Map<string, string> => {
const commentTexts: string[] = [];
let sibling = methodNode.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
commentTexts.unshift(sibling.text);
} else if (sibling.isNamed && !SKIP_NODE_TYPES.has(sibling.type)) {
break;
}
sibling = sibling.previousSibling;
}
if (commentTexts.length === 0) return new Map();
const params = new Map<string, string>();
const commentBlock = commentTexts.join('\n');
PHPDOC_PARAM_RE.lastIndex = 0;
let match: RegExpExecArray | null;
while ((match = PHPDOC_PARAM_RE.exec(commentBlock)) !== null) {
const typeName = normalizePhpType(match[1]);
const paramName = match[2]; // without $ prefix
if (typeName) {
// Store with $ prefix to match how PHP variables appear in the env
params.set('$' + paramName, typeName);
}
}
// Also check alternate PHPDoc order: @param $name Type
PHPDOC_PARAM_ALT_RE.lastIndex = 0;
while ((match = PHPDOC_PARAM_ALT_RE.exec(commentBlock)) !== null) {
const paramName = match[1];
if (params.has('$' + paramName)) continue; // standard format takes priority
const typeName = normalizePhpType(match[2]);
if (typeName) {
params.set('$' + paramName, typeName);
}
}
return params;
};
/**
* PHP: typed class properties (PHP 7.4+): private UserRepo $repo;
* Also: PHPDoc @param annotations on method/function definitions.
*/
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
// PHPDoc @param on methods/functions — pre-populate env with param types
if (node.type === 'method_declaration' || node.type === 'function_definition') {
const phpDocParams = collectPhpDocParams(node);
for (const [paramName, typeName] of phpDocParams) {
if (!env.has(paramName)) env.set(paramName, typeName);
}
return;
}
if (node.type !== 'property_declaration') return;
const typeNode = node.childForFieldName('type');
if (!typeNode) return;
const typeName = extractSimpleTypeName(typeNode);
if (!typeName) return;
// The variable name is inside property_element > variable_name
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'property_element') {
const varNameNode = child.firstNamedChild; // variable_name
if (varNameNode) {
const varName = extractVarName(varNameNode);
if (varName) env.set(varName, typeName);
}
break;
}
}
};
/** PHP: $x = new User() — infer type from object_creation_expression */
const extractInitializer: InitializerExtractor = (node: SyntaxNode, env: Map<string, string>, _classNames: ClassNameLookup): void => {
if (node.type !== 'assignment_expression') return;
const left = node.childForFieldName('left');
const right = node.childForFieldName('right');
if (!left || !right) return;
if (right.type !== 'object_creation_expression') return;
// The class name is the first named child of object_creation_expression
// (tree-sitter-php uses 'name' or 'qualified_name' nodes here)
const ctorType = right.firstNamedChild;
if (!ctorType) return;
const typeName = extractSimpleTypeName(ctorType);
if (!typeName) return;
// Resolve PHP self/static/parent to actual class names
const resolvedType = (typeName === 'self' || typeName === 'static' || typeName === 'parent')
? resolvePhpKeyword(typeName, node)
: typeName;
if (!resolvedType) return;
const varName = extractVarName(left);
if (varName) env.set(varName, resolvedType);
};
/** PHP: simple_parameter → type $name */
@@ -25,12 +199,206 @@ const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string,
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
if (!varName) return;
// Don't overwrite PHPDoc-derived types (e.g. @param User[] $users → User)
// with the less-specific AST type annotation (e.g. array).
if (env.has(varName)) return;
const typeName = extractSimpleTypeName(typeNode);
if (varName && typeName) env.set(varName, typeName);
if (typeName) env.set(varName, typeName);
};
/** PHP: $x = SomeFactory() or $x = $this->getUser() — bind variable to call return type */
const scanConstructorBinding: ConstructorBindingScanner = (node) => {
if (node.type !== 'assignment_expression') return undefined;
const left = node.childForFieldName('left');
const right = node.childForFieldName('right');
if (!left || !right) return undefined;
if (left.type !== 'variable_name') return undefined;
// Skip object_creation_expression (new User()) — handled by extractInitializer
if (right.type === 'object_creation_expression') return undefined;
// Handle both standalone function calls and method calls ($this->getUser())
if (right.type === 'function_call_expression') {
const calleeName = extractCalleeName(right);
if (!calleeName) return undefined;
return { varName: left.text, calleeName };
}
if (right.type === 'member_call_expression') {
const methodName = right.childForFieldName('name');
if (!methodName) return undefined;
// When receiver is $this/self/static, qualify with enclosing class for disambiguation
const receiver = right.childForFieldName('object');
const receiverText = receiver?.text;
let receiverClassName: string | undefined;
if (receiverText === '$this' || receiverText === 'self' || receiverText === 'static') {
const cls = findEnclosingClass(node);
const clsName = cls?.childForFieldName('name');
if (clsName) receiverClassName = clsName.text;
}
return { varName: left.text, calleeName: methodName.text, receiverClassName };
}
return undefined;
};
/** Regex to extract PHPDoc @return annotations: `@return User` */
const PHPDOC_RETURN_RE = /@return\s+(\S+)/;
/**
* Extract return type from PHPDoc `@return Type` annotation preceding a method.
* Walks backwards through preceding siblings looking for comment nodes.
*/
const extractReturnType: ReturnTypeExtractor = (node) => {
let sibling = node.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
const match = PHPDOC_RETURN_RE.exec(sibling.text);
if (match) return normalizePhpType(match[1]);
} else if (sibling.isNamed && !SKIP_NODE_TYPES.has(sibling.type)) break;
sibling = sibling.previousSibling;
}
return undefined;
};
/** PHP: $alias = $user → assignment_expression with variable_name left/right.
* PHP TypeEnv stores variables WITH $ prefix ($user → User), so we keep $ in lhs/rhs. */
const extractPendingAssignment: PendingAssignmentExtractor = (node, scopeEnv) => {
if (node.type !== 'assignment_expression') return undefined;
const left = node.childForFieldName('left');
const right = node.childForFieldName('right');
if (!left || !right) return undefined;
if (left.type !== 'variable_name' || right.type !== 'variable_name') return undefined;
const lhs = left.text;
const rhs = right.text;
if (!lhs || !rhs || scopeEnv.has(lhs)) return undefined;
return { lhs, rhs };
};
const FOR_LOOP_NODE_TYPES: ReadonlySet<string> = new Set([
'foreach_statement',
]);
/** Extract element type from a PHP type annotation AST node.
* PHP has limited AST-level container types — `array` is a primitive_type with no generic args.
* Named types (e.g., `Collection`) are returned as-is (container descriptor lookup handles them). */
const extractPhpElementTypeFromTypeNode = (_typeNode: SyntaxNode): string | undefined => {
// PHP AST type nodes don't carry generic parameters (array<User> is PHPDoc-only).
// primitive_type 'array' and named_type 'Collection' don't encode element types.
return undefined;
};
/** Walk up from a foreach to the enclosing function and search parameter type annotations.
* PHP parameter type hints are limited (array, ClassName) — this extracts element type when possible. */
const findPhpParamElementType = (iterableName: string, startNode: SyntaxNode): string | undefined => {
let current: SyntaxNode | null = startNode.parent;
while (current) {
if (current.type === 'method_declaration' || current.type === 'function_definition') {
const paramsNode = current.childForFieldName('parameters');
if (paramsNode) {
for (let i = 0; i < paramsNode.namedChildCount; i++) {
const param = paramsNode.namedChild(i);
if (!param || param.type !== 'simple_parameter') continue;
const nameNode = param.childForFieldName('name');
if (nameNode?.text !== iterableName) continue;
const typeNode = param.childForFieldName('type');
if (typeNode) return extractPhpElementTypeFromTypeNode(typeNode);
}
}
break;
}
current = current.parent;
}
return undefined;
};
/**
* PHP: foreach ($users as $user) — extract loop variable binding.
*
* AST structure (from tree-sitter-php grammar):
* foreach_statement — no named fields for iterable/value (only 'body')
* children[0]: expression (iterable, e.g. $users)
* children[1]: expression (simple value) OR pair ($key => $value)
* pair children: expression (key), expression (value)
*
* PHP's PHPDoc @param normalizes `User[]` → `User` in the env, so the iterable's
* stored type IS the element type. We first try resolveIterableElementType (for
* constructor-binding cases that retain container types), then fall back to direct
* scopeEnv lookup (for PHPDoc-normalized types).
*/
const extractForLoopBinding: ForLoopExtractor = (
node: SyntaxNode,
scopeEnv: Map<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
): void => {
if (node.type !== 'foreach_statement') return;
// Collect non-body named children: first is the iterable, second is value or pair
const children: SyntaxNode[] = [];
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child && child !== node.childForFieldName('body')) {
children.push(child);
}
}
if (children.length < 2) return;
const iterableNode = children[0];
const valueOrPair = children[1];
// Determine the loop variable node
let loopVarNode: SyntaxNode;
if (valueOrPair.type === 'pair') {
// $key => $value — the value is the last named child of the pair
const lastChild = valueOrPair.namedChild(valueOrPair.namedChildCount - 1);
if (!lastChild) return;
// Handle by_ref: foreach ($arr as $k => &$v)
loopVarNode = lastChild.type === 'by_ref' ? (lastChild.firstNamedChild ?? lastChild) : lastChild;
} else {
// Simple: foreach ($users as $user) or foreach ($users as &$user)
loopVarNode = valueOrPair.type === 'by_ref' ? (valueOrPair.firstNamedChild ?? valueOrPair) : valueOrPair;
}
const varName = extractVarName(loopVarNode);
if (!varName) return;
// Get iterable variable name (PHP vars include $ prefix)
let iterableName: string | undefined;
if (iterableNode.type === 'variable_name') {
iterableName = iterableNode.text;
} else if (iterableNode?.type === 'member_access_expression') {
const name = iterableNode.childForFieldName('name');
// PHP properties are stored in scopeEnv with $ prefix ($users), but
// member_access_expression.name returns without $ (users). Add $ to match.
if (name) iterableName = '$' + name.text;
}
if (!iterableName) return;
// Strategy A: try resolveIterableElementType (handles constructor-binding container types)
const elementType = resolveIterableElementType(
iterableName, node, scopeEnv, declarationTypeNodes, scope,
extractPhpElementTypeFromTypeNode, findPhpParamElementType,
undefined,
);
if (elementType) {
scopeEnv.set(varName, elementType);
return;
}
// Strategy B: direct scopeEnv lookup — PHP normalizePhpType strips User[] → User,
// so the iterable's stored type is already the element type from PHPDoc annotations.
const iterableType = scopeEnv.get(iterableName);
if (iterableType) {
scopeEnv.set(varName, iterableType);
}
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
forLoopNodeTypes: FOR_LOOP_NODE_TYPES,
extractDeclaration,
extractParameter,
extractInitializer,
scanConstructorBinding,
extractReturnType,
extractForLoopBinding,
extractPendingAssignment,
};
@@ -1,20 +1,53 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor, InitializerExtractor, ClassNameLookup, ConstructorBindingScanner, PendingAssignmentExtractor, PatternBindingExtractor, ForLoopExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, extractElementTypeFromString, extractGenericTypeArgs, resolveIterableElementType, methodToTypeArgPosition, type TypeArgPosition } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'assignment',
'named_expression',
'expression_statement',
]);
/** Python: x: Foo = ... (PEP 484 annotations) */
/** Python: x: Foo = ... (PEP 484 annotated assignment) or x: Foo (standalone annotation).
*
* tree-sitter-python grammar produces two distinct shapes:
*
* 1. Annotated assignment with value: `name: str = ""`
* Node type: `assignment`
* Fields: left=identifier, type=identifier/type, right=value
*
* 2. Standalone annotation (no value): `name: str`
* Node type: `expression_statement`
* Child: `type` node with fields name=identifier, type=identifier/type
*
* Both appear at file scope and inside class bodies (PEP 526 class variable annotations).
*/
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
// Python annotated assignment: left : type = value
// tree-sitter represents this differently based on grammar version
if (node.type === 'expression_statement') {
// Standalone annotation: expression_statement > type { name: identifier, type: identifier }
const typeChild = node.firstNamedChild;
if (!typeChild || typeChild.type !== 'type') return;
const nameNode = typeChild.childForFieldName('name');
const typeNode = typeChild.childForFieldName('type');
if (!nameNode || !typeNode) return;
const varName = extractVarName(nameNode);
const inner = typeNode.type === 'type' ? (typeNode.firstNamedChild ?? typeNode) : typeNode;
const typeName = extractSimpleTypeName(inner) ?? inner.text;
if (varName && typeName) env.set(varName, typeName);
return;
}
// Annotated assignment: left : type = value
const left = node.childForFieldName('left');
const typeNode = node.childForFieldName('type');
if (!left || !typeNode) return;
const varName = extractVarName(left);
const typeName = extractSimpleTypeName(typeNode);
// extractSimpleTypeName handles identifiers and qualified names.
// Python 3.10+ union syntax `User | None` is parsed as binary_operator,
// which extractSimpleTypeName doesn't handle. Fall back to raw text so
// stripNullable can process it at lookup time (e.g., "User | None" → "User").
const inner = typeNode.type === 'type' ? (typeNode.firstNamedChild ?? typeNode) : typeNode;
const typeName = extractSimpleTypeName(inner) ?? inner.text;
if (varName && typeName) env.set(varName, typeName);
};
@@ -37,8 +70,315 @@ const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string,
if (varName && typeName) env.set(varName, typeName);
};
/** Python: user = User("alice") — infer type from call when callee is a known class.
* Python constructors are syntactically identical to function calls, so we verify
* against classNames (which may include cross-file SymbolTable lookups).
* Also handles walrus operator: if (user := User("alice")): */
const extractInitializer: InitializerExtractor = (node: SyntaxNode, env: Map<string, string>, classNames: ClassNameLookup): void => {
let left: SyntaxNode | null;
let right: SyntaxNode | null;
if (node.type === 'named_expression') {
// Walrus operator: (user := User("alice"))
// tree-sitter-python: named_expression has 'name' and 'value' fields
left = node.childForFieldName('name');
right = node.childForFieldName('value');
} else if (node.type === 'assignment') {
left = node.childForFieldName('left');
right = node.childForFieldName('right');
// Skip if already has type annotation — extractDeclaration handled it
if (node.childForFieldName('type')) return;
} else {
return;
}
if (!left || !right) return;
const varName = extractVarName(left);
if (!varName || env.has(varName)) return;
if (right.type !== 'call') return;
const func = right.childForFieldName('function');
if (!func) return;
// Support both direct calls (User()) and qualified calls (models.User())
// tree-sitter-python: direct → identifier, qualified → attribute
const calleeName = extractSimpleTypeName(func);
if (!calleeName) return;
if (classNames.has(calleeName)) {
env.set(varName, calleeName);
}
};
/** Python: user = User("alice") — scan assignment/walrus for constructor-like calls.
* Returns {varName, calleeName} without checking classNames (caller validates). */
const scanConstructorBinding: ConstructorBindingScanner = (node) => {
let left: SyntaxNode | null;
let right: SyntaxNode | null;
if (node.type === 'named_expression') {
left = node.childForFieldName('name');
right = node.childForFieldName('value');
} else if (node.type === 'assignment') {
left = node.childForFieldName('left');
right = node.childForFieldName('right');
if (node.childForFieldName('type')) return undefined;
} else {
return undefined;
}
if (!left || !right) return undefined;
if (left.type !== 'identifier') return undefined;
if (right.type !== 'call') return undefined;
const func = right.childForFieldName('function');
if (!func) return undefined;
const calleeName = extractSimpleTypeName(func);
if (!calleeName) return undefined;
return { varName: left.text, calleeName };
};
const FOR_LOOP_NODE_TYPES: ReadonlySet<string> = new Set([
'for_statement',
]);
/** Python function/method node types that carry a parameters list. */
const PY_FUNCTION_NODE_TYPES = new Set([
'function_definition', 'decorated_definition',
]);
/**
* Extract element type from a Python type annotation AST node.
* Handles:
* subscript "List[User]" → extractElementTypeFromString("List[User]") → "User"
* generic_type → extractGenericTypeArgs → first arg
* Falls back to text-based extraction.
*/
const extractPyElementTypeFromAnnotation = (typeNode: SyntaxNode, pos: TypeArgPosition = 'last'): string | undefined => {
// Unwrap 'type' wrapper node to get to the actual type (e.g., type > generic_type)
const inner = typeNode.type === 'type' ? (typeNode.firstNamedChild ?? typeNode) : typeNode;
// Python subscript: List[User], Sequence[User] — use raw text
if (inner.type === 'subscript') {
return extractElementTypeFromString(inner.text, pos);
}
// generic_type: dict[str, User] — tree-sitter-python uses type_parameter child
if (inner.type === 'generic_type') {
// Try standard extractGenericTypeArgs first (handles type_arguments)
const args = extractGenericTypeArgs(inner);
if (args.length >= 1) return pos === 'first' ? args[0] : args[args.length - 1];
// Fallback: look for type_parameter child (tree-sitter-python specific)
for (let i = 0; i < inner.namedChildCount; i++) {
const child = inner.namedChild(i);
if (child?.type === 'type_parameter') {
if (pos === 'first') {
const firstArg = child.firstNamedChild;
if (firstArg) return extractSimpleTypeName(firstArg);
} else {
const lastArg = child.lastNamedChild;
if (lastArg) return extractSimpleTypeName(lastArg);
}
}
}
}
// Fallback: raw text extraction (handles User[], [User], etc.)
return extractElementTypeFromString(inner.text, pos);
};
/**
* Walk up the AST from a for-statement to find the enclosing function definition,
* then search its parameters for one named `iterableName`.
* Returns the element type extracted from its type annotation, or undefined.
*
* Handles both `parameter` and `typed_parameter` node types in tree-sitter-python.
* `typed_parameter` may not expose the name as a `name` field — falls back to
* checking the first identifier-type named child.
*/
const findPyParamElementType = (iterableName: string, startNode: SyntaxNode, pos: TypeArgPosition = 'last'): string | undefined => {
let current: SyntaxNode | null = startNode.parent;
while (current) {
if (current.type === 'function_definition') {
const paramsNode = current.childForFieldName('parameters');
if (paramsNode) {
for (let i = 0; i < paramsNode.namedChildCount; i++) {
const param = paramsNode.namedChild(i);
if (!param) continue;
// Try named `name` field first (parameter node), then first identifier child
// (typed_parameter node may store name as first positional child)
const nameNode = param.childForFieldName('name')
?? (param.firstNamedChild?.type === 'identifier' ? param.firstNamedChild : null);
if (nameNode?.text !== iterableName) continue;
// Try `type` field, then last named child (typed_parameter stores type last)
const typeAnnotation = param.childForFieldName('type')
?? (param.namedChildCount >= 2 ? param.namedChild(param.namedChildCount - 1) : null);
if (typeAnnotation && typeAnnotation !== nameNode) {
return extractPyElementTypeFromAnnotation(typeAnnotation, pos);
}
}
}
break;
}
current = current.parent;
}
return undefined;
};
/**
* Python: for user in users: where users has a known container type annotation.
*
* AST node: `for_statement` with `left` (loop variable) and `right` (iterable).
*
* Tier 1c: resolves the element type via three strategies in priority order:
* 1. declarationTypeNodes — raw type annotation AST node (covers stored container types)
* 2. scopeEnv string — extractElementTypeFromString on the stored type
* 3. AST walk — walks up to the enclosing function's parameters to read List[User] directly
*/
const extractForLoopBinding: ForLoopExtractor = (
node: SyntaxNode,
scopeEnv: Map<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
): void => {
if (node.type !== 'for_statement') return;
// The iterable is the `right` field — may be identifier or call (data.items()/keys()/values()).
const rightNode = node.childForFieldName('right');
let iterableName: string | undefined;
let methodName: string | undefined;
if (rightNode?.type === 'identifier') {
iterableName = rightNode.text;
} else if (rightNode?.type === 'attribute') {
const prop = rightNode.lastNamedChild;
if (prop) iterableName = prop.text;
} else if (rightNode?.type === 'call') {
// data.items() → call > function: attribute > identifier('data') + identifier('items')
const fn = rightNode.childForFieldName('function');
if (fn?.type === 'attribute') {
const obj = fn.firstNamedChild;
if (obj?.type === 'identifier') iterableName = obj.text;
// Extract method name: items, keys, values
const method = fn.lastNamedChild;
if (method?.type === 'identifier' && method !== obj) methodName = method.text;
}
}
if (!iterableName) return;
const containerTypeName = scopeEnv.get(iterableName);
const typeArgPos = methodToTypeArgPosition(methodName, containerTypeName);
const elementType = resolveIterableElementType(
iterableName, node, scopeEnv, declarationTypeNodes, scope,
extractPyElementTypeFromAnnotation, findPyParamElementType,
typeArgPos,
);
if (!elementType) return;
// The loop variable is the `left` field — identifier or pattern_list.
const leftNode = node.childForFieldName('left');
if (!leftNode) return;
// Handle tuple unpacking: for key, value in data.items()
if (leftNode.type === 'pattern_list') {
const lastChild = leftNode.lastNamedChild;
if (lastChild?.type === 'identifier') {
scopeEnv.set(lastChild.text, elementType);
}
return;
}
const loopVarName = extractVarName(leftNode);
if (loopVarName) scopeEnv.set(loopVarName, elementType);
};
/** Python: alias = u → assignment with left/right fields.
* Also handles walrus operator: alias := u → named_expression with name/value fields. */
const extractPendingAssignment: PendingAssignmentExtractor = (node, scopeEnv) => {
let left: SyntaxNode | null;
let right: SyntaxNode | null;
if (node.type === 'assignment') {
left = node.childForFieldName('left');
right = node.childForFieldName('right');
} else if (node.type === 'named_expression') {
left = node.childForFieldName('name');
right = node.childForFieldName('value');
} else {
return undefined;
}
if (!left || !right) return undefined;
const lhs = left.type === 'identifier' ? left.text : undefined;
if (!lhs || scopeEnv.has(lhs)) return undefined;
if (right.type === 'identifier') return { lhs, rhs: right.text };
return undefined;
};
/**
* Python match/case `as` pattern binding: `case User() as u:`
*
* AST structure (tree-sitter-python):
* as_pattern
* alias: as_pattern_target ← the bound variable name (e.g. "u")
* children[0]: case_pattern ← wraps class_pattern (or is class_pattern directly)
* class_pattern
* dotted_name ← the class name (e.g. "User")
*
* The `alias` field is an `as_pattern_target` node whose `.text` is the identifier.
* The class name lives in the first non-alias named child: either a `case_pattern`
* wrapping a `class_pattern`, or a direct `class_pattern`.
*
* Conservative: returns undefined when:
* - The node is not an `as_pattern`
* - The pattern side is not a class_pattern (e.g. guard or literal match)
* - The variable was already bound in scopeEnv
*/
const extractPatternBinding: PatternBindingExtractor = (node, scopeEnv) => {
if (node.type !== 'as_pattern') return undefined;
// as_pattern: `case User() as u:` — binds matched value to a name.
// Try named field first (future grammar versions may expose it), fall back to positional.
if (node.namedChildCount < 2) return undefined;
const patternChild = node.namedChild(0);
const varNameNode = node.childForFieldName('alias')
?? node.namedChild(node.namedChildCount - 1);
if (!patternChild || !varNameNode) return undefined;
if (varNameNode.type !== 'identifier') return undefined;
const varName = varNameNode.text;
if (!varName || scopeEnv.has(varName)) return undefined;
// Find the class_pattern — may be direct or wrapped in case_pattern.
let classPattern: SyntaxNode | null = null;
if (patternChild.type === 'class_pattern') {
classPattern = patternChild;
} else if (patternChild.type === 'case_pattern') {
// Unwrap one level: case_pattern wraps class_pattern
for (let j = 0; j < patternChild.namedChildCount; j++) {
const inner = patternChild.namedChild(j);
if (inner?.type === 'class_pattern') {
classPattern = inner;
break;
}
}
}
if (!classPattern) return undefined;
// class_pattern children: dotted_name (the class name) + optional keyword_pattern args.
const classNameNode = classPattern.firstNamedChild;
if (!classNameNode || (classNameNode.type !== 'dotted_name' && classNameNode.type !== 'identifier')) return undefined;
const typeName = classNameNode.text;
if (!typeName) return undefined;
return { varName, typeName };
};
const PATTERN_BINDING_NODE_TYPES: ReadonlySet<string> = new Set(['as_pattern']);
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
forLoopNodeTypes: FOR_LOOP_NODE_TYPES,
patternBindingNodeTypes: PATTERN_BINDING_NODE_TYPES,
extractDeclaration,
extractParameter,
extractInitializer,
scanConstructorBinding,
extractForLoopBinding,
extractPendingAssignment,
extractPatternBinding,
};
@@ -0,0 +1,411 @@
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor, InitializerExtractor, ClassNameLookup, ConstructorBindingScanner, ReturnTypeExtractor, PendingAssignmentExtractor, ForLoopExtractor } from './types.js';
import { extractRubyConstructorAssignment, extractSimpleTypeName, extractElementTypeFromString, extractVarName, resolveIterableElementType } from './shared.js';
import type { SyntaxNode } from '../utils.js';
/**
* Ruby type extractor — YARD annotation parsing.
*
* Ruby has no static type system, but the YARD documentation convention
* provides de facto type annotations via comments:
*
* # @param name [String] the user's name
* # @param repo [UserRepo] the repository
* # @return [User]
* def create(name, repo)
* repo.save
* end
*
* This extractor parses `@param name [Type]` patterns from comment nodes
* preceding method definitions and binds parameter names to their types.
*
* Resolution tiers:
* - Tier 0: YARD @param annotations (extractDeclaration pre-populates env)
* - Tier 1: Constructor inference via `user = User.new` (handled by scanConstructorBinding in typeConfig)
*/
/** Regex to extract @param annotations: `@param name [Type]` */
const YARD_PARAM_RE = /@param\s+(\w+)\s+\[([^\]]+)\]/g;
/** Alternate YARD order: `@param [Type] name` */
const YARD_PARAM_ALT_RE = /@param\s+\[([^\]]+)\]\s+(\w+)/g;
/** Regex to extract @return annotations: `@return [Type]` */
const YARD_RETURN_RE = /@return\s+\[([^\]]+)\]/;
/**
* Extract the simple type name from a YARD type string.
* Handles:
* - Simple types: "String" → "String"
* - Qualified types: "Models::User" → "User"
* - Generic types: "Array<User>" → "Array"
* - Nullable types: "String, nil" → "String"
* - Union types: "String, Integer" → undefined (ambiguous)
*/
const extractYardTypeName = (yardType: string): string | undefined => {
const trimmed = yardType.trim();
// Handle nullable: "Type, nil" or "nil, Type"
// Use bracket-balanced split to avoid breaking on commas inside generics like Hash<Symbol, User>
const parts: string[] = [];
let depth = 0, start = 0;
for (let i = 0; i < trimmed.length; i++) {
if (trimmed[i] === '<') depth++;
else if (trimmed[i] === '>') depth--;
else if (trimmed[i] === ',' && depth === 0) {
parts.push(trimmed.slice(start, i).trim());
start = i + 1;
}
}
parts.push(trimmed.slice(start).trim());
const filtered = parts.filter(p => p !== '' && p !== 'nil');
if (filtered.length !== 1) return undefined; // ambiguous union
const typePart = filtered[0];
// Handle qualified: "Models::User" → "User"
const segments = typePart.split('::');
const last = segments[segments.length - 1];
// Handle generic: "Array<User>" → "Array"
const genericMatch = last.match(/^(\w+)\s*[<{(]/);
if (genericMatch) return genericMatch[1];
// Simple identifier check
if (/^\w+$/.test(last)) return last;
return undefined;
};
/**
* Collect YARD @param annotations from comment nodes preceding a method definition.
* Returns a map of paramName → typeName.
*
* In tree-sitter-ruby, comments are sibling nodes that appear before the method node.
* We walk backwards through preceding siblings collecting consecutive comment nodes.
*/
const collectYardParams = (methodNode: SyntaxNode): Map<string, string> => {
const params = new Map<string, string>();
// In tree-sitter-ruby, YARD comments preceding a method inside a class body
// are placed as children of the `class` node, NOT as siblings of the `method`
// inside `body_statement`. The AST structure is:
//
// class
// constant = "ClassName"
// comment = "# @param ..." ← sibling of body_statement
// comment = "# @param ..." ← sibling of body_statement
// body_statement
// method ← method is here, no preceding siblings
//
// For top-level methods (outside classes), comments ARE direct siblings.
// We handle both by checking: if method has no preceding comment siblings,
// look at parent (body_statement) siblings instead.
const commentTexts: string[] = [];
const collectComments = (startNode: SyntaxNode): void => {
let sibling = startNode.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
commentTexts.unshift(sibling.text);
} else if (sibling.isNamed) {
break;
}
sibling = sibling.previousSibling;
}
};
// Try method's own siblings first (top-level methods)
collectComments(methodNode);
// If no comments found and parent is body_statement, check parent's siblings
if (commentTexts.length === 0 && methodNode.parent?.type === 'body_statement') {
collectComments(methodNode.parent);
}
// Parse all comment lines for @param annotations
const commentBlock = commentTexts.join('\n');
let match: RegExpExecArray | null;
// Reset regex state
YARD_PARAM_RE.lastIndex = 0;
while ((match = YARD_PARAM_RE.exec(commentBlock)) !== null) {
const paramName = match[1];
const rawType = match[2];
const typeName = extractYardTypeName(rawType);
if (typeName) {
params.set(paramName, typeName);
}
}
// Also check alternate YARD order: @param [Type] name
YARD_PARAM_ALT_RE.lastIndex = 0;
while ((match = YARD_PARAM_ALT_RE.exec(commentBlock)) !== null) {
const rawType = match[1];
const paramName = match[2];
if (params.has(paramName)) continue; // standard format takes priority
const typeName = extractYardTypeName(rawType);
if (typeName) {
params.set(paramName, typeName);
}
}
return params;
};
/**
* Ruby node types that may carry type bindings.
* - `method`/`singleton_method`: YARD @param annotations (via extractDeclaration)
* - `assignment`: Constructor inference like `user = User.new` (via extractInitializer;
* extractDeclaration returns early for these nodes)
*/
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'method',
'singleton_method',
'assignment',
]);
/**
* Extract YARD annotations from method definitions.
* Pre-populates the scope env with parameter types before the
* standard parameter walk (which won't find types since Ruby has none).
*/
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
if (node.type !== 'method' && node.type !== 'singleton_method') return;
const yardParams = collectYardParams(node);
if (yardParams.size === 0) return;
// Pre-populate env with YARD type bindings for each parameter
for (const [paramName, typeName] of yardParams) {
env.set(paramName, typeName);
}
};
/**
* Ruby parameter extraction.
* Ruby parameters (identifiers inside method_parameters) have no inline
* type annotations. YARD types are already populated by extractDeclaration,
* so this is a no-op — the bindings are already in the env.
*
* We still register this to maintain the LanguageTypeConfig contract.
*/
const extractParameter: ParameterExtractor = (_node: SyntaxNode, _env: Map<string, string>): void => {
// Ruby parameters have no type annotations.
// YARD types are pre-populated by extractDeclaration.
};
/**
* Ruby constructor inference: user = User.new or service = Models::User.new
* Uses the shared extractRubyConstructorAssignment helper for AST matching,
* then resolves against locally-known class names.
*/
const extractInitializer: InitializerExtractor = (node, env, classNames): void => {
const result = extractRubyConstructorAssignment(node);
if (!result) return;
if (env.has(result.varName)) return;
if (classNames.has(result.calleeName)) {
env.set(result.varName, result.calleeName);
}
};
/**
* Extract return type from YARD `@return [Type]` annotation preceding a method.
* Reuses the same comment-walking strategy as collectYardParams: try direct
* siblings first, fall back to parent (body_statement) siblings for class methods.
*/
const extractReturnType: ReturnTypeExtractor = (node) => {
const search = (startNode: SyntaxNode): string | undefined => {
let sibling = startNode.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
const match = YARD_RETURN_RE.exec(sibling.text);
if (match) return extractYardTypeName(match[1]);
} else if (sibling.isNamed) {
break;
}
sibling = sibling.previousSibling;
}
return undefined;
};
const result = search(node);
if (result) return result;
if (node.parent?.type === 'body_statement') {
return search(node.parent);
}
return undefined;
};
/**
* Ruby constructor binding scanner: captures both `user = User.new` and
* plain call assignments like `user = get_user()`.
* The `.new` pattern returns the class name directly; plain calls return the
* callee name for return-type inference via SymbolTable lookup.
*/
const scanConstructorBinding: ConstructorBindingScanner = (node) => {
// Try the .new pattern first (returns class name directly)
const newResult = extractRubyConstructorAssignment(node);
if (newResult) return newResult;
// Plain call assignment: user = get_user() / user = Models.create()
if (node.type !== 'assignment') return undefined;
const left = node.childForFieldName('left');
const right = node.childForFieldName('right');
if (!left || !right) return undefined;
if (left.type !== 'identifier' && left.type !== 'constant') return undefined;
if (right.type !== 'call') return undefined;
const method = right.childForFieldName('method');
if (!method) return undefined;
const calleeName = extractSimpleTypeName(method);
if (!calleeName) return undefined;
return { varName: left.text, calleeName };
};
/** Ruby method node types that carry a parameter list. */
const RUBY_METHOD_NODE_TYPES = new Set(['method', 'singleton_method']);
const FOR_LOOP_NODE_TYPES: ReadonlySet<string> = new Set(['for']);
/**
* Collect raw YARD @param type strings from comment nodes preceding a method.
* Unlike collectYardParams which returns simplified type names, this returns the
* raw bracket content (e.g., "Array<User>" not "Array") for element type extraction.
*/
const collectYardRawParams = (methodNode: SyntaxNode): Map<string, string> => {
const params = new Map<string, string>();
const commentTexts: string[] = [];
const collectComments = (startNode: SyntaxNode): void => {
let sibling = startNode.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
commentTexts.unshift(sibling.text);
} else if (sibling.isNamed) {
break;
}
sibling = sibling.previousSibling;
}
};
collectComments(methodNode);
if (commentTexts.length === 0 && methodNode.parent?.type === 'body_statement') {
collectComments(methodNode.parent);
}
const commentBlock = commentTexts.join('\n');
let match: RegExpExecArray | null;
YARD_PARAM_RE.lastIndex = 0;
while ((match = YARD_PARAM_RE.exec(commentBlock)) !== null) {
params.set(match[1], match[2]);
}
YARD_PARAM_ALT_RE.lastIndex = 0;
while ((match = YARD_PARAM_ALT_RE.exec(commentBlock)) !== null) {
if (!params.has(match[2])) params.set(match[2], match[1]);
}
return params;
};
/**
* Walk up the AST from a for-statement to find the enclosing method,
* then search its YARD @param annotations for one named `iterableName`.
* Returns the element type extracted from the raw YARD type string.
*
* Example: `@param users [Array<User>]` → extracts "User" from "Array<User>".
*/
const findRubyParamElementType = (iterableName: string, startNode: SyntaxNode): string | undefined => {
let current: SyntaxNode | null = startNode.parent;
while (current) {
if (RUBY_METHOD_NODE_TYPES.has(current.type)) {
const rawParams = collectYardRawParams(current);
const rawType = rawParams.get(iterableName);
if (rawType) return extractElementTypeFromString(rawType);
break;
}
current = current.parent;
}
return undefined;
};
/**
* Ruby: for user in users ... end
*
* tree-sitter-ruby `for` node structure:
* pattern field: the loop variable (identifier)
* value field: `in` node whose child is the iterable expression
*
* Tier 1c: resolves the element type via:
* 1. scopeEnv string — extractElementTypeFromString on the stored type
* 2. AST walk — walks up to the enclosing method's YARD @param to read Array<User> directly
*
* Ruby has no static types on loop variables, so this mainly works when the
* iterable has a YARD-annotated container type (e.g., `@param users [Array<User>]`).
*/
const extractForLoopBinding: ForLoopExtractor = (
node: SyntaxNode,
scopeEnv: Map<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
): void => {
if (node.type !== 'for') return;
// The loop variable is the `pattern` field (identifier).
const patternNode = node.childForFieldName('pattern');
if (!patternNode) return;
const loopVarName = extractVarName(patternNode);
if (!loopVarName) return;
// The iterable is inside the `value` field which is an `in` node wrapping the expression.
const inNode = node.childForFieldName('value');
if (!inNode) return;
const iterableNode = inNode.firstNamedChild;
let iterableName: string | undefined;
if (iterableNode?.type === 'identifier') {
iterableName = iterableNode.text;
} else if (iterableNode?.type === 'call') {
const method = iterableNode.childForFieldName('method');
if (method) iterableName = method.text;
}
if (!iterableName) return;
// Ruby has no extractFromTypeNode (no AST type annotations), pass a no-op.
const noopExtractFromTypeNode = (): string | undefined => undefined;
const elementType = resolveIterableElementType(
iterableName, node, scopeEnv, declarationTypeNodes, scope,
noopExtractFromTypeNode, findRubyParamElementType,
undefined,
);
if (!elementType) return;
scopeEnv.set(loopVarName, elementType);
};
/**
* Ruby: alias_user = user → assignment with left/right identifier fields.
* Only handles plain identifier RHS (not calls, not literals).
* Skips if LHS already has a resolved type in scopeEnv.
*/
const extractPendingAssignment: PendingAssignmentExtractor = (node, scopeEnv) => {
if (node.type !== 'assignment') return undefined;
const lhsNode = node.childForFieldName('left');
if (!lhsNode || lhsNode.type !== 'identifier') return undefined;
const varName = lhsNode.text;
if (scopeEnv.has(varName)) return undefined;
const rhsNode = node.childForFieldName('right');
if (!rhsNode || rhsNode.type !== 'identifier') return undefined;
return { lhs: varName, rhs: rhsNode.text };
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
forLoopNodeTypes: FOR_LOOP_NODE_TYPES,
extractDeclaration,
extractParameter,
extractInitializer,
scanConstructorBinding,
extractReturnType,
extractForLoopBinding,
extractPendingAssignment,
};
@@ -1,13 +1,91 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor, InitializerExtractor, ClassNameLookup, ConstructorBindingScanner, PendingAssignmentExtractor, PatternBindingExtractor, ForLoopExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, hasTypeAnnotation, unwrapAwait, extractGenericTypeArgs, resolveIterableElementType, methodToTypeArgPosition, type TypeArgPosition } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'let_declaration',
'let_condition',
]);
/** Rust: let x: Foo = ... */
/** Walk up the AST to find the enclosing impl block and extract the implementing type name. */
const findEnclosingImplType = (node: SyntaxNode): string | undefined => {
let current = node.parent;
while (current) {
if (current.type === 'impl_item') {
// The 'type' field holds the implementing type (e.g., `impl User { ... }`)
const typeNode = current.childForFieldName('type');
if (typeNode) return extractSimpleTypeName(typeNode);
}
current = current.parent;
}
return undefined;
};
/**
* Extract the type name from a struct_pattern's 'type' field.
* Handles both simple `User { .. }` and scoped `Message::Data { .. }`.
*/
const extractStructPatternType = (structPattern: SyntaxNode): string | undefined => {
const typeNode = structPattern.childForFieldName('type');
if (!typeNode) return undefined;
return extractSimpleTypeName(typeNode);
};
/**
* Recursively scan a pattern tree for captured_pattern nodes (x @ StructType { .. })
* and extract variable → type bindings from them.
*/
const extractCapturedPatternBindings = (pattern: SyntaxNode, env: Map<string, string>, depth = 0): void => {
if (depth > 50) return;
if (pattern.type === 'captured_pattern') {
// captured_pattern: identifier @ inner_pattern
// The first named child is the identifier, followed by the inner pattern.
const nameNode = pattern.firstNamedChild;
if (!nameNode || nameNode.type !== 'identifier') return;
// Find the struct_pattern child — that gives us the type
for (let i = 0; i < pattern.namedChildCount; i++) {
const child = pattern.namedChild(i);
if (child?.type === 'struct_pattern') {
const typeName = extractStructPatternType(child);
if (typeName) env.set(nameNode.text, typeName);
return;
}
}
return;
}
// Recurse into tuple_struct_pattern children to find nested captured_patterns
// e.g., Some(user @ User { .. })
if (pattern.type === 'tuple_struct_pattern') {
for (let i = 0; i < pattern.namedChildCount; i++) {
const child = pattern.namedChild(i);
if (child) extractCapturedPatternBindings(child, env, depth + 1);
}
}
};
/** Rust: let x: Foo = ... | if let / while let pattern bindings */
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
if (node.type === 'let_condition') {
// if let / while let: extract type bindings from pattern matching.
//
// Supported patterns:
// - captured_pattern: `if let user @ User { .. } = expr` → user: User
// - tuple_struct_pattern with nested captured_pattern:
// `if let Some(user @ User { .. }) = expr` → user: User
//
// NOT supported (requires generic unwrapping — Phase 3):
// - `if let Some(x) = opt` where opt: Option<T> → x: T
//
// struct_pattern without capture (`if let User { name } = expr`)
// destructures fields — individual field types are unknown without
// field-type resolution, so no bindings are extracted.
const pattern = node.childForFieldName('pattern');
if (!pattern) return;
extractCapturedPatternBindings(pattern, env);
return;
}
// Standard let_declaration: let x: Foo = ...
const pattern = node.childForFieldName('pattern');
const typeNode = node.childForFieldName('type');
if (!pattern || !typeNode) return;
@@ -16,6 +94,46 @@ const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<str
if (varName && typeName) env.set(varName, typeName);
};
/** Rust: let x = User::new(), let x = User::default(), or let x = User { ... } */
const extractInitializer: InitializerExtractor = (node: SyntaxNode, env: Map<string, string>, _classNames: ClassNameLookup): void => {
// Skip if there's an explicit type annotation — Tier 0 already handled it
if (node.childForFieldName('type') !== null) return;
const pattern = node.childForFieldName('pattern');
const value = node.childForFieldName('value');
if (!pattern || !value) return;
// Rust struct literal: let user = User { name: "alice", age: 30 }
// tree-sitter-rust: struct_expression with 'name' field holding the type
if (value.type === 'struct_expression') {
const typeNode = value.childForFieldName('name');
if (!typeNode) return;
const rawType = extractSimpleTypeName(typeNode);
if (!rawType) return;
// Resolve Self to the actual struct/enum name from the enclosing impl block
const typeName = rawType === 'Self' ? findEnclosingImplType(node) : rawType;
const varName = extractVarName(pattern);
if (varName && typeName) env.set(varName, typeName);
return;
}
if (value.type !== 'call_expression') return;
const func = value.childForFieldName('function');
if (!func || func.type !== 'scoped_identifier') return;
const nameField = func.childForFieldName('name');
// Only match ::new() and ::default() — the two idiomatic Rust constructors.
// Deliberately excludes ::from(), ::with_capacity(), etc. to avoid false positives
// (e.g. String::from("x") is not necessarily the "String" type we want for method resolution).
if (!nameField || (nameField.text !== 'new' && nameField.text !== 'default')) return;
const pathField = func.childForFieldName('path');
if (!pathField) return;
const rawType = extractSimpleTypeName(pathField);
if (!rawType) return;
// Resolve Self to the actual struct/enum name from the enclosing impl block
const typeName = rawType === 'Self' ? findEnclosingImplType(node) : rawType;
const varName = extractVarName(pattern);
if (varName && typeName) env.set(varName, typeName);
};
/** Rust: parameter → pattern: type */
const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
let nameNode: SyntaxNode | null = null;
@@ -35,8 +153,263 @@ const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string,
if (varName && typeName) env.set(varName, typeName);
};
/** Rust: let user = get_user("alice") — let_declaration with call_expression value, no type annotation.
* Skips `let user: User = ...` (explicit type annotation — handled by extractDeclaration).
* Skips `let user = User::new()` (scoped_identifier callee named "new" — handled by extractInitializer).
* Unwraps `let mut user = get_user()` by looking inside mut_pattern for the inner identifier.
*/
const scanConstructorBinding: ConstructorBindingScanner = (node) => {
if (node.type !== 'let_declaration') return undefined;
if (hasTypeAnnotation(node)) return undefined;
let patternNode = node.childForFieldName('pattern');
if (!patternNode) return undefined;
if (patternNode.type === 'mut_pattern') {
patternNode = patternNode.firstNamedChild;
if (!patternNode) return undefined;
}
if (patternNode.type !== 'identifier') return undefined;
// Unwrap `.await`: `let user = get_user().await` → await_expression wraps call_expression
const value = unwrapAwait(node.childForFieldName('value'));
if (!value || value.type !== 'call_expression') return undefined;
const func = value.childForFieldName('function');
if (!func) return undefined;
if (func.type === 'scoped_identifier') {
const methodName = func.lastNamedChild;
if (methodName?.text === 'new' || methodName?.text === 'default') return undefined;
}
const calleeName = extractSimpleTypeName(func);
if (!calleeName) return undefined;
return { varName: patternNode.text, calleeName };
};
/** Rust: let alias = u; → let_declaration with pattern + value fields */
const extractPendingAssignment: PendingAssignmentExtractor = (node, scopeEnv) => {
if (node.type !== 'let_declaration') return undefined;
const pattern = node.childForFieldName('pattern');
const value = node.childForFieldName('value');
if (!pattern || !value) return undefined;
const lhs = extractVarName(pattern);
if (!lhs || scopeEnv.has(lhs)) return undefined;
if (value.type === 'identifier') return { lhs, rhs: value.text };
return undefined;
};
/**
* Rust pattern binding extractor for `if let` / `while let` constructs that unwrap
* enum variants and introduce new typed variables.
*
* Supported patterns:
* - `if let Some(x) = opt` → x: T (opt: Option<T>, T already in scopeEnv via NULLABLE_WRAPPER_TYPES)
* - `if let Ok(x) = res` → x: T (res: Result<T, E>, T extracted from declarationTypeNodes)
*
* These complement the captured_pattern support in extractDeclaration (which handles
* `if let x @ Struct { .. } = expr` but NOT tuple struct unwrapping like Some(x) / Ok(x)).
*
* Conservative: returns undefined when:
* - The source variable's type is unknown (not in scopeEnv)
* - The wrapper is not a known single-unwrap variant (Some / Ok)
* - The value side is not a simple identifier
*/
const extractPatternBinding: PatternBindingExtractor = (
node,
scopeEnv,
declarationTypeNodes,
scope,
) => {
let patternNode: SyntaxNode | null = null;
let valueNode: SyntaxNode | null = null;
if (node.type === 'let_condition') {
patternNode = node.childForFieldName('pattern');
valueNode = node.childForFieldName('value');
} else if (node.type === 'match_arm') {
// match_arm → pattern field is match_pattern wrapping the actual pattern
const matchPatternNode = node.childForFieldName('pattern');
// Unwrap match_pattern to get the tuple_struct_pattern inside
patternNode = matchPatternNode?.type === 'match_pattern'
? matchPatternNode.firstNamedChild
: matchPatternNode;
// source variable is in the parent match_expression's 'value' field
const matchExpr = node.parent?.parent; // match_arm → match_block → match_expression
if (matchExpr?.type === 'match_expression') {
valueNode = matchExpr.childForFieldName('value');
}
}
if (!patternNode || !valueNode) return undefined;
// Only handle tuple_struct_pattern: Some(x) or Ok(x)
if (patternNode.type !== 'tuple_struct_pattern') return undefined;
// Extract the wrapper type name: Some | Ok
const wrapperTypeNode = patternNode.childForFieldName('type');
if (!wrapperTypeNode) return undefined;
const wrapperName = extractSimpleTypeName(wrapperTypeNode);
if (wrapperName !== 'Some' && wrapperName !== 'Ok' && wrapperName !== 'Err') return undefined;
// Extract the inner variable name from the single child of the tuple_struct_pattern.
// `Some(x)` → the first named child after the type field is the identifier.
// tree-sitter-rust: tuple_struct_pattern has 'type' field + unnamed children for args.
let innerVar: string | undefined;
for (let i = 0; i < patternNode.namedChildCount; i++) {
const child = patternNode.namedChild(i);
if (!child) continue;
// Skip the type node itself
if (child === wrapperTypeNode) continue;
if (child.type === 'identifier') {
innerVar = child.text;
break;
}
}
if (!innerVar) return undefined;
// The value must be a simple identifier so we can look it up in scopeEnv
const sourceVarName = valueNode.type === 'identifier' ? valueNode.text : undefined;
if (!sourceVarName) return undefined;
// For `Some(x)`: Option<T> is already unwrapped to T in scopeEnv (via NULLABLE_WRAPPER_TYPES).
// For `Ok(x)`: Result<T, E> stores "Result" in scopeEnv — must use declarationTypeNodes.
if (wrapperName === 'Some') {
const innerType = scopeEnv.get(sourceVarName);
if (!innerType) return undefined;
return { varName: innerVar, typeName: innerType };
}
// wrapperName === 'Ok' or 'Err': look up the Result<T, E> type AST node.
// Ok(x) → extract T (typeArgs[0]), Err(e) → extract E (typeArgs[1]).
const typeNodeKey = `${scope}\0${sourceVarName}`;
const typeAstNode = declarationTypeNodes.get(typeNodeKey);
if (!typeAstNode) return undefined;
const typeArgs = extractGenericTypeArgs(typeAstNode);
const argIndex = wrapperName === 'Err' ? 1 : 0;
if (typeArgs.length < argIndex + 1) return undefined;
return { varName: innerVar, typeName: typeArgs[argIndex] };
};
// --- For-loop Tier 1c ---
const FOR_LOOP_NODE_TYPES: ReadonlySet<string> = new Set(['for_expression']);
/** Extract element type from a Rust type annotation AST node.
* Handles: generic_type (Vec<User>), reference_type (&[User]), array_type ([User; N]),
* slice_type ([User]). For call-graph purposes, strips references (&User → User). */
const extractRustElementTypeFromTypeNode = (typeNode: SyntaxNode, pos: TypeArgPosition = 'last', depth = 0): string | undefined => {
if (depth > 50) return undefined;
// generic_type: Vec<User>, HashMap<K, V> — extract type arg based on position
if (typeNode.type === 'generic_type') {
const args = extractGenericTypeArgs(typeNode);
if (args.length >= 1) return pos === 'first' ? args[0] : args[args.length - 1];
}
// reference_type: &[User] or &Vec<User> — unwrap the reference and recurse
if (typeNode.type === 'reference_type') {
const inner = typeNode.lastNamedChild;
if (inner) return extractRustElementTypeFromTypeNode(inner, pos, depth + 1);
}
// array_type: [User; N] — element is the first child
if (typeNode.type === 'array_type') {
const elemNode = typeNode.firstNamedChild;
if (elemNode) return extractSimpleTypeName(elemNode);
}
// slice_type: [User] — element is the first child
if (typeNode.type === 'slice_type') {
const elemNode = typeNode.firstNamedChild;
if (elemNode) return extractSimpleTypeName(elemNode);
}
return undefined;
};
/** Walk up from a for-loop to the enclosing function_item and search parameters
* for one named `iterableName`. Returns the element type from its annotation. */
const findRustParamElementType = (iterableName: string, startNode: SyntaxNode, pos: TypeArgPosition = 'last'): string | undefined => {
let current: SyntaxNode | null = startNode.parent;
while (current) {
if (current.type === 'function_item') {
const paramsNode = current.childForFieldName('parameters');
if (paramsNode) {
for (let i = 0; i < paramsNode.namedChildCount; i++) {
const param = paramsNode.namedChild(i);
if (!param || param.type !== 'parameter') continue;
const nameNode = param.childForFieldName('pattern');
if (!nameNode) continue;
// Unwrap reference patterns: &users, &mut users
let identNode = nameNode;
if (identNode.type === 'reference_pattern') {
identNode = identNode.lastNamedChild ?? identNode;
}
if (identNode.type === 'mut_pattern') {
identNode = identNode.firstNamedChild ?? identNode;
}
if (identNode.text !== iterableName) continue;
const typeNode = param.childForFieldName('type');
if (typeNode) return extractRustElementTypeFromTypeNode(typeNode, pos);
}
}
break;
}
current = current.parent;
}
return undefined;
};
/** Rust: for user in &users where users has a known container type.
* Unwraps reference_expression (&users, &mut users) to get the iterable name. */
const extractForLoopBinding: ForLoopExtractor = (
node: SyntaxNode,
scopeEnv: Map<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
): void => {
if (node.type !== 'for_expression') return;
const patternNode = node.childForFieldName('pattern');
const valueNode = node.childForFieldName('value');
if (!patternNode || !valueNode) return;
// Extract iterable name + method — may be &users, users, or users.iter()/keys()/values()
let iterableName: string | undefined;
let methodName: string | undefined;
if (valueNode.type === 'reference_expression') {
const inner = valueNode.lastNamedChild;
if (inner?.type === 'identifier') iterableName = inner.text;
} else if (valueNode.type === 'identifier') {
iterableName = valueNode.text;
} else if (valueNode.type === 'field_expression') {
const prop = valueNode.lastNamedChild;
if (prop) iterableName = prop.text;
} else if (valueNode.type === 'call_expression') {
// users.iter() → call_expression > function: field_expression > identifier + field_identifier
const fieldExpr = valueNode.childForFieldName('function');
if (fieldExpr?.type === 'field_expression') {
const obj = fieldExpr.firstNamedChild;
if (obj?.type === 'identifier') iterableName = obj.text;
// Extract method name: iter, keys, values, into_iter, etc.
const field = fieldExpr.lastNamedChild;
if (field?.type === 'field_identifier') methodName = field.text;
}
}
if (!iterableName) return;
const containerTypeName = scopeEnv.get(iterableName);
const typeArgPos = methodToTypeArgPosition(methodName, containerTypeName);
const elementType = resolveIterableElementType(
iterableName, node, scopeEnv, declarationTypeNodes, scope,
extractRustElementTypeFromTypeNode, findRustParamElementType,
typeArgPos,
);
if (!elementType) return;
const loopVarName = extractVarName(patternNode);
if (loopVarName) scopeEnv.set(loopVarName, elementType);
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
forLoopNodeTypes: FOR_LOOP_NODE_TYPES,
patternBindingNodeTypes: new Set(['let_condition', 'match_arm']),
extractDeclaration,
extractInitializer,
extractParameter,
scanConstructorBinding,
extractForLoopBinding,
extractPendingAssignment,
extractPatternBinding,
};
@@ -1,61 +1,282 @@
import type { SyntaxNode } from '../utils.js';
/** Which type argument to extract from a multi-arg generic container.
* - 'first': key type (e.g., K from Map<K,V>) — used for .keys(), .keySet()
* - 'last': value type (e.g., V from Map<K,V>) — used for .values(), .items(), .iter() */
export type TypeArgPosition = 'first' | 'last';
// ---------------------------------------------------------------------------
// Container type descriptors — maps container base names to type parameter
// semantics per access method. Replaces the simple KEY_METHODS heuristic.
//
// For user-defined generics (MyCache<K,V> extends Map<K,V>), heritage-aware
// fallback can walk the EXTENDS chain to find a matching descriptor.
// ---------------------------------------------------------------------------
/** Describes which type parameter position each access method yields. */
interface ContainerDescriptor {
/** Number of type parameters (1 = single-element, 2 = key-value) */
arity: number;
/** Methods that yield the first type parameter (key type for maps) */
keyMethods: ReadonlySet<string>;
/** Methods that yield the last type parameter (value type) */
valueMethods: ReadonlySet<string>;
}
/** Empty set for containers that have no key-yielding methods */
const NO_KEYS: ReadonlySet<string> = new Set();
/** Standard key-yielding methods across languages */
const STD_KEY_METHODS: ReadonlySet<string> = new Set(['keys']);
const JAVA_KEY_METHODS: ReadonlySet<string> = new Set(['keySet']);
const CSHARP_KEY_METHODS: ReadonlySet<string> = new Set(['Keys']);
/** Standard value-yielding methods across languages */
const STD_VALUE_METHODS: ReadonlySet<string> = new Set(['values', 'get', 'pop', 'remove']);
const CSHARP_VALUE_METHODS: ReadonlySet<string> = new Set(['Values', 'TryGetValue']);
const SINGLE_ELEMENT_METHODS: ReadonlySet<string> = new Set([
'iter', 'into_iter', 'iterator', 'get', 'first', 'last', 'pop',
'peek', 'poll', 'find', 'filter', 'map',
]);
const CONTAINER_DESCRIPTORS: ReadonlyMap<string, ContainerDescriptor> = new Map([
// --- Map / Dict types (arity 2: key + value) ---
['Map', { arity: 2, keyMethods: STD_KEY_METHODS, valueMethods: STD_VALUE_METHODS }],
['WeakMap', { arity: 2, keyMethods: STD_KEY_METHODS, valueMethods: STD_VALUE_METHODS }],
['HashMap', { arity: 2, keyMethods: STD_KEY_METHODS, valueMethods: STD_VALUE_METHODS }],
['BTreeMap', { arity: 2, keyMethods: STD_KEY_METHODS, valueMethods: STD_VALUE_METHODS }],
['LinkedHashMap', { arity: 2, keyMethods: JAVA_KEY_METHODS, valueMethods: STD_VALUE_METHODS }],
['TreeMap', { arity: 2, keyMethods: JAVA_KEY_METHODS, valueMethods: STD_VALUE_METHODS }],
['dict', { arity: 2, keyMethods: STD_KEY_METHODS, valueMethods: STD_VALUE_METHODS }],
['Dict', { arity: 2, keyMethods: STD_KEY_METHODS, valueMethods: STD_VALUE_METHODS }],
['Dictionary', { arity: 2, keyMethods: CSHARP_KEY_METHODS, valueMethods: CSHARP_VALUE_METHODS }],
['SortedDictionary', { arity: 2, keyMethods: CSHARP_KEY_METHODS, valueMethods: CSHARP_VALUE_METHODS }],
['Record', { arity: 2, keyMethods: STD_KEY_METHODS, valueMethods: STD_VALUE_METHODS }],
['OrderedDict', { arity: 2, keyMethods: STD_KEY_METHODS, valueMethods: STD_VALUE_METHODS }],
['ConcurrentHashMap', { arity: 2, keyMethods: JAVA_KEY_METHODS, valueMethods: STD_VALUE_METHODS }],
['ConcurrentDictionary', { arity: 2, keyMethods: CSHARP_KEY_METHODS, valueMethods: CSHARP_VALUE_METHODS }],
// --- Single-element containers (arity 1) ---
['Array', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['List', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['ArrayList', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['LinkedList',{ arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['Vec', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['VecDeque', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['Set', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['HashSet', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['BTreeSet', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['TreeSet', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['Queue', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['Deque', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['Stack', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['Sequence', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['Iterable', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['Iterator', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['IEnumerable', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['IList', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['ICollection', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['Collection', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['ObservableCollection', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['IEnumerator', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['SortedSet', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['Stream', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['MutableList', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['MutableSet', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['LinkedHashSet', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['ArrayDeque', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['PriorityQueue', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['MutableMap', { arity: 2, keyMethods: STD_KEY_METHODS, valueMethods: STD_VALUE_METHODS }],
['list', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['set', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['tuple', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
['frozenset', { arity: 1, keyMethods: NO_KEYS, valueMethods: SINGLE_ELEMENT_METHODS }],
]);
/** Determine which type arg to extract based on container type name and access method.
*
* Resolution order:
* 1. If container is known and method is in keyMethods → 'first'
* 2. If container is known with arity 1 → 'last' (same as 'first' for single-arg)
* 3. If container is unknown → fall back to method name heuristic
* 4. Default: 'last' (value type)
*/
export function methodToTypeArgPosition(methodName: string | undefined, containerTypeName?: string): TypeArgPosition {
if (containerTypeName) {
const desc = CONTAINER_DESCRIPTORS.get(containerTypeName);
if (desc) {
// Single-element container: always 'last' (= only arg)
if (desc.arity === 1) return 'last';
// Multi-element: check if method yields key type
if (methodName && desc.keyMethods.has(methodName)) return 'first';
// Default for multi-element: value type
return 'last';
}
}
// Fallback for unknown containers: simple method name heuristic
if (methodName && (methodName === 'keys' || methodName === 'keySet' || methodName === 'Keys')) {
return 'first';
}
return 'last';
}
/** Look up the container descriptor for a type name. Exported for heritage-chain lookups. */
export function getContainerDescriptor(typeName: string): ContainerDescriptor | undefined {
return CONTAINER_DESCRIPTORS.get(typeName);
}
/**
* Shared 3-strategy fallback for resolving the element type of a container variable.
* Used by all for-loop extractors to resolve the loop variable's type from the iterable.
*
* Strategy 1: declarationTypeNodes — raw AST type annotation node (handles container types
* where extractSimpleTypeName returned undefined, e.g., User[], List[User])
* Strategy 2: scopeEnv string — extractElementTypeFromString on the stored type string
* Strategy 3: AST walk — language-specific upward walk to enclosing function parameters
*
* @param extractFromTypeNode Language-specific function to extract element type from AST node
* @param findParamElementType Optional language-specific AST walk to find parameter type
* @param typeArgPos Which generic type arg to extract: 'first' for keys, 'last' for values (default)
*/
export function resolveIterableElementType(
iterableName: string,
node: SyntaxNode,
scopeEnv: ReadonlyMap<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
extractFromTypeNode: (typeNode: SyntaxNode, pos?: TypeArgPosition) => string | undefined,
findParamElementType?: (name: string, startNode: SyntaxNode, pos?: TypeArgPosition) => string | undefined,
typeArgPos: TypeArgPosition = 'last',
): string | undefined {
// Strategy 1: declarationTypeNodes AST node (check current scope, then file scope)
const typeNode = declarationTypeNodes.get(`${scope}\0${iterableName}`)
?? (scope !== '' ? declarationTypeNodes.get(`\0${iterableName}`) : undefined);
if (typeNode) {
const t = extractFromTypeNode(typeNode, typeArgPos);
if (t) return t;
}
// Strategy 2: scopeEnv string → extractElementTypeFromString
const iterableType = scopeEnv.get(iterableName);
if (iterableType) {
const el = extractElementTypeFromString(iterableType, typeArgPos);
if (el) return el;
}
// Strategy 3: AST walk to function parameters
if (findParamElementType) return findParamElementType(iterableName, node, typeArgPos);
return undefined;
}
/** Known single-arg nullable wrapper types that unwrap to their inner type
* for receiver resolution. Optional<User> → "User", Option<User> → "User".
* Only nullable wrappers — NOT containers (List, Vec) or async wrappers (Promise, Future).
* See call-processor.ts WRAPPER_GENERICS for the full set used in return-type inference. */
const NULLABLE_WRAPPER_TYPES = new Set([
'Optional', // Java
'Option', // Rust, Scala
'Maybe', // Haskell-style, Kotlin Arrow
]);
/**
* Extract the simple type name from a type AST node.
* Handles generic types (e.g., List<User> → List), qualified names
* (e.g., models.User → User), and nullable types (e.g., User? → User).
* Returns undefined for complex types (unions, intersections, function types).
*/
export const extractSimpleTypeName = (typeNode: SyntaxNode): string | undefined => {
// Direct type identifier
export const extractSimpleTypeName = (typeNode: SyntaxNode, depth = 0): string | undefined => {
if (depth > 50 || typeNode.text.length > 2048) return undefined;
// Direct type identifier (includes Ruby 'constant' for class names)
if (typeNode.type === 'type_identifier' || typeNode.type === 'identifier'
|| typeNode.type === 'simple_identifier') {
|| typeNode.type === 'simple_identifier' || typeNode.type === 'constant') {
return typeNode.text;
}
// Qualified/scoped names: take the last segment (e.g., models.User → User)
// Qualified/scoped names: take the last segment (e.g., models.User → User, Models::User → User)
if (typeNode.type === 'scoped_identifier' || typeNode.type === 'qualified_identifier'
|| typeNode.type === 'scoped_type_identifier' || typeNode.type === 'qualified_name'
|| typeNode.type === 'qualified_type'
|| typeNode.type === 'member_expression' || typeNode.type === 'attribute') {
|| typeNode.type === 'member_expression' || typeNode.type === 'member_access_expression'
|| typeNode.type === 'attribute'
|| typeNode.type === 'scope_resolution'
|| typeNode.type === 'selector_expression') {
const last = typeNode.lastNamedChild;
if (last && (last.type === 'type_identifier' || last.type === 'identifier'
|| last.type === 'simple_identifier' || last.type === 'name')) {
|| last.type === 'simple_identifier' || last.type === 'name'
|| last.type === 'constant' || last.type === 'property_identifier'
|| last.type === 'field_identifier')) {
return last.text;
}
}
// C++ template_type (e.g., vector<User>, map<string, User>): extract base name
if (typeNode.type === 'template_type') {
const base = typeNode.childForFieldName('name') ?? typeNode.firstNamedChild;
if (base) return extractSimpleTypeName(base, depth + 1);
}
// Generic types: extract the base type (e.g., List<User> → List)
if (typeNode.type === 'generic_type' || typeNode.type === 'parameterized_type') {
// For nullable wrappers (Optional<User>, Option<User>), unwrap to inner type.
if (typeNode.type === 'generic_type' || typeNode.type === 'parameterized_type'
|| typeNode.type === 'generic_name') {
const base = typeNode.childForFieldName('name')
?? typeNode.childForFieldName('type')
?? typeNode.firstNamedChild;
if (base) return extractSimpleTypeName(base);
if (!base) return undefined;
const baseName = extractSimpleTypeName(base, depth + 1);
// Unwrap known nullable wrappers: Optional<User> → User, Option<User> → User
if (baseName && NULLABLE_WRAPPER_TYPES.has(baseName)) {
const args = extractGenericTypeArgs(typeNode);
if (args.length >= 1) return args[0];
}
return baseName;
}
// Nullable types (Kotlin User?, C# User?)
if (typeNode.type === 'nullable_type') {
const inner = typeNode.firstNamedChild;
if (inner) return extractSimpleTypeName(inner);
if (inner) return extractSimpleTypeName(inner, depth + 1);
}
// Nullable union types (TS/JS: User | null, User | undefined, User | null | undefined)
// Extract the single non-null/undefined type from the union.
if (typeNode.type === 'union_type') {
const nonNullTypes: SyntaxNode[] = [];
for (let i = 0; i < typeNode.namedChildCount; i++) {
const child = typeNode.namedChild(i);
if (!child) continue;
// Skip null/undefined/void literal types
const text = child.text;
if (text === 'null' || text === 'undefined' || text === 'void') continue;
nonNullTypes.push(child);
}
// Only unwrap if exactly one meaningful type remains
if (nonNullTypes.length === 1) {
return extractSimpleTypeName(nonNullTypes[0], depth + 1);
}
}
// Type annotations that wrap the actual type (TS/Python: `: Foo`, Kotlin: user_type)
if (typeNode.type === 'type_annotation' || typeNode.type === 'type'
|| typeNode.type === 'user_type') {
const inner = typeNode.firstNamedChild;
if (inner) return extractSimpleTypeName(inner);
if (inner) return extractSimpleTypeName(inner, depth + 1);
}
// Pointer/reference types (C++, Rust): User*, &User, &mut User
if (typeNode.type === 'pointer_type' || typeNode.type === 'reference_type') {
const inner = typeNode.firstNamedChild;
if (inner) return extractSimpleTypeName(inner);
if (inner) return extractSimpleTypeName(inner, depth + 1);
}
// Primitive/predefined types: string, int, float, bool, number, unknown, any
// PHP: primitive_type; TS/JS: predefined_type
if (typeNode.type === 'primitive_type' || typeNode.type === 'predefined_type') {
return typeNode.text;
}
// PHP named_type / optional_type
if (typeNode.type === 'named_type' || typeNode.type === 'optional_type') {
const inner = typeNode.childForFieldName('name') ?? typeNode.firstNamedChild;
if (inner) return extractSimpleTypeName(inner);
if (inner) return extractSimpleTypeName(inner, depth + 1);
}
// Name node (PHP)
@@ -72,7 +293,8 @@ export const extractSimpleTypeName = (typeNode: SyntaxNode): string | undefined
*/
export const extractVarName = (node: SyntaxNode): string | undefined => {
if (node.type === 'identifier' || node.type === 'simple_identifier'
|| node.type === 'variable_name' || node.type === 'name') {
|| node.type === 'variable_name' || node.type === 'name'
|| node.type === 'constant' || node.type === 'property_identifier') {
return node.text;
}
// variable_declarator (Java/C#): has a 'name' field
@@ -80,6 +302,11 @@ export const extractVarName = (node: SyntaxNode): string | undefined => {
const nameChild = node.childForFieldName('name');
if (nameChild) return extractVarName(nameChild);
}
// Rust: let mut x = ... — mut_pattern wraps an identifier
if (node.type === 'mut_pattern') {
const inner = node.firstNamedChild;
if (inner) return extractVarName(inner);
}
return undefined;
};
@@ -89,10 +316,177 @@ export const TYPED_PARAMETER_TYPES = new Set([
'optional_parameter', // TS: (x?: Foo)
'formal_parameter', // Java/Kotlin
'parameter', // C#/Rust/Go/Python/Swift
'typed_parameter', // Python: def f(x: Foo) — distinct from 'parameter' in tree-sitter-python
'parameter_declaration', // C/C++ void f(Type name)
'simple_parameter', // PHP function(Foo $x)
'property_promotion_parameter', // PHP 8.0+ constructor promotion: __construct(private Foo $x)
'closure_parameter', // Rust: |user: User| — typed closure parameters
]);
/**
* Extract type arguments from a generic type node.
* e.g., List<User, String> → ['User', 'String'], Vec<User> → ['User']
*
* Used by extractSimpleTypeName to unwrap nullable wrappers (Optional<User> → User).
*
* Handles language-specific AST structures:
* - TS/Java/Rust/Go: generic_type > type_arguments > type nodes
* - C#: generic_type > type_argument_list > type nodes
* - Kotlin: generic_type > type_arguments > type_projection > type nodes
*
* Note: Go slices/maps use slice_type/map_type, not generic_type — those are
* NOT handled here. Use language-specific extractors for Go container types.
*
* @param typeNode A generic_type or parameterized_type AST node (or any node —
* returns [] for non-generic types).
* @returns Array of resolved type argument names. Unresolvable arguments are omitted.
*/
export const extractGenericTypeArgs = (typeNode: SyntaxNode, depth = 0): string[] => {
if (depth > 50) return [];
// Unwrap wrapper nodes that may sit above the generic_type
if (typeNode.type === 'type_annotation' || typeNode.type === 'type'
|| typeNode.type === 'user_type' || typeNode.type === 'nullable_type'
|| typeNode.type === 'optional_type') {
const inner = typeNode.firstNamedChild;
if (inner) return extractGenericTypeArgs(inner, depth + 1);
return [];
}
// Only process generic/parameterized type nodes (includes C#'s generic_name)
if (typeNode.type !== 'generic_type' && typeNode.type !== 'parameterized_type'
&& typeNode.type !== 'generic_name') {
return [];
}
// Find the type_arguments / type_argument_list child
let argsNode: SyntaxNode | null = null;
for (let i = 0; i < typeNode.namedChildCount; i++) {
const child = typeNode.namedChild(i);
if (child && (child.type === 'type_arguments' || child.type === 'type_argument_list')) {
argsNode = child;
break;
}
}
if (!argsNode) return [];
const result: string[] = [];
for (let i = 0; i < argsNode.namedChildCount; i++) {
let argNode = argsNode.namedChild(i);
if (!argNode) continue;
// Kotlin: type_arguments > type_projection > user_type > type_identifier
if (argNode.type === 'type_projection') {
argNode = argNode.firstNamedChild;
if (!argNode) continue;
}
const name = extractSimpleTypeName(argNode);
if (name) result.push(name);
}
return result;
};
/**
* Match Ruby constructor assignment: `user = User.new` or `service = Models::User.new`.
* Returns { varName, calleeName } or undefined if the node is not a Ruby constructor assignment.
* Handles both simple constants and scope_resolution (namespaced) receivers.
*/
export const extractRubyConstructorAssignment = (
node: SyntaxNode,
): { varName: string; calleeName: string } | undefined => {
if (node.type !== 'assignment') return undefined;
const left = node.childForFieldName('left');
const right = node.childForFieldName('right');
if (!left || !right) return undefined;
if (left.type !== 'identifier' && left.type !== 'constant') return undefined;
if (right.type !== 'call') return undefined;
const method = right.childForFieldName('method');
if (!method || method.text !== 'new') return undefined;
const receiver = right.childForFieldName('receiver');
if (!receiver) return undefined;
let calleeName: string;
if (receiver.type === 'constant') {
calleeName = receiver.text;
} else if (receiver.type === 'scope_resolution') {
// Models::User → extract last segment "User"
const last = receiver.lastNamedChild;
if (!last || last.type !== 'constant') return undefined;
calleeName = last.text;
} else {
return undefined;
}
return { varName: left.text, calleeName };
};
/**
* Check if an AST node has an explicit type annotation.
* Checks both named fields ('type') and child nodes ('type_annotation').
* Used by constructor binding scanners to skip annotated declarations.
*/
export const hasTypeAnnotation = (node: SyntaxNode): boolean => {
if (node.childForFieldName('type')) return true;
for (let i = 0; i < node.childCount; i++) {
if (node.child(i)?.type === 'type_annotation') return true;
}
return false;
};
/** Bare nullable keywords that should not produce a receiver binding. */
const NULLABLE_KEYWORDS = new Set(['null', 'undefined', 'void', 'None', 'nil']);
/**
* Strip nullable wrappers from a type name string.
* Used by both lookupInEnv (TypeEnv annotations) and extractReturnTypeName
* (return-type text) to normalize types before receiver lookup.
*
* "User | null" → "User"
* "User | undefined" → "User"
* "User | null | undefined" → "User"
* "User?" → "User"
* "User | Repo" → undefined (genuine union — refuse)
* "null" → undefined
*/
export const stripNullable = (typeName: string): string | undefined => {
let text = typeName.trim();
if (!text) return undefined;
if (NULLABLE_KEYWORDS.has(text)) return undefined;
// Strip nullable suffix: User? → User
if (text.endsWith('?')) text = text.slice(0, -1).trim();
// Strip union with null/undefined/None/nil/void
if (text.includes('|')) {
const parts = text.split('|').map(p => p.trim()).filter(p =>
p !== '' && !NULLABLE_KEYWORDS.has(p)
);
if (parts.length === 1) return parts[0];
return undefined; // genuine union or all-nullable — refuse
}
return text || undefined;
};
/**
* Unwrap an await_expression to get the inner value.
* Returns the node itself if not an await_expression, or null if input is null.
*/
export const unwrapAwait = (node: SyntaxNode | null): SyntaxNode | null => {
if (!node) return null;
return node.type === 'await_expression' ? node.firstNamedChild : node;
};
/**
* Extract the callee name from a call_expression node.
* Navigates to the 'function' field (or first named child) and extracts a simple type name.
*/
export const extractCalleeName = (callNode: SyntaxNode): string | undefined => {
const func = callNode.childForFieldName('function') ?? callNode.firstNamedChild;
if (!func) return undefined;
return extractSimpleTypeName(func);
};
/** Find the first named child with the given node type */
export const findChildByType = (node: SyntaxNode, type: string): SyntaxNode | null => {
for (let i = 0; i < node.namedChildCount; i++) {
@@ -101,3 +495,116 @@ export const findChildByType = (node: SyntaxNode, type: string): SyntaxNode | nu
}
return null;
};
// Internal helper: extract the first comma-separated argument from a string,
// respecting nested angle-bracket and square-bracket depth.
function extractFirstArg(args: string): string {
let depth = 0;
for (let i = 0; i < args.length; i++) {
const ch = args[i];
if (ch === '<' || ch === '[') depth++;
else if (ch === '>' || ch === ']') depth--;
else if (ch === ',' && depth === 0) return args.slice(0, i).trim();
}
return args.trim();
}
/**
* Extract element type from a container type string.
* Uses bracket-balanced parsing (no regex) for generic argument extraction.
* Returns undefined for ambiguous or unparseable strings.
*
* Handles:
* - Array<User> → User (generic angle brackets)
* - User[] → User (array suffix)
* - []User → User (Go slice prefix)
* - List[User] → User (Python subscript)
* - [User] → User (Swift array sugar)
* - vector<User> → User (C++ container)
* - Vec<User> → User (Rust container)
*
* For multi-argument generics (Map<K, V>), returns the first or last type arg
* based on `pos` ('first' for keys, 'last' for values — default 'last').
* Returns undefined when the extracted type is not a simple word.
*/
export function extractElementTypeFromString(typeStr: string, pos: TypeArgPosition = 'last'): string | undefined {
if (!typeStr || typeStr.length === 0 || typeStr.length > 2048) return undefined;
// 1. Array suffix: User[] → User
if (typeStr.endsWith('[]')) {
const base = typeStr.slice(0, -2).trim();
return base && /^\w+$/.test(base) ? base : undefined;
}
// 2. Go slice prefix: []User → User
if (typeStr.startsWith('[]')) {
const element = typeStr.slice(2).trim();
return element && /^\w+$/.test(element) ? element : undefined;
}
// 3. Swift array sugar: [User] → User
// Must start with '[', end with ']', and contain no angle brackets
// (to avoid confusing with List[User] handled below).
if (typeStr.startsWith('[') && typeStr.endsWith(']') && !typeStr.includes('<')) {
const element = typeStr.slice(1, -1).trim();
return element && /^\w+$/.test(element) ? element : undefined;
}
// 4. Generic bracket-balanced extraction: Array<User> / List[User] / Vec<User>
// Find the first opening bracket (< or [) and pick the one that appears first.
const openAngle = typeStr.indexOf('<');
const openSquare = typeStr.indexOf('[');
let openIdx = -1;
let openChar = '';
let closeChar = '';
if (openAngle >= 0 && (openSquare < 0 || openAngle < openSquare)) {
openIdx = openAngle;
openChar = '<';
closeChar = '>';
} else if (openSquare >= 0) {
openIdx = openSquare;
openChar = '[';
closeChar = ']';
}
if (openIdx < 0) return undefined;
// Walk bracket-balanced from the character after the opening bracket to find
// the matching close bracket, tracking depth for nested brackets.
// All bracket types (<, >, [, ]) contribute to depth uniformly, but only the
// selected closeChar can match at depth 0 (prevents cross-bracket miscounting).
let depth = 0;
const start = openIdx + 1;
let lastCommaIdx = -1; // Track last top-level comma for 'last' position
for (let i = start; i < typeStr.length; i++) {
const ch = typeStr[i];
if (ch === '<' || ch === '[') {
depth++;
} else if (ch === '>' || ch === ']') {
if (depth === 0) {
// At depth 0 — only match if it is our selected close bracket.
if (ch !== closeChar) return undefined; // mismatched bracket = malformed
if (pos === 'last' && lastCommaIdx >= 0) {
// Return last arg (text after last comma)
const lastArg = typeStr.slice(lastCommaIdx + 1, i).trim();
return lastArg && /^\w+$/.test(lastArg) ? lastArg : undefined;
}
const inner = typeStr.slice(start, i).trim();
const firstArg = extractFirstArg(inner);
return firstArg && /^\w+$/.test(firstArg) ? firstArg : undefined;
}
depth--;
} else if (ch === ',' && depth === 0) {
if (pos === 'first') {
// Return first arg (text before first comma)
const arg = typeStr.slice(start, i).trim();
return arg && /^\w+$/.test(arg) ? arg : undefined;
}
lastCommaIdx = i;
}
}
return undefined;
}
@@ -1,6 +1,6 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, findChildByType } from './shared.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor, InitializerExtractor, ClassNameLookup, ConstructorBindingScanner } from './types.js';
import { extractSimpleTypeName, extractVarName, findChildByType, hasTypeAnnotation } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'property_declaration',
@@ -39,8 +39,89 @@ const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string,
if (varName && typeName) env.set(varName, typeName);
};
/** Swift: let user = User(name: "alice") — infer type from call when callee is a known class.
* Swift initializers are syntactically identical to function calls, so we verify
* against classNames (which may include cross-file SymbolTable lookups). */
const extractInitializer: InitializerExtractor = (node: SyntaxNode, env: Map<string, string>, classNames: ClassNameLookup): void => {
if (node.type !== 'property_declaration') return;
// Skip if has type annotation — extractDeclaration handled it
if (node.childForFieldName('type') || findChildByType(node, 'type_annotation')) return;
// Find pattern (variable name)
const pattern = node.childForFieldName('pattern') ?? findChildByType(node, 'pattern');
if (!pattern) return;
const varName = extractVarName(pattern) ?? pattern.text;
if (!varName || env.has(varName)) return;
// Find call_expression in the value
const callExpr = findChildByType(node, 'call_expression');
if (!callExpr) return;
const callee = callExpr.firstNamedChild;
if (!callee) return;
// Direct call: User(name: "alice")
if (callee.type === 'simple_identifier') {
const calleeName = callee.text;
if (calleeName && classNames.has(calleeName)) {
env.set(varName, calleeName);
}
return;
}
// Explicit init: User.init(name: "alice") — navigation_expression with .init suffix
if (callee.type === 'navigation_expression') {
const receiver = callee.firstNamedChild;
const suffix = callee.lastNamedChild;
if (receiver?.type === 'simple_identifier' && suffix?.text === 'init') {
const calleeName = receiver.text;
if (calleeName && classNames.has(calleeName)) {
env.set(varName, calleeName);
}
}
}
};
/** Swift: let user = User(name: "alice") — scan property_declaration for constructor binding */
const scanConstructorBinding: ConstructorBindingScanner = (node) => {
if (node.type !== 'property_declaration') return undefined;
if (hasTypeAnnotation(node)) return undefined;
const pattern = node.childForFieldName('pattern');
if (!pattern) return undefined;
const varName = pattern.text;
if (!varName) return undefined;
let callExpr: SyntaxNode | null = null;
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'call_expression') { callExpr = child; break; }
}
if (!callExpr) return undefined;
const callee = callExpr.firstNamedChild;
if (!callee) return undefined;
if (callee.type === 'simple_identifier') {
return { varName, calleeName: callee.text };
}
if (callee.type === 'navigation_expression') {
const receiver = callee.firstNamedChild;
const suffix = callee.lastNamedChild;
if (receiver?.type === 'simple_identifier' && suffix?.text === 'init') {
return { varName, calleeName: receiver.text };
}
// General qualified call: service.getUser() → extract method name.
// tree-sitter-swift may wrap the identifier in navigation_suffix, so
// check both direct simple_identifier and navigation_suffix > simple_identifier.
if (suffix?.type === 'simple_identifier') {
return { varName, calleeName: suffix.text };
}
if (suffix?.type === 'navigation_suffix') {
const inner = suffix.lastNamedChild;
if (inner?.type === 'simple_identifier') {
return { varName, calleeName: inner.text };
}
}
}
return undefined;
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
extractDeclaration,
extractParameter,
extractInitializer,
scanConstructorBinding,
};
@@ -6,12 +6,101 @@ export type TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>)
/** Extracts type bindings from a parameter node into the env map */
export type ParameterExtractor = (node: SyntaxNode, env: Map<string, string>) => void;
/** Minimal interface for checking whether a name is a known class/struct.
* Narrower than ReadonlySet — only `.has()` is used by extractors. */
export type ClassNameLookup = { has(name: string): boolean };
/** Extracts type bindings from a constructor-call initializer, with access to known class names */
export type InitializerExtractor = (node: SyntaxNode, env: Map<string, string>, classNames: ClassNameLookup) => void;
/** Scans an AST node for untyped `var = callee()` patterns for return-type inference.
* Returns { varName, calleeName } if the node matches, undefined otherwise.
* `receiverClassName` — optional hint for method calls on known receivers
* (e.g. $this->getUser() in PHP provides the enclosing class name). */
export type ConstructorBindingScanner = (node: SyntaxNode) => { varName: string; calleeName: string; receiverClassName?: string } | undefined;
/** Extracts a return type string from a method/function definition node.
* Used for languages where return types are expressed in comments (e.g. YARD @return [Type])
* rather than in AST fields. Returns undefined if no return type can be determined. */
export type ReturnTypeExtractor = (node: SyntaxNode) => string | undefined;
/** Extracts loop variable type binding from a for-each statement.
* All parameters are required (aligned with PatternBindingExtractor convention)
* to prevent new extractors from silently ignoring declarationTypeNodes/scope. */
export type ForLoopExtractor = (
node: SyntaxNode,
scopeEnv: Map<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
) => void;
/** Extracts a plain-identifier assignment for Tier 2 propagation.
* For `const b = a`, returns { lhs: 'b', rhs: 'a' } when the LHS has no resolved type.
* Returns undefined if the node is not a plain identifier assignment. */
export type PendingAssignmentExtractor = (
node: SyntaxNode,
scopeEnv: ReadonlyMap<string, string>,
) => { lhs: string; rhs: string } | undefined;
/** Extracts a typed variable binding from a pattern-matching construct.
* Returns { varName, typeName } for patterns that introduce NEW variables.
* Examples: `if let Some(user) = opt` (Rust), `x instanceof User user` (Java).
* Conservative: returns undefined when the source variable's type is unknown.
*
* @param scopeEnv Read-only view of already-resolved type bindings in the current scope.
* @param declarationTypeNodes Maps `scope\0varName` to the original declaration's type
* annotation AST node. Allows extracting generic type arguments (e.g., T from Result<T,E>)
* that are stripped during normal TypeEnv extraction.
* @param scope Current scope key (e.g. `"process@42"`) for declarationTypeNodes lookups. */
export type PatternBindingExtractor = (
node: SyntaxNode,
scopeEnv: ReadonlyMap<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
) => { varName: string; typeName: string } | undefined;
/** Per-language type extraction configuration */
export interface LanguageTypeConfig {
/** Allow pattern binding to overwrite existing scopeEnv entries.
* WARNING: Enables function-scope type pollution. Only for languages with
* smart-cast semantics (e.g., Kotlin `when/is`) where the subject variable
* already exists in scopeEnv from its declaration. */
readonly allowPatternBindingOverwrite?: boolean;
/** Node types that represent typed declarations for this language */
declarationNodeTypes: ReadonlySet<string>;
/** AST node types for for-each/for-in statements with explicit element types. */
forLoopNodeTypes?: ReadonlySet<string>;
/** Optional allowlist of AST node types on which extractPatternBinding should run.
* When present, extractPatternBinding is only invoked for nodes whose type is in this set,
* short-circuiting the call for all other node types. When absent, every node is passed to
* extractPatternBinding (legacy behaviour). */
patternBindingNodeTypes?: ReadonlySet<string>;
/** Extract a (varName → typeName) binding from a declaration node */
extractDeclaration: TypeBindingExtractor;
/** Extract a (varName → typeName) binding from a parameter node */
extractParameter: ParameterExtractor;
/** Extract a (varName → typeName) binding from a constructor-call initializer.
* Called as fallback when extractDeclaration produces no binding for a declaration node.
* Only for languages with syntactic constructor markers (new, composite_literal, ::new).
* Receives classNames — the set of class/struct names visible in the current file's AST. */
extractInitializer?: InitializerExtractor;
/** Scan for untyped `var = callee()` assignments for return-type inference.
* Called on every AST node during buildTypeEnv walk; returns undefined for non-matches.
* The callee binding is unverified — the caller must confirm against the SymbolTable. */
scanConstructorBinding?: ConstructorBindingScanner;
/** Extract return type from comment-based annotations (e.g. YARD @return [Type]).
* Called as fallback when extractMethodSignature finds no AST-based return type. */
extractReturnType?: ReturnTypeExtractor;
/** Extract loop variable → type binding from a for-each AST node. */
extractForLoopBinding?: ForLoopExtractor;
/** Extract plain-identifier assignment (e.g. `const b = a`) for Tier 2 chain propagation.
* Called on declaration/assignment nodes; returns {lhs, rhs} when the RHS is a bare identifier
* and the LHS has no resolved type yet. Language-specific because AST shapes differ widely. */
extractPendingAssignment?: PendingAssignmentExtractor;
/** Extract a typed variable binding from a pattern-matching construct.
* Called on every AST node; returns { varName, typeName } when the node introduces a new
* typed variable via pattern matching (e.g. `if let Some(x) = opt`, `x instanceof T t`).
* The extractor receives the current scope's resolved bindings (read-only) to look up the
* source variable's type. Returns undefined for non-matching nodes or unknown source types. */
extractPatternBinding?: PatternBindingExtractor;
}
@@ -1,14 +1,98 @@
import type { SyntaxNode } from '../utils.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName } from './shared.js';
import type { LanguageTypeConfig, ParameterExtractor, TypeBindingExtractor, InitializerExtractor, ClassNameLookup, ConstructorBindingScanner, ReturnTypeExtractor, PendingAssignmentExtractor, ForLoopExtractor, PatternBindingExtractor } from './types.js';
import { extractSimpleTypeName, extractVarName, hasTypeAnnotation, unwrapAwait, extractCalleeName, extractElementTypeFromString, extractGenericTypeArgs, resolveIterableElementType, methodToTypeArgPosition, type TypeArgPosition } from './shared.js';
const DECLARATION_NODE_TYPES: ReadonlySet<string> = new Set([
'lexical_declaration',
'variable_declaration',
'function_declaration', // JSDoc @param on function declarations
'method_definition', // JSDoc @param on class methods
'public_field_definition', // class field: private users: User[]
]);
/** TypeScript: const x: Foo = ..., let x: Foo */
const normalizeJsDocType = (raw: string): string | undefined => {
let type = raw.trim();
// Strip JSDoc nullable/non-nullable prefixes: ?User → User, !User → User
if (type.startsWith('?') || type.startsWith('!')) type = type.slice(1);
// Strip union with null/undefined/void: User|null → User
const parts = type.split('|').map(p => p.trim()).filter(p =>
p !== 'null' && p !== 'undefined' && p !== 'void'
);
if (parts.length !== 1) return undefined; // ambiguous union
type = parts[0];
// Strip module: prefix — module:models.User → models.User
if (type.startsWith('module:')) type = type.slice(7);
// Take last segment of dotted path: models.User → User
const segments = type.split('.');
type = segments[segments.length - 1];
// Strip generic wrapper: Promise<User> → Promise (base type, not inner)
const genericMatch = type.match(/^(\w+)\s*</);
if (genericMatch) type = genericMatch[1];
// Simple identifier check
if (/^\w+$/.test(type)) return type;
return undefined;
};
/** Regex to extract JSDoc @param annotations: `@param {Type} name` */
const JSDOC_PARAM_RE = /@param\s*\{([^}]+)\}\s+\[?(\w+)[\]=]?[^\s]*/g;
/**
* Collect JSDoc @param type bindings from comment nodes preceding a function/method.
* Returns a map of paramName → typeName.
*/
const collectJsDocParams = (funcNode: SyntaxNode): Map<string, string> => {
const commentTexts: string[] = [];
let sibling = funcNode.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
commentTexts.unshift(sibling.text);
} else if (sibling.isNamed && sibling.type !== 'decorator') {
break;
}
sibling = sibling.previousSibling;
}
if (commentTexts.length === 0) return new Map();
const params = new Map<string, string>();
const commentBlock = commentTexts.join('\n');
JSDOC_PARAM_RE.lastIndex = 0;
let match: RegExpExecArray | null;
while ((match = JSDOC_PARAM_RE.exec(commentBlock)) !== null) {
const typeName = normalizeJsDocType(match[1]);
const paramName = match[2];
if (typeName) {
params.set(paramName, typeName);
}
}
return params;
};
/**
* TypeScript: const x: Foo = ..., let x: Foo
* Also: JSDoc @param annotations on function/method definitions (for .js files).
*/
const extractDeclaration: TypeBindingExtractor = (node: SyntaxNode, env: Map<string, string>): void => {
// JSDoc @param on functions/methods — pre-populate env with param types
if (node.type === 'function_declaration' || node.type === 'method_definition') {
const jsDocParams = collectJsDocParams(node);
for (const [paramName, typeName] of jsDocParams) {
if (!env.has(paramName)) env.set(paramName, typeName);
}
return;
}
// Class field: `private users: User[]` — public_field_definition has name + type fields directly.
if (node.type === 'public_field_definition') {
const nameNode = node.childForFieldName('name');
const typeAnnotation = node.childForFieldName('type');
if (!nameNode || !typeAnnotation) return;
const varName = nameNode.text;
if (!varName) return;
const typeName = extractSimpleTypeName(typeAnnotation);
if (typeName) env.set(varName, typeName);
return;
}
for (let i = 0; i < node.namedChildCount; i++) {
const declarator = node.namedChild(i);
if (declarator?.type !== 'variable_declarator') continue;
@@ -41,8 +125,344 @@ const extractParameter: ParameterExtractor = (node: SyntaxNode, env: Map<string,
if (varName && typeName) env.set(varName, typeName);
};
/** TypeScript: const x = new User() — infer type from new_expression */
const extractInitializer: InitializerExtractor = (node: SyntaxNode, env: Map<string, string>, _classNames: ClassNameLookup): void => {
for (let i = 0; i < node.namedChildCount; i++) {
const declarator = node.namedChild(i);
if (declarator?.type !== 'variable_declarator') continue;
// Only activate when there is no explicit type annotation — extractDeclaration already
// handles the annotated case and this function is called as a fallback.
if (declarator.childForFieldName('type') !== null) continue;
let valueNode = declarator.childForFieldName('value');
// Unwrap `new User() as T`, `new User()!`, and double-cast `new User() as unknown as T`
while (valueNode?.type === 'as_expression' || valueNode?.type === 'non_null_expression') {
valueNode = valueNode.firstNamedChild;
}
if (valueNode?.type !== 'new_expression') continue;
const constructorNode = valueNode.childForFieldName('constructor');
if (!constructorNode) continue;
const nameNode = declarator.childForFieldName('name');
if (!nameNode) continue;
const varName = extractVarName(nameNode);
const typeName = extractSimpleTypeName(constructorNode);
if (varName && typeName) env.set(varName, typeName);
}
};
/**
* TypeScript/JavaScript: const user = getUser() — variable_declarator with call_expression value.
* Only matches unannotated declarators; annotated ones are handled by extractDeclaration.
* await is unwrapped: const user = await fetchUser() → callee = 'fetchUser'.
*/
const scanConstructorBinding: ConstructorBindingScanner = (node) => {
if (node.type !== 'variable_declarator') return undefined;
if (hasTypeAnnotation(node)) return undefined;
const nameNode = node.childForFieldName('name');
if (!nameNode || nameNode.type !== 'identifier') return undefined;
const value = unwrapAwait(node.childForFieldName('value'));
if (!value || value.type !== 'call_expression') return undefined;
const calleeName = extractCalleeName(value);
if (!calleeName) return undefined;
return { varName: nameNode.text, calleeName };
};
/** Regex to extract @returns or @return from JSDoc comments: `@returns {Type}` */
const JSDOC_RETURN_RE = /@returns?\s*\{([^}]+)\}/;
/**
* Minimal sanitization for JSDoc return types — preserves generic wrappers
* (e.g. `Promise<User>`) so that extractReturnTypeName in call-processor
* can apply WRAPPER_GENERICS unwrapping. Unlike normalizeJsDocType (which
* strips generics), this only strips JSDoc-specific syntax markers.
*/
const sanitizeReturnType = (raw: string): string | undefined => {
let type = raw.trim();
// Strip JSDoc nullable/non-nullable prefixes: ?User → User, !User → User
if (type.startsWith('?') || type.startsWith('!')) type = type.slice(1);
// Strip module: prefix — module:models.User → models.User
if (type.startsWith('module:')) type = type.slice(7);
// Reject unions (ambiguous)
if (type.includes('|')) return undefined;
if (!type) return undefined;
return type;
};
/**
* Extract return type from JSDoc `@returns {Type}` or `@return {Type}` annotation
* preceding a function/method definition. Walks backwards through preceding siblings
* looking for comment nodes containing the annotation.
*/
const extractReturnType: ReturnTypeExtractor = (node) => {
let sibling = node.previousSibling;
while (sibling) {
if (sibling.type === 'comment') {
const match = JSDOC_RETURN_RE.exec(sibling.text);
if (match) return sanitizeReturnType(match[1]);
} else if (sibling.isNamed && sibling.type !== 'decorator') break;
sibling = sibling.previousSibling;
}
return undefined;
};
const FOR_LOOP_NODE_TYPES: ReadonlySet<string> = new Set([
'for_in_statement',
]);
/** TS function/method node types that carry a parameters list. */
const TS_FUNCTION_NODE_TYPES = new Set([
'function_declaration', 'function_expression', 'arrow_function',
'method_definition', 'generator_function', 'generator_function_declaration',
]);
/**
* Extract element type from a TypeScript type annotation AST node.
* Handles:
* type_annotation ": User[]" → array_type → type_identifier "User"
* type_annotation ": Array<User>" → generic_type → extractGenericTypeArgs → "User"
* Falls back to text-based extraction via extractElementTypeFromString.
*/
const extractTsElementTypeFromAnnotation = (typeAnnotation: SyntaxNode, pos: TypeArgPosition = 'last', depth = 0): string | undefined => {
if (depth > 50) return undefined;
// Unwrap type_annotation (the node text includes ': ' prefix)
const inner = typeAnnotation.type === 'type_annotation'
? (typeAnnotation.firstNamedChild ?? typeAnnotation)
: typeAnnotation;
// readonly User[] — readonly_type wraps array_type: unwrap and recurse
if (inner.type === 'readonly_type') {
const wrapped = inner.firstNamedChild;
if (wrapped) return extractTsElementTypeFromAnnotation(wrapped, pos, depth + 1);
}
// User[] — array_type: first named child is the element type
if (inner.type === 'array_type') {
const elem = inner.firstNamedChild;
if (elem) return extractSimpleTypeName(elem);
}
// Array<User>, Map<string, User> — generic_type
// pos determines which type arg: 'first' for keys, 'last' for values
if (inner.type === 'generic_type') {
const args = extractGenericTypeArgs(inner);
if (args.length >= 1) return pos === 'first' ? args[0] : args[args.length - 1];
}
// Fallback: strip ': ' prefix from type_annotation text and use string extraction
const rawText = inner.text;
return extractElementTypeFromString(rawText, pos);
};
/**
* Search a statement_block (function body) for a variable_declarator named `iterableName`
* that has a type annotation, preceding the given `beforeNode`.
* Returns the element type from the type annotation, or undefined.
*/
const findTsLocalDeclElementType = (
iterableName: string,
blockNode: SyntaxNode,
beforeNode: SyntaxNode,
pos: TypeArgPosition = 'last',
): string | undefined => {
for (let i = 0; i < blockNode.namedChildCount; i++) {
const stmt = blockNode.namedChild(i);
if (!stmt) continue;
// Stop when we reach the for-loop itself
if (stmt === beforeNode || stmt.startIndex >= beforeNode.startIndex) break;
// Look for lexical_declaration or variable_declaration
if (stmt.type !== 'lexical_declaration' && stmt.type !== 'variable_declaration') continue;
for (let j = 0; j < stmt.namedChildCount; j++) {
const decl = stmt.namedChild(j);
if (decl?.type !== 'variable_declarator') continue;
const nameNode = decl.childForFieldName('name');
if (nameNode?.text !== iterableName) continue;
const typeAnnotation = decl.childForFieldName('type');
if (typeAnnotation) return extractTsElementTypeFromAnnotation(typeAnnotation, pos);
}
}
return undefined;
};
/**
* Walk up the AST from a for-loop node to find the enclosing function scope,
* then search (1) its parameter list and (2) local declarations in the body
* for a variable named `iterableName` with a container type annotation.
* Returns the element type extracted from the annotation, or undefined.
*/
const findTsIterableElementType = (iterableName: string, startNode: SyntaxNode, pos: TypeArgPosition = 'last'): string | undefined => {
let current: SyntaxNode | null = startNode.parent;
// Capture the immediate statement_block parent to search local declarations
const blockNode = current?.type === 'statement_block' ? current : null;
while (current) {
if (TS_FUNCTION_NODE_TYPES.has(current.type)) {
// Search function parameters
const paramsNode = current.childForFieldName('parameters')
?? current.childForFieldName('formal_parameters');
if (paramsNode) {
for (let i = 0; i < paramsNode.namedChildCount; i++) {
const param = paramsNode.namedChild(i);
if (!param) continue;
const patternNode = param.childForFieldName('pattern') ?? param.childForFieldName('name');
if (patternNode?.text === iterableName) {
const typeAnnotation = param.childForFieldName('type');
if (typeAnnotation) return extractTsElementTypeFromAnnotation(typeAnnotation, pos);
}
}
}
// Search local declarations in the function body (statement_block)
if (blockNode) {
const result = findTsLocalDeclElementType(iterableName, blockNode, startNode, pos);
if (result) return result;
}
break; // stop at the nearest function boundary
}
current = current.parent;
}
return undefined;
};
/**
* TypeScript/JavaScript: for (const user of users) where users has a known array type.
*
* Both `for...of` and `for...in` use the same `for_in_statement` AST node in tree-sitter.
* We differentiate by checking for the `of` keyword among the unnamed children.
*
* Tier 1c: resolves the element type via three strategies in priority order:
* 1. declarationTypeNodes — raw type annotation AST node (covers Array<User> from declarations)
* 2. scopeEnv string — extractElementTypeFromString on the stored type (covers locally annotated vars)
* 3. AST walk — walks up to the enclosing function's parameters to read User[] annotations directly
* Only handles `for...of`; `for...in` produces string keys, not element types.
*/
const extractForLoopBinding: ForLoopExtractor = (
node: SyntaxNode,
scopeEnv: Map<string, string>,
declarationTypeNodes: ReadonlyMap<string, SyntaxNode>,
scope: string,
): void => {
if (node.type !== 'for_in_statement') return;
// Confirm this is `for...of`, not `for...in`, by scanning unnamed children for the keyword text.
let isForOf = false;
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (child && !child.isNamed && child.text === 'of') {
isForOf = true;
break;
}
}
if (!isForOf) return;
// The iterable is the `right` field — may be identifier or call_expression.
const rightNode = node.childForFieldName('right');
let iterableName: string | undefined;
let methodName: string | undefined;
if (rightNode?.type === 'identifier') {
iterableName = rightNode.text;
} else if (rightNode?.type === 'member_expression') {
const prop = rightNode.childForFieldName('property');
if (prop) iterableName = prop.text;
} else if (rightNode?.type === 'call_expression') {
// entries.values() → call_expression > function: member_expression > object + property
// this.repos.values() → nested member_expression: extract property from inner member
const fn = rightNode.childForFieldName('function');
if (fn?.type === 'member_expression') {
const obj = fn.childForFieldName('object');
const prop = fn.childForFieldName('property');
if (obj?.type === 'identifier') {
iterableName = obj.text;
} else if (obj?.type === 'member_expression') {
// this.repos.values() → obj = this.repos → extract 'repos'
const innerProp = obj.childForFieldName('property');
if (innerProp) iterableName = innerProp.text;
}
if (prop?.type === 'property_identifier') methodName = prop.text;
}
}
if (!iterableName) return;
// Look up the container's base type name for descriptor-aware resolution
const containerTypeName = scopeEnv.get(iterableName);
const typeArgPos = methodToTypeArgPosition(methodName, containerTypeName);
const elementType = resolveIterableElementType(
iterableName, node, scopeEnv, declarationTypeNodes, scope,
extractTsElementTypeFromAnnotation, findTsIterableElementType,
typeArgPos,
);
if (!elementType) return;
// The loop variable is the `left` field.
const leftNode = node.childForFieldName('left');
if (!leftNode) return;
// Handle destructured for-of: for (const [k, v] of entries)
// AST: left = array_pattern directly (no variable_declarator wrapper)
// Bind the LAST identifier to the element type (value in [key, value] patterns)
if (leftNode.type === 'array_pattern') {
const lastChild = leftNode.lastNamedChild;
if (lastChild?.type === 'identifier') {
scopeEnv.set(lastChild.text, elementType);
}
return;
}
if (leftNode.type === 'object_pattern') {
// Object destructuring (e.g., `for (const { id } of users)`) destructures
// into fields of the element type. Without field-level resolution, we cannot
// bind individual properties to their correct types. Skip to avoid false bindings.
return;
}
let loopVarNode: SyntaxNode | null = leftNode;
// `const user` parses as: left → variable_declarator containing an identifier named `user`
if (loopVarNode.type === 'variable_declarator') {
loopVarNode = loopVarNode.childForFieldName('name') ?? loopVarNode.firstNamedChild;
}
if (!loopVarNode) return;
const loopVarName = extractVarName(loopVarNode);
if (loopVarName) scopeEnv.set(loopVarName, elementType);
};
/** TS/JS: const alias = u → variable_declarator with name/value fields */
const extractPendingAssignment: PendingAssignmentExtractor = (node, scopeEnv) => {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (!child || child.type !== 'variable_declarator') continue;
const nameNode = child.childForFieldName('name');
const valueNode = child.childForFieldName('value');
if (!nameNode || !valueNode) continue;
const lhs = nameNode.text;
if (scopeEnv.has(lhs)) continue;
if (valueNode.type === 'identifier') return { lhs, rhs: valueNode.text };
}
return undefined;
};
/** TS instanceof narrowing: `x instanceof User` → bind x to User.
* Only works when x has no prior type binding (e.g. x: unknown, untyped params).
* Typed params (x: Animal) are blocked by the !scopeEnv.has() guard in buildTypeEnv.
* Uses first-writer-wins, same as Rust match arm bindings. */
const extractPatternBinding: PatternBindingExtractor = (node) => {
if (node.type !== 'binary_expression') return undefined;
const op = node.children.find(c => !c.isNamed && c.text === 'instanceof');
if (!op) return undefined;
// binary_expression children are positional — no left/right fields
const left = node.namedChild(0);
const right = node.namedChild(1);
if (left?.type !== 'identifier' || right?.type !== 'identifier') return undefined;
return { varName: left.text, typeName: right.text };
};
export const typeConfig: LanguageTypeConfig = {
declarationNodeTypes: DECLARATION_NODE_TYPES,
forLoopNodeTypes: FOR_LOOP_NODE_TYPES,
patternBindingNodeTypes: new Set(['binary_expression']),
extractDeclaration,
extractParameter,
extractInitializer,
scanConstructorBinding,
extractReturnType,
extractForLoopBinding,
extractPendingAssignment,
extractPatternBinding,
};
+384 -32
View File
@@ -35,7 +35,7 @@ export const DEFINITION_CAPTURE_KEYS = [
] as const;
/** Extract the definition node from a tree-sitter query capture map. */
export const getDefinitionNodeFromCaptures = (captureMap: Record<string, any>): any | null => {
export const getDefinitionNodeFromCaptures = (captureMap: Record<string, any>): SyntaxNode | null => {
for (const key of DEFINITION_CAPTURE_KEYS) {
if (captureMap[key]) return captureMap[key];
}
@@ -77,6 +77,9 @@ export const FUNCTION_NODE_TYPES = new Set([
// Swift
'init_declaration',
'deinit_declaration',
// Ruby
'method', // def foo
'singleton_method', // def self.foo
]);
/**
@@ -119,7 +122,7 @@ export const BUILT_IN_NAMES = new Set([
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'append', 'extend', 'update',
// NOTE: 'open', 'read', 'write', 'close' removed — these are real C POSIX syscalls
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
// Kotlin stdlib
'println', 'print', 'readLine', 'require', 'requireNotNull', 'check', 'assert', 'lazy', 'error',
@@ -236,6 +239,21 @@ export const BUILT_IN_NAMES = new Set([
'lock', 'read', 'write', 'try_lock',
'spawn', 'join', 'sleep',
'Some', 'None', 'Ok', 'Err',
// Ruby built-ins and Kernel methods
'puts', 'p', 'pp', 'raise', 'fail',
'require', 'require_relative', 'load', 'autoload',
'include', 'extend', 'prepend',
'attr_accessor', 'attr_reader', 'attr_writer',
'public', 'private', 'protected', 'module_function',
'lambda', 'proc', 'block_given?',
'nil?', 'is_a?', 'kind_of?', 'instance_of?', 'respond_to?',
'freeze', 'frozen?', 'dup', 'tap', 'yield_self',
// Ruby enumerables
'each', 'select', 'reject', 'detect', 'collect',
'inject', 'flat_map', 'each_with_object', 'each_with_index',
'any?', 'all?', 'none?', 'count', 'first', 'last',
'sort_by', 'min_by', 'max_by',
'group_by', 'partition', 'compact', 'flatten', 'uniq',
]);
/** Check if a name is a built-in function or common noise that should be filtered out */
@@ -250,6 +268,12 @@ export const CLASS_CONTAINER_TYPES = new Set([
'class_definition',
'trait_declaration',
'protocol_declaration',
// Ruby
'class',
'module',
// Kotlin
'object_declaration',
'companion_object',
]);
export const CONTAINER_TYPE_TO_LABEL: Record<string, string> = {
@@ -265,6 +289,10 @@ export const CONTAINER_TYPE_TO_LABEL: Record<string, string> = {
trait_declaration: 'Trait',
record_declaration: 'Record',
protocol_declaration: 'Interface',
class: 'Class',
module: 'Module',
object_declaration: 'Class',
companion_object: 'Class',
};
/** Walk up AST to find enclosing class/struct/interface/impl, return its generateId or null.
@@ -307,7 +335,7 @@ export const findEnclosingClassId = (node: any, filePath: string): string | null
}
const nameNode = current.childForFieldName?.('name')
?? current.children?.find((c: any) =>
c.type === 'type_identifier' || c.type === 'identifier' || c.type === 'name'
c.type === 'type_identifier' || c.type === 'identifier' || c.type === 'name' || c.type === 'constant'
);
if (nameNode) {
const label = CONTAINER_TYPE_TO_LABEL[current.type] || 'Class';
@@ -323,7 +351,7 @@ export const findEnclosingClassId = (node: any, filePath: string): string | null
* Extract function name and label from a function_definition or similar AST node.
* Handles C/C++ qualified_identifier (ClassName::MethodName) and other language patterns.
*/
export const extractFunctionName = (node: any): { funcName: string | null; label: string } => {
export const extractFunctionName = (node: SyntaxNode): { funcName: string | null; label: string } => {
let funcName: string | null = null;
let label = 'Function';
@@ -338,21 +366,40 @@ export const extractFunctionName = (node: any): { funcName: string | null; label
if (FUNCTION_DECLARATION_TYPES.has(node.type)) {
// C/C++: function_definition -> [pointer_declarator ->] function_declarator -> qualified_identifier/identifier
// Unwrap pointer_declarator / reference_declarator wrappers to reach function_declarator
let declarator = node.childForFieldName?.('declarator') ||
node.children?.find((c: any) => c.type === 'function_declarator');
let declarator = node.childForFieldName?.('declarator');
if (!declarator) {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'function_declarator') { declarator = c; break; }
}
}
while (declarator && (declarator.type === 'pointer_declarator' || declarator.type === 'reference_declarator')) {
declarator = declarator.childForFieldName?.('declarator') ||
declarator.children?.find((c: any) =>
c.type === 'function_declarator' || c.type === 'pointer_declarator' || c.type === 'reference_declarator');
let nextDeclarator = declarator.childForFieldName?.('declarator');
if (!nextDeclarator) {
for (let i = 0; i < declarator.childCount; i++) {
const c = declarator.child(i);
if (c?.type === 'function_declarator' || c?.type === 'pointer_declarator' || c?.type === 'reference_declarator') { nextDeclarator = c; break; }
}
}
declarator = nextDeclarator;
}
if (declarator) {
const innerDeclarator = declarator.childForFieldName?.('declarator') ||
declarator.children?.find((c: any) =>
c.type === 'qualified_identifier' || c.type === 'identifier' || c.type === 'parenthesized_declarator');
let innerDeclarator = declarator.childForFieldName?.('declarator');
if (!innerDeclarator) {
for (let i = 0; i < declarator.childCount; i++) {
const c = declarator.child(i);
if (c?.type === 'qualified_identifier' || c?.type === 'identifier' || c?.type === 'parenthesized_declarator') { innerDeclarator = c; break; }
}
}
if (innerDeclarator?.type === 'qualified_identifier') {
const nameNode = innerDeclarator.childForFieldName?.('name') ||
innerDeclarator.children?.find((c: any) => c.type === 'identifier');
let nameNode = innerDeclarator.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < innerDeclarator.childCount; i++) {
const c = innerDeclarator.child(i);
if (c?.type === 'identifier') { nameNode = c; break; }
}
}
if (nameNode?.text) {
funcName = nameNode.text;
label = 'Method';
@@ -360,11 +407,19 @@ export const extractFunctionName = (node: any): { funcName: string | null; label
} else if (innerDeclarator?.type === 'identifier') {
funcName = innerDeclarator.text;
} else if (innerDeclarator?.type === 'parenthesized_declarator') {
const nestedId = innerDeclarator.children?.find((c: any) =>
c.type === 'qualified_identifier' || c.type === 'identifier');
let nestedId: SyntaxNode | null = null;
for (let i = 0; i < innerDeclarator.childCount; i++) {
const c = innerDeclarator.child(i);
if (c?.type === 'qualified_identifier' || c?.type === 'identifier') { nestedId = c; break; }
}
if (nestedId?.type === 'qualified_identifier') {
const nameNode = nestedId.childForFieldName?.('name') ||
nestedId.children?.find((c: any) => c.type === 'identifier');
let nameNode = nestedId.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < nestedId.childCount; i++) {
const c = nestedId.child(i);
if (c?.type === 'identifier') { nameNode = c; break; }
}
}
if (nameNode?.text) {
funcName = nameNode.text;
label = 'Method';
@@ -377,35 +432,74 @@ export const extractFunctionName = (node: any): { funcName: string | null; label
// Fallback for other languages (Kotlin uses simple_identifier, Swift uses simple_identifier)
if (!funcName) {
const nameNode = node.childForFieldName?.('name') ||
node.children?.find((c: any) => c.type === 'identifier' || c.type === 'property_identifier' || c.type === 'simple_identifier');
let nameNode = node.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'identifier' || c?.type === 'property_identifier' || c?.type === 'simple_identifier') { nameNode = c; break; }
}
}
funcName = nameNode?.text;
}
} else if (node.type === 'impl_item') {
const funcItem = node.children?.find((c: any) => c.type === 'function_item');
let funcItem: SyntaxNode | null = null;
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'function_item') { funcItem = c; break; }
}
if (funcItem) {
const nameNode = funcItem.childForFieldName?.('name') ||
funcItem.children?.find((c: any) => c.type === 'identifier');
let nameNode = funcItem.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < funcItem.childCount; i++) {
const c = funcItem.child(i);
if (c?.type === 'identifier') { nameNode = c; break; }
}
}
funcName = nameNode?.text;
label = 'Method';
}
} else if (node.type === 'method_definition') {
const nameNode = node.childForFieldName?.('name') ||
node.children?.find((c: any) => c.type === 'property_identifier');
let nameNode = node.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'property_identifier') { nameNode = c; break; }
}
}
funcName = nameNode?.text;
label = 'Method';
} else if (node.type === 'method_declaration' || node.type === 'constructor_declaration') {
const nameNode = node.childForFieldName?.('name') ||
node.children?.find((c: any) => c.type === 'identifier');
let nameNode = node.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'identifier') { nameNode = c; break; }
}
}
funcName = nameNode?.text;
label = 'Method';
} else if (node.type === 'arrow_function' || node.type === 'function_expression') {
const parent = node.parent;
if (parent?.type === 'variable_declarator') {
const nameNode = parent.childForFieldName?.('name') ||
parent.children?.find((c: any) => c.type === 'identifier');
let nameNode = parent.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < parent.childCount; i++) {
const c = parent.child(i);
if (c?.type === 'identifier') { nameNode = c; break; }
}
}
funcName = nameNode?.text;
}
} else if (node.type === 'method' || node.type === 'singleton_method') {
let nameNode = node.childForFieldName?.('name');
if (!nameNode) {
for (let i = 0; i < node.childCount; i++) {
const c = node.child(i);
if (c?.type === 'identifier') { nameNode = c; break; }
}
}
funcName = nameNode?.text;
label = 'Method';
}
return { funcName, label };
@@ -417,6 +511,9 @@ export const extractFunctionName = (node: any): { funcName: string | null; label
*/
export const yieldToEventLoop = (): Promise<void> => new Promise(resolve => setImmediate(resolve));
/** Ruby extensionless filenames recognised as Ruby source */
const RUBY_EXTENSIONLESS_FILES = new Set(['Rakefile', 'Gemfile', 'Guardfile', 'Vagrantfile', 'Brewfile']);
/**
* Find a child of `childType` within a sibling node of `siblingType`.
* Used for Kotlin AST traversal where visibility_modifier lives inside a modifiers sibling.
@@ -469,6 +566,16 @@ export const getLanguageFromFilename = (filename: string): SupportedLanguages |
filename.endsWith('.php5') || filename.endsWith('.php8')) {
return SupportedLanguages.PHP;
}
// Ruby (extensions)
if (filename.endsWith('.rb') || filename.endsWith('.rake') || filename.endsWith('.gemspec')) {
return SupportedLanguages.Ruby;
}
// Ruby (extensionless files)
const basename = filename.split('/').pop() || filename;
if (RUBY_EXTENSIONLESS_FILES.has(basename)) {
return SupportedLanguages.Ruby;
}
// Swift (extensions)
if (filename.endsWith('.swift')) return SupportedLanguages.Swift;
return null;
};
@@ -572,9 +679,19 @@ export const extractMethodSignature = (node: SyntaxNode | null | undefined): Met
// Go: 'result' field is either a type_identifier or parameter_list (multi-return)
const goResult = node.childForFieldName?.('result');
if (goResult) {
returnType = goResult.type === 'parameter_list'
? goResult.text // multi-return: "(string, error)"
: goResult.text; // single return: "int"
if (goResult.type === 'parameter_list') {
// Multi-return: extract first parameter's type only (e.g. (*User, error) → *User)
const firstParam = goResult.firstNamedChild;
if (firstParam?.type === 'parameter_declaration') {
const typeNode = firstParam.childForFieldName('type');
if (typeNode) returnType = typeNode.text;
} else if (firstParam) {
// Unnamed return types: (string, error) — first child is a bare type node
returnType = firstParam.text;
}
} else {
returnType = goResult.text;
}
}
// Rust: 'return_type' field — the value IS the type node (e.g. primitive_type, type_identifier).
@@ -594,6 +711,14 @@ export const extractMethodSignature = (node: SyntaxNode | null | undefined): Met
}
}
// C#: 'returns' field on method_declaration
if (!returnType) {
const csReturn = node.childForFieldName?.('returns');
if (csReturn && csReturn.text !== 'void') {
returnType = csReturn.text;
}
}
// TS/Rust/Python/C#/Kotlin: type_annotation or return_type child
if (!returnType) {
for (const child of node.children) {
@@ -604,6 +729,25 @@ export const extractMethodSignature = (node: SyntaxNode | null | undefined): Met
}
}
// Kotlin: fun getUser(): User — return type is a bare user_type child of
// function_declaration. The Kotlin grammar does NOT wrap it in type_annotation
// or return_type; it appears as a direct child after function_value_parameters.
// Note: Kotlin uses function_value_parameters (not a field), so we find it by type.
if (!returnType) {
let paramsEnd = -1;
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (!child) continue;
if (child.type === 'function_value_parameters' || child.type === 'value_parameters') {
paramsEnd = child.endIndex;
}
if (paramsEnd >= 0 && child.type === 'user_type' && child.startIndex > paramsEnd) {
returnType = child.text;
break;
}
}
}
if (isVariadic) parameterCount = undefined;
return { parameterCount, returnType };
@@ -655,6 +799,7 @@ const MEMBER_ACCESS_NODE_TYPES = new Set([
'field_expression', // Rust/C++: obj.method() / ptr->method()
'selector_expression', // Go: obj.Method()
'navigation_suffix', // Kotlin/Swift: obj.method() — nameNode sits inside navigation_suffix
'member_binding_expression', // C#: user?.Method() — null-conditional access
]);
/**
@@ -715,6 +860,11 @@ export const inferCallForm = (
return 'member';
}
// 4b. Ruby call with receiver: obj.method
if (callNode.type === 'call' && callNode.childForFieldName('receiver')) {
return 'member';
}
// 5. Scoped calls (Rust Foo::new(), C++ ns::func()): treat as free
// The receiver is a type, not an instance — handled differently in Phase 3
if (nameParent && SCOPED_CALL_NODE_TYPES.has(nameParent.type)) {
@@ -741,6 +891,11 @@ const SIMPLE_RECEIVER_TYPES = new Set([
'name', // PHP name node
'this', // TS/JS/Java/C# this.method()
'self', // Rust/Python self.method()
'super', // TS/JS/Java/Kotlin/Ruby super.method()
'super_expression', // Kotlin wraps super in super_expression
'base', // C# base.Method()
'parent', // PHP parent::method()
'constant', // Ruby CONSTANT.method() (uppercase identifiers)
]);
export const extractReceiverName = (
@@ -773,6 +928,30 @@ export const extractReceiverName = (
receiver = callNode.childForFieldName('object');
}
// Ruby: call node has 'receiver' field
if (!receiver && parent.type === 'call') {
receiver = parent.childForFieldName('receiver');
}
// PHP scoped_call_expression (parent::method(), self::method()):
// nameNode's direct parent IS the scoped_call_expression (name is a direct child)
if (!receiver && (parent.type === 'scoped_call_expression' || callNode.type === 'scoped_call_expression')) {
const scopedCall = parent.type === 'scoped_call_expression' ? parent : callNode;
receiver = scopedCall.childForFieldName('scope');
// relative_scope wraps 'parent'/'self'/'static' — unwrap to get the keyword
if (receiver?.type === 'relative_scope') {
receiver = receiver.firstChild;
}
}
// C# null-conditional: user?.Save() → conditional_access_expression wraps member_binding_expression
if (!receiver && parent.type === 'member_binding_expression') {
const condAccess = parent.parent;
if (condAccess?.type === 'conditional_access_expression') {
receiver = condAccess.firstNamedChild;
}
}
// Kotlin/Swift: navigation_expression target is the first child
if (!receiver && parent.type === 'navigation_suffix') {
const navExpr = parent.parent;
@@ -794,9 +973,82 @@ export const extractReceiverName = (
return receiver.text;
}
// Python super().method(): receiver is a call node `super()` — extract the function name
if (receiver.type === 'call') {
const func = receiver.childForFieldName('function');
if (func?.text === 'super') return 'super';
}
return undefined;
};
/**
* Extract the raw receiver AST node for a member call.
* Unlike extractReceiverName, this returns the receiver node regardless of its type —
* including call_expression / method_invocation nodes that appear in chained calls
* like `svc.getUser().save()`.
*
* Returns undefined when the call is not a member call or when no receiver node
* can be found (e.g. top-level free calls).
*/
export const extractReceiverNode = (
nameNode: SyntaxNode,
): SyntaxNode | undefined => {
const parent = nameNode.parent;
if (!parent) return undefined;
const callNode = parent.parent ?? parent;
let receiver: SyntaxNode | null = null;
receiver = parent.childForFieldName('object')
?? parent.childForFieldName('value')
?? parent.childForFieldName('operand')
?? parent.childForFieldName('expression')
?? parent.childForFieldName('argument');
if (!receiver && callNode.type === 'method_invocation') {
receiver = callNode.childForFieldName('object');
}
if (!receiver && (callNode.type === 'member_call_expression' || callNode.type === 'nullsafe_member_call_expression')) {
receiver = callNode.childForFieldName('object');
}
if (!receiver && parent.type === 'call') {
receiver = parent.childForFieldName('receiver');
}
if (!receiver && (parent.type === 'scoped_call_expression' || callNode.type === 'scoped_call_expression')) {
const scopedCall = parent.type === 'scoped_call_expression' ? parent : callNode;
receiver = scopedCall.childForFieldName('scope');
if (receiver?.type === 'relative_scope') {
receiver = receiver.firstChild;
}
}
if (!receiver && parent.type === 'member_binding_expression') {
const condAccess = parent.parent;
if (condAccess?.type === 'conditional_access_expression') {
receiver = condAccess.firstNamedChild;
}
}
if (!receiver && parent.type === 'navigation_suffix') {
const navExpr = parent.parent;
if (navExpr?.type === 'navigation_expression') {
for (const child of navExpr.children) {
if (child.isNamed && child !== parent) {
receiver = child;
break;
}
}
}
}
return receiver ?? undefined;
};
export const isVerboseIngestionEnabled = (): boolean => {
const raw = process.env.GITNEXUS_VERBOSE;
if (!raw) return false;
@@ -804,6 +1056,106 @@ export const isVerboseIngestionEnabled = (): boolean => {
return value === '1' || value === 'true' || value === 'yes';
};
// ── Chained-call extraction ───────────────────────────────────────────────
/** Node types representing call expressions across supported languages. */
export const CALL_EXPRESSION_TYPES = new Set([
'call_expression', // TS/JS/C/C++/Go/Rust
'method_invocation', // Java
'member_call_expression', // PHP
'nullsafe_member_call_expression', // PHP ?.
'call', // Python/Ruby
'invocation_expression', // C#
]);
/**
* Hard limit on chain depth to prevent runaway recursion.
* For `a.b().c().d()`, the chain has depth 2 (b and c before d).
*/
export const MAX_CHAIN_DEPTH = 3;
/**
* Walk a receiver AST node that is itself a call expression, accumulating the
* chain of intermediate method names up to MAX_CHAIN_DEPTH.
*
* For `svc.getUser().save()`, called with the receiver of `save` (getUser() call):
* returns { chain: ['getUser'], baseReceiverName: 'svc' }
*
* For `a.b().c().d()`, called with the receiver of `d` (c() call):
* returns { chain: ['b', 'c'], baseReceiverName: 'a' }
*/
export function extractCallChain(
receiverCallNode: SyntaxNode,
): { chain: string[]; baseReceiverName: string | undefined } | undefined {
const chain: string[] = [];
let current: SyntaxNode = receiverCallNode;
while (CALL_EXPRESSION_TYPES.has(current.type) && chain.length < MAX_CHAIN_DEPTH) {
// Extract the method name from this call node.
const funcNode = current.childForFieldName?.('function')
?? current.childForFieldName?.('name')
?? current.childForFieldName?.('method'); // Ruby `call` node
let methodName: string | undefined;
let innerReceiver: SyntaxNode | null = null;
if (funcNode) {
// member_expression / attribute: last named child is the method identifier
methodName = funcNode.lastNamedChild?.text ?? funcNode.text;
}
// Kotlin/Swift: call_expression exposes callee as firstNamedChild, not a field.
// navigation_expression: method name is in navigation_suffix → simple_identifier.
if (!funcNode && current.type === 'call_expression') {
const callee = current.firstNamedChild;
if (callee?.type === 'navigation_expression') {
const suffix = callee.lastNamedChild;
if (suffix?.type === 'navigation_suffix') {
methodName = suffix.lastNamedChild?.text;
// The receiver is the part of navigation_expression before the suffix
for (let i = 0; i < callee.namedChildCount; i++) {
const child = callee.namedChild(i);
if (child && child.type !== 'navigation_suffix') {
innerReceiver = child;
break;
}
}
}
}
}
if (!methodName) break;
chain.unshift(methodName); // build chain outermost-last
// Walk into the receiver of this call to continue the chain
if (!innerReceiver && funcNode) {
innerReceiver = funcNode.childForFieldName?.('object')
?? funcNode.childForFieldName?.('value')
?? funcNode.childForFieldName?.('operand')
?? funcNode.childForFieldName?.('expression');
}
// Java method_invocation: object field is on the call node
if (!innerReceiver && current.type === 'method_invocation') {
innerReceiver = current.childForFieldName?.('object');
}
// PHP member_call_expression
if (!innerReceiver && (current.type === 'member_call_expression' || current.type === 'nullsafe_member_call_expression')) {
innerReceiver = current.childForFieldName?.('object');
}
// Ruby `call` node: receiver field is on the call node itself
if (!innerReceiver && current.type === 'call') {
innerReceiver = current.childForFieldName?.('receiver');
}
if (!innerReceiver) break;
if (CALL_EXPRESSION_TYPES.has(innerReceiver.type)) {
current = innerReceiver; // continue walking
} else {
// Reached a simple identifier — the base receiver
return { chain, baseReceiverName: innerReceiver.text || undefined };
}
}
return chain.length > 0 ? { chain, baseReceiverName: undefined } : undefined;
}
@@ -9,8 +9,8 @@ import CPP from 'tree-sitter-cpp';
import CSharp from 'tree-sitter-c-sharp';
import Go from 'tree-sitter-go';
import Rust from 'tree-sitter-rust';
import Kotlin from 'tree-sitter-kotlin';
import PHP from 'tree-sitter-php';
import Ruby from 'tree-sitter-ruby';
import { createRequire } from 'node:module';
import { SupportedLanguages } from '../../../config/supported-languages.js';
import { LANGUAGE_QUERIES } from '../tree-sitter-queries.js';
@@ -20,7 +20,11 @@ import { getTreeSitterBufferSize, TREE_SITTER_MAX_BUFFER } from '../constants.js
const _require = createRequire(import.meta.url);
let Swift: any = null;
try { Swift = _require('tree-sitter-swift'); } catch {}
import {
// tree-sitter-kotlin is an optionalDependency — may not be installed
let Kotlin: any = null;
try { Kotlin = _require('tree-sitter-kotlin'); } catch {}
import {
getLanguageFromFilename,
FUNCTION_NODE_TYPES,
extractFunctionName,
@@ -30,14 +34,20 @@ import {
extractMethodSignature,
countCallArguments,
inferCallForm,
extractReceiverName
extractReceiverName,
extractReceiverNode,
CALL_EXPRESSION_TYPES,
extractCallChain,
} from '../utils.js';
import { buildTypeEnv, lookupTypeEnv } from '../type-env.js';
import { buildTypeEnv } from '../type-env.js';
import type { ConstructorBinding } from '../type-env.js';
import { isNodeExported } from '../export-detection.js';
import { detectFrameworkFromAST } from '../framework-detection.js';
import { typeConfigs } from '../type-extractors/index.js';
import { generateId } from '../../../lib/utils.js';
import { extractNamedBindings } from '../named-binding-extraction.js';
import { appendKotlinWildcard } from '../resolvers/index.js';
import { callRouters } from '../call-routing.js';
// ============================================================================
// Types for serializable results
@@ -76,6 +86,7 @@ interface ParsedSymbol {
nodeId: string;
type: string;
parameterCount?: number;
returnType?: string;
ownerId?: string;
}
@@ -99,13 +110,21 @@ export interface ExtractedCall {
receiverName?: string;
/** Resolved type name of the receiver (e.g., 'User' for user.save() when user: User) */
receiverTypeName?: string;
/**
* Chained call names when the receiver is itself a call expression.
* For `svc.getUser().save()`, the `save` ExtractedCall gets receiverCallChain = ['getUser']
* with receiverName = 'svc'. The chain is ordered outermost-last, e.g.:
* `a.b().c().d()` → calledName='d', receiverCallChain=['b','c'], receiverName='a'
* Length is capped at MAX_CHAIN_DEPTH (3).
*/
receiverCallChain?: string[];
}
export interface ExtractedHeritage {
filePath: string;
className: string;
parentName: string;
/** 'extends' | 'implements' | 'trait-impl' */
/** 'extends' | 'implements' | 'trait-impl' | 'include' | 'extend' | 'prepend' */
kind: string;
}
@@ -120,6 +139,12 @@ export interface ExtractedRoute {
lineNumber: number;
}
/** Constructor bindings keyed by filePath for cross-file type resolution */
export interface FileConstructorBindings {
filePath: string;
bindings: ConstructorBinding[];
}
export interface ParseWorkerResult {
nodes: ParsedNode[];
relationships: ParsedRelationship[];
@@ -128,6 +153,8 @@ export interface ParseWorkerResult {
calls: ExtractedCall[];
heritage: ExtractedHeritage[];
routes: ExtractedRoute[];
constructorBindings: FileConstructorBindings[];
skippedLanguages: Record<string, number>;
fileCount: number;
}
@@ -153,11 +180,25 @@ const languageMap: Record<string, any> = {
[SupportedLanguages.CSharp]: CSharp,
[SupportedLanguages.Go]: Go,
[SupportedLanguages.Rust]: Rust,
[SupportedLanguages.Kotlin]: Kotlin,
...(Kotlin ? { [SupportedLanguages.Kotlin]: Kotlin } : {}),
[SupportedLanguages.PHP]: PHP.php_only,
[SupportedLanguages.Ruby]: Ruby,
...(Swift ? { [SupportedLanguages.Swift]: Swift } : {}),
};
/**
* Check if a language grammar is available in this worker.
* Duplicated from parser-loader.ts because workers can't import from the main thread.
* Extra filePath parameter needed to distinguish .tsx from .ts (different grammars
* under the same SupportedLanguages.TypeScript key).
*/
const isLanguageAvailable = (language: SupportedLanguages, filePath: string): boolean => {
const key = language === SupportedLanguages.TypeScript && filePath.endsWith('.tsx')
? `${language}:tsx`
: language;
return key in languageMap && languageMap[key] != null;
};
const setLanguage = (language: SupportedLanguages, filePath: string): void => {
const key = language === SupportedLanguages.TypeScript && filePath.endsWith('.tsx')
? `${language}:tsx`
@@ -238,6 +279,8 @@ const processBatch = (files: ParseWorkerInput[], onProgress?: (filesProcessed: n
calls: [],
heritage: [],
routes: [],
constructorBindings: [],
skippedLanguages: {},
fileCount: 0,
};
@@ -288,21 +331,29 @@ const processBatch = (files: ParseWorkerInput[], onProgress?: (filesProcessed: n
// Process regular files for this language
if (regularFiles.length > 0) {
try {
setLanguage(language, regularFiles[0].path);
processFileGroup(regularFiles, language, queryString, result, onFileProcessed);
} catch {
// parser unavailable — skip this language group
if (isLanguageAvailable(language, regularFiles[0].path)) {
try {
setLanguage(language, regularFiles[0].path);
processFileGroup(regularFiles, language, queryString, result, onFileProcessed);
} catch {
// parser unavailable — skip this language group
}
} else {
result.skippedLanguages[language] = (result.skippedLanguages[language] || 0) + regularFiles.length;
}
}
// Process tsx files separately (different grammar)
if (tsxFiles.length > 0) {
try {
setLanguage(language, tsxFiles[0].path);
processFileGroup(tsxFiles, language, queryString, result, onFileProcessed);
} catch {
// parser unavailable — skip this language group
if (isLanguageAvailable(language, tsxFiles[0].path)) {
try {
setLanguage(language, tsxFiles[0].path);
processFileGroup(tsxFiles, language, queryString, result, onFileProcessed);
} catch {
// parser unavailable — skip this language group
}
} else {
result.skippedLanguages[language] = (result.skippedLanguages[language] || 0) + tsxFiles.length;
}
}
}
@@ -821,8 +872,14 @@ const processFileGroup = (
result.fileCount++;
onFileProcessed?.();
// Build per-file TypeEnv from explicit type annotations (for receiver resolution)
// Build per-file type environment + constructor bindings in a single AST walk.
// Constructor bindings are verified against the SymbolTable in processCallsFromExtracted.
const typeEnv = buildTypeEnv(tree, language);
const callRouter = callRouters[language];
if (typeEnv.constructorBindings.length > 0) {
result.constructorBindings.push({ filePath: file.path, bindings: [...typeEnv.constructorBindings] });
}
let matches;
try {
@@ -858,13 +915,119 @@ const processFileGroup = (
const callNameNode = captureMap['call.name'];
if (callNameNode) {
const calledName = callNameNode.text;
// Dispatch: route language-specific calls (heritage, properties, imports)
const routed = callRouter(calledName, captureMap['call']);
if (routed) {
if (routed.kind === 'skip') continue;
if (routed.kind === 'import') {
result.imports.push({
filePath: file.path,
rawImportPath: routed.importPath,
language,
});
continue;
}
if (routed.kind === 'heritage') {
for (const item of routed.items) {
result.heritage.push({
filePath: file.path,
className: item.enclosingClass,
parentName: item.mixinName,
kind: item.heritageKind,
});
}
continue;
}
if (routed.kind === 'properties') {
const propEnclosingClassId = findEnclosingClassId(captureMap['call'], file.path);
for (const item of routed.items) {
const nodeId = generateId('Property', `${file.path}:${item.propName}`);
result.nodes.push({
id: nodeId,
label: 'Property',
properties: {
name: item.propName,
filePath: file.path,
startLine: item.startLine,
endLine: item.endLine,
language,
isExported: true,
description: item.accessorType,
},
});
result.symbols.push({
filePath: file.path,
name: item.propName,
nodeId,
type: 'Property',
...(propEnclosingClassId ? { ownerId: propEnclosingClassId } : {}),
});
const fileId = generateId('File', file.path);
const relId = generateId('DEFINES', `${fileId}->${nodeId}`);
result.relationships.push({
id: relId,
sourceId: fileId,
targetId: nodeId,
type: 'DEFINES',
confidence: 1.0,
reason: '',
});
if (propEnclosingClassId) {
result.relationships.push({
id: generateId('HAS_METHOD', `${propEnclosingClassId}->${nodeId}`),
sourceId: propEnclosingClassId,
targetId: nodeId,
type: 'HAS_METHOD',
confidence: 1.0,
reason: '',
});
}
}
continue;
}
// kind === 'call' — fall through to normal call processing below
}
if (!isBuiltInOrNoise(calledName)) {
const callNode = captureMap['call'];
const sourceId = findEnclosingFunctionId(callNode, file.path)
|| generateId('File', file.path);
const callForm = inferCallForm(callNode, callNameNode);
const receiverName = callForm === 'member' ? extractReceiverName(callNameNode) : undefined;
const receiverTypeName = receiverName ? lookupTypeEnv(typeEnv, receiverName, callNode) : undefined;
let receiverName = callForm === 'member' ? extractReceiverName(callNameNode) : undefined;
let receiverTypeName = receiverName ? typeEnv.lookup(receiverName, callNode) : undefined;
let receiverCallChain: string[] | undefined;
// When the receiver is a call_expression (e.g. svc.getUser().save()),
// extractReceiverName returns undefined because it refuses complex expressions.
// Instead, walk the receiver node to build a call chain for deferred resolution.
// We capture the base receiver name so processCallsFromExtracted can look it up
// from constructor bindings. receiverTypeName is intentionally left unset here —
// the chain resolver in processCallsFromExtracted needs the base type as input and
// produces the final receiver type as output.
if (callForm === 'member' && receiverName === undefined && !receiverTypeName) {
const receiverNode = extractReceiverNode(callNameNode);
if (receiverNode && CALL_EXPRESSION_TYPES.has(receiverNode.type)) {
const extracted = extractCallChain(receiverNode);
if (extracted) {
receiverCallChain = extracted.chain;
// Set receiverName to the base object so Step 1 in processCallsFromExtracted
// can resolve it via constructor bindings to a base type for the chain.
receiverName = extracted.baseReceiverName;
// Also try the type environment immediately (covers explicitly-typed locals
// and annotated parameters like `fn process(svc: &UserService)`).
// This sets a base type that chain resolution (Step 2) will use as input.
if (receiverName) {
receiverTypeName = typeEnv.lookup(receiverName, callNode);
}
}
}
}
result.calls.push({
filePath: file.path,
calledName,
@@ -873,6 +1036,7 @@ const processFileGroup = (
...(callForm !== undefined ? { callForm } : {}),
...(receiverName !== undefined ? { receiverName } : {}),
...(receiverTypeName !== undefined ? { receiverTypeName } : {}),
...(receiverCallChain !== undefined ? { receiverCallChain } : {}),
});
}
}
@@ -949,6 +1113,14 @@ const processFileGroup = (
const sig = extractMethodSignature(definitionNode);
parameterCount = sig.parameterCount;
returnType = sig.returnType;
// Language-specific return type fallback (e.g. Ruby YARD @return [Type])
if (!returnType && definitionNode) {
const tc = typeConfigs[language as keyof typeof typeConfigs];
if (tc?.extractReturnType) {
returnType = tc.extractReturnType(definitionNode);
}
}
}
result.nodes.push({
@@ -982,6 +1154,7 @@ const processFileGroup = (
nodeId,
type: nodeLabel,
...(parameterCount !== undefined ? { parameterCount } : {}),
...(returnType !== undefined ? { returnType } : {}),
...(enclosingClassId ? { ownerId: enclosingClassId } : {}),
});
@@ -1024,7 +1197,7 @@ const processFileGroup = (
/** Accumulated result across sub-batches */
let accumulated: ParseWorkerResult = {
nodes: [], relationships: [], symbols: [],
imports: [], calls: [], heritage: [], routes: [], fileCount: 0,
imports: [], calls: [], heritage: [], routes: [], constructorBindings: [], skippedLanguages: {}, fileCount: 0,
};
let cumulativeProcessed = 0;
@@ -1036,6 +1209,10 @@ const mergeResult = (target: ParseWorkerResult, src: ParseWorkerResult) => {
target.calls.push(...src.calls);
target.heritage.push(...src.heritage);
target.routes.push(...src.routes);
target.constructorBindings.push(...src.constructorBindings);
for (const [lang, count] of Object.entries(src.skippedLanguages)) {
target.skippedLanguages[lang] = (target.skippedLanguages[lang] || 0) + count;
}
target.fileCount += src.fileCount;
};
@@ -1057,7 +1234,7 @@ parentPort!.on('message', (msg: any) => {
if (msg && msg.type === 'flush') {
parentPort!.postMessage({ type: 'result', data: accumulated });
// Reset for potential reuse
accumulated = { nodes: [], relationships: [], symbols: [], imports: [], calls: [], heritage: [], routes: [], fileCount: 0 };
accumulated = { nodes: [], relationships: [], symbols: [], imports: [], calls: [], heritage: [], routes: [], constructorBindings: [], skippedLanguages: {}, fileCount: 0 };
cumulativeProcessed = 0;
return;
}
@@ -1,5 +1,5 @@
/**
* CSV Generator for KuzuDB Hybrid Schema
* CSV Generator for LadybugDB Hybrid Schema
*
* Streams CSV rows directly to disk files in a single pass over graph nodes.
* File contents are lazy-read from disk per-node to avoid holding the entire
@@ -2,7 +2,7 @@ import fs from 'fs/promises';
import { createReadStream } from 'fs';
import { createInterface } from 'readline';
import path from 'path';
import kuzu from 'kuzu';
import lbug from '@ladybugdb/core';
import { KnowledgeGraph } from '../graph/types.js';
import {
NODE_TABLES,
@@ -13,12 +13,12 @@ import {
} from './schema.js';
import { streamAllCSVsToDisk } from './csv-generator.js';
let db: kuzu.Database | null = null;
let conn: kuzu.Connection | null = null;
let db: lbug.Database | null = null;
let conn: lbug.Connection | null = null;
let currentDbPath: string | null = null;
let ftsLoaded = false;
// Global session lock for operations that touch module-level kuzu globals.
// Global session lock for operations that touch module-level lbug globals.
// This guarantees no DB switch can happen while an operation is running.
let sessionLock: Promise<void> = Promise.resolve();
@@ -39,30 +39,30 @@ const runWithSessionLock = async <T>(operation: () => Promise<T>): Promise<T> =>
const normalizeCopyPath = (filePath: string): string => filePath.replace(/\\/g, '/');
export const initKuzu = async (dbPath: string) => {
return runWithSessionLock(() => ensureKuzuInitialized(dbPath));
export const initLbug = async (dbPath: string) => {
return runWithSessionLock(() => ensureLbugInitialized(dbPath));
};
/**
* Execute multiple queries against one repo DB atomically.
* While the callback runs, no other request can switch the active DB.
*/
export const withKuzuDb = async <T>(dbPath: string, operation: () => Promise<T>): Promise<T> => {
export const withLbugDb = async <T>(dbPath: string, operation: () => Promise<T>): Promise<T> => {
return runWithSessionLock(async () => {
await ensureKuzuInitialized(dbPath);
await ensureLbugInitialized(dbPath);
return operation();
});
};
const ensureKuzuInitialized = async (dbPath: string) => {
const ensureLbugInitialized = async (dbPath: string) => {
if (conn && currentDbPath === dbPath) {
return { db, conn };
}
await doInitKuzu(dbPath);
await doInitLbug(dbPath);
return { db, conn };
};
const doInitKuzu = async (dbPath: string) => {
const doInitLbug = async (dbPath: string) => {
// Different database requested — close the old one first
if (conn || db) {
try { if (conn) await conn.close(); } catch {}
@@ -73,32 +73,36 @@ const doInitKuzu = async (dbPath: string) => {
ftsLoaded = false;
}
// kuzu v0.11 stores the database as a single file (not a directory).
// If the path already exists, it must be a valid kuzu database file.
// LadybugDB stores the database as a single file (not a directory).
// If the path already exists, it must be a valid LadybugDB database file.
// Remove stale empty directories or files from older versions.
try {
const stat = await fs.stat(dbPath);
if (stat.isDirectory()) {
// Old-style directory database or empty leftover - remove it
const files = await fs.readdir(dbPath);
if (files.length === 0) {
await fs.rmdir(dbPath);
} else {
// Non-empty directory from older kuzu version - remove entire directory
await fs.rm(dbPath, { recursive: true, force: true });
const stat = await fs.lstat(dbPath);
if (stat.isSymbolicLink()) {
// Never follow symlinks — just remove the link itself
await fs.unlink(dbPath);
} else if (stat.isDirectory()) {
// Verify path is within expected storage directory before deleting
const realPath = await fs.realpath(dbPath);
const parentDir = path.dirname(dbPath);
const realParent = await fs.realpath(parentDir);
if (!realPath.startsWith(realParent + path.sep) && realPath !== realParent) {
throw new Error(`Refusing to delete ${dbPath}: resolved path ${realPath} is outside storage directory`);
}
// Old-style directory database or empty leftover - remove it
await fs.rm(dbPath, { recursive: true, force: true });
}
// If it's a file, assume it's an existing kuzu database - kuzu will open it
// If it's a file, assume it's an existing LadybugDB database - LadybugDB will open it
} catch {
// Path doesn't exist, which is what kuzu wants for a new database
// Path doesn't exist, which is what LadybugDB wants for a new database
}
// Ensure parent directory exists
const parentDir = path.dirname(dbPath);
await fs.mkdir(parentDir, { recursive: true });
db = new kuzu.Database(dbPath);
conn = new kuzu.Connection(db);
db = new lbug.Database(dbPath);
conn = new lbug.Connection(db);
for (const schemaQuery of SCHEMA_QUERIES) {
try {
@@ -116,16 +120,16 @@ const doInitKuzu = async (dbPath: string) => {
return { db, conn };
};
export type KuzuProgressCallback = (message: string) => void;
export type LbugProgressCallback = (message: string) => void;
export const loadGraphToKuzu = async (
export const loadGraphToLbug = async (
graph: KnowledgeGraph,
repoPath: string,
storagePath: string,
onProgress?: KuzuProgressCallback
onProgress?: LbugProgressCallback
) => {
if (!conn) {
throw new Error('KuzuDB not initialized. Call initKuzu first.');
throw new Error('LadybugDB not initialized. Call initLbug first.');
}
const log = onProgress || (() => {});
@@ -142,7 +146,7 @@ export const loadGraphToKuzu = async (
return nodeId.split(':')[0];
};
// Bulk COPY all node CSVs (sequential — KuzuDB allows only one write txn at a time)
// Bulk COPY all node CSVs (sequential — LadybugDB allows only one write txn at a time)
const nodeFiles = [...csvResult.nodeFiles.entries()];
const totalSteps = nodeFiles.length + 1; // +1 for relationships
let stepsDone = 0;
@@ -167,7 +171,7 @@ export const loadGraphToKuzu = async (
}
}
// Bulk COPY relationships — split by FROM→TO label pair (KuzuDB requires it)
// Bulk COPY relationships — split by FROM→TO label pair (LadybugDB requires it)
// Stream-read the relation CSV line by line to avoid exceeding V8 max string length
let relHeader = '';
const relsByPair = new Map<string, string[]>();
@@ -258,10 +262,10 @@ export const loadGraphToKuzu = async (
return { success: true, insertedRels, skippedRels, warnings };
};
// KuzuDB default ESCAPE is '\' (backslash), but our CSV uses RFC 4180 escaping ("" for literal quotes).
// LadybugDB default ESCAPE is '\' (backslash), but our CSV uses RFC 4180 escaping ("" for literal quotes).
// Source code content is full of backslashes which confuse the auto-detection.
// We MUST explicitly set ESCAPE='"' to use RFC 4180 escaping, and disable auto_detect to prevent
// KuzuDB from overriding our settings based on sample rows.
// LadybugDB from overriding our settings based on sample rows.
const COPY_CSV_OPTS = `(HEADER=true, ESCAPE='"', DELIM=',', QUOTE='"', PARALLEL=false, auto_detect=false)`;
// Multi-language table names that were created with backticks in CODE_ELEMENT_BASE
@@ -300,10 +304,11 @@ const fallbackRelationshipInserts = async (
const confidence = parseFloat(confidenceStr) || 1.0;
const step = parseInt(stepStr) || 0;
const esc = (s: string) => s.replace(/'/g, "''").replace(/\\/g, '\\\\').replace(/\n/g, '\\n').replace(/\r/g, '\\r');
await conn.query(`
MATCH (a:${escapeLabel(fromLabel)} {id: '${fromId.replace(/'/g, "''")}' }),
(b:${escapeLabel(toLabel)} {id: '${toId.replace(/'/g, "''")}' })
CREATE (a)-[:${REL_TABLE_NAME} {type: '${relType}', confidence: ${confidence}, reason: '${reason.replace(/'/g, "''")}', step: ${step}}]->(b)
MATCH (a:${escapeLabel(fromLabel)} {id: '${esc(fromId)}' }),
(b:${escapeLabel(toLabel)} {id: '${esc(toId)}' })
CREATE (a)-[:${REL_TABLE_NAME} {type: '${esc(relType)}', confidence: ${confidence}, reason: '${esc(reason)}', step: ${step}}]->(b)
`);
} catch {
// skip
@@ -340,12 +345,12 @@ const getCopyQuery = (table: NodeTableName, filePath: string): string => {
};
/**
* Insert a single node to KuzuDB
* Insert a single node to LadybugDB
* @param label - Node type (File, Function, Class, etc.)
* @param properties - Node properties
* @param dbPath - Path to KuzuDB database (optional if already initialized)
* @param dbPath - Path to LadybugDB database (optional if already initialized)
*/
export const insertNodeToKuzu = async (
export const insertNodeToLbug = async (
label: string,
properties: Record<string, any>,
dbPath?: string
@@ -353,7 +358,7 @@ export const insertNodeToKuzu = async (
// Use provided dbPath or fall back to module-level db
const targetDbPath = dbPath || (db ? undefined : null);
if (!targetDbPath && !db) {
throw new Error('KuzuDB not initialized. Provide dbPath or call initKuzu first.');
throw new Error('LadybugDB not initialized. Provide dbPath or call initLbug first.');
}
try {
@@ -361,7 +366,7 @@ export const insertNodeToKuzu = async (
if (v === null || v === undefined) return 'NULL';
if (typeof v === 'number') return String(v);
// Escape backslashes first (for Windows paths), then single quotes
return `'${String(v).replace(/\\/g, '\\\\').replace(/'/g, "''")}'`;
return `'${String(v).replace(/\\/g, '\\\\').replace(/'/g, "''").replace(/\n/g, '\\n').replace(/\r/g, '\\r')}'`;
};
// Build INSERT query based on node type
@@ -380,11 +385,11 @@ export const insertNodeToKuzu = async (
const descPart = properties.description ? `, description: ${escapeValue(properties.description)}` : '';
query = `CREATE (n:${t} {id: ${escapeValue(properties.id)}, name: ${escapeValue(properties.name)}, filePath: ${escapeValue(properties.filePath)}, startLine: ${properties.startLine || 0}, endLine: ${properties.endLine || 0}, content: ${escapeValue(properties.content || '')}${descPart}})`;
}
// Use per-query connection if dbPath provided (avoids lock conflicts)
if (targetDbPath) {
const tempDb = new kuzu.Database(targetDbPath);
const tempConn = new kuzu.Connection(tempDb);
const tempDb = new lbug.Database(targetDbPath);
const tempConn = new lbug.Connection(tempDb);
try {
await tempConn.query(query);
return true;
@@ -397,7 +402,7 @@ export const insertNodeToKuzu = async (
await conn.query(query);
return true;
}
return false;
} catch (e: any) {
// Node may already exist or other error
@@ -407,36 +412,36 @@ export const insertNodeToKuzu = async (
};
/**
* Batch insert multiple nodes to KuzuDB using a single connection
* Batch insert multiple nodes to LadybugDB using a single connection
* @param nodes - Array of {label, properties} to insert
* @param dbPath - Path to KuzuDB database
* @param dbPath - Path to LadybugDB database
* @returns Object with success count and error count
*/
export const batchInsertNodesToKuzu = async (
export const batchInsertNodesToLbug = async (
nodes: Array<{ label: string; properties: Record<string, any> }>,
dbPath: string
): Promise<{ inserted: number; failed: number }> => {
if (nodes.length === 0) return { inserted: 0, failed: 0 };
const escapeValue = (v: any): string => {
if (v === null || v === undefined) return 'NULL';
if (typeof v === 'number') return String(v);
// Escape backslashes first (for Windows paths), then single quotes
return `'${String(v).replace(/\\/g, '\\\\').replace(/'/g, "''")}'`;
// Escape backslashes first (for Windows paths), then single quotes, then newlines
return `'${String(v).replace(/\\/g, '\\\\').replace(/'/g, "''").replace(/\n/g, '\\n').replace(/\r/g, '\\r')}'`;
};
// Open a single connection for all inserts
const tempDb = new kuzu.Database(dbPath);
const tempConn = new kuzu.Connection(tempDb);
const tempDb = new lbug.Database(dbPath);
const tempConn = new lbug.Connection(tempDb);
let inserted = 0;
let failed = 0;
try {
for (const { label, properties } of nodes) {
try {
let query: string;
// Use MERGE instead of CREATE for upsert behavior (handles duplicates gracefully)
const t = escapeTableName(label);
if (label === 'File') {
@@ -450,7 +455,7 @@ export const batchInsertNodesToKuzu = async (
const descPart = properties.description ? `, n.description = ${escapeValue(properties.description)}` : '';
query = `MERGE (n:${t} {id: ${escapeValue(properties.id)}}) SET n.name = ${escapeValue(properties.name)}, n.filePath = ${escapeValue(properties.filePath)}, n.startLine = ${properties.startLine || 0}, n.endLine = ${properties.endLine || 0}, n.content = ${escapeValue(properties.content || '')}${descPart}`;
}
await tempConn.query(query);
inserted++;
} catch (e: any) {
@@ -462,17 +467,17 @@ export const batchInsertNodesToKuzu = async (
try { await tempConn.close(); } catch {}
try { await tempDb.close(); } catch {}
}
return { inserted, failed };
};
export const executeQuery = async (cypher: string): Promise<any[]> => {
if (!conn) {
throw new Error('KuzuDB not initialized. Call initKuzu first.');
throw new Error('LadybugDB not initialized. Call initLbug first.');
}
const queryResult = await conn.query(cypher);
// kuzu v0.11 uses getAll() instead of hasNext()/getNext()
// LadybugDB uses getAll() instead of hasNext()/getNext()
// Query returns QueryResult for single queries, QueryResult[] for multi-statement
const result = Array.isArray(queryResult) ? queryResult[0] : queryResult;
const rows = await result.getAll();
@@ -484,7 +489,7 @@ export const executeWithReusedStatement = async (
paramsList: Array<Record<string, any>>
): Promise<void> => {
if (!conn) {
throw new Error('KuzuDB not initialized. Call initKuzu first.');
throw new Error('LadybugDB not initialized. Call initLbug first.');
}
if (paramsList.length === 0) return;
@@ -504,11 +509,11 @@ export const executeWithReusedStatement = async (
// Log the error and continue with next batch
console.warn('Batch execution error:', e);
}
// Note: kuzu 0.8.2 PreparedStatement doesn't require explicit close()
// Note: LadybugDB PreparedStatement doesn't require explicit close()
}
};
export const getKuzuStats = async (): Promise<{ nodes: number; edges: number }> => {
export const getLbugStats = async (): Promise<{ nodes: number; edges: number }> => {
if (!conn) return { nodes: 0, edges: 0 };
let totalNodes = 0;
@@ -541,7 +546,7 @@ export const getKuzuStats = async (): Promise<{ nodes: number; edges: number }>
};
/**
* Load cached embeddings from KuzuDB before a rebuild.
* Load cached embeddings from LadybugDB before a rebuild.
* Returns all embedding vectors so they can be re-inserted after the graph is reloaded,
* avoiding expensive re-embedding of unchanged nodes.
*/
@@ -575,7 +580,7 @@ export const loadCachedEmbeddings = async (): Promise<{
return { embeddingNodeIds, embeddings };
};
export const closeKuzu = async (): Promise<void> => {
export const closeLbug = async (): Promise<void> => {
if (conn) {
try {
await conn.close();
@@ -592,41 +597,41 @@ export const closeKuzu = async (): Promise<void> => {
ftsLoaded = false;
};
export const isKuzuReady = (): boolean => conn !== null && db !== null;
export const isLbugReady = (): boolean => conn !== null && db !== null;
/**
* Delete all nodes (and their relationships) for a specific file from KuzuDB
* Delete all nodes (and their relationships) for a specific file from LadybugDB
* @param filePath - The file path to delete nodes for
* @param dbPath - Optional path to KuzuDB for per-query connection
* @param dbPath - Optional path to LadybugDB for per-query connection
* @returns Object with counts of deleted nodes
*/
export const deleteNodesForFile = async (filePath: string, dbPath?: string): Promise<{ deletedNodes: number }> => {
const usePerQuery = !!dbPath;
// Set up connection (either use existing or create per-query)
let tempDb: kuzu.Database | null = null;
let tempConn: kuzu.Connection | null = null;
let targetConn: kuzu.Connection | null = conn;
let tempDb: lbug.Database | null = null;
let tempConn: lbug.Connection | null = null;
let targetConn: lbug.Connection | null = conn;
if (usePerQuery) {
tempDb = new kuzu.Database(dbPath);
tempConn = new kuzu.Connection(tempDb);
tempDb = new lbug.Database(dbPath);
tempConn = new lbug.Connection(tempDb);
targetConn = tempConn;
} else if (!conn) {
throw new Error('KuzuDB not initialized. Provide dbPath or call initKuzu first.');
throw new Error('LadybugDB not initialized. Provide dbPath or call initLbug first.');
}
try {
let deletedNodes = 0;
const escapedPath = filePath.replace(/'/g, "''");
// Delete nodes from each table that has filePath
// DETACH DELETE removes the node and all its relationships
for (const tableName of NODE_TABLES) {
// Skip tables that don't have filePath (Community, Process)
if (tableName === 'Community' || tableName === 'Process') continue;
try {
// First count how many we'll delete
const tn = escapeTableName(tableName);
@@ -648,7 +653,7 @@ export const deleteNodesForFile = async (filePath: string, dbPath?: string): Pro
// Some tables may not support this query, skip
}
}
// Also delete any embeddings for nodes in this file
try {
await targetConn!.query(
@@ -657,7 +662,7 @@ export const deleteNodesForFile = async (filePath: string, dbPath?: string): Pro
} catch {
// Embedding table may not exist or nodeId format may differ
}
return { deletedNodes };
} finally {
// Close per-query connection if used
@@ -683,7 +688,7 @@ export const getEmbeddingTableName = (): string => EMBEDDING_TABLE_NAME;
export const loadFTSExtension = async (): Promise<void> => {
if (ftsLoaded) return;
if (!conn) {
throw new Error('KuzuDB not initialized. Call initKuzu first.');
throw new Error('LadybugDB not initialized. Call initLbug first.');
}
try {
await conn.query('INSTALL fts');
@@ -713,7 +718,7 @@ export const createFTSIndex = async (
stemmer: string = 'porter'
): Promise<void> => {
if (!conn) {
throw new Error('KuzuDB not initialized. Call initKuzu first.');
throw new Error('LadybugDB not initialized. Call initLbug first.');
}
await loadFTSExtension();
@@ -747,24 +752,24 @@ export const queryFTS = async (
conjunctive: boolean = false
): Promise<Array<{ nodeId: string; name: string; filePath: string; score: number; [key: string]: any }>> => {
if (!conn) {
throw new Error('KuzuDB not initialized. Call initKuzu first.');
throw new Error('LadybugDB not initialized. Call initLbug first.');
}
// Escape backslashes and single quotes to prevent Cypher injection
const escapedQuery = query.replace(/\\/g, '\\\\').replace(/'/g, "''");
const cypher = `
CALL QUERY_FTS_INDEX('${tableName}', '${indexName}', '${escapedQuery}', conjunctive := ${conjunctive})
RETURN node, score
ORDER BY score DESC
LIMIT ${limit}
`;
try {
const queryResult = await conn.query(cypher);
const result = Array.isArray(queryResult) ? queryResult[0] : queryResult;
const rows = await result.getAll();
return rows.map((row: any) => {
const node = row.node || row[0] || {};
const score = row.score ?? row[1] ?? 0;
@@ -790,9 +795,9 @@ export const queryFTS = async (
*/
export const dropFTSIndex = async (tableName: string, indexName: string): Promise<void> => {
if (!conn) {
throw new Error('KuzuDB not initialized. Call initKuzu first.');
throw new Error('LadybugDB not initialized. Call initLbug first.');
}
try {
await conn.query(`CALL DROP_FTS_INDEX('${tableName}', '${indexName}')`);
} catch {
@@ -1,5 +1,5 @@
/**
* KuzuDB Schema Definitions
* LadybugDB Schema Definitions
*
* Hybrid Schema:
* - Separate node tables for each code element type (File, Function, Class, etc.)
+13 -13
View File
@@ -1,11 +1,11 @@
/**
* Full-Text Search via KuzuDB FTS
*
* Uses KuzuDB's built-in full-text search indexes for keyword-based search.
* Full-Text Search via LadybugDB FTS
*
* Uses LadybugDB's built-in full-text search indexes for keyword-based search.
* Always reads from the database (no cached state to drift).
*/
import { queryFTS } from '../kuzu/kuzu-adapter.js';
import { queryFTS } from '../lbug/lbug-adapter.js';
export interface BM25SearchResult {
filePath: string;
@@ -15,7 +15,7 @@ export interface BM25SearchResult {
/**
* Execute a single FTS query via a custom executor (for MCP connection pool).
* Returns the same shape as core queryFTS.
* Returns the same shape as core queryFTS (from LadybugDB adapter).
*/
async function queryFTSViaExecutor(
executor: (cypher: string) => Promise<any[]>,
@@ -48,24 +48,24 @@ async function queryFTSViaExecutor(
}
/**
* Search using KuzuDB's built-in FTS (always fresh, reads from disk)
*
* Search using LadybugDB's built-in FTS (always fresh, reads from disk)
*
* Queries multiple node tables (File, Function, Class, Method) in parallel
* and merges results by filePath, summing scores for the same file.
*
*
* @param query - Search query string
* @param limit - Maximum results
* @param repoId - If provided, queries will be routed via the MCP connection pool
* @returns Ranked search results from FTS indexes
*/
export const searchFTSFromKuzu = async (query: string, limit: number = 20, repoId?: string): Promise<BM25SearchResult[]> => {
export const searchFTSFromLbug = async (query: string, limit: number = 20, repoId?: string): Promise<BM25SearchResult[]> => {
let fileResults: any[], functionResults: any[], classResults: any[], methodResults: any[], interfaceResults: any[];
if (repoId) {
// Use MCP connection pool via dynamic import
// IMPORTANT: KuzuDB uses a single connection per repo — queries must be sequential
// to avoid deadlocking. Do NOT use Promise.all here.
const { executeQuery } = await import('../../mcp/core/kuzu-adapter.js');
// IMPORTANT: FTS queries run sequentially to avoid connection contention.
// The MCP pool supports multiple connections, but FTS is best run serially.
const { executeQuery } = await import('../../mcp/core/lbug-adapter.js');
const executor = (cypher: string) => executeQuery(repoId, cypher);
fileResults = await queryFTSViaExecutor(executor, 'File', 'file_fts', query, limit);
functionResults = await queryFTSViaExecutor(executor, 'Function', 'function_fts', query, limit);
@@ -73,7 +73,7 @@ export const searchFTSFromKuzu = async (query: string, limit: number = 20, repoI
methodResults = await queryFTSViaExecutor(executor, 'Method', 'method_fts', query, limit);
interfaceResults = await queryFTSViaExecutor(executor, 'Interface', 'interface_fts', query, limit);
} else {
// Use core kuzu adapter (CLI / pipeline context) — also sequential for safety
// Use core lbug adapter (CLI / pipeline context) — also sequential for safety
fileResults = await queryFTS('File', 'file_fts', query, limit, false).catch(() => []);
functionResults = await queryFTS('Function', 'function_fts', query, limit, false).catch(() => []);
classResults = await queryFTS('Class', 'class_fts', query, limit, false).catch(() => []);
+6 -6
View File
@@ -8,7 +8,7 @@
* production search systems.
*/
import { searchFTSFromKuzu, type BM25SearchResult } from './bm25-index.js';
import { searchFTSFromLbug, type BM25SearchResult } from './bm25-index.js';
import type { SemanticSearchResult } from '../embeddings/types.js';
/**
@@ -114,11 +114,11 @@ export const mergeWithRRF = (
/**
* Check if hybrid search is available
* KuzuDB FTS is always available once the database is initialized.
* LadybugDB FTS is always available once the database is initialized.
* Semantic search is optional - hybrid works with just FTS if embeddings aren't ready.
*/
export const isHybridSearchReady = (): boolean => {
return true; // FTS is always available via KuzuDB when DB is open
return true; // FTS is always available via LadybugDB when DB is open
};
/**
@@ -146,7 +146,7 @@ export const formatHybridResults = (results: HybridSearchResult[]): string => {
/**
* Execute BM25 + semantic search and merge with RRF.
* Uses KuzuDB FTS for always-fresh BM25 results (no cached data).
* Uses LadybugDB FTS for always-fresh BM25 results (no cached data).
* The semanticSearch function is injected to keep this module environment-agnostic.
*/
export const hybridSearch = async (
@@ -155,8 +155,8 @@ export const hybridSearch = async (
executeQuery: (cypher: string) => Promise<any[]>,
semanticSearch: (executeQuery: (cypher: string) => Promise<any[]>, query: string, k?: number) => Promise<SemanticSearchResult[]>
): Promise<HybridSearchResult[]> => {
// Use KuzuDB FTS for always-fresh BM25 results
const bm25Results = await searchFTSFromKuzu(query, limit);
// Use LadybugDB FTS for always-fresh BM25 results
const bm25Results = await searchFTSFromLbug(query, limit);
const semanticResults = await semanticSearch(executeQuery, query, limit);
return mergeWithRRF(bm25Results, semanticResults, limit);
};
@@ -8,8 +8,8 @@ import CPP from 'tree-sitter-cpp';
import CSharp from 'tree-sitter-c-sharp';
import Go from 'tree-sitter-go';
import Rust from 'tree-sitter-rust';
import Kotlin from 'tree-sitter-kotlin';
import PHP from 'tree-sitter-php';
import Ruby from 'tree-sitter-ruby';
import { createRequire } from 'node:module';
import { SupportedLanguages } from '../../config/supported-languages.js';
@@ -18,6 +18,10 @@ const _require = createRequire(import.meta.url);
let Swift: any = null;
try { Swift = _require('tree-sitter-swift'); } catch {}
// tree-sitter-kotlin is an optionalDependency — may not be installed
let Kotlin: any = null;
try { Kotlin = _require('tree-sitter-kotlin'); } catch {}
let parser: Parser | null = null;
const languageMap: Record<string, any> = {
@@ -31,8 +35,9 @@ const languageMap: Record<string, any> = {
[SupportedLanguages.CSharp]: CSharp,
[SupportedLanguages.Go]: Go,
[SupportedLanguages.Rust]: Rust,
[SupportedLanguages.Kotlin]: Kotlin,
...(Kotlin ? { [SupportedLanguages.Kotlin]: Kotlin } : {}),
[SupportedLanguages.PHP]: PHP.php_only,
[SupportedLanguages.Ruby]: Ruby,
...(Swift ? { [SupportedLanguages.Swift]: Swift } : {}),
};
+4 -4
View File
@@ -93,7 +93,7 @@ export class WikiGenerator {
private repoPath: string;
private storagePath: string;
private wikiDir: string;
private kuzuPath: string;
private lbugPath: string;
private llmConfig: LLMConfig;
private maxTokensPerModule: number;
private concurrency: number;
@@ -104,7 +104,7 @@ export class WikiGenerator {
constructor(
repoPath: string,
storagePath: string,
kuzuPath: string,
lbugPath: string,
llmConfig: LLMConfig,
options: WikiOptions = {},
onProgress?: ProgressCallback,
@@ -112,7 +112,7 @@ export class WikiGenerator {
this.repoPath = repoPath;
this.storagePath = storagePath;
this.wikiDir = path.join(storagePath, WIKI_DIR);
this.kuzuPath = kuzuPath;
this.lbugPath = lbugPath;
this.options = options;
this.llmConfig = llmConfig;
this.maxTokensPerModule = options.maxTokensPerModule ?? DEFAULT_MAX_TOKENS_PER_MODULE;
@@ -171,7 +171,7 @@ export class WikiGenerator {
// Init graph
this.onProgress('init', 2, 'Connecting to knowledge graph...');
await initWikiDb(this.kuzuPath);
await initWikiDb(this.lbugPath);
let result: { pagesGenerated: number; mode: 'full' | 'incremental' | 'up-to-date'; failedModules: string[] };
try {
+8 -8
View File
@@ -1,11 +1,11 @@
/**
* Graph Queries for Wiki Generation
*
*
* Encapsulated Cypher queries against the GitNexus knowledge graph.
* Uses the MCP-style pooled kuzu-adapter for connection management.
* Uses the MCP-style pooled lbug-adapter for connection management.
*/
import { initKuzu, executeQuery, closeKuzu } from '../../mcp/core/kuzu-adapter.js';
import { initLbug, executeQuery, closeLbug } from '../../mcp/core/lbug-adapter.js';
const REPO_ID = '__wiki__';
@@ -35,17 +35,17 @@ export interface ProcessInfo {
}
/**
* Initialize the KuzuDB connection for wiki generation.
* Initialize the LadybugDB connection for wiki generation.
*/
export async function initWikiDb(kuzuPath: string): Promise<void> {
await initKuzu(REPO_ID, kuzuPath);
export async function initWikiDb(lbugPath: string): Promise<void> {
await initLbug(REPO_ID, lbugPath);
}
/**
* Close the KuzuDB connection.
* Close the LadybugDB connection.
*/
export async function closeWikiDb(): Promise<void> {
await closeKuzu(REPO_ID);
await closeLbug(REPO_ID);
}
/**

Some files were not shown because too many files have changed in this diff Show More