Compare commits

...
Author SHA1 Message Date
Gergo Magyar 1003d8b6a5 test: add coverage for perf optimizations — fastStripNullable, skipGraphPhases, AST pruning
- 6 new unit tests for fastStripNullable branches (simple id, nullable union, bare keyword)
- 4 new integration tests for skipGraphPhases pipeline option
- Tests for SKIP_SUBTREE_TYPES and interestingNodeTypes code paths
2026-03-17 17:31:24 +00:00
Gergo Magyar 74b9701509 chore: bump version to 1.4.5, add CHANGELOG.md 2026-03-17 17:18:35 +00:00
Gergő Magyar f0132c1077 feat: Phase 6 type resolution — for-loop Tier 1c, pattern matching, container descriptors, 10-language coverage (#318)
* feat: Phase 6 type resolution — pattern matching, for-loop Tier 1c, coverage completion

- Add patternBindingNodeTypes gate to LanguageTypeConfig for 50% perf improvement
- Expand ForLoopExtractor signature with optional declarationTypeNodes + scope
- Add extractElementTypeFromString shared utility for container type parsing
- Python match/case: extractPatternBinding for `case User() as u:` pattern
- C# refactor: move is_pattern_expression from extractDeclaration to extractPatternBinding
- Ruby: add extractPendingAssignment for assignment chain propagation
- TS/JS: add for-loop Tier 1c for `for (const user of users)` with User[] inference
- Python: add for-loop Tier 1c for `for user in users:` with type annotation inference
- Go: add for-loop Tier 1c for `for _, user := range users` with []User inference
- Fix 'Property' as any stale cast in call-processor.ts
- Add dual return-type string length cap (2048 pre-cap, 512 post-cap)
- Add chain call integration tests for C#, Go, Rust, Python, JS, C++
- Add Python match/case integration test fixtures
- 27 new extractElementTypeFromString unit tests
- 3 for-loop edge cases skipped (declarationTypeNodes scope key lookup)

* fix: address code review findings for Phase 6

- Add missing patternBindingNodeTypes to C# typeConfig (perf gate)
- Add 2048-char input length guard to extractElementTypeFromString
- Skip Python match/case integration tests (call extraction needs query updates)

* reorganise

* fix: Phase 1 bug fixes — Go range semantics, typed_parameter, bracket depth

- Go single-var range correctly returns early for slices/maps (index, not element)
- Go single-var range on channels correctly resolves element type
- Added map_type and channel_type to extractGoElementTypeFromTypeNode
- Added isChannelType helper for channel detection before skip decision
- Added 'typed_parameter' to TYPED_PARAMETER_TYPES for Python annotated params
- Fixed bracket depth tracking in extractElementTypeFromString — only match
  selected closeChar at depth 0, return undefined for mismatched brackets
- Un-skipped 3 prematurely skipped tests (TS local const, Python List/Sequence)
- Added tests for map range, single-var range semantics, bracket edge cases

* refactor: Phase 2 architecture — shared helper, required params, decoupled type nodes

- Extract resolveIterableElementType shared helper in shared.ts implementing
  3-strategy fallback (declarationTypeNodes → scopeEnv string → AST walk)
- Refactor TS, Python, Go extractors to use shared helper (eliminates 3x duplication)
- Make ForLoopExtractor params required (aligned with PatternBindingExtractor)
- Update Java, Kotlin, C# extractor signatures to accept required params
- Decouple declarationTypeNodes from scopeEnv — capture raw type annotation
  nodes BEFORE extractDeclaration for container types (User[], []User, List[User])
- Hybrid approach: direct name extraction + keysBefore fallback for multi-declarator
- Document declarationTypeNodes invariant change (superset of scopeEnv)

* feat: Phase 3 partial — Rust for-loop + C# var foreach Tier 1c

- Rust: add extractForLoopBinding with for_expression support
  - Handles &users, &mut users via reference_expression unwrapping
  - extractRustElementTypeFromTypeNode: generic_type, reference_type, slice/array
  - findRustParamElementType: AST walk with reference/mut pattern unwrapping
  - 4 unit tests (Vec<User>, &[User], range expr negative, no-annotation negative)

- C#: upgrade foreach to handle var (implicit_type) via Tier 1c
  - extractCSharpElementTypeFromTypeNode: generic_name, array_type, nullable_type
  - findCSharpParamElementType: AST walk to method_declaration parameters
  - 3 unit tests (var foreach, explicit type regression, no-annotation negative)

* feat: Phase 3 complete — all language gaps + pattern matching

Kotlin Tier 1c:
- Unannotated for-loop resolves via shared helper
- extractKotlinElementTypeFromTypeNode handles type_projection unwrapping
- findKotlinParamElementType walks to function_declaration

Java Tier 1c:
- var foreach resolves via shared helper
- extractJavaElementTypeFromTypeNode handles generic_type, array_type
- findJavaParamElementType walks to method_declaration

TypeScript:
- readonly User[] unwrapped via readonly_type → array_type recursion

C# switch patterns:
- declaration_pattern added to patternBindingNodeTypes
- extractPatternBinding handles standalone declaration_pattern (switch case/expr)

Rust match arms:
- match_arm added to patternBindingNodeTypes
- extractPatternBinding extended with match_arm → match_expression parent traversal

Python:
- as_pattern tries childForFieldName('alias') before positional fallback

Tests: 237 pass (was 224), 13 new tests added

* feat: Phase 4 — known limitation tests, match arm fix, final verification

- Fix Rust match_arm pattern extraction: unwrap match_pattern to get
  tuple_struct_pattern inside (tree-sitter-rust wraps in match_pattern node)
- Add first-writer-wins regression test for match arm scope leakage
- Add 5 documented skip tests for known limitations:
  - TS destructured for-of (tuple destructuring)
  - Python tuple unpacking in for-loops
  - TS instanceof narrowing (block-level scoping)
  - Rust for with .iter() (method call iterable)
  - Ruby block parameters (closure param inference)

Final: 238 passed, 5 skipped (documented limitations), tsc clean

* test: integration tests for all Phase 6 language gaps + fix Rust param pattern field

Integration test fixtures and tests (30 new tests, all with exact match + negative):

Rust for-loop (5 tests):
- for user in &users with Vec<User> → User#save, negative Repo#save
- for repo in &repos with Vec<Repo> → Repo#save, negative User#save

Rust match arm (5 tests):
- match opt { Some(user) => user.save() } → User#save, negative Repo#save
- if let Ok(repo) = res → Repo#save, negative User#save

C# var foreach (5 tests):
- foreach (var user in users) with List<User> → User#Save, negative Repo#Save
- foreach (var repo in repos) with List<Repo> → Repo#Save

C# switch pattern (4 tests):
- is User user → User#Save, case Repo repo → Repo#Save

Kotlin unannotated for (4 tests):
- for (user in users) with List<User> → user.save, negative repo.save

Go map range (3 tests):
- for _, user := range userMap with map[string]User → User#Save, negative

TypeScript readonly (4 tests):
- for (const user of users) with readonly User[] → user.save, negative

Bug fix: type-env.ts parameter branch now falls back to childForFieldName('pattern')
for Rust parameters (Rust uses 'pattern' not 'name' for parameter names)

* test: add assertion bodies to known limitation skip tests

Convert empty skip test stubs to proper tests with parse/buildTypeEnv/expect
assertions following the codebase convention (e.g., call-processor.test.ts:319).
Each skip test now documents the exact expected behavior, so removing .skip
will cause a meaningful failure when the limitation is eventually fixed.

Also clarify Python integration skip tests as call-extraction issues (not
type-env) and Swift integration skips as build-dep issues (self/super
resolution code already exists in type-env.ts).

* feat: resolve 4 known limitation skip tests + method-aware type arg selection

Unskip 4 of 5 type-env known limitations with full integration test coverage:

1. TS destructured for-of: handle array_pattern by binding last named child
   to element type. Fix Map<K,V> to return last generic arg (value type).
2. Python dict.items() loop: handle `call` iterables + `pattern_list` left
   side. Fix dict[K,V] extraction via type_parameter with last-arg heuristic.
   Unwrap `type` wrapper in extractPyElementTypeFromAnnotation.
3. TS instanceof narrowing: add extractPatternBinding for binary_expression
   with positional child access. First-writer-wins (not block-scoped).
4. Rust .iter() for-loops: handle call_expression in for_expression value
   node by extracting receiver from field_expression.

Method-aware type arg resolution:
- Add TypeArgPosition ('first'|'last') to resolveIterableElementType
- .keys()/.keySet()/.Keys → first type arg (key); all else → last (value)
- Thread position through all 3 strategy callbacks in TS/Rust/Python
- Add predefined_type to extractSimpleTypeName for TS primitives (string etc)

New fixtures: rust-iter-for-loop, typescript-destructured-for-of,
typescript-instanceof-narrowing, python-dict-items-loop.
248 unit tests pass (6 new), 1 skip (Ruby block params).

* feat: container descriptor table for generic type arg resolution

Replace simple KEY_METHODS heuristic with CONTAINER_DESCRIPTORS table
that maps 30+ container types across all languages to their type parameter
semantics per access method.

Key improvements:
- Container-aware resolution: HashMap.iter() correctly yields V (arity 2),
  while Vec.iter() yields T (arity 1) — same method, different semantics
- Cross-language coverage: Map/HashMap/BTreeMap/dict/Dict/Dictionary/
  ConcurrentHashMap + List/Vec/Set/HashSet/Queue/Deque/Stack etc.
- Method categorization: keyMethods (keys/keySet/Keys) vs valueMethods
  (values/get/pop/iter/first/last) per container type
- Fallback for unknown containers: still uses method name heuristic,
  so MyCache<K,V>.keys() correctly returns first arg
- Exported getContainerDescriptor() for future heritage-chain lookups

Each language extractor now passes containerTypeName from scopeEnv to
methodToTypeArgPosition for descriptor-aware resolution.

252 unit tests pass (4 new descriptor tests), 1 skip (Ruby).

* feat: method-aware for-loop extractors + integration tests for all languages

Upgrade 4 existing extractors + create 3 new ones for full cross-language
coverage of call_expression iterables and container descriptor resolution:

Upgraded (add call expr iterable + methodToTypeArgPosition):
- Java: method_invocation (data.keySet(), data.values())
- Kotlin: navigation_expression + call_expression (data.keys, data.values())
- C#: member_access_expression + invocation_expression (data.Keys, data.Values)
- Go: TypeArgPosition threading for Go 1.18+ generics

New for-loop extractors:
- C++: for_range_loop with auto& unwrapping, template_type + qualified_identifier
  (std::vector<User>) extraction, explicit vs auto type handling
- PHP: foreach_statement with simple/key-value/by-reference forms, PHPDoc
  @param priority over AST array type
- Ruby: for-in with YARD @param type resolution via comment parsing

Integration test fixtures + tests for all 6 languages:
- java-map-keys-values (Map.values() + List iteration)
- kotlin-map-keys-values (HashMap.values + List iteration)
- csharp-dictionary-keys-values (Dictionary.Values foreach)
- cpp-range-for (auto& + const auto& range-based for)
- php-foreach-loop (foreach with PHPDoc @param User[])
- ruby-for-in-loop (for-in with YARD @param Array<User>)

Bugs fixed during integration testing:
- C++: qualified_identifier (std::vector) not unwrapped to template_type
- PHP: extractParameter overwrote PHPDoc-derived types with bare 'array'

252 unit tests pass, 201 integration tests pass across 6 languages.

* fix: update extractElementTypeFromString tests for last-arg default

TypeArgPosition change (default 'last') broke 5 existing tests expecting
first arg from multi-arg generics. Updated expectations and added explicit
pos='first' tests for key type extraction.

* fix: rename C++ fixture files to correct case for case-sensitive CI

On case-sensitive filesystems (Linux/macOS CI), git tracked both the old
lowercase files (app.cpp, user.h) and the new uppercase files (App.cpp,
User.h) as separate files. The pipeline processed both, causing the old
app.cpp (with explicit User& type) to interfere with the new auto& test.

Removes old lowercase entries and re-adds with uppercase casing to match
the #include directives in the fixture.

* feat: PR #318 review findings — pattern bindings, member access iterables, structured bindings

Address all 7 genuine gaps identified in PR #318 deep code review:

- Kotlin: add extractKotlinPatternBinding for when/is (type_test AST node)
  with allowPatternBindingOverwrite for smart-cast semantics
- Java: add type_pattern branch for Java 17+ switch pattern variables
- TypeScript: explicit object_pattern skip in for-of (no false bindings)
- Cross-language: member access iterables (self.users, this.users, repo.users)
  across all 10 language extractors
- C++: structured_binding_declarator handling in range-for (last-child heuristic)
- Rust: closure_parameter added to TYPED_PARAMETER_TYPES
- PHP: normalizePhpType handles angle-bracket generics (Collection<User>)

Code review fixes applied:
- Remove 4 debug console.log statements (c-cpp.ts, call-processor.ts)
- Hoist KNOWN_CONTAINER_PROPS to module scope (csharp.ts)
- Guard keysBefore allocation behind typeNode check (type-env.ts)
- Add depth limits (50) to 7 recursive type extraction functions
- Add 2048-char length cap to extractSimpleTypeName
- Fix PHP/Ruby missing typeArgPos parameter in resolveIterableElementType

Integration test fixtures: kotlin-when-pattern, java-switch-pattern,
cpp-structured-binding, typescript-member-access-for-loop,
python-member-access-for-loop

* fix: position-indexed when/is bindings, Kotlin param extraction, HashMap.values for-loop

Three root causes for failing Kotlin integration tests:

1. When/is multi-arm resolution: flat scopeEnv stored only the last arm's
   type (last-writer-wins). Added PatternOverrides with AST range indexing
   so each when arm resolves to its narrowed type independently.

2. HashMap.values for-loop: navigation_expression without call_suffix was
   classified as bare property access (iterableName='values' instead of
   'data'). Now tries object-as-iterable + property-as-method first, with
   fallback to property-as-iterable for this.users patterns.

3. Kotlin parameter extraction: tree-sitter-kotlin parameter nodes use
   positional children (simple_identifier, user_type) not named fields
   (name, type). Added fallback to findChildByType in both
   extractKotlinParameter and extractTypeBinding.

Integration tests added for .keys/.values/Set/MutableMap iteration,
3-arm when/is, multi-call within arms, and when+else branch.

* feat: enhance PHP type resolution for generics and member access in foreach loops

* feat: Phase 6.1 type resolution gap closure — container descriptors, recursive_pattern, class fields

Add 13 missing container type descriptors (Collection, MutableMap, Stream, SortedSet, etc.)
to CONTAINER_DESCRIPTORS for correct element type extraction across C#, Kotlin, and Java.

Extend C# pattern binding to handle recursive_pattern (obj is User { Name: "Alice" } u)
in both is-expression and switch expression contexts.

Add TypeScript class field declaration support (public_field_definition) so for-loop
iteration over this.fieldName resolves element types from class field type annotations.
Includes file-scope fallback in resolveIterableElementType and nested member_expression
handling for this.field.method() patterns.

* docs: add type resolution system documentation with roadmap

Covers the full architecture, resolution tiers (0-2), scope model,
language feature matrix, container descriptors, pipeline integration,
and the Phase 7-9 roadmap for cross-scope propagation, field-type
resolution, and return-type-aware binding.

* feat: Phase 6.2 review findings — C# nested member foreach, C++ deref range-for, Java field_access

Close two gaps found during fourth-pass review of PR #318:

- C# foreach (var user in this.data.Values): nested member_access_expression
  now extracts intermediate property name for scopeEnv lookup
- C++ for (auto& user : *ptr): pointer_expression dereference now recognized
  as range-for iterable

Root causes fixed in shared infrastructure:
- extractSimpleTypeName: add template_type (C++) and generic_name (C#)
- extractGenericTypeArgs: add generic_name for consistency
- type-env.ts: unwrap variable_declaration wrapper in field_declaration
  for declarationTypeNodes capture (zero-allocation manual loop)

Additional review findings addressed:
- Java: add field_access handler for this.data.values() in method_invocation
- C++ pointer_expression: document limitation (*identifier only)
- TypeScript: fix stale comment about property_identifier

All 525 tests pass (278 unit + 247 integration).

* perf: optimize type resolution pipeline — worker threshold, skip graph phases, AST pruning

- Skip worker pool creation for small repos (<15 files or <512KB) — saves 100-400ms
- Add skipGraphPhases option to runPipelineFromRepo to skip MRO/community/process phases
- Add conservative SKIP_SUBTREE_TYPES for leaf-only AST nodes (string, comment, number)
- Pre-compute interestingNodeTypes set — single Set.has() replaces 3 checks per node
- Add fastStripNullable — skip full stripNullable for simple identifiers (90%+ case)
- Replace .children?.find() with manual for loops in extractFunctionName (no array alloc)
- Add hookTimeout: 120000 to vitest.config.ts for CI beforeAll hooks

* fix: review findings — remove template_string from SKIP_SUBTREE_TYPES, handle bare nullable keywords

- Remove template_string and concatenated_string from SKIP_SUBTREE_TYPES
  (template literals contain interpolated expressions with typed code)
- Add FAST_NULLABLE_KEYWORDS check to fastStripNullable for behavioral
  parity with stripNullable on bare null/undefined/void/None/nil
- Add explanatory comment on extractPendingAssignment scopeEnv guard

* feat: add type resolution system and roadmap documentation
2026-03-17 17:10:22 +00:00
Chirag Nighutandchirag-nighut f6b92d4f13 fix(resolver): fix for same-directory python imports (#328)
* fix(resolver): prefer same-directory file for Python bare imports

Python's sys.path searches the importing script's own directory first,
so `import user` from services/auth.py should resolve to services/user.py
even if models/user.py was indexed first in the suffix index.

Add a proximity check in resolveImportPath that consults the existing
dirMap index (O(1)) before falling back to global suffix matching, for
single-segment bare Python imports only.

Made-with: Cursor

* refactor(resolver): replace dirMap scan with O(1) allFiles.has() for proximity check

The previous implementation used index.getFilesInDir() + siblings.find()
which had two issues:
- dirMap stores all suffix levels, so getFilesInDir('services') matched
  files from every directory named 'services/' across the repo — false
  positives in monorepos
- siblings.find() was an O(n) linear scan despite the O(1) claim

Replace with a direct allFiles.has(importerDir + '/' + name + '.py') lookup.
allFiles is a Set<string> of full repo-relative paths, so the lookup is
truly O(1) and exact — no suffix ambiguity possible.

Also fixes: dead code (the '.rb' branch was unreachable since the outer if
gates on Python), and Windows backslash handling via normalize before split.

Made-with: Cursor

* test: remove flag-based demo from unit tests

Made-with: Cursor

* fix(resolver): cover package __init__.py in proximity check and add end-to-end CALLS test

- Also try importerDir/name/__init__.py as a second O(1) candidate so that
  `import user` resolves to services/user/__init__.py when the target is a
  package rather than a bare module file
- Add unit tests for package proximity, __init__.py fallback, and Windows
  backslash path handling
- Add end-to-end CALLS assertion to the bare-import integration test:
  svc.execute() must resolve to UserService#execute in services/user.py,
  proving the fix propagates correctly through the type inference pipeline

Made-with: Cursor

* refactor: extract Python import resolution into resolvers/python.ts

- Move PEP 328 relative import and proximity-based bare import logic
  from standard.ts into a dedicated resolvers/python.ts (resolvePythonImport)
- Dispatch Python imports from resolveLanguageImport in import-processor.ts,
  consistent with how Ruby, PHP, and other languages are handled
- standard.ts is now language-agnostic (TS/JS aliases, Rust paths, suffix fallback)
- Add inline comment on __init__.py vs .py resolution order edge case
- Update unit tests to call resolvePythonImport directly

Made-with: Cursor

* docs: add PEP 302/328/451 references to python.ts comments

Made-with: Cursor

* fix(python): address reviewer comments on PEP compliance

- Guard dirParts.pop() against over-traversal: return null when dot
  count exceeds directory depth, matching CPython's ImportError for
  'attempted relative import beyond top-level package' (PEP 328)
- Swap __init__.py / .py check order to match CPython's finder
  precedence (PEP 451 §4); coexistence is physically impossible so
  order only matters for spec compliance
- Fix overstated PEP 302 comment: proximity check is a static
  heuristic, not a sys.path[0] implementation
- Acknowledge namespace package gap (PEP 420) in docstring
- Add unit test for over-traversal guard

Made-with: Cursor

* test(python): document namespace package resolution behaviour

Add two unit tests for PEP 420 namespace packages (directory with no
__init__.py): bare import returns null (expected — no file exists to
resolve to, CPython sets __file__ = None), while the submodule form
(import user.model) resolves correctly via suffixResolve fallback.

Made-with: Cursor

---------

Co-authored-by: chirag-nighut <chiragnighut@gmail.com>
2026-03-17 16:32:34 +00:00
Zak 64b7ff0061 docs: add Codex MCP configuration to README (#236)
- Add Codex to Editor Support table
- Add Codex manual config example (~/.codex/config.toml)
- Update editor list in usage table

Fixes #131

Made-with: Cursor
2026-03-16 21:23:14 +00:00
Gergő Magyar f2d3df48f6 feat: Phase 5 type resolution — chained calls, pattern matching, class-as-receiver (#315)
* feat: Phase 5 type resolution — chained calls, pattern matching, class-as-receiver, code review fixes

Phase 5.1: Chained method call resolution (depth-capped at 3)
- resolveChainedReceiver() resolves a.getUser().save() by walking the chain
  and looking up intermediate return types from the SymbolTable
- extractReceiverNode() + extractCallChain() shared in utils.ts
- receiverCallChain on ExtractedCall for worker path parity
- MAX_CHAIN_DEPTH=3 enforced in both extraction and resolution

Phase 5.2: Pattern matching binding extractors
- PatternBindingExtractor type added to LanguageTypeConfig
- declarationTypeNodes map tracks original type AST nodes for generic unwrapping
- Rust: if let Some(x)/Ok(x) unwrapping with extractGenericTypeArgs
- Java: instanceof pattern variables (Java 16+)
- C#: is-pattern disambiguation fixture (already working via extractDeclaration)

Phase 5.5d: Python standalone type annotations (name: str)
- expression_statement with type child now captured in DECLARATION_NODE_TYPES

Phase 5.5e: ReceiverKey collision fix for overloaded methods
- receiverKey preserves @startIndex to prevent same-name method collisions
- lookupReceiverType does prefix scan with ambiguity refusal

Class-as-receiver for static method calls (#289)
- UserService.find_user() now resolves via ctx.resolve() tiered lookup
- Respects import scoping — no false positives from unrelated packages

Code review fixes:
- Extracted CALL_EXPRESSION_TYPES + extractCallChain to utils.ts (eliminated duplication)
- Converted resolveChainedReceiver from recursion to loop (no exposed depth param)
- Added depth cap to extractReturnTypeName (defense against nested wrapper types)
- Replaced lookupFuzzy with ctx.resolve for class-as-receiver (architecturally consistent)

Closes #289

Test coverage: 6 new fixtures, 12+ new unit tests, 7 new integration test suites

* fix: Ruby chain calls, Rust Err(x) unwrap, Enum class-as-receiver (#315)

Address three per-language gaps identified in Phase 5 code review:

- Ruby: add `method`/`receiver` field fallbacks to extractCallChain
  (tree-sitter-ruby uses different field names than other grammars)
- Rust: handle `Err(e)` pattern binding via typeArgs[1] from Result<T,E>
- Enum: include Enum type in class-as-receiver filter (both paths)

Integration tests added for all three fixes.

* fix: chain base type resolution parity between serial and worker paths (#315)

- Worker path: add typeEnv.lookup for chain base receiver after extraction
  (typed parameters like `fn process(svc: &UserService)` were silently lost)
- Serial path: add ctx.resolve class-as-receiver fallback for chain base
  (class-name chains like `UserService.find_user().save()` failed)
- Fix misleading comment in parse-worker.ts that described unimplemented logic
- Integration tests: typed-parameter chain, static class-name chain

* fix: Kotlin chain call extraction, createClassNameLookup Enum/Struct (#315)

- Kotlin: extractCallChain now handles navigation_expression → navigation_suffix
  AST structure (Kotlin's call_expression has no 'function' field)
- createClassNameLookup: include Enum and Struct alongside Class for consistent
  constructor recognition in extractInitializer
- Integration test: kotlin-chain-call fixture verifying svc.getUser().save()
2026-03-16 19:38:09 +00:00
Gergő Magyar 5fa73bafdf feat: Phase 4 type resolution — nullable unwrapping, for-loop typing, assignment chains, code review fixes (#310)
* feat: Phase 4 type resolution — nullable unwrapping, for-loop typing, assignment chains, Kotlin return types

Phase 4.1: Nullable/optional chain unwrapping
- Add stripNullable utility in shared.ts for stripping nullable wrappers
- Apply in lookupInEnv to unwrap User | null → User, User? → User before receiver lookup
- Handles TS union, Kotlin/C#/Swift nullable suffix, Python Union[T, None], Rust Option<T>
- Enables receiver-type disambiguation through ?. optional chaining

Phase 4.2: For-loop element typing (Tier 0 — Java/C#/Kotlin)
- Add ForLoopExtractor type and forLoopNodeTypes to LanguageTypeConfig
- Java enhanced_for_statement, C# foreach_statement, Kotlin for_statement extractors
- Only explicit element types in AST (Tier 0); inference-based languages deferred

Phase 4.3: Assignment chain propagation (single-pass, depth-1)
- Add PendingAssignmentExtractor to LanguageTypeConfig with per-language implementations
- Handles TS/JS variable_declarator, Rust let_declaration, Python assignment,
  Go short_var_declaration, C# equals_value_clause, Java/Kotlin variable_declarator
- Single post-walk propagation pass (no fixpoint iteration per Sorbet/Pyright design)
- Resolves const b = a; b.save() when a has known type from Tier 0/1/1b

Phase 4.5: Kotlin return type extraction (bug fix)
- Fix extractMethodSignature to handle Kotlin user_type after function_value_parameters
- Remove lenient test assertions, add strict disambiguation proof

Integration tests across 10+ languages with competing same-name methods
and negative assertions proving disambiguation.

* fix: per-language assignment chain gaps from code review

- Kotlin: new extractKotlinPendingAssignment for property_declaration →
  variable_declaration AST (Java's variable_declarator doesn't exist in Kotlin)
- Go: handle var_spec (var b = u) alongside short_var_declaration (:=)
- PHP: add extractPendingAssignment for $alias = $user with $ prefix preserved

Integration tests added for all three languages with competing
same-name methods and negative disambiguation assertions.

* fix: code review fixes — DRY nullable keywords, avoid array allocations, clarify depth comment

Addresses findings from 6-agent code review on PR #310:

- Move stripNullable JSDoc to correct position (was orphaned above NULLABLE_KEYWORDS)
- DRY: reuse NULLABLE_KEYWORDS set in pipe-split filter instead of inline strings
- Replace node.children.find() with findChildByType/manual loops in jvm.ts,
  go.ts, csharp.ts to avoid unnecessary array allocations per tree-sitter call
- Clarify "depth-1" comment in type-env.ts: single-pass resolves multi-hop
  chains when forward-declared; reverse-order is depth-1 only
- Annotate extractGenericTypeArgs as Phase 5 infrastructure (zero production callers)
- Re-export PendingAssignmentExtractor from index.ts for API consistency
- Add explicit return undefined in Go extractPendingAssignment
- Remove redundant child.text === '=' check in Kotlin extractor

Test coverage:
- 20 new unit tests: stripNullable edge cases, per-language assignment chains,
  reverse-order depth limitation, nullable lookup resolution
- 15 new integration tests: multi-hop chains (a→b→c), nullable+chain combined
  (User|null + alias), Python User|None through stripNullable path
- 3 new fixtures: ts-multi-hop-chain, ts-nullable-chain, python-nullable-chain

* fix: third-pass review — walrus chain, scanner allocations, Kotlin variable_declaration, C# type guard

Addresses 4 new findings from third-pass CI review:

1. Python walrus operator (:=) now handled by extractPendingAssignment —
   named_expression nodes propagate alias chains alongside regular assignment
2. Scanner .namedChildren.find()/.some() in jvm.ts replaced with
   findChildByType() — consistent with 98daed4 code review fixes
3. Kotlin extractPendingAssignment extended to handle variable_declaration
   nodes in addition to property_declaration (function-local val/var)
4. C# extractPendingAssignment early-returns for is_pattern_expression and
   field_declaration nodes (never contain variable_declarator children)

Integration tests:
- Python: walrus chain (alias := u) with disambiguation (5 tests, 1 fixture)
- Kotlin: assignment chain with typed declarations (5 tests, 1 fixture)
- C#: assignment chain + is-pattern coexistence (6 tests, 1 fixture)
- Unit: Python walrus propagation (1 test)

* feat: nullable wrapper unwrapping + C++ assignment chains

Gaps 1, 2, 4 from code review — architectural changes to type resolution:

1. extractSimpleTypeName now unwraps nullable wrapper generics:
   - Optional<User> → "User" (Java), Option<User> → "User" (Rust),
     Maybe<User> → "User" (Kotlin Arrow/Haskell-style)
   - Containers (List, Map) and async wrappers (Promise, Future) are NOT
     unwrapped — methods are called on the container, not the inner type
   - Uses existing extractGenericTypeArgs (now production-active, was dead code)
   - NULLABLE_WRAPPER_TYPES set: Optional, Option, Maybe

2. C++ extractPendingAssignment added for auto alias chains:
   - auto alias = user; alias.save() now propagates User type
   - Handles pointer/reference declarators, auto/decltype(auto)

3. Updated existing Rust test: Option<User> parameter now correctly
   stores "User" instead of "Option" in TypeEnv

Integration tests with fixtures for Java Optional, Rust Option, C++ auto
chain. Full pipeline resolution marked .todo — requires call-processor
enhancement (TypeEnv stores correct types but call-processor needs
additional work to produce CALLS edges for these patterns).

Unit tests: 196 passed (7 new). Integration: all 9 languages green.

* fix: resolve .todo tests — stale dist/ was the root cause

The Rust Option<User> and C++ auto assignment chain integration tests
were marked .todo because the pipeline didn't produce CALLS edges.
Root cause: dist/ was compiled from pre-Phase 4 source and lacked:
- NULLABLE_WRAPPER_TYPES unwrapping in extractSimpleTypeName
- C++ extractPendingAssignment

After npm run build, all tests pass as real assertions:
- Rust: alias.save() resolves to User#save via Option<User> unwrap + chain
- C++: alias.save() and rAlias.save() resolve via auto assignment chain
  with correct disambiguation (User vs Repo)

Only remaining .todo: Rust user.unwrap().save() (Phase 5 — chained
return type inference, not a TypeEnv issue).
2026-03-16 15:21:54 +00:00
fbff6d08c0 feat(ingestion): respect .gitignore and .gitnexusignore during file discovery (#231)
* feat(ingestion): respect .gitignore and .gitnexusignore during file discovery

Add support for excluding files from indexing based on .gitignore and
.gitnexusignore patterns. Previously, GitNexus used only a hardcoded
ignore list, causing significant index pollution in repositories with
git-ignored directories containing code (e.g., Docker-mounted volumes).

Changes:
- Add `ignore` package for gitignore-spec pattern matching
- Add `loadIgnoreRules()` to parse .gitignore + .gitnexusignore
- Add `createIgnoreFilter()` returning glob-compatible IgnoreLike object
- Integrate filter into glob's `ignore` option for directory-level pruning
- Remove post-glob `.filter()` call (now handled during traversal)

The hardcoded DEFAULT_IGNORE_LIST remains as fallback for non-git repos.

Closes #228

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ingestion): address review feedback on ignore filtering

- Distinguish ENOENT vs EACCES in loadIgnoreRules (warn on permission errors)
- Add GITNEXUS_NO_GITIGNORE env var to bypass .gitignore parsing
- Fix bare-name pattern matching in childrenIgnored (check both with/without trailing slash)
- Rename isIgnoredDirectory to isHardcodedIgnoredDirectory for clarity
- Add clarifying comments for design decisions (D2 negation, D3 dot:false redundancy)
- Add tests for bare-name patterns, file-glob patterns, EACCES handling, env var

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ingestion): address second round of review feedback

- G1: Document GITNEXUS_NO_GITIGNORE in `analyze --help` and log when active
- G2: Add comment clarifying path-scurry POSIX normalization contract
- G3: Add IgnoreOptions interface — env var now falls back, callers can
  pass `noGitignore` explicitly for testability and future CLI flag
- G4: Add integration test verifying walkRepositoryPaths respects the env var

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(ingestion): gracefully skip files with unavailable tree-sitter grammars

Port unsupported language resilience from PR #301 by @jecanore.
- Make Kotlin import optional (like Swift) in parser-loader and parse-worker
- Add worker-local isLanguageAvailable() with filePath param for tsx distinction
- Track and log skipped files per language in both sequential and worker paths
- Add skippedLanguages to ParseWorkerResult for worker→main aggregation
- Add isLanguageAvailable unit tests

Refs: #301, #155, #228

Co-Authored-By: jecanore <juan@housingbase.io>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test(e2e): add ignore + language-skip end-to-end test with fixture repo

Add a fixture repo (test/fixtures/ignore-and-skip-repo/) with .gitignore,
.gitnexusignore, TypeScript source files, and a Swift file to exercise
all three features end-to-end:

- File discovery: verifies .gitignore excludes data/ and *.log,
  .gitnexusignore excludes vendor/, source files are discovered
- Parsing: verifies TypeScript files produce Function nodes and DEFINES
  relationships, Swift files are skipped gracefully when grammar is
  unavailable

Add the test to the standalone group in ci-integration.yml and coverage job.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ci): move ignore-and-skip-e2e test to e2e group per review feedback

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(test): use temp directory instead of fixture for e2e ignore test

The fixture's .gitignore prevented data/seed.json and debug.log from
being committed — these files would be missing after checkout in CI.

Switch to creating the entire test structure in a temp directory via
beforeAll (matching filesystem-walker.test.ts pattern). This ensures
all files exist regardless of git ignore rules.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(test): correct graph API usage in e2e ignore test

Use graph.nodes property getter instead of graph.getNodes(), and check
Function node filePath instead of non-existent File nodes (File nodes
are created by processStructure, not processParsing).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* ci: add workflows permission to ci-integration.yml

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* ci: change workflows permission to write per review

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* ci: move workflows permission from ci-integration.yml to ci.yml caller

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ci): fix Claude workflows for fork PRs, remove misplaced workflows perm

Three issues prevented Claude from running on fork PRs:

1. claude-code-review.yml lacked workflows:write — push failed when
   fork PRs modify .github/workflows/ files
2. claude.yml had no fork PR support — checked out main and couldn't
   fetch the fork's branch from origin
3. Cleanup step unconditionally deleted branches even when push failed,
   breaking the concurrent claude.yml workflow

Also removes workflows:write from ci.yml's integration job — CI tests
don't need that permission. The permission belongs on the claude
workflows that push fork branches.

Changes:
- Add workflows:write to both claude workflow permissions blocks
- Add fork PR detection + branch push/cleanup to claude.yml
- Add step id to push-fork; cleanup only runs if push succeeded
- Pass branch names via env vars to prevent shell injection (security)
- Add concurrency groups to prevent race conditions between workflows
- Remove misplaced workflows:write from ci.yml integration job

* fix(ci): use GitHub API for fork branch refs instead of git push

GITHUB_TOKEN cannot have 'workflows' permission — it's only valid for
PATs and GitHub Apps. This means git push fails whenever a fork PR
modifies .github/workflows/ files.

Replace git push with the GitHub REST API (POST/PATCH /git/refs) to
create temporary branch refs. The API creates a pointer to the
already-existing PR head commit without triggering the workflow file
push protection. Similarly, cleanup uses DELETE /git/refs instead of
git push --delete.

Also removes the invalid 'workflows: write' from permissions blocks.

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: jecanore <juan@housingbase.io>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-03-16 13:26:20 +00:00
Gergő Magyar 6c18ae08f7 feat: return type inference, doc-comment parsing, and per-language type extractors (#284)
* feat: Phase 3 — return type inference, generic args extraction, Ruby YARD type extractor

Three architectural improvements to the type resolution system:

1. Return type inference — wire extractMethodSignature returnType through
   SymbolDefinition into call-processor. When var = callee() and callee
   has a known return type, bind var to that type. Handles Promise<T>
   unwrapping, nullable stripping, pointer/reference removal.

2. Generic type argument extraction — new extractGenericTypeArgs() utility
   that extracts type parameters from List<User> → ['User']. Handles
   TS/Java/Kotlin/C#/Rust generic syntax. Building block for for-loop
   variable typing.

3. Ruby dedicated type extractor — replaces the stub with YARD annotation
   parsing (@param name [Type]), handling qualified types, nullable types,
   and singleton methods. Ruby now has real type resolution.

Unit tests: 127 → 192+ (type-env) + 65 (symbol-table, call-processor) + 18 (generics)
Integration tests: 8+ new test cases with fixtures across TS/Python/Go/Java/Ruby

* fix: Phase 3 gaps — WRAPPER_GENERICS correctness, Ruby :: qualifier, namespaced constructors

- Remove collection types (List, Array, Vec, Set) from WRAPPER_GENERICS to prevent
  false CALLS edges (e.g. List<User> no longer unwraps to User)
- Add :: qualifier handling in extractReturnTypeName for Ruby/C++/Rust namespaced types
- Add Ruby `constant` and `scope_resolution` node types to shared extractors
- Extract shared extractRubyConstructorAssignment helper (dedup type-env.ts + ruby.ts)
- Add integration tests for return type inference: Python, TypeScript, Go, Java, Ruby
- Add Ruby namespaced constructor fixture (Models::UserService.new)
- Add unit tests for collection reclassification and :: qualifiers

* feat: Phase 4 — CONSTRUCTOR_BINDING_SCANNERS for all languages + return type inference tests

Add CONSTRUCTOR_BINDING_SCANNERS for 6 missing languages, completing
return type inference coverage across all 11 supported languages:

- TypeScript/JS: variable_declarator with call_expression, unwraps await
- Go: short_var_declaration single-assignment (skips multi-return, new/make)
- Java: local_variable_declaration with `var` type + method_invocation
- C#: variable_declaration with implicit_type (var) + invocation_expression
- Rust: let_declaration without type annotation, handles mut_pattern
- PHP: assignment_expression with function_call_expression

Also adds property_identifier to extractSimpleTypeName for qualified
member calls (repo.getUser → getUser), fixing namespaced constructor
inference that was previously a known limitation.

Integration tests added for all 11 languages with correct label
assertions (Function vs Method per language's tree-sitter queries).

* refactor: merge CONSTRUCTOR_BINDING_SCANNERS into per-language LanguageTypeConfig

Eliminates the parallel dispatch map in type-env.ts by moving all 11
constructor binding scanners into their respective type-extractors/*.ts
files as `scanConstructorBinding` on LanguageTypeConfig.

- Add ConstructorBindingScanner type to types.ts
- Add shared helpers: hasTypeAnnotation, unwrapAwait, extractCalleeName
- Move scanners to typescript.ts, jvm.ts, python.ts, php.ts, go.ts,
  rust.ts, swift.ts, c-cpp.ts, csharp.ts, ruby.ts
- Fix `any` types in C# scanner → SyntaxNode | null
- Delete ~300 lines from type-env.ts (CONSTRUCTOR_BINDING_SCANNERS map)
- Update buildTypeEnv to use config.scanConstructorBinding

All 143 type-env unit tests and all 10 language integration suites pass.

* fix: remove unused import, fix any type in Java scanner, update stale comment

- Remove unused extractCalleeName import from jvm.ts
- Fix (c: any) → (c: SyntaxNode) in Java scanner
- Update stale CONSTRUCTOR_BINDING_SCANNERS reference in ruby.ts comment

* fix: C# and PHP return type inference — scanner fixes, method signature extraction, and cross-file resolution

Addresses code review findings on PR #284:

C# scanner (csharp.ts):
- Fix type node lookup: iterate children instead of childForFieldName('type')
  which returns undefined in tree-sitter-c-sharp
- Fix initializer lookup: handle direct invocation_expression children
  (no equals_value_clause wrapper in tree-sitter-c-sharp)

C# return type extraction (utils.ts):
- Add 'returns' field check to extractMethodSignature — tree-sitter-c-sharp
  uses 'returns', not 'type', for method return types

C# cross-file resolution (call-processor.ts + fixture):
- Add constructor binding verification to sequential processCalls path
  (was only in the worker processCallsFromExtracted path)
- Add ReturnType.csproj to csharp-return-type fixture
- Update fixture namespaces to use ReturnType.Models/ReturnType.Services
  prefix (matches real C# project conventions)

PHP scanner (php.ts):
- Extend scanConstructorBinding to handle member_call_expression
  ($this->getUser() patterns), not just function_call_expression

Shared (shared.ts):
- Add member_access_expression to extractSimpleTypeName qualified-names
  block (C# method calls like svc.GetUser())

Tests:
- Add Repo.cs/Repo.php disambiguation fixtures (two Save methods)
- Strengthen C# and PHP return type tests with hard disambiguation assertions
- Add C# scanner unit tests and return type extraction test

* feat: per-language ReturnTypeExtractor + doc-comment @param parsing for PHP, JS, Ruby

Add ReturnTypeExtractor to LanguageTypeConfig interface with implementations
for Ruby (YARD @return), PHP (PHPDoc @return), and JS/TS (JSDoc @returns).
The fallback is wired in both parsing-processor and parse-worker paths,
activating only when extractMethodSignature finds no AST-based return type.

Also add doc-comment @param type extraction for PHP and JS/TS, following
Ruby's existing collectYardParams pattern. This enables parameter.method()
resolution in loosely-typed codebases using PHPDoc @param or JSDoc @param.

Additional fixes from PR #284 code review:
- Go: add selector_expression + field_identifier to extractSimpleTypeName
  (enables package-qualified factory calls like models.NewUser())
- Ruby: broaden scanConstructorBinding to capture plain call assignments
  (user = get_user()) in addition to Class.new patterns
- Ruby: harden return-type fixture with disambiguation (two save methods)

Test coverage: +14 new integration tests across Go, Ruby, PHP, JS/TS

* fix: JSDoc async return type, PHP attribute walkers, and $this receiver disambiguation

Three fixes from fourth-pass code review on PR #284:

1. JSDoc `@returns {Promise<User>}` no longer stripped to `Promise` — extractReturnType
   now uses sanitizeReturnType (preserves generics) instead of normalizeJsDocType
   (which stripped them before extractReturnTypeName could unwrap WRAPPER_GENERICS).

2. PHP 8+ `#[Attribute]` and JS `@decorator` nodes no longer break doc-comment walkers.
   Both extractReturnType and collect*Params functions now skip attribute_list/decorator
   nodes instead of breaking on them as named siblings.

3. PHP `$this->method()` now provides receiverClassName for disambiguation.
   When two classes define the same method, the enclosing class narrows candidates
   via ownerId matching in call-processor, preventing false no-binding results.

* fix: sanitizeReturnType dot corruption, JS test assertions, Ruby constant receiver

- Remove redundant dot-path stripping from sanitizeReturnType that corrupted
  qualified names inside generics (e.g. Promise<models.User> → User>)
- Split JS async fixture into separate files and add negative assertions
  to properly verify disambiguation (mirroring PHP test pattern)
- Accept 'constant' node type in Ruby scanConstructorBinding for factory
  call assignments (SERVICE = build_service())
- Add 'constant' to SIMPLE_RECEIVER_TYPES so extractReceiverName handles
  Ruby constant receivers (SERVICE.process)

* fix: nested generic arg splitting, JS/Ruby test false positives

- Replace naive comma split in extractReturnTypeName with bracket-balanced
  extractFirstGenericArg so nested types like Future<Result<User, Error>>
  unwrap correctly instead of producing malformed "Result<User"
- Add CompletableFuture to WRAPPER_GENERICS for Java async unwrapping
- Split js-jsdoc-return-type fixture models.js into user.js/repo.js and
  add negative assertions to prove disambiguation (not just file match)
- Split ruby-constant-factory-call fixture into separate service files
  and add negative assertions against AdminService resolution

* fix: review findings — receiverClassName parity, Rust wrappers, Go multi-return, Kotlin/Swift qualified calls

P1: Sequential path now includes receiverClassName narrowing for PHP
$this->method() disambiguation (was missing vs worker path).

P2: Added Rc/Arc/Weak/MutexGuard/Cow + 6 more Rust Deref types to
WRAPPER_GENERICS (Box excluded — Java Swing collision). Extended
Kotlin/Swift scanners to handle navigation_expression callees.
Added Go multi-return support (user, err := f()) with blank/_/err/ok
guard + AST-level first-return extraction in extractMethodSignature.

P3: Extracted shared verifyConstructorBindings() eliminating 60 lines
of duplication between sequential and worker paths. Added return-type
inference integration tests for C++, Rust, Swift with competing
methods and negative disambiguation assertions.

* fix: Swift navigation_suffix unwrapping, Rust lifetime skipping, Kotlin disambiguation tests

- Swift scanConstructorBinding: handle tree-sitter wrapping qualified
  identifiers in navigation_suffix nodes
- Add extractFirstTypeArg to skip Rust lifetime parameters ('a, '_)
  when unwrapping wrapper generics like Ref<'_, User>
- Kotlin tests: add Repo class fixture with competing save() methods
  to prove disambiguation; assert no spurious edges on known gap
- Remove tree-sitter-kotlin from optionalDependencies (now regular dep)

* fix: C# null-conditional calls, Ruby YARD bracket-balanced split, PHPDoc alternate order, escapeValue hardening

- Add C# null-conditional call support (user?.Save()): tree-sitter query for
  conditional_access_expression, member_binding_expression in MEMBER_ACCESS_NODE_TYPES,
  receiver extraction via conditional_access_expression parent walk
- Fix Ruby YARD type parsing for nested generics (Hash<Symbol, User>): replace
  naive split(',') with bracket-balanced splitter respecting <> depth
- Add alternate YARD format (@param [Type] name) alongside standard (@param name [Type])
- Add alternate PHPDoc format (@param $name Type) alongside standard (@param Type $name)
- Harden escapeValue in kuzu-adapter.ts: escape \n and \r to prevent Cypher injection
- Integration tests: C# null-conditional fixture (5 tests), Ruby YARD generics fixture (6 tests)
- Unit tests: PHPDoc alternate order (2 tests), C# null-conditional call-form (updated)

* test: add Python static/classmethod integration tests (issue #289)

Verifies that classes using only @staticmethod/@classmethod have HAS_METHOD
edges connecting them to their child methods. This was the root cause of
issue #289 where context() and impact() returned empty for such classes.

Tests cover: HAS_METHOD edge emission, unique static method resolution
(create_user, delete_user), and ambiguous same-named method handling
(find_user on both UserService and AdminService — safely refused).

* fix: lbug batch escapeValue newline hardening, Rust ::default() scanner exclusion

- Apply \n/\r escaping to batch upsert escapeValue in lbug-adapter.ts:429
  (missed instance of the CREATE-path fix from ec4dca4)
- Exclude Rust ::default() from scanConstructorBinding to match
  extractInitializer behavior — avoids wasted cross-file lookups on
  the broadly-implemented Default trait
- Unit tests: 2 new scanner exclusion tests (::default and ::new)
- Integration tests: 6 new Rust ::default() constructor resolution tests
  with disambiguation fixture (User::default vs Repo::default)

* fix: C#/Rust async await unwrap, PHP backslash namespace, fallback escaping

- C# scanConstructorBinding: unwrap await_expression to find invocation_expression
  (var user = await svc.GetUserAsync() now produces constructor binding)
- Rust scanConstructorBinding: unwrap .await postfix via shared unwrapAwait helper
  (let user = get_user().await now produces constructor binding)
- extractReturnTypeName: handle PHP backslash namespace separator (\App\Models\User → User)
- fallbackRelationshipInserts: match batch escapeValue hardening with \n/\r escaping

Tests: 2 unit (type-env), 3 unit (call-processor), 7 integration (csharp+rust), 7 fixtures

* fix: C#/Rust async-binding test false positives — add competing types and negative assertions

C# fixture: add Order.cs with Order.Save(), change OrderService to return
Task<Order> via GetOrderAsync, add negative assertion proving user.Save()
does not resolve to Order#Save.

Rust fixture: split models.rs into user.rs/repo.rs, make process_user and
process_repo async fn, add bidirectional negative assertions proving no
cross-contamination between User#save and Repo#save.

* fix: C# async-binding broken assertion, bare wrapper type leak, JSDoc optional params

- Split Program.cs Main into ProcessUser/ProcessOrder so negative
  assertions use strict toBeUndefined() (matching Rust pattern)
- Guard bare wrapper types (Task, Promise, Option…) in
  extractReturnTypeName — return undefined instead of the wrapper name
- Update JSDOC_PARAM_RE to capture @param {Type} [optionalName] syntax

* fix: update symbol and relationship counts in documentation
2026-03-15 18:49:40 +00:00
Candido Sales GomesandClaude Opus 4.6 5a5850832c refactor: migrate from KuzuDB to LadybugDB v0.15 (#275)
* refactor: migrate from KuzuDB to LadybugDB v0.15

KuzuDB was archived (Apple acquisition, Oct 2025). LadybugDB is the
community fork with full API compatibility.

- Package swap: kuzu → @ladybugdb/core, kuzu-wasm → @ladybugdb/wasm-core
- Rename all internal paths: kuzu → lbug (adapters, schema, storage)
- Storage path: .gitnexus/kuzu → .gitnexus/lbug (with auto-cleanup)
- Add explicit VECTOR extension loading (required in v0.15)
- Update CI workflow, documentation, and all tests
- 1151 unit + 27 integration tests passing

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address code review findings (P1-P3)

P1: Fix WASM adapter to use getAll() API, wire cleanupOldKuzuFiles
into analyze command, add symlink path traversal protection.
P2: Cache VECTOR extension load state, batch augmentation engine
queries (20→4), fix web getCopyQuery for multi-language tables,
fix stale KuzuDB references, correct brainstorm package names.
P3: Complete lbug-wasm.d.ts type declarations, batch semantic
search per-label, update stale BM25 comment.

* chore: remove outdated KuzuDB migration brainstorming document

* fix: load FTS extension in MCP pool adapter on init

The read-only pool adapter never loaded the FTS extension, so all
QUERY_FTS_INDEX calls failed silently. This broke search-pool and
augmentation integration tests, and caused empty results in the
web UI server mode.

* feat: implement shared Database caching and connection reference counting

* feat: enhance KuzuDB migration handling and status reporting

* fix: mock cleanupOldKuzuFiles in local backend callTool tests

* fix: update mock for cleanupOldKuzuFiles and adjust imports in callTool tests

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 15:53:01 +00:00
Gergő Magyar 62242d5f44 feat: TypeEnvironment API with constructor inference, self/this/super resolution (#274)
* feat(type-env): constructor-call type inference for TypeEnv (Phase 1)

Add extractInitializer as a Tier 1 fallback in buildTypeEnv: when a
declaration node has no explicit type annotation, infer the type from
constructor-call patterns (new X(), X::new(), X::default(), $x = new X()).

Languages covered: TypeScript/JS, Java (var), Rust, PHP, C++ (auto).
Python/Kotlin/Swift deferred — need symbol-table access to distinguish
class constructors from function calls.

Adds 20 new unit tests covering constructor inference, annotation
precedence, and known limitations across all supported languages.

* fix(type-env): class-aware constructor resolution, multi-declarator fix

- Add collectClassNames pre-scan: walks AST to build Set<string> of
  class/struct names defined in the file
- C++ extractInitializer uses classNames.has() to verify identifier is
  a known class before inferring (auto x = User() resolves, auto x =
  getUser() does not — no false positives)
- Add InitializerExtractor type that receives classNames parameter
- Fix env.size gating: always call extractInitializer when available,
  so mixed declarators like const a: A = x, b = new B() resolve both
- Add env.has() guard in Java extractInitializer to skip already-bound vars
- Document Rust new/default whitelist rationale
- Pin all test assertions, add mixed multi-declarator test case

* fix(type-env): resolve Self/self/static/parent to actual type names

- Rust: Self::new()/Self::default() resolves to enclosing impl type
- PHP: new self()/static() resolves to enclosing class, parent() to superclass
- Rust: Tier 0 annotation guard prevents overwrite by constructor inference
- Rust: mut_pattern handling in extractVarName for let mut bindings
- TS: fix misleading comment in extractInitializer
- 58 tests passing (3 new Self/self resolution tests)

* perf(type-env): single-pass AST walk with closure-scoped state

Refactors buildTypeEnv to use closures instead of passing mutable state
as parameters. classNames, env, and config are captured by the inner
walk and extractTypeBinding functions — no parameter mutation.

- Eliminates separate collectClassNames pre-scan (O(2n) → O(n))
- config looked up once per file instead of per-node
- 29 fewer lines

* feat(type-env): constructor-inferred type resolution for all languages

Add cross-file constructor type inference to the ingestion pipeline,
enabling receiver-type disambiguation for member calls like
`user.save()` when the variable is assigned from a constructor without
explicit type annotations.

Pipeline changes:
- Add extractInitializer to Python and Swift type extractors
- Add CONSTRUCTOR_BINDING_SCANNERS for Python, Swift, C/C++ in type-env
- Wire constructorBindings through parse-worker → parsing-processor →
  pipeline → processCallsFromExtracted
- Rewrite resolveCallTarget receiver-type filtering (step D) to use
  tiered import resolution (same-file → import-scoped → global) before
  falling back to fuzzy ownerId matching
- Use collectTieredCandidates for constructor binding verification
  instead of raw lookupFuzzy

Bug fixes:
- Fix C++ inline method query: @definition.method was captured on
  field_declaration_list instead of function_definition, causing wrong
  parameterCount for all inline class methods
- Fix parse-worker accumulated/flush results missing constructorBindings

CI changes:
- Add swift.test.ts to ci-integration pipeline group and coverage job
- Update ci-report to fetch base branch (main) coverage for delta
  reporting instead of showing config thresholds
- Add per-suite timing breakdown table (unit/integration/total)
- Add expandable skipped test details section

Tests: 288 passed, 4 skipped (swift — macOS only) across 10 languages
- 36 new constructor-inferred integration tests (4 per language)
- 10 fixture directories with cross-file constructor patterns
- TypeScript, JavaScript, Java, Kotlin, Python, PHP, Rust, Go, C++, Swift

* fix(type-extractors): add type assertion for LanguageTypeConfig

* feat(ruby): constructor-inferred type resolution and self-receiver mapping

Add Ruby User.new constructor binding scanner to type-env, enabling
receiver-type disambiguation for member calls like user.save vs repo.save.
Add self/this → enclosing class resolution in lookupTypeEnv so self.method()
calls resolve to the correct class even when the method name is ambiguous.

* docs: update README with constructor inference and self/this resolution details

* refactor(ingestion): unified ResolutionContext replaces fragmented map passing

Introduce createResolutionContext() as the single resolution API for all
processors. Eliminates duplicated tier-selection logic, fixes heritage
namedImportMap bug, and adds per-file resolution caching.

- NEW resolution-context.ts: closure-factory with resolve(), per-file cache,
  TIER_CONFIDENCE constant, and shared ResolutionTier type
- DELETE symbol-resolver.ts: zero production importers, logic now in
  resolution-context.ts
- call-processor: all functions take ctx instead of 6 separate maps,
  collectTieredCandidates removed (ctx.resolve replaces it),
  D4 redundant re-resolve eliminated
- heritage-processor: takes ctx, resolveHeritageId helper extracts
  repeated 14-line fallback pattern, namedImportMap now included
- import-processor: takes ctx, dead createImportMap/createPackageMap/
  createNamedImportMap factories removed
- pipeline: creates single ctx, wires onProgress to all processors,
  logs cache hit rate in dev mode
- Tier renamed: unique-global → global (honest about returning all candidates)
- Tests migrated: 1178 unit + 84 integration passing

* feat(type-env): self/this/super resolution, TypeEnvironment API, and review fixes

Add cross-language receiver keyword resolution:
- self/this/$this → enclosing class name via AST walk
- super/base/parent → parent class name via heritage AST extraction
  (8 grammar variants: TS/JS, Java, Python, Ruby, C#, PHP, Kotlin, C++, Swift)
- D-phase widening in resolveCallTarget for super→parent method dispatch

Introduce TypeEnvironment API replacing loose TypeEnvResult + lookupTypeEnv:
- buildTypeEnv() returns TypeEnvironment with .lookup() method
- Single-pass AST walk merges constructor binding scan (was separate traversal)
- ClassNameLookup type replaces over-broad ReadonlySet<string> facade
- Memoized class name lookups to avoid redundant SymbolTable scans

Code review fixes (6 agents, 11 findings):
- Replace ctx.resolve(name, '') hack with direct symbols.lookupFuzzy()
- Extract scope key helpers (extractFuncNameFromScope, receiverKey)
- Simplify D-phase from 5 steps to 4 with deduped typeNodeIds
- Remove C from CONSTRUCTOR_BINDING_SCANNERS (YAGNI — C has no constructors)
- Cache Map reuse in ResolutionContext to reduce GC pressure
- Remove unused TieredCandidates import

Integration tests for self/this, parent, and super resolution across all
12 supported languages with per-language fixture directories.

* fix(type-env): generic parent resolution, TS cast inference, C++ brace-init

Fix generic parent class breaking super resolution:
- extractParentClassFromNode now uses extractSimpleTypeName to strip
  generic params (Base<T> → Base) and qualified names (models.Model → Model)
- Affects TS, Java, Python, C# heritage extraction

Fix TypeScript new X() as T / new X()! missed inference:
- Unwrap as_expression and non_null_expression before checking for
  new_expression in extractInitializer

Fix C++ brace-init User{} missed inference:
- Handle compound_literal_expression with type_identifier child
  in extractInitializer

Clean up deprecated lookupTypeEnv:
- Remove standalone lookupTypeEnv export, migrate all callers to
  TypeEnvironment.lookup() method
- Update all 80+ test assertions to use the new API

Integration test fixtures added:
- typescript-cast-constructor-inference (new X() as T, new X()!)
- typescript/java/csharp/kotlin-generic-parent-resolution
- cpp-brace-init-inference (auto x = User{})

* fix(type-extractors): Go &User{}, TS double-cast, Swift .init inference

Fix Go pointer-to-struct literal not inferred:
- Unwrap unary_expression (address-of &) before composite_literal check
- user := &User{} now correctly infers type User

Fix TypeScript double-cast only unwrapping one level:
- Change if to while loop for nested as_expression/non_null_expression
- new User() as unknown as Admin now correctly infers type User

Fix Swift User.init(name:) explicit init call missed:
- Handle navigation_expression callee with .init suffix in extractInitializer

Integration test fixtures:
- go-pointer-constructor-inference (&User{}, &Repo{})
- typescript-double-cast-inference (as unknown as T)

* feat: Rust struct literal, Python qualified ctor, Go new(), Swift .init scanner

- Rust: handle struct_expression in extractInitializer (User { name: "alice" })
- Python: support attribute nodes in extractInitializer (models.User("alice"))
  and the cross-file scanner — extractSimpleTypeName handles qualified names
- Go: handle new(User) built-in in extractGoShortVarDeclaration
- Swift: extend CONSTRUCTOR_BINDING_SCANNERS to handle navigation_expression
  callee for User.init(name:) cross-file resolution

Unit tests: 87 → 96 (Rust struct literal, Go new(), Python qualified ctor,
Python scanner qualified, plus edge cases)
Integration tests: 4 new describe blocks with fixtures

* fix: Rust Self{} resolution, C++ scoped brace-init, PHP promotion params, Ruby constants

- Rust: resolve Self {} struct literal to enclosing impl type (was stored as "Self")
- C++: replace type_identifier guard with extractSimpleTypeName for compound_literal_expression,
  enabling ns::User{} scoped brace-init (closes previously deferred gap)
- PHP: add property_promotion_parameter to TYPED_PARAMETER_TYPES for PHP 8.0+
  constructor property promotion (__construct(private Foo $x))
- Ruby: extend extractRubyConstructorBinding to accept constant left-hand side
  (REPO = Repo.new)

Unit tests: 96 → 101 (+5: Rust Self{} ×2, C++ ns::User{} ×1, PHP promotion ×1,
Ruby constant ×1)
Integration tests: 4 new describe blocks with fixtures

* feat: Phase 1 type resolution gaps — walrus, PHP properties, nullable, Go make/assert

Phase 1 quick wins from the type resolution gap analysis:

1. Python walrus operator := (named_expression) — extractInitializer + scanner
2. PHP 7.4+ typed class properties — property_declaration in extractDeclaration
3. Nullable union unwrapping — User | null → User in extractSimpleTypeName
4. Go make() builtin — slice/map element type extraction
5. Go type assertions — iface.(User) type extraction

Also: PHP primitive_type handling in extractSimpleTypeName (string, int, etc.)

Unit tests: 101 → 114 (+13)
Integration tests: 8 new describe blocks with fixtures

* feat: Phase 2 type resolution gaps — C++ range-for, Rust if-let, C# pattern matching, Python class annotations

Phase 2 medium-effort improvements:

1. C++ range-for with explicit type — for (User& u : vec) binds u: User
2. Rust if-let/while-let captured_pattern — user @ User { .. } binds user: User
3. C# is-pattern matching — if (obj is User user) binds user: User
4. Python class-level annotations — confirmed already working, added tests

Unit tests: 114 → 127 (+13)
Integration tests: 11 new test cases with fixtures
2026-03-14 19:05:49 +00:00
Gergő Magyar 6e38db879e fix(ruby): method-level call resolution, HAS_METHOD edges, and dispatch table (#278)
* fix(ruby): method-level call resolution, HAS_METHOD edges, and dispatch table refactoring

- Replace all `if (language === Ruby)` checks in processors with a
  `callRouters` dispatch table in call-routing.ts (renamed from
  ruby-call-routing.ts to preserve git history)
- Add Ruby `method` and `singleton_method` to FUNCTION_NODE_TYPES so
  findEnclosingFunction produces Method-level CALLS sources
- Add Ruby `class` and `module` to CLASS_CONTAINER_TYPES for HAS_METHOD
  edge generation
- Add bare call capture via tree-sitter query `(body_statement (identifier))`
  for Ruby methods called without parentheses
- Add Ruby member call detection (`call` node with `receiver` field) to
  inferCallForm and extractReceiverName
- Wire resolveRubyImport into resolveLanguageImport
- Add 24 integration tests across 5 suites: heritage/properties, arity
  filtering, member calls, ambiguous disambiguation, local shadow
- Add ruby.test.ts to CI integration workflow

* fix(ruby): resolve 6 Ruby resolution gaps from PR review

- Fix singleton_method label mismatch: @definition.function → @definition.method
  so CALLS edges from `def self.foo` bodies get correct sourceId
- Add ownerId and HAS_METHOD edges to attr_* Property nodes by calling
  findEnclosingClassId in both parse-worker and call-processor property branches
- Distinguish include/extend/prepend heritage: add heritageKind to
  RubyHeritageItem, propagate through heritage pipeline as IMPLEMENTS reason
- Document bare call over-capture limitation in tree-sitter query comment
- Add bare `require` (non-relative) import test coverage
- Add prepend/extend test coverage with distinct Loggable/Cacheable modules

31 Ruby integration tests passing, no regressions in other language resolvers.

* fix(ruby): web package parity — heritage reasons, property HAS_METHOD, singleton_method label

- Web call-processor: use item.heritageKind as IMPLEMENTS reason instead of
  hardcoded 'trait-impl', add :${kind} suffix to edge ID for uniqueness
- Web call-processor: port findEnclosingClassId, add HAS_METHOD edges for
  attr_* Property nodes to match CLI fix
- Web call-processor: singleton_method label 'Function' → 'Method' to match
  CLI tree-sitter query fix
- CLI parse-worker: update stale ExtractedHeritage.kind JSDoc to include
  'include' | 'extend' | 'prepend'

* fix(web): add HAS_METHOD to RelationshipType union

Web package was missing HAS_METHOD in the RelationshipType union,
causing a type mismatch with the HAS_METHOD edges emitted by the
attr_* property fix in call-processor.ts.
2026-03-14 11:34:30 +00:00
Candido Sales Gomes 0999595444 feat(ruby): Add Ruby language support for CLI and web (#111) 2026-03-13 21:46:59 +00:00
Chirag Nighut 649ad80dbb fix(cli): dynamically discover and install agent skills (#270) 2026-03-13 19:09:03 +00:00
Gergo Magyar 3dbe08fab6 chore: release v1.4.0 2026-03-13 13:18:14 +00:00
Gergő Magyar 1afe9166aa feat: language-aware code intelligence — symbol resolution, MRO, constructor discrimination (#238)
* feat: add Method Resolution Order (MRO) with language-specific rules

Implement full MRO computation for multi-language inheritance hierarchies:

- HAS_METHOD edges: Class→Method ownership edges emitted during parsing
  (both worker pool and sequential fallback paths)
- Method signatures: extract parameterCount and returnType from AST nodes
- C# heritage fix: distinguish EXTENDS vs IMPLEMENTS for base_list captures
  using symbol table lookup + I[A-Z] naming heuristic fallback
- MRO processor (Phase 4.5): walks inheritance DAG, detects method-name
  collisions across parents, applies language-specific resolution:
  - C++: leftmost base class in declaration order wins
  - C#/Java: class method wins over interface default
  - Python: C3 linearization with cycle detection
  - Rust: no auto-resolution (requires qualified syntax)
  - Default: first definition in BFS order wins
- OVERRIDES edges emitted for resolved method collisions
- KuzuDB schema: Method table extended with parameterCount/returnType;
  dedicated CSV writer and COPY query for 10-column Method rows
- MCP tools: updated Cypher examples for HAS_METHOD, OVERRIDES, diamond

72 tests across 5 test files covering MRO resolution, HAS_METHOD edges,
method signature extraction, C# heritage resolution, and integration
tests across C#/Rust/Python/TS/Java/C++.

* feat: add scope-based symbol resolution replacing raw lookupFuzzy

Introduces a shared 3-tier resolveSymbol function used by both
heritage-processor and call-processor:
1. Same-file (lookupExactFull — authoritative)
2. Import-scoped (filtered by ImportMap — high confidence)
3. Global fuzzy (first match — low confidence fallback)

Adds lookupExactFull to SymbolTable returning full SymbolDefinition
with type info needed for heritage Class/Interface disambiguation.

* refactor: tighten symbol resolution — Tier 3 refuses ambiguous matches

- lookupExactFull now O(1) via direct SymbolDefinition storage in fileIndex
  (shared object references with globalIndex — zero additional memory)
- Added resolveSymbolInternal() preserving { definition, tier, candidateCount }
  for test assertions and logging
- Tier 3 now returns null when multiple global candidates exist instead of
  arbitrary allDefs[0] — a wrong edge is worse than no edge
- call-processor: renamed fuzzy-global → unique-global, removed dead branch
- 12 new tests: tier assertions, ambiguous refusal per language family,
  heritage false-positive guard, O(1) shared reference verification

* fix: critical language support bugs in import resolution and MRO

Phase 5 critical fixes from all-language analysis:
- Python: add relative_import query capture (PEP 328) — `.models`, `..utils`
  were silently dropped, producing zero ImportMap entries
- Rust: extract prefix from grouped imports `crate::module::{A, B}` — brace
  groups previously failed resolution entirely
- Swift: use normalizedFileList for Windows path compatibility in module
  import resolution (matches Go's resolveGoPackage pattern)
- MRO: fix c_sharp → csharp language name mismatch (enum is 'csharp'),
  add Kotlin to C#/Java resolution rules (class method wins over interface)

* feat: add strict multi-language integration tests + fix C/C++ import resolution

Add 32 integration tests across 6 language fixtures (TypeScript, C#, C++,
Java, Python, Rust) with exact toBe/toEqual assertions validating heritage
edges, import resolution, and trait implementations.

Fix C/C++ import resolution bug where dot-to-slash conversion mangled
include paths (e.g. "animal.h" became "animal/h"). Now skips conversion
for C/C++ languages which use actual file paths in #include directives.

* fix: language-gate heritage heuristic, add Swift extension heritage, handle Rust grouped imports

- Gate I[A-Z] naming heuristic to C#/Java only (was firing for all languages)
- Swift unresolved types default to IMPLEMENTS (protocol conformance is the norm)
- Add tree-sitter query for Swift extension protocol conformance (extension Foo: Protocol)
- Handle Rust top-level grouped imports (use {crate::a, crate::b}) in both import loops
- Add 4 new heritage-processor tests (TypeScript refusal, Swift default, Swift Tier 1)

* feat: add Go struct embedding heritage + PackageMap optimization

Add Go struct embedding detection (anonymous fields → EXTENDS edges) via
new tree-sitter heritage query with named-field filtering in both
parse-worker and heritage-processor paths.

Implement PackageMap optimization for Go cross-package resolution:
replace O(N) file-level ImportMap expansion with directory-level suffix
matching (Tier 2b in symbol resolver). Graph IMPORTS edges are preserved
via addImportGraphEdge split.

Remove overly broad @definition.type from GO_QUERIES that was
double-matching structs/interfaces as TypeAlias nodes, breaking Tier 3
unique-global resolution.

Add Go fixture (go-pkg) with Admin→User embedding, cross-package calls,
and 7 integration tests covering structs, functions, imports, calls,
and heritage edges.

* test: add Kotlin heritage integration tests

Adds a kotlin-heritage fixture and 7 integration tests validating
class inheritance, interface implementation, JVM-style import
resolution, and symbol-table-driven EXTENDS/IMPLEMENTS disambiguation
via Kotlin delegation specifiers.

* feat: extract resolvers, add PHP tests, ambiguous tests for all languages

- Extract language-specific resolvers from import-processor.ts into
  resolvers/ directory (P7): jvm, go, csharp, php, rust, standard, utils
- import-processor.ts reduced from 1412 to 711 lines (50% reduction)
- Add comprehensive PHP integration tests: PSR-4 imports, traits, enums,
  heritage edges, method calls, MRO overrides
- Add ambiguous symbol resolution tests for all 9 languages verifying
  correct disambiguation via import chains
- Split monolithic lang-resolution.test.ts (1080 lines) into 9 per-language
  files under test/integration/resolvers/ with shared helpers

* feat: update integration tests to include resolver tests for multiple languages

* fix: address code review — schema gap, Rust impl name, Property OVERRIDES

Bugs fixed:
- Add 13 missing FROM/TO pairs in RELATION_SCHEMA for HAS_METHOD edges
  (Class/Interface/Struct/Trait/Impl/Record to Method/Constructor/Property)
- Fix findEnclosingClassId to pick implementing type for Rust
  impl Trait for Struct blocks (was picking trait name)
- Exclude Property nodes from MRO OVERRIDES collision detection
- Change MRO language fallback from typescript to unknown

Tests added:
- Unit: Property OVERRIDES exclusion (2 tests), Rust impl Trait for
  Struct name resolution (2 tests), schema HAS_METHOD pair coverage
- Integration: no OVERRIDES targets Property nodes across all 9 languages
- PHP fixture: added shared $status property to both traits to create
  real collision scenario for Property OVERRIDES exclusion test

Documentation:
- OVERRIDES edge direction (Class to Method), Go return type gap,
  BFS first-reach heuristic limitation

* feat: harden CALLS-edge resolution — Phase 0 validation

- Fix same-file confidence (0.85 → 0.95) to correctly outrank import-scoped (0.9)
- Fix Tier 1 overload preservation: use globalIndex filter instead of fileIndex lookup
- Add callable-kind guard: refuse CALLS edges to Interface and Enum symbols
- Fix Kotlin countCallArguments: handle call_suffix → value_arguments nesting
- Fix Kotlin extractFunctionName: add simple_identifier to fallback search
- Strictly type findParameterList and countCallArguments (remove all `any`)
- Add arity-based call resolution integration tests for 9 languages
- Add unit regression tests for Interface/Enum CALLS refusal

* chore: remove C# build artifacts from fixtures

* feat: add call-form discrimination and ownerId to symbol table (Phase 1)

Add inferCallForm() and extractReceiverName() to distinguish free/member/constructor
calls at the AST level across all 9 languages. Add ownerId field to SymbolDefinition
linking Method/Constructor/Property to their owning class. Includes 36 unit tests
and member-call integration tests for all 9 languages (132 tests, 0 failures).

* feat: constructor/struct-literal resolution across all languages (Phase 2)

Add constructor discrimination to CALLS-edge resolution: new Foo(),
User{...} struct literals, and C# primary constructors now resolve to
Constructor/Class/Struct/Record nodes instead of being filtered out.

Queries: new_expression (C++), object_creation_expression (PHP),
composite_literal (Go), struct_expression (Rust), primary constructor
and implicit_object_creation_expression (C#).

Relaxes global tier in collectTieredCandidates to pass all candidates
through filterCallableCandidates, allowing kind/arity narrowing to
disambiguate at lower confidence.

* feat: receiver-constrained resolution with integration tests for all 9 languages

Add receiver-type filtering (Phase 3): when a member call like `user.save()`
has a known receiver type from TypeEnv, filter candidates by ownerId to
disambiguate methods with the same name across different classes.

Key changes:
- call-processor: build per-file TypeEnv, pass receiverTypeName to resolveCallTarget
- parse-worker: extract receiverTypeName from TypeEnv in worker thread
- resolveCallTarget: new step D filters by ownerId matching receiver type
- utils: extractReceiverName supports C++ field_expression (argument field)
- utils: findEnclosingClassId extracts Go method receiver types
- type-env: handle Go qualified_type, Kotlin user_type/variable_declaration
- parse-worker + parsing-processor: Function added to needsOwner for
  Kotlin/Rust/Python class methods captured as Function nodes

Integration tests added for receiver-constrained resolution across all 9
languages: TypeScript, Java, Python, Go, Rust, C++, C#, Kotlin, PHP.

* feat: NamedImportMap, scoped TypeEnv, broadened signatures + TS rest-param variadic fix

Address all 4 PR #238 review items:
1. Remove redundant lookupFuzzy in processRoutesFromExtracted
2. Add NamedImportMap for TS/Python symbol-level import tracking (Tier 2a)
3. Make TypeEnv scope-aware (Map<scopeKey, Map<varName, type>>) to fix
   non-deterministic receiver resolution across functions
4. Broaden extractMethodSignature: Go/Rust/C++ return types, variadic
   detection for Go/Java/Python/C++/Kotlin/TypeScript rest params

Discovered and fixed: TS rest params (...args) were not detected as
variadic — added rest_pattern detection inside required_parameter nodes.

Integration tests added: scoped receiver, named import disambiguation,
and variadic call resolution for both TypeScript and Python.

* fix: alias import resolution, Go multi-assign TypeEnv, dead code removal

- NamedImportMap now stores {sourcePath, exportedName} so aliased imports
  (import { User as U }) resolve U → User in the source file
- Named binding check moved before empty-allDefs early return in both
  call-processor and symbol-resolver, fixing constructor calls via aliases
- Go extractFromGoShortVarDeclaration iterates all LHS/RHS pairs for
  multi-assignment (user, repo := User{}, Repo{}) instead of only first
- Remove unused TYPED_DECLARATION_TYPES set (TYPED_PARAMETER_TYPES kept)
- Integration tests for both fixes (go-multi-assign, typescript-alias-imports)

* feat: alias import extraction for Kotlin, Rust, PHP, C# + integration tests

Add named import alias extraction to both pipeline paths
(import-processor.ts and parse-worker.ts) for Kotlin, Rust, PHP,
and C#. Add integration test fixtures and tests for all 5 languages
(Python alias extraction already worked, just needed the test).

Each test verifies: class detection, member call resolution through
aliases to correct target files, and IMPORTS edge emission.

* refactor: use SupportedLanguages enum everywhere instead of raw strings

Replace all raw language string literals and `language: string` types
with the SupportedLanguages enum across 10 files. This ensures
compile-time safety for language dispatch and eliminates dead
`language === 'tsx'` checks (tsx maps to TypeScript in the enum).

* fix: tier-ordering bug, re-export chains, PHP grouped imports, Java named imports

- Fix collectTieredCandidates tier-ordering: same-file now checked before
  named bindings, preventing imports from shadowing local definitions
  (matches resolveSymbolInternal priority order)
- Add re-export chain resolution for TypeScript/JavaScript barrel files:
  export { X } from './base' and export type { X } from './base' now
  followed up to 5 hops through NamedImportMap
- Fix PHP grouped import alias extraction: use App\Models\{User, Repo as R}
  now correctly handled in both parse-worker and import-processor
- Add Java NamedImportMap support: import com.example.models.User now
  records User as a named binding for precise disambiguation
- Add 16 new integration tests across TypeScript, PHP, and Java resolvers
  (220 total resolver tests, all passing)

* refactor: consolidate alias extraction + add variadic/constructor/shadow integration tests

- Extract shared named-binding-extraction.ts from duplicate logic in
  import-processor.ts and parse-worker.ts (net -200 lines)
- Deduplicate appendKotlinWildcard (now imported from resolvers/index.ts)
- Add integration tests: constructor calls (Kotlin, Python), variadic
  resolution (Go, Java, C#, C++, Kotlin), re-export chains (Python),
  local definition shadowing (Python, Go)
- Add TODO(stack-graph) for TypeEnv scope key collision
- 225 integration tests passing (was 223)

* fix: PHP non-aliased imports, Python node identity, re-export chain dedup + local-shadow tests

- PHP flat non-aliased imports (use App\Models\User) now stored in NamedImportMap
- PHP grouped non-aliased imports ({User} in {User, Repo as R}) now stored in NamedImportMap
- Python: replace non-public child.id with child.startIndex for node identity
- Extract shared walkBindingChain() from symbol-resolver and call-processor
- Add PHP variadic resolution fixture + test (variadic_parameter already covers PHP)
- Add local-shadow integration tests for Java, C#, Kotlin, Rust, PHP, C++ (6 languages)

* feat: Rust non-aliased use bindings, Kotlin non-aliased imports, re-export chain resolution

Extend NamedImportMap coverage for Rust and Kotlin non-aliased imports:

- Rust: rename collectUseAsClauses → collectRustBindings, extract terminal
  scoped_identifier (use crate::models::User) and identifier in use_list
  (use crate::models::{User, Repo}) into NamedImportMap. This also enables
  pub use re-export chain following via walkBindingChain.
- Kotlin: extend extractKotlinNamedBindings to handle non-aliased imports
  (import com.example.User), skipping wildcard imports.
- Add rust-reexport-chain fixture + 3 integration tests verifying Handler{}
  resolves through mod.rs pub use to handler.rs.
- Add Kotlin heritage + constructor-calls reason assertions for non-aliased
  import-resolved resolution.
- Add C# heritage test documenting namespace import tier behavior.

* fix: skip Kotlin lowercase member imports in NamedImportMap

Member imports like `import util.OneArg.writeAudit` (lowercase last
segment) must not populate NamedImportMap — same-named function imports
from different classes collide, breaking arity-based disambiguation.
Apply the same guard Java already uses: skip lowercase last segments.

* fix: skip spurious path-prefix bindings in Rust grouped imports

collectRustBindings was extracting the path segment (e.g. "models") from
`use crate::models::{User, Repo}` as a spurious NamedImportMap entry.
Skip scoped_identifier nodes that are direct children of scoped_use_list
since they are path prefixes, not importable symbols.

Adds rust-grouped-imports fixture and 4 integration tests verifying both
symbols resolve correctly and no spurious binding leaks through.

* fix: use startIndex in TypeEnv scope key to prevent same-name method collision

Two methods named identically in different classes within the same file
previously shared a scope key, causing non-deterministic type resolution.
Now keys use funcName@startIndex for uniqueness.

Also adds tests documenting destructuring assignment extraction gap.

* test: document C# namespace-level import limitation in named binding extraction

* test: document same-arity overload discrimination limitation in call processor

* perf: parallelize calls/heritage/routes processing in worker path

Worker path now runs processCallsFromExtracted, processHeritageFromExtracted,
and processRoutesFromExtracted via Promise.all instead of sequentially.
Safe because all three only read shared state and write via addRelationship's
dedup guard. Sequential fallback path stays sequential (shared LRU astCache).

Also fixes Rust collectRustBindings spurious path-prefix bindings for 3+ level
grouped imports, and adds @param JSDoc for walkBindingChain's allDefs invariant.

* docs: improve Promise.all safety comment and walkBindingChain JSDoc

Clarify that the parallelization safety comes from disjoint relationship
types + idempotent id-keyed Maps, not from lack of shared state (the
graph is shared). Strengthen allDefs JSDoc to describe silent-miss
consequence of passing pre-filtered results.

* refactor: extract language-specific processing into modular dispatch tables

Phase 1: Extract type binding logic from type-env.ts (635→125 LOC) into
type-extractors/ directory with per-language files and Record<SupportedLanguages,
LanguageTypeConfig> + satisfies dispatch.

Phase 2: Extract 5 config loaders from import-processor.ts into
language-config.ts (removed ~196 LOC of inline loaders).

Phase 3: Convert export-detection.ts switch/case to exhaustive
Record<SupportedLanguages, ExportChecker> + satisfies dispatch table,
fix node: any → SyntaxNode.

Also adds language feature matrix to README.

All 1146 unit tests and 433 integration tests pass.

* refactor: extract type binding logic into type-extractors/ directory (Phase 1)

Extract per-language type extraction from type-env.ts (635→125 LOC) into
type-extractors/ with Record<SupportedLanguages, LanguageTypeConfig> + satisfies
dispatch. 9 per-language files, shared helpers, and barrel index.

* refactor: extract config loaders to language-config.ts (Phase 2)

Move 5 language-specific config loaders and their type interfaces from
import-processor.ts into standalone language-config.ts module.
2026-03-13 13:12:23 +00:00
Zander Raycraft 03bfa3c4d9 FEAT: Added support for optional skill generation based on KuzuDB after initial repo analysis (npx gitnexus analyze --skills) (#171)
* calm fix 4 adding skills to repo [ISSUE #140]

* inspect

* unit and integration tests

* fixed hardcoded cohesion miss

* e2e tests for --skills flag for langauge/repo support

* Cohesion test e2e tests
2026-03-13 08:29:13 +00:00
Subham Kundu 74c0e462c3 Merge pull request #217 from JasonOA888/feat/issue-215-deepseek-model
feat(models): add DeepSeek model configurations
2026-03-13 00:31:02 +05:30
Gergő Magyar 7376e92063 fix: consolidate C/C++/C#/Rust language support from 6 overlapping PRs (#237)
* fix: consolidate C/C++/C#/Rust language support from 6 overlapping PRs

Merges fixes from PRs #163, #170, #178, #216, #227, #234 into a single
coherent changeset with shared modules and deduplication.

Phase 0 — Pre-merge consolidation:
- Extract isNodeExported to shared export-detection.ts module
- Extract TREE_SITTER_BUFFER_SIZE to shared constants.ts with adaptive sizing
- Consolidate FUNCTION_NODE_TYPES, extractFunctionName, isBuiltInOrNoise
  from duplicated call-processor.ts and parse-worker.ts into shared utils.ts
- Add query compilation smoke tests for all 12 languages

Language fixes:
- fix(c/cpp): isExported checks static linkage instead of returning false
- fix(c/cpp): .h files parsed as C++ (tree-sitter-cpp is superset of C)
- fix(c/cpp): expanded entry point patterns (~30 new for C, ~18 for C++)
- fix(cpp): add typedef, union, macro, prototype, inline method queries
- fix(c#): isExported scans sibling modifiers instead of parent walk
- fix(c#): heritage queries use correct base_list AST structure
- fix(c#): add framework detection, import resolution, entry point scoring
- fix(rust): isExported scans sibling visibility_modifier in declaration
- fix(builtins): remove open/read/write/close (real C POSIX syscalls)
- fix(buffer): adaptive bufferSize (2x fileSize, 512KB-32MB range)
- feat(ts/js): add call_expression query patterns for const assignments

Deduplication:
- call-processor.ts: -226 lines (uses shared utils)
- parse-worker.ts: -320 lines (uses shared utils)
- parsing-processor.ts: -156 lines (uses shared export-detection)

* perf: fix review findings — hoist Sets, deduplicate DEFINITION_CAPTURE_KEYS

- Hoist CSHARP_DECL_TYPES and RUST_DECL_TYPES to module-level constants
  in export-detection.ts (was allocating new Set on every isNodeExported call)
- Extract DEFINITION_CAPTURE_KEYS and getDefinitionNodeFromCaptures to
  shared utils.ts (was duplicated in parsing-processor.ts and parse-worker.ts)
- Pre-compute merged entry point patterns to avoid per-call array spread
  in calculateEntryPointScore

* test: add C, C++, and Tree-sitter buffer size tests

* fix: C/C++/Rust review findings + comprehensive test coverage (+72 tests)

Source fixes:
- Add Rust built-in noise (unwrap, clone, into, collect, panic, etc.)
- C++ anonymous namespace → internal linkage (not exported)
- Replace .text regex with storage_class_specifier child scan (perf)
- Raise file skip threshold from 512KB to 32MB (TREE_SITTER_MAX_BUFFER)
- Export TREE_SITTER_MAX_BUFFER from constants.ts
- Add C++ double pointer query patterns to CPP_QUERIES
- Add C#: record_struct, record_class, file_scoped_namespace to decl types
- Add Rust: union_item to visibility scanning set

Tests (214 → 286):
- ingestion-utils: +24 (Rust/C# noise, pointer/ref/destructor extraction, buffer)
- parsing: +36 (real AST C/C++ static/namespace, Rust/C#/Java/PHP/Swift edge cases)
- tree-sitter-languages: +12 (query accuracy for C/C++/C#/Rust captures)
2026-03-10 23:03:32 +00:00
RyanbaandGergo Magyar 1be910f54a fix: skip unavailable native Swift parsers in sequential ingestion (#188)
* fix: skip unavailable native Swift parsers in sequential ingestion

* fix: warn when ingestion skips languages in verbose mode

* test: cover verbose skip warnings

* docs: update analyze flags

* docs: clarify verbose default

---------

Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-03-10 14:26:36 +00:00
Gergo Magyar fa9ba8925c fix(ci): support fork PRs in Claude Code Review workflow
claude-code-action fetches branches by name from origin, which fails
for fork PRs since the branch only exists on the fork remote. Work
around by detecting fork PRs and temporarily pushing the branch to
origin before the action runs, then cleaning up afterwards.

Also changed trigger from automatic (every push) to on-demand only
(label "claude-review" or comment "@claude" / "/review").
2026-03-09 15:33:41 +00:00
8efc272609 fix(ci): move PR report to workflow_run for fork PR support (#225)
* Initial plan

* fix: add pull-requests write permissions to GitHub Actions workflows

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(ci): remove ineffective job-level permissions from reusable workflow

* fix(ci): pass PR write permission from caller to reusable unit-tests workflow

* fix(ci): harden CI/CD workflows with security fixes and reliability improvements

- Pin all actions to commit SHAs to prevent supply-chain attacks
- Fix shell injection in ci-integration.yml by using env vars instead of direct interpolation
- Scope permissions per-job in publish.yml (was granting pull-requests:write to publish job)
- Restrict claude-code-review to trusted contributors only (OWNER/MEMBER/COLLABORATOR)
- Switch claude-code-review to pull_request_target for fork PR support
- Fix fail-fast: false in ci-unit-tests.yml cross-platform matrix
- Remove duplicate ubuntu-latest from unit test matrix
- Add timeouts to all workflow jobs
- Improve kuzu-db test loop to continue on failure and report per-file errors

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ci): read thresholds from `vitest.config.ts`

* fix(ci): move PR report to workflow_run for fork PR support

The sticky-pull-request-comment and vitest-coverage-report-action
both fail on fork PRs because pull_request events receive a read-only
GITHUB_TOKEN. This extracts PR reporting into a separate ci-report.yml
workflow triggered by workflow_run, which always gets read/write tokens.

Changes:
- ci.yml: replace pr-report job with save-pr-meta artifact upload
- ci-unit-tests.yml: remove davelosert/vitest-coverage-report-action,
  add coverage-final.json to artifact for merging
- ci-integration.yml: add ubuntu coverage job for non-kuzu groups
- ci-report.yml (new): workflow_run handler that downloads artifacts,
  merges unit + integration coverage via Istanbul, and posts combined
  PR comment with sticky-pull-request-comment

* feat(ci): show unit, integration, and merged coverage in PR report

- Disable coverage thresholds for integration-only run (partial coverage)
- Display combined coverage as the primary metric
- Show per-suite breakdown (unit / integration) in expandable details
- Thresholds applied against combined coverage, not individual suites

* fix(ci): add coverage collection input for PR reports and validate job results

* fix(ci): refine Claude Code Review workflow to support issue comments and enhance trusted contributor checks

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 15:07:46 +00:00
c990d7e6c6 fix(ci): harden CI/CD workflows with security fixes and reliability improvements (#222)
* Initial plan

* fix: add pull-requests write permissions to GitHub Actions workflows

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(ci): remove ineffective job-level permissions from reusable workflow

* fix(ci): pass PR write permission from caller to reusable unit-tests workflow

* fix(ci): harden CI/CD workflows with security fixes and reliability improvements

- Pin all actions to commit SHAs to prevent supply-chain attacks
- Fix shell injection in ci-integration.yml by using env vars instead of direct interpolation
- Scope permissions per-job in publish.yml (was granting pull-requests:write to publish job)
- Restrict claude-code-review to trusted contributors only (OWNER/MEMBER/COLLABORATOR)
- Switch claude-code-review to pull_request_target for fork PR support
- Fix fail-fast: false in ci-unit-tests.yml cross-platform matrix
- Remove duplicate ubuntu-latest from unit test matrix
- Add timeouts to all workflow jobs
- Improve kuzu-db test loop to continue on failure and report per-file errors

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(ci): read thresholds from `vitest.config.ts`

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 08:25:03 +00:00
Gergo Magyar b4fbf33bd6 fix(ci): remove workflow-level permissions from reusable workflows 2026-03-08 18:20:05 +00:00
Gergő MagyarandClaude Opus 4.6 892e1d6088 test: add integration test coverage and fix KuzuDB fork crashes (#209)
* ci: add macOS to cross-platform test matrix

* ci: run integration tests on all platforms, add macOS to matrix

* ci: add build step before cross-platform integration tests

Worker pool requires compiled parse-worker.js in dist/.
Without build, falls back to sequential parsing which times out
on macOS runners.

* fix(pipeline): resolve worker path to dist/ when running under vitest

import.meta.url points to src/ under vitest where no .js exists.
Fall back to dist/core/ingestion/workers/parse-worker.js so worker
threads spawn correctly on all platforms instead of sequential fallback
that times out on slower macOS CI runners.

* ci: split cross-platform unit and integration tests into parallel jobs

* test: add integration tests for worker pool and hooks e2e

- worker-pool.test.ts: 7 tests verifying dist/ worker spawning,
  multi-file parsing, progress reporting, and clean termination
- hooks-e2e.test.ts: 28 tests with real git repos testing staleness
  detection, embeddings flag, mutation regex, cwd validation,
  and .gitnexus directory discovery

* refactor: extract shared hook test helpers and simplify worker fallback

- Extract runHook/parseHookOutput into test/utils/hook-test-helpers.ts
- Deduplicate fileURLToPath calls in pipeline.ts worker resolution
- Add isDev logging for worker pool creation failures

* fix(test): accept timeout as valid outcome for PreToolUse CLI spawn

The Plugin hook spawns `gitnexus augment` which may hang on macOS
when the CLI is unavailable, causing a 10s timeout (status=null)
instead of a clean exit (status=0). Accept both as non-crash outcomes.

* test: add integration test coverage and fix KuzuDB fork crashes

- Add new integration tests: search, enrichment, CLI e2e (968 total tests)
- Fix KuzuDB native destructor segfault in vitest fork pool by adding
  detachKuzu() that nulls refs without calling .close()
- Merge core adapter test blocks to share one coreHandle (prevents
  multiple coreInitKuzu calls that re-open native DB handles)
- Fix FTS Cypher injection: escape backslashes in bm25-index.ts and
  kuzu-adapter.ts queryFTS
- Add worker script existence check in worker-pool.ts to prevent
  MODULE_NOT_FOUND crashes in worker threads
- Add test/setup.ts global teardown that detaches native refs
- Add test/helpers/test-indexed-db.ts shared KuzuDB test lifecycle helper

* fix(test): update worker-pool test to expect throw on invalid path

The fs.existsSync validation in createWorkerPool now throws
synchronously for missing worker scripts. Update the test assertion
from .not.toThrow() to .toThrow(/Worker script not found/).

* fix(test): use fileParallelism instead of deprecated singleFork

vitest 4.x removed poolOptions.forks.singleFork. The top-level
singleFork was silently ignored, causing multiple forks to spawn
and timeout during KuzuDB native cleanup on CI.

* fix(test): add maxWorkers: 1 to prevent per-file kuzu native addon reload

On Ubuntu CI, vitest forks pool creates a new child process per test
file. Each fork loads the KuzuDB native addon (~40s on Ubuntu runners),
causing 12 files × 40s = 8 minutes of overhead that exceeds the
10-minute CI timeout.

maxWorkers: 1 forces vitest to reuse a single fork process, loading
the native addon once. Combined with fileParallelism: false, all test
files run sequentially in that single fork.

* fix(test): prevent KuzuDB native destructor hangs on fork worker exit

- setup.ts: closeKuzu() first (marks native handles closed so destructors
  are no-ops), then detachKuzu() as safety net
- test-indexed-db.ts: use detachKuzu() in per-test cleanup instead of
  closeKuzu() which could hang during teardown

* refactor(test): add withTestKuzuDB lifecycle wrapper with declarative options

withTestKuzuDB now manages the full KuzuDB test lifecycle so test files
never call initKuzu/closeCoreKuzu/poolInitKuzu/loadFTSExtension directly.

Options: seed, ftsIndexes, poolAdapter, afterSetup, timeout.
Each call is wrapped in its own describe block to isolate lifecycle hooks.

Migrated search.test.ts, enrichment-and-augmentation.test.ts, and
kuzu-pool.test.ts core adapter block to use the wrapper.

* refactor(test): migrate all integration tests to withTestKuzuDB

- Split enrichment-and-augmentation.test.ts into enrichment.test.ts
  and augmentation.test.ts for focused test isolation
- Migrate kuzu-pool.test.ts pool lifecycle tests to withTestKuzuDB
- Migrate local-backend.test.ts to two withTestKuzuDB blocks
  (pool queries + callTool dispatch)
- Zero direct kuzu.Database/Connection usage remains in test files

* refactor(test): enforce one describe per test file

- Split search.test.ts → search-core.test.ts + search-pool.test.ts
- Split kuzu-pool.test.ts → kuzu-pool.test.ts + kuzu-core-adapter.test.ts
- Split local-backend.test.ts → local-backend.test.ts + local-backend-calltool.test.ts
- Wrap enrichment.test.ts in single top-level describe
- Wrap parsing.test.ts in single top-level describe
- Every integration test file now has exactly 1 top-level block

* refactor(test): extract shared seed data into fixture files

- Create test/fixtures/search-seed.ts with SEARCH_SEED_DATA and SEARCH_FTS_INDEXES
- Create test/fixtures/local-backend-seed.ts with LOCAL_BACKEND_SEED_DATA and LOCAL_BACKEND_FTS_INDEXES
- Remove duplicated constants from split test files
- Remove dead vi.mock from local-backend.test.ts
- Prefix unused handle param with underscore in search-core.test.ts

* fix(test): prevent KuzuDB C++ destructor hang on Ubuntu CI

Add process.on('beforeExit', () => process.exit(0)) to force
immediate exit before GC can trigger native C++ destructors on
orphaned KuzuDB Database/Connection objects.

Root cause: detachKuzu() nulls JS refs but native C++ objects
remain in V8 heap. During fork worker exit, GC runs finalizers
that invoke C++ destructors on a torn-down runtime — hangs on
Ubuntu, segfaults on Windows.

The beforeExit event fires when the event loop has drained
(test results already sent via IPC), so process.exit(0) is safe.

Also simplifies afterAll: removes closeKuzu() calls (always
no-ops since withTestKuzuDB detaches first) — only detachKuzu().

* perf(test): share single KuzuDB instance across integration tests

Create schema once in globalSetup instead of per-file, eliminating
29 DDL queries × 7 test files. Each file now only clears and reseeds
data via DETACH DELETE, reducing DB open/close cycles significantly.

* fix(test): improve KuzuDB cleanup to prevent C++ destructor hangs on exit

* fix(test): replace async close calls with synchronous counterparts to prevent potential hangs

* feat(ci): enhance integration test matrix with detailed test groups and improved reporting

* test: add diagnostic output to analyze CLI e2e assertion for CI debugging

* fix: pass NODE_OPTIONS in runCli to prevent ensureHeap re-exec in tests

* update gitnexus analysis md files

* feat(ci): modular workflow architecture with artifact reporting

Refactor monolithic ci.yml into orchestrator calling three reusable
workflows (quality, unit-tests, integration) via workflow_call.

- Add composite action for shared Node.js 20 setup and npm ci
- Add ci-quality.yml for TypeScript typecheck
- Add ci-unit-tests.yml with coverage reporting, JSON test results,
  and artifact upload for PR summary comments
- Add ci-integration.yml with 4 test groups x 3 OS matrix (12 jobs)
- Add PR report job with sticky comment showing coverage metrics
- Add unified CI Gate status check for branch protection
- Add explicit permissions blocks to all child workflows

* test: add comprehensive unhappy path coverage across all 16 integration test files

Add 80+ error handling, edge case, and unhappy path tests covering:
- KuzuDB core adapter: invalid Cypher, duplicate FTS index, empty queries, missing paths
- CLI e2e: non-git dirs, non-indexed repos, unknown commands, help flag
- Local backend callTool: missing params, invalid Cypher, nonexistent symbols
- Tree-sitter: unsupported languages, malformed code, empty content, binary files
- Worker pool: dispatch after terminate, double terminate, empty content, zero-size pool
- Pipeline: empty content parsing, flexible file count assertions
- Search, enrichment, augmentation, CSV, hooks, filesystem: various edge cases

Also fixes pre-existing test issues:
- isWriteQuery CREATED test (CYPHER_WRITE_RE uses \b word boundaries)
- KuzuDB throws Binder exception for unknown tables (not empty result)
- runPipelineFromRepo requires onProgress callback

All 1,086 tests pass (53 files).

* fix: prevent KuzuDB worker hang with handle unref strategy and safety-net timer

Replace beforeExit force-exit with per-file handle unref + safety-net timer
that doesn't leak across files in single-fork mode.

* refactor: improve KuzuDB test isolation and cleanup strategy

* fix: prevent KuzuDB N-API destructor hang on Linux/macOS

Pool adapter closeOne() now just deletes the pool entry without calling
native close methods — read-only DBs have no WAL to flush, so GC/process
exit safely reclaims native resources without triggering the C++ destructor
segfault.

withTestKuzuDB wrapper handles core adapter close platform-conditionally:
Windows needs explicit closeKuzu() due to file locks, Linux/macOS skips
it to avoid deadlock. kuzu-pool.test.ts now uses poolAdapter: true instead
of manual afterSetup. pipeline.test.ts assertion fixed to match actual
behavior (resolves with empty result, not rejects).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: restore vitest safety nets and skip globalSetup close on Linux

- Restore dangerouslyIgnoreUnhandledErrors and teardownTimeout in
  vitest.config.ts — KuzuDB N-API destructor segfaults on fork exit
  are not real test failures (all 839 unit tests pass).
- Skip conn.close()/db.close() in globalSetup on Linux/macOS to
  prevent N-API destructor crash that kills the vitest process before
  fork workers can start (fixes search-core.test.ts EPIPE on Ubuntu CI).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: enable coverage auto-ratcheting with bumped thresholds

- Bump vitest coverage thresholds to match actual CI values (26/23/28/27)
- Enable thresholds.autoUpdate for automatic local ratcheting

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(ci): rich PR report with coverage bars, test counts, and threshold tracking

- Fix coverage N/A bug: use find instead of hardcoded artifact path
- Add emoji status icons and overall pass/fail banner
- Show covered/total counts alongside percentages
- Add visual progress bars with green/red threshold indicators
- Show test suite count and duration
- Add collapsible auto-ratchet explainer
- Graceful fallback when coverage data is unavailable

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore: bump version to 1.3.11, update CHANGELOG, add release.yml

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-08 18:00:45 +00:00
Jason c2bd8667a3 feat(models): add DeepSeek model configurations
Add configurations for DeepSeek-V3 and DeepSeek-Chat models
via OpenRouter integration.

- DeepSeek-V3: Reasoning model (input: /usr/bin/bash.27, output: .10)
- DeepSeek-Chat: Chat model (input: /usr/bin/bash.14, output: /usr/bin/bash.28)

Fixes #215
2026-03-08 15:50:25 +08:00
Gergő Magyar 1952c2c346 ci: add macOS to cross-platform test matrix (#208)
* ci: add macOS to cross-platform test matrix

* ci: run integration tests on all platforms, add macOS to matrix

* ci: add build step before cross-platform integration tests

Worker pool requires compiled parse-worker.js in dist/.
Without build, falls back to sequential parsing which times out
on macOS runners.

* fix(pipeline): resolve worker path to dist/ when running under vitest

import.meta.url points to src/ under vitest where no .js exists.
Fall back to dist/core/ingestion/workers/parse-worker.js so worker
threads spawn correctly on all platforms instead of sequential fallback
that times out on slower macOS CI runners.

* ci: split cross-platform unit and integration tests into parallel jobs

* test: add integration tests for worker pool and hooks e2e

- worker-pool.test.ts: 7 tests verifying dist/ worker spawning,
  multi-file parsing, progress reporting, and clean termination
- hooks-e2e.test.ts: 28 tests with real git repos testing staleness
  detection, embeddings flag, mutation regex, cwd validation,
  and .gitnexus directory discovery

* refactor: extract shared hook test helpers and simplify worker fallback

- Extract runHook/parseHookOutput into test/utils/hook-test-helpers.ts
- Deduplicate fileURLToPath calls in pipeline.ts worker resolution
- Add isDev logging for worker pool creation failures

* fix(test): accept timeout as valid outcome for PreToolUse CLI spawn

The Plugin hook spawns `gitnexus augment` which may hang on macOS
when the CLI is unavailable, causing a 10s timeout (status=null)
instead of a clean exit (status=0). Accept both as non-crash outcomes.
2026-03-07 10:38:32 +00:00
Linus Beckhaus c4eaf45ab1 feat(hooks): auto-reindex notification with cross-platform hardening (#205)
Adds PostToolUse hook that detects stale GitNexus index after git mutations (commit, merge, rebase, cherry-pick, pull) and notifies the agent to reindex. Uses lightweight staleness check (git rev-parse HEAD vs meta.json) instead of running gitnexus analyze synchronously, avoiding KuzuDB corruption and 120s blocks. Security and cross-platform hardening: remove shell:true from all spawnSync calls, use .cmd extensions on Windows, add path.isAbsolute(cwd) guards, fix setup.ts path escaping with JSON.stringify, use sendHookResponse() consistently. Includes 73 regression tests.
2026-03-07 08:59:54 +00:00
Gergo Magyar 0796e1e68c chore: bump version to 1.3.10 and add CHANGELOG
Add CHANGELOG.md with release notes for v1.3.10 covering MCP transport
security hardening, dual-framing compatibility, lazy CLI loading, and
bug fixes from recent PRs.
2026-03-07 08:04:55 +00:00
ShockangandGergo Magyar 9d5ec5d19a Improve MCP startup compatibility and lazy-load CLI commands (#207)
* Fix MCP startup transport compatibility

* Preserve CLI flags in MCP startup fix

* Harden MCP transport error handling

* Harden transport security and improve type safety

Transport hardening:
- Add MAX_BUFFER_SIZE (10 MB) cap to prevent OOM from oversized
  Content-Length or unbounded newline-delimited input
- Replace recursive readNewlineMessage with iterative loop to prevent
  stack overflow from consecutive empty lines
- Tighten looksLikeContentLength to require 14+ bytes before matching
- Add closed-state guard and error handling to send()
- Simplify processReadBuffer loop to break on error
- Fix loose equality (==) to strict (===)
- Widen constructor param types to ReadableStream/WritableStream

Type safety:
- Constrain createLazyAction generics so export name is validated
  against the module's actual exports at compile time
- Use proper type guard instead of lint suppression
- Fix test tsconfig type errors

Regression tests for all hardening fixes (13 tests passing).

---------

Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-03-07 07:47:09 +00:00
abhigyanpatwariandClaude Opus 4.6 4de40e4011 chore: update AI context files with inline imperative instructions
Regenerated CLAUDE.md and AGENTS.md using gitnexus@1.3.9 which replaces
the old skill-router format with inline imperative instructions (PR #190).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 15:12:00 +05:30
Gergő Magyar 20e8c52028 Merge pull request #144 from magyargergo/fix/lru-cache-zero-max-crash
fix: guard createASTCache against zero maxSize to prevent LRU cache crash
2026-03-06 09:22:17 +00:00
Gary Magyar 2868da5ddb Merge remote-tracking branch 'origin/main' into fix/lru-cache-zero-max-crash 2026-03-06 09:02:50 +00:00
Abhigyan PatwariandClaude Opus 4.6 3db47f7ee5 fix(ingestion): align CALLS edge sourceId with node ID format (#194)
findEnclosingFunctionId generated IDs without :startLine suffix,
but node creation includes it. This caused every CALLS edge to
reference a non-existent source node, making the process detector
find 0 entry points and produce 0 execution flows.

Bumps to 1.3.9.

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 13:57:20 +05:30
Gary Magyar 76e0e5a35a fix: guard createASTCache against zero maxSize to prevent LRU cache crash
When a repo has no parseable files (e.g., unsupported languages or all
files filtered out), chunks.reduce returns 0, causing createASTCache(0)
to pass max:0 to LRUCache which throws TypeError. This clamps maxSize
to at least 1 and adds a progress message when no parseable files exist.
2026-03-02 08:47:20 +00:00
1044 changed files with 53220 additions and 4700 deletions
@@ -22,7 +22,7 @@ Run from the project root. This parses all source files, builds the knowledge gr
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale.
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook runs `analyze` automatically after `git commit` and `git merge`, preserving embeddings if previously generated.
### status — Check index freshness
@@ -1,89 +1,89 @@
---
name: gitnexus-debugging
description: "Use when the user is debugging a bug, tracing an error, or asking why something fails. Examples: \"Why is X failing?\", \"Where does this error come from?\", \"Trace this bug\""
---
# Debugging with GitNexus
## When to Use
- "Why is this function failing?"
- "Trace where this error comes from"
- "Who calls this method?"
- "This endpoint returns 500"
- Investigating bugs, errors, or unexpected behavior
## Workflow
```
1. gitnexus_query({query: "<error or symptom>"}) → Find related execution flows
2. gitnexus_context({name: "<suspect>"}) → See callers/callees/processes
3. READ gitnexus://repo/{name}/process/{name} → Trace execution flow
4. gitnexus_cypher({query: "MATCH path..."}) → Custom traces if needed
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] Understand the symptom (error message, unexpected behavior)
- [ ] gitnexus_query for error text or related code
- [ ] Identify the suspect function from returned processes
- [ ] gitnexus_context to see callers and callees
- [ ] Trace execution flow via process resource if applicable
- [ ] gitnexus_cypher for custom call chain traces if needed
- [ ] Read source files to confirm root cause
```
## Debugging Patterns
| Symptom | GitNexus Approach |
| -------------------- | ---------------------------------------------------------- |
| Error message | `gitnexus_query` for error text → `context` on throw sites |
| Wrong return value | `context` on the function → trace callees for data flow |
| Intermittent failure | `context` → look for external calls, async deps |
| Performance issue | `context` → find symbols with many callers (hot paths) |
| Recent regression | `detect_changes` to see what your changes affect |
## Tools
**gitnexus_query** — find code related to error:
```
gitnexus_query({query: "payment validation error"})
→ Processes: CheckoutFlow, ErrorHandling
→ Symbols: validatePayment, handlePaymentError, PaymentException
```
**gitnexus_context** — full context for a suspect:
```
gitnexus_context({name: "validatePayment"})
→ Incoming calls: processCheckout, webhookHandler
→ Outgoing calls: verifyCard, fetchRates (external API!)
→ Processes: CheckoutFlow (step 3/7)
```
**gitnexus_cypher** — custom call chain traces:
```cypher
MATCH path = (a)-[:CodeRelation {type: 'CALLS'}*1..2]->(b:Function {name: "validatePayment"})
RETURN [n IN nodes(path) | n.name] AS chain
```
## Example: "Payment endpoint returns 500 intermittently"
```
1. gitnexus_query({query: "payment error handling"})
→ Processes: CheckoutFlow, ErrorHandling
→ Symbols: validatePayment, handlePaymentError
2. gitnexus_context({name: "validatePayment"})
→ Outgoing calls: verifyCard, fetchRates (external API!)
3. READ gitnexus://repo/my-app/process/CheckoutFlow
→ Step 3: validatePayment → calls fetchRates (external)
4. Root cause: fetchRates calls external API without proper timeout
```
---
name: gitnexus-debugging
description: "Use when the user is debugging a bug, tracing an error, or asking why something fails. Examples: \"Why is X failing?\", \"Where does this error come from?\", \"Trace this bug\""
---
# Debugging with GitNexus
## When to Use
- "Why is this function failing?"
- "Trace where this error comes from"
- "Who calls this method?"
- "This endpoint returns 500"
- Investigating bugs, errors, or unexpected behavior
## Workflow
```
1. gitnexus_query({query: "<error or symptom>"}) → Find related execution flows
2. gitnexus_context({name: "<suspect>"}) → See callers/callees/processes
3. READ gitnexus://repo/{name}/process/{name} → Trace execution flow
4. gitnexus_cypher({query: "MATCH path..."}) → Custom traces if needed
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] Understand the symptom (error message, unexpected behavior)
- [ ] gitnexus_query for error text or related code
- [ ] Identify the suspect function from returned processes
- [ ] gitnexus_context to see callers and callees
- [ ] Trace execution flow via process resource if applicable
- [ ] gitnexus_cypher for custom call chain traces if needed
- [ ] Read source files to confirm root cause
```
## Debugging Patterns
| Symptom | GitNexus Approach |
| -------------------- | ---------------------------------------------------------- |
| Error message | `gitnexus_query` for error text → `context` on throw sites |
| Wrong return value | `context` on the function → trace callees for data flow |
| Intermittent failure | `context` → look for external calls, async deps |
| Performance issue | `context` → find symbols with many callers (hot paths) |
| Recent regression | `detect_changes` to see what your changes affect |
## Tools
**gitnexus_query** — find code related to error:
```
gitnexus_query({query: "payment validation error"})
→ Processes: CheckoutFlow, ErrorHandling
→ Symbols: validatePayment, handlePaymentError, PaymentException
```
**gitnexus_context** — full context for a suspect:
```
gitnexus_context({name: "validatePayment"})
→ Incoming calls: processCheckout, webhookHandler
→ Outgoing calls: verifyCard, fetchRates (external API!)
→ Processes: CheckoutFlow (step 3/7)
```
**gitnexus_cypher** — custom call chain traces:
```cypher
MATCH path = (a)-[:CodeRelation {type: 'CALLS'}*1..2]->(b:Function {name: "validatePayment"})
RETURN [n IN nodes(path) | n.name] AS chain
```
## Example: "Payment endpoint returns 500 intermittently"
```
1. gitnexus_query({query: "payment error handling"})
→ Processes: CheckoutFlow, ErrorHandling
→ Symbols: validatePayment, handlePaymentError
2. gitnexus_context({name: "validatePayment"})
→ Outgoing calls: verifyCard, fetchRates (external API!)
3. READ gitnexus://repo/my-app/process/CheckoutFlow
→ Step 3: validatePayment → calls fetchRates (external)
4. Root cause: fetchRates calls external API without proper timeout
```
@@ -1,78 +1,78 @@
---
name: gitnexus-exploring
description: "Use when the user asks how code works, wants to understand architecture, trace execution flows, or explore unfamiliar parts of the codebase. Examples: \"How does X work?\", \"What calls this function?\", \"Show me the auth flow\""
---
# Exploring Codebases with GitNexus
## When to Use
- "How does authentication work?"
- "What's the project structure?"
- "Show me the main components"
- "Where is the database logic?"
- Understanding code you haven't seen before
## Workflow
```
1. READ gitnexus://repos → Discover indexed repos
2. READ gitnexus://repo/{name}/context → Codebase overview, check staleness
3. gitnexus_query({query: "<what you want to understand>"}) → Find related execution flows
4. gitnexus_context({name: "<symbol>"}) → Deep dive on specific symbol
5. READ gitnexus://repo/{name}/process/{name} → Trace full execution flow
```
> If step 2 says "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] READ gitnexus://repo/{name}/context
- [ ] gitnexus_query for the concept you want to understand
- [ ] Review returned processes (execution flows)
- [ ] gitnexus_context on key symbols for callers/callees
- [ ] READ process resource for full execution traces
- [ ] Read source files for implementation details
```
## Resources
| Resource | What you get |
| --------------------------------------- | ------------------------------------------------------- |
| `gitnexus://repo/{name}/context` | Stats, staleness warning (~150 tokens) |
| `gitnexus://repo/{name}/clusters` | All functional areas with cohesion scores (~300 tokens) |
| `gitnexus://repo/{name}/cluster/{name}` | Area members with file paths (~500 tokens) |
| `gitnexus://repo/{name}/process/{name}` | Step-by-step execution trace (~200 tokens) |
## Tools
**gitnexus_query** — find execution flows related to a concept:
```
gitnexus_query({query: "payment processing"})
→ Processes: CheckoutFlow, RefundFlow, WebhookHandler
→ Symbols grouped by flow with file locations
```
**gitnexus_context** — 360-degree view of a symbol:
```
gitnexus_context({name: "validateUser"})
→ Incoming calls: loginHandler, apiMiddleware
→ Outgoing calls: checkToken, getUserById
→ Processes: LoginFlow (step 2/5), TokenRefresh (step 1/3)
```
## Example: "How does payment processing work?"
```
1. READ gitnexus://repo/my-app/context → 918 symbols, 45 processes
2. gitnexus_query({query: "payment processing"})
→ CheckoutFlow: processPayment → validateCard → chargeStripe
→ RefundFlow: initiateRefund → calculateRefund → processRefund
3. gitnexus_context({name: "processPayment"})
→ Incoming: checkoutHandler, webhookHandler
→ Outgoing: validateCard, chargeStripe, saveTransaction
4. Read src/payments/processor.ts for implementation details
```
---
name: gitnexus-exploring
description: "Use when the user asks how code works, wants to understand architecture, trace execution flows, or explore unfamiliar parts of the codebase. Examples: \"How does X work?\", \"What calls this function?\", \"Show me the auth flow\""
---
# Exploring Codebases with GitNexus
## When to Use
- "How does authentication work?"
- "What's the project structure?"
- "Show me the main components"
- "Where is the database logic?"
- Understanding code you haven't seen before
## Workflow
```
1. READ gitnexus://repos → Discover indexed repos
2. READ gitnexus://repo/{name}/context → Codebase overview, check staleness
3. gitnexus_query({query: "<what you want to understand>"}) → Find related execution flows
4. gitnexus_context({name: "<symbol>"}) → Deep dive on specific symbol
5. READ gitnexus://repo/{name}/process/{name} → Trace full execution flow
```
> If step 2 says "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] READ gitnexus://repo/{name}/context
- [ ] gitnexus_query for the concept you want to understand
- [ ] Review returned processes (execution flows)
- [ ] gitnexus_context on key symbols for callers/callees
- [ ] READ process resource for full execution traces
- [ ] Read source files for implementation details
```
## Resources
| Resource | What you get |
| --------------------------------------- | ------------------------------------------------------- |
| `gitnexus://repo/{name}/context` | Stats, staleness warning (~150 tokens) |
| `gitnexus://repo/{name}/clusters` | All functional areas with cohesion scores (~300 tokens) |
| `gitnexus://repo/{name}/cluster/{name}` | Area members with file paths (~500 tokens) |
| `gitnexus://repo/{name}/process/{name}` | Step-by-step execution trace (~200 tokens) |
## Tools
**gitnexus_query** — find execution flows related to a concept:
```
gitnexus_query({query: "payment processing"})
→ Processes: CheckoutFlow, RefundFlow, WebhookHandler
→ Symbols grouped by flow with file locations
```
**gitnexus_context** — 360-degree view of a symbol:
```
gitnexus_context({name: "validateUser"})
→ Incoming calls: loginHandler, apiMiddleware
→ Outgoing calls: checkToken, getUserById
→ Processes: LoginFlow (step 2/5), TokenRefresh (step 1/3)
```
## Example: "How does payment processing work?"
```
1. READ gitnexus://repo/my-app/context → 918 symbols, 45 processes
2. gitnexus_query({query: "payment processing"})
→ CheckoutFlow: processPayment → validateCard → chargeStripe
→ RefundFlow: initiateRefund → calculateRefund → processRefund
3. gitnexus_context({name: "processPayment"})
→ Incoming: checkoutHandler, webhookHandler
→ Outgoing: validateCard, chargeStripe, saveTransaction
4. Read src/payments/processor.ts for implementation details
```
@@ -1,97 +1,97 @@
---
name: gitnexus-impact-analysis
description: "Use when the user wants to know what will break if they change something, or needs safety analysis before editing code. Examples: \"Is it safe to change X?\", \"What depends on this?\", \"What will break?\""
---
# Impact Analysis with GitNexus
## When to Use
- "Is it safe to change this function?"
- "What will break if I modify X?"
- "Show me the blast radius"
- "Who uses this code?"
- Before making non-trivial code changes
- Before committing — to understand what your changes affect
## Workflow
```
1. gitnexus_impact({target: "X", direction: "upstream"}) → What depends on this
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
3. gitnexus_detect_changes() → Map current git changes to affected flows
4. Assess risk and report to user
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] gitnexus_impact({target, direction: "upstream"}) to find dependents
- [ ] Review d=1 items first (these WILL BREAK)
- [ ] Check high-confidence (>0.8) dependencies
- [ ] READ processes to check affected execution flows
- [ ] gitnexus_detect_changes() for pre-commit check
- [ ] Assess risk level and report to user
```
## Understanding Output
| Depth | Risk Level | Meaning |
| ----- | ---------------- | ------------------------ |
| d=1 | **WILL BREAK** | Direct callers/importers |
| d=2 | LIKELY AFFECTED | Indirect dependencies |
| d=3 | MAY NEED TESTING | Transitive effects |
## Risk Assessment
| Affected | Risk |
| ------------------------------ | -------- |
| <5 symbols, few processes | LOW |
| 5-15 symbols, 2-5 processes | MEDIUM |
| >15 symbols or many processes | HIGH |
| Critical path (auth, payments) | CRITICAL |
## Tools
**gitnexus_impact** — the primary tool for symbol blast radius:
```
gitnexus_impact({
target: "validateUser",
direction: "upstream",
minConfidence: 0.8,
maxDepth: 3
})
→ d=1 (WILL BREAK):
- loginHandler (src/auth/login.ts:42) [CALLS, 100%]
- apiMiddleware (src/api/middleware.ts:15) [CALLS, 100%]
→ d=2 (LIKELY AFFECTED):
- authRouter (src/routes/auth.ts:22) [CALLS, 95%]
```
**gitnexus_detect_changes** — git-diff based impact analysis:
```
gitnexus_detect_changes({scope: "staged"})
→ Changed: 5 symbols in 3 files
→ Affected: LoginFlow, TokenRefresh, APIMiddlewarePipeline
→ Risk: MEDIUM
```
## Example: "What breaks if I change validateUser?"
```
1. gitnexus_impact({target: "validateUser", direction: "upstream"})
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
2. READ gitnexus://repo/my-app/processes
→ LoginFlow and TokenRefresh touch validateUser
3. Risk: 2 direct callers, 2 processes = MEDIUM
```
---
name: gitnexus-impact-analysis
description: "Use when the user wants to know what will break if they change something, or needs safety analysis before editing code. Examples: \"Is it safe to change X?\", \"What depends on this?\", \"What will break?\""
---
# Impact Analysis with GitNexus
## When to Use
- "Is it safe to change this function?"
- "What will break if I modify X?"
- "Show me the blast radius"
- "Who uses this code?"
- Before making non-trivial code changes
- Before committing — to understand what your changes affect
## Workflow
```
1. gitnexus_impact({target: "X", direction: "upstream"}) → What depends on this
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
3. gitnexus_detect_changes() → Map current git changes to affected flows
4. Assess risk and report to user
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] gitnexus_impact({target, direction: "upstream"}) to find dependents
- [ ] Review d=1 items first (these WILL BREAK)
- [ ] Check high-confidence (>0.8) dependencies
- [ ] READ processes to check affected execution flows
- [ ] gitnexus_detect_changes() for pre-commit check
- [ ] Assess risk level and report to user
```
## Understanding Output
| Depth | Risk Level | Meaning |
| ----- | ---------------- | ------------------------ |
| d=1 | **WILL BREAK** | Direct callers/importers |
| d=2 | LIKELY AFFECTED | Indirect dependencies |
| d=3 | MAY NEED TESTING | Transitive effects |
## Risk Assessment
| Affected | Risk |
| ------------------------------ | -------- |
| <5 symbols, few processes | LOW |
| 5-15 symbols, 2-5 processes | MEDIUM |
| >15 symbols or many processes | HIGH |
| Critical path (auth, payments) | CRITICAL |
## Tools
**gitnexus_impact** — the primary tool for symbol blast radius:
```
gitnexus_impact({
target: "validateUser",
direction: "upstream",
minConfidence: 0.8,
maxDepth: 3
})
→ d=1 (WILL BREAK):
- loginHandler (src/auth/login.ts:42) [CALLS, 100%]
- apiMiddleware (src/api/middleware.ts:15) [CALLS, 100%]
→ d=2 (LIKELY AFFECTED):
- authRouter (src/routes/auth.ts:22) [CALLS, 95%]
```
**gitnexus_detect_changes** — git-diff based impact analysis:
```
gitnexus_detect_changes({scope: "staged"})
→ Changed: 5 symbols in 3 files
→ Affected: LoginFlow, TokenRefresh, APIMiddlewarePipeline
→ Risk: MEDIUM
```
## Example: "What breaks if I change validateUser?"
```
1. gitnexus_impact({target: "validateUser", direction: "upstream"})
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
2. READ gitnexus://repo/my-app/processes
→ LoginFlow and TokenRefresh touch validateUser
3. Risk: 2 direct callers, 2 processes = MEDIUM
```
@@ -1,121 +1,121 @@
---
name: gitnexus-refactoring
description: "Use when the user wants to rename, extract, split, move, or restructure code safely. Examples: \"Rename this function\", \"Extract this into a module\", \"Refactor this class\", \"Move this to a separate file\""
---
# Refactoring with GitNexus
## When to Use
- "Rename this function safely"
- "Extract this into a module"
- "Split this service"
- "Move this to a new file"
- Any task involving renaming, extracting, splitting, or restructuring code
## Workflow
```
1. gitnexus_impact({target: "X", direction: "upstream"}) → Map all dependents
2. gitnexus_query({query: "X"}) → Find execution flows involving X
3. gitnexus_context({name: "X"}) → See all incoming/outgoing refs
4. Plan update order: interfaces → implementations → callers → tests
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklists
### Rename Symbol
```
- [ ] gitnexus_rename({symbol_name: "oldName", new_name: "newName", dry_run: true}) — preview all edits
- [ ] Review graph edits (high confidence) and ast_search edits (review carefully)
- [ ] If satisfied: gitnexus_rename({..., dry_run: false}) — apply edits
- [ ] gitnexus_detect_changes() — verify only expected files changed
- [ ] Run tests for affected processes
```
### Extract Module
```
- [ ] gitnexus_context({name: target}) — see all incoming/outgoing refs
- [ ] gitnexus_impact({target, direction: "upstream"}) — find all external callers
- [ ] Define new module interface
- [ ] Extract code, update imports
- [ ] gitnexus_detect_changes() — verify affected scope
- [ ] Run tests for affected processes
```
### Split Function/Service
```
- [ ] gitnexus_context({name: target}) — understand all callees
- [ ] Group callees by responsibility
- [ ] gitnexus_impact({target, direction: "upstream"}) — map callers to update
- [ ] Create new functions/services
- [ ] Update callers
- [ ] gitnexus_detect_changes() — verify affected scope
- [ ] Run tests for affected processes
```
## Tools
**gitnexus_rename** — automated multi-file rename:
```
gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
→ 12 edits across 8 files
→ 10 graph edits (high confidence), 2 ast_search edits (review)
→ Changes: [{file_path, edits: [{line, old_text, new_text, confidence}]}]
```
**gitnexus_impact** — map all dependents first:
```
gitnexus_impact({target: "validateUser", direction: "upstream"})
→ d=1: loginHandler, apiMiddleware, testUtils
→ Affected Processes: LoginFlow, TokenRefresh
```
**gitnexus_detect_changes** — verify your changes after refactoring:
```
gitnexus_detect_changes({scope: "all"})
→ Changed: 8 files, 12 symbols
→ Affected processes: LoginFlow, TokenRefresh
→ Risk: MEDIUM
```
**gitnexus_cypher** — custom reference queries:
```cypher
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(f:Function {name: "validateUser"})
RETURN caller.name, caller.filePath ORDER BY caller.filePath
```
## Risk Rules
| Risk Factor | Mitigation |
| ------------------- | ----------------------------------------- |
| Many callers (>5) | Use gitnexus_rename for automated updates |
| Cross-area refs | Use detect_changes after to verify scope |
| String/dynamic refs | gitnexus_query to find them |
| External/public API | Version and deprecate properly |
## Example: Rename `validateUser` to `authenticateUser`
```
1. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
→ 12 edits: 10 graph (safe), 2 ast_search (review)
→ Files: validator.ts, login.ts, middleware.ts, config.json...
2. Review ast_search edits (config.json: dynamic reference!)
3. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: false})
→ Applied 12 edits across 8 files
4. gitnexus_detect_changes({scope: "all"})
→ Affected: LoginFlow, TokenRefresh
→ Risk: MEDIUM — run tests for these flows
```
---
name: gitnexus-refactoring
description: "Use when the user wants to rename, extract, split, move, or restructure code safely. Examples: \"Rename this function\", \"Extract this into a module\", \"Refactor this class\", \"Move this to a separate file\""
---
# Refactoring with GitNexus
## When to Use
- "Rename this function safely"
- "Extract this into a module"
- "Split this service"
- "Move this to a new file"
- Any task involving renaming, extracting, splitting, or restructuring code
## Workflow
```
1. gitnexus_impact({target: "X", direction: "upstream"}) → Map all dependents
2. gitnexus_query({query: "X"}) → Find execution flows involving X
3. gitnexus_context({name: "X"}) → See all incoming/outgoing refs
4. Plan update order: interfaces → implementations → callers → tests
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklists
### Rename Symbol
```
- [ ] gitnexus_rename({symbol_name: "oldName", new_name: "newName", dry_run: true}) — preview all edits
- [ ] Review graph edits (high confidence) and ast_search edits (review carefully)
- [ ] If satisfied: gitnexus_rename({..., dry_run: false}) — apply edits
- [ ] gitnexus_detect_changes() — verify only expected files changed
- [ ] Run tests for affected processes
```
### Extract Module
```
- [ ] gitnexus_context({name: target}) — see all incoming/outgoing refs
- [ ] gitnexus_impact({target, direction: "upstream"}) — find all external callers
- [ ] Define new module interface
- [ ] Extract code, update imports
- [ ] gitnexus_detect_changes() — verify affected scope
- [ ] Run tests for affected processes
```
### Split Function/Service
```
- [ ] gitnexus_context({name: target}) — understand all callees
- [ ] Group callees by responsibility
- [ ] gitnexus_impact({target, direction: "upstream"}) — map callers to update
- [ ] Create new functions/services
- [ ] Update callers
- [ ] gitnexus_detect_changes() — verify affected scope
- [ ] Run tests for affected processes
```
## Tools
**gitnexus_rename** — automated multi-file rename:
```
gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
→ 12 edits across 8 files
→ 10 graph edits (high confidence), 2 ast_search edits (review)
→ Changes: [{file_path, edits: [{line, old_text, new_text, confidence}]}]
```
**gitnexus_impact** — map all dependents first:
```
gitnexus_impact({target: "validateUser", direction: "upstream"})
→ d=1: loginHandler, apiMiddleware, testUtils
→ Affected Processes: LoginFlow, TokenRefresh
```
**gitnexus_detect_changes** — verify your changes after refactoring:
```
gitnexus_detect_changes({scope: "all"})
→ Changed: 8 files, 12 symbols
→ Affected processes: LoginFlow, TokenRefresh
→ Risk: MEDIUM
```
**gitnexus_cypher** — custom reference queries:
```cypher
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(f:Function {name: "validateUser"})
RETURN caller.name, caller.filePath ORDER BY caller.filePath
```
## Risk Rules
| Risk Factor | Mitigation |
| ------------------- | ----------------------------------------- |
| Many callers (>5) | Use gitnexus_rename for automated updates |
| Cross-area refs | Use detect_changes after to verify scope |
| String/dynamic refs | gitnexus_query to find them |
| External/public API | Version and deprecate properly |
## Example: Rename `validateUser` to `authenticateUser`
```
1. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
→ 12 edits: 10 graph (safe), 2 ast_search (review)
→ Files: validator.ts, login.ts, middleware.ts, config.json...
2. Review ast_search edits (config.json: dynamic reference!)
3. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: false})
→ Applied 12 edits across 8 files
4. gitnexus_detect_changes({scope: "all"})
→ Affected: LoginFlow, TokenRefresh
→ Risk: MEDIUM — run tests for these flows
```
+28
View File
@@ -0,0 +1,28 @@
name: Setup GitNexus
description: Setup Node.js 20, install dependencies, and optionally build
inputs:
build:
description: Whether to run npm run build after install
required: false
default: 'false'
runs:
using: composite
steps:
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus/package-lock.json
- name: Install dependencies
run: npm ci
shell: bash
working-directory: gitnexus
- name: Build
if: ${{ inputs.build == 'true' }}
run: npm run build
shell: bash
working-directory: gitnexus
+45
View File
@@ -0,0 +1,45 @@
changelog:
exclude:
labels:
- chore
authors:
- dependabot
- dependabot[bot]
categories:
- title: "\U0001F6A8 Security"
labels:
- security
- title: "\U0001F4A5 Breaking Changes"
labels:
- breaking
- title: "\U0001F680 Features"
labels:
- enhancement
- title: "\U0001F41B Bug Fixes"
labels:
- bug
- title: "\U0001F3CE\uFE0F Performance"
labels:
- performance
- title: "\U0001F9EA Tests"
labels:
- test
- title: "\U0001F504 Refactoring"
labels:
- refactor
- title: "\U0001F477 CI/CD"
labels:
- ci
- title: "\U0001F4E6 Dependencies"
labels:
- dependencies
- title: "\U0001F4DD Other Changes"
labels:
- "*"
exclude:
labels:
- dependencies
- ci
- test
- refactor
- chore
+198
View File
@@ -0,0 +1,198 @@
name: Integration Tests
on:
workflow_call:
inputs:
collect-coverage:
description: 'Whether to run the coverage collection job (only needed for PR reports)'
required: false
default: true
type: boolean
jobs:
# ── Integration test matrix ─────────────────────────────────────────
# Each test-group runs on a SEPARATE runner per OS, giving full process
# isolation for the LadybugDB native C++ addon.
# 3 OS x 4 groups = 12 parallel jobs.
#
# Groups:
# lbug-db — 7 files using withTestLbugDB / lbug-adapter (native addon)
# Each file runs as its own `vitest run` invocation for full
# process isolation. LadybugDB's native N-API addon registers
# persistent handles that prevent fork workers from exiting
# on Linux, and its C++ destructors segfault during
# process.exit(). Running each file in its own process lets
# the OS reclaim all resources cleanly.
# pipeline — 12 files: ingestion pipeline + csv + 9 resolver tests
# e2e — 4 files: child-process only (spawnSync), no in-process lbug
# standalone — 4 files: pure logic, no lbug, no child processes
test-matrix:
name: integration (${{ matrix.os }} / ${{ matrix.test-group }})
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, windows-latest, macos-latest]
test-group: [lbug-db, pipeline, e2e, standalone]
include:
- test-group: lbug-db
# Marker — actual files are listed in the run step below
test-glob: ''
- test-group: pipeline
test-glob: >-
test/integration/pipeline.test.ts
test/integration/csv-pipeline.test.ts
test/integration/parsing.test.ts
test/integration/resolvers/typescript.test.ts
test/integration/resolvers/csharp.test.ts
test/integration/resolvers/cpp.test.ts
test/integration/resolvers/java.test.ts
test/integration/resolvers/python.test.ts
test/integration/resolvers/rust.test.ts
test/integration/resolvers/go.test.ts
test/integration/resolvers/kotlin.test.ts
test/integration/resolvers/php.test.ts
test/integration/resolvers/ruby.test.ts
test/integration/resolvers/swift.test.ts
- test-group: e2e
test-glob: >-
test/integration/cli-e2e.test.ts
test/integration/hooks-e2e.test.ts
test/integration/skills-e2e.test.ts
test/integration/ignore-and-skip-e2e.test.ts
- test-group: standalone
test-glob: >-
test/integration/filesystem-walker.test.ts
test/integration/enrichment.test.ts
test/integration/tree-sitter-languages.test.ts
test/integration/worker-pool.test.ts
runs-on: ${{ matrix.os }}
timeout-minutes: 25
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
with:
build: 'true'
# lbug-db: run each file in its own vitest process for full isolation.
# LadybugDB's native addon hangs fork workers on Linux — process isolation
# is the only reliable fix boundary.
- name: Run integration tests — lbug-db (process-isolated)
if: matrix.test-group == 'lbug-db'
working-directory: gitnexus
shell: bash
run: |
set -e
files=(
test/integration/lbug-core-adapter.test.ts
test/integration/lbug-pool.test.ts
test/integration/local-backend.test.ts
test/integration/local-backend-calltool.test.ts
test/integration/search-core.test.ts
test/integration/search-pool.test.ts
test/integration/augmentation.test.ts
)
exit_code=0
for f in "${files[@]}"; do
echo "::group::$f"
if ! npx vitest run --reporter=verbose --pool=forks "$f"; then
exit_code=1
echo "::error::Test file failed: $f"
fi
echo "::endgroup::"
done
exit $exit_code
# Non-lbug groups: run all files in a single vitest invocation
- name: Run integration tests — ${{ matrix.test-group }}
if: matrix.test-group != 'lbug-db'
shell: bash
env:
TEST_GLOB: ${{ matrix.test-glob }}
run: npx vitest run --reporter=verbose $TEST_GLOB
working-directory: gitnexus
# ── Coverage collection (ubuntu only) ─────────────────────────────────
# Runs non-lbug integration tests with coverage enabled so the PR report
# can merge integration + unit coverage for a combined view.
# lbug-db tests are excluded because each file must run in its own vitest
# process (native addon isolation) which prevents single-run coverage merge.
coverage:
name: integration (ubuntu / coverage)
if: inputs.collect-coverage
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
with:
build: 'true'
- name: Run integration tests with coverage
working-directory: gitnexus
run: >-
npx vitest run
--reporter=default
--reporter=json
--outputFile=integration-results.json
--coverage
--coverage.reporter=json-summary
--coverage.reporter=json
--coverage.reporter=text
--coverage.thresholdAutoUpdate=false
--coverage.reportOnFailure=true
--coverage.thresholds.statements=0
--coverage.thresholds.branches=0
--coverage.thresholds.functions=0
--coverage.thresholds.lines=0
test/integration/pipeline.test.ts
test/integration/csv-pipeline.test.ts
test/integration/parsing.test.ts
test/integration/cli-e2e.test.ts
test/integration/hooks-e2e.test.ts
test/integration/filesystem-walker.test.ts
test/integration/enrichment.test.ts
test/integration/tree-sitter-languages.test.ts
test/integration/worker-pool.test.ts
test/integration/ignore-and-skip-e2e.test.ts
test/integration/resolvers/typescript.test.ts
test/integration/resolvers/csharp.test.ts
test/integration/resolvers/cpp.test.ts
test/integration/resolvers/java.test.ts
test/integration/resolvers/python.test.ts
test/integration/resolvers/rust.test.ts
test/integration/resolvers/go.test.ts
test/integration/resolvers/kotlin.test.ts
test/integration/resolvers/php.test.ts
test/integration/resolvers/ruby.test.ts
test/integration/resolvers/swift.test.ts
- name: Upload integration coverage
if: always()
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
name: integration-reports
path: |
gitnexus/coverage/coverage-summary.json
gitnexus/coverage/coverage-final.json
gitnexus/integration-results.json
retention-days: 5
# ── Unified status gate ──────────────────────────────────────────────
# Branch protection should require THIS job, not the matrix jobs directly.
# ci.yml's needs.integration.result aggregates through this gate.
status:
name: integration (all groups)
needs: test-matrix
if: always()
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- name: Check all matrix jobs passed
shell: bash
env:
RESULT: ${{ needs.test-matrix.result }}
run: |
if [[ "$RESULT" != "success" ]]; then
echo "::error::Integration matrix failed or cancelled: $RESULT"
exit 1
fi
+14
View File
@@ -0,0 +1,14 @@
name: Quality Checks
on:
workflow_call:
jobs:
typecheck:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
- run: npx tsc --noEmit
working-directory: gitnexus
+522
View File
@@ -0,0 +1,522 @@
name: CI Report
# Triggered after the CI workflow completes. Because workflow_run
# always runs code from the *default branch*, it receives a read/write
# GITHUB_TOKEN — even when the triggering PR comes from a fork.
on:
workflow_run:
workflows: ["CI"]
types: [completed]
permissions:
actions: read # needed to list/download workflow run artifacts
contents: read # needed for sparse checkout of vitest.config.ts
pull-requests: write # needed to post sticky PR comment
jobs:
pr-report:
name: PR Report
# Only run for pull-request CI runs
if: >-
github.event.workflow_run.event == 'pull_request' &&
github.event.workflow_run.conclusion != 'cancelled'
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
# ── Download artifacts from the CI run ────────────────────────
- name: Download artifacts
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
with:
script: |
const fs = require('fs');
const path = require('path');
const runId = context.payload.workflow_run.id;
const allArtifacts = await github.rest.actions.listWorkflowRunArtifacts({
owner: context.repo.owner,
repo: context.repo.repo,
run_id: runId,
});
async function downloadArtifact(name, dest) {
const match = allArtifacts.data.artifacts.find(a => a.name === name);
if (!match) {
core.warning(`Artifact "${name}" not found`);
return false;
}
const zip = await github.rest.actions.downloadArtifact({
owner: context.repo.owner,
repo: context.repo.repo,
artifact_id: match.id,
archive_format: 'zip',
});
fs.mkdirSync(dest, { recursive: true });
fs.writeFileSync(path.join(dest, `${name}.zip`), Buffer.from(zip.data));
return true;
}
const temp = process.env.RUNNER_TEMP;
await downloadArtifact('pr-meta', path.join(temp, 'dl'));
await downloadArtifact('test-reports', path.join(temp, 'dl'));
await downloadArtifact('integration-reports', path.join(temp, 'dl'));
- name: Extract artifacts
shell: bash
run: |
cd "$RUNNER_TEMP/dl"
# Extract each artifact into its own directory to avoid filename collisions
for z in *.zip; do
[ -f "$z" ] || continue
name="${z%.zip}"
mkdir -p "$RUNNER_TEMP/artifacts/$name"
unzip -o "$z" -d "$RUNNER_TEMP/artifacts/$name"
done
- name: Read PR metadata
id: meta
shell: bash
run: |
DIR="$RUNNER_TEMP/artifacts/pr-meta"
if [ ! -f "$DIR/pr_number" ]; then
echo "skip=true" >> "$GITHUB_OUTPUT"
echo "::warning::pr_number artifact missing — skipping report"
exit 0
fi
# Validate PR number is a positive integer (artifact comes from
# untrusted fork code, so treat contents defensively).
PR_NUM=$(cat "$DIR/pr_number" | tr -d '[:space:]')
if ! [[ "$PR_NUM" =~ ^[0-9]+$ ]]; then
echo "skip=true" >> "$GITHUB_OUTPUT"
echo "::error::Invalid PR number in artifact: '$PR_NUM'"
exit 0
fi
echo "skip=false" >> "$GITHUB_OUTPUT"
echo "pr_number=$PR_NUM" >> "$GITHUB_OUTPUT"
# Validate job-result strings against known GitHub Actions values.
# Artifact contents come from the PR workflow (potentially untrusted
# fork code), so we whitelist to prevent newline injection into
# GITHUB_OUTPUT.
validate_result() {
local val
val=$(cat "$1" | tr -d '[:space:]')
case "$val" in
success|failure|cancelled|skipped) echo "$val" ;;
*) echo "unknown" ;;
esac
}
echo "quality=$(validate_result "$DIR/quality_result")" >> "$GITHUB_OUTPUT"
echo "unit=$(validate_result "$DIR/unit_result")" >> "$GITHUB_OUTPUT"
echo "integration=$(validate_result "$DIR/integration_result")" >> "$GITHUB_OUTPUT"
- name: Checkout (for vitest config)
if: steps.meta.outputs.skip != 'true'
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
with:
sparse-checkout: gitnexus/vitest.config.ts
sparse-checkout-cone-mode: false
# ── Fetch base branch coverage for delta reporting ───────────
- name: Fetch base branch coverage
if: steps.meta.outputs.skip != 'true'
id: base-coverage
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
with:
script: |
const fs = require('fs');
const path = require('path');
// Find the latest successful CI run on main
const runs = await github.rest.actions.listWorkflowRuns({
owner: context.repo.owner,
repo: context.repo.repo,
workflow_id: 'ci.yml',
branch: 'main',
status: 'success',
per_page: 1,
});
if (runs.data.workflow_runs.length === 0) {
core.setOutput('found', 'false');
core.info('No successful main branch CI runs found');
return;
}
const mainRunId = runs.data.workflow_runs[0].id;
const artifacts = await github.rest.actions.listWorkflowRunArtifacts({
owner: context.repo.owner,
repo: context.repo.repo,
run_id: mainRunId,
});
const testReports = artifacts.data.artifacts.find(a => a.name === 'test-reports');
if (!testReports) {
core.setOutput('found', 'false');
core.info('No test-reports artifact on main branch');
return;
}
const zip = await github.rest.actions.downloadArtifact({
owner: context.repo.owner,
repo: context.repo.repo,
artifact_id: testReports.id,
archive_format: 'zip',
});
const dest = path.join(process.env.RUNNER_TEMP, 'base-coverage');
fs.mkdirSync(dest, { recursive: true });
fs.writeFileSync(path.join(dest, 'base.zip'), Buffer.from(zip.data));
core.setOutput('found', 'true');
core.setOutput('dir', dest);
- name: Extract base coverage
if: steps.meta.outputs.skip != 'true' && steps.base-coverage.outputs.found == 'true'
shell: bash
run: |
cd "${{ steps.base-coverage.outputs.dir }}"
unzip -o base.zip -d base
# ── Merge coverage from unit + integration ─────────────────────
- name: Setup Node.js
if: steps.meta.outputs.skip != 'true'
uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
- name: Install coverage merge tools
if: steps.meta.outputs.skip != 'true'
run: npm install --no-save istanbul-lib-coverage istanbul-lib-report istanbul-reports
- name: Merge coverage reports
if: steps.meta.outputs.skip != 'true'
id: coverage
shell: bash
run: |
DIR="$RUNNER_TEMP/artifacts"
UNIT_COV=$(find "$DIR/test-reports" -name "coverage-final.json" -type f 2>/dev/null | head -1)
INTEG_COV=$(find "$DIR/integration-reports" -name "coverage-final.json" -type f 2>/dev/null | head -1)
MERGED_DIR="$RUNNER_TEMP/merged-coverage"
mkdir -p "$MERGED_DIR"
if [ -n "$UNIT_COV" ] && [ -n "$INTEG_COV" ]; then
echo "has_merged=true" >> "$GITHUB_OUTPUT"
# Merge using Node.js + istanbul-lib-coverage.
# Paths are passed via env vars to avoid shell interpolation
# inside the script string.
UNIT_COV_PATH="$UNIT_COV" \
INTEG_COV_PATH="$INTEG_COV" \
MERGED_OUT_DIR="$MERGED_DIR" \
node -e "
const libCoverage = require('istanbul-lib-coverage');
const libReport = require('istanbul-lib-report');
const reports = require('istanbul-reports');
const fs = require('fs');
const map = libCoverage.createCoverageMap({});
map.merge(JSON.parse(fs.readFileSync(process.env.UNIT_COV_PATH, 'utf8')));
map.merge(JSON.parse(fs.readFileSync(process.env.INTEG_COV_PATH, 'utf8')));
const context = libReport.createContext({
coverageMap: map,
dir: process.env.MERGED_OUT_DIR,
});
reports.create('json-summary').execute(context);
console.log('Merged coverage written to ' + process.env.MERGED_OUT_DIR + '/coverage-summary.json');
"
elif [ -n "$UNIT_COV" ]; then
echo "has_merged=false" >> "$GITHUB_OUTPUT"
echo "::warning::Integration coverage not found — using unit coverage only"
else
echo "has_merged=false" >> "$GITHUB_OUTPUT"
echo "::warning::No coverage data found"
fi
- name: Build report
if: steps.meta.outputs.skip != 'true'
id: report
shell: bash
env:
QUALITY: ${{ steps.meta.outputs.quality }}
UNIT: ${{ steps.meta.outputs.unit }}
INTEG: ${{ steps.meta.outputs.integration }}
HAS_MERGED: ${{ steps.coverage.outputs.has_merged }}
BASE_FOUND: ${{ steps.base-coverage.outputs.found }}
BASE_DIR: ${{ steps.base-coverage.outputs.dir }}
RUN_URL: ${{ github.event.workflow_run.html_url }}
run: |
DIR="$RUNNER_TEMP/artifacts"
MERGED_DIR="$RUNNER_TEMP/merged-coverage"
# ── Helper: read coverage summary into prefixed vars ──
read_cov() {
local prefix=$1 file=$2
if [ -n "$file" ] && [ -f "$file" ]; then
local val
val=$(jq -r '.total.statements.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_STMTS" '%s' "$val"
val=$(jq -r '.total.branches.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_BRANCH" '%s' "$val"
val=$(jq -r '.total.functions.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_FUNCS" '%s' "$val"
val=$(jq -r '.total.lines.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_LINES" '%s' "$val"
val=$(jq -r '"\(.total.statements.covered)/\(.total.statements.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_STMTS_COV" '%s' "$val"
val=$(jq -r '"\(.total.branches.covered)/\(.total.branches.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_BRANCH_COV" '%s' "$val"
val=$(jq -r '"\(.total.functions.covered)/\(.total.functions.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_FUNCS_COV" '%s' "$val"
val=$(jq -r '"\(.total.lines.covered)/\(.total.lines.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_LINES_COV" '%s' "$val"
return 0
else
printf -v "${prefix}_STMTS" '%s' "N/A"
printf -v "${prefix}_BRANCH" '%s' "N/A"
printf -v "${prefix}_FUNCS" '%s' "N/A"
printf -v "${prefix}_LINES" '%s' "N/A"
printf -v "${prefix}_STMTS_COV" '%s' ""
printf -v "${prefix}_BRANCH_COV" '%s' ""
printf -v "${prefix}_FUNCS_COV" '%s' ""
printf -v "${prefix}_LINES_COV" '%s' ""
return 1
fi
}
# ── Read all coverage reports ──
UNIT_SUMMARY=$(find "$DIR/test-reports" -name "coverage-summary.json" -type f 2>/dev/null | head -1)
INTEG_SUMMARY=$(find "$DIR/integration-reports" -name "coverage-summary.json" -type f 2>/dev/null | head -1)
MERGED_SUMMARY="$MERGED_DIR/coverage-summary.json"
read_cov "U" "$UNIT_SUMMARY"
HAS_UNIT=$?
read_cov "I" "$INTEG_SUMMARY"
HAS_INTEG=$?
read_cov "M" "$MERGED_SUMMARY"
# ── Read base branch coverage (main) ──
BASE_SUMMARY=""
if [ "$BASE_FOUND" = "true" ] && [ -n "$BASE_DIR" ]; then
BASE_SUMMARY=$(find "$BASE_DIR/base" -name "coverage-summary.json" -type f 2>/dev/null | head -1)
fi
read_cov "B" "$BASE_SUMMARY"
# ── Locate test results ──
RESULTS_FILE=$(find "$DIR/test-reports" -name "test-results.json" -type f 2>/dev/null | head -1)
INTEG_RESULTS=$(find "$DIR/integration-reports" -name "integration-results.json" -type f 2>/dev/null | head -1)
if [ -n "$RESULTS_FILE" ]; then
U_TOTAL=$(jq -r '.numTotalTests' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_PASSED=$(jq -r '.numPassedTests' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_FAILED=$(jq -r '.numFailedTests' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_SKIPPED=$(jq -r '.numPendingTests' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_SUITES=$(jq -r '.numTotalTestSuites' "$RESULTS_FILE" 2>/dev/null || echo 0)
U_DURATION=$(jq -r '((.testResults | map(.endTime) | max) - (.startTime)) / 1000 | floor' "$RESULTS_FILE" 2>/dev/null || echo 0)
else
U_TOTAL=0; U_PASSED=0; U_FAILED=0; U_SKIPPED=0; U_SUITES=0; U_DURATION=0
fi
if [ -n "$INTEG_RESULTS" ]; then
I_TOTAL=$(jq -r '.numTotalTests' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_PASSED=$(jq -r '.numPassedTests' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_FAILED=$(jq -r '.numFailedTests' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_SKIPPED=$(jq -r '.numPendingTests' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_SUITES=$(jq -r '.numTotalTestSuites' "$INTEG_RESULTS" 2>/dev/null || echo 0)
I_DURATION=$(jq -r '((.testResults | map(.endTime) | max) - (.startTime)) / 1000 | floor' "$INTEG_RESULTS" 2>/dev/null || echo 0)
else
I_TOTAL=0; I_PASSED=0; I_FAILED=0; I_SKIPPED=0; I_SUITES=0; I_DURATION=0
fi
# ── Sum test results ──
TOTAL=$((U_TOTAL + I_TOTAL))
PASSED=$((U_PASSED + I_PASSED))
FAILED=$((U_FAILED + I_FAILED))
SKIPPED=$((U_SKIPPED + I_SKIPPED))
SUITES=$((U_SUITES + I_SUITES))
# ── Status helpers ──
status_icon() {
case "$1" in
success) echo "✅" ;;
failure) echo "❌" ;;
cancelled) echo "⏭️" ;;
*) echo "❓" ;;
esac
}
cov_delta() {
local pct=$1 base=$2
if [ "$pct" = "N/A" ] || [ "$base" = "N/A" ]; then echo "—"; return; fi
local diff
diff=$(awk "BEGIN { printf \"%.1f\", $pct - $base }")
if [ "$(awk "BEGIN { print ($pct > $base) ? 1 : 0 }")" = "1" ]; then
echo "📈 +${diff}"
elif [ "$(awk "BEGIN { print ($pct < $base) ? 1 : 0 }")" = "1" ]; then
echo "📉 ${diff}"
else
echo "= ${diff}"
fi
}
cov_bar() {
local pct=$1 base=$2
if [ "$pct" = "N/A" ]; then echo "—"; return; fi
local filled
filled=$(awk "BEGIN { printf \"%d\", $pct / 5 }")
(( filled < 0 )) && filled=0
(( filled > 20 )) && filled=20
local empty=$((20 - filled))
local bar=""
for ((i=0; i<filled; i++)); do bar+="█"; done
for ((i=0; i<empty; i++)); do bar+="░"; done
# Green if >= base (or base unavailable), red if dropped
if [ "$base" = "N/A" ] || [ "$(awk "BEGIN { print ($pct >= $base) ? 1 : 0 }")" = "1" ]; then
echo "🟢 ${bar}"
else
echo "🔴 ${bar}"
fi
}
# ── Overall status ──
if [[ "$QUALITY" == "success" && "$UNIT" == "success" && "$INTEG" == "success" ]]; then
OVERALL="✅ **All checks passed**"
else
OVERALL="❌ **Some checks failed**"
fi
# ── Build markdown ──
{
echo "body<<GITNEXUS_CI_REPORT_EOF_7f3a"
echo "## CI Report"
echo ""
echo "${OVERALL}"
echo ""
echo "### Pipeline Status"
echo ""
echo "| Stage | Status | Details |"
echo "|-------|--------|---------|"
echo "| $(status_icon "$QUALITY") Typecheck | \`${QUALITY}\` | tsc --noEmit |"
echo "| $(status_icon "$UNIT") Unit Tests | \`${UNIT}\` | 3 platforms |"
echo "| $(status_icon "$INTEG") Integration | \`${INTEG}\` | 3 OS x 4 groups = 12 jobs |"
echo ""
if [ "$TOTAL" -gt 0 ] 2>/dev/null; then
echo "### Test Results"
echo ""
echo "| Suite | Tests | Passed | Failed | Skipped | Duration |"
echo "|-------|-------|--------|--------|---------|----------|"
if [ "$U_TOTAL" -gt 0 ] 2>/dev/null; then
echo "| Unit | ${U_TOTAL} | ${U_PASSED} | ${U_FAILED} | ${U_SKIPPED} | ${U_DURATION}s |"
fi
if [ "$I_TOTAL" -gt 0 ] 2>/dev/null; then
echo "| Integration | ${I_TOTAL} | ${I_PASSED} | ${I_FAILED} | ${I_SKIPPED} | ${I_DURATION}s |"
fi
echo "| **Total** | **${TOTAL}** | **${PASSED}** | **${FAILED}** | **${SKIPPED}** | **$((U_DURATION + I_DURATION))s** |"
echo ""
if [ "$FAILED" = "0" ]; then
echo "✅ All **${PASSED}** tests passed"
else
echo "❌ **${FAILED}** failed / **${PASSED}** passed"
fi
if [ "$SKIPPED" != "0" ]; then
echo ""
echo "<details>"
echo "<summary>${SKIPPED} test(s) skipped — expand for details</summary>"
echo ""
# Extract skipped test names from integration results
if [ -n "$INTEG_RESULTS" ] && [ "$I_SKIPPED" -gt 0 ] 2>/dev/null; then
echo "**Integration:**"
jq -r '
.testResults[]
| .assertionResults[]?
| select(.status == "pending" or .status == "skipped")
| "- \(.ancestorTitles | join(" > ")) > \(.title)"
' "$INTEG_RESULTS" 2>/dev/null || echo "- _(unable to parse skipped test details)_"
fi
# Extract skipped test names from unit results
if [ -n "$RESULTS_FILE" ] && [ "$U_SKIPPED" -gt 0 ] 2>/dev/null; then
echo ""
echo "**Unit:**"
jq -r '
.testResults[]
| .assertionResults[]?
| select(.status == "pending" or .status == "skipped")
| "- \(.ancestorTitles | join(" > ")) > \(.title)"
' "$RESULTS_FILE" 2>/dev/null || echo "- _(unable to parse skipped test details)_"
fi
echo ""
echo "</details>"
fi
echo ""
fi
# ── Coverage table helper ──
cov_table() {
local label=$1 s=$2 b=$3 f=$4 l=$5 sc=$6 bc=$7 fc=$8 lc=$9
shift 9
local bs=$1 bb=$2 bf=$3 bl=$4
echo "#### ${label}"
echo ""
echo "| Metric | Coverage | Covered | Base | Delta | Status |"
echo "|--------|----------|---------|------|-------|--------|"
echo "| Statements | **${s}%** | ${sc} | ${bs}% | $(cov_delta "$s" "$bs") | $(cov_bar "$s" "$bs") |"
echo "| Branches | **${b}%** | ${bc} | ${bb}% | $(cov_delta "$b" "$bb") | $(cov_bar "$b" "$bb") |"
echo "| Functions | **${f}%** | ${fc} | ${bf}% | $(cov_delta "$f" "$bf") | $(cov_bar "$f" "$bf") |"
echo "| Lines | **${l}%** | ${lc} | ${bl}% | $(cov_delta "$l" "$bl") | $(cov_bar "$l" "$bl") |"
echo ""
}
if [ "$M_STMTS" != "N/A" ]; then
echo "### Code Coverage"
echo ""
cov_table "Combined (Unit + Integration)" \
"$M_STMTS" "$M_BRANCH" "$M_FUNCS" "$M_LINES" \
"$M_STMTS_COV" "$M_BRANCH_COV" "$M_FUNCS_COV" "$M_LINES_COV" \
"$B_STMTS" "$B_BRANCH" "$B_FUNCS" "$B_LINES"
echo "<details>"
echo "<summary>Coverage breakdown by test suite</summary>"
echo ""
if [ "$U_STMTS" != "N/A" ]; then
cov_table "Unit Tests" \
"$U_STMTS" "$U_BRANCH" "$U_FUNCS" "$U_LINES" \
"$U_STMTS_COV" "$U_BRANCH_COV" "$U_FUNCS_COV" "$U_LINES_COV" \
"$B_STMTS" "$B_BRANCH" "$B_FUNCS" "$B_LINES"
fi
if [ "$I_STMTS" != "N/A" ]; then
cov_table "Integration Tests" \
"$I_STMTS" "$I_BRANCH" "$I_FUNCS" "$I_LINES" \
"$I_STMTS_COV" "$I_BRANCH_COV" "$I_FUNCS_COV" "$I_LINES_COV" \
"$B_STMTS" "$B_BRANCH" "$B_FUNCS" "$B_LINES"
fi
echo "</details>"
echo ""
elif [ "$U_STMTS" != "N/A" ]; then
echo "### Code Coverage (Unit only)"
echo ""
cov_table "Unit Tests" \
"$U_STMTS" "$U_BRANCH" "$U_FUNCS" "$U_LINES" \
"$U_STMTS_COV" "$U_BRANCH_COV" "$U_FUNCS_COV" "$U_LINES_COV" \
"$B_STMTS" "$B_BRANCH" "$B_FUNCS" "$B_LINES"
else
echo "### Code Coverage"
echo ""
echo "⚠️ Coverage data unavailable - check the [unit test job](${RUN_URL}) for details."
echo ""
fi
echo "---"
echo "<sub>📋 [View full run](${RUN_URL}) · Generated by CI</sub>"
echo "GITNEXUS_CI_REPORT_EOF_7f3a"
} >> "$GITHUB_OUTPUT"
- name: Comment on PR
if: steps.meta.outputs.skip != 'true'
uses: marocchino/sticky-pull-request-comment@773744901bac0e8cbb5a0dc842800d45e9b2b405 # v2
with:
header: ci-report
number: ${{ steps.meta.outputs.pr_number }}
message: ${{ steps.report.outputs.body }}
+53
View File
@@ -0,0 +1,53 @@
name: Unit Tests
on:
workflow_call:
jobs:
unit-tests:
name: unit (ubuntu / coverage)
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
- name: Run unit tests with coverage
run: >-
npx vitest run test/unit
--reporter=default
--reporter=json
--outputFile=test-results.json
--coverage
--coverage.reporter=json-summary
--coverage.reporter=json
--coverage.reporter=text
--coverage.thresholdAutoUpdate=false
--coverage.reportOnFailure=true
working-directory: gitnexus
- name: Upload test reports
if: always()
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
name: test-reports
path: |
gitnexus/coverage/coverage-summary.json
gitnexus/coverage/coverage-final.json
gitnexus/test-results.json
retention-days: 5
cross-platform:
name: unit (${{ matrix.os }})
strategy:
fail-fast: false
matrix:
# Ubuntu already covered by the coverage job above
os: [windows-latest, macos-latest]
runs-on: ${{ matrix.os }}
timeout-minutes: 15
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: ./.github/actions/setup-gitnexus
- run: npx vitest run test/unit
working-directory: gitnexus
+83 -53
View File
@@ -3,67 +3,97 @@ name: CI
on:
push:
branches: [main]
paths-ignore: ['**.md', 'docs/**', 'LICENSE']
pull_request:
branches: [main]
paths-ignore: ['**.md', 'docs/**', 'LICENSE']
workflow_call:
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
# ── Reusable workflow orchestration ─────────────────────────────────
# Each concern lives in its own workflow file for maintainability:
# ci-quality.yml — typecheck (tsc --noEmit)
# ci-unit-tests.yml — unit tests with coverage + cross-platform
# ci-integration.yml — integration test matrix (3 OS x 4 groups)
#
# Shared setup is DRY via .github/actions/setup-gitnexus composite action.
jobs:
typecheck:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus/package-lock.json
- run: npm ci
working-directory: gitnexus
- run: npx tsc --noEmit
working-directory: gitnexus
quality:
uses: ./.github/workflows/ci-quality.yml
permissions:
contents: read
unit-tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus/package-lock.json
- run: npm ci
working-directory: gitnexus
- run: npx vitest run test/unit --coverage --coverage.thresholdAutoUpdate=false
working-directory: gitnexus
uses: ./.github/workflows/ci-unit-tests.yml
permissions:
contents: read
integration-tests:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus/package-lock.json
- run: npm ci
working-directory: gitnexus
- run: npx vitest run test/integration
working-directory: gitnexus
integration:
uses: ./.github/workflows/ci-integration.yml
with:
collect-coverage: ${{ github.event_name == 'pull_request' }}
permissions:
contents: read
cross-platform:
strategy:
matrix:
os: [ubuntu-latest, windows-latest]
runs-on: ${{ matrix.os }}
# ── Save PR metadata for the reporting workflow ─────────────────
# The ci-report.yml workflow (triggered by workflow_run) needs the
# PR number and job results to post a comment. We save them as an
# artifact because workflow_run context doesn't reliably carry PR
# info for fork PRs.
save-pr-meta:
name: Save PR Metadata
if: always() && github.event_name == 'pull_request'
needs: [quality, unit-tests, integration]
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
- name: Write metadata
shell: bash
env:
PR_NUMBER: ${{ github.event.number }}
QUALITY: ${{ needs.quality.result }}
UNIT: ${{ needs.unit-tests.result }}
INTEG: ${{ needs.integration.result }}
run: |
mkdir -p pr-meta
echo "$PR_NUMBER" > pr-meta/pr_number
echo "$QUALITY" > pr-meta/quality_result
echo "$UNIT" > pr-meta/unit_result
echo "$INTEG" > pr-meta/integration_result
- name: Upload PR metadata
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus/package-lock.json
- run: npm ci
working-directory: gitnexus
- run: npx vitest run test/unit
working-directory: gitnexus
name: pr-meta
path: pr-meta/
retention-days: 1
# ── Unified CI gate ──────────────────────────────────────────────
# Single required check for branch protection.
ci-status:
name: CI Gate
needs: [quality, unit-tests, integration]
if: always()
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- name: Check all jobs passed
shell: bash
env:
QUALITY: ${{ needs.quality.result }}
UNIT: ${{ needs.unit-tests.result }}
INTEG: ${{ needs.integration.result }}
run: |
echo "Quality: $QUALITY"
echo "Unit Tests: $UNIT"
echo "Integration: $INTEG"
if [[ "$QUALITY" != "success" ]] ||
[[ "$UNIT" != "success" ]] ||
[[ "$INTEG" != "success" ]]; then
echo "::error::One or more CI jobs failed"
exit 1
fi
+99 -22
View File
@@ -1,44 +1,121 @@
name: Claude Code Review
# Uses pull_request_target so the workflow runs as defined on the default branch,
# which allows access to secrets for posting review comments on fork PRs.
# SECURITY: The checkout below uses the PR head SHA to review the correct code.
# The claude-code-action sandboxes execution — it does NOT run arbitrary code
# from the checked-out source.
on:
pull_request:
types: [opened, synchronize, ready_for_review, reopened]
# Optional: Only run on specific file changes
# paths:
# - "src/**/*.ts"
# - "src/**/*.tsx"
# - "src/**/*.js"
# - "src/**/*.jsx"
# Trigger only when explicitly requested:
# - Add the "claude-review" label to a PR, OR
# - Comment "@claude" or "/review" on a PR
pull_request_target:
types: [labeled]
issue_comment:
types: [created]
# Serialize per-PR so concurrent @claude comments don't race on the
# temporary fork branch push/delete.
concurrency:
group: claude-review-${{ github.event.issue.number || github.event.pull_request.number }}
cancel-in-progress: false
jobs:
claude-review:
# Optional: Filter by PR author
# if: |
# github.event.pull_request.user.login == 'external-contributor' ||
# github.event.pull_request.user.login == 'new-developer' ||
# github.event.pull_request.author_association == 'FIRST_TIME_CONTRIBUTOR'
# Run only when:
# 1. The "claude-review" label is added to a non-draft PR by a trusted contributor, OR
# 2. A trusted contributor comments "@claude" or "/review" on a PR
if: |
(
github.event_name == 'pull_request_target' &&
github.event.label.name == 'claude-review' &&
github.event.pull_request.draft == false &&
(github.event.pull_request.author_association == 'OWNER' ||
github.event.pull_request.author_association == 'MEMBER' ||
github.event.pull_request.author_association == 'COLLABORATOR')
) ||
(
github.event_name == 'issue_comment' &&
github.event.issue.pull_request &&
(contains(github.event.comment.body, '@claude') ||
contains(github.event.comment.body, '/review')) &&
(github.event.comment.author_association == 'OWNER' ||
github.event.comment.author_association == 'MEMBER' ||
github.event.comment.author_association == 'COLLABORATOR')
)
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
contents: read
pull-requests: read
contents: write # needed to create fork branch ref via API
pull-requests: write
issues: read
id-token: write
steps:
- name: Checkout repository
uses: actions/checkout@v4
# For issue_comment triggers, resolve the PR number, head SHA, and branch name
- name: Resolve PR context
id: pr
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
with:
script: |
let pr;
if (context.eventName === 'issue_comment') {
const resp = await github.rest.pulls.get({
owner: context.repo.owner,
repo: context.repo.repo,
pull_number: context.payload.issue.number,
});
pr = resp.data;
} else {
pr = context.payload.pull_request;
}
core.setOutput('number', pr.number);
core.setOutput('sha', pr.head.sha);
core.setOutput('branch', pr.head.ref);
core.setOutput('is_fork', String(pr.head.repo.full_name !== pr.base.repo.full_name));
- name: Checkout PR head
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
with:
ref: ${{ steps.pr.outputs.sha }}
fetch-depth: 1
# claude-code-action fetches branches by name from origin, which fails
# for fork PRs. Create a temporary branch ref via the API so the action
# can find it. Using the API (not git push) avoids the GITHUB_TOKEN
# restriction that blocks pushing commits containing workflow file changes.
- name: Create fork branch ref on origin
id: push-fork
if: steps.pr.outputs.is_fork == 'true'
env:
FORK_BRANCH: ${{ steps.pr.outputs.branch }}
FORK_SHA: ${{ steps.pr.outputs.sha }}
GH_TOKEN: ${{ github.token }}
run: |
gh api "repos/${{ github.repository }}/git/refs" \
--method POST \
-f ref="refs/heads/$FORK_BRANCH" \
-f sha="$FORK_SHA" \
|| gh api "repos/${{ github.repository }}/git/refs/heads/$FORK_BRANCH" \
--method PATCH \
-f sha="$FORK_SHA" \
-F force=true
- name: Run Claude Code Review
id: claude-review
uses: anthropics/claude-code-action@v1
uses: anthropics/claude-code-action@9469d113c6afd29550c402740f22d1a97dd1209b # v1
with:
claude_code_oauth_token: ${{ secrets.CLAUDE_CODE_OAUTH_TOKEN }}
plugin_marketplaces: 'https://github.com/anthropics/claude-code.git'
plugins: 'code-review@claude-code-plugins'
prompt: '/code-review:code-review ${{ github.repository }}/pull/${{ github.event.pull_request.number }}'
# See https://github.com/anthropics/claude-code-action/blob/main/docs/usage.md
# or https://code.claude.com/docs/en/cli-reference for available options
prompt: '/code-review:code-review ${{ github.repository }}/pull/${{ steps.pr.outputs.number }}'
# Clean up the temporary branch ref we created for fork PRs.
# Only delete if the create step actually succeeded.
- name: Delete fork branch ref from origin
if: always() && steps.push-fork.outcome == 'success'
env:
FORK_BRANCH: ${{ steps.pr.outputs.branch }}
GH_TOKEN: ${{ github.token }}
run: gh api "repos/${{ github.repository }}/git/refs/heads/$FORK_BRANCH" --method DELETE || true
+80 -15
View File
@@ -10,6 +10,12 @@ on:
pull_request_review:
types: [submitted]
# Serialize per-PR so concurrent @claude comments don't race on the
# temporary fork branch push/delete.
concurrency:
group: claude-code-${{ github.event.issue.number || github.event.pull_request.number || github.event.issue.id }}
cancel-in-progress: false
jobs:
claude:
if: |
@@ -18,21 +24,80 @@ jobs:
(github.event_name == 'pull_request_review' && contains(github.event.review.body, '@claude')) ||
(github.event_name == 'issues' && (contains(github.event.issue.body, '@claude') || contains(github.event.issue.title, '@claude')))
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
contents: read
pull-requests: read
issues: read
contents: write # needed to create fork branch ref via API
pull-requests: write
issues: write
id-token: write
actions: read # Required for Claude to read CI results on PRs
actions: read # required for Claude to read CI results on PRs
steps:
- name: Checkout repository
uses: actions/checkout@v4
# For PR-related triggers, resolve fork context so we can create a
# temporary branch ref (claude-code-action fetches by branch name).
- name: Resolve PR context
id: pr
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
with:
script: |
// Determine if this event is PR-related
let prNumber = null;
if (context.eventName === 'issue_comment' && context.payload.issue.pull_request) {
prNumber = context.payload.issue.number;
} else if (context.eventName === 'pull_request_review_comment') {
prNumber = context.payload.pull_request.number;
} else if (context.eventName === 'pull_request_review') {
prNumber = context.payload.pull_request.number;
}
if (!prNumber) {
core.setOutput('is_pr', 'false');
core.setOutput('is_fork', 'false');
return;
}
const resp = await github.rest.pulls.get({
owner: context.repo.owner,
repo: context.repo.repo,
pull_number: prNumber,
});
const pr = resp.data;
const isFork = pr.head.repo.full_name !== pr.base.repo.full_name;
core.setOutput('is_pr', 'true');
core.setOutput('is_fork', String(isFork));
core.setOutput('branch', pr.head.ref);
core.setOutput('sha', pr.head.sha);
- name: Checkout repository
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
with:
ref: ${{ steps.pr.outputs.is_fork == 'true' && steps.pr.outputs.sha || '' }}
fetch-depth: 1
# claude-code-action fetches branches by name from origin, which fails
# for fork PRs. Create a temporary branch ref via the API so the action
# can find it. Using the API (not git push) avoids the GITHUB_TOKEN
# restriction that blocks pushing commits containing workflow file changes.
- name: Create fork branch ref on origin
id: push-fork
if: steps.pr.outputs.is_fork == 'true'
env:
FORK_BRANCH: ${{ steps.pr.outputs.branch }}
FORK_SHA: ${{ steps.pr.outputs.sha }}
GH_TOKEN: ${{ github.token }}
run: |
gh api "repos/${{ github.repository }}/git/refs" \
--method POST \
-f ref="refs/heads/$FORK_BRANCH" \
-f sha="$FORK_SHA" \
|| gh api "repos/${{ github.repository }}/git/refs/heads/$FORK_BRANCH" \
--method PATCH \
-f sha="$FORK_SHA" \
-F force=true
- name: Run Claude Code
id: claude
uses: anthropics/claude-code-action@v1
uses: anthropics/claude-code-action@9469d113c6afd29550c402740f22d1a97dd1209b # v1
with:
claude_code_oauth_token: ${{ secrets.CLAUDE_CODE_OAUTH_TOKEN }}
@@ -40,11 +105,11 @@ jobs:
additional_permissions: |
actions: read
# Optional: Give a custom prompt to Claude. If this is not specified, Claude will perform the instructions specified in the comment that tagged it.
# prompt: 'Update the pull request description to include a summary of changes.'
# Optional: Add claude_args to customize behavior and configuration
# See https://github.com/anthropics/claude-code-action/blob/main/docs/usage.md
# or https://code.claude.com/docs/en/cli-reference for available options
# claude_args: '--allowed-tools Bash(gh pr:*)'
# Clean up the temporary branch ref we created for fork PRs.
# Only delete if the create step actually succeeded.
- name: Delete fork branch ref from origin
if: always() && steps.push-fork.outcome == 'success'
env:
FORK_BRANCH: ${{ steps.pr.outputs.branch }}
GH_TOKEN: ${{ github.token }}
run: gh api "repos/${{ github.repository }}/git/refs/heads/$FORK_BRANCH" --method DELETE || true
+14 -3
View File
@@ -5,19 +5,25 @@ on:
tags:
- 'v*'
# No workflow-level permissions — scoped per job below.
jobs:
ci:
uses: ./.github/workflows/ci.yml
permissions:
contents: read
pull-requests: write
publish:
needs: ci
runs-on: ubuntu-latest
timeout-minutes: 15
permissions:
contents: write
id-token: write
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
registry-url: https://registry.npmjs.org
@@ -27,8 +33,13 @@ jobs:
working-directory: gitnexus
- name: Verify version consistency
shell: bash
run: |
TAG_VERSION="${GITHUB_REF#refs/tags/v}"
if ! [[ "$TAG_VERSION" =~ ^[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$ ]]; then
echo "::error::Tag does not follow semver: v$TAG_VERSION"
exit 1
fi
PKG_VERSION=$(node -p "require('./package.json').version")
if [ "$TAG_VERSION" != "$PKG_VERSION" ]; then
echo "::error::Tag version (v$TAG_VERSION) does not match package.json version ($PKG_VERSION)"
@@ -52,6 +63,6 @@ jobs:
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Create GitHub Release
uses: softprops/action-gh-release@v2
uses: softprops/action-gh-release@a06a81a03ee405af7f2048a818ed3f03bbf83c7b # v2
with:
generate_release_notes: true
+12
View File
@@ -48,6 +48,9 @@ coverage/
# Claude Code worktrees
.claude/worktrees/
# Claude code skills
.claude/skills/generated/
# Assets (screenshots, images)
assets/
@@ -56,3 +59,12 @@ repomix-output*
# Design docs (local only)
docs/plans/
gitnexus/test/fixtures/mini-repo/*.md
gitnexus/test/fixtures/mini-repo/.claude
gitnexus/test/fixtures/mini-repo/.gitignore
# Ignore csharp generated obj and bin folders
gitnexus/test/fixtures/lang-resolution/**/obj
gitnexus/test/fixtures/lang-resolution/**/bin
GitNexus.sln
+84 -8
View File
@@ -1,17 +1,93 @@
<!-- gitnexus:start -->
# GitNexus MCP
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitnexusV2** (1444 symbols, 3700 relationships, 111 execution flows).
This project is indexed by GitNexus as **GitNexus** (2071 symbols, 4727 relationships, 154 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
## Always Start Here
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
1. **Read `gitnexus://repo/{name}/context`** — codebase overview + check index freshness
2. **Match your task to a skill below** and **read that skill file**
3. **Follow the skill's workflow and checklist**
## Always Do
> If step 1 warns the index is stale, run `npx gitnexus analyze` in the terminal first.
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Skills
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find execution flows related to the issue
2. `gitnexus_context({name: "<suspect function>"})` — see all callers, callees, and process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace the full execution flow step by step
4. For regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})` — see what your branch changed
## When Refactoring
- **Renaming**: MUST use `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Review the preview — graph edits are safe, text_search edits need manual review. Then run with `dry_run: false`.
- **Extracting/Splitting**: MUST run `gitnexus_context({name: "target"})` to see all incoming/outgoing refs, then `gitnexus_impact({target: "target", direction: "upstream"})` to find all external callers before moving code.
- After any refactor: run `gitnexus_detect_changes({scope: "all"})` to verify only expected files changed.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Tools Quick Reference
| Tool | When to use | Command |
|------|-------------|---------|
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update these |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## Self-Check Before Finishing
Before completing any code modification task, verify:
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL risk warnings were ignored
3. `gitnexus_detect_changes()` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## Keeping the Index Fresh
After committing code changes, the GitNexus index becomes stale. Re-run analyze to update it:
```bash
npx gitnexus analyze
```
If the index previously included embeddings, preserve them by adding `--embeddings`:
```bash
npx gitnexus analyze --embeddings
```
To check whether embeddings exist, inspect `.gitnexus/meta.json` — the `stats.embeddings` field shows the count (0 means no embeddings). **Running analyze without `--embeddings` will delete any previously generated embeddings.**
> Claude Code users: A PostToolUse hook handles this automatically after `git commit` and `git merge`.
## CLI
| Task | Read this skill file |
|------|---------------------|
+109
View File
@@ -0,0 +1,109 @@
# Changelog
All notable changes to GitNexus will be documented in this file.
## [Unreleased]
### Changed
- Migrated from KuzuDB to LadybugDB v0.15 (`@ladybugdb/core`, `@ladybugdb/wasm-core`)
- Renamed all internal paths from `kuzu` to `lbug` (storage: `.gitnexus/kuzu` → `.gitnexus/lbug`)
- Added automatic cleanup of stale KuzuDB index files
- LadybugDB v0.15 requires explicit VECTOR extension loading for semantic search
## [1.4.0] - 2026-03-13
### Added
- **Language-aware symbol resolution engine** with 3-tier resolver: exact FQN → scope-walk → guarded fuzzy fallback that refuses ambiguous matches (#238) — @magyargergo
- **Method Resolution Order (MRO)** with 5 language-specific strategies: C++ leftmost-base, C#/Java class-over-interface, Python C3 linearization, Rust qualified syntax, default BFS (#238) — @magyargergo
- **Constructor & struct literal resolution** across all languages — `new Foo()`, `User{...}`, C# primary constructors, target-typed new (#238) — @magyargergo
- **Receiver-constrained resolution** using per-file TypeEnv — disambiguates `user.save()` vs `repo.save()` via `ownerId` matching (#238) — @magyargergo
- **Heritage & ownership edges** — HAS_METHOD, OVERRIDES, Go struct embedding, Swift extension heritage, method signatures (`parameterCount`, `returnType`) (#238) — @magyargergo
- **Language-specific resolver directory** (`resolvers/`) — extracted JVM, Go, C#, PHP, Rust resolvers from monolithic import-processor (#238) — @magyargergo
- **Type extractor directory** (`type-extractors/`) — per-language type binding extraction with `Record<SupportedLanguages, Handler>` + `satisfies` dispatch (#238) — @magyargergo
- **Export detection dispatch table** — compile-time exhaustive `Record` + `satisfies` pattern replacing switch/if chains (#238) — @magyargergo
- **Language config module** (`language-config.ts`) — centralized tsconfig, go.mod, composer.json, .csproj, Swift package config loaders (#238) — @magyargergo
- **Optional skill generation** via `npx gitnexus analyze --skills` — generates AI agent skills from KuzuDB knowledge graph (#171) — @zander-raycraft
- **First-class C# support** — sibling-based modifier scanning, record/delegate/property/field/event declaration types (#163, #170, #178 via #237) — @Alice523, @benny-yamagata, @jnMetaCode
- **C/C++ support fixes** — `.h` → C++ mapping, static-linkage export detection, qualified/parenthesized declarators, 48 entry point patterns (#163, #227 via #237) — @Alice523, @bitgineer
- **Rust support fixes** — sibling-based `visibility_modifier` scanning for `pub` detection (#227 via #237) — @bitgineer
- **Adaptive tree-sitter buffer sizing** — `Math.min(Math.max(contentLength * 2, 512KB), 32MB)` (#216 via #237) — @JasonOA888
- **Call expression matching** in tree-sitter queries (#234 via #237) — @ex-nihilo-jg
- **DeepSeek model configurations** (#217) — @JasonOA888
- 282+ new unit tests, 178 integration resolver tests across 9 languages, 53 test files, 1146 total tests passing
### Fixed
- Skip unavailable native Swift parsers in sequential ingestion (#188) — @Gujiassh
- Heritage heuristic language-gated — no longer applies class/interface rules to wrong languages (#238) — @magyargergo
- C# `base_list` distinguishes EXTENDS vs IMPLEMENTS via symbol table + `I[A-Z]` heuristic (#238) — @magyargergo
- Go `qualified_type` (`models.User`) correctly unwrapped in TypeEnv (#238) — @magyargergo
- Global tier no longer blocks resolution when kind/arity filtering can narrow to 1 candidate (#238) — @magyargergo
### Changed
- `import-processor.ts` reduced from 1412 → 711 lines (50% reduction) via resolver and config extraction (#238) — @magyargergo
- `type-env.ts` reduced from 635 → ~125 lines via type-extractor extraction (#238) — @magyargergo
- CI/CD workflows hardened with security fixes and fork PR support (#222, #225) — @magyargergo
## [1.3.11] - 2026-03-08
### Security
- Fix FTS Cypher injection by escaping backslashes in search queries (#209) — @magyargergo
### Added
- Auto-reindex hook that runs `gitnexus analyze` after commits and merges, with automatic embeddings preservation (#205) — @L1nusB
- 968 integration tests (up from ~840) covering unhappy paths across search, enrichment, CLI, pipeline, worker pool, and KuzuDB (#209) — @magyargergo
- Coverage auto-ratcheting so thresholds bump automatically on CI (#209) — @magyargergo
- Rich CI PR report with coverage bars, test counts, and threshold tracking (#209) — @magyargergo
- Modular CI workflow architecture with separate unit-test, integration-test, and orchestrator jobs (#209) — @magyargergo
### Fixed
- KuzuDB native addon crashes on Linux/macOS by running integration tests in isolated vitest processes with `--pool=forks` (#209) — @magyargergo
- Worker pool `MODULE_NOT_FOUND` crash when script path is invalid (#209) — @magyargergo
### Changed
- Added macOS to the cross-platform CI test matrix (#208) — @magyargergo
## [1.3.10] - 2026-03-07
### Security
- **MCP transport buffer cap**: Added 10 MB `MAX_BUFFER_SIZE` limit to prevent out-of-memory attacks via oversized `Content-Length` headers or unbounded newline-delimited input
- **Content-Length validation**: Reject `Content-Length` values exceeding the buffer cap before allocating memory
- **Stack overflow prevention**: Replaced recursive `readNewlineMessage` with iterative loop to prevent stack overflow from consecutive empty lines
- **Ambiguous prefix hardening**: Tightened `looksLikeContentLength` to require 14+ bytes before matching, preventing false framing detection on short input
- **Closed transport guard**: `send()` now rejects with a clear error when called after `close()`, with proper write-error propagation
### Added
- **Dual-framing MCP transport** (`CompatibleStdioServerTransport`): Auto-detects Content-Length (Codex/OpenCode) and newline-delimited JSON (Cursor/Claude Code) framing on the first message, responds in the same format (#207)
- **Lazy CLI module loading**: All CLI subcommands now use `createLazyAction()` to defer heavy imports (tree-sitter, ONNX, KuzuDB) until invocation, significantly improving `gitnexus mcp` startup time (#207)
- **Type-safe lazy actions**: `createLazyAction` uses constrained generics to validate export names against module types at compile time
- **Regression test suite**: 13 unit tests covering transport framing, security hardening, buffer limits, and lazy action loading
### Fixed
- **CALLS edge sourceId alignment**: `findEnclosingFunctionId` now generates IDs with `:startLine` suffix matching node creation format, fixing process detector finding 0 entry points (#194)
- **LRU cache zero maxSize crash**: Guard `createASTCache` against `maxSize=0` when repos have no parseable files (#144)
### Changed
- Transport constructor accepts `NodeJS.ReadableStream` / `NodeJS.WritableStream` (widened from concrete `ReadStream`/`WriteStream`)
- `processReadBuffer` simplified to break on first error instead of stale-buffer retry loop
## [1.3.9] - 2026-03-06
### Fixed
- Aligned CALLS edge sourceId with node ID format in parse worker (#194)
## [1.3.8] - 2026-03-05
### Fixed
- Force-exit after analyze to prevent KuzuDB native cleanup hang (#192)
+84 -8
View File
@@ -1,17 +1,93 @@
<!-- gitnexus:start -->
# GitNexus MCP
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitnexusV2** (1444 symbols, 3700 relationships, 111 execution flows).
This project is indexed by GitNexus as **GitNexus** (2071 symbols, 4727 relationships, 154 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
## Always Start Here
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
1. **Read `gitnexus://repo/{name}/context`** — codebase overview + check index freshness
2. **Match your task to a skill below** and **read that skill file**
3. **Follow the skill's workflow and checklist**
## Always Do
> If step 1 warns the index is stale, run `npx gitnexus analyze` in the terminal first.
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Skills
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find execution flows related to the issue
2. `gitnexus_context({name: "<suspect function>"})` — see all callers, callees, and process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace the full execution flow step by step
4. For regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})` — see what your branch changed
## When Refactoring
- **Renaming**: MUST use `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Review the preview — graph edits are safe, text_search edits need manual review. Then run with `dry_run: false`.
- **Extracting/Splitting**: MUST run `gitnexus_context({name: "target"})` to see all incoming/outgoing refs, then `gitnexus_impact({target: "target", direction: "upstream"})` to find all external callers before moving code.
- After any refactor: run `gitnexus_detect_changes({scope: "all"})` to verify only expected files changed.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Tools Quick Reference
| Tool | When to use | Command |
|------|-------------|---------|
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update these |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## Self-Check Before Finishing
Before completing any code modification task, verify:
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL risk warnings were ignored
3. `gitnexus_detect_changes()` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## Keeping the Index Fresh
After committing code changes, the GitNexus index becomes stale. Re-run analyze to update it:
```bash
npx gitnexus analyze
```
If the index previously included embeddings, preserve them by adding `--embeddings`:
```bash
npx gitnexus analyze --embeddings
```
To check whether embeddings exist, inspect `.gitnexus/meta.json` — the `stats.embeddings` field shows the count (0 means no embeddings). **Running analyze without `--embeddings` will delete any previously generated embeddings.**
> Claude Code users: A PostToolUse hook handles this automatically after `git commit` and `git merge`.
## CLI
| Task | Read this skill file |
|------|---------------------|
+46 -13
View File
@@ -48,10 +48,10 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
| | **CLI + MCP** | **Web UI** |
| ----------------- | -------------------------------------------------------------- | ------------------------------------------------------------ |
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
| **For** | Daily development with Cursor, Claude Code, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
| **For** | Daily development with Cursor, Claude Code, Windsurf, OpenCode, Codex | Quick exploration, demos, one-off analysis |
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
| **Install** | `npm install -g gitnexus` | No install —[gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Storage** | KuzuDB native (fast, persistent) | KuzuDB WASM (in-memory, per session) |
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Privacy** | Everything local, no network | Everything in-browser, no server |
@@ -82,12 +82,13 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
| --------------------- | --- | ------ | -------------------- | -------------- |
| **Claude Code** | Yes | Yes | Yes (PreToolUse) | **Full** |
| **Claude Code** | Yes | Yes | Yes (PreToolUse + PostToolUse) | **Full** |
| **Cursor** | Yes | Yes | — | MCP + Skills |
| **Windsurf** | Yes | — | — | MCP |
| **OpenCode** | Yes | Yes | — | MCP + Skills |
| **Codex** | Yes | — | — | MCP |
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that automatically enrich grep/glob/bash calls with knowledge graph context.
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that enrich searches with graph context + PostToolUse hooks that auto-reindex after commits.
### Community Integrations
@@ -129,13 +130,24 @@ claude mcp add gitnexus -- npx -y gitnexus@latest mcp
}
```
**Codex** (`~/.codex/config.toml` for system scope, or `.codex/config.toml` for project scope):
```toml
[mcp_servers.gitnexus]
command = "npx"
args = ["-y", "gitnexus@latest", "mcp"]
```
### CLI Commands
```bash
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
gitnexus serve # Start local HTTP server (multi-repo) for web UI connection
gitnexus list # List all indexed repositories
@@ -189,6 +201,10 @@ gitnexus wiki --base-url <url> # Wiki with custom LLM API base URL
- **Impact Analysis** — Analyze blast radius before changes
- **Refactoring** — Plan safe refactors using dependency mapping
**Repo-specific skills** generated with `--skills`:
When you run `gitnexus analyze --skills`, GitNexus detects the functional areas of your codebase (via Leiden community detection) and generates a `SKILL.md` file for each one under `.claude/skills/generated/`. Each skill describes a module's key files, entry points, execution flows, and cross-area connections — so your AI agent gets targeted context for the exact area of code you're working in. Skills are regenerated on each `--skills` run to stay current with the codebase.
---
## Multi-Repo MCP Architecture
@@ -217,8 +233,8 @@ flowchart TD
Server["server.ts"]
Backend["LocalBackend"]
Pool["Connection Pool"]
ConnA["KuzuDB conn A"]
ConnB["KuzuDB conn B"]
ConnA["LadybugDB conn A"]
ConnB["LadybugDB conn B"]
end
Setup -->|"writes global MCP config"| CursorConfig["~/.cursor/mcp.json"]
@@ -235,7 +251,7 @@ flowchart TD
ConnB -->|"queries"| RepoB
```
**How it works:** Each `gitnexus analyze` stores the index in `.gitnexus/` inside the repo (portable, gitignored) and registers a pointer in `~/.gitnexus/registry.json`. When an AI agent starts, the MCP server reads the registry and can serve any indexed repo. KuzuDB connections are opened lazily on first query and evicted after 5 minutes of inactivity (max 5 concurrent). If only one repo is indexed, the `repo` parameter is optional on all tools — agents don't need to change anything.
**How it works:** Each `gitnexus analyze` stores the index in `.gitnexus/` inside the repo (portable, gitignored) and registers a pointer in `~/.gitnexus/registry.json`. When an AI agent starts, the MCP server reads the registry and can serve any indexed repo. LadybugDB connections are opened lazily on first query and evicted after 5 minutes of inactivity (max 5 concurrent). If only one repo is indexed, the `repo` parameter is optional on all tools — agents don't need to change anything.
---
@@ -256,7 +272,7 @@ npm install
npm run dev
```
The web UI uses the same indexing pipeline as the CLI but runs entirely in WebAssembly (Tree-sitter WASM, KuzuDB WASM, in-browser embeddings). It's great for quick exploration but limited by browser memory for larger repos.
The web UI uses the same indexing pipeline as the CLI but runs entirely in WebAssembly (Tree-sitter WASM, LadybugDB WASM, in-browser embeddings). It's great for quick exploration but limited by browser memory for larger repos.
**Local Backend Mode:** Run `gitnexus serve` and open the web UI locally — it auto-detects the server and shows all your indexed repos, with full AI chat support. No need to re-upload or re-index. The agent's tools (Cypher queries, search, code navigation) route through the backend HTTP API automatically.
@@ -313,14 +329,30 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
1. **Structure** — Walks the file tree and maps folder/file relationships
2. **Parsing** — Extracts functions, classes, methods, and interfaces using Tree-sitter ASTs
3. **Resolution** — Resolves imports and function calls across files with language-aware logic
3. **Resolution** — Resolves imports, function calls, heritage, constructor inference, and `self`/`this` receiver types across files with language-aware logic
4. **Clustering** — Groups related symbols into functional communities
5. **Processes** — Traces execution flows from entry points through call chains
6. **Search** — Builds hybrid search indexes for fast retrieval
### Supported Languages
TypeScript, JavaScript, Python, Java, Kotlin, C, C++, C#, Go, Rust, PHP, Swift
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
|----------|---------|----------------|---------|----------|-----------------|---------------------|--------|------------|-------------|
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
---
@@ -459,7 +491,7 @@ The wiki generator reads the indexed graph structure, groups files into modules
| ------------------------- | ------------------------------------- | --------------------------------------- |
| **Runtime** | Node.js (native) | Browser (WASM) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Database** | KuzuDB native | KuzuDB WASM |
| **Database** | LadybugDB native | LadybugDB WASM |
| **Embeddings** | HuggingFace transformers.js (GPU/CPU) | transformers.js (WebGPU/WASM) |
| **Search** | BM25 + semantic + RRF | BM25 + semantic + RRF |
| **Agent Interface** | MCP (stdio) | LangChain ReAct agent |
@@ -480,9 +512,10 @@ The wiki generator reads the indexed graph structure, groups files into modules
### Recently Completed
- [X] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
- [X] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
- [X] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
- [X] Multi-Repo MCP, Zero-Config Setup, 11 Language Support
- [X] Multi-Repo MCP, Zero-Config Setup, 13 Language Support
- [X] Community Detection, Process Detection, Confidence Scoring
- [X] Hybrid Search, Vector Index
@@ -499,7 +532,7 @@ The wiki generator reads the indexed graph structure, groups files into modules
## Acknowledgments
- [Tree-sitter](https://tree-sitter.github.io/) — AST parsing
- [KuzuDB](https://kuzudb.com/) — Embedded graph database with vector support
- [LadybugDB](https://ladybugdb.com/) — Embedded graph database with vector support (formerly KuzuDB)
- [Sigma.js](https://www.sigmajs.org/) — WebGL graph rendering
- [transformers.js](https://huggingface.co/docs/transformers.js) — Browser ML
- [Graphology](https://graphology.github.io/) — Graph data structures
+2 -2
View File
@@ -148,7 +148,7 @@ Each mode has a `system_{mode}.jinja` + `instance_{mode}.jinja` pair. The agent
1. Docker container starts with SWE-bench instance (repo at specific commit)
2. **GitNexus setup**: Node.js + gitnexus installed, `gitnexus analyze` runs (or restores from cache)
3. **Eval-server starts**: `gitnexus eval-server` daemon (persistent HTTP server, keeps KuzuDB warm)
3. **Eval-server starts**: `gitnexus eval-server` daemon (persistent HTTP server, keeps LadybugDB warm)
4. **Standalone tool scripts installed** in `/usr/local/bin/` — works with `subprocess.run` (no `.bashrc` needed)
5. Agent runs with the configured model + system prompt + GitNexus tools
6. Agent's patch is extracted as a git diff
@@ -167,7 +167,7 @@ Each tool script in `/usr/local/bin/` is standalone — no sourcing, no env inhe
### Eval-server
The eval-server is a lightweight HTTP daemon that:
- Keeps KuzuDB warm in memory (no cold start per tool call)
- Keeps LadybugDB warm in memory (no cold start per tool call)
- Returns LLM-friendly text (not raw JSON — saves tokens)
- Includes next-step hints to guide tool chaining (query → context → impact → fix)
- Auto-shuts down after idle timeout
+13
View File
@@ -0,0 +1,13 @@
model: deepseek-ai/deepseek-chat
provider: openrouter
cost:
input: 0.14 # per 1M tokens
output: 0.28 # per 1M tokens
# Native DeepSeek API (direct)
api_key: null
base_url: null
# For OpenRouter, uncomment below and comment out direct config above
# api_key: \${OPENROUTER_API_KEY}
# base_url: https://openrouter.ai/api/v1
+15
View File
@@ -0,0 +1,15 @@
model: deepseek-ai/DeepSeek-V3
provider: openrouter
cost:
input: 0.27 # per 1M tokens
output: 1.10 # per 1M tokens
# Native DeepSeek API (direct)
# Get your API key at: https://platform.deepseek.com/
# Or use OpenRouter with: OPENROUTER_API_KEY
api_key: null
base_url: null
# For OpenRouter, uncomment below and comment out direct config above
# api_key: \${OPENROUTER_API_KEY}
# base_url: https://openrouter.ai/api/v1
+146 -64
View File
@@ -2,8 +2,10 @@
/**
* GitNexus Claude Code Plugin Hook
*
* PreToolUse handler — intercepts Grep/Glob/Bash searches
* and augments with graph context from the GitNexus index.
* PreToolUse — intercepts Grep/Glob/Bash searches and augments
* with graph context from the GitNexus index.
* PostToolUse — detects stale index after git mutations and notifies
* the agent to reindex.
*
* NOTE: SessionStart hooks are broken on Windows (Claude Code bug #23576).
* Session context is injected via CLAUDE.md / skills instead.
@@ -26,19 +28,19 @@ function readInput() {
}
/**
* Check if a directory (or ancestor) has a .gitnexus index.
* Find the .gitnexus directory by walking up from startDir.
* Returns the path to .gitnexus/ or null if not found.
*/
function findGitNexusIndex(startDir) {
function findGitNexusDir(startDir) {
let dir = startDir || process.cwd();
for (let i = 0; i < 5; i++) {
if (fs.existsSync(path.join(dir, '.gitnexus'))) {
return true;
}
const candidate = path.join(dir, '.gitnexus');
if (fs.existsSync(candidate)) return candidate;
const parent = path.dirname(dir);
if (parent === dir) break;
dir = parent;
}
return false;
return null;
}
/**
@@ -83,66 +85,146 @@ function extractPattern(toolName, toolInput) {
return null;
}
/**
* Spawn a gitnexus CLI command synchronously.
* Detects binary on PATH once, then runs exactly once.
*
* SECURITY: Never use shell: true with user-controlled arguments.
* On Windows, invoke gitnexus.cmd directly (no shell needed).
*/
function runGitNexusCli(args, cwd, timeout) {
const isWin = process.platform === 'win32';
// Detect whether 'gitnexus' is on PATH (cheap check, no execution)
let useDirectBinary = false;
try {
const which = spawnSync(
isWin ? 'where' : 'which', ['gitnexus'],
{ encoding: 'utf-8', timeout: 3000, stdio: ['pipe', 'pipe', 'pipe'] }
);
useDirectBinary = which.status === 0;
} catch { /* not on PATH */ }
if (useDirectBinary) {
return spawnSync(
isWin ? 'gitnexus.cmd' : 'gitnexus', args,
{ encoding: 'utf-8', timeout, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
}
// npx fallback needs shell on Windows since npx is a .cmd script
return spawnSync(
isWin ? 'npx.cmd' : 'npx', ['-y', 'gitnexus', ...args],
{ encoding: 'utf-8', timeout: timeout + 5000, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
}
/**
* Emit a hook response with additional context for the agent.
*/
function sendHookResponse(hookEventName, message) {
console.log(JSON.stringify({
hookSpecificOutput: { hookEventName, additionalContext: message }
}));
}
/**
* PreToolUse handler — augment searches with graph context.
*/
function handlePreToolUse(input) {
const cwd = input.cwd || process.cwd();
if (!path.isAbsolute(cwd)) return;
if (!findGitNexusDir(cwd)) return;
const toolName = input.tool_name || '';
const toolInput = input.tool_input || {};
if (toolName !== 'Grep' && toolName !== 'Glob' && toolName !== 'Bash') return;
const pattern = extractPattern(toolName, toolInput);
if (!pattern || pattern.length < 3) return;
let result = '';
try {
const child = runGitNexusCli(['augment', '--', pattern], cwd, 7000);
if (!child.error && child.status === 0) {
result = child.stderr || '';
}
} catch { /* graceful failure */ }
if (result && result.trim()) {
sendHookResponse('PreToolUse', result.trim());
}
}
/**
* PostToolUse handler — detect index staleness after git mutations.
*
* Instead of spawning a full `gitnexus analyze` synchronously (which blocks
* the agent for up to 120s and risks LadybugDB corruption on timeout), we do a
* lightweight staleness check: compare `git rev-parse HEAD` against the
* lastCommit stored in `.gitnexus/meta.json`. If they differ, notify the
* agent so it can decide when to reindex.
*/
function handlePostToolUse(input) {
const toolName = input.tool_name || '';
if (toolName !== 'Bash') return;
const command = (input.tool_input || {}).command || '';
if (!/\bgit\s+(commit|merge|rebase|cherry-pick|pull)(\s|$)/.test(command)) return;
// Only proceed if the command succeeded
const toolOutput = input.tool_output || {};
if (toolOutput.exit_code !== undefined && toolOutput.exit_code !== 0) return;
const cwd = input.cwd || process.cwd();
if (!path.isAbsolute(cwd)) return;
const gitNexusDir = findGitNexusDir(cwd);
if (!gitNexusDir) return;
// Compare HEAD against last indexed commit — skip if unchanged
let currentHead = '';
try {
const headResult = spawnSync('git', ['rev-parse', 'HEAD'], {
encoding: 'utf-8', timeout: 3000, cwd, stdio: ['pipe', 'pipe', 'pipe'],
});
currentHead = (headResult.stdout || '').trim();
} catch { return; }
if (!currentHead) return;
let lastCommit = '';
let hadEmbeddings = false;
try {
const meta = JSON.parse(fs.readFileSync(path.join(gitNexusDir, 'meta.json'), 'utf-8'));
lastCommit = meta.lastCommit || '';
hadEmbeddings = (meta.stats && meta.stats.embeddings > 0);
} catch { /* no meta — treat as stale */ }
// If HEAD matches last indexed commit, no reindex needed
if (currentHead && currentHead === lastCommit) return;
const analyzeCmd = `npx gitnexus analyze${hadEmbeddings ? ' --embeddings' : ''}`;
sendHookResponse('PostToolUse',
`GitNexus index is stale (last indexed: ${lastCommit ? lastCommit.slice(0, 7) : 'never'}). ` +
`Run \`${analyzeCmd}\` to update the knowledge graph.`
);
}
// Dispatch map for hook events
const handlers = {
PreToolUse: handlePreToolUse,
PostToolUse: handlePostToolUse,
};
function main() {
try {
const input = readInput();
const hookEvent = input.hook_event_name || '';
if (hookEvent !== 'PreToolUse') return;
const cwd = input.cwd || process.cwd();
if (!findGitNexusIndex(cwd)) return;
const toolName = input.tool_name || '';
const toolInput = input.tool_input || {};
if (toolName !== 'Grep' && toolName !== 'Glob' && toolName !== 'Bash') return;
const pattern = extractPattern(toolName, toolInput);
if (!pattern || pattern.length < 3) return;
// augment CLI writes result to stderr (KuzuDB's native module captures
// stdout fd at OS level, making it unusable in subprocess contexts).
let result = '';
const isWin = process.platform === 'win32';
// Try direct gitnexus binary first (faster if globally installed)
try {
const child = spawnSync(
'gitnexus',
['augment', pattern],
{ encoding: 'utf-8', timeout: 8000, cwd, stdio: ['pipe', 'pipe', 'pipe'], shell: isWin }
);
if (child.status === 0 && child.stderr && child.stderr.trim()) {
result = child.stderr;
}
} catch { /* not on PATH */ }
// Fallback to npx if direct binary didn't produce output
if (!result || !result.trim()) {
try {
const child = spawnSync(
'npx',
['-y', 'gitnexus', 'augment', pattern],
{ encoding: 'utf-8', timeout: 15000, cwd, stdio: ['pipe', 'pipe', 'pipe'], shell: isWin }
);
if (child.status === 0 && child.stderr && child.stderr.trim()) {
result = child.stderr;
}
} catch { /* graceful failure */ }
const handler = handlers[input.hook_event_name || ''];
if (handler) handler(input);
} catch (err) {
if (process.env.GITNEXUS_DEBUG) {
console.error('GitNexus hook error:', (err.message || '').slice(0, 200));
}
if (result && result.trim()) {
console.log(JSON.stringify({
hookSpecificOutput: {
hookEventName: 'PreToolUse',
additionalContext: result.trim()
}
}));
}
} catch {
// Graceful failure
}
}
+13
View File
@@ -12,6 +12,19 @@
}
]
}
],
"PostToolUse": [
{
"matcher": "Bash",
"hooks": [
{
"type": "command",
"command": "node ${CLAUDE_PLUGIN_ROOT}/hooks/gitnexus-hook.js",
"timeout": 10,
"statusMessage": "Checking GitNexus index freshness..."
}
]
}
]
}
}
+25 -26
View File
@@ -10,6 +10,7 @@
"dependencies": {
"@huggingface/transformers": "^3.0.0",
"@isomorphic-git/lightning-fs": "^4.6.2",
"@ladybugdb/wasm-core": "^0.15.1",
"@langchain/anthropic": "^1.3.10",
"@langchain/core": "^1.1.15",
"@langchain/google-genai": "^2.1.10",
@@ -30,7 +31,6 @@
"graphology-utils": "^2.3.0",
"isomorphic-git": "^1.36.1",
"jszip": "^3.10.1",
"kuzu-wasm": "^0.11.1",
"langchain": "^1.2.10",
"lru-cache": "^11.2.4",
"lucide-react": "^0.562.0",
@@ -1643,6 +1643,30 @@
"@jridgewell/sourcemap-codec": "^1.4.14"
}
},
"node_modules/@ladybugdb/wasm-core": {
"version": "0.15.1",
"resolved": "https://registry.npmjs.org/@ladybugdb/wasm-core/-/wasm-core-0.15.1.tgz",
"integrity": "sha512-dHEq8inJQBkHnJrqZMKGdltSfeSv9OHECkzWQixqDLApXXGlbJ5Ugq5rRfk2PLJuZ74LVHT0cZvcn4JLmsnAIA==",
"license": "MIT",
"dependencies": {
"threads": "^1.7.0",
"tiny-worker": "^2.3.0",
"uuid": "^11.0.3"
}
},
"node_modules/@ladybugdb/wasm-core/node_modules/uuid": {
"version": "11.1.0",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-11.1.0.tgz",
"integrity": "sha512-0/A9rDy9P7cJ+8w1c9WD9V//9Wj15Ce2MPz8Ri6032usz+NfePxx5AcN3bN+r6ZL6jEo066/yNYB3tn4pQEx+A==",
"funding": [
"https://github.com/sponsors/broofa",
"https://github.com/sponsors/ctavan"
],
"license": "MIT",
"bin": {
"uuid": "dist/esm/bin/uuid"
}
},
"node_modules/@langchain/anthropic": {
"version": "1.3.10",
"resolved": "https://registry.npmjs.org/@langchain/anthropic/-/anthropic-1.3.10.tgz",
@@ -6194,31 +6218,6 @@
"resolved": "https://registry.npmjs.org/khroma/-/khroma-2.1.0.tgz",
"integrity": "sha512-Ls993zuzfayK269Svk9hzpeGUKob/sIgZzyHYdjQoAdQetRKpOLj+k/QQQ/6Qi0Yz65mlROrfd+Ev+1+7dz9Kw=="
},
"node_modules/kuzu-wasm": {
"version": "0.11.3",
"resolved": "https://registry.npmjs.org/kuzu-wasm/-/kuzu-wasm-0.11.3.tgz",
"integrity": "sha512-+bLOqXgYZJJ2dHJG1y9LTLyb9ZB73eLxErRZahZz2rPokfIdyLaktTJFzJH7wX39hgyukKn8QxeRNobH6gl27g==",
"deprecated": "Package no longer supported. Contact Support at https://www.npmjs.com/support for more info.",
"license": "MIT",
"dependencies": {
"threads": "^1.7.0",
"tiny-worker": "^2.3.0",
"uuid": "^11.0.3"
}
},
"node_modules/kuzu-wasm/node_modules/uuid": {
"version": "11.1.0",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-11.1.0.tgz",
"integrity": "sha512-0/A9rDy9P7cJ+8w1c9WD9V//9Wj15Ce2MPz8Ri6032usz+NfePxx5AcN3bN+r6ZL6jEo066/yNYB3tn4pQEx+A==",
"funding": [
"https://github.com/sponsors/broofa",
"https://github.com/sponsors/ctavan"
],
"license": "MIT",
"bin": {
"uuid": "dist/esm/bin/uuid"
}
},
"node_modules/langchain": {
"version": "1.2.10",
"resolved": "https://registry.npmjs.org/langchain/-/langchain-1.2.10.tgz",
+1 -1
View File
@@ -33,7 +33,7 @@
"graphology-layout-noverlap": "^0.4.2",
"isomorphic-git": "^1.36.1",
"jszip": "^3.10.1",
"kuzu-wasm": "^0.11.1",
"@ladybugdb/wasm-core": "^0.15.1",
"langchain": "^1.2.10",
"lru-cache": "^11.2.4",
"lucide-react": "^0.562.0",
Binary file not shown.
Binary file not shown.
@@ -5,6 +5,42 @@ import { vscDarkPlus } from 'react-syntax-highlighter/dist/esm/styles/prism';
import { useAppState } from '../hooks/useAppState';
import { NODE_COLORS } from '../lib/constants';
/** Map file extension to Prism syntax highlighter language identifier */
const getSyntaxLanguage = (filePath: string | undefined): string => {
if (!filePath) return 'text';
const ext = filePath.split('.').pop()?.toLowerCase();
switch (ext) {
case 'js': case 'jsx': case 'mjs': case 'cjs': return 'javascript';
case 'ts': case 'tsx': case 'mts': case 'cts': return 'typescript';
case 'py': case 'pyw': return 'python';
case 'rb': case 'rake': case 'gemspec': return 'ruby';
case 'java': return 'java';
case 'go': return 'go';
case 'rs': return 'rust';
case 'c': case 'h': return 'c';
case 'cpp': case 'cc': case 'cxx': case 'hpp': case 'hxx': case 'hh': return 'cpp';
case 'cs': return 'csharp';
case 'php': return 'php';
case 'kt': case 'kts': return 'kotlin';
case 'swift': return 'swift';
case 'json': return 'json';
case 'yaml': case 'yml': return 'yaml';
case 'md': case 'mdx': return 'markdown';
case 'html': case 'htm': case 'erb': return 'markup';
case 'css': case 'scss': case 'sass': return 'css';
case 'sh': case 'bash': case 'zsh': return 'bash';
case 'sql': return 'sql';
case 'xml': return 'xml';
default: break;
}
// Handle extensionless Ruby files
const basename = filePath.split('/').pop() || '';
if (['Rakefile', 'Gemfile', 'Guardfile', 'Vagrantfile', 'Brewfile'].includes(basename)) return 'ruby';
if (['Makefile'].includes(basename)) return 'makefile';
if (['Dockerfile'].includes(basename)) return 'docker';
return 'text';
};
// Match the code theme used elsewhere in the app
const customTheme = {
...vscDarkPlus,
@@ -267,12 +303,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
<div className="flex-1 min-h-0 overflow-auto scrollbar-thin">
{selectedFileContent ? (
<SyntaxHighlighter
language={
selectedFilePath?.endsWith('.py') ? 'python' :
selectedFilePath?.endsWith('.js') || selectedFilePath?.endsWith('.jsx') ? 'javascript' :
selectedFilePath?.endsWith('.ts') || selectedFilePath?.endsWith('.tsx') ? 'typescript' :
'text'
}
language={getSyntaxLanguage(selectedFilePath)}
style={customTheme as any}
showLineNumbers
startingLineNumber={1}
@@ -339,11 +370,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
const hasRange = typeof ref.startLine === 'number';
const startDisplay = hasRange ? (ref.startLine ?? 0) + 1 : undefined;
const endDisplay = hasRange ? (ref.endLine ?? ref.startLine ?? 0) + 1 : undefined;
const language =
ref.filePath.endsWith('.py') ? 'python' :
ref.filePath.endsWith('.js') || ref.filePath.endsWith('.jsx') ? 'javascript' :
ref.filePath.endsWith('.ts') || ref.filePath.endsWith('.tsx') ? 'typescript' :
'text';
const language = getSyntaxLanguage(ref.filePath);
const isGlowing = glowRefId === ref.id;
@@ -83,7 +83,7 @@ export const EmbeddingStatus = () => {
<button
onClick={handleTestArrayParams}
className="flex items-center gap-1 px-2 py-1.5 bg-surface border border-border-subtle rounded-lg text-xs text-text-muted hover:bg-hover hover:text-text-secondary transition-all"
title="Test if KuzuDB supports array params"
title="Test if LadybugDB supports array params"
>
<FlaskConical className="w-3 h-3" />
{testResult || 'Test'}
@@ -9,6 +9,6 @@ export enum SupportedLanguages {
Go = 'go',
Rust = 'rust',
PHP = 'php',
// Ruby = 'ruby',
Ruby = 'ruby',
Swift = 'swift',
}
+1 -1
View File
@@ -275,7 +275,7 @@ export const embedBatch = async (texts: string[]): Promise<Float32Array[]> => {
};
/**
* Convert Float32Array to regular number array (for KuzuDB storage)
* Convert Float32Array to regular number array (for LadybugDB storage)
*/
export const embeddingToArray = (embedding: Float32Array): number[] => {
return Array.from(embedding);
@@ -2,10 +2,10 @@
* Embedding Pipeline Module
*
* Orchestrates the background embedding process:
* 1. Query embeddable nodes from KuzuDB
* 1. Query embeddable nodes from LadybugDB
* 2. Generate text representations
* 3. Batch embed using transformers.js
* 4. Update KuzuDB with embeddings
* 4. Update LadybugDB with embeddings
* 5. Create vector index for semantic search
*/
@@ -27,7 +27,7 @@ import {
export type EmbeddingProgressCallback = (progress: EmbeddingProgress) => void;
/**
* Query all embeddable nodes from KuzuDB
* Query all embeddable nodes from LadybugDB
* Uses table-specific queries (File has different schema than code elements)
*/
const queryEmbeddableNodes = async (
@@ -102,9 +102,23 @@ const batchInsertEmbeddings = async (
* Create the vector index for semantic search
* Now indexes the separate CodeEmbedding table
*/
let vectorExtensionLoaded = false;
const createVectorIndex = async (
executeQuery: (cypher: string) => Promise<any[]>
): Promise<void> => {
// LadybugDB v0.15+ requires explicit VECTOR extension loading (once per session)
if (!vectorExtensionLoaded) {
try {
await executeQuery('INSTALL VECTOR');
await executeQuery('LOAD EXTENSION VECTOR');
vectorExtensionLoaded = true;
} catch {
// Extension may already be loaded — CREATE_VECTOR_INDEX will fail clearly if not
vectorExtensionLoaded = true;
}
}
const cypher = `
CALL CREATE_VECTOR_INDEX('CodeEmbedding', 'code_embedding_idx', 'embedding', metric := 'cosine')
`;
@@ -122,7 +136,7 @@ const createVectorIndex = async (
/**
* Run the embedding pipeline
*
* @param executeQuery - Function to execute Cypher queries against KuzuDB
* @param executeQuery - Function to execute Cypher queries against LadybugDB
* @param executeWithReusedStatement - Function to execute with reused prepared statement
* @param onProgress - Callback for progress updates
* @param config - Optional configuration override
@@ -206,7 +220,7 @@ export const runEmbeddingPipeline = async (
// Embed the batch
const embeddings = await embedBatch(texts);
// Update KuzuDB with embeddings
// Update LadybugDB with embeddings
const updates = batch.map((node, i) => ({
id: node.id,
embedding: embeddingToArray(embeddings[i]),
@@ -313,51 +327,64 @@ export const semanticSearch = async (
return [];
}
// Get metadata for each result by querying each node table
const results: SemanticSearchResult[] = [];
// Group results by label for batched metadata queries
const byLabel = new Map<string, Array<{ nodeId: string; distance: number }>>();
for (const embRow of embResults) {
const nodeId = embRow.nodeId ?? embRow[0];
const distance = embRow.distance ?? embRow[1];
// Extract label from node ID (format: Label:path:name)
const labelEndIdx = nodeId.indexOf(':');
const label = labelEndIdx > 0 ? nodeId.substring(0, labelEndIdx) : 'Unknown';
// Query the specific table for this node
// File nodes don't have startLine/endLine
if (!byLabel.has(label)) byLabel.set(label, []);
byLabel.get(label)!.push({ nodeId, distance });
}
// Batch-fetch metadata per label
const results: SemanticSearchResult[] = [];
for (const [label, items] of byLabel) {
const idList = items.map(i => `'${i.nodeId.replace(/'/g, "''")}'`).join(', ');
try {
let nodeQuery: string;
if (label === 'File') {
nodeQuery = `
MATCH (n:File {id: '${nodeId.replace(/'/g, "''")}'})
RETURN n.name AS name, n.filePath AS filePath
MATCH (n:File) WHERE n.id IN [${idList}]
RETURN n.id AS id, n.name AS name, n.filePath AS filePath
`;
} else {
nodeQuery = `
MATCH (n:${label} {id: '${nodeId.replace(/'/g, "''")}'})
RETURN n.name AS name, n.filePath AS filePath,
MATCH (n:${label}) WHERE n.id IN [${idList}]
RETURN n.id AS id, n.name AS name, n.filePath AS filePath,
n.startLine AS startLine, n.endLine AS endLine
`;
}
const nodeRows = await executeQuery(nodeQuery);
if (nodeRows.length > 0) {
const nodeRow = nodeRows[0];
results.push({
nodeId,
name: nodeRow.name ?? nodeRow[0] ?? '',
label,
filePath: nodeRow.filePath ?? nodeRow[1] ?? '',
distance,
startLine: label !== 'File' ? (nodeRow.startLine ?? nodeRow[2]) : undefined,
endLine: label !== 'File' ? (nodeRow.endLine ?? nodeRow[3]) : undefined,
});
const rowMap = new Map<string, any>();
for (const row of nodeRows) {
const id = row.id ?? row[0];
rowMap.set(id, row);
}
for (const item of items) {
const nodeRow = rowMap.get(item.nodeId);
if (nodeRow) {
results.push({
nodeId: item.nodeId,
name: nodeRow.name ?? nodeRow[1] ?? '',
label,
filePath: nodeRow.filePath ?? nodeRow[2] ?? '',
distance: item.distance,
startLine: label !== 'File' ? (nodeRow.startLine ?? nodeRow[3]) : undefined,
endLine: label !== 'File' ? (nodeRow.endLine ?? nodeRow[4]) : undefined,
});
}
}
} catch {
// Table might not exist, skip
}
}
// Re-sort by distance since batch queries may have mixed order
results.sort((a, b) => a.distance - b.distance);
return results;
};
+1 -1
View File
@@ -92,7 +92,7 @@ export interface SemanticSearchResult {
}
/**
* Node data for embedding (minimal structure from KuzuDB query)
* Node data for embedding (minimal structure from LadybugDB query)
*/
export interface EmbeddableNode {
id: string;
+1
View File
@@ -54,6 +54,7 @@ export type RelationshipType =
| 'DECORATES'
| 'IMPLEMENTS'
| 'EXTENDS'
| 'HAS_METHOD'
| 'MEMBER_OF'
| 'STEP_IN_PROCESS'
+190 -34
View File
@@ -6,6 +6,7 @@ import { loadParser, loadLanguage } from '../tree-sitter/parser-loader';
import { LANGUAGE_QUERIES } from './tree-sitter-queries';
import { generateId } from '../../lib/utils';
import { getLanguageFromFilename } from './utils';
import { callRouters } from './call-routing';
/**
* Node types that represent function/method definitions across languages.
@@ -35,6 +36,9 @@ const FUNCTION_NODE_TYPES = new Set([
// Rust
'function_item',
'impl_item', // Methods inside impl blocks
// Ruby
'method', // def foo
'singleton_method', // def self.foo
]);
/**
@@ -92,6 +96,18 @@ const findEnclosingFunction = (
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method'; // Treat constructors as methods for process detection
} else if (current.type === 'method') {
// Ruby instance method: def foo
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'singleton_method') {
// Ruby class method: def self.foo
const nameNode = current.childForFieldName?.('name') ||
current.children?.find((c: any) => c.type === 'identifier');
funcName = nameNode?.text;
label = 'Method';
} else if (current.type === 'arrow_function' || current.type === 'function_expression') {
// Arrow/expression: const foo = () => {} - check parent variable declarator
const parent = current.parent;
@@ -126,6 +142,47 @@ const findEnclosingFunction = (
return null; // Top-level call (not inside any function)
};
/** AST node types that represent a class-like container */
const CLASS_CONTAINER_TYPES = new Set([
'class_declaration', 'abstract_class_declaration',
'interface_declaration', 'struct_declaration', 'record_declaration',
'class_specifier', 'struct_specifier',
'impl_item', 'trait_item',
'class_definition',
'trait_declaration',
'protocol_declaration',
'class', 'module', // Ruby
]);
const CONTAINER_TYPE_TO_LABEL: Record<string, string> = {
class_declaration: 'Class', abstract_class_declaration: 'Class',
interface_declaration: 'Interface',
struct_declaration: 'Struct', struct_specifier: 'Struct',
class_specifier: 'Class', class_definition: 'Class',
impl_item: 'Impl', trait_item: 'Trait', trait_declaration: 'Trait',
record_declaration: 'Record', protocol_declaration: 'Interface',
class: 'Class', module: 'Module',
};
/** Walk up AST to find enclosing class/struct/interface, return its generateId or null. */
const findEnclosingClassId = (node: any, filePath: string): string | null => {
let current = node.parent;
while (current) {
if (CLASS_CONTAINER_TYPES.has(current.type)) {
const nameNode = current.childForFieldName?.('name')
?? current.children?.find((c: any) =>
c.type === 'type_identifier' || c.type === 'identifier' || c.type === 'name' || c.type === 'constant'
);
if (nameNode) {
const label = CONTAINER_TYPE_TO_LABEL[current.type] || 'Class';
return generateId(label, `${filePath}:${nameNode.text}`);
}
}
current = current.parent;
}
return null;
};
export const processCalls = async (
graph: KnowledgeGraph,
files: { path: string; content: string }[],
@@ -171,6 +228,8 @@ export const processCalls = async (
continue;
}
const callRouter = callRouters[language];
// 3. Process each call match
matches.forEach(match => {
const captureMap: Record<string, any> = {};
@@ -184,6 +243,68 @@ export const processCalls = async (
const calledName = nameNode.text;
// Dispatch: route language-specific calls (heritage, properties, imports)
const routed = callRouter(calledName, captureMap['call']);
if (routed) {
switch (routed.kind) {
case 'skip':
case 'import': // handled by import-processor
return;
case 'heritage':
for (const item of routed.items) {
const childId = symbolTable.lookupExact(file.path, item.enclosingClass) ||
symbolTable.lookupFuzzy(item.enclosingClass)[0]?.nodeId ||
generateId('Class', `${file.path}:${item.enclosingClass}`);
const parentId = symbolTable.lookupFuzzy(item.mixinName)[0]?.nodeId ||
generateId('Module', `${item.mixinName}`);
if (childId && parentId) {
const relId = generateId('IMPLEMENTS', `${childId}->${parentId}:${item.heritageKind}`);
graph.addRelationship({
id: relId, sourceId: childId, targetId: parentId,
type: 'IMPLEMENTS', confidence: 1.0, reason: item.heritageKind,
});
}
}
return;
case 'properties': {
const fileId = generateId('File', file.path);
const propEnclosingClassId = findEnclosingClassId(captureMap['call'], file.path);
for (const item of routed.items) {
const nodeId = generateId('Property', `${file.path}:${item.propName}`);
graph.addNode({
id: nodeId,
label: 'Property' as any, // TODO: add 'Property' to graph node label union
properties: {
name: item.propName, filePath: file.path,
startLine: item.startLine, endLine: item.endLine,
language, isExported: true,
description: item.accessorType,
},
});
symbolTable.add(file.path, item.propName, nodeId, 'Property');
const relId = generateId('DEFINES', `${fileId}->${nodeId}`);
graph.addRelationship({
id: relId, sourceId: fileId, targetId: nodeId,
type: 'DEFINES', confidence: 1.0, reason: '',
});
if (propEnclosingClassId) {
graph.addRelationship({
id: generateId('HAS_METHOD', `${propEnclosingClassId}->${nodeId}`),
sourceId: propEnclosingClassId, targetId: nodeId,
type: 'HAS_METHOD', confidence: 1.0, reason: '',
});
}
}
return;
}
case 'call':
break; // fall through to normal call processing below
}
}
// Skip common built-ins and noise
if (isBuiltInOrNoise(calledName)) return;
@@ -200,10 +321,10 @@ export const processCalls = async (
// 5. Find the enclosing function (caller)
const callNode = captureMap['call'];
const enclosingFuncId = findEnclosingFunction(callNode, file.path, symbolTable);
// Use enclosing function as source, fallback to file for top-level calls
const sourceId = enclosingFuncId || generateId('File', file.path);
const relId = generateId('CALLS', `${sourceId}:${calledName}->${resolved.nodeId}`);
graph.addRelationship({
@@ -711,37 +832,72 @@ const resolveCallTarget = (
* Filter out common built-in functions and noise
* that shouldn't be tracked as calls
*/
const isBuiltInOrNoise = (name: string): boolean => {
const builtIns = new Set([
// JavaScript/TypeScript built-ins
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export',
'fetch', 'Response', 'Request',
// React hooks and common functions
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
// Common array/object methods
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python built-ins
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'open', 'read', 'write', 'close', 'append', 'extend', 'update',
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
]);
/** Pre-built set (module-level singleton) to avoid re-creating per call */
const BUILT_IN_NAMES = new Set([
// JavaScript/TypeScript built-ins
'console', 'log', 'warn', 'error', 'info', 'debug',
'setTimeout', 'setInterval', 'clearTimeout', 'clearInterval',
'parseInt', 'parseFloat', 'isNaN', 'isFinite',
'encodeURI', 'decodeURI', 'encodeURIComponent', 'decodeURIComponent',
'JSON', 'parse', 'stringify',
'Object', 'Array', 'String', 'Number', 'Boolean', 'Symbol', 'BigInt',
'Map', 'Set', 'WeakMap', 'WeakSet',
'Promise', 'resolve', 'reject', 'then', 'catch', 'finally',
'Math', 'Date', 'RegExp', 'Error',
'require', 'import', 'export',
'fetch', 'Response', 'Request',
// React hooks and common functions
'useState', 'useEffect', 'useCallback', 'useMemo', 'useRef', 'useContext',
'useReducer', 'useLayoutEffect', 'useImperativeHandle', 'useDebugValue',
'createElement', 'createContext', 'createRef', 'forwardRef', 'memo', 'lazy',
// Common array/object methods
'map', 'filter', 'reduce', 'forEach', 'find', 'findIndex', 'some', 'every',
'includes', 'indexOf', 'slice', 'splice', 'concat', 'join', 'split',
'push', 'pop', 'shift', 'unshift', 'sort', 'reverse',
'keys', 'values', 'entries', 'assign', 'freeze', 'seal',
'hasOwnProperty', 'toString', 'valueOf',
// Python built-ins
'print', 'len', 'range', 'str', 'int', 'float', 'list', 'dict', 'set', 'tuple',
'open', 'read', 'write', 'close', 'append', 'extend', 'update',
'super', 'type', 'isinstance', 'issubclass', 'getattr', 'setattr', 'hasattr',
'enumerate', 'zip', 'sorted', 'reversed', 'min', 'max', 'sum', 'abs',
// C/C++ standard library and common kernel helpers
'printf', 'fprintf', 'sprintf', 'snprintf', 'vprintf', 'vfprintf', 'vsprintf', 'vsnprintf',
'scanf', 'fscanf', 'sscanf',
'malloc', 'calloc', 'realloc', 'free', 'memcpy', 'memmove', 'memset', 'memcmp',
'strlen', 'strcpy', 'strncpy', 'strcat', 'strncat', 'strcmp', 'strncmp', 'strstr', 'strchr', 'strrchr',
'atoi', 'atol', 'atof', 'strtol', 'strtoul', 'strtoll', 'strtoull', 'strtod',
'sizeof', 'offsetof', 'typeof',
'assert', 'abort', 'exit', '_exit',
'fopen', 'fclose', 'fread', 'fwrite', 'fseek', 'ftell', 'rewind', 'fflush', 'fgets', 'fputs',
// Linux kernel common macros/helpers (not real call targets)
'likely', 'unlikely', 'BUG', 'BUG_ON', 'WARN', 'WARN_ON', 'WARN_ONCE',
'IS_ERR', 'PTR_ERR', 'ERR_PTR', 'IS_ERR_OR_NULL',
'ARRAY_SIZE', 'container_of', 'list_for_each_entry', 'list_for_each_entry_safe',
'min', 'max', 'clamp', 'abs', 'swap',
'pr_info', 'pr_warn', 'pr_err', 'pr_debug', 'pr_notice', 'pr_crit', 'pr_emerg',
'printk', 'dev_info', 'dev_warn', 'dev_err', 'dev_dbg',
'GFP_KERNEL', 'GFP_ATOMIC',
'spin_lock', 'spin_unlock', 'spin_lock_irqsave', 'spin_unlock_irqrestore',
'mutex_lock', 'mutex_unlock', 'mutex_init',
'kfree', 'kmalloc', 'kzalloc', 'kcalloc', 'krealloc', 'kvmalloc', 'kvfree',
'get', 'put',
// Ruby built-ins and Kernel methods
'puts', 'print', 'p', 'pp', 'warn', 'raise', 'fail',
'require', 'require_relative', 'load', 'autoload',
'include', 'extend', 'prepend',
'attr_accessor', 'attr_reader', 'attr_writer',
'public', 'private', 'protected', 'module_function',
'lambda', 'proc', 'block_given?',
'nil?', 'is_a?', 'kind_of?', 'instance_of?', 'respond_to?',
'freeze', 'frozen?', 'dup', 'clone', 'tap', 'then', 'yield_self',
// Ruby enumerables
'each', 'map', 'select', 'reject', 'find', 'detect', 'collect',
'inject', 'reduce', 'flat_map', 'each_with_object', 'each_with_index',
'any?', 'all?', 'none?', 'count', 'first', 'last',
'sort', 'sort_by', 'min', 'max', 'min_by', 'max_by',
'group_by', 'partition', 'zip', 'compact', 'flatten', 'uniq',
]);
return builtIns.has(name);
};
const isBuiltInOrNoise = (name: string): boolean => BUILT_IN_NAMES.has(name);
@@ -0,0 +1,148 @@
/**
* Shared Ruby call routing logic.
*
* Ruby expresses imports, heritage (mixins), and property definitions as
* method calls rather than syntax-level constructs. This module provides a
* routing function used by the CLI call-processor, CLI parse-worker, and
* the web call-processor so that the classification logic lives in one place.
*
* NOTE: This file is intentionally duplicated in gitnexus-web/ because the
* two packages have separate build targets (Node native vs WASM/browser).
* Keep both copies in sync until a shared package is introduced.
*/
import { SupportedLanguages } from '../../config/supported-languages';
// ── Call routing dispatch table ─────────────────────────────────────────────
/** null = this call was not routed; fall through to default call handling */
export type CallRoutingResult = RubyCallRouting | null;
export type CallRouter = (
calledName: string,
callNode: any,
) => CallRoutingResult;
/** No-op router: returns null for every call (passthrough to normal processing) */
const noRouting: CallRouter = () => null;
/** Per-language call routing. noRouting = no special routing (normal call processing) */
export const callRouters: Record<SupportedLanguages, CallRouter> = {
[SupportedLanguages.JavaScript]: noRouting,
[SupportedLanguages.TypeScript]: noRouting,
[SupportedLanguages.Python]: noRouting,
[SupportedLanguages.Java]: noRouting,
[SupportedLanguages.Go]: noRouting,
[SupportedLanguages.Rust]: noRouting,
[SupportedLanguages.CSharp]: noRouting,
[SupportedLanguages.PHP]: noRouting,
[SupportedLanguages.Swift]: noRouting,
[SupportedLanguages.CPlusPlus]: noRouting,
[SupportedLanguages.C]: noRouting,
[SupportedLanguages.Ruby]: routeRubyCall,
};
// ── Result types ────────────────────────────────────────────────────────────
export type RubyCallRouting =
| { kind: 'import'; importPath: string; isRelative: boolean }
| { kind: 'heritage'; items: RubyHeritageItem[] }
| { kind: 'properties'; items: RubyPropertyItem[] }
| { kind: 'call' }
| { kind: 'skip' };
export interface RubyHeritageItem {
enclosingClass: string;
mixinName: string;
heritageKind: 'include' | 'extend' | 'prepend';
}
export type RubyAccessorType = 'attr_accessor' | 'attr_reader' | 'attr_writer';
export interface RubyPropertyItem {
propName: string;
accessorType: RubyAccessorType;
startLine: number;
endLine: number;
}
// ── Pre-allocated singletons for common return values ────────────────────────
const CALL_RESULT: RubyCallRouting = { kind: 'call' };
const SKIP_RESULT: RubyCallRouting = { kind: 'skip' };
/** Max depth for parent-walking loops to prevent pathological AST traversals */
const MAX_PARENT_DEPTH = 50;
// ── Routing function ────────────────────────────────────────────────────────
/**
* Classify a Ruby call node and extract its semantic payload.
*
* @param calledName - The method name (e.g. 'require', 'include', 'attr_accessor')
* @param callNode - The tree-sitter `call` AST node
* @returns A discriminated union describing the call's semantic role
*/
export function routeRubyCall(calledName: string, callNode: any): RubyCallRouting {
// ── require / require_relative → import ─────────────────────────────────
if (calledName === 'require' || calledName === 'require_relative') {
const argList = callNode.childForFieldName?.('arguments');
const stringNode = argList?.children?.find((c: any) => c.type === 'string');
const contentNode = stringNode?.children?.find((c: any) => c.type === 'string_content');
if (!contentNode) return SKIP_RESULT;
let importPath: string = contentNode.text;
// Validate: reject null bytes, control chars, excessively long paths
if (!importPath || importPath.length > 1024 || /[\x00-\x1f]/.test(importPath)) {
return SKIP_RESULT;
}
const isRelative = calledName === 'require_relative';
if (isRelative && !importPath.startsWith('.')) {
importPath = './' + importPath;
}
return { kind: 'import', importPath, isRelative };
}
// ── include / extend / prepend → heritage (mixin) ──────────────────────
if (calledName === 'include' || calledName === 'extend' || calledName === 'prepend') {
let enclosingClass: string | null = null;
let current = callNode.parent;
let depth = 0;
while (current && ++depth <= MAX_PARENT_DEPTH) {
if (current.type === 'class' || current.type === 'module') {
const nameNode = current.childForFieldName?.('name');
if (nameNode) { enclosingClass = nameNode.text; break; }
}
current = current.parent;
}
if (!enclosingClass) return SKIP_RESULT;
const items: RubyHeritageItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of (argList?.children ?? [])) {
if (arg.type === 'constant' || arg.type === 'scope_resolution') {
items.push({ enclosingClass, mixinName: arg.text, heritageKind: calledName as 'include' | 'extend' | 'prepend' });
}
}
return items.length > 0 ? { kind: 'heritage', items } : SKIP_RESULT;
}
// ── attr_accessor / attr_reader / attr_writer → property definitions ───
if (calledName === 'attr_accessor' || calledName === 'attr_reader' || calledName === 'attr_writer') {
const items: RubyPropertyItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of (argList?.children ?? [])) {
if (arg.type === 'simple_symbol') {
items.push({
propName: arg.text.startsWith(':') ? arg.text.slice(1) : arg.text,
accessorType: calledName as RubyAccessorType,
startLine: arg.startPosition.row,
endLine: arg.endPosition.row,
});
}
}
return items.length > 0 ? { kind: 'properties', items } : SKIP_RESULT;
}
// ── Everything else → regular call ─────────────────────────────────────
return CALL_RESULT;
}
@@ -330,25 +330,20 @@ const calculateCohesion = (memberIds: string[], graph: Graph): number => {
const memberSet = new Set(memberIds);
let internalEdges = 0;
// Count edges within the community
let totalEdges = 0;
// Count internal vs total edges for community members
memberIds.forEach(nodeId => {
if (graph.hasNode(nodeId)) {
graph.forEachNeighbor(nodeId, neighbor => {
totalEdges++;
if (memberSet.has(neighbor)) {
internalEdges++;
}
});
}
});
// Each edge is counted twice (once from each end), so divide by 2
internalEdges = internalEdges / 2;
// Maximum possible internal edges for n nodes: n*(n-1)/2
const maxPossibleEdges = (memberIds.length * (memberIds.length - 1)) / 2;
if (maxPossibleEdges === 0) return 1.0;
return Math.min(1.0, internalEdges / maxPossibleEdges);
if (totalEdges === 0) return 1.0;
return Math.min(1.0, internalEdges / totalEdges);
};
@@ -13,7 +13,7 @@
import { detectFrameworkFromPath } from './framework-detection';
// ============================================================================
// NAME PATTERNS - All 9 supported languages
// NAME PATTERNS - All 11 supported languages
// ============================================================================
/**
@@ -143,6 +143,13 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
/^save$/, // Repository::save()
/^delete$/, // Repository::delete()
],
// Ruby
'ruby': [
/^call$/, // Service objects (MyService.call)
/^perform$/, // Background jobs (Sidekiq, ActiveJob)
/^execute$/, // Command pattern
],
};
// ============================================================================
@@ -302,7 +309,12 @@ export function isTestFile(filePath: string): boolean {
p.endsWith('test.php') ||
p.endsWith('spec.php') ||
p.includes('/tests/feature/') ||
p.includes('/tests/unit/')
p.includes('/tests/unit/') ||
// Ruby test patterns
p.endsWith('_spec.rb') ||
p.endsWith('_test.rb') ||
p.includes('/spec/') ||
p.includes('/test/fixtures/')
);
}
@@ -257,6 +257,17 @@ export function detectFrameworkFromPath(filePath: string): FrameworkHint | null
return { framework: 'laravel', entryPointMultiplier: 1.5, reason: 'laravel-repository' };
}
// ========== RUBY ==========
// Ruby: bin/ or exe/ (CLI entry points)
if ((p.includes('/bin/') || p.includes('/exe/')) && p.endsWith('.rb')) {
return { framework: 'ruby', entryPointMultiplier: 2.5, reason: 'ruby-executable' };
}
// Ruby: Rakefile or *.rake (task definitions)
if (p.endsWith('/rakefile') || p.endsWith('.rake')) {
return { framework: 'ruby', entryPointMultiplier: 1.5, reason: 'ruby-rake' };
}
// ========== SWIFT / iOS ==========
// iOS App entry points (highest priority)
@@ -4,6 +4,7 @@ import { loadParser, loadLanguage } from '../tree-sitter/parser-loader';
import { LANGUAGE_QUERIES } from './tree-sitter-queries';
import { generateId } from '../../lib/utils';
import { getLanguageFromFilename } from './utils';
import { callRouters } from './call-routing';
// Type: Map<FilePath, Set<ResolvedFilePath>>
// Stores all files that a given file imports from
@@ -53,7 +54,9 @@ const resolveImportPath = (
// Go
'.go',
// Rust
'.rs', '/mod.rs'
'.rs', '/mod.rs',
// Ruby
'.rb', '.rake',
];
if (importPath.startsWith('.')) {
@@ -220,6 +223,35 @@ export const processImports = async (
importMap.get(file.path)!.add(resolvedPath);
}
}
// ---- Language-specific call-as-import routing (Ruby require, etc.) ----
if (captureMap['call']) {
const callNameNode = captureMap['call.name'];
if (callNameNode) {
const callRouter = callRouters[language];
const routed = callRouter(callNameNode.text, captureMap['call']);
if (routed && routed.kind === 'import') {
totalImportsFound++;
const resolvedPath = resolveImportPath(
file.path, routed.importPath, allFilePaths, allFileList, resolveCache
);
if (resolvedPath) {
const sourceId = generateId('File', file.path);
const targetId = generateId('File', resolvedPath);
const relId = generateId('IMPORTS', `${file.path}->${resolvedPath}`);
totalImportsResolved++;
graph.addRelationship({
id: relId, sourceId, targetId,
type: 'IMPORTS', confidence: 1.0, reason: '',
});
if (!importMap.has(file.path)) {
importMap.set(file.path, new Set());
}
importMap.get(file.path)!.add(resolvedPath);
}
}
}
}
});
// If re-parsed just for this, delete the tree to save memory
@@ -14,7 +14,7 @@ export type FileProgressCallback = (current: number, total: number, filePath: st
/**
* Check if a symbol (function, class, etc.) is exported/public
* Handles all 9 supported languages with explicit logic
* Handles all 11 supported languages with explicit logic
*
* @param node - The AST node for the symbol name
* @param name - The symbol name
@@ -104,7 +104,11 @@ const isNodeExported = (node: any, name: string, language: string): boolean => {
case 'c':
case 'cpp':
return false;
// Ruby: All top-level definitions are public by default
case 'ruby':
return true;
default:
return false;
}
@@ -396,6 +396,40 @@ export const PHP_QUERIES = `
[(name) (qualified_name)] @heritage.trait))) @heritage
`;
// Ruby queries - works with tree-sitter-ruby
// NOTE: Ruby uses `call` for require, include, extend, prepend, attr_* etc.
// These are all captured as @call and routed in JS post-processing:
// - require/require_relative → import extraction
// - include/extend/prepend → heritage (mixin) extraction
// - attr_accessor/attr_reader/attr_writer → property definition extraction
// - everything else → regular call extraction
export const RUBY_QUERIES = `
; ── Modules ──────────────────────────────────────────────────────────────────
(module
name: (constant) @name) @definition.module
; ── Classes ──────────────────────────────────────────────────────────────────
(class
name: (constant) @name) @definition.class
; ── Instance methods ─────────────────────────────────────────────────────────
(method
name: (identifier) @name) @definition.method
; ── Singleton (class-level) methods ──────────────────────────────────────────
(singleton_method
name: (identifier) @name) @definition.function
; ── All calls (require, include, attr_*, and regular calls routed in JS) ─────
(call
method: (identifier) @call.name) @call
; ── Heritage: class < SuperClass ─────────────────────────────────────────────
(class
name: (constant) @heritage.class
superclass: (superclass
(constant) @heritage.extends)) @heritage`;
// Swift queries - works with tree-sitter-swift
export const SWIFT_QUERIES = `
; Classes
@@ -460,6 +494,7 @@ export const LANGUAGE_QUERIES: Record<SupportedLanguages, string> = {
[SupportedLanguages.CSharp]: CSHARP_QUERIES,
[SupportedLanguages.Rust]: RUST_QUERIES,
[SupportedLanguages.PHP]: PHP_QUERIES,
[SupportedLanguages.Ruby]: RUBY_QUERIES,
[SupportedLanguages.Swift]: SWIFT_QUERIES,
};
+12
View File
@@ -1,5 +1,8 @@
import { SupportedLanguages } from '../../config/supported-languages';
/** Ruby extensionless filenames recognised as Ruby source */
const RUBY_EXTENSIONLESS_FILES = new Set(['Rakefile', 'Gemfile', 'Guardfile', 'Vagrantfile', 'Brewfile']);
/**
* Map file extension to SupportedLanguage enum
*/
@@ -31,6 +34,15 @@ export const getLanguageFromFilename = (filename: string): SupportedLanguages |
filename.endsWith('.php5') || filename.endsWith('.php8')) {
return SupportedLanguages.PHP;
}
// Ruby (extensions)
if (filename.endsWith('.rb') || filename.endsWith('.rake') || filename.endsWith('.gemspec')) {
return SupportedLanguages.Ruby;
}
// Ruby (extensionless files)
const basename = filename.split('/').pop() || filename;
if (RUBY_EXTENSIONLESS_FILES.has(basename)) {
return SupportedLanguages.Ruby;
}
// Swift
if (filename.endsWith('.swift')) return SupportedLanguages.Swift;
return null;
@@ -1,5 +1,5 @@
/**
* CSV Generator for KuzuDB Hybrid Schema
* CSV Generator for LadybugDB Hybrid Schema
*
* Generates separate CSV files for each node table and one relation CSV.
* This enables efficient bulk loading via COPY FROM for hybrid schema.
@@ -18,10 +18,10 @@ import { NODE_TABLES, NodeTableName } from './schema';
// ============================================================================
/**
* Sanitize string to ensure valid UTF-8 and safe CSV content for KuzuDB
* Sanitize string to ensure valid UTF-8 and safe CSV content for LadybugDB
* Removes or replaces invalid characters that would break CSV parsing.
*
* Critical: KuzuDB's CSV parser can misinterpret \r\n inside quoted fields.
* Critical: LadybugDB's CSV parser can misinterpret \r\n inside quoted fields.
* We normalize all line endings to \n only.
*/
const sanitizeUTF8 = (str: string): string => {
@@ -213,7 +213,7 @@ const generateCommunityCSV = (nodes: GraphNode[]): string => {
for (const node of nodes) {
if (node.label !== 'Community') continue;
// Handle keywords array - convert to KuzuDB array format
// Handle keywords array - convert to LadybugDB array format
const keywords = (node.properties as any).keywords || [];
const keywordsStr = `[${keywords.map((k: string) => `'${k.replace(/'/g, "''")}'`).join(',')}]`;
@@ -221,7 +221,7 @@ const generateCommunityCSV = (nodes: GraphNode[]): string => {
escapeCSVField(node.id),
escapeCSVField(node.properties.name || ''), // label is stored in name
escapeCSVField(node.properties.heuristicLabel || ''),
keywordsStr, // Array format for KuzuDB
keywordsStr, // Array format for LadybugDB
escapeCSVField((node.properties as any).description || ''),
escapeCSVField((node.properties as any).enrichedBy || 'heuristic'),
escapeCSVNumber(node.properties.cohesion, 0),
@@ -1,51 +1,51 @@
/**
* KuzuDB Adapter
*
* Manages the KuzuDB WASM instance for client-side graph database operations.
* LadybugDB Adapter
*
* Manages the LadybugDB WASM instance for client-side graph database operations.
* Uses the "Snapshot / Bulk Load" pattern with COPY FROM for performance.
*
*
* Multi-table schema: separate tables for File, Function, Class, etc.
*/
import { KnowledgeGraph } from '../graph/types';
import {
NODE_TABLES,
import {
NODE_TABLES,
REL_TABLE_NAME,
SCHEMA_QUERIES,
SCHEMA_QUERIES,
EMBEDDING_TABLE_NAME,
NodeTableName,
} from './schema';
import { generateAllCSVs } from './csv-generator';
// Holds the reference to the dynamically loaded module
let kuzu: any = null;
let lbug: any = null;
let db: any = null;
let conn: any = null;
/**
* Initialize KuzuDB WASM module and create in-memory database
* Initialize LadybugDB WASM module and create in-memory database
*/
export const initKuzu = async () => {
if (conn) return { db, conn, kuzu };
export const initLbug = async () => {
if (conn) return { db, conn, lbug };
try {
if (import.meta.env.DEV) console.log('🚀 Initializing KuzuDB...');
if (import.meta.env.DEV) console.log('🚀 Initializing LadybugDB...');
// 1. Dynamic Import (Fixes the "not a function" bundler issue)
const kuzuModule = await import('kuzu-wasm');
const lbugModule = await import('@ladybugdb/wasm-core');
// 2. Handle Vite/Webpack "default" wrapping
kuzu = kuzuModule.default || kuzuModule;
lbug = lbugModule.default || lbugModule;
// 3. Initialize WASM
await kuzu.init();
// 4. Create Database with 512MB buffer pool
await lbug.init();
// 4. Create Database with 512MB buffer manager
const BUFFER_POOL_SIZE = 512 * 1024 * 1024; // 512MB
db = new kuzu.Database(':memory:', BUFFER_POOL_SIZE);
conn = new kuzu.Connection(db);
if (import.meta.env.DEV) console.log('✅ KuzuDB WASM Initialized');
db = new lbug.Database(':memory:', BUFFER_POOL_SIZE);
conn = new lbug.Connection(db);
if (import.meta.env.DEV) console.log('✅ LadybugDB WASM Initialized');
// 5. Initialize Schema (all node tables, then rel tables, then embedding table)
for (const schemaQuery of SCHEMA_QUERIES) {
@@ -58,60 +58,60 @@ export const initKuzu = async () => {
}
}
}
if (import.meta.env.DEV) console.log('✅ KuzuDB Multi-Table Schema Created');
return { db, conn, kuzu };
if (import.meta.env.DEV) console.log('✅ LadybugDB Multi-Table Schema Created');
return { db, conn, lbug };
} catch (error) {
if (import.meta.env.DEV) console.error('❌ KuzuDB Initialization Failed:', error);
if (import.meta.env.DEV) console.error('❌ LadybugDB Initialization Failed:', error);
throw error;
}
};
/**
* Load a KnowledgeGraph into KuzuDB using COPY FROM (bulk load)
* Load a KnowledgeGraph into LadybugDB using COPY FROM (bulk load)
* Uses batched CSV writes and COPY statements for optimal performance
*/
export const loadGraphToKuzu = async (
graph: KnowledgeGraph,
export const loadGraphToLbug = async (
graph: KnowledgeGraph,
fileContents: Map<string, string>
) => {
const { conn, kuzu } = await initKuzu();
const { conn, lbug } = await initLbug();
try {
if (import.meta.env.DEV) console.log(`KuzuDB: Generating CSVs for ${graph.nodeCount} nodes...`);
if (import.meta.env.DEV) console.log(`LadybugDB: Generating CSVs for ${graph.nodeCount} nodes...`);
// 1. Generate all CSVs (per-table)
const csvData = generateAllCSVs(graph, fileContents);
const fs = kuzu.FS;
const fs = lbug.FS;
// 2. Write all node CSVs to virtual filesystem
const nodeFiles: Array<{ table: NodeTableName; path: string }> = [];
for (const [tableName, csv] of csvData.nodes.entries()) {
// Skip empty CSVs (only header row)
if (csv.split('\n').length <= 1) continue;
const path = `/${tableName.toLowerCase()}.csv`;
try { await fs.unlink(path); } catch {}
await fs.writeFile(path, csv);
nodeFiles.push({ table: tableName, path });
}
// 3. Parse relation CSV and prepare for INSERT (COPY FROM doesn't work with multi-pair tables)
const relLines = csvData.relCSV.split('\n').slice(1).filter(line => line.trim());
const relCount = relLines.length;
if (import.meta.env.DEV) {
console.log(`KuzuDB: Wrote ${nodeFiles.length} node CSVs, ${relCount} relations to insert`);
console.log(`LadybugDB: Wrote ${nodeFiles.length} node CSVs, ${relCount} relations to insert`);
}
// 4. COPY all node tables (must complete before rels due to FK constraints)
for (const { table, path } of nodeFiles) {
const copyQuery = getCopyQuery(table, path);
await conn.query(copyQuery);
}
// 5. INSERT relations one by one (COPY doesn't work with multi-pair REL tables)
// Build a set of valid table names for fast lookup
const validTables = new Set<string>(NODE_TABLES as readonly string[]);
@@ -135,13 +135,13 @@ export const loadGraphToKuzu = async (
// Format: "from","to","type",confidence,"reason",step
const match = line.match(/"([^"]*)","([^"]*)","([^"]*)",([0-9.]+),"([^"]*)",([0-9-]+)/);
if (!match) continue;
const [, fromId, toId, relType, confidenceStr, reason, stepStr] = match;
const fromLabel = getNodeLabel(fromId);
const toLabel = getNodeLabel(toId);
// Skip relationships where either node's label doesn't have a table in KuzuDB
// Skip relationships where either node's label doesn't have a table in LadybugDB
// Querying a non-existent table causes a fatal native crash
if (!validTables.has(fromLabel) || !validTables.has(toLabel)) {
skippedRels++;
@@ -150,7 +150,7 @@ export const loadGraphToKuzu = async (
const confidence = parseFloat(confidenceStr) || 1.0;
const step = parseInt(stepStr) || 0;
const insertQuery = `
MATCH (a:${escapeLabel(fromLabel)} {id: '${fromId.replace(/'/g, "''")}'}),
(b:${escapeLabel(toLabel)} {id: '${toId.replace(/'/g, "''")}'})
@@ -167,38 +167,39 @@ export const loadGraphToKuzu = async (
const toLabel = getNodeLabel(toId);
const key = `${relType}:${fromLabel}->` + toLabel;
skippedRelStats.set(key, (skippedRelStats.get(key) || 0) + 1);
if (import.meta.env.DEV) {
console.warn(`⚠️ Skipped: ${key} | "${fromId}" → "${toId}" | ${err instanceof Error ? err.message : String(err)}`);
}
}
}
}
if (import.meta.env.DEV) {
console.log(`KuzuDB: Inserted ${insertedRels}/${relCount} relations`);
console.log(`LadybugDB: Inserted ${insertedRels}/${relCount} relations`);
if (skippedRels > 0) {
const topSkipped = Array.from(skippedRelStats.entries())
.sort((a, b) => b[1] - a[1])
.slice(0, 10);
console.warn(`KuzuDB: Skipped ${skippedRels}/${relCount} relations (top by kind/pair):`, topSkipped);
console.warn(`LadybugDB: Skipped ${skippedRels}/${relCount} relations (top by kind/pair):`, topSkipped);
}
}
// 6. Verify results
let totalNodes = 0;
for (const tableName of NODE_TABLES) {
try {
const countRes = await conn.query(`MATCH (n:${tableName}) RETURN count(n) AS cnt`);
const countRow = await countRes.getNext();
const countRows = await countRes.getAll();
const countRow = countRows[0];
const count = countRow ? (countRow.cnt ?? countRow[0] ?? 0) : 0;
totalNodes += Number(count);
} catch {
// Table might be empty, skip
}
}
if (import.meta.env.DEV) console.log(`✅ KuzuDB Bulk Load Complete. Total nodes: ${totalNodes}, edges: ${insertedRels}`);
if (import.meta.env.DEV) console.log(`✅ LadybugDB Bulk Load Complete. Total nodes: ${totalNodes}, edges: ${insertedRels}`);
// 7. Cleanup CSV files
for (const { path } of nodeFiles) {
@@ -208,12 +209,12 @@ export const loadGraphToKuzu = async (
return { success: true, count: totalNodes };
} catch (error) {
if (import.meta.env.DEV) console.error('❌ KuzuDB Bulk Load Failed:', error);
if (import.meta.env.DEV) console.error('❌ LadybugDB Bulk Load Failed:', error);
return { success: false, count: 0 };
}
};
// KuzuDB default ESCAPE is '\' (backslash), but our CSV uses RFC 4180 escaping ("" for literal quotes).
// LadybugDB default ESCAPE is '\' (backslash), but our CSV uses RFC 4180 escaping ("" for literal quotes).
// Source code content is full of backslashes which confuse the auto-detection.
// We MUST explicitly set ESCAPE='"' and disable auto_detect.
const COPY_CSV_OPTS = `(HEADER=true, ESCAPE='"', DELIM=',', QUOTE='"', PARALLEL=false, auto_detect=false)`;
@@ -229,6 +230,9 @@ const escapeTableName = (table: string): string => {
return BACKTICK_TABLES.has(table) ? `\`${table}\`` : table;
};
/** Tables with isExported column (TypeScript/JS-native types) */
const TABLES_WITH_EXPORTED = new Set<string>(['Function', 'Class', 'Interface', 'Method', 'CodeElement']);
/**
* Get the COPY query for a node table with correct column mapping
*/
@@ -246,8 +250,12 @@ const getCopyQuery = (table: NodeTableName, path: string): string => {
if (table === 'Process') {
return `COPY ${t}(id, label, heuristicLabel, processType, stepCount, communities, entryPointId, terminalId) FROM "${path}" ${COPY_CSV_OPTS}`;
}
// Code element tables (Function, Class, Interface, Method, CodeElement, and multi-language)
return `COPY ${t}(id, name, filePath, startLine, endLine, isExported, content) FROM "${path}" ${COPY_CSV_OPTS}`;
// TypeScript/JS code element tables have isExported; multi-language tables do not
if (TABLES_WITH_EXPORTED.has(table)) {
return `COPY ${t}(id, name, filePath, startLine, endLine, isExported, content) FROM "${path}" ${COPY_CSV_OPTS}`;
}
// Multi-language tables (Struct, Impl, Trait, Macro, etc.)
return `COPY ${t}(id, name, filePath, startLine, endLine, content) FROM "${path}" ${COPY_CSV_OPTS}`;
};
/**
@@ -256,12 +264,12 @@ const getCopyQuery = (table: NodeTableName, path: string): string => {
*/
export const executeQuery = async (cypher: string): Promise<any[]> => {
if (!conn) {
await initKuzu();
await initLbug();
}
try {
const result = await conn.query(cypher);
// Extract column names from RETURN clause
const returnMatch = cypher.match(/RETURN\s+(.+?)(?:\s+ORDER|\s+LIMIT|\s+SKIP|\s*$)/is);
let columnNames: string[] = [];
@@ -284,12 +292,11 @@ export const executeQuery = async (cypher: string): Promise<any[]> => {
return col.replace(/[^a-zA-Z0-9_]/g, '_');
});
}
// Collect all rows
const allRows = await result.getAll();
const rows: any[] = [];
while (await result.hasNext()) {
const row = await result.getNext();
for (const row of allRows) {
// Convert tuple to named object if we have column names and row is array
if (Array.isArray(row) && columnNames.length === row.length) {
const namedRow: Record<string, any> = {};
@@ -302,7 +309,7 @@ export const executeQuery = async (cypher: string): Promise<any[]> => {
rows.push(row);
}
}
return rows;
} catch (error) {
if (import.meta.env.DEV) console.error('Query execution failed:', error);
@@ -313,7 +320,7 @@ export const executeQuery = async (cypher: string): Promise<any[]> => {
/**
* Get database statistics
*/
export const getKuzuStats = async (): Promise<{ nodes: number; edges: number }> => {
export const getLbugStats = async (): Promise<{ nodes: number; edges: number }> => {
if (!conn) {
return { nodes: 0, edges: 0 };
}
@@ -324,43 +331,45 @@ export const getKuzuStats = async (): Promise<{ nodes: number; edges: number }>
for (const tableName of NODE_TABLES) {
try {
const nodeResult = await conn.query(`MATCH (n:${tableName}) RETURN count(n) AS cnt`);
const nodeRow = await nodeResult.getNext();
const nodeRows = await nodeResult.getAll();
const nodeRow = nodeRows[0];
totalNodes += Number(nodeRow?.cnt ?? nodeRow?.[0] ?? 0);
} catch {
// Table might not exist or be empty
}
}
// Count edges from single relation table
let totalEdges = 0;
try {
const edgeResult = await conn.query(`MATCH ()-[r:${REL_TABLE_NAME}]->() RETURN count(r) AS cnt`);
const edgeRow = await edgeResult.getNext();
const edgeRows = await edgeResult.getAll();
const edgeRow = edgeRows[0];
totalEdges = Number(edgeRow?.cnt ?? edgeRow?.[0] ?? 0);
} catch {
// Table might not exist or be empty
}
return { nodes: totalNodes, edges: totalEdges };
} catch (error) {
if (import.meta.env.DEV) {
console.warn('Failed to get Kuzu stats:', error);
console.warn('Failed to get LadybugDB stats:', error);
}
return { nodes: 0, edges: 0 };
}
};
/**
* Check if KuzuDB is initialized and has data
* Check if LadybugDB is initialized and has data
*/
export const isKuzuReady = (): boolean => {
export const isLbugReady = (): boolean => {
return conn !== null && db !== null;
};
/**
* Close the database connection (cleanup)
*/
export const closeKuzu = async (): Promise<void> => {
export const closeLbug = async (): Promise<void> => {
if (conn) {
try {
await conn.close();
@@ -373,7 +382,7 @@ export const closeKuzu = async (): Promise<void> => {
} catch {}
db = null;
}
kuzu = null;
lbug = null;
};
/**
@@ -387,24 +396,20 @@ export const executePrepared = async (
params: Record<string, any>
): Promise<any[]> => {
if (!conn) {
await initKuzu();
await initLbug();
}
try {
const stmt = await conn.prepare(cypher);
if (!stmt.isSuccess()) {
const errMsg = await stmt.getErrorMessage();
throw new Error(`Prepare failed: ${errMsg}`);
}
const result = await conn.execute(stmt, params);
const rows: any[] = [];
while (await result.hasNext()) {
const row = await result.getNext();
rows.push(row);
}
const rows = await result.getAll();
await stmt.close();
return rows;
} catch (error) {
@@ -421,22 +426,22 @@ export const executeWithReusedStatement = async (
paramsList: Array<Record<string, any>>
): Promise<void> => {
if (!conn) {
await initKuzu();
await initLbug();
}
if (paramsList.length === 0) return;
const SUB_BATCH_SIZE = 4;
for (let i = 0; i < paramsList.length; i += SUB_BATCH_SIZE) {
const subBatch = paramsList.slice(i, i + SUB_BATCH_SIZE);
const stmt = await conn.prepare(cypher);
if (!stmt.isSuccess()) {
const errMsg = await stmt.getErrorMessage();
throw new Error(`Prepare failed: ${errMsg}`);
}
try {
for (const params of subBatch) {
await conn.execute(stmt, params);
@@ -444,7 +449,7 @@ export const executeWithReusedStatement = async (
} finally {
await stmt.close();
}
if (i + SUB_BATCH_SIZE < paramsList.length) {
await new Promise(r => setTimeout(r, 0));
}
@@ -456,65 +461,67 @@ export const executeWithReusedStatement = async (
*/
export const testArrayParams = async (): Promise<{ success: boolean; error?: string }> => {
if (!conn) {
await initKuzu();
await initLbug();
}
try {
const testEmbedding = new Array(384).fill(0).map((_, i) => i / 384);
// Get any node ID to test with (try File first, then others)
let testNodeId: string | null = null;
for (const tableName of NODE_TABLES) {
try {
const nodeResult = await conn.query(`MATCH (n:${tableName}) RETURN n.id AS id LIMIT 1`);
const nodeRow = await nodeResult.getNext();
const nodeRows = await nodeResult.getAll();
const nodeRow = nodeRows[0];
if (nodeRow) {
testNodeId = nodeRow.id ?? nodeRow[0];
break;
}
} catch {}
}
if (!testNodeId) {
return { success: false, error: 'No nodes found to test with' };
}
if (import.meta.env.DEV) {
console.log('🧪 Testing array params with node:', testNodeId);
}
// First create an embedding entry
const createQuery = `CREATE (e:${EMBEDDING_TABLE_NAME} {nodeId: $nodeId, embedding: $embedding})`;
const stmt = await conn.prepare(createQuery);
if (!stmt.isSuccess()) {
const errMsg = await stmt.getErrorMessage();
return { success: false, error: `Prepare failed: ${errMsg}` };
}
await conn.execute(stmt, {
nodeId: testNodeId,
embedding: testEmbedding,
});
await stmt.close();
// Verify it was stored
const verifyResult = await conn.query(
`MATCH (e:${EMBEDDING_TABLE_NAME} {nodeId: '${testNodeId}'}) RETURN e.embedding AS emb`
);
const verifyRow = await verifyResult.getNext();
const verifyRows = await verifyResult.getAll();
const verifyRow = verifyRows[0];
const storedEmb = verifyRow?.emb ?? verifyRow?.[0];
if (storedEmb && Array.isArray(storedEmb) && storedEmb.length === 384) {
if (import.meta.env.DEV) {
console.log('✅ Array params WORK! Stored embedding length:', storedEmb.length);
}
return { success: true };
} else {
return {
success: false,
error: `Embedding not stored correctly. Got: ${typeof storedEmb}, length: ${storedEmb?.length}`
return {
success: false,
error: `Embedding not stored correctly. Got: ${typeof storedEmb}, length: ${storedEmb?.length}`
};
}
} catch (error) {
@@ -1,5 +1,5 @@
/**
* KuzuDB Schema Definitions
* LadybugDB Schema Definitions
*
* Hybrid Schema:
* - Separate node tables for each code element type (File, Function, Class, etc.)
+2 -2
View File
@@ -17,7 +17,7 @@ import { z } from 'zod';
import { WebGPUNotAvailableError, embedText, embeddingToArray, initEmbedder, isEmbedderReady } from '../embeddings/embedder';
/**
* Tool factory - creates tools bound to the KuzuDB query functions
* Tool factory - creates tools bound to the LadybugDB query functions
*/
export const createGraphRAGTools = (
executeQuery: (cypher: string) => Promise<any[]>,
@@ -975,7 +975,7 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`,
// For code elements (Function, Class, etc.), use the direct id
const isFileTarget = targetType === 'File';
// Query each depth level separately (KuzuDB doesn't support list comprehensions on paths)
// Query each depth level separately (LadybugDB doesn't support list comprehensions on paths)
// For depth 1: direct connections only
// For depth 2+: chain multiple single-hop queries
const depthQueries: Promise<any[]>[] = [];
+1 -1
View File
@@ -224,7 +224,7 @@ export interface AgentStep {
* Graph schema information for LLM context
*/
export const GRAPH_SCHEMA_DESCRIPTION = `
KUZU GRAPH DATABASE SCHEMA (Multi-Table):
LADYBUG GRAPH DATABASE SCHEMA (Multi-Table):
NODE TABLES:
1. File - Source files
@@ -40,6 +40,7 @@ const getWasmPath = (language: SupportedLanguages, filePath?: string): string =>
[SupportedLanguages.Go]: '/wasm/go/tree-sitter-go.wasm',
[SupportedLanguages.Rust]: '/wasm/rust/tree-sitter-rust.wasm',
[SupportedLanguages.PHP]: '/wasm/php/tree-sitter-php.wasm',
[SupportedLanguages.Ruby]: '/wasm/ruby/tree-sitter-ruby.wasm',
[SupportedLanguages.Swift]: '/wasm/swift/tree-sitter-swift.wasm',
};
@@ -1,28 +1,35 @@
declare module 'kuzu-wasm' {
declare module '@ladybugdb/wasm-core' {
export function init(): Promise<void>;
export class Database {
constructor(path: string);
constructor(path: string, bufferPoolSize?: number);
close(): Promise<void>;
}
export class Connection {
constructor(db: Database);
query(cypher: string): Promise<QueryResult>;
prepare(cypher: string): Promise<PreparedStatement>;
execute(stmt: PreparedStatement, params?: Record<string, any>): Promise<QueryResult>;
close(): Promise<void>;
}
export interface QueryResult {
getAll(): Promise<any[]>;
hasNext(): Promise<boolean>;
getNext(): Promise<any>;
}
export interface PreparedStatement {
isSuccess(): boolean;
getErrorMessage(): Promise<string>;
close(): Promise<void>;
}
export const FS: {
writeFile(path: string, data: string): Promise<void>;
unlink(path: string): Promise<void>;
};
const kuzu: {
const lbug: {
init: typeof init;
Database: typeof Database;
Connection: typeof Connection;
FS: typeof FS;
};
export default kuzu;
export default lbug;
}
+66 -66
View File
@@ -26,13 +26,13 @@ import {
type HybridSearchResult,
} from '../core/search';
// Lazy import for Kuzu to avoid breaking worker if SharedArrayBuffer unavailable
let kuzuAdapter: typeof import('../core/kuzu/kuzu-adapter') | null = null;
const getKuzuAdapter = async () => {
if (!kuzuAdapter) {
kuzuAdapter = await import('../core/kuzu/kuzu-adapter');
// Lazy import for LadybugDB to avoid breaking worker if SharedArrayBuffer unavailable
let lbugAdapter: typeof import('../core/lbug/lbug-adapter') | null = null;
const getLbugAdapter = async () => {
if (!lbugAdapter) {
lbugAdapter = await import('../core/lbug/lbug-adapter');
}
return kuzuAdapter;
return lbugAdapter;
};
// Embedding state
@@ -172,52 +172,52 @@ const workerApi = {
console.log(`🔍 BM25 index built: ${bm25DocCount} documents`);
}
// Load graph into KuzuDB for querying (optional - gracefully degrades)
// Load graph into LadybugDB for querying (optional - gracefully degrades)
try {
onProgress({
phase: 'complete',
percent: 98,
message: 'Loading into KuzuDB...',
message: 'Loading into LadybugDB...',
stats: {
filesProcessed: result.graph.nodeCount,
totalFiles: result.graph.nodeCount,
nodesCreated: result.graph.nodeCount,
},
});
const kuzu = await getKuzuAdapter();
await kuzu.loadGraphToKuzu(result.graph, result.fileContents);
const lbug = await getLbugAdapter();
await lbug.loadGraphToLbug(result.graph, result.fileContents);
if (import.meta.env.DEV) {
const stats = await kuzu.getKuzuStats();
console.log('KuzuDB loaded:', stats);
const stats = await lbug.getLbugStats();
console.log('LadybugDB loaded:', stats);
console.log('📁 Stored', storedFileContents.size, 'files for grep/read tools');
}
} catch {
// KuzuDB is optional - silently continue without it
// LadybugDB is optional - silently continue without it
}
// Store clustering config for background enrichment (runs after graph loads)
if (clusteringConfig) {
pendingEnrichmentConfig = clusteringConfig;
console.log('📋 Clustering config saved for background enrichment');
}
// Convert to serializable format for transfer back to main thread
return serializePipelineResult(result);
},
/**
* Execute a Cypher query against the KuzuDB database
* Execute a Cypher query against the LadybugDB database
* @param cypher - The Cypher query string
* @returns Query results as an array of objects
*/
async runQuery(cypher: string): Promise<any[]> {
const kuzu = await getKuzuAdapter();
if (!kuzu.isKuzuReady()) {
const lbug = await getLbugAdapter();
if (!lbug.isLbugReady()) {
throw new Error('Database not ready. Please load a repository first.');
}
return kuzu.executeQuery(cypher);
return lbug.executeQuery(cypher);
},
/**
@@ -225,8 +225,8 @@ const workerApi = {
*/
async isReady(): Promise<boolean> {
try {
const kuzu = await getKuzuAdapter();
return kuzu.isKuzuReady();
const lbug = await getLbugAdapter();
return lbug.isLbugReady();
} catch {
return false;
}
@@ -237,8 +237,8 @@ const workerApi = {
*/
async getStats(): Promise<{ nodes: number; edges: number }> {
try {
const kuzu = await getKuzuAdapter();
return kuzu.getKuzuStats();
const lbug = await getLbugAdapter();
return lbug.getLbugStats();
} catch {
return { nodes: 0, edges: 0 };
}
@@ -276,29 +276,29 @@ const workerApi = {
console.log(`🔍 BM25 index built: ${bm25DocCount} documents`);
}
// Load graph into KuzuDB for querying (optional - gracefully degrades)
// Load graph into LadybugDB for querying (optional - gracefully degrades)
try {
onProgress({
phase: 'complete',
percent: 98,
message: 'Loading into KuzuDB...',
message: 'Loading into LadybugDB...',
stats: {
filesProcessed: result.graph.nodeCount,
totalFiles: result.graph.nodeCount,
nodesCreated: result.graph.nodeCount,
},
});
const kuzu = await getKuzuAdapter();
await kuzu.loadGraphToKuzu(result.graph, result.fileContents);
const lbug = await getLbugAdapter();
await lbug.loadGraphToLbug(result.graph, result.fileContents);
if (import.meta.env.DEV) {
const stats = await kuzu.getKuzuStats();
console.log('KuzuDB loaded:', stats);
const stats = await lbug.getLbugStats();
console.log('LadybugDB loaded:', stats);
console.log('📁 Stored', storedFileContents.size, 'files for grep/read tools');
}
} catch {
// KuzuDB is optional - silently continue without it
// LadybugDB is optional - silently continue without it
}
// Store clustering config for background enrichment (runs after graph loads)
@@ -325,8 +325,8 @@ const workerApi = {
onProgress: (progress: EmbeddingProgress) => void,
forceDevice?: 'webgpu' | 'wasm'
): Promise<void> {
const kuzu = await getKuzuAdapter();
if (!kuzu.isKuzuReady()) {
const lbug = await getLbugAdapter();
if (!lbug.isLbugReady()) {
throw new Error('Database not ready. Please load a repository first.');
}
@@ -343,8 +343,8 @@ const workerApi = {
};
await runEmbeddingPipeline(
kuzu.executeQuery,
kuzu.executeWithReusedStatement,
lbug.executeQuery,
lbug.executeWithReusedStatement,
progressCallback,
forceDevice ? { device: forceDevice } : {}
);
@@ -400,15 +400,15 @@ const workerApi = {
k: number = 10,
maxDistance: number = 0.5
): Promise<SemanticSearchResult[]> {
const kuzu = await getKuzuAdapter();
if (!kuzu.isKuzuReady()) {
const lbug = await getLbugAdapter();
if (!lbug.isLbugReady()) {
throw new Error('Database not ready. Please load a repository first.');
}
if (!isEmbeddingComplete) {
throw new Error('Embeddings not ready. Please wait for embedding pipeline to complete.');
}
return doSemanticSearch(kuzu.executeQuery, query, k, maxDistance);
return doSemanticSearch(lbug.executeQuery, query, k, maxDistance);
},
/**
@@ -424,15 +424,15 @@ const workerApi = {
k: number = 5,
hops: number = 2
): Promise<any[]> {
const kuzu = await getKuzuAdapter();
if (!kuzu.isKuzuReady()) {
const lbug = await getLbugAdapter();
if (!lbug.isLbugReady()) {
throw new Error('Database not ready. Please load a repository first.');
}
if (!isEmbeddingComplete) {
throw new Error('Embeddings not ready. Please wait for embedding pipeline to complete.');
}
return doSemanticSearchWithContext(kuzu.executeQuery, query, k, hops);
return doSemanticSearchWithContext(lbug.executeQuery, query, k, hops);
},
/**
@@ -458,9 +458,9 @@ const workerApi = {
let semanticResults: SemanticSearchResult[] = [];
if (isEmbeddingComplete) {
try {
const kuzu = await getKuzuAdapter();
if (kuzu.isKuzuReady()) {
semanticResults = await doSemanticSearch(kuzu.executeQuery, query, k * 3, 0.5);
const lbug = await getLbugAdapter();
if (lbug.isLbugReady()) {
semanticResults = await doSemanticSearch(lbug.executeQuery, query, k * 3, 0.5);
}
} catch {
// Semantic search failed, continue with BM25 only
@@ -516,15 +516,15 @@ const workerApi = {
},
/**
* Test if KuzuDB supports array parameters in prepared statements
* Test if LadybugDB supports array parameters in prepared statements
* This is a diagnostic function
*/
async testArrayParams(): Promise<{ success: boolean; error?: string }> {
const kuzu = await getKuzuAdapter();
if (!kuzu.isKuzuReady()) {
const lbug = await getLbugAdapter();
if (!lbug.isLbugReady()) {
return { success: false, error: 'Database not ready' };
}
return kuzu.testArrayParams();
return lbug.testArrayParams();
},
// ============================================================
@@ -539,8 +539,8 @@ const workerApi = {
*/
async initializeAgent(config: ProviderConfig, projectName?: string): Promise<{ success: boolean; error?: string }> {
try {
const kuzu = await getKuzuAdapter();
if (!kuzu.isKuzuReady()) {
const lbug = await getLbugAdapter();
if (!lbug.isLbugReady()) {
return { success: false, error: 'Database not ready. Please load a repository first.' };
}
@@ -549,31 +549,31 @@ const workerApi = {
if (!isEmbeddingComplete) {
throw new Error('Embeddings not ready');
}
return doSemanticSearch(kuzu.executeQuery, query, k, maxDistance);
return doSemanticSearch(lbug.executeQuery, query, k, maxDistance);
};
const semanticSearchWithContextWrapper = async (query: string, k?: number, hops?: number) => {
if (!isEmbeddingComplete) {
throw new Error('Embeddings not ready');
}
return doSemanticSearchWithContext(kuzu.executeQuery, query, k, hops);
return doSemanticSearchWithContext(lbug.executeQuery, query, k, hops);
};
// Hybrid search wrapper - combines BM25 + semantic
const hybridSearchWrapper = async (query: string, k?: number) => {
// Get BM25 results (always available after ingestion)
const bm25Results = searchBM25(query, (k ?? 10) * 3);
// Get semantic results if embeddings are ready
let semanticResults: any[] = [];
if (isEmbeddingComplete) {
try {
semanticResults = await doSemanticSearch(kuzu.executeQuery, query, (k ?? 10) * 3, 0.5);
semanticResults = await doSemanticSearch(lbug.executeQuery, query, (k ?? 10) * 3, 0.5);
} catch {
// Semantic search failed, continue with BM25 only
}
}
// Merge with RRF
return mergeWithRRF(bm25Results, semanticResults, k ?? 10);
};
@@ -586,7 +586,7 @@ const workerApi = {
let codebaseContext;
try {
codebaseContext = await buildCodebaseContext(kuzu.executeQuery, resolvedProjectName);
codebaseContext = await buildCodebaseContext(lbug.executeQuery, resolvedProjectName);
if (import.meta.env.DEV) {
console.log('📊 Codebase context built:', {
files: codebaseContext.stats.fileCount,
@@ -600,7 +600,7 @@ const workerApi = {
currentAgent = createGraphRAGAgent(
config,
kuzu.executeQuery,
lbug.executeQuery,
semanticSearchWrapper,
semanticSearchWithContextWrapper,
hybridSearchWrapper,
@@ -627,7 +627,7 @@ const workerApi = {
/**
* Initialize the Graph RAG agent in backend mode (HTTP-backed tools).
* Uses HTTP wrappers instead of local KuzuDB for all tool queries.
* Uses HTTP wrappers instead of local LadybugDB for all tool queries.
* @param config - Provider configuration for the LLM
* @param backendUrl - Base URL of the gitnexus serve backend
* @param repoName - Repository name on the backend
@@ -848,9 +848,9 @@ const workerApi = {
}
});
// Update KuzuDB with new data
// Update LadybugDB with new data
try {
const kuzu = await getKuzuAdapter();
const lbug = await getLbugAdapter();
onProgress(enrichments.size, enrichments.size); // Done
@@ -872,11 +872,11 @@ const workerApi = {
c.enrichedBy = "llm"
`;
await kuzu.executeQuery(query);
await lbug.executeQuery(query);
}
} catch (err) {
console.error('Failed to update KuzuDB with enrichment:', err);
console.error('Failed to update LadybugDB with enrichment:', err);
}
// Convert Map to Record for serialization
+5 -5
View File
@@ -12,11 +12,11 @@ export default defineConfig({
tailwindcss(),
wasm(),
topLevelAwait(),
// Copy kuzu-wasm worker file to assets folder for production
// Copy lbug-wasm worker file to assets folder for production
viteStaticCopy({
targets: [
{
src: 'node_modules/kuzu-wasm/kuzu_wasm_worker.js',
src: 'node_modules/@ladybugdb/wasm-core/lbug_wasm_worker.js',
dest: 'assets'
}
]
@@ -35,12 +35,12 @@ export default defineConfig({
define: {
global: 'globalThis',
},
// Optimize deps - exclude kuzu-wasm from pre-bundling (it has WASM files)
// Optimize deps - exclude lbug-wasm from pre-bundling (it has WASM files)
optimizeDeps: {
exclude: ['kuzu-wasm'],
exclude: ['@ladybugdb/wasm-core'],
include: ['buffer'],
},
// Required for KuzuDB WASM (SharedArrayBuffer needs Cross-Origin Isolation)
// Required for LadybugDB WASM (SharedArrayBuffer needs Cross-Origin Isolation)
server: {
headers: {
'Cross-Origin-Opener-Policy': 'same-origin',
+7
View File
@@ -0,0 +1,7 @@
{
"permissions": {
"allow": [
"mcp__plugin_claude-mem_mcp-search__get_observations"
]
}
}
+45
View File
@@ -0,0 +1,45 @@
# Changelog
All notable changes to GitNexus will be documented in this file.
## [1.4.5] - 2026-03-17
### Added
- **Ruby language support** for CLI and web (#111)
- **TypeEnvironment API** with constructor inference, self/this/super resolution (#274)
- **Return type inference** with doc-comment parsing (JSDoc, PHPDoc, YARD) and per-language type extractors (#284)
- **Phase 4 type resolution** — nullable unwrapping, for-loop typing, assignment chain propagation (#310)
- **Phase 5 type resolution** — chained calls, pattern matching, class-as-receiver (#315)
- **Phase 6 type resolution** — for-loop Tier 1c, pattern matching, container descriptors, 10-language coverage (#318)
- Container descriptor table for generic type argument resolution (Map keys vs values)
- Method-aware for-loop extractors with integration tests for all languages
- Recursive pattern binding (C# `is` patterns, Kotlin `when/is` smart casts)
- Class field declaration unwrapping for C#/Java
- PHP `$this->property` foreach member access
- C++ pointer dereference range-for
- Java `this.data.values()` field access patterns
- Position-indexed when/is bindings for branch-local narrowing
- **Type resolution system documentation** with architecture guide and roadmap
- `.gitignore` and `.gitnexusignore` support during file discovery (#231)
- Codex MCP configuration documentation in README (#236)
- `skipGraphPhases` pipeline option to skip MRO/community/process phases for faster test runs
- `hookTimeout: 120000` in vitest config for CI beforeAll hooks
### Changed
- **Migrated from KuzuDB to LadybugDB v0.15** (#275)
- Dynamically discover and install agent skills in CLI (#270)
### Performance
- Worker pool threshold — skip worker creation for small repos (<15 files or <512KB total)
- AST walk pruning via `SKIP_SUBTREE_TYPES` for leaf-only nodes (string, comment, number literals)
- Pre-computed `interestingNodeTypes` set — single Set.has() replaces 3 checks per AST node
- `fastStripNullable` — skip full nullable parsing for simple identifiers (90%+ case)
- Replace `.children?.find()` with manual for loops in `extractFunctionName` to eliminate array allocations
### Fixed
- Same-directory Python import resolution (#328)
- Ruby method-level call resolution, HAS_METHOD edges, and dispatch table (#278)
- C++ fixture file casing for case-sensitive CI
- Template string incorrectly included in AST pruning set (contains interpolated expressions)
## [1.4.0] - Previous release
+9
View File
@@ -0,0 +1,9 @@
FROM node:22-bookworm
WORKDIR /app
RUN apt-get update && apt-get install -y python3 make g++ && rm -rf /var/lib/apt/lists/*
COPY . .
RUN npm ci --ignore-scripts \
&& node scripts/patch-tree-sitter-swift.cjs \
&& (npm rebuild 2>&1 || true) \
&& cd node_modules/tree-sitter-kotlin && npx --yes node-gyp rebuild 2>&1
CMD ["npx", "vitest", "run", "test/integration", "--reporter=verbose"]
+24 -3
View File
@@ -96,7 +96,7 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
5. **Processes** — Traces execution flows from entry points through call chains
6. **Search** — Builds hybrid search indexes for fast retrieval
The result is a **KuzuDB graph database** stored locally in `.gitnexus/` with full-text search and semantic embeddings.
The result is a **LadybugDB graph database** stored locally in `.gitnexus/` with full-text search and semantic embeddings.
## MCP Tools
@@ -139,7 +139,8 @@ Your AI agent gets these tools automatically:
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
gitnexus serve # Start local HTTP server (multi-repo) for web UI
gitnexus list # List all indexed repositories
@@ -156,7 +157,27 @@ GitNexus supports indexing multiple repositories. Each `gitnexus analyze` regist
## Supported Languages
TypeScript, JavaScript, Python, Java, C, C++, C#, Go, Rust, PHP, Swift
TypeScript, JavaScript, Python, Java, C, C++, C#, Go, Rust, PHP, Kotlin, Swift, Ruby
### Language Feature Matrix
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
|----------|---------|----------------|---------|----------|-----------------|---------------------|--------|------------|-------------|
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
## Agent Skills
+154 -71
View File
@@ -2,8 +2,10 @@
/**
* GitNexus Claude Code Hook
*
* PreToolUse handler — intercepts Grep/Glob/Bash searches
* and augments with graph context from the GitNexus index.
* PreToolUse — intercepts Grep/Glob/Bash searches and augments
* with graph context from the GitNexus index.
* PostToolUse — detects stale index after git mutations and notifies
* the agent to reindex.
*
* NOTE: SessionStart hooks are broken on Windows (Claude Code bug).
* Session context is injected via CLAUDE.md / skills instead.
@@ -11,7 +13,7 @@
const fs = require('fs');
const path = require('path');
const { execFileSync } = require('child_process');
const { spawnSync } = require('child_process');
/**
* Read JSON input from stdin synchronously.
@@ -26,19 +28,19 @@ function readInput() {
}
/**
* Check if a directory (or ancestor) has a .gitnexus index.
* Find the .gitnexus directory by walking up from startDir.
* Returns the path to .gitnexus/ or null if not found.
*/
function findGitNexusIndex(startDir) {
function findGitNexusDir(startDir) {
let dir = startDir || process.cwd();
for (let i = 0; i < 5; i++) {
if (fs.existsSync(path.join(dir, '.gitnexus'))) {
return true;
}
const candidate = path.join(dir, '.gitnexus');
if (fs.existsSync(candidate)) return candidate;
const parent = path.dirname(dir);
if (parent === dir) break;
dir = parent;
}
return false;
return null;
}
/**
@@ -83,72 +85,153 @@ function extractPattern(toolName, toolInput) {
return null;
}
/**
* Resolve the gitnexus CLI path.
* 1. Relative path (works when script is inside npm package)
* 2. require.resolve (works when gitnexus is globally installed)
* 3. Fall back to npx (returns empty string)
*/
function resolveCliPath() {
let cliPath = path.resolve(__dirname, '..', '..', 'dist', 'cli', 'index.js');
if (!fs.existsSync(cliPath)) {
try {
cliPath = require.resolve('gitnexus/dist/cli/index.js');
} catch {
cliPath = '';
}
}
return cliPath;
}
/**
* Spawn a gitnexus CLI command synchronously.
* Returns the stderr output (KuzuDB captures stdout at OS level).
*/
function runGitNexusCli(cliPath, args, cwd, timeout) {
const isWin = process.platform === 'win32';
if (cliPath) {
return spawnSync(
process.execPath,
[cliPath, ...args],
{ encoding: 'utf-8', timeout, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
}
// On Windows, invoke npx.cmd directly (no shell needed)
return spawnSync(
isWin ? 'npx.cmd' : 'npx',
['-y', 'gitnexus', ...args],
{ encoding: 'utf-8', timeout: timeout + 5000, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
}
/**
* PreToolUse handler — augment searches with graph context.
*/
function handlePreToolUse(input) {
const cwd = input.cwd || process.cwd();
if (!path.isAbsolute(cwd)) return;
if (!findGitNexusDir(cwd)) return;
const toolName = input.tool_name || '';
const toolInput = input.tool_input || {};
if (toolName !== 'Grep' && toolName !== 'Glob' && toolName !== 'Bash') return;
const pattern = extractPattern(toolName, toolInput);
if (!pattern || pattern.length < 3) return;
const cliPath = resolveCliPath();
let result = '';
try {
const child = runGitNexusCli(cliPath, ['augment', '--', pattern], cwd, 7000);
if (!child.error && child.status === 0) {
result = child.stderr || '';
}
} catch { /* graceful failure */ }
if (result && result.trim()) {
sendHookResponse('PreToolUse', result.trim());
}
}
/**
* Emit a PostToolUse hook response with additional context for the agent.
*/
function sendHookResponse(hookEventName, message) {
console.log(JSON.stringify({
hookSpecificOutput: { hookEventName, additionalContext: message }
}));
}
/**
* PostToolUse handler — detect index staleness after git mutations.
*
* Instead of spawning a full `gitnexus analyze` synchronously (which blocks
* the agent for up to 120s and risks KuzuDB corruption on timeout), we do a
* lightweight staleness check: compare `git rev-parse HEAD` against the
* lastCommit stored in `.gitnexus/meta.json`. If they differ, notify the
* agent so it can decide when to reindex.
*/
function handlePostToolUse(input) {
const toolName = input.tool_name || '';
if (toolName !== 'Bash') return;
const command = (input.tool_input || {}).command || '';
if (!/\bgit\s+(commit|merge|rebase|cherry-pick|pull)(\s|$)/.test(command)) return;
// Only proceed if the command succeeded
const toolOutput = input.tool_output || {};
if (toolOutput.exit_code !== undefined && toolOutput.exit_code !== 0) return;
const cwd = input.cwd || process.cwd();
if (!path.isAbsolute(cwd)) return;
const gitNexusDir = findGitNexusDir(cwd);
if (!gitNexusDir) return;
// Compare HEAD against last indexed commit — skip if unchanged
let currentHead = '';
try {
const headResult = spawnSync('git', ['rev-parse', 'HEAD'], {
encoding: 'utf-8', timeout: 3000, cwd, stdio: ['pipe', 'pipe', 'pipe'],
});
currentHead = (headResult.stdout || '').trim();
} catch { return; }
if (!currentHead) return;
let lastCommit = '';
let hadEmbeddings = false;
try {
const meta = JSON.parse(fs.readFileSync(path.join(gitNexusDir, 'meta.json'), 'utf-8'));
lastCommit = meta.lastCommit || '';
hadEmbeddings = (meta.stats && meta.stats.embeddings > 0);
} catch { /* no meta — treat as stale */ }
// If HEAD matches last indexed commit, no reindex needed
if (currentHead && currentHead === lastCommit) return;
const analyzeCmd = `npx gitnexus analyze${hadEmbeddings ? ' --embeddings' : ''}`;
sendHookResponse('PostToolUse',
`GitNexus index is stale (last indexed: ${lastCommit ? lastCommit.slice(0, 7) : 'never'}). ` +
`Run \`${analyzeCmd}\` to update the knowledge graph.`
);
}
// Dispatch map for hook events
const handlers = {
PreToolUse: handlePreToolUse,
PostToolUse: handlePostToolUse,
};
function main() {
try {
const input = readInput();
const hookEvent = input.hook_event_name || '';
if (hookEvent !== 'PreToolUse') return;
const cwd = input.cwd || process.cwd();
if (!findGitNexusIndex(cwd)) return;
const toolName = input.tool_name || '';
const toolInput = input.tool_input || {};
if (toolName !== 'Grep' && toolName !== 'Glob' && toolName !== 'Bash') return;
const pattern = extractPattern(toolName, toolInput);
if (!pattern || pattern.length < 3) return;
// Resolve CLI path — try multiple strategies:
// 1. Relative path (works when script is inside npm package)
// 2. require.resolve (works when gitnexus is globally installed)
// 3. Fall back to npx (works when neither is available)
let cliPath = path.resolve(__dirname, '..', '..', 'dist', 'cli', 'index.js');
if (!fs.existsSync(cliPath)) {
try {
cliPath = require.resolve('gitnexus/dist/cli/index.js');
} catch {
cliPath = ''; // will use npx fallback
}
}
// augment CLI writes result to stderr (KuzuDB's native module captures
// stdout fd at OS level, making it unusable in subprocess contexts).
const { spawnSync } = require('child_process');
let result = '';
try {
let child;
if (cliPath) {
child = spawnSync(
process.execPath,
[cliPath, 'augment', pattern],
{ encoding: 'utf-8', timeout: 8000, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
} else {
// npx fallback
const cmd = process.platform === 'win32' ? 'npx.cmd' : 'npx';
child = spawnSync(
cmd,
['-y', 'gitnexus', 'augment', pattern],
{ encoding: 'utf-8', timeout: 15000, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
}
result = child.stderr || '';
} catch { /* graceful failure */ }
if (result && result.trim()) {
console.log(JSON.stringify({
hookSpecificOutput: {
hookEventName: 'PreToolUse',
additionalContext: result.trim()
}
}));
}
const handler = handlers[input.hook_event_name || ''];
if (handler) handler(input);
} catch (err) {
// Graceful failure — log to stderr for debugging
console.error('GitNexus hook error:', err.message);
if (process.env.GITNEXUS_DEBUG) {
console.error('GitNexus hook error:', (err.message || '').slice(0, 200));
}
}
}
+174 -503
View File
@@ -1,16 +1,17 @@
{
"name": "gitnexus",
"version": "1.3.8",
"version": "1.4.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "gitnexus",
"version": "1.3.8",
"version": "1.4.0",
"hasInstallScript": true,
"license": "PolyForm-Noncommercial-1.0.0",
"dependencies": {
"@huggingface/transformers": "^3.0.0",
"@ladybugdb/core": "^0.15.1",
"@modelcontextprotocol/sdk": "^1.0.0",
"cli-progress": "^3.12.0",
"commander": "^12.0.0",
@@ -20,7 +21,7 @@
"graphology": "^0.25.4",
"graphology-indices": "^0.17.0",
"graphology-utils": "^2.3.0",
"kuzu": "^0.11.3",
"ignore": "^7.0.5",
"lru-cache": "^11.0.0",
"mnemonist": "^0.39.0",
"pandemonium": "^2.4.0",
@@ -31,10 +32,11 @@
"tree-sitter-go": "^0.21.0",
"tree-sitter-java": "^0.21.0",
"tree-sitter-javascript": "^0.21.0",
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-php": "^0.23.12",
"tree-sitter-python": "^0.21.0",
"tree-sitter-ruby": "^0.23.1",
"tree-sitter-rust": "^0.21.0",
"tree-sitter-swift": "^0.6.0",
"tree-sitter-typescript": "^0.21.0",
"uuid": "^13.0.0"
},
@@ -56,6 +58,7 @@
"node": ">=18.0.0"
},
"optionalDependencies": {
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-swift": "^0.6.0"
}
},
@@ -1147,6 +1150,133 @@
"@jridgewell/sourcemap-codec": "^1.4.14"
}
},
"node_modules/@ladybugdb/core": {
"version": "0.15.1",
"resolved": "https://registry.npmjs.org/@ladybugdb/core/-/core-0.15.1.tgz",
"integrity": "sha512-a+jhzIlS2+57Y2YWXlta7Dq5A3577dQ8YO7DzPCFZxozeiGIZn0K9v0ROO+ws4PW9BwuQYI5BXQxTEtaa1Otlg==",
"hasInstallScript": true,
"license": "MIT",
"dependencies": {
"cmake-js": "^8.0.0",
"node-addon-api": "^6.0.0"
}
},
"node_modules/@ladybugdb/core/node_modules/chownr": {
"version": "3.0.0",
"resolved": "https://registry.npmjs.org/chownr/-/chownr-3.0.0.tgz",
"integrity": "sha512-+IxzY9BZOQd/XuYPRmrvEVjF/nqj5kgT4kEq7VofrDoM1MxoRjEWkrCC3EtLi59TVawxTAn+orJwFQcrqEN1+g==",
"license": "BlueOak-1.0.0",
"engines": {
"node": ">=18"
}
},
"node_modules/@ladybugdb/core/node_modules/cmake-js": {
"version": "8.0.0",
"resolved": "https://registry.npmjs.org/cmake-js/-/cmake-js-8.0.0.tgz",
"integrity": "sha512-YbUP88RDwCvoQkZhRtGURYm9RIpWdtvZuhT87fKNoLjk8kIFIFeARpKfuZQGdwfH99GZpUmqSfcDrK62X7lTgg==",
"license": "MIT",
"dependencies": {
"debug": "^4.4.3",
"fs-extra": "^11.3.3",
"node-api-headers": "^1.8.0",
"rc": "1.2.8",
"semver": "^7.7.3",
"tar": "^7.5.6",
"url-join": "^4.0.1",
"which": "^6.0.0",
"yargs": "^17.7.2"
},
"bin": {
"cmake-js": "bin/cmake-js"
},
"engines": {
"node": "^20.17.0 || >=22.9.0"
}
},
"node_modules/@ladybugdb/core/node_modules/debug": {
"version": "4.4.3",
"resolved": "https://registry.npmjs.org/debug/-/debug-4.4.3.tgz",
"integrity": "sha512-RGwwWnwQvkVfavKVt22FGLw+xYSdzARwm0ru6DhTVA3umU5hZc28V3kO4stgYryrTlLpuvgI9GiijltAjNbcqA==",
"license": "MIT",
"dependencies": {
"ms": "^2.1.3"
},
"engines": {
"node": ">=6.0"
},
"peerDependenciesMeta": {
"supports-color": {
"optional": true
}
}
},
"node_modules/@ladybugdb/core/node_modules/isexe": {
"version": "4.0.0",
"resolved": "https://registry.npmjs.org/isexe/-/isexe-4.0.0.tgz",
"integrity": "sha512-FFUtZMpoZ8RqHS3XeXEmHWLA4thH+ZxCv2lOiPIn1Xc7CxrqhWzNSDzD+/chS/zbYezmiwWLdQC09JdQKmthOw==",
"license": "BlueOak-1.0.0",
"engines": {
"node": ">=20"
}
},
"node_modules/@ladybugdb/core/node_modules/minizlib": {
"version": "3.1.0",
"resolved": "https://registry.npmjs.org/minizlib/-/minizlib-3.1.0.tgz",
"integrity": "sha512-KZxYo1BUkWD2TVFLr0MQoM8vUUigWD3LlD83a/75BqC+4qE0Hb1Vo5v1FgcfaNXvfXzr+5EhQ6ing/CaBijTlw==",
"license": "MIT",
"dependencies": {
"minipass": "^7.1.2"
},
"engines": {
"node": ">= 18"
}
},
"node_modules/@ladybugdb/core/node_modules/ms": {
"version": "2.1.3",
"resolved": "https://registry.npmjs.org/ms/-/ms-2.1.3.tgz",
"integrity": "sha512-6FlzubTLZG3J2a/NVCAleEhjzq5oxgHyaCU9yYXvcLsvoVaHJq/s5xXI6/XXP6tz7R9xAOtHnSO/tXtF3WRTlA==",
"license": "MIT"
},
"node_modules/@ladybugdb/core/node_modules/tar": {
"version": "7.5.11",
"resolved": "https://registry.npmjs.org/tar/-/tar-7.5.11.tgz",
"integrity": "sha512-ChjMH33/KetonMTAtpYdgUFr0tbz69Fp2v7zWxQfYZX4g5ZN2nOBXm1R2xyA+lMIKrLKIoKAwFj93jE/avX9cQ==",
"license": "BlueOak-1.0.0",
"dependencies": {
"@isaacs/fs-minipass": "^4.0.0",
"chownr": "^3.0.0",
"minipass": "^7.1.2",
"minizlib": "^3.1.0",
"yallist": "^5.0.0"
},
"engines": {
"node": ">=18"
}
},
"node_modules/@ladybugdb/core/node_modules/which": {
"version": "6.0.1",
"resolved": "https://registry.npmjs.org/which/-/which-6.0.1.tgz",
"integrity": "sha512-oGLe46MIrCRqX7ytPUf66EAYvdeMIZYn3WaocqqKZAxrBpkqHfL/qvTyJ/bTk5+AqHCjXmrv3CEWgy368zhRUg==",
"license": "ISC",
"dependencies": {
"isexe": "^4.0.0"
},
"bin": {
"node-which": "bin/which.js"
},
"engines": {
"node": "^20.17.0 || >=22.9.0"
}
},
"node_modules/@ladybugdb/core/node_modules/yallist": {
"version": "5.0.0",
"resolved": "https://registry.npmjs.org/yallist/-/yallist-5.0.0.tgz",
"integrity": "sha512-YgvUTfwqyc7UXVMrB+SImsVYSmTS8X/tSrtdNZMImM+n7+QTriRXyXim0mBrTXNeqzVF0KWGgHPeiyViFFrNDw==",
"license": "BlueOak-1.0.0",
"engines": {
"node": ">=18"
}
},
"node_modules/@modelcontextprotocol/sdk": {
"version": "1.25.3",
"resolved": "https://registry.npmjs.org/@modelcontextprotocol/sdk/-/sdk-1.25.3.tgz",
@@ -2273,26 +2403,6 @@
"url": "https://github.com/chalk/ansi-styles?sponsor=1"
}
},
"node_modules/aproba": {
"version": "2.1.0",
"resolved": "https://registry.npmjs.org/aproba/-/aproba-2.1.0.tgz",
"integrity": "sha512-tLIEcj5GuR2RSTnxNKdkK0dJ/GrC7P38sUkiDmDuHfsHmbagTFAxDVIBltoklXEVIQ/f14IL8IMJ5pn9Hez1Ew==",
"license": "ISC"
},
"node_modules/are-we-there-yet": {
"version": "3.0.1",
"resolved": "https://registry.npmjs.org/are-we-there-yet/-/are-we-there-yet-3.0.1.tgz",
"integrity": "sha512-QZW4EDmGwlYur0Yyf/b2uGucHQMa8aFUP7eu9ddR73vvhFyt4V0Vl3QHPcTNJ8l6qYOBdxgXdnBXQrHilfRQBg==",
"deprecated": "This package is no longer supported.",
"license": "ISC",
"dependencies": {
"delegates": "^1.0.0",
"readable-stream": "^3.6.0"
},
"engines": {
"node": "^12.13.0 || ^14.15.0 || >=16.0.0"
}
},
"node_modules/array-flatten": {
"version": "1.1.1",
"resolved": "https://registry.npmjs.org/array-flatten/-/array-flatten-1.1.1.tgz",
@@ -2321,23 +2431,6 @@
"js-tokens": "^10.0.0"
}
},
"node_modules/asynckit": {
"version": "0.4.0",
"resolved": "https://registry.npmjs.org/asynckit/-/asynckit-0.4.0.tgz",
"integrity": "sha512-Oei9OH4tRh0YqU3GxhX79dM/mwVgvbZJaSNaRk+bshkj0S5cfHcgYakreBjrHwatXKbz+IoIdYLxrKim2MjW0Q==",
"license": "MIT"
},
"node_modules/axios": {
"version": "1.13.4",
"resolved": "https://registry.npmjs.org/axios/-/axios-1.13.4.tgz",
"integrity": "sha512-1wVkUaAO6WyaYtCkcYCOx12ZgpGf9Zif+qXa4n+oYzK558YryKqiL6UWwd5DqiH3VRW0GYhTZQ/vlgJrCoNQlg==",
"license": "MIT",
"dependencies": {
"follow-redirects": "^1.15.6",
"form-data": "^4.0.4",
"proxy-from-env": "^1.1.0"
}
},
"node_modules/body-parser": {
"version": "1.20.4",
"resolved": "https://registry.npmjs.org/body-parser/-/body-parser-1.20.4.tgz",
@@ -2432,15 +2525,6 @@
"node": ">=18"
}
},
"node_modules/chownr": {
"version": "2.0.0",
"resolved": "https://registry.npmjs.org/chownr/-/chownr-2.0.0.tgz",
"integrity": "sha512-bIomtDF5KGpdogkLd9VspvFzk9KfpyyGlS8YFVZl7TGPBHL5snIOnxeshwVgPteQ9b4Eydl+pVbIyE1DcvCWgQ==",
"license": "ISC",
"engines": {
"node": ">=10"
}
},
"node_modules/cli-progress": {
"version": "3.12.0",
"resolved": "https://registry.npmjs.org/cli-progress/-/cli-progress-3.12.0.tgz",
@@ -2581,55 +2665,6 @@
"url": "https://github.com/chalk/wrap-ansi?sponsor=1"
}
},
"node_modules/cmake-js": {
"version": "7.4.0",
"resolved": "https://registry.npmjs.org/cmake-js/-/cmake-js-7.4.0.tgz",
"integrity": "sha512-Lw0JxEHrmk+qNj1n9W9d4IvkDdYTBn7l2BW6XmtLj7WPpIo2shvxUy+YokfjMxAAOELNonQwX3stkPhM5xSC2Q==",
"license": "MIT",
"dependencies": {
"axios": "^1.6.5",
"debug": "^4",
"fs-extra": "^11.2.0",
"memory-stream": "^1.0.0",
"node-api-headers": "^1.1.0",
"npmlog": "^6.0.2",
"rc": "^1.2.7",
"semver": "^7.5.4",
"tar": "^6.2.0",
"url-join": "^4.0.1",
"which": "^2.0.2",
"yargs": "^17.7.2"
},
"bin": {
"cmake-js": "bin/cmake-js"
},
"engines": {
"node": ">= 14.15.0"
}
},
"node_modules/cmake-js/node_modules/debug": {
"version": "4.4.3",
"resolved": "https://registry.npmjs.org/debug/-/debug-4.4.3.tgz",
"integrity": "sha512-RGwwWnwQvkVfavKVt22FGLw+xYSdzARwm0ru6DhTVA3umU5hZc28V3kO4stgYryrTlLpuvgI9GiijltAjNbcqA==",
"license": "MIT",
"dependencies": {
"ms": "^2.1.3"
},
"engines": {
"node": ">=6.0"
},
"peerDependenciesMeta": {
"supports-color": {
"optional": true
}
}
},
"node_modules/cmake-js/node_modules/ms": {
"version": "2.1.3",
"resolved": "https://registry.npmjs.org/ms/-/ms-2.1.3.tgz",
"integrity": "sha512-6FlzubTLZG3J2a/NVCAleEhjzq5oxgHyaCU9yYXvcLsvoVaHJq/s5xXI6/XXP6tz7R9xAOtHnSO/tXtF3WRTlA==",
"license": "MIT"
},
"node_modules/color-convert": {
"version": "2.0.1",
"resolved": "https://registry.npmjs.org/color-convert/-/color-convert-2.0.1.tgz",
@@ -2648,27 +2683,6 @@
"integrity": "sha512-dOy+3AuW3a2wNbZHIuMZpTcgjGuLU/uBL/ubcZF9OXbDo8ff4O8yVp5Bf0efS8uEoYo5q4Fx7dY9OgQGXgAsQA==",
"license": "MIT"
},
"node_modules/color-support": {
"version": "1.1.3",
"resolved": "https://registry.npmjs.org/color-support/-/color-support-1.1.3.tgz",
"integrity": "sha512-qiBjkpbMLO/HL68y+lh4q0/O1MZFj2RX6X/KmMa3+gJD3z+WwI1ZzDHysvqHGS3mP6mznPckpXmw1nI9cJjyRg==",
"license": "ISC",
"bin": {
"color-support": "bin.js"
}
},
"node_modules/combined-stream": {
"version": "1.0.8",
"resolved": "https://registry.npmjs.org/combined-stream/-/combined-stream-1.0.8.tgz",
"integrity": "sha512-FQN4MRfuJeHf7cBbBMJFXhKSDq+2kAArBlmRBvcvFE5BB1HZKXtSFASDhdlz9zOYwxh8lDdnvmMOe/+5cdoEdg==",
"license": "MIT",
"dependencies": {
"delayed-stream": "~1.0.0"
},
"engines": {
"node": ">= 0.8"
}
},
"node_modules/commander": {
"version": "12.1.0",
"resolved": "https://registry.npmjs.org/commander/-/commander-12.1.0.tgz",
@@ -2678,12 +2692,6 @@
"node": ">=18"
}
},
"node_modules/console-control-strings": {
"version": "1.1.0",
"resolved": "https://registry.npmjs.org/console-control-strings/-/console-control-strings-1.1.0.tgz",
"integrity": "sha512-ty/fTekppD2fIwRvnZAVdeOiGd1c7YXEixbgJTNzqcxJWKQnjJ/V1bNEEE6hygpM3WjwHFUVK6HTjWSzV4a8sQ==",
"license": "ISC"
},
"node_modules/content-disposition": {
"version": "0.5.4",
"resolved": "https://registry.npmjs.org/content-disposition/-/content-disposition-0.5.4.tgz",
@@ -2803,21 +2811,6 @@
"url": "https://github.com/sponsors/ljharb"
}
},
"node_modules/delayed-stream": {
"version": "1.0.0",
"resolved": "https://registry.npmjs.org/delayed-stream/-/delayed-stream-1.0.0.tgz",
"integrity": "sha512-ZySD7Nf91aLB0RxL4KGrKHBXl7Eds1DAmEdcoVawXnLD7SDhpNgtuII2aAkg7a7QS41jxPSZ17p4VdGnMHk3MQ==",
"license": "MIT",
"engines": {
"node": ">=0.4.0"
}
},
"node_modules/delegates": {
"version": "1.0.0",
"resolved": "https://registry.npmjs.org/delegates/-/delegates-1.0.0.tgz",
"integrity": "sha512-bd2L678uiWATM6m5Z1VzNCErI3jiGzt6HGY8OVICs40JQq/HALfbyNJmp0UDakEY4pMMaN0Ly5om/B1VI/+xfQ==",
"license": "MIT"
},
"node_modules/depd": {
"version": "2.0.0",
"resolved": "https://registry.npmjs.org/depd/-/depd-2.0.0.tgz",
@@ -2930,21 +2923,6 @@
"node": ">= 0.4"
}
},
"node_modules/es-set-tostringtag": {
"version": "2.1.0",
"resolved": "https://registry.npmjs.org/es-set-tostringtag/-/es-set-tostringtag-2.1.0.tgz",
"integrity": "sha512-j6vWzfrGVfyXxge+O0x5sh6cvxAog0a/4Rdd2K36zCMV5eJ+/+tOAngRO8cODMNWbVRdVlmGZQL2YS3yR8bIUA==",
"license": "MIT",
"dependencies": {
"es-errors": "^1.3.0",
"get-intrinsic": "^1.2.6",
"has-tostringtag": "^1.0.2",
"hasown": "^2.0.2"
},
"engines": {
"node": ">= 0.4"
}
},
"node_modules/es6-error": {
"version": "4.1.1",
"resolved": "https://registry.npmjs.org/es6-error/-/es6-error-4.1.1.tgz",
@@ -3204,26 +3182,6 @@
"integrity": "sha512-MI1qs7Lo4Syw0EOzUl0xjs2lsoeqFku44KpngfIduHBYvzm8h2+7K8YMQh1JtVVVrUvhLpNwqVi4DERegUJhPQ==",
"license": "Apache-2.0"
},
"node_modules/follow-redirects": {
"version": "1.15.11",
"resolved": "https://registry.npmjs.org/follow-redirects/-/follow-redirects-1.15.11.tgz",
"integrity": "sha512-deG2P0JfjrTxl50XGCDyfI97ZGVCxIpfKYmfyrQ54n5FO/0gfIES8C/Psl6kWVDolizcaaxZJnTS0QSMxvnsBQ==",
"funding": [
{
"type": "individual",
"url": "https://github.com/sponsors/RubenVerborgh"
}
],
"license": "MIT",
"engines": {
"node": ">=4.0"
},
"peerDependenciesMeta": {
"debug": {
"optional": true
}
}
},
"node_modules/foreground-child": {
"version": "3.3.1",
"resolved": "https://registry.npmjs.org/foreground-child/-/foreground-child-3.3.1.tgz",
@@ -3240,22 +3198,6 @@
"url": "https://github.com/sponsors/isaacs"
}
},
"node_modules/form-data": {
"version": "4.0.5",
"resolved": "https://registry.npmjs.org/form-data/-/form-data-4.0.5.tgz",
"integrity": "sha512-8RipRLol37bNs2bhoV67fiTEvdTrbMUYcFTiy3+wuuOnUog2QBHCZWXDRijWQfAkhBj2Uf5UnVaiWwA5vdd82w==",
"license": "MIT",
"dependencies": {
"asynckit": "^0.4.0",
"combined-stream": "^1.0.8",
"es-set-tostringtag": "^2.1.0",
"hasown": "^2.0.2",
"mime-types": "^2.1.12"
},
"engines": {
"node": ">= 6"
}
},
"node_modules/forwarded": {
"version": "0.2.0",
"resolved": "https://registry.npmjs.org/forwarded/-/forwarded-0.2.0.tgz",
@@ -3288,30 +3230,6 @@
"node": ">=14.14"
}
},
"node_modules/fs-minipass": {
"version": "2.1.0",
"resolved": "https://registry.npmjs.org/fs-minipass/-/fs-minipass-2.1.0.tgz",
"integrity": "sha512-V/JgOLFCS+R6Vcq0slCuaeWEdNC3ouDlJMNIsacH2VtALiu9mV4LPrHc5cDl8k5aw6J8jwgWWpiTo5RYhmIzvg==",
"license": "ISC",
"dependencies": {
"minipass": "^3.0.0"
},
"engines": {
"node": ">= 8"
}
},
"node_modules/fs-minipass/node_modules/minipass": {
"version": "3.3.6",
"resolved": "https://registry.npmjs.org/minipass/-/minipass-3.3.6.tgz",
"integrity": "sha512-DxiNidxSEK+tHG6zOIklvNOwm3hvCrbUrdtzY74U6HKTJxvIDfOUL5W5P2Ghd3DTkhhKPYGqeNUIh5qcM4YBfw==",
"license": "ISC",
"dependencies": {
"yallist": "^4.0.0"
},
"engines": {
"node": ">=8"
}
},
"node_modules/fsevents": {
"version": "2.3.3",
"resolved": "https://registry.npmjs.org/fsevents/-/fsevents-2.3.3.tgz",
@@ -3336,73 +3254,6 @@
"url": "https://github.com/sponsors/ljharb"
}
},
"node_modules/gauge": {
"version": "4.0.4",
"resolved": "https://registry.npmjs.org/gauge/-/gauge-4.0.4.tgz",
"integrity": "sha512-f9m+BEN5jkg6a0fZjleidjN51VE1X+mPFQ2DJ0uv1V39oCLCbsGe6yjbBnp7eK7z/+GAon99a3nHuqbuuthyPg==",
"deprecated": "This package is no longer supported.",
"license": "ISC",
"dependencies": {
"aproba": "^1.0.3 || ^2.0.0",
"color-support": "^1.1.3",
"console-control-strings": "^1.1.0",
"has-unicode": "^2.0.1",
"signal-exit": "^3.0.7",
"string-width": "^4.2.3",
"strip-ansi": "^6.0.1",
"wide-align": "^1.1.5"
},
"engines": {
"node": "^12.13.0 || ^14.15.0 || >=16.0.0"
}
},
"node_modules/gauge/node_modules/ansi-regex": {
"version": "5.0.1",
"resolved": "https://registry.npmjs.org/ansi-regex/-/ansi-regex-5.0.1.tgz",
"integrity": "sha512-quJQXlTSUGL2LH9SUXo8VwsY4soanhgo6LNSm84E1LBcE8s3O0wpdiRzyR9z/ZZJMlMWv37qOOb9pdJlMUEKFQ==",
"license": "MIT",
"engines": {
"node": ">=8"
}
},
"node_modules/gauge/node_modules/emoji-regex": {
"version": "8.0.0",
"resolved": "https://registry.npmjs.org/emoji-regex/-/emoji-regex-8.0.0.tgz",
"integrity": "sha512-MSjYzcWNOA0ewAHpz0MxpYFvwg6yjy1NG3xteoqz644VCo/RPgnr1/GGt+ic3iJTzQ8Eu3TdM14SawnVUmGE6A==",
"license": "MIT"
},
"node_modules/gauge/node_modules/signal-exit": {
"version": "3.0.7",
"resolved": "https://registry.npmjs.org/signal-exit/-/signal-exit-3.0.7.tgz",
"integrity": "sha512-wnD2ZE+l+SPC/uoS0vXeE9L1+0wuaMqKlfz9AMUo38JsyLSBWSFcHR1Rri62LZc12vLr1gb3jl7iwQhgwpAbGQ==",
"license": "ISC"
},
"node_modules/gauge/node_modules/string-width": {
"version": "4.2.3",
"resolved": "https://registry.npmjs.org/string-width/-/string-width-4.2.3.tgz",
"integrity": "sha512-wKyQRQpjJ0sIp62ErSZdGsjMJWsap5oRNihHhu6G7JVO/9jIB6UyevL+tXuOqrng8j/cxKTWyWUwvSTriiZz/g==",
"license": "MIT",
"dependencies": {
"emoji-regex": "^8.0.0",
"is-fullwidth-code-point": "^3.0.0",
"strip-ansi": "^6.0.1"
},
"engines": {
"node": ">=8"
}
},
"node_modules/gauge/node_modules/strip-ansi": {
"version": "6.0.1",
"resolved": "https://registry.npmjs.org/strip-ansi/-/strip-ansi-6.0.1.tgz",
"integrity": "sha512-Y38VPSHcqkFrCpFnQ9vuSXmquuv5oXOKpGeT6aGrr3o3Gc9AlVa6JBfUSOCnbxGGZF+/0ooI7KrPuUSztUdU5A==",
"license": "MIT",
"dependencies": {
"ansi-regex": "^5.0.1"
},
"engines": {
"node": ">=8"
}
},
"node_modules/get-caller-file": {
"version": "2.0.5",
"resolved": "https://registry.npmjs.org/get-caller-file/-/get-caller-file-2.0.5.tgz",
@@ -3618,27 +3469,6 @@
"url": "https://github.com/sponsors/ljharb"
}
},
"node_modules/has-tostringtag": {
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/has-tostringtag/-/has-tostringtag-1.0.2.tgz",
"integrity": "sha512-NqADB8VjPFLM2V0VvHUewwwsw0ZWBaIdgo+ieHtK3hasLz4qeCRjYcqfB6AQrBggRKppKF8L52/VqdVsO47Dlw==",
"license": "MIT",
"dependencies": {
"has-symbols": "^1.0.3"
},
"engines": {
"node": ">= 0.4"
},
"funding": {
"url": "https://github.com/sponsors/ljharb"
}
},
"node_modules/has-unicode": {
"version": "2.0.1",
"resolved": "https://registry.npmjs.org/has-unicode/-/has-unicode-2.0.1.tgz",
"integrity": "sha512-8Rf9Y83NBReMnx0gFzA8JImQACstCYWUplepDa9xprwwtmgEZUF0h/i5xSA625zB/I37EtrswSST6OXxwaaIJQ==",
"license": "ISC"
},
"node_modules/hasown": {
"version": "2.0.2",
"resolved": "https://registry.npmjs.org/hasown/-/hasown-2.0.2.tgz",
@@ -3700,6 +3530,15 @@
"node": ">=0.10.0"
}
},
"node_modules/ignore": {
"version": "7.0.5",
"resolved": "https://registry.npmjs.org/ignore/-/ignore-7.0.5.tgz",
"integrity": "sha512-Hs59xBNfUIunMFgWAbGX5cq6893IbWg4KnrjbYwX3tx0ztorVgTDA6B2sxf8ejHJ4wz8BqGUMYlnzNBer5NvGg==",
"license": "MIT",
"engines": {
"node": ">= 4"
}
},
"node_modules/inherits": {
"version": "2.0.4",
"resolved": "https://registry.npmjs.org/inherits/-/inherits-2.0.4.tgz",
@@ -3842,18 +3681,6 @@
"graceful-fs": "^4.1.6"
}
},
"node_modules/kuzu": {
"version": "0.11.3",
"resolved": "https://registry.npmjs.org/kuzu/-/kuzu-0.11.3.tgz",
"integrity": "sha512-4+hD3Y+YMV3e0uiqTv1/GUal47D04l8qluw1WFWg8Nx3k7rLsHG1Pmq9WHIOlf1742svxQvTYQiuY6oS1qxAZA==",
"deprecated": "Package no longer supported. Contact Support at https://www.npmjs.com/support for more info.",
"hasInstallScript": true,
"license": "MIT",
"dependencies": {
"cmake-js": "^7.3.0",
"node-addon-api": "^6.0.0"
}
},
"node_modules/long": {
"version": "5.3.2",
"resolved": "https://registry.npmjs.org/long/-/long-5.3.2.tgz",
@@ -3937,15 +3764,6 @@
"node": ">= 0.6"
}
},
"node_modules/memory-stream": {
"version": "1.0.0",
"resolved": "https://registry.npmjs.org/memory-stream/-/memory-stream-1.0.0.tgz",
"integrity": "sha512-Wm13VcsPIMdG96dzILfij09PvuS3APtcKNh7M28FsCA/w6+1mjR7hhPmfFNoilX9xU7wTdhsH5lJAm6XNzdtww==",
"license": "MIT",
"dependencies": {
"readable-stream": "^3.4.0"
}
},
"node_modules/merge-descriptors": {
"version": "1.0.3",
"resolved": "https://registry.npmjs.org/merge-descriptors/-/merge-descriptors-1.0.3.tgz",
@@ -4030,43 +3848,6 @@
"node": ">=16 || 14 >=14.17"
}
},
"node_modules/minizlib": {
"version": "2.1.2",
"resolved": "https://registry.npmjs.org/minizlib/-/minizlib-2.1.2.tgz",
"integrity": "sha512-bAxsR8BVfj60DWXHE3u30oHzfl4G7khkSuPW+qvpd7jFRHm7dLxOjUk1EHACJ/hxLY8phGJ0YhYHZo7jil7Qdg==",
"license": "MIT",
"dependencies": {
"minipass": "^3.0.0",
"yallist": "^4.0.0"
},
"engines": {
"node": ">= 8"
}
},
"node_modules/minizlib/node_modules/minipass": {
"version": "3.3.6",
"resolved": "https://registry.npmjs.org/minipass/-/minipass-3.3.6.tgz",
"integrity": "sha512-DxiNidxSEK+tHG6zOIklvNOwm3hvCrbUrdtzY74U6HKTJxvIDfOUL5W5P2Ghd3DTkhhKPYGqeNUIh5qcM4YBfw==",
"license": "ISC",
"dependencies": {
"yallist": "^4.0.0"
},
"engines": {
"node": ">=8"
}
},
"node_modules/mkdirp": {
"version": "1.0.4",
"resolved": "https://registry.npmjs.org/mkdirp/-/mkdirp-1.0.4.tgz",
"integrity": "sha512-vVqVZQyf3WLx2Shd0qJ9xuvqgAyKPLAiqITEtqW0oIUjzo3PePDd6fW9iFz30ef7Ysp/oiWqbhszeGWW2T6Gzw==",
"license": "MIT",
"bin": {
"mkdirp": "bin/cmd.js"
},
"engines": {
"node": ">=10"
}
},
"node_modules/mnemonist": {
"version": "0.39.8",
"resolved": "https://registry.npmjs.org/mnemonist/-/mnemonist-0.39.8.tgz",
@@ -4133,22 +3914,6 @@
"node-gyp-build-test": "build-test.js"
}
},
"node_modules/npmlog": {
"version": "6.0.2",
"resolved": "https://registry.npmjs.org/npmlog/-/npmlog-6.0.2.tgz",
"integrity": "sha512-/vBvz5Jfr9dT/aFWd0FIRf+T/Q2WBsLENygUaFUqstqsycmZAP/t5BvFJTK0viFmSUxiUKTUplWy5vt+rvKIxg==",
"deprecated": "This package is no longer supported.",
"license": "ISC",
"dependencies": {
"are-we-there-yet": "^3.0.0",
"console-control-strings": "^1.1.0",
"gauge": "^4.0.3",
"set-blocking": "^2.0.0"
},
"engines": {
"node": "^12.13.0 || ^14.15.0 || >=16.0.0"
}
},
"node_modules/object-assign": {
"version": "4.1.1",
"resolved": "https://registry.npmjs.org/object-assign/-/object-assign-4.1.1.tgz",
@@ -4469,12 +4234,6 @@
"node": ">= 0.10"
}
},
"node_modules/proxy-from-env": {
"version": "1.1.0",
"resolved": "https://registry.npmjs.org/proxy-from-env/-/proxy-from-env-1.1.0.tgz",
"integrity": "sha512-D+zkORCbA9f1tdWRK0RaCR3GPv50cMxcrz4X8k5LTSUD1Dkw47mKJEZQNunItRTkWwgtaUSo1RVFRIG9ZXiFYg==",
"license": "MIT"
},
"node_modules/qs": {
"version": "6.14.1",
"resolved": "https://registry.npmjs.org/qs/-/qs-6.14.1.tgz",
@@ -4545,20 +4304,6 @@
"rc": "cli.js"
}
},
"node_modules/readable-stream": {
"version": "3.6.2",
"resolved": "https://registry.npmjs.org/readable-stream/-/readable-stream-3.6.2.tgz",
"integrity": "sha512-9u/sniCrY3D5WdsERHzHE4G2YCXqoG5FTHUiCC4SIbr6XcLZBY05ya9EKjYek9O5xOAwjGq+1JdGBAS7Q9ScoA==",
"license": "MIT",
"dependencies": {
"inherits": "^2.0.3",
"string_decoder": "^1.1.1",
"util-deprecate": "^1.0.1"
},
"engines": {
"node": ">= 6"
}
},
"node_modules/require-directory": {
"version": "2.1.1",
"resolved": "https://registry.npmjs.org/require-directory/-/require-directory-2.1.1.tgz",
@@ -4802,12 +4547,6 @@
"node": ">= 0.8.0"
}
},
"node_modules/set-blocking": {
"version": "2.0.0",
"resolved": "https://registry.npmjs.org/set-blocking/-/set-blocking-2.0.0.tgz",
"integrity": "sha512-KiKBS8AnWGEyLzofFfmvKwpdPzqiy16LvQfK3yv/fVH7Bj13/wl3JSR1J+rfgRE9q7xUJK4qvgS8raSOeLUehw==",
"license": "ISC"
},
"node_modules/setprototypeof": {
"version": "1.2.0",
"resolved": "https://registry.npmjs.org/setprototypeof/-/setprototypeof-1.2.0.tgz",
@@ -5009,15 +4748,6 @@
"dev": true,
"license": "MIT"
},
"node_modules/string_decoder": {
"version": "1.3.0",
"resolved": "https://registry.npmjs.org/string_decoder/-/string_decoder-1.3.0.tgz",
"integrity": "sha512-hkRX8U1WjJFd8LsDJ2yQ/wWWxaopEsABU1XfkM8A+j0+85JAGppt16cr1Whg6KIbb4okU6Mql6BOj+uup/wKeA==",
"license": "MIT",
"dependencies": {
"safe-buffer": "~5.2.0"
}
},
"node_modules/string-width": {
"version": "5.1.2",
"resolved": "https://registry.npmjs.org/string-width/-/string-width-5.1.2.tgz",
@@ -5136,33 +4866,6 @@
"node": ">=8"
}
},
"node_modules/tar": {
"version": "6.2.1",
"resolved": "https://registry.npmjs.org/tar/-/tar-6.2.1.tgz",
"integrity": "sha512-DZ4yORTwrbTj/7MZYq2w+/ZFdI6OZ/f9SFHR+71gIVUZhOQPHzVCLpvRnPgyaMpfWxxk/4ONva3GQSyNIKRv6A==",
"deprecated": "Old versions of tar are not supported, and contain widely publicized security vulnerabilities, which have been fixed in the current version. Please update. Support for old versions may be purchased (at exhorbitant rates) by contacting i@izs.me",
"license": "ISC",
"dependencies": {
"chownr": "^2.0.0",
"fs-minipass": "^2.0.0",
"minipass": "^5.0.0",
"minizlib": "^2.1.1",
"mkdirp": "^1.0.3",
"yallist": "^4.0.0"
},
"engines": {
"node": ">=10"
}
},
"node_modules/tar/node_modules/minipass": {
"version": "5.0.0",
"resolved": "https://registry.npmjs.org/minipass/-/minipass-5.0.0.tgz",
"integrity": "sha512-3FnjYuehv9k6ovOEbyOswadCDPX1piCfhV8ncmYtHOjuPwylVWsghTLo7rabjC3Rx5xD4HDx8Wm1xnMF7S5qFQ==",
"license": "ISC",
"engines": {
"node": ">=8"
}
},
"node_modules/tinybench": {
"version": "2.9.0",
"resolved": "https://registry.npmjs.org/tinybench/-/tinybench-2.9.0.tgz",
@@ -5415,6 +5118,7 @@
"integrity": "sha512-A4obq6bjzmYrA+F0JLLoheFPcofFkctNaZSpnDd+GPn1SfVZLY4/GG4C0cYVBTOShuPBGGAOPLM1JWLZQV4m1g==",
"hasInstallScript": true,
"license": "MIT",
"optional": true,
"dependencies": {
"node-addon-api": "^7.1.0",
"node-gyp-build": "^4.8.0"
@@ -5432,7 +5136,8 @@
"version": "7.1.1",
"resolved": "https://registry.npmjs.org/node-addon-api/-/node-addon-api-7.1.1.tgz",
"integrity": "sha512-5m3bsyrjFWE1xf7nz7YXdN4udnVtXK6/Yfgn5qnahL6bCkf2yKt4k3nuTKAtT4r3IG8JNR2ncsIMdZuAzJjHQQ==",
"license": "MIT"
"license": "MIT",
"optional": true
},
"node_modules/tree-sitter-php": {
"version": "0.23.12",
@@ -5487,6 +5192,34 @@
"integrity": "sha512-5m3bsyrjFWE1xf7nz7YXdN4udnVtXK6/Yfgn5qnahL6bCkf2yKt4k3nuTKAtT4r3IG8JNR2ncsIMdZuAzJjHQQ==",
"license": "MIT"
},
"node_modules/tree-sitter-ruby": {
"version": "0.23.1",
"resolved": "https://registry.npmjs.org/tree-sitter-ruby/-/tree-sitter-ruby-0.23.1.tgz",
"integrity": "sha512-d9/RXgWjR6HanN7wTYhS5bpBQLz1VkH048Vm3CodPGyJVnamXMGb8oEhDypVCBq4QnHui9sTXuJBBP3WtCw5RA==",
"hasInstallScript": true,
"license": "MIT",
"dependencies": {
"node-addon-api": "^8.2.2",
"node-gyp-build": "^4.8.2"
},
"peerDependencies": {
"tree-sitter": "^0.21.1"
},
"peerDependenciesMeta": {
"tree-sitter": {
"optional": true
}
}
},
"node_modules/tree-sitter-ruby/node_modules/node-addon-api": {
"version": "8.6.0",
"resolved": "https://registry.npmjs.org/node-addon-api/-/node-addon-api-8.6.0.tgz",
"integrity": "sha512-gBVjCaqDlRUk0EwoPNKzIr9KkS9041G/q31IBShPs1Xz6UTA+EXdZADbzqAJQrpDRq71CIMnOP5VMut3SL0z5Q==",
"license": "MIT",
"engines": {
"node": "^18 || ^20 || >= 21"
}
},
"node_modules/tree-sitter-rust": {
"version": "0.21.0",
"resolved": "https://registry.npmjs.org/tree-sitter-rust/-/tree-sitter-rust-0.21.0.tgz",
@@ -5677,12 +5410,6 @@
"integrity": "sha512-jk1+QP6ZJqyOiuEI9AEWQfju/nB2Pw466kbA0LEZljHwKeMgd9WrAEgEGxjPDD2+TNbbb37rTyhEfrCXfuKXnA==",
"license": "MIT"
},
"node_modules/util-deprecate": {
"version": "1.0.2",
"resolved": "https://registry.npmjs.org/util-deprecate/-/util-deprecate-1.0.2.tgz",
"integrity": "sha512-EPD5q1uXyFxJpCrLnCc1nHnq3gOa6DZBocAIiI2TaSCA7VCJ1UJDMagCzIkXNsUYfD1daK//LTEQ8xiIbrHtcw==",
"license": "MIT"
},
"node_modules/utils-merge": {
"version": "1.0.1",
"resolved": "https://registry.npmjs.org/utils-merge/-/utils-merge-1.0.1.tgz",
@@ -5899,56 +5626,6 @@
"node": ">=8"
}
},
"node_modules/wide-align": {
"version": "1.1.5",
"resolved": "https://registry.npmjs.org/wide-align/-/wide-align-1.1.5.tgz",
"integrity": "sha512-eDMORYaPNZ4sQIuuYPDHdQvf4gyCF9rEEV/yPxGfwPkRodwEgiMUUXTx/dex+Me0wxx53S+NgUHaP7y3MGlDmg==",
"license": "ISC",
"dependencies": {
"string-width": "^1.0.2 || 2 || 3 || 4"
}
},
"node_modules/wide-align/node_modules/ansi-regex": {
"version": "5.0.1",
"resolved": "https://registry.npmjs.org/ansi-regex/-/ansi-regex-5.0.1.tgz",
"integrity": "sha512-quJQXlTSUGL2LH9SUXo8VwsY4soanhgo6LNSm84E1LBcE8s3O0wpdiRzyR9z/ZZJMlMWv37qOOb9pdJlMUEKFQ==",
"license": "MIT",
"engines": {
"node": ">=8"
}
},
"node_modules/wide-align/node_modules/emoji-regex": {
"version": "8.0.0",
"resolved": "https://registry.npmjs.org/emoji-regex/-/emoji-regex-8.0.0.tgz",
"integrity": "sha512-MSjYzcWNOA0ewAHpz0MxpYFvwg6yjy1NG3xteoqz644VCo/RPgnr1/GGt+ic3iJTzQ8Eu3TdM14SawnVUmGE6A==",
"license": "MIT"
},
"node_modules/wide-align/node_modules/string-width": {
"version": "4.2.3",
"resolved": "https://registry.npmjs.org/string-width/-/string-width-4.2.3.tgz",
"integrity": "sha512-wKyQRQpjJ0sIp62ErSZdGsjMJWsap5oRNihHhu6G7JVO/9jIB6UyevL+tXuOqrng8j/cxKTWyWUwvSTriiZz/g==",
"license": "MIT",
"dependencies": {
"emoji-regex": "^8.0.0",
"is-fullwidth-code-point": "^3.0.0",
"strip-ansi": "^6.0.1"
},
"engines": {
"node": ">=8"
}
},
"node_modules/wide-align/node_modules/strip-ansi": {
"version": "6.0.1",
"resolved": "https://registry.npmjs.org/strip-ansi/-/strip-ansi-6.0.1.tgz",
"integrity": "sha512-Y38VPSHcqkFrCpFnQ9vuSXmquuv5oXOKpGeT6aGrr3o3Gc9AlVa6JBfUSOCnbxGGZF+/0ooI7KrPuUSztUdU5A==",
"license": "MIT",
"dependencies": {
"ansi-regex": "^5.0.1"
},
"engines": {
"node": ">=8"
}
},
"node_modules/wrap-ansi": {
"version": "8.1.0",
"resolved": "https://registry.npmjs.org/wrap-ansi/-/wrap-ansi-8.1.0.tgz",
@@ -6055,12 +5732,6 @@
"node": ">=10"
}
},
"node_modules/yallist": {
"version": "4.0.0",
"resolved": "https://registry.npmjs.org/yallist/-/yallist-4.0.0.tgz",
"integrity": "sha512-3wdGidZyq5PB084XLES5TpOSRA3wjXAlIWMhum2kRcv/41Sn2emQ0dycQW4uZXLejwKvg6EsvbdlVL+FYEct7A==",
"license": "ISC"
},
"node_modules/yargs": {
"version": "17.7.2",
"resolved": "https://registry.npmjs.org/yargs/-/yargs-17.7.2.tgz",
+5 -3
View File
@@ -1,6 +1,6 @@
{
"name": "gitnexus",
"version": "1.3.8",
"version": "1.4.5",
"description": "Graph-powered code intelligence for AI agents. Index any codebase, query via MCP or CLI.",
"author": "Abhigyan Patwari",
"license": "PolyForm-Noncommercial-1.0.0",
@@ -58,7 +58,8 @@
"graphology": "^0.25.4",
"graphology-indices": "^0.17.0",
"graphology-utils": "^2.3.0",
"kuzu": "^0.11.3",
"@ladybugdb/core": "^0.15.1",
"ignore": "^7.0.5",
"lru-cache": "^11.0.0",
"mnemonist": "^0.39.0",
"pandemonium": "^2.4.0",
@@ -69,14 +70,15 @@
"tree-sitter-go": "^0.21.0",
"tree-sitter-java": "^0.21.0",
"tree-sitter-javascript": "^0.21.0",
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-php": "^0.23.12",
"tree-sitter-python": "^0.21.0",
"tree-sitter-ruby": "^0.23.1",
"tree-sitter-rust": "^0.21.0",
"tree-sitter-typescript": "^0.21.0",
"uuid": "^13.0.0"
},
"optionalDependencies": {
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-swift": "^0.6.0"
},
"devDependencies": {
+1 -1
View File
@@ -22,7 +22,7 @@ Run from the project root. This parses all source files, builds the knowledge gr
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale.
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook runs `analyze` automatically after `git commit` and `git merge`, preserving embeddings if previously generated.
### status — Check index freshness
+39 -6
View File
@@ -9,6 +9,7 @@
import fs from 'fs/promises';
import path from 'path';
import { fileURLToPath } from 'url';
import { type GeneratedSkillInfo } from './skill-gen.js';
// ESM equivalent of __dirname
const __filename = fileURLToPath(import.meta.url);
@@ -37,7 +38,22 @@ const GITNEXUS_END_MARKER = '<!-- gitnexus:end -->';
* - Exact tool commands with parameters — vague directives get ignored
* - Self-review checklist — forces model to verify its own work
*/
function generateGitNexusContent(projectName: string, stats: RepoStats): string {
function generateGitNexusContent(projectName: string, stats: RepoStats, generatedSkills?: GeneratedSkillInfo[]): string {
const generatedRows = (generatedSkills && generatedSkills.length > 0)
? generatedSkills.map(s =>
`| Work in the ${s.label} area (${s.symbolCount} symbols) | \`.claude/skills/generated/${s.name}/SKILL.md\` |`
).join('\n')
: '';
const skillsTable = `| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | \`.claude/skills/gitnexus/gitnexus-exploring/SKILL.md\` |
| Blast radius / "What breaks if I change X?" | \`.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md\` |
| Trace bugs / "Why is X failing?" | \`.claude/skills/gitnexus/gitnexus-debugging/SKILL.md\` |
| Rename / extract / split / refactor | \`.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md\` |
| Tools, resources, schema reference | \`.claude/skills/gitnexus/gitnexus-guide/SKILL.md\` |
| Index, status, clean, wiki CLI commands | \`.claude/skills/gitnexus/gitnexus-cli/SKILL.md\` |${generatedRows ? '\n' + generatedRows : ''}`;
return `${GITNEXUS_START_MARKER}
# GitNexus — Code Intelligence
@@ -109,11 +125,27 @@ Before completing any code modification task, verify:
3. \`gitnexus_detect_changes()\` confirms changes match expected scope
4. All d=1 (WILL BREAK) dependents were updated
## Keeping the Index Fresh
After committing code changes, the GitNexus index becomes stale. Re-run analyze to update it:
\`\`\`bash
npx gitnexus analyze
\`\`\`
If the index previously included embeddings, preserve them by adding \`--embeddings\`:
\`\`\`bash
npx gitnexus analyze --embeddings
\`\`\`
To check whether embeddings exist, inspect \`.gitnexus/meta.json\` — the \`stats.embeddings\` field shows the count (0 means no embeddings). **Running analyze without \`--embeddings\` will delete any previously generated embeddings.**
> Claude Code users: A PostToolUse hook handles this automatically after \`git commit\` and \`git merge\`.
## CLI
- Re-index: \`npx gitnexus analyze\`
- Check freshness: \`npx gitnexus status\`
- Generate docs: \`npx gitnexus wiki\`
${skillsTable}
${GITNEXUS_END_MARKER}`;
}
@@ -252,9 +284,10 @@ export async function generateAIContextFiles(
repoPath: string,
_storagePath: string,
projectName: string,
stats: RepoStats
stats: RepoStats,
generatedSkills?: GeneratedSkillInfo[]
): Promise<{ files: string[] }> {
const content = generateGitNexusContent(projectName, stats);
const content = generateGitNexusContent(projectName, stats, generatedSkills);
const createdFiles: string[] = [];
// Create AGENTS.md (standard for Cursor, Windsurf, OpenCode, Cline, etc.)
+64 -30
View File
@@ -9,14 +9,15 @@ import { execFileSync } from 'child_process';
import v8 from 'v8';
import cliProgress from 'cli-progress';
import { runPipelineFromRepo } from '../core/ingestion/pipeline.js';
import { initKuzu, loadGraphToKuzu, getKuzuStats, executeQuery, executeWithReusedStatement, closeKuzu, createFTSIndex, loadCachedEmbeddings } from '../core/kuzu/kuzu-adapter.js';
import { initLbug, loadGraphToLbug, getLbugStats, executeQuery, executeWithReusedStatement, closeLbug, createFTSIndex, loadCachedEmbeddings } from '../core/lbug/lbug-adapter.js';
// Embedding imports are lazy (dynamic import) so onnxruntime-node is never
// loaded when embeddings are not requested. This avoids crashes on Node
// versions whose ABI is not yet supported by the native binary (#89).
// disposeEmbedder intentionally not called — ONNX Runtime segfaults on cleanup (see #38)
import { getStoragePaths, saveMeta, loadMeta, addToGitignore, registerRepo, getGlobalRegistryPath } from '../storage/repo-manager.js';
import { getStoragePaths, saveMeta, loadMeta, addToGitignore, registerRepo, getGlobalRegistryPath, cleanupOldKuzuFiles } from '../storage/repo-manager.js';
import { getCurrentCommit, isGitRepo, getGitRoot } from '../storage/git.js';
import { generateAIContextFiles } from './ai-context.js';
import { generateSkillFiles, type GeneratedSkillInfo } from './skill-gen.js';
import fs from 'fs/promises';
@@ -45,6 +46,8 @@ function ensureHeap(): boolean {
export interface AnalyzeOptions {
force?: boolean;
embeddings?: boolean;
skills?: boolean;
verbose?: boolean;
}
/** Threshold: auto-skip embeddings for repos with more nodes than this */
@@ -60,7 +63,7 @@ const PHASE_LABELS: Record<string, string> = {
communities: 'Detecting communities',
processes: 'Detecting processes',
complete: 'Pipeline complete',
kuzu: 'Loading into KuzuDB',
lbug: 'Loading into LadybugDB',
fts: 'Creating search indexes',
embeddings: 'Generating embeddings',
done: 'Done',
@@ -72,6 +75,10 @@ export const analyzeCommand = async (
) => {
if (ensureHeap()) return;
if (options?.verbose) {
process.env.GITNEXUS_VERBOSE = '1';
}
console.log('\n GitNexus Analyzer\n');
let repoPath: string;
@@ -93,15 +100,27 @@ export const analyzeCommand = async (
return;
}
const { storagePath, kuzuPath } = getStoragePaths(repoPath);
const { storagePath, lbugPath } = getStoragePaths(repoPath);
// Clean up stale KuzuDB files from before the LadybugDB migration.
// If kuzu existed but lbug doesn't, we're doing a migration re-index — say so.
const kuzuResult = await cleanupOldKuzuFiles(storagePath);
if (kuzuResult.found && kuzuResult.needsReindex) {
console.log(' Migrating from KuzuDB to LadybugDB — rebuilding index...\n');
}
const currentCommit = getCurrentCommit(repoPath);
const existingMeta = await loadMeta(storagePath);
if (existingMeta && !options?.force && existingMeta.lastCommit === currentCommit) {
if (existingMeta && !options?.force && !options?.skills && existingMeta.lastCommit === currentCommit) {
console.log(' Already up to date\n');
return;
}
if (process.env.GITNEXUS_NO_GITIGNORE) {
console.log(' GITNEXUS_NO_GITIGNORE is set — skipping .gitignore (still reading .gitnexusignore)\n');
}
// Single progress bar for entire pipeline
const bar = new cliProgress.SingleBar({
format: ' {bar} {percentage}% | {phase}',
@@ -123,7 +142,7 @@ export const analyzeCommand = async (
aborted = true;
bar.stop();
console.log('\n Interrupted — cleaning up...');
closeKuzu().catch(() => {}).finally(() => process.exit(130));
closeLbug().catch(() => {}).finally(() => process.exit(130));
};
process.on('SIGINT', sigintHandler);
@@ -173,13 +192,13 @@ export const analyzeCommand = async (
if (options?.embeddings && existingMeta && !options?.force) {
try {
updateBar(0, 'Caching embeddings...');
await initKuzu(kuzuPath);
await initLbug(lbugPath);
const cached = await loadCachedEmbeddings();
cachedEmbeddingNodeIds = cached.embeddingNodeIds;
cachedEmbeddings = cached.embeddings;
await closeKuzu();
await closeLbug();
} catch {
try { await closeKuzu(); } catch {}
try { await closeLbug(); } catch {}
}
}
@@ -190,25 +209,25 @@ export const analyzeCommand = async (
updateBar(scaled, phaseLabel);
});
// ── Phase 2: KuzuDB (60–85%) ──────────────────────────────────────
updateBar(60, 'Loading into KuzuDB...');
// ── Phase 2: LadybugDB (60–85%) ──────────────────────────────────────
updateBar(60, 'Loading into LadybugDB...');
await closeKuzu();
const kuzuFiles = [kuzuPath, `${kuzuPath}.wal`, `${kuzuPath}.lock`];
for (const f of kuzuFiles) {
await closeLbug();
const lbugFiles = [lbugPath, `${lbugPath}.wal`, `${lbugPath}.lock`];
for (const f of lbugFiles) {
try { await fs.rm(f, { recursive: true, force: true }); } catch {}
}
const t0Kuzu = Date.now();
await initKuzu(kuzuPath);
let kuzuMsgCount = 0;
const kuzuResult = await loadGraphToKuzu(pipelineResult.graph, pipelineResult.repoPath, storagePath, (msg) => {
kuzuMsgCount++;
const progress = Math.min(84, 60 + Math.round((kuzuMsgCount / (kuzuMsgCount + 10)) * 24));
const t0Lbug = Date.now();
await initLbug(lbugPath);
let lbugMsgCount = 0;
const lbugResult = await loadGraphToLbug(pipelineResult.graph, pipelineResult.repoPath, storagePath, (msg) => {
lbugMsgCount++;
const progress = Math.min(84, 60 + Math.round((lbugMsgCount / (lbugMsgCount + 10)) * 24));
updateBar(progress, msg);
});
const kuzuTime = ((Date.now() - t0Kuzu) / 1000).toFixed(1);
const kuzuWarnings = kuzuResult.warnings;
const lbugTime = ((Date.now() - t0Lbug) / 1000).toFixed(1);
const lbugWarnings = lbugResult.warnings;
// ── Phase 3: FTS (85–90%) ─────────────────────────────────────────
updateBar(85, 'Creating search indexes...');
@@ -242,7 +261,7 @@ export const analyzeCommand = async (
}
// ── Phase 4: Embeddings (90–98%) ──────────────────────────────────
const stats = await getKuzuStats();
const stats = await getLbugStats();
let embeddingTime = '0.0';
let embeddingSkipped = true;
let embeddingSkipReason = 'off (use --embeddings to enable)';
@@ -276,6 +295,13 @@ export const analyzeCommand = async (
// ── Phase 5: Finalize (98–100%) ───────────────────────────────────
updateBar(98, 'Saving metadata...');
// Count embeddings in the index (cached + newly generated)
let embeddingCount = 0;
try {
const embResult = await executeQuery(`MATCH (e:CodeEmbedding) RETURN count(e) AS cnt`);
embeddingCount = embResult?.[0]?.cnt ?? 0;
} catch { /* table may not exist if embeddings never ran */ }
const meta = {
repoPath,
lastCommit: currentCommit,
@@ -286,6 +312,7 @@ export const analyzeCommand = async (
edges: stats.edges,
communities: pipelineResult.communityResult?.stats.totalCommunities,
processes: pipelineResult.processResult?.stats.totalProcesses,
embeddings: embeddingCount,
},
};
await saveMeta(storagePath, meta);
@@ -303,6 +330,13 @@ export const analyzeCommand = async (
aggregatedClusterCount = Array.from(groups.values()).filter(count => count >= 5).length;
}
let generatedSkills: GeneratedSkillInfo[] = [];
if (options?.skills && pipelineResult.communityResult) {
updateBar(99, 'Generating skill files...');
const skillResult = await generateSkillFiles(repoPath, projectName, pipelineResult);
generatedSkills = skillResult.skills;
}
const aiContext = await generateAIContextFiles(repoPath, storagePath, projectName, {
files: pipelineResult.totalFileCount,
nodes: stats.nodes,
@@ -310,9 +344,9 @@ export const analyzeCommand = async (
communities: pipelineResult.communityResult?.stats.totalCommunities,
clusters: aggregatedClusterCount,
processes: pipelineResult.processResult?.stats.totalProcesses,
});
}, generatedSkills);
await closeKuzu();
await closeLbug();
// Note: we intentionally do NOT call disposeEmbedder() here.
// ONNX Runtime's native cleanup segfaults on macOS and some Linux configs.
// Since the process exits immediately after, Node.js reclaims everything.
@@ -333,7 +367,7 @@ export const analyzeCommand = async (
const embeddingsCached = cachedEmbeddings.length > 0;
console.log(`\n Repository indexed successfully (${totalTime}s)${embeddingsCached ? ` [${cachedEmbeddings.length} embeddings cached]` : ''}\n`);
console.log(` ${stats.nodes.toLocaleString()} nodes | ${stats.edges.toLocaleString()} edges | ${pipelineResult.communityResult?.stats.totalCommunities || 0} clusters | ${pipelineResult.processResult?.stats.totalProcesses || 0} flows`);
console.log(` KuzuDB ${kuzuTime}s | FTS ${ftsTime}s | Embeddings ${embeddingSkipped ? embeddingSkipReason : embeddingTime + 's'}`);
console.log(` LadybugDB ${lbugTime}s | FTS ${ftsTime}s | Embeddings ${embeddingSkipped ? embeddingSkipReason : embeddingTime + 's'}`);
console.log(` ${repoPath}`);
if (aiContext.files.length > 0) {
@@ -341,12 +375,12 @@ export const analyzeCommand = async (
}
// Show a quiet summary if some edge types needed fallback insertion
if (kuzuWarnings.length > 0) {
const totalFallback = kuzuWarnings.reduce((sum, w) => {
if (lbugWarnings.length > 0) {
const totalFallback = lbugWarnings.reduce((sum, w) => {
const m = w.match(/\((\d+) edges\)/);
return sum + (m ? parseInt(m[1]) : 0);
}, 0);
console.log(` Note: ${totalFallback} edges across ${kuzuWarnings.length} types inserted via fallback (schema will be updated in next release)`);
console.log(` Note: ${totalFallback} edges across ${lbugWarnings.length} types inserted via fallback (schema will be updated in next release)`);
}
try {
@@ -357,7 +391,7 @@ export const analyzeCommand = async (
console.log('');
// KuzuDB's native module holds open handles that prevent Node from exiting.
// LadybugDB's native module holds open handles that prevent Node from exiting.
// ONNX Runtime also registers native atexit hooks that segfault on some
// platforms (#38, #40). Force-exit to ensure clean termination.
process.exit(0);
+1 -1
View File
@@ -23,7 +23,7 @@ export async function augmentCommand(pattern: string): Promise<void> {
if (result) {
// IMPORTANT: Write to stderr, NOT stdout.
// KuzuDB's native module captures stdout fd at OS level during init,
// LadybugDB's native module captures stdout fd at OS level during init,
// which makes stdout permanently broken in subprocess contexts.
// stderr is never captured, so it works reliably everywhere.
// The hook reads from the subprocess's stderr.
+1 -1
View File
@@ -1,7 +1,7 @@
/**
* Eval Server — Lightweight HTTP server for SWE-bench evaluation
*
* Keeps KuzuDB warm in memory so tool calls from the agent are near-instant.
* Keeps LadybugDB warm in memory so tool calls from the agent are near-instant.
* Designed to run inside Docker containers during SWE-bench evaluation.
*
* KEY DESIGN: Returns LLM-friendly text, not raw JSON.
+19 -25
View File
@@ -4,18 +4,9 @@
// Removing it from here improves MCP server startup time significantly.
import { Command } from 'commander';
import { analyzeCommand } from './analyze.js';
import { serveCommand } from './serve.js';
import { listCommand } from './list.js';
import { statusCommand } from './status.js';
import { mcpCommand } from './mcp.js';
import { cleanCommand } from './clean.js';
import { setupCommand } from './setup.js';
import { augmentCommand } from './augment.js';
import { wikiCommand } from './wiki.js';
import { queryCommand, contextCommand, impactCommand, cypherCommand } from './tool.js';
import { evalServerCommand } from './eval-server.js';
import { createRequire } from 'node:module';
import { createLazyAction } from './lazy-action.js';
const _require = createRequire(import.meta.url);
const pkg = _require('../../package.json');
const program = new Command();
@@ -28,43 +19,46 @@ program
program
.command('setup')
.description('One-time setup: configure MCP for Cursor, Claude Code, OpenCode')
.action(setupCommand);
.action(createLazyAction(() => import('./setup.js'), 'setupCommand'));
program
.command('analyze [path]')
.description('Index a repository (full analysis)')
.option('-f, --force', 'Force full re-index even if up to date')
.option('--embeddings', 'Enable embedding generation for semantic search (off by default)')
.action(analyzeCommand);
.option('--skills', 'Generate repo-specific skill files from detected communities')
.option('-v, --verbose', 'Enable verbose ingestion warnings (default: false)')
.addHelpText('after', '\nEnvironment variables:\n GITNEXUS_NO_GITIGNORE=1 Skip .gitignore parsing (still reads .gitnexusignore)')
.action(createLazyAction(() => import('./analyze.js'), 'analyzeCommand'));
program
.command('serve')
.description('Start local HTTP server for web UI connection')
.option('-p, --port <port>', 'Port number', '4747')
.option('--host <host>', 'Bind address (default: 127.0.0.1, use 0.0.0.0 for remote access)')
.action(serveCommand);
.action(createLazyAction(() => import('./serve.js'), 'serveCommand'));
program
.command('mcp')
.description('Start MCP server (stdio) — serves all indexed repos')
.action(mcpCommand);
.action(createLazyAction(() => import('./mcp.js'), 'mcpCommand'));
program
.command('list')
.description('List all indexed repositories')
.action(listCommand);
.action(createLazyAction(() => import('./list.js'), 'listCommand'));
program
.command('status')
.description('Show index status for current repo')
.action(statusCommand);
.action(createLazyAction(() => import('./status.js'), 'statusCommand'));
program
.command('clean')
.description('Delete GitNexus index for current repo')
.option('-f, --force', 'Skip confirmation prompt')
.option('--all', 'Clean all indexed repos')
.action(cleanCommand);
.action(createLazyAction(() => import('./clean.js'), 'cleanCommand'));
program
.command('wiki [path]')
@@ -75,12 +69,12 @@ program
.option('--api-key <key>', 'LLM API key (saved to ~/.gitnexus/config.json)')
.option('--concurrency <n>', 'Parallel LLM calls (default: 3)', '3')
.option('--gist', 'Publish wiki as a public GitHub Gist after generation')
.action(wikiCommand);
.action(createLazyAction(() => import('./wiki.js'), 'wikiCommand'));
program
.command('augment <pattern>')
.description('Augment a search pattern with knowledge graph context (used by hooks)')
.action(augmentCommand);
.action(createLazyAction(() => import('./augment.js'), 'augmentCommand'));
// ─── Direct Tool Commands (no MCP overhead) ────────────────────────
// These invoke LocalBackend directly for use in eval, scripts, and CI.
@@ -93,7 +87,7 @@ program
.option('-g, --goal <text>', 'What you want to find')
.option('-l, --limit <n>', 'Max processes to return (default: 5)')
.option('--content', 'Include full symbol source code')
.action(queryCommand);
.action(createLazyAction(() => import('./tool.js'), 'queryCommand'));
program
.command('context [name]')
@@ -102,7 +96,7 @@ program
.option('-u, --uid <uid>', 'Direct symbol UID (zero-ambiguity lookup)')
.option('-f, --file <path>', 'File path to disambiguate common names')
.option('--content', 'Include full symbol source code')
.action(contextCommand);
.action(createLazyAction(() => import('./tool.js'), 'contextCommand'));
program
.command('impact <target>')
@@ -111,13 +105,13 @@ program
.option('-r, --repo <name>', 'Target repository')
.option('--depth <n>', 'Max relationship depth (default: 3)')
.option('--include-tests', 'Include test files in results')
.action(impactCommand);
.action(createLazyAction(() => import('./tool.js'), 'impactCommand'));
program
.command('cypher <query>')
.description('Execute raw Cypher query against the knowledge graph')
.option('-r, --repo <name>', 'Target repository')
.action(cypherCommand);
.action(createLazyAction(() => import('./tool.js'), 'cypherCommand'));
// ─── Eval Server (persistent daemon for SWE-bench) ─────────────────
@@ -126,6 +120,6 @@ program
.description('Start lightweight HTTP server for fast tool calls during evaluation')
.option('-p, --port <port>', 'Port number', '4848')
.option('--idle-timeout <seconds>', 'Auto-shutdown after N seconds idle (0 = disabled)', '0')
.action(evalServerCommand);
.action(createLazyAction(() => import('./eval-server.js'), 'evalServerCommand'));
program.parse(process.argv);
+26
View File
@@ -0,0 +1,26 @@
/**
* Creates a lazy-loaded CLI action that defers module import until invocation.
* The generic constraints ensure the export name is a valid key of the module
* at compile time — catching typos when used with concrete module imports.
*/
function isCallable(value: unknown): value is (...args: unknown[]) => unknown {
return typeof value === 'function';
}
export function createLazyAction<
TModule extends Record<string, unknown>,
TKey extends string & keyof TModule,
>(
loader: () => Promise<TModule>,
exportName: TKey,
): (...args: unknown[]) => Promise<void> {
return async (...args: unknown[]): Promise<void> => {
const module = await loader();
const action = module[exportName];
if (!isCallable(action)) {
throw new Error(`Lazy action export not found: ${exportName}`);
}
await action(...args);
};
}
+1 -1
View File
@@ -11,7 +11,7 @@ import { LocalBackend } from '../mcp/local/local-backend.js';
export const mcpCommand = async () => {
// Prevent unhandled errors from crashing the MCP server process.
// KuzuDB lock conflicts and transient errors should degrade gracefully.
// LadybugDB lock conflicts and transient errors should degrade gracefully.
process.on('uncaughtException', (err) => {
console.error(`GitNexus MCP: uncaught exception — ${err.message}`);
// Process is in an undefined state after uncaughtException — exit after flushing
+53 -33
View File
@@ -10,6 +10,7 @@ import fs from 'fs/promises';
import path from 'path';
import os from 'os';
import { fileURLToPath } from 'url';
import { glob } from 'glob';
import { getGlobalDir } from '../storage/repo-manager.js';
const __filename = fileURLToPath(import.meta.url);
@@ -168,16 +169,18 @@ async function installClaudeCodeHooks(result: SetupResult): Promise<void> {
// even when it's no longer inside the npm package tree
const resolvedCli = path.join(__dirname, '..', 'cli', 'index.js');
const normalizedCli = path.resolve(resolvedCli).replace(/\\/g, '/');
const jsonCli = JSON.stringify(normalizedCli);
content = content.replace(
"let cliPath = path.resolve(__dirname, '..', '..', 'dist', 'cli', 'index.js');",
`let cliPath = '${normalizedCli}';`
`let cliPath = ${jsonCli};`
);
await fs.writeFile(dest, content, 'utf-8');
} catch {
// Script not found in source — skip
}
const hookCmd = `node "${path.join(destHooksDir, 'gitnexus-hook.cjs').replace(/\\/g, '/')}"`;
const hookPath = path.join(destHooksDir, 'gitnexus-hook.cjs').replace(/\\/g, '/');
const hookCmd = `node "${hookPath.replace(/"/g, '\\"')}"`;
// Merge hook config into ~/.claude/settings.json
const existing = await readJsonFile(settingsPath) || {};
@@ -186,25 +189,31 @@ async function installClaudeCodeHooks(result: SetupResult): Promise<void> {
// NOTE: SessionStart hooks are broken on Windows (Claude Code bug #23576).
// Session context is delivered via CLAUDE.md / skills instead.
// Add PreToolUse hook if not already present
if (!existing.hooks.PreToolUse) existing.hooks.PreToolUse = [];
const hasPreToolHook = existing.hooks.PreToolUse.some(
(h: any) => h.hooks?.some((hh: any) => hh.command?.includes('gitnexus'))
);
if (!hasPreToolHook) {
existing.hooks.PreToolUse.push({
matcher: 'Grep|Glob|Bash',
hooks: [{
type: 'command',
command: hookCmd,
timeout: 8000,
statusMessage: 'Enriching with GitNexus graph context...',
}],
});
// Helper: add a hook entry if one with 'gitnexus-hook' isn't already registered
interface HookEntry { hooks?: Array<{ command?: string }> }
function ensureHookEntry(
eventName: string,
matcher: string,
timeout: number,
statusMessage: string,
) {
if (!existing.hooks[eventName]) existing.hooks[eventName] = [];
const hasHook = existing.hooks[eventName].some(
(h: HookEntry) => h.hooks?.some(hh => hh.command?.includes('gitnexus-hook'))
);
if (!hasHook) {
existing.hooks[eventName].push({
matcher,
hooks: [{ type: 'command', command: hookCmd, timeout, statusMessage }],
});
}
}
ensureHookEntry('PreToolUse', 'Grep|Glob|Bash', 10, 'Enriching with GitNexus graph context...');
ensureHookEntry('PostToolUse', 'Bash', 10, 'Checking GitNexus index freshness...');
await writeJsonFile(settingsPath, existing);
result.configured.push('Claude Code hooks (PreToolUse)');
result.configured.push('Claude Code hooks (PreToolUse, PostToolUse)');
} catch (err: any) {
result.errors.push(`Claude Code hooks: ${err.message}`);
}
@@ -232,8 +241,6 @@ async function setupOpenCode(result: SetupResult): Promise<void> {
// ─── Skill Installation ───────────────────────────────────────────
const SKILL_NAMES = ['gitnexus-exploring', 'gitnexus-debugging', 'gitnexus-impact-analysis', 'gitnexus-refactoring', 'gitnexus-guide', 'gitnexus-cli'];
/**
* Install GitNexus skills to a target directory.
* Each skill is installed as {targetDir}/gitnexus-{skillName}/SKILL.md
@@ -247,25 +254,38 @@ async function installSkillsTo(targetDir: string): Promise<string[]> {
const installed: string[] = [];
const skillsRoot = path.join(__dirname, '..', '..', 'skills');
for (const skillName of SKILL_NAMES) {
let flatFiles: string[] = [];
let dirSkillFiles: string[] = [];
try {
[flatFiles, dirSkillFiles] = await Promise.all([
glob('*.md', { cwd: skillsRoot }),
glob('*/SKILL.md', { cwd: skillsRoot }),
]);
} catch {
return [];
}
const skillSources = new Map<string, { isDirectory: boolean }>();
for (const relPath of dirSkillFiles) {
skillSources.set(path.dirname(relPath), { isDirectory: true });
}
for (const relPath of flatFiles) {
const skillName = path.basename(relPath, '.md');
if (!skillSources.has(skillName)) {
skillSources.set(skillName, { isDirectory: false });
}
}
for (const [skillName, source] of skillSources) {
const skillDir = path.join(targetDir, skillName);
try {
// Try directory-based skill first (skills/{name}/SKILL.md)
const dirSource = path.join(skillsRoot, skillName);
const dirSkillFile = path.join(dirSource, 'SKILL.md');
let isDirectory = false;
try {
const stat = await fs.stat(dirSource);
isDirectory = stat.isDirectory();
} catch { /* not a directory */ }
if (isDirectory) {
if (source.isDirectory) {
const dirSource = path.join(skillsRoot, skillName);
await copyDirRecursive(dirSource, skillDir);
installed.push(skillName);
} else {
// Fall back to flat file (skills/{name}.md)
const flatSource = path.join(skillsRoot, `${skillName}.md`);
const content = await fs.readFile(flatSource, 'utf-8');
await fs.mkdir(skillDir, { recursive: true });
+712
View File
@@ -0,0 +1,712 @@
/**
* Skill File Generator
*
* Generates repo-specific SKILL.md files from detected Leiden communities.
* Each significant community becomes a skill that describes a functional area
* of the codebase, including key files, entry points, execution flows, and
* cross-community connections.
*/
import fs from 'fs/promises';
import path from 'path';
import { PipelineResult } from '../types/pipeline.js';
import { CommunityNode, CommunityMembership } from '../core/ingestion/community-processor.js';
import { ProcessNode } from '../core/ingestion/process-processor.js';
import { GraphNode, KnowledgeGraph } from '../core/graph/types.js';
// ============================================================================
// TYPES
// ============================================================================
export interface GeneratedSkillInfo {
name: string;
label: string;
symbolCount: number;
fileCount: number;
}
interface AggregatedCommunity {
label: string;
rawIds: string[];
symbolCount: number;
cohesion: number;
}
interface MemberSymbol {
id: string;
name: string;
label: string;
filePath: string;
startLine: number;
isExported: boolean;
}
interface FileInfo {
relativePath: string;
symbols: string[];
}
interface CrossConnection {
targetLabel: string;
count: number;
}
// ============================================================================
// MAIN EXPORT
// ============================================================================
/**
* @brief Generate repo-specific skill files from detected communities
* @param {string} repoPath - Absolute path to the repository root
* @param {string} projectName - Human-readable project name
* @param {PipelineResult} pipelineResult - In-memory pipeline data with communities, processes, graph
* @returns {Promise<{ skills: GeneratedSkillInfo[], outputPath: string }>} Generated skill metadata
*/
export const generateSkillFiles = async (
repoPath: string,
projectName: string,
pipelineResult: PipelineResult
): Promise<{ skills: GeneratedSkillInfo[]; outputPath: string }> => {
const { communityResult, processResult, graph } = pipelineResult;
const outputDir = path.join(repoPath, '.claude', 'skills', 'generated');
if (!communityResult || !communityResult.memberships.length) {
console.log('\n Skills: no communities detected, skipping skill generation');
return { skills: [], outputPath: outputDir };
}
console.log('\n Generating repo-specific skills...');
// Step 1: Build communities from memberships (not the filtered communities array).
// The community processor skips singletons from its communities array but memberships
// include ALL assignments. For repos with sparse CALLS edges, the communities array
// can be empty while memberships still has useful groupings.
const communities = communityResult.communities.length > 0
? communityResult.communities
: buildCommunitiesFromMemberships(communityResult.memberships, graph, repoPath);
const aggregated = aggregateCommunities(communities);
// Step 2: Filter to significant communities
// Keep communities with >= 3 symbols after aggregation.
const significant = aggregated
.filter(c => c.symbolCount >= 3)
.sort((a, b) => b.symbolCount - a.symbolCount)
.slice(0, 20);
if (significant.length === 0) {
console.log('\n Skills: no significant communities found (all below 3-symbol threshold)');
return { skills: [], outputPath: outputDir };
}
// Step 3: Build lookup maps
const membershipsByComm = buildMembershipMap(communityResult.memberships);
const nodeIdToCommunityLabel = buildNodeCommunityLabelMap(
communityResult.memberships,
communities
);
// Step 4: Clear and recreate output directory
try {
await fs.rm(outputDir, { recursive: true, force: true });
} catch { /* may not exist */ }
await fs.mkdir(outputDir, { recursive: true });
// Step 5: Generate skill files
const skills: GeneratedSkillInfo[] = [];
const usedNames = new Set<string>();
for (const community of significant) {
// Gather member symbols
const members = gatherMembers(community.rawIds, membershipsByComm, graph);
if (members.length === 0) continue;
// Gather file info
const files = gatherFiles(members, repoPath);
// Gather entry points
const entryPoints = gatherEntryPoints(members);
// Gather execution flows
const flows = gatherFlows(community.rawIds, processResult?.processes || []);
// Gather cross-community connections
const connections = gatherCrossConnections(
community.rawIds,
community.label,
membershipsByComm,
nodeIdToCommunityLabel,
graph
);
// Generate kebab name
const kebabName = toKebabName(community.label, usedNames);
usedNames.add(kebabName);
// Generate SKILL.md content
const content = renderSkillMarkdown(
community,
projectName,
members,
files,
entryPoints,
flows,
connections,
kebabName
);
// Write file
const skillDir = path.join(outputDir, kebabName);
await fs.mkdir(skillDir, { recursive: true });
await fs.writeFile(path.join(skillDir, 'SKILL.md'), content, 'utf-8');
const info: GeneratedSkillInfo = {
name: kebabName,
label: community.label,
symbolCount: community.symbolCount,
fileCount: files.length,
};
skills.push(info);
console.log(` \u2713 ${community.label} (${community.symbolCount} symbols, ${files.length} files)`);
}
console.log(`\n ${skills.length} skills generated \u2192 .claude/skills/generated/`);
return { skills, outputPath: outputDir };
};
// ============================================================================
// FALLBACK COMMUNITY BUILDER
// ============================================================================
/**
* @brief Build CommunityNode-like objects from raw memberships when the community
* processor's communities array is empty (all singletons were filtered out)
* @param {CommunityMembership[]} memberships - All node-to-community assignments
* @param {KnowledgeGraph} graph - The knowledge graph for resolving node metadata
* @param {string} repoPath - Repository root for path normalization
* @returns {CommunityNode[]} Synthetic community nodes built from membership data
*/
const buildCommunitiesFromMemberships = (
memberships: CommunityMembership[],
graph: KnowledgeGraph,
repoPath: string
): CommunityNode[] => {
// Group memberships by communityId
const groups = new Map<string, string[]>();
for (const m of memberships) {
const arr = groups.get(m.communityId);
if (arr) {
arr.push(m.nodeId);
} else {
groups.set(m.communityId, [m.nodeId]);
}
}
const communities: CommunityNode[] = [];
for (const [commId, nodeIds] of groups) {
// Derive a heuristic label from the most common parent directory
const folderCounts = new Map<string, number>();
for (const nodeId of nodeIds) {
const node = graph.getNode(nodeId);
if (!node?.properties.filePath) continue;
const normalized = node.properties.filePath.replace(/\\/g, '/');
const parts = normalized.split('/').filter(Boolean);
if (parts.length >= 2) {
const folder = parts[parts.length - 2];
if (!['src', 'lib', 'core', 'utils', 'common', 'shared', 'helpers'].includes(folder.toLowerCase())) {
folderCounts.set(folder, (folderCounts.get(folder) || 0) + 1);
}
}
}
let bestFolder = '';
let bestCount = 0;
for (const [folder, count] of folderCounts) {
if (count > bestCount) {
bestCount = count;
bestFolder = folder;
}
}
const label = bestFolder
? bestFolder.charAt(0).toUpperCase() + bestFolder.slice(1)
: `Cluster_${commId.replace('comm_', '')}`;
// Compute cohesion as internal-edge ratio (matches backend calculateCohesion).
// For each member node, count edges that stay inside the community vs total.
const nodeSet = new Set(nodeIds);
let internalEdges = 0;
let totalEdges = 0;
graph.forEachRelationship(rel => {
if (nodeSet.has(rel.sourceId)) {
totalEdges++;
if (nodeSet.has(rel.targetId)) internalEdges++;
}
});
const cohesion = totalEdges > 0 ? Math.min(1.0, internalEdges / totalEdges) : 1.0;
communities.push({
id: commId,
label,
heuristicLabel: label,
cohesion,
symbolCount: nodeIds.length,
});
}
return communities.sort((a, b) => b.symbolCount - a.symbolCount);
};
// ============================================================================
// AGGREGATION
// ============================================================================
/**
* @brief Aggregate raw Leiden communities by heuristicLabel
* @param {CommunityNode[]} communities - Raw community nodes from Leiden detection
* @returns {AggregatedCommunity[]} Aggregated communities grouped by label
*/
const aggregateCommunities = (communities: CommunityNode[]): AggregatedCommunity[] => {
const groups = new Map<string, {
rawIds: string[];
totalSymbols: number;
weightedCohesion: number;
}>();
for (const c of communities) {
const label = c.heuristicLabel || c.label || 'Unknown';
const symbols = c.symbolCount || 0;
const cohesion = c.cohesion || 0;
const existing = groups.get(label);
if (!existing) {
groups.set(label, {
rawIds: [c.id],
totalSymbols: symbols,
weightedCohesion: cohesion * symbols,
});
} else {
existing.rawIds.push(c.id);
existing.totalSymbols += symbols;
existing.weightedCohesion += cohesion * symbols;
}
}
return Array.from(groups.entries()).map(([label, g]) => ({
label,
rawIds: g.rawIds,
symbolCount: g.totalSymbols,
cohesion: g.totalSymbols > 0 ? g.weightedCohesion / g.totalSymbols : 0,
}));
};
// ============================================================================
// LOOKUP MAP BUILDERS
// ============================================================================
/**
* @brief Build a map from communityId to member nodeIds
* @param {CommunityMembership[]} memberships - All membership records
* @returns {Map<string, string[]>} Map of communityId -> nodeId[]
*/
const buildMembershipMap = (memberships: CommunityMembership[]): Map<string, string[]> => {
const map = new Map<string, string[]>();
for (const m of memberships) {
const arr = map.get(m.communityId);
if (arr) {
arr.push(m.nodeId);
} else {
map.set(m.communityId, [m.nodeId]);
}
}
return map;
};
/**
* @brief Build a map from nodeId to aggregated community label
* @param {CommunityMembership[]} memberships - All membership records
* @param {CommunityNode[]} communities - Community nodes with labels
* @returns {Map<string, string>} Map of nodeId -> community label
*/
const buildNodeCommunityLabelMap = (
memberships: CommunityMembership[],
communities: CommunityNode[]
): Map<string, string> => {
const commIdToLabel = new Map<string, string>();
for (const c of communities) {
commIdToLabel.set(c.id, c.heuristicLabel || c.label || 'Unknown');
}
const map = new Map<string, string>();
for (const m of memberships) {
const label = commIdToLabel.get(m.communityId);
if (label) {
map.set(m.nodeId, label);
}
}
return map;
};
// ============================================================================
// DATA GATHERING
// ============================================================================
/**
* @brief Gather member symbols for an aggregated community
* @param {string[]} rawIds - Raw community IDs belonging to this aggregated community
* @param {Map<string, string[]>} membershipsByComm - communityId -> nodeIds
* @param {KnowledgeGraph} graph - The knowledge graph
* @returns {MemberSymbol[]} Array of member symbol information
*/
const gatherMembers = (
rawIds: string[],
membershipsByComm: Map<string, string[]>,
graph: KnowledgeGraph
): MemberSymbol[] => {
const seen = new Set<string>();
const members: MemberSymbol[] = [];
for (const commId of rawIds) {
const nodeIds = membershipsByComm.get(commId) || [];
for (const nodeId of nodeIds) {
if (seen.has(nodeId)) continue;
seen.add(nodeId);
const node = graph.getNode(nodeId);
if (!node) continue;
members.push({
id: node.id,
name: node.properties.name,
label: node.label,
filePath: node.properties.filePath || '',
startLine: node.properties.startLine || 0,
isExported: node.properties.isExported === true,
});
}
}
return members;
};
/**
* @brief Gather deduplicated file info with per-file symbol names
* @param {MemberSymbol[]} members - Member symbols
* @param {string} repoPath - Repository root for relative path computation
* @returns {FileInfo[]} Sorted by symbol count descending
*/
const gatherFiles = (members: MemberSymbol[], repoPath: string): FileInfo[] => {
const fileMap = new Map<string, string[]>();
for (const m of members) {
if (!m.filePath) continue;
const rel = toRelativePath(m.filePath, repoPath);
const arr = fileMap.get(rel);
if (arr) {
arr.push(m.name);
} else {
fileMap.set(rel, [m.name]);
}
}
return Array.from(fileMap.entries())
.map(([relativePath, symbols]) => ({ relativePath, symbols }))
.sort((a, b) => b.symbols.length - a.symbols.length);
};
/**
* @brief Gather exported entry points prioritized by type
* @param {MemberSymbol[]} members - Member symbols
* @returns {MemberSymbol[]} Exported symbols sorted by type priority
*/
const gatherEntryPoints = (members: MemberSymbol[]): MemberSymbol[] => {
const typePriority: Record<string, number> = {
Function: 0,
Class: 1,
Method: 2,
Interface: 3,
};
return members
.filter(m => m.isExported)
.sort((a, b) => {
const pa = typePriority[a.label] ?? 99;
const pb = typePriority[b.label] ?? 99;
return pa - pb;
});
};
/**
* @brief Gather execution flows touching this community
* @param {string[]} rawIds - Raw community IDs for this aggregated community
* @param {ProcessNode[]} processes - All detected processes
* @returns {ProcessNode[]} Processes whose communities intersect rawIds, sorted by stepCount
*/
const gatherFlows = (rawIds: string[], processes: ProcessNode[]): ProcessNode[] => {
const rawIdSet = new Set(rawIds);
return processes
.filter(proc => proc.communities.some(cid => rawIdSet.has(cid)))
.sort((a, b) => b.stepCount - a.stepCount);
};
/**
* @brief Gather cross-community call connections
* @param {string[]} rawIds - Raw community IDs for this aggregated community
* @param {string} ownLabel - This community's aggregated label
* @param {Map<string, string[]>} membershipsByComm - communityId -> nodeIds
* @param {Map<string, string>} nodeIdToCommunityLabel - nodeId -> community label
* @param {KnowledgeGraph} graph - The knowledge graph
* @returns {CrossConnection[]} Aggregated cross-community connections sorted by count
*/
const gatherCrossConnections = (
rawIds: string[],
ownLabel: string,
membershipsByComm: Map<string, string[]>,
nodeIdToCommunityLabel: Map<string, string>,
graph: KnowledgeGraph
): CrossConnection[] => {
// Collect all node IDs in this aggregated community
const ownNodeIds = new Set<string>();
for (const commId of rawIds) {
const nodeIds = membershipsByComm.get(commId) || [];
for (const nid of nodeIds) {
ownNodeIds.add(nid);
}
}
// Count outgoing CALLS to nodes in different communities
const targetCounts = new Map<string, number>();
graph.forEachRelationship(rel => {
if (rel.type !== 'CALLS') return;
if (!ownNodeIds.has(rel.sourceId)) return;
if (ownNodeIds.has(rel.targetId)) return; // same community
const targetLabel = nodeIdToCommunityLabel.get(rel.targetId);
if (!targetLabel || targetLabel === ownLabel) return;
targetCounts.set(targetLabel, (targetCounts.get(targetLabel) || 0) + 1);
});
return Array.from(targetCounts.entries())
.map(([targetLabel, count]) => ({ targetLabel, count }))
.sort((a, b) => b.count - a.count);
};
// ============================================================================
// MARKDOWN RENDERING
// ============================================================================
/**
* @brief Render SKILL.md content for a single community
* @param {AggregatedCommunity} community - The aggregated community data
* @param {string} projectName - Project name for the description
* @param {MemberSymbol[]} members - All member symbols
* @param {FileInfo[]} files - File info with symbol names
* @param {MemberSymbol[]} entryPoints - Exported entry point symbols
* @param {ProcessNode[]} flows - Execution flows touching this community
* @param {CrossConnection[]} connections - Cross-community connections
* @param {string} kebabName - Kebab-case name for the skill
* @returns {string} Full SKILL.md content
*/
const renderSkillMarkdown = (
community: AggregatedCommunity,
projectName: string,
members: MemberSymbol[],
files: FileInfo[],
entryPoints: MemberSymbol[],
flows: ProcessNode[],
connections: CrossConnection[],
kebabName: string
): string => {
const cohesionPct = Math.round(community.cohesion * 100);
// Dominant directory: most common top-level directory
const dominantDir = getDominantDirectory(files);
// Top symbol names for "When to Use"
const topNames = entryPoints.slice(0, 3).map(e => e.name);
if (topNames.length === 0) {
// Fallback to any members
topNames.push(...members.slice(0, 3).map(m => m.name));
}
const lines: string[] = [];
// Frontmatter
lines.push('---');
lines.push(`name: ${kebabName}`);
lines.push(`description: "Skill for the ${community.label} area of ${projectName}. ${community.symbolCount} symbols across ${files.length} files."`);
lines.push('---');
lines.push('');
// Title
lines.push(`# ${community.label}`);
lines.push('');
lines.push(`${community.symbolCount} symbols | ${files.length} files | Cohesion: ${cohesionPct}%`);
lines.push('');
// When to Use
lines.push('## When to Use');
lines.push('');
if (dominantDir) {
lines.push(`- Working with code in \`${dominantDir}/\``);
}
if (topNames.length > 0) {
lines.push(`- Understanding how ${topNames.join(', ')} work`);
}
lines.push(`- Modifying ${community.label.toLowerCase()}-related functionality`);
lines.push('');
// Key Files (top 10)
lines.push('## Key Files');
lines.push('');
lines.push('| File | Symbols |');
lines.push('|------|---------|');
for (const f of files.slice(0, 10)) {
const symbolList = f.symbols.slice(0, 5).join(', ');
const suffix = f.symbols.length > 5 ? ` (+${f.symbols.length - 5})` : '';
lines.push(`| \`${f.relativePath}\` | ${symbolList}${suffix} |`);
}
lines.push('');
// Entry Points (top 5)
if (entryPoints.length > 0) {
lines.push('## Entry Points');
lines.push('');
lines.push('Start here when exploring this area:');
lines.push('');
for (const ep of entryPoints.slice(0, 5)) {
lines.push(`- **\`${ep.name}\`** (${ep.label}) \u2014 \`${ep.filePath}:${ep.startLine}\``);
}
lines.push('');
}
// Key Symbols (top 20, exported first, then by type)
lines.push('## Key Symbols');
lines.push('');
lines.push('| Symbol | Type | File | Line |');
lines.push('|--------|------|------|------|');
const sortedMembers = [...members].sort((a, b) => {
if (a.isExported !== b.isExported) return a.isExported ? -1 : 1;
return a.label.localeCompare(b.label);
});
for (const m of sortedMembers.slice(0, 20)) {
lines.push(`| \`${m.name}\` | ${m.label} | \`${m.filePath}\` | ${m.startLine} |`);
}
lines.push('');
// Execution Flows
if (flows.length > 0) {
lines.push('## Execution Flows');
lines.push('');
lines.push('| Flow | Type | Steps |');
lines.push('|------|------|-------|');
for (const f of flows.slice(0, 10)) {
lines.push(`| \`${f.heuristicLabel}\` | ${f.processType} | ${f.stepCount} |`);
}
lines.push('');
}
// Connected Areas
if (connections.length > 0) {
lines.push('## Connected Areas');
lines.push('');
lines.push('| Area | Connections |');
lines.push('|------|-------------|');
for (const c of connections.slice(0, 8)) {
lines.push(`| ${c.targetLabel} | ${c.count} calls |`);
}
lines.push('');
}
// How to Explore
const firstEntry = entryPoints.length > 0 ? entryPoints[0].name : (members.length > 0 ? members[0].name : community.label);
lines.push('## How to Explore');
lines.push('');
lines.push(`1. \`gitnexus_context({name: "${firstEntry}"})\` \u2014 see callers and callees`);
lines.push(`2. \`gitnexus_query({query: "${community.label.toLowerCase()}"})\` \u2014 find related execution flows`);
lines.push('3. Read key files listed above for implementation details');
lines.push('');
return lines.join('\n');
};
// ============================================================================
// UTILITY HELPERS
// ============================================================================
/**
* @brief Convert a community label to a kebab-case directory name
* @param {string} label - The community label
* @param {Set<string>} usedNames - Already-used names for collision detection
* @returns {string} Unique kebab-case name capped at 50 characters
*/
const toKebabName = (label: string, usedNames: Set<string>): string => {
let name = label
.toLowerCase()
.replace(/[^a-z0-9]+/g, '-')
.replace(/^-+|-+$/g, '')
.slice(0, 50);
if (!name) name = 'skill';
let candidate = name;
let counter = 2;
while (usedNames.has(candidate)) {
candidate = `${name}-${counter}`;
counter++;
}
return candidate;
};
/**
* @brief Convert an absolute or repo-relative file path to a clean relative path
* @param {string} filePath - The file path from the graph node
* @param {string} repoPath - Repository root path
* @returns {string} Relative path using forward slashes
*/
const toRelativePath = (filePath: string, repoPath: string): string => {
// Normalize to forward slashes for cross-platform consistency
const normalizedFile = filePath.replace(/\\/g, '/');
const normalizedRepo = repoPath.replace(/\\/g, '/');
if (normalizedFile.startsWith(normalizedRepo)) {
return normalizedFile.slice(normalizedRepo.length).replace(/^\//, '');
}
// Already relative or different root
return normalizedFile.replace(/^\//, '');
};
/**
* @brief Find the dominant (most common) top-level directory across files
* @param {FileInfo[]} files - File info entries
* @returns {string | null} Most common directory or null
*/
const getDominantDirectory = (files: FileInfo[]): string | null => {
const dirCounts = new Map<string, number>();
for (const f of files) {
const parts = f.relativePath.split('/');
if (parts.length >= 2) {
const dir = parts[0];
dirCounts.set(dir, (dirCounts.get(dir) || 0) + f.symbols.length);
}
}
let best: string | null = null;
let bestCount = 0;
for (const [dir, count] of dirCounts) {
if (count > bestCount) {
bestCount = count;
best = dir;
}
}
return best;
};
+13 -5
View File
@@ -4,12 +4,12 @@
* Shows the indexing status of the current repository.
*/
import { findRepo } from '../storage/repo-manager.js';
import { getCurrentCommit, isGitRepo } from '../storage/git.js';
import { findRepo, getStoragePaths, hasKuzuIndex } from '../storage/repo-manager.js';
import { getCurrentCommit, isGitRepo, getGitRoot } from '../storage/git.js';
export const statusCommand = async () => {
const cwd = process.cwd();
if (!isGitRepo(cwd)) {
console.log('Not a git repository.');
return;
@@ -17,8 +17,16 @@ export const statusCommand = async () => {
const repo = await findRepo(cwd);
if (!repo) {
console.log('Repository not indexed.');
console.log('Run: gitnexus analyze');
// Check if there's a stale KuzuDB index that needs migration
const repoRoot = getGitRoot(cwd) ?? cwd;
const { storagePath } = getStoragePaths(repoRoot);
if (await hasKuzuIndex(storagePath)) {
console.log('Repository has a stale KuzuDB index from a previous version.');
console.log('Run: gitnexus analyze (rebuilds the index with LadybugDB)');
} else {
console.log('Repository not indexed.');
console.log('Run: gitnexus analyze');
}
return;
}
+2 -2
View File
@@ -10,7 +10,7 @@
* gitnexus impact --target "AuthService" --direction upstream
* gitnexus cypher "MATCH (n:Function) RETURN n.name LIMIT 10"
*
* Note: Output goes to stderr because KuzuDB's native module captures stdout
* Note: Output goes to stderr because LadybugDB's native module captures stdout
* at the OS level during init. This is consistent with augment.ts.
*/
@@ -31,7 +31,7 @@ async function getBackend(): Promise<LocalBackend> {
function output(data: any): void {
const text = typeof data === 'string' ? data : JSON.stringify(data, null, 2);
// stderr because KuzuDB captures stdout at OS level
// stderr because LadybugDB captures stdout at OS level
process.stderr.write(text + '\n');
}
+2 -2
View File
@@ -101,7 +101,7 @@ export const wikiCommand = async (
}
// ── Check for existing index ────────────────────────────────────────
const { storagePath, kuzuPath } = getStoragePaths(repoPath);
const { storagePath, lbugPath } = getStoragePaths(repoPath);
const meta = await loadMeta(storagePath);
if (!meta) {
@@ -247,7 +247,7 @@ export const wikiCommand = async (
const generator = new WikiGenerator(
repoPath,
storagePath,
kuzuPath,
lbugPath,
llmConfig,
wikiOptions,
(phase, percent, detail) => {
+92
View File
@@ -1,3 +1,8 @@
import ignore, { type Ignore } from 'ignore';
import fs from 'fs/promises';
import nodePath from 'path';
import type { Path } from 'path-scurry';
const DEFAULT_IGNORE_LIST = new Set([
// Version Control
'.git',
@@ -186,6 +191,10 @@ const IGNORED_FILES = new Set([
// NOTE: Negation patterns in .gitnexusignore (e.g. `!vendor/`) cannot override
// entries in DEFAULT_IGNORE_LIST — this is intentional. The hardcoded list protects
// against indexing directories that are almost never source code (node_modules, .git, etc.).
// Users who need to include such directories should remove them from the hardcoded list.
export const shouldIgnorePath = (filePath: string): boolean => {
const normalizedPath = filePath.replace(/\\/g, '/');
const parts = normalizedPath.split('/');
@@ -237,3 +246,86 @@ export const shouldIgnorePath = (filePath: string): boolean => {
return false;
}
/** Check if a directory name is in the hardcoded ignore list */
export const isHardcodedIgnoredDirectory = (name: string): boolean => {
return DEFAULT_IGNORE_LIST.has(name);
};
/**
* Load .gitignore and .gitnexusignore rules from the repo root.
* Returns an `ignore` instance with all patterns, or null if no files found.
*/
export interface IgnoreOptions {
/** Skip .gitignore parsing, only read .gitnexusignore. Defaults to GITNEXUS_NO_GITIGNORE env var. */
noGitignore?: boolean;
}
export const loadIgnoreRules = async (
repoPath: string,
options?: IgnoreOptions
): Promise<Ignore | null> => {
const ig = ignore();
let hasRules = false;
// Allow users to bypass .gitignore parsing (e.g. when .gitignore accidentally excludes source files)
const skipGitignore = options?.noGitignore ?? !!process.env.GITNEXUS_NO_GITIGNORE;
const filenames = skipGitignore
? ['.gitnexusignore']
: ['.gitignore', '.gitnexusignore'];
for (const filename of filenames) {
try {
const content = await fs.readFile(nodePath.join(repoPath, filename), 'utf-8');
ig.add(content);
hasRules = true;
} catch (err: unknown) {
const code = (err as NodeJS.ErrnoException).code;
if (code !== 'ENOENT') {
console.warn(` Warning: could not read ${filename}: ${(err as Error).message}`);
}
}
}
return hasRules ? ig : null;
};
/**
* Create a glob-compatible ignore filter combining:
* - .gitignore / .gitnexusignore patterns (via `ignore` package)
* - Hardcoded DEFAULT_IGNORE_LIST, IGNORED_EXTENSIONS, IGNORED_FILES
*
* Returns an IgnoreLike object for glob's `ignore` option,
* enabling directory-level pruning during traversal.
*/
export const createIgnoreFilter = async (repoPath: string, options?: IgnoreOptions) => {
const ig = await loadIgnoreRules(repoPath, options);
return {
ignored(p: Path): boolean {
// path-scurry's Path.relative() returns POSIX paths on all platforms,
// which is what the `ignore` package expects. No explicit normalization needed.
const rel = p.relative();
if (!rel) return false;
// Check .gitignore / .gitnexusignore patterns
if (ig && ig.ignores(rel)) return true;
// Fall back to hardcoded rules
return shouldIgnorePath(rel);
},
childrenIgnored(p: Path): boolean {
// Fast path: check directory name against hardcoded list.
// Note: dot-directories (.git, .vscode, etc.) are primarily excluded by
// glob's `dot: false` option in filesystem-walker.ts. This check is
// defense-in-depth — do not remove `dot: false` assuming this covers it.
if (DEFAULT_IGNORE_LIST.has(p.name)) return true;
// Check against .gitignore / .gitnexusignore patterns.
// Test both bare path and path with trailing slash to handle
// bare-name patterns (e.g. `local`) and dir-only patterns (e.g. `local/`).
if (ig) {
const rel = p.relative();
if (rel && (ig.ignores(rel) || ig.ignores(rel + '/'))) return true;
}
return false;
},
};
};
+1 -1
View File
@@ -7,9 +7,9 @@ export enum SupportedLanguages {
CPlusPlus = 'cpp',
CSharp = 'csharp',
Go = 'go',
Ruby = 'ruby',
Rust = 'rust',
PHP = 'php',
Kotlin = 'kotlin',
// Ruby = 'ruby',
Swift = 'swift',
}
+102 -77
View File
@@ -24,7 +24,7 @@ import { listRegisteredRepos } from '../../storage/repo-manager.js';
async function findRepoForCwd(cwd: string): Promise<{
name: string;
storagePath: string;
kuzuPath: string;
lbugPath: string;
} | null> {
try {
const entries = await listRegisteredRepos({ validate: true });
@@ -66,7 +66,7 @@ async function findRepoForCwd(cwd: string): Promise<{
return {
name: bestMatch.name,
storagePath: bestMatch.storagePath,
kuzuPath: path.join(bestMatch.storagePath, 'kuzu'),
lbugPath: path.join(bestMatch.storagePath, 'lbug'),
};
} catch {
return null;
@@ -92,19 +92,19 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
const repo = await findRepoForCwd(workDir);
if (!repo) return '';
// Lazy-load kuzu adapter (skip unnecessary init)
const { initKuzu, executeQuery, isKuzuReady } = await import('../../mcp/core/kuzu-adapter.js');
const { searchFTSFromKuzu } = await import('../search/bm25-index.js');
// Lazy-load lbug adapter (skip unnecessary init)
const { initLbug, executeQuery, isLbugReady } = await import('../../mcp/core/lbug-adapter.js');
const { searchFTSFromLbug } = await import('../search/bm25-index.js');
const repoId = repo.name.toLowerCase();
// Init KuzuDB if not already
if (!isKuzuReady(repoId)) {
await initKuzu(repoId, repo.kuzuPath);
// Init LadybugDB if not already
if (!isLbugReady(repoId)) {
await initLbug(repoId, repo.lbugPath);
}
// Step 1: BM25 search (fast, no embeddings)
const bm25Results = await searchFTSFromKuzu(pattern, 10, repoId);
const bm25Results = await searchFTSFromLbug(pattern, 10, repoId);
if (bm25Results.length === 0) return '';
@@ -140,8 +140,90 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
if (symbolMatches.length === 0) return '';
// Step 3: For top matches, fetch callers/callees/processes
// Also get cluster cohesion internally for ranking
// Step 3: Batch-fetch callers/callees/processes/cohesion for top matches
// Uses batched WHERE n.id IN [...] queries instead of per-symbol queries
const uniqueSymbols = symbolMatches.slice(0, 5).filter((sym, i, arr) =>
arr.findIndex(s => s.nodeId === sym.nodeId) === i
);
if (uniqueSymbols.length === 0) return '';
const idList = uniqueSymbols.map(s => `'${s.nodeId.replace(/'/g, "''")}'`).join(', ');
// Batch fetch callers
const callersMap = new Map<string, string[]>();
try {
const rows = await executeQuery(repoId, `
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(n)
WHERE n.id IN [${idList}]
RETURN n.id AS targetId, caller.name AS name
LIMIT 15
`);
for (const r of rows) {
const tid = r.targetId || r[0];
const name = r.name || r[1];
if (tid && name) {
if (!callersMap.has(tid)) callersMap.set(tid, []);
callersMap.get(tid)!.push(name);
}
}
} catch { /* skip */ }
// Batch fetch callees
const calleesMap = new Map<string, string[]>();
try {
const rows = await executeQuery(repoId, `
MATCH (n)-[:CodeRelation {type: 'CALLS'}]->(callee)
WHERE n.id IN [${idList}]
RETURN n.id AS sourceId, callee.name AS name
LIMIT 15
`);
for (const r of rows) {
const sid = r.sourceId || r[0];
const name = r.name || r[1];
if (sid && name) {
if (!calleesMap.has(sid)) calleesMap.set(sid, []);
calleesMap.get(sid)!.push(name);
}
}
} catch { /* skip */ }
// Batch fetch processes
const processesMap = new Map<string, string[]>();
try {
const rows = await executeQuery(repoId, `
MATCH (n)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
WHERE n.id IN [${idList}]
RETURN n.id AS nodeId, p.heuristicLabel AS label, r.step AS step, p.stepCount AS stepCount
`);
for (const r of rows) {
const nid = r.nodeId || r[0];
const label = r.label || r[1];
const step = r.step || r[2];
const stepCount = r.stepCount || r[3];
if (nid && label) {
if (!processesMap.has(nid)) processesMap.set(nid, []);
processesMap.get(nid)!.push(`${label} (step ${step}/${stepCount})`);
}
}
} catch { /* skip */ }
// Batch fetch cohesion
const cohesionMap = new Map<string, number>();
try {
const rows = await executeQuery(repoId, `
MATCH (n)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
WHERE n.id IN [${idList}]
RETURN n.id AS nodeId, c.cohesion AS cohesion
`);
for (const r of rows) {
const nid = r.nodeId || r[0];
const coh = r.cohesion ?? r[1] ?? 0;
if (nid) cohesionMap.set(nid, coh);
}
} catch { /* skip */ }
// Assemble enriched results
const enriched: Array<{
name: string;
filePath: string;
@@ -150,72 +232,15 @@ export async function augment(pattern: string, cwd?: string): Promise<string> {
processes: string[];
cohesion: number;
}> = [];
const seen = new Set<string>();
for (const sym of symbolMatches.slice(0, 5)) {
if (seen.has(sym.nodeId)) continue;
seen.add(sym.nodeId);
const escaped = sym.nodeId.replace(/'/g, "''");
// Callers
let callers: string[] = [];
try {
const rows = await executeQuery(repoId, `
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(n {id: '${escaped}'})
RETURN caller.name AS name
LIMIT 3
`);
callers = rows.map((r: any) => r.name || r[0]).filter(Boolean);
} catch { /* skip */ }
// Callees
let callees: string[] = [];
try {
const rows = await executeQuery(repoId, `
MATCH (n {id: '${escaped}'})-[:CodeRelation {type: 'CALLS'}]->(callee)
RETURN callee.name AS name
LIMIT 3
`);
callees = rows.map((r: any) => r.name || r[0]).filter(Boolean);
} catch { /* skip */ }
// Processes
let processes: string[] = [];
try {
const rows = await executeQuery(repoId, `
MATCH (n {id: '${escaped}'})-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
RETURN p.heuristicLabel AS label, r.step AS step, p.stepCount AS stepCount
`);
processes = rows.map((r: any) => {
const label = r.label || r[0];
const step = r.step || r[1];
const stepCount = r.stepCount || r[2];
return `${label} (step ${step}/${stepCount})`;
}).filter(Boolean);
} catch { /* skip */ }
// Cluster cohesion (internal ranking signal)
let cohesion = 0;
try {
const rows = await executeQuery(repoId, `
MATCH (n {id: '${escaped}'})-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
RETURN c.cohesion AS cohesion
LIMIT 1
`);
if (rows.length > 0) {
cohesion = (rows[0].cohesion ?? rows[0][0]) || 0;
}
} catch { /* skip */ }
for (const sym of uniqueSymbols) {
enriched.push({
name: sym.name,
filePath: sym.filePath,
callers,
callees,
processes,
cohesion,
callers: (callersMap.get(sym.nodeId) || []).slice(0, 3),
callees: (calleesMap.get(sym.nodeId) || []).slice(0, 3),
processes: processesMap.get(sym.nodeId) || [],
cohesion: cohesionMap.get(sym.nodeId) || 0,
});
}
+1 -1
View File
@@ -262,7 +262,7 @@ export const embedBatch = async (texts: string[]): Promise<Float32Array[]> => {
};
/**
* Convert Float32Array to regular number array (for KuzuDB storage)
* Convert Float32Array to regular number array (for LadybugDB storage)
*/
export const embeddingToArray = (embedding: Float32Array): number[] => {
return Array.from(embedding);
@@ -2,10 +2,10 @@
* Embedding Pipeline Module
*
* Orchestrates the background embedding process:
* 1. Query embeddable nodes from KuzuDB
* 1. Query embeddable nodes from LadybugDB
* 2. Generate text representations
* 3. Batch embed using transformers.js
* 4. Update KuzuDB with embeddings
* 4. Update LadybugDB with embeddings
* 5. Create vector index for semantic search
*/
@@ -29,7 +29,7 @@ const isDev = process.env.NODE_ENV === 'development';
export type EmbeddingProgressCallback = (progress: EmbeddingProgress) => void;
/**
* Query all embeddable nodes from KuzuDB
* Query all embeddable nodes from LadybugDB
* Uses table-specific queries (File has different schema than code elements)
*/
const queryEmbeddableNodes = async (
@@ -104,9 +104,23 @@ const batchInsertEmbeddings = async (
* Create the vector index for semantic search
* Now indexes the separate CodeEmbedding table
*/
let vectorExtensionLoaded = false;
const createVectorIndex = async (
executeQuery: (cypher: string) => Promise<any[]>
): Promise<void> => {
// LadybugDB v0.15+ requires explicit VECTOR extension loading (once per session)
if (!vectorExtensionLoaded) {
try {
await executeQuery('INSTALL VECTOR');
await executeQuery('LOAD EXTENSION VECTOR');
vectorExtensionLoaded = true;
} catch {
// Extension may already be loaded — CREATE_VECTOR_INDEX will fail clearly if not
vectorExtensionLoaded = true;
}
}
const cypher = `
CALL CREATE_VECTOR_INDEX('CodeEmbedding', 'code_embedding_idx', 'embedding', metric := 'cosine')
`;
@@ -124,7 +138,7 @@ const createVectorIndex = async (
/**
* Run the embedding pipeline
*
* @param executeQuery - Function to execute Cypher queries against KuzuDB
* @param executeQuery - Function to execute Cypher queries against LadybugDB
* @param executeWithReusedStatement - Function to execute with reused prepared statement
* @param onProgress - Callback for progress updates
* @param config - Optional configuration override
@@ -219,7 +233,7 @@ export const runEmbeddingPipeline = async (
// Embed the batch
const embeddings = await embedBatch(texts);
// Update KuzuDB with embeddings
// Update LadybugDB with embeddings
const updates = batch.map((node, i) => ({
id: node.id,
embedding: embeddingToArray(embeddings[i]),
@@ -326,51 +340,64 @@ export const semanticSearch = async (
return [];
}
// Get metadata for each result by querying each node table
const results: SemanticSearchResult[] = [];
// Group results by label for batched metadata queries
const byLabel = new Map<string, Array<{ nodeId: string; distance: number }>>();
for (const embRow of embResults) {
const nodeId = embRow.nodeId ?? embRow[0];
const distance = embRow.distance ?? embRow[1];
// Extract label from node ID (format: Label:path:name)
const labelEndIdx = nodeId.indexOf(':');
const label = labelEndIdx > 0 ? nodeId.substring(0, labelEndIdx) : 'Unknown';
// Query the specific table for this node
// File nodes don't have startLine/endLine
if (!byLabel.has(label)) byLabel.set(label, []);
byLabel.get(label)!.push({ nodeId, distance });
}
// Batch-fetch metadata per label
const results: SemanticSearchResult[] = [];
for (const [label, items] of byLabel) {
const idList = items.map(i => `'${i.nodeId.replace(/'/g, "''")}'`).join(', ');
try {
let nodeQuery: string;
if (label === 'File') {
nodeQuery = `
MATCH (n:File {id: '${nodeId.replace(/'/g, "''")}'})
RETURN n.name AS name, n.filePath AS filePath
MATCH (n:File) WHERE n.id IN [${idList}]
RETURN n.id AS id, n.name AS name, n.filePath AS filePath
`;
} else {
nodeQuery = `
MATCH (n:${label} {id: '${nodeId.replace(/'/g, "''")}'})
RETURN n.name AS name, n.filePath AS filePath,
MATCH (n:${label}) WHERE n.id IN [${idList}]
RETURN n.id AS id, n.name AS name, n.filePath AS filePath,
n.startLine AS startLine, n.endLine AS endLine
`;
}
const nodeRows = await executeQuery(nodeQuery);
if (nodeRows.length > 0) {
const nodeRow = nodeRows[0];
results.push({
nodeId,
name: nodeRow.name ?? nodeRow[0] ?? '',
label,
filePath: nodeRow.filePath ?? nodeRow[1] ?? '',
distance,
startLine: label !== 'File' ? (nodeRow.startLine ?? nodeRow[2]) : undefined,
endLine: label !== 'File' ? (nodeRow.endLine ?? nodeRow[3]) : undefined,
});
const rowMap = new Map<string, any>();
for (const row of nodeRows) {
const id = row.id ?? row[0];
rowMap.set(id, row);
}
for (const item of items) {
const nodeRow = rowMap.get(item.nodeId);
if (nodeRow) {
results.push({
nodeId: item.nodeId,
name: nodeRow.name ?? nodeRow[1] ?? '',
label,
filePath: nodeRow.filePath ?? nodeRow[2] ?? '',
distance: item.distance,
startLine: label !== 'File' ? (nodeRow.startLine ?? nodeRow[3]) : undefined,
endLine: label !== 'File' ? (nodeRow.endLine ?? nodeRow[4]) : undefined,
});
}
}
} catch {
// Table might not exist, skip
}
}
// Re-sort by distance since batch queries may have mixed order
results.sort((a, b) => a.distance - b.distance);
return results;
};
+1 -1
View File
@@ -92,7 +92,7 @@ export interface SemanticSearchResult {
}
/**
* Node data for embedding (minimal structure from KuzuDB query)
* Node data for embedding (minimal structure from LadybugDB query)
*/
export interface EmbeddableNode {
id: string;
+12 -6
View File
@@ -35,12 +35,14 @@ export type NodeLabel =
| 'Template';
import { SupportedLanguages } from '../../config/supported-languages.js';
export type NodeProperties = {
name: string,
filePath: string,
startLine?: number,
endLine?: number,
language?: string,
language?: SupportedLanguages,
isExported?: boolean,
// Optional AST-derived framework hint (e.g. @Controller, @GetMapping)
astFrameworkMultiplier?: number,
@@ -61,19 +63,23 @@ export type NodeProperties = {
// Entry point scoring (computed by process detection)
entryPointScore?: number,
entryPointReason?: string,
// Method signature (for MRO disambiguation)
parameterCount?: number,
returnType?: string,
}
export type RelationshipType =
| 'CONTAINS'
| 'CALLS'
| 'INHERITS'
| 'OVERRIDES'
export type RelationshipType =
| 'CONTAINS'
| 'CALLS'
| 'INHERITS'
| 'OVERRIDES'
| 'IMPORTS'
| 'USES'
| 'DEFINES'
| 'DECORATES'
| 'IMPLEMENTS'
| 'EXTENDS'
| 'HAS_METHOD'
| 'MEMBER_OF'
| 'STEP_IN_PROCESS'
+3 -2
View File
@@ -10,10 +10,11 @@ export interface ASTCache {
}
export const createASTCache = (maxSize: number = 50): ASTCache => {
const effectiveMax = Math.max(maxSize, 1);
// Initialize the cache with a 'dispose' handler
// This is the magic: When an item is evicted (dropped), this runs automatically.
const cache = new LRUCache<string, Parser.Tree>({
max: maxSize,
max: effectiveMax,
dispose: (tree) => {
try {
// NOTE: web-tree-sitter has tree.delete(); native tree-sitter trees are GC-managed.
@@ -41,7 +42,7 @@ export const createASTCache = (maxSize: number = 50): ASTCache => {
stats: () => ({
size: cache.size,
maxSize: maxSize
maxSize: effectiveMax
})
};
};
File diff suppressed because it is too large Load Diff
+149
View File
@@ -0,0 +1,149 @@
/**
* Shared Ruby call routing logic.
*
* Ruby expresses imports, heritage (mixins), and property definitions as
* method calls rather than syntax-level constructs. This module provides a
* routing function used by the CLI call-processor, CLI parse-worker, and
* the web call-processor so that the classification logic lives in one place.
*
* NOTE: This file is intentionally duplicated in gitnexus-web/ because the
* two packages have separate build targets (Node native vs WASM/browser).
* Keep both copies in sync until a shared package is introduced.
*/
import { SupportedLanguages } from '../../config/supported-languages.js';
// ── Call routing dispatch table ─────────────────────────────────────────────
/** null = this call was not routed; fall through to default call handling */
export type CallRoutingResult = RubyCallRouting | null;
export type CallRouter = (
calledName: string,
callNode: any,
) => CallRoutingResult;
/** No-op router: returns null for every call (passthrough to normal processing) */
const noRouting: CallRouter = () => null;
/** Per-language call routing. noRouting = no special routing (normal call processing) */
export const callRouters: Record<SupportedLanguages, CallRouter> = {
[SupportedLanguages.JavaScript]: noRouting,
[SupportedLanguages.TypeScript]: noRouting,
[SupportedLanguages.Python]: noRouting,
[SupportedLanguages.Java]: noRouting,
[SupportedLanguages.Kotlin]: noRouting,
[SupportedLanguages.Go]: noRouting,
[SupportedLanguages.Rust]: noRouting,
[SupportedLanguages.CSharp]: noRouting,
[SupportedLanguages.PHP]: noRouting,
[SupportedLanguages.Swift]: noRouting,
[SupportedLanguages.CPlusPlus]: noRouting,
[SupportedLanguages.C]: noRouting,
[SupportedLanguages.Ruby]: routeRubyCall,
};
// ── Result types ────────────────────────────────────────────────────────────
export type RubyCallRouting =
| { kind: 'import'; importPath: string; isRelative: boolean }
| { kind: 'heritage'; items: RubyHeritageItem[] }
| { kind: 'properties'; items: RubyPropertyItem[] }
| { kind: 'call' }
| { kind: 'skip' };
export interface RubyHeritageItem {
enclosingClass: string;
mixinName: string;
heritageKind: 'include' | 'extend' | 'prepend';
}
export type RubyAccessorType = 'attr_accessor' | 'attr_reader' | 'attr_writer';
export interface RubyPropertyItem {
propName: string;
accessorType: RubyAccessorType;
startLine: number;
endLine: number;
}
// ── Pre-allocated singletons for common return values ────────────────────────
const CALL_RESULT: RubyCallRouting = { kind: 'call' };
const SKIP_RESULT: RubyCallRouting = { kind: 'skip' };
/** Max depth for parent-walking loops to prevent pathological AST traversals */
const MAX_PARENT_DEPTH = 50;
// ── Routing function ────────────────────────────────────────────────────────
/**
* Classify a Ruby call node and extract its semantic payload.
*
* @param calledName - The method name (e.g. 'require', 'include', 'attr_accessor')
* @param callNode - The tree-sitter `call` AST node
* @returns A discriminated union describing the call's semantic role
*/
export function routeRubyCall(calledName: string, callNode: any): RubyCallRouting {
// ── require / require_relative → import ─────────────────────────────────
if (calledName === 'require' || calledName === 'require_relative') {
const argList = callNode.childForFieldName?.('arguments');
const stringNode = argList?.children?.find((c: any) => c.type === 'string');
const contentNode = stringNode?.children?.find((c: any) => c.type === 'string_content');
if (!contentNode) return SKIP_RESULT;
let importPath: string = contentNode.text;
// Validate: reject null bytes, control chars, excessively long paths
if (!importPath || importPath.length > 1024 || /[\x00-\x1f]/.test(importPath)) {
return SKIP_RESULT;
}
const isRelative = calledName === 'require_relative';
if (isRelative && !importPath.startsWith('.')) {
importPath = './' + importPath;
}
return { kind: 'import', importPath, isRelative };
}
// ── include / extend / prepend → heritage (mixin) ──────────────────────
if (calledName === 'include' || calledName === 'extend' || calledName === 'prepend') {
let enclosingClass: string | null = null;
let current = callNode.parent;
let depth = 0;
while (current && ++depth <= MAX_PARENT_DEPTH) {
if (current.type === 'class' || current.type === 'module') {
const nameNode = current.childForFieldName?.('name');
if (nameNode) { enclosingClass = nameNode.text; break; }
}
current = current.parent;
}
if (!enclosingClass) return SKIP_RESULT;
const items: RubyHeritageItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of (argList?.children ?? [])) {
if (arg.type === 'constant' || arg.type === 'scope_resolution') {
items.push({ enclosingClass, mixinName: arg.text, heritageKind: calledName as 'include' | 'extend' | 'prepend' });
}
}
return items.length > 0 ? { kind: 'heritage', items } : SKIP_RESULT;
}
// ── attr_accessor / attr_reader / attr_writer → property definitions ───
if (calledName === 'attr_accessor' || calledName === 'attr_reader' || calledName === 'attr_writer') {
const items: RubyPropertyItem[] = [];
const argList = callNode.childForFieldName?.('arguments');
for (const arg of (argList?.children ?? [])) {
if (arg.type === 'simple_symbol') {
items.push({
propName: arg.text.startsWith(':') ? arg.text.slice(1) : arg.text,
accessorType: calledName as RubyAccessorType,
startLine: arg.startPosition.row,
endLine: arg.endPosition.row,
});
}
}
return items.length > 0 ? { kind: 'properties', items } : SKIP_RESULT;
}
// ── Everything else → regular call ─────────────────────────────────────
return CALL_RESULT;
}
+19
View File
@@ -0,0 +1,19 @@
/**
* Default minimum buffer size for tree-sitter parsing (512 KB).
* tree-sitter requires bufferSize >= file size in bytes.
*/
export const TREE_SITTER_BUFFER_SIZE = 512 * 1024;
/**
* Maximum buffer size cap (32 MB) to prevent OOM on huge files.
* Also used as the file-size skip threshold — files larger than this are not parsed.
*/
export const TREE_SITTER_MAX_BUFFER = 32 * 1024 * 1024;
/**
* Compute adaptive buffer size for tree-sitter parsing.
* Uses 2× file size, clamped between 512 KB and 32 MB.
* Previous 256 KB fixed limit silently skipped files > ~200 KB (e.g., imgui.h at 411 KB).
*/
export const getTreeSitterBufferSize = (contentLength: number): number =>
Math.min(Math.max(contentLength * 2, TREE_SITTER_BUFFER_SIZE), TREE_SITTER_MAX_BUFFER);
@@ -11,9 +11,10 @@
*/
import { detectFrameworkFromPath } from './framework-detection.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
// ============================================================================
// NAME PATTERNS - All 9 supported languages
// NAME PATTERNS - All 11 supported languages
// ============================================================================
/**
@@ -38,39 +39,47 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
],
// JavaScript/TypeScript
'javascript': [
[SupportedLanguages.JavaScript]: [
/^use[A-Z]/, // React hooks (useEffect, etc.)
],
'typescript': [
[SupportedLanguages.TypeScript]: [
/^use[A-Z]/, // React hooks
],
// Python
'python': [
[SupportedLanguages.Python]: [
/^app$/, // Flask/FastAPI app
/^(get|post|put|delete|patch)_/i, // REST conventions
/^api_/, // API functions
/^view_/, // Django views
],
// Java
'java': [
[SupportedLanguages.Java]: [
/^do[A-Z]/, // doGet, doPost (Servlets)
/^create[A-Z]/, // Factory patterns
/^build[A-Z]/, // Builder patterns
/Service$/, // UserService
],
// C#
'csharp': [
/^(Get|Post|Put|Delete)/, // ASP.NET conventions
/Action$/, // MVC actions
/^On[A-Z]/, // Event handlers
/Async$/, // Async entry points
[SupportedLanguages.CSharp]: [
/^(Get|Post|Put|Delete|Patch)/, // ASP.NET action methods
/Action$/, // MVC actions
/^On[A-Z]/, // Event handlers / Blazor lifecycle
/Async$/, // Async entry points
/^Configure$/, // Startup.Configure
/^ConfigureServices$/, // Startup.ConfigureServices
/^Handle$/, // MediatR / generic handler
/^Execute$/, // Command pattern
/^Invoke$/, // Middleware Invoke
/^Map[A-Z]/, // Minimal API MapGet, MapPost
/Service$/, // Service classes
/^Seed/, // Database seeding
],
// Go
'go': [
[SupportedLanguages.Go]: [
/Handler$/, // http.Handler pattern
/^Serve/, // ServeHTTP
/^New[A-Z]/, // Constructor pattern (returns new instance)
@@ -78,7 +87,7 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
],
// Rust
'rust': [
[SupportedLanguages.Rust]: [
/^(get|post|put|delete)_handler$/i,
/^handle_/, // handle_request
/^new$/, // Constructor pattern
@@ -86,25 +95,64 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
/^spawn/, // Async spawn
],
// C - explicit main() boost (critical for C programs)
'c': [
// C - explicit main() boost plus common C entry point conventions
[SupportedLanguages.C]: [
/^main$/, // THE entry point
/^init_/, // Initialization functions
/^start_/, // Start functions
/^run_/, // Run functions
/^init_/, // init_server, init_client
/_init$/, // module_init, server_init
/^start_/, // start_server
/_start$/, // thread_start
/^run_/, // run_loop
/_run$/, // event_run
/^stop_/, // stop_server
/_stop$/, // service_stop
/^open_/, // open_connection
/_open$/, // file_open
/^close_/, // close_connection
/_close$/, // socket_close
/^create_/, // create_session
/_create$/, // object_create
/^destroy_/, // destroy_session
/_destroy$/, // object_destroy
/^handle_/, // handle_request
/_handler$/, // signal_handler
/_callback$/, // event_callback
/^cmd_/, // tmux: cmd_new_window, cmd_attach_session
/^server_/, // server_start, server_loop
/^client_/, // client_connect
/^session_/, // session_create
/^window_/, // window_resize (tmux)
/^key_/, // key_press
/^input_/, // input_parse
/^output_/, // output_write
/^notify_/, // notify_client
/^control_/, // control_start
],
// C++ - same as C plus class patterns
'cpp': [
// C++ - same as C plus OOP/template patterns
[SupportedLanguages.CPlusPlus]: [
/^main$/, // THE entry point
/^init_/,
/_init$/,
/^Create[A-Z]/, // Factory patterns
/^create_/,
/^Run$/, // Run methods
/^run$/,
/^Start$/, // Start methods
/^start$/,
/^handle_/,
/_handler$/,
/_callback$/,
/^OnEvent/, // Event callbacks
/^on_/,
/::Run$/, // Class::Run
/::Start$/, // Class::Start
/::Init$/, // Class::Init
/::Execute$/, // Class::Execute
],
// Swift / iOS
'swift': [
[SupportedLanguages.Swift]: [
/^viewDidLoad$/, // UIKit lifecycle
/^viewWillAppear$/, // UIKit lifecycle
/^viewDidAppear$/, // UIKit lifecycle
@@ -124,7 +172,7 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
],
// PHP / Laravel
'php': [
[SupportedLanguages.PHP]: [
/Controller$/, // UserController (class name convention)
/^handle$/, // Job::handle(), Listener::handle()
/^execute$/, // Command::execute()
@@ -143,8 +191,23 @@ const ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {
/^save$/, // Repository::save()
/^delete$/, // Repository::delete()
],
// Ruby
[SupportedLanguages.Ruby]: [
/^call$/, // Service objects (MyService.call)
/^perform$/, // Background jobs (Sidekiq, ActiveJob)
/^execute$/, // Command pattern
],
};
/** Pre-computed merged patterns (universal + language-specific) to avoid per-call array allocation. */
const MERGED_ENTRY_POINT_PATTERNS: Record<string, RegExp[]> = {};
const UNIVERSAL_PATTERNS = ENTRY_POINT_PATTERNS['*'] || [];
for (const [lang, patterns] of Object.entries(ENTRY_POINT_PATTERNS)) {
if (lang === '*') continue;
MERGED_ENTRY_POINT_PATTERNS[lang] = [...UNIVERSAL_PATTERNS, ...patterns];
}
// ============================================================================
// UTILITY PATTERNS - Functions that should be penalized
// ============================================================================
@@ -199,7 +262,7 @@ export interface EntryPointScoreResult {
*/
export function calculateEntryPointScore(
name: string,
language: string,
language: SupportedLanguages,
isExported: boolean,
callerCount: number,
calleeCount: number,
@@ -232,9 +295,7 @@ export function calculateEntryPointScore(
reasons.push('utility-pattern');
} else {
// Check positive patterns
const universalPatterns = ENTRY_POINT_PATTERNS['*'] || [];
const langPatterns = ENTRY_POINT_PATTERNS[language] || [];
const allPatterns = [...universalPatterns, ...langPatterns];
const allPatterns = MERGED_ENTRY_POINT_PATTERNS[language] || UNIVERSAL_PATTERNS;
if (allPatterns.some(p => p.test(name))) {
nameMultiplier = 1.5; // Bonus for matching entry point pattern
@@ -296,13 +357,23 @@ export function isTestFile(filePath: string): boolean {
p.endsWith('test.swift') ||
p.includes('uitests/') ||
// C# test patterns
p.endsWith('tests.cs') ||
p.endsWith('test.cs') ||
p.includes('.tests/') ||
p.includes('tests.cs') ||
p.includes('.test/') ||
p.includes('.integrationtests/') ||
p.includes('.unittests/') ||
p.includes('/testproject/') ||
// PHP/Laravel test patterns
p.endsWith('test.php') ||
p.endsWith('spec.php') ||
p.includes('/tests/feature/') ||
p.includes('/tests/unit/')
p.includes('/tests/unit/') ||
// Ruby test patterns
p.endsWith('_spec.rb') ||
p.endsWith('_test.rb') ||
p.includes('/spec/') ||
p.includes('/test/fixtures/')
);
}
@@ -0,0 +1,243 @@
/**
* Export Detection
*
* Determines whether a symbol (function, class, etc.) is exported/public
* in its language. This is a pure function — safe for use in worker threads.
*
* Shared between parse-worker.ts (worker pool) and parsing-processor.ts (sequential fallback).
*/
import { findSiblingChild, SyntaxNode } from './utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
/** Handler type: given a node and symbol name, return true if the symbol is exported/public. */
type ExportChecker = (node: SyntaxNode, name: string) => boolean;
// ============================================================================
// Per-language export checkers
// ============================================================================
/** JS/TS: walk ancestors looking for export_statement or export_specifier. */
const tsExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
const type = current.type;
if (type === 'export_statement' ||
type === 'export_specifier' ||
(type === 'lexical_declaration' && current.parent?.type === 'export_statement')) {
return true;
}
// Fallback: check if node text starts with 'export ' for edge cases
if (current.text?.startsWith('export ')) {
return true;
}
current = current.parent;
}
return false;
};
/** Python: public if no leading underscore (convention). */
const pythonExportChecker: ExportChecker = (_node, name) => !name.startsWith('_');
/** Java: check for 'public' modifier — modifiers are siblings of the name node, not parents. */
const javaExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.parent) {
const parent = current.parent;
for (let i = 0; i < parent.childCount; i++) {
const child = parent.child(i);
if (child?.type === 'modifiers' && child.text?.includes('public')) {
return true;
}
}
if (parent.type === 'method_declaration' || parent.type === 'constructor_declaration') {
if (parent.text?.trimStart().startsWith('public')) {
return true;
}
}
}
current = current.parent;
}
return false;
};
/** C# declaration node types for sibling modifier scanning. */
const CSHARP_DECL_TYPES = new Set([
'method_declaration', 'local_function_statement', 'constructor_declaration',
'class_declaration', 'interface_declaration', 'struct_declaration',
'enum_declaration', 'record_declaration', 'record_struct_declaration',
'record_class_declaration', 'delegate_declaration',
'property_declaration', 'field_declaration', 'event_declaration',
'namespace_declaration', 'file_scoped_namespace_declaration',
]);
/**
* C#: modifier nodes are SIBLINGS of the name node inside the declaration.
* Walk up to the declaration node, then scan its direct children.
*/
const csharpExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (CSHARP_DECL_TYPES.has(current.type)) {
for (let i = 0; i < current.childCount; i++) {
const child = current.child(i);
if (child?.type === 'modifier' && child.text === 'public') return true;
}
return false;
}
current = current.parent;
}
return false;
};
/** Go: uppercase first letter = exported. */
const goExportChecker: ExportChecker = (_node, name) => {
if (name.length === 0) return false;
const first = name[0];
return first === first.toUpperCase() && first !== first.toLowerCase();
};
/** Rust declaration node types for sibling visibility_modifier scanning. */
const RUST_DECL_TYPES = new Set([
'function_item', 'struct_item', 'enum_item', 'trait_item', 'impl_item',
'union_item', 'type_item', 'const_item', 'static_item', 'mod_item',
'use_declaration', 'associated_type', 'function_signature_item',
]);
/**
* Rust: visibility_modifier is a SIBLING of the name node within the declaration node
* (function_item, struct_item, etc.), not a parent. Walk up to the declaration node,
* then scan its direct children.
*/
const rustExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (RUST_DECL_TYPES.has(current.type)) {
for (let i = 0; i < current.childCount; i++) {
const child = current.child(i);
if (child?.type === 'visibility_modifier' && child.text?.startsWith('pub')) return true;
}
return false;
}
current = current.parent;
}
return false;
};
/**
* Kotlin: default visibility is public (unlike Java).
* visibility_modifier is inside modifiers, a sibling of the name node within the declaration.
*/
const kotlinExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.parent) {
const visMod = findSiblingChild(current.parent, 'modifiers', 'visibility_modifier');
if (visMod) {
const text = visMod.text;
if (text === 'private' || text === 'internal' || text === 'protected') return false;
if (text === 'public') return true;
}
}
current = current.parent;
}
// No visibility modifier = public (Kotlin default)
return true;
};
/**
* C/C++: functions without 'static' storage class have external linkage by default,
* making them globally accessible (equivalent to exported). Only functions explicitly
* marked 'static' are file-scoped (not exported). C++ anonymous namespaces
* (namespace { ... }) also give internal linkage.
*/
const cCppExportChecker: ExportChecker = (node, _name) => {
let cur: SyntaxNode | null = node;
while (cur) {
if (cur.type === 'function_definition' || cur.type === 'declaration') {
// Check for 'static' storage class specifier as a direct child node.
// This avoids reading the full function text (which can be very large).
for (let i = 0; i < cur.childCount; i++) {
const child = cur.child(i);
if (child?.type === 'storage_class_specifier' && child.text === 'static') return false;
}
}
// C++ anonymous namespace: namespace_definition with no name child = internal linkage
if (cur.type === 'namespace_definition') {
const hasName = cur.childForFieldName?.('name');
if (!hasName) return false;
}
cur = cur.parent;
}
return true; // Top-level C/C++ functions default to external linkage
};
/** PHP: check for visibility modifier or top-level scope. */
const phpExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.type === 'class_declaration' ||
current.type === 'interface_declaration' ||
current.type === 'trait_declaration' ||
current.type === 'enum_declaration') {
return true;
}
if (current.type === 'visibility_modifier') {
return current.text === 'public';
}
current = current.parent;
}
// Top-level functions are globally accessible
return true;
};
/** Swift: check for 'public' or 'open' access modifiers. */
const swiftExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.type === 'modifiers' || current.type === 'visibility_modifier') {
const text = current.text || '';
if (text.includes('public') || text.includes('open')) return true;
}
current = current.parent;
}
return false;
};
// ============================================================================
// Exhaustive dispatch table — satisfies enforces all SupportedLanguages are covered
// ============================================================================
const exportCheckers = {
[SupportedLanguages.JavaScript]: tsExportChecker,
[SupportedLanguages.TypeScript]: tsExportChecker,
[SupportedLanguages.Python]: pythonExportChecker,
[SupportedLanguages.Java]: javaExportChecker,
[SupportedLanguages.CSharp]: csharpExportChecker,
[SupportedLanguages.Go]: goExportChecker,
[SupportedLanguages.Rust]: rustExportChecker,
[SupportedLanguages.Kotlin]: kotlinExportChecker,
[SupportedLanguages.C]: cCppExportChecker,
[SupportedLanguages.CPlusPlus]: cCppExportChecker,
[SupportedLanguages.PHP]: phpExportChecker,
[SupportedLanguages.Swift]: swiftExportChecker,
[SupportedLanguages.Ruby]: (_node, _name) => true,
} satisfies Record<SupportedLanguages, ExportChecker>;
// ============================================================================
// Public API
// ============================================================================
/**
* Check if a tree-sitter node is exported/public in its language.
* @param node - The tree-sitter AST node
* @param name - The symbol name
* @param language - The programming language
* @returns true if the symbol is exported/public
*/
export const isNodeExported = (node: SyntaxNode, name: string, language: SupportedLanguages): boolean => {
const checker = exportCheckers[language];
if (!checker) return false;
return checker(node, name);
};
@@ -1,7 +1,7 @@
import fs from 'fs/promises';
import path from 'path';
import { glob } from 'glob';
import { shouldIgnorePath } from '../../config/ignore-service.js';
import { createIgnoreFilter } from '../../config/ignore-service.js';
export interface FileEntry {
path: string;
@@ -32,13 +32,14 @@ export const walkRepositoryPaths = async (
repoPath: string,
onProgress?: (current: number, total: number, filePath: string) => void
): Promise<ScannedFile[]> => {
const files = await glob('**/*', {
const ignoreFilter = await createIgnoreFilter(repoPath);
const filtered = await glob('**/*', {
cwd: repoPath,
nodir: true,
dot: false,
ignore: ignoreFilter,
});
const filtered = files.filter(file => !shouldIgnorePath(file));
const entries: ScannedFile[] = [];
let processed = 0;
let skippedLarge = 0;
@@ -183,7 +183,35 @@ export function detectFrameworkFromPath(filePath: string): FrameworkHint | null
if (p.endsWith('controller.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 3.0, reason: 'aspnet-controller-file' };
}
// ASP.NET Services
if ((p.includes('/services/') || p.includes('/service/')) && p.endsWith('.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 1.8, reason: 'aspnet-service' };
}
// ASP.NET Middleware
if (p.includes('/middleware/') && p.endsWith('.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 2.5, reason: 'aspnet-middleware' };
}
// SignalR Hubs
if (p.includes('/hubs/') && p.endsWith('.cs')) {
return { framework: 'signalr', entryPointMultiplier: 2.5, reason: 'signalr-hub' };
}
if (p.endsWith('hub.cs')) {
return { framework: 'signalr', entryPointMultiplier: 2.5, reason: 'signalr-hub-file' };
}
// Minimal API / Program.cs / Startup.cs
if (p.endsWith('/program.cs') || p.endsWith('/startup.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 3.0, reason: 'aspnet-entry' };
}
// Background services / Hosted services
if ((p.includes('/backgroundservices/') || p.includes('/hostedservices/')) && p.endsWith('.cs')) {
return { framework: 'aspnet', entryPointMultiplier: 2.0, reason: 'aspnet-background-service' };
}
// Blazor pages
if (p.includes('/pages/') && p.endsWith('.razor')) {
return { framework: 'blazor', entryPointMultiplier: 2.5, reason: 'blazor-page' };
@@ -302,6 +330,18 @@ export function detectFrameworkFromPath(filePath: string): FrameworkHint | null
return { framework: 'laravel', entryPointMultiplier: 1.5, reason: 'laravel-repository' };
}
// ========== RUBY ==========
// Ruby: bin/ or exe/ (CLI entry points)
if ((p.includes('/bin/') || p.includes('/exe/')) && p.endsWith('.rb')) {
return { framework: 'ruby', entryPointMultiplier: 2.5, reason: 'ruby-executable' };
}
// Ruby: Rakefile or *.rake (task definitions)
if (p.endsWith('/rakefile') || p.endsWith('.rake')) {
return { framework: 'ruby', entryPointMultiplier: 1.5, reason: 'ruby-rake' };
}
// ========== SWIFT / iOS ==========
// iOS App entry points (highest priority)
@@ -385,7 +425,11 @@ export const FRAMEWORK_AST_PATTERNS = {
'jaxrs': ['@Path', '@GET', '@POST', '@PUT', '@DELETE'],
// C# attributes
'aspnet': ['[ApiController]', '[HttpGet]', '[HttpPost]', '[Route]'],
'aspnet': ['[ApiController]', '[HttpGet]', '[HttpPost]', '[HttpPut]', '[HttpDelete]',
'[Route]', '[Authorize]', '[AllowAnonymous]'],
'signalr': ['[HubMethodName]', ': Hub', ': Hub<'],
'blazor': ['@page', '[Parameter]', '@inject'],
'efcore': ['DbContext', 'DbSet<', 'OnModelCreating'],
// Go patterns (function signatures)
'go-http': ['http.Handler', 'http.HandlerFunc', 'ServeHTTP'],
@@ -405,6 +449,8 @@ export const FRAMEWORK_AST_PATTERNS = {
'combine': ['sink', 'assign', 'Publisher', 'Subscriber'],
};
import { SupportedLanguages } from '../../config/supported-languages.js';
interface AstFrameworkPatternConfig {
framework: string;
entryPointMultiplier: number;
@@ -413,30 +459,33 @@ interface AstFrameworkPatternConfig {
}
const AST_FRAMEWORK_PATTERNS_BY_LANGUAGE: Record<string, AstFrameworkPatternConfig[]> = {
javascript: [
[SupportedLanguages.JavaScript]: [
{ framework: 'nestjs', entryPointMultiplier: 3.2, reason: 'nestjs-decorator', patterns: FRAMEWORK_AST_PATTERNS.nestjs },
],
typescript: [
[SupportedLanguages.TypeScript]: [
{ framework: 'nestjs', entryPointMultiplier: 3.2, reason: 'nestjs-decorator', patterns: FRAMEWORK_AST_PATTERNS.nestjs },
],
python: [
[SupportedLanguages.Python]: [
{ framework: 'fastapi', entryPointMultiplier: 3.0, reason: 'fastapi-decorator', patterns: FRAMEWORK_AST_PATTERNS.fastapi },
{ framework: 'flask', entryPointMultiplier: 2.8, reason: 'flask-decorator', patterns: FRAMEWORK_AST_PATTERNS.flask },
],
java: [
[SupportedLanguages.Java]: [
{ framework: 'spring', entryPointMultiplier: 3.2, reason: 'spring-annotation', patterns: FRAMEWORK_AST_PATTERNS.spring },
{ framework: 'jaxrs', entryPointMultiplier: 3.0, reason: 'jaxrs-annotation', patterns: FRAMEWORK_AST_PATTERNS.jaxrs },
],
kotlin: [
[SupportedLanguages.Kotlin]: [
{ framework: 'spring-kotlin', entryPointMultiplier: 3.2, reason: 'spring-kotlin-annotation', patterns: FRAMEWORK_AST_PATTERNS.spring },
{ framework: 'jaxrs', entryPointMultiplier: 3.0, reason: 'jaxrs-annotation', patterns: FRAMEWORK_AST_PATTERNS.jaxrs },
{ framework: 'ktor', entryPointMultiplier: 2.8, reason: 'ktor-routing', patterns: ['routing', 'embeddedServer', 'Application.module'] },
{ framework: 'android-kotlin', entryPointMultiplier: 2.5, reason: 'android-annotation', patterns: ['@AndroidEntryPoint', 'AppCompatActivity', 'Fragment('] },
],
csharp: [
[SupportedLanguages.CSharp]: [
{ framework: 'aspnet', entryPointMultiplier: 3.2, reason: 'aspnet-attribute', patterns: FRAMEWORK_AST_PATTERNS.aspnet },
{ framework: 'signalr', entryPointMultiplier: 2.8, reason: 'signalr-attribute', patterns: FRAMEWORK_AST_PATTERNS.signalr },
{ framework: 'blazor', entryPointMultiplier: 2.5, reason: 'blazor-attribute', patterns: FRAMEWORK_AST_PATTERNS.blazor },
{ framework: 'efcore', entryPointMultiplier: 2.0, reason: 'efcore-pattern', patterns: FRAMEWORK_AST_PATTERNS.efcore },
],
php: [
[SupportedLanguages.PHP]: [
{ framework: 'laravel', entryPointMultiplier: 3.0, reason: 'php-route-attribute', patterns: FRAMEWORK_AST_PATTERNS.laravel },
],
};
@@ -456,7 +505,7 @@ const AST_PATTERNS_LOWERED: Record<string, Array<{ framework: string; entryPoint
* Note: callers should slice definitionText to ~300 chars since annotations appear at the start.
*/
export function detectFrameworkFromAST(
language: string,
language: SupportedLanguages,
definitionText: string
): FrameworkHint | null {
if (!language || !definitionText) return null;
+130 -68
View File
@@ -1,29 +1,99 @@
/**
* Heritage Processor
*
*
* Extracts class inheritance relationships:
* - EXTENDS: Class extends another Class (TS, JS, Python)
* - IMPLEMENTS: Class implements an Interface (TS only)
* - EXTENDS: Class extends another Class (TS, JS, Python, C#, C++)
* - IMPLEMENTS: Class implements an Interface (TS, C#, Java, Kotlin, PHP)
*
* Languages like C# use a single `base_list` for both class and interface parents.
* We resolve the correct edge type by checking the symbol table: if the parent is
* registered as an Interface, we emit IMPLEMENTS; otherwise EXTENDS. For unresolved
* external symbols, the fallback heuristic is language-gated:
* - C# / Java: apply the `I[A-Z]` naming convention (e.g. IDisposable → IMPLEMENTS)
* - Swift: default to IMPLEMENTS (protocol conformance is more common than class inheritance)
* - All other languages: default to EXTENDS
*/
import { KnowledgeGraph } from '../graph/types.js';
import { ASTCache } from './ast-cache.js';
import { SymbolTable } from './symbol-table.js';
import Parser from 'tree-sitter';
import { loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { isLanguageAvailable, loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { generateId } from '../../lib/utils.js';
import { getLanguageFromFilename, yieldToEventLoop } from './utils.js';
import { getLanguageFromFilename, isVerboseIngestionEnabled, yieldToEventLoop } from './utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
import { getTreeSitterBufferSize } from './constants.js';
import type { ExtractedHeritage } from './workers/parse-worker.js';
import type { ResolutionContext } from './resolution-context.js';
/** C#/Java convention: interfaces start with I followed by an uppercase letter */
const INTERFACE_NAME_RE = /^I[A-Z]/;
/**
* Determine whether a heritage.extends capture is actually an IMPLEMENTS relationship.
* Uses the symbol table first (authoritative — Tier 1); falls back to a language-gated
* heuristic for external symbols not present in the graph:
* - C# / Java: `I[A-Z]` naming convention
* - Swift: default IMPLEMENTS (protocol conformance is the norm)
* - All others: default EXTENDS
*/
const resolveExtendsType = (
parentName: string,
currentFilePath: string,
ctx: ResolutionContext,
language: SupportedLanguages,
): { type: 'EXTENDS' | 'IMPLEMENTS'; idPrefix: string } => {
const resolved = ctx.resolve(parentName, currentFilePath);
if (resolved && resolved.candidates.length > 0) {
const isInterface = resolved.candidates[0].type === 'Interface';
return isInterface
? { type: 'IMPLEMENTS', idPrefix: 'Interface' }
: { type: 'EXTENDS', idPrefix: 'Class' };
}
// Unresolved symbol — fall back to language-specific heuristic
if (language === SupportedLanguages.CSharp || language === SupportedLanguages.Java) {
if (INTERFACE_NAME_RE.test(parentName)) {
return { type: 'IMPLEMENTS', idPrefix: 'Interface' };
}
} else if (language === SupportedLanguages.Swift) {
// Protocol conformance is far more common than class inheritance in Swift
return { type: 'IMPLEMENTS', idPrefix: 'Interface' };
}
return { type: 'EXTENDS', idPrefix: 'Class' };
};
/**
* Resolve a symbol ID for heritage, with fallback to generated ID.
* Uses ctx.resolve() → pick first candidate's nodeId → generate synthetic ID.
*/
const resolveHeritageId = (
name: string,
filePath: string,
ctx: ResolutionContext,
fallbackLabel: string,
fallbackKey?: string,
): string => {
const resolved = ctx.resolve(name, filePath);
if (resolved && resolved.candidates.length > 0) {
// For global with multiple candidates, refuse (a wrong edge is worse than no edge)
if (resolved.tier === 'global' && resolved.candidates.length > 1) {
return generateId(fallbackLabel, fallbackKey ?? name);
}
return resolved.candidates[0].nodeId;
}
return generateId(fallbackLabel, fallbackKey ?? name);
};
export const processHeritage = async (
graph: KnowledgeGraph,
files: { path: string; content: string }[],
astCache: ASTCache,
symbolTable: SymbolTable,
onProgress?: (current: number, total: number) => void
ctx: ResolutionContext,
onProgress?: (current: number, total: number) => void,
) => {
const parser = await loadParser();
const logSkipped = isVerboseIngestionEnabled();
const skippedByLang = logSkipped ? new Map<string, number>() : null;
for (let i = 0; i < files.length; i++) {
const file = files[i];
@@ -33,6 +103,12 @@ export const processHeritage = async (
// 1. Check language support
const language = getLanguageFromFilename(file.path);
if (!language) continue;
if (!isLanguageAvailable(language)) {
if (skippedByLang) {
skippedByLang.set(language, (skippedByLang.get(language) ?? 0) + 1);
}
continue;
}
const queryStr = LANGUAGE_QUERIES[language];
if (!queryStr) continue;
@@ -42,17 +118,14 @@ export const processHeritage = async (
// 3. Get AST
let tree = astCache.get(file.path);
let wasReparsed = false;
if (!tree) {
// Use larger bufferSize for files > 32KB
try {
tree = parser.parse(file.content, undefined, { bufferSize: 1024 * 256 });
tree = parser.parse(file.content, undefined, { bufferSize: getTreeSitterBufferSize(file.content.length) });
} catch (parseError) {
// Skip files that can't be parsed
continue;
}
wasReparsed = true;
// Cache re-parsed tree for potential future use
astCache.set(file.path, tree);
}
@@ -75,27 +148,30 @@ export const processHeritage = async (
captureMap[c.name] = c.node;
});
// EXTENDS: Class extends another Class
// EXTENDS or IMPLEMENTS: resolve via symbol table for languages where
// the tree-sitter query can't distinguish classes from interfaces (C#, Java)
if (captureMap['heritage.class'] && captureMap['heritage.extends']) {
// Go struct embedding: skip named fields (only anonymous fields are embedded)
const extendsNode = captureMap['heritage.extends'];
const fieldDecl = extendsNode.parent;
if (fieldDecl?.type === 'field_declaration' && fieldDecl.childForFieldName('name')) {
return; // Named field, not struct embedding
}
const className = captureMap['heritage.class'].text;
const parentClassName = captureMap['heritage.extends'].text;
// Resolve both class IDs
const childId = symbolTable.lookupExact(file.path, className) ||
symbolTable.lookupFuzzy(className)[0]?.nodeId ||
generateId('Class', `${file.path}:${className}`);
const parentId = symbolTable.lookupFuzzy(parentClassName)[0]?.nodeId ||
generateId('Class', `${parentClassName}`);
const { type: relType, idPrefix } = resolveExtendsType(parentClassName, file.path, ctx, language);
const childId = resolveHeritageId(className, file.path, ctx, 'Class', `${file.path}:${className}`);
const parentId = resolveHeritageId(parentClassName, file.path, ctx, idPrefix);
if (childId && parentId && childId !== parentId) {
const relId = generateId('EXTENDS', `${childId}->${parentId}`);
graph.addRelationship({
id: relId,
id: generateId(relType, `${childId}->${parentId}`),
sourceId: childId,
targetId: parentId,
type: 'EXTENDS',
type: relType,
confidence: 1.0,
reason: '',
});
@@ -107,19 +183,12 @@ export const processHeritage = async (
const className = captureMap['heritage.class'].text;
const interfaceName = captureMap['heritage.implements'].text;
// Resolve class and interface IDs
const classId = symbolTable.lookupExact(file.path, className) ||
symbolTable.lookupFuzzy(className)[0]?.nodeId ||
generateId('Class', `${file.path}:${className}`);
const interfaceId = symbolTable.lookupFuzzy(interfaceName)[0]?.nodeId ||
generateId('Interface', `${interfaceName}`);
const classId = resolveHeritageId(className, file.path, ctx, 'Class', `${file.path}:${className}`);
const interfaceId = resolveHeritageId(interfaceName, file.path, ctx, 'Interface');
if (classId && interfaceId) {
const relId = generateId('IMPLEMENTS', `${classId}->${interfaceId}`);
graph.addRelationship({
id: relId,
id: generateId('IMPLEMENTS', `${classId}->${interfaceId}`),
sourceId: classId,
targetId: interfaceId,
type: 'IMPLEMENTS',
@@ -134,19 +203,12 @@ export const processHeritage = async (
const structName = captureMap['heritage.class'].text;
const traitName = captureMap['heritage.trait'].text;
// Resolve struct and trait IDs
const structId = symbolTable.lookupExact(file.path, structName) ||
symbolTable.lookupFuzzy(structName)[0]?.nodeId ||
generateId('Struct', `${file.path}:${structName}`);
const traitId = symbolTable.lookupFuzzy(traitName)[0]?.nodeId ||
generateId('Trait', `${traitName}`);
const structId = resolveHeritageId(structName, file.path, ctx, 'Struct', `${file.path}:${structName}`);
const traitId = resolveHeritageId(traitName, file.path, ctx, 'Trait');
if (structId && traitId) {
const relId = generateId('IMPLEMENTS', `${structId}->${traitId}`);
graph.addRelationship({
id: relId,
id: generateId('IMPLEMENTS', `${structId}->${traitId}`),
sourceId: structId,
targetId: traitId,
type: 'IMPLEMENTS',
@@ -159,6 +221,14 @@ export const processHeritage = async (
// Tree is now owned by the LRU cache — no manual delete needed
}
if (skippedByLang && skippedByLang.size > 0) {
for (const [lang, count] of skippedByLang.entries()) {
console.warn(
`[ingestion] Skipped ${count} ${lang} file(s) in heritage processing — ${lang} parser not available.`
);
}
}
};
/**
@@ -168,8 +238,8 @@ export const processHeritage = async (
export const processHeritageFromExtracted = async (
graph: KnowledgeGraph,
extractedHeritage: ExtractedHeritage[],
symbolTable: SymbolTable,
onProgress?: (current: number, total: number) => void
ctx: ResolutionContext,
onProgress?: (current: number, total: number) => void,
) => {
const total = extractedHeritage.length;
@@ -182,30 +252,26 @@ export const processHeritageFromExtracted = async (
const h = extractedHeritage[i];
if (h.kind === 'extends') {
const childId = symbolTable.lookupExact(h.filePath, h.className) ||
symbolTable.lookupFuzzy(h.className)[0]?.nodeId ||
generateId('Class', `${h.filePath}:${h.className}`);
const fileLanguage = getLanguageFromFilename(h.filePath);
if (!fileLanguage) continue;
const { type: relType, idPrefix } = resolveExtendsType(h.parentName, h.filePath, ctx, fileLanguage);
const parentId = symbolTable.lookupFuzzy(h.parentName)[0]?.nodeId ||
generateId('Class', `${h.parentName}`);
const childId = resolveHeritageId(h.className, h.filePath, ctx, 'Class', `${h.filePath}:${h.className}`);
const parentId = resolveHeritageId(h.parentName, h.filePath, ctx, idPrefix);
if (childId && parentId && childId !== parentId) {
graph.addRelationship({
id: generateId('EXTENDS', `${childId}->${parentId}`),
id: generateId(relType, `${childId}->${parentId}`),
sourceId: childId,
targetId: parentId,
type: 'EXTENDS',
type: relType,
confidence: 1.0,
reason: '',
});
}
} else if (h.kind === 'implements') {
const classId = symbolTable.lookupExact(h.filePath, h.className) ||
symbolTable.lookupFuzzy(h.className)[0]?.nodeId ||
generateId('Class', `${h.filePath}:${h.className}`);
const interfaceId = symbolTable.lookupFuzzy(h.parentName)[0]?.nodeId ||
generateId('Interface', `${h.parentName}`);
const classId = resolveHeritageId(h.className, h.filePath, ctx, 'Class', `${h.filePath}:${h.className}`);
const interfaceId = resolveHeritageId(h.parentName, h.filePath, ctx, 'Interface');
if (classId && interfaceId) {
graph.addRelationship({
@@ -217,22 +283,18 @@ export const processHeritageFromExtracted = async (
reason: '',
});
}
} else if (h.kind === 'trait-impl') {
const structId = symbolTable.lookupExact(h.filePath, h.className) ||
symbolTable.lookupFuzzy(h.className)[0]?.nodeId ||
generateId('Struct', `${h.filePath}:${h.className}`);
const traitId = symbolTable.lookupFuzzy(h.parentName)[0]?.nodeId ||
generateId('Trait', `${h.parentName}`);
} else if (h.kind === 'trait-impl' || h.kind === 'include' || h.kind === 'extend' || h.kind === 'prepend') {
const structId = resolveHeritageId(h.className, h.filePath, ctx, 'Struct', `${h.filePath}:${h.className}`);
const traitId = resolveHeritageId(h.parentName, h.filePath, ctx, 'Trait');
if (structId && traitId) {
graph.addRelationship({
id: generateId('IMPLEMENTS', `${structId}->${traitId}`),
id: generateId('IMPLEMENTS', `${structId}->${traitId}:${h.kind}`),
sourceId: structId,
targetId: traitId,
type: 'IMPLEMENTS',
confidence: 1.0,
reason: 'trait-impl',
reason: h.kind,
});
}
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,215 @@
import fs from 'fs/promises';
import path from 'path';
const isDev = process.env.NODE_ENV === 'development';
// ============================================================================
// LANGUAGE-SPECIFIC CONFIG TYPES
// ============================================================================
/** TypeScript path alias config parsed from tsconfig.json */
export interface TsconfigPaths {
/** Map of alias prefix -> target prefix (e.g., "@/" -> "src/") */
aliases: Map<string, string>;
/** Base URL for path resolution (relative to repo root) */
baseUrl: string;
}
/** Go module config parsed from go.mod */
export interface GoModuleConfig {
/** Module path (e.g., "github.com/user/repo") */
modulePath: string;
}
/** PHP Composer PSR-4 autoload config */
export interface ComposerConfig {
/** Map of namespace prefix -> directory (e.g., "App\\" -> "app/") */
psr4: Map<string, string>;
}
/** C# project config parsed from .csproj files */
export interface CSharpProjectConfig {
/** Root namespace from <RootNamespace> or assembly name (default: project directory name) */
rootNamespace: string;
/** Directory containing the .csproj file */
projectDir: string;
}
/** Swift Package Manager module config */
export interface SwiftPackageConfig {
/** Map of target name -> source directory path (e.g., "SiuperModel" -> "Package/Sources/SiuperModel") */
targets: Map<string, string>;
}
// ============================================================================
// LANGUAGE-SPECIFIC CONFIG LOADERS
// ============================================================================
/**
* Parse tsconfig.json to extract path aliases.
* Tries tsconfig.json, tsconfig.app.json, tsconfig.base.json in order.
*/
export async function loadTsconfigPaths(repoRoot: string): Promise<TsconfigPaths | null> {
const candidates = ['tsconfig.json', 'tsconfig.app.json', 'tsconfig.base.json'];
for (const filename of candidates) {
try {
const tsconfigPath = path.join(repoRoot, filename);
const raw = await fs.readFile(tsconfigPath, 'utf-8');
// Strip JSON comments (// and /* */ style) for robustness
const stripped = raw.replace(/\/\/.*$/gm, '').replace(/\/\*[\s\S]*?\*\//g, '');
const tsconfig = JSON.parse(stripped);
const compilerOptions = tsconfig.compilerOptions;
if (!compilerOptions?.paths) continue;
const baseUrl = compilerOptions.baseUrl || '.';
const aliases = new Map<string, string>();
for (const [pattern, targets] of Object.entries(compilerOptions.paths)) {
if (!Array.isArray(targets) || targets.length === 0) continue;
const target = targets[0] as string;
// Convert glob patterns: "@/*" -> "@/", "src/*" -> "src/"
const aliasPrefix = pattern.endsWith('/*') ? pattern.slice(0, -1) : pattern;
const targetPrefix = target.endsWith('/*') ? target.slice(0, -1) : target;
aliases.set(aliasPrefix, targetPrefix);
}
if (aliases.size > 0) {
if (isDev) {
console.log(`📦 Loaded ${aliases.size} path aliases from ${filename}`);
}
return { aliases, baseUrl };
}
} catch {
// File doesn't exist or isn't valid JSON - try next
}
}
return null;
}
/**
* Parse go.mod to extract module path.
*/
export async function loadGoModulePath(repoRoot: string): Promise<GoModuleConfig | null> {
try {
const goModPath = path.join(repoRoot, 'go.mod');
const content = await fs.readFile(goModPath, 'utf-8');
const match = content.match(/^module\s+(\S+)/m);
if (match) {
if (isDev) {
console.log(`📦 Loaded Go module path: ${match[1]}`);
}
return { modulePath: match[1] };
}
} catch {
// No go.mod
}
return null;
}
/** Parse composer.json to extract PSR-4 autoload mappings (including autoload-dev). */
export async function loadComposerConfig(repoRoot: string): Promise<ComposerConfig | null> {
try {
const composerPath = path.join(repoRoot, 'composer.json');
const raw = await fs.readFile(composerPath, 'utf-8');
const composer = JSON.parse(raw);
const psr4Raw = composer.autoload?.['psr-4'] ?? {};
const psr4Dev = composer['autoload-dev']?.['psr-4'] ?? {};
const merged = { ...psr4Raw, ...psr4Dev };
const psr4 = new Map<string, string>();
for (const [ns, dir] of Object.entries(merged)) {
const nsNorm = (ns as string).replace(/\\+$/, '');
const dirNorm = (dir as string).replace(/\\/g, '/').replace(/\/+$/, '');
psr4.set(nsNorm, dirNorm);
}
if (isDev) {
console.log(`📦 Loaded ${psr4.size} PSR-4 mappings from composer.json`);
}
return { psr4 };
} catch {
return null;
}
}
/**
* Parse .csproj files to extract RootNamespace.
* Scans the repo root for .csproj files and returns configs for each.
*/
export async function loadCSharpProjectConfig(repoRoot: string): Promise<CSharpProjectConfig[]> {
const configs: CSharpProjectConfig[] = [];
// BFS scan for .csproj files up to 5 levels deep, cap at 100 dirs to avoid runaway scanning
const scanQueue: { dir: string; depth: number }[] = [{ dir: repoRoot, depth: 0 }];
const maxDepth = 5;
const maxDirs = 100;
let dirsScanned = 0;
while (scanQueue.length > 0 && dirsScanned < maxDirs) {
const { dir, depth } = scanQueue.shift()!;
dirsScanned++;
try {
const entries = await fs.readdir(dir, { withFileTypes: true });
for (const entry of entries) {
if (entry.isDirectory() && depth < maxDepth) {
// Skip common non-project directories
if (entry.name === 'node_modules' || entry.name === '.git' || entry.name === 'bin' || entry.name === 'obj') continue;
scanQueue.push({ dir: path.join(dir, entry.name), depth: depth + 1 });
}
if (entry.isFile() && entry.name.endsWith('.csproj')) {
try {
const csprojPath = path.join(dir, entry.name);
const content = await fs.readFile(csprojPath, 'utf-8');
const nsMatch = content.match(/<RootNamespace>\s*([^<]+)\s*<\/RootNamespace>/);
const rootNamespace = nsMatch
? nsMatch[1].trim()
: entry.name.replace(/\.csproj$/, '');
const projectDir = path.relative(repoRoot, dir).replace(/\\/g, '/');
configs.push({ rootNamespace, projectDir });
if (isDev) {
console.log(`📦 Loaded C# project: ${entry.name} (namespace: ${rootNamespace}, dir: ${projectDir})`);
}
} catch {
// Can't read .csproj
}
}
}
} catch {
// Can't read directory
}
}
return configs;
}
export async function loadSwiftPackageConfig(repoRoot: string): Promise<SwiftPackageConfig | null> {
// Swift imports are module-name based (e.g., `import SiuperModel`)
// SPM convention: Sources/<TargetName>/ or Package/Sources/<TargetName>/
// We scan for these directories to build a target map
const targets = new Map<string, string>();
const sourceDirs = ['Sources', 'Package/Sources', 'src'];
for (const sourceDir of sourceDirs) {
try {
const fullPath = path.join(repoRoot, sourceDir);
const entries = await fs.readdir(fullPath, { withFileTypes: true });
for (const entry of entries) {
if (entry.isDirectory()) {
targets.set(entry.name, sourceDir + '/' + entry.name);
}
}
} catch {
// Directory doesn't exist
}
}
if (targets.size > 0) {
if (isDev) {
console.log(`📦 Loaded ${targets.size} Swift package targets`);
}
return { targets };
}
return null;
}
@@ -0,0 +1,465 @@
/**
* MRO (Method Resolution Order) Processor
*
* Walks the inheritance DAG (EXTENDS/IMPLEMENTS edges), collects methods from
* each ancestor via HAS_METHOD edges, detects method-name collisions across
* parents, and applies language-specific resolution rules to emit OVERRIDES edges.
*
* Language-specific rules:
* - C++: leftmost base class in declaration order wins
* - C#/Java: class method wins over interface default; multiple interface
* methods with same name are ambiguous (null resolution)
* - Python: C3 linearization determines MRO; first in linearized order wins
* - Rust: no auto-resolution — requires qualified syntax, resolvedTo = null
* - Default: single inheritance — first definition wins
*
* OVERRIDES edge direction: Class → Method (not Method → Method).
* The source is the child class that inherits conflicting methods,
* the target is the winning ancestor method node.
* Cypher: MATCH (c:Class)-[r:CodeRelation {type: 'OVERRIDES'}]->(m:Method)
*/
import { KnowledgeGraph, GraphRelationship } from '../graph/types.js';
import { generateId } from '../../lib/utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
// ---------------------------------------------------------------------------
// Public types
// ---------------------------------------------------------------------------
export interface MROEntry {
classId: string;
className: string;
language: SupportedLanguages;
mro: string[]; // linearized parent names
ambiguities: MethodAmbiguity[];
}
export interface MethodAmbiguity {
methodName: string;
definedIn: Array<{ classId: string; className: string; methodId: string }>;
resolvedTo: string | null; // winning methodId or null if truly ambiguous
reason: string;
}
export interface MROResult {
entries: MROEntry[];
overrideEdges: number;
ambiguityCount: number;
}
// ---------------------------------------------------------------------------
// Internal helpers
// ---------------------------------------------------------------------------
/** Collect EXTENDS, IMPLEMENTS, and HAS_METHOD adjacency from the graph. */
function buildAdjacency(graph: KnowledgeGraph) {
// parentMap: childId → parentIds[] (in insertion / declaration order)
const parentMap = new Map<string, string[]>();
// methodMap: classId → methodIds[]
const methodMap = new Map<string, string[]>();
// Track which edge type each parent link came from
const parentEdgeType = new Map<string, Map<string, 'EXTENDS' | 'IMPLEMENTS'>>();
graph.forEachRelationship((rel) => {
if (rel.type === 'EXTENDS' || rel.type === 'IMPLEMENTS') {
let parents = parentMap.get(rel.sourceId);
if (!parents) {
parents = [];
parentMap.set(rel.sourceId, parents);
}
parents.push(rel.targetId);
let edgeTypes = parentEdgeType.get(rel.sourceId);
if (!edgeTypes) {
edgeTypes = new Map();
parentEdgeType.set(rel.sourceId, edgeTypes);
}
edgeTypes.set(rel.targetId, rel.type);
}
if (rel.type === 'HAS_METHOD') {
let methods = methodMap.get(rel.sourceId);
if (!methods) {
methods = [];
methodMap.set(rel.sourceId, methods);
}
methods.push(rel.targetId);
}
});
return { parentMap, methodMap, parentEdgeType };
}
/**
* Gather all ancestor IDs in BFS / topological order.
* Returns the linearized list of ancestor IDs (excluding the class itself).
*/
function gatherAncestors(
classId: string,
parentMap: Map<string, string[]>,
): string[] {
const visited = new Set<string>();
const order: string[] = [];
const queue: string[] = [...(parentMap.get(classId) ?? [])];
while (queue.length > 0) {
const id = queue.shift()!;
if (visited.has(id)) continue;
visited.add(id);
order.push(id);
const grandparents = parentMap.get(id);
if (grandparents) {
for (const gp of grandparents) {
if (!visited.has(gp)) queue.push(gp);
}
}
}
return order;
}
// ---------------------------------------------------------------------------
// C3 linearization (Python MRO)
// ---------------------------------------------------------------------------
/**
* Compute C3 linearization for a class given a parentMap.
* Returns an array of ancestor IDs in C3 order (excluding the class itself),
* or null if linearization fails (inconsistent or cyclic hierarchy).
*/
function c3Linearize(
classId: string,
parentMap: Map<string, string[]>,
cache: Map<string, string[] | null>,
inProgress?: Set<string>,
): string[] | null {
if (cache.has(classId)) return cache.get(classId)!;
// Cycle detection: if we're already computing this class, the hierarchy is cyclic
const visiting = inProgress ?? new Set<string>();
if (visiting.has(classId)) {
cache.set(classId, null);
return null;
}
visiting.add(classId);
const directParents = parentMap.get(classId);
if (!directParents || directParents.length === 0) {
visiting.delete(classId);
cache.set(classId, []);
return [];
}
// Compute linearization for each parent first
const parentLinearizations: string[][] = [];
for (const pid of directParents) {
const pLin = c3Linearize(pid, parentMap, cache, visiting);
if (pLin === null) {
visiting.delete(classId);
cache.set(classId, null);
return null;
}
parentLinearizations.push([pid, ...pLin]);
}
// Add the direct parents list as the final sequence
const sequences = [...parentLinearizations, [...directParents]];
const result: string[] = [];
while (sequences.some(s => s.length > 0)) {
// Find a good head: one that doesn't appear in the tail of any other sequence
let head: string | null = null;
for (const seq of sequences) {
if (seq.length === 0) continue;
const candidate = seq[0];
const inTail = sequences.some(
other => other.length > 1 && other.indexOf(candidate, 1) !== -1
);
if (!inTail) {
head = candidate;
break;
}
}
if (head === null) {
// Inconsistent hierarchy
visiting.delete(classId);
cache.set(classId, null);
return null;
}
result.push(head);
// Remove the chosen head from all sequences
for (const seq of sequences) {
if (seq.length > 0 && seq[0] === head) {
seq.shift();
}
}
}
visiting.delete(classId);
cache.set(classId, result);
return result;
}
// ---------------------------------------------------------------------------
// Language-specific resolution
// ---------------------------------------------------------------------------
type MethodDef = { classId: string; className: string; methodId: string };
type Resolution = { resolvedTo: string | null; reason: string };
/** Resolve by MRO order — first ancestor in linearized order wins. */
function resolveByMroOrder(
methodName: string,
defs: MethodDef[],
mroOrder: string[],
reasonPrefix: string,
): Resolution {
for (const ancestorId of mroOrder) {
const match = defs.find(d => d.classId === ancestorId);
if (match) {
return {
resolvedTo: match.methodId,
reason: `${reasonPrefix}: ${match.className}::${methodName}`,
};
}
}
return { resolvedTo: defs[0].methodId, reason: `${reasonPrefix} fallback: first definition` };
}
function resolveCsharpJava(
methodName: string,
defs: MethodDef[],
parentEdgeTypes: Map<string, 'EXTENDS' | 'IMPLEMENTS'> | undefined,
): Resolution {
const classDefs: MethodDef[] = [];
const interfaceDefs: MethodDef[] = [];
for (const def of defs) {
const edgeType = parentEdgeTypes?.get(def.classId);
if (edgeType === 'IMPLEMENTS') {
interfaceDefs.push(def);
} else {
classDefs.push(def);
}
}
if (classDefs.length > 0) {
return {
resolvedTo: classDefs[0].methodId,
reason: `class method wins: ${classDefs[0].className}::${methodName}`,
};
}
if (interfaceDefs.length > 1) {
return {
resolvedTo: null,
reason: `ambiguous: ${methodName} defined in multiple interfaces: ${interfaceDefs.map(d => d.className).join(', ')}`,
};
}
if (interfaceDefs.length === 1) {
return {
resolvedTo: interfaceDefs[0].methodId,
reason: `single interface default: ${interfaceDefs[0].className}::${methodName}`,
};
}
return { resolvedTo: null, reason: 'no resolution found' };
}
// ---------------------------------------------------------------------------
// Main entry point
// ---------------------------------------------------------------------------
export function computeMRO(graph: KnowledgeGraph): MROResult {
const { parentMap, methodMap, parentEdgeType } = buildAdjacency(graph);
const c3Cache = new Map<string, string[] | null>();
const entries: MROEntry[] = [];
let overrideEdges = 0;
let ambiguityCount = 0;
// Process every class that has at least one parent
for (const [classId, directParents] of parentMap) {
if (directParents.length === 0) continue;
const classNode = graph.getNode(classId);
if (!classNode) continue;
const language = classNode.properties.language;
if (!language) continue;
const className = classNode.properties.name;
// Compute linearized MRO depending on language
let mroOrder: string[];
if (language === SupportedLanguages.Python) {
const c3Result = c3Linearize(classId, parentMap, c3Cache);
mroOrder = c3Result ?? gatherAncestors(classId, parentMap);
} else {
mroOrder = gatherAncestors(classId, parentMap);
}
// Get the parent names for the MRO entry
const mroNames: string[] = mroOrder
.map(id => graph.getNode(id)?.properties.name)
.filter((n): n is string => n !== undefined);
// Collect methods from all ancestors, grouped by method name
const methodsByName = new Map<string, MethodDef[]>();
for (const ancestorId of mroOrder) {
const ancestorNode = graph.getNode(ancestorId);
if (!ancestorNode) continue;
const methods = methodMap.get(ancestorId) ?? [];
for (const methodId of methods) {
const methodNode = graph.getNode(methodId);
if (!methodNode) continue;
// Properties don't participate in method resolution order
if (methodNode.label === 'Property') continue;
const methodName = methodNode.properties.name;
let defs = methodsByName.get(methodName);
if (!defs) {
defs = [];
methodsByName.set(methodName, defs);
}
// Avoid duplicates (same method seen via multiple paths)
if (!defs.some(d => d.methodId === methodId)) {
defs.push({
classId: ancestorId,
className: ancestorNode.properties.name,
methodId,
});
}
}
}
// Detect collisions: methods defined in 2+ different ancestors
const ambiguities: MethodAmbiguity[] = [];
// Compute transitive edge types once per class (only needed for C#/Java)
const needsEdgeTypes = language === SupportedLanguages.CSharp || language === SupportedLanguages.Java || language === SupportedLanguages.Kotlin;
const classEdgeTypes = needsEdgeTypes
? buildTransitiveEdgeTypes(classId, parentMap, parentEdgeType)
: undefined;
for (const [methodName, defs] of methodsByName) {
if (defs.length < 2) continue;
// Own method shadows inherited — no ambiguity
const ownMethods = methodMap.get(classId) ?? [];
const ownDefinesIt = ownMethods.some(mid => {
const mn = graph.getNode(mid);
return mn?.properties.name === methodName;
});
if (ownDefinesIt) continue;
let resolution: Resolution;
switch (language) {
case SupportedLanguages.CPlusPlus:
resolution = resolveByMroOrder(methodName, defs, mroOrder, 'C++ leftmost base');
break;
case SupportedLanguages.CSharp:
case SupportedLanguages.Java:
case SupportedLanguages.Kotlin:
resolution = resolveCsharpJava(methodName, defs, classEdgeTypes);
break;
case SupportedLanguages.Python:
resolution = resolveByMroOrder(methodName, defs, mroOrder, 'Python C3 MRO');
break;
case SupportedLanguages.Rust:
resolution = {
resolvedTo: null,
reason: `Rust requires qualified syntax: <Type as Trait>::${methodName}()`,
};
break;
default:
resolution = resolveByMroOrder(methodName, defs, mroOrder, 'first definition');
break;
}
const ambiguity: MethodAmbiguity = {
methodName,
definedIn: defs,
resolvedTo: resolution.resolvedTo,
reason: resolution.reason,
};
ambiguities.push(ambiguity);
if (resolution.resolvedTo === null) {
ambiguityCount++;
}
// Emit OVERRIDES edge if resolution found
if (resolution.resolvedTo !== null) {
graph.addRelationship({
id: generateId('OVERRIDES', `${classId}->${resolution.resolvedTo}`),
sourceId: classId,
targetId: resolution.resolvedTo,
type: 'OVERRIDES',
confidence: 1.0,
reason: resolution.reason,
});
overrideEdges++;
}
}
entries.push({
classId,
className,
language,
mro: mroNames,
ambiguities,
});
}
return { entries, overrideEdges, ambiguityCount };
}
/**
* Build transitive edge types for a class using BFS from the class to all ancestors.
*
* Known limitation: BFS first-reach heuristic can misclassify an interface as
* EXTENDS if it's reachable via a class chain before being seen via IMPLEMENTS.
* E.g. if BaseClass also implements IFoo, IFoo may be classified as EXTENDS.
* This affects C#/Java/Kotlin conflict resolution in rare diamond hierarchies.
*/
function buildTransitiveEdgeTypes(
classId: string,
parentMap: Map<string, string[]>,
parentEdgeType: Map<string, Map<string, 'EXTENDS' | 'IMPLEMENTS'>>,
): Map<string, 'EXTENDS' | 'IMPLEMENTS'> {
const result = new Map<string, 'EXTENDS' | 'IMPLEMENTS'>();
const directEdges = parentEdgeType.get(classId);
if (!directEdges) return result;
// BFS: propagate edge type from direct parents
const queue: Array<{ id: string; edgeType: 'EXTENDS' | 'IMPLEMENTS' }> = [];
const directParents = parentMap.get(classId) ?? [];
for (const pid of directParents) {
const et = directEdges.get(pid) ?? 'EXTENDS';
if (!result.has(pid)) {
result.set(pid, et);
queue.push({ id: pid, edgeType: et });
}
}
while (queue.length > 0) {
const { id, edgeType } = queue.shift()!;
const grandparents = parentMap.get(id) ?? [];
for (const gp of grandparents) {
if (!result.has(gp)) {
result.set(gp, edgeType);
queue.push({ id: gp, edgeType });
}
}
}
return result;
}
@@ -0,0 +1,384 @@
import { SupportedLanguages } from '../../config/supported-languages.js';
import type { SymbolTable, SymbolDefinition } from './symbol-table.js';
import type { NamedImportMap } from './import-processor.js';
/**
* Walk a named-binding re-export chain through NamedImportMap.
*
* When file A imports { User } from B, and B re-exports { User } from C,
* the NamedImportMap for A points to B, but B has no User definition.
* This function follows the chain: A→B→C until a definition is found.
*
* Returns the definitions found at the end of the chain, or null if the
* chain breaks (missing binding, circular reference, or depth exceeded).
* Max depth 5 to prevent infinite loops.
*
* @param allDefs Pre-computed `symbolTable.lookupFuzzy(name)` result — must be the
* complete unfiltered result. Passing a file-filtered subset will cause
* silent misses at depth=0 for non-aliased bindings.
*/
export function walkBindingChain(
name: string,
currentFilePath: string,
symbolTable: SymbolTable,
namedImportMap: NamedImportMap,
allDefs: SymbolDefinition[],
): SymbolDefinition[] | null {
let lookupFile = currentFilePath;
let lookupName = name;
const visited = new Set<string>();
for (let depth = 0; depth < 5; depth++) {
const bindings = namedImportMap.get(lookupFile);
if (!bindings) return null;
const binding = bindings.get(lookupName);
if (!binding) return null;
const key = `${binding.sourcePath}:${binding.exportedName}`;
if (visited.has(key)) return null; // circular
visited.add(key);
const targetName = binding.exportedName;
const resolvedDefs = targetName !== lookupName || depth > 0
? symbolTable.lookupFuzzy(targetName).filter(def => def.filePath === binding.sourcePath)
: allDefs.filter(def => def.filePath === binding.sourcePath);
if (resolvedDefs.length > 0) return resolvedDefs;
// No definition in source file → follow re-export chain
lookupFile = binding.sourcePath;
lookupName = targetName;
}
return null;
}
/**
* Extract named bindings from an import AST node.
* Returns undefined if the import is not a named import (e.g., import * or default).
*
* TS: import { User, Repo as R } from './models'
* → [{local:'User', exported:'User'}, {local:'R', exported:'Repo'}]
*
* Python: from models import User, Repo as R
* → [{local:'User', exported:'User'}, {local:'R', exported:'Repo'}]
*/
export function extractNamedBindings(
importNode: any,
language: SupportedLanguages,
): { local: string; exported: string }[] | undefined {
if (language === SupportedLanguages.TypeScript || language === SupportedLanguages.JavaScript) {
return extractTsNamedBindings(importNode);
}
if (language === SupportedLanguages.Python) {
return extractPythonNamedBindings(importNode);
}
if (language === SupportedLanguages.Kotlin) {
return extractKotlinNamedBindings(importNode);
}
if (language === SupportedLanguages.Rust) {
return extractRustNamedBindings(importNode);
}
if (language === SupportedLanguages.PHP) {
return extractPhpNamedBindings(importNode);
}
if (language === SupportedLanguages.CSharp) {
return extractCsharpNamedBindings(importNode);
}
if (language === SupportedLanguages.Java) {
return extractJavaNamedBindings(importNode);
}
return undefined;
}
export function extractTsNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// import_statement > import_clause > named_imports > import_specifier*
const importClause = findChild(importNode, 'import_clause');
if (importClause) {
const namedImports = findChild(importClause, 'named_imports');
if (!namedImports) return undefined; // default import, namespace import, or side-effect
const bindings: { local: string; exported: string }[] = [];
for (let i = 0; i < namedImports.namedChildCount; i++) {
const specifier = namedImports.namedChild(i);
if (specifier?.type !== 'import_specifier') continue;
const identifiers: string[] = [];
for (let j = 0; j < specifier.namedChildCount; j++) {
const child = specifier.namedChild(j);
if (child?.type === 'identifier') identifiers.push(child.text);
}
if (identifiers.length === 1) {
bindings.push({ local: identifiers[0], exported: identifiers[0] });
} else if (identifiers.length === 2) {
// import { Foo as Bar } → exported='Foo', local='Bar'
bindings.push({ local: identifiers[1], exported: identifiers[0] });
}
}
return bindings.length > 0 ? bindings : undefined;
}
// Re-export: export { X } from './y' → export_statement > export_clause > export_specifier
const exportClause = findChild(importNode, 'export_clause');
if (exportClause) {
const bindings: { local: string; exported: string }[] = [];
for (let i = 0; i < exportClause.namedChildCount; i++) {
const specifier = exportClause.namedChild(i);
if (specifier?.type !== 'export_specifier') continue;
const identifiers: string[] = [];
for (let j = 0; j < specifier.namedChildCount; j++) {
const child = specifier.namedChild(j);
if (child?.type === 'identifier') identifiers.push(child.text);
}
if (identifiers.length === 1) {
// export { User } from './base' → re-exports User as User
bindings.push({ local: identifiers[0], exported: identifiers[0] });
} else if (identifiers.length === 2) {
// export { Repo as Repository } from './models' → name=Repo, alias=Repository
// For re-exports, the first id is the source name, second is what's exported
// When another file imports { Repository }, they get Repo from the source
bindings.push({ local: identifiers[1], exported: identifiers[0] });
}
}
return bindings.length > 0 ? bindings : undefined;
}
return undefined;
}
export function extractPythonNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// Only from import_from_statement, not plain import_statement
if (importNode.type !== 'import_from_statement') return undefined;
const bindings: { local: string; exported: string }[] = [];
for (let i = 0; i < importNode.namedChildCount; i++) {
const child = importNode.namedChild(i);
if (!child) continue;
if (child.type === 'dotted_name') {
// Skip the module_name (first dotted_name is the source module)
const fieldName = importNode.childForFieldName?.('module_name');
if (fieldName && child.startIndex === fieldName.startIndex) continue;
// This is an imported name: from x import User
const name = child.text;
if (name) bindings.push({ local: name, exported: name });
}
if (child.type === 'aliased_import') {
// from x import Repo as R
const dottedName = findChild(child, 'dotted_name');
const aliasIdent = findChild(child, 'identifier');
if (dottedName && aliasIdent) {
bindings.push({ local: aliasIdent.text, exported: dottedName.text });
}
}
}
return bindings.length > 0 ? bindings : undefined;
}
export function extractKotlinNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// import_header > identifier + import_alias > simple_identifier
if (importNode.type !== 'import_header') return undefined;
const fullIdent = findChild(importNode, 'identifier');
if (!fullIdent) return undefined;
const fullText = fullIdent.text;
const exportedName = fullText.includes('.') ? fullText.split('.').pop()! : fullText;
const importAlias = findChild(importNode, 'import_alias');
if (importAlias) {
// Aliased: import com.example.User as U
const aliasIdent = findChild(importAlias, 'simple_identifier');
if (!aliasIdent) return undefined;
return [{ local: aliasIdent.text, exported: exportedName }];
}
// Non-aliased: import com.example.User → local="User", exported="User"
// Skip wildcard imports (ending in *)
if (fullText.endsWith('.*') || fullText.endsWith('*')) return undefined;
// Skip lowercase last segments — those are member/function imports (e.g.,
// import util.OneArg.writeAudit), not class imports. Multiple member imports
// with the same function name would collide in NamedImportMap, breaking
// arity-based disambiguation.
if (exportedName[0] && exportedName[0] === exportedName[0].toLowerCase()) return undefined;
return [{ local: exportedName, exported: exportedName }];
}
export function extractRustNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// use_declaration may contain use_as_clause at any depth
if (importNode.type !== 'use_declaration') return undefined;
const bindings: { local: string; exported: string }[] = [];
collectRustBindings(importNode, bindings);
return bindings.length > 0 ? bindings : undefined;
}
function collectRustBindings(node: any, bindings: { local: string; exported: string }[]): void {
if (node.type === 'use_as_clause') {
// First identifier = exported name, second identifier = local alias
const idents: string[] = [];
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'identifier') idents.push(child.text);
// For scoped_identifier, extract the last segment
if (child?.type === 'scoped_identifier') {
const nameNode = child.childForFieldName?.('name');
if (nameNode) idents.push(nameNode.text);
}
}
if (idents.length === 2) {
bindings.push({ local: idents[1], exported: idents[0] });
}
return;
}
// Terminal identifier in a use_list: use crate::models::{User, Repo}
if (node.type === 'identifier' && node.parent?.type === 'use_list') {
bindings.push({ local: node.text, exported: node.text });
return;
}
// Skip scoped_identifier that serves as path prefix in scoped_use_list
// e.g. use crate::models::{User, Repo} — the path node "crate::models" is not an importable symbol
if (node.type === 'scoped_identifier' && node.parent?.type === 'scoped_use_list') {
return; // path prefix — the use_list sibling handles the actual symbols
}
// Terminal scoped_identifier: use crate::models::User;
// Only extract if this is a leaf (no deeper use_list/use_as_clause/scoped_use_list)
if (node.type === 'scoped_identifier') {
let hasDeeper = false;
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === 'use_list' || child?.type === 'use_as_clause' || child?.type === 'scoped_use_list') {
hasDeeper = true;
break;
}
}
if (!hasDeeper) {
const nameNode = node.childForFieldName?.('name');
if (nameNode) {
bindings.push({ local: nameNode.text, exported: nameNode.text });
}
return;
}
}
// Recurse into children
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child) collectRustBindings(child, bindings);
}
}
export function extractPhpNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// namespace_use_declaration > namespace_use_clause* (flat)
// namespace_use_declaration > namespace_use_group > namespace_use_clause* (grouped)
if (importNode.type !== 'namespace_use_declaration') return undefined;
const bindings: { local: string; exported: string }[] = [];
// Collect all clauses — from direct children AND from namespace_use_group
const clauses: any[] = [];
for (let i = 0; i < importNode.namedChildCount; i++) {
const child = importNode.namedChild(i);
if (child?.type === 'namespace_use_clause') {
clauses.push(child);
} else if (child?.type === 'namespace_use_group') {
for (let j = 0; j < child.namedChildCount; j++) {
const groupChild = child.namedChild(j);
if (groupChild?.type === 'namespace_use_clause') clauses.push(groupChild);
}
}
}
for (const clause of clauses) {
// Flat imports: qualified_name + name (alias)
let qualifiedName: any = null;
const names: any[] = [];
for (let j = 0; j < clause.namedChildCount; j++) {
const child = clause.namedChild(j);
if (child?.type === 'qualified_name') qualifiedName = child;
else if (child?.type === 'name') names.push(child);
}
if (qualifiedName && names.length > 0) {
// Flat aliased import: use App\Models\Repo as R;
const fullText = qualifiedName.text;
const exportedName = fullText.includes('\\') ? fullText.split('\\').pop()! : fullText;
bindings.push({ local: names[0].text, exported: exportedName });
} else if (qualifiedName && names.length === 0) {
// Flat non-aliased import: use App\Models\User;
const fullText = qualifiedName.text;
const lastSegment = fullText.includes('\\') ? fullText.split('\\').pop()! : fullText;
bindings.push({ local: lastSegment, exported: lastSegment });
} else if (!qualifiedName && names.length >= 2) {
// Grouped aliased import: {Repo as R} — first name = exported, second = alias
bindings.push({ local: names[1].text, exported: names[0].text });
} else if (!qualifiedName && names.length === 1) {
// Grouped non-aliased import: {User} in use App\Models\{User, Repo as R}
bindings.push({ local: names[0].text, exported: names[0].text });
}
}
return bindings.length > 0 ? bindings : undefined;
}
export function extractCsharpNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// using_directive with identifier (alias) + qualified_name (target)
if (importNode.type !== 'using_directive') return undefined;
let aliasIdent: any = null;
let qualifiedName: any = null;
for (let i = 0; i < importNode.namedChildCount; i++) {
const child = importNode.namedChild(i);
if (child?.type === 'identifier' && !aliasIdent) aliasIdent = child;
else if (child?.type === 'qualified_name') qualifiedName = child;
}
if (!aliasIdent || !qualifiedName) return undefined;
const fullText = qualifiedName.text;
const exportedName = fullText.includes('.') ? fullText.split('.').pop()! : fullText;
return [{ local: aliasIdent.text, exported: exportedName }];
}
export function extractJavaNamedBindings(importNode: any): { local: string; exported: string }[] | undefined {
// import_declaration > scoped_identifier "com.example.models.User"
// Wildcard imports (.*) don't produce named bindings
if (importNode.type !== 'import_declaration') return undefined;
// Check for asterisk (wildcard import) — skip those
for (let i = 0; i < importNode.childCount; i++) {
const child = importNode.child(i);
if (child?.type === 'asterisk') return undefined;
}
const scopedId = findChild(importNode, 'scoped_identifier');
if (!scopedId) return undefined;
const fullText = scopedId.text;
const lastDot = fullText.lastIndexOf('.');
if (lastDot === -1) return undefined;
const className = fullText.slice(lastDot + 1);
// Skip lowercase names — those are package imports, not class imports
if (className[0] && className[0] === className[0].toLowerCase()) return undefined;
return [{ local: className, exported: className }];
}
function findChild(node: any, type: string): any {
for (let i = 0; i < node.namedChildCount; i++) {
const child = node.namedChild(i);
if (child?.type === type) return child;
}
return null;
}
+91 -189
View File
@@ -1,14 +1,17 @@
import { KnowledgeGraph, GraphNode, GraphRelationship } from '../graph/types.js';
import Parser from 'tree-sitter';
import { loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { loadParser, loadLanguage, isLanguageAvailable } from '../tree-sitter/parser-loader.js';
import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { generateId } from '../../lib/utils.js';
import { SymbolTable } from './symbol-table.js';
import { ASTCache } from './ast-cache.js';
import { findSiblingChild, getLanguageFromFilename, yieldToEventLoop } from './utils.js';
import { getLanguageFromFilename, yieldToEventLoop, DEFINITION_CAPTURE_KEYS, getDefinitionNodeFromCaptures, findEnclosingClassId, extractMethodSignature } from './utils.js';
import { isNodeExported } from './export-detection.js';
import { detectFrameworkFromAST } from './framework-detection.js';
import { typeConfigs } from './type-extractors/index.js';
import { WorkerPool } from './workers/worker-pool.js';
import type { ParseWorkerResult, ParseWorkerInput, ExtractedImport, ExtractedCall, ExtractedHeritage, ExtractedRoute } from './workers/parse-worker.js';
import type { ParseWorkerResult, ParseWorkerInput, ExtractedImport, ExtractedCall, ExtractedHeritage, ExtractedRoute, FileConstructorBindings } from './workers/parse-worker.js';
import { getTreeSitterBufferSize, TREE_SITTER_MAX_BUFFER } from './constants.js';
export type FileProgressCallback = (current: number, total: number, filePath: string) => void;
@@ -17,185 +20,12 @@ export interface WorkerExtractedData {
calls: ExtractedCall[];
heritage: ExtractedHeritage[];
routes: ExtractedRoute[];
constructorBindings: FileConstructorBindings[];
}
const DEFINITION_CAPTURE_KEYS = [
'definition.function',
'definition.class',
'definition.interface',
'definition.method',
'definition.struct',
'definition.enum',
'definition.namespace',
'definition.module',
'definition.trait',
'definition.impl',
'definition.type',
'definition.const',
'definition.static',
'definition.typedef',
'definition.macro',
'definition.union',
'definition.property',
'definition.record',
'definition.delegate',
'definition.annotation',
'definition.constructor',
'definition.template',
] as const;
const getDefinitionNodeFromCaptures = (captureMap: Record<string, any>): any | null => {
for (const key of DEFINITION_CAPTURE_KEYS) {
if (captureMap[key]) return captureMap[key];
}
return null;
};
// ============================================================================
// EXPORT DETECTION - Language-specific visibility detection
// ============================================================================
/**
* Check if a symbol (function, class, etc.) is exported/public
* Handles all 9 supported languages with explicit logic
*
* @param node - The AST node for the symbol name
* @param name - The symbol name
* @param language - The programming language
* @returns true if the symbol is exported/public
*/
export const isNodeExported = (node: any, name: string, language: string): boolean => {
let current = node;
switch (language) {
// JavaScript/TypeScript: Check for export keyword in ancestors
case 'javascript':
case 'typescript':
while (current) {
const type = current.type;
if (type === 'export_statement' ||
type === 'export_specifier' ||
type === 'lexical_declaration' && current.parent?.type === 'export_statement') {
return true;
}
// Also check if text starts with 'export '
if (current.text?.startsWith('export ')) {
return true;
}
current = current.parent;
}
return false;
// Python: Public if no leading underscore (convention)
case 'python':
return !name.startsWith('_');
// Java: Check for 'public' modifier
// In tree-sitter Java, modifiers are siblings of the name node, not parents
case 'java':
while (current) {
// Check if this node or any sibling is a 'modifiers' node containing 'public'
if (current.parent) {
const parent = current.parent;
// Check all children of the parent for modifiers
for (let i = 0; i < parent.childCount; i++) {
const child = parent.child(i);
if (child?.type === 'modifiers' && child.text?.includes('public')) {
return true;
}
}
// Also check if the parent's text starts with 'public' (fallback)
if (parent.type === 'method_declaration' || parent.type === 'constructor_declaration') {
if (parent.text?.trimStart().startsWith('public')) {
return true;
}
}
}
current = current.parent;
}
return false;
// C#: Check for 'public' modifier in ancestors
case 'csharp':
while (current) {
if (current.type === 'modifier' || current.type === 'modifiers') {
if (current.text?.includes('public')) return true;
}
current = current.parent;
}
return false;
// Go: Uppercase first letter = exported
case 'go':
if (name.length === 0) return false;
const first = name[0];
// Must be uppercase letter (not a number or symbol)
return first === first.toUpperCase() && first !== first.toLowerCase();
// Rust: Check for 'pub' visibility modifier
case 'rust':
while (current) {
if (current.type === 'visibility_modifier') {
if (current.text?.includes('pub')) return true;
}
current = current.parent;
}
return false;
// Kotlin: Default visibility is public (unlike Java)
// visibility_modifier is inside modifiers, a sibling of the name node within the declaration
case 'kotlin':
while (current) {
if (current.parent) {
const visMod = findSiblingChild(current.parent, 'modifiers', 'visibility_modifier');
if (visMod) {
const text = visMod.text;
if (text === 'private' || text === 'internal' || text === 'protected') return false;
if (text === 'public') return true;
}
}
current = current.parent;
}
// No visibility modifier = public (Kotlin default)
return true;
// C/C++: No native export concept at language level
// Entry points will be detected via name patterns (main, etc.)
case 'c':
case 'cpp':
return false;
// Swift: Check for 'public' or 'open' access modifiers
case 'swift':
while (current) {
if (current.type === 'modifiers' || current.type === 'visibility_modifier') {
const text = current.text || '';
if (text.includes('public') || text.includes('open')) return true;
}
current = current.parent;
}
return false;
// PHP: Check for visibility modifier or top-level scope
case 'php':
while (current) {
if (current.type === 'class_declaration' ||
current.type === 'interface_declaration' ||
current.type === 'trait_declaration' ||
current.type === 'enum_declaration') {
return true;
}
if (current.type === 'visibility_modifier') {
return current.text === 'public';
}
current = current.parent;
}
return true; // Top-level functions are globally accessible
default:
return false;
}
};
// isNodeExported imported from ./export-detection.js (shared module)
// Re-export for backward compatibility with any external consumers
export { isNodeExported } from './export-detection.js';
// ============================================================================
// Worker-based parallel parsing
@@ -216,7 +46,7 @@ const processParsingWithWorkers = async (
if (lang) parseableFiles.push({ path: file.path, content: file.content });
}
if (parseableFiles.length === 0) return { imports: [], calls: [], heritage: [], routes: [] };
if (parseableFiles.length === 0) return { imports: [], calls: [], heritage: [], routes: [], constructorBindings: [] };
const total = files.length;
@@ -233,6 +63,7 @@ const processParsingWithWorkers = async (
const allCalls: ExtractedCall[] = [];
const allHeritage: ExtractedHeritage[] = [];
const allRoutes: ExtractedRoute[] = [];
const allConstructorBindings: FileConstructorBindings[] = [];
for (const result of chunkResults) {
for (const node of result.nodes) {
graph.addNode({
@@ -247,18 +78,37 @@ const processParsingWithWorkers = async (
}
for (const sym of result.symbols) {
symbolTable.add(sym.filePath, sym.name, sym.nodeId, sym.type);
symbolTable.add(sym.filePath, sym.name, sym.nodeId, sym.type, {
parameterCount: sym.parameterCount,
returnType: sym.returnType,
ownerId: sym.ownerId,
});
}
allImports.push(...result.imports);
allCalls.push(...result.calls);
allHeritage.push(...result.heritage);
allRoutes.push(...result.routes);
allConstructorBindings.push(...result.constructorBindings);
}
// Merge and log skipped languages from workers
const skippedLanguages = new Map<string, number>();
for (const result of chunkResults) {
for (const [lang, count] of Object.entries(result.skippedLanguages)) {
skippedLanguages.set(lang, (skippedLanguages.get(lang) || 0) + count);
}
}
if (skippedLanguages.size > 0) {
const summary = Array.from(skippedLanguages.entries())
.map(([lang, count]) => `${lang}: ${count}`)
.join(', ');
console.warn(` Skipped unsupported languages: ${summary}`);
}
// Final progress
onFileProgress?.(total, total, 'done');
return { imports: allImports, calls: allCalls, heritage: allHeritage, routes: allRoutes };
return { imports: allImports, calls: allCalls, heritage: allHeritage, routes: allRoutes, constructorBindings: allConstructorBindings };
};
// ============================================================================
@@ -274,6 +124,7 @@ const processParsingSequential = async (
) => {
const parser = await loadParser();
const total = files.length;
const skippedLanguages = new Map<string, number>();
for (let i = 0; i < files.length; i++) {
const file = files[i];
@@ -286,18 +137,24 @@ const processParsingSequential = async (
if (!language) continue;
// Skip very large files — they can crash tree-sitter or cause OOM
if (file.content.length > 512 * 1024) continue;
// Skip unsupported languages (e.g. Swift when tree-sitter-swift not installed)
if (!isLanguageAvailable(language)) {
skippedLanguages.set(language, (skippedLanguages.get(language) || 0) + 1);
continue;
}
// Skip files larger than the max tree-sitter buffer (32 MB)
if (file.content.length > TREE_SITTER_MAX_BUFFER) continue;
try {
await loadLanguage(language, file.path);
} catch {
continue; // parser unavailable — already warned in pipeline
continue; // parser unavailable — safety net
}
let tree;
try {
tree = parser.parse(file.content, undefined, { bufferSize: 1024 * 256 });
tree = parser.parse(file.content, undefined, { bufferSize: getTreeSitterBufferSize(file.content.length) });
} catch (parseError) {
console.warn(`Skipping unparseable file: ${file.path}`);
continue;
@@ -368,13 +225,26 @@ const processParsingSequential = async (
const definitionNodeForRange = getDefinitionNodeFromCaptures(captureMap);
const startLine = definitionNodeForRange ? definitionNodeForRange.startPosition.row : (nameNode ? nameNode.startPosition.row : 0);
const nodeId = generateId(nodeLabel, `${file.path}:${nodeName}:${startLine}`);
const nodeId = generateId(nodeLabel, `${file.path}:${nodeName}`);
const definitionNode = getDefinitionNodeFromCaptures(captureMap);
const frameworkHint = definitionNode
? detectFrameworkFromAST(language, (definitionNode.text || '').slice(0, 300))
: null;
// Extract method signature for Method/Constructor nodes
const methodSig = (nodeLabel === 'Function' || nodeLabel === 'Method' || nodeLabel === 'Constructor')
? extractMethodSignature(definitionNode)
: undefined;
// Language-specific return type fallback (e.g. Ruby YARD @return [Type])
if (methodSig && !methodSig.returnType && definitionNode) {
const tc = typeConfigs[language as keyof typeof typeConfigs];
if (tc?.extractReturnType) {
methodSig.returnType = tc.extractReturnType(definitionNode);
}
}
const node: GraphNode = {
id: nodeId,
label: nodeLabel as any,
@@ -389,12 +259,25 @@ const processParsingSequential = async (
astFrameworkMultiplier: frameworkHint.entryPointMultiplier,
astFrameworkReason: frameworkHint.reason,
} : {}),
...(methodSig ? {
parameterCount: methodSig.parameterCount,
returnType: methodSig.returnType,
} : {}),
},
};
graph.addNode(node);
symbolTable.add(file.path, nodeName, nodeId, nodeLabel);
// Compute enclosing class for Method/Constructor/Property/Function — used for both ownerId and HAS_METHOD
// Function is included because Kotlin/Rust/Python capture class methods as Function nodes
const needsOwner = nodeLabel === 'Method' || nodeLabel === 'Constructor' || nodeLabel === 'Property' || nodeLabel === 'Function';
const enclosingClassId = needsOwner ? findEnclosingClassId(nameNode || definitionNodeForRange, file.path) : null;
symbolTable.add(file.path, nodeName, nodeId, nodeLabel, {
parameterCount: methodSig?.parameterCount,
returnType: methodSig?.returnType,
ownerId: enclosingClassId ?? undefined,
});
const fileId = generateId('File', file.path);
@@ -410,8 +293,27 @@ const processParsingSequential = async (
};
graph.addRelationship(relationship);
// ── HAS_METHOD: link method/constructor/property to enclosing class ──
if (enclosingClassId) {
graph.addRelationship({
id: generateId('HAS_METHOD', `${enclosingClassId}->${nodeId}`),
sourceId: enclosingClassId,
targetId: nodeId,
type: 'HAS_METHOD',
confidence: 1.0,
reason: '',
});
}
});
}
if (skippedLanguages.size > 0) {
const summary = Array.from(skippedLanguages.entries())
.map(([lang, count]) => `${lang}: ${count}`)
.join(', ');
console.warn(` Skipped unsupported languages: ${summary}`);
}
};
// ============================================================================
+236 -128
View File
@@ -1,18 +1,26 @@
import { createKnowledgeGraph } from '../graph/graph.js';
import { processStructure } from './structure-processor.js';
import { processParsing } from './parsing-processor.js';
import { processImports, processImportsFromExtracted, createImportMap, buildImportResolutionContext } from './import-processor.js';
import {
processImports,
processImportsFromExtracted,
buildImportResolutionContext
} from './import-processor.js';
import { processCalls, processCallsFromExtracted, processRoutesFromExtracted } from './call-processor.js';
import { processHeritage, processHeritageFromExtracted } from './heritage-processor.js';
import { computeMRO } from './mro-processor.js';
import { processCommunities } from './community-processor.js';
import { processProcesses } from './process-processor.js';
import { createSymbolTable } from './symbol-table.js';
import { createResolutionContext } from './resolution-context.js';
import { createASTCache } from './ast-cache.js';
import { PipelineProgress, PipelineResult } from '../../types/pipeline.js';
import { walkRepositoryPaths, readFileContents } from './filesystem-walker.js';
import { getLanguageFromFilename } from './utils.js';
import { isLanguageAvailable } from '../tree-sitter/parser-loader.js';
import { createWorkerPool, WorkerPool } from './workers/worker-pool.js';
import fs from 'node:fs';
import path from 'node:path';
import { fileURLToPath, pathToFileURL } from 'node:url';
const isDev = process.env.NODE_ENV === 'development';
@@ -25,18 +33,24 @@ const CHUNK_BYTE_BUDGET = 20 * 1024 * 1024; // 20MB
/** Max AST trees to keep in LRU cache */
const AST_CACHE_CAP = 50;
export interface PipelineOptions {
/** Skip MRO, community detection, and process extraction for faster test runs. */
skipGraphPhases?: boolean;
}
export const runPipelineFromRepo = async (
repoPath: string,
onProgress: (progress: PipelineProgress) => void
onProgress: (progress: PipelineProgress) => void,
options?: PipelineOptions,
): Promise<PipelineResult> => {
const graph = createKnowledgeGraph();
const symbolTable = createSymbolTable();
const ctx = createResolutionContext();
const symbolTable = ctx.symbols;
let astCache = createASTCache(AST_CACHE_CAP);
const importMap = createImportMap();
const cleanup = () => {
astCache.clear();
symbolTable.clear();
ctx.clear();
};
try {
@@ -108,6 +122,15 @@ export const runPipelineFromRepo = async (
const totalParseable = parseableScanned.length;
if (totalParseable === 0) {
onProgress({
phase: 'parsing',
percent: 82,
message: 'No parseable files found — skipping parsing phase',
stats: { filesProcessed: 0, totalFiles: 0, nodesCreated: graph.nodeCount },
});
}
// Build byte-budget chunks
const chunks: string[][] = [];
let currentChunk: string[] = [];
@@ -137,13 +160,29 @@ export const runPipelineFromRepo = async (
stats: { filesProcessed: 0, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
});
// Don't spawn workers for tiny repos — overhead exceeds benefit
const MIN_FILES_FOR_WORKERS = 15;
const MIN_BYTES_FOR_WORKERS = 512 * 1024;
const totalBytes = parseableScanned.reduce((s, f) => s + f.size, 0);
// Create worker pool once, reuse across chunks
let workerPool: WorkerPool | undefined;
try {
const workerUrl = new URL('./workers/parse-worker.js', import.meta.url);
workerPool = createWorkerPool(workerUrl);
} catch (err) {
// Worker pool creation failed — sequential fallback
if (totalParseable >= MIN_FILES_FOR_WORKERS || totalBytes >= MIN_BYTES_FOR_WORKERS) {
try {
let workerUrl = new URL('./workers/parse-worker.js', import.meta.url);
// When running under vitest, import.meta.url points to src/ where no .js exists.
// Fall back to the compiled dist/ worker so the pool can spawn real worker threads.
const thisDir = fileURLToPath(new URL('.', import.meta.url));
if (!fs.existsSync(fileURLToPath(workerUrl))) {
const distWorker = path.resolve(thisDir, '..', '..', '..', 'dist', 'core', 'ingestion', 'workers', 'parse-worker.js');
if (fs.existsSync(distWorker)) {
workerUrl = pathToFileURL(distWorker) as URL;
}
}
workerPool = createWorkerPool(workerUrl);
} catch (err) {
if (isDev) console.warn('Worker pool creation failed, using sequential fallback:', (err as Error).message);
}
}
let filesParsedSoFar = 0;
@@ -190,23 +229,69 @@ export const runPipelineFromRepo = async (
workerPool,
);
const chunkBasePercent = 20 + ((filesParsedSoFar / totalParseable) * 62);
if (chunkWorkerData) {
// Imports
await processImportsFromExtracted(graph, allPathObjects, chunkWorkerData.imports, importMap, undefined, repoPath, importCtx);
// Calls — resolve immediately, then free the array
if (chunkWorkerData.calls.length > 0) {
await processCallsFromExtracted(graph, chunkWorkerData.calls, symbolTable, importMap);
}
// Heritage — resolve immediately, then free
if (chunkWorkerData.heritage.length > 0) {
await processHeritageFromExtracted(graph, chunkWorkerData.heritage, symbolTable);
}
// Routes — resolve immediately (Laravel route→controller CALLS edges)
if (chunkWorkerData.routes && chunkWorkerData.routes.length > 0) {
await processRoutesFromExtracted(graph, chunkWorkerData.routes, symbolTable, importMap);
}
await processImportsFromExtracted(graph, allPathObjects, chunkWorkerData.imports, ctx, (current, total) => {
onProgress({
phase: 'parsing',
percent: Math.round(chunkBasePercent),
message: `Resolving imports (chunk ${chunkIdx + 1}/${numChunks})...`,
detail: `${current}/${total} files`,
stats: { filesProcessed: filesParsedSoFar, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
});
}, repoPath, importCtx);
// Calls + Heritage + Routes — resolve in parallel (no shared mutable state between them)
// This is safe because each writes disjoint relationship types into idempotent id-keyed Maps,
// and the single-threaded event loop prevents races between synchronous addRelationship calls.
await Promise.all([
processCallsFromExtracted(
graph,
chunkWorkerData.calls,
ctx,
(current, total) => {
onProgress({
phase: 'parsing',
percent: Math.round(chunkBasePercent),
message: `Resolving calls (chunk ${chunkIdx + 1}/${numChunks})...`,
detail: `${current}/${total} files`,
stats: { filesProcessed: filesParsedSoFar, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
});
},
chunkWorkerData.constructorBindings,
),
processHeritageFromExtracted(
graph,
chunkWorkerData.heritage,
ctx,
(current, total) => {
onProgress({
phase: 'parsing',
percent: Math.round(chunkBasePercent),
message: `Resolving heritage (chunk ${chunkIdx + 1}/${numChunks})...`,
detail: `${current}/${total} records`,
stats: { filesProcessed: filesParsedSoFar, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
});
},
),
processRoutesFromExtracted(
graph,
chunkWorkerData.routes ?? [],
ctx,
(current, total) => {
onProgress({
phase: 'parsing',
percent: Math.round(chunkBasePercent),
message: `Resolving routes (chunk ${chunkIdx + 1}/${numChunks})...`,
detail: `${current}/${total} routes`,
stats: { filesProcessed: filesParsedSoFar, totalFiles: totalParseable, nodesCreated: graph.nodeCount },
});
},
),
]);
} else {
await processImports(graph, chunkFiles, astCache, importMap, undefined, repoPath, allPaths);
await processImports(graph, chunkFiles, astCache, ctx, undefined, repoPath, allPaths);
sequentialChunkPaths.push(chunkPaths);
}
@@ -227,11 +312,22 @@ export const runPipelineFromRepo = async (
.filter(p => chunkContents.has(p))
.map(p => ({ path: p, content: chunkContents.get(p)! }));
astCache = createASTCache(chunkFiles.length);
await processCalls(graph, chunkFiles, astCache, symbolTable, importMap);
await processHeritage(graph, chunkFiles, astCache, symbolTable);
const rubyHeritage = await processCalls(graph, chunkFiles, astCache, ctx);
await processHeritage(graph, chunkFiles, astCache, ctx);
if (rubyHeritage.length > 0) {
await processHeritageFromExtracted(graph, rubyHeritage, ctx);
}
astCache.clear();
}
// Log resolution cache stats
if (isDev) {
const rcStats = ctx.getStats();
const total = rcStats.cacheHits + rcStats.cacheMisses;
const hitRate = total > 0 ? ((rcStats.cacheHits / total) * 100).toFixed(1) : '0';
console.log(`🔍 Resolution cache: ${rcStats.cacheHits} hits, ${rcStats.cacheMisses} misses (${hitRate}% hit rate)`);
}
// Free import resolution context — suffix index + resolve cache no longer needed
// (allPathObjects and importCtx hold ~94MB+ for large repos)
allPathObjects.length = 0;
@@ -239,125 +335,137 @@ export const runPipelineFromRepo = async (
(importCtx as any).suffixIndex = null;
(importCtx as any).normalizedFileList = null;
if (isDev) {
let importsCount = 0;
for (const r of graph.iterRelationships()) {
if (r.type === 'IMPORTS') importsCount++;
}
console.log(`📊 Pipeline: graph has ${importsCount} IMPORTS, ${graph.relationshipCount} total relationships`);
}
let communityResult: Awaited<ReturnType<typeof processCommunities>> | undefined;
let processResult: Awaited<ReturnType<typeof processProcesses>> | undefined;
// ── Phase 5: Communities ───────────────────────────────────────────
onProgress({
phase: 'communities',
percent: 82,
message: 'Detecting code communities...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
const communityResult = await processCommunities(graph, (message, progress) => {
const communityProgress = 82 + (progress * 0.10);
if (!options?.skipGraphPhases) {
// ── Phase 4.5: Method Resolution Order ──────────────────────────────
onProgress({
phase: 'communities',
percent: Math.round(communityProgress),
message,
phase: 'parsing',
percent: 81,
message: 'Computing method resolution order...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
});
if (isDev) {
console.log(`🏘️ Community detection: ${communityResult.stats.totalCommunities} communities found (modularity: ${communityResult.stats.modularity.toFixed(3)})`);
}
const mroResult = computeMRO(graph);
if (isDev && mroResult.entries.length > 0) {
console.log(`🔀 MRO: ${mroResult.entries.length} classes analyzed, ${mroResult.ambiguityCount} ambiguities found, ${mroResult.overrideEdges} OVERRIDES edges`);
}
communityResult.communities.forEach(comm => {
graph.addNode({
id: comm.id,
label: 'Community' as const,
properties: {
name: comm.label,
filePath: '',
heuristicLabel: comm.heuristicLabel,
cohesion: comm.cohesion,
symbolCount: comm.symbolCount,
}
// ── Phase 5: Communities ───────────────────────────────────────────
onProgress({
phase: 'communities',
percent: 82,
message: 'Detecting code communities...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
});
communityResult.memberships.forEach(membership => {
graph.addRelationship({
id: `${membership.nodeId}_member_of_${membership.communityId}`,
type: 'MEMBER_OF',
sourceId: membership.nodeId,
targetId: membership.communityId,
confidence: 1.0,
reason: 'leiden-algorithm',
});
});
// ── Phase 6: Processes ─────────────────────────────────────────────
onProgress({
phase: 'processes',
percent: 94,
message: 'Detecting execution flows...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
let symbolCount = 0;
graph.forEachNode(n => { if (n.label !== 'File') symbolCount++; });
const dynamicMaxProcesses = Math.max(20, Math.min(300, Math.round(symbolCount / 10)));
const processResult = await processProcesses(
graph,
communityResult.memberships,
(message, progress) => {
const processProgress = 94 + (progress * 0.05);
communityResult = await processCommunities(graph, (message, progress) => {
const communityProgress = 82 + (progress * 0.10);
onProgress({
phase: 'processes',
percent: Math.round(processProgress),
phase: 'communities',
percent: Math.round(communityProgress),
message,
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
},
{ maxProcesses: dynamicMaxProcesses, minSteps: 3 }
);
});
if (isDev) {
console.log(`🔄 Process detection: ${processResult.stats.totalProcesses} processes found (${processResult.stats.crossCommunityCount} cross-community)`);
if (isDev) {
console.log(`🏘️ Community detection: ${communityResult.stats.totalCommunities} communities found (modularity: ${communityResult.stats.modularity.toFixed(3)})`);
}
communityResult.communities.forEach(comm => {
graph.addNode({
id: comm.id,
label: 'Community' as const,
properties: {
name: comm.label,
filePath: '',
heuristicLabel: comm.heuristicLabel,
cohesion: comm.cohesion,
symbolCount: comm.symbolCount,
}
});
});
communityResult.memberships.forEach(membership => {
graph.addRelationship({
id: `${membership.nodeId}_member_of_${membership.communityId}`,
type: 'MEMBER_OF',
sourceId: membership.nodeId,
targetId: membership.communityId,
confidence: 1.0,
reason: 'leiden-algorithm',
});
});
// ── Phase 6: Processes ─────────────────────────────────────────────
onProgress({
phase: 'processes',
percent: 94,
message: 'Detecting execution flows...',
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
let symbolCount = 0;
graph.forEachNode(n => { if (n.label !== 'File') symbolCount++; });
const dynamicMaxProcesses = Math.max(20, Math.min(300, Math.round(symbolCount / 10)));
processResult = await processProcesses(
graph,
communityResult.memberships,
(message, progress) => {
const processProgress = 94 + (progress * 0.05);
onProgress({
phase: 'processes',
percent: Math.round(processProgress),
message,
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
},
{ maxProcesses: dynamicMaxProcesses, minSteps: 3 }
);
if (isDev) {
console.log(`🔄 Process detection: ${processResult.stats.totalProcesses} processes found (${processResult.stats.crossCommunityCount} cross-community)`);
}
processResult.processes.forEach(proc => {
graph.addNode({
id: proc.id,
label: 'Process' as const,
properties: {
name: proc.label,
filePath: '',
heuristicLabel: proc.heuristicLabel,
processType: proc.processType,
stepCount: proc.stepCount,
communities: proc.communities,
entryPointId: proc.entryPointId,
terminalId: proc.terminalId,
}
});
});
processResult.steps.forEach(step => {
graph.addRelationship({
id: `${step.nodeId}_step_${step.step}_${step.processId}`,
type: 'STEP_IN_PROCESS',
sourceId: step.nodeId,
targetId: step.processId,
confidence: 1.0,
reason: 'trace-detection',
step: step.step,
});
});
}
processResult.processes.forEach(proc => {
graph.addNode({
id: proc.id,
label: 'Process' as const,
properties: {
name: proc.label,
filePath: '',
heuristicLabel: proc.heuristicLabel,
processType: proc.processType,
stepCount: proc.stepCount,
communities: proc.communities,
entryPointId: proc.entryPointId,
terminalId: proc.terminalId,
}
});
});
processResult.steps.forEach(step => {
graph.addRelationship({
id: `${step.nodeId}_step_${step.step}_${step.processId}`,
type: 'STEP_IN_PROCESS',
sourceId: step.nodeId,
targetId: step.processId,
confidence: 1.0,
reason: 'trace-detection',
step: step.step,
});
});
onProgress({
phase: 'complete',
percent: 100,
message: `Graph complete! ${communityResult.stats.totalCommunities} communities, ${processResult.stats.totalProcesses} processes detected.`,
message: communityResult && processResult
? `Graph complete! ${communityResult.stats.totalCommunities} communities, ${processResult.stats.totalProcesses} processes detected.`
: 'Graph complete! (graph phases skipped)',
stats: {
filesProcessed: totalFiles,
totalFiles,
@@ -13,6 +13,7 @@
import { KnowledgeGraph, GraphNode, GraphRelationship, NodeLabel } from '../graph/types.js';
import { CommunityMembership } from './community-processor.js';
import { calculateEntryPointScore, isTestFile } from './entry-point-scoring.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
const isDev = process.env.NODE_ENV === 'development';
@@ -287,7 +288,7 @@ const findEntryPoints = (
// Calculate entry point score using new scoring system
const { score: baseScore, reasons } = calculateEntryPointScore(
node.properties.name,
node.properties.language || 'javascript',
node.properties.language ?? SupportedLanguages.JavaScript,
node.properties.isExported ?? false,
callers.length,
callees.length,
@@ -0,0 +1,192 @@
/**
* Resolution Context
*
* Single implementation of tiered name resolution. Replaces the duplicated
* tier-selection logic previously split between symbol-resolver.ts and
* call-processor.ts.
*
* Resolution tiers (highest confidence first):
* 1. Same file (lookupExactFull — authoritative)
* 2a-named. Named binding chain (walkBindingChain via NamedImportMap)
* 2a. Import-scoped (lookupFuzzy filtered by ImportMap)
* 2b. Package-scoped (lookupFuzzy filtered by PackageMap)
* 3. Global (all candidates — consumers must check candidate count)
*/
import type { SymbolTable, SymbolDefinition } from './symbol-table.js';
import { createSymbolTable } from './symbol-table.js';
import type { NamedImportBinding } from './import-processor.js';
import { isFileInPackageDir } from './import-processor.js';
import { walkBindingChain } from './named-binding-extraction.js';
/** Resolution tier for tracking, logging, and test assertions. */
export type ResolutionTier = 'same-file' | 'import-scoped' | 'global';
/** Tier-selected candidates with metadata. */
export interface TieredCandidates {
readonly candidates: readonly SymbolDefinition[];
readonly tier: ResolutionTier;
}
/** Confidence scores per resolution tier. */
export const TIER_CONFIDENCE: Record<ResolutionTier, number> = {
'same-file': 0.95,
'import-scoped': 0.9,
'global': 0.5,
};
// --- Map types ---
export type ImportMap = Map<string, Set<string>>;
export type PackageMap = Map<string, Set<string>>;
export type NamedImportMap = Map<string, Map<string, NamedImportBinding>>;
export interface ResolutionContext {
/**
* The only resolution API. Returns all candidates at the winning tier.
*
* Tier 3 ('global') returns ALL candidates regardless of count —
* consumers must check candidates.length and refuse ambiguous matches.
*/
resolve(name: string, fromFile: string): TieredCandidates | null;
// --- Data access (for pipeline wiring, not resolution) ---
/** Symbol table — used by parsing-processor to populate symbols. */
readonly symbols: SymbolTable;
/** Raw maps — used by import-processor to populate import data. */
readonly importMap: ImportMap;
readonly packageMap: PackageMap;
readonly namedImportMap: NamedImportMap;
// --- Per-file cache lifecycle ---
enableCache(filePath: string): void;
clearCache(): void;
// --- Operational ---
getStats(): { fileCount: number; globalSymbolCount: number; cacheHits: number; cacheMisses: number };
clear(): void;
}
export const createResolutionContext = (): ResolutionContext => {
const symbols = createSymbolTable();
const importMap: ImportMap = new Map();
const packageMap: PackageMap = new Map();
const namedImportMap: NamedImportMap = new Map();
// Per-file cache state
let cacheFile: string | null = null;
let cache: Map<string, TieredCandidates | null> | null = null;
let cacheHits = 0;
let cacheMisses = 0;
// --- Core resolution (single implementation of tier logic) ---
const resolveUncached = (name: string, fromFile: string): TieredCandidates | null => {
// Tier 1: Same file — authoritative match
const localDef = symbols.lookupExactFull(fromFile, name);
if (localDef) {
return { candidates: [localDef], tier: 'same-file' };
}
// Get all global definitions for subsequent tiers
const allDefs = symbols.lookupFuzzy(name);
// Tier 2a-named: Check named bindings BEFORE empty-allDefs early return
// because aliased imports mean lookupFuzzy('U') returns empty but we
// can resolve via the exported name.
const chainResult = walkBindingChain(name, fromFile, symbols, namedImportMap, allDefs);
if (chainResult && chainResult.length > 0) {
return { candidates: chainResult, tier: 'import-scoped' };
}
if (allDefs.length === 0) return null;
// Tier 2a: Import-scoped — definition in a file imported by fromFile
const importedFiles = importMap.get(fromFile);
if (importedFiles) {
const importedDefs = allDefs.filter(def => importedFiles.has(def.filePath));
if (importedDefs.length > 0) {
return { candidates: importedDefs, tier: 'import-scoped' };
}
}
// Tier 2b: Package-scoped — definition in a package dir imported by fromFile
const importedPackages = packageMap.get(fromFile);
if (importedPackages) {
const packageDefs = allDefs.filter(def => {
for (const dirSuffix of importedPackages) {
if (isFileInPackageDir(def.filePath, dirSuffix)) return true;
}
return false;
});
if (packageDefs.length > 0) {
return { candidates: packageDefs, tier: 'import-scoped' };
}
}
// Tier 3: Global — pass all candidates through.
// Consumers must check candidate count and refuse ambiguous matches.
return { candidates: allDefs, tier: 'global' };
};
const resolve = (name: string, fromFile: string): TieredCandidates | null => {
// Check cache (only when enabled AND fromFile matches cached file)
if (cache && cacheFile === fromFile) {
if (cache.has(name)) {
cacheHits++;
return cache.get(name)!;
}
cacheMisses++;
}
const result = resolveUncached(name, fromFile);
// Store in cache if active and file matches
if (cache && cacheFile === fromFile) {
cache.set(name, result);
}
return result;
};
// --- Cache lifecycle ---
const enableCache = (filePath: string): void => {
cacheFile = filePath;
if (!cache) cache = new Map();
else cache.clear();
};
const clearCache = (): void => {
cacheFile = null;
// Reuse the Map instance — just clear entries to reduce GC pressure at scale.
cache?.clear();
};
const getStats = () => ({
...symbols.getStats(),
cacheHits,
cacheMisses,
});
const clear = (): void => {
symbols.clear();
importMap.clear();
packageMap.clear();
namedImportMap.clear();
clearCache();
cacheHits = 0;
cacheMisses = 0;
};
return {
resolve,
symbols,
importMap,
packageMap,
namedImportMap,
enableCache,
clearCache,
getStats,
clear,
};
};
@@ -0,0 +1,128 @@
/**
* C# namespace import resolution.
* Handles using-directive resolution via .csproj root namespace stripping.
*/
import type { SuffixIndex } from './utils.js';
import { suffixResolve } from './utils.js';
/** C# project config parsed from .csproj files */
export interface CSharpProjectConfig {
/** Root namespace from <RootNamespace> or assembly name (default: project directory name) */
rootNamespace: string;
/** Directory containing the .csproj file */
projectDir: string;
}
/**
* Resolve a C# using-directive import path to matching .cs files.
* Tries single-file match first, then directory match for namespace imports.
*/
export function resolveCSharpImport(
importPath: string,
csharpConfigs: CSharpProjectConfig[],
normalizedFileList: string[],
allFileList: string[],
index?: SuffixIndex,
): string[] {
const namespacePath = importPath.replace(/\./g, '/');
const results: string[] = [];
for (const config of csharpConfigs) {
const nsPath = config.rootNamespace.replace(/\./g, '/');
let relative: string;
if (namespacePath.startsWith(nsPath + '/')) {
relative = namespacePath.slice(nsPath.length + 1);
} else if (namespacePath === nsPath) {
// The import IS the root namespace — resolve to all .cs files in project root
relative = '';
} else {
continue;
}
const dirPrefix = config.projectDir
? (relative ? config.projectDir + '/' + relative : config.projectDir)
: relative;
// 1. Try as single file: relative.cs (e.g., "Models/DlqMessage.cs")
if (relative) {
const candidate = dirPrefix + '.cs';
if (index) {
const result = index.get(candidate) || index.getInsensitive(candidate);
if (result) return [result];
}
// Also try suffix match
const suffixResult = index?.get(relative + '.cs') || index?.getInsensitive(relative + '.cs');
if (suffixResult) return [suffixResult];
}
// 2. Try as directory: all .cs files directly inside (namespace import)
if (index) {
const dirFiles = index.getFilesInDir(dirPrefix, '.cs');
for (const f of dirFiles) {
const normalized = f.replace(/\\/g, '/');
// Check it's a direct child by finding the dirPrefix and ensuring no deeper slashes
const prefixIdx = normalized.indexOf(dirPrefix + '/');
if (prefixIdx < 0) continue;
const afterDir = normalized.substring(prefixIdx + dirPrefix.length + 1);
if (!afterDir.includes('/')) {
results.push(f);
}
}
if (results.length > 0) return results;
}
// 3. Linear scan fallback for directory matching
if (results.length === 0) {
const dirTrail = dirPrefix + '/';
for (let i = 0; i < normalizedFileList.length; i++) {
const normalized = normalizedFileList[i];
if (!normalized.endsWith('.cs')) continue;
const prefixIdx = normalized.indexOf(dirTrail);
if (prefixIdx < 0) continue;
const afterDir = normalized.substring(prefixIdx + dirTrail.length);
if (!afterDir.includes('/')) {
results.push(allFileList[i]);
}
}
if (results.length > 0) return results;
}
}
// Fallback: suffix matching without namespace stripping (single file)
const pathParts = namespacePath.split('/').filter(Boolean);
const fallback = suffixResolve(pathParts, normalizedFileList, allFileList, index);
return fallback ? [fallback] : [];
}
/**
* Compute the directory suffix for a C# namespace import (for PackageMap).
* Returns a suffix like "/ProjectDir/Models/" or null if no config matches.
*/
export function resolveCSharpNamespaceDir(
importPath: string,
csharpConfigs: CSharpProjectConfig[],
): string | null {
const namespacePath = importPath.replace(/\./g, '/');
for (const config of csharpConfigs) {
const nsPath = config.rootNamespace.replace(/\./g, '/');
let relative: string;
if (namespacePath.startsWith(nsPath + '/')) {
relative = namespacePath.slice(nsPath.length + 1);
} else if (namespacePath === nsPath) {
relative = '';
} else {
continue;
}
const dirPrefix = config.projectDir
? (relative ? config.projectDir + '/' + relative : config.projectDir)
: relative;
if (!dirPrefix) continue;
return '/' + dirPrefix + '/';
}
return null;
}
@@ -0,0 +1,58 @@
/**
* Go package import resolution.
* Handles Go module path-based package imports.
*/
/** Go module config parsed from go.mod */
export interface GoModuleConfig {
/** Module path (e.g., "github.com/user/repo") */
modulePath: string;
}
/**
* Extract the package directory suffix from a Go import path.
* Returns the suffix string (e.g., "/internal/auth/") or null if invalid.
*/
export function resolveGoPackageDir(
importPath: string,
goModule: GoModuleConfig,
): string | null {
if (!importPath.startsWith(goModule.modulePath)) return null;
const relativePkg = importPath.slice(goModule.modulePath.length + 1);
if (!relativePkg) return null;
return '/' + relativePkg + '/';
}
/**
* Resolve a Go internal package import to all .go files in the package directory.
* Returns an array of file paths.
*/
export function resolveGoPackage(
importPath: string,
goModule: GoModuleConfig,
normalizedFileList: string[],
allFileList: string[],
): string[] {
if (!importPath.startsWith(goModule.modulePath)) return [];
// Strip module path to get relative package path
const relativePkg = importPath.slice(goModule.modulePath.length + 1); // e.g., "internal/auth"
if (!relativePkg) return [];
const pkgSuffix = '/' + relativePkg + '/';
const matches: string[] = [];
for (let i = 0; i < normalizedFileList.length; i++) {
// Prepend '/' so paths like "internal/auth/service.go" match suffix "/internal/auth/"
const normalized = '/' + normalizedFileList[i];
// File must be directly in the package directory (not a subdirectory)
if (normalized.includes(pkgSuffix) && normalized.endsWith('.go') && !normalized.endsWith('_test.go')) {
const afterPkg = normalized.substring(normalized.indexOf(pkgSuffix) + pkgSuffix.length);
if (!afterPkg.includes('/')) {
matches.push(allFileList[i]);
}
}
}
return matches;
}

Some files were not shown because too many files have changed in this diff Show More