Compare commits

..
Author SHA1 Message Date
Gergo Magyar e898e19714 refactor: rename FileAllScopeBindings to FileScopeBindings and update related references
- Updated the interface name from FileAllScopeBindings to FileScopeBindings to better reflect its purpose.
- Adjusted all occurrences of the renamed interface across multiple files including parsing-processor.ts, pipeline.ts, parse-worker.ts, and type-env.ts.
- Enhanced comments and documentation to clarify the narrowing of scope bindings and the rationale behind the changes.
- Improved error handling and validation in the pipeline for file-scope bindings.
- Added integration tests to ensure the correct behavior of the BindingAccumulator and its interaction with the TypeEnv flush process.
2026-04-09 19:43:59 +01:00
Gergo Magyar 448a4b22a6 fix(SM-14): address PR #743 post-fix review findings
Four items from the deep review on commit d3c25d20 — three Low
findings plus one informational note.

Low #1 — Expose `get disposed(): boolean` for API symmetry
- BindingAccumulator's `_disposed` field was set but never read or
  exposed. `_finalized` had a public getter (`get finalized()`) but
  `_disposed` did not. Added the matching `get disposed()` getter so
  debug tooling and future Phase 9 consumers can detect a disposed
  accumulator without inspecting empty state heuristically.
- JSDoc notes that disposal and finalization are orthogonal lifecycle
  dimensions — a disposed accumulator may or may not be finalized.

Low #2 — Test for Tier 0 "don't overwrite" protection
- Production enrichment loop at pipeline.ts:1104-1108 has a priority
  guard:
    if (!fileExports.has(name)) { fileExports.set(name, type); }
  preventing a worker-path binding from clobbering a higher-quality
  Tier 0 SymbolTable entry. The existing `runEnrichmentLoop` test
  helper in binding-accumulator.test.ts was missing this guard, and
  no test exercised the priority branch.
- Fixed the helper to mirror the production guard.
- Added a new test: "does not overwrite existing SymbolTable entry
  (Tier 0 priority)" — pre-populates exportedTypeMap with an
  "SymbolTableAuthoritativeType" entry, runs the enrichment loop
  against an accumulator with "WorkerInferredType" for the same name,
  asserts the authoritative type survives.

Low #3 — Move finalize() to before the enrichment loop
- Previously, finalize() was called at pipeline.ts:1715 (inside
  runPipelineFromRepo, AFTER runChunkedParseAndResolve had already
  returned). The enrichment loop at pipeline.ts:1087 (inside
  runChunkedParseAndResolve) consumed the still-mutable accumulator.
  The `finalized` state was therefore not a reliable "all reads are
  done" signal — it was a "no more writes" signal that arrived later
  than the actual last read.
- Moved finalize() to immediately before the enrichment loop at line
  1087. By that point all worker-path appends (line 934) and all
  sequential-path flushes (via processCalls at line 1051, also inside
  runChunkedParseAndResolve) have completed. Grep confirmed no further
  `bindingAccumulator.appendFile` calls exist outside runChunkedParseAndResolve.
- Lifecycle contract is now explicit:
    append phase → finalize → consume → dispose
- Replaced the old finalize() call at line 1715 with an explanatory
  comment pointing to the new seam.

Informational — parsing-processor.ts TypeEnv clarification
- parsing-processor.ts builds a FieldExtractor-only TypeEnv that is
  intentionally NOT flushed into the accumulator — the accumulator
  feed happens later in call-processor.ts via its own flush() call.
  A future reader might see `buildTypeEnv()` here and try to add a
  flush call, double-counting entries and tripping the single-use
  invariant.
- Added a multi-line comment explaining the ownership rule and
  cross-referencing PR #743 and plan 2026-04-09-005.

Verification
- `tsc --noEmit` clean
- 3110 unit tests pass (+1 new Tier 0 priority test)
- 1766 resolver integration tests pass — critically, the finalize()
  relocation did not regress any real pipeline path, proving all
  writes complete before the new finalize point
- Zero regressions

Plan: docs/plans/2026-04-09-005-fix-sm14-sequential-path-memory-regression-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/743#issuecomment-4216262583
2026-04-09 18:52:51 +01:00
Gergo Magyar d3c25d2093 fix(SM-14): close sequential-path memory regression (Codex adversarial review)
Addresses the medium-severity finding from Codex's adversarial review of
commit 803631fe: the sequential path's `typeEnv.flush()` was still
writing every scope (file + function) into the BindingAccumulator, and
the accumulator stayed alive through Phase 14 and runGraphAnalysisPhases
with no reader. On fallback runs (workers disabled or unavailable),
large repos accumulated heap for nothing.

Applies BOTH Codex remediations — narrowing AND disposal:

R1 — Narrow typeEnv.flush() to file-scope only
- The sequential path now mirrors the worker-path narrowing from
  commit 803631fe. `flush()` iterates only `env.get(FILE_SCOPE)`
  instead of the nested `for (scope, scopeMap) of env` loop, writing
  entries with `scope: ''` hardcoded. Function-scope bindings never
  reach the accumulator from either execution path until a Phase 9
  consumer lands.
- The earlier rationale for keeping sequential-path full-scope data
  ("preserve a Phase 9 prototyping sample") did not survive the
  Codex challenge — Phase 9 authors use synthetic fixtures, and a
  live-repo sample from the sequential-only path isn't representative
  of production worker-dominant runs.
- Phase 9 reversion path documented inline at the flush() seam.

R2 — Add BindingAccumulator.dispose()
- New public method clears `_allByFile`, `_fileScopeByFile`, and
  explicitly resets `_totalBindings = 0` (feasibility review caught
  that clearing the maps alone would leave the `totalBindings` getter
  reporting stale counts). Idempotent and orthogonal to finalize() —
  calling dispose() doesn't change the finalized state.
- Post-dispose contract: all read methods return empty/undefined
  state matching a never-appended accumulator. Documented in class
  JSDoc with the lifecycle sequence.
- Used before finalize(): accumulator behaves like a fresh one,
  appends still succeed.
- Used after finalize(): reads return empty but appends still throw
  the existing "finalized" error.

R3 — Wire dispose() into the pipeline
- Inserted immediately after the dev telemetry log at pipeline.ts
  line ~1723, before `runCrossFileBindingPropagation` (Phase 14) and
  `runGraphAnalysisPhases`. Sequence is:
    enrichment loop → finalize → telemetry → dispose → Phase 14
  The telemetry log captures peak state before disposal, then the
  heap footprint is released for the long tail of graph analysis.
- Verified both runCrossFileBindingPropagation and
  runGraphAnalysisPhases signatures do NOT take a bindingAccumulator
  parameter — grep confirmed the last usage is at line 1720.

Tests (+6 scenarios)
- test/unit/type-env.test.ts: existing "flushes function-scoped
  bindings into accumulator" test was rewritten as a negative
  assertion ("does NOT flush function-scoped bindings, narrowed per
  PR #743 Codex review"). Plus a new "narrows mixed file-scope and
  function-scope env to file-scope only" test that builds a
  realistic TypeScript file with both scopes and asserts only the
  file-scope entry lands in the accumulator. This is the R1 red/
  green signal — both tests were written test-first and failed
  against the pre-narrowing flush() body.
- test/unit/binding-accumulator.test.ts: new `describe('dispose', ...)`
  block with 5 scenarios — empty all read methods, idempotency, pre-
  finalize behavior, post-finalize behavior, and
  `estimateMemoryBytes() === 0` guard.

Verification
- `tsc --noEmit` clean
- 3109 unit tests pass (+6 net: +3 narrowing tests — 2 new + 1
  rewritten — plus +5 dispose scenarios − 1 pre-existing test
  replaced = net +6)
- 1766 resolver integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-005-fix-sm14-sequential-path-memory-regression-plan.md
Codex review: branch diff against main, verdict needs-attention
Previous commit: 803631fe (worker-path narrowing)
2026-04-09 18:39:37 +01:00
Gergo Magyar 803631fef7 fix(SM-14): address PR #743 BindingAccumulator review findings
Addresses the 5 findings from PR #743 review comment 4211636245.

Critical (R1) — Strip function-scope bindings from worker IPC
- parse-worker.ts previously serialized typeEnv.allScopes() over the
  worker IPC boundary on every batch, pushing ~4.9 MB of function-scope
  bindings (e.g. `handleRequest@15 → db: Database`) into the accumulator
  with zero downstream consumers. The only reader is the ExportedTypeMap
  enrichment loop in pipeline.ts, which calls fileScopeEntries() —
  the `scope = ''` subset only.
- Narrowed parse-worker.ts to use typeEnv.fileScope() and emit
  [varName, typeName] pairs. FileAllScopeBindings.bindings type narrowed
  from [string, string, string][] to [string, string][].
- pipeline.ts adapter updated to unpack the new two-element tuples and
  construct BindingEntry with scope: '' hardcoded.
- Sequential path (call-processor.ts → typeEnv.flush()) is UNCHANGED
  and still writes all scopes — preserves a working sample of higher-
  quality bindings for Phase 9 prototyping without IPC cost.
- Phase 9 reversion path documented inline on both FileAllScopeBindings
  and the pipeline adapter: change fileScope() → allScopes(), widen the
  tuple back to 3 elements, done. Field name `allScopeBindings` kept to
  keep that revert mechanically trivial.

Medium #1 (R2) — Worker-vs-sequential quality asymmetry
- BindingAccumulator class JSDoc now documents that entries are NOT
  homogeneous in resolution quality: sequential path has SymbolTable
  + importedBindings access (Tier 2 cross-file propagation); worker
  path has Tier 0 + local Tier 1 only. Phase 9 consumers that trust
  every entry equally will silently produce worse results for large
  repos (worker path dominant) than for small ones.

Medium #2 (R3) — Integration test for ExportedTypeMap enrichment
- Added a 4-scenario test suite mocking KnowledgeGraph nodes and
  running the exact enrichment loop from pipeline.ts:1082-1110 inline:
    (a) Exported Function node → enriched
    (b) Non-exported Variable → filtered out (isExported gate)
    (c) Exported Const node → enriched
    (d) No matching graph node → silently skipped via continue
- Locks in the `{Label}:{filePath}:{name}` node-ID format contract.
  If the ID format drifts for any language, this test fires.

Low #1 (R4) — Storage split for O(n_file_scope) reads
- BindingAccumulator now stores two parallel maps:
    _allByFile:       Map<string, BindingEntry[]>     — full entry list
    _fileScopeByFile: Map<string, [string, string][]> — scope='' fast path
- appendFile iterates input once, populates both maps synchronously.
  fileScopeEntries becomes O(1) map lookup + O(n_file_scope) return —
  no longer walks function-scope entries to filter.
- 4 new tests: mixed scopes correctness, only-function-scope file still
  visible via files()/fileCount, multiple appends accumulate
  consistently, 1001-entry performance guard.

Low #2 (R5) — Duplicate iteration logic
- Resolved as a side effect of R1: after the worker narrowing, the
  parse-worker loop iterates `fileScope()` (flat map) and
  typeEnv.flush() iterates `allScopes()` (nested map). Different data
  shapes — no common helper to extract.

Swift CI gap (R7)
- Acknowledged as out of scope. Not SM-14 specific — 95 Swift tests
  skipped across the broader test file.

Verification
- `tsc --noEmit` clean
- 3103 unit tests pass (+10 new scenarios in binding-accumulator.test.ts)
- 1766 resolver integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-004-fix-sm14-binding-accumulator-review-findings-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/743#issuecomment-4211636245
2026-04-09 18:05:20 +01:00
Gergo Magyar 47edb754d1 Merge remote-tracking branch 'origin/main' into sm14-binding-accumulator 2026-04-09 17:42:15 +01:00
d09078925e Extract resolveFreeCall from resolveCallTarget (SM-13) (#756)
* Initial plan

* feat(SM-13): extract resolveFreeCall from resolveCallTarget

Extract the free-function call resolution path into a dedicated
`resolveFreeCall(calledName, filePath, ctx)` function that uses
`lookupExact` + import-scoped resolution via `ctx.resolve()`.

- Free function calls (foo()) now route through `resolveFreeCall`
- Swift/Kotlin implicit constructors (User()) delegate to
  `resolveStaticCall` within `resolveFreeCall`
- `resolveCallTarget` dispatches `callForm === 'free'` early,
  removing the inline freeFormHasClassTarget logic
- S0 block simplified to only handle `callForm === 'constructor'`
- Global (Tier 3) fallthrough preserved via ctx.resolve() until Phase 5
- 9 new unit tests for resolveFreeCall
- All 163 unit tests pass, all 1199 integration resolver tests pass

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c5f2e73a-259a-438c-b5c8-286b82e3c215

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore: revert unrelated package-lock.json change

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c5f2e73a-259a-438c-b5c8-286b82e3c215

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-13): address PR #756 review findings on resolveFreeCall

Addresses all 7 findings from the PR #756 review comment.

Code (R1, finding #1)
- Replace the literal `'Class' | 'Struct' | 'Record'` check in
  `hasClassTarget` with `INSTANTIABLE_CLASS_TYPES.has(c.type)`. Converts
  an invariant that was previously comment-enforced ("keep this list
  aligned with INSTANTIABLE_CLASS_TYPES") into one enforced structurally.
  Any future extension of the set propagates here automatically. The
  narrower Swift extension dedup block below still uses literal
  `'Class' | 'Struct'` by design — Swift extensions only produce Class
  duplicates in practice, Record is deliberately excluded there, and
  the inline comment now documents that asymmetry.

Tests (+12 regression scenarios)

Finding #2 — language coverage
- Go free function (doStuff())
- Python free function (def helper(): ... helper())
- Rust free function outside any impl block
- Java statically-imported function
- JavaScript module-level function
Each exercises `_resolveCallTargetForTesting` with `callForm='free'`
and the language-specific file extension. `resolveFreeCall` has no
file-extension branching, so these guard the dispatch chain per
language without assuming extractor-specific symbol shapes.

Finding #3 — argCount threading
- 2-arg overload selected when argCount=2
- 0-arg overload selected when argCount=0

Finding #5 — Tier 3 (global) resolution
- Function globally visible but not imported. Asserts exact
  `TIER_CONFIDENCE.global === 0.5` and `reason === 'global'` to catch
  silent drift if the tier table is ever refactored.

Finding #6 — preComputedArgTypes worker path
- String overload matched via preComputedArgTypes=['String']
- Int overload matched via preComputedArgTypes=['int'] (lowercase,
  mirroring the parse-worker's inferred-literal shape; stored 'Int' is
  normalized via normalizeJvmTypeName at comparison time)

Finding #7 — Enum null-route documentation
- Enum-only free call asserts `toBeNull()` with an explanatory comment
  linking to the INSTANTIABLE_CLASS_TYPES rationale. NOT marked skipped
  — current behavior is intentional, not broken.

Finding #4 — Swift extension dedup guard
- Two same-name Class entries at different path lengths; exercises the
  full dispatch chain:
    1. filterCallableCandidates with 'free' strips Class → length 0
    2. hasClassTarget triggers resolveStaticCall
    3. Homonym ambiguity null-routes per SM-12 round-1 contract
    4. Constructor-form retry repopulates with both Classes
    5. Dedup block sorts by filePath.length → shortest path wins

Verification
- `tsc --noEmit` clean
- 3064 unit tests pass (+12)
- 1766 integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-003-fix-sm13-resolve-free-call-review-findings-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/756#issuecomment-4213879002

* refactor(SM-13): extract dedupSwiftExtensionCandidates shared helper

Follow-up to the PR #756 review fix. SM-13 duplicated the Swift
extension same-name collision dedup block between `resolveCallTarget`
and `resolveFreeCall` — two copies of identical 15-line logic with the
same heuristic (`filePath.length` sort, Class/Struct-only, `length > 1`
guard). Extract a single shared helper so the two sites cannot drift.

Changes
- New `dedupSwiftExtensionCandidates(candidates, tier)` helper defined
  alongside `tryOverloadDisambiguation`, with JSDoc documenting:
  - The Swift extension scenario it addresses
  - Why it is intentionally narrower than INSTANTIABLE_CLASS_TYPES
    (Class/Struct only, not Record — C#/Kotlin records don't exhibit
    the multi-file definition pattern, widening risks accidental
    dedup of legitimately distinct record types)
  - The return-null-on-no-match contract so callers can fall through
- `resolveCallTarget` tail dedup (was lines 1593-1610): replaced with
  a single `dedupSwiftExtensionCandidates` call
- `resolveFreeCall` tail dedup (was lines 1994-2012): same replacement
- Net line count: -32 insertions, -9 deletions in the consumer sites,
  +36 for the shared helper + JSDoc

Verification
- `tsc --noEmit` clean
- 3064 unit tests pass (including the R7 Swift dedup guard test added
  in the previous commit that exercises the full free-form retry
  chain through this helper)
- 1766 integration tests pass
- Zero regressions

Follows-up on: https://github.com/abhigyanpatwari/GitNexus/pull/756

* docs(SM-13): address PR #756 final review — comment cleanup only

Three documentation-only findings from the approval review. No
behavior change, no new tests, no code path modifications.

Finding #1 — stale line-number comment
- The comment inside `resolveFreeCall` at the `hasClassTarget` site
  referenced "lines ~1994-2008" for the Swift extension dedup block.
  Those lines were the inlined pre-SM-13 version; the block has since
  been extracted to `dedupSwiftExtensionCandidates`. Replaced the line
  reference with the helper name so future readers don't chase dead
  line numbers.

Finding #2 — fuzzy-widening asymmetry undocumented
- `resolveFreeCall` intentionally has no `widenCache` parameter and no
  D2 fuzzy-widening pass (unlike `resolveCallTarget`'s member-call
  path). Added an explicit "Asymmetry vs `resolveCallTarget`" paragraph
  to the JSDoc so a caller comparing the two signatures knows the
  skipped pass is deliberate and tied to Phase 5.

Finding #3 — constructor-form retry reasons undocumented
- `resolveStaticCall` can return null for three distinct reasons
  (empty instantiable pool, homonym ambiguity, ownerless Constructor
  nodes). The retry below it unconditionally re-filters with
  `'constructor'` form, which is correct for all three but not
  obvious. Added a structured three-case comment enumerating each
  reason and linking (a) to the SM-12 null-route contract, (b) to
  the R7 dedup test, and (c) to the currently-uncovered ownerless-
  Constructor path (noted as a future test candidate).

Verification
- `tsc --noEmit` clean
- 175 `resolveFreeCall` + `resolveStaticCall` + sibling tests pass
  (sanity check — no behavior change expected)
- No regressions

Follows-up on: https://github.com/abhigyanpatwari/GitNexus/pull/756#issuecomment-4215739052

* test(SM-13): cover ownerless-Constructor retry + PHP free function

Two low-severity test gaps from PR #756 review comment 4215739052 —
previously addressed doc-only, now have concrete test coverage.

Finding #3 low — ownerless-Constructor retry path (previously comment-only)
- The retry after resolveStaticCall returns null handles three distinct
  null-return reasons. Cases (a) and (b) were already tested (Interface/
  Trait null-route from SM-12, Swift shadowing dedup from R7). Case (c) —
  resolveStaticCall step-4 bailout when the tiered pool contains
  ownerless Constructor nodes — was only covered by a comment.
- New test: Class + ownerless Constructor in tiered pool, callForm='free'.
  Exercises the full chain:
    1. resolveStaticCall step 3 walks classCandidates via
       lookupMethodByOwner — ownerless Constructor not in methodByOwner,
       nothing found.
    2. Step 4 detects Constructor in tiered pool, bails with null.
    3. resolveFreeCall retry re-runs filterCallableCandidates with
       'constructor' form, which prefers Constructor over Class per
       CONSTRUCTOR_TARGET_TYPES ordering.
    4. Single survivor returned.
- Asserts the Constructor node (not the Class) is the resolved target.

Low — PHP free function coverage gap
- The language coverage table in the same review flagged PHP free
  functions (top-level `function helper()` outside any class) as
  uncovered. Added a test mirroring the existing Go/Python/Rust/Java/
  JS language tests — exercises the `.php` dispatch path for free
  calls. Ruby and C/C++ remain uncovered; deferred to a future round
  since those languages also have other gaps in the broader test file.

Verification
- `tsc --noEmit` clean
- 3066 unit tests pass (+2 new regression tests)
- 1766 integration tests pass
- Zero regressions

Follows-up on: https://github.com/abhigyanpatwari/GitNexus/pull/756#issuecomment-4215739052

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-09 17:41:28 +01:00
JaysonAlbertandgfwangjie 338cb01ee0 [codex] fix large repository graph loading (#732)
* fix(web): stream large graph responses

* fix(server): harden graph streaming

* fix(ci): stabilize graph loading coverage

---------

Co-authored-by: gfwangjie <gfwangjie@gf.com.cn>
2026-04-09 17:40:24 +01:00
4450a14b98 feat(SM-12): Extract resolveStaticCall from resolveCallTarget (#754)
* Initial plan

* feat(SM-12): extract resolveStaticCall from resolveCallTarget

- Add resolveStaticCall(className, methodName, currentFile, ctx, argCount?) using
  lookupClassByName + lookupMethodByOwner for O(1) constructor/static resolution
- Add S0 fast path in resolveCallTarget for constructor/free-form class calls
- Export resolveStaticCall from call-processor.ts
- Add 11 unit tests covering constructor resolution, confidence tiers,
  arity disambiguation, and resolveCallTarget delegation

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c9471ca9-57ff-4dae-956e-e7ffdc326bc4

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore: revert unrelated package-lock.json change

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c9471ca9-57ff-4dae-956e-e7ffdc326bc4

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor: shorten verbose test name per code review feedback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c9471ca9-57ff-4dae-956e-e7ffdc326bc4

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-12): address PR #754 review findings

Addresses Claude's review comments on PR #754:

Performance
- Pass pre-computed `tiered` result into `resolveStaticCall` as optional
  `tieredOverride` parameter, eliminating the duplicate `ctx.resolve(className,
  currentFile)` on every constructor call path.
- Cache `freeFormHasClassTarget` in `resolveCallTarget` so the S0 fast path
  and the free-form constructor retry share a single `.some()` scan.

Architecture
- Reconcile `CLASS_LIKE_TYPES` (call-processor) with `CLASS_TYPES`
  (symbol-table): `CLASS_LIKE_TYPES = [...CLASS_TYPES, 'Impl']`. This makes
  the relationship explicit — the call resolver's set is a strict superset
  of the heritage-index set, guaranteeing anything reachable via
  `lookupClassByName` also passes the resolver filter. Trait is now included
  (harmless: traits have no Constructor nodes, so step-3 returns undefined
  and step-5 still returns the class-like node when unique). Documented
  the Interface inclusion rationale (static methods + MRO walker).
- Collapse `resolveStaticCall`'s `methodName` parameter into `className` —
  all call sites passed identical values. Named constructors (Dart
  `User.fromJson()`) arrive as member calls and go through
  `resolveMemberCall`. Documented the reserved path for when a language
  surfaces a static-method-shaped call with a distinct member name.
- Document the known gap: `callForm === 'member'` constructor patterns
  (e.g. Python `models.User()`) are handled by the tail fallback, not S0.

Tests
- Add tiered-override test asserting `ctx.resolve` is not re-invoked when
  a pre-computed result is passed in.
- Add language-specific `_resolveCallTargetForTesting` integration tests
  for Java (`new User()`), Python (`User()`), and Kotlin (`User()`).

Verification: 3031 unit + 1766 integration tests pass, zero regressions.

* fix(SM-12): restrict resolveStaticCall fallback to instantiable kinds

Addresses the high-severity finding from the Codex adversarial review of
PR #754: `resolveStaticCall`'s step-5 "return the class itself when no
Constructor node is found" fallback reused `CLASS_LIKE_TYPES`, which —
after SM-11 and PR #754's reconciliation — now includes `Interface`,
`Trait`, and `Impl`. That is the method-dispatch set, not the
instantiable set, so constructor-shaped calls could resolve to
non-instantiable nodes and emit false `CALLS` edges.

Concrete failure: Rust same-file `impl User { ... }` alongside
`struct User { ... }` — both land at same-file tier, the Impl is not
filtered out, and the step-5 fallback produces a `CALLS` edge to the
`Impl` block instead of the `Struct`. The same widening exposed
Interface / Trait targets in Java / C# / PHP / Scala.

Fix
- Introduce `INSTANTIABLE_CLASS_TYPES = {'Class', 'Struct', 'Record'}`
  as a sibling to `CLASS_LIKE_TYPES`, documenting the contract
  explicitly and cross-referencing `CONSTRUCTOR_TARGET_TYPES`.
- Update `CLASS_LIKE_TYPES` JSDoc to clarify it is the method-dispatch
  set and add an anti-pattern warning against reusing it for
  constructor-fallback filtering.
- Tighten `resolveStaticCall` step 5: filter `classCandidates` through
  `INSTANTIABLE_CLASS_TYPES` before the `length === 1` check. This
  strips `Impl` from the Rust shadowing scenario (leaving `Struct` as
  the sole instantiable target) and null-routes Interface / Trait /
  `Impl`-alone scenarios, matching the SM-10 R3 null-route precedent.
- Step 3 (explicit Constructor lookup via `lookupMethodByOwner`) is
  intentionally unchanged — its `def.type === 'Constructor'` check is
  the correct contract, and legitimate Constructor nodes attached to
  `Impl` owners still resolve correctly.

Tests (+10 regression scenarios)
- Positive guards: Struct, Record fallback paths.
- Null-route: Interface (Java/C#/TS), PHP Trait, Rust Trait.
- Rust same-file shadowing: Struct wins over Impl.
- Rust Impl-alone: null-routes (no Struct present).
- Step-3 preservation: Constructor owned by Impl still resolves to the
  Constructor node, proving step-5 tightening doesn't leak into step 3.
- Full cascade via `_resolveCallTargetForTesting` for Interface and
  Trait — confirms no downstream path silently re-introduces the edge.

Verification
- `tsc --noEmit` clean
- 3041 unit tests pass (+10)
- 1766 integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-002-fix-sm12-constructor-fallback-instantiable-only-plan.md
Codex review job: review-mnrao7fr-nv9y0e

* fix(SM-12): address PR #754 second review round

Addresses the 9 findings from the follow-up review on PR #754.

Performance
- Align `freeFormHasClassTarget` with `INSTANTIABLE_CLASS_TYPES`: drop
  `Enum` (S0 would always return null for it — wasted lookup work) and
  add `Record` (C# records and Kotlin data classes were bypassing S0
  entirely). The trigger set and the fallback filter set now agree by
  construction, documented inline.

Documentation
- Remove stale single-line JSDoc on `CLASS_LIKE_TYPES` (line 57) that
  duplicated the full multi-line block immediately below it — tooling
  picks up the first block so the old one-liner was shadowing the
  current explanation.
- Rewrite the `resolveStaticCall` JSDoc step list to match the actual
  step boundaries in the implementation (steps 3, 4, 5 were blurred in
  the old description).
- Add inline comment on step 3 documenting the same-name lookup
  assumption (`${candidate.nodeId}\0${className}`) and the symmetric
  miss case for Python `__init__`-style constructors.
- Add inline comment on step 4 documenting that it also catches the
  ambiguous-step-3 case, and warning against removing the check
  without handling that path explicitly.
- Add inline comment on step 5 enumerating the three length outcomes
  (0 / 1 / >1) so future readers see the dominant null-route case.
- Document Ruby `User.new` as a known gap alongside Python
  `models.User()` in the S0 header comment.

Tests (+2 scenarios)
- Record free-form constructor call via `_resolveCallTargetForTesting`
  exercises the aligned `freeFormHasClassTarget` trigger end-to-end,
  closing the gap where the direct `resolveStaticCall` test passed
  but the integration path was silently bypassing S0.
- Arity threading via `_resolveCallTargetForTesting` asserts that
  `call.argCount` flows through resolveCallTarget → S0 →
  resolveStaticCall → lookupMethodByOwner, catching any future
  regression where the argCount is dropped at the S0 call site.

Verification
- `tsc --noEmit` clean
- 3043 unit tests pass (+2)
- 1766 integration tests pass
- Zero regressions

Plan: docs/plans/2026-04-09-002-fix-sm12-constructor-fallback-instantiable-only-plan.md
Review: https://github.com/abhigyanpatwari/GitNexus/pull/754#issuecomment-4213536094

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-09 12:09:07 +01:00
Kunal Hemnani 3f0b8c1a5b fix(ingestion): replace lookupExact with lookupExactAll in named-binding-processor (#755) 2026-04-09 11:57:36 +01:00
bb68cc1eb0 Extract resolveMemberCall from resolveCallTarget (SM-11) (#744)
* Initial plan

* feat(SM-11): extract resolveMemberCall from resolveCallTarget

- Create resolveMemberCall(ownerType, methodName, currentFile, ctx, heritageMap?)
  that uses owner-scoped + MRO resolution only (no fuzzy lookup)
- resolveCallTarget delegates member calls (D0 path) to resolveMemberCall
- walkMixedChain uses resolveMemberCall for owner-scoped member-call resolution
- Add 7 unit tests for resolveMemberCall covering direct, inherited, MRO,
  null cases, and confidence tier assertions
- Export resolveMemberCall for external use

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3b7889a9-5f2f-4572-8904-45084210f10d

* fix(SM-11): address PR #744 review

Blocking fixes:

- B1: Revert unrelated package-lock.json gitnexus-shared addition

- B2: Document confidence-tier semantic change on resolveMemberCall

Performance / coupling fixes:

- S1: walkMixedChain now calls resolveMethodByOwner directly (hot path) to avoid throwaway ResolveResult allocation per chain step

- S2: Thread tier from resolveMethodByOwner via { def, tier } tuple; eliminates double ctx.resolve

Alignment with semantic-model plan (Phase 3 target):

- resolveMethodByOwner now iterates ALL class-like candidates from ctx.resolve, deduplicating matches by nodeId. Absorbs D4's ownerId-filtering into the owner-scoped path.

- Handles homonym classes (two Users in different files) without falling through to D1-D4 fuzzy widening

- Shared-ancestor MRO walks automatically dedup (both homonyms walk to same base method)

- Unified direct-vs-MRO lookup under a single canWalkMRO check

Tests added:

- T1: Three D0 skip-condition tests via new _resolveCallTargetForTesting internal export (overloadHints, preComputedArgTypes, hasActiveModuleAlias)

- T2: Rust qualified-syntax null test (trait-inherited method) + direct impl control

- T3: C++ leftmost-base diamond inheritance test

- B2 lock-in: cross-file class tier assertion

- Homonym disambiguation: only-one-owns-method, both-own-method ambiguity, shared-ancestor MRO convergence

Verification:

- tsc --noEmit: clean

- vitest run test/unit/: 3014 passed

- vitest run test/integration/resolvers/: 1746 passed

* test(SM-11): address second PR #744 review round + per-language integration tests

Review fixes (https://github.com/abhigyanpatwari/GitNexus/pull/744#issuecomment-4211877593):

P1 (Performance): Replace Map allocation in resolveMethodByOwner with a firstDef+ambiguous flag pattern. Zero allocation for the common single-candidate case on the hot path — the previous Map approach allocated on every member call regardless of whether deduplication was needed.

P2 (Test gap): Strengthen the module-alias D0 skip test with a homonym fixture (two Users in different files). Previously the test passed whether or not D0 was actually bypassed; the new version proves D0 must be skipped by showing that resolveMemberCall directly returns null (ambiguous) but D1-D4 with alias narrowing picks the right one. Also fixes the underlying D2-vs-alias widening interaction: when filteredCandidates was narrowed by module-alias disambiguation, D2 no longer widens back to the full fuzzy pool (introduces aliasNarrowed boolean flag).

L1 (Language coverage): Add C# and Kotlin implements-split tests at the resolveMemberCall layer.

L2 (Maintainability): Export OverloadHints as @internal so the test can use a direct cast instead of fragile Parameters<...> type inference.

Per-language integration tests:

- rust-child-extends-parent: Direct impl method resolution via D0 (with honest documentation of the trait-method-as-Function gap that is Phase 5 / SM-16 scope)

- java-interface-default-method: User implements Validator with default method resolved via implements-split MRO

- csharp-interface-default-method: Same pattern for C# 8.0+ default interface methods

- kotlin-interface-default-method: Same pattern for Kotlin interfaces with default implementations

- python-multi-level-mro: 3-level C3 linearization (Grandparent ← Parent ← Child)

- cpp-diamond-inheritance: Classic diamond (Base ← A, B ← Derived) via leftmost-base MRO

Verification:

- tsc --noEmit: clean

- vitest run test/unit/: 3015 passed

- vitest run test/integration/resolvers/: 1763 passed (+17 new per-language tests)

* fix(SM-11): Codex adversarial review corrections + deeper D0 fixes

Addresses the three high-severity findings from the Codex adversarial review of PR #744 (https://github.com/abhigyanpatwari/GitNexus/pull/744#issuecomment-4212075120), plus four deeper fixes discovered during regression triage. All discovered issues are now addressed end-to-end rather than papered over with tail-return fallbacks.

Codex review findings:

R1 (C++ diamond): The cpp-diamond-inheritance fixture used non-virtual inheritance, which is genuinely ambiguous in real C++ (two Base subobjects). Changed A and B to use 'virtual public Base' so there's a single shared Base subobject and d.method() is an unambiguous call that the leftmost-base MRO walk correctly resolves.

R2 (C# default-interface): The csharp-interface-default-method fixture called user.Validate() via a User-typed variable, but C# does not inherit default interface methods as callable class members — the call is only valid through an interface-typed variable. Changed App.cs to 'IValidator user = new User(...)' which is the idiomatic dispatch pattern.

R3 (resolveCallTarget tail-return): When D1-D4 receiver filtering produced zero file-matched and zero owner-matched candidates for a member call, the function fell through to the permissive single-candidate tail return — silently emitting CALLS edges for methods that don't belong to the receiver. Added an explicit null-route inside the D1-D4 block that fires only when both filters yielded 0.

R4 (Rust negative assertion): Added the c.trait_only() negative integration test in rust.test.ts demonstrating that direct member calls on Rust structs do not walk trait ancestry. The test now passes because of R3 (previously fell through to the tail return).

Regression triage discoveries:

1. D0 was dead code on the sequential pipeline. The sequential path sets overloadHints for every call regardless of whether the method is overloaded, and the original D0 skip condition '!overloadHints && !preComputedArgTypes' was therefore always false. The Java/C#/C++ SM-9/SM-10 inheritance tests were passing ONLY via the tail-return fallback. Fix: narrow the skip to 'overloadHints && filteredCandidates.length > 1' — skip D0 only when there are actually multiple candidates that need overload disambiguation.

2. lookupMethodByOwner couldn't disambiguate arity-differing overloads (e.g. C++ greet() vs greet(string)). With D0 now firing on the sequential path, same-name/different-arity overloads would collapse to an arbitrary first pick. Fix: added an optional argCount parameter to lookupMethodByOwner + lookupMethodByOwnerWithMRO that filters the overload set by parameterCount/requiredParameterCount before the returnType dedup.

3. Python and Rust class methods are captured as Function nodes (not Method) with ownerId set to the class. The methodByOwner index only accepted 'Method' and 'Constructor' types, so Python class methods and Rust trait methods were invisible to D0. Fix: extended the methodByOwner indexing condition to include 'Function' when ownerId is set. This also unlocks the Rust trait-method negative assertion by ensuring the qualified-syntax MRO strategy has something to return null for.

4. D0 was being skipped when a local variable shadowed an imported module name (Python 'from models.c import C; c = C()' creates both a module alias 'c → models/c.py' AND a typed local 'c'). Fix: the D0 skip now gates on 'aliasNarrowed' (a new boolean tracking whether the alias block actually narrowed filteredCandidates) instead of 'hasActiveModuleAlias'. If the method isn't in the aliased module, the receiver is a typed local variable and D0 should run.

5. PHP trait walk missed the HasTimestamps trait because lookupClassByName did not include 'Trait' type. buildHeritageMap uses lookupClassByName to resolve parent names, so 'BaseModel use HasTimestamps' was failing to register an ancestor edge for BaseModel → HasTimestamps. Fix: added 'Trait' to CLASS_TYPES. The trait is now a valid class-like type for heritage resolution (PHP use, Rust impl Trait for Struct, Scala traits).

Test updates:

- Updated the 'no heritageMap' unit test in call-processor.test.ts to assert the correct null-route behavior instead of the old tail-return fallback.

- Added a new unit test asserting Trait inclusion in the class set.

- Updated the 'does NOT include other type-like labels' test to remove Trait from its rejection set.

Verification:

- tsc --noEmit: clean

- vitest run test/unit/: 3016 passed (+1 new Trait inclusion test)

- vitest run test/integration/resolvers/: 1764 passed (+1 new Rust negative assertion)

- Zero regressions

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergo Magyar <magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-09 09:52:12 +01:00
Roshan Warrierandtxhno d6debf3324 fix(symbol-table): index constructors in methodByOwner (#753)
Co-authored-by: txhno <198242577+txhno@users.noreply.github.com>
2026-04-09 08:26:15 +01:00
Pratyush Sharma 9ab92a97d0 fix(deps): pin tree-sitter-c override to resolve peer dep conflict (#720) (#723) 2026-04-09 06:40:17 +01:00
Murat Çelik 4fde5f241b feat: print skipped large file paths in verbose analyze output (#745) 2026-04-09 06:18:32 +01:00
evolution 9f9bbcd744 feat: support GITNEXUS_HOME env var to customize global directory (#746) 2026-04-09 06:18:17 +01:00
Cocoon-Break fd67cfd5a7 docs: fix web UI install link spacing in README (#731) 2026-04-09 06:16:14 +01:00
Pratyush Sharma 3f28f7ead5 fix(web): correct dev-mode serve command in OnboardingGuide (#725) 2026-04-09 06:15:14 +01:00
d9ba9aa998 SM-10: Add MRO fast path before D2 fuzzy widening in resolveCallTarget (#741)
* Initial plan

* Add MRO fast path before D2 fuzzy widening in resolveCallTarget

When receiverTypeName is known, try resolveMethodByOwner (owner-scoped
+ MRO lookup) before falling back to the expensive lookupFuzzy in D2.
This short-circuits cross-file member call resolution for the common
non-overloaded case.

The fast path is skipped when overload disambiguation hints are
available (overloadHints or preComputedArgTypes) to avoid picking the
wrong overload for same-return-type overloaded methods.

Passes heritageMap to resolveCallTarget from all 4 call sites:
- Language seed path (processCalls)
- Sequential path (processCalls)
- walkMixedChain fallback
- Worker path (processCallsFromExtracted)

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9e49521f-2472-47bc-96e9-be4a46b073f0

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-10): address PR #741 review

Correctness:
- Module-alias guard for D0. When call.receiverName matches an active
  entry in ctx.moduleAliasMap for the current file, D0 is now skipped
  and resolution falls through to D1-D4 which respects the
  alias-narrowed candidate pool. Prevents a homonymous class in a
  different file from being picked by ctx.resolve(receiverTypeName)
  inside resolveMethodByOwner. New unit test pins the contract.

Unit tests (call-processor.test.ts — 3 new):
- D0 hit: child.parentMethod() resolves via MRO walk when
  heritageMap is provided.
- D0 skipped: same scenario still resolves via D1-D4 when heritageMap
  is undefined (backward-compat guard).
- Module-alias guard: two files both define class User with a save()
  method; 'import auth_mod as auth' in app.py must resolve
  auth.user.save() to auth_mod.py, not user_mod.py.

Integration language coverage (+3 fixtures/tests):
- swift-child-extends-parent — first-wins, gated on swiftAvailable.
- ruby-child-extends-parent   — first-wins.
- php-child-extends-parent    — first-wins (uses ParentClass since
  'Parent' is a PHP reserved word).

* test(SM-10): address second PR #741 review round

Unit tests (call-processor.test.ts, +2 new):
- overloadHints guard: Java source with two same-return-type overloads
  method(int) and method(String), int added first so lookupMethodByOwner
  would return it. processCalls auto-generates overloadHints for Java,
  forcing D0 to be skipped. o.method("hello") must resolve to
  method(String) via literal-inferred disambiguation.
- preComputedArgTypes guard: worker-path equivalent via
  processCallsFromExtracted with ExtractedCall.argTypes=['String'].
  Same two overloads, same correctness guarantee.

Integration tests (+2 fixtures + test blocks):
- go-child-extends-parent    — struct embedding, first-wins
  (Go structs are labeled 'Struct' not 'Class' in GitNexus).
- dart-child-extends-parent  — extends, first-wins, gated on
  dartAvailable like other Dart tests.

Documentation:
- Expanded the fallthrough comment in resolveMethodByOwner to clarify
  that unknown-extension paths land on plain lookupMethodByOwner
  without an ancestor walk, and that D1-D4 still runs on D0 miss.

* test(SM-10): D0 miss with heritageMap present falls through to D1-D4

Closes the last remaining gap from PR #741 review round 3. The existing
'D0 skipped' test only covered the heritageMap=undefined case, leaving
the miss-with-heritageMap path implicitly covered by integration tests
only. This adds a focused unit test where:

- Class Obj has a method doWork findable via tiered resolution
  (import-scoped) but intentionally NOT registered in methodByOwner
  (no ownerId), so lookupMethodByOwner misses.
- heritageMap is provided but built from an empty heritage array, so
  getAncestors(class:Obj) returns []. The MRO walk yields no parents.
- lookupMethodByOwnerWithMRO therefore returns undefined → D0 miss.
- D1 resolves the receiver type; D2 widens via lookupFuzzy;
  D3 file-filter picks the single matching candidate.
- A CALLS edge must still be emitted — D0 miss must not swallow
  the call.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-08 23:24:07 +01:00
abhigyanpatwariandClaude Opus 4.6 89feea744d refactor(sm-14): address PR review feedback
- Remove redundant typeEnvBindings worker payload — allScopeBindings is a strict
  superset (file-scope entries with scope=''). Removes duplicate IPC data and
  the dead fallback branch in pipeline.ts.
- Add single-use guard to TypeEnvironment.flush() — throws on second call to
  prevent silent duplication. Update JSDoc to clarify "copy" semantics
  (env is not actually drained — TypeEnv is per-file and discarded immediately).
- Remove redundant inner guard in parse-worker.ts allScopeBindings serialization.
- Rename allAllScopeBindings -> allScopeBindingsByFile for clarity.
- Document estimateMemoryBytes pessimistic ASCII assumption (V8 uses Latin-1
  for all-ASCII strings, so actual heap cost is ~half).
- Add test for single-use flush() guard.

Addresses PR #743 review feedback.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 03:20:18 +05:30
abhigyanpatwariandClaude Opus 4.6 ec85c1e8bc style: fix prettier formatting
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 02:48:54 +05:30
abhigyanpatwariandClaude Opus 4.6 0f83912636 test(sm-14): add pipeline integration simulation for BindingAccumulator
Simulates the worker deserialization -> accumulator -> fileScopeEntries flow
to verify end-to-end correctness.

Part of #679.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 02:31:23 +05:30
abhigyanpatwariandClaude Opus 4.6 750f4336f6 feat(ingestion): wire flush() into sequential processCalls path
Pass bindingAccumulator to processCalls on the sequential code path so
TypeEnv scopes are flushed into the accumulator for Phase 9+ cross-file
type propagation, matching the worker path wired in Task 4.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 02:31:23 +05:30
abhigyanpatwariandClaude Opus 4.6 a7fdb36113 refactor(pipeline): wire BindingAccumulator into chunked parse pipeline
Replace ad-hoc workerTypeEnvBindings array with BindingAccumulator class.
Add allScopeBindings to WorkerExtractedData for multi-scope binding
collection with backward-compat fallback to old typeEnvBindings format.
Return BindingAccumulator from runChunkedParseAndResolve and finalize
before Phase 14 cross-file binding propagation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 02:30:14 +05:30
abhigyanpatwariandClaude Sonnet 4.6 2f72fcc8d1 feat(parse-worker): extend worker serialization to include all scopes
Add FileAllScopeBindings interface and allScopeBindings field to
ParseWorkerResult, serializing all TypeEnv scopes (including
function-local) for BindingAccumulator cross-file type propagation.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-09 02:30:14 +05:30
abhigyanpatwariandClaude Sonnet 4.6 56e2e7f520 feat(type-env): add flush() method to TypeEnvironment
Adds flush(filePath, accumulator) to the TypeEnvironment interface and
buildTypeEnv return object, draining all scoped bindings into a
BindingAccumulator. Adds 4 unit tests covering file-scope, function-scope,
empty env, and multi-file accumulation.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-09 02:30:13 +05:30
abhigyanpatwariandClaude Opus 4.6 3111f4dcd3 feat(sm-14): add BindingAccumulator class with unit tests
Read-append-only accumulator that collects (filePath, scope, varName) -> typeName
bindings from TypeEnv outputs across all files. Supports finalization, file-scope
filtering, iteration, and memory estimation.

Part of #679.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 02:30:13 +05:30
c19e76a4a3 feat(SM-9): Add lookupMethodByOwnerWithMRO using HeritageMap (#740)
* Initial plan

* feat(SM-9): add lookupMethodByOwnerWithMRO with HeritageMap parent chain walking

- Export c3Linearize from mro-processor.ts for reuse
- Add lookupMethodByOwnerWithMRO in call-processor.ts with MRO strategy support
- Update resolveMethodByOwner to fall back to MRO walk when HeritageMap available
- Thread heritageMap through walkMixedChain for chain resolution
- Add 10 unit tests covering all acceptance criteria

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc58249b-42f1-45a9-89fb-e3917e4d0171

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat(SM-9): add Java integration test with class Child extends Parent fixture

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc58249b-42f1-45a9-89fb-e3917e4d0171

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: address code review comments on MRO strategy documentation

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/cc58249b-42f1-45a9-89fb-e3917e4d0171

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* perf(SM-9): address PR #740 review comments

- Eliminate double direct lookup in resolveMethodByOwner: delegate
  straight to lookupMethodByOwnerWithMRO when a HeritageMap is
  available (the MRO helper already does the direct lookup before
  walking ancestors). Fallback path handles the no-HeritageMap case.
- Memoize C3 linearization per HeritageMap via a WeakMap keyed cache.
  HeritageMap is immutable after build, so C3 results are stable for
  its lifetime; WeakMap lets the cache auto-drain when the HeritageMap
  is GC'd. Null sentinel caches linearization failures so cyclic
  hierarchies are not reprocessed. Eliminates per-call buildParentMap +
  c3Linearize on Python codebases.
- ancestors variable typed as readonly to accept the cached result
  without copying.
- Add four missing MRO unit tests: Kotlin implements-split, C#
  implements-split, JavaScript first-wins (separate provider from TS),
  and C++ leftmost-base diamond (first diamond test for C++).

* fix(SM-9): CI prettier + address PR #740 follow-up review

- Fix CI prettier failure in test/integration/resolvers/java.test.ts
  (auto-formatted — was introduced in 37563a31 before my first fix
  commit but had not been caught locally).
- Pin caller on the SM-9 Java integration test (parentMethodCall.source
  === 'run') so a regression that misattributes the CALLS edge fails.
- Add two implements-split unit tests:
  * Ambiguous default from two interfaces → BFS first-wins. Pins the
    contract that lookupMethodByOwnerWithMRO returns a defined result
    (full ambiguity detection is deferred to computeMRO graph pass).
  * Class method precedence over interface default: Child extends Base
    implements IFoo where both define handle() — documents that BFS
    visits the extends edge first, matching Java's class-wins rule.
- Add @internal JSDoc on lookupMethodByOwnerWithMRO clarifying it is
  exported only for testing; resolveMethodByOwner is the proper entry
  point for callers.

* test(SM-9): per-language integration fixtures and tests for inherited method resolution

Extends the SM-9 integration coverage beyond Java with six new
child-extends-parent fixtures, one per MRO strategy:

- python-child-extends-parent       → C3 strategy
- typescript-child-extends-parent   → first-wins
- javascript-child-extends-parent   → first-wins (separate provider)
- kotlin-child-extends-parent       → implements-split
- csharp-child-extends-parent       → implements-split
- cpp-child-extends-parent          → leftmost-base

Each fixture follows the java-child-extends-parent pattern:
- Parent class with a single method
- Child class extending Parent, no override
- App class/function that instantiates Child and calls
  the parent method — exercises the full ingestion pipeline,
  HeritageMap construction, and lookupMethodByOwnerWithMRO walk.

For every fixture the matching integration test asserts:
- Parent and Child classes are detected
- Child → Parent EXTENDS edge is emitted
- The parent-method call resolves to the correct target file
- The caller is pinned (source === 'run' / 'Run') to catch
  edge misattribution regressions

Rust is intentionally omitted — its qualified-syntax strategy
returns undefined from lookupMethodByOwnerWithMRO by design, so
there is no inherited-method resolution to assert against.

All 1739 integration resolver tests pass (+18 new SM-9 tests
across 6 languages).

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-08 20:46:45 +01:00
b75e76d44a feat(SM-8): Build HeritageMap from accumulated ExtractedHeritage[] (#739)
* Initial plan

* feat(SM-8): add HeritageMap with MRO-aware parent/ancestor lookup

- New heritage-map.ts: HeritageMap interface with getParents() and getAncestors()
- buildHeritageMap() consumes ExtractedHeritage[], resolves names via lookupClassByName
- Cycle protection and bounded depth (MAX_ANCESTOR_DEPTH=32) in getAncestors
- Worker path: HeritageMap built from deferredWorkerHeritage, threaded into processCallsFromExtracted
- Sequential path: Heritage accumulated across chunks, HeritageMap built after all chunks, passed to processCalls
- 18 unit tests covering parent lookup, multi-level, diamond, cycles, missing parent, bounded depth

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c413e0a3-5d63-4ddb-8ece-02fe6ed99efd

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: rename cycle test for clarity per code review

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/c413e0a3-5d63-4ddb-8ece-02fe6ed99efd

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor(SM-8): merge implementor map into heritage map

- Add `getImplementorFiles(interfaceName)` to HeritageMap interface
- Build implementor index (interface name → file paths) alongside parent
  lookup in `buildHeritageMap`, using same `resolveExtendsType` logic
- Remove `ImplementorMap` type, `buildImplementorMap`, `mergeImplementorMaps`
  from call-processor.ts
- Update `findInterfaceDispatchTargets`, `processCalls`, and
  `processCallsFromExtracted` to use HeritageMap for both parent
  lookup and implementor dispatch
- Pipeline: single `buildHeritageMap` call replaces separate
  buildImplementorMap + buildHeritageMap for both worker and
  sequential paths
- Migrate implementor tests from call-processor.test.ts to
  heritage-map.test.ts (4 new getImplementorFiles tests)
- Update interface dispatch test to use buildHeritageMap instead
  of hand-constructed ImplementorMap

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/085dffb4-b31e-4aa5-9aa3-4314bc0010e7

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: rename implementor test for clarity per code review

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/085dffb4-b31e-4aa5-9aa3-4314bc0010e7

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(SM-8): address PR #739 review comments

- pipeline.ts: cache chunk file contents from Pass 1 to eliminate
  double-read of sequential chunks in Pass 2. Peak memory drains
  incrementally as Pass 2 processes each chunk.
- heritage-map.ts: document Rust trait-impl omission from implementor
  index and the interface-name collision limitation.
- heritage-map.test.ts: add six tests covering the extends->IMPLEMENTS
  path across C# (interfaceNamePattern), Swift (heritageDefaultEdge),
  Java (symbol-table Interface lookup), Kotlin, PHP, and the Rust
  trait-impl omission.
- pipeline.ts: comment why the heritage accumulation uses a manual
  push loop instead of spread (ref #650).

* test(SM-8): address second PR #739 review pass

- Add TypeScript implements test to getImplementorFiles (closes
  the .ts coverage gap flagged by the bot reviewer).
- Tighten deep-chain boundary assertion from toBeLessThanOrEqual(32)
  to toBe(32) so a future regression returning fewer ancestors
  fails loudly. Added an ancestors[31] === 'class:Level32' check
  to pin the upper boundary.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-04-08 19:00:08 +01:00
MyShiningand许恩宁 83b5bec293 [cli] Replace owner-filtered method lookups in type-env (#736)
* refactor(type-env): use owner method lookup

* test(type-env): cover owner lookup edge cases

* test(type-env): cover inherited overload ambiguity

---------

Co-authored-by: 许恩宁 <xuenning@qiyi.com>
2026-04-08 17:41:09 +01:00
MyShiningand许恩宁 3388ae16d7 [cli] Replace Phase P class checks with class lookup index (#734)
* refactor(call-processor): use class lookup index in phase p

* test(call-processor): cover class lookup fallback

---------

Co-authored-by: 许恩宁 <xuenning@qiyi.com>
2026-04-08 14:48:54 +01:00
MyShiningand许恩宁 d784f591b2 [cli] Replace class-type fuzzy lookups in type-env.ts (#733)
* refactor(type-env): use class lookup index for type resolution

* test(type-env): add lookupClassByName regression coverage

* test(type-env): expand class lookup regression coverage

---------

Co-authored-by: 许恩宁 <xuenning@qiyi.com>
2026-04-08 14:12:01 +01:00
Kunal Hemnani 0f43190543 feat(symbol-table): add fuzzy lookup counters (#708) 2026-04-08 07:53:30 +01:00
Roshan Warrier fe87ff8f74 fix(symbol-table): index constructors in methodByOwner (#694) 2026-04-08 06:32:02 +01:00
Deepak Chauhan be2401061e [cli] Add qualified class lookups to SymbolTable (#716) 2026-04-07 22:57:18 +01:00
Deepak Chauhan b73233d232 feat(symbol-table): add class name lookup index (#707) 2026-04-07 13:29:50 +01:00
Tushar Dhawas (Kyo) 1c8ae5eb46 refactor: extract CLASS_LIKE_TYPES constant (#693)
* refactor: extract CLASS_LIKE_TYPES constant

* chore: apply prettier formatting
2026-04-07 11:56:03 +01:00
Zander Raycraft b73928f732 scarf (#688) 2026-04-06 18:39:43 -05:00
Dmytro Semchuk 19faf3b326 fix(docs): fix codex duplicate typo in main readme file (#687) 2026-04-06 22:20:35 +01:00
Gergő Magyar cb772b9e29 feat: lookupMethodByOwner index for O(1) cross-class chain resolution (#665)
Add eagerly-populated methodByOwner index to SymbolTable, keyed by
ownerNodeId\0methodName. Used by walkMixedChain as a fast path for
resolving intermediate method calls in cross-class chains like
user.getAddress().getCity().getZipCode(), avoiding expensive fuzzy
lookups when the owner type is already known.

Handles overloaded methods: returns the first match when all overloads
share the same returnType, undefined when return types differ (ambiguous).

- Add lookupMethodByOwner to SymbolTable interface + implementation
- Add resolveMethodByOwner helper in call-processor.ts
- Add fast path in walkMixedChain before resolveCallTarget fallback
- Add Java cross-class chain fixture + 6 integration tests
- Add 148 unit tests for methodByOwner index behavior
2026-04-06 10:20:04 +01:00
ivkondandClaude Opus 4.6 10f8815639 fix(ignore): respect negation patterns in .gitnexusignore (#654)
* fix(ignore): respect negation patterns in .gitnexusignore childrenIgnored

childrenIgnored checked `ig.ignores(rel) || ig.ignores(rel + '/')` which
short-circuited on the bare path — directory-only negation patterns like
`!iOS/` were missed because `ig.ignores('iOS')` treats the path as a file.
Now only checks with trailing slash since childrenIgnored is only called
for directories. Bare-name patterns (e.g. `local`) still match per gitignore spec.

Fixes #596

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(ignore): add edge-case for bare `!dir` negation pattern

Verifies that `!iOS` (without trailing slash) also un-ignores the iOS/
directory — confirms the `ignore` package normalizes both `!dir` and
`!dir/` forms consistently when tested with a trailing-slash path.

Addresses non-blocking review suggestion on #654.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs(ignore): link ignore package docs for bare-name normalization

Adds references to the `ignore` package documentation in both the
childrenIgnored comment and the bare-negation test, explaining why
`!iOS` (without trailing slash) also re-includes the iOS/ directory.

Addresses non-blocking review suggestion on #654.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-06 08:32:47 +01:00
Abhigyan Patwari 6ead5e5986 fix(setup): prefer global gitnexus binary over npx for MCP config (#653) 2026-04-06 07:46:10 +01:00
Abhigyan Patwari 14791ded4a fix(server): return clean CORS rejection instead of 500 error (#646) 2026-04-06 07:45:19 +01:00
tantkandtian kian tan 9eeb20bb04 fix: replace Array.push(...spread) with loop to prevent stack overflow (#650)
* fix: replace Array.push(...spread) with loop to prevent stack overflow

On large codebases (78K+ C files), deferred arrays in
runChunkedParseAndResolve accumulate 100K+ entries. The spread
operator in push(...array) puts every element on the call stack
as a function argument, exceeding the maximum call stack size.

Replace all 11 occurrences of `arr.push(...other)` with
`for (const _item of other) arr.push(_item)` which uses
constant stack space regardless of array size.

Fixes #649

* fix: also replace push(...spread) in parsing-processor.ts

* style: format pipeline.ts to match prettier config

---------

Co-authored-by: tian kian tan <tan@example.com>
2026-04-05 21:54:09 +01:00
Gergő Magyar 5a7c0fdbb1 feat: same-arity overload disambiguation via type-hash suffix (#651) (#658)
* feat: same-arity overload disambiguation via type-hash suffix (#651)

Add ~type1,type2 suffix to Method/Constructor node IDs when same-arity
overloads with different parameter types exist in the same class. Also add
$const suffix for C++ const-qualified method overloads via new isConst field.

Key changes:
- typeTagForId() detects same-arity collisions and appends ~typeTag
- constTagForId() detects const/non-const collisions and appends $const
- TS/JS excluded from type-hashing (overload signatures collapse to impl body)
- Sequential findEnclosingFunction fixed: falls through on ambiguous same-class
  candidates instead of picking first; fallback path includes typeTag + constTag
- Per-call-site integration tests across Java, C#, Kotlin, C++, TypeScript
- Cross-file + chain resolution tests for all 5 languages
- C++ isConst extraction via tree-sitter type_qualifier in function_declarator

1710 integration + 18 unit tests pass.

* fix: preserve generic/template args in type-hash, perf + type safety fixes

- Add rawType field to ParameterInfo preserving full type text (vector<int>)
  while type stays simplified (vector). typeTagForId uses rawType for tags.
- Populate rawType in all 11 language method extractors
- Add buildCollisionGroups() to pre-group methods by name#arity (O(N) once
  per class instead of O(N) per method call)
- Cache method extraction in call-processor findEnclosingFunction fallback
- Fix null guards on getLanguageFromFilename in all findEnclosing paths
- Tighten SKIP_TYPE_HASH_LANGUAGES to ReadonlySet<SupportedLanguages>
- Document ID stability invariant on first overload introduction
- C++ integration tests: template overloads (vector<int> vs vector<string>),
  cross-file template + chain resolution, out-of-class method definitions

1718 integration + 20 unit tests pass.

* fix: add rawType to method-extraction unit test assertions

All 26 parameter .toEqual() assertions in method-extraction.test.ts
needed the new rawType field added to match ParameterInfo schema change.

* perf: cache tempMap/groups per class, consolidate extractFromNode

- Cache derived method map + collision groups per classNode.id in
  parsing-processor (avoids rebuild per method in same class)
- Replace per-call extractFromNode with cached class extraction +
  funcName:line lookup in call-processor fallback (avoids AST walk
  per call site)
- Remove dead clearEnclosingFunctionCache export, fix JSDoc

* test: add sequential-path integration test for same-arity overloads

Add skipWorkers option to PipelineOptions to force sequential parsing.
New test suite verifies type-hash disambiguation produces identical
results through the sequential path (parsing-processor + call-processor
findEnclosingFunction) as the worker path.
2026-04-05 21:51:55 +01:00
Gergő Magyar 0561d24efd feat: METHOD_IMPLEMENTS edges, overload disambiguation, MethodExtractor unification (#574) (#642) 2026-04-04 18:41:47 +01:00
Abhigyan Patwari 153262304c fix(mcp): unify stdout silencing to prevent embedder/pool-adapter conflicts (#645) 2026-04-04 11:56:49 +01:00
Abhigyan Patwari 16cf4c503e fix(web): replace aggressive heartbeat disconnect with graceful reconnection (#643) 2026-04-04 11:56:35 +01:00
Abhigyan Patwari 57951a197b fix(web): scope all backend calls to the active repo, not always the first (#644) 2026-04-04 11:55:34 +01:00
Gergő MagyarandClaude Opus 4.6 63fc4c795f feat: MethodExtractor configs for Python, PHP, Swift, Dart, Rust, Ruby (#624)
* feat: MethodExtractor configs for Python, PHP, Swift, Dart, Rust, Ruby with exhaustive integration tests

Add per-language MethodExtractionConfig for all remaining tree-sitter languages
(RFC #568 PR 2). Each config follows the established createMethodExtractor()
factory pattern — no new types, no parse-worker changes.

Configs:
- Python: @abstractmethod, @staticmethod/@classmethod, *args/**kwargs, type hints, _/__ visibility
- PHP: abstract/final/static keywords, PHP 8 #[] attributes, __construct/__destruct
- Swift: 5-level visibility, protocol-as-abstract, static/class methods, @ attributes
- Dart: _ convention visibility, abstract (no body), method_signature unwrapping
- Rust: pub visibility, &self receiver, trait_item + impl_item, #[] attributes
- Ruby: positional visibility via sibling-walk, singleton_method as static

Integration fixtures (18 directories) covering 3 resolution patterns:
- Method enrichment: parameterTypes, isAbstract, isFinal, annotations on graph nodes
- Overload dispatch: arity-based CALLS resolution via parameterTypes
- Abstract dispatch: abstract/concrete method distinction (Python, PHP, Rust, Swift)

Go deferred — requires factory changes for receiver-based method extraction.

Closes #571

* fix: address code review findings across 6 MethodExtractor configs

Fix all actionable items from the PR #624 deep-dive review:

Dart (critical — fixes 6 CI failures):
- isDartStatic: check children first, siblings as fallback
- isDartAbstract: handle declaration nodes for abstract methods
- extractSingleParam: detect required keyword as sibling token
- Add declaration to methodNodeTypes, mixin_declaration to typeDeclarationNodes
- Add member call query for variable assignments in tree-sitter-queries

Python:
- hasDecorator now matches dotted paths (e.g. @abc.abstractmethod)
- Fix version comment from ^0.23.6 to 0.23.4

PHP:
- Add enum_declaration to typeDeclarationNodes (PHP 8.1+)
- Add version comment for 0.23.12

Swift:
- Add isOverride using hasKeyword/hasModifier pattern

Rust:
- Fix version comment from ^0.23.2 to 0.23.1

Also: identifier fallback in generic.ts for mixin owner names,
Dart integration test label fix (Method vs Function), version
comment for tree-sitter-dart 1.0.0.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: Dart extension_declaration and Ruby module_function support

Dart:
- Add extension_declaration to typeDeclarationNodes and extension_body
  to bodyNodeTypes — extension methods are now extracted into the graph
- Add extension_declaration and mixin_declaration to CLASS_CONTAINER_TYPES
  for HAS_METHOD edge resolution

Ruby:
- module_function now maps to visibility 'private' in extractRubyVisibility
- module_function methods marked isStatic via backward-walk in isStatic
- Override semantics: private/public after module_function resets isStatic

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(go): Go MethodExtractor config with receiver-based extraction

Add Go as the 13th language with a per-language MethodExtractor config.
Go methods are top-level (not nested in struct bodies), so this adds
extractFromNode() to the MethodExtractor interface for direct method
node extraction without an enclosing class.

Config extracts:
- Name from field_identifier (methods) / identifier (functions)
- Return type including multi-return (first type from parameter_list)
- Parameters with variadic support
- Visibility via uppercase/lowercase convention
- Receiver type with pointer unwrapping (*User → User)
- isStatic for functions (no receiver)

Infrastructure:
- extractOwnerName optional hook on MethodExtractionConfig
- extractFromNode on MethodExtractor (factory auto-implements)
- Parse-worker uses extractFromNode when no enclosing class found
- method_declaration added to CLASS_CONTAINER_TYPES

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test: method enrichment integration tests for 7 languages + TS abstract class fix

Add method-enrichment integration test fixtures and test blocks for
Go, C++, Java, Kotlin, TypeScript, JavaScript, and C#. Each fixture
tests: class detection, HAS_METHOD edges, EXTENDS edges, isAbstract,
isStatic, annotations, parameterTypes, and CALLS edge resolution.

Fixes found during testing:
- Remove method_declaration from CLASS_CONTAINER_TYPES (added for Go
  but broke Java/C# HAS_METHOD edge resolution — method_declaration
  is also Java's method node type)
- Add abstract_class_declaration query to TypeScript tree-sitter
  queries (was missing, so abstract classes were invisible to pipeline)

1699 integration tests pass across 20 test files, 0 regressions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: format typeDeclarationNodes array for better readability in PHP config

* fix: Go interface methods + Rust impl-for-Struct owner resolution

Go:
- Add method_elem to methodNodeTypes so interface method signatures
  are extractable as abstract methods
- Integration test: Animal interface detected, Speak isAbstract,
  CALLS edges from app.go

Rust:
- Add extractOwnerName to resolve impl Trait for Struct to the
  concrete Struct (not the Trait) — fixes method misattribution
- Fix findEnclosingClassId to generate Struct: label (not Impl:)
  for impl blocks so HAS_METHOD edges resolve to struct nodes
- Tighten abstract-dispatch test: assert SqlRepo owns find/save

generic.ts:
- Fix extractOwnerName fallback: when hook returns a value, skip
  both name-field and type_identifier scan (was overwriting result)

1703 integration tests pass, 0 regressions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: code review response — Rust impl label, Swift params, Dart async, sequential methodExtractor

Address code review findings from PR #624:

- ast-helpers: Rust `impl Trait for Struct` uses Struct label (matches existing
  graph node), plain `impl Struct` uses Impl label (matches definition.impl)
- swift: fix parameter type extraction (user_type not type_annotation), detect
  default values as function_declaration siblings, add version comment
- dart: isDartAsync now detects async*/sync* generators, add clarifying comment
  for declaration nodes in extension bodies
- python: correct isFinal comment (PEP 591 @typing.final exists, just not modeled)
- parsing-processor: port methodExtractor enrichment to sequential path so
  isAbstract/isStatic/visibility/annotations/isFinal populate on <15-file repos
- tests: remove silent `if (prop !== undefined)` guards, assert properties
  directly, fix label queries (Dart Method vs Function, Swift Method for protocol
  methods), add Rust HAS_METHOD sourceLabel tests, Swift parameterTypes tests,
  and Dart async/sync* integration tests with fixture

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: Rust grammar gap + qualified method IDs to resolve same-file collisions

Phase 1 — Rust grammar:
- Add function_signature_item query to RUST_QUERIES so abstract trait methods
  (fn speak(&self) -> String;) become graph nodes with isAbstract=true

Phase 2 — Qualified method IDs:
- findEnclosingClassInfo returns {classId, className} for AST-based class lookup
- Both parsing paths (sequential + worker) qualify method/property IDs with
  enclosing class: Method:file:ClassName.method instead of Method:file:method
- extractFuncNameFromSourceId handles ClassName.method format
- Fixes silent data loss when same-name methods in different classes shared a
  file (e.g., Animal.speak and Dog.speak both now exist as distinct graph nodes)

Test updates:
- Rust: abstract+concrete trait methods both verified, function count adjusted
- Python: static method disambiguation now emits 2 CALLS edges (correct — no
  more ID collision masking the second call)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: owner-aware resolution for qualified method IDs

Address Codex adversarial review findings after qualified ID change:

- findEnclosingFunction: disambiguate candidates by ownerId when multiple
  same-name methods exist in file; qualify fallback-generated IDs
- findEnclosingFunctionId (worker): qualify sourceIds with enclosing class
  name so CALLS source attribution matches definition-phase node IDs
- buildExportedTypeMapFromGraph: use lookupExactAll + nodeId match instead
  of lookupExactFull which returns first definition for bare name

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: methodExtractor variadic arity, return type preservation, PHP abstract dispatch

Three bugs in the methodExtractor enrichment path broke 17 integration tests:

1. Variadic parameterCount: buildMethodProps and parse-worker set
   parameterCount = info.parameters.length even for variadic functions,
   causing arity filtering to reject valid calls. Now checks isVariadic
   and sets parameterCount = undefined (matching extractMethodSignature).

2. C++ bare `...` token: extractCppParameters only iterated named
   children, missing the unnamed `...` token in C-style variadics like
   log_entry(const char* fmt, ...). Added fallback scan of all children.

3. Return type stripping: All 11 language extractReturnType functions
   used extractSimpleTypeName() which strips generic parameters
   (List<User> → "List", Task<User> → "Task"). Changed to .text?.trim()
   to preserve full generic types needed for for-loop iterable resolution,
   async-await binding, and return-type inference.

Also fixes PHP abstract dispatch test that matched SqlRepository instead
of the interface due to ambiguous filePath.includes('Repository') filter,
and adds parent-walk fallback in PHP isAbstract for extractFromNode path.

* chore: remove plan and review artifacts from PR

* fix: address Round 4 review findings + infrastructure improvements

- Ruby: add singleton_class support for class << self methods (4 new tests)
- PHP: add enum_declaration to CLASS_CONTAINER_TYPES
- Dart: add mixin/extension labels to CONTAINER_TYPE_TO_LABEL
- Swift: add TODO for unverifiable struct/enum node types on Node 22
- C#: add grammar version comment (0.23.1)
- Ruby: fix version comment range to pin (0.23.1)
- Rust/ast-helpers: add cross-reference comments for impl_item duplication
- ast-helpers: document CLASS_CONTAINER_TYPES ↔ typeDeclarationNodes invariant
- generic.ts: replace Array.includes with Set for O(1) dedup in addNestedBodies
- Go/Python/Ruby: align isAbstract signature with 2-param interface contract
- CLAUDE.md: fix malformed backtick around gitnexus:start HTML comment
- parsing-processor: add per-class method extraction cache (eliminates O(N*M))
- ast-helpers: add scoped_type_identifier to impl_item resolution
- call-processor: add dev-mode warnings at silent candidates[0] fallbacks
- MCP context(): surface methodMetadata for Method/Function/Constructor nodes
- resources.ts: update schema to list all stored Method properties

* fix: singleton_class HAS_METHOD edge regression in findEnclosingClassInfo

singleton_class (class << self) was added to CLASS_CONTAINER_TYPES but
has no name field — its receiver `self` has node type 'self', not
'identifier'. findEnclosingClassInfo now walks up to the enclosing
class/module to inherit its name, matching ruby.ts:extractOwnerName.

Also fixes findEnclosingClassNode in parse-worker.ts to skip
singleton_class and return the actual class/module node.

Adds integration test assertions for from_habitat (class << self method):
HAS_METHOD edge from Animal, isStatic=true, parameterCount=1.

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 16:11:31 +01:00
Abhigyan Patwari 5c4fca21c3 Merge pull request #626 from ivkond/feat/intra-repo-service-tracking-clean
[group] Intra-repo service communication tracking
2026-04-03 17:04:24 +05:30
Nguyen Hai SonandClaude Opus 4.6 dd0f5eed7d feat(vue): Vue SFC support + destructured call result tracking (#604)
* feat(vue): add Vue SFC (.vue) support for indexing

Vue Single File Components are now fully supported in the indexing pipeline.
The implementation extracts <script> / <script setup> blocks from .vue files
and parses them using the existing TypeScript tree-sitter grammar — no new
npm dependencies required.

Key changes:
- SFC script extractor: regex-based extraction of <script setup lang="ts">
  blocks with correct line offset mapping back to the .vue file
- Vue language provider: reuses TypeScript queries, type config, field
  extractors, and named binding extraction
- Import resolution: .vue added to EXTENSIONS so `import Foo from './Foo'`
  resolves to Foo.vue; Vue import resolver delegates to TS resolver for
  tsconfig path alias support
- Export detection: <script setup> top-level bindings are implicitly exported
- Template component detection: PascalCase tags in <template> emit CALLS edges
- Line offsets applied to all emitted positions (startLine, endLine, route
  lineNumbers, decorator positions) in both worker and sequential paths

Validated on a 3,553-file Vue project:
  Before: 24,693 nodes | 73,614 edges | 0 symbols from .vue
  After:  30,495 nodes | 112,324 edges | 5,213 symbols from .vue
          18,682 imports from .vue | 5,826 vue-to-vue imports

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(typescript): track destructured call results in TypeEnv

Extend `extractPendingAssignment` to handle object destructuring from
function calls and await expressions:

  const { isMaker } = useUserRole()
  const { data } = await fetchData()
  const { name } = repo.getProfile()

Previously, only `const { x } = someVariable` (identifier RHS) produced
TypeEnv bindings. Call-expression RHS was silently skipped, leaving
destructured properties untracked.

The fix emits a synthetic `callResult` item plus N `fieldAccess` items
per destructured property, which the existing fixpoint resolver processes
in 2 iterations. No changes needed to type-env.ts, PendingAssignment
types, or call-processor — the existing infrastructure handles it.

Also extracts a `collectDestructuredFields` helper to share the
object_pattern property iteration logic between the identifier and
call-expression branches.

Note: Full property-type resolution requires the callee to have a
declared returnType in the SymbolTable. Arrow-function composables
without type annotations (common in Vue/React) won't resolve property
types until return-type inference is added in a future change.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(vue): address PR review issues for Vue SFC support

- Extract duplicated isVueSetupTopLevel to vue-sfc-extractor.ts shared
  utility, removing identical copies from parse-worker.ts and
  parsing-processor.ts
- Fix VUE_BUILT_INS to be a superset of TS BUILT_INS by importing and
  spreading the TypeScript set, preventing spurious unresolved calls for
  standard built-ins (Symbol, BigInt, WeakMap, array methods, etc.)
- Add Vue template component CALLS edge resolution in both sequential
  and worker paths (call-processor.ts), matching PascalCase template
  tags against imported .vue file basenames via the import map
- Add integration test for template PascalCase CALLS edges
  (App.vue → Button.vue)
- Add integration test for isExported: false on non-setup <script>
  blocks (OldStyle.vue options API)
- Add comment explaining TEMPLATE_RE greedy regex behavior for nested
  template tags
- Fix stale language count comment (14 → 15) and remove dead code
  branch in test

Made-with: Cursor

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 14:18:55 +05:30
Chirag Nighut e3d73a7aed Java method reference (#622) 2026-04-02 15:16:35 +01:00
ivkondandClaude Opus 4.6 255e3e79eb fix(group): address 4 HIGH-priority issues from PR #626 review
1. Path traversal via group name — add validateGroupName() with regex
   [a-zA-Z0-9][a-zA-Z0-9_-]*, called in getGroupDir (defense in depth)

2. gRPC proto regex can't handle nested braces — replace serviceRe with
   extractServiceBlocks() brace-depth counter (init depth=1, skip
   malformed protos)

3. Service boundary detector directory exclusions — add EXCLUDED_DIRS
   set (vendor, target, build, dist, __pycache__, .venv, venv, .tox,
   .mypy_cache, .gradle, .mvn, out, bin) replacing inline node_modules

4. Double-close of LadybugDB pools — remove blanket closeLbug() from
   cli/group.ts; sync.ts per-id cleanup is sufficient

Tests: 22 new tests across 5 files. Full suite: 4706 passed, 0 failed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 12:55:33 +03:00
ivkondandClaude Opus 4.6 be4458b650 docs: add design spec and implementation plan for PR #626 HIGH fixes
Spec covers 4 HIGH-priority issues from review: path traversal via
group name, gRPC proto regex nested braces, service boundary detector
directory exclusions, double-close of LadybugDB pools.

Plan: 6 tasks with TDD, ordered by complexity.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 12:55:18 +03:00
ivkondandClaude Opus 4.6 4fed097abb feat(group): add sync pipeline, CLI, MCP tools, and monorepo fixture
Wire extractors into the sync pipeline with service boundary detection.
GroupService provides high-level API for all group operations.

- Sync pipeline: orchestrates extraction (HTTP, gRPC, topics) with
  service boundary assignment and exact matching
- GroupService: groupList, groupSync, groupContracts, groupQuery,
  groupStatus (groupImpact deferred to cross-repo follow-up PR)
- CLI: group create/add/remove/list/sync/contracts/query/status
- MCP tools: group_list, group_sync, group_contracts, group_query,
  group_status
- Monorepo fixture: 3 services (auth/orders/gateway) connected via
  gRPC + Kafka + HTTP — all intra-repo cross-links discovered
- Documentation: CLI commands and MCP tools added to both READMEs

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 00:40:31 +03:00
ivkondandClaude Opus 4.6 4fa395f4b6 feat(group): add service boundary detection and contract extractors
Service communication detection for microservice monorepos:

- ServiceBoundaryDetector: auto-detects service boundaries via markers
  (package.json, go.mod, Dockerfile, pom.xml, Cargo.toml, build.gradle,
  pyproject.toml, etc.)
- HttpRouteExtractor: graph-assisted (Strategy A) with source-scan
  fallback (Strategy B) for Spring, Express, Laravel, FastAPI providers
  and fetch/axios consumers
- GrpcExtractor: parses .proto files, detects Go/Java/Python/TS gRPC
  servers (RegisterXxxServer, @GrpcService, add_XxxServicer_to_server,
  @GrpcMethod) and clients (NewXxxClient, newBlockingStub, XxxStub)
- TopicExtractor: Kafka (@KafkaListener, producer.send), RabbitMQ
  (@RabbitListener, channel.publish/consume), NATS (nc.Subscribe/Publish)
  across Java, Node, Go, and Python

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 00:40:12 +03:00
ivkondandClaude Opus 4.6 52277247fe feat(group): add group infrastructure and contract matching
Core foundation for repository group analysis:
- Type system: ContractType, ExtractedContract, StoredContract, CrossLink
  with optional `service` field for intra-repo matching
- Config parser for group.yaml (repos, detection flags, matching thresholds)
- Contract registry storage with atomic writes
- Exact matching engine with per-type normalization (HTTP, gRPC, topic)
  and intra-repo support (different services within same repo can match)
- Extract LadybugDB pool-adapter from MCP backend for reuse by sync pipeline
- Git staleness checker for group status reporting

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 00:39:43 +03:00
Gergő Magyar ba5de0bde4 feat(cpp): C/C++ MethodExtractor config with pure virtual detection (#617)
* feat(cpp): C/C++ MethodExtractor config with pure virtual detection (#572)

- Pure virtual (= 0) detected as isAbstract via token scanning
- virtual/final/override via hasKeyword and virtual_specifier children
- Access specifier visibility via backward sibling walk (public:/private:/protected:)
- Pointer/reference parameter types extracted correctly
- Constructor and destructor support via declaration node type
- Static detection via storage_class_specifier
- 16 new tests covering all acceptance criteria

* fix(cpp): isVirtual infers from override/final + out-of-class resolution

- isVirtual returns true for override/final methods (C++ mandates these
  are virtual)
- Add findClassNodeByQualifiedName to parse-worker: resolves Foo::bar()
  back to the Foo class declaration for method extractor enrichment
- Handles pointer/ref return types, constructors, destructors
- Integration test for virtual/static/constructor inline methods
- 233 unit+integration tests pass, 97 C++ resolver tests pass

* fix(cpp): address review — deep pointers, templates, unions, trailing returns

- Fix extractParamName: recursive unwrap for int** ptr → "ptr" (not "**ptr")
- Fix findFunctionDeclarator: recursive unwrap for multi-level pointer chains
- Template methods: generic extractor unwraps template_declaration to inner node
- union_specifier: added to typeDeclarationNodes, visibility defaults to public
- Trailing return type: auto foo() -> T now extracts T instead of "auto"
- Fix version comment: ^0.22.4 → ^0.23.4 to match package.json
- 4 new tests: double pointer params, template methods, union methods, trailing returns

* fix(cpp): template method visibility + union isTypeDeclaration test

extractCppVisibility now walks from the template_declaration parent
when the node is wrapped by a template, restoring correct access-
specifier resolution for templated class methods.

Also adds missing isTypeDeclaration assertion for union_specifier and
expands the template method test with explicit visibility checks.

* fix(cpp): address deep gap analysis review findings

- findClassNodeByQualifiedName: recursive pointer/reference
  declarator unwrap, fixing out-of-class linking for deep pointer
  return types (e.g. int** Foo::bar())
- findClassNodeByQualifiedName: recurse into namespace_definition
  blocks so namespace-wrapped classes resolve correctly
- Suppress = delete / = default special members from extraction
  via delete_method_clause / default_method_clause node detection
- Update known-gaps: namespace-wrapped classes, const-overload collapse
- Add tree-sitter-c version comment for consistency
- toBeFalsy() → toBe(undefined) for precise isVirtual assertion
- Tests: = delete, = default, = 0 non-regression, operator overloads,
  deep pointer return types, default visibility (class vs struct),
  multiple access specifier sections
2026-04-01 18:07:11 +01:00
Abhigyan PatwariandClaude Opus 4.6 dc86ea96dc chore: release v1.5.3 — update CHANGELOG and package-lock
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 21:35:24 +05:30
Abhigyan PatwariandClaude Opus 4.6 d1adc8331a chore: release v1.6.0 — update CHANGELOG and package-lock
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 21:32:11 +05:30
80d363f145 fix(wiki): Azure OpenAI compat and HTML viewer script injection (#618)
* fix(wiki): Azure OpenAI compat and HTML viewer script injection

- Use max_completion_tokens instead of deprecated max_tokens for all models
- Skip sending temperature for Azure provider (some models reject non-default values)
- Simplify Azure interactive setup: endpoint + deployment + key (3 prompts instead of 7)
- Escape </script> in embedded JSON to prevent premature script tag closure

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): align wiki-llm-client test with max_completion_tokens change

The test expected max_tokens for non-reasoning models, but the source
now uses max_completion_tokens for all models since max_tokens is
deprecated by newer OpenAI models.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Abhigyan Patwari <abhigyan@Abhigyans-MacBook-Air.local>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 21:30:43 +05:30
Gergő Magyar 12be2025f1 feat(ts,js): TypeScript/JavaScript MethodExtractor config (#588)
* feat(ts,js): MethodExtractor config for TypeScript and JavaScript (#570)

Add per-language method extraction config following the established
JVM and C# patterns. Shared config base mirrors the field extractor's
typescript-javascript.ts pattern — TS-only node types are harmless
no-ops for JS.

Key features:
- isAbstract for abstract class methods and interface methods
- Parameter extraction with isOptional (?:, defaults) and isVariadic (...)
- Decorator extraction from preceding body-level siblings
- isAsync and isOverride detection
- Visibility via accessibility_modifier two-pass pattern
- Return type extraction unwrapping type_annotation

* test(ts,js): add override, getter/setter, destructured param tests

Address code review findings:
- Add override method detection test
- Add getter/setter extraction test
- Add destructured parameter with type annotation test
- Tighten constructor and private method assertions

* refactor(ts,js): address code review findings

- Replace O(M*N) decorator index scan with previousNamedSibling walk
- Remove dead findVisibility 'modifiers' fallback (TS uses
  accessibility_modifier, not a modifiers wrapper)
- Document call_signature/construct_signature as known gaps
- Document that TS constructors are method_definition nodes
- Remove unused findVisibility import

* fix(ts,js): type guard before cast, add generator/computed/overload tests

- Use type guard pattern (Set.has check before as-cast) in visibility
  extraction to ensure string is validated before narrowing
- Add generator method test (*items()) — confirms extraction works
- Add computed property name test ([Symbol.iterator]) — documents
  bracket-in-name behavior as intentional
- Add class-level method overload test — verifies overload signatures
  + implementation are all extracted

* fix(ts,js): detect #private methods as visibility 'private'

ES2022 private class methods (#name) use private_property_identifier
as their name node type. Detect this and return 'private' visibility
instead of the default 'public'.

* fix(ts,js): address review findings + close ingestion gaps

- hasKeyword/findVisibility: skip name field child to prevent false
  positives on soft-keyword method names (e.g. `abstract()`, `static()`)
- extractTsJsParameters: filter TS `this` parameter (compile-time only)
- extractMethodSignature: mirror `this`-param skip in fallback path
- tree-sitter queries: capture abstract_method_signature,
  method_signature, and private_property_identifier for TS; add
  private_property_identifier for JS
- Remove dead childForFieldName('name') fallbacks and typeFromAnnotation
  fallback
- Add 10+ unit tests, 4 integration tests through query pipeline

* test(ts): update HAS_METHOD count for interface method_signature capture

The new method_signature query now captures ILogger.log() as a Method
node with a HAS_METHOD edge, increasing the expected count from 4 to 5.

* fix(ts,js): address second review — async generator test, declare module gap

- Add async generator method test (async *values() → isAsync: true)
- Document declare module/global augmentation as known gap
2026-04-01 14:09:59 +01:00
Abhigyan PatwariandClaude Opus 4.6 dcf40da3a8 fix: ensure import rewrites survive npm publish lifecycle
npm runs `prepare` after `prepack` during publish, so the previous
`prepare: tsc` overwrote the rewritten imports before packing.
Both `prepare` and `prepack` now run the full build script so the
tarball always contains rewritten relative imports.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 17:28:14 +05:30
Abhigyan PatwariandClaude Opus 4.6 c617350c58 chore: release v1.5.1 — update CHANGELOG and package-lock
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 17:05:08 +05:30
71255a096c fix: bundle gitnexus-shared into CLI dist (#613)
* fix(ci): build gitnexus-shared before publish, use CHANGELOG for release notes

The publish workflow was missing the gitnexus-shared build step that
the setup-gitnexus composite action provides. Since PR #536 unified
the ingestion pipeline, gitnexus imports types from gitnexus-shared,
so it must be built first.

Also replaces generate_release_notes with CHANGELOG.md extraction so
GitHub Releases use the reviewed changelog entry instead of a flat
PR title list.

Made-with: Cursor

* fix: bundle gitnexus-shared into CLI dist to fix module resolution

gitnexus-shared was declared as a file: dependency but never published
to npm, causing ERR_MODULE_NOT_FOUND for users installing gitnexus
globally. The build script now copies gitnexus-shared/dist into
dist/_shared/ and rewrites bare specifiers to relative paths.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: move gitnexus-shared to devDependencies, use tsc for prepare

gitnexus-shared must remain available for tsc to resolve imports during
development/CI, but is not needed at runtime since it's bundled into
dist/_shared/. Moving it to devDependencies keeps it out of production
installs while allowing compilation. The prepare script now runs plain
tsc (no shared bundling needed for local dev).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Abhigyan Patwari <abhigyan@Abhigyans-MacBook-Air.local>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 17:04:39 +05:30
Gergo Magyar 39322e37df fix(ci): build gitnexus-shared before npm ci in publish workflow
The publish job failed because npm ci triggers the prepare script (tsc)
before gitnexus-shared types are available. Add the same build step
that the setup-gitnexus composite action uses in CI.
2026-04-01 07:35:46 +01:00
Abhigyan PatwariandAbhigyan Patwari 0e6f8f1720 chore: release v1.5.0 — update CHANGELOG and package-lock (#611)
Made-with: Cursor

Co-authored-by: Abhigyan Patwari <abhigyan@Abhigyans-MacBook-Air.local>
2026-04-01 11:24:38 +05:30
Abhigyan Patwari 113dadf945 chore: bump version to 1.5.0
Made-with: Cursor
2026-04-01 10:34:46 +05:30
Abhigyan PatwariandAbhigyan Patwari af421cc9a3 feat(web): repo landing screen with selectable repo cards (#607)
* feat(web): add repo landing screen with selectable repo cards

Instead of auto-loading the first indexed repo when the backend is
detected, show a landing screen that lets users choose which repo
to explore or analyze a new one. This addresses the UX gap where
users with multiple indexed repos had no way to pick—they were
always sent to the first one found.

- New RepoLanding component with clickable repo cards (name, stats,
  indexed date) and an embedded RepoAnalyzer for new repos
- DropZone gains a 'landing' phase between server detection and
  graph loading
- Shared connectToRepo handler replaces the old handleAnalyzeComplete
  for both repo selection and post-analysis connection

Made-with: Cursor

* fix(web): update e2e flow for repo landing screen

The new landing screen intentionally stops auto-loading the first indexed
repo, so the existing Playwright tests were still waiting for the explorer
to appear automatically. Update the specs to select a repo from the landing
screen before asserting on the graph, and add a stable test id for repo cards.

Also format DropZone to satisfy the Prettier CI check.

Made-with: Cursor

* fix(e2e): use waitFor instead of instant isVisible for landing card

locator.isVisible() is a non-retrying instant check — the landing card
hadn't rendered yet when it was called, causing the click to be silently
skipped. Switch to waitFor which properly polls until the element appears.

Made-with: Cursor

---------

Co-authored-by: Abhigyan Patwari <abhigyan@Abhigyans-MacBook-Air.local>
2026-04-01 10:34:10 +05:30
Gergő Magyar c72890d59d feat(csharp): C# MethodExtractor config (#582)
* feat(csharp): add C# MethodExtractor config (#573)

Add C# method extraction config mirroring the JVM pattern from PR #576.
Wire csharpMethodConfig into the C# language provider and add 18 tests
covering classes, interfaces, abstract classes, structs, records,
constructors, params/out/ref/optional parameters, sealed methods,
attributes, and visibility modifiers.

* fix(csharp): add destructor, operator, conversion operator, and in-param support

- Add destructor_declaration, operator_declaration, and
  conversion_operator_declaration to methodNodeTypes
- Custom extractName for operators (e.g., "operator +", "implicit operator double")
- Fix extractReturnType for operator declarations (use type field, not returns)
- Add in modifier to parameter extraction (alongside out/ref)
- Add 4 new tests: destructor, operator+, implicit conversion, in parameter

* fix(csharp): add ref param test and document compound visibility limitation

- Add test for ref parameter modifier (was only testing out)
- Document that protected internal / private protected resolve to first modifier

* feat(csharp): support compound visibilities (protected internal, private protected)

- Add 'protected internal' and 'private protected' to FieldVisibility union
- Detect compound modifiers in both C# method and field extractors via
  collectModifierTexts helper scanning adjacent modifier nodes
- Add 2 tests for compound visibility detection

* feat(csharp): primary constructors, virtual/override/async, primary fields

Address all known limitations from review:

- Primary constructor support (C# 12): add extractPrimaryConstructor to
  MethodExtractionConfig and extractPrimaryFields to FieldExtractionConfig.
  Record params become public readonly properties; class params become
  private captured fields.
- Add isVirtual, isOverride, isAsync optional fields to MethodInfo,
  MethodExtractionConfig, NodeProperties, and parse-worker propagation.
- Detect virtual/override/async modifiers in C# method config.
- Move collectModifierTexts to shared helpers.ts (deduplicate).
- Fix destructor name to ~ClassName (disambiguates from constructor).
- Add expression-bodied method test.
- 118 tests total across method + field extraction suites, all passing.

* fix(csharp): review round 2 — annotations, record_struct, grammar pin

- Fix primary constructor annotations: use [] instead of extracting
  class-level attributes (C# has no syntax for ctor-specific attributes)
- Add record_struct_declaration to typeDeclarationNodes in both method
  and field extractors, CLASS_CONTAINER_TYPES, and isRecord visibility check
- Pin tree-sitter-c-sharp version (^0.23.1) in params comment

* fix(csharp): complete record_struct query + label mapping, sealed override test

- Add record_struct_declaration capture patterns to tree-sitter-queries.ts
  (type definition + primary constructor)
- Add record_struct_declaration → 'Struct' in CONTAINER_TYPE_TO_LABEL
- Assert isOverride: true alongside isFinal in sealed override test

* fix(csharp): record_struct label mismatch, add record struct + documented limitation tests

- Fix record_struct_declaration query tag: @definition.struct (not @definition.record)
  to match CONTAINER_TYPE_TO_LABEL and prevent broken HAS_METHOD edges
- Add 3 record struct tests: isTypeDeclaration, method extraction, primary constructor
- Add documented limitation tests: partial method (isAbstract: false), generic type
  parameter stripping (name excludes <T>)

* fix(csharp): remove record_struct_declaration — not a real tree-sitter node type

tree-sitter-c-sharp 0.23.1 parses 'record struct' as record_declaration
(absorbs the 'struct' keyword as an unnamed child token). The non-existent
record_struct_declaration in queries caused TSQueryErrorNodeType, breaking
ALL C# file processing.

Remove from: tree-sitter-queries.ts, typeDeclarationNodes in both
extractors, CLASS_CONTAINER_TYPES, and CONTAINER_TYPE_TO_LABEL.
Record struct types are already handled via record_declaration.

* feat(csharp): add isPartial support, filter targeted attributes, static ctor test

- Add isPartial optional field to MethodInfo, MethodExtractionConfig,
  NodeProperties, and parse-worker propagation pipeline
- Detect partial modifier in C# config — marks both declaration-only
  and implemented partial methods
- Filter targeted attribute lists (e.g. [return: MarshalAs(...)]) in
  extractCSharpAnnotations — only untargeted attributes collected
- Add static constructor test (isStatic: true, same name as class)
- Add 3 partial method tests: declaration-only, with body, coexisting pair
- Document record_struct/record_class as defensive dead code in
  export-detection.ts (grammar absorbs keywords into record_declaration)

* fix(csharp): this param for extension methods, dedup visibility, test fixes

- Handle this modifier on extension method parameters (type prefixed
  as 'this string', consistent with out/ref/in handling)
- Deduplicate visibility logic in extractPrimaryConstructor — reuse
  csharpMethodConfig.extractVisibility instead of inline compound check
- Fix record struct test title to reflect actual grammar behavior
- Add conversion operator returnType assertion
- Add extension method this parameter test

* fix(csharp): primary constructor line points to param list, empty name guard

- Use paramList.startPosition instead of ownerNode.startPosition for
  primary constructor line number (avoids methodInfoCache key collision)
- Guard against empty param names from tree-sitter error recovery nodes
2026-03-30 08:41:17 +01:00
agung widhiatmojoandClaude Opus 4.6 be210c7c61 docs: add gitnexus-shared build step before gitnexus-web (#585)
gitnexus-web imports from gitnexus-shared, which requires npm run build
to generate dist/. Without this step, npm run dev fails with module
resolution errors.

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-30 08:40:47 +01:00
Subham Kundu 8d7aa94381 chore: add enterprise offering section to README, ignore local_docs/ (#579)
* chore: add enterprise offering section to README, ignore local_docs/

* chore: normalize em dash to hyphen in README enterprise section

* docs: expand enterprise section with commercial licensing and feature details
2026-03-29 14:55:17 +05:30
Subham Kundu d90d1ba96f fix(eval): exclude litellm 1.82.7 and 1.82.8 due to compatibility issues (#580) 2026-03-29 05:23:37 +01:00
Gergő Magyar 313b13fade feat(java,kotlin): MethodExtractor abstraction with per-language configs (#576) 2026-03-28 21:31:08 +00:00
Gabriel J Campbell b03413dcf9 feat: added skip-agents-md cli flag (#517)
* feat: added skip-agents-md cli flag

* fix: apply prettier formatting

* feat: added skip-agents-md cli flag

* fix: apply prettier formatting

* feat: add skipAgentsMd option to skip AGENTS.md and CLAUDE.md updates

* fixed bad merge
2026-03-28 21:23:59 +00:00
Abhigyan PatwariandClaude Opus 4.6 9f69c43100 feat(wiki): Azure OpenAI support for wiki command (#562)
* feat(wiki): extend LLMConfig/CLIConfig with Azure and reasoning model fields

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(wiki): restore cursor model resolution, fix LLMProvider type, clean up regex

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(wiki): remove stale LLMProvider type alias from repo-manager

* fix(wiki): fix Azure auth header, api-version param, reasoning model params, content_filter error

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(wiki): tighten Azure detection, reasoning model regex, content_filter gating

- isReasoningModel: new regex matches only o1/o3 bare + any oN-mini/oN-preview; bare o4/o5/etc now return false
- isAzureProvider: use URL hostname matching to block spoofed subdomain URLs
- callLLM: warn on Azure legacy /deployments/ URL without api-version
- callLLM: gate content_filter error to azure===true; also catch ResponsibleAIPolicyViolation
- tests: add afterEach stub cleanup, spoofed-URL, bare-o4, non-Azure content_filter, and URL-only Azure auto-detect tests

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(wiki): detect content_filter finish_reason in SSE stream and throw clear error

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(wiki): skip delta accumulation after content_filter, use provider-neutral error message

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(wiki): add Azure OpenAI option to interactive setup wizard

Inserts Azure as option [3] in the provider menu (shifting Custom to [4]
and Cursor to [5]), adds guided Azure setup flow with resource/deployment
prompts, v1/legacy URL format selection, reasoning-model flag, and
content_filter error handling in the catch block.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(wiki): store explicit false for non-reasoning Azure deployments, trim resource name inputs

- isReasoningModelDeployment now stores false (not undefined) when user says no
- Always include isReasoningModel in saved azureConfig (no conditional guard needed)
- Trim resourceName and deploymentName prompt inputs to avoid whitespace issues
- Improve reasoning model note to mention Azure requirement

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(wiki): add --api-version and --reasoning-model CLI flags for Azure

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(wiki): include apiVersion and reasoningModel in hasCLIOverrides guard

* fix(wiki): default provider to 'openai' in resolveLLMConfig when not configured

* style: apply prettier formatting

* fix(wiki): address PR review — remove unrelated files, harden inputs

- Remove evidence/, fix-adapter.js, and planning doc accidentally included
- URL-encode apiVersion in buildRequestUrl to prevent query string injection
- Add --no-reasoning-model flag to allow CLI override of saved config
- Simplify verbose ternary in Azure wizard prompt

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(wiki): use execFileSync for EDITOR to prevent shell injection

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-29 00:23:41 +05:30
Gergő Magyar ddb6a704b3 refactor: reduce explicit any types (#566)
* refactor: replace NodeProperties index signature any with unknown

Change [key: string]: any to [key: string]: unknown in NodeProperties.
Remove 19 redundant (node.properties as any) casts in csv-generator.ts
— all accessed properties are already declared on the type.

* refactor: replace any with SyntaxNode across ingestion layer

Mechanical substitution — all tree-sitter AST node parameters and
variables typed as any are now properly typed as SyntaxNode.

- ast-helpers.ts: 13 any → SyntaxNode
- parsing-processor.ts: 8 any → SyntaxNode
- parse-worker.ts: 40 any → SyntaxNode/TreeSitterLanguage/Parser.Query
- php.ts: 11 any → SyntaxNode

Also adds TreeSitterLanguage type alias for optional grammar loading.

* refactor: eliminate remaining any in ingestion layer

- call-processor, call-routing, c-cpp: SyntaxNode substitutions
- parse-worker: typed WorkerIncomingMessage discriminated union
- worker-pool: typed WorkerOutgoingMessage + Error handler
- ast-cache, import-processor: targeted cast for Tree.delete()
- community-processor: graphology AbstractGraph types, LeidenModule
  interface for vendored leiden code

Ingestion layer: 130 → 7 any warnings remaining.
2026-03-28 17:11:24 +00:00
Chirag Nighut e2de9271fc feat(java): method references, worker overload disambiguation, interface dispatch (#540)
* feat(java): method references + worker overload disambiguation (TypeEnv + argTypes)

Fix two Java gaps: (1) method references (obj::method) via
  tree-sitter @call + parseJavaMethodReference wired through extractLanguageCallSiteSeed for parse-worker and call-processor; (2) overloaded calls with typed
  non-literal args by extending OverloadHints with TypeEnv for identifiers and adding ExtractedCall.argTypes from extractCallArgTypes on the worker path with
  matchCandidatesByArgTypes (inferJvmLiteralType remains for literals).

* test(csharp): expect interface-dispatch edge for IRepository.Save in heritage fixture

* refactor(ingestion): move parseJavaMethodReference to call-sites/java.ts

* refactor(ingestion): defer worker call resolution until implementor map is complete

* style: prettier + remove unused import for CI quality checks

Made-with: Cursor

* fix(ingestion): implementor map for C# base_list + sequential pipeline path

- buildImplementorMap: treat extends rows as implements when resolveExtendsType
  says IMPLEMENTS (worker heritage mirrors parse-worker, all base_list as extends)
- Worker path: pass ctx into buildImplementorMap(deferredWorkerHeritage, ctx)
- Sequential path: extract heritage before processCalls and pass implementor map
  so small repos get interface-dispatch CALLS (fixes csharp-proj integration test)

Made-with: Cursor

* perf(pipeline): accumulate sequential implementor map without O(E) per chunk

- Merge buildImplementorMap(chunk heritage) into one map each sequential chunk so
  work is O(heritage) per chunk and interface dispatch sees prior chunks (worker parity)
- Drop unused globalImplementorMap + redundant merge after worker pass

Made-with: Cursor
2026-03-28 16:59:28 +00:00
Gergő Magyar acf6fbdd39 feat: configure eslint with unused import removal (#564)
* feat: configure eslint with unused import removal

Add ESLint v9 (flat config) for code quality:
- eslint-plugin-unused-imports for auto-removing dead imports
- @typescript-eslint for TypeScript-aware linting
- eslint-plugin-react-hooks for React hooks rules
- eslint-config-prettier to avoid formatting conflicts
- lint-staged runs eslint --fix before prettier on .ts/.tsx
- CI lint job added to ci-quality.yml

* refactor: remove unused imports via eslint --fix

Auto-fixed by eslint-plugin-unused-imports. No logic changes.

* chore: add eslint fix commit to .git-blame-ignore-revs
2026-03-28 15:28:09 +00:00
Gergő Magyar bf09eab95b feat: configure prettier with pre-commit hook (#563)
* feat: configure prettier with pre-commit hook integration

Add prettier, lint-staged, and prettier-plugin-tailwindcss at the repo
root with husky pre-commit hook integration. Moves husky from
gitnexus/ to root package.json for reliable hook installation.

- Root package.json with prepare/format/format:check scripts
- .prettierrc with endOfLine:lf and tailwindStylesheet for TW v4
- .prettierignore excluding fixtures, vendor, generated, *.d.ts, *.md
- .gitattributes enforcing LF line endings for Windows consistency
- Pre-commit hook uses direct node_modules/.bin/ paths (no npx)

* style: apply prettier formatting to entire codebase

One-time bulk format. No logic changes.
Use .git-blame-ignore-revs to skip this commit in git blame.

* chore: add .git-blame-ignore-revs for prettier format commit

* perf: pre-commit hook runs only tests related to staged files

Use vitest --related to scope test execution to tests that import
the changed files, instead of running the full suite on every commit.

* perf: remove vitest from pre-commit hook, keep in CI only

Pre-commit now runs lint-staged + tsc only. Tests run in CI
(ci-tests.yml) where they belong — keeps commits fast.

* ci: add prettier format check to quality workflow

PRs will now fail if code isn't formatted with prettier.
2026-03-28 14:58:04 +00:00
Gergő Magyar fd7fb5bf1f feat: unify web and cli ingestion pipeline (#536)
* feat: add server-side ingestion API (POST /api/analyze, SSE progress)

Extract core analysis orchestration from CLI into shared run-analyze.ts
module. Add server-side analyze endpoints so the web app can trigger
ingestion via HTTP instead of running the full pipeline in-browser.

New files:
- src/core/run-analyze.ts — shared runFullAnalysis() orchestrator
- src/server/analyze-job.ts — job manager (single-slot, dedup, SSE events)
- src/server/analyze-worker.ts — forked child process (8GB heap, IPC)
- src/server/git-clone.ts — shallow clone/pull with SSRF protection

API endpoints:
- POST /api/analyze — start analysis (returns 202 + jobId)
- GET /api/analyze/:jobId — poll job status
- GET /api/analyze/:jobId/progress — SSE progress stream

Security: URL validation blocks private IPs and non-HTTP schemes.
Path validation requires absolute paths. Git stderr not leaked to API.

* feat(web): add server-side analyze UI (Phase 2)

Add "Analyze on Server" flow to the web app's Server tab so users
can trigger server-side ingestion from the browser. On completion,
the graph is automatically loaded via the existing connectToServer flow.

New files:
- AnalyzeProgress.tsx — progress bar with phase label, elapsed time, cancel

Modified files:
- backend.ts — startAnalyze(), streamAnalyzeProgress() SSE client
- DropZone.tsx — analyze URL input + button below Connect section
- App.tsx — onServerAnalyze handler wires analyze -> connect flow

* feat: add job cancellation, timeout, and child process tracking (Phase 3)

- DELETE /api/analyze/:jobId — cancel running analysis (SIGTERM to worker)
- 30-minute timeout kills long-running workers automatically
- Child process refs tracked in JobManager for cleanup on shutdown
- dispose() kills all active children on SIGINT/SIGTERM
- Web cancel button now calls server DELETE endpoint
- cancelAnalyze() added to web backend client

* refactor(web): remove browser ingestion pipeline (Phase 4)

Delete 16 duplicated ingestion files, 2 unused service files
(git-clone, zip), and tree-sitter parser-loader from gitnexus-web.
All ingestion now runs server-side via POST /api/analyze.

Deleted (18 files, ~5,000 lines):
- core/ingestion/*.ts (16 pipeline processors)
- core/tree-sitter/parser-loader.ts (WASM tree-sitter loader)
- services/git-clone.ts (isomorphic-git client-side clone)
- services/zip.ts (JSZip extraction)

Simplified:
- DropZone.tsx — server-only (removed ZIP/GitHub tabs)
- ingestion.worker.ts — removed runPipeline/runPipelineFromFiles
- useAppState.tsx — removed pipeline callbacks
- App.tsx — removed handleFileSelect/handleGitClone
- main.tsx — removed Buffer polyfill for isomorphic-git
- types/pipeline.ts — removed PipelineResult/serialize helpers

Kept: cluster-enricher.ts (LLM enrichment, still used by worker)

Dependencies now removable: web-tree-sitter, isomorphic-git,
@isomorphic-git/lightning-fs, jszip (estimated 3-4MB bundle savings)

* refactor(web): sync graph schema from CLI + delete WASM grammars

Sync graph/types.ts and lbug/schema.ts from the CLI (source of truth)
to the web module so the browser LadybugDB can handle all node and
relationship types the server pipeline produces.

Synced types: Route, Tool, Section node labels; HANDLES_ROUTE, FETCHES,
HANDLES_TOOL, ENTRY_POINT_OF, WRAPS, QUERIES relationship types;
description fields on Function/Class/Interface/Method/CodeElement.

Deleted: public/wasm/ directory (14 tree-sitter WASM grammars + core).
Removed deps: web-tree-sitter, isomorphic-git, @isomorphic-git/lightning-fs,
jszip, buffer, @types/jszip (~3-4MB bundle savings).

* feat: create gitnexus-shared package for unified type definitions

Create a new gitnexus-shared package that is the single source of truth
for types shared between the CLI and web modules:

- SupportedLanguages enum (15 languages)
- Graph types: NodeLabel, NodeProperties, RelationshipType, GraphNode, GraphRelationship
- Schema constants: NODE_TABLES, REL_TYPES, REL_TABLE_NAME, EMBEDDING_TABLE_NAME
- Pipeline types: PipelinePhase, PipelineProgress

Both gitnexus (CLI) and gitnexus-web import from gitnexus-shared via
file: dependency. Each package re-exports and extends with platform-specific
additions (CLI: KnowledgeGraph with mutation methods; Web: simpler KnowledgeGraph).

This ensures types can never drift between packages — adding a new
language, node type, or relationship type in gitnexus-shared automatically
propagates to both consumers.

* refactor: import shared types directly from gitnexus-shared at call sites

Replace all re-export patterns with direct imports from gitnexus-shared.
72 files updated across CLI and web:

- SupportedLanguages: 49 CLI files now import from 'gitnexus-shared'
  instead of '../config/supported-languages.js'
- GraphNode, GraphRelationship, NodeLabel: 22 CLI + 10 web files now
  import from 'gitnexus-shared' instead of local re-export wrappers
- NODE_TABLES: api.ts imports from 'gitnexus-shared'
- PipelineProgress: useAppState.tsx imports from 'gitnexus-shared'

Local types.ts files now only define platform-specific KnowledgeGraph
(CLI has mutation methods, web has add-only). No more re-exports.

* fix: update lock files for gitnexus-shared, remove stale vite polyfills

Add gitnexus-shared@1.0.0 to lock files so npm ci succeeds in CI.
Remove buffer polyfill and global define from vite.config.ts (isomorphic-git was removed).

* fix(security): add write guard to HTTP /api/query, fix CORS proxy bypass

- Add isWriteQuery() check to POST /api/query handler — blocks CREATE,
  DELETE, SET, MERGE, DROP, etc. via HTTP API (guard was only in MCP
  pool adapter and browser-side, not the HTTP server path)
- Extend CYPHER_WRITE_RE with CALL, INSTALL, LOAD keywords
- Fix CORS proxy subdomain bypass: endsWith('github.com') allowed
  'evil-github.com'. Now requires exact match or '.github.com' suffix

* feat(server): enhance /api/search with enrichment, add /api/grep, strip graph content

- POST /api/search: add mode param (hybrid|semantic|bm25), server-side
  enrichment returns connections/cluster/processes per result in one call
  (collapses 31 sequential HTTP calls to 1 for the agent search tool)
- GET /api/grep: regex search across indexed file contents, eliminates
  need to transfer all file contents to browser
- GET /api/graph: strip content field by default (80-95% payload
  reduction). Use ?includeContent=true for backward compat
- Add LRU cache invalidation hook point for future caching

* feat(server): add /api/embed endpoint for server-side embedding generation

- POST /api/embed: triggers embedding pipeline via onnxruntime-node
  with JobManager for single-slot concurrency, timeout, and dedup
- GET /api/embed/:jobId: poll job status
- GET /api/embed/:jobId/progress: SSE stream with heartbeat, event IDs,
  and X-Accel-Buffering:no header for proxy compatibility
- DELETE /api/embed/:jobId: cancel running embedding job
- Maps embedding pipeline phases (ready→complete, error→failed) to
  JobManager status conventions

* feat(web): create consolidated BackendClient module

Single HTTP client replacing backend.ts, server-connection.ts, and
worker HTTP helpers. Includes:
- Typed methods: runQuery, search (enriched), grep, readFile, connect
- Generic streamSSE<T> utility extracted from analyze progress pattern
- BackendError with discriminated code field (network/server/client/timeout)
- Embed API: startEmbeddings, streamEmbeddingProgress, cancelEmbeddings
- Search with mode param (hybrid|semantic|bm25) and enrichment

* refactor(web): rewrite Graph RAG tools for backend-only HTTP queries

- Search tool: uses enriched /api/search (1 call replaces 31 sequential queries)
- Cypher tool: removes browser-side embedding; {{QUERY_VECTOR}} routes to
  /api/search with mode:'semantic' instead of local transformers.js
- Grep tool: uses /api/grep instead of in-memory fileContents map
- Read tool: uses /api/file instead of fileContents map lookup
- Impact tool: getCallSiteSnippet now async via /api/file
- createGraphRAGTools now accepts GraphRAGBackend interface instead of
  7 separate function params + fileContents map
- createGraphRAGAgent simplified to (config, backend, context?)
- Removed imports: embedder, lbug/schema (replaced with gitnexus-shared)
- Net: -205 lines

* refactor(web): delete WASM infrastructure, remove 7 packages (-5242 lines)

Delete browser-side LadybugDB, embeddings, search, and worker:
- gitnexus-web/src/core/lbug/ (adapter, csv-generator, schema, query-result)
- gitnexus-web/src/core/embeddings/ (embedder, pipeline, text-gen, types)
- gitnexus-web/src/core/search/ (bm25-index, hybrid-search)
- gitnexus-web/src/workers/ingestion.worker.ts (828 lines)
- gitnexus-web/src/services/server-connection.ts (merged into backend-client)
- gitnexus-web/src/types/lbug-wasm.d.ts

Remove packages: @ladybugdb/wasm-core, @huggingface/transformers,
comlink, minisearch, vite-plugin-wasm, vite-plugin-top-level-await,
vite-plugin-static-copy

Update vite.config.ts: remove WASM plugins, COOP/COEP headers,
worker config, optimizeDeps exclude

Update imports: App.tsx, DropZone, Header, AnalyzeProgress,
BackendRepoSelector, useBackend → backend-client

* refactor(web): replace Worker/Comlink with direct BackendClient calls

- useAppState: remove Worker instantiation, Comlink.wrap, apiRef.
  All queries now go through BackendClient HTTP functions directly.
- Agent runs on main thread (I/O-bound LLM streaming, not CPU-bound)
- initializeAgent: creates GraphRAGAgent with GraphRAGBackend interface
  bound to BackendClient methods (runQuery, search, grep, readFile)
- startEmbeddings: calls POST /api/embed + SSE progress instead of
  running browser-side transformers.js pipeline
- switchRepo: no longer loads graph into WASM DB or extracts fileContents
- App.tsx: handleServerConnect simplified (no fileContents, no loadServerGraph)
- Delete old backend.ts (replaced by backend-client.ts)
- Net: -396 lines

* fix(web): fix await-in-map build error in agent streaming

Move dynamic import of AIMessage outside .map() callback to avoid
"await can only be used inside an async function" build error.

* fix(web): remove stale apiRef references that broke chat functionality

sendChatMessage referenced apiRef.current (deleted Worker ref) which
would throw TypeError. Replaced with agentRef.current guard since agent
now runs on main thread.

* fix(server): dispose embedJobManager on shutdown, fix job mutation

- Add embedJobManager.dispose() to shutdown handler (was missing,
  causing cleanup timer to keep Node process alive)
- Replace direct job.repoName/status mutation with updateJob() to
  ensure SSE event emission for initial status change

* fix(server): parameterize Cypher, harden grep, unify SSE endpoints

- Search enrichment: replace string interpolation with executePrepared()
  using $nid parameter binding to prevent Cypher injection
- Add executePrepared() to core lbug-adapter (prepare/execute pattern)
- /api/grep: add 200-char pattern length limit (ReDoS protection),
  search files on disk instead of loading entire corpus into memory
  (constant memory usage regardless of repo size)
- Extract mountSSEProgress() shared helper for SSE streaming — both
  analyze and embed endpoints now have consistent heartbeat (30s),
  event IDs (reconnection support), and X-Accel-Buffering header

* refactor(web): remove dead code from Worker-era architecture

- Remove loadServerGraph no-op function, interface member, and all consumers
- Remove testArrayParams stub and interface member
- Remove fileContents state from GraphStateProvider (never populated in
  server-side architecture)
- Remove forceDevice parameter from startEmbeddings (server-side, no device choice)
- Replace phantom EmbeddingProgress type with inline { phase, percent }
- Replace resolvePathFromContents (needed fileContents Map) with graph-based
  file path resolution using filePathIndex built from graph nodes
- Fix: AI citation grounding ([[file.ts:10]]) now works via graph node lookup
  instead of broken fileContents-based resolution

* fix(web): use streamAgentResponse for full tool_call/reasoning streaming

Replace naive agent.stream() loop that only handled content chunks with
streamAgentResponse() generator from agent.ts. This properly routes:
- reasoning tokens (before/between tool calls)
- tool_call events (name, args, status)
- tool_result events (completed tool output)
- content tokens (final answer after all tools done)

Previously the onChunk handler for tool_call/tool_result/reasoning was
dead code since the streaming loop only emitted content events.

* fix(web): resolve CI type errors from dead code removal

- Import GraphNode/GraphRelationship from gitnexus-shared in graph.ts
  (not re-exported from local types.ts)
- Add Route, Tool entries to NODE_COLORS and NODE_SIZES constants
- Add PipelineResult type to web types/pipeline.ts
- Remove fileContents from CodeReferencesPanel and RightPanel
- Remove testArrayParams and forceDevice from EmbeddingStatus
- Remove forceDevice args from startEmbeddings() calls in App.tsx
- Fix embeddingProgress property accesses for simplified type

* fix(ci): add setup-gitnexus-web action, build shared once per job

- Remove prepare script from gitnexus-shared (tsc not available during
  npm ci of consuming packages)
- Create .github/actions/setup-gitnexus-web composite action: builds
  gitnexus-shared then runs npm ci for gitnexus-web
- setup-gitnexus action: already builds gitnexus-shared for CLI jobs
- ci-quality typecheck-web: uses setup-gitnexus-web (DRY)
- ci-e2e: uses setup-gitnexus-web (DRY)
- ci-tests: gitnexus-shared already built by setup-gitnexus, just
  install web deps without rebuilding

* fix(ci): use prepare script so gitnexus-shared builds during npm ci

Move typescript from devDependencies to dependencies in gitnexus-shared
so the prepare script (tsc) works when npm resolves file: deps during
npm ci. No GHA modifications needed — npm handles the build lifecycle
automatically.

Remove manual gitnexus-shared build steps from setup-gitnexus and
setup-gitnexus-web actions.

* fix(ci): build gitnexus-shared explicitly in setup actions

The file: dependency protocol doesn't reliably run prepare scripts
because devDependencies aren't installed first. Instead of fragile
lifecycle hacks, build gitnexus-shared explicitly in both setup actions:
- setup-gitnexus: npm install && npm run build in gitnexus-shared/
- setup-gitnexus-web: same, before npm ci in gitnexus-web/
- ci-tests: shared already built by setup-gitnexus, web just npm ci

No prepare script, no dist in git, no typescript as a prod dependency.

* fix: remove CALL from CYPHER_WRITE_RE — breaks FTS and vector search

CALL is used by read-only procedures: CALL QUERY_FTS_INDEX(...) and
CALL QUERY_VECTOR_INDEX(...). Adding it to the write guard blocked all
FTS search, causing 3 test failures. The database is opened in read-only
mode as defense-in-depth against write procedures via CALL.

Keep INSTALL and LOAD in the blocklist (genuinely dangerous).

* fix(web): update vercel.json for gitnexus-shared, remove COOP/COEP

- Add installCommand that builds gitnexus-shared before installing
  web deps (Vercel doesn't know about the monorepo file: dependency)
- Remove Cross-Origin-Opener-Policy and Cross-Origin-Embedder-Policy
  headers (no longer needed — WASM LadybugDB removed)

* fix(web): update tests for deleted modules

- Delete csv-generator.test.ts (tests deleted WASM-only csv-generator)
- Update security-guards.test.ts: import NODE_TABLES/REL_TYPES from
  gitnexus-shared instead of deleted src/core/lbug/schema
- Update server-connection.test.ts: import normalizeServerUrl from
  backend-client, remove extractFileContents tests (function deleted)

* fix(e2e): remove Server tab click — UI is now server-only

The DropZone no longer has ZIP/GitHub/Server tabs (browser ingestion
was removed). The server URL input is directly visible on the landing
page. Update e2e test to skip the tab click and go straight to input.

All 5 e2e tests pass locally.

* refactor: use gitnexus-shared for PipelinePhase/PipelineProgress types

CLI was duplicating PipelinePhase and PipelineProgress locally instead
of importing from gitnexus-shared. Updated all consumers to import
directly. Also removed dead code: SerializablePipelineResult,
serializePipelineResult(), deserializePipelineResult().

* fix(server): address PR #536 review — security, race conditions, dead code

- Fix path traversal in POST /api/analyze: split into isAbsolute + normalize check
- Add shared repo lock (activeRepoPaths) preventing concurrent analyze+embed on same repo
- Fix 202 response returning actual job.status instead of hardcoded 'queued'
- Add 30-minute timeout for embedding jobs (was missing unlike analyze jobs)
- Fix DropZone calling startAnalyze without setting backend URL first
- Add SSE reconnect with exponential backoff (3 retries) and Last-Event-ID
- Fix normalizeServerUrl to return base URL (no /api suffix) — clear contract
- Delete dead code: proxy.ts, server-graph-hydration.ts, pipeline.ts re-export barrel
- Update LoadingOverlay to import PipelineProgress directly from gitnexus-shared

* fix(server): fix repo lock key mismatch and embed cancel race

- Use getStoragePath(targetPath) as lock key in analyze handler to match
  embed handler's entry.storagePath — keys now always align
- Guard embed completion: don't overwrite 'failed' with 'complete' when
  job was cancelled while pipeline was still running
- Remove unused jobType parameter from acquireRepoLock
- Log backend.init() errors instead of silently swallowing

* fix: add gitnexus-shared as a local dependency in package-lock.json

* refactor: move language detection to gitnexus-shared, add syntax highlighting for all 15 languages

Move getLanguageFromFilename() from CLI to gitnexus-shared with COBOL
support added. Add getSyntaxLanguageFromFilename() for Prism-compatible
syntax highlighting covering all 15 code languages plus auxiliary
formats (json, yaml, markdown, html, css, bash, sql, xml).

Refactor CodeReferencesPanel to use shared function instead of a local
30-line switch. Delete dead gitnexus-web/src/config/supported-languages.ts
(web already imports SupportedLanguages from gitnexus-shared).

* feat(web): add first-time user onboarding with auto server detection

Replace the manual "Connect to Server" panel with an automatic onboarding
flow that guides first-time users through starting the GitNexus server.

Server detection:
- useBackend hook polls via setTimeout chain (3s, no overlap)
- Page Visibility API pauses polling when tab is hidden
- SSE heartbeat (/api/heartbeat) for instant disconnect detection

Onboarding UI (OnboardingGuide.tsx):
- Step-by-step flow: copy command → run → auto-connect
- Smart command: shows `gitnexus serve` in dev, `npx gitnexus@latest serve` in prod
- Node.js version auto-detected from package.json via Vite define
- Faux terminal windows with copy-to-clipboard, platform tabs, polling indicator

Transitions (DropZone.tsx):
- Crossfade wrapper with snapshot pattern for smooth phase transitions
- Three phases: onboarding → success (1.2s hold) → loading → graph
- Auto-recovery: falls back to onboarding if server dies or connect fails

Server changes:
- GET /api/heartbeat: SSE endpoint for liveness detection
- GET /api/info: version, launch context, Node.js version
- npm run serve script for local development
- app.disable('x-powered-by') hardening

* feat(web): add repo analysis UI, SSE heartbeat, and review fixes

Repo analysis:
- AnalyzeOnboarding: empty-state card when server has zero repos
- RepoAnalyzer: GitHub URL + Local Folder tabs with browse button
- Header repo dropdown: click project badge to switch repos or analyze new
- DropZone 'analyze' phase integrated into Crossfade transitions

Reliability fixes from 5-agent review:
- Polling: stop scheduling timers when tab hidden, restart on visibility return
- Heartbeat: exponential backoff (1s/2s/4s, 3 retries) prevents graph loss on blip
- RepoAnalyzer: completion timer tracked in ref, cleaned up on unmount
- DropZone: standardized card padding (p-7), heading sizes (text-lg)

Accessibility:
- prefers-reduced-motion global CSS rule (WCAG 2.3.3)
- focus-visible rings on CopyButton
- cursor-pointer on all Header buttons
- Consistent rounded-xl on all dropdowns

Cleanup:
- Deleted dead AnalyzeSheet.tsx (219 LOC) and BackendRepoSelector.tsx (89 LOC)
- Fixed AnalyzeProgress lucide import (lucide-react → @/lib/lucide-icons)

* fix(server): resolve analyze worker fork crash in dev mode

The forked analyze worker was crashing immediately with exit code 1
when running via `npm run serve` (tsx). Two issues:

1. Worker path resolved to `analyze-worker.js` but only `.ts` exists
   in the source directory — the `.js` file is only in `dist/`.

2. On Windows, bare `--import tsx` in execArgv fails because Node's
   ESM resolver for --import uses the child's CWD, not the parent's
   node_modules. Windows also rejects raw paths as `d:` is not a
   valid URL scheme.

Fix: detect dev vs prod via `import.meta.url` extension. In dev mode,
resolve `tsx/esm` to an absolute `file://` URL via `pathToFileURL()`
anchored to the parent's `createRequire` context. This works on all
platforms and doesn't depend on the child's CWD or PATH.

Also captures child stderr for better crash diagnostics.

Verified: `POST /api/analyze` with GitHub URL completes successfully
in dev mode (tsx) — status goes from cloning → analyzing → complete.

* fix(server): add worker auto-retry, error handling, and crash diagnostics

Worker resilience:
- Auto-retry up to 2 times with exponential backoff (1s, 2s) on crash
- SSE progress shows "Retrying after crash (1/2)..." during retry
- Captures child stderr for crash diagnostics in failure message
- AnalyzeJob tracks retryCount per job

Server error handling:
- app.listen wrapped in Promise so EADDRINUSE/EACCES propagate cleanly
- serve.ts catches startup errors with friendly messages and exit code 1
- EADDRINUSE gets actionable guidance (stop other process or --port flag)
- Global uncaughtException/unhandledRejection handlers prevent silent exits
- DEBUG=1 env var shows full stack traces

* feat: add e2e tests for onboarding flows, worker retry, and error handling

E2E tests (onboarding.spec.ts — 11 tests):
- Flow 1: OnboardingGuide shown when server unreachable (6 tests)
- Flow 2: Auto-connect with success card, analyze phase for zero repos
- Flow 3: Analyze form — GitHub URL validation, Local Folder tab, tab switching
- Flow 4: Repo dropdown in exploring view (skipped without live server)

Updated server-connect.spec.ts:
- Replaced manual Connect button flow with auto-connect waitForGraphLoaded

Server resilience:
- Worker auto-retry (2 attempts with exponential backoff) on crash
- Friendly error messages for serve startup failures (EADDRINUSE etc.)
- Global uncaughtException/unhandledRejection handlers prevent silent exits
- app.listen wrapped in Promise for proper error propagation

* refactor(shared): enforce exhaustive language coverage via Record types

Replace the if/else chain in getLanguageFromFilename with two exhaustive
Record<SupportedLanguages, ...> maps:

- EXTENSION_MAP: every language → its file extensions
- SYNTAX_MAP: every language → its Prism syntax identifier

Adding a new member to the SupportedLanguages enum without adding it to
both maps now produces a TypeScript compile error:

  Property '[SupportedLanguages.NewLang]' is missing in type...

This matches the existing pattern in languages/index.ts (providers table)
which already uses `satisfies Record<SupportedLanguages, LanguageProvider>`.

Three compile-time enforcement points now exist:
1. EXTENSION_MAP in language-detection.ts (file extensions)
2. SYNTAX_MAP in language-detection.ts (Prism syntax identifiers)
3. providers in languages/index.ts (LanguageProvider instances)

* feat(web): load source code from server and scroll to selected line

CodeReferencesPanel now fetches file content via GET /api/file when a
node is selected, instead of showing "Code not available in memory".

- Fetches via readFile() from backend-client when selectedFilePath changes
- Shows loading spinner while fetching
- After content loads, auto-scrolls to the selected node's startLine
- Highlights the selected line range with a cyan left border
- Cancels in-flight fetch if selection changes before it completes

Also: refactored language-detection.ts to use exhaustive Record types
(EXTENSION_MAP and SYNTAX_MAP) so adding a new SupportedLanguages enum
member without implementing extensions/syntax is a compile error.

* feat: buffered file reading for Code Inspector

Server: GET /api/file now supports ?startLine=N&endLine=M for reading
a line range instead of the entire file. Returns { content, startLine,
endLine, totalLines }.

Client: readFile() returns ReadFileResult with metadata. When selecting
a symbol (function, class, method), fetches only ±50 lines around the
symbol's startLine/endLine instead of the full file. File nodes still
fetch the entire file.

SyntaxHighlighter startingLineNumber set from the buffer offset so line
numbers are correct even for partial reads.

* fix: adapt readFile callers to new ReadFileResult return type

tools.ts: readFile comes from GraphRAGBackend interface which returns
Promise<string> (the adapter in useAppState extracts .content), so
revert the { content } destructuring back to plain string assignment.

useAppState.tsx: wrap backendReadFile with { repo } options object
and extract .content to satisfy the GraphRAGBackend interface.

* fix(web): ensure new repos appear in list immediately after analysis

Two fixes:

1. DropZone: handleAnalyzeComplete now passes the repoName through to
   connectToServer so the specific newly-analyzed repo loads — not the
   server's default first repo.

2. App.tsx: fetchRepos() is now awaited BEFORE handleServerConnect in
   both the DropZone and Header flows. This ensures the repo list is
   populated before the exploring view renders, so the new repo appears
   in the header dropdown immediately without a page reload.

* feat: delete repos, re-analyze with force, select after analysis

Server — DELETE /api/repo:
- Acquires repo lock first (409 if analyze/embed in flight)
- Closes LadybugDB, deletes index + clone dir, unregisters, re-inits
- Lock released in finally block

Server — analyze complete:
- backend.init() must succeed before SSE complete fires
- If backend.init() fails, job is marked failed (not complete)

Web — Header repo dropdown:
- Re-analyze: calls POST /api/analyze with force=true, shows spinning
  icon + inline progress bar via SSE
- Delete: acquires lock, aborts any running re-analysis SSE for same
  repo, refreshes list, switches to next repo
- After analysis completes: refreshes repo list, connects to the
  specific repo by name, loads graph, shows in explorer
- Retry with 1.5s backoff on 404 (server may still be reinitializing)

Type safety:
- err: any → err: unknown + instanceof BackendError in retry loop
- Added missing BackendRepo + BackendError imports in App.tsx
2026-03-28 14:07:11 +00:00
Juan Pablo Maddoni c27c9f612c fix/opencode mcp gitnexus timeout (#363) 2026-03-27 12:18:01 -05:00
Gergo Magyar 9edf06bda2 chore: bump version to 1.4.10, update CHANGELOG 2026-03-27 13:52:59 +00:00
Copilot 01ddc3e756 fix: resolve tree-sitter peer dependency conflicts (#538)
Downgrade tree-sitter from ^0.25.0 to ^0.21.1 and align all parser versions to eliminate ERESOLVE peer dependency conflicts that break MCP server install via npx. Also corrects hallucinated tree-sitter-dart SHA. Fixes #537
2026-03-27 13:49:28 +00:00
Gergo Magyar f373a09ebc chore: bump version to 1.4.9, add CHANGELOG.md 2026-03-26 16:40:02 +00:00
Oleg DeezusandGergo Magyar cde3858e03 refactor: Phase 8 & 9 — Field Types and Return-Type Binding (#494)
* feat(phase8): add field type data structures and extractor interface

* feat(phase8): implement TypeScript field extractor

* feat-phase9-add-call-result-binding

* test-phase8-add-field-extraction-unit-tests

* docs: update documentation for Phase 8 and Phase 9

* feat(swift): Phase 8/9 integration tests for field-type and call-result binding

Add Swift field-type resolution and call-result binding integration tests
with fixtures, plus merge-conflict fixes for the FieldExtractor code.

**Swift integration tests:**
- `swift-field-types/` fixture (Models.swift + App.swift) — tests
  HAS_PROPERTY edges, field-chain CALLS resolution (user.address.save()
  → Address#save), and ACCESSES edges for field reads.
- `swift-call-result-binding/` fixture — tests call-result binding
  (let user = getUser(); user.save() → User#save).
- 2 new describe blocks in swift.test.ts with skipIf(!swiftAvailable).

**Swift arity fix:**
- extractMethodSignature fallback counts direct `parameter` children
  when no wrapper list node exists (Swift's tree-sitter grammar places
  parameters as direct children of function_declaration).  Without this,
  all Swift functions had parameterCount: 0 and the arity filter rejected
  valid call targets.

**FieldExtractor merge-conflict fixes:**
- field-extractor.ts: update import from removed ./utils.js to
  ./utils/ast-helpers.js; use typeEnv.fileScope() instead of .get('').
- field-extractors/typescript.ts: same import fix.
- field-types.ts: alias TypeEnvironment as TypeEnv (renamed on main).
- field-extraction.test.ts: mock TypeEnvironment interface properly.

* feat(field-extractors): generic table-driven field extractors for all 14 languages, wired into pipeline

Implements field extractors for all supported languages and integrates
them into the ingestion pipeline as the single source of truth for
Property node metadata.

**Generic field extractor factory** — `field-extractors/generic.ts`
defines a `createFieldExtractor(config)` factory that generates
FieldExtractor instances from a per-language `FieldExtractionConfig`.
Each config specifies AST node types, name/type/visibility extraction
functions, and static/readonly detection — typically 20-40 lines per
language vs 300+ for a hand-written extractor.

**Per-language configs** — `field-extractors/configs/` has 11 config
files covering 13 languages (TS/JS share, Java/Kotlin share).
TypeScript keeps its hand-written extractor for richer handling.

**LanguageProvider integration** — New optional `fieldExtractor` property
on LanguageProviderConfig, set via defineLanguage() in each language
file. Follows the same strategy pattern as typeConfig, exportChecker,
and labelOverride.  Removed the separate FieldExtractorRegistry class
and field-extractors/index.ts — extractors are accessed via
getProvider(lang).fieldExtractor.

**Pipeline wiring** — Both parse-worker.ts (worker pool) and
parsing-processor.ts (sequential fallback) now call the FieldExtractor
during Property node creation. Results are cached per class node.
Property nodes are enriched with: declaredType, visibility, isStatic,
isReadonly.

**extractPropertyDeclaredType removed** — The 100-line multi-strategy
function in type-extractors/shared.ts is replaced by the FieldExtractor.
All 14 languages register an extractor, eliminating the need for a
generic fallback.  The Python config's extractType was fixed to handle
annotation-without-value patterns (address: Address).

**Integration tests** — Each language's resolver test file gains
pipeline-based assertions verifying visibility/isStatic/isReadonly on
Property nodes via getNodesByLabelFull.  Tests run through
runPipelineFromRepo with real fixtures — no direct extractor calls.

* fix(type-env): thread enclosingFunctionFinder through scope resolution, unskip Dart ACCESSES test

The type-env's findEnclosingScopeKey had the same Dart sibling problem
as findEnclosingFunction — it walked parents but never found
function_signature because the call lives inside function_body (a
sibling).  Instead of hardcoding a function_body check, thread the
provider's enclosingFunctionFinder hook through BuildTypeEnvOptions →
lookupInEnv → findEnclosingScopeKey.  All three buildTypeEnv call sites
(call-processor, parsing-processor, parse-worker) now pass the hook.

This enables the type-env to resolve scoped parameter bindings for Dart
(e.g., `user: User` in processUser), which lets the chain-resolution
tier (Step 1c) walk `user.address` and emit ACCESSES edges.

Dart integration test unskipped — 10/10 passing including ACCESSES.
Reverted CHANGELOG.md to origin/main.

* fix: resolve all PR #494 review findings (10 items)

CRITICAL:
- parse-worker.ts: classNode: any → SyntaxNode on getFieldInfo
  and findEnclosingClassNode; removed redundant as number casts
- parsing-processor.ts: classNode: any → SyntaxNode on seqGetFieldInfo

HIGH:
- ruby.ts: attr_accessor now extracts ALL symbol arguments via
  extractNames hook in generic factory (was firstNamedChild only)
- typescript.ts: added JSDoc explaining why hand-written extractor
  coexists with config-based typescript-javascript.ts

MEDIUM:
- field-types.ts: FieldVisibility union type replaces string
  ('public'|'private'|'protected'|'internal'|'package'|'fileprivate'|'open')
  Propagated through field-extractor.ts, generic.ts, all 7 config files
- typescript.ts: extractFullType collapsed from 12 branches to 3 lines
- generic.ts: added extractNames? optional hook + buildField refactor

LOW:
- ruby.ts: extractVisibility(node) → extractVisibility(_node)
- python.ts: fixed misleading isStatic comment

TypeScript compiles cleanly.

* test: add 24 field extraction tests for generic factory + 5 languages

Generic factory (4 tests):
- createFieldExtractor with TypeScript config validates factory itself
- Body discovery for interfaces, static/readonly modifiers
- Non-type node rejection

Python (4 tests):
- Annotated class field extraction
- Underscore-based visibility: _name=protected, __name=private

Go (5 tests):
- isTypeDeclaration on type_declaration nodes
- Config functions: uppercase=public, lowercase=package visibility
- extractType, isStatic, isReadonly

C++ (5 tests):
- public/private/protected access specifier backward-sibling walk
- Default visibility: class=private, struct=public
- static/const modifier detection

Ruby (6 tests):
- attr_accessor multi-symbol: :name, :email, :age → 3 fields
- attr_reader=readonly, attr_writer=non-readonly
- Multiple attr_* calls in one class

Total: 46 tests passing

* chore: remove plan doc from PR

---------

Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-03-26 16:37:34 +00:00
Gergő Magyar d2cd0b676f feat: add COBOL language support with regex extraction pipeline (#498)
* feat: add COBOL language support with regex extraction pipeline

Standalone COBOL processor following the markdown-processor.ts pattern:
- No LanguageProvider modification — COBOL uses regex, not tree-sitter
- No SupportedLanguages enum change — standalone processor pattern

New files:
- cobol-processor.ts — orchestrator (processCobol, isCobolFile, isJclFile)
- cobol/cobol-preprocessor.ts — regex state machine extraction (~888 LOC)
- cobol/cobol-copy-expander.ts — COPY statement expansion with circular detection
- cobol/jcl-parser.ts — JCL job/step/DD extraction
- cobol/jcl-processor.ts — JCL graph node creation

Extraction produces:
- Module nodes (PROGRAM-ID)
- Function nodes (paragraphs)
- Namespace nodes (sections)
- Property nodes (data items)
- CALLS edges (PERFORM intra-file, CALL cross-program)
- IMPORTS edges (COPY statements)
- CONTAINS edges (section → paragraph hierarchy)

Pipeline integration: single processCobol() call in Phase 2.6

54 new tests (33 COBOL + 21 JCL), all 3889 tests pass.

* docs: document custom processor pattern in pipeline.ts

Add comment block at the custom processor integration point
documenting the pattern for future non-tree-sitter language additions.

* feat(cobol): enrich graph with EXEC SQL/CICS, ENTRY points, MOVE data flow, PERFORM THRU

Maps the remaining 60% of CobolRegexResults to the graph:
- EXEC SQL blocks → CodeElement nodes + ACCESSES edges to DB tables
- EXEC CICS LINK/XCTL → CodeElement nodes + cross-program CALLS edges
- ENTRY points → Constructor nodes (registered for cross-program resolution)
- MOVE statements → ACCESSES edges (read/write data flow tracking)
- PERFORM THRU → expanded CALLS edges for range targets
- File declarations → Record nodes with assignment metadata
- Cross-program CALL 2nd pass: resolves unresolved targets after all programs processed

* test(cobol): add 26 integration tests with exact assertions + fix CICS resolution bug

Integration tests (test/integration/resolvers/cobol.test.ts):
- 26 tests covering full COBOL system extraction
- ALL assertions use exact toBe(N) — zero fuzzy assertions
- Fixtures: CUSTUPDT.cbl, AUDITLOG.cbl, CUSTDAT.cpy, RPTGEN.cbl, RUNJOBS.jcl

Bug fix (cobol-processor.ts):
- CICS LINK/XCTL cross-program resolution was broken — edges were
  created with "resolved" reason but pointing to <unresolved> targets
- Fix: use cics-link-unresolved / cics-xctl-unresolved suffix pattern
  matching the existing cobol-call-unresolved pattern
- Second-pass resolver now patches both CALL and CICS unresolved edges

All 3915 tests pass, 0 failures.

* test(cobol): exhaustive 57-test suite with strict exact assertions

Complete rewrite of COBOL integration tests using ground-truth approach:
dump the full graph, then assert EVERY node and EVERY edge.

57 tests across 9 sections:
- Node completeness: Module(3), Function(13), Namespace(2), Property(21),
  Record(1), CodeElement(8), Constructor(1) — exact sorted arrays
- Edge completeness: 22 tests covering every type+reason combination
  with exact source→target pairs
- Cross-program resolution: 6 tests verifying CALL, CICS LINK/XCTL, JCL
- COPY expansion: copybook data items in RPTGEN
- Section hierarchy: exact paragraph membership per section
- Data item ownership: exact per-module breakdown
- MOVE data flow: exact read/write pairs
- JCL integration: job/step/dataset containment
- Grand totals: CALLS(22), CONTAINS(48), IMPORTS(1), ACCESSES(7)

Fixture enhancements:
- CUSTUPDT.cbl: added INIT-SECTION + PROCESSING-SECTION, PERFORM THRU
- AUDITLOG.cbl: added ENTRY "AUDITLOG-BATCH"
- RPTGEN.cbl: added EXEC CICS XCTL

Zero fuzzy assertions — every expect uses toBe(N) or toEqual([...sorted]).

* fix(cobol): add removeRelationship API + single-quote CALL/COPY/ENTRY, PERFORM keyword skip

Phase 0A: Add removeRelationship(id) to KnowledgeGraph interface and
implementation (trivial Map.delete wrapper). Required for orphan edge
cleanup in next commit.

Phase 1A (from PR #500 review, modified):
- RE_CALL and RE_COPY_QUOTED now match both "double" and 'single' quotes
- parseSingleCopyStatement in copy-expander updated for single quotes
- PERFORM_KEYWORD_SKIP set prevents UNTIL/VARYING/WITH/TEST/FOREVER
  from being stored as false-positive perform targets
- Sequence number stripping uses /[^0-9 ]/ (preserves numeric seq numbers
  unlike PR #500's /\S/ which stripped them)
- Normalized || to ?? for regex group extraction in copy-expander

5 new graph unit tests, all 57 COBOL integration tests pass.

* fix(cobol): RE_ENTRY single-quote + remove orphan unresolved CALLS edges

Phase 1B: RE_ENTRY regex now supports both "double" and 'single' quoted
ENTRY targets. Uses named intermediates (entryName, usingClause) with ??
operator. USING capture group shifted from [2] to [3].

Phase 1C: Second-pass resolution now collects resolved orphan edge IDs
during iteration and removes them after the loop completes, using the new
graph.removeRelationship() API. Graph no longer contains phantom
<unresolved>: edges alongside their resolved replacements. CALLS count
drops from 22 to 18 (4 orphan edges removed).

* fix(cobol): Property ID collisions + O(1) Map lookup for MOVE edges

Phase 1D+3C (atomic): Property node IDs now use composite key
filePath:section:level:name instead of filePath:name. This prevents
duplicate data item names in different sections (e.g., STATUS in both
WORKING-STORAGE and LINKAGE) from silently colliding.

New generatePropertyId() helper ensures both node creation and MOVE
edge lookup use the identical key formula. buildDataItemMap() replaces
the O(n) findDataItemNode linear scan with O(1) Map lookup, built once
per file before MOVE processing.

* feat(cobol): MOVE multi-target extraction with OF/IN qualifier filtering

MOVE X TO A B C now produces write edges for all targets, not just the
first. extractMoveTargets() helper handles OF/IN qualified names
(WS-NAME OF WS-RECORD -> target is WS-NAME), subscript stripping
(WS-TABLE(I) -> WS-TABLE), and MOVE_SKIP filtering on targets.

Data model: CobolRegexResults.moves.to:string -> targets:string[]
MOVE CORRESPONDING stays single-target per COBOL standard.
Processor MOVE loop now iterates move.targets.

* feat(cobol): COPY IN/OF library, pseudotext REPLACING, dynamic CALL, PERFORM TIMES, CICS MAP unquoted

Phase 2B: COPY ... IN/OF library-name now captured as metadata in
CopyResolution (IN and OF are synonyms per COBOL-85 standard).

Phase 2C: COPY REPLACING ==pseudotext== support. Tokenizer handles
==...== delimiters alongside "quoted" strings. Pseudotext forces EXACT
type. Two-pass applyReplacing: first pass handles space-containing/
non-identifier pseudotext via global string replace; second pass handles
identifier-level LEADING/TRAILING/EXACT. New test file
cobol-copy-expander.test.ts with 10 tests.

Phase 2E: PERFORM WS-COUNT TIMES no longer produces a false-positive
perform target (checks for TIMES keyword after captured identifier).

Phase 2F: Dynamic CALL via data item (CALL WS-PROG-NAME without quotes)
now emits a CodeElement annotation node with description 'dynamic-call'
instead of silently ignoring. Adds isQuoted:boolean to call results.

Phase 3A: CICS MAP(WS-MAP-NAME) unquoted identifiers now captured.
Phase 3B: Normalized || to ?? in copy-expander (done in Phase 1A).

* feat(cobol): nested program support — capture multiple PROGRAM-IDs per file

Phase 2D: The state machine now captures all PROGRAM-IDs, not just the
first. The primary program name stays in programName; additional nested
programs go into nestedPrograms[]. The processor creates separate Module
nodes for each nested program, contained by the outer module, and
registers them in moduleNodeIds for cross-program CALL resolution.

Paragraphs/data items are not yet scoped per-program (attributed to the
outer module) — full per-program scoping is a future enhancement that
requires END PROGRAM boundary tracking in the state machine.

* test(cobol): expand integration tests for all new language features

New fixtures:
- NESTED.cbl — two PROGRAM-IDs (OUTER-PROG, INNER-PROG) for nested
  program support testing
- COPYLIB.cpy — copybook for pseudotext REPLACING test target

Modified fixtures:
- CUSTUPDT.cbl — single-quoted ENTRY 'ALTENTRY', multi-target MOVE
  (WS-AMT TO FIELD-A FIELD-B), dynamic CALL WS-PROG-NAME, COPY COPYLIB
  with pseudotext REPLACING, LINKAGE SECTION with LS-PARAM
- RPTGEN.cbl — PERFORM WS-COUNT TIMES (false-positive guard), unquoted
  MAP(WS-MAP-NAME), additional data items WS-COUNT WS-MAP-NAME

Integration test rewritten with 62 exact assertions covering:
- 5 Module, 17 Function, 33 Property, 9 CodeElement, 2 Constructor nodes
- Nested program containment (OUTER-PROG -> INNER-PROG)
- Dynamic CALL annotation (CodeElement with cobol-dynamic-call)
- Multi-target MOVE (UPDATE-BALANCE: 2 reads, 3 writes)
- Single-quoted ENTRY (ALTENTRY under CUSTUPDT)
- PERFORM TIMES guard (WS-COUNT not in CALLS)
- Orphan unresolved edge removal (zero -unresolved edges)
- Grand totals: 21 CALLS, 68 CONTAINS, 2 IMPORTS, 10 ACCESSES

* fix(cobol): pseudotext REPLACING now applies correctly via isPseudotext flag

Root cause: ==PREFIX-== matched /^[A-Z][A-Z0-9-]*$/i (trailing hyphens
allowed), routing it to the second-pass EXACT identifier match where
PREFIX-RECORD !== PREFIX- failed silently.

Fix: Propagate isPseudotext from parseReplacingClause to CopyReplacing
interface, then use it in applyReplacing first-pass condition to force
global string replacement for all pseudotext entries regardless of
whether the content looks like an identifier.

Result: COPY COPYLIB REPLACING ==PREFIX-== BY ==WS-==. now correctly
transforms PREFIX-RECORD → WS-RECORD, PREFIX-CODE → WS-CODE, etc.

* refactor(cobol): per-program scoping via boundary tracking + line-range grouping

State machine changes (minimal, ~30 lines):
- Add RE_END_PROGRAM regex for END PROGRAM program-name. detection
- Replace nestedPrograms[] with programs[] containing startLine/endLine/
  nestingDepth metadata for each PROGRAM-ID in the file
- Reset division/section/paragraph state on new PROGRAM-ID boundary
- EOF finalization flushes remaining stack entries (single-program files)
- Programs sorted by startLine (outer before inner)

Processor changes:
- Uses programs[] with line-range containment to find enclosing parent
  Module for nested programs (replaces hardcoded nestedParent logic)
- programModuleIds Map tracks Module node IDs per program name

Fixture: NESTED.cbl now includes END PROGRAM lines for both programs.

Integration test: PREFIX-* Property nodes now correctly appear as WS-*
after the pseudotext REPLACING fix from the previous commit.

* feat(cobol): free-format COBOL support (>>source free)

Auto-detects >>SOURCE FREE directive in the first 500 chars and switches
to free-format line processing:
- No column-position rules (cols 1-6 are program text, not sequence area)
- Comments use *> prefix instead of col 7 indicator
- No continuation line indicator
- Strip inline *> comments
- Skip >>SOURCE directive lines

preprocessCobolSource() skips col-1-6 stripping for free-format files.

Paragraph/section regexes relaxed from fixed 7-space prefix to flexible
whitespace with case-insensitivity (/^\s*([A-Z][A-Z0-9-]+)\.\s*$/i).
EXCLUDED_PARA_NAMES expanded with COBOL verbs (GOBACK, END-READ, etc.)
to prevent false-positive paragraph detection in free-format.

Also fixes: entry-point-scoring.ts crash when language is 'cobol'
(MERGED_ENTRY_POINT_PATTERNS[language] was undefined → optional chaining).

Benchmark on ACAS 3.01 (268 GnuCOBOL free-format programs, 10MB):
- Before: 407 nodes, 393 edges (near-empty, only file nodes)
- After:  4,297 nodes, 3,612 edges, 542 clusters, 11 flows

* fix(cobol): relax data item regexes for free-format (^\s+ to ^\s*)

RE_FD, RE_DATA_ITEM, RE_ANONYMOUS_REDEFINES, and RE_88_LEVEL all used
^\s+ which requires at least 1 leading space. In free-format mode, lines
are trimmed before processing, so data items like "01 WS-FIELD PIC X."
have no leading whitespace after trimming.

Changed to ^\s* (zero or more spaces) which works for both fixed-format
(indented lines still have spaces) and free-format (trimmed lines).

ACAS benchmark (268 GnuCOBOL programs):
- Before: 4,297 nodes, 3,612 edges (paragraphs only)
- After:  13,832 nodes, 8,615 edges (+ data items, FDs, 88-levels)

* feat(cobol): 100% structural feature coverage — GO TO, SCREEN, SD/RD, SORT, SEARCH, CANCEL, Level 66

New extractions: GO TO (CALLS edges), SCREEN SECTION data items,
SD/RD alongside FD (Record nodes), SORT/MERGE USING/GIVING (ACCESSES),
SEARCH (ACCESSES), CANCEL (CALLS), Level 66 RENAMES (Property),
IS EXTERNAL/IS GLOBAL (Property description enrichment).

ACAS: 13,951 nodes | 13,193 edges | 685 clusters | 150 flows
(+53% edges from new GO TO/SORT/SEARCH/CANCEL extractions)

* feat(cobol): enriched CICS extraction — file I/O, dynamic PROGRAM, queues, HANDLE ABEND

EXEC CICS blocks now extract:
- FILE/DATASET clause: captures VSAM file name (literal or data item ref)
  for READ/WRITE/REWRITE/DELETE/STARTBR/READNEXT/READPREV → ACCESSES edges
- PROGRAM clause: now handles unquoted variable references (dynamic CICS
  program transfer) → CodeElement annotation with cics-dynamic-program reason
- QUEUE clause: captures TS/TD queue names from WRITEQ/READQ → ACCESSES edges
- LABEL clause: captures HANDLE ABEND error handler targets → CALLS edges
- TRANSID: now handles unquoted variable references

CodeElement descriptions enriched with all captured fields (map, program,
transid, file, queue, label).

CardDemo benchmark: +49 nodes, +33 edges from enriched CICS extraction.

* feat(cobol): complete CICS command extraction — all 7 expert recommendations

From COBOL expert agent analysis:
1. ENDBR added to isRead file command list
2. LOAD added to PROGRAM edge commands (alongside LINK/XCTL)
3. Two-word commands expanded: WRITEQ/READQ/DELETEQ TS/TD, HANDLE
   ABEND/AID/CONDITION, START TRANSID
4. Queue reason differentiated: cics-queue-read/-write/-delete
5. RETURN/START TRANSID → CALLS edges to synthetic <transid> target
6. MAP → ACCESSES edges for screen traceability
7. INTO/FROM data fields extracted → ACCESSES edges to data items

Also: dataItemMap built before CICS block processing (was declared after),
CodeElement descriptions enriched with all captured CICS fields.

* test(cobol): strict exhaustive integration tests with exact edgeSet assertions

Every edge reason has exact sorted pair assertions via edgeSet(), not
just counts. Any change to extraction that adds, removes, or reorders
edges will produce a precise, descriptive failure.

Updated RPTGEN.cbl fixture with:
- GO TO EXIT-PARAGRAPH, SORT USING/GIVING, SEARCH table
- EXEC CICS READ FILE INTO, WRITEQ TS QUEUE FROM, SEND MAP FROM
- EXEC CICS HANDLE ABEND LABEL, RETURN TRANSID, XCTL PROGRAM(variable)
- ABEND-HANDLER and EXIT-PARAGRAPH paragraphs

46 tests covering 24 CALLS + 79 CONTAINS + 18 ACCESSES + 2 IMPORTS edges
across 15 distinct edge reason codes, all with exact sorted pair lists.

* fix(cobol): address 5 findings from second Claude review (compiler front-end perspective)

Finding #2: Numeric sequence numbers now stripped (changed /[^0-9 ]/ to
/\S/ in preprocessCobolSource). Lines like "000100 MAIN-PARAGRAPH." now
have cols 1-6 blanked so paragraph regex matches correctly.

Finding #11: JCL in-stream PROC ordering fixed — pre-register all PROCs
into moduleNames before step processing. Steps that EXEC a PROC defined
later in the same file now get CALLS edges.

Finding #A: PROCEDURE DIVISION USING no longer captures calling-convention
keywords (BY, VALUE, REFERENCE, CONTENT, ADDRESS, OF) as parameter names.

Finding #C: SORT/MERGE USING/GIVING now captures ALL file references
(multi-file), not just the first. Changed from single-match to section
extraction with split.

Finding #D: Section headers no longer set currentParagraph, preventing
PERFORM caller misattribution to Namespace instead of Function nodes.

* fix(cobol): address code review findings — ReDoS fix, perf, cleanup

P1 CRITICAL — ReDoS in SORT USING/GIVING:
Replaced nested-quantifier regex with safe indexOf+substring+split
approach. No backtracking possible on crafted input.

P2 — readCopy O(M) linear scan:
Added copybookByPath reverse Map for O(1) path-to-content lookup.

P3 — Dead code removal:
Deleted unused RE_SORT_USING and RE_SORT_GIVING constants.

P3 — EXCLUDED_PARA_NAMES simplification:
Replaced 20 END-* entries with startsWith('END-') prefix check.
Auto-covers future END-* verbs.

P3 — Misplaced JSDoc on removeRelationship:
Fixed comment that described removeNodesByFile instead.
Added missing JSDoc to removeNodesByFile.

Review agents: architecture-strategist, performance-oracle,
security-sentinel, code-simplicity-reviewer.

* refactor: add Cobol to SupportedLanguages with parseStrategy: standalone

New languages/cobol.ts — standalone regex processor provider with no-op
tree-sitter fields. Declares parseStrategy: 'standalone' to distinguish
from tree-sitter-based languages.

Added parseStrategy: 'tree-sitter' | 'standalone' to LanguageProviderConfig
for languages that use their own processor instead of tree-sitter.

Removed all 11 'cobol' as any casts — now uses SupportedLanguages.Cobol.
Added empty Cobol entries to entry-point-scoring and framework-detection.

* fix(cobol): 5 fixes from third Claude review + 3 regression tests

Fixes:
- Line numbers now 1-indexed in fixed-format (was 0-indexed, off-by-one
  in jump-to-definition links)
- Copybook content preprocessed before COPY expansion (sequence numbers
  and patch markers in copybooks no longer survive into expanded source)
- ENTRY USING filters calling-convention keywords (BY, VALUE, REFERENCE,
  CONTENT, ADDRESS, OF) — same fix as PROCEDURE DIVISION USING
- SORT/MERGE trailing period stripped from USING/GIVING file tokens
- Paragraph exclusion uses exact match for SECTION/DIVISION (was substring
  match that excluded valid names like CROSS-SECTION-ANALYSIS)

USING_KEYWORDS moved to module scope for reuse by both PROCEDURE DIVISION
USING and ENTRY USING handlers.

New unit tests:
- ENTRY USING BY VALUE filtering
- Paragraph names containing SECTION not excluded
- Numeric sequence numbers stripped enabling paragraph detection

* fix(cobol): address 6 findings from fourth Claude review + tests

Fourth review findings fixed:
- New #IV: PERFORM TIMES guard uses perfMatch.index instead of
  line.indexOf (prevents wrong match when target appears earlier in line)
- New #V: 88-level condition values now handle single-quoted literals
  ('Y' no longer stored with embedded quotes)
- New #I: CANCEL edges use two-pass resolution like CALL (no longer
  silently dropped when target indexed after source)
- New #3: Multi-line SORT/MERGE accumulation — sortAccum state variable
  accumulates lines until period, then extracts USING/GIVING from full
  statement (95% of production SORT statements span multiple lines)
- New #II: PROCEDURE DIVISION USING on split lines — pendingProcUsing
  flag defers parameter capture to next line if USING not on same line
- New #6 (prior): EXCLUDED_PARA_NAMES exact match for SECTION/DIVISION

Updated fixture: RPTGEN.cbl SORT now uses multi-line format with GIVING
on separate line (period-terminated). New sort-giving integration test.
ACCESSES total: 18 → 19 (new sort-giving edge from multi-line capture).

* fix(cobol): address 4 findings from fifth Claude review

Finding #B (5 reviews old): Section/paragraph node IDs now include
enclosing program name to prevent collision when nested programs share
section/paragraph names. New findOwningProgramName() helper uses
programs[] line ranges to find the innermost enclosing program.

Finding #α: pendingProcUsing now reset in the if(procUsingMatch) branch
(was only set in else branch, could leak across nested programs).

Finding #β: RE_CALL_DYNAMIC uses negative lookbehind (?<![A-Z0-9-]) to
prevent false-positive on compound identifiers like WS-CALL OCCURS.

Finding #γ: sortAccum flushed at EOF (parallel to flushSelect and
pendingFdName EOF cleanup). Prevents silent loss of SORT USING/GIVING
relationships in truncated files.

* fix(cobol): address findings from reviews 5+6 with full test coverage

Review 5 fixes:
- #α: pendingProcUsing reset in if(procUsingMatch) branch
- #β: RE_CALL_DYNAMIC negative lookbehind prevents WS-CALL false positive
- #γ: sortAccum flushed at EOF for truncated files
- #B: Section/paragraph IDs include owning program name

Review 6 fixes:
- #P: sectionNodeIds/paraNodeIds maps use program-scoped keys
  (PROGNAME:NAME). New scopedParaLookup/scopedCallerLookup helpers.
  findContainingSection updated with programs parameter.
- #Q: RETURNING added to USING_KEYWORDS for COBOL 2002+
- #R: RE_PERFORM matches both THRU and THROUGH via alternation

New unit tests (6):
- PERFORM THROUGH captures thruTarget
- PROCEDURE DIVISION USING RETURNING filters keyword
- RE_CALL_DYNAMIC no false-match on WS-CALL compound identifier
- Multi-line SORT captures USING/GIVING from continuation lines
- PROCEDURE DIVISION USING on split line via pendingProcUsing
- Copybook preprocessing strips sequence numbers

* fix(cobol): address findings from seventh Claude review + 3 tests

Review 7 fixes:
- #i: findContainingSection only updates best when lookup succeeds
  (prevents undefined overwriting valid parent section)
- #ii: RE_PROC_SECTION handles segment numbers (SECTION 30.)
- #III: procedureUsing now stored per-program on boundary stack
  entries, propagated to programs[] output. Inner programs no longer
  overwrite outer program's parameters.
- #δ: Dynamic CANCEL (CANCEL variable) now creates CodeElement
  annotation node, matching dynamic CALL behavior. RE_CANCEL_DYNAMIC
  with negative lookbehind. cancels[] gains isQuoted field.
- #Q: RETURNING added to USING_KEYWORDS (already in prev commit)
- #R: PERFORM THROUGH already fixed (THRU|THROUGH alternation)

New unit tests:
- Nested programs carry per-program procedureUsing
- SECTION with segment number detected
- Dynamic CANCEL via data item captured with isQuoted=false

* feat(cobol): link PROCEDURE DIVISION USING to LINKAGE data items + close 4 findings

Finding #10 FIXED: procedureUsing parameters now create ACCESSES edges
with reason 'cobol-procedure-using' from Module to matching LINKAGE
SECTION Property nodes. This exposes the program's parameter contract
in the graph (e.g., AUDITLOG → LS-CUST-ID, AUDITLOG → LS-AMOUNT).

Findings closed by expert agent consensus:
- #6 COPY IN library: WONTFIX — captured metadata, no universal
  library-to-directory mapping exists. Field costs nothing and is useful
  for library queries.
- #14 SQL DELETE: WONTFIX — DB2 requires FROM; existing FROM pattern
  handles it. Bare DELETE would risk false positives.
- #E OCCURS DEPENDING ON: WONTFIX — runtime sizing concern, not
  structural. The static occurs count is sufficient for indexing.

All 39 findings from 7 Claude reviews now resolved or closed.

* fix(cobol): resolve 48 review findings across 9 review cycles

Ninth deep review resolved all remaining COBOL parser gaps identified
by 5 specialist agents (COBOL expert, architecture strategist,
TypeScript reviewer, security sentinel, code simplicity reviewer).

Fixes (P1 — critical):
- SELECT OPTIONAL now correctly skips OPTIONAL keyword (C1)
- RETURNING params excluded from PROCEDURE DIVISION USING list (C7)
- SORT GIVING no longer captures clause keywords as file names (C5)
- Extract flushSort() helper eliminating 40-line duplication (S2)
- Flush unclosed EXEC blocks at EOF matching SORT/SELECT pattern (S3)
- Guard undefined map key in jcl-processor moduleNames (S1)
- Add MAX_TOTAL_EXPANSIONS=500 to prevent exponential COPY breadth (S4)

Fixes (P2 — important):
- Quote-aware stripInlineComment for | and *> in string literals (C2+C3)
- Fixed-format literal continuation now handles quoted strings (C6)
- PROGRAM-ID detected regardless of division state for siblings (C9)

Fixes (P3 — cleanup):
- EXEC SQL INTO restricted to INSERT INTO to avoid FETCH false-pos (C8)
- Copy expander line numbers fixed from 0-based to 1-based (C11)
- Remove dead code: inInStreamProc, fileIsLiteral, expansionDepth (S7-S10)

Also fixes 8th-review findings: nested program CONTAINS attribution,
multi-PERFORM on same line, INPUT/OUTPUT PROCEDURE IS in SORT,
GO TO DEPENDING ON multi-target, MOVE CORR abbreviation, per-program
procedureUsing ACCESSES edges.

Tests: 145 COBOL tests passing (59 integration + 86 unit)
Benchmarks: CardDemo 12,323 nodes/8,893 edges (7.4s)
            ACAS 14,016 nodes/15,452 edges (9.3s, -9% faster)

* docs(cobol): update documentation for ninth review cycle fixes

Update all 4 COBOL documentation files to reflect the 16 fixes
from the ninth review cycle:

- regex-extraction.md: quote-aware comment stripping, SELECT OPTIONAL,
  RETURNING exclusion, SORT_CLAUSE_NOISE filter, flushSort() helper,
  GO TO multi-target, PROGRAM-ID division-independent detection
- copy-expansion.md: MAX_TOTAL_EXPANSIONS=500 breadth guard, 1-based
  line numbers, removed expansionDepth/warnedCircular param
- deep-indexing.md: GO TO DEPENDING ON, INPUT/OUTPUT PROCEDURE IS,
  MOVE CORR edge reasons, INSERT INTO restriction, literal continuation
- performance.md: updated benchmarks (CardDemo 12,323n/8,893e/7.4s,
  ACAS 14,016n/15,452e/9.3s), COPY breadth guard

* fix(cobol): resolve 10th review findings — nested program edge attribution

Fix 6 findings from the 10th review (PR #498 comment #4132201110):

#A+#F: All CALL/CANCEL/CICS/ENTRY/SQL/SEARCH/file-declaration edges
now use owningModuleId() for nested program attribution instead of
the outer program's parentId. Added helper function owningModuleId()
to centralize the pattern.

#B: Added USING and GIVING to SORT_CLAUSE_NOISE set to prevent MERGE
USING + OUTPUT PROCEDURE from capturing clause keywords as file names.

#C: INPUT/OUTPUT PROCEDURE regex now captures optional THRU/THROUGH
range end paragraph, mirroring RE_PERFORM's THRU support.

#D: scopedCallerLookup fallback now uses programModuleIds.get(pgm)
instead of parentId, so PERFORM/MOVE/GOTO in nested programs with
unresolvable paragraphs fall back to the correct inner module.

#E: pendingProcUsing only set when PROCEDURE DIVISION line is NOT
period-terminated, preventing false USING expectation.

Tests: 145 passing | TypeScript clean

* fix(cobol): resolve 10th review findings — nested program edge attribution

Fix 6 findings from the 10th review (PR #498 comment #4132201110):

#A+#F: All CALL/CANCEL/CICS/ENTRY/SQL/SEARCH/file-declaration edges
now use owningModuleId() for nested program attribution instead of
the outer program's parentId. Added helper function owningModuleId()
to centralize the pattern.

#B: Added USING and GIVING to SORT_CLAUSE_NOISE set to prevent MERGE
USING + OUTPUT PROCEDURE from capturing clause keywords as file names.

#C: INPUT/OUTPUT PROCEDURE regex now captures optional THRU/THROUGH
range end paragraph, mirroring RE_PERFORM's THRU support.

#D: scopedCallerLookup fallback now uses programModuleIds.get(pgm)
instead of parentId, so PERFORM/MOVE/GOTO in nested programs with
unresolvable paragraphs fall back to the correct inner module.

#E: pendingProcUsing only set when PROCEDURE DIVISION line is NOT
period-terminated, preventing false USING expectation.

Tests: 145 passing | TypeScript clean

* fix(cobol): resolve 11th review findings — final nested program + multi-CALL gaps

#1: scopedCallerLookup(null) now uses owningModuleId(lineNum) instead
of parentId, fixing PERFORM/MOVE/GOTO before first paragraph in nested
programs.

#2+#3: CALL and CANCEL extraction now uses matchAll (global flag) to
capture multiple occurrences on the same line. Dynamic CALL/CANCEL
checked independently instead of in else branch.

#4: SORT/MERGE ACCESSES edge IDs now use owningModuleId(sort.line)
instead of parentId for nested program correctness.

#5: preprocessCobolSource free-format detection now uses first 10 lines
(consistent with extractCobolSymbolsWithRegex threshold).

#6: EXCLUDED_PARA_NAMES expanded with DISPLAY, ACCEPT, WRITE, READ,
REWRITE, DELETE, OPEN, CLOSE, RETURN, RELEASE, SORT, MERGE to prevent
false-positive paragraph detection on isolated verbs.

Also removed unused GraphNode import from cobol-processor.ts.

Tests: 145 passing | TypeScript clean

* docs(cobol): deepened full language coverage plan with research findings

3 research agents analyzed Phase 1-2 features and graph value ranking.

Key findings: cobol-call-using is #1 edge type (9.2/10); multi-line
accumulation is dominant challenge; DECLARATIVES is lowest-risk Phase 2
item; SET TO TRUE covers 80-90% of SET usage.

* feat(cobol): implement Phase 1 — high-value data flow edges

4 new extraction features that create new ACCESSES and IMPORTS edges:

1.1: EXEC SQL INCLUDE -> IMPORTS edges with reason 'sql-include'
     Handles unquoted (SQLCA), quoted ('DBRMLIB.MEMBER'), and
     underscored (CUST_TBL_DCL) member names.

1.2: CALL USING parameter extraction -> ACCESSES edges
     Extracts parameters from CALL USING clause, filtering BY/REFERENCE/
     CONTENT/VALUE/ADDRESS/OF/LENGTH/OMITTED keywords. Creates
     'cobol-call-using' ACCESSES edges (graph value: 9.2/10).

1.4: OCCURS DEPENDING ON -> ACCESSES edges with reason 'cobol-depends-on'
     Extended OCCURS regex captures DEPENDING ON field with subscript
     stripping. Creates dependency edge from table to controlling field.

1.5: VALUE clause for standard data items
     Extracts VALUE from data item clauses: quoted strings with type
     prefix (X/N/G/B), ALL literals, numerics (incl negative/decimal),
     and figurative constants. Populates Property node values.

Tests: 145 passing (+2 ACCESSES from CALL USING) | TypeScript clean

* feat(cobol): implement Phase 2 — DECLARATIVES, SET, INSPECT, EXEC DLI

4 new extraction features for error handling, data flow, and IMS/DB:

2.1: EXEC DLI (IMS/DB) -> CodeElement + ACCESSES edges
     Accumulates EXEC DLI blocks like EXEC SQL. Parses DLI verbs
     (GU, GN, ISRT, REPL, DLET, CHKP, SCHD, TERM). Extracts
     SEGMENT, PCB, INTO/FROM, PSB. Creates dli-{verb} ACCESSES
     edges to <ims>:segment Record nodes.

2.2: DECLARATIVES / USE AFTER EXCEPTION -> ACCESSES edges
     Tracks inDeclaratives state. Detects USE AFTER STANDARD
     EXCEPTION ON file-name. Creates cobol-error-handler ACCESSES
     edge from handler section to file Record.

2.3: SET statement -> ACCESSES edges
     Detects SET TO TRUE (80-90% of SET usage) and SET index
     TO/UP BY/DOWN BY. Creates cobol-set-condition / cobol-set-index
     write edges + cobol-set-read for identifier values.

2.4: INSPECT -> ACCESSES edges with multi-line accumulator
     Accumulates INSPECT until period (like SORT). Extracts inspected
     field + tally counters. Creates cobol-inspect-read/write/tally
     edges. Form detection: tallying/replacing/converting/combined.

Preprocessor: 1398 -> 1597 LOC (+199). Tests: 145 passing.

* feat(cobol): implement Phase 3 — completeness fixes

6 partial features fixed to first-class support:

3.1: CALL RETURNING -> ACCESSES write edge (cobol-call-returning)
3.2: SELECT OPTIONAL flag preserved in FileDeclaration + Record node
3.3: ALTERNATE RECORD KEY extraction (matchAll for multiple keys)
3.4: COMMON attribute on nested programs (RE_PROGRAM_ID extended)
3.5: IS EXTERNAL / IS GLOBAL as first-class boolean properties
     (removed usage string hack)
3.6: AUTHOR / DATE-WRITTEN mapped to Module node description

Tests: 145 passing | TypeScript clean

* feat(cobol): implement Phase 4 — INITIALIZE + metadata completeness

4.1: INITIALIZE statement -> ACCESSES write edge (cobol-initialize)
4.2: DATE-COMPILED and INSTALLATION paragraphs extracted and mapped
     to Module node description alongside existing AUTHOR/DATE-WRITTEN

All 4 plan phases complete. Coverage: ~95% (up from 71.9%).
Tests: 145 passing | TypeScript clean

* test(cobol): add 24 unit tests for Phase 1-4 features

Coverage for all new extraction features:

Phase 1 (8 tests):
- EXEC SQL INCLUDE (unquoted, quoted, underscored)
- CALL USING (simple, mixed modes, ADDRESS OF, OMITTED)
- CALL RETURNING
- OCCURS DEPENDING ON
- VALUE clause (string, numeric, figurative constant)

Phase 2 (10 tests):
- EXEC DLI GU/ISRT/SCHD (verb, segment, PCB, INTO, FROM, PSB)
- DECLARATIVES USE AFTER EXCEPTION (single + multiple sections)
- SET TO TRUE, SET index UP BY
- INSPECT TALLYING, INSPECT REPLACING

Phase 3-4 (6 tests):
- SELECT OPTIONAL flag
- ALTERNATE RECORD KEY
- PROGRAM-ID IS COMMON
- IS EXTERNAL / IS GLOBAL booleans
- INITIALIZE extraction
- Full programMetadata (AUTHOR, DATE-WRITTEN, DATE-COMPILED, INSTALLATION)

Total: 168 tests passing (145 + 24 - 1 removed duplicate)

* fix(cobol): use /\r?\n/ split for Windows CRLF compatibility

All 4 COBOL source files now split on /\r?\n/ instead of '\n' to
handle CRLF line endings on Windows. Previously, trailing \r in
lines caused RE_GOTO's $ anchor to fail on multi-line GO TO
DEPENDING ON statements, producing only 1 goto edge instead of 4.

Files fixed: cobol-preprocessor.ts (2 sites), cobol-processor.ts,
jcl-parser.ts, cobol-copy-expander.ts

Tests: 168 passing | TypeScript clean

* fix(cobol): resolve 12th review — dynamic CALL/CANCEL dedup + trailing anchors

#1+#2: Removed incorrect hasQuotedCall/hasQuotedCancel deduplication
guards. RE_CALL_DYNAMIC and RE_CANCEL_DYNAMIC require [A-Z] after
CALL/CANCEL, so they CANNOT match quoted targets — the guards were
both unnecessary and actively harmful, suppressing dynamic CALL/CANCEL
in ON EXCEPTION patterns.

#3+#5: Changed RE_CALL_DYNAMIC and RE_CANCEL_DYNAMIC trailing anchor
from (?:\s|\.) to (?=\s|\.|$) (lookahead). The consuming anchor
failed when the identifier was the last token on a physical line.

Tests: 168 passing | TypeScript clean

* feat(cobol): add CALL accumulator + fix SORT double-statement (#4, #6)

Finding #4: Multi-line CALL USING accumulator
Added callAccum state variable that accumulates CALL statements
spanning multiple physical lines until period or END-CALL is found.
Uses flushCallAccum() to re-extract CALL target + USING parameters
from the full accumulated statement. This fixes the silent loss of
ACCESSES parameter edges when USING appears on lines after CALL.

Finding #6: SORT double-statement on same line
After flushSort(), the code now falls through to re-check the
current line for a new SORT/MERGE start (was previously blocked
by the sortAccum === null check evaluating before flushSort ran).

Also fixed: used non-global regex for CALL detection test to avoid
the classic global regex .test() lastIndex bug.

Tests: 168 passing (+1 ACCESSES from multi-line CALL USING)

* fix(cobol): resolve 13th review — CICS LOAD, USING extraction, file scoping

#1: CICS LOAD unresolved edge no longer silently deleted in second pass.
    Changed narrow cics-link/cics-xctl check to catch-all pattern:
    rel.reason?.startsWith('cics-') && rel.reason.endsWith('-unresolved')

#2: flushCallAccum USING extraction now stops before COBOL statement
    verbs (INSPECT, SEARCH, SORT, MERGE, DISPLAY, ACCEPT, MOVE, PERFORM,
    GO TO, CALL, IF, EVALUATE). Prevents absorbing adjacent statements
    as false USING parameters in legacy pre-COBOL-85 code without END-CALL.

#3: CICS FILE Record nodes now globally-scoped (<cics-file>:FILENAME)
    instead of per-file-scoped. Enables cross-program CICS file access
    analysis, consistent with SQL table scoping (<db>:TABLE).

#4: callAccum pre-check regex now has (?<![A-Z0-9-]) lookbehind to
    prevent false activation on compound identifiers like WS-CALL-FLAG.

Tests: 168 passing | TypeScript clean

* fix(cobol): resolve 14th review — callAccum false paragraph + Area A guard

#1: callAccum continuation lines now check for COBOL statement verb
    starts (GO TO, PERFORM, MOVE, etc.) and paragraph/section headers.
    If detected, the CALL is flushed as-is and the line processed
    normally — prevents false paragraph detection and currentParagraph
    corruption from lines like "WS-ADDR." being treated as paragraphs.

#4: callAccum pre-check now guarded by currentDivision === 'procedure'
    to prevent unnecessary activations in DATA DIVISION.

#5: Fixed-format paragraph detection now rejects lines with >7 leading
    spaces (Area B indentation) as paragraph candidates. Paragraph
    names in fixed-format must start in Area A (col 8-11, max 7 spaces).
    Free-format mode is unaffected.

Tests: 168 passing | TypeScript clean

* fix(cobol): resolve 15th review — callAccum Area A + verb boundary fixes

#A: Column-position-aware paragraph detection in callAccum flush.
#B: inspectAccum early-flush on paragraph/section/verb headers.
#C: Verb boundary \b → (?:\s|$) prevents MOVE-COUNT false flush.

* test(cobol): add 17 edge-case regression tests + fix USING verb boundary

17 new tests covering all recurring review patterns:

Multi-line CALL USING (7 tests):
- Parameters on separate continuation lines (IBM mainframe style)
- No absorption of INSPECT/GO TO/paragraphs following CALL
- END-CALL scope terminator
- Hyphenated identifiers (MOVE-COUNT) not triggering false flush
- Dual quoted+dynamic CALL on same line (ON EXCEPTION)

Nested program attribution (2 tests):
- CALL in inner program within inner line range
- PERFORM before first paragraph has null caller

CRLF compatibility (1 test):
- GO TO DEPENDING ON with \r\n line endings

Area A paragraph detection (2 tests):
- Area B (>7 spaces) rejected; Area A (7 spaces) accepted

SORT/MERGE (1 test): COLLATING SEQUENCE keywords not captured
PROCEDURE USING (2 tests): RETURNING excluded, period-terminated
Comment stripping (1 test): pipe in quoted string preserved
SELECT OPTIONAL (1 test): correct file name, not OPTIONAL keyword

Bug fix: USING extraction regex verb terminators changed from
\bVERB\b to \bVERB(?=\s|$) in flushCallAccum — prevents truncation
on hyphenated identifiers like MOVE-COUNT, PERFORM-LIMIT.

Total: 185 tests passing

* test(cobol): add 32 comprehensive edge-case regression tests

13 new describe blocks covering all extraction features:

- EXEC DLI: no-SEGMENT, multi-line accumulation (2 tests)
- SET: multiple targets, DOWN BY, TO numeric (3 tests)
- INSPECT: CONVERTING, multiple counters, tallying-replacing,
  paragraph flush during accumulation (4 tests)
- DECLARATIVES: no-STANDARD keyword, I-O mode, post-END paragraphs (3)
- COPY REPLACING: pseudotext deletion ==OLD== BY ==== (1 test)
- VALUE: hex literal, negative numeric, ALL literal (3 tests)
- OCCURS: TO range, fixed-size without DEPENDING ON (2 tests)
- Dynamic CALL/CANCEL: end-of-line, multiple CANCELs (3 tests)
- EXEC SQL: INCLUDE skips tables, SELECT INTO host vars, host
  variable extraction (3 tests)
- INITIALIZE: target and caller context (1 test)
- Nested programs: sibling scoping, PROGRAM-ID without ID DIV (2)
- EXEC EOF flush: unclosed EXEC SQL flushed (1 test)
- Multi-PERFORM: IF/ELSE dual PERFORM on single line (1 test)
- IS EXTERNAL: USAGE not polluted by external flag (1 test)

Total: 215 tests passing

* fix(cobol): resolve 16th review — CANCEL in CALL block + USING boundary

#1: flushCallAccum now extracts CANCEL statements from within CALL
    ON EXCEPTION blocks. Adds RE_CANCEL + RE_CANCEL_DYNAMIC matchAll
    passes alongside existing CALL extraction.

#2: Added \bCANCEL(?=\s|$) to USING lookahead regex to prevent CANCEL
    keyword being captured as false USING parameter.

#3: Multi-line CALL start now returns immediately to prevent the CALL
    start line from simultaneously feeding sortAccum/inspectAccum.

#6: Division transitions now flush all active accumulators (callAccum,
    sortAccum, inspectAccum) to prevent state leakage across programs.

Also added CANCEL to callAccum flush trigger verb list.

Tests: 215 passing | TypeScript clean

* refactor(cobol): extract shared verb constants + resolve 17th review

Extract COBOL_STATEMENT_VERBS, RE_STATEMENT_VERB_START, and
RE_USING_PARAMS as shared constants — eliminates 4 duplicated
25-verb regex patterns.

17th review: #1 flushCallAccum before EXEC entry, #2 inspectAccum
verb parity via shared constant.

Tests: 215 passing | TypeScript clean

* test(cobol): replace all fuzzy assertions with exact toBe checks

Replaced 7 toBeGreaterThan/toBeLessThan/toBeGreaterThanOrEqual
assertions with exact toBe values:

- dataItems.length: >= 3 → toBe(3)
- calls.length: >= 1 → toBe(1)
- calls[0].line: range check → toBe(10)
- programs[].startLine/endLine: comparison → exact values
- innerA.endLine/innerB.startLine: comparison → exact values

Also added 11 new edge-case tests (accumulator flush on EXEC/division
transitions, free-format, CANCEL in CALL block, SORT THRU, verb
flush, integration).

226 tests passing — zero fuzzy assertions remain.

* fix(cobol): resolve 19th review + 15 accumulator flush tests

Fixes:
#1: END PROGRAM flushes callAccum/sortAccum/inspectAccum
#2: PROGRAM-ID sibling path flushes all accumulators
#3: Added COMPUTE/ADD/SUBTRACT/MULTIPLY/DIVIDE/STRING/UNSTRING
    to COBOL_STATEMENT_VERBS (now 32 verbs)

Tests (15 new):
- END PROGRAM flush: single + nested programs (2)
- PROGRAM-ID sibling flush (1)
- Arithmetic verb flush: COMPUTE/ADD/SUBTRACT/MULTIPLY/DIVIDE (5)
- String verb flush: STRING/UNSTRING (2)
- Arithmetic not captured as false USING params (1)
- SORT flushed at END PROGRAM (1)
- INSPECT flushed at END PROGRAM (1)
- All with exact toBe assertions (2)

Total: 239 tests passing | Zero fuzzy assertions

* fix(cobol): resolve 20th review — INITIALIZE multi-target + 2 tests

Finding 1: INITIALIZE now captures multiple targets with REPLACING
clause keyword filtering. Regex changed to lazy match stopping at
REPLACING/WITH/period boundary. Targets split on whitespace and
filtered against INITIALIZE_CLAUSE_KEYWORDS set.

Tests (2 new):
- INITIALIZE multi-target: WS-CUSTOMER WS-ORDER WS-LINE-ITEM → 3
- INITIALIZE with REPLACING: only WS-RECORD captured, not keywords

Total: 241 tests passing | TypeScript clean
2026-03-26 14:03:48 +00:00
3c896cdbcd fix: close remaining Dart language support gaps (#524)
* fix: close remaining Dart language support gaps

Four issues that were not addressed in PR #204:

1. extractFunctionName: add function_signature/method_signature handlers
   and add both to FUNCTION_NODE_TYPES. Without this, findEnclosingFunctionId
   cannot resolve Dart function scopes — all calls inside Dart functions
   have no sourceId, breaking CALLS edge attribution.

2. formal_parameter_list: add to paramListTypes in extractMethodSignature.
   Dart's tree-sitter grammar uses this node type (not formal_parameters),
   so parameter counting returns 0 for all Dart functions.

3. Write-access queries: add @assignment patterns for obj.field = value
   and this.field = value. Without these, no ACCESSES write edges are
   emitted for Dart code.

4. initialized_identifier guard in extractDartDeclaration: comma-separated
   declarations (String a, b, c) produce initialized_identifier nodes
   which are in DART_DECLARATION_NODE_TYPES but were unhandled — the type
   lives on the parent node.

Also adds Dart column to the feature matrix in type-resolution-system.md.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(dart): field-type resolution, call attribution, import resolution, and integration tests

Fixes five Dart language support gaps with integration tests and
architectural alignment:

**Tree-sitter queries** — Add field declaration patterns for typed and
nullable class fields (`String name = ''`, `String? name`).  Without
these, Dart class fields were invisible to the pipeline (zero Property
nodes, zero HAS_PROPERTY edges).

**Import resolution** — Dart relative imports (`import 'models.dart'`)
don't use a leading `./`.  The standard resolver only recognises paths
starting with `.` as relative; bare paths fell through to a Java-style
dot-to-slash conversion that mangled `models.dart` into `models/dart`.
Fix: prepend `./` before calling resolveStandard.

**Call attribution** — Dart's tree-sitter grammar places `function_body`
as a sibling of `function_signature`, not as a child wrapping both.  The
`findEnclosingFunction` parent-walk never found the function because the
call lives inside `function_body` which is a sibling of the signature.
Fix: add `enclosingFunctionFinder` hook to LanguageProvider interface
(following the same strategy pattern as `labelOverride`), with the
Dart-specific logic in `languages/dart.ts`.  Both `parse-worker.ts` and
`call-processor.ts` consume the hook generically — no Dart-specific code
in the generic processors.

**Receiver chain extraction** — Add `unconditional_assignable_selector`
to `MEMBER_ACCESS_NODE_TYPES` so `inferCallForm` returns `'member'` for
Dart method calls.  Add Dart-specific receiver extraction blocks in
`extractReceiverName`, `extractReceiverNode`, and a `selector` handler
in `extractMixedChain` for Dart's flat sibling-selector model (vs the
nested member-expression model used by all other languages).

**Integration tests** — New `dart.test.ts` with field-type resolution
and call-result-binding describe blocks.  Fixtures: `dart-field-types/`
(models.dart + app.dart) and `dart-call-result-binding/` (models.dart +
app.dart).  9 passing tests, 1 skipped (ACCESSES edges for field reads
depend on type-env parameter binding propagation — tracked for follow-up).

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-03-26 13:45:24 +00:00
Wojciech Guziak 546128cdcb refactor: split global BUILT_IN_NAMES into per-language provider fields (#523)
* refactor: make isBuiltInOrNoise provider-aware, remove global BUILT_IN_NAMES

Add builtInNames field to LanguageProviderConfig. Rewrite noise-filter.ts
to accept a LanguageProvider and check provider.builtInNames instead of
a global Set. Update all 3 call sites to pass their existing provider.

Built-in entries will be added per-language in subsequent commits.

* refactor(js/ts): add per-language builtInNames to JS/TS providers

* refactor(python): add per-language builtInNames

* refactor(kotlin): add per-language builtInNames

* refactor(c/cpp): add per-language builtInNames

* refactor(csharp): add per-language builtInNames

* refactor(php): add per-language builtInNames

* refactor(swift): add per-language builtInNames

* refactor(rust): add per-language builtInNames

* refactor(ruby): add per-language builtInNames

* refactor(dart): add per-language builtInNames

* test: update noise-filter tests for per-language API, add isolation tests

- Update ingestion-utils.test.ts to pass provider to isBuiltInOrNoise
- Add noise-filter.test.ts with 15 cross-language isolation tests
- Fix Java heritage test: serialize() is now correctly unfiltered for Java
  (was false-positive noise from global PHP serialize entry)

* refactor: remove noise-filter.ts, add provider.isBuiltInName() method

Per review feedback: delete noise-filter.ts entirely and move the check
into LanguageProvider as isBuiltInName(name) method, generated by
defineLanguage() from the builtInNames set.

Call sites now use provider.isBuiltInName(calledName) directly.
2026-03-26 12:15:29 +00:00
Flavius MironandClaude Opus 4.6 c6ed25de8b docs: add Dart to supported languages table in README (#525)
Dart was added as the 14th supported language in PR #204 but the README
was not updated. Adds Dart row to the supported languages table and
updates the language count from 13 to 14.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-26 10:46:10 +00:00
ThinhKVT aec9a3f216 [CLI] Fixes a false-positive in the Cypher write-detection regex and improves the Impact tool's enrichment path by using batched chunking and entry-point grouping (#496) (#507) 2026-03-26 10:39:27 +00:00
marxo126 a047a08f54 feat: add ORM dataflow detection (Prisma + Supabase) (#511) 2026-03-26 08:48:40 +00:00
Wojciech Guziak 8864a0290c feat: add Dart language support (#204) 2026-03-26 08:42:49 +00:00
Chirag Nighut 1a0784befe feat: add more node types in filter panel (#519)
* fix(web): enable Cypher queries when connected to backend server

Route queries through HTTP API in backend mode instead of checking local WASM database.

Made-with: Cursor

* chore: add Maven/Gradle wrapper files to default ignore list

Add build wrapper scripts and directories to hardcoded ignore lists:
- Directories: .mvn, .gradle, gradle
- Files: mvnw, mvnw.cmd, gradlew, gradlew.bat

These are build infrastructure files, not source code.

Made-with: Cursor

* ci: re-trigger CI (Windows flaky timeout)

Made-with: Cursor

* feat: add more node types in filter panel

* feat: add more node types in filter panel

* revert additional changes

* test(web): add unit tests for filter panel node types

- FILTERABLE_LABELS: verify new types (Enum, Type, Decorator, Variable)
  have colors, sizes, and no duplicates
- Filter panel icons: verify every filterable label has an icon mapped
  and all icons are exported from lucide-icons
- Color legend: verify new types are included, ordered correctly, and
  are a subset of FILTERABLE_LABELS

Made-with: Cursor
2026-03-26 06:50:40 +00:00
Gergő Magyar 33225c2fce fix(ci): move shape-check-regression test to lbug-db project (#518)
The shape-check-regression test uses withTestLbugDB but was running in
the default vitest project with parallel forks, causing LadybugDB
file-lock conflicts on Windows CI. Move it to the lbug-db project
(sequential execution) and exclude from default.

Follows up on #501.
2026-03-26 06:23:02 +00:00
marxo126 5a7ac218df fix: shape_check false positives — quoted keys, DOM leaks, errorKeys (#501) 2026-03-26 05:43:37 +00:00
jelsco b959b9933b fix(python): resolve two remaining alias gaps (#417) (#505) 2026-03-26 05:23:27 +00:00
Zander Raycraft 6fabd7a2df Merge pull request #402 from adonisdoda/feat/index-cli 2026-03-25 20:22:21 -05:00
marxo126 f860653a69 feat(routes): link Next.js project-level middleware.ts to routes (#504) 2026-03-25 22:07:58 +00:00
Wojciech Guziak 77dcb06a8d chore: upgrade tree-sitter to 0.25.0 and all grammar packages (#516) 2026-03-25 21:55:14 +00:00
Zander Raycraft e7e26d6345 Merge pull request #381 from cnighut/feat/cursor-cli-wiki-provider 2026-03-25 08:20:23 -05:00
marxo126 95f97c884c feat: add Expo Router file-based route detection (#503) 2026-03-25 11:05:55 +00:00
marxo126andClaude Opus 4.6 4bc4815bd2 feat: PHP response shape extraction for json_encode patterns (#502)
* feat: add PHP response shape extraction for json_encode patterns

Adds extractPHPResponseShapes() to detect response keys from PHP
json_encode() calls with associative array literals. Supports:
- Short array syntax: json_encode(['key' => value])
- Long array syntax: json_encode(array('key' => value))
- Error classification via http_response_code() and header() status
- exit;/die; boundary detection to prevent cross-block status leaking
- Nested array filtering (only top-level keys extracted)

Pipeline integration dispatches PHP files to the new extractor.
Verified on collector project: 10 PHP routes now show responseKeys.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address review — exit boundary, die; offset, CGI Status header

- Replace lastIndexOf('exit;')/lastIndexOf('die;') with regex that
  matches exit(N), exit(0), die('msg'), die($var) as boundaries
- Fixes die; off-by-one (was slicing at +5 for a 4-char keyword)
- Add header('Status: NNN') CGI/FastCGI format detection
- Add 3 regression tests for the fixed bugs

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: extract shared helpers, remove duplicate test block

- Extract lastMatchGroup() and buildShapeResult() to eliminate repeated
  patterns in both JS/TS and PHP extractors
- Simplify detectPHPStatusCode to use ?? chaining with lastMatchGroup
- Remove duplicate 9-test PHP describe block (kept the 12-test version)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test: add PHP response shape integration tests

Adds a PHP fixture (api/items.php, api/submit.php) with multiple
json_encode patterns and a pipeline integration test verifying:
- Route nodes created for PHP endpoints
- responseKeys/errorKeys correctly extracted and separated
- exit(N)/die() boundaries respected
- HANDLES_ROUTE edges point to correct PHP handler files

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 08:27:53 +00:00
c68d7975e6 docs: agent development framework, GitHub templates, eval refactor (#479)
* ci: E2E workflow, web typecheck job, pre-commit hook, test suite

CI:
- ci.yml consolidated to reference ci-tests.yml
- ci-quality.yml: add typecheck-web job for gitnexus-web/
- ci-e2e.yml: E2E workflow with dorny/paths-filter (web changes only)
- ci-report.yml: remove dead integration-reports references
- CI gate allows skipped E2E status
- .gitignore: playwright artifacts, eval test artifacts

Pre-commit hook:
- .githooks/pre-commit: typecheck + unit tests for both packages
- Activated via git config core.hooksPath in prepare script

Test infrastructure:
- Vitest + React Testing Library: 58 unit tests
  (graph, server-connection, mermaid, settings, constants, utils, paths)
- Playwright E2E: 5 tests + manual recording harness
- vitest.config from vitest/config, engines.node >= 20
- Playwright artifacts retain-on-failure
- wait-on in devDependencies
- vitest/coverage-v8 aligned with vitest 4.x

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: update gitnexus-web package-lock.json

Reflects devDependency additions (vitest, playwright, wait-on,
@testing-library, etc.) from package.json changes in this PR.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(e2e): add missing process-list-loaded testid, increase CI timeouts

- Add data-testid="process-list-loaded" to ProcessesPanel (E2E tests
  were waiting for an element that didn't exist)
- Increase server connect timeouts from 5s to 10s for slower CI

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): run gitnexus-web unit tests in CI, remove unused variable

- Add gitnexus-web npm ci + vitest run to ci-tests.yml so web unit
  tests are gated by the CI status check (were only running locally)
- Remove unused IS_PLAYWRIGHT_AUTOMATION variable from E2E spec

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(e2e): add process-row testid, wait for networkidle on page load

- Add data-testid="process-row" to ProcessItem component (E2E tests
  referenced it but it didn't exist in the source)
- Use waitUntil: 'networkidle' on page.goto to ensure Vite dev server
  is fully ready before interacting (fixes first-test timeout in CI)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(e2e): add process-view-button and process-highlight-button testids

E2E tests referenced these data-testid attributes but they didn't
exist in ProcessItem. All 6 E2E testids now have matching source
elements: status-ready, process-list-loaded, process-row,
process-view-button, process-highlight-button, server-url-input.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(e2e): remove networkidle — Vite HMR WebSocket prevents it from resolving

networkidle waits for zero network activity for 500ms, but Vite's HMR
WebSocket stays open permanently, causing page.goto to timeout at 60s
on all tests after the first. The explicit toBeVisible waits on UI
elements are sufficient and deterministic.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(e2e): wait for Server button visibility, add CI retry, all 5 tests pass locally

Root cause: test 1 clicked the Server button before React hydrated,
so the tab content never rendered and the input wasn't found.

Fixes:
- Wait for Server button toBeVisible before clicking
- Increase input wait to 15s
- Remove networkidle (Vite HMR WebSocket prevents it from resolving)
- Add retries: 1 in CI for transient cold-start flakiness

Verified locally: all 5 E2E tests pass, 198 unit tests pass, typecheck clean.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): tolerate LadybugDB native crash during analyze step

gitnexus analyze can crash with "double free or corruption" (known
issue #273) during the LadybugDB native addon shutdown. The index is
usually written successfully before the crash. The workflow now:
1. Allows analyze to exit non-zero with a warning
2. Verifies .gitnexus index was actually created
3. Only fails if no index exists (real failure)

All tests verified locally: 198 unit, 5 E2E pass, typecheck clean.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): fix shell quoting in analyze step, simplify to || true

The previous echo string had special characters that broke bash
quoting in GitHub Actions. Simplified to: analyze || true, then
check if .gitnexus exists.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: add agent development framework, GitHub templates, eval refactor

Agent framework (layered docs for AI-assisted contributions):
- AGENTS.md: canonical instructions, impact analysis, MCP tools
- CLAUDE.md: Claude Code-specific deltas and hooks
- GUARDRAILS.md: safety boundaries, non-negotiables, escalation
- ARCHITECTURE.md: monorepo layout, data flow map
- TESTING.md: test structure, commands, categories
- RUNBOOK.md: copy-paste operations for dev/CI/MCP
- llms.txt: minimal LLM context pointer

Editor integration:
- .cursor/index.mdc + rules/100-monorepo.mdc

GitHub templates:
- PR template with areas-touched checkboxes
- Bug report + feature request issue forms

Eval harness:
- Refactored mcp_bridge, tool_registry, constants
- Error sanitization utilities
- Property-based tests via Hypothesis

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(eval): use format_exception instead of format_exc in sanitize_exception

format_exc() returns the currently handled exception traceback, which
may be unrelated if called outside an active except block. Using
format_exception(type(exc), exc, exc.__traceback__) reliably captures
the passed exception's traceback.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: update CONTRIBUTING.md and TESTING.md for current CI/hook setup

- CONTRIBUTING.md: add gitnexus-web typecheck command, pre-commit hook
  checklist item
- TESTING.md: add gitnexus-web typecheck command, pre-commit hook
  section (husky), update CI integration to list actual workflow files
  (ci-quality, ci-tests, ci-e2e)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: update testing docs to reflect CI/E2E changes from PR #486

- AGENTS.md: update test counts (CLI ~2000 unit, ~1850 integration),
  add gitnexus-web testing section (198 unit, 5 E2E with commands)
- RUNBOOK.md: fix Node requirement to >=20, fix E2E local repro command
- TESTING.md: E2E uses data-testid selectors + real servers, not mocks
- .cursor/rules/100-monorepo.mdc: add web test/E2E commands

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: address context engineering review — deduplicate tokens, expand Cursor rules

- Remove ~100-line gitnexus:start block from CLAUDE.md (was duplicated from AGENTS.md)
- Fix gitnexus:start block inlined inside AGENTS.md Reference Docs bullet (doubled)
- Replace CLAUDE.md scope table with pointer to AGENTS.md (single source of truth)
- Expand .cursor/index.mdc with 5 non-negotiable safety rules for always-on context
- Add .cursor/rules/200-eval.mdc with Python/eval commands (glob-scoped to eval/**)
- Improve llms.txt with priority annotations and descriptions
- Bump version headers to 1.2.0, last-reviewed to 2026-03-24

Saves ~1,400 tokens/session with zero information loss.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-03-25 06:48:41 +00:00
chirag-nighut 048347df84 fix: address PR review — TTY guard, test rename, unify debug env var
- Add process.stdin.isTTY guard before --review prompt to prevent CI hangs
- Rename misleading --verbose e2e test to reflect it checks help output
- Replace DEBUG with GITNEXUS_VERBOSE for error stack traces

Made-with: Cursor
2026-03-25 11:39:37 +05:30
chirag-nighut f17bde69ba ci: re-trigger CI (Windows flaky timeout in skills-e2e)
Made-with: Cursor
2026-03-25 11:39:37 +05:30
chirag-nighut 7e66ec3a4f test: add e2e CLI tests for wiki flags (--provider, --review, --verbose)
Spawn actual CLI process to verify:
- wiki --help surfaces all new flags
- wiki on non-git directory exits with code 1
- wiki on non-indexed repo fails with "No GitNexus index"
- --provider cursor skips API key prompt in non-TTY mode
- --verbose is accepted as valid flag

Made-with: Cursor
2026-03-25 11:39:37 +05:30
chirag-nighut 76a581b295 test: add unit tests for wiki CLI flags (--provider, --review, --verbose)
17 tests covering:
- detectCursorCLI caching (avoids repeated spawns)
- resolveCursorConfig defaults
- resolveLLMConfig provider routing (cursor vs openai)
- --verbose env var propagation
- WikiGenerator reviewOnly mode (early return with moduleTree)
- CLI config round-trip for cursor and openai providers
- invokeLLM dispatch (cursor → callCursorLLM, openai → callLLM)
- callCursorLLM error when CLI not found
- estimateTokens heuristic

Made-with: Cursor
2026-03-25 11:39:37 +05:30
chirag-nighut 1e309c1ec1 fix: address PR review — remove redundancies and add wiki help test
- Cache detectCursorCLI() result to avoid spawning `agent --version`
  on every LLM call
- Fix stale JSDoc in cursor-client.ts (no longer uses stdin or
  stream-json)
- Remove unused WikiOptions fields (model, baseUrl, apiKey) that were
  passed but never read by WikiGenerator
- Fix inconsistent progress callback phase tracking in --review
  continuation path
- Add wiki CLI help test covering --provider, --review, --verbose flags

Made-with: Cursor
2026-03-25 11:39:37 +05:30
chirag-nighut fca7815c16 feat(wiki): add Cursor CLI as LLM provider option
Add Cursor headless CLI as a 4th provider option for wiki generation,
allowing users to leverage their Cursor subscription for wiki pages.

- Add --provider cursor flag and cursor-client.ts
- Add --review flag for interactive module tree editing
- Add --verbose flag for debugging
- Improve module tree generation (flatten single-child, unique slugs)
- Prevent DB timeout during long LLM calls

Usage: gitnexus wiki --provider cursor --model claude-4.5-opus-high
Made-with: Cursor
2026-03-25 11:39:37 +05:30
Yash a191c26571 fix(#480): resolve impact/context returning empty results for Java cl… (#489) 2026-03-24 21:46:46 +00:00
adonis 1bb80c3858 feat(cli): enhance index command with single path validation and add to .gitignore 2026-03-23 00:19:26 -03:00
adonis 6fe450d206 test(cli): refactor indexCommand tests for consistency and add new cases 2026-03-22 22:05:22 -03:00
adonis 3567ef5794 feat(cli): add --allow-non-git option to index command and update tests 2026-03-20 19:46:47 -03:00
adonis 680cefd00c feat(cli): refactor indexCommand for improved error handling and add unit tests 2026-03-20 19:36:07 -03:00
adonis c1996f45d8 feat(cli): add 'index' command to register existing .gitnexus/ folder 2026-03-20 19:16:09 -03:00
833 changed files with 89789 additions and 33583 deletions
+21
View File
@@ -0,0 +1,21 @@
---
alwaysApply: true
---
# GitNexus — Cursor project rules
Last reviewed: 2026-03-24
Canonical agent instructions: **[AGENTS.md](../AGENTS.md)** (GitNexus MCP rules, monorepo commands, Cursor Cloud notes). **[CLAUDE.md](../CLAUDE.md)** adds Claude Code-specific notes and points back to AGENTS.md for GitNexus.
## Non-negotiables (always apply)
- NEVER edit a function/class/method without running `gitnexus_impact` first.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename`.
- NEVER commit without running `gitnexus_detect_changes()`.
- NEVER ignore HIGH/CRITICAL risk warnings from impact analysis.
- NEVER run `npx gitnexus analyze` without `--embeddings` if `.gitnexus/meta.json` shows stored embeddings.
Full rules: **[AGENTS.md](../AGENTS.md)** (`gitnexus:start` block, Cursor Cloud section).
**Rule architecture:** Prefer this file plus optional `.cursor/rules/*.mdc` globs (YAML `globs` in frontmatter). Legacy `.cursorrules` is deprecated; content lives here.
+12
View File
@@ -0,0 +1,12 @@
---
globs:
- "gitnexus/**"
- "gitnexus-web/**"
---
# GitNexus build/test quick refs
- CLI (`gitnexus/`): `npm test`; `npm run test:integration`; `npx tsc --noEmit`.
- Web (`gitnexus-web/`): `npm test`; `npm run dev`; `npx tsc -b --noEmit`; `E2E=1 npx playwright test` (needs servers).
- `npm install` in `gitnexus/` runs `prepare` (tsc build) and `postinstall` (tree-sitter patches); needs `python3`, `make`, `g++`.
- LadybugDB locking tests may fail in containerized environments because of `/tmp` file locks (known issue, not a code bug).
+14
View File
@@ -0,0 +1,14 @@
---
globs:
- "eval/**"
---
# GitNexus eval harness (Python)
- **Run tests**: `cd eval && uv run pytest tests/`
- **Run with coverage**: `cd eval && uv run coverage run -m pytest tests/ && uv run coverage report`
- **Lint**: `cd eval && uv run ruff check .`
- **Run eval**: `cd eval && uv run python run_eval.py --config configs/<config>.yaml`
- Shared constants live in `eval/constants.py`; tool specs in `eval/tool_registry.py`.
- Error logging uses `utils/errors.py` — set `GITNEXUS_EVAL_DEBUG=1` for full tracebacks.
- Property-based tests use Hypothesis (`eval/tests/test_property_based.py`).
+3 -3
View File
@@ -1,5 +1,5 @@
# AI Agent Rules
# Deprecated for Cursor Agent Mode
Follow .gitnexus/RULES.md for all project context and coding guidelines.
Use **`.cursor/index.mdc`** (`alwaysApply: true`) for project rules. See [AGENTS.md](AGENTS.md).
This project uses GitNexus MCP for code intelligence. See .gitnexus/RULES.md for available tools and best practices.
This file is kept only as a breadcrumb for older workflows.
+5
View File
@@ -0,0 +1,5 @@
# Prettier initial formatting (2026-03-28)
afcc3d1523f99c77ff67c4fd1af12334660113f6
# ESLint unused import removal (2026-03-28)
1491826bc8da5436d3b1eb092d9274f2c9f028a7
+2
View File
@@ -0,0 +1,2 @@
* text=auto eol=lf
.husky/* text eol=lf
+86
View File
@@ -0,0 +1,86 @@
name: Bug report
description: Report unexpected behavior or a regression
labels: [bug]
body:
- type: markdown
attributes:
value: |
**Goal:** capture enough context to reproduce and fix the issue quickly.
Use **one issue per bug**; split unrelated problems.
- type: dropdown
id: area
attributes:
label: Area
description: Where does the problem show up?
options:
- gitnexus (CLI / core / indexing / MCP server)
- gitnexus-web (browser UI / WASM / workers)
- CI / GitHub Actions
- Documentation / developer experience
- Other
validations:
required: true
- type: textarea
id: summary
attributes:
label: Summary
description: One sentence — what went wrong?
validations:
required: true
- type: textarea
id: context
attributes:
label: Context
description: What were you trying to do? Any relevant links, PRs, or commits?
validations:
required: false
- type: textarea
id: expected
attributes:
label: Expected behavior
validations:
required: true
- type: textarea
id: actual
attributes:
label: Actual behavior
validations:
required: true
- type: textarea
id: reproduce
attributes:
label: Steps to reproduce
description: Ordered steps, sample repo or minimal case, commands run.
placeholder: |
1. …
2. …
3. …
validations:
required: true
- type: textarea
id: environment
attributes:
label: Environment
description: OS, Node version, browser (if web), GitNexus version or commit SHA.
placeholder: |
- OS:
- Node:
- Browser (if applicable):
- Commit / version:
validations:
required: false
- type: textarea
id: logs
attributes:
label: Logs / screenshots
description: Paste errors, stack traces, or attach screenshots (redact secrets).
validations:
required: false
+1
View File
@@ -0,0 +1 @@
blank_issues_enabled: true
@@ -0,0 +1,74 @@
name: Feature request
description: Propose a new capability or improvement
labels: [enhancement]
body:
- type: markdown
attributes:
value: |
**Goal:** describe the problem and desired outcome so maintainers can size and prioritize.
Prefer **small, shippable** requests; split large ideas into phases.
- type: dropdown
id: area
attributes:
label: Area
description: Primary part of the monorepo this relates to.
options:
- gitnexus (CLI / core / indexing / MCP server)
- gitnexus-web (browser UI / WASM / workers)
- CI / release / packaging
- Documentation / developer experience
- Other
validations:
required: true
- type: textarea
id: problem
attributes:
label: Problem or opportunity
description: What pain point or gap exists today?
validations:
required: true
- type: textarea
id: proposal
attributes:
label: Proposed solution
description: What should happen instead? User-visible behavior, APIs, or UX.
validations:
required: true
- type: textarea
id: alternatives
attributes:
label: Alternatives considered
description: Other approaches you considered and why this one is preferred.
validations:
required: false
- type: textarea
id: acceptance
attributes:
label: Acceptance criteria
description: Testable conditions for “done” (bullets or checkboxes in prose).
placeholder: |
- When … then …
- Documentation / tests updated where appropriate
validations:
required: false
- type: textarea
id: constraints
attributes:
label: Constraints
description: Compatibility, performance, security, or “must not change” boundaries.
validations:
required: false
- type: checkboxes
id: willing
attributes:
label: Contribution
options:
- label: I am willing to open a PR for this (may need design discussion first).
required: false
+52
View File
@@ -0,0 +1,52 @@
## Summary
<!-- One or two sentences: what does this PR change? -->
## Motivation / context
<!-- Why is this change needed? Link issues, ADRs, or prior discussion. -->
## Areas touched
<!-- Check all that apply -->
- [ ] `gitnexus/` (CLI / core / MCP server)
- [ ] `gitnexus-web/` (Vite / React UI)
- [ ] `.github/` (workflows, actions)
- [ ] `eval/` or other tooling
- [ ] Docs / agent config only (`AGENTS.md`, `CLAUDE.md`, `.cursor/`, `llms.txt`, etc.)
## Scope & constraints
**In scope**
- <!-- bullets -->
**Explicitly out of scope / not done here**
- <!-- bullets — prevents reviewers assuming missing work is an oversight -->
## Implementation notes
<!-- Optional: design choices, tradeoffs, follow-ups -->
## Testing & verification
<!-- What you ran; paste commands. Omit sections that do not apply. -->
- [ ] `cd gitnexus && npm test`
- [ ] `cd gitnexus && npm run test:integration` *(if core/indexing/MCP paths changed)*
- [ ] `cd gitnexus && npx tsc --noEmit`
- [ ] `cd gitnexus-web && npm test` *(if web changed)*
- [ ] `cd gitnexus-web && npx tsc -b --noEmit` *(if web changed)*
- [ ] Manual / Playwright E2E *(note environment — see `gitnexus-web/e2e/`)*
## Risk & rollout
<!-- Breaking changes, migrations, index refresh (`npx gitnexus analyze`), release notes -->
## Checklist
- [ ] PR body meets repo minimum length (workflow may label short descriptions)
- [ ] If `AGENTS.md` / overlays changed: headers, scope block, and changelog updated per project conventions
- [ ] No secrets, tokens, or machine-specific paths committed
@@ -0,0 +1,21 @@
name: Setup GitNexus Web
description: Setup Node.js 20, build gitnexus-shared, install web dependencies
runs:
using: composite
steps:
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus-web/package-lock.json
- name: Build gitnexus-shared
run: npm install && npm run build
shell: bash
working-directory: gitnexus-shared
- name: Install web dependencies
run: npm ci
shell: bash
working-directory: gitnexus-web
@@ -16,6 +16,11 @@ runs:
cache: npm
cache-dependency-path: gitnexus/package-lock.json
- name: Build gitnexus-shared
run: npm install && npm run build
shell: bash
working-directory: gitnexus-shared
- name: Install dependencies
run: npm ci
shell: bash
+1 -1
View File
@@ -35,7 +35,7 @@ changelog:
- dependencies
- title: "\U0001F4DD Other Changes"
labels:
- "*"
- '*'
exclude:
labels:
- dependencies
+1 -9
View File
@@ -28,15 +28,7 @@ jobs:
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus-web/package-lock.json
- name: Install frontend dependencies
run: npm ci
working-directory: gitnexus-web
- uses: ./.github/actions/setup-gitnexus-web
- name: Install Playwright browsers
run: npx playwright install --with-deps chromium
+27 -7
View File
@@ -4,6 +4,32 @@ on:
workflow_call:
jobs:
format:
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
cache: npm
cache-dependency-path: package-lock.json
- run: npm ci
- run: npx prettier --check .
lint:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
cache: npm
cache-dependency-path: package-lock.json
- run: npm ci
- run: npx eslint .
typecheck:
runs-on: ubuntu-latest
timeout-minutes: 10
@@ -18,12 +44,6 @@ jobs:
timeout-minutes: 10
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus-web/package-lock.json
- run: npm ci
working-directory: gitnexus-web
- uses: ./.github/actions/setup-gitnexus-web
- run: npx tsc -b --noEmit
working-directory: gitnexus-web
+3 -3
View File
@@ -6,12 +6,12 @@ name: CI Report
on:
workflow_run:
workflows: ["CI"]
workflows: ['CI']
types: [completed]
permissions:
actions: read # needed to list/download workflow run artifacts
contents: read # needed for sparse checkout of vitest.config.ts
actions: read # needed to list/download workflow run artifacts
contents: read # needed for sparse checkout of vitest.config.ts
pull-requests: write # needed to post sticky PR comment
jobs:
+1
View File
@@ -28,6 +28,7 @@ jobs:
--coverage.reportOnFailure=true
working-directory: gitnexus
# gitnexus-shared already built by setup-gitnexus action above
- name: Install gitnexus-web dependencies
run: npm ci
working-directory: gitnexus-web
+1 -1
View File
@@ -53,7 +53,7 @@ jobs:
pull-requests: write
issues: write
id-token: write
actions: read # required for Claude to read CI results on PRs
actions: read # required for Claude to read CI results on PRs
steps:
# For PR-related triggers, resolve the fork repo so we can checkout correctly.
- name: Resolve PR context
+20 -1
View File
@@ -30,6 +30,10 @@ jobs:
registry-url: https://registry.npmjs.org
cache: npm
cache-dependency-path: gitnexus/package-lock.json
- name: Build gitnexus-shared
run: npm install && npm run build
working-directory: gitnexus-shared
- run: npm ci
working-directory: gitnexus
@@ -63,7 +67,22 @@ jobs:
env:
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Extract release notes from CHANGELOG
id: changelog
shell: bash
run: |
VERSION="${GITHUB_REF#refs/tags/v}"
NOTES=$(awk "/^## \\[$VERSION\\]/{found=1; next} /^## \\[/{if(found) exit} found" gitnexus/CHANGELOG.md)
if [ -z "$NOTES" ]; then
echo "::warning::No CHANGELOG entry found for v$VERSION, falling back to auto-generated notes"
echo "fallback=true" >> "$GITHUB_OUTPUT"
else
echo "$NOTES" > /tmp/release-notes.md
echo "fallback=false" >> "$GITHUB_OUTPUT"
fi
- name: Create GitHub Release
uses: softprops/action-gh-release@a06a81a03ee405af7f2048a818ed3f03bbf83c7b # v2
with:
generate_release_notes: true
body_path: ${{ steps.changelog.outputs.fallback == 'false' && '/tmp/release-notes.md' || '' }}
generate_release_notes: ${{ steps.changelog.outputs.fallback == 'true' }}
+3 -1
View File
@@ -93,4 +93,6 @@ GitNexus.sln
.history/
.swarm/
.swarm/
local_docs/
+9 -16
View File
@@ -1,33 +1,26 @@
#!/usr/bin/env bash
# Pre-commit hook (husky): typecheck + unit tests for both packages.
# Mirrors CI checks from ci-quality.yml and ci-tests.yml.
# Pre-commit hook: format staged files + typecheck.
# Tests run in CI (ci-tests.yml), not here.
# Skip with: git commit --no-verify
#
# CI coverage:
# quality / typecheck → tsc --noEmit in gitnexus/
# quality / typecheck-web → tsc -b --noEmit in gitnexus-web/
# tests / ubuntu+coverage → vitest run in gitnexus/ (all projects)
# e2e / chromium → playwright (requires servers — skipped)
ROOT="$(git rev-parse --show-toplevel)"
# 1. Format staged files with prettier via lint-staged
echo "pre-commit: formatting staged files..."
"$ROOT/node_modules/.bin/lint-staged" || exit 1
# 2. Typecheck changed packages
WEB_CHANGED=$(git diff --cached --name-only -- 'gitnexus-web/' | head -1)
CLI_CHANGED=$(git diff --cached --name-only -- 'gitnexus/' | head -1)
if [ -n "$WEB_CHANGED" ]; then
echo "pre-commit: typechecking gitnexus-web (tsc -b)..."
cd "$ROOT/gitnexus-web" && npx tsc -b --noEmit
echo "pre-commit: running gitnexus-web unit tests..."
npx vitest run --reporter=dot
cd "$ROOT/gitnexus-web" && ./node_modules/.bin/tsc -b --noEmit || exit 1
fi
if [ -n "$CLI_CHANGED" ]; then
echo "pre-commit: typechecking gitnexus..."
cd "$ROOT/gitnexus" && npx tsc --noEmit
echo "pre-commit: running gitnexus unit tests (default project)..."
npx vitest run --project default --reporter=dot
cd "$ROOT/gitnexus" && ./node_modules/.bin/tsc --noEmit || exit 1
fi
echo "pre-commit: all checks passed"
+16
View File
@@ -0,0 +1,16 @@
dist/
coverage/
gitnexus/vendor/
gitnexus/test/fixtures/
gitnexus-web/playwright-report/
gitnexus-web/test-results/
*.d.ts
*.snap
*.wasm
*.md
.gitnexus/
.vercel/
.claude-flow/
.swarm/
assets/
repomix-output*
+10
View File
@@ -0,0 +1,10 @@
{
"semi": true,
"singleQuote": true,
"trailingComma": "all",
"printWidth": 100,
"tabWidth": 2,
"endOfLine": "lf",
"plugins": ["prettier-plugin-tailwindcss"],
"tailwindStylesheet": "./gitnexus-web/src/index.css"
}
+104 -1
View File
@@ -1,7 +1,69 @@
<!-- version: 1.2.0 -->
<!--
Metadata: version, last reviewed, scope, model policy, reference docs, changelog.
Last updated: 2026-03-22
-->
Last reviewed: 2026-03-24
**Project:** GitNexus · **Environment:** dev · **Maintainer:** repository maintainers (see GitHub)
This file uses a standard agent header (version, scope, model policy, reference docs, changelog), adapted for this **TypeScript/JavaScript monorepo**.
## Scope
| | |
|--|--|
| **Reads** | Repository tree as needed for the task: `gitnexus/`, `gitnexus-web/`, `eval/`, plugin packages, `.github/`, `.gitnexus/` when present, and docs. |
| **Writes** | Only paths required for the requested change; keep diffs minimal. Update lockfiles when dependencies change. |
| **Executes** | `npm`, `npx`, `node` under `gitnexus/` and `gitnexus-web/`; `uv run` for Python under `eval/` when applicable; shell utilities for documented CI/dev workflows. |
| **Off-limits** | User secrets (e.g. real `.env`), production deployment credentials, unrelated repositories, destructive git history operations without explicit human confirmation. |
## Model Configuration
- **Primary:** Pin in **Cursor** (Settings → model). Use a **named** model (e.g. GPT-5.2, Claude Sonnet 4.x). Avoid relying on **Auto** when reproducibility or audit trail matters.
- **Fallback:** As configured in Cursor or your organization (do not encode `latest` or wildcards in automation configs).
- **Notes:** The open-source GitNexus CLI indexer does not call an LLM. Optional Nexus AI in the web UI uses end-user provider keys and models.
## Execution Sequence (complex tasks)
Long sessions dilute instructions. For **multi-step** work, state up front:
1. Which rules in this file and **[GUARDRAILS.md](GUARDRAILS.md)** apply (and any relevant Signs).
2. Current **Scope** boundaries (Reads / Writes / Off-limits).
3. Which **validation commands** you will run (e.g. `cd gitnexus && npm test`, `npx tsc --noEmit`).
On very long threads, the human may add *“Remember: apply all AGENTS.md rules”* to re-weight rule tokens against context dilution.
## Claude Code hooks
Hooks enforce gates that prompts cannot. In **Claude Code**, **PreToolUse** hooks can block tools such as `git_commit` until checks pass. Adapt to this repo: e.g. `cd gitnexus && npm test` before commit.
## Context budget (Cursor / standards)
Generic “core standards” playbooks are often long and stack-specific. For this monorepo, commands and gotchas live under **Cursor Cloud specific instructions** below and in **[CONTRIBUTING.md](CONTRIBUTING.md)**. If always-on rules grow, split domain rules into **`.cursor/rules/*.mdc`** (globs). **Cursor:** project-wide rules live in **`.cursor/index.mdc`** (YAML frontmatter with `alwaysApply: true`). **Claude Code:** optionally load a **`STANDARDS.md`** only when needed (e.g. *“When writing new code, read STANDARDS.md”*) to save context.
## Reference Documentation
- **This repository:** **[ARCHITECTURE.md](ARCHITECTURE.md)**, **[CONTRIBUTING.md](CONTRIBUTING.md)**, **[GUARDRAILS.md](GUARDRAILS.md)**.
- **Cursor:** `.cursor/index.mdc` (always-on rules); optional `.cursor/rules/*.mdc` (glob-scoped). Legacy `.cursorrules` is deprecated — see `.cursor/index.mdc`.
- **Optional local files:** `NOTES.md` (short vendor-neutral project snapshot). For handoffs, keep notes local (e.g., a scratch file outside the repo) rather than committing `HANDOFF.md`.
- **GitNexus:** skills under `.claude/skills/gitnexus/`; machine-oriented rules in the `gitnexus:start` … `gitnexus:end` block below.
## Changelog
| Date | Version | Change |
|------|---------|--------|
| 2026-03-24 | 1.2.0 | Fixed gitnexus:start block duplication (was inlined in Reference Docs bullet). |
| 2026-03-23 | 1.1.0 | Updated agent instructions (sections, references, Cursor layout). |
| 2026-03-22 | 1.0.0 | Added structured agent header and changelog. |
---
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (2487 symbols, 6056 relationships, 188 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (3298 symbols, 7954 relationships, 185 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
@@ -99,3 +161,44 @@ To check whether embeddings exist, inspect `.gitnexus/meta.json` — the `stats.
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
<!-- gitnexus:end -->
## Cursor Cloud specific instructions
### Repository structure
This is a monorepo with two main products and supporting config packages:
| Component | Path | Purpose |
|-----------|------|---------|
| **GitNexus CLI/Core** | `gitnexus/` | Main product — TypeScript CLI, indexing pipeline, MCP server. Published to npm. |
| **GitNexus Web UI** | `gitnexus-web/` | React/Vite browser app — graph explorer + AI chat. Runs entirely in WASM. |
| Claude Plugin | `gitnexus-claude-plugin/` | Static config for Claude marketplace (no build). |
| Cursor Integration | `gitnexus-cursor-integration/` | Static config for Cursor editor (no build). |
| SWE-bench Eval | `eval/` | Python evaluation harness (optional; needs Docker + LLM API keys). |
### Running services
- **CLI/Core**: `cd gitnexus && npm run dev` (tsx watch mode) or `npm run build && node dist/cli/index.js <command>`
- **Web UI**: `cd gitnexus-web && npm run dev` (Vite on port 5173)
- **Backend mode**: `cd <indexed-repo> && node /workspace/gitnexus/dist/cli/index.js serve` (HTTP API on port 3741 by default)
### Testing
**CLI / Core (`gitnexus/`)**
- **Unit tests**: `cd gitnexus && npm test` (vitest, ~2000 tests)
- **Integration tests**: `cd gitnexus && npm run test:integration` (vitest, ~1850 tests). Two LadybugDB file-locking tests (`lbug-core-adapter`, `search-core`) may fail in containerized environments due to `/tmp` locking limitations — this is a known environment issue, not a code bug.
- **TypeScript check**: `cd gitnexus && npx tsc --noEmit`
**Web UI (`gitnexus-web/`)**
- **Unit tests**: `cd gitnexus-web && npm test` (vitest, ~200 tests)
- **E2E tests**: `cd gitnexus-web && E2E=1 npx playwright test` (Playwright, 5 tests — requires `gitnexus serve` + `npm run dev` running)
- **TypeScript check**: `cd gitnexus-web && npx tsc -b --noEmit`
No separate lint command is configured; TypeScript strict checking serves as the primary static analysis.
### Gotchas
- `npm install` in `gitnexus/` triggers `prepare` (builds via `tsc`) and `postinstall` (patches tree-sitter-swift). Native tree-sitter bindings require `python3`, `make`, and `g++` to be present.
- `tree-sitter-kotlin` and `tree-sitter-swift` are optional dependencies — install warnings for these are expected and non-blocking.
- The Web UI uses `vite-plugin-wasm` and requires `Cross-Origin-Opener-Policy`/`Cross-Origin-Embedder-Policy` headers for `SharedArrayBuffer` (handled automatically by Vite dev server).
- There is no ESLint/Prettier configuration in this repo.
+123
View File
@@ -0,0 +1,123 @@
# Architecture — GitNexus
This repository is a **monorepo** with two main products: the **CLI / MCP package** (`gitnexus/`) and the **browser UI** (`gitnexus-web/`). Supporting folders ship editor integrations and plugins without changing the core graph engine.
## Repository layout
| Path | Role |
|------|------|
| `gitnexus/` | Published npm package `gitnexus`: CLI, MCP server (stdio), local HTTP API for bridge mode, ingestion pipeline, LadybugDB graph, embeddings (optional). |
| `gitnexus-web/` | Vite + React UI: in-browser indexing (WASM), graph visualization, optional connection to `gitnexus serve`. |
| `.claude/`, `gitnexus-claude-plugin/`, `gitnexus-cursor-integration/` | Packaged **skills** and plugin metadata so agents discover the same workflows as documented in `AGENTS.md`. |
| `eval/` | Evaluation harnesses and docs for benchmarking tool usage. |
| `.github/` | CI workflows (quality, unit, integration, E2E) and composite actions. |
## End-to-end flow: index → graph → tools
1. **Ingestion** (`gitnexus analyze`)
- Entry: `gitnexus/src/cli/analyze.ts` → `runPipelineFromRepo` in `gitnexus/src/core/ingestion/pipeline.ts`.
- Walks the git working tree, parses supported languages via **Tree-sitter**, resolves imports/calls/inheritance, detects **communities** and **processes** (execution flows), and builds an in-memory **knowledge graph** (`gitnexus/src/core/graph/`).
- Output is loaded into **LadybugDB** under **`.gitnexus/`** at the repo root (`lbug/`, `meta.json`, etc.). Optional **FTS** indexes and **embeddings** attach to the same store.
- The repo is registered in **`~/.gitnexus/registry.json`** so MCP can find it from any working directory.
2. **Persistence & metadata**
- `gitnexus/src/storage/repo-manager.ts` — paths, registry, cleanup of legacy Kuzu artifacts.
- `gitnexus/src/core/lbug/lbug-adapter.ts` — graph load, queries, embedding restore batches.
3. **Query & agents**
- **MCP (stdio):** `gitnexus/src/cli/mcp.ts` → `startMCPServer` → `LocalBackend` (`gitnexus/src/mcp/local/local-backend.ts`) opens registered repos and serves **tools** from `gitnexus/src/mcp/tools.ts` and **resources** from `gitnexus/src/mcp/resources.ts`.
- **Bridge HTTP:** `gitnexus/src/cli/serve.ts` → Express app in `gitnexus/src/server/api.ts` (CORS-limited) exposes REST + MCP-over-HTTP for the web UI.
- **CLI tools (no MCP):** `gitnexus query`, `context`, `impact`, `cypher` in `gitnexus/src/cli/tool.ts` call the same backend for scripts and CI.
4. **Staleness**
- `gitnexus/src/mcp/staleness.ts` compares indexed `lastCommit` to `HEAD` and surfaces hints when the graph is behind git.
## MCP tools (summary)
| Tool | Purpose |
|------|---------|
| `list_repos` | Discover indexed repositories when more than one is registered. |
| `query` | Natural-language / keyword search over the graph (hybrid BM25 + optional vectors). |
| `cypher` | Ad hoc **Cypher** against the schema (see resource `gitnexus://repo/{name}/schema`). |
| `context` | Callers, callees, processes for one symbol (with disambiguation). |
| `impact` | Blast radius (upstream/downstream) with depth and risk summary. |
| `detect_changes` | Map git diffs to affected symbols and processes. |
| `rename` | Graph-assisted rename with `dry_run` preview (`graph` vs `text_search` confidence). |
## Where to change what
| If you are changing… | Start in… |
|----------------------|-----------|
| CLI commands / flags | `gitnexus/src/cli/` (`index.ts`, per-command modules). |
| Parsing or graph construction | `gitnexus/src/core/ingestion/` (pipeline, processors, resolvers, type-extractors). |
| Graph schema / DB access | `gitnexus/src/core/lbug/` (`schema.ts`, `lbug-adapter.ts`), `gitnexus/src/mcp/core/lbug-adapter.ts` if MCP-specific. |
| MCP protocol, tools, resources | `gitnexus/src/mcp/server.ts`, `tools.ts`, `resources.ts`. |
| Search ranking | `gitnexus/src/core/search/` (BM25, hybrid fusion). |
| Embeddings | `gitnexus/src/core/embeddings/`, phases in `analyze.ts`. |
| Wiki generation | `gitnexus/src/core/wiki/`. |
| Web UI behavior | `gitnexus-web/src/` (components, workers, graph client). |
| CI | `.github/workflows/*.yml`, `.github/actions/setup-gitnexus/`. |
## Known limitations
### Overloaded method resolution
Method and Constructor node IDs include an arity suffix (`#<paramCount>`) to
disambiguate overloaded methods. Two overloads with different parameter counts
produce distinct graph nodes: `Method:file:Class.method#1` vs
`Method:file:Class.method#2`.
**Same-arity overload disambiguation:** When two overloads share the same
parameter count but differ in types (e.g. `save(int)` vs `save(String)`), a
type-hash suffix `~type1,type2` is appended to produce distinct node IDs:
`Method:file:Class.save#1~int` vs `Method:file:Class.save#1~String`. The suffix
is only added when a same-arity collision is detected within a class and all
parameters have non-null type annotations. Languages without type info (Python,
Ruby, JS) fall back to arity-only IDs. TypeScript/JavaScript overload signatures
are intentionally excluded from type-hashing because they are declaration-only
contracts that should collapse to the implementation body's node ID. See issue
\#651.
**C++ const-qualified overload disambiguation:** Methods overloaded by const
qualification (e.g. `begin()` vs `begin() const`) are disambiguated via an
`isConst` property and a `$const` ID suffix appended to the const-qualified
variant when a non-const collision exists. The `$const` suffix appears after the
type-hash suffix: e.g. `Method:file:Container.begin#0$const`.
**Generic/template type preservation in type-hash:** The type-hash suffix uses
`rawType` (full AST text including generic/template args) rather than the
simplified `type` from `extractSimpleTypeName`. This means C++ template overloads
like `process(vector<int>)` vs `process(vector<string>)` produce distinct IDs:
`~vector<int>` vs `~vector<std::string>`. Java generic overloads like
`process(List<String>)` vs `process(List<Integer>)` are a compile error due to
type erasure, so this gap is theoretical for Java.
**ID stability on first overload:** Type and const tags are collision-only. When
a class has `save(int)` as its only `save` method, the ID is `save#1` (no tag).
Adding `save(String)` changes the original to `save#1~int`. This is correct for
fresh analysis but means IDs are not stable across overload additions. Future
incremental re-analysis should account for this.
**Variadic method matching:** When one side is variadic (`parameterCount`
undefined) and the other has a fixed count, `METHOD_IMPLEMENTS` edges are
emitted with confidence 0.7 instead of 1.0. Variadic methods like
`foo(String... args)` may superficially match `foo(String s)` by type but
are not guaranteed to be interchangeable across all languages (Java/Kotlin
accept this via varargs sugar; TypeScript, C#, Rust do not).
**Confidence tiering** for `METHOD_IMPLEMENTS` edges:
| Match quality | Confidence | When |
|---|---|---|
| Exact parameter types match | 1.0 | Both sides have `parameterTypes` arrays and they match |
| Arity (count) matches | 1.0 | Both sides have `parameterCount`, types unavailable |
| Variadic vs fixed | 0.7 | One side is variadic, other has fixed count |
| Lenient (insufficient info) | 0.7 | One or both sides lack type and count data |
## Related docs
- [MIGRATION.md](MIGRATION.md) — breaking changes and migration guidance.
- [RUNBOOK.md](RUNBOOK.md) — operational commands and recovery.
- [GUARDRAILS.md](GUARDRAILS.md) — safety boundaries for humans and agents.
- [TESTING.md](TESTING.md) — how to run tests.
- `AGENTS.md` / `CLAUDE.md` — agent workflows and tool usage expectations for **this** repo when indexed by GitNexus.
+13
View File
@@ -10,6 +10,19 @@ All notable changes to GitNexus will be documented in this file.
- Added automatic cleanup of stale KuzuDB index files
- LadybugDB v0.15 requires explicit VECTOR extension loading for semantic search
## [1.5.3] - 2026-04-01
### Added
- **TypeScript/JavaScript MethodExtractor config** — shared extraction config covering abstract methods, visibility modifiers, async/override keywords, decorators, rest/optional/destructured parameters, and return types (#588) — @compound-ai
### Fixed
- **Azure OpenAI compatibility** — use `max_completion_tokens` instead of deprecated `max_tokens` (newer models reject `max_tokens`); skip `temperature` for Azure provider (some models reject non-default values) (#618)
- **Simplified Azure interactive setup** — 3 prompts (endpoint, deployment, key) instead of 7 (#618)
- **Wiki HTML viewer script injection** — escape `</script>` in embedded JSON so LLM-generated markdown no longer breaks the viewer (#618)
- Ensure import rewrites survive npm publish lifecycle
## [1.4.0] - 2026-03-13
### Added
+54 -1
View File
@@ -1,7 +1,60 @@
<!-- version: 1.2.0 -->
<!--
Metadata: version, last reviewed, scope, model policy, reference docs, changelog.
Last updated: 2026-03-22
-->
Last reviewed: 2026-03-24
**Project:** GitNexus · **Environment:** dev · **Maintainer:** repository maintainers (see GitHub)
Follow **AGENTS.md** for the canonical rules; this file adds Claude Code–specific deltas. Cursor-specific notes live only in `AGENTS.md`.
## Scope
See the **Scope** table in [AGENTS.md](AGENTS.md) for read/write/execute/off-limits boundaries. Cursor-specific workflow notes also live only in AGENTS.md.
## Model Configuration
- **Primary:** Pin per **Claude Code** / Anthropic org policy (explicit model id). Do not rely on an unversioned `latest` alias for governed workflows.
- **Fallback:** As configured in Claude Code (organization default or user override).
- **Notes:** The GitNexus CLI analyzer does not call an LLM.
## Execution Sequence (complex tasks)
Same discipline as [AGENTS.md](AGENTS.md): before large multi-step work, state which **AGENTS.md** / **GUARDRAILS.md** rules apply, current **Scope**, and planned validation commands (`npm test`, `tsc`, etc.). When pausing, summarize progress in the chat or a **local** scratch file (do not add `HANDOFF.md` to the repo), then `/clear` and resume with that summary.
## Claude Code hooks
Prefer **PreToolUse** hooks for hard gates (e.g. tests before `git_commit`). Adapt hook commands to `gitnexus/` npm scripts.
## Context budget
If always-on instructions grow, load deep conventions via conditional reads (e.g. *“When writing new code, read STANDARDS.md”*) instead of pasting long blocks here. In Cursor, prefer `.cursor/index.mdc` plus optional `.cursor/rules/*.mdc` globs (see [AGENTS.md](AGENTS.md) § Context budget).
## Reference Documentation
- **This repository:** [AGENTS.md](AGENTS.md) (Cursor + monorepo notes), [ARCHITECTURE.md](ARCHITECTURE.md), [CONTRIBUTING.md](CONTRIBUTING.md), [GUARDRAILS.md](GUARDRAILS.md).
- **GitNexus:** `.claude/skills/gitnexus/`; MCP and indexed-repo rules live only in [AGENTS.md](AGENTS.md) (`gitnexus:start` … `gitnexus:end`). See **GitNexus rules** below.
## Changelog
| Date | Version | Change |
|------|---------|--------|
| 2026-03-24 | 1.2.0 | Removed duplicated gitnexus:start block and scope table; replaced with pointers to AGENTS.md. |
| 2026-03-23 | 1.1.0 | Updated agent instructions to match AGENTS.md. |
| 2026-03-22 | 1.0.0 | Added structured header and changelog. |
---
## GitNexus rules
GitNexus MCP rules are in the `<!-- gitnexus:start -->` … `<!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** — load that section when working with MCP tools or the graph index.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (2487 symbols, 6056 relationships, 188 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (3298 symbols, 7954 relationships, 185 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
+50
View File
@@ -0,0 +1,50 @@
# Contributing to GitNexus
How to propose changes, run checks locally, and open pull requests.
## License
This project uses the [PolyForm Noncommercial License 1.0.0](https://polyformproject.org/licenses/noncommercial/1.0.0/). By contributing, you agree your contributions are licensed under the same terms unless stated otherwise.
## Where to discuss
- **Issues & feature ideas:** use [GitHub Issues](https://github.com/abhigyanpatwari/GitNexus/issues) for the upstream repo, or your fork’s tracker if you work from a fork.
- **Community:** see the Discord link in the root [README.md](README.md).
## Development setup
1. Clone the repository.
2. **CLI / MCP package:** `cd gitnexus && npm install && npm run build`
3. **Web UI (if needed):** `cd gitnexus-web && npm install`
4. Run tests as described in [TESTING.md](TESTING.md).
## Branch and pull requests
- Use short-lived branches off the default branch of the repo you are targeting.
- Prefer **conventional commits** (short prefix + description), for example:
```text
feat: add graph export option
fix: correct MCP tool schema for query
test: cover cluster merge edge case
docs: clarify analyze flags
```
- **PR title:** `[area] Short description` (e.g. `[cli] Fix index refresh race`).
- **PR description:** what changed, why, how to verify (commands), and any risk or rollback notes.
## Before you open a PR
- [ ] Tests pass for the packages you touched (`gitnexus` and/or `gitnexus-web`).
- [ ] Typecheck passes: `npx tsc --noEmit` in `gitnexus/` and `npx tsc -b --noEmit` in `gitnexus-web/`.
- [ ] No secrets, tokens, or machine-specific paths committed.
- [ ] Documentation updated if behavior or public CLI/MCP contract changes.
- [ ] Pre-commit hook runs clean (`.husky/pre-commit` — typecheck + unit tests for staged packages).
## Code review
Maintainers may request changes for correctness, tests, performance, or consistency with existing patterns. Keeping diffs focused makes review faster.
## AI-assisted contributions
If you use coding agents, follow project context files (e.g. `AGENTS.md`, `CLAUDE.md`) and avoid drive-by refactors unrelated to the issue. Prefer incremental, test-backed changes.
+88
View File
@@ -0,0 +1,88 @@
# Guardrails — GitNexus (repo + agents)
Rules for **human contributors** and **AI agents** working on this codebase or publishing artifacts. These complement `AGENTS.md` / `CLAUDE.md` (which focus on GitNexus-in-GitNexus workflows).
## Scope (typical agent session)
When automating changes in this repository, treat scope as **least privilege**:
- **Read:** Source, tests, docs, public config as needed for the task.
- **Write:** Only files required for the requested fix or feature; avoid unrelated formatting or refactors.
- **Execute:** Tests, typecheck, and documented CLI commands; do not run destructive commands on user data outside the repo without explicit approval.
- **Off-limits:** Other people’s machines, production deployments you don’t own, and credentials you didn’t receive permission to use.
Adjust explicitly if the maintainer defines a different scope for a task.
---
## Non-negotiables
1. **Never commit secrets** — API keys, tokens, `.env` with real values, private URLs, or session cookies. Use `.env.example` with placeholders only.
2. **Never rename symbols with blind find-and-replace** when working in a GitNexus-indexed project — use the **`rename` MCP tool** with **`dry_run: true` first**, then review `graph` vs `text_search` edits. (There is no separate `gitnexus rename` CLI; renaming goes through MCP or editor integration.)
3. **Run impact analysis before editing shared symbols** — use **`impact`** (upstream) for functions/classes/methods others call; do not ignore **HIGH** / **CRITICAL** risk without maintainer sign-off.
4. **Prefer `detect_changes` before commit** — confirm diffs map to expected symbols/processes when the graph is available.
5. **Preserve embeddings** — if `.gitnexus/meta.json` shows embeddings, run `npx gitnexus analyze --embeddings` when refreshing the index; plain `analyze` can drop them.
---
## Signs (recurring failure patterns)
Use this format: **Trigger → Instruction → Reason**.
Append new Signs here when the same mistake repeats (e.g. CI broken twice the same way).
### Sign: Stale graph after edits
- **Trigger:** MCP or resources warn the index is behind `HEAD`, or code search doesn’t match latest commit.
- **Instruction:** Run `npx gitnexus analyze` from the repo root (plus `--embeddings` if the project used them).
- **Reason:** Tools query LadybugDB built at last analyze; git changes are invisible until re-indexed.
### Sign: Embeddings vanished after analyze
- **Trigger:** Semantic search quality drops; `stats.embeddings` in `.gitnexus/meta.json` is 0 after a refresh.
- **Instruction:** Re-run `npx gitnexus analyze --embeddings` and confirm `meta.json` reflects stored embeddings.
- **Reason:** Embedding generation is opt-in; analyze without the flag does not preserve prior vectors.
### Sign: MCP lists no repos
- **Trigger:** MCP stderr says no indexed repos.
- **Instruction:** Run `npx gitnexus analyze` in the target repository; verify `npx gitnexus list` shows it.
- **Reason:** The MCP server discovers repos via `~/.gitnexus/registry.json`, populated by analyze.
### Sign: Wrong repo in multi-repo setups
- **Trigger:** Query/impact results clearly belong to another project.
- **Instruction:** Call `list_repos`, then pass **`repo`** on subsequent tools (or use per-workspace MCP config).
- **Reason:** Default target may be ambiguous when multiple repos are registered.
### Sign: LadybugDB lock / “database busy”
- **Trigger:** Errors opening `.gitnexus/lbug` while MCP and analyze both run.
- **Instruction:** Stop overlapping processes; one writer at a time. Retry analyze or restart MCP.
- **Reason:** Embedded DB expects single-process ownership of the store.
---
## Publishing & supply chain
- **npm:** Do not publish from unreviewed automation; follow maintainer release process. Bump version intentionally; tag releases to match `package.json`.
- **Dependencies:** Prefer minimal, auditable changes to `package.json`; run tests and CI after lockfile updates.
- **License:** This project ships under **PolyForm Noncommercial 1.0.0** — do not relicense or imply a different license in docs or metadata without maintainer approval.
---
## Escalation
Stop and ask a **human maintainer** when:
- Impact analysis shows **HIGH** / **CRITICAL** risk and the task still requires the change.
- You need to alter **CI**, **release**, or **security-sensitive** config.
- Requirements conflict (e.g. “speed up analyze” vs “must keep all embeddings on huge repo”).
- You are unsure whether data loss is acceptable (`clean`, forced migrations, schema changes).
---
## Related docs
- [ARCHITECTURE.md](ARCHITECTURE.md) — components and data flow.
- [RUNBOOK.md](RUNBOOK.md) — commands for recovery.
- [CONTRIBUTING.md](CONTRIBUTING.md) — PR and commit expectations.
+27
View File
@@ -0,0 +1,27 @@
# Migration Guide
## OVERRIDES → METHOD_OVERRIDES (PR #642)
The `OVERRIDES` relationship type has been renamed to `METHOD_OVERRIDES` for
consistency with the new `METHOD_IMPLEMENTS` edge type.
### Do I need to migrate?
**No.** Backward compatibility is handled automatically at runtime:
- `local-backend.ts` dual-reads both `OVERRIDES` and `METHOD_OVERRIDES` in all
impact-analysis and context queries. Existing stored graphs with `OVERRIDES`
edges continue to return correct results without any manual intervention.
- The `REL_TYPES` array in `schema-constants.ts` includes both names so Cypher
queries that reference either will work.
### What happens on re-index?
Running `npx gitnexus analyze` on a repository produces `METHOD_OVERRIDES`
edges going forward. The old `OVERRIDES` edges are replaced as part of the
normal full re-index.
### When will the legacy alias be removed?
The `OVERRIDES` compat alias will remain until a future major version. Removal
will be announced in this file and in the changelog before it happens.
+62 -14
View File
@@ -19,6 +19,8 @@
<img src="https://img.shields.io/badge/License-PolyForm%20Noncommercial-blue.svg" alt="License: PolyForm Noncommercial"/>
</a>
<p><strong>Enterprise (SaaS & Self-hosted)</strong> - <a href="https://akonlabs.com">akonlabs.com</a></p>
</div>
**Building nervous system for agent context.**
@@ -50,7 +52,7 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
| **For** | Daily development with Cursor, Claude Code, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
| **Install** | `npm install -g gitnexus` | No install —[gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Privacy** | Everything local, no network | Everything in-browser, no server |
@@ -59,6 +61,36 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
---
## Enterprise
GitNexus is available as an **enterprise offering** - either as a fully managed **SaaS** or a **self-hosted** deployment. Also available for **commercial use** of the OSS version with proper licensing.
Enterprise includes:
- **PR Review** - automated blast radius analysis on pull requests
- **Auto-updating Code Wiki** - always up-to-date documentation (Code Wiki is also available in OSS)
- **Auto-reindexing** - knowledge graph stays fresh automatically
- **Multi-repo support** - unified graph across repositories
- **OCaml support** - additional language coverage
- **Priority feature/language support** - request new languages or features
**Upcoming:**
- Auto regression forensics
- End-to-end test generation
👉 Learn more at [akonlabs.com](https://akonlabs.com)
💬 For commercial licensing or enterprise inquiries, ping us on [Discord](https://discord.gg/AAsRVT6fGb) or drop an email at founders@akonlabs.com
---
## Development
- [ARCHITECTURE.md](ARCHITECTURE.md) — packages, index → graph → MCP flow, where to change code
- [RUNBOOK.md](RUNBOOK.md) — analyze, embeddings, stale index, MCP recovery, CI snippets
- [GUARDRAILS.md](GUARDRAILS.md) — safety rules and operational “Signs” for contributors and agents
- [CONTRIBUTING.md](CONTRIBUTING.md) — license, setup, commits, and pull requests
- [TESTING.md](TESTING.md) — test commands for `gitnexus` and `gitnexus-web`
## CLI + MCP (recommended)
The CLI indexes your repository and runs an MCP server that gives AI agents deep codebase awareness.
@@ -87,7 +119,6 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
| **Codex** | Yes | Yes | — | MCP + Skills |
| **Windsurf** | Yes | — | — | MCP |
| **OpenCode** | Yes | Yes | — | MCP + Skills |
| **Codex** | Yes | — | — | MCP |
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that enrich searches with graph context + PostToolUse hooks that auto-reindex after commits.
@@ -139,8 +170,8 @@ codex mcp add gitnexus -- npx -y gitnexus@latest mcp
{
"mcp": {
"gitnexus": {
"command": "npx",
"args": ["-y", "gitnexus@latest", "mcp"]
"type": "local",
"command": ["gitnexus", "mcp"]
}
}
}
@@ -157,13 +188,14 @@ args = ["-y", "gitnexus@latest", "mcp"]
### CLI Commands
```bash
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
gitnexus serve # Start local HTTP server (multi-repo) for web UI connection
gitnexus list # List all indexed repositories
@@ -173,11 +205,21 @@ gitnexus clean --all --force # Delete all indexes
gitnexus wiki [path] # Generate repository wiki from knowledge graph
gitnexus wiki --model <model> # Wiki with custom LLM model (default: gpt-4o-mini)
gitnexus wiki --base-url <url> # Wiki with custom LLM API base URL
# Repository groups (multi-repo / monorepo service tracking)
gitnexus group create <name> # Create a repository group
gitnexus group add <name> <repo> # Add a repo to a group
gitnexus group remove <name> <repo> # Remove a repo from a group
gitnexus group list [name] # List groups, or show one group's config
gitnexus group sync <name> # Extract contracts and match across repos/services
gitnexus group contracts <name> # Inspect extracted contracts and cross-links
gitnexus group query <name> <q> # Search execution flows across all repos in a group
gitnexus group status <name> # Check staleness of repos in a group
```
### What Your AI Agent Gets
**7 tools** exposed via MCP:
**16 tools** exposed via MCP (11 per-repo + 5 group):
| Tool | What It Does | `repo` Param |
| ------------------ | ----------------------------------------------------------------- | -------------- |
@@ -188,6 +230,11 @@ gitnexus wiki --base-url <url> # Wiki with custom LLM API base URL
| `detect_changes` | Git-diff impact — maps changed lines to affected processes | Optional |
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
| `cypher` | Raw Cypher graph queries | Optional |
| `group_list` | List configured repository groups | — |
| `group_sync` | Extract contracts and match across repos/services | — |
| `group_contracts`| Inspect extracted contracts and cross-links | — |
| `group_query` | Search execution flows across all repos in a group | — |
| `group_status` | Check staleness of repos in a group | — |
> When only one repo is indexed, the `repo` parameter is optional. With multiple repos, specify which one: `query({query: "auth", repo: "my-app"})`.
@@ -283,8 +330,8 @@ Or run locally:
```bash
git clone https://github.com/abhigyanpatwari/gitnexus.git
cd gitnexus/gitnexus-web
npm install
cd gitnexus/gitnexus-shared && npm install && npm run build
cd ../gitnexus-web && npm install
npm run dev
```
@@ -367,6 +414,7 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
@@ -531,7 +579,7 @@ The wiki generator reads the indexed graph structure, groups files into modules
- [X] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
- [X] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
- [X] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
- [X] Multi-Repo MCP, Zero-Config Setup, 13 Language Support
- [X] Multi-Repo MCP, Zero-Config Setup, 14 Language Support
- [X] Community Detection, Process Detection, Confidence Scoring
- [X] Hybrid Search, Vector Index
+163
View File
@@ -0,0 +1,163 @@
# Runbook — GitNexus
Short, copy-paste operations for **local development**, **MCP**, and **CI**. Commands assume a Unix shell; on Windows use Git Bash or equivalent paths.
## Prerequisites
- **Node.js** ≥ 20 (`gitnexus-web/package.json` `engines`).
- **Git** (analyze requires a git repository).
- From repo root, install and build the CLI package:
```bash
cd gitnexus
npm install
npm run build
```
Use `npx gitnexus …` from any path after global/published install, or `node dist/cli/index.js …` when developing from `gitnexus/` with a local build.
---
## Index out of date / “stale” tools
**Symptom:** MCP or resources warn the index is behind `HEAD`, or results don’t reflect recent commits.
**Fix (from the target repo root):**
```bash
npx gitnexus analyze
```
**Force full rebuild** (same commit but suspect corruption or changed ignore rules):
```bash
npx gitnexus analyze --force
```
**Check status:**
```bash
npx gitnexus status
```
**List what MCP knows about:**
```bash
npx gitnexus list
```
---
## Embeddings
**First time with vectors** (slower, more disk/RAM):
```bash
npx gitnexus analyze --embeddings
```
**Important:** If you already had embeddings, **always** pass `--embeddings` on later analyzes, or they can be dropped. See `stats.embeddings` in `.gitnexus/meta.json` (0 means none).
**Large repos:** Analyze may skip or limit embedding work when node counts are very high; watch CLI output.
---
## MCP: no repos / empty tools
**Symptom:** `GitNexus: No indexed repos yet` on stderr when starting MCP.
**Fix:** In each project you want indexed:
```bash
cd /path/to/repo
npx gitnexus analyze
```
Restart the editor MCP session if needed. The server **refreshes the registry lazily**; new analyzes are picked up without necessarily reinstalling MCP.
**Symptom:** Wrong repo when multiple are indexed — pass `repo` on tools or use `list_repos` first.
---
## Clean slate (corrupt or huge `.gitnexus`)
**Current repo only** (prompts for confirmation):
```bash
npx gitnexus clean
```
**Skip confirmation:**
```bash
npx gitnexus clean --force
```
**All registered repos:**
```bash
npx gitnexus clean --all --force
```
Then re-run `npx gitnexus analyze` (and `--embeddings` if you need vectors).
---
## Local bridge for the web UI
```bash
cd gitnexus
npx gitnexus serve
# default http://127.0.0.1:4747 — see serve --help for port/host
```
Use when the browser UI should talk to **local** indexed repos instead of WASM-only mode.
---
## CLI equivalents of MCP tools
Useful for debugging without an editor:
```bash
cd gitnexus
npx gitnexus query "authentication flow" --repo MyRepo
npx gitnexus context SomeSymbol --repo MyRepo
npx gitnexus impact SomeSymbol --direction upstream --repo MyRepo
npx gitnexus cypher "MATCH (n) RETURN count(n) LIMIT 1" --repo MyRepo
```
---
## CI failures (contributors)
Orchestrator: `.github/workflows/ci.yml`.
| Job | Typical local repro |
|-----|---------------------|
| **quality** | `cd gitnexus && npx tsc --noEmit` |
| **unit-tests** | `cd gitnexus && npx vitest run test/unit` |
| **integration** | `cd gitnexus && npx vitest run test/integration` (see workflow matrix for groups) |
| **e2e** | Triggered when `gitnexus-web/` changes; `cd gitnexus-web && E2E=1 npx playwright test` (requires `gitnexus serve` + `npm run dev`) |
**Note:** Pushes that touch only certain markdown paths may be skipped by `paths-ignore` in CI — see workflow file for exact patterns.
---
## Memory / analyze crashes
Analyze re-execs Node with a **large old-space heap** when needed (`analyze.ts`). If you still OOM on huge repos, close other processes, avoid `--embeddings` for a first pass, or analyze a smaller path if supported by your workflow.
---
## LadybugDB / lock errors
Only one process should open a repo’s `.gitnexus/lbug` store at a time. If MCP and a second `analyze` run conflict, stop one process, then retry `analyze` or restart MCP.
---
## Where to dig deeper
- Architecture overview: [ARCHITECTURE.md](ARCHITECTURE.md)
- Agent safety rules: [GUARDRAILS.md](GUARDRAILS.md)
- Tests: [TESTING.md](TESTING.md)
+95
View File
@@ -0,0 +1,95 @@
# Testing — GitNexus
How we structure tests and which commands to run locally and in CI.
## Packages
| Package | Path | Runner | Notes |
| -------------- | -------------- | -------- | ------------------------------ |
| CLI + MCP core | `gitnexus/` | Vitest | Primary test surface in CI |
| Web UI | `gitnexus-web/`| Vitest | Unit/component tests |
| Web UI E2E | `gitnexus-web/`| Playwright | Run when changing UI flows |
## Commands (local)
From repository root, unless noted:
**`gitnexus` (CLI / library)**
```bash
cd gitnexus
npm install
npm run build
npm test # unit: vitest run test/unit
npm run test:integration # integration suite
npm run test:all
npm run test:coverage
npx tsc --noEmit # typecheck (matches CI)
```
**`gitnexus-web`**
```bash
cd gitnexus-web
npm install
npm test # unit tests (vitest)
npx tsc -b --noEmit # typecheck (matches CI)
npm run test:coverage
npm run test:e2e # Playwright (requires gitnexus serve + npm run dev)
```
## Pre-commit hook
A husky pre-commit hook (`.husky/pre-commit`) runs automatically on every `git commit`:
- **`gitnexus-web/` files staged** → `tsc -b --noEmit` + `vitest run`
- **`gitnexus/` files staged** → `tsc --noEmit` + `vitest run --project default`
Skip with `git commit --no-verify` (use sparingly).
## Test categories
- **Unit** — Pure logic, parsers, graph/query helpers; fast; no network.
- **Integration** — Real combinations (filesystem, MCP wiring, larger pipelines) as already organized under `gitnexus/test/integration`.
- **Eval-style / golden sets** — For agent- or classification-style behavior, keep labeled inputs and expected outputs (JSON or table-driven tests) and run them in CI when relevant.
- **E2E (web)** — Critical user paths only; prefer `data-testid` attributes for stable selectors. Tests run against real backend (`gitnexus serve`) and Vite dev server.
## Performance metrics (targets)
Set targets to match team expectations, then tune to this repo’s CI reality:
| Metric | Target (initial) | Notes |
| ------------------- | ---------------- | ------------------------------------------ |
| Unit coverage | Align with CI | CI runs Vitest with coverage in `gitnexus` |
| Unit wall time | Fast PR feedback | Use `vitest run test/unit` for tight loop |
| Integration duration| &lt; few minutes | Guard heavy tests with env flags if needed |
## Regression testing
Re-run the full relevant suite when:
- Prompt or agent-behavior documentation changes (if tests encode behavior)
- Model or embedding-related code paths change
- Graph schema, query contracts, or MCP tool shapes change
- Dependencies with parsing or runtime impact upgrade
## CI integration
GitHub Actions (`.github/workflows/ci.yml`) orchestrate:
- **`ci-quality.yml`** — `tsc --noEmit` for `gitnexus/` + `tsc -b --noEmit` for `gitnexus-web/`
- **`ci-tests.yml`** — `vitest run` with coverage (ubuntu) + cross-platform (macOS, Windows)
- **`ci-e2e.yml`** — Playwright E2E tests, gated on `gitnexus-web/**` changes
Local checks before pushing:
```bash
cd gitnexus && npx tsc --noEmit && npm test
cd ../gitnexus-web && npx tsc -b --noEmit && npm test
```
Or rely on the pre-commit hook which runs these automatically for staged files.
## User acceptance / beta (optional)
For staged releases or UI betas: deploy to a staging environment, collect structured feedback, watch errors and latency, then iterate before a wider release.
+13 -1
View File
@@ -84,13 +84,17 @@ sequenceDiagram
end
```
The return type `CopyExpansionResult` contains `expandedContent` and `copyResolutions`. The `expansionDepth` field has been removed from the return type (it was unused by callers).
COPY statement line numbers in `CopyResolution` are 1-based (consistent with the preprocessor's line numbering). The splice operation that replaces COPY lines with expanded content adjusts for 0-based array indexing internally.
## Cycle Detection
Circular COPY references (e.g., copybook A includes copybook B which includes copybook A) are detected and handled:
1. Each expansion chain maintains a `visited` set of resolved copybook paths
2. If a copybook path is already in the visited set, the expansion is skipped
3. A `warnedCircular` set (shared across all files in a chunk) deduplicates warning messages
3. A `warnedCircular` set (internal to `expandCopies()`, not a parameter) deduplicates warning messages within a single file expansion
Known circular copybooks in PROJECT-NAME: `ANAZI`, `ANDIP`, `QDIPE` (self-referential includes).
@@ -98,6 +102,10 @@ Known circular copybooks in PROJECT-NAME: `ANAZI`, `ANDIP`, `QDIPE` (self-refere
Nested COPY expansion is limited to **10 levels** (`DEFAULT_MAX_DEPTH`). If a COPY chain exceeds this depth, a warning is logged and the remaining COPY statements are left unexpanded.
## Max Total Expansions
A breadth amplification guard caps the total number of COPY expansions across all branches within a single file to **500** (`MAX_TOTAL_EXPANSIONS`). This prevents exponential blowup from diamond-shaped COPY graphs where N copybooks each include N other copybooks. Once the limit is reached, further COPY statements in that file are left unexpanded and a single warning is logged.
## REPLACING Application Detail
The REPLACING engine works by scanning all COBOL identifiers (matching `\b[A-Z][A-Z0-9-]*\b`) in the copybook content and applying each replacement rule:
@@ -139,6 +147,10 @@ The expansion runs **per chunk**, after file content is read but before dispatch
3. Only programs (not copybooks themselves) undergo expansion
4. The expanded content replaces the original content in-place before worker dispatch
## Inline Comment Handling
The copy expander's `stripInlineComment()` helper is quote-aware: pipe characters (`|`) inside single- or double-quoted strings are preserved. This matches the same quote-aware logic used by the preprocessor.
## Source Files
- `gitnexus/src/core/ingestion/cobol-copy-expander.ts` -- `expandCopies()`, `parseReplacingClause()`, `applyReplacing()`
+51 -4
View File
@@ -100,10 +100,12 @@ Tables are extracted from SQL clauses:
| Clause Pattern | Example |
|----------------|---------|
| `FROM <table>` | `SELECT * FROM EMPLOYEES` |
| `INTO <table>` | `INSERT INTO EMPLOYEES` |
| `INSERT INTO <table>` | `INSERT INTO EMPLOYEES` |
| `UPDATE <table>` | `UPDATE EMPLOYEES SET ...` |
| `JOIN <table>` | `LEFT JOIN DEPARTMENTS ON ...` |
Note: The `INTO` pattern is restricted to `INSERT INTO` to avoid false positives from `FETCH ... INTO :host-var` and `SELECT ... INTO :host-var` statements, where `INTO` introduces host variables rather than table names.
### Cursor Detection
```cobol
@@ -242,21 +244,66 @@ The USING clause identifies parameters received by the program from its caller.
## MOVE Statements
MOVE statements are extracted but currently only stored in the regex results (not emitted as graph edges):
MOVE statements produce `ACCESSES` edges in the graph:
```cobol
MOVE WK-NAME TO OUT-NAME.
MOVE CORRESPONDING WK-INPUT TO WK-OUTPUT.
MOVE CORR WK-IN TO WK-OUT.
```
### Extraction Details
- Source and target identifiers are captured
- `CORRESPONDING` keyword is tracked (bulk field-by-field move)
- `CORRESPONDING` and its abbreviation `CORR` are both recognized (bulk field-by-field move)
- Figurative constants (SPACES, ZEROS, LOW-VALUES, HIGH-VALUES, QUOTES, ALL) are skipped
- The enclosing paragraph (`caller`) is tracked for context
DATA_FLOW edges from MOVE statements are reserved for a future release.
### MOVE CORRESPONDING / CORR Edge Reasons
MOVE CORRESPONDING (and CORR) produces distinct edge reasons to differentiate from simple MOVE:
| Edge | Reason (simple MOVE) | Reason (CORRESPONDING/CORR) |
|------|---------------------|-----------------------------|
| Read (source) | `cobol-move-read` | `cobol-move-corresponding-read` |
| Write (target) | `cobol-move-write` | `cobol-move-corresponding-write` |
This distinction allows queries to find bulk field-by-field moves separately from simple variable assignments.
## GO TO DEPENDING ON
The `GO TO` statement with multiple targets and a `DEPENDING ON` clause is a computed branch:
```cobol
GO TO PARA-1 PARA-2 PARA-3
DEPENDING ON WK-SELECTOR.
```
All target paragraph names are extracted and emitted as separate `gotos` entries. Each target produces a `CALLS` edge in the graph (same semantics as PERFORM). The `DEPENDING ON` variable is not currently tracked as a data-flow dependency.
## SORT INPUT/OUTPUT PROCEDURE
SORT and MERGE statements can specify procedural entry points instead of file-based I/O:
```cobol
SORT SORT-FILE ON ASCENDING KEY SORT-KEY
INPUT PROCEDURE IS PREPARE-INPUT
OUTPUT PROCEDURE IS FORMAT-OUTPUT.
```
`INPUT PROCEDURE IS` and `OUTPUT PROCEDURE IS` targets are extracted as control-flow targets (same as PERFORM). They produce `performs` entries and corresponding `CALLS` edges in the graph.
## Fixed-Format Literal Continuation
In fixed-format COBOL, string literals can span multiple lines using the continuation indicator (`-` in column 7). When a continuation line starts with a quote character, the extractor joins it with the predecessor by removing the trailing quote from the previous line and the opening quote from the continuation:
```
Line N: MOVE "THIS IS A LONG STRI
Line N+1 (cont): - "NG VALUE" TO WK-FIELD.
Merged: MOVE "THIS IS A LONG STRING VALUE" TO WK-FIELD.
```
The trailing `"` on line N and the opening `"` on line N+1 are both removed, producing a seamless literal. If no matching quote is found on the predecessor line, the continuation is appended as-is.
## Source Files
+29
View File
@@ -45,6 +45,31 @@ GITNEXUS_COBOL_DIRS=s,c,wfproc GITNEXUS_VERBOSE=1 node --max-old-space-size=8192
/path/to/gitnexus/dist/cli/index.js analyze --force
```
## Open-Source Benchmarks
### CardDemo (AWS)
| Metric | Value |
| ------ | ----- |
| Graph nodes | 12,323 |
| Graph edges | 8,893 |
| Total time | 7.4s |
### ACAS
| Metric | Value |
| ------ | ----- |
| Graph nodes | 14,016 |
| Graph edges | 15,452 |
| Total time | 9.3s |
### Micro-Benchmark (Single-File Extraction)
| Metric | Value |
| ------ | ----- |
| Per-iteration | 0.65ms |
| Throughput | ~382K lines/sec |
## Worker Pool Tuning
### Sub-Batch Size
@@ -109,6 +134,10 @@ To increase the cap for specific needs, modify the `MAX_DATA_ITEMS_PER_FILE` con
## Memory Management
### COPY Expansion Breadth Guard
A per-file `MAX_TOTAL_EXPANSIONS = 500` limit prevents exponential blowup from diamond-shaped COPY graphs (e.g., N copybooks each containing N COPY statements). Once the limit is reached, further COPY statements in that file are left unexpanded. See [copy-expansion.md](copy-expansion.md) for details.
### COPY Expansion Memory
All copybook content is loaded upfront into a Map before chunk processing begins. For PROJECT-NAME:
+24 -4
View File
@@ -71,11 +71,13 @@ Indicator col 7
buffer as new pending logical line
```
After all lines are processed, the final pending line is flushed, along with any accumulated SELECT statement.
After all lines are processed, the final pending line is flushed, along with any accumulated SELECT statement, SORT/MERGE accumulator, and any open EXEC block (truncated file without `END-EXEC`).
### Inline Comment Stripping
Enterprise COBOL (particularly Italian dialect) uses the pipe character `|` as an inline comment marker. Everything from `|` to end of line is stripped before processing.
Enterprise COBOL (particularly Italian dialect) uses the pipe character `|` as an inline comment marker. The `stripInlineComment()` helper is **quote-aware**: it tracks whether the scan position is inside a single- or double-quoted string and only treats `|` as a comment marker when outside quotes. Pipe characters inside string literals are preserved.
Free-format `*>` inline comment stripping uses the same quote-aware approach: the scanner walks character by character, toggling quote state, and only recognizes `*>` as a comment marker when not inside a quoted string.
### Patch Marker Handling
@@ -111,7 +113,7 @@ All patterns are compiled once as module-level constants and reused across calls
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_SELECT_START` | `\bSELECT\s+([A-Z][A-Z0-9-]+)` | File SELECT start | `SELECT MASTER-FILE` |
| `RE_SELECT_START` | `\bSELECT\s+(?:OPTIONAL\s+)?([A-Z][A-Z0-9-]+)` | File SELECT start (with optional `SELECT OPTIONAL` support) | `SELECT MASTER-FILE`, `SELECT OPTIONAL TRANS-FILE` |
SELECT statements are accumulated across multiple lines until a period terminator is found, then parsed for ASSIGN, ORGANIZATION, ACCESS, RECORD KEY, and FILE STATUS clauses.
@@ -135,7 +137,9 @@ The trailing clauses of `RE_DATA_ITEM` are parsed by `parseDataItemClauses()` fo
| `RE_PERFORM` | `\bPERFORM\s+([A-Z][A-Z0-9-]+)(?:\s+THRU\s+([A-Z][A-Z0-9-]+))?` | PERFORM call | `PERFORM CALC-TAX THRU CALC-TAX-EXIT` |
| `RE_PROC_USING` | `\bPROCEDURE\s+DIVISION\s+USING\s+([\s\S]*?)(?:\.\|$)` | USING parameters | `PROCEDURE DIVISION USING WK-PARAM` |
| `RE_ENTRY` | `\bENTRY\s+"([^"]+)"(?:\s+USING\s+([\s\S]*?))?(?:\.\|$)` | ENTRY point | `ENTRY "SUBPROG" USING WK-DATA` |
| `RE_MOVE` | `\bMOVE\s+(CORRESPONDING\s+)?([A-Z][A-Z0-9-]+)\s+TO\s+([A-Z][A-Z0-9-]+)` | MOVE statement | `MOVE WK-NAME TO OUT-NAME` |
| `RE_MOVE` | `\bMOVE\s+((?:CORRESPONDING\|CORR)\s+)?([A-Z][A-Z0-9-]+)\s+TO\s+(.+)` | MOVE statement (supports CORR abbreviation and multi-target) | `MOVE WK-NAME TO OUT-NAME`, `MOVE CORR WK-IN TO WK-OUT` |
The USING parameter list (`RE_PROC_USING`) is split on `\bRETURNING\b` before tokenization -- any RETURNING clause and everything after it is excluded from the parameter list (`.split(/\bRETURNING\b/i)[0]`).
Note: `RE_PROC_SECTION` and `RE_PROC_PARAGRAPH` require exactly 7 spaces of leading indentation (COBOL area A starting at column 8). This is the standard COBOL paragraph indentation.
@@ -149,6 +153,22 @@ These patterns are checked regardless of current division:
| `RE_COPY_UNQUOTED` | `\bCOPY\s+([A-Z][A-Z0-9-]+)(?:\s\|\.)` | COPY (unquoted) | `COPY CPSESP.` |
| `RE_COPY_QUOTED` | `\bCOPY\s+"([^"]+)"(?:\s\|\.)` | COPY (quoted) | `COPY "WORKGRID.CPY".` |
### SORT/MERGE Support
| Constant | Purpose |
|----------|---------|
| `SORT_CLAUSE_NOISE` | Set of SORT/MERGE clause keywords filtered from USING/GIVING file lists: `ON`, `ASCENDING`, `DESCENDING`, `KEY`, `WITH`, `DUPLICATES`, `IN`, `ORDER`, `COLLATING`, `SEQUENCE`, `IS`, `THROUGH`, `THRU`, `INPUT`, `OUTPUT`, `PROCEDURE` |
SORT and MERGE statements are accumulated across multiple lines (like SELECT) until a period terminator is found, then parsed for USING/GIVING file lists and INPUT/OUTPUT PROCEDURE targets. The `flushSort()` helper encapsulates the flush-and-parse logic, mirroring the existing `flushSelect()` pattern. Both helpers are called at EOF to handle truncated files.
### GO TO Multi-Target
`RE_GOTO` captures all paragraph names in a `GO TO` statement, including the multi-target form `GO TO p1 p2 p3 DEPENDING ON x`. The captured group contains all target names (space-separated), which are split into individual targets. Each target produces a separate `gotos` entry.
### PROGRAM-ID Detection
PROGRAM-ID is detected regardless of the current division state. This handles sibling programs that appear after `END PROGRAM` and omit the `IDENTIFICATION DIVISION` header -- the extractor will still capture the PROGRAM-ID and push a new program boundary.
### EXEC Block Patterns
| Constant | Pattern | Purpose | Example Match |
@@ -0,0 +1,326 @@
---
title: "feat: Complete COBOL language feature coverage for maximum knowledge graph value"
type: feat
status: active
date: 2026-03-26
origin: Feature audit from v3-integration-architect agent (session 8642401e)
---
## Enhancement Summary
**Deepened on:** 2026-03-26
**Research agents used:** COBOL expert (Phase 1+2), graph value analyst, codebase explorer
**Sections enhanced:** Phase 1 (5 features), Phase 2 (4 features), graph value ranking
### Key Improvements from Research
1. **CALL USING** is the #1 highest-value edge type (9.2/10) — fixes ~40% of missing caller references
2. **EXEC DLI** requires dual-interface support (EXEC DLI + CBLTDLI CALL) for full IMS coverage
3. **DECLARATIVES** is lowest-risk Phase 2 item — existing section/paragraph detection already captures structure
4. **SET TO TRUE** accounts for 80-90% of all SET statements — prioritize this form
5. **INSPECT** needs multi-line accumulator (like SORT) — can span 5+ continuation lines
6. **Graph value ranking**: cobol-call-using (9.2) > cobol-error-handler (9.0) > dli-gu (8.2) > cobol-string (6.2)
### New Edge Cases Discovered
- CALL USING supports mixed modes: `USING BY REFERENCE WS-A BY CONTENT WS-B BY VALUE WS-C`
- CALL USING `ADDRESS OF` and `OMITTED` must be filtered from parameter lists
- EXEC DLI can have multiple SEGMENT levels in hierarchical retrieval (use matchAll)
- DECLARATIVES can have multiple USE sections (one per file + catch-all for INPUT/OUTPUT/I-O/EXTEND)
- INSPECT TALLYING can have multiple counters in a single statement
- STRING/UNSTRING can span multiple lines (need accumulator pattern)
---
# Complete COBOL Language Feature Coverage
## Overview
Implement the remaining 25 unhandled COBOL language features and fix 10 partial features to achieve ~95% coverage (up from 71.9%). The goal is to build the richest possible knowledge graph from COBOL codebases, enabling a future `modernize` MCP command (out of scope for this plan) that would use the graph to assist with COBOL-to-modern-language migration.
## Problem Statement
The COBOL processor currently handles 54 of 89 applicable language features (71.9%). The 25 unhandled features represent real data loss in the knowledge graph:
- **Cross-program data flow** is invisible (CALL ... USING parameters not extracted)
- **IMS/DB programs** produce empty graphs (EXEC DLI not recognized)
- **String transformation logic** is invisible (STRING/UNSTRING/INSPECT not tracked)
- **SQL copybook dependencies** are missing (EXEC SQL INCLUDE not mapped)
- **Error handling flows** are lost (DECLARATIVES/USE AFTER not captured)
## Proposed Solution
Implement features in 4 phases, ordered by graph value density (edges created per LOC of implementation). Each phase is independently shippable and testable.
## Technical Approach
### Phase 1: High-Value Data Flow Edges (~150 LOC, ~8 new edge types)
The highest-ROI features: they create new ACCESSES and IMPORTS edges that directly improve impact analysis.
**Critical research finding**: Multi-line statement accumulation is the dominant challenge. CALL USING, STRING/UNSTRING, and multi-line data item clauses all span multiple lines in production COBOL. The free-format path processes each line independently — these features need statement accumulators (like SORT/SELECT) or the free-format path needs multi-line awareness. Estimated LOC increased from 110 to 150 to account for accumulator infrastructure.
#### 1.1 EXEC SQL INCLUDE -> IMPORTS edges
- **File:** `cobol-preprocessor.ts` (parseExecSqlBlock)
- **What:** Detect `INCLUDE` as the operation, extract member name, emit as a `copies[]` entry
- **Graph:** IMPORTS edge from File to included copybook/SQLCA with reason `sql-include`
- **Tests:** Unit test for `EXEC SQL INCLUDE SQLCA END-EXEC` and `EXEC SQL INCLUDE CUSTCOPY END-EXEC`
**Research insights (EXEC SQL INCLUDE):**
- DB2 member names can contain underscores: `EXEC SQL INCLUDE CUST_TBL_DCL END-EXEC` — regex must use `[A-Z][A-Z0-9_-]+`
- Quoted literal form: `EXEC SQL INCLUDE 'DBRMLIB.MEMBER' END-EXEC` (z/OS PDS qualified name)
- SQLCA/SQLDA are DB2 builtins — won't resolve to repo files. Emit unresolved IMPORTS edge (still valuable)
- No REPLACING support on EXEC SQL INCLUDE (unlike COPY)
- Add `INCLUDE` to `OP_MAP` in `parseExecSqlBlock`; extract member via `RE_SQL_INCLUDE = /^INCLUDE\s+(?:'([^']+)'|"([^"]+)"|([A-Z][A-Z0-9_-]+))/i`
#### 1.2 CALL ... USING parameter extraction -> ACCESSES edges (Graph value: 9.2/10)
- **File:** `cobol-preprocessor.ts` (processLogicalLine CALL section)
- **What:** After capturing CALL target, scan for USING clause. Extract parameter names (reuse USING_KEYWORDS filter). Store as `calls[].parameters: string[]`
- **Interface:** Add `parameters?: string[]` to calls array type in CobolRegexResults
- **File:** `cobol-processor.ts` (CALL edge block)
- **Graph:** For each USING parameter, create ACCESSES edge from caller to data item Property node with reason `cobol-call-using`
- **Tests:** `CALL 'AUDITLOG' USING CUST-ID WS-AMOUNT` -> 2 ACCESSES edges
**Research insights (CALL USING forms):**
- Mixed modes: `CALL 'PGM' USING BY REFERENCE WS-A BY CONTENT WS-B BY VALUE WS-C`
- Pointer passing: `CALL 'PGM' USING ADDRESS OF WS-A`
- Placeholder: `CALL 'PGM' USING OMITTED WS-B`
- Filter keywords: add `ADDRESS`, `OMITTED`, `LENGTH` to USING_KEYWORDS (already has BY/VALUE/REFERENCE/CONTENT)
- **Impact tool enhancement:** CALL-USING edges enable BFS traversal through parameter data flow — single most impactful edge type for COBOL impact analysis
#### 1.3 STRING/UNSTRING data flow -> ACCESSES edges
- **File:** `cobol-preprocessor.ts` (new section in extractProcedure)
- **What:** Accumulate multi-line STRING/UNSTRING until period or END-STRING/END-UNSTRING. Extract sources and INTO targets.
- **Interface:** Add `strings: Array<{ sources: string[]; target: string; type: 'string' | 'unstring'; line: number; caller: string | null }>` to CobolRegexResults
- **Graph:** read-ACCESSES on sources, write-ACCESSES on INTO target with reason `cobol-string-read` / `cobol-string-write`
- **Tests:** 2 unit tests + integration test assertions
**Research insights (STRING/UNSTRING):**
- **Needs statement accumulator** — STRING/UNSTRING always span multiple lines in production
- Terminate accumulation at: period, END-STRING/END-UNSTRING, or start of next COBOL verb
- STRING sources: identifiers before each `DELIMITED BY`. Filter: STRING, DELIMITED, BY, SIZE, ALL, INTO, WITH, POINTER, ON, OVERFLOW, NOT, END-STRING
- UNSTRING: source is first identifier after UNSTRING; INTO targets are identifiers after INTO. Filter: DELIMITER, IN, COUNT, TALLYING, OR
- WITH POINTER field is both read AND written (starting position updated)
- TALLYING IN / COUNT IN fields are write targets
- Literal sources (`'text'`) must be filtered — quote-aware tokenization needed
- **Edge case**: STRING terminated by next verb, not period — existing fixture has `STRING ... DISPLAY` without period between them
#### 1.4 OCCURS DEPENDING ON -> ACCESSES edge
- **File:** `cobol-preprocessor.ts` (parseDataItemClauses)
- **What:** Extend OCCURS regex to capture DEPENDING ON field, KEY fields, and INDEXED BY names
- **Interface:** Add `dependingOn?: string`, `occursMax?: number`, `occursKeys?: Array<{direction: string; fields: string[]}>`, `indexedBy?: string[]` to data items
- **Graph:** ACCESSES edge from table item to controlling field with reason `cobol-depends-on`
- **Tests:** `05 WS-TABLE OCCURS 100 DEPENDING ON WS-COUNT` -> edge
**Research insights (OCCURS):**
- IBM allows `OCCURS 0 TO n DEPENDING ON` (zero minimum) and `OCCURS UNBOUNDED DEPENDING ON` (V6.4)
- Subscripted controlling fields: `DEPENDING ON WS-COUNT(WS-IDX)` — strip subscripts before storing
- **Pre-existing gap**: Multi-line data item clauses without continuation indicator are NOT captured. `05 WS-TABLE\n OCCURS 100\n DEPENDING ON WS-COUNT.` — the current RE_DATA_ITEM only gets the first line, `rest` is empty. Fixing properly requires a data item accumulator (like SELECT). **Defer full fix to Phase 3; implement same-line capture now.**
- KEY IS fields: `ASCENDING KEY IS WS-KEY-1 WS-KEY-2` — capture for SEARCH ALL resolution
- INDEXED BY: `INDEXED BY IDX-1 IDX-2` — capture for SET/SEARCH context
#### 1.5 VALUE clause for standard data items
- **File:** `cobol-preprocessor.ts` (parseDataItemClauses)
- **What:** Extract VALUE using a pragmatic function that handles quoted strings, numerics, figurative constants, hex/national literals
- **Interface:** Already exists as `values?: string[]` on data items (currently only populated for 88-level)
- **Graph:** Stored in Property node description (no new edges)
- **Tests:** `01 WS-STATUS PIC X VALUE 'A'` -> values: ['A']
**Research insights (VALUE forms):**
- Hex literals: `VALUE X'F1F2F3F4'`, National: `VALUE N'text'`, DBCS: `VALUE G'text'`
- Figurative constants: SPACES, ZEROS, ZEROES, LOW-VALUES, HIGH-VALUES, QUOTES, NULL, NULLS
- ALL literal: `VALUE ALL '*'`
- Numeric with sign/decimal: `VALUE -123.45`, `VALUE +1`
- `VALUE IS` optional — both `VALUE 'A'` and `VALUE IS 'A'` valid
- **Decimal vs period ambiguity**: `VALUE 100.` — is `.` decimal or terminator? `parseDataItemClauses` already strips trailing period, so this is handled
- IBM V6.4: floating-point `VALUE 1.0E5` — extend numeric regex if needed
- Implementation: use a pragmatic `extractValue(rest)` function, not a single complex regex
### Phase 2: EXEC DLI + DECLARATIVES (~90 LOC, ~4 new edge types)
IMS/DB support and error handling flows.
#### 2.1 EXEC DLI (IMS/DB) -> ACCESSES edges (Graph value: 8.2/10)
- **File:** `cobol-preprocessor.ts` (processLogicalLine — add RE_EXEC_DLI_START check alongside SQL/CICS)
- **What:** Accumulate EXEC DLI blocks like EXEC SQL. Parse DLI verbs (GU, GN, GNP, GHU, GHN, GHNP, ISRT, DLET, REPL, CHKP, SCHD, TERM). Extract segment name, PCB number, INTO/FROM areas, WHERE fields, PSB name.
- **Interface:** Add `execDliBlocks: Array<{ line: number; verb: string; pcbNumber?: number; segmentName?: string; intoField?: string; fromField?: string; whereField?: string; psbName?: string }>` to CobolRegexResults
- **Graph:** CodeElement node + ACCESSES edge to `<ims>:<segmentName>` Record node with reason `dli-{verb}`; ACCESSES edges to INTO/FROM data areas; PSB ACCESSES for SCHD
- **Tests:** `EXEC DLI GU USING PCB(1) SEGMENT(CUSTOMER) INTO(WS-CUST) END-EXEC`
**Research insights (dual IMS interface):**
- **EXEC DLI**: Embedded command interface for CICS-DL/I programs only
- **CBLTDLI CALL**: Batch interface via `CALL 'CBLTDLI' USING function-code PCB io-area SSA1..SSA15`
- CBLTDLI is already captured as a CALL to 'CBLTDLI' — enrich with USING parameter semantics later
- Multiple SEGMENT levels in hierarchical retrieval — use `matchAll` on segment regex
- DLI verbs: GU (most common), GN, GNP, GHU, GHN, GHNP, ISRT, REPL, DLET, CHKP, SCHD, TERM, ROLL, ROLB
- **Edge case**: DLET/REPL have no SEGMENT clause (operate on current position)
- **Recommended order**: Implement AFTER DECLARATIVES and SET (lower risk, higher frequency)
#### 2.2 DECLARATIVES / USE AFTER STANDARD EXCEPTION (Graph value: 9.0/10)
- **File:** `cobol-preprocessor.ts` (processLogicalLine — detect DECLARATIVES keyword, track USE AFTER blocks)
- **What:** When `DECLARATIVES.` is encountered, switch to declaratives mode. Extract USE statements binding sections to files/modes.
- **Interface:** Add `declaratives: Array<{ sectionName: string; useType: 'error' | 'debug' | 'label' | 'reporting'; target: string; line: number }>` to CobolRegexResults
- **Graph:** ACCESSES edge from declarative Namespace to file Record with reason `cobol-declarative-error-handler`
- **Tests:** Unit test with DECLARATIVES section, integration test for error flow
**Research insights (DECLARATIVES syntax):**
- `USE AFTER STANDARD {EXCEPTION|ERROR} ON {file-name|INPUT|OUTPUT|I-O|EXTEND}`
- EXCEPTION and ERROR are synonymous; STANDARD is optional in IBM dialects
- Multiple USE sections allowed (one per file + catch-all for I/O modes)
- `END DECLARATIVES.` must NOT reset PROCEDURE DIVISION state
- `DECLARATIVES` is already in EXCLUDED_PARA_NAMES — no false paragraph risk
- Existing section/paragraph detection already captures structural elements — just need USE binding
- **Lowest risk Phase 2 item** — implement first
#### 2.3 SET statement -> ACCESSES edges
- **File:** `cobol-preprocessor.ts` (extractProcedure — new RE_SET regex)
- **Interface:** Add `sets: Array<{ targets: string[]; form: 'to-true'|'to-value'|'up-by'|'down-by'|'address-of'|'to-null'|'to-entry'; value?: string; entryTarget?: string; entryIsLiteral?: boolean; line: number; caller: string | null }>` to CobolRegexResults
- **Graph:** ACCESSES write edge with reason `cobol-set-condition` (TO TRUE), `cobol-set-index` (TO/UP/DOWN), `cobol-set-address` (ADDRESS OF). SET ENTRY with literal -> CALLS edge.
- **Tests:** `SET WS-EOF TO TRUE`, `SET IDX-1 TO 5`, `SET IDX-1 UP BY 1`
**Research insights (SET forms by frequency):**
- `SET condition TO TRUE` — 80-90% of all SET usage. Multiple targets: `SET COND-A COND-B TO TRUE`
- `SET index TO/UP BY/DOWN BY` — ~8%. Multiple indices: `SET IDX-1 IDX-2 UP BY 1`
- `SET pointer TO ADDRESS OF data-item` / `SET ADDRESS OF data-item TO pointer` — ~2%
- `SET proc-ptr TO ENTRY "PROGNAME"` — rare but creates CALLS edge (like dynamic CALL)
- Filter OF/IN qualifiers: `SET COND-A OF WS-RECORD TO TRUE` (strip OF WS-RECORD)
- **Prioritize**: SET TO TRUE alone covers 80-90% — implement this form first
#### 2.4 INSPECT -> ACCESSES edges
- **File:** `cobol-preprocessor.ts` (extractProcedure — new `inspectAccum` accumulator like SORT)
- **What:** Accumulate multi-line INSPECT until period. Extract inspected field + tally counters.
- **Interface:** Add `inspects: Array<{ inspectedField: string; counters: string[]; form: 'tallying'|'replacing'|'converting'|'tallying-replacing'; line: number; caller: string | null }>` to CobolRegexResults
- **Graph:** ACCESSES read on inspected field always; write if REPLACING/CONVERTING. Write edges for tally counters. Reason: `cobol-inspect-read`/`cobol-inspect-write`/`cobol-inspect-tally`
- **Tests:** `INSPECT WS-FIELD TALLYING WS-COUNT FOR ALL 'A'` -> read on WS-FIELD, write on WS-COUNT
**Research insights (INSPECT forms by frequency):**
- REPLACING (~60%): `INSPECT WS-STR REPLACING ALL 'A' BY 'B'`
- TALLYING (~25%): `INSPECT WS-STR TALLYING WS-CNT FOR ALL 'A'` — multiple counters possible
- CONVERTING (~10%): `INSPECT WS-STR CONVERTING 'abc' TO 'ABC'`
- Combined (~5%): TALLYING + REPLACING in single statement
- **Needs multi-line accumulator** — INSPECT frequently spans 3-5 lines in production
- Extract tally counters with `([A-Z][A-Z0-9-]+)\s+FOR\b` matchAll pattern
- Filter figurative constants (SPACES, ZEROS) using existing MOVE_SKIP set
### Phase 3: Completeness Fixes (~60 LOC)
Fix the 10 partial features and small gaps.
#### 3.1 CALL ... RETURNING extraction
- Extend RE_CALL processing to capture RETURNING target after the USING clause
- Store as `calls[].returning?: string`
- Graph: ACCESSES write edge with reason `cobol-call-returning`
#### 3.2 SELECT OPTIONAL flag preservation
- Store `isOptional: boolean` in FileDeclaration interface
- Include in Record node description
#### 3.3 ALTERNATE RECORD KEY extraction
- Add regex in parseSelectStatement: `/\bALTERNATE\s+RECORD\s+KEY\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)/i`
- Store as `alternateKeys?: string[]`
#### 3.4 COMMON attribute on nested programs
- Extend RE_PROGRAM_ID: `/\bPROGRAM-ID\.\s*([A-Z][A-Z0-9-]+)(?:\s+IS\s+COMMON)?/i`
- Store `isCommon: boolean` on Module node
- Affects cross-program CALL resolution scope
#### 3.5 IS EXTERNAL / IS GLOBAL as first-class properties
- Change from usage string hack to proper boolean fields on data items
- Add `isExternal?: boolean`, `isGlobal?: boolean` to data item interface
#### 3.6 AUTHOR / DATE-WRITTEN mapped to Module node
- Already extracted as programMetadata — map to Module node properties
- `graph.addNode({ ..., properties: { ..., author, dateWritten } })`
#### 3.7 REPLACE statement
- Track REPLACE / REPLACE OFF state in preprocessor
- Apply text substitutions during preprocessing (before regex extraction)
- Complex: requires careful scoping rules
### Phase 4: Niche Features (~30 LOC)
Low-priority but nice for completeness.
#### 4.1 INITIALIZE statement -> write ACCESSES
- `/\bINITIALIZE\s+([A-Z][A-Z0-9-]+)/i`
- ACCESSES write edge with reason `cobol-initialize`
#### 4.2 Remaining IDENTIFICATION DIVISION paragraphs
- DATE-COMPILED, INSTALLATION, SECURITY, REMARKS
- Map to Module node description properties
#### 4.3 EXEC SQL INCLUDE -> IMPORTS edge (expansion)
- For EXEC SQL INCLUDE inside EXEC blocks that reference copybooks containing SQL
- Create IMPORTS edge similar to COPY
## Acceptance Criteria
### Functional Requirements
- [ ] Phase 1: All 5 features implemented with unit + integration tests
- [ ] Phase 2: All 4 features implemented with unit + integration tests
- [ ] Phase 3: All 7 partial features fixed
- [ ] Phase 4: At least 2 of 3 niche features implemented
- [ ] All existing 145 tests continue to pass
- [ ] TypeScript compiles cleanly
### Non-Functional Requirements
- [ ] No performance regression: CardDemo benchmark stays under 8s
- [ ] No file exceeds 1500 LOC (preprocessor currently 1326)
- [ ] ACAS benchmark shows increased node/edge counts (more data extracted)
- [ ] CardDemo benchmark shows increased edge counts (CALL USING, STRING, etc.)
### Quality Gates
- [ ] Each phase has its own commit
- [ ] Integration test assertions updated with exact counts per phase
- [ ] Benchmark run after each phase to track graph growth
## Dependencies & Risks
### Dependencies
- None. All changes are additive to existing COBOL processor code.
- No LanguageProvider changes needed.
- No graph schema changes needed (all new constructs map to existing node labels + edge types).
### Risks
- **preprocessor.ts size**: Currently 1326 LOC. Phase 1+2 adds ~200 LOC -> 1526 LOC. May need to extract helpers into a separate `cobol-data-flow.ts` module if it exceeds 1500.
- **REPLACE statement** (Phase 3.7) is the most complex feature — requires tracking text substitution state across logical lines. Consider deferring to a separate PR if it takes >100 LOC.
- **EXEC DLI** (Phase 2.1) is only testable against IMS codebases. Need fixture data or synthetic test cases.
## Graph Value Ranking by MCP Tool Impact
Research agent analyzed all 5 MCP tools (query, context, impact, detect_changes, rename) against planned edge types:
| Edge Type | QUERY | CONTEXT | IMPACT | DETECT | RENAME | **Overall** |
|-----------|-------|---------|--------|--------|--------|-------------|
| `cobol-call-using` | 4/5 | 5/5 | 5/5 | 4/5 | 4/5 | **9.2/10** |
| `cobol-error-handler` | 5/5 | 4/5 | 5/5 | 5/5 | 2/5 | **9.0/10** |
| `dli-*` (IMS verbs) | 4/5 | 4/5 | 5/5 | 4/5 | 2/5 | **8.2/10** |
| `cobol-string-*` | 4/5 | 3/5 | 3/5 | 3/5 | 2/5 | **6.2/10** |
**Key finding**: `cobol-call-using` alone would fix ~40% of missing caller references in COBOL graphs.
## Future Considerations
This plan provides the graph data foundation for a future `modernize` MCP command (out of scope) that would:
- Use CALL USING edges to map data contracts between programs
- Use STRING/UNSTRING edges to identify data transformation logic
- Use EXEC SQL/DLI edges to map database access patterns
- Use DECLARATIVES to understand error handling architecture
- Use the complete knowledge graph to generate migration plans
**MCP tool enhancements needed** (after this plan ships):
- Add `cobol-call-using`, `cobol-error-handler`, `dli-*` to IMPACT tool's default `relationTypes` for COBOL repos
- Add confidence floors for new edge types in `IMPACT_RELATION_CONFIDENCE`
- Register new edge types in `VALID_RELATION_TYPES` set (`local-backend.ts:52`)
## Sources & References
### Internal References
- Feature audit: session 8642401e (COBOL expert agent, 123 features audited)
- Prior plans: `docs/plans/2026-03-25-feat-cobol-100-percent-feature-coverage-plan.md`
- Architecture: `docs/code-indexing/cobol/` (7 documentation files)
### External References
- COBOL features reference: mainframestechhelp.com/tutorials/cobol/features.htm
- COBOL-85 standard: ISO/IEC 1989:1985
- IBM Enterprise COBOL reference
@@ -0,0 +1,725 @@
# PR #626 HIGH-Priority Fixes Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Fix 4 HIGH-priority issues from PR #626 code review before merge.
**Architecture:** Minimal targeted fixes — each task is independent. TDD: tests first, then implementation. No refactoring beyond what's needed.
**Tech Stack:** TypeScript, Vitest, Node.js fs/path APIs
**Spec:** `docs/superpowers/specs/2026-04-02-pr626-high-fixes-design.md`
**Paths:** All file paths are relative to the monorepo root (`GitNexus/`). Git commands run from the root. The `gitnexus/` prefix is a package subdirectory, not a separate repo.
---
### Task 1: Path Traversal — Validate Group Name
**Files:**
- Modify: `gitnexus/src/core/group/storage.ts:17-19` (getGroupDir) and `:63-68` (createGroupDir)
- Test: `gitnexus/test/unit/group/storage.test.ts`
- [ ] **Step 1: Write failing tests for validateGroupName**
In `gitnexus/test/unit/group/storage.test.ts`, add `createGroupDir` and `validateGroupName` to the existing import from `'../../../src/core/group/storage.js'` (line 6-11). Then add these describe blocks at the end of the outer `describe('Group storage', ...)`:
```typescript
describe('validateGroupName', () => {
it('test_validateGroupName_traversal_path_throws', () => {
expect(() => validateGroupName('../../evil')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_slash_in_name_throws', () => {
expect(() => validateGroupName('foo/bar')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_empty_string_throws', () => {
expect(() => validateGroupName('')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_starts_with_dash_throws', () => {
expect(() => validateGroupName('-leading-dash')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_starts_with_underscore_throws', () => {
expect(() => validateGroupName('_leading')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_dots_throws', () => {
expect(() => validateGroupName('com.example')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_valid_alphanumeric_passes', () => {
expect(() => validateGroupName('my-group_01')).not.toThrow();
});
it('test_validateGroupName_single_char_passes', () => {
expect(() => validateGroupName('A')).not.toThrow();
});
it('test_validateGroupName_all_digits_passes', () => {
expect(() => validateGroupName('123')).not.toThrow();
});
});
describe('getGroupDir rejects invalid names', () => {
it('test_getGroupDir_traversal_throws', () => {
expect(() => getGroupDir(tmpDir, '../../etc')).toThrow(/Invalid group name/);
});
it('test_getGroupDir_valid_name_returns_path', () => {
const dir = getGroupDir(tmpDir, 'company');
expect(dir).toBe(path.join(tmpDir, 'groups', 'company'));
});
});
describe('createGroupDir rejects invalid names', () => {
it('test_createGroupDir_traversal_throws', async () => {
await expect(createGroupDir(tmpDir, '../evil')).rejects.toThrow(/Invalid group name/);
});
});
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `cd gitnexus && npx vitest run test/unit/group/storage.test.ts`
Expected: FAIL — `validateGroupName` is not exported, `getGroupDir` does not throw.
- [ ] **Step 3: Implement validateGroupName and wire into getGroupDir and createGroupDir**
In `gitnexus/src/core/group/storage.ts`, add the validation function before `getGroupDir` and call it:
```typescript
const GROUP_NAME_RE = /^[a-zA-Z0-9][a-zA-Z0-9_-]*$/;
export function validateGroupName(name: string): void {
if (!GROUP_NAME_RE.test(name)) {
throw new Error(
`Invalid group name "${name}". Names must start with a letter or digit and contain only [a-zA-Z0-9_-].`,
);
}
}
export function getGroupDir(gitnexusDir: string, groupName: string): string {
validateGroupName(groupName);
return path.join(gitnexusDir, 'groups', groupName);
}
```
`createGroupDir` already calls `getGroupDir` at line 68, so it inherits validation automatically. No change needed in `createGroupDir`.
- [ ] **Step 4: Run tests to verify they pass**
Run: `cd gitnexus && npx vitest run test/unit/group/storage.test.ts`
Expected: ALL PASS
- [ ] **Step 5: Commit**
```bash
cd gitnexus && git add src/core/group/storage.ts test/unit/group/storage.test.ts
git commit -m "fix(group): validate group name to prevent path traversal
Add validateGroupName() with regex [a-zA-Z0-9][a-zA-Z0-9_-]*.
Called in getGroupDir (defense in depth) which covers all CLI entry
points: create, add, remove, status, sync.
Addresses PR #626 review item 1 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
### Task 2: Directory Exclusions in Service Boundary Detector
**Files:**
- Modify: `gitnexus/src/core/group/service-boundary-detector.ts:24-51` (add constant), `:78` (walkForBoundaries), `:130` (hasSourceFilesInSubdirs)
- Test: `gitnexus/test/unit/group/service-boundary-detector.test.ts`
- [ ] **Step 1: Write failing tests for excluded directories**
Add this describe block inside the existing `detectServiceBoundaries` describe in `gitnexus/test/unit/group/service-boundary-detector.test.ts`:
```typescript
it('test_detect_skips_vendor_directory', async () => {
writeFile('services/auth/package.json', '{}');
writeFile('services/auth/src/index.ts', '');
// vendor should be skipped — its contents should not create a boundary
writeFile('vendor/some-dep/package.json', '{}');
writeFile('vendor/some-dep/src/lib.go', '');
const boundaries = await detectServiceBoundaries(tmpDir);
const paths = boundaries.map((b) => b.servicePath);
expect(paths).toContain('services/auth');
expect(paths).not.toContain('vendor/some-dep');
});
it('test_detect_skips_target_directory', async () => {
writeFile('services/api/go.mod', 'module api');
writeFile('services/api/main.go', '');
writeFile('target/classes/Main.java', '');
writeFile('target/pom.xml', '<project/>');
const boundaries = await detectServiceBoundaries(tmpDir);
const paths = boundaries.map((b) => b.servicePath);
expect(paths).toContain('services/api');
expect(paths).not.toContain('target');
});
it('test_detect_skips_pycache_directory', async () => {
writeFile('services/ml/pyproject.toml', '[project]');
writeFile('services/ml/model.py', '');
// __pycache__ with a marker + source files — would be detected as
// a boundary if not excluded, since it has package.json + .py file
writeFile('__pycache__/package.json', '{}');
writeFile('__pycache__/cached.py', '');
const boundaries = await detectServiceBoundaries(tmpDir);
const paths = boundaries.map((b) => b.servicePath);
expect(paths).toContain('services/ml');
expect(paths.every((p) => !p.includes('__pycache__'))).toBe(true);
});
it('test_detect_skips_dotfile_directories_regression', async () => {
writeFile('services/api/package.json', '{}');
writeFile('services/api/src/index.ts', '');
writeFile('.hidden/package.json', '{}');
writeFile('.hidden/src/index.ts', '');
const boundaries = await detectServiceBoundaries(tmpDir);
const paths = boundaries.map((b) => b.servicePath);
expect(paths).toContain('services/api');
expect(paths).not.toContain('.hidden');
});
it('test_detect_does_not_skip_regular_source_directories', async () => {
writeFile('services/api/package.json', '{}');
writeFile('services/api/src/index.ts', '');
const boundaries = await detectServiceBoundaries(tmpDir);
expect(boundaries).toHaveLength(1);
expect(boundaries[0].serviceName).toBe('api');
});
```
- [ ] **Step 2: Run tests to verify `vendor` and `target` tests fail**
Run: `cd gitnexus && npx vitest run test/unit/group/service-boundary-detector.test.ts`
Expected: `test_detect_skips_vendor_directory` and `test_detect_skips_target_directory` FAIL (vendor/target not excluded). Other new tests may pass since dotfile exclusion already exists.
- [ ] **Step 3: Add EXCLUDED_DIRS constant and update both walking functions**
In `gitnexus/src/core/group/service-boundary-detector.ts`:
After `SOURCE_EXTENSIONS` (after line 51), add:
```typescript
const EXCLUDED_DIRS = new Set([
'node_modules',
'vendor',
'target',
'build',
'dist',
'__pycache__',
'.venv',
'venv',
'.tox',
'.mypy_cache',
'.gradle',
'.mvn',
'out',
'bin',
]);
```
In `walkForBoundaries`, replace line 78:
```typescript
if (entry.name.startsWith('.') || entry.name === 'node_modules') continue;
```
with:
```typescript
if (entry.name.startsWith('.') || EXCLUDED_DIRS.has(entry.name)) continue;
```
In `hasSourceFilesInSubdirs`, replace line 130:
```typescript
if (entry.isDirectory() && !entry.name.startsWith('.') && entry.name !== 'node_modules') {
```
with:
```typescript
if (entry.isDirectory() && !entry.name.startsWith('.') && !EXCLUDED_DIRS.has(entry.name)) {
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `cd gitnexus && npx vitest run test/unit/group/service-boundary-detector.test.ts`
Expected: ALL PASS
- [ ] **Step 5: Commit**
```bash
cd gitnexus && git add src/core/group/service-boundary-detector.ts test/unit/group/service-boundary-detector.test.ts
git commit -m "fix(group): add directory exclusions to service boundary detector
Add EXCLUDED_DIRS set: vendor, target, build, dist, __pycache__,
.venv, venv, .tox, .mypy_cache, .gradle, .mvn, out, bin.
Applied in walkForBoundaries and hasSourceFilesInSubdirs.
Replaces inline node_modules check.
Addresses PR #626 review item 3 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
### Task 3: Remove Double-Close of LadybugDB Pools
**Files:**
- Modify: `gitnexus/src/cli/group.ts:160` (remove import), `:187-189` (remove finally block body)
- Test: `gitnexus/test/unit/group/sync.test.ts` (add pool cleanup test)
- Test: `gitnexus/test/integration/group/group-cli.test.ts` (verify no blanket close in source)
- [ ] **Step 1: Write unit tests for per-id pool cleanup in sync.ts**
Add to `gitnexus/test/unit/group/sync.test.ts`, inside the existing `describe('syncGroup', ...)`:
```typescript
it('test_syncGroup_closes_only_opened_pools', async () => {
const config = makeConfig({
'app/backend': 'backend-repo',
'app/frontend': 'frontend-repo',
});
const closedIds: string[] = [];
// Mock initLbug/closeLbug via per-repo override that tracks pool lifecycle
const { vi } = await import('vitest');
const poolAdapter = await import('../../../src/core/lbug/pool-adapter.js');
const initSpy = vi.spyOn(poolAdapter, 'initLbug').mockResolvedValue(undefined);
const closeSpy = vi.spyOn(poolAdapter, 'closeLbug').mockImplementation(async (id?: string) => {
if (id) closedIds.push(id);
});
try {
await syncGroup(config, {
resolveRepoHandle: async (_name, groupPath) => ({
id: groupPath.replace(/\//g, '-'),
path: groupPath,
repoPath: '/tmp/' + groupPath,
storagePath: '/tmp/' + groupPath + '/.gitnexus',
}),
skipWrite: true,
}).catch(() => {});
// Regardless of extraction errors, closeLbug should be called per id
// closeLbug should only receive specific pool ids, never undefined/empty
for (const id of closedIds) {
expect(id).toBeTruthy();
expect(typeof id).toBe('string');
}
// No blanket close (no-arg call)
const blanketCalls = closeSpy.mock.calls.filter((args) => args.length === 0 || !args[0]);
expect(blanketCalls).toHaveLength(0);
} finally {
initSpy.mockRestore();
closeSpy.mockRestore();
}
});
```
- [ ] **Step 2: Run sync unit test to verify it passes (sync.ts already does per-id cleanup)**
Run: `cd gitnexus && npx vitest run test/unit/group/sync.test.ts`
Expected: PASS — sync.ts already cleans up correctly. This test locks the behavior.
- [ ] **Step 3: Write test verifying CLI source has no blanket closeLbug()**
Add to `gitnexus/test/integration/group/group-cli.test.ts`:
```typescript
it('test_sync_command_source_does_not_call_blanket_closeLbug', () => {
const cliGroupPath = path.join(repoRoot, 'src', 'cli', 'group.ts');
const source = fs.readFileSync(cliGroupPath, 'utf-8');
// closeLbug() without arguments (blanket close) must not appear.
// closeLbug(id) with argument is fine (that's in sync.ts, not here).
// Match closeLbug() but not closeLbug(someArg)
const blanketClosePattern = /closeLbug\s*\(\s*\)/;
expect(source).not.toMatch(blanketClosePattern);
});
```
- [ ] **Step 4: Run test to verify it fails**
Run: `cd gitnexus && npx vitest run test/integration/group/group-cli.test.ts`
Expected: FAIL — `closeLbug()` (no args) exists at line 188.
- [ ] **Step 5: Remove blanket closeLbug() from cli/group.ts**
In `gitnexus/src/cli/group.ts`:
Remove the `closeLbug` import at line 160:
```typescript
const { closeLbug } = await import('../core/lbug/pool-adapter.js');
```
Replace the try/finally wrapper (lines 162-189):
```typescript
try {
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
console.log(`Syncing group "${name}" (${Object.keys(config.repos).length} repos)...\n`);
const result = await syncGroup(config, {
groupDir,
allowStale: Boolean(opts.allowStale),
verbose: Boolean(opts.verbose),
skipEmbeddings: Boolean(opts.skipEmbeddings),
exactOnly: Boolean(opts.exactOnly),
});
if (opts.json) {
console.log(JSON.stringify(result, null, 2));
} else {
console.log(`\nMatching cascade:`);
const exactLinks = result.crossLinks.filter((l) => l.matchType === 'exact');
console.log(` exact: ${exactLinks.length} cross-links (confidence 1.0)`);
console.log(` unmatched: ${result.unmatched.length} contracts`);
console.log(
`\nWrote contracts.json (${result.contracts.length} contracts, ${result.crossLinks.length} cross-links)`,
);
}
} finally {
await closeLbug().catch(() => {});
}
```
Becomes (remove try/finally entirely, since sync.ts handles its own cleanup):
```typescript
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
console.log(`Syncing group "${name}" (${Object.keys(config.repos).length} repos)...\n`);
const result = await syncGroup(config, {
groupDir,
allowStale: Boolean(opts.allowStale),
verbose: Boolean(opts.verbose),
skipEmbeddings: Boolean(opts.skipEmbeddings),
exactOnly: Boolean(opts.exactOnly),
});
if (opts.json) {
console.log(JSON.stringify(result, null, 2));
} else {
console.log(`\nMatching cascade:`);
const exactLinks = result.crossLinks.filter((l) => l.matchType === 'exact');
console.log(` exact: ${exactLinks.length} cross-links (confidence 1.0)`);
console.log(` unmatched: ${result.unmatched.length} contracts`);
console.log(
`\nWrote contracts.json (${result.contracts.length} contracts, ${result.crossLinks.length} cross-links)`,
);
}
```
- [ ] **Step 6: Run tests to verify they pass**
Run: `cd gitnexus && npx vitest run test/integration/group/group-cli.test.ts test/unit/group/sync.test.ts`
Expected: ALL PASS
- [ ] **Step 7: Commit**
```bash
cd gitnexus && git add src/cli/group.ts test/integration/group/group-cli.test.ts test/unit/group/sync.test.ts
git commit -m "fix(group): remove blanket closeLbug() from CLI sync command
sync.ts already closes pools per-id in its finally block.
The blanket closeLbug() in cli/group.ts tears down ALL active pools
including unrelated ones in MCP server context.
Addresses PR #626 review item 4 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
### Task 4: gRPC Proto Regex — Brace-Depth Counter
**Files:**
- Modify: `gitnexus/src/core/group/extractors/grpc-extractor.ts:101-130` (parseProtoFile)
- Test: `gitnexus/test/unit/group/grpc-extractor.test.ts`
- [ ] **Step 1: Write failing tests for nested braces in proto services**
Add this describe block inside the existing `proto file parsing` describe in `gitnexus/test/unit/group/grpc-extractor.test.ts`:
```typescript
it('test_extract_proto_with_google_api_http_nested_braces', async () => {
writeFile(
'api/gateway.proto',
`syntax = "proto3";
package gateway.v1;
import "google/api/annotations.proto";
service GatewayService {
rpc GetUser (GetUserRequest) returns (UserResponse) {
option (google.api.http) = {
get: "/v1/users/{user_id}"
};
}
rpc CreateUser (CreateUserRequest) returns (UserResponse) {
option (google.api.http) = {
post: "/v1/users"
body: "*"
};
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter(
(c) => c.role === 'provider' && c.symbolRef.filePath === 'api/gateway.proto',
);
expect(providers).toHaveLength(2);
const ids = providers.map((c) => c.contractId).sort();
expect(ids).toEqual([
'grpc::gateway.v1.GatewayService/CreateUser',
'grpc::gateway.v1.GatewayService/GetUser',
]);
});
it('test_extract_proto_with_multiple_services', async () => {
writeFile(
'api/multi.proto',
`syntax = "proto3";
package multi;
service ServiceA {
rpc MethodA (Req) returns (Res);
}
service ServiceB {
rpc MethodB1 (Req) returns (Res);
rpc MethodB2 (Req) returns (Res);
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter(
(c) => c.role === 'provider' && c.symbolRef.filePath === 'api/multi.proto',
);
expect(providers).toHaveLength(3);
const ids = providers.map((c) => c.contractId).sort();
expect(ids).toEqual([
'grpc::multi.ServiceA/MethodA',
'grpc::multi.ServiceB/MethodB1',
'grpc::multi.ServiceB/MethodB2',
]);
});
it('test_extract_proto_with_nested_option_blocks_in_rpc', async () => {
writeFile(
'api/nested.proto',
`syntax = "proto3";
package nested;
service DeepService {
rpc DeepMethod (Req) returns (Res) {
option (google.api.http) = {
post: "/v1/deep"
body: "*"
additional_bindings {
get: "/v1/deep/{id}"
}
};
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter(
(c) => c.role === 'provider' && c.symbolRef.filePath === 'api/nested.proto',
);
expect(providers).toHaveLength(1);
expect(providers[0].contractId).toBe('grpc::nested.DeepService/DeepMethod');
});
it('test_extract_proto_malformed_unclosed_brace_skips_service', async () => {
writeFile(
'api/broken.proto',
`syntax = "proto3";
package broken;
service IncompleteService {
rpc SomeMethod (Req) returns (Res);
// Missing closing brace — EOF before depth returns to 0
`,
);
// Should not throw; incomplete service is silently skipped
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter(
(c) => c.role === 'provider' && c.symbolRef.filePath === 'api/broken.proto',
);
// The old regex would find partial match; the new parser should skip it
expect(providers).toHaveLength(0);
});
```
- [ ] **Step 2: Run tests to verify the nested brace test fails**
Run: `cd gitnexus && npx vitest run test/unit/group/grpc-extractor.test.ts`
Expected: `test_extract_proto_with_google_api_http_nested_braces` FAIL — regex stops at first `}` inside the `option` block.
- [ ] **Step 3: Replace serviceRe regex with extractServiceBlocks function**
In `gitnexus/src/core/group/extractors/grpc-extractor.ts`, replace the `parseProtoFile` method (lines 101-130):
```typescript
private parseProtoFile(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const pkgMatch = content.match(/^package\s+([\w.]+)\s*;/m);
const pkg = pkgMatch ? pkgMatch[1] : '';
for (const { name: serviceName, body } of extractServiceBlocks(content)) {
const rpcRe = /rpc\s+(\w+)\s*\(/g;
let rpcMatch: RegExpExecArray | null;
while ((rpcMatch = rpcRe.exec(body)) !== null) {
const methodName = rpcMatch[1];
const cid = contractId(pkg, serviceName, methodName);
out.push(
makeContract(cid, 'provider', filePath, `${serviceName}.${methodName}`, 0.85, {
package: pkg,
service: serviceName,
method: methodName,
source: 'proto',
}),
);
}
}
return out;
}
```
Add this function before the class (e.g. after `serviceOnlyContractId`, around line 26):
```typescript
function extractServiceBlocks(content: string): Array<{ name: string; body: string }> {
const results: Array<{ name: string; body: string }> = [];
const headerRe = /service\s+(\w+)\s*\{/g;
let headerMatch: RegExpExecArray | null;
while ((headerMatch = headerRe.exec(content)) !== null) {
const serviceName = headerMatch[1];
const bodyStart = headerMatch.index + headerMatch[0].length;
let depth = 1;
let pos = bodyStart;
while (pos < content.length && depth > 0) {
const ch = content[pos];
if (ch === '{') depth++;
else if (ch === '}') depth--;
pos++;
}
// If EOF before depth returns to 0, skip incomplete service
if (depth !== 0) continue;
// body is between opening { (consumed by regex) and closing } (pos is one past it)
const body = content.slice(bodyStart, pos - 1);
results.push({ name: serviceName, body });
}
return results;
}
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `cd gitnexus && npx vitest run test/unit/group/grpc-extractor.test.ts`
Expected: ALL PASS (including existing regression tests)
- [ ] **Step 5: Commit**
```bash
cd gitnexus && git add src/core/group/extractors/grpc-extractor.ts test/unit/group/grpc-extractor.test.ts
git commit -m "fix(group): replace gRPC proto regex with brace-depth counter
The serviceRe regex used [^}]* which stopped at the first '}'.
Proto services with google.api.http annotations contain nested {}
blocks, causing methods to be missed.
New extractServiceBlocks() uses a brace-depth counter (init depth=1
after opening {, scan char-by-char). Malformed protos with unclosed
braces are silently skipped.
Addresses PR #626 review item 2 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
### Task 5: Run Full Test Suite
- [ ] **Step 1: Run all group-related tests**
Run: `cd gitnexus && npx vitest run test/unit/group/ test/integration/group/`
Expected: ALL PASS
- [ ] **Step 2: Run full test suite to catch regressions**
Run: `cd gitnexus && npx vitest run`
Expected: ALL PASS, 0 failures
- [ ] **Step 3: Run typecheck**
Run: `cd gitnexus && npx tsc --noEmit`
Expected: No errors
---
### Task 6: CLI Integration Smoke Test
- [ ] **Step 1: Add CLI smoke test for path traversal**
Add to `gitnexus/test/integration/group/group-cli.test.ts` inside the existing `group CLI` describe:
```typescript
it('test_create_with_invalid_name_fails', () => {
const result = runGroup(['create', '../../evil']);
expect(result.status).not.toBe(0);
expect(result.stderr).toContain('Invalid group name');
});
```
- [ ] **Step 2: Run test**
Run: `cd gitnexus && npx vitest run test/integration/group/group-cli.test.ts`
Expected: ALL PASS
- [ ] **Step 3: Commit**
```bash
cd gitnexus && git add test/integration/group/group-cli.test.ts
git commit -m "test(group): add CLI smoke test for path traversal rejection
Verifies that 'group create ../../evil' fails with Invalid group name.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
@@ -0,0 +1,175 @@
# PR #626 HIGH-Priority Fixes Design
**Date:** 2026-04-02
**PR:** abhigyanpatwari/GitNexus#626 — Intra-repo service communication tracking
**Scope:** 4 HIGH-priority issues identified by abhigyanpatwari and xkonjin
**Approach:** Minimal targeted fixes (option A) — no refactoring, no scope creep
---
## Fix 1: Path Traversal via Group Name
**File:** `gitnexus/src/core/group/storage.ts`
**Risk:** A group name like `../../etc` creates directories outside the intended path.
### Solution
Add `validateGroupName(name: string): void` that enforces `/^[a-zA-Z0-9][a-zA-Z0-9_-]*$/`.
- Call in `createGroupDir` (primary entry point)
- Call in `getGroupDir` (defense in depth)
- Throw descriptive error on invalid names
**Legacy:** Groups already on disk with names outside this pattern are not auto-renamed; only new `create` / resolved paths are validated.
### Why regex over path.resolve + startsWith
- abhigyanpatwari explicitly requested `[a-zA-Z0-9_-]`
- Stricter: disallows spaces, dots, Unicode edge cases
- Simpler to reason about
### Tests
- `../../evil` throws
- `foo/bar` throws
- Empty string throws
- `my-group_01` passes
- `A` (single char) passes
- CLI smoke: one integration test that hits `getGroupDir` / `createGroupDir` (e.g. `group create` or `group add`) with an invalid name proves wiring for every subcommand that resolves a group through storage
### CLI/API entry points accepting groupName
All paths flow through `getGroupDir` (which validates), so coverage is implicit. For reference:
| Command | Entry | Calls |
|---------|-------|-------|
| `group create` | `cli/group.ts` action | `createGroupDir` -> `getGroupDir` |
| `group add` | `cli/group.ts` action | `getGroupDir` |
| `group remove` | `cli/group.ts` action | `getGroupDir` |
| `group list` | `cli/group.ts` action | reads `groups/` dir directly — no traversal risk (reads, not writes) |
| `group status` | `cli/group.ts` action | `getGroupDir` |
| `group sync` | `cli/group.ts` action | `getGroupDir` |
**`listGroups`:** Reads directory names from disk without validation. Not a write path, so no traversal risk. May surface manually-created directories with non-conforming names — accepted as-is, not in scope.
---
## Fix 2: gRPC Proto Regex -> Brace-Depth Counter
**File:** `gitnexus/src/core/group/extractors/grpc-extractor.ts`
**Risk:** `serviceRe = /service\s+(\w+)\s*\{([^}]*)}/gs` stops at first `}`. Proto services with `google.api.http` annotations inside RPCs contain nested `{ }` blocks.
### Solution
Replace `serviceRe` regex with `extractServiceBlocks(content: string): Array<{ name: string; body: string }>`:
1. Use regex only to find `service <Name> {` start positions (regex consumes the opening `{`)
2. Initialise depth to 1 immediately after the opening `{`
3. Scan forward char by char: `{` -> depth++, `}` -> depth--; collect into body
4. Stop when depth reaches 0 (the matching closing `}`)
5. Return name + body pairs
Inner `rpcRe` regex remains unchanged — it operates on the already-extracted body.
**Malformed input:** If EOF is reached before `depth` returns to 0, skip the incomplete service (do not add to results). Lock this in the test.
**Scope limitation (v1):** Brace-depth only — no lexer for string literals or comments containing `{`/`}`. Sufficient for `google.api.http` annotations. Known false positive: braces inside `//` comments or quoted strings within proto options. Accepted for v1; a proper proto lexer is out of scope.
### Tests
- Proto with single service, no nesting (regression)
- Proto with `google.api.http` nested braces inside RPC options
- Proto with multiple services
- Proto with nested `option` blocks inside RPC (e.g. `google.api.http`)
- Malformed proto with unclosed brace (graceful handling)
---
## Fix 3: Directory Exclusions in Service Boundary Detector
**File:** `gitnexus/src/core/group/service-boundary-detector.ts`
**Risk:** Walks entire repo tree, only skipping dotfiles and `node_modules`. Extremely slow on repos with `vendor/`, `target/`, `__pycache__/`, `.venv/`.
### Solution
Create `EXCLUDED_DIRS` as a `Set<string>` (alongside existing `SERVICE_MARKERS`, `SOURCE_EXTENSIONS`), for example:
```text
node_modules, vendor, target, build, dist,
__pycache__, .venv, venv, .tox, .mypy_cache,
.gradle, .mvn, out, bin
```
(Implement as `new Set([...])` — the list above is the membership, not a string literal.)
Apply in both:
- `walkForBoundaries` (line 77-78) — replace current inline `=== 'node_modules'` check with `EXCLUDED_DIRS.has(entry.name)`
- `hasSourceFilesInSubdirs` (line 130) — replace `entry.name !== 'node_modules'` with `!EXCLUDED_DIRS.has(entry.name)`
Note: remove the old `=== 'node_modules'` literal from both locations — it is covered by `EXCLUDED_DIRS`.
Dotfile exclusion (`.` prefix) remains as a separate check since it's a pattern, not a name.
Exclusions apply only to `isDirectory()` entries — file names are never checked against `EXCLUDED_DIRS`.
**Tradeoff:** Rare layouts that keep source under names like `out/` or `bin/` will be skipped; accepted for performance on typical monorepos.
**Case sensitivity:** `Set.has` is case-sensitive (matches current `=== 'node_modules'` behavior). Windows case-insensitive FS not handled — accepted as-is, consistent with existing code.
### Tests
- Directory named `vendor/` is skipped
- Directory named `target/` is skipped
- Directory named `__pycache__/` is skipped
- Regular source directories are NOT skipped
- Dotfile directories still skipped (regression)
---
## Fix 4: Double-Close of LadybugDB Pools
**Files:**
- `gitnexus/src/core/group/sync.ts` (lines 155-157) — per-id cleanup (KEEP)
- `gitnexus/src/cli/group.ts` (line 188) — blanket `closeLbug()` (REMOVE)
**Risk:** In MCP server context, `closeLbug()` without arguments tears down ALL active pools, including ones from unrelated operations.
### Solution
Remove the `closeLbug()` call (no arguments) from `cli/group.ts` finally block. The per-id cleanup in `sync.ts` is sufficient:
```typescript
// sync.ts — KEEP: cleans up only pools opened by this sync
finally {
for (const id of [...new Set(openPoolIds)]) {
await closeLbug(id).catch(() => {});
}
}
```
```typescript
// cli/group.ts — REMOVE: blanket close that kills all pools
finally {
await closeLbug().catch(() => {}); // DELETE THIS
}
```
Remove the `closeLbug` import from `cli/group.ts` — after removing the `finally` call it has no remaining usages.
### Tests (unit level — mock pool adapter)
- `syncGroup` closes only the pools it opened (mock `closeLbug`, assert called with specific ids)
- Two-pool scenario: sync opens pools A and B, both closed in finally; pool C (opened elsewhere) not touched
- CLI `sync` command does not call blanket `closeLbug()` (verify no zero-arg call in source — static check or grep-based test)
---
## Out of Scope
- JSON -> LadybugDB migration (tracked in #606)
- MEDIUM/LOW issues (items 5-10 from review summary)
- Test gap coverage beyond what's needed for these 4 fixes
- Any refactoring or architectural changes
## Execution Order
Fixes are independent — can be implemented in parallel or any order.
Recommended order for review clarity: 1 -> 3 -> 4 -> 2 (simplest to most complex).
+83
View File
@@ -0,0 +1,83 @@
import tsPlugin from '@typescript-eslint/eslint-plugin';
import tsParser from '@typescript-eslint/parser';
import unusedImports from 'eslint-plugin-unused-imports';
import reactHooks from 'eslint-plugin-react-hooks';
import prettierConfig from 'eslint-config-prettier';
export default [
// Global ignores
{
ignores: [
'**/dist/**',
'**/node_modules/**',
'**/coverage/**',
'gitnexus/vendor/**',
'gitnexus-web/src/vendor/**',
'gitnexus/test/fixtures/**',
'gitnexus-web/playwright-report/**',
'gitnexus-web/test-results/**',
'**/*.d.ts',
'.claude/**',
'.history/**',
],
},
// Base TypeScript config for all packages
{
files: ['**/*.{ts,tsx}'],
languageOptions: {
parser: tsParser,
parserOptions: {
ecmaVersion: 2022,
sourceType: 'module',
},
},
plugins: {
'@typescript-eslint': tsPlugin,
'unused-imports': unusedImports,
},
rules: {
// Unused imports — auto-fixable
'unused-imports/no-unused-imports': 'error',
'unused-imports/no-unused-vars': [
'warn',
{ vars: 'all', varsIgnorePattern: '^_', args: 'after-used', argsIgnorePattern: '^_' },
],
// TypeScript quality
'@typescript-eslint/no-unused-vars': 'off', // handled by unused-imports plugin
'no-unused-vars': 'off', // handled by unused-imports plugin
'@typescript-eslint/no-explicit-any': 'warn',
'@typescript-eslint/no-non-null-assertion': 'warn',
// General quality
'no-debugger': 'error',
'prefer-const': 'error',
'no-var': 'error',
eqeqeq: ['error', 'always', { null: 'ignore' }],
},
},
// CLI package — allow console.log (it's a CLI tool)
{
files: ['gitnexus/src/cli/**/*.ts', 'gitnexus/src/server/**/*.ts'],
rules: {
'no-console': 'off',
},
},
// React-specific rules for gitnexus-web
{
files: ['gitnexus-web/src/**/*.{ts,tsx}'],
plugins: {
'react-hooks': reactHooks,
},
rules: {
'react-hooks/rules-of-hooks': 'error',
'react-hooks/exhaustive-deps': 'warn',
},
},
// Disable formatting rules (prettier handles those)
prettierConfig,
];
+4
View File
@@ -50,6 +50,10 @@ All models are routed through **OpenRouter** by default, so a single `OPENROUTER
docker pull swebench/sweb.eval.x86_64.django_1776_django-16527:latest
```
### Debug logging
Set `GITNEXUS_EVAL_DEBUG=1` to include full Python tracebacks in run summaries and logs. By default, errors are sanitized to avoid leaking host paths or stack traces.
## Quick Start
### Debug a single instance
+8 -18
View File
@@ -22,8 +22,10 @@ import time
from enum import Enum
from pathlib import Path
from constants import AUGMENT_TIMEOUT_SECONDS
from minisweagent import Environment, Model
from minisweagent.agents.default import AgentConfig, DefaultAgent
from tool_registry import BINARIES_BY_KEY, TOOL_METRIC_KEYS
logger = logging.getLogger("gitnexus_agent")
@@ -40,7 +42,7 @@ class GitNexusMode(str, Enum):
class GitNexusAgentConfig(AgentConfig):
"""Extended config for GitNexus evaluation agent."""
gitnexus_mode: GitNexusMode = GitNexusMode.BASELINE
augment_timeout: float = 5.0
augment_timeout: float = AUGMENT_TIMEOUT_SECONDS
augment_min_pattern_length: int = 3
track_gitnexus_usage: bool = True
@@ -152,16 +154,10 @@ class GitNexusAgent(DefaultAgent):
"""Track which GitNexus tools the agent uses."""
for action in message.get("extra", {}).get("actions", []):
command = action.get("command", "")
if "gitnexus-query" in command:
self.gitnexus_metrics.tool_calls["query"] += 1
elif "gitnexus-context" in command:
self.gitnexus_metrics.tool_calls["context"] += 1
elif "gitnexus-impact" in command:
self.gitnexus_metrics.tool_calls["impact"] += 1
elif "gitnexus-cypher" in command:
self.gitnexus_metrics.tool_calls["cypher"] += 1
elif "gitnexus-overview" in command:
self.gitnexus_metrics.tool_calls["overview"] += 1
for key, binary in BINARIES_BY_KEY.items():
if binary in command and key in self.gitnexus_metrics.tool_calls:
self.gitnexus_metrics.tool_calls[key] += 1
break
def serialize(self, *extra_dicts) -> dict:
"""Serialize with GitNexus-specific metrics."""
@@ -180,13 +176,7 @@ class GitNexusMetrics:
"""Tracks GitNexus-specific metrics during evaluation."""
def __init__(self):
self.tool_calls: dict[str, int] = {
"query": 0,
"context": 0,
"impact": 0,
"cypher": 0,
"overview": 0,
}
self.tool_calls: dict[str, int] = {key: 0 for key in TOOL_METRIC_KEYS}
self.augmentation_calls: int = 0
self.augmentation_hits: int = 0
self.augmentation_errors: int = 0
+31 -22
View File
@@ -27,6 +27,8 @@ import typer
from rich.console import Console
from rich.table import Table
from tool_registry import TOOL_METRIC_KEYS
logger = logging.getLogger("analyze_results")
console = Console()
app = typer.Typer(rich_markup_mode="rich", add_completion=False)
@@ -77,13 +79,20 @@ def load_run_results(results_dir: Path) -> dict[str, dict]:
def parse_run_id(run_id: str) -> tuple[str, str]:
"""Parse 'model_mode' into (model, mode)."""
# Handle multi-word model names like 'minimax-2.5'
# Modes are: baseline, mcp, augment, full
known_modes = {"baseline", "mcp", "augment", "full"}
parts = run_id.rsplit("_", 1)
if len(parts) == 2 and parts[1] in known_modes:
return parts[0], parts[1]
"""Parse 'model_mode' into (model, mode) using known suffixes."""
# Match the longest known suffix first to avoid hyphen collisions in model names.
known_modes = [
"native_augment",
"native",
"baseline",
"mcp",
"augment",
"full",
]
for mode in known_modes:
suffix = f"_{mode}"
if run_id.endswith(suffix):
return run_id[: -len(suffix)], mode
return run_id, "unknown"
@@ -266,12 +275,17 @@ def compare_modes(
_, mode = parse_run_id(run_id)
metrics[mode] = compute_metrics(run_data)
mode_order = [
mode
for mode in ["baseline", "native", "native_augment", "mcp", "augment", "full"]
if mode in metrics
] or sorted(metrics.keys())
# Print comparison table
table = Table(title=f"Mode Comparison: {model}")
table.add_column("Metric", style="bold")
for mode in ["baseline", "mcp", "augment", "full"]:
if mode in metrics:
table.add_column(mode, justify="right")
for mode in mode_order:
table.add_column(mode, justify="right")
rows = [
("Instances", "n_instances", "d"),
@@ -288,7 +302,7 @@ def compare_modes(
for label, key, fmt in rows:
values = []
for mode in ["baseline", "mcp", "augment", "full"]:
for mode in mode_order:
if mode in metrics:
v = metrics[mode].get(key, 0)
if fmt == ".1%":
@@ -307,8 +321,8 @@ def compare_modes(
baseline_calls = metrics["baseline"]["avg_api_calls"]
table.add_section()
for mode in ["mcp", "augment", "full"]:
if mode not in metrics:
for mode in mode_order:
if mode == "baseline":
continue
mode_cost = metrics[mode]["avg_cost"]
mode_calls = metrics[mode]["avg_api_calls"]
@@ -340,10 +354,8 @@ def gitnexus_usage(
table = Table(title="Tool Usage by Run")
table.add_column("Run", style="bold")
table.add_column("query", justify="right")
table.add_column("context", justify="right")
table.add_column("impact", justify="right")
table.add_column("cypher", justify="right")
for key in TOOL_METRIC_KEYS:
table.add_column(key, justify="right")
table.add_column("Total", justify="right")
table.add_column("Augment Hits", justify="right")
@@ -353,7 +365,7 @@ def gitnexus_usage(
continue
# Aggregate tool calls across trajectories
tool_totals: dict[str, int] = {"query": 0, "context": 0, "impact": 0, "cypher": 0, "overview": 0}
tool_totals: dict[str, int] = {key: 0 for key in TOOL_METRIC_KEYS}
augment_hits = 0
for traj in run_data.get("trajectories", {}).values():
@@ -373,10 +385,7 @@ def gitnexus_usage(
if total > 0 or augment_hits > 0:
table.add_row(
run_id,
str(tool_totals.get("query", 0)),
str(tool_totals.get("context", 0)),
str(tool_totals.get("impact", 0)),
str(tool_totals.get("cypher", 0)),
*[str(tool_totals.get(key, 0)) for key in TOOL_METRIC_KEYS],
str(total),
str(augment_hits),
)
+77 -35
View File
@@ -17,6 +17,14 @@ import time
from pathlib import Path
from typing import Any
from constants import (
MCP_FIND_GITNEXUS_FALLBACK_TIMEOUT_SECONDS,
MCP_FIND_GITNEXUS_TIMEOUT_SECONDS,
MCP_READ_TIMEOUT_SECONDS,
MCP_STOP_WAIT_SECONDS,
)
from utils.errors import is_debug_enabled, log_safe_exception
logger = logging.getLogger("mcp_bridge")
@@ -78,7 +86,7 @@ class MCPBridge:
return True
except Exception as e:
logger.error(f"Failed to start MCP bridge: {e}")
log_safe_exception(logger, "Failed to start MCP bridge", e, include_debug=is_debug_enabled())
self.stop()
return False
@@ -86,9 +94,14 @@ class MCPBridge:
"""Stop the MCP server subprocess."""
if self.process:
try:
self.process.stdin.close()
if self.process.stdin:
self.process.stdin.close()
if self.process.stdout:
self.process.stdout.close()
if self.process.stderr:
self.process.stderr.close()
self.process.terminate()
self.process.wait(timeout=5)
self.process.wait(timeout=MCP_STOP_WAIT_SECONDS)
except Exception:
try:
self.process.kill()
@@ -146,7 +159,9 @@ class MCPBridge:
try:
result = subprocess.run(
[cmd, "gitnexus", "--version"],
capture_output=True, text=True, timeout=15,
capture_output=True,
text=True,
timeout=MCP_FIND_GITNEXUS_TIMEOUT_SECONDS,
cwd=self.repo_path,
)
if result.returncode == 0:
@@ -158,7 +173,9 @@ class MCPBridge:
try:
result = subprocess.run(
["gitnexus", "--version"],
capture_output=True, text=True, timeout=10,
capture_output=True,
text=True,
timeout=MCP_FIND_GITNEXUS_FALLBACK_TIMEOUT_SECONDS,
)
if result.returncode == 0:
return "gitnexus"
@@ -194,7 +211,7 @@ class MCPBridge:
self.process.stdin.flush()
# Read response
response = self._read_response(timeout=30)
response = self._read_response(timeout=MCP_READ_TIMEOUT_SECONDS)
if response and response.get("id") == request_id:
if "error" in response:
logger.error(f"MCP error: {response['error']}")
@@ -203,7 +220,7 @@ class MCPBridge:
return None
except Exception as e:
logger.error(f"MCP request failed: {e}")
log_safe_exception(logger, "MCP request failed", e, include_debug=is_debug_enabled())
return None
def _send_notification(self, method: str, params: dict):
@@ -224,42 +241,67 @@ class MCPBridge:
self.process.stdin.write(message.encode("utf-8"))
self.process.stdin.flush()
except Exception as e:
logger.error(f"MCP notification failed: {e}")
log_safe_exception(logger, "MCP notification failed", e, include_debug=is_debug_enabled())
def _read_response(self, timeout: float = 30) -> dict | None:
def _read_content_length(self, deadline: float) -> int | None:
"""Read Content-Length header, returning the byte length or None."""
if not self.process or not self.process.stdout:
return None
header_line = b""
while time.time() < deadline:
byte = self.process.stdout.read(1)
if not byte:
return None
header_line += byte
if header_line.endswith(b"\r\n\r\n") or header_line.endswith(b"\n\n"):
break
if not header_line:
return None
header_str = header_line.decode("utf-8").strip()
for line in header_str.split("\r\n"):
if line.lower().startswith("content-length:"):
try:
return int(line.split(":", 1)[1].strip())
except (ValueError, IndexError):
return None
return None
def _read_body(self, content_length: int, deadline: float) -> bytes | None:
"""Read a response body of the expected length before deadline."""
if not self.process or not self.process.stdout:
return None
remaining = content_length
chunks: list[bytes] = []
while remaining > 0 and time.time() < deadline:
chunk = self.process.stdout.read(remaining)
if not chunk:
return None
chunks.append(chunk)
remaining -= len(chunk)
if remaining > 0:
return None
return b"".join(chunks)
def _read_response(self, timeout: float = MCP_READ_TIMEOUT_SECONDS) -> dict | None:
"""Read a JSON-RPC response from the MCP server."""
if not self.process or not self.process.stdout:
return None
start = time.time()
try:
while time.time() - start < timeout:
# Read Content-Length header
header_line = b""
while True:
byte = self.process.stdout.read(1)
if not byte:
return None
header_line += byte
if header_line.endswith(b"\r\n\r\n"):
break
if header_line.endswith(b"\n\n"):
break
# Parse content length
header_str = header_line.decode("utf-8").strip()
content_length = None
for line in header_str.split("\r\n"):
if line.lower().startswith("content-length:"):
content_length = int(line.split(":")[1].strip())
break
deadline = time.time() + timeout
while time.time() < deadline:
content_length = self._read_content_length(deadline)
if content_length is None:
continue
# Read body
body = self.process.stdout.read(content_length)
body = self._read_body(content_length, deadline)
if not body:
return None
@@ -272,7 +314,7 @@ class MCPBridge:
return None
except Exception as e:
logger.error(f"Error reading MCP response: {e}")
log_safe_exception(logger, "Error reading MCP response", e, include_debug=is_debug_enabled())
return None
+2 -2
View File
@@ -1,8 +1,8 @@
# Claude Haiku 4.5 — fast, cheap, good baseline
# Via OpenRouter (set OPENROUTER_API_KEY in .env)
model:
model_name: "openrouter/anthropic/claude-haiku-4.5"
cost_tracking: "ignore_errors"
model_name: 'openrouter/anthropic/claude-haiku-4.5'
cost_tracking: 'ignore_errors'
model_kwargs:
max_tokens: 8192
temperature: 0
+2 -2
View File
@@ -2,8 +2,8 @@
# Via OpenRouter (set OPENROUTER_API_KEY in .env)
# To use Anthropic directly, change to: anthropic/claude-opus-4-20250514
model:
model_name: "openrouter/anthropic/claude-opus-4"
cost_tracking: "ignore_errors"
model_name: 'openrouter/anthropic/claude-opus-4'
cost_tracking: 'ignore_errors'
model_kwargs:
max_tokens: 16384
temperature: 0
+2 -2
View File
@@ -2,8 +2,8 @@
# Via OpenRouter (set OPENROUTER_API_KEY in .env)
# To use Anthropic directly, change to: anthropic/claude-sonnet-4-20250514
model:
model_name: "openrouter/anthropic/claude-sonnet-4"
cost_tracking: "ignore_errors"
model_name: 'openrouter/anthropic/claude-sonnet-4'
cost_tracking: 'ignore_errors'
model_kwargs:
max_tokens: 16384
temperature: 0
+3 -3
View File
@@ -1,9 +1,9 @@
model: deepseek-ai/deepseek-chat
provider: openrouter
cost:
input: 0.14 # per 1M tokens
output: 0.28 # per 1M tokens
input: 0.14 # per 1M tokens
output: 0.28 # per 1M tokens
# Native DeepSeek API (direct)
api_key: null
base_url: null
+3 -3
View File
@@ -1,9 +1,9 @@
model: deepseek-ai/DeepSeek-V3
provider: openrouter
cost:
input: 0.27 # per 1M tokens
output: 1.10 # per 1M tokens
input: 0.27 # per 1M tokens
output: 1.10 # per 1M tokens
# Native DeepSeek API (direct)
# Get your API key at: https://platform.deepseek.com/
# Or use OpenRouter with: OPENROUTER_API_KEY
+2 -2
View File
@@ -1,7 +1,7 @@
# GLM 4.7 — via OpenRouter (set OPENROUTER_API_KEY in .env)
model:
model_name: "openrouter/zhipuai/glm-4.7"
cost_tracking: "ignore_errors"
model_name: 'openrouter/zhipuai/glm-4.7'
cost_tracking: 'ignore_errors'
model_kwargs:
max_tokens: 8192
temperature: 0
+2 -2
View File
@@ -1,7 +1,7 @@
# GLM 5 — via OpenRouter (set OPENROUTER_API_KEY in .env)
model:
model_name: "openrouter/zhipuai/glm-5"
cost_tracking: "ignore_errors"
model_name: 'openrouter/zhipuai/glm-5'
cost_tracking: 'ignore_errors'
model_kwargs:
max_tokens: 8192
temperature: 0
+2 -2
View File
@@ -1,7 +1,7 @@
# MiniMax M1 2.5 — via OpenRouter (set OPENROUTER_API_KEY in .env)
model:
model_name: "openrouter/minimax/minimax-m1-2.5"
cost_tracking: "ignore_errors"
model_name: 'openrouter/minimax/minimax-m1-2.5'
cost_tracking: 'ignore_errors'
model_kwargs:
max_tokens: 8192
temperature: 0
+2 -2
View File
@@ -3,9 +3,9 @@
# The action_regex tells mini-swe-agent to parse ```bash blocks from responses.
model:
model_class: litellm_textbased
model_name: "openrouter/minimax/minimax-m2.5"
model_name: 'openrouter/minimax/minimax-m2.5'
action_regex: "```(?:bash|mswea_bash_command)\\s*\\n(.*?)\\n```"
cost_tracking: "ignore_errors"
cost_tracking: 'ignore_errors'
model_kwargs:
max_tokens: 8192
temperature: 0
+3 -3
View File
@@ -1,9 +1,9 @@
# Baseline mode — no GitNexus, pure mini-swe-agent (control group)
agent:
agent_class: "eval.agents.gitnexus_agent.GitNexusAgent"
gitnexus_mode: "baseline"
agent_class: 'eval.agents.gitnexus_agent.GitNexusAgent'
gitnexus_mode: 'baseline'
step_limit: 30
cost_limit: 3.0
environment:
environment_class: "docker"
environment_class: 'docker'
+3 -3
View File
@@ -5,14 +5,14 @@
#
# Use this mode to isolate the value of explicit tools without grep augmentation.
agent:
agent_class: "eval.agents.gitnexus_agent.GitNexusAgent"
gitnexus_mode: "native"
agent_class: 'eval.agents.gitnexus_agent.GitNexusAgent'
gitnexus_mode: 'native'
step_limit: 30
cost_limit: 3.0
track_gitnexus_usage: true
environment:
environment_class: "eval.environments.gitnexus_docker.GitNexusDockerEnvironment"
environment_class: 'eval.environments.gitnexus_docker.GitNexusDockerEnvironment'
enable_gitnexus: true
skip_embeddings: true
gitnexus_timeout: 120
+3 -3
View File
@@ -8,8 +8,8 @@
#
# The agent decides when to use explicit tools vs rely on enriched grep results.
agent:
agent_class: "eval.agents.gitnexus_agent.GitNexusAgent"
gitnexus_mode: "native_augment"
agent_class: 'eval.agents.gitnexus_agent.GitNexusAgent'
gitnexus_mode: 'native_augment'
step_limit: 30
cost_limit: 3.0
augment_timeout: 5.0
@@ -17,7 +17,7 @@ agent:
track_gitnexus_usage: true
environment:
environment_class: "eval.environments.gitnexus_docker.GitNexusDockerEnvironment"
environment_class: 'eval.environments.gitnexus_docker.GitNexusDockerEnvironment'
enable_gitnexus: true
skip_embeddings: true
gitnexus_timeout: 120
+15
View File
@@ -0,0 +1,15 @@
DEBUG_ENV_VAR = "GITNEXUS_EVAL_DEBUG"
# GitNexus eval-server health checks
EVAL_SERVER_HEALTH_RETRIES = 30
EVAL_SERVER_HEALTH_INTERVAL_SECONDS = 0.5
EVAL_SERVER_HEALTH_TIMEOUT_SECONDS = 3
# MCP bridge timeouts
MCP_FIND_GITNEXUS_TIMEOUT_SECONDS = 15
MCP_FIND_GITNEXUS_FALLBACK_TIMEOUT_SECONDS = 10
MCP_READ_TIMEOUT_SECONDS = 30
MCP_STOP_WAIT_SECONDS = 5
# Agent defaults
AUGMENT_TIMEOUT_SECONDS = 5.0
+75 -81
View File
@@ -26,72 +26,20 @@ import shutil
import time
from pathlib import Path
from constants import (
EVAL_SERVER_HEALTH_INTERVAL_SECONDS,
EVAL_SERVER_HEALTH_RETRIES,
EVAL_SERVER_HEALTH_TIMEOUT_SECONDS,
)
from minisweagent.environments.docker import DockerEnvironment
from tool_registry import TOOL_SPECS, ToolScriptSpec
from utils.errors import is_debug_enabled, log_safe_exception
logger = logging.getLogger("gitnexus_docker")
DEFAULT_CACHE_DIR = Path.home() / ".gitnexus-eval-cache"
EVAL_SERVER_PORT = 4848
# Standalone tool scripts installed into /usr/local/bin/ inside the container.
# Each script calls the eval-server via curl, with a CLI fallback.
# These are standalone — no sourcing, no env inheritance needed.
TOOL_SCRIPT_QUERY = r'''#!/bin/bash
PORT="${GITNEXUS_EVAL_PORT:-__PORT__}"
query="$1"; task_ctx="${2:-}"; goal="${3:-}"
[ -z "$query" ] && echo "Usage: gitnexus-query <query> [task_context] [goal]" && exit 1
args="{\"query\": \"$query\""
[ -n "$task_ctx" ] && args="$args, \"task_context\": \"$task_ctx\""
[ -n "$goal" ] && args="$args, \"goal\": \"$goal\""
args="$args}"
result=$(curl -sf -X POST "http://127.0.0.1:${PORT}/tool/query" -H "Content-Type: application/json" -d "$args" 2>/dev/null)
if [ $? -eq 0 ] && [ -n "$result" ]; then echo "$result"; exit 0; fi
cd /testbed && npx gitnexus query "$query" 2>&1
'''
TOOL_SCRIPT_CONTEXT = r'''#!/bin/bash
PORT="${GITNEXUS_EVAL_PORT:-__PORT__}"
name="$1"; file_path="${2:-}"
[ -z "$name" ] && echo "Usage: gitnexus-context <symbol_name> [file_path]" && exit 1
args="{\"name\": \"$name\""
[ -n "$file_path" ] && args="$args, \"file_path\": \"$file_path\""
args="$args}"
result=$(curl -sf -X POST "http://127.0.0.1:${PORT}/tool/context" -H "Content-Type: application/json" -d "$args" 2>/dev/null)
if [ $? -eq 0 ] && [ -n "$result" ]; then echo "$result"; exit 0; fi
cd /testbed && npx gitnexus context "$name" 2>&1
'''
TOOL_SCRIPT_IMPACT = r'''#!/bin/bash
PORT="${GITNEXUS_EVAL_PORT:-__PORT__}"
target="$1"; direction="${2:-upstream}"
[ -z "$target" ] && echo "Usage: gitnexus-impact <symbol_name> [upstream|downstream]" && exit 1
result=$(curl -sf -X POST "http://127.0.0.1:${PORT}/tool/impact" -H "Content-Type: application/json" -d "{\"target\": \"$target\", \"direction\": \"$direction\"}" 2>/dev/null)
if [ $? -eq 0 ] && [ -n "$result" ]; then echo "$result"; exit 0; fi
cd /testbed && npx gitnexus impact "$target" --direction "$direction" 2>&1
'''
TOOL_SCRIPT_CYPHER = r'''#!/bin/bash
PORT="${GITNEXUS_EVAL_PORT:-__PORT__}"
query="$1"
[ -z "$query" ] && echo "Usage: gitnexus-cypher <cypher_query>" && exit 1
result=$(curl -sf -X POST "http://127.0.0.1:${PORT}/tool/cypher" -H "Content-Type: application/json" -d "{\"query\": \"$query\"}" 2>/dev/null)
if [ $? -eq 0 ] && [ -n "$result" ]; then echo "$result"; exit 0; fi
cd /testbed && npx gitnexus cypher "$query" 2>&1
'''
TOOL_SCRIPT_AUGMENT = r'''#!/bin/bash
cd /testbed && npx gitnexus augment "$1" 2>&1 || true
'''
TOOL_SCRIPT_OVERVIEW = r'''#!/bin/bash
PORT="${GITNEXUS_EVAL_PORT:-__PORT__}"
echo "=== Code Knowledge Graph Overview ==="
result=$(curl -sf -X POST "http://127.0.0.1:${PORT}/tool/list_repos" -H "Content-Type: application/json" -d "{}" 2>/dev/null)
if [ $? -eq 0 ] && [ -n "$result" ]; then echo "$result"; exit 0; fi
cd /testbed && npx gitnexus list 2>&1
'''
class GitNexusDockerEnvironment(DockerEnvironment):
"""
@@ -133,7 +81,13 @@ class GitNexusDockerEnvironment(DockerEnvironment):
try:
self._setup_gitnexus()
except Exception as e:
logger.warning(f"GitNexus setup failed, continuing without it: {e}")
log_safe_exception(
logger,
"GitNexus setup failed, continuing without it",
e,
include_debug=is_debug_enabled(),
level="warning",
)
self._gitnexus_ready = False
return result
@@ -222,27 +176,59 @@ class GitNexusDockerEnvironment(DockerEnvironment):
"timeout": 5,
})
# Wait for the server to be ready (up to 15s for KuzuDB init)
for i in range(30):
time.sleep(0.5)
# Wait for the server to be ready (up to ~15s for KuzuDB init)
for i in range(EVAL_SERVER_HEALTH_RETRIES):
time.sleep(EVAL_SERVER_HEALTH_INTERVAL_SECONDS)
health = self.execute({
"command": f"curl -sf http://127.0.0.1:{self.eval_server_port}/health 2>/dev/null || echo 'NOT_READY'",
"timeout": 3,
"timeout": EVAL_SERVER_HEALTH_TIMEOUT_SECONDS,
})
output = health.get("output", "").strip()
if "NOT_READY" not in output and "ok" in output:
logger.info(f"Eval-server ready after {(i + 1) * 0.5:.1f}s")
logger.info(
f"Eval-server ready after {(i + 1) * EVAL_SERVER_HEALTH_INTERVAL_SECONDS:.1f}s"
)
return
log_output = self.execute({
"command": "cat /tmp/gitnexus-eval-server.log 2>/dev/null | tail -20",
})
logger.warning(
f"Eval-server didn't become ready in 15s. "
f"Eval-server didn't become ready in "
f"{EVAL_SERVER_HEALTH_RETRIES * EVAL_SERVER_HEALTH_INTERVAL_SECONDS:.1f}s. "
f"Tools will fall back to direct CLI.\n"
f"Server log: {log_output.get('output', 'N/A')}"
)
@staticmethod
def _render_tool_script(spec: ToolScriptSpec, port: str) -> str:
"""
Render a standalone bash script for a GitNexus tool.
Scripts call the eval-server fast path when an endpoint is present,
and fall back to the CLI otherwise.
"""
lines = ["#!/bin/bash"]
if spec.endpoint:
lines.append(f'PORT="${{GITNEXUS_EVAL_PORT:-{port}}}"')
if spec.header:
lines.append(spec.header.strip())
if spec.payload_builder:
lines.append(spec.payload_builder.strip())
if spec.endpoint:
lines.append(
f'result=$(curl -sf -X POST "http://127.0.0.1:${{PORT}}{spec.endpoint}" '
'-H "Content-Type: application/json" -d "$payload" 2>/dev/null)'
)
lines.append('if [ $? -eq 0 ] && [ -n "$result" ]; then echo "$result"; exit 0; fi')
lines.append(spec.fallback.strip())
return "\n".join(lines)
def _install_tools(self):
"""
Install standalone GitNexus tool scripts in /usr/local/bin/.
@@ -259,24 +245,20 @@ class GitNexusDockerEnvironment(DockerEnvironment):
"""
port = str(self.eval_server_port)
tools = {
"gitnexus-query": TOOL_SCRIPT_QUERY,
"gitnexus-context": TOOL_SCRIPT_CONTEXT,
"gitnexus-impact": TOOL_SCRIPT_IMPACT,
"gitnexus-cypher": TOOL_SCRIPT_CYPHER,
"gitnexus-augment": TOOL_SCRIPT_AUGMENT,
"gitnexus-overview": TOOL_SCRIPT_OVERVIEW,
}
for name, script in tools.items():
script_content = script.replace("__PORT__", port).strip()
for spec in TOOL_SPECS.values():
script_content = self._render_tool_script(spec, port).strip()
# Use heredoc with quoted delimiter — prevents all variable expansion and quoting issues
self.execute({
"command": f"cat << 'GITNEXUS_SCRIPT_EOF' > /usr/local/bin/{name}\n{script_content}\nGITNEXUS_SCRIPT_EOF\nchmod +x /usr/local/bin/{name}",
"command": (
f"cat << 'GITNEXUS_SCRIPT_EOF' > /usr/local/bin/{spec.bin_name}\n"
f"{script_content}\n"
"GITNEXUS_SCRIPT_EOF\n"
f"chmod +x /usr/local/bin/{spec.bin_name}"
),
"timeout": 5,
})
logger.info(f"Installed {len(tools)} GitNexus tool scripts in /usr/local/bin/")
logger.info(f"Installed {len(TOOL_SPECS)} GitNexus tool scripts in /usr/local/bin/")
def _get_repo_info(self) -> dict:
"""Get repository identity info from the container."""
@@ -325,7 +307,13 @@ class GitNexusDockerEnvironment(DockerEnvironment):
logger.info(f"Cached GitNexus index: {cache_path}")
except Exception as e:
logger.warning(f"Failed to cache GitNexus index: {e}")
log_safe_exception(
logger,
"Failed to cache GitNexus index",
e,
include_debug=is_debug_enabled(),
level="warning",
)
if cache_path.exists():
shutil.rmtree(cache_path, ignore_errors=True)
@@ -361,7 +349,13 @@ class GitNexusDockerEnvironment(DockerEnvironment):
logger.info("GitNexus index restored from cache")
except Exception as e:
logger.warning(f"Failed to restore cache, re-indexing: {e}")
log_safe_exception(
logger,
"Failed to restore cache, re-indexing",
e,
include_debug=is_debug_enabled(),
level="warning",
)
self._index_repository()
def stop(self) -> dict:
+5 -3
View File
@@ -6,7 +6,7 @@ readme = "README.md"
requires-python = ">=3.11"
dependencies = [
"mini-swe-agent>=2.0.0",
"litellm>=1.50.0",
"litellm>=1.50.0,!=1.82.7,!=1.82.8",
"datasets>=3.0.0",
"typer>=0.12.0",
"rich>=13.0.0",
@@ -20,6 +20,8 @@ dependencies = [
dev = [
"pytest>=8.0.0",
"ruff>=0.5.0",
"hypothesis>=6.88.0",
"coverage>=7.6.0",
]
[project.scripts]
@@ -31,8 +33,8 @@ requires = ["hatchling"]
build-backend = "hatchling.build"
[tool.hatch.build.targets.wheel]
packages = ["agents", "environments", "analysis", "bridge"]
extra-files = ["run_eval.py"]
packages = ["agents", "environments", "analysis", "bridge", "utils"]
extra-files = ["run_eval.py", "tool_registry.py", "constants.py"]
[tool.ruff]
line-length = 120
+73 -37
View File
@@ -25,7 +25,6 @@ import logging
import os
import threading
import time
import traceback
from itertools import product
from pathlib import Path
from typing import Any
@@ -36,6 +35,8 @@ from rich.console import Console
from rich.live import Live
from rich.table import Table
from utils.errors import is_debug_enabled, log_safe_exception
# Load .env file from eval/ directory
_env_file = Path(__file__).parent / ".env"
if _env_file.exists():
@@ -138,6 +139,65 @@ def get_swebench_docker_image(instance: dict) -> str:
return image_name
def _build_model(config: dict):
"""Construct the model from config."""
from minisweagent.models import get_model
return get_model(config=config.get("model", {}))
def _build_environment(config: dict, instance: dict):
"""Construct the environment for the instance."""
env_config = dict(config.get("environment", {}))
env_class_name = env_config.pop("environment_class", "docker")
if env_class_name == "eval.environments.gitnexus_docker.GitNexusDockerEnvironment":
from environments.gitnexus_docker import GitNexusDockerEnvironment
env_config["image"] = get_swebench_docker_image(instance)
return GitNexusDockerEnvironment(**env_config)
from minisweagent.environments.docker import DockerEnvironment
return DockerEnvironment(image=get_swebench_docker_image(instance), **env_config)
def _build_agent(config: dict, model, env, instance_dir: Path, instance_id: str):
"""Construct the GitNexus agent with trajectory output configured."""
from agents.gitnexus_agent import GitNexusAgent
agent_config = dict(config.get("agent", {}))
agent_config.pop("agent_class", "eval.agents.gitnexus_agent.GitNexusAgent")
traj_path = instance_dir / f"{instance_id}.traj.json"
agent_config["output_path"] = traj_path
return GitNexusAgent(model, env, **agent_config)
def _extract_submission(env, info: dict, run_id: str) -> str:
"""Pull the git diff patch from the container, falling back to the agent submission."""
try:
patch_output = env.execute({"command": "cd /testbed && git diff"})
return patch_output.get("output", "").strip()
except Exception as patch_err:
logger.warning(f"[{run_id}] Failed to extract patch: {patch_err}")
return info.get("submission", "")
def _record_failure(run_id: str, instance_id: str, result: dict, error: Exception):
sanitized = log_safe_exception(
logger,
f"[{run_id}] Error on {instance_id}",
error,
include_debug=is_debug_enabled(),
)
result["exit_status"] = sanitized["error_type"]
result["error_type"] = sanitized["error_type"]
result["error_message"] = sanitized["error_message"]
result["error"] = sanitized["error_message"]
if "error_detail_debug" in sanitized:
result["error_detail_debug"] = sanitized["error_detail_debug"]
def process_instance(
instance: dict,
config: dict,
@@ -149,8 +209,6 @@ def process_instance(
Process a single SWE-bench instance with the given config.
Returns result dict with instance_id, exit_status, submission, metrics.
"""
from minisweagent.models import get_model
instance_id = instance["instance_id"]
run_id = f"{model_name}_{mode_name}"
instance_dir = output_dir / run_id / instance_id
@@ -168,31 +226,12 @@ def process_instance(
}
agent = None
env = None
try:
# Build model
model = get_model(config=config.get("model", {}))
# Build environment
env_config = dict(config.get("environment", {}))
env_class_name = env_config.pop("environment_class", "docker")
if env_class_name == "eval.environments.gitnexus_docker.GitNexusDockerEnvironment":
from environments.gitnexus_docker import GitNexusDockerEnvironment
env_config["image"] = get_swebench_docker_image(instance)
env = GitNexusDockerEnvironment(**env_config)
else:
from minisweagent.environments.docker import DockerEnvironment
env = DockerEnvironment(image=get_swebench_docker_image(instance), **env_config)
# Build agent
agent_config = dict(config.get("agent", {}))
agent_class_name = agent_config.pop("agent_class", "eval.agents.gitnexus_agent.GitNexusAgent")
from agents.gitnexus_agent import GitNexusAgent
traj_path = instance_dir / f"{instance_id}.traj.json"
agent_config["output_path"] = traj_path
agent = GitNexusAgent(model, env, **agent_config)
model = _build_model(config)
env = _build_environment(config, instance)
agent = _build_agent(config, model, env, instance_dir, instance_id)
# Run
logger.info(f"[{run_id}] Starting {instance_id}")
@@ -204,18 +243,10 @@ def process_instance(
result["gitnexus_metrics"] = agent.gitnexus_metrics.to_dict()
# Extract git diff patch from the container (SWE-bench needs the model_patch)
try:
patch_output = env.execute({"command": "cd /testbed && git diff"})
result["submission"] = patch_output.get("output", "").strip()
except Exception as patch_err:
logger.warning(f"[{run_id}] Failed to extract patch: {patch_err}")
result["submission"] = info.get("submission", "")
result["submission"] = _extract_submission(env, info, run_id)
except Exception as e:
logger.error(f"[{run_id}] Error on {instance_id}: {e}")
result["exit_status"] = type(e).__name__
result["error"] = str(e)
result["traceback"] = traceback.format_exc()
_record_failure(run_id, instance_id, result, e)
finally:
if agent:
@@ -287,7 +318,12 @@ def run_configuration(
results.append(future.result())
except Exception as e:
iid = futures[future]
logger.error(f"[{run_id}] Uncaught error for {iid}: {e}")
log_safe_exception(
logger,
f"[{run_id}] Uncaught error for {iid}",
e,
include_debug=is_debug_enabled(),
)
# Save run summary
summary = {
+1
View File
@@ -0,0 +1 @@
"""Tests for the GitNexus eval harness."""
+6
View File
@@ -0,0 +1,6 @@
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
if str(ROOT) not in sys.path:
sys.path.insert(0, str(ROOT))
+29
View File
@@ -0,0 +1,29 @@
from utils.errors import sanitize_exception
def _raise_value_error():
raise ValueError("boom")
def test_sanitize_exception_without_debug(monkeypatch):
monkeypatch.delenv("GITNEXUS_EVAL_DEBUG", raising=False)
try:
_raise_value_error()
except Exception as exc: # noqa: BLE001
data = sanitize_exception(exc, include_debug=False)
assert data["error_type"] == "ValueError"
assert data["error_message"] == "boom"
assert "error_detail_debug" not in data
def test_sanitize_exception_with_debug(monkeypatch):
monkeypatch.setenv("GITNEXUS_EVAL_DEBUG", "1")
try:
_raise_value_error()
except Exception as exc: # noqa: BLE001
data = sanitize_exception(exc)
assert data["error_type"] == "ValueError"
assert "error_detail_debug" in data
assert "ValueError" in data["error_detail_debug"]
+20
View File
@@ -0,0 +1,20 @@
import pytest
analysis_module = pytest.importorskip("analysis.analyze_results")
parse_run_id = analysis_module.parse_run_id
def test_parse_run_id_native_augment():
model, mode = parse_run_id("claude-sonnet_native_augment")
assert model == "claude-sonnet"
assert mode == "native_augment"
def test_parse_run_id_hyphenated_model():
model, mode = parse_run_id("glm-4.7_native")
assert model == "glm-4.7"
assert mode == "native"
def test_parse_run_id_unknown():
assert parse_run_id("custom_model") == ("custom_model", "unknown")
+63
View File
@@ -0,0 +1,63 @@
from __future__ import annotations
from hypothesis import given
from hypothesis import strategies as st
from analysis.analyze_results import parse_run_id
from environments.gitnexus_docker import GitNexusDockerEnvironment
from tool_registry import TOOL_SPECS
from utils.errors import sanitize_exception
KNOWN_MODES = [
"native_augment",
"native",
"baseline",
"mcp",
"augment",
"full",
]
def _model_strategy():
base_chars = st.characters(
blacklist_categories=("Cs",),
blacklist_characters={" ", "\n", "\t"},
)
text = st.text(alphabet=base_chars, min_size=1)
return text.filter(lambda s: not any(s.endswith(f"_{m}") for m in KNOWN_MODES))
@given(_model_strategy(), st.sampled_from(KNOWN_MODES))
def test_parse_run_id_round_trip(model: str, mode: str) -> None:
run_id = f"{model}_{mode}"
parsed_model, parsed_mode = parse_run_id(run_id)
assert parsed_model == model
assert parsed_mode == mode
@given(st.text())
def test_sanitize_exception_respects_debug_flag(message: str) -> None:
exc = ValueError(message)
data = sanitize_exception(exc, include_debug=False)
assert data["error_type"] == "ValueError"
assert data["error_message"] == (message or "ValueError")
assert "error_detail_debug" not in data
data_debug = sanitize_exception(exc, include_debug=True)
assert data_debug["error_type"] == "ValueError"
assert "error_detail_debug" in data_debug
assert data_debug["error_detail_debug"]
@given(st.sampled_from(list(TOOL_SPECS.values())), st.integers(min_value=1, max_value=99999))
def test_render_tool_script_contains_expected_paths(spec, port: int) -> None:
script = GitNexusDockerEnvironment._render_tool_script(spec, str(port))
assert spec.fallback.strip() in script
if spec.endpoint:
assert spec.endpoint in script
assert f"${{GITNEXUS_EVAL_PORT:-{port}}}" in script
assert "curl" in script
else:
assert "curl" not in script
+21
View File
@@ -0,0 +1,21 @@
import pytest
GitNexusDockerEnvironment = pytest.importorskip(
"environments.gitnexus_docker"
).GitNexusDockerEnvironment
tool_registry = pytest.importorskip("tool_registry")
TOOL_SPECS = tool_registry.TOOL_SPECS
def test_render_query_script_uses_endpoint_and_fallback():
script = GitNexusDockerEnvironment._render_tool_script(TOOL_SPECS["query"], "4848")
assert "/tool/query" in script
assert "gitnexus query" in script
assert "GITNEXUS_EVAL_PORT" in script
def test_render_augment_script_skips_curl():
script = GitNexusDockerEnvironment._render_tool_script(TOOL_SPECS["augment"], "4848")
assert "/tool/" not in script
assert "curl" not in script
assert "gitnexus augment" in script
+79
View File
@@ -0,0 +1,79 @@
from __future__ import annotations
from dataclasses import dataclass
from typing import Dict, Tuple
@dataclass(frozen=True)
class ToolScriptSpec:
key: str
bin_name: str
endpoint: str | None
payload_builder: str
fallback: str
header: str | None = None
TOOL_METRIC_KEYS: Tuple[str, ...] = ("query", "context", "impact", "cypher", "overview")
TOOL_SPECS: Dict[str, ToolScriptSpec] = {
"query": ToolScriptSpec(
key="query",
bin_name="gitnexus-query",
endpoint="/tool/query",
payload_builder=r'''query="$1"; task_ctx="${2:-}"; goal="${3:-}"
[ -z "$query" ] && echo "Usage: gitnexus-query <query> [task_context] [goal]" && exit 1
payload="{\"query\": \"$query\""
[ -n "$task_ctx" ] && payload="$payload, \"task_context\": \"$task_ctx\""
[ -n "$goal" ] && payload="$payload, \"goal\": \"$goal\""
payload="$payload}"''',
fallback='cd /testbed && npx gitnexus query "$query" 2>&1',
),
"context": ToolScriptSpec(
key="context",
bin_name="gitnexus-context",
endpoint="/tool/context",
payload_builder=r'''name="$1"; file_path="${2:-}"
[ -z "$name" ] && echo "Usage: gitnexus-context <symbol_name> [file_path]" && exit 1
payload="{\"name\": \"$name\""
[ -n "$file_path" ] && payload="$payload, \"file_path\": \"$file_path\""
payload="$payload}"''',
fallback='cd /testbed && npx gitnexus context "$name" 2>&1',
),
"impact": ToolScriptSpec(
key="impact",
bin_name="gitnexus-impact",
endpoint="/tool/impact",
payload_builder=r'''target="$1"; direction="${2:-upstream}"
[ -z "$target" ] && echo "Usage: gitnexus-impact <symbol_name> [upstream|downstream]" && exit 1
payload="{\"target\": \"$target\", \"direction\": \"$direction\"}"''',
fallback='cd /testbed && npx gitnexus impact "$target" --direction "$direction" 2>&1',
),
"cypher": ToolScriptSpec(
key="cypher",
bin_name="gitnexus-cypher",
endpoint="/tool/cypher",
payload_builder=r'''query="$1"
[ -z "$query" ] && echo "Usage: gitnexus-cypher <cypher_query>" && exit 1
payload="{\"query\": \"$query\"}"''',
fallback='cd /testbed && npx gitnexus cypher "$query" 2>&1',
),
"overview": ToolScriptSpec(
key="overview",
bin_name="gitnexus-overview",
endpoint="/tool/list_repos",
header='echo "=== Code Knowledge Graph Overview ==="',
payload_builder='payload="{}"',
fallback='cd /testbed && npx gitnexus list 2>&1',
),
"augment": ToolScriptSpec(
key="augment",
bin_name="gitnexus-augment",
endpoint=None,
payload_builder="",
fallback='cd /testbed && npx gitnexus augment "$1" 2>&1 || true',
),
}
BINARIES_BY_KEY: Dict[str, str] = {spec.key: spec.bin_name for spec in TOOL_SPECS.values()}
ENDPOINTS_BY_KEY: Dict[str, str | None] = {spec.key: spec.endpoint for spec in TOOL_SPECS.values()}
+1
View File
@@ -0,0 +1 @@
"""Utility package for the eval harness."""
+60
View File
@@ -0,0 +1,60 @@
from __future__ import annotations
import os
import traceback
from typing import Any, Callable
from constants import DEBUG_ENV_VAR
def is_debug_enabled() -> bool:
"""Return True when debug output (full tracebacks) should be emitted."""
return os.getenv(DEBUG_ENV_VAR, "").strip().lower() in {"1", "true", "yes", "on"}
def sanitize_exception(exc: BaseException, *, include_debug: bool | None = None) -> dict[str, str]:
"""
Produce a log-safe, JSON-friendly view of an exception.
- Always returns error_type and error_message.
- Only includes error_detail_debug (full traceback) when debug is enabled.
"""
debug = is_debug_enabled() if include_debug is None else include_debug
error_type = type(exc).__name__
message = str(exc) or error_type
data: dict[str, str] = {
"error_type": error_type,
"error_message": message,
}
if debug:
tb = "".join(traceback.format_exception(type(exc), exc, exc.__traceback__))
if tb:
data["error_detail_debug"] = tb
return data
def log_safe_exception(
logger: Any,
prefix: str,
exc: BaseException,
*,
include_debug: bool | None = None,
level: str = "error",
) -> dict[str, str]:
"""
Log an exception without leaking stack traces unless debug is enabled.
Returns the sanitized dict so callers can persist it.
"""
data = sanitize_exception(exc, include_debug=include_debug)
debug = "error_detail_debug" in data
log_fn: Callable[..., None] = getattr(logger, level, logger.error)
message = f"{prefix}: {data['error_type']}: {data['error_message']}"
log_kwargs = {"exc_info": True} if debug else {}
log_fn(message, **log_kwargs)
return data
Generated
+2529
View File
File diff suppressed because it is too large Load Diff
+60 -25
View File
@@ -64,10 +64,26 @@ function extractPattern(toolName, toolInput) {
const tokens = cmd.split(/\s+/);
let foundCmd = false;
let skipNext = false;
const flagsWithValues = new Set(['-e', '-f', '-m', '-A', '-B', '-C', '-g', '--glob', '-t', '--type', '--include', '--exclude']);
const flagsWithValues = new Set([
'-e',
'-f',
'-m',
'-A',
'-B',
'-C',
'-g',
'--glob',
'-t',
'--type',
'--include',
'--exclude',
]);
for (const token of tokens) {
if (skipNext) { skipNext = false; continue; }
if (skipNext) {
skipNext = false;
continue;
}
if (!foundCmd) {
if (/\brg$|\bgrep$/.test(token)) foundCmd = true;
continue;
@@ -98,33 +114,42 @@ function runGitNexusCli(args, cwd, timeout) {
// Detect whether 'gitnexus' is on PATH (cheap check, no execution)
let useDirectBinary = false;
try {
const which = spawnSync(
isWin ? 'where' : 'which', ['gitnexus'],
{ encoding: 'utf-8', timeout: 3000, stdio: ['pipe', 'pipe', 'pipe'] }
);
const which = spawnSync(isWin ? 'where' : 'which', ['gitnexus'], {
encoding: 'utf-8',
timeout: 3000,
stdio: ['pipe', 'pipe', 'pipe'],
});
useDirectBinary = which.status === 0;
} catch { /* not on PATH */ }
} catch {
/* not on PATH */
}
if (useDirectBinary) {
return spawnSync(
isWin ? 'gitnexus.cmd' : 'gitnexus', args,
{ encoding: 'utf-8', timeout, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
return spawnSync(isWin ? 'gitnexus.cmd' : 'gitnexus', args, {
encoding: 'utf-8',
timeout,
cwd,
stdio: ['pipe', 'pipe', 'pipe'],
});
}
// npx fallback needs shell on Windows since npx is a .cmd script
return spawnSync(
isWin ? 'npx.cmd' : 'npx', ['-y', 'gitnexus', ...args],
{ encoding: 'utf-8', timeout: timeout + 5000, cwd, stdio: ['pipe', 'pipe', 'pipe'] }
);
return spawnSync(isWin ? 'npx.cmd' : 'npx', ['-y', 'gitnexus', ...args], {
encoding: 'utf-8',
timeout: timeout + 5000,
cwd,
stdio: ['pipe', 'pipe', 'pipe'],
});
}
/**
* Emit a hook response with additional context for the agent.
*/
function sendHookResponse(hookEventName, message) {
console.log(JSON.stringify({
hookSpecificOutput: { hookEventName, additionalContext: message }
}));
console.log(
JSON.stringify({
hookSpecificOutput: { hookEventName, additionalContext: message },
}),
);
}
/**
@@ -149,7 +174,9 @@ function handlePreToolUse(input) {
if (!child.error && child.status === 0) {
result = child.stderr || '';
}
} catch { /* graceful failure */ }
} catch {
/* graceful failure */
}
if (result && result.trim()) {
sendHookResponse('PreToolUse', result.trim());
@@ -185,10 +212,15 @@ function handlePostToolUse(input) {
let currentHead = '';
try {
const headResult = spawnSync('git', ['rev-parse', 'HEAD'], {
encoding: 'utf-8', timeout: 3000, cwd, stdio: ['pipe', 'pipe', 'pipe'],
encoding: 'utf-8',
timeout: 3000,
cwd,
stdio: ['pipe', 'pipe', 'pipe'],
});
currentHead = (headResult.stdout || '').trim();
} catch { return; }
} catch {
return;
}
if (!currentHead) return;
@@ -197,16 +229,19 @@ function handlePostToolUse(input) {
try {
const meta = JSON.parse(fs.readFileSync(path.join(gitNexusDir, 'meta.json'), 'utf-8'));
lastCommit = meta.lastCommit || '';
hadEmbeddings = (meta.stats && meta.stats.embeddings > 0);
} catch { /* no meta — treat as stale */ }
hadEmbeddings = meta.stats && meta.stats.embeddings > 0;
} catch {
/* no meta — treat as stale */
}
// If HEAD matches last indexed commit, no reindex needed
if (currentHead && currentHead === lastCommit) return;
const analyzeCmd = `npx gitnexus analyze${hadEmbeddings ? ' --embeddings' : ''}`;
sendHookResponse('PostToolUse',
sendHookResponse(
'PostToolUse',
`GitNexus index is stale (last indexed: ${lastCommit ? lastCommit.slice(0, 7) : 'never'}). ` +
`Run \`${analyzeCmd}\` to update the knowledge graph.`
`Run \`${analyzeCmd}\` to update the knowledge graph.`,
);
}
+29
View File
@@ -0,0 +1,29 @@
{
"name": "gitnexus-shared",
"version": "1.0.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "gitnexus-shared",
"version": "1.0.0",
"devDependencies": {
"typescript": "^6.0.2"
}
},
"node_modules/typescript": {
"version": "6.0.2",
"resolved": "https://registry.npmjs.org/typescript/-/typescript-6.0.2.tgz",
"integrity": "sha512-bGdAIrZ0wiGDo5l8c++HWtbaNCWTS4UTv7RaTH/ThVIgjkveJt83m74bBHMJkuCbslY8ixgLBVZJIOiQlQTjfQ==",
"dev": true,
"license": "Apache-2.0",
"bin": {
"tsc": "bin/tsc",
"tsserver": "bin/tsserver"
},
"engines": {
"node": ">=14.17"
}
}
}
}
+25
View File
@@ -0,0 +1,25 @@
{
"name": "gitnexus-shared",
"version": "1.0.0",
"private": true,
"description": "Shared type definitions for GitNexus CLI and web",
"type": "module",
"main": "dist/index.js",
"types": "dist/index.d.ts",
"exports": {
".": {
"types": "./dist/index.d.ts",
"default": "./dist/index.js"
}
},
"scripts": {
"build": "tsc"
},
"files": [
"dist",
"src"
],
"devDependencies": {
"typescript": "^6.0.2"
}
}
+134
View File
@@ -0,0 +1,134 @@
/**
* Graph type definitions — single source of truth.
*
* Both gitnexus (CLI) and gitnexus-web import from this package.
* Do NOT add Node.js-specific or browser-specific imports here.
*/
import { SupportedLanguages } from '../languages.js';
export type NodeLabel =
| 'Project'
| 'Package'
| 'Module'
| 'Folder'
| 'File'
| 'Class'
| 'Function'
| 'Method'
| 'Variable'
| 'Interface'
| 'Enum'
| 'Decorator'
| 'Import'
| 'Type'
| 'CodeElement'
| 'Community'
| 'Process'
// Multi-language node types
| 'Struct'
| 'Macro'
| 'Typedef'
| 'Union'
| 'Namespace'
| 'Trait'
| 'Impl'
| 'TypeAlias'
| 'Const'
| 'Static'
| 'Property'
| 'Record'
| 'Delegate'
| 'Annotation'
| 'Constructor'
| 'Template'
| 'Section'
| 'Route'
| 'Tool';
export type NodeProperties = {
name: string;
filePath: string;
startLine?: number;
endLine?: number;
language?: SupportedLanguages | string;
isExported?: boolean;
astFrameworkMultiplier?: number;
astFrameworkReason?: string;
// Community
heuristicLabel?: string;
cohesion?: number;
symbolCount?: number;
keywords?: string[];
description?: string;
enrichedBy?: 'heuristic' | 'llm';
// Process
processType?: 'intra_community' | 'cross_community';
stepCount?: number;
communities?: string[];
entryPointId?: string;
terminalId?: string;
entryPointScore?: number;
entryPointReason?: string;
// Method/property
parameterCount?: number;
level?: number;
returnType?: string;
declaredType?: string;
visibility?: string;
isStatic?: boolean;
isReadonly?: boolean;
isAbstract?: boolean;
isFinal?: boolean;
isVirtual?: boolean;
isOverride?: boolean;
isAsync?: boolean;
isPartial?: boolean;
annotations?: string[];
// Route/response
responseKeys?: string[];
errorKeys?: string[];
middleware?: string[];
// Extensible
[key: string]: unknown;
};
export type RelationshipType =
| 'CONTAINS'
| 'CALLS'
| 'INHERITS'
| 'METHOD_OVERRIDES'
| 'METHOD_IMPLEMENTS'
| 'IMPORTS'
| 'USES'
| 'DEFINES'
| 'DECORATES'
| 'IMPLEMENTS'
| 'EXTENDS'
| 'HAS_METHOD'
| 'HAS_PROPERTY'
| 'ACCESSES'
| 'MEMBER_OF'
| 'STEP_IN_PROCESS'
| 'HANDLES_ROUTE'
| 'FETCHES'
| 'HANDLES_TOOL'
| 'ENTRY_POINT_OF'
| 'WRAPS'
| 'QUERIES';
export interface GraphNode {
id: string;
label: NodeLabel;
properties: NodeProperties;
}
export interface GraphRelationship {
id: string;
sourceId: string;
targetId: string;
type: RelationshipType;
confidence: number;
reason: string;
step?: number;
}
+24
View File
@@ -0,0 +1,24 @@
// Graph types
export type {
NodeLabel,
NodeProperties,
RelationshipType,
GraphNode,
GraphRelationship,
} from './graph/types.js';
// Schema constants
export {
NODE_TABLES,
REL_TABLE_NAME,
REL_TYPES,
EMBEDDING_TABLE_NAME,
} from './lbug/schema-constants.js';
export type { NodeTableName, RelType } from './lbug/schema-constants.js';
// Language support
export { SupportedLanguages } from './languages.js';
export { getLanguageFromFilename, getSyntaxLanguageFromFilename } from './language-detection.js';
// Pipeline progress
export type { PipelinePhase, PipelineProgress } from './pipeline.js';
+148
View File
@@ -0,0 +1,148 @@
/**
* Language Detection — maps file paths to SupportedLanguages enum values.
*
* Shared between CLI (ingestion pipeline) and web (syntax highlighting).
*
* ADDING A NEW LANGUAGE:
* 1. Add enum member to SupportedLanguages in languages.ts
* 2. Add file extensions to EXTENSION_MAP below
* 3. TypeScript will error if you miss either step (exhaustive Record)
*/
import { SupportedLanguages } from './languages.js';
/** Ruby extensionless filenames recognised as Ruby source */
const RUBY_EXTENSIONLESS_FILES = new Set([
'Rakefile',
'Gemfile',
'Guardfile',
'Vagrantfile',
'Brewfile',
]);
/**
* Exhaustive map: every SupportedLanguages member → its file extensions.
*
* If a new language is added to the enum without adding an entry here,
* TypeScript emits a compile error: "Property 'NewLang' is missing in type..."
*/
const EXTENSION_MAP: Record<SupportedLanguages, readonly string[]> = {
[SupportedLanguages.JavaScript]: ['.js', '.jsx', '.mjs', '.cjs'],
[SupportedLanguages.TypeScript]: ['.ts', '.tsx', '.mts', '.cts'],
[SupportedLanguages.Python]: ['.py'],
[SupportedLanguages.Java]: ['.java'],
[SupportedLanguages.C]: ['.c'],
[SupportedLanguages.CPlusPlus]: ['.cpp', '.cc', '.cxx', '.h', '.hpp', '.hxx', '.hh'],
[SupportedLanguages.CSharp]: ['.cs'],
[SupportedLanguages.Go]: ['.go'],
[SupportedLanguages.Ruby]: ['.rb', '.rake', '.gemspec'],
[SupportedLanguages.Rust]: ['.rs'],
[SupportedLanguages.PHP]: ['.php', '.phtml', '.php3', '.php4', '.php5', '.php8'],
[SupportedLanguages.Kotlin]: ['.kt', '.kts'],
[SupportedLanguages.Swift]: ['.swift'],
[SupportedLanguages.Dart]: ['.dart'],
[SupportedLanguages.Vue]: ['.vue'],
[SupportedLanguages.Cobol]: ['.cbl', '.cob', '.cpy', '.cobol'],
} satisfies Record<SupportedLanguages, readonly string[]>; // Ensure exhaustiveness
/** Pre-built reverse lookup: extension → language (built once at module load). */
const extToLang = new Map<string, SupportedLanguages>();
for (const [lang, exts] of Object.entries(EXTENSION_MAP) as [
SupportedLanguages,
readonly string[],
][]) {
for (const ext of exts) {
extToLang.set(ext, lang);
}
}
/**
* Map file extension to SupportedLanguage enum.
* Returns null if the file extension is not recognized.
*/
export const getLanguageFromFilename = (filename: string): SupportedLanguages | null => {
// Fast path: check the extension map
const lastDot = filename.lastIndexOf('.');
if (lastDot >= 0) {
const ext = filename.slice(lastDot).toLowerCase();
const lang = extToLang.get(ext);
if (lang !== undefined) return lang;
}
// Ruby extensionless files (Rakefile, Gemfile, etc.)
const basename = filename.split('/').pop() || filename;
if (RUBY_EXTENSIONLESS_FILES.has(basename)) {
return SupportedLanguages.Ruby;
}
return null;
};
/**
* Exhaustive map: every SupportedLanguages member → Prism syntax identifier.
*
* If a new language is added to the enum without adding an entry here,
* TypeScript emits a compile error.
*/
const SYNTAX_MAP: Record<SupportedLanguages, string> = {
[SupportedLanguages.JavaScript]: 'javascript',
[SupportedLanguages.TypeScript]: 'typescript',
[SupportedLanguages.Python]: 'python',
[SupportedLanguages.Java]: 'java',
[SupportedLanguages.C]: 'c',
[SupportedLanguages.CPlusPlus]: 'cpp',
[SupportedLanguages.CSharp]: 'csharp',
[SupportedLanguages.Go]: 'go',
[SupportedLanguages.Ruby]: 'ruby',
[SupportedLanguages.Rust]: 'rust',
[SupportedLanguages.PHP]: 'php',
[SupportedLanguages.Kotlin]: 'kotlin',
[SupportedLanguages.Swift]: 'swift',
[SupportedLanguages.Dart]: 'dart',
[SupportedLanguages.Vue]: 'typescript',
[SupportedLanguages.Cobol]: 'cobol',
} satisfies Record<SupportedLanguages, string>; // Ensure exhaustiveness
/** Non-code file extensions → Prism-compatible syntax identifiers */
const AUXILIARY_SYNTAX_MAP: Record<string, string> = {
json: 'json',
yaml: 'yaml',
yml: 'yaml',
md: 'markdown',
mdx: 'markdown',
html: 'markup',
htm: 'markup',
erb: 'markup',
xml: 'markup',
css: 'css',
scss: 'css',
sass: 'css',
sh: 'bash',
bash: 'bash',
zsh: 'bash',
sql: 'sql',
toml: 'toml',
ini: 'ini',
dockerfile: 'docker',
};
/** Extensionless filenames → Prism-compatible syntax identifiers */
const AUXILIARY_BASENAME_MAP: Record<string, string> = {
Makefile: 'makefile',
Dockerfile: 'docker',
};
/**
* Map file path to a Prism-compatible syntax highlight language string.
* Covers all SupportedLanguages (code files) plus common non-code formats.
* Returns 'text' for unrecognised files.
*/
export const getSyntaxLanguageFromFilename = (filePath: string): string => {
const lang = getLanguageFromFilename(filePath);
if (lang) return SYNTAX_MAP[lang];
const ext = filePath.split('.').pop()?.toLowerCase();
if (ext && ext in AUXILIARY_SYNTAX_MAP) return AUXILIARY_SYNTAX_MAP[ext];
const basename = filePath.split('/').pop() || '';
if (basename in AUXILIARY_BASENAME_MAP) return AUXILIARY_BASENAME_MAP[basename];
return 'text';
};
+25
View File
@@ -0,0 +1,25 @@
/**
* Supported language enum — single source of truth.
*
* Both CLI and web use this to identify which language a file/node belongs to.
* The CLI uses it throughout the ingestion pipeline; the web uses it for display.
*/
export enum SupportedLanguages {
JavaScript = 'javascript',
TypeScript = 'typescript',
Python = 'python',
Java = 'java',
C = 'c',
CPlusPlus = 'cpp',
CSharp = 'csharp',
Go = 'go',
Ruby = 'ruby',
Rust = 'rust',
PHP = 'php',
Kotlin = 'kotlin',
Swift = 'swift',
Dart = 'dart',
Vue = 'vue',
/** Standalone regex processor — no tree-sitter, no LanguageProvider. */
Cobol = 'cobol',
}
@@ -0,0 +1,73 @@
/**
* LadybugDB schema constants — single source of truth.
*
* NODE_TABLES and REL_TYPES define what the knowledge graph can contain.
* Both CLI and web must agree on these for data compatibility.
*
* Full DDL schemas remain in each package's own schema.ts because
* the CLI uses native LadybugDB and the web uses WASM.
*/
export const NODE_TABLES = [
'File',
'Folder',
'Function',
'Class',
'Interface',
'Method',
'CodeElement',
'Community',
'Process',
'Section',
'Struct',
'Enum',
'Macro',
'Typedef',
'Union',
'Namespace',
'Trait',
'Impl',
'TypeAlias',
'Const',
'Static',
'Property',
'Record',
'Delegate',
'Annotation',
'Constructor',
'Template',
'Module',
'Route',
'Tool',
] as const;
export type NodeTableName = (typeof NODE_TABLES)[number];
export const REL_TABLE_NAME = 'CodeRelation';
export const REL_TYPES = [
'CONTAINS',
'DEFINES',
'IMPORTS',
'CALLS',
'EXTENDS',
'IMPLEMENTS',
'HAS_METHOD',
'HAS_PROPERTY',
'ACCESSES',
'METHOD_OVERRIDES',
'OVERRIDES', // Legacy compat alias — kept until all stored indexes are migrated
'METHOD_IMPLEMENTS',
'MEMBER_OF',
'STEP_IN_PROCESS',
'HANDLES_ROUTE',
'FETCHES',
'HANDLES_TOOL',
'ENTRY_POINT_OF',
'WRAPS',
'QUERIES',
] as const;
export type RelType = (typeof REL_TYPES)[number];
export const EMBEDDING_TABLE_NAME = 'CodeEmbedding';
+29
View File
@@ -0,0 +1,29 @@
/**
* Pipeline progress types — shared between CLI and web.
*/
export type PipelinePhase =
| 'idle'
| 'extracting'
| 'structure'
| 'parsing'
| 'imports'
| 'calls'
| 'heritage'
| 'communities'
| 'processes'
| 'enriching'
| 'complete'
| 'error';
export interface PipelineProgress {
phase: PipelinePhase;
percent: number;
message: string;
detail?: string;
stats?: {
filesProcessed: number;
totalFiles: number;
nodesCreated: number;
};
}
+18
View File
@@ -0,0 +1,18 @@
{
"compilerOptions": {
"target": "ES2022",
"module": "NodeNext",
"moduleResolution": "NodeNext",
"strict": true,
"declaration": true,
"declarationMap": true,
"sourceMap": true,
"esModuleInterop": true,
"skipLibCheck": true,
"forceConsistentCasingInFileNames": true,
"composite": true,
"rootDir": "src",
"outDir": "dist"
},
"include": ["src"]
}
-104
View File
@@ -1,104 +0,0 @@
import type { VercelRequest, VercelResponse } from '@vercel/node';
/**
* CORS Proxy for isomorphic-git
*
* isomorphic-git calls: /api/proxy?url=https://github.com/...
*/
export default async function handler(req: VercelRequest, res: VercelResponse) {
// Handle CORS preflight
if (req.method === 'OPTIONS') {
res.setHeader('Access-Control-Allow-Origin', '*');
res.setHeader('Access-Control-Allow-Methods', 'GET, POST, OPTIONS');
res.setHeader('Access-Control-Allow-Headers', 'Content-Type, Authorization, Git-Protocol, Accept');
res.status(200).end();
return;
}
// Get URL from query parameter
const { url } = req.query;
if (!url || typeof url !== 'string') {
res.status(400).json({ error: 'Missing url query parameter' });
return;
}
// Only allow GitHub URLs for security
const allowedHosts = ['github.com', 'raw.githubusercontent.com'];
let parsedUrl: URL;
try {
parsedUrl = new URL(url);
} catch {
res.status(400).json({ error: 'Invalid URL' });
return;
}
if (!allowedHosts.some(host => parsedUrl.hostname.endsWith(host))) {
res.status(403).json({ error: 'Only GitHub URLs are allowed' });
return;
}
try {
const headers: Record<string, string> = {
'User-Agent': 'git/isomorphic-git',
};
// Forward relevant headers
if (req.headers.authorization) {
headers['Authorization'] = req.headers.authorization as string;
}
if (req.headers['content-type']) {
headers['Content-Type'] = req.headers['content-type'] as string;
}
if (req.headers['git-protocol']) {
headers['Git-Protocol'] = req.headers['git-protocol'] as string;
}
if (req.headers.accept) {
headers['Accept'] = req.headers.accept as string;
}
// Get request body for POST requests
let body: Buffer | undefined;
if (req.method === 'POST') {
const chunks: Buffer[] = [];
for await (const chunk of req) {
chunks.push(typeof chunk === 'string' ? Buffer.from(chunk) : chunk);
}
body = Buffer.concat(chunks);
}
const response = await fetch(url, {
method: req.method || 'GET',
headers,
body: body ? new Uint8Array(body) : undefined,
});
// Set CORS headers
res.setHeader('Access-Control-Allow-Origin', '*');
res.setHeader('Access-Control-Expose-Headers', '*');
// Forward response headers (except ones that cause issues)
const skipHeaders = [
'content-encoding',
'transfer-encoding',
'connection',
'www-authenticate', // IMPORTANT: Strip this to prevent browser's native auth popup!
];
response.headers.forEach((value, key) => {
if (!skipHeaders.includes(key.toLowerCase())) {
res.setHeader(key, value);
}
});
res.status(response.status);
const buffer = await response.arrayBuffer();
res.send(Buffer.from(buffer));
} catch (error) {
console.error('Proxy error:', error);
res.status(500).json({ error: 'Proxy request failed', details: String(error) });
}
}
+10 -4
View File
@@ -1,4 +1,4 @@
import { test, expect, type TestInfo } from '@playwright/test';
import { test, expect } from '@playwright/test';
/**
* Debug harnesses for investigating specific UI issues.
@@ -9,7 +9,7 @@ const BACKEND_URL = process.env.BACKEND_URL ?? 'http://localhost:4747';
const debugTest = process.env.DEBUG_E2E ? test : test.skip;
async function connectToServer(page: import('@playwright/test').Page) {
page.on('console', msg => {
page.on('console', (msg) => {
if (msg.type() === 'error') console.log(`[error] ${msg.text()}`);
});
@@ -86,7 +86,10 @@ debugTest('debug: process view Reset View button', async ({ page }, testInfo) =>
const transformAfterReset = await diagramDiv.getAttribute('style');
console.log('Transform AFTER reset:', transformAfterReset);
await page.screenshot({ path: testInfo.outputPath('debug-modal-after-reset.png'), fullPage: true });
await page.screenshot({
path: testInfo.outputPath('debug-modal-after-reset.png'),
fullPage: true,
});
// Verify transform actually changed back
expect(transformAfterZoom).not.toBe(transformBefore);
@@ -123,5 +126,8 @@ debugTest('debug: lightbulb clears node selection dimming', async ({ page }, tes
// Click it again to toggle back on
await lightbulbBtn.click();
await page.waitForTimeout(500);
await page.screenshot({ path: testInfo.outputPath('debug-after-lightbulb-toggle-back.png'), fullPage: true });
await page.screenshot({
path: testInfo.outputPath('debug-after-lightbulb-toggle-back.png'),
fullPage: true,
});
});
@@ -0,0 +1,109 @@
import { test, expect } from '@playwright/test';
/**
* E2E tests for heartbeat disconnect/reconnect behavior.
*
* Verifies the key regression: when the heartbeat fails, the UI shows a
* "reconnecting" banner instead of resetting to the onboarding screen.
*
* Strategy: block /api/heartbeat via route interception BEFORE loading the
* graph. The heartbeat EventSource can never connect, so onReconnecting
* fires on the first retry attempt. This reliably tests the banner behavior
* without depending on setOffline timing (which varies across CI environments).
*/
const BACKEND_URL = process.env.BACKEND_URL ?? 'http://localhost:4747';
const FRONTEND_URL = process.env.FRONTEND_URL ?? 'http://localhost:5173';
test.beforeAll(async () => {
if (process.env.E2E) return;
try {
const [backendRes, frontendRes] = await Promise.allSettled([
fetch(`${BACKEND_URL}/api/repos`),
fetch(FRONTEND_URL),
]);
if (
backendRes.status === 'rejected' ||
(backendRes.status === 'fulfilled' && !backendRes.value.ok)
) {
test.skip(true, 'gitnexus serve not available');
return;
}
if (
frontendRes.status === 'rejected' ||
(frontendRes.status === 'fulfilled' && !frontendRes.value.ok)
) {
test.skip(true, 'Vite dev server not available');
return;
}
if (backendRes.status === 'fulfilled') {
const repos = await backendRes.value.json();
if (!repos.length) {
test.skip(true, 'No indexed repos');
return;
}
}
} catch {
test.skip(true, 'servers not available');
}
});
test.describe('Heartbeat Reconnect', () => {
test('shows reconnecting banner instead of onboarding reset when heartbeat is unavailable', async ({
page,
}) => {
// Block the heartbeat BEFORE navigating — the EventSource will fail
// immediately on every connection attempt, triggering onReconnecting.
await page.route('**/api/heartbeat', (route) => route.abort('connectionrefused'));
// Load the app and connect to a repo (all other endpoints work normally)
await page.goto('/');
const landingCard = page.locator('[data-testid="landing-repo-card"]').first();
try {
await landingCard.waitFor({ state: 'visible', timeout: 15_000 });
await landingCard.click();
} catch {
// auto-connect may skip the landing screen
}
// Wait for graph to load (heartbeat is blocked, but graph loads fine)
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// The reconnecting banner should appear (heartbeat is failing)
const banner = page.getByText('Server connection lost');
await expect(banner).toBeVisible({ timeout: 15_000 });
// The graph canvas should STILL be visible — NOT reset to onboarding
await expect(page.locator('canvas').first()).toBeVisible();
});
test('banner clears when heartbeat becomes available', async ({ page }) => {
// Start with heartbeat blocked
await page.route('**/api/heartbeat', (route) => route.abort('connectionrefused'));
await page.goto('/');
const landingCard = page.locator('[data-testid="landing-repo-card"]').first();
try {
await landingCard.waitFor({ state: 'visible', timeout: 15_000 });
await landingCard.click();
} catch {
// auto-connect may skip the landing screen
}
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// Verify banner appears
const banner = page.getByText('Server connection lost');
await expect(banner).toBeVisible({ timeout: 15_000 });
// Unblock heartbeat — the real server is running, so reconnect will succeed
await page.unroute('**/api/heartbeat');
// Banner should disappear as heartbeat reconnects
await expect(banner).not.toBeVisible({ timeout: 30_000 });
// Graph should still be there
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible();
});
});
+3 -3
View File
@@ -12,16 +12,16 @@ import { test } from '@playwright/test';
*/
test.skip(
!!process.env.CI || process.env.PWDEBUG !== '1',
'Manual recording requires --headed and PWDEBUG=1. Run: PWDEBUG=1 npx playwright test e2e/manual-record.spec.ts --headed --timeout=0'
'Manual recording requires --headed and PWDEBUG=1. Run: PWDEBUG=1 npx playwright test e2e/manual-record.spec.ts --headed --timeout=0',
);
test('manual recording session', async ({ page }) => {
page.on('console', msg => {
page.on('console', (msg) => {
if (msg.type() === 'error' || msg.type() === 'warning') {
console.log(`[${msg.type()}] ${msg.text()}`);
}
});
page.on('pageerror', err => console.log(`[crash] ${err.message}`));
page.on('pageerror', (err) => console.log(`[crash] ${err.message}`));
await page.goto('http://localhost:5173');
await page.pause();
+112
View File
@@ -0,0 +1,112 @@
import { test, expect } from '@playwright/test';
/**
* E2E tests for multi-repo scoping and URL persistence.
*
* Verifies that:
* - Connecting via ?server= loads data and sets ?project= in the URL
* - The repo name appears in the UI after connecting
* - F5 with ?server=&project= reconnects to the correct repo
*
* Runs against the single indexed repo in CI — validates the plumbing
* works end-to-end even with one repo.
*/
const BACKEND_URL = process.env.BACKEND_URL ?? 'http://localhost:4747';
const FRONTEND_URL = process.env.FRONTEND_URL ?? 'http://localhost:5173';
let firstRepoName: string;
test.beforeAll(async () => {
if (process.env.E2E) {
// Still need to fetch the repo name for assertions
try {
const res = await fetch(`${BACKEND_URL}/api/repos`);
const repos = await res.json();
firstRepoName = repos[0]?.name ?? '';
} catch {
firstRepoName = '';
}
return;
}
try {
const [backendRes, frontendRes] = await Promise.allSettled([
fetch(`${BACKEND_URL}/api/repos`),
fetch(FRONTEND_URL),
]);
if (
backendRes.status === 'rejected' ||
(backendRes.status === 'fulfilled' && !backendRes.value.ok)
) {
test.skip(true, 'gitnexus serve not available');
return;
}
if (
frontendRes.status === 'rejected' ||
(frontendRes.status === 'fulfilled' && !frontendRes.value.ok)
) {
test.skip(true, 'Vite dev server not available');
return;
}
if (backendRes.status === 'fulfilled') {
const repos = await backendRes.value.json();
if (!repos.length) {
test.skip(true, 'No indexed repos');
return;
}
firstRepoName = repos[0].name;
}
} catch {
test.skip(true, 'servers not available');
}
});
test.describe('Multi-Repo Scoping', () => {
test('auto-connect via ?server= sets ?project= in URL', async ({ page }) => {
// Navigate with ?server= param (the bookmarkable shortcut)
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
// Wait for graph to load
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// URL should now contain ?project= with the repo name
const url = new URL(page.url());
const project = url.searchParams.get('project');
expect(project).toBeTruthy();
expect(project).toBe(firstRepoName);
});
test('?server= is preserved in URL for F5 recovery', async ({ page }) => {
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// URL should still have ?server=
const url = new URL(page.url());
expect(url.searchParams.get('server')).toBeTruthy();
// F5 should reconnect (not show onboarding)
await page.reload();
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
});
test('node count in status bar matches backend data', async ({ page }) => {
await page.goto(`/?server=${encodeURIComponent(BACKEND_URL)}`);
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// Fetch expected node count from backend
const res = await fetch(`${BACKEND_URL}/api/repo?repo=${encodeURIComponent(firstRepoName)}`);
const repoInfo = await res.json();
const expectedNodes = repoInfo.stats?.nodes;
if (expectedNodes) {
// Status bar shows node count — use the status-ready area to avoid
// matching multiple elements (file tree, header may also show counts)
const statusBar = page.locator('footer');
const nodeText = statusBar.getByText(/\d+ nodes/).first();
await expect(nodeText).toBeVisible({ timeout: 10_000 });
const text = await nodeText.textContent();
const displayedNodes = parseInt(text?.match(/(\d+)\s*nodes/)?.[1] ?? '0', 10);
expect(displayedNodes).toBeGreaterThan(0);
}
});
});
+300
View File
@@ -0,0 +1,300 @@
import { test, expect } from '@playwright/test';
/**
* E2E tests for the onboarding and analysis user flows.
*
* These tests cover:
* - Flow 1: OnboardingGuide shown when no server is running
* - Flow 2: Analyze form when server has zero repos
* - Flow 3: Auto-connect when server has repos
* - Flow 4: Repo dropdown in exploring view
*
* Most tests mock the backend at the network level so they don't
* require a live gitnexus server.
*/
const BACKEND_URL = 'http://localhost:4747';
async function enterExploringView(page: import('@playwright/test').Page) {
await page.goto('/');
const landingCard = page.locator('[data-testid="landing-repo-card"]').first();
try {
await landingCard.waitFor({ state: 'visible', timeout: 15_000 });
await landingCard.click();
} catch {
// Landing screen may not appear (e.g. ?server auto-connect)
}
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
}
// ── Flow 1: Onboarding (no server running) ─────────────────────────────────
test.describe('Flow 1: Onboarding — no server', () => {
test('shows OnboardingGuide when backend is unreachable', async ({ page }, testInfo) => {
// Block all requests to the backend so the probe fails
await page.route(`${BACKEND_URL}/**`, (route) => route.abort('connectionrefused'));
await page.goto('/');
// Wait for initial probe to complete and onboarding to appear
await expect(page.getByText('Start your local server')).toBeVisible({ timeout: 10_000 });
await page.screenshot({ path: testInfo.outputPath('onboarding-visible.png') });
});
test('shows step-by-step instructions', async ({ page }) => {
await page.route(`${BACKEND_URL}/**`, (route) => route.abort('connectionrefused'));
await page.goto('/');
// Step 1 is active (done once polling starts)
await expect(page.getByText('Copy the command')).toBeAttached({ timeout: 10_000 });
// Step 2 title changes to "Waiting for server to start" once polling begins
await expect(page.getByText('Waiting for server to start')).toBeAttached({ timeout: 10_000 });
// Step 3 is always rendered
await expect(page.getByText('Auto-connects and opens the graph')).toBeAttached({
timeout: 5_000,
});
});
test('shows terminal window with command', async ({ page }) => {
await page.route(`${BACKEND_URL}/**`, (route) => route.abort('connectionrefused'));
await page.goto('/');
// Should show either dev or prod command in a terminal block
const terminal = page.locator('code');
await expect(terminal.first()).toBeVisible({ timeout: 10_000 });
// The $ prompt should be present
await expect(page.getByText('$')).toBeVisible();
});
test('shows polling indicator', async ({ page }) => {
await page.route(`${BACKEND_URL}/**`, (route) => route.abort('connectionrefused'));
await page.goto('/');
// Polling starts after initial probe fails
await expect(page.getByText('Listening for server')).toBeVisible({ timeout: 10_000 });
});
test('shows Node.js version requirement', async ({ page }) => {
await page.route(`${BACKEND_URL}/**`, (route) => route.abort('connectionrefused'));
await page.goto('/');
await expect(page.getByText(/Node\.js.*\d+/)).toBeVisible({ timeout: 10_000 });
await expect(page.getByText('Port 4747')).toBeVisible();
});
test('copy button has accessible label', async ({ page }) => {
await page.route(`${BACKEND_URL}/**`, (route) => route.abort('connectionrefused'));
await page.goto('/');
await expect(page.getByText('Copy the command')).toBeVisible({ timeout: 10_000 });
const copyBtn = page.getByLabel('Copy to clipboard').first();
await expect(copyBtn).toBeVisible();
});
});
// ── Flow 2: Server detected → success → auto-connect ──────────────────────
test.describe('Flow 2: Server detected — auto-connect', () => {
test('shows success card when server becomes reachable', async ({ page }, testInfo) => {
// Start with server unreachable
let blockBackend = true;
await page.route(`${BACKEND_URL}/**`, (route) => {
if (blockBackend) return route.abort('connectionrefused');
// Let it through to the real handler below
return route.fallback();
});
// Mock the backend responses for when we "start" the server
await page.route(`${BACKEND_URL}/api/repos`, async (route) => {
if (blockBackend) return route.abort('connectionrefused');
await route.fulfill({ json: [{ name: 'test-repo', path: '/tmp/test' }] });
});
await page.route(`${BACKEND_URL}/api/repo`, async (route) => {
if (blockBackend) return route.abort('connectionrefused');
await route.fulfill({
json: { name: 'test-repo', path: '/tmp/test', repoPath: '/tmp/test' },
});
});
await page.route(`${BACKEND_URL}/api/graph**`, async (route) => {
if (blockBackend) return route.abort('connectionrefused');
await route.fulfill({ json: { nodes: [], relationships: [] } });
});
await page.route(`${BACKEND_URL}/api/heartbeat`, async (route) => {
if (blockBackend) return route.abort('connectionrefused');
// SSE response
await route.fulfill({
status: 200,
headers: { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-cache' },
body: ':ok\n\n',
});
});
await page.goto('/');
// Verify onboarding is shown first
await expect(page.getByText('Start your local server')).toBeVisible({ timeout: 10_000 });
await page.screenshot({ path: testInfo.outputPath('before-server-start.png') });
// "Start" the server by unblocking requests
blockBackend = false;
// Wait for success card
await expect(page.getByText('Server Connected')).toBeVisible({ timeout: 15_000 });
await page.screenshot({ path: testInfo.outputPath('success-card.png') });
});
test('transitions to analyze phase when server has zero repos', async ({ page }, testInfo) => {
// Mock server with zero repos — repos endpoint returns empty array
await page.route(`${BACKEND_URL}/api/repos`, (route) => route.fulfill({ json: [] }));
await page.route(`${BACKEND_URL}/api/info`, (route) =>
route.fulfill({ json: { version: '1.0.0', launchContext: 'npx', nodeVersion: 'v22.0.0' } }),
);
await page.route(`${BACKEND_URL}/api/heartbeat`, (route) =>
route.fulfill({
status: 200,
headers: { 'Content-Type': 'text/event-stream' },
body: ':ok\n\n',
}),
);
await page.goto('/');
// Should transition: onboarding → success → analyze (zero repos)
// The analyze form tabs should be visible
await expect(page.getByRole('tab', { name: 'GitHub URL' })).toBeVisible({ timeout: 20_000 });
await expect(page.getByRole('tab', { name: 'Local Folder' })).toBeVisible();
await page.screenshot({ path: testInfo.outputPath('analyze-empty-state.png') });
});
});
// ── Flow 3: Analyze form ───────────────────────────────────────────────────
test.describe('Flow 3: Analyze form', () => {
test.beforeEach(async ({ page }) => {
// Mock server with zero repos to show the analyze form
await page.route(`${BACKEND_URL}/api/repos`, (route) => route.fulfill({ json: [] }));
await page.route(`${BACKEND_URL}/api/info`, (route) =>
route.fulfill({ json: { version: '1.0.0', launchContext: 'npx', nodeVersion: 'v22.0.0' } }),
);
await page.route(`${BACKEND_URL}/api/heartbeat`, (route) =>
route.fulfill({
status: 200,
headers: { 'Content-Type': 'text/event-stream' },
body: ':ok\n\n',
}),
);
});
test('GitHub URL tab validates input', async ({ page }, testInfo) => {
await page.goto('/');
// Wait for analyze form (transition: onboarding → success → analyze)
await expect(page.getByRole('tab', { name: 'GitHub URL' })).toBeVisible({ timeout: 20_000 });
// Type an invalid URL
const input = page.locator('input[type="url"]');
await input.fill('not-a-url');
// Analyze button should be visible but disabled
const analyzeBtn = page.getByRole('button', { name: /Analyze Repository/ });
await expect(analyzeBtn).toBeVisible();
// Type a valid GitHub URL
await input.fill('https://github.com/anthropics/courses');
await page.screenshot({ path: testInfo.outputPath('valid-github-url.png') });
});
test('Local Folder tab shows browse button', async ({ page }, testInfo) => {
await page.goto('/');
await expect(page.getByRole('tab', { name: 'Local Folder' })).toBeVisible({ timeout: 20_000 });
// Switch to Local Folder tab
await page.getByRole('tab', { name: 'Local Folder' }).click();
// Browse button should be visible
await expect(page.getByText('Browse for folder')).toBeVisible();
await page.screenshot({ path: testInfo.outputPath('local-folder-tab.png') });
});
test('switching tabs clears input', async ({ page }) => {
await page.goto('/');
await expect(page.getByRole('tab', { name: 'GitHub URL' })).toBeVisible({ timeout: 20_000 });
// Type in GitHub URL
const urlInput = page.locator('input[type="url"]');
await urlInput.fill('https://github.com/test/repo');
// Switch to Local Folder
await page.getByRole('tab', { name: 'Local Folder' }).click();
// Switch back to GitHub URL — input should be empty
await page.getByRole('tab', { name: 'GitHub URL' }).click();
const newUrlInput = page.locator('input[type="url"]');
await expect(newUrlInput).toHaveValue('');
});
});
// ── Flow 4: Repo dropdown (requires running server) ────────────────────────
test.describe('Flow 4: Repo dropdown in exploring view', () => {
const SKIP_MSG = 'Requires running gitnexus server with indexed repos';
test.beforeAll(async () => {
if (process.env.E2E) return;
try {
const res = await fetch(`${BACKEND_URL}/api/repos`);
if (!res.ok) {
test.skip(true, SKIP_MSG);
return;
}
const repos = await res.json();
if (!repos.length) {
test.skip(true, 'Server has no indexed repos');
return;
}
} catch {
test.skip(true, SKIP_MSG);
}
});
test('project badge opens repo dropdown', async ({ page }, testInfo) => {
await enterExploringView(page);
await page.screenshot({ path: testInfo.outputPath('exploring-loaded.png') });
// Click the project badge (has a chevron)
const badge = page
.locator('header button')
.filter({ has: page.locator('svg') })
.first();
await badge.click();
// Repo dropdown should be visible
await expect(page.getByText('Repositories')).toBeVisible({ timeout: 5_000 });
await expect(page.getByText('Analyze a new repository')).toBeVisible();
await page.screenshot({ path: testInfo.outputPath('repo-dropdown-open.png') });
});
test('analyze option opens inline form', async ({ page }, testInfo) => {
await enterExploringView(page);
// Open repo dropdown
const badge = page
.locator('header button')
.filter({ has: page.locator('svg') })
.first();
await badge.click();
// Click "Analyze a new repository..."
await page.getByText('Analyze a new repository').click();
// Should show the analyze form inline
await expect(page.getByText('GitHub URL')).toBeVisible({ timeout: 5_000 });
await expect(page.getByText('Local Folder')).toBeVisible();
await page.screenshot({ path: testInfo.outputPath('inline-analyze-form.png') });
});
});
+72 -74
View File
@@ -1,9 +1,10 @@
import { test, expect, type TestInfo } from '@playwright/test';
import { test, expect } from '@playwright/test';
/**
* E2E tests for the GitNexus web UI.
* E2E tests for the GitNexus web UI — exploring view features.
*
* Requires:
* - gitnexus serve running on localhost:4747
* - gitnexus serve running on localhost:4747 with at least one indexed repo
* - gitnexus-web dev server running on localhost:5173
*
* Skipped when servers aren't available (CI without services, etc.).
@@ -12,95 +13,103 @@ import { test, expect, type TestInfo } from '@playwright/test';
const BACKEND_URL = process.env.BACKEND_URL ?? 'http://localhost:4747';
const FRONTEND_URL = process.env.FRONTEND_URL ?? 'http://localhost:5173';
// Skip all tests if the gitnexus server or Vite dev server isn't reachable
test.beforeAll(async () => {
if (process.env.E2E) return; // force-run
if (process.env.E2E) return;
try {
const [backendRes, frontendRes] = await Promise.allSettled([
fetch(`${BACKEND_URL}/api/repos`),
fetch(FRONTEND_URL),
]);
if (backendRes.status === 'rejected' || (backendRes.status === 'fulfilled' && !backendRes.value.ok)) {
if (
backendRes.status === 'rejected' ||
(backendRes.status === 'fulfilled' && !backendRes.value.ok)
) {
test.skip(true, 'gitnexus serve not available on :4747');
return;
}
if (frontendRes.status === 'rejected' || (frontendRes.status === 'fulfilled' && !frontendRes.value.ok)) {
if (
frontendRes.status === 'rejected' ||
(frontendRes.status === 'fulfilled' && !frontendRes.value.ok)
) {
test.skip(true, 'Vite dev server not available on :5173');
return;
}
// Check there's at least one indexed repo
if (backendRes.status === 'fulfilled') {
const repos = await backendRes.value.json();
if (!repos.length) {
test.skip(true, 'No indexed repos — run gitnexus analyze first');
return;
}
}
} catch {
test.skip(true, 'servers not available');
}
});
/** Shared helper: connect to the local server and wait for the graph to load */
async function connectAndWaitForGraph(page: import('@playwright/test').Page, testInfo: TestInfo) {
// Signal to the app that we are running under Playwright (used to skip heavy Ladybug loads).
await page.addInitScript(() => {
(window as unknown as { __PLAYWRIGHT_TEST__?: boolean }).__PLAYWRIGHT_TEST__ = true;
});
/**
* Wait for the server-detection flow to complete.
*
* The app auto-detects the server, then either:
* - shows the landing screen when indexed repos exist, or
* - goes straight into analyze onboarding when there are zero repos.
*
* For these tests we require at least one indexed repo, so pick the first
* landing card when present and then wait for the exploring view.
*/
async function waitForGraphLoaded(page: import('@playwright/test').Page) {
await page.goto('/');
// Wait for the app to fully render before interacting
const serverTab = page.getByRole('button', { name: 'Server' });
await expect(serverTab).toBeVisible({ timeout: 15_000 });
await page.screenshot({ path: testInfo.outputPath('step-1-landing.png') });
const landingCards = page.locator('[data-testid="landing-repo-card"]');
const preferredLandingCard = landingCards
.filter({ hasText: /GitNexus|local-integration/ })
.first();
try {
await landingCards.first().waitFor({ state: 'visible', timeout: 15_000 });
const landingCard =
(await preferredLandingCard.count()) > 0 ? preferredLandingCard : landingCards.first();
await landingCard.click();
} catch {
// Landing screen may not appear (e.g. ?server auto-connect)
}
// Click "Server" tab and wait for the input to appear
await serverTab.click();
const serverInput = page.locator('input[name="server-url-input"]');
await expect(serverInput).toBeVisible({ timeout: 15_000 });
await page.screenshot({ path: testInfo.outputPath('step-2-server-tab.png') });
await serverInput.fill(BACKEND_URL);
await page.screenshot({ path: testInfo.outputPath('step-3-url-filled.png') });
await page.getByRole('button', { name: /Connect/ }).click();
// Wait for graph to load — status bar shows "Ready"
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.getByText(/\d+ nodes/).first()).toBeVisible();
await page.screenshot({ path: testInfo.outputPath('step-4-graph-loaded.png') });
const statusBar = page.getByRole('contentinfo');
await expect(statusBar.getByText('Ready', { exact: true })).toBeVisible({ timeout: 45_000 });
await expect(statusBar).toContainText(/nodes/, {
timeout: 20_000,
});
}
test.describe('Server Connection & Graph Loading', () => {
test('connects to server and loads graph', async ({ page }, testInfo) => {
await connectAndWaitForGraph(page, testInfo);
await page.screenshot({ path: testInfo.outputPath('graph-loaded.png'), fullPage: true });
test('selects a repo from landing and loads graph', async ({ page }) => {
await waitForGraphLoaded(page);
});
});
test.describe('Nexus AI', () => {
test('panel opens and agent initializes without error', async ({ page }, testInfo) => {
await connectAndWaitForGraph(page, testInfo);
test('panel opens and agent initializes without error', async ({ page }) => {
await waitForGraphLoaded(page);
// Click Nexus AI button to open the panel
await page.getByRole('button', { name: 'Nexus AI' }).click();
// Should see the Nexus AI tab content
await expect(page.getByText('Ask me anything')).toBeVisible({ timeout: 15_000 });
await page.screenshot({ path: testInfo.outputPath('nexus-ai-panel.png'), fullPage: true });
// "Database not ready" should NOT be visible
const errorBanner = page.getByText('Database not ready');
expect(await errorBanner.isVisible().catch(() => false)).toBe(false);
});
});
test.describe('Processes Panel', () => {
test('shows process list and View button works', async ({ page }, testInfo) => {
await connectAndWaitForGraph(page, testInfo);
test('shows process list and View button works', async ({ page }) => {
await waitForGraphLoaded(page);
// Open Nexus AI panel, switch to Processes tab
await page.getByRole('button', { name: 'Nexus AI' }).click();
await page.getByText('Processes').click();
// Should show process count — wait for data-testid instead of fixed timeout
await expect(page.locator('[data-testid="process-list-loaded"]')).toBeVisible({ timeout: 15_000 });
await page.screenshot({ path: testInfo.outputPath('processes-panel.png'), fullPage: true });
await expect(page.locator('[data-testid="process-list-loaded"]')).toBeVisible({
timeout: 15_000,
});
// Hover first process item to reveal View button, then click it
const processRow = page.locator('[data-testid="process-row"]').first();
await expect(processRow).toBeVisible({ timeout: 10_000 });
await processRow.hover();
@@ -108,21 +117,18 @@ test.describe('Processes Panel', () => {
const viewBtn = processRow.locator('[data-testid="process-view-button"]');
await viewBtn.waitFor({ state: 'visible', timeout: 5_000 });
await viewBtn.click();
// Wait for modal to appear
await expect(page.locator('[data-testid="process-modal"]')).toBeVisible({ timeout: 5_000 });
await page.screenshot({ path: testInfo.outputPath('process-view-clicked.png'), fullPage: true });
});
test('lightbulb highlights nodes in graph', async ({ page }, testInfo) => {
await connectAndWaitForGraph(page, testInfo);
test('lightbulb highlights nodes in graph', async ({ page }) => {
await waitForGraphLoaded(page);
await page.getByRole('button', { name: 'Nexus AI' }).click();
await page.getByText('Processes').click();
await expect(page.locator('[data-testid="process-list-loaded"]')).toBeVisible({ timeout: 15_000 });
await expect(page.locator('[data-testid="process-list-loaded"]')).toBeVisible({
timeout: 15_000,
});
await page.screenshot({ path: testInfo.outputPath('before-highlight.png'), fullPage: true });
// Hover first process to reveal lightbulb
const processRow = page.locator('[data-testid="process-row"]').first();
await expect(processRow).toBeVisible({ timeout: 10_000 });
await processRow.hover();
@@ -130,36 +136,28 @@ test.describe('Processes Panel', () => {
const lightbulb = processRow.locator('[data-testid="process-highlight-button"]');
await lightbulb.waitFor({ state: 'visible', timeout: 5_000 });
await lightbulb.click();
// Wait for highlight to apply — the process row gets amber styling when focused
await expect(processRow).toHaveClass(/bg-amber-950/, { timeout: 5_000 });
await page.screenshot({ path: testInfo.outputPath('after-highlight.png'), fullPage: true });
});
});
test.describe('Turn Off All Highlights', () => {
test('selecting a node dims others, button clears it', async ({ page }, testInfo) => {
await connectAndWaitForGraph(page, testInfo);
test('selecting a node dims others, button clears it', async ({ page }) => {
await waitForGraphLoaded(page);
// Wait for graph to fully render by checking for canvas element
await expect(page.locator('canvas').first()).toBeVisible({ timeout: 10_000 });
await page.screenshot({ path: testInfo.outputPath('before-select.png'), fullPage: true });
// Click a file in the file tree to select a node
const fileItem = page.getByText('package.json').first();
await expect(fileItem).toBeVisible({ timeout: 10_000 });
await fileItem.click();
// Wait for highlight toggle to show "Turn off" (indicates highlights are active)
const highlightToggle = page.locator('[data-testid="ai-highlights-toggle"]');
await expect(highlightToggle).toHaveAttribute('title', 'Turn off all highlights', { timeout: 5_000 });
await page.screenshot({ path: testInfo.outputPath('node-selected.png'), fullPage: true });
await expect(highlightToggle).toHaveAttribute('title', 'Turn off all highlights', {
timeout: 5_000,
});
// Click the toggle to clear all highlights
await highlightToggle.click();
// Verify highlights are now off — button title changes to "Turn on"
await expect(highlightToggle).toHaveAttribute('title', 'Turn on AI highlights', { timeout: 5_000 });
await page.screenshot({ path: testInfo.outputPath('highlights-cleared.png'), fullPage: true });
await expect(highlightToggle).toHaveAttribute('title', 'Turn on AI highlights', {
timeout: 5_000,
});
});
});
+6 -3
View File
@@ -4,9 +4,12 @@
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<title>GitNexus</title>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=JetBrains+Mono:wght@400;500;600&family=Outfit:wght@300;400;500;600;700&display=swap" rel="stylesheet">
<link rel="preconnect" href="https://fonts.googleapis.com" />
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin />
<link
href="https://fonts.googleapis.com/css2?family=JetBrains+Mono:wght@400;500;600&family=Outfit:wght@300;400;500;600;700&display=swap"
rel="stylesheet"
/>
</head>
<body>
<div id="root"></div>
+54 -1835
View File
File diff suppressed because it is too large Load Diff
+1 -13
View File
@@ -18,9 +18,7 @@
"test:e2e:report": "playwright show-report"
},
"dependencies": {
"@huggingface/transformers": "^3.0.0",
"@isomorphic-git/lightning-fs": "^4.6.2",
"@ladybugdb/wasm-core": "^0.15.1",
"gitnexus-shared": "file:../gitnexus-shared",
"@langchain/anthropic": "^1.3.10",
"@langchain/core": "^1.1.15",
"@langchain/google-genai": "^2.1.10",
@@ -30,8 +28,6 @@
"@sigma/edge-curve": "^3.1.0",
"@tailwindcss/vite": "^4.1.18",
"axios": "^1.13.2",
"buffer": "^6.0.3",
"comlink": "^4.4.2",
"d3": "^7.9.0",
"dompurify": "^3.3.3",
"graphology": "^0.26.0",
@@ -40,13 +36,10 @@
"graphology-layout-forceatlas2": "^0.10.1",
"graphology-layout-noverlap": "^0.4.2",
"graphology-utils": "^2.3.0",
"isomorphic-git": "^1.36.1",
"jszip": "^3.10.1",
"langchain": "^1.2.10",
"lru-cache": "^11.2.4",
"lucide-react": "^0.562.0",
"mermaid": "^11.12.2",
"minisearch": "^7.2.0",
"mnemonist": "^0.39.0",
"pandemonium": "^2.4.0",
"react": "^18.3.1",
@@ -58,9 +51,6 @@
"sigma": "^3.0.2",
"tailwindcss": "^4.1.18",
"uuid": "^13.0.0",
"vite-plugin-top-level-await": "^1.6.0",
"vite-plugin-wasm": "^3.5.0",
"web-tree-sitter": "^0.20.8",
"zod": "^3.25.76"
},
"devDependencies": {
@@ -70,7 +60,6 @@
"@testing-library/react": "^16.3.2",
"@testing-library/user-event": "^14.6.1",
"@types/dompurify": "^3.0.5",
"@types/jszip": "^3.4.0",
"@types/node": "^24.10.1",
"@types/react": "^18.3.5",
"@types/react-dom": "^18.3.0",
@@ -82,7 +71,6 @@
"tree-sitter-wasms": "^0.1.13",
"typescript": "^5.4.5",
"vite": "^5.2.0",
"vite-plugin-static-copy": "^3.1.4",
"vitest": "^3.2.4",
"wait-on": "^8.0.5"
}
+1 -4
View File
@@ -40,9 +40,6 @@ export default defineConfig({
use: { browserName: 'chromium' },
},
],
reporter: [
['list'],
['html', { open: 'never', outputFolder: 'playwright-report' }],
],
reporter: [['list'], ['html', { open: 'never', outputFolder: 'playwright-report' }]],
outputDir: 'test-results',
});
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.

Some files were not shown because too many files have changed in this diff Show More