Compare commits

...
Author SHA1 Message Date
copilot-swe-agent[bot]andmagyargergo 1d7782e4e1 fix(cobol): single-quote CALL/COPY, sequence number stripping, PERFORM keyword filtering
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/afed41bd-98b4-48e8-a9e5-ebf0b97deaa8
2026-03-24 17:10:16 +00:00
copilot-swe-agent[bot]andmagyargergo 7ce1371ad8 Initial plan: review COBOL processor completeness
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/afed41bd-98b4-48e8-a9e5-ebf0b97deaa8
2026-03-24 16:59:25 +00:00
copilot-swe-agent[bot] ae437dc30e Initial plan 2026-03-24 16:48:12 +00:00
Gergo Magyar 832789f288 test(cobol): exhaustive 57-test suite with strict exact assertions
Complete rewrite of COBOL integration tests using ground-truth approach:
dump the full graph, then assert EVERY node and EVERY edge.

57 tests across 9 sections:
- Node completeness: Module(3), Function(13), Namespace(2), Property(21),
  Record(1), CodeElement(8), Constructor(1) — exact sorted arrays
- Edge completeness: 22 tests covering every type+reason combination
  with exact source→target pairs
- Cross-program resolution: 6 tests verifying CALL, CICS LINK/XCTL, JCL
- COPY expansion: copybook data items in RPTGEN
- Section hierarchy: exact paragraph membership per section
- Data item ownership: exact per-module breakdown
- MOVE data flow: exact read/write pairs
- JCL integration: job/step/dataset containment
- Grand totals: CALLS(22), CONTAINS(48), IMPORTS(1), ACCESSES(7)

Fixture enhancements:
- CUSTUPDT.cbl: added INIT-SECTION + PROCESSING-SECTION, PERFORM THRU
- AUDITLOG.cbl: added ENTRY "AUDITLOG-BATCH"
- RPTGEN.cbl: added EXEC CICS XCTL

Zero fuzzy assertions — every expect uses toBe(N) or toEqual([...sorted]).
2026-03-24 16:10:29 +00:00
Gergo Magyar 41b0d8dfad test(cobol): add 26 integration tests with exact assertions + fix CICS resolution bug
Integration tests (test/integration/resolvers/cobol.test.ts):
- 26 tests covering full COBOL system extraction
- ALL assertions use exact toBe(N) — zero fuzzy assertions
- Fixtures: CUSTUPDT.cbl, AUDITLOG.cbl, CUSTDAT.cpy, RPTGEN.cbl, RUNJOBS.jcl

Bug fix (cobol-processor.ts):
- CICS LINK/XCTL cross-program resolution was broken — edges were
  created with "resolved" reason but pointing to <unresolved> targets
- Fix: use cics-link-unresolved / cics-xctl-unresolved suffix pattern
  matching the existing cobol-call-unresolved pattern
- Second-pass resolver now patches both CALL and CICS unresolved edges

All 3915 tests pass, 0 failures.
2026-03-24 15:41:00 +00:00
Gergo Magyar 9760f966cb feat(cobol): enrich graph with EXEC SQL/CICS, ENTRY points, MOVE data flow, PERFORM THRU
Maps the remaining 60% of CobolRegexResults to the graph:
- EXEC SQL blocks → CodeElement nodes + ACCESSES edges to DB tables
- EXEC CICS LINK/XCTL → CodeElement nodes + cross-program CALLS edges
- ENTRY points → Constructor nodes (registered for cross-program resolution)
- MOVE statements → ACCESSES edges (read/write data flow tracking)
- PERFORM THRU → expanded CALLS edges for range targets
- File declarations → Record nodes with assignment metadata
- Cross-program CALL 2nd pass: resolves unresolved targets after all programs processed
2026-03-24 15:20:26 +00:00
Gergo Magyar 88c89c42e6 docs: document custom processor pattern in pipeline.ts
Add comment block at the custom processor integration point
documenting the pattern for future non-tree-sitter language additions.
2026-03-24 14:59:09 +00:00
Gergo Magyar 4af677e637 feat: add COBOL language support with regex extraction pipeline
Standalone COBOL processor following the markdown-processor.ts pattern:
- No LanguageProvider modification — COBOL uses regex, not tree-sitter
- No SupportedLanguages enum change — standalone processor pattern

New files:
- cobol-processor.ts — orchestrator (processCobol, isCobolFile, isJclFile)
- cobol/cobol-preprocessor.ts — regex state machine extraction (~888 LOC)
- cobol/cobol-copy-expander.ts — COPY statement expansion with circular detection
- cobol/jcl-parser.ts — JCL job/step/DD extraction
- cobol/jcl-processor.ts — JCL graph node creation

Extraction produces:
- Module nodes (PROGRAM-ID)
- Function nodes (paragraphs)
- Namespace nodes (sections)
- Property nodes (data items)
- CALLS edges (PERFORM intra-file, CALL cross-program)
- IMPORTS edges (COPY statements)
- CONTAINS edges (section → paragraph hierarchy)

Pipeline integration: single processCobol() call in Phase 2.6

54 new tests (33 COBOL + 21 JCL), all 3889 tests pass.
2026-03-24 14:39:25 +00:00
Gergő Magyar 7999b6ba7b refactor: SICP-informed LanguageProvider architecture (#488)
* refactor: SICP-informed LanguageProvider architecture for ingestion pipeline

Consolidate 16 scattered dispatch surfaces into a single LanguageProvider
Strategy interface per language. Processors are now fully language-agnostic —
zero SupportedLanguages.X enum access, zero dispatch table imports.

Architecture (5-layer DAG, zero circular dependencies):
  L0: Capability modules (dispatch tables, single source of truth)
  L1: LanguageProvider interface + createLanguageProvider factory
  L2: 13 per-language provider files (Strategy objects)
  L3: Registry with satisfies Record<SL, LP> + pre-built lookup maps
  L4: Processors (language-agnostic, all behavior via provider.*)

Key changes:
- Add LanguageProvider interface with 15 properties (6 required, 9 optional)
- Create 13 provider files in languages/ + php-helpers.ts
- Migrate all processors to getProvider(language) — cached once per scope
- Replace heritage if-checks with provider.interfaceNamePattern/heritageDefaultEdge
- Replace MRO switch(language) with switch(provider.mroStrategy)
- Replace isNodeExported with provider.exportChecker
- Move PHP description extraction behind provider.descriptionExtractor
- Move Swift implicit imports behind provider.implicitImportWirer
- Move PHP route detection behind provider.isRouteFile
- Move Kotlin wildcard append behind provider.importPathPreprocessor
- Remove deprecated TypeEnvironment.env, add fileScope()/allScopes()
- De-export TypeEnv type (module-private)
- Pre-build extensionMap, WILDCARD_LANGUAGES, SYNTHESIS_LANGUAGES at load
- Remove dead entryPointPatterns/frameworkPatterns from interface
- Derive createLanguageProvider config type via Pick/Partial/Omit
- Tighten callback types from any to SyntaxNode
- Migrate 270+ test call sites from .env to TypeEnvironment API

Adding a new language: 3 files (enum + provider + registry line).
No processor file touched. Ever.

* refactor: clean architecture for LanguageProvider with O(1) AST cache

Address all PR #488 review comments and achieve pristine SICP layer separation:

Interface redesign:
- Split LanguageProvider into Config (input) + Provider (runtime with defaults)
- Rename createLanguageProvider → defineLanguage with explicit DEFAULTS constant
- Add MroStrategy, ImportSemantics named type aliases for better IDE tooltips
- Tighten labelOverride signature: string|null → NodeLabel|null (compile-time safety)
- Tighten descriptionExtractor nodeLabel: string → NodeLabel
- Un-export LanguageProviderConfig (internal to defineLanguage)

CI fixes (all 4 failures resolved):
- isNodeExported: add null guard for unknown languages
- preprocessImportPath tests: pass getProvider() instead of raw enum
- MRO tests: update expected strings to match language-agnostic prefixes

Code deduplication:
- Extract findDescendant/extractStringContent to ast-helpers.ts (single source of truth)
- Unify Kotlin method detection: remove duplicate from extractFunctionName,
  use provider.labelOverride as single source of truth via findEnclosingFunctionId
- extractFunctionName return type: string → NodeLabel

Performance (O(1) AST node access):
- Add per-file Map-based memoization in parse-worker for parent-chain walks
- Cache enclosingClassId, enclosingFunctionId, exportStatus per SyntaxNode
- Clear caches before each file parse (not after — handles parse failures)

Architecture (pristine languages/ folder):
- Move php-helpers.ts → helpers/php.ts (L0 capability, not L2 config)
- Create helpers/swift.ts from extracted Swift provider logic
- Extract cppLabelOverride AST walk → isCppInsideClassOrStruct in ast-helpers.ts
- Extract isPhpRouteFile → helpers/php.ts
- All 13 provider files are now pure configuration — zero implementation logic
- Ruby: remove no-op namedBindingExtractor assignment (undefined from dispatch table)

* refactor: eliminate LANGUAGE_QUERIES, typeConfigs, namedBindingExtractors dispatch tables

Phase 1 of L0 dispatch table elimination. Providers now import capabilities
directly instead of indexing into redundant Record<SL, T> dispatch tables:

- LANGUAGE_QUERIES: providers import named query constants directly
  (TYPESCRIPT_QUERIES, PYTHON_QUERIES, etc.). Table kept in tree-sitter-queries.ts
  for call-processor.ts dynamic lookup + test consumers.

- typeConfigs: providers import from individual type-extractor files
  (typescriptConfig from typescript.ts, javaTypeConfig from jvm.ts, etc.).
  Dispatch table fully removed from type-extractors/index.ts.

- namedBindingExtractors: providers import extractors directly from
  named-binding-extraction.ts (extractTsNamedBindings, etc.).
  Dispatch table fully removed from import-resolution.ts.

Net: -48 LOC of dispatch table indirection. L3 satisfies Record<SL, LP>
remains the single exhaustiveness check.

* refactor: eliminate exportCheckers, callRouters, importResolvers dispatch tables

Phase 2 of L0 dispatch table elimination. All 6 dispatch tables are now gone:

- exportCheckers: individual checkers exported directly (tsExportChecker,
  pythonExportChecker, etc.). isNodeExported uses a local checkersByLanguage
  map to avoid circular dependency with languages/index.ts.

- callRouters: table removed. Providers import noRouting or routeRubyCall
  directly. noRouting now exported. Dead import removed from call-processor.ts.

- importResolvers: resolver functions exported with clean names
  (resolveTypescriptImport, resolveJavaImport, etc.). Inline lambdas
  extracted to named exports. Dispatch functions renamed from *Dispatch
  suffix to clean resolve*Import pattern.

Combined with Phase 1, all 6 L0 dispatch tables have been eliminated.
L3 satisfies Record<SL, LanguageProvider> is the single exhaustiveness check.
Providers are now fully self-contained — each imports its capabilities directly.

* perf+refactor: type-env caching, sequential fallback caching, utils.ts split

Phase 3 — performance optimizations and barrel cleanup:

Type-env parent-walk caching:
- Memoize findEnclosingClassName and findEnclosingParentClassName with
  per-file Map<SyntaxNode, string|undefined> caches
- Eliminates O(n*m) repeated child scanning in extractParentClassFromNode
- Caches cleared in buildTypeEnv before each file's walk phase

Sequential fallback caching:
- Add classIdCache + exportCache Maps to parsing-processor.ts
- Mirrors the O(1) memoization pattern from parse-worker.ts
- Both paths now have identical caching for parent-chain walks

Split utils.ts barrel into focused modules:
- noise-filter.ts: BUILT_IN_NAMES + isBuiltInOrNoise (167 LOC)
- language-detection.ts: getLanguageFromFilename (58 LOC)
- utils.ts slimmed to re-exports + yieldToEventLoop + isVerboseIngestionEnabled
- Backward compatible — existing imports from utils.ts still work

* refactor: rename resolvers/ → import-resolvers/, restructure tests per-concern

Directory renames (git mv — history preserved):
- src/core/ingestion/resolvers/ → import-resolvers/ (10 files)
- test/unit/call-routing.test.ts → call-routing/ruby.test.ts
- test/unit/named-binding-extraction.test.ts → named-bindings/csharp.test.ts
- test/unit/import-resolution.test.ts → import-resolution/preprocessing.test.ts

All 11 import paths updated to reference new import-resolvers/ location.
Test imports updated for new subdirectory depth.

Note: test/integration/resolvers/ NOT renamed — those tests cover the full
ingestion pipeline per-language, not just import resolution.

* refactor: eliminate utils.ts barrel — all 33 consumers now import directly

Migrated 65 import sites across 33 files to import from the focused source
module instead of the utils.ts barrel:

- ast-helpers.js: SyntaxNode, extractFunctionName, findEnclosingClassId, etc.
- call-analysis.js: inferCallForm, extractReceiverName, countCallArguments, etc.
- noise-filter.js: BUILT_IN_NAMES, isBuiltInOrNoise
- language-detection.js: getLanguageFromFilename

utils.ts reduced to 2 original functions only:
- yieldToEventLoop
- isVerboseIngestionEnabled

Zero re-exports remain. Every import is now direct to its source module.

* refactor: create utils/ folder, move all shared utilities, delete utils.ts barrel

Final phase of module structure migration:

- git mv ast-helpers.ts, call-analysis.ts, noise-filter.ts,
  language-detection.ts → utils/ subdirectory (history preserved)
- Extract yieldToEventLoop → utils/event-loop.ts
- Extract isVerboseIngestionEnabled → utils/verbose.ts
- Delete utils.ts (zero re-exports, zero functions remain)
- Update 38 import paths across source and test files

The ingestion/ root is now clean — only processors, capability modules,
and the pipeline orchestrator live at the top level. All shared utilities
are in utils/, all language-specific helpers in helpers/, all import
resolvers in import-resolvers/.

* refactor: move findChild from import-resolvers/utils.ts to utils/ast-helpers.ts

findChild is a generic AST helper (find first named child by type) — it
belongs with the other AST traversal utilities, not in the import resolver
module. 4 consumers updated to import from utils/ast-helpers.js.

* refactor: split named-binding-extraction.ts into per-language files

Rename named-binding-extraction.ts → named-binding-processor.ts (git mv,
history preserved), keeping only walkBindingChain for re-export chain resolution.

7 per-language extractor functions moved to named-bindings/ subdirectory:
- named-bindings/typescript.ts (extractTsNamedBindings — TS + JS)
- named-bindings/python.ts (extractPythonNamedBindings)
- named-bindings/kotlin.ts (extractKotlinNamedBindings)
- named-bindings/rust.ts (extractRustNamedBindings + collectRustBindings)
- named-bindings/php.ts (extractPhpNamedBindings)
- named-bindings/csharp.ts (extractCsharpNamedBindings)
- named-bindings/java.ts (extractJavaNamedBindings)

Each provider now imports its binding extractor from the per-language file.

* refactor: eliminate import-resolution.ts — distribute to natural homes

Split per-language resolvers into import-resolvers/ per-language files and
eliminate the import-resolution.ts catch-all module entirely:

Per-language resolvers moved to import-resolvers/:
- standard.ts: resolveStandard, resolveJavascriptImport, resolveTypescriptImport,
  resolveCImport, resolveCppImport
- jvm.ts: resolveJavaImport, resolveKotlinImport
- go.ts: resolveGoImport
- csharp.ts: resolveCSharpImport (helper renamed to Internal)
- php.ts, python.ts, ruby.ts, rust.ts: same pattern
- swift.ts: new file for resolveSwiftImport

Types distributed to their concern directories:
- import-resolvers/types.ts: ImportResult, ImportConfigs, ResolveCtx, ImportResolverFn
- named-bindings/types.ts: NamedBinding, NamedBindingExtractorFn

preprocessImportPath moved to import-processor.ts (its primary consumer).

import-resolution.ts deleted — zero catch-all modules remain.

* refactor: tighten SPR — eliminate re-exports, dead code, type holes, and redundant patterns

12 review findings resolved across the ingestion layer:

Type safety:
- CallRouter callNode: any → SyntaxNode (closes type hole)
- CaptureMap type alias replaces Record<string, any>
- providersWithImplicitWiring filter now type-narrowed (removes ! assertions)
- Ruby exportChecker: unnecessary as-cast removed, named export created

Architecture:
- Circular type dependency eliminated (ImportResolutionContext moved to types.ts)
- LANGUAGE_QUERIES residual dispatch replaced with provider.treeSitterQueries
- noRouting sentinel deleted — callRouter now properly optional on 12 providers
- All 6 re-exports from import-processor/pipeline/languages eliminated

Pattern cleanup:
- Dead checkersByLanguage table + isNodeExported removed from export-detection
- 4 duplicated config interfaces consolidated to language-config.ts
- extractCsharpNamedBindings → extractCSharpNamedBindings (casing consistency)

Simplification:
- import-resolvers/index.ts barrel deleted (dead re-exports)
- helpers/ inlined into languages/ (php.ts, swift.ts) — 1 directory removed

Verified: tsc --noEmit clean, 3837 tests pass, 0 failures.

* refactor: address review — remove LANGUAGE_QUERIES table, type-extractors barrel, fix Windows timeout

Review comment fixes (github.com/abhigyanpatwari/GitNexus/pull/488#issuecomment-4117817648):

1. LANGUAGE_QUERIES dispatch table removed from tree-sitter-queries.ts
   — 5 test files migrated to getProvider(lang).treeSitterQueries
   — eliminates last parallel dispatch surface

2. type-extractors/index.ts barrel deleted
   — type-env.ts now imports TYPED_PARAMETER_TYPES from shared.js directly

3. Windows CI timeout fix: afterAll cleanup hook in test-indexed-db.ts
   now passes explicit 120s timeout to prevent KuzuDB C++ destructor
   hang from hitting vitest's default 30s testTimeout on Windows

Verified: tsc --noEmit clean, 3835 tests pass, 0 failures.

* refactor: eliminate chained getProvider property access — assign to variable first

All getProvider(lang).property calls now follow the pattern:
  const provider = getProvider(language);
  const x = provider.property;

5 source files + 4 test files updated (~35 occurrences).
This ensures consistent provider variable usage and avoids
repeated lookups in hot paths.

* refactor: remove last 4 re-exports from import-resolvers, fix stale CaptureMap comment

- Remove `export type { TsconfigPaths }` from standard.ts
- Remove `export type { GoModuleConfig }` from go.ts
- Remove `export type { ComposerConfig }` from php.ts
- Remove `export type { CSharpProjectConfig }` from csharp.ts
  All 4 types are canonically defined in language-config.ts;
  zero consumers imported via the resolver re-exports.

- Fix stale CaptureMap JSDoc: said "Uses any" but type is SyntaxNode | undefined
2026-03-24 13:42:39 +00:00
John R. Eakin f0540b33fb ci: E2E workflow, web typecheck job, pre-commit hook, test suite (#486) 2026-03-24 06:01:08 +00:00
harryphung fec47cbc32 feat(web): add GLM (Z.AI) as LLM provider (#468)
Add GLM support using OpenAI-compatible API via ChatOpenAI from LangChain.
Defaults to the Z.AI coding endpoint (https://api.z.ai/api/coding/paas/v4)
with configurable base URL. Supported models: GLM-5, GLM-5-Turbo, GLM-4.7, GLM-4.5.
2026-03-24 06:10:37 +05:30
B. Lawrence Sharma a9d43d5680 fix(ui): use actual repository name instead of top-level folder for GitHub clones (#491) 2026-03-24 05:54:38 +05:30
Zander Raycraft e6aea1c82b Merge pull request #490 from zander-raycraft/main
Updating claude mcp init for window and adding OS detection when installing
2026-03-23 17:47:03 -05:00
Zander Raycraft 3961ad01cf updating windows support 2026-03-23 17:33:32 -05:00
marxo126 c437acf6bb feat: deep flow detection — consumer access tracking, middleware chains, error shapes, api_impact tool (#482) 2026-03-23 22:13:48 +00:00
Gergő Magyar 2ede01dbff Merge pull request #477 from jreakin/perf/web-o1-lookups-memo-bundle
perf(web): O(1) lookups, memoization, bundle optimizations, React fixes
2026-03-23 17:01:09 +00:00
jreakinandClaude Opus 4.6 fd61cd2990 fix(web): allow foreignObject in DOMPurify SVG, remove leftover loop code
- Add ADD_TAGS: ['foreignObject'] to all DOMPurify.sanitize calls —
  Mermaid uses foreignObject for HTML text labels inside flowchart
  nodes. The SVG profile was stripping them, causing empty boxes.
- Remove leftover sub-batch loop lines from prepared statement hoist

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 10:16:34 -05:00
jreakinandClaude Opus 4.6 83ec256cec fix(web): address React rendering review — stale closure, ref dep, O(1) lookups
- MarkdownRenderer: wrap handleLinkClick in useCallback, add to
  markdownComponents useMemo deps (fixes stale closure)
- GraphCanvas: remove sigmaRef from useEffect deps (ref identity
  never changes), extract handleToggleAIHighlights to useCallback
- CodeReferencesPanel: add nodeById Map for O(1) focus-in-graph
  lookup (was O(N) graph.nodes.find on every click)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 10:16:34 -05:00
jreakinandClaude Opus 4.6 29795bb86e fix(web): switch lucide-icons from deep ESM paths to standard re-exports
Deep imports from lucide-react/dist/esm/icons/*.js are internal paths
that broke the Vercel production build. Replaced with standard named
re-exports from lucide-react — keeps the centralized module pattern
without relying on fragile internal paths.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 10:16:34 -05:00
jreakinandClaude Opus 4.6 c8cd733815 perf(web): prepare statement once per label pair, not per sub-batch
Hoists conn.prepare(cypher) outside the sub-batch loop so it's called
once per (fromLabel, toLabel) pair instead of ceil(N/4) times. The
statement is reused for all rows in the group, then closed in finally.

Yields to event loop every 500 relations instead of every sub-batch.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 10:16:33 -05:00
jreakinandClaude Opus 4.6 5616f71fd2 perf(web): switch remaining components to deep lucide imports, extract ProviderConfigCard
- BackendRepoSelector, EmbeddingStatus, Header, MermaidDiagram,
  QueryFAB, RightPanel, ToolCallCard, WebGPUFallbackDialog: switch
  from lucide-react barrel imports to @/lib/lucide-icons deep imports
- MermaidDiagram: lazy-load ProcessFlowModal via React.lazy
- Extract ProviderConfigCard from SettingsPanel for cleaner separation

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 10:16:33 -05:00
jreakinandClaude Opus 4.6 2bd04fe2c4 test: add positive/negative tests for O(1) lookup optimizations
8 tests verifying the data structures underlying performance changes:

nodeById Map (O(1) lookup):
  + Map.get returns correct node by ID
  + duplicate IDs: last wins
  - non-existent ID returns undefined
  - empty Map returns undefined

Set.has (O(1) highlight matching):
  + present IDs return true
  - absent IDs return false
  + handles IDs with colons, dots, slashes
  - case-sensitive matching

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 10:16:33 -05:00
jreakinandClaude Opus 4.6 4199c5e4f3 perf(web): O(1) lookups, memoization, bundle optimizations, React fixes
Performance:
- nodeById Map in GraphCanvas for O(1) click/hover lookups
- fileNodeByPath Map in useAppState for O(1) file path lookups
- Set.has() for HIGHLIGHT_NODES/IMPACT matching (was O(N²))
- useMemo for primaryLanguage in StatusBar
- useCallback on toggleLabelVisibility/toggleEdgeVisibility
- Recursive FileTreePanel search (full subtree, not 1 level)

React fixes:
- Remove stale queryResult dep from clearAICodeReferences
- Cancel RAF chains in CodeReferencesPanel on cleanup
- Clean up timeouts in SettingsPanel and MarkdownRenderer on unmount
- try/catch on localStorage in DropZone (private browsing)
- Await handleServerConnect in App.tsx auto-connect
- pendingToolCalls counter replaces allToolsDone boolean in agent.ts
- JSON.parse try/catch in agent streaming
- sessionStorage JSDoc fix in settings-service

Bundle:
- Centralized lucide icon deep imports

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 10:16:33 -05:00
jreakinandClaude Opus 4.6 0d9a2ca6b6 fix(web): allow foreignObject in DOMPurify SVG sanitization
Mermaid uses foreignObject for HTML text labels inside flowchart
nodes. The SVG profile strips them by default, causing empty boxes.
ADD_TAGS: ['foreignObject'] preserves text while still sanitizing
against XSS.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 10:16:33 -05:00
Gergő Magyar 88fa1a4e09 Merge pull request #476 from jreakin/fix/web-csv-schema-correctness
fix(web): 18 multi-language CSV tables, RFC 4180, schema sync
2026-03-23 14:12:28 +00:00
jreakinandClaude Opus 4.6 8956da61bb fix(web): add 15 missing NodeLabel colors/sizes, escape table names, schema sync
- Add NODE_COLORS and NODE_SIZES entries for all 15 multi-language labels
  (Struct, Trait, Impl, TypeAlias, Const, Static, Namespace, Union,
  Typedef, Macro, Property, Record, Delegate, Annotation, Constructor,
  Template) — fixes Record<NodeLabel, ...> completeness
- Escape table names in count/lookup queries with escapeTableName() to
  prevent silent failures for backtick-required tables
- Add HAS_PROPERTY and ACCESSES to RelationshipType union
- Update initPromise after db recreation in loadGraphToLbug so subsequent
  initLbug() calls return fresh db/conn refs

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 09:03:40 -05:00
Abhigyan PatwariandClaude Opus 4.6 9ba2b22ac6 fix(web): resolve Vercel build errors from security hardening PR (#483)
- ProcessFlowModal: revert import from non-existent @/lib/lucide-icons back to lucide-react
- embedding-pipeline: remove extra argument in executeQuery call (signature only accepts 1 arg)

Both issues were introduced in PR #475 (web-security-hardening).

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 19:31:08 +05:30
Gergő Magyar db6b302fee Merge pull request #408 from marxo126/fix/swift-query-and-patch-script
feat: complete Swift support — query fix, export detection, implicit imports, constructor resolution
2026-03-23 13:53:25 +00:00
marxo126andClaude Opus 4.6 fcd2c3ff38 docs: add macro declarations to Swift ingestion gaps
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 14:21:32 +01:00
jreakinandClaude Opus 4.6 f4bb78c03a test: add negative tests for CSV generation
- empty graph produces header-only CSVs (no data rows)
- empty graph relCSV has only header
- double quotes in node names are RFC 4180 escaped
- file node without fileContents gets empty content (no crash)
- community with empty keywords array produces valid CSV
- unknown node labels are silently skipped

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 08:21:31 -05:00
jreakinandClaude Opus 4.6 47f9b6c48d fix: add multi-language NodeLabel types, close old db/conn, CSV tests
- Add 15 multi-language labels to NodeLabel union (Struct, Trait, Impl,
  TypeAlias, Const, Static, Namespace, Union, Typedef, Macro, Property,
  Record, Delegate, Annotation, Constructor, Template) — eliminates
  unsafe casts in csv-generator
- Close previous conn/db before recreating in loadGraphToLbug to prevent
  WASM resource leaks across repo switches
- Add 7 CSV generation tests: multi-language tables, column count,
  keyword comma escaping, file content, relation CSV, all NODE_TABLES

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 08:21:31 -05:00
jreakinandClaude Opus 4.6 433f403fe2 fix(web): 18 multi-language CSV tables, RFC 4180 escaping, schema sync
- Generate CSVs for Struct, Enum, Macro, Trait, Impl, TypeAlias, Const,
  Static, Property, Record, Delegate, Annotation, Constructor, Template,
  Module, Namespace, Union, Typedef (were defined in schema but silently
  dropped during CSV generation — non-JS/TS repos lost all nodes)
- Escape community keywords array in CSV to prevent comma breakage
- Add Section to NodeLabel type, NODE_COLORS, NODE_SIZES
- Add 7 missing REL_TYPES: HAS_METHOD, HAS_PROPERTY, OVERRIDES, ACCESSES,
  INHERITS, USES, DECORATES
- RFC 4180 doubled-quote regex for relation CSV parsing

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 08:21:31 -05:00
Gergő Magyar 519dc3eccc Merge pull request #475 from jreakin/fix/web-security-hardening
fix(web): Cypher injection guards, DOMPurify SVG, readOnly executeQuery
2026-03-23 13:21:03 +00:00
marxo126andClaude Opus 4.6 fcf8fb9bdf docs: add Swift ingestion gaps tracker and update feature matrix
- Create swift-ingestion-gaps.md with prioritized gap tracker (High/Medium/Low)
- Update type-resolution-system.md feature matrix: 5 Swift entries corrected
  (for-loop→Yes, pattern binding→Partial, call-result/field/method→Yes)
- Add footnotes explaining Swift-specific semantics
- Document resolved items with commit references

Addresses @magyargergo's request to document missing Swift features.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 14:16:43 +01:00
marxo126 47fdad14ed Merge remote-tracking branch 'upstream/main' into fix/swift-query-and-patch-script
# Conflicts:
#	gitnexus/src/core/ingestion/call-processor.ts
2026-03-23 12:21:53 +01:00
Gergo Magyar ffabe857a3 fix(docs): update symbol and relationship counts in AGENTS.md and CLAUDE.md 2026-03-23 11:15:56 +00:00
marxo126andClaude Opus 4.6 1f4c4e77ab refactor: simplify Swift support code after review
- Move `pattern` node handling into shared extractVarName (like mut_pattern)
  instead of inline fallback in type-env — benefits all callers
- Remove non-null assertion (!) on firstNamedChild — defensive null check
- Avoid 100K wrapper object allocation: addSwiftImplicitImports now accepts
  string[] directly, eliminating allFileList.map(p => ({ path: p }))
- Remove duplicate comment block in for-loop test

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 11:44:11 +01:00
marxo126andClaude Opus 4.6 956dfd0bb4 feat: add Swift integration tests for if-let, await/try, for-loop + fix cross-chunk imports
- Add 3 new test fixtures: swift-if-let-guard-let, swift-await-try, swift-for-loop-inference
- Add integration tests for if let/guard let binding resolution (4 assertions)
- Add integration tests for await/try expression unwrapping (3 assertions)
- Add for-loop-inference fixture (documented as known gap — type-env infrastructure
  is in place but call-processor re-parse path doesn't propagate the binding yet)
- Fix cross-chunk Swift implicit imports: standard processImports path now passes
  allFileList instead of chunk-only files to addSwiftImplicitImports, matching
  the fast-path behavior
- Add Swift type_annotation fallback in type-env declarationTypeNodes population
  (handles [User] array sugar where childForFieldName('type') returns null)
- Handle Swift 'pattern' node in extractVarName fallback (pattern wraps simple_identifier)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 11:40:24 +01:00
jreakinandClaude Opus 4.6 57e087de49 fix(web): strip quoted strings before readOnly keyword check
The readOnly guard was matching keywords inside string literals,
blocking legitimate queries like WHERE n.name CONTAINS "delete".
Now strips single/double-quoted strings before checking, so only
actual Cypher write keywords outside strings are blocked.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 05:09:44 -05:00
jreakinandClaude Opus 4.6 fc6c114570 fix(web): MermaidDiagram XSS, isSafeId allows @, rawMermaid typed
- Add DOMPurify.sanitize() to MermaidDiagram.tsx dangerouslySetInnerHTML
  (was rendering AI-generated SVG unsanitized — the highest-risk surface)
- Add @ to isSafeId regex for scoped npm packages (@angular/core, etc.)
- Add rawMermaid to ProcessData interface, remove (process as any) casts
- Update security tests: @scope/pkg now accepted, @angular/core added

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 05:01:57 -05:00
jreakinandClaude Opus 4.6 96181fba05 fix: pass readOnly=false for CREATE_VECTOR_INDEX in embedding pipeline
The readOnly=true default on executeQuery blocks queries containing
CREATE, which includes CALL CREATE_VECTOR_INDEX. The embedding pipeline
needs write access for this setup step.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 04:45:09 -05:00
jreakinandClaude Opus 4.6 208ade9b11 fix: add / back to isSafeId — node IDs contain file paths
Node IDs are generated as Label:filePath (e.g., Function:src/foo.ts:bar),
so forward slashes are expected in legitimate IDs. The over-tightened
regex was dropping all path-based IDs from Cypher queries.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 04:45:09 -05:00
jreakinandClaude Opus 4.6 9f14edc226 fix(web): security hardening — Cypher injection, XSS, readOnly, prepared statements
- DOMPurify.sanitize() on mermaid SVG in ProcessFlowModal
- validLabel()/validRelType() guards on all Cypher interpolation in tools.ts
- isSafeId() strict regex in ProcessesPanel (no spaces/slashes/metacharacters)
- executeQuery defaults readOnly=true (rejects CREATE/DELETE/DROP/etc.)
- Singleton promise for initLbug (prevents concurrent init races)
- Prepared statements for relation inserts (batched by label pair)
- executeWithReusedStatement for enrichment updates
- try/finally on all PreparedStatement cleanup
- 99 security guard tests (validLabel, validRelType, isSafeId, readOnly regex)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 04:45:09 -05:00
marxo126andClaude Opus 4.6 884b4acf84 fix: regenerate package-lock.json for CI compatibility
npm ci was failing with "Missing: hono@4.12.8" and
"Missing: graphology-types@0.24.8" because the lock file was
out of sync after rebase. Regenerated from clean state.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 18:59:55 +01:00
marxo126andClaude Opus 4.6 babf0f90d3 fix: pin tree-sitter versions and add npm overrides
Pin exact versions (no ^) to prevent surprise upgrades:
- tree-sitter: "0.22.4" (was "^0.22.4")
- tree-sitter-swift: "0.7.1" (was "^0.7.1")

Add npm overrides to suppress peer dependency warnings from grammar
packages that declare ^0.21.x but work fine with 0.22.4.

Note: tree-sitter-swift 0.6.0 fails to build on current Node (needs
node-gyp + Swift toolchain). 0.7.1 with prebuilt binaries is required
for Swift support to work at all.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 11:47:21 +01:00
marxo126andClaude Opus 4.6 0a3cdce00e fix: address Copilot review — private(set) export, for-loop tuple pattern
1. Export detection: exclude private(set)/fileprivate(set) from
   unexported check. Only the setter is restricted — the symbol
   itself is still readable cross-file.

2. For-loop binding: use extractVarName() instead of raw .text
   to avoid polluting scopeEnv with non-identifier keys from
   tuple destructuring patterns (e.g. `for (a, b) in ...`).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 11:47:09 +01:00
marxo126andClaude Opus 4.6 16b1a63134 feat: 7 Swift features — if/guard let, await/try, for-in, enum cases, self/super, optional chaining, multi-inheritance
Covers all high and medium impact gaps from the Swift feature coverage
analysis:

1. if let / guard let bindings: add if_statement and guard_statement
   to DECLARATION_NODE_TYPES, extract varName and value for Tier 2
   return-type propagation (callResult, copy, fieldAccess, methodCallResult)

2. await / try expression unwrapping: add unwrapSwiftExpression() that
   strips await_expression and try_expression wrappers before checking
   for call_expression. Applied in extractPendingAssignment,
   extractInitializer, and scanConstructorBinding.

3. for item in collection: add extractForLoopBinding for Swift with
   extractSwiftElementTypeFromTypeNode that handles [User] array sugar
   and Array<User> generic types. Registered in typeConfig.

4. Multiple inheritance specifiers: already working — tree-sitter
   queries match all inheritance_specifier occurrences automatically.
   Verified, no code changes needed.

5. Enum case extraction: add (enum_entry (simple_identifier) @name)
   @definition.property query to SWIFT_QUERIES.

6. self/super resolution: unskipped both describe.skip test suites
   (tree-sitter-swift 0.7.1 ships prebuilds, Node 22 build issue
   resolved). Both pass — 5 previously-skipped tests now running.

7. Optional chaining obj?.method(): already working — tree-sitter-swift
   parses the ? transparently. Verified, no code changes needed.

Tests: 3,603 → 3,608 (5 unskipped self/super tests)
Swift tests: 23 → 28 passing

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 11:47:09 +01:00
marxo126andClaude Opus 4.6 99f0aaaea5 test: add integration tests for Swift implicit imports, extension dedup, constructor fallback, export visibility
Addresses reviewer feedback: the new Swift behaviors (implicit imports,
constructor fallback, extension dedup, export detection) had no dedicated
integration tests. Adds 4 fixture directories and 11 new test assertions:

1. swift-implicit-imports: two files, no explicit import, cross-file
   constructor + member call resolves via addSwiftImplicitImports
2. swift-extension-dedup: extension creates duplicate Class node,
   constructor still resolves to primary definition
3. swift-constructor-fallback: ClassName() without `new` resolves as
   constructor via free→constructor retry
4. swift-export-visibility: internal symbols visible cross-file,
   public/open visible, private/fileprivate noted as Tier 3 limitation

All 3,603 tests pass (11 new, 0 regressions).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 11:47:09 +01:00
marxo126andClaude Opus 4.6 53b576776a feat: add extractPendingAssignment for Swift return-type inference
Swift was missing the extractPendingAssignment extractor, which meant
return-type-based variable bindings like `let user = getUser()` couldn't
propagate the return type of `getUser()` to `user`. This broke member
call resolution: `user.save()` couldn't resolve to `User.save()` when
there were competing methods (both User and Repo have save()).

Handles four Swift patterns:
- let user = getUser()        → callResult (Tier 2 propagation)
- let result = user.save()    → methodCallResult
- let name = user.name        → fieldAccess
- let copy = user             → copy

All 3,592 tests pass — including the 2 previously-failing Swift
return-type inference tests.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 11:47:09 +01:00
marxo126andClaude Opus 4.6 712598a055 fix: update Swift queries and tests for tree-sitter-swift 0.7.1
- Fix assignment query: tree-sitter-swift 0.7.1 uses named fields
  (target:/result:/suffix:) instead of positional children
- Update export detection tests: Swift `internal` (default) is now
  correctly treated as exported (module-scoped visibility)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 11:47:09 +01:00
marxo126andClaude Opus 4.6 90c9153ccc feat: upgrade tree-sitter to 0.22.4 and tree-sitter-swift to 0.7.1
Upgrades:
- tree-sitter: ^0.21.0 → ^0.22.4
- tree-sitter-swift: ^0.6.0 → ^0.7.1

Benefits:
- tree-sitter-swift@0.7.1 ships prebuilds (no manual node-gyp needed)
- Fixes parsing of Swift 5.9+ features: #Predicate, typed throws, ~Copyable
- All 13 language parsers verified compatible with tree-sitter@0.22.4

Tested:
- All 141 query/import/call tests pass
- All 13 parsers (C, C++, C#, Go, Java, JS, Kotlin, PHP, Python, Ruby,
  Rust, TypeScript, Swift) parse correctly with 0.22.4
- PricePal (61-file iOS 26 project) indexes fully: 3,094 nodes, 10,459 edges

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 11:47:09 +01:00
marxo126andClaude Opus 4.6 dfe83f333e fix: deduplicate Swift extension class nodes in call resolution
When Swift extensions create multiple Class nodes with the same name
(e.g. Product.swift + ProductMatchableConformance.swift), the call
resolver gets multiple candidates and refuses to emit a CALLS edge.

Add dedup: when all candidates share the same type (Class/Struct) and
differ only by file, prefer the primary definition (shortest filepath).

Note: This fix is partial — some constructor calls inside function
bodies may still be consumed by the type-env constructor binding
scanner before reaching resolveCallTarget. Filed as known limitation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 11:46:34 +01:00
marxo126andClaude Opus 4.6 561a54a154 refactor: simplify Swift fixes after code review
- Extract shared addSwiftImplicitImports() helper (DRY — was duplicated
  in processImports and processImportsFromExtracted)
- Cache importMap.get(srcFile) outside inner loop (avoids redundant
  Map lookups per iteration)
- Fix export-detection: use \bprivate\b regex instead of includes()
  to avoid substring false positives
- Fix groupSwiftFilesByTarget: check path boundary with indexOf + char
  check instead of loose includes()

All 141 relevant tests pass (queries + imports + calls).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 11:46:34 +01:00
marxo126andClaude Opus 4.6 65bc99c448 feat: full Swift cross-file resolution (export, imports, constructors)
Three changes that together enable cross-file call resolution for Swift:

1. export-detection.ts: Treat internal (default) Swift symbols as exported.
   Swift's default access level is `internal` (module-scoped, visible to
   all files in the same target). Only private/fileprivate are file-scoped.
   Previously all non-public/open symbols were marked unexported.

2. import-processor.ts: Add implicit import edges between all Swift files
   in the same module/target. Swift has no file-level imports — all files
   see each other automatically. Without these edges, the tiered resolver
   can't find cross-file symbols at Tier 2a (import-scoped).
   Supports SPM targets via Package.swift; falls back to single-module
   for Xcode projects without SPM.

3. call-processor.ts: Add constructor fallback for free-form calls.
   Swift constructors look like free function calls (no `new` keyword):
   `let ocr = OCRService()`. The call form is inferred as `free`, which
   filters out Class/Struct targets. Now retries with `constructor` form
   when free-form finds no callable but the name resolves to a type.

Tested on 61-file iOS 26 project (PricePal):
- Before: 0 cross-file CALLS edges
- After: full cross-file resolution (OCRService traced from ScanViewModel)
- 3,099 nodes, 10,449 edges, 246 clusters, 243 flows

Related: #406, #407

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 11:46:34 +01:00
marxo126andClaude Opus 4.6 0c8ec952ee fix: handle trailing commas in tree-sitter-swift binding.gyp patch
The patch script fails to parse tree-sitter-swift@0.6.0's binding.gyp
because the file contains both Python-style # comments AND trailing
commas in JSON arrays. The existing regex strips # comments but leaves
trailing commas, causing JSON.parse() to fail with:

  "Unexpected token ']'"

This silently prevents tree-sitter-swift from building, which means
Swift files are skipped entirely during analysis.

Fix: add a second regex pass to strip trailing commas before ] or }
after comment removal.

Fixes #386, #406

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 11:46:34 +01:00
279 changed files with 18805 additions and 4658 deletions
+91
View File
@@ -0,0 +1,91 @@
name: E2E Tests
on:
workflow_call:
jobs:
check-changes:
name: Check web module changes
runs-on: ubuntu-latest
timeout-minutes: 5
outputs:
web_changed: ${{ steps.filter.outputs.web }}
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: dorny/paths-filter@de90cc6fb38fc0963ad72b210f1f284cd68cea36 # v3
id: filter
with:
filters: |
web:
- 'gitnexus-web/**'
e2e:
name: e2e (chromium)
needs: check-changes
if: needs.check-changes.result == 'success' && needs.check-changes.outputs.web_changed == 'true'
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus-web/package-lock.json
- name: Install frontend dependencies
run: npm ci
working-directory: gitnexus-web
- name: Install Playwright browsers
run: npx playwright install --with-deps chromium
working-directory: gitnexus-web
- name: Install backend dependencies
run: npm ci
working-directory: gitnexus
- name: Build backend
run: npm run build
working-directory: gitnexus
- name: Analyze repository (index for backend)
run: |
node gitnexus/dist/cli/index.js analyze || true
if [ ! -d ".gitnexus" ]; then
echo "::error::No .gitnexus index created"
exit 1
fi
- name: Start backend server
run: node dist/cli/index.js serve &
working-directory: gitnexus
- name: Wait for backend readiness
run: npx wait-on http://localhost:4747/api/repos --timeout 30000
working-directory: gitnexus-web
- name: Start Vite dev server
run: npm run dev &
working-directory: gitnexus-web
- name: Wait for Vite dev server
run: npx wait-on http://localhost:5173 --timeout 30000
working-directory: gitnexus-web
- name: Run E2E tests
run: npx playwright test
working-directory: gitnexus-web
env:
E2E: '1'
- name: Upload test results
if: always()
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
name: e2e-results
path: |
gitnexus-web/test-results/
gitnexus-web/playwright-report/
retention-days: 5
+15
View File
@@ -12,3 +12,18 @@ jobs:
- uses: ./.github/actions/setup-gitnexus
- run: npx tsc --noEmit
working-directory: gitnexus
typecheck-web:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
with:
node-version: 20
cache: npm
cache-dependency-path: gitnexus-web/package-lock.json
- run: npm ci
working-directory: gitnexus-web
- run: npx tsc -b --noEmit
working-directory: gitnexus-web
+314 -249
View File
@@ -1,132 +1,126 @@
name: CI Report
# Triggered after the CI workflow completes. Because workflow_run
# always runs code from the *default branch*, it receives a read/write
# GITHUB_TOKEN — even when the triggering PR comes from a fork.
on:
workflow_run:
workflows: ['CI']
workflows: ["CI"]
types: [completed]
permissions:
actions: read
contents: read
pull-requests: write
actions: read # needed to list/download workflow run artifacts
contents: read # needed for sparse checkout of vitest.config.ts
pull-requests: write # needed to post sticky PR comment
jobs:
pr-report:
name: PR Report
# Only run for pull-request CI runs
if: >-
github.event.workflow_run.event == 'pull_request' &&
github.event.workflow_run.conclusion != 'cancelled'
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- name: Download PR metadata
# ── Download artifacts from the CI run ────────────────────────
- name: Download artifacts
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
with:
script: |
const fs = require('fs');
const path = require('path');
const runId = context.payload.workflow_run.id;
const artifacts = await github.rest.actions.listWorkflowRunArtifacts({
const allArtifacts = await github.rest.actions.listWorkflowRunArtifacts({
owner: context.repo.owner,
repo: context.repo.repo,
run_id: ${{ github.event.workflow_run.id }},
run_id: runId,
});
const meta = artifacts.data.artifacts.find(a => a.name === 'pr-meta');
if (!meta) {
core.setFailed('pr-meta artifact not found — skipping report');
return;
async function downloadArtifact(name, dest) {
const match = allArtifacts.data.artifacts.find(a => a.name === name);
if (!match) {
core.warning(`Artifact "${name}" not found`);
return false;
}
const zip = await github.rest.actions.downloadArtifact({
owner: context.repo.owner,
repo: context.repo.repo,
artifact_id: match.id,
archive_format: 'zip',
});
fs.mkdirSync(dest, { recursive: true });
fs.writeFileSync(path.join(dest, `${name}.zip`), Buffer.from(zip.data));
return true;
}
const zip = await github.rest.actions.downloadArtifact({
owner: context.repo.owner,
repo: context.repo.repo,
artifact_id: meta.id,
archive_format: 'zip',
});
const temp = process.env.RUNNER_TEMP;
await downloadArtifact('pr-meta', path.join(temp, 'dl'));
await downloadArtifact('test-reports', path.join(temp, 'dl'));
const dest = path.join(process.env.RUNNER_TEMP, 'pr-meta');
fs.mkdirSync(dest, { recursive: true });
fs.writeFileSync(path.join(dest, 'pr-meta.zip'), Buffer.from(zip.data));
- name: Extract artifacts
shell: bash
run: |
cd "$RUNNER_TEMP/dl"
# Extract each artifact into its own directory to avoid filename collisions
for z in *.zip; do
[ -f "$z" ] || continue
name="${z%.zip}"
mkdir -p "$RUNNER_TEMP/artifacts/$name"
unzip -o "$z" -d "$RUNNER_TEMP/artifacts/$name"
done
- name: Extract PR metadata
- name: Read PR metadata
id: meta
shell: bash
run: |
cd "$RUNNER_TEMP/pr-meta"
unzip -o pr-meta.zip
PR_NUMBER=$(cat pr-number | tr -d '[:space:]')
if ! [[ "$PR_NUMBER" =~ ^[0-9]+$ ]]; then
echo "::error::Invalid PR number: '$PR_NUMBER'"
exit 1
DIR="$RUNNER_TEMP/artifacts/pr-meta"
if [ ! -f "$DIR/pr_number" ]; then
echo "skip=true" >> "$GITHUB_OUTPUT"
echo "::warning::pr_number artifact missing — skipping report"
exit 0
fi
echo "pr-number=$PR_NUMBER" >> "$GITHUB_OUTPUT"
echo "quality=$(cat quality-result | tr -d '[:space:]')" >> "$GITHUB_OUTPUT"
echo "tests=$(cat tests-result | tr -d '[:space:]')" >> "$GITHUB_OUTPUT"
# Validate PR number is a positive integer (artifact comes from
# untrusted fork code, so treat contents defensively).
PR_NUM=$(cat "$DIR/pr_number" | tr -d '[:space:]')
if ! [[ "$PR_NUM" =~ ^[0-9]+$ ]]; then
echo "skip=true" >> "$GITHUB_OUTPUT"
echo "::error::Invalid PR number in artifact: '$PR_NUM'"
exit 0
fi
- name: Download test reports
id: download-test-reports
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
echo "skip=false" >> "$GITHUB_OUTPUT"
echo "pr_number=$PR_NUM" >> "$GITHUB_OUTPUT"
# Validate job-result strings against known GitHub Actions values.
# Artifact contents come from the PR workflow (potentially untrusted
# fork code), so we whitelist to prevent newline injection into
# GITHUB_OUTPUT.
validate_result() {
local val
val=$(cat "$1" | tr -d '[:space:]')
case "$val" in
success|failure|cancelled|skipped) echo "$val" ;;
*) echo "unknown" ;;
esac
}
echo "quality=$(validate_result "$DIR/quality_result")" >> "$GITHUB_OUTPUT"
echo "tests=$(validate_result "$DIR/tests_result")" >> "$GITHUB_OUTPUT"
echo "e2e=$(validate_result "$DIR/e2e_result")" >> "$GITHUB_OUTPUT"
- name: Checkout (for vitest config)
if: steps.meta.outputs.skip != 'true'
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
with:
script: |
const fs = require('fs');
const path = require('path');
const artifacts = await github.rest.actions.listWorkflowRunArtifacts({
owner: context.repo.owner,
repo: context.repo.repo,
run_id: ${{ github.event.workflow_run.id }},
});
const reports = artifacts.data.artifacts.find(a => a.name === 'test-reports');
if (!reports) {
core.warning('test-reports artifact not found');
return;
}
const zip = await github.rest.actions.downloadArtifact({
owner: context.repo.owner,
repo: context.repo.repo,
artifact_id: reports.id,
archive_format: 'zip',
});
const dest = path.join(process.env.RUNNER_TEMP, 'test-reports');
fs.mkdirSync(dest, { recursive: true });
fs.writeFileSync(path.join(dest, 'test-reports.zip'), Buffer.from(zip.data));
- name: Extract test reports
if: steps.download-test-reports.outcome == 'success'
shell: bash
run: |
cd "$RUNNER_TEMP/test-reports"
unzip -o test-reports.zip || true
- name: Fetch cross-platform job results
id: jobs
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
with:
script: |
const jobs = await github.rest.actions.listJobsForWorkflowRun({
owner: context.repo.owner,
repo: context.repo.repo,
run_id: ${{ github.event.workflow_run.id }},
per_page: 50,
});
const results = {};
for (const job of jobs.data.jobs) {
if (job.name.includes('ubuntu')) results.ubuntu = job.conclusion || 'pending';
else if (job.name.includes('windows')) results.windows = job.conclusion || 'pending';
else if (job.name.includes('macos')) results.macos = job.conclusion || 'pending';
}
core.setOutput('ubuntu', results.ubuntu || 'unknown');
core.setOutput('windows', results.windows || 'unknown');
core.setOutput('macos', results.macos || 'unknown');
sparse-checkout: gitnexus/vitest.config.ts
sparse-checkout-cone-mode: false
# ── Fetch base branch coverage for delta reporting ───────────
- name: Fetch base branch coverage
if: steps.meta.outputs.skip != 'true'
id: base-coverage
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
with:
@@ -134,6 +128,7 @@ jobs:
const fs = require('fs');
const path = require('path');
// Find the latest successful CI run on main
const runs = await github.rest.actions.listWorkflowRuns({
owner: context.repo.owner,
repo: context.repo.repo,
@@ -145,6 +140,7 @@ jobs:
if (runs.data.workflow_runs.length === 0) {
core.setOutput('found', 'false');
core.info('No successful main branch CI runs found');
return;
}
@@ -158,6 +154,7 @@ jobs:
const testReports = artifacts.data.artifacts.find(a => a.name === 'test-reports');
if (!testReports) {
core.setOutput('found', 'false');
core.info('No test-reports artifact on main branch');
return;
}
@@ -175,174 +172,242 @@ jobs:
core.setOutput('dir', dest);
- name: Extract base coverage
if: steps.base-coverage.outputs.found == 'true'
if: steps.meta.outputs.skip != 'true' && steps.base-coverage.outputs.found == 'true'
shell: bash
run: |
cd "${{ steps.base-coverage.outputs.dir }}"
mkdir -p base
unzip -o base.zip -d base
- name: Build and post report
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7
- name: Build report
if: steps.meta.outputs.skip != 'true'
id: report
shell: bash
env:
PR_NUMBER: ${{ steps.meta.outputs.pr-number }}
QUALITY: ${{ steps.meta.outputs.quality }}
TESTS: ${{ steps.meta.outputs.tests }}
UBUNTU: ${{ steps.jobs.outputs.ubuntu }}
WINDOWS: ${{ steps.jobs.outputs.windows }}
MACOS: ${{ steps.jobs.outputs.macos }}
E2E: ${{ steps.meta.outputs.e2e }}
BASE_FOUND: ${{ steps.base-coverage.outputs.found }}
BASE_DIR: ${{ steps.base-coverage.outputs.dir }}
RUN_ID: ${{ github.event.workflow_run.id }}
HEAD_SHA: ${{ github.event.workflow_run.head_sha }}
RUN_URL: ${{ github.event.workflow_run.html_url }}
run: |
DIR="$RUNNER_TEMP/artifacts"
# ── Helper: read coverage summary into prefixed vars ──
read_cov() {
local prefix=$1 file=$2
if [ -n "$file" ] && [ -f "$file" ]; then
local val
val=$(jq -r '.total.statements.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_STMTS" '%s' "$val"
val=$(jq -r '.total.branches.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_BRANCH" '%s' "$val"
val=$(jq -r '.total.functions.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_FUNCS" '%s' "$val"
val=$(jq -r '.total.lines.pct // "N/A"' "$file" 2>/dev/null) || val="N/A"
printf -v "${prefix}_LINES" '%s' "$val"
val=$(jq -r '"\(.total.statements.covered)/\(.total.statements.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_STMTS_COV" '%s' "$val"
val=$(jq -r '"\(.total.branches.covered)/\(.total.branches.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_BRANCH_COV" '%s' "$val"
val=$(jq -r '"\(.total.functions.covered)/\(.total.functions.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_FUNCS_COV" '%s' "$val"
val=$(jq -r '"\(.total.lines.covered)/\(.total.lines.total)"' "$file" 2>/dev/null) || val=""
printf -v "${prefix}_LINES_COV" '%s' "$val"
return 0
else
printf -v "${prefix}_STMTS" '%s' "N/A"
printf -v "${prefix}_BRANCH" '%s' "N/A"
printf -v "${prefix}_FUNCS" '%s' "N/A"
printf -v "${prefix}_LINES" '%s' "N/A"
printf -v "${prefix}_STMTS_COV" '%s' ""
printf -v "${prefix}_BRANCH_COV" '%s' ""
printf -v "${prefix}_FUNCS_COV" '%s' ""
printf -v "${prefix}_LINES_COV" '%s' ""
return 1
fi
}
# ── Read coverage reports ──
UNIT_SUMMARY=$(find "$DIR/test-reports" -name "coverage-summary.json" -type f 2>/dev/null | head -1)
read_cov "U" "$UNIT_SUMMARY"
# ── Read base branch coverage (main) ──
BASE_SUMMARY=""
if [ "$BASE_FOUND" = "true" ] && [ -n "$BASE_DIR" ]; then
BASE_SUMMARY=$(find "$BASE_DIR/base" -name "coverage-summary.json" -type f 2>/dev/null | head -1)
fi
read_cov "B" "$BASE_SUMMARY"
# ── Locate test results ──
RESULTS_FILE=$(find "$DIR/test-reports" -name "test-results.json" -type f 2>/dev/null | head -1)
WEB_RESULTS_FILE=$(find "$DIR/test-reports" -name "web-test-results.json" -type f 2>/dev/null | head -1)
sum_results() {
local file=$1
if [ -n "$file" ] && [ -f "$file" ]; then
jq -r '"\(.numTotalTests) \(.numPassedTests) \(.numFailedTests) \(.numPendingTests) \(.numTotalTestSuites) \(((.testResults | map(.endTime) | max) - (.startTime)) / 1000 | floor)"' "$file" 2>/dev/null || echo "0 0 0 0 0 0"
else
echo "0 0 0 0 0 0"
fi
}
read CLI_T CLI_P CLI_F CLI_S CLI_SU CLI_D <<< "$(sum_results "$RESULTS_FILE")"
read WEB_T WEB_P WEB_F WEB_S WEB_SU WEB_D <<< "$(sum_results "$WEB_RESULTS_FILE")"
TOTAL=$((CLI_T + WEB_T))
PASSED=$((CLI_P + WEB_P))
FAILED=$((CLI_F + WEB_F))
SKIPPED=$((CLI_S + WEB_S))
SUITES=$((CLI_SU + WEB_SU))
DURATION=$((CLI_D > WEB_D ? CLI_D : WEB_D))
# ── Status helpers ──
status_icon() {
case "$1" in
success) echo "✅" ;;
failure) echo "❌" ;;
cancelled) echo "⏭️" ;;
*) echo "❓" ;;
esac
}
# Validate a value looks like a number (integer or decimal, optional
# leading minus). Returns 1 for anything else — guards against awk
# injection when artifact values come from untrusted fork code.
is_numeric() { [[ "$1" =~ ^-?[0-9]+(\.[0-9]+)?$ ]]; }
cov_delta() {
local pct=$1 base=$2
if [ "$pct" = "N/A" ] || [ "$base" = "N/A" ]; then echo "—"; return; fi
if ! is_numeric "$pct" || ! is_numeric "$base"; then echo "—"; return; fi
local diff
diff=$(awk -v p="$pct" -v b="$base" 'BEGIN { printf "%.1f", p - b }')
if [ "$(awk -v p="$pct" -v b="$base" 'BEGIN { print (p > b) ? 1 : 0 }')" = "1" ]; then
echo "📈 +${diff}"
elif [ "$(awk -v p="$pct" -v b="$base" 'BEGIN { print (p < b) ? 1 : 0 }')" = "1" ]; then
echo "📉 ${diff}"
else
echo "= ${diff}"
fi
}
cov_bar() {
local pct=$1 base=$2
if [ "$pct" = "N/A" ] || ! is_numeric "$pct"; then echo "—"; return; fi
local filled
filled=$(awk -v p="$pct" 'BEGIN { printf "%d", p / 5 }')
(( filled < 0 )) && filled=0
(( filled > 20 )) && filled=20
local empty=$((20 - filled))
local bar=""
for ((i=0; i<filled; i++)); do bar+="█"; done
for ((i=0; i<empty; i++)); do bar+="░"; done
# Green if >= base (or base unavailable), red if dropped
if [ "$base" = "N/A" ] || ! is_numeric "$base" || [ "$(awk -v p="$pct" -v b="$base" 'BEGIN { print (p >= b) ? 1 : 0 }')" = "1" ]; then
echo "🟢 ${bar}"
else
echo "🔴 ${bar}"
fi
}
# ── Overall status ──
if [[ "$QUALITY" == "success" && "$TESTS" == "success" && ("$E2E" == "success" || "$E2E" == "skipped") ]]; then
OVERALL="✅ **All checks passed**"
else
OVERALL="❌ **Some checks failed**"
fi
# ── Build markdown ──
{
echo "body<<GITNEXUS_CI_REPORT_EOF_7f3a"
echo "## CI Report"
echo ""
echo "${OVERALL}"
echo ""
echo "### Pipeline Status"
echo ""
echo "| Stage | Status | Details |"
echo "|-------|--------|---------|"
echo "| $(status_icon "$QUALITY") Typecheck | \`${QUALITY}\` | tsc --noEmit |"
echo "| $(status_icon "$TESTS") Tests | \`${TESTS}\` | unit tests, 3 platforms |"
echo "| $(status_icon "$E2E") E2E | \`${E2E}\` | gitnexus-web changes only |"
echo ""
if [ "$TOTAL" -gt 0 ] 2>/dev/null; then
echo "### Test Results"
echo ""
echo "| Tests | Passed | Failed | Skipped | Duration |"
echo "|-------|--------|--------|---------|----------|"
echo "| ${TOTAL} | ${PASSED} | ${FAILED} | ${SKIPPED} | ${DURATION}s |"
echo ""
if [ "$FAILED" = "0" ]; then
echo "✅ All **${PASSED}** tests passed"
else
echo "❌ **${FAILED}** failed / **${PASSED}** passed"
fi
if [ "$SKIPPED" != "0" ]; then
echo ""
echo "<details>"
echo "<summary>${SKIPPED} test(s) skipped — expand for details</summary>"
echo ""
for rf in "$RESULTS_FILE" "$WEB_RESULTS_FILE"; do
if [ -n "$rf" ] && [ -f "$rf" ]; then
jq -r '
.testResults[]
| .assertionResults[]?
| select(.status == "pending" or .status == "skipped")
| "- \(.ancestorTitles | join(" > ")) > \(.title)"
' "$rf" 2>/dev/null || true
fi
done
echo ""
echo "</details>"
fi
echo ""
fi
# ── Coverage table helper ──
cov_table() {
local label=$1 s=$2 b=$3 f=$4 l=$5 sc=$6 bc=$7 fc=$8 lc=$9
shift 9
local bs=$1 bb=$2 bf=$3 bl=$4
echo "#### ${label}"
echo ""
echo "| Metric | Coverage | Covered | Base | Delta | Status |"
echo "|--------|----------|---------|------|-------|--------|"
echo "| Statements | **${s}%** | ${sc} | ${bs}% | $(cov_delta "$s" "$bs") | $(cov_bar "$s" "$bs") |"
echo "| Branches | **${b}%** | ${bc} | ${bb}% | $(cov_delta "$b" "$bb") | $(cov_bar "$b" "$bb") |"
echo "| Functions | **${f}%** | ${fc} | ${bf}% | $(cov_delta "$f" "$bf") | $(cov_bar "$f" "$bf") |"
echo "| Lines | **${l}%** | ${lc} | ${bl}% | $(cov_delta "$l" "$bl") | $(cov_bar "$l" "$bl") |"
echo ""
}
if [ "$U_STMTS" != "N/A" ]; then
echo "### Code Coverage"
echo ""
cov_table "Tests" \
"$U_STMTS" "$U_BRANCH" "$U_FUNCS" "$U_LINES" \
"$U_STMTS_COV" "$U_BRANCH_COV" "$U_FUNCS_COV" "$U_LINES_COV" \
"$B_STMTS" "$B_BRANCH" "$B_FUNCS" "$B_LINES"
else
echo "### Code Coverage"
echo ""
echo "⚠️ Coverage data unavailable - check the [unit test job](${RUN_URL}) for details."
echo ""
fi
echo "---"
echo "<sub>📋 [View full run](${RUN_URL}) · Generated by CI</sub>"
echo "GITNEXUS_CI_REPORT_EOF_7f3a"
} >> "$GITHUB_OUTPUT"
- name: Comment on PR
if: steps.meta.outputs.skip != 'true'
uses: marocchino/sticky-pull-request-comment@773744901bac0e8cbb5a0dc842800d45e9b2b405 # v2
with:
script: |
const fs = require('fs');
const path = require('path');
const icon = (s) => ({ success: '✅', failure: '❌', cancelled: '⏭️' }[s] || '❓');
const temp = process.env.RUNNER_TEMP;
// ── Read coverage ──
function readCov(dir) {
const out = { stmts: 'N/A', branch: 'N/A', funcs: 'N/A', lines: 'N/A',
stmtsCov: '', branchCov: '', funcsCov: '', linesCov: '' };
try {
const files = require('child_process')
.execSync(`find "${dir}" -name coverage-summary.json -type f`, { encoding: 'utf8' })
.trim().split('\n').filter(Boolean);
if (!files.length) return out;
const d = JSON.parse(fs.readFileSync(files[0], 'utf8')).total;
out.stmts = d.statements.pct; out.branch = d.branches.pct;
out.funcs = d.functions.pct; out.lines = d.lines.pct;
out.stmtsCov = `${d.statements.covered}/${d.statements.total}`;
out.branchCov = `${d.branches.covered}/${d.branches.total}`;
out.funcsCov = `${d.functions.covered}/${d.functions.total}`;
out.linesCov = `${d.lines.covered}/${d.lines.total}`;
} catch {}
return out;
}
const cov = readCov(path.join(temp, 'test-reports'));
const base = process.env.BASE_FOUND === 'true'
? readCov(path.join(process.env.BASE_DIR, 'base'))
: { stmts: 'N/A', branch: 'N/A', funcs: 'N/A', lines: 'N/A' };
// ── Read test results ──
let total = 0, passed = 0, failed = 0, skipped = 0, suites = 0, duration = '0s';
let skippedTests = [];
try {
const files = require('child_process')
.execSync(`find "${path.join(temp, 'test-reports')}" -name test-results.json -type f`, { encoding: 'utf8' })
.trim().split('\n').filter(Boolean);
if (files.length) {
const r = JSON.parse(fs.readFileSync(files[0], 'utf8'));
total = r.numTotalTests || 0;
passed = r.numPassedTests || 0;
failed = r.numFailedTests || 0;
skipped = r.numPendingTests || 0;
suites = r.numTotalTestSuites || 0;
const durS = Math.floor((Math.max(...r.testResults.map(t => t.endTime)) - r.startTime) / 1000);
duration = durS >= 60 ? `${Math.floor(durS / 60)}m ${durS % 60}s` : `${durS}s`;
// Collect skipped test names
for (const suite of r.testResults) {
for (const t of (suite.assertionResults || [])) {
if (t.status === 'pending' || t.status === 'skipped') {
skippedTests.push(`- ${t.ancestorTitles.join(' > ')} > ${t.title}`);
}
}
}
}
} catch {}
// ── Coverage delta ──
function delta(pct, basePct) {
if (pct === 'N/A' || basePct === 'N/A') return '—';
const d = (pct - basePct).toFixed(1);
const dNum = parseFloat(d);
if (dNum > 0) return `📈 +${d}%`;
if (dNum < 0) return `📉 ${d}%`;
return '=';
}
// ── Build markdown ──
const { PR_NUMBER, QUALITY, TESTS, UBUNTU, WINDOWS, MACOS, RUN_ID, HEAD_SHA } = process.env;
const prNumber = parseInt(PR_NUMBER, 10);
const overall = (QUALITY === 'success' && TESTS === 'success')
? '✅ **All checks passed**' : '❌ **Some checks failed**';
const sha = HEAD_SHA.slice(0, 7);
let body = `## CI Report\n\n${overall} &ensp; \`${sha}\`\n\n`;
body += `### Pipeline\n\n`;
body += `| Stage | Status | Ubuntu | Windows | macOS |\n`;
body += `|-------|--------|--------|---------|-------|\n`;
body += `| Typecheck | ${icon(QUALITY)} \`${QUALITY}\` | — | — | — |\n`;
body += `| Tests | ${icon(TESTS)} \`${TESTS}\` | ${icon(UBUNTU)} | ${icon(WINDOWS)} | ${icon(MACOS)} |\n\n`;
if (total > 0) {
body += `### Tests\n\n`;
body += `| Metric | Value |\n|--------|-------|\n`;
body += `| Total | **${total}** |\n`;
body += `| Passed | **${passed}** |\n`;
if (failed > 0) body += `| Failed | **${failed}** |\n`;
if (skipped > 0) body += `| Skipped | ${skipped} |\n`;
body += `| Files | ${suites} |\n`;
body += `| Duration | ${duration} |\n\n`;
if (failed === 0) {
body += `✅ All **${passed}** tests passed across **${suites}** files\n`;
} else {
body += `❌ **${failed}** failed / **${passed}** passed\n`;
}
if (skippedTests.length > 0) {
body += `\n<details>\n<summary>${skipped} test(s) skipped</summary>\n\n`;
body += skippedTests.join('\n') + '\n\n</details>\n';
}
body += '\n';
}
if (cov.stmts !== 'N/A') {
body += `### Coverage\n\n`;
body += `| Metric | Coverage | Covered | Base (main) | Delta |\n`;
body += `|--------|----------|---------|-------------|-------|\n`;
body += `| Statements | **${cov.stmts}%** | ${cov.stmtsCov} | ${base.stmts}% | ${delta(cov.stmts, base.stmts)} |\n`;
body += `| Branches | **${cov.branch}%** | ${cov.branchCov} | ${base.branch}% | ${delta(cov.branch, base.branch)} |\n`;
body += `| Functions | **${cov.funcs}%** | ${cov.funcsCov} | ${base.funcs}% | ${delta(cov.funcs, base.funcs)} |\n`;
body += `| Lines | **${cov.lines}%** | ${cov.linesCov} | ${base.lines}% | ${delta(cov.lines, base.lines)} |\n\n`;
} else {
const runUrl = `${context.serverUrl}/${context.repo.owner}/${context.repo.repo}/actions/runs/${RUN_ID}`;
body += `### Coverage\n\n⚠️ Coverage data unavailable — check the [test job](${runUrl}) for details.\n\n`;
}
const runUrl = `${context.serverUrl}/${context.repo.owner}/${context.repo.repo}/actions/runs/${RUN_ID}`;
body += `---\n<sub>📋 [Full run](${runUrl}) · Coverage from Ubuntu · Generated by CI</sub>`;
// ── Post sticky comment ──
const { data: comments } = await github.rest.issues.listComments({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: prNumber,
per_page: 100,
direction: 'desc',
});
const marker = '<!-- ci-report -->';
const existing = comments.find(c => c.body?.includes(marker));
const fullBody = marker + '\n' + body;
if (existing) {
await github.rest.issues.updateComment({
owner: context.repo.owner,
repo: context.repo.repo,
comment_id: existing.id,
body: fullBody,
});
} else {
await github.rest.issues.createComment({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: prNumber,
body: fullBody,
});
}
header: ci-report
number: ${{ steps.meta.outputs.pr_number }}
message: ${{ steps.report.outputs.body }}
+13
View File
@@ -28,6 +28,18 @@ jobs:
--coverage.reportOnFailure=true
working-directory: gitnexus
- name: Install gitnexus-web dependencies
run: npm ci
working-directory: gitnexus-web
- name: Run gitnexus-web unit tests
run: >-
npx vitest run
--reporter=default
--reporter=json
--outputFile=web-test-results.json
working-directory: gitnexus-web
- name: Upload test reports
if: always()
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
@@ -37,6 +49,7 @@ jobs:
gitnexus/coverage/coverage-summary.json
gitnexus/coverage/coverage-final.json
gitnexus/test-results.json
gitnexus-web/web-test-results.json
retention-days: 5
cross-platform:
+62 -35
View File
@@ -16,8 +16,10 @@ concurrency:
# ── Reusable workflow orchestration ─────────────────────────────────
# Each concern lives in its own workflow file for maintainability:
# ci-quality.yml — typecheck (tsc --noEmit)
# ci-tests.yml — all tests with coverage (ubuntu) + cross-platform
# ci-report.yml — PR comment (workflow_run trigger for fork write access)
# ci-tests.yml — unit + integration tests with coverage + cross-platform
# ci-e2e.yml — E2E tests (only when gitnexus-web/ changes)
#
# Shared setup is DRY via .github/actions/setup-gitnexus composite action.
jobs:
quality:
@@ -30,11 +32,59 @@ jobs:
permissions:
contents: read
e2e:
uses: ./.github/workflows/ci-e2e.yml
permissions:
contents: read
# ── Save PR metadata for the reporting workflow ─────────────────
# The ci-report.yml workflow (triggered by workflow_run) needs the
# PR number and job results to post a comment. We save them as an
# artifact because workflow_run context doesn't reliably carry PR
# info for fork PRs.
save-pr-meta:
name: Save PR Metadata
if: always() && github.event_name == 'pull_request'
needs: [quality, tests, e2e]
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- name: Write metadata
shell: bash
env:
PR_NUMBER: ${{ github.event.number }}
QUALITY: ${{ needs.quality.result }}
TESTS: ${{ needs.tests.result }}
E2E: ${{ needs.e2e.result }}
run: |
mkdir -p pr-meta
echo "$PR_NUMBER" > pr-meta/pr_number
echo "$QUALITY" > pr-meta/quality_result
echo "$TESTS" > pr-meta/tests_result
echo "$E2E" > pr-meta/e2e_result
# TODO(post-merge): remove backward-compat copies once ci-report.yml
# on main reads underscore names.
# Backward-compat: ci-report.yml on main still reads hyphenated
# names. workflow_run always executes from the default branch, so
# the main-branch reader won't find the underscore variants until
# this PR is merged. Write both until then.
cp pr-meta/pr_number pr-meta/pr-number
cp pr-meta/quality_result pr-meta/quality-result
cp pr-meta/tests_result pr-meta/tests-result
cp pr-meta/e2e_result pr-meta/e2e-result
- name: Upload PR metadata
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
name: pr-meta
path: pr-meta/
retention-days: 1
# ── Unified CI gate ──────────────────────────────────────────────
# Single required check for branch protection.
ci-status:
name: CI Gate
needs: [quality, tests]
needs: [quality, tests, e2e]
if: always()
runs-on: ubuntu-latest
timeout-minutes: 5
@@ -44,40 +94,17 @@ jobs:
env:
QUALITY: ${{ needs.quality.result }}
TESTS: ${{ needs.tests.result }}
E2E: ${{ needs.e2e.result }}
run: |
echo "Quality: $QUALITY"
echo "Tests: $TESTS"
echo "Quality: $QUALITY"
echo "Tests: $TESTS"
echo "E2E: $E2E"
if [[ "$QUALITY" != "success" ]] ||
[[ "$TESTS" != "success" ]]; then
echo "::error::One or more CI jobs failed"
echo "::error::Quality or test jobs failed"
exit 1
fi
if [[ "$E2E" != "success" && "$E2E" != "skipped" ]]; then
echo "::error::E2E job failed"
exit 1
fi
# ── PR metadata for ci-report.yml ────────────────────────────────
# Saves PR number and job results so the workflow_run-triggered
# report can post comments with a write token (works for forks).
save-pr-meta:
name: Save PR Metadata
if: always() && github.event_name == 'pull_request'
needs: [quality, tests]
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- name: Write PR metadata
shell: bash
env:
PR_NUMBER: ${{ github.event.pull_request.number }}
QUALITY: ${{ needs.quality.result }}
TESTS: ${{ needs.tests.result }}
run: |
mkdir -p pr-meta
echo "$PR_NUMBER" > pr-meta/pr-number
echo "$QUALITY" > pr-meta/quality-result
echo "$TESTS" > pr-meta/tests-result
- name: Upload PR metadata
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
name: pr-meta
path: pr-meta/
retention-days: 1
+23 -1
View File
@@ -33,6 +33,8 @@ coverage/
# Misc
*.local
HANDOFF.md
HANDOFF*.md
.vercel
@@ -57,6 +59,14 @@ assets/
# Generated files (should not be indexed)
repomix-output*
# Playwright artifacts
gitnexus-web/playwright-report/
gitnexus-web/test-results/
# Python test artifacts
eval/.coverage
eval/.hypothesis/
# Design docs (local only)
docs/plans/
@@ -71,4 +81,16 @@ GitNexus.sln
# Git worktrees
.worktrees/
/github/scripts/triage/__pycache__/
/github/scripts/triage/__pycache__/
.claude-flow/
.claude/agents/
.claude/commands/
.claude/helpers
.claude/skills/
!.claude/skills/gitnexus/
.history/
.swarm/
+1 -1
View File
@@ -1,7 +1,7 @@
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (2273 symbols, 5419 relationships, 174 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (2487 symbols, 6056 relationships, 188 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
+1 -1
View File
@@ -1,7 +1,7 @@
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (2273 symbols, 5419 relationships, 174 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (2487 symbols, 6056 relationships, 188 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
+4
View File
@@ -107,7 +107,11 @@ If you prefer manual configuration:
**Claude Code** (full support — MCP + skills + hooks):
```bash
# macOS / Linux
claude mcp add gitnexus -- npx -y gitnexus@latest mcp
# Windows
claude mcp add gitnexus -- cmd /c npx -y gitnexus@latest mcp
```
**Codex** (full support — MCP + skills):
+100
View File
@@ -0,0 +1,100 @@
# COBOL Code Indexing
GitNexus indexes COBOL codebases using a **regex-only extraction** strategy, bypassing tree-sitter entirely. This document explains why, how the pipeline works, and links to detailed sub-documents.
## Why Regex-Only?
The tree-sitter-cobol grammar (v0.0.1) has three critical limitations that make it unusable for production indexing:
| Issue | Impact | Severity |
|-------|--------|----------|
| External scanner hangs on ~5% of files | No timeout mechanism exists for the C scanner; the process blocks indefinitely | **Blocking** |
| Only ~15% of paragraph headers detected | Most procedure-division paragraphs are invisible to the grammar | High |
| Patch markers in cols 1-6 cause parse errors | Enterprise COBOL uses non-standard sequence area content (e.g., `mzADD`, `estero`, `#FIX`) | High |
Because the external scanner hang cannot be interrupted (there is no `setTimeoutMicros` equivalent for tree-sitter), using tree-sitter-cobol would hang the indexing pipeline on a non-trivial fraction of real-world files.
The regex-only approach provides:
- **Speed**: ~1ms per file average extraction time
- **Reliability**: zero hangs, zero crashes across 13,000+ files
- **Coverage**: captures all critical symbols -- program name, paragraphs, sections, CALL, PERFORM, COPY, data items (01-77, 88-level), file declarations, FD entries, EXEC SQL/CICS blocks, ENTRY points, and MOVE statements
## Architecture
```mermaid
flowchart TD
A[Repository Scan] --> B{File Detection}
B -->|Extension match| C[COBOL file]
B -->|GITNEXUS_COBOL_DIRS match| C
B -->|No match| Z[Skip]
C --> D{Copybook?}
D -->|Yes| E[Add to Copybook Map]
D -->|No| F[Source Program]
E --> G[COPY Expansion Engine]
F --> G
G -->|Inline copybook content| H[Expanded Source]
H --> I[Patch Marker Cleanup]
I --> J[Regex State Machine]
J --> K[Extracted Symbols]
K --> L[Graph Model Builder]
L --> M[Knowledge Graph]
subgraph "Per-Chunk Processing"
G
H
I
J
K
L
end
subgraph "Post-Processing"
M --> N[Community Detection]
M --> O[Process Detection]
M --> P[Contract Detection]
end
style J fill:#e8f5e9,stroke:#2e7d32
style G fill:#e3f2fd,stroke:#1565c0
```
## COBOL vs Tree-Sitter Languages
| Feature | COBOL (Regex) | Tree-Sitter Languages |
|---------|--------------|----------------------|
| Parser | Single-pass regex state machine | tree-sitter grammar + queries |
| Speed | ~1ms/file | ~5ms/file |
| AST available | No | Yes |
| COPY expansion | Yes (pre-processing step) | N/A |
| Deep indexing | Data items, SQL, CICS, FD, ENTRY | Type annotations, generics, etc. |
| Call extraction | PERFORM (intra-file) + CALL (cross-program) | AST-based call site detection |
| Import extraction | COPY statements | `import`/`require`/`use`/`#include` |
| Coverage | All critical symbols | Language-dependent query coverage |
| Failure mode | Never hangs | External scanner can hang (COBOL only) |
## Sub-Documents
| Document | Description |
|----------|-------------|
| [File Detection](./file-detection.md) | Extension mapping, `GITNEXUS_COBOL_DIRS`, copybook classification |
| [COPY Expansion](./copy-expansion.md) | Copybook inlining, REPLACING transformations, cycle detection |
| [Regex Extraction](./regex-extraction.md) | State machine, regex patterns, line processing |
| [Deep Indexing](./deep-indexing.md) | Data items, EXEC SQL/CICS, file declarations, FD, ENTRY, MOVE |
| [Graph Model](./graph-model.md) | COBOL-specific node types, edge types, full annotated example |
| [Performance](./performance.md) | Benchmarks, worker pool tuning, caps, troubleshooting |
## Key Source Files
| File | Purpose |
|------|---------|
| `gitnexus/src/core/ingestion/cobol-preprocessor.ts` | Patch marker cleanup + regex extraction engine |
| `gitnexus/src/core/ingestion/cobol-copy-expander.ts` | COPY statement expansion with REPLACING |
| `gitnexus/src/core/ingestion/utils.ts` | `getLanguageFromPath`, `getLanguageFromFilename` |
| `gitnexus/src/core/ingestion/pipeline.ts` | `isCobolCopybook`, `expandCobolCopies`, `detectCrossProgamContracts` |
| `gitnexus/src/core/ingestion/workers/parse-worker.ts` | `processCobolRegexOnly` -- graph model builder |
| `gitnexus/src/core/ingestion/workers/worker-pool.ts` | Configurable sub-batch size for COBOL |
+145
View File
@@ -0,0 +1,145 @@
# COBOL COPY Expansion
The COPY statement is COBOL's include mechanism -- analogous to `#include` in C or `import` in modern languages. GitNexus expands COPY statements **before** regex extraction so that symbols defined inside copybooks (data items, paragraphs, etc.) are visible in the program's extracted graph.
## Supported Syntax
### Basic COPY
```cobol
COPY CPSESP.
COPY "WORKGRID.CPY".
```
Inlines the content of the named copybook, replacing the COPY line(s).
### COPY with REPLACING
```cobol
COPY CPSESP REPLACING "ANAZI-KEY" BY "LK-KEY".
COPY CPSESP REPLACING LEADING "ESP-" BY "LK-ESP-"
LEADING "KPSESPL" BY "LK-KPSESPL".
COPY LINKAGE REPLACING TRAILING "-IN" BY "-OUT".
```
Three REPLACING types are supported:
| Type | Syntax | Behavior | Example |
| ------------ | ------------------------------------ | --------------------------------------- | -------------------------------- |
| **EXACT** | `REPLACING "OLD" BY "NEW"` | Replace exact identifier matches | `ANAZI-KEY` becomes `LK-KEY` |
| **LEADING** | `REPLACING LEADING "PFX-" BY "NEW-"` | Replace prefix on all COBOL identifiers | `ESP-NAME` becomes `LK-ESP-NAME` |
| **TRAILING** | `REPLACING TRAILING "-IN" BY "-OUT"` | Replace suffix on all COBOL identifiers | `DATA-IN` becomes `DATA-OUT` |
Multiple REPLACING clauses can appear in a single COPY statement. They are applied in order to each COBOL identifier in the copybook content.
### Multi-Line COPY
COPY statements can span multiple lines (standard COBOL continuation rules apply):
```cobol
COPY CPSESP REPLACING
- LEADING "ESP-" BY "LK-ESP-"
- LEADING "KPSESPL" BY "LK-KPSESPL".
```
Continuation lines (indicator `-` in column 7) are merged before COPY statement scanning.
## Expansion Flow
```mermaid
sequenceDiagram
participant Pipeline
participant Expander as COPY Expander
participant Resolver
participant Reader
Pipeline->>Pipeline: Identify all COBOL files
Pipeline->>Pipeline: Classify copybooks vs programs
Pipeline->>Reader: Read all copybook content upfront
Reader-->>Pipeline: Copybook content map (name -> content)
loop For each source file in chunk
Pipeline->>Expander: expandCopies(content, filePath, resolveFile, readFile)
Expander->>Expander: Merge continuation lines
Expander->>Expander: Detect COPY statements via regex
loop For each COPY statement (reverse order)
Expander->>Resolver: resolveFile(copyTarget)
Resolver-->>Expander: Copybook key or null
alt Resolved successfully
Expander->>Reader: readFile(resolvedKey)
Reader-->>Expander: Copybook content
Expander->>Expander: Apply REPLACING transformations
Expander->>Expander: Recurse for nested COPYs (depth + 1)
Expander->>Expander: Splice expanded content into output
else Not resolved
Expander->>Expander: Keep original COPY line
end
end
Expander-->>Pipeline: Expanded content + resolution metadata
Pipeline->>Pipeline: Replace file content with expanded content
end
```
## Cycle Detection
Circular COPY references (e.g., copybook A includes copybook B which includes copybook A) are detected and handled:
1. Each expansion chain maintains a `visited` set of resolved copybook paths
2. If a copybook path is already in the visited set, the expansion is skipped
3. A `warnedCircular` set (shared across all files in a chunk) deduplicates warning messages
Known circular copybooks in PROJECT-NAME: `ANAZI`, `ANDIP`, `QDIPE` (self-referential includes).
## Max Depth
Nested COPY expansion is limited to **10 levels** (`DEFAULT_MAX_DEPTH`). If a COPY chain exceeds this depth, a warning is logged and the remaining COPY statements are left unexpanded.
## REPLACING Application Detail
The REPLACING engine works by scanning all COBOL identifiers (matching `\b[A-Z][A-Z0-9-]*\b`) in the copybook content and applying each replacement rule:
```
Original copybook content:
05 ESP-NAME PIC X(30).
05 ESP-CODE PIC X(10).
05 KPSESPL-FLAG PIC X(01).
After REPLACING LEADING "ESP-" BY "LK-ESP-" LEADING "KPSESPL" BY "LK-KPSESPL":
05 LK-ESP-NAME PIC X(30).
05 LK-ESP-CODE PIC X(10).
05 LK-KPSESPL-FLAG PIC X(01).
```
For LEADING replacements, the engine checks if each identifier starts with the `from` prefix (case-insensitive) and replaces only the prefix portion, preserving the rest of the identifier.
For TRAILING replacements, the same logic applies to suffixes.
For EXACT replacements, only identifiers that match the `from` value exactly (case-insensitive) are replaced.
## Copybook Resolution
The resolver tries multiple strategies to match a COPY target name to a copybook file:
1. **Exact match**: `COPY CPSESP` resolves to copybook named `CPSESP`
2. **Strip extension**: `COPY WORKGRID.CPY` strips `.CPY` and resolves to `WORKGRID`
3. **Add extension**: `COPY CPSESP` tries `CPSESP.CPY` and `CPSESP.COPY`
If no match is found, the COPY statement is left in place (unexpanded) and a resolution record with `resolvedPath: null` is created.
## Pipeline Integration
The expansion runs **per chunk**, after file content is read but before dispatch to worker threads:
1. All copybook files are read upfront (they are typically small, collectively under 100MB)
2. Per chunk, the copybook map is merged with chunk content (in case a chunk contains copybooks)
3. Only programs (not copybooks themselves) undergo expansion
4. The expanded content replaces the original content in-place before worker dispatch
## Source Files
- `gitnexus/src/core/ingestion/cobol-copy-expander.ts` -- `expandCopies()`, `parseReplacingClause()`, `applyReplacing()`
- `gitnexus/src/core/ingestion/pipeline.ts` -- `expandCobolCopies()`, copybook map construction, chunk integration
+265
View File
@@ -0,0 +1,265 @@
# COBOL Deep Indexing
Beyond basic symbol extraction (program name, paragraphs, CALL, PERFORM, COPY), GitNexus performs deep indexing of COBOL-specific constructs: data items, EXEC SQL/CICS blocks, file declarations, FD entries, ENTRY points, and MOVE statements.
## Data Items
### Level Numbers
| Level Range | Meaning | Graph Node Type |
|-------------|---------|-----------------|
| 01 | Record (group item) | `Record` |
| 02-49 | Elementary/group items | `Property` |
| 66 | RENAMES | `Property` |
| 77 | Independent item | `Property` |
| 88 | Condition name | `Const` |
FILLER items are skipped (no useful name for the graph).
### Clauses Parsed
The `parseDataItemClauses()` function extracts these clauses from the trailing text of a data item declaration:
| Clause | Pattern | Example |
|--------|---------|---------|
| `PIC` / `PICTURE` | `\bPIC(?:TURE)?\s+(?:IS\s+)?(\S+)` | `PIC X(30)`, `PICTURE IS 9(5)V99` |
| `USAGE` | `\bUSAGE\s+(?:IS\s+)?(COMP\|BINARY\|...)` | `USAGE IS COMP-3`, `BINARY` |
| `REDEFINES` | `\bREDEFINES\s+([A-Z][A-Z0-9-]+)` | `REDEFINES WK-DATE-NUM` |
| `OCCURS` | `\bOCCURS\s+(\d+)` | `OCCURS 12 TIMES` |
Standalone COMP variants (without the `USAGE` keyword) are also detected: `COMP`, `COMP-1` through `COMP-6`, `COMP-X`, `BINARY`, `PACKED-DECIMAL`.
### Data Hierarchy
Data items form a hierarchical structure based on level numbers. The extractor uses a **stack algorithm**:
```
Processing order:
01 WK-RECORD -> push {01, WK-RECORD} -> parent: Module
05 WK-NAME -> push {05, WK-NAME} -> parent: WK-RECORD (01 < 05)
10 WK-FIRST -> push {10, WK-FIRST} -> parent: WK-NAME (05 < 10)
10 WK-LAST -> pop WK-FIRST, push -> parent: WK-NAME (05 < 10)
05 WK-CODE -> pop WK-LAST, WK-NAME -> parent: WK-RECORD (01 < 05)
88 WK-ACTIVE -> (88 handled separately) -> parent: WK-CODE
```
The stack maintains items where each entry's level is strictly less than the next. When a new item arrives with a level <= the top of stack, items are popped until the stack top has a smaller level. A `CONTAINS` edge is created from the stack top to the new item.
For 88-level condition names, the parent is the immediately preceding non-88 data item (found by scanning backwards).
### Annotated Example
```cobol
01 WK-EMPLOYEE.
05 WK-EMP-ID PIC 9(6).
05 WK-EMP-NAME PIC X(30).
05 WK-EMP-STATUS PIC X(01).
88 WK-ACTIVE VALUE "A".
88 WK-INACTIVE VALUE "I".
05 WK-SALARY PIC 9(7)V99 COMP-3.
05 WK-DEPT PIC X(04) OCCURS 3 TIMES.
```
Produces:
- `Record` node: `WK-EMPLOYEE` (level 01, section: working-storage)
- `Property` nodes: `WK-EMP-ID`, `WK-EMP-NAME`, `WK-EMP-STATUS`, `WK-SALARY`, `WK-DEPT`
- `Const` nodes: `WK-ACTIVE` (values: `A`), `WK-INACTIVE` (values: `I`)
- `CONTAINS` edges: `WK-EMPLOYEE -> WK-EMP-ID`, `WK-EMPLOYEE -> WK-EMP-NAME`, etc.
- `CONTAINS` edges: `WK-EMP-STATUS -> WK-ACTIVE`, `WK-EMP-STATUS -> WK-INACTIVE`
### Data Item Cap
A maximum of **500 data items per file** (`MAX_DATA_ITEMS_PER_FILE`) are processed. Some COBOL programs (especially after COPY expansion) can have 10,000+ data items, which would cause graph bloat and push the V8 relationship Map past its 16.7M entry limit across thousands of files.
The cap applies after extraction: the first 500 items in source order are kept. Since 01-level records appear first, critical top-level structure is preserved.
## EXEC SQL
EXEC SQL blocks are accumulated across lines between `EXEC SQL` and `END-EXEC`, then parsed as a unit.
### Operation Classification
The first SQL keyword determines the operation:
| First Keyword | Operation |
|---------------|-----------|
| `SELECT` | SELECT |
| `INSERT` | INSERT |
| `UPDATE` | UPDATE |
| `DELETE` | DELETE |
| `DECLARE` | DECLARE |
| `OPEN` | OPEN |
| `CLOSE` | CLOSE |
| `FETCH` | FETCH |
| *(anything else)* | OTHER |
### Table Extraction
Tables are extracted from SQL clauses:
| Clause Pattern | Example |
|----------------|---------|
| `FROM <table>` | `SELECT * FROM EMPLOYEES` |
| `INTO <table>` | `INSERT INTO EMPLOYEES` |
| `UPDATE <table>` | `UPDATE EMPLOYEES SET ...` |
| `JOIN <table>` | `LEFT JOIN DEPARTMENTS ON ...` |
### Cursor Detection
```cobol
EXEC SQL
DECLARE C-EMPLOYEES CURSOR FOR
SELECT EMP-ID, EMP-NAME FROM EMPLOYEES
WHERE DEPT = :WK-DEPT
END-EXEC
```
Extracts: cursor `C-EMPLOYEES`, table `EMPLOYEES`, host variable `WK-DEPT`.
### Host Variables
Host variables are COBOL variables referenced in SQL with a `:` prefix. The colon is stripped:
```sql
WHERE EMP-ID = :WK-EMP-ID AND DEPT = :WK-DEPT
```
Extracts: `WK-EMP-ID`, `WK-DEPT`.
### Graph Output
- `CodeElement` node per table, with description `sql-table op:{OP}`
- `CodeElement` node per cursor, with description `sql-cursor`
- `ACCESSES` edge from Module to each CodeElement
- Deduplication: if the same table appears in multiple SQL blocks, only one node is created
## EXEC CICS
EXEC CICS blocks are accumulated and parsed similarly to SQL blocks.
### Command Detection
Two-word commands are detected first (matched against the block start):
```
SEND MAP, RECEIVE MAP, SEND TEXT, SEND CONTROL, READ NEXT, READ PREV
```
If no two-word command matches, the first word is used (e.g., `LINK`, `XCTL`, `RETURN`, `READ`, `WRITE`).
### Extraction
| Element | Pattern | Example |
|---------|---------|---------|
| MAP name | `MAP('name')` or `MAP("name")` | `EXEC CICS SEND MAP('EMPMENU')` |
| PROGRAM name | `PROGRAM('name')` or `PROGRAM("name")` | `EXEC CICS LINK PROGRAM('BGTABUP')` |
| TRANSID | `TRANSID('name')` or `TRANSID("name")` | `EXEC CICS START TRANSID('EMP1')` |
### Graph Output
- MAP: `CodeElement` node with description `cics-map cmd:{CMD}` + `ACCESSES` edge from Module
- PROGRAM: `CALLS` edge (cross-program call via CICS LINK/XCTL)
- TRANSID: `CodeElement` node with description `cics-transid cmd:{CMD}` + `ACCESSES` edge from Module
### Annotated Example
```cobol
EXEC CICS
SEND MAP('EMPMENU')
MAPSET('EMPSET')
FROM(WK-MAP-DATA)
ERASE
END-EXEC
```
Produces:
- `CodeElement` node: `EMPMENU` (description: `cics-map cmd:SEND MAP`)
- `ACCESSES` edge: Module -> `EMPMENU`
## File Declarations
SELECT statements in the INPUT-OUTPUT SECTION are accumulated across multiple lines (until a period terminator) and parsed for:
| Clause | Pattern | Example |
|--------|---------|---------|
| SELECT | `SELECT <name>` | `SELECT MASTER-FILE` |
| ASSIGN | `ASSIGN TO <file>` | `ASSIGN TO "MASTER.DAT"` |
| ORGANIZATION | `ORGANIZATION IS <type>` | `ORGANIZATION IS INDEXED` |
| ACCESS | `ACCESS MODE IS <mode>` | `ACCESS MODE IS DYNAMIC` |
| RECORD KEY | `RECORD KEY IS <field>` | `RECORD KEY IS WK-EMP-ID` |
| FILE STATUS | `FILE STATUS IS <field>` | `FILE STATUS IS WK-FILE-STATUS` |
### Graph Output
- `CodeElement` node with description containing all parsed clauses (e.g., `select org:INDEXED access:DYNAMIC key:WK-EMP-ID status:WK-FILE-STATUS assign:MASTER.DAT`)
- `RECORD_KEY_OF` edge: from Property node to CodeElement (confidence 0.8)
- `FILE_STATUS_OF` edge: from Property node to CodeElement (confidence 0.8)
## FD Entries
FD (File Description) entries associate a file name with its record layout:
```cobol
FD MASTER-FILE.
01 MASTER-RECORD.
05 MR-EMP-ID PIC 9(6).
05 MR-EMP-NAME PIC X(30).
```
The extractor tracks `pendingFdName` state: when an `FD` line is seen, the next 01-level data item becomes its record.
### Graph Output
- `CodeElement` node with description `fd record:{recordName}`
- `CONTAINS` edge: FD CodeElement -> Record node
- `CONTAINS` edge: SELECT CodeElement -> FD CodeElement (linking file declaration to file description)
## ENTRY Points
The `ENTRY` statement defines additional entry points into a COBOL program (in addition to the main program entry):
```cobol
ENTRY "SUBPROG" USING WK-PARAM-1 WK-PARAM-2.
```
### Graph Output
- `Constructor` node with description `entry params:{param1},{param2}` (or just `entry` if no parameters)
- `CONTAINS` edge: Module -> Constructor
- Symbol table entry (so the entry point is discoverable by name)
## PROCEDURE DIVISION USING
```cobol
PROCEDURE DIVISION USING WK-INPUT-REC WK-OUTPUT-REC.
```
The USING clause identifies parameters received by the program from its caller.
### Graph Output
- `RECEIVES` edge: Module -> Property (for each parameter name, confidence 0.8)
## MOVE Statements
MOVE statements are extracted but currently only stored in the regex results (not emitted as graph edges):
```cobol
MOVE WK-NAME TO OUT-NAME.
MOVE CORRESPONDING WK-INPUT TO WK-OUTPUT.
```
### Extraction Details
- Source and target identifiers are captured
- `CORRESPONDING` keyword is tracked (bulk field-by-field move)
- Figurative constants (SPACES, ZEROS, LOW-VALUES, HIGH-VALUES, QUOTES, ALL) are skipped
- The enclosing paragraph (`caller`) is tracked for context
DATA_FLOW edges from MOVE statements are reserved for a future release.
## Source Files
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- All extraction logic, clause parsers, EXEC block parsers
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `processCobolRegexOnly()`, graph node/edge emission
- `gitnexus/src/core/ingestion/parsing-processor.ts` -- Sequential fallback with same `MAX_DATA_ITEMS_PER_FILE` cap
+126
View File
@@ -0,0 +1,126 @@
# COBOL File Detection
GitNexus detects COBOL files through two mechanisms: extension-based mapping and directory-based override for extensionless files. This document covers both, plus the copybook/program classification logic.
## Extension Mapping
### Program Extensions
| Extension | Type |
|-----------|------|
| `.cbl` | COBOL program |
| `.cob` | COBOL program |
| `.cobol` | COBOL program |
### Copybook Extensions
| Extension | Type | Notes |
|-----------|------|-------|
| `.cpy` | Copybook | Standard |
| `.copy` | Copybook | Standard |
| `.gnm` / `.GNM` | Copybook | Enterprise (GnuCOBOL naming) |
| `.fd` / `.FD` | Copybook | File Description fragment |
| `.wrk` / `.WRK` | Copybook | Working-Storage fragment |
| `.sel` / `.SEL` | Copybook | SELECT clause fragment |
| `.open` / `.OPEN` | Copybook | File OPEN fragment |
| `.close` / `.CLOSE` | Copybook | File CLOSE fragment |
| `.ini` / `.INI` | Copybook | Initialization fragment |
| `.def` / `.DEF` | Copybook | Definition fragment |
All extension matching is case-sensitive in `getLanguageFromFilename` (the extensions above are matched as written, including uppercase variants like `.GNM`).
## Extensionless File Detection: `GITNEXUS_COBOL_DIRS`
Many enterprise COBOL repositories use extensionless files -- the filename alone identifies the program (e.g., `s/BGTABFL` is the source for program `BGTABFL`). GitNexus handles this via the `GITNEXUS_COBOL_DIRS` environment variable.
### Configuration
Set `GITNEXUS_COBOL_DIRS` to a comma-separated list of directory names:
```bash
# Files in s/, c/, and wfproc/ directories (at any depth) are treated as COBOL
export GITNEXUS_COBOL_DIRS=s,c,wfproc
```
The matching is **case-insensitive** and checks all path segments:
- `/repo/s/BGTABFL` -- matches segment `s` -- COBOL
- `/repo/src/c/CPSESP` -- matches segment `c` -- COBOL
- `/repo/wfproc/WF001` -- matches segment `wfproc` -- COBOL
- `/repo/docs/README` -- no matching segment -- skipped
### Decision Tree
```mermaid
flowchart TD
A[getLanguageFromPath] --> B[getLanguageFromFilename]
B --> C{Known extension?}
C -->|Yes .cbl/.cob/.cobol/.cpy/...| D[Return COBOL]
C -->|Yes .ts/.py/.java/...| E[Return other language]
C -->|No match| F{Has extension?}
F -->|"Has dot in basename"| G[Return null]
F -->|"No dot = extensionless"| H{GITNEXUS_COBOL_DIRS set?}
H -->|No| G
H -->|Yes| I{Any path segment<br/>matches a configured dir?}
I -->|Yes| D
I -->|No| G
style D fill:#e8f5e9,stroke:#2e7d32
style G fill:#ffebee,stroke:#c62828
```
### Implementation Detail
The `GITNEXUS_COBOL_DIRS` value is parsed once (on first call) and cached in a `Set<string>`:
```typescript
// From gitnexus/src/core/ingestion/utils.ts
const getCobolDirs = (): Set<string> => {
if (_cobolDirs) return _cobolDirs;
const raw = process.env.GITNEXUS_COBOL_DIRS;
_cobolDirs = raw
? new Set(raw.split(',').map(d => d.trim().toLowerCase()))
: new Set();
return _cobolDirs;
};
```
The path segment check splits the full path on `/` and tests each segment against the cached set.
## Copybook vs Program Classification
After a file is identified as COBOL, it must be classified as either a **program** (to be parsed for symbols) or a **copybook** (to be loaded into the copybook map for COPY expansion).
### Classification Rules
A COBOL file is classified as a **copybook** if ANY of these conditions is true:
1. It has a recognized copybook extension (`.cpy`, `.copy`, `.gnm`, `.fd`, `.wrk`, `.sel`, `.open`, `.close`, `.ini`, `.def`)
2. It is an extensionless file whose path contains a directory segment matching one of: `c`, `copy`, `copybooks`, `copylib`, `cpy`
A file is classified as a **program** if:
1. It has a program extension (`.cbl`, `.cob`, `.cobol`), OR
2. It is extensionless and does NOT match any copybook directory pattern
### Copybook Name Resolution
Copybook names are derived from the filename:
- Strip the extension (if any)
- Convert to uppercase
Examples:
- `c/CPSESP` -- name: `CPSESP`
- `copy/workgrid.cpy` -- name: `WORKGRID`
- `c/ANAZI.GNM` -- name: `ANAZI`
This name is used to resolve `COPY CPSESP.` statements during expansion.
## Source Files
- `gitnexus/src/core/ingestion/utils.ts` -- `getLanguageFromPath()`, `getLanguageFromFilename()`, `getCobolDirs()`
- `gitnexus/src/core/ingestion/pipeline.ts` -- `isCobolCopybook()`, `getCopybookName()`, `COPYBOOK_EXTENSIONS`, `COBOL_PROGRAM_EXTENSIONS`
+193
View File
@@ -0,0 +1,193 @@
# COBOL Graph Model
This document describes the graph nodes and edges that GitNexus creates for COBOL codebases. The COBOL graph model is richer than most tree-sitter languages because it captures domain-specific constructs: file declarations, FD entries, data hierarchies, SQL tables, CICS maps, and cross-program contracts.
## Entity-Relationship Diagram
```mermaid
erDiagram
File ||--o{ Module : DEFINES
File ||--o{ Function : DEFINES
File ||--o{ Namespace : DEFINES
File ||--o{ Record : DEFINES
File ||--o{ Property : DEFINES
File ||--o{ Const : DEFINES
File ||--o{ CodeElement : DEFINES
File ||--o{ Constructor : DEFINES
File }o--o{ File : IMPORTS
Module ||--o{ Record : CONTAINS
Module ||--o{ Constructor : CONTAINS
Module }o--o{ CodeElement : ACCESSES
Module }o--o{ Module : CALLS
Module }o--o{ Module : CONTRACTS
Module }o--o{ Property : RECEIVES
Record ||--o{ Property : CONTAINS
Record ||--o{ Const : CONTAINS
Record }o--o{ Record : REDEFINES
Property ||--o{ Property : CONTAINS
Property ||--o{ Const : CONTAINS
Property }o--o{ Property : REDEFINES
Property }o--o{ CodeElement : RECORD_KEY_OF
Property }o--o{ CodeElement : FILE_STATUS_OF
CodeElement ||--o{ CodeElement : CONTAINS
CodeElement ||--o{ Record : CONTAINS
Function }o--o{ Function : CALLS
```
## Node Types
| Node Type | COBOL Concept | Created From | Example |
|-----------|--------------|--------------|---------|
| `Module` | PROGRAM-ID | `PROGRAM-ID. BGTABFL` | Name: `BGTABFL`, description may include author and date |
| `Function` | Paragraph | `PROCESS-RECORD.` at column 8 | Name: `PROCESS-RECORD` |
| `Namespace` | Procedure section | `MAIN-LOGIC SECTION.` at column 8 | Name: `MAIN-LOGIC` |
| `Record` | 01-level data item | `01 WK-EMPLOYEE.` | Description: `level:01 section:working-storage` |
| `Property` | 02-49/66/77 data item | `05 WK-NAME PIC X(30).` | Description: `level:05 pic:X(30) section:working-storage` |
| `Const` | 88-level condition | `88 WK-ACTIVE VALUE "A".` | Description: `level:88 values:A` |
| `CodeElement` | SELECT, FD, SQL table, CICS map, cursor, transid | Various | Description varies by subtype |
| `Constructor` | ENTRY point | `ENTRY "SUBPROG" USING WK-DATA` | Description: `entry params:WK-DATA` |
### CodeElement Subtypes
CodeElement is used for multiple COBOL constructs, distinguished by their description prefix:
| Subtype | ID Pattern | Description Format | Example |
|---------|-----------|-------------------|---------|
| File SELECT | `CodeElement:{path}:SELECT:{name}` | `select org:INDEXED access:DYNAMIC ...` | `SELECT MASTER-FILE` |
| FD entry | `CodeElement:{path}:FD:{name}` | `fd record:{recordName}` | `FD MASTER-FILE` |
| SQL table | `CodeElement:{path}:sql-table:{name}` | `sql-table op:SELECT` | Table `EMPLOYEES` |
| SQL cursor | `CodeElement:{path}:sql-cursor:{name}` | `sql-cursor` | Cursor `C-EMPLOYEES` |
| CICS map | `CodeElement:{path}:cics-map:{name}` | `cics-map cmd:SEND MAP` | Map `EMPMENU` |
| CICS transid | `CodeElement:{path}:cics-transid:{name}` | `cics-transid cmd:START` | Transid `EMP1` |
## Edge Types
| Edge Type | Source | Target | Created By | Confidence | Example |
|-----------|--------|--------|-----------|------------|---------|
| `DEFINES` | File | any node | File defines its symbols | 1.0 | File -> Module `BGTABFL` |
| `CALLS` | Function | Function | `PERFORM X [THRU Y]` | (via call-processor) | `PROCESS-RECORD` -> `CALC-TAX` |
| `CALLS` | Module | Module | `CALL "BGTABUP"` | (via call-processor) | `BGTABFL` -> `BGTABUP` |
| `CALLS` | Module | Module | `EXEC CICS LINK PROGRAM('X')` | (via call-processor) | `BGTABFL` -> `BGTABUP` |
| `IMPORTS` | File | File | `COPY copybook` | (via import-processor) | Source file -> Copybook file |
| `CONTAINS` | Module | Record | Data hierarchy root | 1.0 | `BGTABFL` -> `WK-EMPLOYEE` |
| `CONTAINS` | Record | Property | Data hierarchy | 1.0 | `WK-EMPLOYEE` -> `WK-NAME` |
| `CONTAINS` | Property | Property | Nested data items | 1.0 | `WK-ADDRESS` -> `WK-CITY` |
| `CONTAINS` | Record/Property | Const | 88-level parent | 1.0 | `WK-STATUS` -> `WK-ACTIVE` |
| `CONTAINS` | CodeElement (FD) | Record | FD record link | 1.0 | `FD:MASTER-FILE` -> `MASTER-RECORD` |
| `CONTAINS` | CodeElement (SELECT) | CodeElement (FD) | SELECT-FD link | 0.9 | `SELECT:MASTER-FILE` -> `FD:MASTER-FILE` |
| `CONTAINS` | Module | Constructor | ENTRY in module | 1.0 | `BGTABFL` -> `SUBPROG` |
| `REDEFINES` | Record | Record | `01 X REDEFINES Y` | 1.0 | `WK-DATE-NUM` -> `WK-DATE-ALPHA` |
| `REDEFINES` | Property | Property | `05 X REDEFINES Y` | 1.0 | `WK-CODE-NUM` -> `WK-CODE-ALPHA` |
| `RECORD_KEY_OF` | Property | CodeElement (SELECT) | `RECORD KEY IS field` | 0.8 | `WK-EMP-ID` -> `SELECT:MASTER-FILE` |
| `FILE_STATUS_OF` | Property | CodeElement (SELECT) | `FILE STATUS IS field` | 0.8 | `WK-FS` -> `SELECT:MASTER-FILE` |
| `ACCESSES` | Module | CodeElement | EXEC SQL/CICS | 0.9 | `BGTABFL` -> `sql-table:EMPLOYEES` |
| `RECEIVES` | Module | Property | `PROCEDURE USING` | 0.8 | `BGTABFL` -> `WK-INPUT-REC` |
| `CONTRACTS` | Module | Module | Shared copybook detection | 0.9 | `BGTABFL` -> `BGTABUP` (via `CPSESP`) |
## Full Annotated Example
Given this COBOL program:
```cobol
IDENTIFICATION DIVISION.
PROGRAM-ID. EMPMAINT.
AUTHOR. Development Team.
ENVIRONMENT DIVISION.
INPUT-OUTPUT SECTION.
FILE-CONTROL.
SELECT EMP-FILE
ASSIGN TO "EMPLOYEE.DAT"
ORGANIZATION IS INDEXED
ACCESS MODE IS DYNAMIC
RECORD KEY IS EMP-ID
FILE STATUS IS WS-FILE-STATUS.
DATA DIVISION.
FILE SECTION.
FD EMP-FILE.
01 EMP-RECORD.
05 EMP-ID PIC 9(6).
05 EMP-NAME PIC X(30).
WORKING-STORAGE SECTION.
01 WS-FLAGS.
05 WS-FILE-STATUS PIC X(02).
05 WS-EOF-FLAG PIC X(01).
88 WS-EOF VALUE "Y".
LINKAGE SECTION.
01 LK-SEARCH-KEY PIC 9(6).
PROCEDURE DIVISION USING LK-SEARCH-KEY.
MAIN-LOGIC SECTION.
MAIN-START.
PERFORM OPEN-FILE
PERFORM PROCESS-RECORDS
PERFORM CLOSE-FILE
STOP RUN.
OPEN-FILE.
OPEN I-O EMP-FILE.
PROCESS-RECORDS.
MOVE LK-SEARCH-KEY TO EMP-ID
EXEC SQL
SELECT EMP_SALARY INTO :WS-SALARY
FROM EMPLOYEES
WHERE EMP_ID = :EMP-ID
END-EXEC
CALL "EMPREPORT".
CLOSE-FILE.
CLOSE EMP-FILE.
```
The graph produced contains:
**Nodes:**
- `Module`: EMPMAINT (description: `author:Development Team`)
- `Namespace`: MAIN-LOGIC
- `Function`: MAIN-START, OPEN-FILE, PROCESS-RECORDS, CLOSE-FILE
- `Record`: EMP-RECORD, WS-FLAGS, LK-SEARCH-KEY
- `Property`: EMP-ID, EMP-NAME, WS-FILE-STATUS, WS-EOF-FLAG
- `Const`: WS-EOF (values: Y)
- `CodeElement`: SELECT:EMP-FILE, FD:EMP-FILE, sql-table:EMPLOYEES
- (COPY imports, if any, would produce File IMPORTS edges)
**Edges:**
- `DEFINES`: File -> all nodes
- `CONTAINS`: EMPMAINT -> EMP-RECORD, EMPMAINT -> WS-FLAGS, EMPMAINT -> LK-SEARCH-KEY
- `CONTAINS`: EMP-RECORD -> EMP-ID, EMP-RECORD -> EMP-NAME
- `CONTAINS`: WS-FLAGS -> WS-FILE-STATUS, WS-FLAGS -> WS-EOF-FLAG
- `CONTAINS`: WS-EOF-FLAG -> WS-EOF
- `CONTAINS`: FD:EMP-FILE -> EMP-RECORD
- `CONTAINS`: SELECT:EMP-FILE -> FD:EMP-FILE
- `CALLS`: MAIN-START -> OPEN-FILE, MAIN-START -> PROCESS-RECORDS, MAIN-START -> CLOSE-FILE
- `CALLS`: EMPMAINT -> EMPREPORT (external CALL)
- `ACCESSES`: EMPMAINT -> sql-table:EMPLOYEES
- `RECEIVES`: EMPMAINT -> LK-SEARCH-KEY (PROCEDURE USING)
- `RECORD_KEY_OF`: EMP-ID -> SELECT:EMP-FILE
- `FILE_STATUS_OF`: WS-FILE-STATUS -> SELECT:EMP-FILE
## How COBOL Differs from Tree-Sitter Languages
| Aspect | COBOL | Tree-Sitter Languages |
|--------|-------|----------------------|
| Node variety | 8 types (Module, Function, Namespace, Record, Property, Const, CodeElement, Constructor) | Typically 4-6 (Function, Class, Method, Interface, Module, Const) |
| Domain edges | RECORD_KEY_OF, FILE_STATUS_OF, ACCESSES, RECEIVES, CONTRACTS, REDEFINES | Primarily CALLS, IMPORTS, EXTENDS, IMPLEMENTS |
| Data hierarchy | Deep CONTAINS chains (01 -> 05 -> 10 -> 88) | Flat class members |
| Cross-program calls | CALL "name" + CICS LINK PROGRAM | Import-based resolution |
| Contract detection | Shared COPY copybook between caller/callee | Not applicable |
| Metadata | AUTHOR, DATE-WRITTEN on Module | JSDoc/docstring (not indexed) |
## Source Files
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `processCobolRegexOnly()`, node/edge emission logic
- `gitnexus/src/core/ingestion/pipeline.ts` -- `detectCrossProgamContracts()` for CONTRACTS edges
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- `CobolRegexResults` interface (all extracted data)
+232
View File
@@ -0,0 +1,232 @@
# COBOL Performance and Tuning
This document covers real-world benchmarks, worker pool configuration, memory management, known limitations, and troubleshooting for COBOL indexing.
## PROJECT-NAME Benchmark
The PROJECT-NAME project is a large Italian payroll system written in COBOL. It serves as the primary benchmark for COBOL indexing performance.
### Input
| Metric | Value |
| --------------------------- | ---------------------------------------------------------------------------- |
| Paths scanned | 14,217 |
| Parseable files | 13,129 |
| Total source size | 224 MB |
| Chunks | 12 (at 20 MB budget) |
| Copybooks loaded | 2,976 |
| Copybooks used in expansion | 2,955 |
| Key directories | `s/` (7773 programs), `c/` (3036 copybooks), `wfproc/` (1973 workflow files) |
### Output
| Metric | Value |
| ---------------------- | ------ |
| Graph nodes | 2.79M |
| Graph edges | 5.67M |
| Clusters (communities) | 16,679 |
| Execution flows | 300 |
### Timing
| Phase | Duration |
| ------------------------------- | ----------------- |
| Total | ~251s |
| KuzuDB write | 132s |
| Full-text search indexing | 6.7s |
| Regex extraction (avg per file) | ~1ms |
| COPY expansion + deep indexing | Remainder (~112s) |
### Indexing Command
```bash
cd /path/to/PROJECT-NAME
GITNEXUS_COBOL_DIRS=s,c,wfproc GITNEXUS_VERBOSE=1 node --max-old-space-size=8192 \
/path/to/gitnexus/dist/cli/index.js analyze --force
```
## Worker Pool Tuning
### Sub-Batch Size
The worker pool splits each worker's chunk into sub-batches to bound peak memory per `postMessage` serialization. COBOL repos use a smaller sub-batch size than the default:
| Parameter | Default | COBOL Mode |
| --------------------- | ----------- | ------------------- |
| Sub-batch size | 1,500 files | 200 files |
| Per sub-batch timeout | 120s | 120s (configurable) |
**Why 200?** COBOL regex extraction + preprocessing takes ~1ms per file on average, but with COPY expansion and deep indexing the effective time is ~150ms per file. At sub-batch size 1500, that would be ~225s per sub-batch, exceeding the 120s timeout.
COBOL mode is activated automatically when `GITNEXUS_COBOL_DIRS` is set:
```typescript
// From pipeline.ts
const cobolSubBatch = process.env.GITNEXUS_COBOL_DIRS ? 200 : undefined;
workerPool = createWorkerPool(workerUrl, undefined, cobolSubBatch);
```
### Worker Count
Workers default to `min(8, cpus - 1)`. For COBOL repos, this is usually sufficient since regex extraction is CPU-bound but fast. The bottleneck is typically KuzuDB write, not extraction.
### Timeout Configuration
| Environment Variable | Default | Purpose |
| ------------------------------------ | --------------- | --------------------------------------------------- |
| `GITNEXUS_WORKER_TIMEOUT_MS` | 120,000 (2 min) | Per sub-batch processing timeout |
| `GITNEXUS_WORKER_STARTUP_TIMEOUT_MS` | 60,000 (1 min) | Worker initialization timeout (tree-sitter loading) |
For COBOL-only repos, worker startup is faster because tree-sitter native modules are loaded lazily (skipped entirely if only COBOL files are present).
## Data Item Cap
### Configuration
```typescript
const MAX_DATA_ITEMS_PER_FILE = 500;
```
This constant appears in both `parse-worker.ts` (worker path) and `parsing-processor.ts` (sequential fallback).
### Rationale
Some COBOL programs, especially after COPY expansion, can have 10,000+ data items. At that scale:
- The in-memory relationship Map (for CONTAINS, REDEFINES, etc.) approaches the V8 16.7M entry limit across thousands of files
- KuzuDB write time increases linearly with edge count
- Most deep-nested items (level 20+) are rarely queried individually
### Impact
The cap truncates data items beyond the 500th in source order. Since 01-level Records appear first in COBOL source, the cap preserves:
- All 01-level record definitions
- The most important 02-49 level items (those closest to the record root)
- 88-level conditions associated with early items
To increase the cap for specific needs, modify the `MAX_DATA_ITEMS_PER_FILE` constant in both files.
## Memory Management
### COPY Expansion Memory
All copybook content is loaded upfront into a Map before chunk processing begins. For PROJECT-NAME:
- 2,976 copybooks, typically under 100MB total
- The Map is shared (read-only) across chunk iterations
- Per-chunk, the copybook map is merged with chunk file content (in case a chunk contains copybooks not in the pre-loaded set)
- After all chunks are processed, the copybook map is freed (`cobolCopybookContents = undefined`)
### Chunk Budget
Source files are grouped into chunks of max 20MB (`CHUNK_BYTE_BUDGET`). Each chunk's lifecycle:
1. Read file content into memory
2. Expand COPY statements (mutates content in-place)
3. Dispatch to workers for extraction
4. Workers return serialized results
5. Merge results into graph
6. Chunk content goes out of scope (GC reclaims)
This ensures only ~20MB of source + ~200-400MB of working memory (ASTs, extracted records, serialization) is active at any time.
### Shared Warning Deduplication
The `warnedCircular` set (used by the COPY expansion engine) is shared across all files in a chunk. This prevents the same circular copybook warning (e.g., `ANAZI includes itself`) from being logged thousands of times.
## Known Limitations
| Limitation | Impact | Workaround |
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------- |
| tree-sitter-cobol hangs on ~5% of files | Cannot use tree-sitter for COBOL | Regex-only extraction (current approach) |
| Data item cap (500/file) | May miss deeply nested items in large programs | Increase `MAX_DATA_ITEMS_PER_FILE` in source |
| Circular copybooks (ANAZI, ANDIP, QDIPE) | Self-referential includes cannot be expanded | Detected and skipped with warning |
| wfproc/ files may not be pure COBOL | Workflow files may produce extraction noise | Exclude `wfproc` from `GITNEXUS_COBOL_DIRS` if problematic |
| No MOVE DATA_FLOW edges yet | Data flow between variables not in graph | Reserved for future release |
| Continuation line handling | Some complex multi-line continuations (especially in string literals spanning 3+ lines) may not merge correctly | Known edge case; affects <0.1% of lines |
| Single-line EXEC blocks | `EXEC SQL SELECT ... END-EXEC` on one line is handled, but pathological nesting is not | Extremely rare in practice |
| Extension case sensitivity | `.GNM` and `.gnm` are matched differently | Use the exact case from the codebase |
## Troubleshooting
### "COPY expansion failed"
```
[pipeline] COPY expansion failed for s/BGTABFL: Cannot read properties of null
```
**Cause:** A copybook referenced by a COPY statement cannot be found.
**Fix:**
1. Verify `GITNEXUS_COBOL_DIRS` includes the directory containing copybooks (typically `c`)
2. Check that copybook filenames match the COPY target (case-insensitive, after stripping extensions)
3. Ensure copybook files are not in `.gitignore`
### Worker sub-batch timeout
```
Worker 3 sub-batch timed out after 120s (chunk: 200 items)
```
**Cause:** A sub-batch took longer than the timeout. Typically happens when one file is extremely large (50,000+ lines after COPY expansion).
**Fix:** Increase the timeout:
```bash
GITNEXUS_WORKER_TIMEOUT_MS=300000 gitnexus analyze
```
### Memory errors (heap out of memory)
```
FATAL ERROR: CALL_AND_RETRY_LAST Allocation failed - JavaScript heap out of memory
```
**Fix:** Increase Node.js heap size:
```bash
node --max-old-space-size=16384 /path/to/gitnexus/dist/cli/index.js analyze
```
For very large repos (>500MB source), consider `--max-old-space-size=32768`.
### Concurrent analyze corruption
**Rule:** Only ONE `gitnexus analyze` process should run at a time per repository. Concurrent writes to KuzuDB corrupt the database.
If corruption occurs:
```bash
# Remove the KuzuDB directory and re-index
rm -rf .gitnexus/kuzu
gitnexus analyze --force
```
### Slow KuzuDB write phase
The KuzuDB write phase (132s for PROJECT-NAME) is the bottleneck for large COBOL repos. This is proportional to the number of nodes and edges being written. Reducing `MAX_DATA_ITEMS_PER_FILE` or excluding non-essential directories from `GITNEXUS_COBOL_DIRS` can help.
### Verbose output
Enable verbose logging to see per-phase timing and statistics:
```bash
GITNEXUS_VERBOSE=1 gitnexus analyze
```
This outputs:
- Scan statistics (paths, parseable files, chunk count)
- Worker pool configuration (worker count, sub-batch size)
- COPY expansion statistics (copybooks loaded, files expanded)
- Community and process detection results
- Contract detection results
## Source Files
- `gitnexus/src/core/ingestion/workers/worker-pool.ts` -- `DEFAULT_SUB_BATCH_SIZE`, `SUB_BATCH_TIMEOUT_MS`, `WORKER_STARTUP_TIMEOUT_MS`
- `gitnexus/src/core/ingestion/pipeline.ts` -- `CHUNK_BYTE_BUDGET`, COBOL sub-batch configuration, chunk lifecycle
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `MAX_DATA_ITEMS_PER_FILE`, `processCobolRegexOnly()`
- `gitnexus/src/core/ingestion/parsing-processor.ts` -- Sequential fallback `MAX_DATA_ITEMS_PER_FILE`
@@ -0,0 +1,186 @@
# COBOL Regex Extraction
The `extractCobolSymbolsWithRegex()` function in `cobol-preprocessor.ts` performs single-pass, state-machine-driven extraction of all COBOL symbols. This document describes the state machine, line processing flow, and every regex pattern used.
## State Machine: Division Tracking
The extractor tracks which COBOL division is currently being processed. Division transitions are detected by the `RE_DIVISION` pattern.
```mermaid
stateDiagram-v2
[*] --> null : Start of file
null --> identification : IDENTIFICATION DIVISION
identification --> environment : ENVIRONMENT DIVISION
environment --> data : DATA DIVISION
data --> procedure : PROCEDURE DIVISION
note right of identification
Extracts: PROGRAM-ID, AUTHOR, DATE-WRITTEN
end note
note right of environment
Extracts: SELECT ... ASSIGN ... (file declarations)
end note
note right of data
Extracts: FD entries, data items (01-77, 88), COPY
end note
note right of procedure
Extracts: paragraphs, sections, PERFORM, CALL,
ENTRY, MOVE, EXEC SQL/CICS
end note
```
## State Machine: Data Section Tracking
Within the DATA DIVISION, a secondary state machine tracks the current section to tag data items with their origin.
```mermaid
stateDiagram-v2
[*] --> unknown : DATA DIVISION entered
unknown --> working_storage : WORKING-STORAGE SECTION
unknown --> linkage : LINKAGE SECTION
unknown --> file : FILE SECTION
unknown --> local_storage : LOCAL-STORAGE SECTION
working_storage --> linkage : LINKAGE SECTION
working_storage --> file : FILE SECTION
linkage --> working_storage : WORKING-STORAGE SECTION
file --> working_storage : WORKING-STORAGE SECTION
file --> linkage : LINKAGE SECTION
local_storage --> working_storage : WORKING-STORAGE SECTION
```
Within the ENVIRONMENT DIVISION, the `currentEnvSection` tracks whether we are in `INPUT-OUTPUT` or `CONFIGURATION` section. SELECT statement accumulation only occurs in `INPUT-OUTPUT`.
## Line Processing Flow
Each raw source line goes through this pipeline:
```
Raw line
|
v
Length < 7? ---------> Skip (flush pending if any)
|
v
Indicator col 7
|
+-- '*' or '/' -----> Comment: skip entirely
|
+-- '-' ------------> Continuation: append to pending line
|
+-- other ----------> Normal: flush pending, strip inline comments (|),
buffer as new pending logical line
```
After all lines are processed, the final pending line is flushed, along with any accumulated SELECT statement.
### Inline Comment Stripping
Enterprise COBOL (particularly Italian dialect) uses the pipe character `|` as an inline comment marker. Everything from `|` to end of line is stripped before processing.
### Patch Marker Handling
The `preprocessCobolSource()` function (run before extraction in the worker) replaces non-standard content in columns 1-6. Standard COBOL expects spaces or digit sequence numbers in this area. If any letter or `#` character is found, the entire sequence area is replaced with 6 spaces:
```
Before: mzADD MOVE WK-AMT TO WK-TOTAL
After: MOVE WK-AMT TO WK-TOTAL
```
This preserves exact line count for position mapping.
## Regex Pattern Reference
All patterns are compiled once as module-level constants and reused across calls.
### Division and Section Detection
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_DIVISION` | `\b(IDENTIFICATION\|ENVIRONMENT\|DATA\|PROCEDURE)\s+DIVISION\b` | Division boundary | `PROCEDURE DIVISION` |
| `RE_SECTION` | `\b(WORKING-STORAGE\|LINKAGE\|FILE\|LOCAL-STORAGE\|INPUT-OUTPUT\|CONFIGURATION)\s+SECTION\b` | Section boundary | `WORKING-STORAGE SECTION` |
### IDENTIFICATION DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_PROGRAM_ID` | `\bPROGRAM-ID\.\s*([A-Z][A-Z0-9-]*)` | Program name | `PROGRAM-ID. BGTABFL` |
| `RE_AUTHOR` | `^\s+AUTHOR\.\s*(.+)` | Author metadata | `AUTHOR. D. Smith` |
| `RE_DATE_WRITTEN` | `^\s+DATE-WRITTEN\.\s*(.+)` | Date metadata | `DATE-WRITTEN. 2024-01-15` |
### ENVIRONMENT DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_SELECT_START` | `\bSELECT\s+([A-Z][A-Z0-9-]+)` | File SELECT start | `SELECT MASTER-FILE` |
SELECT statements are accumulated across multiple lines until a period terminator is found, then parsed for ASSIGN, ORGANIZATION, ACCESS, RECORD KEY, and FILE STATUS clauses.
### DATA DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_FD` | `^\s+FD\s+([A-Z][A-Z0-9-]+)` | File description | `FD MASTER-FILE` |
| `RE_DATA_ITEM` | `^\s+(\d{1,2})\s+([A-Z][A-Z0-9-]+)\s*(.*)` | Data item (01-77) | `05 WK-NAME PIC X(30)` |
| `RE_ANONYMOUS_REDEFINES` | `^\s+(\d{1,2})\s+REDEFINES\s+([A-Z][A-Z0-9-]+)` | Anonymous REDEFINES | `01 REDEFINES WK-REC` |
| `RE_88_LEVEL` | `^\s+88\s+([A-Z][A-Z0-9-]+)\s+VALUES?\s+(?:ARE\s+)?(.+)` | Condition name | `88 WK-ACTIVE VALUE "Y"` |
The trailing clauses of `RE_DATA_ITEM` are parsed by `parseDataItemClauses()` for PIC, USAGE, OCCURS, and REDEFINES.
### PROCEDURE DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_PROC_SECTION` | `^ ([A-Z][A-Z0-9-]+)\s+SECTION\.\s*$` | Procedure section header | ` MAIN-LOGIC SECTION.` |
| `RE_PROC_PARAGRAPH` | `^ ([A-Z][A-Z0-9-]+)\.\s*$` | Paragraph header | ` PROCESS-RECORD.` |
| `RE_PERFORM` | `\bPERFORM\s+([A-Z][A-Z0-9-]+)(?:\s+THRU\s+([A-Z][A-Z0-9-]+))?` | PERFORM call | `PERFORM CALC-TAX THRU CALC-TAX-EXIT` |
| `RE_PROC_USING` | `\bPROCEDURE\s+DIVISION\s+USING\s+([\s\S]*?)(?:\.\|$)` | USING parameters | `PROCEDURE DIVISION USING WK-PARAM` |
| `RE_ENTRY` | `\bENTRY\s+"([^"]+)"(?:\s+USING\s+([\s\S]*?))?(?:\.\|$)` | ENTRY point | `ENTRY "SUBPROG" USING WK-DATA` |
| `RE_MOVE` | `\bMOVE\s+(CORRESPONDING\s+)?([A-Z][A-Z0-9-]+)\s+TO\s+([A-Z][A-Z0-9-]+)` | MOVE statement | `MOVE WK-NAME TO OUT-NAME` |
Note: `RE_PROC_SECTION` and `RE_PROC_PARAGRAPH` require exactly 7 spaces of leading indentation (COBOL area A starting at column 8). This is the standard COBOL paragraph indentation.
### All-Division Patterns
These patterns are checked regardless of current division:
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_CALL` | `\bCALL\s+"([^"]+)"` | External program call | `CALL "BGTABUP"` |
| `RE_COPY_UNQUOTED` | `\bCOPY\s+([A-Z][A-Z0-9-]+)(?:\s\|\.)` | COPY (unquoted) | `COPY CPSESP.` |
| `RE_COPY_QUOTED` | `\bCOPY\s+"([^"]+)"(?:\s\|\.)` | COPY (quoted) | `COPY "WORKGRID.CPY".` |
### EXEC Block Patterns
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_EXEC_SQL_START` | `\bEXEC\s+SQL\b` | Start of EXEC SQL block | `EXEC SQL` |
| `RE_EXEC_CICS_START` | `\bEXEC\s+CICS\b` | Start of EXEC CICS block | `EXEC CICS` |
| `RE_END_EXEC` | `\bEND-EXEC\b` | End of EXEC block | `END-EXEC` |
EXEC blocks accumulate all lines between `EXEC SQL/CICS` and `END-EXEC`, then delegate to `parseExecSqlBlock()` or `parseExecCicsBlock()` for detailed extraction.
## Excluded Paragraph Names
The following names are excluded from paragraph detection to avoid false positives from division/section headers:
```
DECLARATIVES, END, PROCEDURE, IDENTIFICATION,
ENVIRONMENT, DATA, WORKING-STORAGE, LINKAGE,
FILE, LOCAL-STORAGE, COMMUNICATION, REPORT,
SCREEN, INPUT-OUTPUT, CONFIGURATION
```
Additionally, paragraph candidates containing `DIVISION` or `SECTION` as substrings are excluded.
## MOVE Skip List (Figurative Constants)
MOVE statements where the source is a figurative constant are skipped:
```
SPACES, ZEROS, ZEROES, LOW-VALUES, LOW-VALUE,
HIGH-VALUES, HIGH-VALUE, QUOTES, QUOTE, ALL
```
## Source Files
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- `preprocessCobolSource()`, `extractCobolSymbolsWithRegex()`, all regex constants
+127
View File
@@ -0,0 +1,127 @@
import { test, expect, type TestInfo } from '@playwright/test';
/**
* Debug harnesses for investigating specific UI issues.
* Excluded from `npm run test:e2e` via testIgnore in playwright.config.ts.
* Run directly: DEBUG_E2E=1 npx playwright test e2e/debug-issues.spec.ts
*/
const BACKEND_URL = process.env.BACKEND_URL ?? 'http://localhost:4747';
const debugTest = process.env.DEBUG_E2E ? test : test.skip;
async function connectToServer(page: import('@playwright/test').Page) {
page.on('console', msg => {
if (msg.type() === 'error') console.log(`[error] ${msg.text()}`);
});
await page.goto('/');
await page.getByText('Server').click();
const serverInput = page.locator('input[name="server-url-input"]');
await serverInput.fill(BACKEND_URL);
await page.getByRole('button', { name: /Connect/ }).click();
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
// Wait for LadybugDB to finish loading — poll isDatabaseReady via process list visibility
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
}
debugTest('debug: process view Reset View button', async ({ page }, testInfo) => {
await connectToServer(page);
// Open Processes tab
await page.getByRole('button', { name: 'Nexus AI' }).click();
await page.getByText('Processes').click();
await expect(page.getByText(/\d+ processes detected/)).toBeVisible({ timeout: 10_000 });
// Click View on the first Cross-Community process
// The View button has opacity-0 by default, use JS click to bypass
const viewButtons = page.locator('button:has-text("View")');
const count = await viewButtons.count();
console.log(`Found ${count} View buttons`);
// Use evaluate to click the first one regardless of visibility
await page.evaluate(() => {
const btns = document.querySelectorAll('button');
for (const btn of btns) {
if (btn.textContent?.trim() === 'View') {
btn.click();
return;
}
}
});
// Wait for modal to appear
const modal = page.locator('[data-testid="process-modal"]');
await expect(modal).toBeVisible({ timeout: 5_000 });
// Screenshot: modal should be open with flowchart
await page.screenshot({ path: testInfo.outputPath('debug-modal-open.png'), fullPage: true });
// Get the diagram's current transform
const diagramDiv = modal.locator('[style*="transform"]');
const transformBefore = await diagramDiv.getAttribute('style');
console.log('Transform BEFORE zoom:', transformBefore);
// Zoom in using the + button
const zoomInBtn = modal.getByRole('button', { name: /Zoom in/ });
await zoomInBtn.click();
await zoomInBtn.click();
await zoomInBtn.click();
// Wait for zoom animation to settle
await expect(async () => {
const t = await diagramDiv.getAttribute('style');
expect(t).not.toBe(transformBefore);
}).toPass({ timeout: 2_000 });
const transformAfterZoom = await diagramDiv.getAttribute('style');
console.log('Transform AFTER zoom:', transformAfterZoom);
await page.screenshot({ path: testInfo.outputPath('debug-modal-zoomed.png'), fullPage: true });
// Click Reset View
const resetBtn = modal.getByRole('button', { name: 'Reset View' });
await resetBtn.click();
// Wait for reset animation to settle
await expect(async () => {
const t = await diagramDiv.getAttribute('style');
expect(t).toBe(transformBefore);
}).toPass({ timeout: 2_000 });
const transformAfterReset = await diagramDiv.getAttribute('style');
console.log('Transform AFTER reset:', transformAfterReset);
await page.screenshot({ path: testInfo.outputPath('debug-modal-after-reset.png'), fullPage: true });
// Verify transform actually changed back
expect(transformAfterZoom).not.toBe(transformBefore);
expect(transformAfterReset).toBe(transformBefore);
});
debugTest('debug: lightbulb clears node selection dimming', async ({ page }, testInfo) => {
await connectToServer(page);
// Wait for graph canvas to render
await expect(page.locator('canvas').first()).toBeVisible({ timeout: 10_000 });
await page.screenshot({ path: testInfo.outputPath('debug-before-select.png'), fullPage: true });
// Click a file in the tree to select a node (causes dimming)
const fileItem = page.getByText('start.sh');
await fileItem.click();
await page.waitForTimeout(500);
await page.screenshot({ path: testInfo.outputPath('debug-node-selected.png'), fullPage: true });
// Check the lightbulb button state
const lightbulbBtn = page.locator('button[title*="Turn off"], button[title*="Turn on"]');
const title = await lightbulbBtn.getAttribute('title');
console.log('Lightbulb title before click:', title);
// Click the lightbulb
await lightbulbBtn.click();
await page.waitForTimeout(500);
const titleAfter = await lightbulbBtn.getAttribute('title');
console.log('Lightbulb title after click:', titleAfter);
await page.screenshot({ path: testInfo.outputPath('debug-after-lightbulb.png'), fullPage: true });
// Click it again to toggle back on
await lightbulbBtn.click();
await page.waitForTimeout(500);
await page.screenshot({ path: testInfo.outputPath('debug-after-lightbulb-toggle-back.png'), fullPage: true });
});
+28
View File
@@ -0,0 +1,28 @@
import { test } from '@playwright/test';
/**
* Manual recording session for interactive debugging.
* Opens the app and pauses so you can interact with the UI.
* Trace, video, and screenshots are saved automatically on close.
*
* Run with: npx playwright test e2e/manual-record.spec.ts --headed --timeout=0
*
* Excluded from `npm run test:e2e` via testIgnore in playwright.config.ts.
* Also skipped when PWDEBUG is not set or in CI, as a safety net.
*/
test.skip(
!!process.env.CI || process.env.PWDEBUG !== '1',
'Manual recording requires --headed and PWDEBUG=1. Run: PWDEBUG=1 npx playwright test e2e/manual-record.spec.ts --headed --timeout=0'
);
test('manual recording session', async ({ page }) => {
page.on('console', msg => {
if (msg.type() === 'error' || msg.type() === 'warning') {
console.log(`[${msg.type()}] ${msg.text()}`);
}
});
page.on('pageerror', err => console.log(`[crash] ${err.message}`));
await page.goto('http://localhost:5173');
await page.pause();
});
+165
View File
@@ -0,0 +1,165 @@
import { test, expect, type TestInfo } from '@playwright/test';
/**
* E2E tests for the GitNexus web UI.
* Requires:
* - gitnexus serve running on localhost:4747
* - gitnexus-web dev server running on localhost:5173
*
* Skipped when servers aren't available (CI without services, etc.).
* Set E2E=1 to force-run even without the availability check.
*/
const BACKEND_URL = process.env.BACKEND_URL ?? 'http://localhost:4747';
const FRONTEND_URL = process.env.FRONTEND_URL ?? 'http://localhost:5173';
// Skip all tests if the gitnexus server or Vite dev server isn't reachable
test.beforeAll(async () => {
if (process.env.E2E) return; // force-run
try {
const [backendRes, frontendRes] = await Promise.allSettled([
fetch(`${BACKEND_URL}/api/repos`),
fetch(FRONTEND_URL),
]);
if (backendRes.status === 'rejected' || (backendRes.status === 'fulfilled' && !backendRes.value.ok)) {
test.skip(true, 'gitnexus serve not available on :4747');
return;
}
if (frontendRes.status === 'rejected' || (frontendRes.status === 'fulfilled' && !frontendRes.value.ok)) {
test.skip(true, 'Vite dev server not available on :5173');
return;
}
} catch {
test.skip(true, 'servers not available');
}
});
/** Shared helper: connect to the local server and wait for the graph to load */
async function connectAndWaitForGraph(page: import('@playwright/test').Page, testInfo: TestInfo) {
// Signal to the app that we are running under Playwright (used to skip heavy Ladybug loads).
await page.addInitScript(() => {
(window as unknown as { __PLAYWRIGHT_TEST__?: boolean }).__PLAYWRIGHT_TEST__ = true;
});
await page.goto('/');
// Wait for the app to fully render before interacting
const serverTab = page.getByRole('button', { name: 'Server' });
await expect(serverTab).toBeVisible({ timeout: 15_000 });
await page.screenshot({ path: testInfo.outputPath('step-1-landing.png') });
// Click "Server" tab and wait for the input to appear
await serverTab.click();
const serverInput = page.locator('input[name="server-url-input"]');
await expect(serverInput).toBeVisible({ timeout: 15_000 });
await page.screenshot({ path: testInfo.outputPath('step-2-server-tab.png') });
await serverInput.fill(BACKEND_URL);
await page.screenshot({ path: testInfo.outputPath('step-3-url-filled.png') });
await page.getByRole('button', { name: /Connect/ }).click();
// Wait for graph to load — status bar shows "Ready"
await expect(page.locator('[data-testid="status-ready"]')).toBeVisible({ timeout: 30_000 });
await expect(page.getByText(/\d+ nodes/).first()).toBeVisible();
await page.screenshot({ path: testInfo.outputPath('step-4-graph-loaded.png') });
}
test.describe('Server Connection & Graph Loading', () => {
test('connects to server and loads graph', async ({ page }, testInfo) => {
await connectAndWaitForGraph(page, testInfo);
await page.screenshot({ path: testInfo.outputPath('graph-loaded.png'), fullPage: true });
});
});
test.describe('Nexus AI', () => {
test('panel opens and agent initializes without error', async ({ page }, testInfo) => {
await connectAndWaitForGraph(page, testInfo);
// Click Nexus AI button to open the panel
await page.getByRole('button', { name: 'Nexus AI' }).click();
// Should see the Nexus AI tab content
await expect(page.getByText('Ask me anything')).toBeVisible({ timeout: 15_000 });
await page.screenshot({ path: testInfo.outputPath('nexus-ai-panel.png'), fullPage: true });
// "Database not ready" should NOT be visible
const errorBanner = page.getByText('Database not ready');
expect(await errorBanner.isVisible().catch(() => false)).toBe(false);
});
});
test.describe('Processes Panel', () => {
test('shows process list and View button works', async ({ page }, testInfo) => {
await connectAndWaitForGraph(page, testInfo);
// Open Nexus AI panel, switch to Processes tab
await page.getByRole('button', { name: 'Nexus AI' }).click();
await page.getByText('Processes').click();
// Should show process count — wait for data-testid instead of fixed timeout
await expect(page.locator('[data-testid="process-list-loaded"]')).toBeVisible({ timeout: 15_000 });
await page.screenshot({ path: testInfo.outputPath('processes-panel.png'), fullPage: true });
// Hover first process item to reveal View button, then click it
const processRow = page.locator('[data-testid="process-row"]').first();
await expect(processRow).toBeVisible({ timeout: 10_000 });
await processRow.hover();
const viewBtn = processRow.locator('[data-testid="process-view-button"]');
await viewBtn.waitFor({ state: 'visible', timeout: 5_000 });
await viewBtn.click();
// Wait for modal to appear
await expect(page.locator('[data-testid="process-modal"]')).toBeVisible({ timeout: 5_000 });
await page.screenshot({ path: testInfo.outputPath('process-view-clicked.png'), fullPage: true });
});
test('lightbulb highlights nodes in graph', async ({ page }, testInfo) => {
await connectAndWaitForGraph(page, testInfo);
await page.getByRole('button', { name: 'Nexus AI' }).click();
await page.getByText('Processes').click();
await expect(page.locator('[data-testid="process-list-loaded"]')).toBeVisible({ timeout: 15_000 });
await page.screenshot({ path: testInfo.outputPath('before-highlight.png'), fullPage: true });
// Hover first process to reveal lightbulb
const processRow = page.locator('[data-testid="process-row"]').first();
await expect(processRow).toBeVisible({ timeout: 10_000 });
await processRow.hover();
const lightbulb = processRow.locator('[data-testid="process-highlight-button"]');
await lightbulb.waitFor({ state: 'visible', timeout: 5_000 });
await lightbulb.click();
// Wait for highlight to apply — the process row gets amber styling when focused
await expect(processRow).toHaveClass(/bg-amber-950/, { timeout: 5_000 });
await page.screenshot({ path: testInfo.outputPath('after-highlight.png'), fullPage: true });
});
});
test.describe('Turn Off All Highlights', () => {
test('selecting a node dims others, button clears it', async ({ page }, testInfo) => {
await connectAndWaitForGraph(page, testInfo);
// Wait for graph to fully render by checking for canvas element
await expect(page.locator('canvas').first()).toBeVisible({ timeout: 10_000 });
await page.screenshot({ path: testInfo.outputPath('before-select.png'), fullPage: true });
// Click a file in the file tree to select a node
const fileItem = page.getByText('package.json').first();
await expect(fileItem).toBeVisible({ timeout: 10_000 });
await fileItem.click();
// Wait for highlight toggle to show "Turn off" (indicates highlights are active)
const highlightToggle = page.locator('[data-testid="ai-highlights-toggle"]');
await expect(highlightToggle).toHaveAttribute('title', 'Turn off all highlights', { timeout: 5_000 });
await page.screenshot({ path: testInfo.outputPath('node-selected.png'), fullPage: true });
// Click the toggle to clear all highlights
await highlightToggle.click();
// Verify highlights are now off — button title changes to "Turn on"
await expect(highlightToggle).toHaveAttribute('title', 'Turn on AI highlights', { timeout: 5_000 });
await page.screenshot({ path: testInfo.outputPath('highlights-cleared.png'), fullPage: true });
});
});
+1882 -838
View File
File diff suppressed because it is too large Load Diff
+23 -6
View File
@@ -2,16 +2,25 @@
"name": "gitnexus",
"private": true,
"version": "0.0.0",
"engines": {
"node": ">=20.0.0"
},
"type": "module",
"scripts": {
"dev": "vite",
"build": "tsc -b && vite build",
"preview": "vite preview",
"test": "vitest run"
"test": "vitest run",
"test:watch": "vitest",
"test:coverage": "vitest run --coverage",
"test:e2e": "playwright test",
"test:e2e:ui": "playwright test --ui",
"test:e2e:report": "playwright show-report"
},
"dependencies": {
"@huggingface/transformers": "^3.0.0",
"@isomorphic-git/lightning-fs": "^4.6.2",
"@ladybugdb/wasm-core": "^0.15.1",
"@langchain/anthropic": "^1.3.10",
"@langchain/core": "^1.1.15",
"@langchain/google-genai": "^2.1.10",
@@ -24,22 +33,22 @@
"buffer": "^6.0.3",
"comlink": "^4.4.2",
"d3": "^7.9.0",
"dompurify": "^3.3.3",
"graphology": "^0.26.0",
"graphology-indices": "^0.17.0",
"graphology-utils": "^2.3.0",
"mnemonist": "^0.39.0",
"pandemonium": "^2.4.0",
"graphology-layout-force": "^0.2.4",
"graphology-layout-forceatlas2": "^0.10.1",
"graphology-layout-noverlap": "^0.4.2",
"graphology-utils": "^2.3.0",
"isomorphic-git": "^1.36.1",
"jszip": "^3.10.1",
"@ladybugdb/wasm-core": "^0.15.2",
"langchain": "^1.2.10",
"lru-cache": "^11.2.4",
"lucide-react": "^0.562.0",
"mermaid": "^11.12.2",
"minisearch": "^7.2.0",
"mnemonist": "^0.39.0",
"pandemonium": "^2.4.0",
"react": "^18.3.1",
"react-dom": "^18.3.1",
"react-markdown": "^10.1.0",
@@ -56,6 +65,11 @@
},
"devDependencies": {
"@babel/types": "^7.28.5",
"@playwright/test": "^1.58.2",
"@testing-library/jest-dom": "^6.9.1",
"@testing-library/react": "^16.3.2",
"@testing-library/user-event": "^14.6.1",
"@types/dompurify": "^3.0.5",
"@types/jszip": "^3.4.0",
"@types/node": "^24.10.1",
"@types/react": "^18.3.5",
@@ -63,10 +77,13 @@
"@types/react-syntax-highlighter": "^15.5.13",
"@vercel/node": "^5.5.16",
"@vitejs/plugin-react": "^5.1.0",
"@vitest/coverage-v8": "^3.2.4",
"jsdom": "^29.0.0",
"tree-sitter-wasms": "^0.1.13",
"typescript": "^5.4.5",
"vite": "^5.2.0",
"vite-plugin-static-copy": "^3.1.4",
"vitest": "^4.0.18"
"vitest": "^3.2.4",
"wait-on": "^8.0.5"
}
}
+48
View File
@@ -0,0 +1,48 @@
import { defineConfig } from '@playwright/test';
// Enable insecure browser config (disabled security + CSP bypass) only when explicitly requested.
// Example: PLAYWRIGHT_INSECURE=1 npx playwright test
const insecureE2E = process.env.PLAYWRIGHT_INSECURE === '1';
// Base launch args: always enable software WebGL for sigma.js graph rendering in headless mode.
const launchArgs = [
'--use-gl=angle',
'--use-angle=swiftshader',
'--enable-webgl',
'--enable-unsafe-swiftshader',
];
if (insecureE2E) {
// Allow cross-origin requests to gitnexus serve on a different port when explicitly enabled.
launchArgs.unshift('--disable-web-security', '--disable-site-isolation-trials');
}
export default defineConfig({
testDir: './e2e',
testIgnore: ['**/manual-record.spec.ts', '**/debug-issues.spec.ts'],
timeout: 60_000,
retries: process.env.CI ? 1 : 0,
use: {
baseURL: 'http://localhost:5173',
trace: 'retain-on-failure',
screenshot: 'retain-on-failure',
video: 'retain-on-failure',
launchOptions: {
args: launchArgs,
},
// Vite dev server sets COEP require-corp for SharedArrayBuffer (LadybugDB WASM).
// Only bypass CSP when explicitly running in insecure E2E mode.
bypassCSP: insecureE2E,
},
projects: [
{
name: 'chromium',
use: { browserName: 'chromium' },
},
],
reporter: [
['list'],
['html', { open: 'never', outputFolder: 'playwright-report' }],
],
outputDir: 'test-results',
});
+19 -43
View File
@@ -13,7 +13,7 @@ import { FileEntry } from './services/zip';
import { getActiveProviderConfig } from './core/llm/settings-service';
import { createKnowledgeGraph } from './core/graph/graph';
import { connectToServer, fetchRepos, normalizeServerUrl, type ConnectToServerResult } from './services/server-connection';
import { HelpPanel } from './components/HelpPanel';
import { ERROR_RESET_DELAY_MS } from './config/ui-constants';
const AppContent = () => {
const {
@@ -29,11 +29,9 @@ const AppContent = () => {
runPipelineFromFiles,
isSettingsPanelOpen,
setSettingsPanelOpen,
isHelpDialogBoxOpen,
setHelpDialogBoxOpen,
refreshLLMSettings,
initializeAgent,
startEmbeddings,
startEmbeddingsWithFallback,
embeddingStatus,
codeReferences,
selectedNode,
@@ -44,7 +42,6 @@ const AppContent = () => {
setAvailableRepos,
switchRepo,
loadServerGraph,
graph
} = useAppState();
const graphCanvasRef = useRef<GraphCanvasHandle>(null);
@@ -72,13 +69,7 @@ const AppContent = () => {
// Auto-start embeddings pipeline in background
// Uses WebGPU if available, falls back to WASM
startEmbeddings().catch((err) => {
if (err?.name === 'WebGPUNotAvailableError' || err?.message?.includes('WebGPU')) {
startEmbeddings('wasm').catch(console.warn);
} else {
console.warn('Embeddings auto-start failed:', err);
}
});
startEmbeddingsWithFallback();
} catch (error) {
console.error('Pipeline error:', error);
setProgress({
@@ -90,13 +81,16 @@ const AppContent = () => {
setTimeout(() => {
setViewMode('onboarding');
setProgress(null);
}, 3000);
}, ERROR_RESET_DELAY_MS);
}
}, [setViewMode, setGraph, setFileContents, setProgress, setProjectName, runPipeline, startEmbeddings, initializeAgent]);
}, [setViewMode, setGraph, setFileContents, setProgress, setProjectName, runPipeline, startEmbeddingsWithFallback, initializeAgent]);
const handleGitClone = useCallback(async (files: FileEntry[]) => {
const firstPath = files[0]?.path || 'repository';
const projectName = firstPath.split('/')[0].replace(/-\d+$/, '') || 'repository';
const handleGitClone = useCallback(async (files: FileEntry[], repoName?: string) => {
let projectName = repoName;
if (!projectName) {
const firstPath = files[0]?.path || 'repository';
projectName = firstPath.split('/')[0].replace(/-\d+$/, '') || 'repository';
}
setProjectName(projectName);
setProgress({ phase: 'extracting', percent: 0, message: 'Starting...', detail: 'Preparing to process files' });
@@ -115,13 +109,7 @@ const AppContent = () => {
initializeAgent(projectName);
}
startEmbeddings().catch((err) => {
if (err?.name === 'WebGPUNotAvailableError' || err?.message?.includes('WebGPU')) {
startEmbeddings('wasm').catch(console.warn);
} else {
console.warn('Embeddings auto-start failed:', err);
}
});
startEmbeddingsWithFallback();
} catch (error) {
console.error('Pipeline error:', error);
setProgress({
@@ -133,14 +121,15 @@ const AppContent = () => {
setTimeout(() => {
setViewMode('onboarding');
setProgress(null);
}, 3000);
}, ERROR_RESET_DELAY_MS);
}
}, [setViewMode, setGraph, setFileContents, setProgress, setProjectName, runPipelineFromFiles, startEmbeddings, initializeAgent]);
}, [setViewMode, setGraph, setFileContents, setProgress, setProjectName, runPipelineFromFiles, startEmbeddingsWithFallback, initializeAgent]);
const handleServerConnect = useCallback((result: ConnectToServerResult): Promise<void> => {
// Extract project name from repoPath
const repoPath = result.repoInfo.repoPath;
const projectName = repoPath.split('/').pop() || 'server-project';
const parts = repoPath.split('/').filter(p => p && !p.startsWith('.'));
const projectName = parts[parts.length - 1] || parts[0] || 'server-project';
setProjectName(projectName);
// Build KnowledgeGraph from server data for visualization
@@ -172,13 +161,7 @@ const AppContent = () => {
}
})
.then(() => {
startEmbeddings().catch((err) => {
if (err?.name === 'WebGPUNotAvailableError' || err?.message?.includes('WebGPU')) {
startEmbeddings('wasm').catch(console.warn);
} else {
console.warn('Embeddings auto-start failed:', err);
}
});
startEmbeddingsWithFallback();
})
.catch((err) => {
console.warn('Failed to load graph into LadybugDB:', err);
@@ -186,7 +169,7 @@ const AppContent = () => {
});
return loadGraphPromise;
}, [setViewMode, setGraph, setFileContents, setProjectName, loadServerGraph, initializeAgent, startEmbeddings]);
}, [setViewMode, setGraph, setFileContents, setProjectName, loadServerGraph, initializeAgent, startEmbeddingsWithFallback]);
// Auto-connect when ?server query param is present (bookmarkable shortcut)
const autoConnectRan = useRef(false);
@@ -235,7 +218,7 @@ const AppContent = () => {
setTimeout(() => {
setViewMode('onboarding');
setProgress(null);
}, 3000);
}, ERROR_RESET_DELAY_MS);
});
}, [handleServerConnect, setProgress, setViewMode, setServerBaseUrl, setAvailableRepos]);
@@ -309,13 +292,6 @@ const AppContent = () => {
onSettingsSaved={handleSettingsSaved}
/>
<HelpPanel
isOpen={isHelpDialogBoxOpen}
onClose={() => setHelpDialogBoxOpen(false)}
nodeCount={graph!.nodes.length}
edgeCount={graph!.relationships.length}
/>
</div>
);
};
@@ -1,4 +1,4 @@
import { Server, ArrowRight } from 'lucide-react';
import { Server, ArrowRight } from '@/lib/lucide-icons';
import { BackendRepo } from '../services/backend';
interface BackendRepoSelectorProps {
@@ -1,8 +1,9 @@
import { useCallback, useEffect, useMemo, useRef, useState } from 'react';
import { Code, PanelLeftClose, PanelLeft, Trash2, X, Target, FileCode, Sparkles, MousePointerClick } from 'lucide-react';
import { Code, PanelLeftClose, PanelLeft, Trash2, X, Target, FileCode, Sparkles, MousePointerClick } from '@/lib/lucide-icons';
import { Prism as SyntaxHighlighter } from 'react-syntax-highlighter';
import { vscDarkPlus } from 'react-syntax-highlighter/dist/esm/styles/prism';
import { useAppState } from '../hooks/useAppState';
import type { GraphNode } from '../core/graph/types';
import { NODE_COLORS } from '../lib/constants';
/** Map file extension to Prism syntax highlighter language identifier */
@@ -75,6 +76,11 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
codeReferenceFocus,
} = useAppState();
const nodeById = useMemo(() => {
if (!graph) return new Map<string, GraphNode>();
return new Map(graph.nodes.map(n => [n.id, n]));
}, [graph]);
const [isCollapsed, setIsCollapsed] = useState(false);
const [glowRefId, setGlowRefId] = useState<string | null>(null);
const panelRef = useRef<HTMLElement | null>(null);
@@ -161,8 +167,9 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
if (!target) return;
// Double rAF: wait for collapse state + list DOM to render.
requestAnimationFrame(() => {
requestAnimationFrame(() => {
const rafIds: number[] = [];
const outerRafId = requestAnimationFrame(() => {
const innerRafId = requestAnimationFrame(() => {
const el = refCardEls.current.get(target.id);
if (!el) return;
@@ -177,7 +184,13 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
glowTimerRef.current = null;
}, 1200);
});
rafIds.push(innerRafId);
});
rafIds.push(outerRafId);
return () => {
rafIds.forEach(id => cancelAnimationFrame(id));
};
}, [codeReferenceFocus?.ts, aiReferences]);
const refsWithSnippets = useMemo(() => {
@@ -414,7 +427,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
const nodeId = ref.nodeId!;
// Sync selection + focus graph
if (graph) {
const node = graph.nodes.find((n) => n.id === nodeId);
const node = nodeById.get(nodeId);
if (node) setSelectedNode(node);
}
onFocusNode(nodeId);
+15 -7
View File
@@ -1,12 +1,12 @@
import { useState, useCallback, useRef, DragEvent } from 'react';
import { Upload, FileArchive, Github, Loader2, ArrowRight, Key, Eye, EyeOff, Globe, X } from 'lucide-react';
import { Upload, FileArchive, Github, Loader2, ArrowRight, Key, Eye, EyeOff, Globe, X } from '@/lib/lucide-icons';
import { cloneRepository, parseGitHubUrl } from '../services/git-clone';
import { connectToServer, type ConnectToServerResult } from '../services/server-connection';
import { FileEntry } from '../services/zip';
interface DropZoneProps {
onFileSelect: (file: File) => void;
onGitClone?: (files: FileEntry[]) => void;
onGitClone?: (files: FileEntry[], repoName?: string) => void;
onServerConnect?: (result: ConnectToServerResult, serverUrl?: string) => void;
}
@@ -27,9 +27,13 @@ export const DropZone = ({ onFileSelect, onGitClone, onServerConnect }: DropZone
const [error, setError] = useState<string | null>(null);
// Server tab state
const [serverUrl, setServerUrl] = useState(() =>
localStorage.getItem('gitnexus-server-url') || ''
);
const [serverUrl, setServerUrl] = useState(() => {
try {
return localStorage.getItem('gitnexus-server-url') || '';
} catch {
return '';
}
});
const [isConnecting, setIsConnecting] = useState(false);
const [serverProgress, setServerProgress] = useState<{
phase: string;
@@ -104,7 +108,7 @@ export const DropZone = ({ onFileSelect, onGitClone, onServerConnect }: DropZone
setGithubToken('');
if (onGitClone) {
onGitClone(files);
onGitClone(files, parsed.repo);
}
} catch (err) {
console.error('Clone failed:', err);
@@ -133,7 +137,11 @@ export const DropZone = ({ onFileSelect, onGitClone, onServerConnect }: DropZone
}
// Persist URL to localStorage
localStorage.setItem('gitnexus-server-url', serverUrl);
try {
localStorage.setItem('gitnexus-server-url', serverUrl);
} catch {
// localStorage may be unavailable (e.g. private browsing, quota exceeded)
}
setError(null);
setIsConnecting(true);
@@ -1,4 +1,4 @@
import { Brain, Loader2, Check, AlertCircle, Zap, FlaskConical } from 'lucide-react';
import { Brain, Loader2, Check, AlertCircle, Zap, FlaskConical } from '@/lib/lucide-icons';
import { useAppState } from '../hooks/useAppState';
import { useState } from 'react';
import { WebGPUFallbackDialog } from './WebGPUFallbackDialog';
@@ -14,7 +14,7 @@ import {
Variable,
Hash,
Target,
} from 'lucide-react';
} from '@/lib/lucide-icons';
import { useAppState } from '../hooks/useAppState';
import { FILTERABLE_LABELS, NODE_COLORS, ALL_EDGE_TYPES, EDGE_INFO, type EdgeType } from '../lib/constants';
import { GraphNode, NodeLabel } from '../core/graph/types';
@@ -98,13 +98,15 @@ const TreeItem = ({
const isSelected = selectedPath === node.path;
const hasChildren = node.children.length > 0;
// Filter children based on search
// Filter children based on search (recursive)
const filteredChildren = useMemo(() => {
if (!searchQuery) return node.children;
return node.children.filter(child =>
child.name.toLowerCase().includes(searchQuery.toLowerCase()) ||
child.children.some(c => c.name.toLowerCase().includes(searchQuery.toLowerCase()))
);
const searchLower = searchQuery.toLowerCase();
const matchesSearch = (node: TreeNode, query: string): boolean => {
if (node.name.toLowerCase().includes(query)) return true;
return node.children?.some(child => matchesSearch(child, query)) ?? false;
};
return node.children.filter(child => matchesSearch(child, searchLower));
}, [node.children, searchQuery]);
// Check if this node matches search
+34 -26
View File
@@ -1,8 +1,9 @@
import { useEffect, useCallback, useMemo, useState, forwardRef, useImperativeHandle } from 'react';
import { ZoomIn, ZoomOut, Maximize2, Focus, RotateCcw, Play, Pause, Lightbulb, LightbulbOff } from 'lucide-react';
import { ZoomIn, ZoomOut, Maximize2, Focus, RotateCcw, Play, Pause, Lightbulb, LightbulbOff } from '@/lib/lucide-icons';
import { useSigma } from '../hooks/useSigma';
import { useAppState } from '../hooks/useAppState';
import { knowledgeGraphToGraphology, filterGraphByDepth, SigmaNodeAttributes, SigmaEdgeAttributes } from '../lib/graph-adapter';
import type { GraphNode } from '../core/graph/types';
import { QueryFAB } from './QueryFAB';
import Graph from 'graphology';
@@ -54,30 +55,44 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
return animatedNodes;
}, [animatedNodes, isAIHighlightsEnabled]);
const nodeById = useMemo(() => {
if (!graph) return new Map<string, GraphNode>();
return new Map(graph.nodes.map(n => [n.id, n]));
}, [graph]);
const handleNodeClick = useCallback((nodeId: string) => {
if (!graph) return;
const node = graph.nodes.find(n => n.id === nodeId);
const node = nodeById.get(nodeId);
if (node) {
setSelectedNode(node);
openCodePanel();
}
}, [graph, setSelectedNode, openCodePanel]);
}, [graph, nodeById, setSelectedNode, openCodePanel]);
const handleNodeHover = useCallback((nodeId: string | null) => {
if (!nodeId || !graph) {
setHoveredNodeName(null);
return;
}
const node = graph.nodes.find(n => n.id === nodeId);
if (node) {
setHoveredNodeName(node.properties.name);
}
}, [graph]);
const node = nodeById.get(nodeId);
setHoveredNodeName(node ? node.properties.name : null);
}, [graph, nodeById]);
const handleStageClick = useCallback(() => {
setSelectedNode(null);
}, [setSelectedNode]);
const handleToggleAIHighlights = useCallback(() => {
if (isAIHighlightsEnabled) {
clearAIToolHighlights();
clearAICitationHighlights();
clearBlastRadius();
setSelectedNode(null);
setSigmaSelectedNode(null);
}
toggleAIHighlights();
}, [isAIHighlightsEnabled, clearAIToolHighlights, clearAICitationHighlights, clearBlastRadius, setSelectedNode, toggleAIHighlights]);
const {
containerRef,
sigmaRef,
@@ -106,7 +121,7 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
focusNode: (nodeId: string) => {
// Also update app state so the selection syncs properly
if (graph) {
const node = graph.nodes.find(n => n.id === nodeId);
const node = nodeById.get(nodeId);
if (node) {
setSelectedNode(node);
openCodePanel();
@@ -114,7 +129,7 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
}
focusNode(nodeId);
}
}), [focusNode, graph, setSelectedNode, openCodePanel]);
}), [focusNode, graph, nodeById, setSelectedNode, openCodePanel]);
// Update Sigma graph when KnowledgeGraph changes
useEffect(() => {
@@ -126,10 +141,11 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
graph.relationships.forEach(rel => {
if (rel.type === 'MEMBER_OF') {
// Find the community node to get its index
const communityNode = graph.nodes.find(n => n.id === rel.targetId && n.label === 'Community');
if (communityNode) {
const communityNode = nodeById.get(rel.targetId);
if (communityNode && communityNode.label === 'Community') {
// Extract community index from id (e.g., "comm_5" -> 5)
const communityIdx = parseInt(rel.targetId.replace('comm_', ''), 10) || 0;
const numericPart = rel.targetId.replace('comm_', '');
const communityIdx = /^\d+$/.test(numericPart) ? parseInt(numericPart, 10) : 0;
communityMemberships.set(rel.sourceId, communityIdx);
}
}
@@ -137,7 +153,7 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
const sigmaGraph = knowledgeGraphToGraphology(graph, communityMemberships);
setSigmaGraph(sigmaGraph);
}, [graph, setSigmaGraph]);
}, [graph, nodeById, setSigmaGraph]);
// Update node visibility when filters change
useEffect(() => {
@@ -149,7 +165,8 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
filterGraphByDepth(sigmaGraph, appSelectedNode?.id || null, depthFilter, visibleLabels);
sigma.refresh();
}, [visibleLabels, depthFilter, appSelectedNode, sigmaRef]);
// eslint-disable-next-line react-hooks/exhaustive-deps -- sigmaRef identity never changes
}, [visibleLabels, depthFilter, appSelectedNode]);
// Sync app selected node with sigma
useEffect(() => {
@@ -307,23 +324,14 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
{/* AI Highlights toggle - Top Right */}
<div className="absolute top-4 right-4 z-20">
<button
onClick={() => {
if (isAIHighlightsEnabled) {
// Turning off — clear AI highlights and selection (preserve user query highlights)
clearAIToolHighlights();
clearAICitationHighlights();
clearBlastRadius();
setSelectedNode(null);
setSigmaSelectedNode(null);
}
toggleAIHighlights();
}}
onClick={handleToggleAIHighlights}
className={
isAIHighlightsEnabled
? 'w-10 h-10 flex items-center justify-center bg-cyan-500/15 border border-cyan-400/40 rounded-lg text-cyan-200 hover:bg-cyan-500/20 hover:border-cyan-300/60 transition-colors'
: 'w-10 h-10 flex items-center justify-center bg-elevated border border-border-subtle rounded-lg text-text-muted hover:bg-hover hover:text-text-primary transition-colors'
}
title={isAIHighlightsEnabled ? 'Turn off all highlights' : 'Turn on AI highlights'}
data-testid="ai-highlights-toggle"
>
{isAIHighlightsEnabled ? <Lightbulb className="w-4 h-4" /> : <LightbulbOff className="w-4 h-4" />}
</button>
+1 -1
View File
@@ -1,4 +1,4 @@
import { Search, Settings, HelpCircle, Sparkles, Github, Star, ChevronDown } from 'lucide-react';
import { Search, Settings, HelpCircle, Sparkles, Github, Star, ChevronDown } from '@/lib/lucide-icons';
import { useAppState } from '../hooks/useAppState';
import type { RepoSummary } from '../services/server-connection';
import { useState, useMemo, useRef, useEffect, useCallback } from 'react';
@@ -1,11 +1,11 @@
import React, { useState } from 'react';
import React, { useState, useRef, useEffect } from 'react';
import ReactMarkdown from 'react-markdown';
import remarkGfm from 'remark-gfm';
import { Prism as SyntaxHighlighter } from 'react-syntax-highlighter';
import { vscDarkPlus } from 'react-syntax-highlighter/dist/esm/styles/prism';
import { MermaidDiagram } from './MermaidDiagram';
import { ToolCallCard } from './ToolCallCard';
import { Copy, Check } from 'lucide-react';
import { Copy, Check } from '@/lib/lucide-icons';
// Custom syntax theme
const customTheme = {
@@ -39,12 +39,24 @@ export const MarkdownRenderer: React.FC<MarkdownRendererProps> = ({
showCopyButton = false
}) => {
const [copied, setCopied] = useState(false);
const copyTimerRef = useRef<ReturnType<typeof setTimeout>>(undefined);
useEffect(() => {
return () => {
if (copyTimerRef.current) {
clearTimeout(copyTimerRef.current);
}
};
}, []);
const handleCopy = async () => {
try {
await navigator.clipboard.writeText(content);
setCopied(true);
setTimeout(() => setCopied(false), 2000);
if (copyTimerRef.current) {
clearTimeout(copyTimerRef.current);
}
copyTimerRef.current = setTimeout(() => setCopied(false), 2000);
} catch (err) {
console.error('Failed to copy:', err);
}
@@ -78,13 +90,13 @@ export const MarkdownRenderer: React.FC<MarkdownRendererProps> = ({
return parts.join('```');
};
const handleLinkClick = (e: React.MouseEvent<HTMLAnchorElement>, href: string) => {
const handleLinkClick = React.useCallback((e: React.MouseEvent<HTMLAnchorElement>, href: string) => {
if (href.startsWith('code-ref:') || href.startsWith('node-ref:')) {
e.preventDefault();
onLinkClick?.(href);
}
// External links open in new tab (default behavior)
};
}, [onLinkClick]);
const formattedContent = React.useMemo(() => formatMarkdownForDisplay(content), [content]);
@@ -164,7 +176,7 @@ export const MarkdownRenderer: React.FC<MarkdownRendererProps> = ({
);
},
pre: ({ children }: any) => <>{children}</>,
}), [onLinkClick]); // Removed handleLinkClick dependency as it is defined inside component but depends on onLinkClick
}), [handleLinkClick]);
return (
<div className="text-text-primary text-sm">
+20 -9
View File
@@ -1,9 +1,13 @@
import { useEffect, useRef, useState } from 'react';
import { Suspense, useEffect, useRef, useState, lazy } from 'react';
import mermaid from 'mermaid';
import { AlertTriangle, Maximize2 } from 'lucide-react';
import { ProcessFlowModal } from './ProcessFlowModal';
import DOMPurify from 'dompurify';
import { AlertTriangle, Maximize2 } from '@/lib/lucide-icons';
import type { ProcessData } from '../lib/mermaid-generator';
const ProcessFlowModal = lazy(() =>
import('./ProcessFlowModal').then((m) => ({ default: m.ProcessFlowModal })),
);
// Initialize mermaid with cyan theme matching ProcessFlowModal
mermaid.initialize({
startOnLoad: false,
@@ -67,7 +71,8 @@ export const MermaidDiagram = ({ code }: MermaidDiagramProps) => {
// Render the diagram
const { svg: renderedSvg } = await mermaid.render(id, code.trim());
setSvg(renderedSvg);
const sanitizedSvg = DOMPurify.sanitize(renderedSvg, { USE_PROFILES: { svg: true, svgFilters: true }, ADD_TAGS: ['foreignObject'] });
setSvg(sanitizedSvg);
setError(null);
} catch (err) {
// Silent catch for streaming:
@@ -140,17 +145,23 @@ export const MermaidDiagram = ({ code }: MermaidDiagramProps) => {
<div
ref={containerRef}
className="flex items-center justify-center p-4 overflow-auto max-h-[400px]"
dangerouslySetInnerHTML={{ __html: svg }}
dangerouslySetInnerHTML={{ __html: DOMPurify.sanitize(svg, { USE_PROFILES: { svg: true, svgFilters: true }, ADD_TAGS: ['foreignObject'] }) }}
/>
</div>
</div>
{/* Use ProcessFlowModal for expansion */}
{showModal && processData && (
<ProcessFlowModal
process={processData}
onClose={() => setShowModal(false)}
/>
<Suspense
fallback={
<div className="p-4 text-sm text-text-muted">Loading diagram…</div>
}
>
<ProcessFlowModal
process={processData}
onClose={() => setShowModal(false)}
/>
</Suspense>
)}
</>
);
@@ -7,6 +7,7 @@
import { useEffect, useRef, useCallback, useState } from 'react';
import { X, GitBranch, Copy, Focus, Layers, ZoomIn, ZoomOut } from 'lucide-react';
import mermaid from 'mermaid';
import DOMPurify from 'dompurify';
import { ProcessData, generateProcessMermaid } from '../lib/mermaid-generator';
interface ProcessFlowModalProps {
@@ -90,6 +91,7 @@ export const ProcessFlowModal = ({ process, onClose, onFocusInGraph, isFullScree
// Handle keyboard zoom
useEffect(() => {
const handleKeyDown = (e: KeyboardEvent) => {
if (e.target instanceof HTMLInputElement || e.target instanceof HTMLTextAreaElement) return;
if (e.key === '+' || e.key === '=') {
setZoom(prev => Math.min(prev + 0.2, maxZoom));
} else if (e.key === '-' || e.key === '_') {
@@ -136,8 +138,8 @@ export const ProcessFlowModal = ({ process, onClose, onFocusInGraph, isFullScree
const renderDiagram = async () => {
try {
// Check if we have raw mermaid code (from AI chat) or need to generate it
const mermaidCode = (process as any).rawMermaid
? (process as any).rawMermaid
const mermaidCode = process.rawMermaid
? process.rawMermaid
: generateProcessMermaid(process);
const id = `mermaid-${Date.now()}`;
@@ -145,7 +147,8 @@ export const ProcessFlowModal = ({ process, onClose, onFocusInGraph, isFullScree
diagramRef.current!.innerHTML = '';
const { svg } = await mermaid.render(id, mermaidCode);
diagramRef.current!.innerHTML = svg;
if (!diagramRef.current) return;
diagramRef.current!.innerHTML = DOMPurify.sanitize(svg, { USE_PROFILES: { svg: true, svgFilters: true }, ADD_TAGS: ['foreignObject'] });
} catch (error) {
console.error('Mermaid render error:', error);
const errorMessage = error instanceof Error ? error.message : String(error);
@@ -208,6 +211,7 @@ export const ProcessFlowModal = ({ process, onClose, onFocusInGraph, isFullScree
ref={containerRef}
className="fixed inset-0 z-50 flex items-center justify-center bg-black/20 animate-fade-in"
onClick={handleBackdropClick}
data-testid="process-modal"
>
{/* Glassmorphism Modal */}
<div className={`bg-slate-900/60 backdrop-blur-2xl border border-white/10 rounded-3xl shadow-2xl shadow-cyan-500/10 flex flex-col animate-scale-in overflow-hidden relative ${isFullScreen
+12 -5
View File
@@ -11,6 +11,9 @@ import { useAppState } from '../hooks/useAppState';
import { ProcessFlowModal } from './ProcessFlowModal';
import type { ProcessData, ProcessStep } from '../lib/mermaid-generator';
/** Validate that an ID contains only expected node identifier characters (no Cypher metacharacters or spaces) */
const isSafeId = (id: string): boolean => /^[a-zA-Z0-9_:.\-/@]+$/.test(id);
export const ProcessesPanel = () => {
const { graph, runQuery, setHighlightedNodeIds, highlightedNodeIds } = useAppState();
const [searchQuery, setSearchQuery] = useState('');
@@ -79,7 +82,7 @@ export const ProcessesPanel = () => {
setLoadingProcess('all');
try {
const allProcessIds = [...processes.cross, ...processes.intra].map(p => p.id);
const allProcessIds = [...processes.cross, ...processes.intra].map(p => p.id).filter(isSafeId);
if (allProcessIds.length === 0) return;
@@ -110,7 +113,7 @@ export const ProcessesPanel = () => {
}
const allSteps = Array.from(allStepsMap.values());
const stepIds = allSteps.map(s => s.id);
const stepIds = allSteps.map(s => s.id).filter(isSafeId);
// Query for all CALLS edges between the combined steps
if (stepIds.length > 0) {
@@ -155,6 +158,7 @@ export const ProcessesPanel = () => {
// Load process steps and open modal
const handleViewProcess = useCallback(async (processId: string, label: string, processType: string) => {
if (!isSafeId(processId)) return;
setLoadingProcess(processId);
try {
@@ -175,7 +179,7 @@ export const ProcessesPanel = () => {
}));
// Get step IDs for edge query
const stepIds = steps.map(s => s.id);
const stepIds = steps.map(s => s.id).filter(isSafeId);
// Query for CALLS edges between the steps in this process
let edges: Array<{ from: string; to: string; type: string }> = [];
@@ -228,6 +232,7 @@ export const ProcessesPanel = () => {
// Toggle focus for any process - loads steps on demand
const handleToggleFocusForProcess = useCallback(async (processId: string) => {
if (!isSafeId(processId)) return;
// If already focused on this process, turn off
if (focusedProcessId === processId) {
setHighlightedNodeIds(new Set());
@@ -321,7 +326,7 @@ export const ProcessesPanel = () => {
/>
</div>
</div>
<div className="flex items-center gap-2 text-xs text-text-muted">
<div className="flex items-center gap-2 text-xs text-text-muted" data-testid="process-list-loaded">
<span>{totalCount} processes detected</span>
</div>
</div>
@@ -457,7 +462,7 @@ const ProcessItem = ({ process, isLoading, isSelected, isFocused, onView, onTogg
: '';
return (
<div className={`flex items-center gap-2 px-4 py-2 mx-2 rounded-lg hover:bg-hover group transition-all ${rowClass}`}>
<div data-testid="process-row" className={`flex items-center gap-2 px-4 py-2 mx-2 rounded-lg hover:bg-hover group transition-all ${rowClass}`}>
<GitBranch className="w-4 h-4 text-text-muted flex-shrink-0" />
<div className="flex-1 min-w-0">
<div className="text-sm text-text-primary truncate">{process.label}</div>
@@ -479,12 +484,14 @@ const ProcessItem = ({ process, isLoading, isSelected, isFocused, onView, onTogg
: 'text-text-muted hover:text-cyan-400 bg-white/5 hover:bg-cyan-500/20 border border-white/10 hover:border-cyan-400/40 opacity-0 group-hover:opacity-100'
}`}
title={isFocused ? 'Click to remove highlight from graph' : 'Click to highlight in graph'}
data-testid="process-highlight-button"
>
<Lightbulb className="w-4 h-4" />
</button>
<button
onClick={onView}
disabled={isLoading}
data-testid="process-view-button"
className={`flex items-center gap-1.5 px-2.5 py-1.5 text-xs font-medium rounded-md transition-all disabled:opacity-50 shadow-sm ${isSelected
? 'text-cyan-300 bg-cyan-900/60 border border-cyan-400/60 opacity-100'
: 'text-cyan-400 hover:text-cyan-300 bg-cyan-950/30 hover:bg-cyan-900/50 border border-cyan-500/30 hover:border-cyan-400/50 opacity-0 group-hover:opacity-100 shadow-cyan-900/20'
+1 -1
View File
@@ -1,5 +1,5 @@
import { useState, useRef, useEffect, useCallback } from 'react';
import { Terminal, Play, X, ChevronDown, ChevronUp, Loader2, Sparkles, Table } from 'lucide-react';
import { Terminal, Play, X, ChevronDown, ChevronUp, Loader2, Sparkles, Table } from '@/lib/lucide-icons';
import { useAppState } from '../hooks/useAppState';
const EXAMPLE_QUERIES = [
+1 -1
View File
@@ -2,7 +2,7 @@ import { useState, useRef, useEffect, useCallback } from 'react';
import {
Send, Square, Sparkles, User,
PanelRightClose, Loader2, AlertTriangle, GitBranch
} from 'lucide-react';
} from '@/lib/lucide-icons';
import { useAppState } from '../hooks/useAppState';
import { ToolCallCard } from './ToolCallCard';
import { isProviderConfigured } from '../core/llm/settings-service';
+178 -220
View File
@@ -1,12 +1,15 @@
import { useState, useEffect, useCallback, useRef, useMemo } from 'react';
import { X, Key, Server, Brain, Check, AlertCircle, Eye, EyeOff, RefreshCw, ChevronDown, Loader2, Search } from 'lucide-react';
import { X, Key, Server, Brain, Check, AlertCircle, Eye, EyeOff, RefreshCw, ChevronDown, Loader2, Search } from '@/lib/lucide-icons';
import {
loadSettings,
saveSettings,
getProviderDisplayName,
getAvailableModels,
fetchOpenRouterModels,
} from '../core/llm/settings-service';
import type { LLMSettings, LLMProvider } from '../core/llm/types';
import { DEFAULT_OLLAMA_BASE_URL } from '../config/ui-constants';
import { ProviderConfigCard } from './settings/ProviderConfigCard';
interface SettingsPanelProps {
isOpen: boolean;
@@ -216,6 +219,7 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
const [settings, setSettings] = useState<LLMSettings>(loadSettings);
const [showApiKey, setShowApiKey] = useState<Record<string, boolean>>({});
const [saveStatus, setSaveStatus] = useState<'idle' | 'saved' | 'error'>('idle');
const saveTimerRef = useRef<ReturnType<typeof setTimeout>>(undefined);
// Ollama connection state
const [ollamaError, setOllamaError] = useState<string | null>(null);
const [isCheckingOllama, setIsCheckingOllama] = useState(false);
@@ -223,6 +227,15 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
const [openRouterModels, setOpenRouterModels] = useState<Array<{ id: string; name: string }>>([]);
const [isLoadingModels, setIsLoadingModels] = useState(false);
// Clean up save timer on unmount
useEffect(() => {
return () => {
if (saveTimerRef.current) {
clearTimeout(saveTimerRef.current);
}
};
}, []);
// Load settings when panel opens
useEffect(() => {
if (isOpen) {
@@ -252,7 +265,7 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
useEffect(() => {
if (settings.activeProvider === 'ollama') {
const baseUrl = settings.ollama?.baseUrl ?? 'http://localhost:11434';
const baseUrl = settings.ollama?.baseUrl ?? DEFAULT_OLLAMA_BASE_URL;
const timer = setTimeout(() => {
checkOllamaConnection(baseUrl);
}, 300);
@@ -269,7 +282,10 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
saveSettings(settings);
setSaveStatus('saved');
onSettingsSaved?.();
setTimeout(() => setSaveStatus('idle'), 2000);
if (saveTimerRef.current) {
clearTimeout(saveTimerRef.current);
}
saveTimerRef.current = setTimeout(() => setSaveStatus('idle'), 2000);
} catch {
setSaveStatus('error');
}
@@ -281,7 +297,7 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
if (!isOpen) return null;
const providers: LLMProvider[] = ['openai', 'gemini', 'anthropic', 'azure-openai', 'ollama', 'openrouter', 'minimax'];
const providers: LLMProvider[] = ['openai', 'gemini', 'anthropic', 'azure-openai', 'ollama', 'openrouter', 'minimax', 'glm'];
return (
@@ -366,7 +382,7 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
w-8 h-8 rounded-lg flex items-center justify-center text-lg
${settings.activeProvider === provider ? 'bg-accent/20' : 'bg-surface'}
`}>
{provider === 'openai' ? '🤖' : provider === 'gemini' ? '💎' : provider === 'anthropic' ? '🧠' : provider === 'ollama' ? '🦙' : provider === 'openrouter' ? '🌐' : provider === 'minimax' ? '⚡' : '☁️'}
{provider === 'openai' ? '🤖' : provider === 'gemini' ? '💎' : provider === 'anthropic' ? '🧠' : provider === 'ollama' ? '🦙' : provider === 'openrouter' ? '🌐' : provider === 'minimax' ? '⚡' : provider === 'glm' ? '🔮' : '☁️'}
</div>
<span className="font-medium">{getProviderDisplayName(provider)}</span>
</button>
@@ -374,60 +390,36 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
</div>
</div>
<div className="p-3 bg-amber-500/10 border border-amber-500/30 rounded-xl text-xs text-amber-200">
API keys are stored in session storage and will be cleared when you close this tab.
</div>
{/* OpenAI Settings */}
{settings.activeProvider === 'openai' && (
<div className="space-y-4 animate-fade-in">
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
<Key className="w-4 h-4" />
API Key
</label>
<div className="relative">
<input
type={showApiKey['openai'] ? 'text' : 'password'}
value={settings.openai?.apiKey ?? ''}
onChange={e => setSettings(prev => ({
...prev,
openai: { ...prev.openai!, apiKey: e.target.value }
}))}
placeholder="Enter your OpenAI API key"
className="w-full px-4 py-3 pr-12 bg-elevated border border-border-subtle rounded-xl text-text-primary placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all"
/>
<button
type="button"
onClick={() => toggleApiKeyVisibility('openai')}
className="absolute right-3 top-1/2 -translate-y-1/2 p-1 text-text-muted hover:text-text-primary transition-colors"
>
{showApiKey['openai'] ? <EyeOff className="w-4 h-4" /> : <Eye className="w-4 h-4" />}
</button>
</div>
<p className="text-xs text-text-muted">
Get your API key from{' '}
<a
href="https://platform.openai.com/api-keys"
target="_blank"
rel="noopener noreferrer"
className="text-accent hover:underline"
>
OpenAI Platform
</a>
</p>
</div>
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">Model</label>
<input
type="text"
value={settings.openai?.model ?? 'gpt-5.2-chat'}
onChange={e => setSettings(prev => ({
...prev,
openai: { ...prev.openai!, model: e.target.value }
}))}
placeholder="e.g., gpt-4o, gpt-4-turbo, gpt-3.5-turbo"
className="w-full px-4 py-3 bg-elevated border border-border-subtle rounded-xl text-text-primary placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all font-mono text-sm"
/>
</div>
<ProviderConfigCard
title="OpenAI"
apiKey={{
value: settings.openai?.apiKey ?? '',
placeholder: 'Enter your OpenAI API key',
helperText: 'Get your API key from',
helperLink: 'https://platform.openai.com/api-keys',
helperLinkLabel: 'OpenAI Platform',
isVisible: !!showApiKey['openai'],
onChange: (value) => setSettings(prev => ({
...prev,
openai: { ...prev.openai!, apiKey: value }
})),
onToggleVisibility: () => toggleApiKeyVisibility('openai'),
}}
model={{
value: settings.openai?.model ?? 'gpt-5.2-chat',
placeholder: 'e.g., gpt-4o, gpt-4-turbo, gpt-3.5-turbo',
onChange: (value) => setSettings(prev => ({
...prev,
openai: { ...prev.openai!, model: value }
})),
}}
>
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
<Server className="w-4 h-4" />
@@ -447,119 +439,63 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
Leave empty to use the default OpenAI API. Set a custom URL for proxies or compatible APIs.
</p>
</div>
</div>
</ProviderConfigCard>
)}
{/* Gemini Settings */}
{settings.activeProvider === 'gemini' && (
<div className="space-y-4 animate-fade-in">
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
<Key className="w-4 h-4" />
API Key
</label>
<div className="relative">
<input
type={showApiKey['gemini'] ? 'text' : 'password'}
value={settings.gemini?.apiKey ?? ''}
onChange={e => setSettings(prev => ({
...prev,
gemini: { ...prev.gemini!, apiKey: e.target.value }
}))}
placeholder="Enter your Google AI API key"
className="w-full px-4 py-3 pr-12 bg-elevated border border-border-subtle rounded-xl text-text-primary placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all"
/>
<button
type="button"
onClick={() => toggleApiKeyVisibility('gemini')}
className="absolute right-3 top-1/2 -translate-y-1/2 p-1 text-text-muted hover:text-text-primary transition-colors"
>
{showApiKey['gemini'] ? <EyeOff className="w-4 h-4" /> : <Eye className="w-4 h-4" />}
</button>
</div>
<p className="text-xs text-text-muted">
Get your API key from{' '}
<a
href="https://aistudio.google.com/app/apikey"
target="_blank"
rel="noopener noreferrer"
className="text-accent hover:underline"
>
Google AI Studio
</a>
</p>
</div>
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">Model</label>
<input
type="text"
value={settings.gemini?.model ?? 'gemini-2.0-flash'}
onChange={e => setSettings(prev => ({
...prev,
gemini: { ...prev.gemini!, model: e.target.value }
}))}
placeholder="e.g., gemini-2.0-flash, gemini-1.5-pro"
className="w-full px-4 py-3 bg-elevated border border-border-subtle rounded-xl text-text-primary placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all font-mono text-sm"
/>
</div>
</div>
<ProviderConfigCard
title="Google Gemini"
apiKey={{
value: settings.gemini?.apiKey ?? '',
placeholder: 'Enter your Google AI API key',
helperText: 'Get your API key from',
helperLink: 'https://aistudio.google.com/app/apikey',
helperLinkLabel: 'Google AI Studio',
isVisible: !!showApiKey['gemini'],
onChange: (value) => setSettings(prev => ({
...prev,
gemini: { ...prev.gemini!, apiKey: value }
})),
onToggleVisibility: () => toggleApiKeyVisibility('gemini'),
}}
model={{
value: settings.gemini?.model ?? 'gemini-2.0-flash',
placeholder: 'e.g., gemini-2.0-flash, gemini-1.5-pro',
onChange: (value) => setSettings(prev => ({
...prev,
gemini: { ...prev.gemini!, model: value }
})),
}}
/>
)}
{/* Anthropic Settings */}
{settings.activeProvider === 'anthropic' && (
<div className="space-y-4 animate-fade-in">
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
<Key className="w-4 h-4" />
API Key
</label>
<div className="relative">
<input
type={showApiKey['anthropic'] ? 'text' : 'password'}
value={settings.anthropic?.apiKey ?? ''}
onChange={e => setSettings(prev => ({
...prev,
anthropic: { ...prev.anthropic!, apiKey: e.target.value }
}))}
placeholder="Enter your Anthropic API key"
className="w-full px-4 py-3 pr-12 bg-elevated border border-border-subtle rounded-xl text-text-primary placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all"
/>
<button
type="button"
onClick={() => toggleApiKeyVisibility('anthropic')}
className="absolute right-3 top-1/2 -translate-y-1/2 p-1 text-text-muted hover:text-text-primary transition-colors"
>
{showApiKey['anthropic'] ? <EyeOff className="w-4 h-4" /> : <Eye className="w-4 h-4" />}
</button>
</div>
<p className="text-xs text-text-muted">
Get your API key from{' '}
<a
href="https://console.anthropic.com/settings/keys"
target="_blank"
rel="noopener noreferrer"
className="text-accent hover:underline"
>
Anthropic Console
</a>
</p>
</div>
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">Model</label>
<input
type="text"
value={settings.anthropic?.model ?? 'claude-sonnet-4-20250514'}
onChange={e => setSettings(prev => ({
...prev,
anthropic: { ...prev.anthropic!, model: e.target.value }
}))}
placeholder="e.g., claude-sonnet-4-20250514, claude-3-opus"
className="w-full px-4 py-3 bg-elevated border border-border-subtle rounded-xl text-text-primary placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all font-mono text-sm"
/>
</div>
</div>
<ProviderConfigCard
title="Anthropic"
apiKey={{
value: settings.anthropic?.apiKey ?? '',
placeholder: 'Enter your Anthropic API key',
helperText: 'Get your API key from',
helperLink: 'https://console.anthropic.com/settings/keys',
helperLinkLabel: 'Anthropic Console',
isVisible: !!showApiKey['anthropic'],
onChange: (value) => setSettings(prev => ({
...prev,
anthropic: { ...prev.anthropic!, apiKey: value }
})),
onToggleVisibility: () => toggleApiKeyVisibility('anthropic'),
}}
model={{
value: settings.anthropic?.model ?? 'claude-sonnet-4-20250514',
placeholder: 'e.g., claude-sonnet-4-20250514, claude-3-opus',
onChange: (value) => setSettings(prev => ({
...prev,
anthropic: { ...prev.anthropic!, model: value }
})),
}}
/>
)}
{/* Azure OpenAI Settings */}
@@ -695,17 +631,17 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
<div className="flex gap-2">
<input
type="url"
value={settings.ollama?.baseUrl ?? 'http://localhost:11434'}
value={settings.ollama?.baseUrl ?? DEFAULT_OLLAMA_BASE_URL}
onChange={e => setSettings(prev => ({
...prev,
ollama: { ...prev.ollama!, baseUrl: e.target.value }
}))}
placeholder="http://localhost:11434"
placeholder={DEFAULT_OLLAMA_BASE_URL}
className="flex-1 px-4 py-3 bg-elevated border border-border-subtle rounded-xl text-text-primary placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all font-mono text-sm"
/>
<button
type="button"
onClick={() => checkOllamaConnection(settings.ollama?.baseUrl ?? 'http://localhost:11434')}
onClick={() => checkOllamaConnection(settings.ollama?.baseUrl ?? DEFAULT_OLLAMA_BASE_URL)}
disabled={isCheckingOllama}
className="px-3 py-3 bg-elevated border border-border-subtle rounded-xl text-text-secondary hover:text-text-primary hover:border-accent/50 transition-colors disabled:opacity-50"
title="Check connection"
@@ -749,44 +685,22 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
{/* OpenRouter Settings */}
{settings.activeProvider === 'openrouter' && (
<div className="space-y-4 animate-fade-in">
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
<Key className="w-4 h-4" />
API Key
</label>
<div className="relative">
<input
type={showApiKey['openrouter'] ? 'text' : 'password'}
value={settings.openrouter?.apiKey ?? ''}
onChange={e => setSettings(prev => ({
...prev,
openrouter: { ...prev.openrouter!, apiKey: e.target.value }
}))}
placeholder="Enter your OpenRouter API key"
className="w-full px-4 py-3 pr-12 bg-elevated border border-border-subtle rounded-xl text-text-primary placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all"
/>
<button
type="button"
onClick={() => toggleApiKeyVisibility('openrouter')}
className="absolute right-3 top-1/2 -translate-y-1/2 p-1 text-text-muted hover:text-text-primary transition-colors"
>
{showApiKey['openrouter'] ? <EyeOff className="w-4 h-4" /> : <Eye className="w-4 h-4" />}
</button>
</div>
<p className="text-xs text-text-muted">
Get your API key from{' '}
<a
href="https://openrouter.ai/keys"
target="_blank"
rel="noopener noreferrer"
className="text-accent hover:underline"
>
OpenRouter Keys
</a>
</p>
</div>
<ProviderConfigCard
title="OpenRouter"
apiKey={{
value: settings.openrouter?.apiKey ?? '',
placeholder: 'Enter your OpenRouter API key',
helperText: 'Get your API key from',
helperLink: 'https://openrouter.ai/keys',
helperLinkLabel: 'OpenRouter Keys',
isVisible: !!showApiKey['openrouter'],
onChange: (value) => setSettings(prev => ({
...prev,
openrouter: { ...prev.openrouter!, apiKey: value }
})),
onToggleVisibility: () => toggleApiKeyVisibility('openrouter'),
}}
>
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">Model</label>
<OpenRouterModelCombobox
@@ -811,11 +725,40 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
</a>
</p>
</div>
</div>
</ProviderConfigCard>
)}
{/* MiniMax Settings */}
{settings.activeProvider === 'minimax' && (
<ProviderConfigCard
title="MiniMax"
apiKey={{
value: settings.minimax?.apiKey ?? '',
placeholder: 'Enter your MiniMax API key',
helperText: 'Get your API key from',
helperLink: 'https://platform.minimax.io',
helperLinkLabel: 'MiniMax Platform',
isVisible: !!showApiKey['minimax'],
onChange: (value) => setSettings(prev => ({
...prev,
minimax: { ...prev.minimax!, apiKey: value }
})),
onToggleVisibility: () => toggleApiKeyVisibility('minimax'),
}}
model={{
value: settings.minimax?.model ?? 'MiniMax-M2.5',
placeholder: 'e.g., MiniMax-M2.5, MiniMax-M2.5-highspeed',
onChange: (value) => setSettings(prev => ({
...prev,
minimax: { ...prev.minimax!, model: value }
})),
helperText: 'Available: MiniMax-M2.5 (default), MiniMax-M2.5-highspeed (faster)',
}}
/>
)}
{/* GLM Settings */}
{settings.activeProvider === 'glm' && (
<div className="space-y-4 animate-fade-in">
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
@@ -824,50 +767,66 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
</label>
<div className="relative">
<input
type={showApiKey['minimax'] ? 'text' : 'password'}
value={settings.minimax?.apiKey ?? ''}
type={showApiKey['glm'] ? 'text' : 'password'}
value={settings.glm?.apiKey ?? ''}
onChange={e => setSettings(prev => ({
...prev,
minimax: { ...prev.minimax!, apiKey: e.target.value }
glm: { ...prev.glm!, apiKey: e.target.value }
}))}
placeholder="Enter your MiniMax API key"
placeholder="Enter your Z.AI API key"
className="w-full px-4 py-3 pr-12 bg-elevated border border-border-subtle rounded-xl text-text-primary placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all"
/>
<button
type="button"
onClick={() => toggleApiKeyVisibility('minimax')}
onClick={() => toggleApiKeyVisibility('glm')}
className="absolute right-3 top-1/2 -translate-y-1/2 p-1 text-text-muted hover:text-text-primary transition-colors"
>
{showApiKey['minimax'] ? <EyeOff className="w-4 h-4" /> : <Eye className="w-4 h-4" />}
{showApiKey['glm'] ? <EyeOff className="w-4 h-4" /> : <Eye className="w-4 h-4" />}
</button>
</div>
<p className="text-xs text-text-muted">
Get your API key from{' '}
<a
href="https://platform.minimax.io"
href="https://docs.z.ai"
target="_blank"
rel="noopener noreferrer"
className="text-accent hover:underline"
>
MiniMax Platform
Z.AI Platform
</a>
</p>
</div>
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">Model</label>
<input
type="text"
value={settings.minimax?.model ?? 'MiniMax-M2.5'}
<select
value={settings.glm?.model ?? 'GLM-5'}
onChange={e => setSettings(prev => ({
...prev,
minimax: { ...prev.minimax!, model: e.target.value }
glm: { ...prev.glm!, model: e.target.value }
}))}
placeholder="e.g., MiniMax-M2.5, MiniMax-M2.5-highspeed"
className="w-full px-4 py-3 bg-elevated border border-border-subtle rounded-xl text-text-primary focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all font-mono text-sm"
>
{getAvailableModels('glm').map(model => (
<option key={model} value={model}>{model}</option>
))}
</select>
</div>
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">Base URL</label>
<input
type="text"
value={settings.glm?.baseUrl ?? 'https://api.z.ai/api/coding/paas/v4'}
onChange={e => setSettings(prev => ({
...prev,
glm: { ...prev.glm!, baseUrl: e.target.value }
}))}
placeholder="https://api.z.ai/api/coding/paas/v4"
className="w-full px-4 py-3 bg-elevated border border-border-subtle rounded-xl text-text-primary placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all font-mono text-sm"
/>
<p className="text-xs text-text-muted">
Available models: MiniMax-M2.5 (default), MiniMax-M2.5-highspeed (faster)
Coding API (default). Use https://api.z.ai/api/paas/v4 for the general API.
</p>
</div>
</div>
@@ -880,8 +839,7 @@ export const SettingsPanel = ({ isOpen, onClose, onSettingsSaved, backendUrl, is
🔒
</div>
<div className="text-xs text-text-muted leading-relaxed">
<span className="text-text-secondary font-medium">Privacy:</span> Your API keys are stored only in your browser's local storage.
They're sent directly to the LLM provider when you chat. Your code never leaves your machine.
<span className="text-text-secondary font-medium">Privacy:</span> Your API keys are stored only in your browser's session storage and are cleared when the tab closes. They're sent directly to the LLM provider when you chat. Your code never leaves your machine.
</div>
</div>
</div>
+5 -4
View File
@@ -1,4 +1,5 @@
import { Heart } from 'lucide-react';
import { useMemo } from 'react';
import { Heart } from '@/lib/lucide-icons';
import { useAppState } from '../hooks/useAppState';
export const StatusBar = () => {
@@ -8,7 +9,7 @@ export const StatusBar = () => {
const edgeCount = graph?.relationships.length ?? 0;
// Detect primary language
const primaryLanguage = (() => {
const primaryLanguage = useMemo(() => {
if (!graph) return null;
const languages = graph.nodes
.map(n => n.properties.language)
@@ -21,7 +22,7 @@ export const StatusBar = () => {
}, {} as Record<string, number>);
return Object.entries(counts).sort((a, b) => b[1] - a[1])[0]?.[0];
})();
}, [graph]);
return (
<footer className="flex items-center justify-between px-5 py-2 bg-deep border-t border-dashed border-border-subtle text-[11px] text-text-muted">
@@ -38,7 +39,7 @@ export const StatusBar = () => {
<span>{progress.message}</span>
</>
) : (
<div className="flex items-center gap-1.5">
<div className="flex items-center gap-1.5" data-testid="status-ready">
<span className="w-1.5 h-1.5 bg-node-function rounded-full" />
<span>Ready</span>
</div>
+1 -1
View File
@@ -6,7 +6,7 @@
*/
import { useState } from 'react';
import { ChevronDown, ChevronRight, Sparkles, Check, Loader2, AlertCircle } from 'lucide-react';
import { ChevronDown, ChevronRight, Sparkles, Check, Loader2, AlertCircle } from '@/lib/lucide-icons';
import type { ToolCallInfo } from '../core/llm/types';
interface ToolCallCardProps {
@@ -1,5 +1,5 @@
import { useState, useEffect } from 'react';
import { X, Snail, Rocket, SkipForward } from 'lucide-react';
import { X, Snail, Rocket, SkipForward } from '@/lib/lucide-icons';
interface WebGPUFallbackDialogProps {
isOpen: boolean;
@@ -0,0 +1,110 @@
import { ReactNode } from 'react';
import { Eye, EyeOff, Key } from '@/lib/lucide-icons';
type ApiKeyField = {
value: string;
placeholder: string;
helperText?: string;
helperLink?: string;
helperLinkLabel?: string;
isVisible: boolean;
onChange: (value: string) => void;
onToggleVisibility: () => void;
};
type ModelField = {
value: string;
placeholder: string;
label?: string;
helperText?: string;
onChange: (value: string) => void;
};
interface ProviderConfigCardProps {
title: string;
description?: string;
apiKey?: ApiKeyField;
model?: ModelField;
children?: ReactNode;
}
export const ProviderConfigCard = ({
title,
description,
apiKey,
model,
children,
}: ProviderConfigCardProps) => {
return (
<div className="space-y-4 animate-fade-in">
<div className="flex items-center justify-between">
<div>
<h3 className="text-sm font-semibold text-text-primary">{title}</h3>
{description ? (
<p className="text-xs text-text-muted">{description}</p>
) : null}
</div>
</div>
{apiKey && (
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
<Key className="w-4 h-4" />
API Key
</label>
<div className="relative">
<input
type={apiKey.isVisible ? 'text' : 'password'}
value={apiKey.value}
onChange={e => apiKey.onChange(e.target.value)}
placeholder={apiKey.placeholder}
className="w-full px-4 py-3 pr-12 bg-elevated border border-border-subtle rounded-xl text-text-primary placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all"
/>
<button
type="button"
onClick={apiKey.onToggleVisibility}
className="absolute right-3 top-1/2 -translate-y-1/2 p-1 text-text-muted hover:text-text-primary transition-colors"
>
{apiKey.isVisible ? <EyeOff className="w-4 h-4" /> : <Eye className="w-4 h-4" />}
</button>
</div>
{apiKey.helperText && (
<p className="text-xs text-text-muted">
{apiKey.helperText}{' '}
{apiKey.helperLink ? (
<a
href={apiKey.helperLink}
target="_blank"
rel="noopener noreferrer"
className="text-accent hover:underline"
>
{apiKey.helperLinkLabel ?? 'Learn more'}
</a>
) : null}
</p>
)}
</div>
)}
{model && (
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">
{model.label ?? 'Model'}
</label>
<input
type="text"
value={model.value}
onChange={e => model.onChange(e.target.value)}
placeholder={model.placeholder}
className="w-full px-4 py-3 bg-elevated border border-border-subtle rounded-xl text-text-primary placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20 outline-none transition-all font-mono text-sm"
/>
{model.helperText ? (
<p className="text-xs text-text-muted">{model.helperText}</p>
) : null}
</div>
)}
{children}
</div>
);
};
+7
View File
@@ -0,0 +1,7 @@
// Centralized UI and provider defaults to reduce magic numbers and duplicated URLs.
export const ERROR_RESET_DELAY_MS = 3000;
export const BACKEND_URL_DEBOUNCE_MS = 500;
export const DEFAULT_BACKEND_URL = 'http://localhost:4747';
export const DEFAULT_OLLAMA_BASE_URL = 'http://localhost:11434';
export const DEFAULT_OPENROUTER_BASE_URL = 'https://openrouter.ai/api/v1';
+20 -1
View File
@@ -15,7 +15,24 @@ export type NodeLabel =
| 'Type'
| 'CodeElement'
| 'Community'
| 'Process';
| 'Process'
| 'Section'
| 'Struct'
| 'Trait'
| 'Impl'
| 'TypeAlias'
| 'Const'
| 'Static'
| 'Namespace'
| 'Union'
| 'Typedef'
| 'Macro'
| 'Property'
| 'Record'
| 'Delegate'
| 'Annotation'
| 'Constructor'
| 'Template';
export type NodeProperties = {
@@ -55,6 +72,8 @@ export type RelationshipType =
| 'IMPLEMENTS'
| 'EXTENDS'
| 'HAS_METHOD'
| 'HAS_PROPERTY'
| 'ACCESSES'
| 'MEMBER_OF'
| 'STEP_IN_PROCESS'
+41 -2
View File
@@ -202,6 +202,35 @@ const generateCodeElementCSV = (
return rows.join('\n');
};
/**
* Generate CSV for multi-language code element nodes (Struct, Enum, Macro, etc.)
* These do NOT have isExported column.
* Headers: id,name,filePath,startLine,endLine,content
*/
const generateMultiLangCSV = (
nodes: GraphNode[],
label: NodeLabel,
fileContents: Map<string, string>
): string => {
const headers = ['id', 'name', 'filePath', 'startLine', 'endLine', 'content'];
const rows: string[] = [headers.join(',')];
for (const node of nodes) {
if (node.label !== label) continue;
const content = extractContent(node, fileContents);
rows.push([
escapeCSVField(node.id),
escapeCSVField(node.properties.name || ''),
escapeCSVField(node.properties.filePath || ''),
escapeCSVNumber(node.properties.startLine, -1),
escapeCSVNumber(node.properties.endLine, -1),
escapeCSVField(content),
].join(','));
}
return rows.join('\n');
};
/**
* Generate CSV for Community nodes (from Leiden algorithm)
* Headers: id,label,heuristicLabel,keywords,description,enrichedBy,cohesion,symbolCount
@@ -221,7 +250,7 @@ const generateCommunityCSV = (nodes: GraphNode[]): string => {
escapeCSVField(node.id),
escapeCSVField(node.properties.name || ''), // label is stored in name
escapeCSVField(node.properties.heuristicLabel || ''),
keywordsStr, // Array format for LadybugDB
escapeCSVField(keywordsStr), // Array format for LadybugDB, needs CSV escaping for commas
escapeCSVField((node.properties as any).description || ''),
escapeCSVField((node.properties as any).enrichedBy || 'heuristic'),
escapeCSVNumber(node.properties.cohesion, 0),
@@ -312,7 +341,17 @@ export const generateAllCSVs = (
nodeCSVs.set('CodeElement', generateCodeElementCSV(nodes, 'CodeElement', fileContents));
nodeCSVs.set('Community', generateCommunityCSV(nodes));
nodeCSVs.set('Process', generateProcessCSV(nodes));
// Generate CSVs for remaining multi-language tables (no isExported column)
const handledTables = new Set<NodeTableName>([
'File', 'Folder', 'Function', 'Class', 'Interface', 'Method', 'CodeElement', 'Community', 'Process',
]);
for (const table of NODE_TABLES) {
if (!handledTables.has(table)) {
nodeCSVs.set(table, generateMultiLangCSV(nodes, table as NodeLabel, fileContents));
}
}
// Generate single relation CSV
const relCSV = generateRelationCSV(graph);
+226 -99
View File
@@ -16,55 +16,62 @@ import {
NodeTableName,
} from './schema';
import { generateAllCSVs } from './csv-generator';
import { getQueryRows } from './query-result';
// Holds the reference to the dynamically loaded module
let lbug: any = null;
let db: any = null;
let conn: any = null;
let initPromise: Promise<{ db: any; conn: any; lbug: any }> | null = null;
/**
* Initialize LadybugDB WASM module and create in-memory database
*/
export const initLbug = async () => {
if (conn) return { db, conn, lbug };
if (initPromise) return initPromise;
initPromise = (async () => {
try {
if (import.meta.env.DEV) console.log('🚀 Initializing LadybugDB...');
try {
if (import.meta.env.DEV) console.log('🚀 Initializing LadybugDB...');
// 1. Dynamic Import (Fixes the "not a function" bundler issue)
const lbugModule = await import('@ladybugdb/wasm-core');
// 1. Dynamic Import (Fixes the "not a function" bundler issue)
const lbugModule = await import('@ladybugdb/wasm-core');
// 2. Handle Vite/Webpack "default" wrapping
lbug = lbugModule.default || lbugModule;
// 2. Handle Vite/Webpack "default" wrapping
lbug = lbugModule.default || lbugModule;
// 3. Initialize WASM
await lbug.init();
// 3. Initialize WASM
await lbug.init();
// 4. Create Database with 512MB buffer manager
const BUFFER_POOL_SIZE = 512 * 1024 * 1024; // 512MB
db = new lbug.Database(':memory:', BUFFER_POOL_SIZE);
conn = new lbug.Connection(db);
// 4. Create Database with 512MB buffer manager
const BUFFER_POOL_SIZE = 512 * 1024 * 1024; // 512MB
db = new lbug.Database(':memory:', BUFFER_POOL_SIZE);
conn = new lbug.Connection(db);
if (import.meta.env.DEV) console.log('✅ LadybugDB WASM Initialized');
if (import.meta.env.DEV) console.log('✅ LadybugDB WASM Initialized');
// 5. Initialize Schema (all node tables, then rel tables, then embedding table)
for (const schemaQuery of SCHEMA_QUERIES) {
try {
await conn.query(schemaQuery);
} catch (e) {
// Schema might already exist, skip
if (import.meta.env.DEV) {
console.warn('Schema creation skipped (may already exist):', e);
// 5. Initialize Schema (all node tables, then rel tables, then embedding table)
for (let i = 0; i < SCHEMA_QUERIES.length; i++) {
try {
await conn.query(SCHEMA_QUERIES[i]);
} catch (e) {
// Schema might already exist, skip
if (import.meta.env.DEV) {
console.warn(`Schema query ${i + 1}/${SCHEMA_QUERIES.length} skipped (may already exist):`, e);
}
}
}
if (import.meta.env.DEV) console.log('✅ LadybugDB Multi-Table Schema Created');
return { db, conn, lbug };
} catch (error) {
if (import.meta.env.DEV) console.error('❌ LadybugDB Initialization Failed:', error);
throw error;
}
if (import.meta.env.DEV) console.log('✅ LadybugDB Multi-Table Schema Created');
return { db, conn, lbug };
})();
try {
return await initPromise;
} catch (error) {
if (import.meta.env.DEV) console.error('❌ LadybugDB Initialization Failed:', error);
initPromise = null; // Reset on failure so retry is possible
throw error;
}
};
@@ -73,11 +80,60 @@ export const initLbug = async () => {
* Load a KnowledgeGraph into LadybugDB using COPY FROM (bulk load)
* Uses batched CSV writes and COPY statements for optimal performance
*/
const isTestEnv = () => {
// Browser-friendly check: Vite only exposes VITE_* vars at runtime; fall back to a window flag if injected by tests.
if (typeof import.meta !== 'undefined' && typeof import.meta.env !== 'undefined') {
if (import.meta.env.VITE_PLAYWRIGHT_TEST || import.meta.env.MODE === 'test') return true;
}
if (typeof window !== 'undefined' && (window as unknown as { __PLAYWRIGHT_TEST__?: boolean }).__PLAYWRIGHT_TEST__) {
return true;
}
if (typeof navigator !== 'undefined' && navigator.webdriver) {
return true;
}
return typeof process !== 'undefined' && (process.env.PLAYWRIGHT_TEST || process.env.NODE_ENV === 'test');
};
export const loadGraphToLbug = async (
graph: KnowledgeGraph,
fileContents: Map<string, string>
) => {
const { conn, lbug } = await initLbug();
// In headless Playwright, skip heavy bulk load to avoid hangs; UI still functions with empty DB.
if (isTestEnv()) {
if (import.meta.env.DEV) console.log('🧪 Skipping LadybugDB bulk load in test mode');
await initLbug(); // ensure module initialized for downstream calls
return { success: true, count: 0 };
}
const { lbug: lbugModule } = await initLbug();
// Close previous connection/database to avoid leaking WASM resources across repo switches
if (conn) {
try { await conn.close(); } catch {}
conn = null;
}
if (db) {
try { await db.close(); } catch {}
db = null;
}
// Recreate a fresh in-memory DB each load to avoid cleanup/quoting issues with reserved names
const BUFFER_POOL_SIZE = 512 * 1024 * 1024; // 512MB (mirror init)
db = new lbugModule.Database(':memory:', BUFFER_POOL_SIZE);
conn = new lbugModule.Connection(db);
// Update initPromise so subsequent initLbug() calls return the fresh db/conn
initPromise = Promise.resolve({ db, conn, lbug: lbugModule });
// Re-run schema creation
for (let i = 0; i < SCHEMA_QUERIES.length; i++) {
try {
await conn.query(SCHEMA_QUERIES[i]);
} catch (e) {
if (import.meta.env.DEV) {
console.warn(`Schema query ${i + 1}/${SCHEMA_QUERIES.length} skipped (may already exist):`, e);
}
}
}
try {
if (import.meta.env.DEV) console.log(`LadybugDB: Generating CSVs for ${graph.nodeCount} nodes...`);
@@ -131,48 +187,85 @@ export const loadGraphToLbug = async (
let insertedRels = 0;
let skippedRels = 0;
const skippedRelStats = new Map<string, number>();
// Group relations by (fromLabel, toLabel) pair for prepared statement reuse
const relsByLabelPair = new Map<string, Array<{ fromId: string; toId: string; relType: string; confidence: number; reason: string; step: number }>>();
// RFC 4180 regex: handles doubled quotes ("") inside quoted fields
const csvRegex = /"((?:[^"]|"")*)","((?:[^"]|"")*)","((?:[^"]|"")*)",([0-9.]+),"((?:[^"]|"")*)",([0-9-]+)/;
for (const line of relLines) {
try {
// Format: "from","to","type",confidence,"reason",step
const match = line.match(/"([^"]*)","([^"]*)","([^"]*)",([0-9.]+),"([^"]*)",([0-9-]+)/);
if (!match) continue;
const match = line.match(csvRegex);
if (!match) continue;
const [, fromId, toId, relType, confidenceStr, reason, stepStr] = match;
// Unescape RFC 4180 doubled quotes
const fromId = match[1].replace(/""/g, '"');
const toId = match[2].replace(/""/g, '"');
const relType = match[3].replace(/""/g, '"');
const reason = match[5].replace(/""/g, '"');
const fromLabel = getNodeLabel(fromId);
const toLabel = getNodeLabel(toId);
const fromLabel = getNodeLabel(fromId);
const toLabel = getNodeLabel(toId);
// Skip relationships where either node's label doesn't have a table in LadybugDB
// Querying a non-existent table causes a fatal native crash
if (!validTables.has(fromLabel) || !validTables.has(toLabel)) {
skippedRels++;
continue;
}
const confidence = parseFloat(confidenceStr) || 1.0;
const step = parseInt(stepStr) || 0;
const insertQuery = `
MATCH (a:${escapeLabel(fromLabel)} {id: '${fromId.replace(/'/g, "''")}'}),
(b:${escapeLabel(toLabel)} {id: '${toId.replace(/'/g, "''")}'})
CREATE (a)-[:${REL_TABLE_NAME} {type: '${relType}', confidence: ${confidence}, reason: '${reason.replace(/'/g, "''")}', step: ${step}}]->(b)
`;
await conn.query(insertQuery);
insertedRels++;
} catch (err) {
// Skip relationships where either node's label doesn't have a table in LadybugDB
// Querying a non-existent table causes a fatal native crash
if (!validTables.has(fromLabel) || !validTables.has(toLabel)) {
skippedRels++;
const match = line.match(/"([^"]*)","([^"]*)","([^"]*)",([0-9.]+),"([^"]*)"/);
if (match) {
const [, fromId, toId, relType] = match;
const fromLabel = getNodeLabel(fromId);
const toLabel = getNodeLabel(toId);
const key = `${relType}:${fromLabel}->` + toLabel;
skippedRelStats.set(key, (skippedRelStats.get(key) || 0) + 1);
continue;
}
if (import.meta.env.DEV) {
console.warn(`⚠️ Skipped: ${key} | "${fromId}" → "${toId}" | ${err instanceof Error ? err.message : String(err)}`);
const key = `${fromLabel}:${toLabel}`;
if (!relsByLabelPair.has(key)) relsByLabelPair.set(key, []);
relsByLabelPair.get(key)!.push({
fromId,
toId,
relType,
confidence: parseFloat(match[4]) || 1.0,
reason,
step: parseInt(match[6]) || 0,
});
}
// Execute batched prepared statements per label pair
// Prepare once per (fromLabel, toLabel) pair and reuse across all rows
for (const [key, rels] of relsByLabelPair) {
const [fromLabel, toLabel] = key.split(':');
const cypher = `
MATCH (a:${escapeLabel(fromLabel)} {id: $fromId}),
(b:${escapeLabel(toLabel)} {id: $toId})
CREATE (a)-[:${REL_TABLE_NAME} {type: $relType, confidence: $confidence, reason: $reason, step: $step}]->(b)
`;
const stmt = await conn.prepare(cypher);
if (!stmt.isSuccess()) {
const errMsg = await stmt.getErrorMessage();
if (import.meta.env.DEV) console.warn(`Prepare failed for ${key}: ${errMsg}`);
skippedRels += rels.length;
await stmt.close();
continue;
}
try {
for (let i = 0; i < rels.length; i++) {
try {
await conn.execute(stmt, rels[i]);
insertedRels++;
} catch (err) {
skippedRels++;
const r = rels[i];
const statKey = `${r.relType}:${fromLabel}->${toLabel}`;
skippedRelStats.set(statKey, (skippedRelStats.get(statKey) || 0) + 1);
if (import.meta.env.DEV) {
console.warn(`⚠️ Skipped: ${statKey} | "${r.fromId}" → "${r.toId}" | ${err instanceof Error ? err.message : String(err)}`);
}
}
// Yield to event loop every 500 relations
if (i > 0 && i % 500 === 0) {
await new Promise(r => setTimeout(r, 0));
}
}
} finally {
await stmt.close();
}
}
@@ -190,8 +283,8 @@ export const loadGraphToLbug = async (
let totalNodes = 0;
for (const tableName of NODE_TABLES) {
try {
const countRes = await conn.query(`MATCH (n:${tableName}) RETURN count(n) AS cnt`);
const countRows = await getQueryRows(countRes);
const countRes = await conn.query(`MATCH (n:${escapeTableName(tableName)}) RETURN count(n) AS cnt`);
const countRows = await countRes.getAllRows();
const countRow = countRows[0];
const count = countRow ? (countRow.cnt ?? countRow[0] ?? 0) : 0;
totalNodes += Number(count);
@@ -225,12 +318,20 @@ const BACKTICK_TABLES = new Set([
'Struct', 'Enum', 'Macro', 'Typedef', 'Union', 'Namespace', 'Trait', 'Impl',
'TypeAlias', 'Const', 'Static', 'Property', 'Record', 'Delegate', 'Annotation',
'Constructor', 'Template', 'Module',
// Reserved/ambiguous identifiers that need quoting
'File',
]);
const escapeTableName = (table: string): string => {
return BACKTICK_TABLES.has(table) ? `\`${table}\`` : table;
};
// LadybugDB DELETE needs standard quoted identifiers for reserved names (e.g., File)
const escapeTableForDelete = (table: string): string => {
if (table === 'File') return `"${table}"`;
return escapeTableName(table);
};
/** Tables with isExported column (TypeScript/JS-native types) */
const TABLES_WITH_EXPORTED = new Set<string>(['Function', 'Class', 'Interface', 'Method', 'CodeElement']);
@@ -263,11 +364,20 @@ const getCopyQuery = (table: NodeTableName, path: string): string => {
* Execute a Cypher query against the database
* Returns results as named objects (not tuples) for better usability
*/
export const executeQuery = async (cypher: string): Promise<any[]> => {
export const executeQuery = async (cypher: string, readOnly = true): Promise<any[]> => {
if (!conn) {
await initLbug();
}
if (readOnly) {
// Strip quoted strings before checking for write keywords, so that
// queries like WHERE n.name CONTAINS "delete" are not blocked.
const stripped = cypher.replace(/'[^']*'|"[^"]*"/g, '').toUpperCase();
if (/\b(CREATE|DELETE|SET|MERGE|REMOVE|DROP|DETACH)\b/.test(stripped)) {
throw new Error('Read-only query attempted a write operation');
}
}
try {
const result = await conn.query(cypher);
@@ -295,7 +405,7 @@ export const executeQuery = async (cypher: string): Promise<any[]> => {
}
// Collect all rows
const allRows = await getQueryRows(result);
const allRows = await result.getAllRows();
const rows: any[] = [];
for (const row of allRows) {
// Convert tuple to named object if we have column names and row is array
@@ -331,8 +441,8 @@ export const getLbugStats = async (): Promise<{ nodes: number; edges: number }>
let totalNodes = 0;
for (const tableName of NODE_TABLES) {
try {
const nodeResult = await conn.query(`MATCH (n:${tableName}) RETURN count(n) AS cnt`);
const nodeRows = await getQueryRows(nodeResult);
const nodeResult = await conn.query(`MATCH (n:${escapeTableName(tableName)}) RETURN count(n) AS cnt`);
const nodeRows = await nodeResult.getAllRows();
const nodeRow = nodeRows[0];
totalNodes += Number(nodeRow?.cnt ?? nodeRow?.[0] ?? 0);
} catch {
@@ -344,7 +454,7 @@ export const getLbugStats = async (): Promise<{ nodes: number; edges: number }>
let totalEdges = 0;
try {
const edgeResult = await conn.query(`MATCH ()-[r:${REL_TABLE_NAME}]->() RETURN count(r) AS cnt`);
const edgeRows = await getQueryRows(edgeResult);
const edgeRows = await edgeResult.getAllRows();
const edgeRow = edgeRows[0];
totalEdges = Number(edgeRow?.cnt ?? edgeRow?.[0] ?? 0);
} catch {
@@ -384,6 +494,7 @@ export const closeLbug = async (): Promise<void> => {
db = null;
}
lbug = null;
initPromise = null;
};
/**
@@ -402,17 +513,18 @@ export const executePrepared = async (
try {
const stmt = await conn.prepare(cypher);
if (!stmt.isSuccess()) {
const errMsg = await stmt.getErrorMessage();
throw new Error(`Prepare failed: ${errMsg}`);
try {
if (!stmt.isSuccess()) {
const errMsg = await stmt.getErrorMessage();
throw new Error(`Prepare failed: ${errMsg}`);
}
const result = await conn.execute(stmt, params);
const rows = await result.getAllRows();
return rows;
} finally {
await stmt.close();
}
const result = await conn.execute(stmt, params);
const rows = await getQueryRows(result);
await stmt.close();
return rows;
} catch (error) {
if (import.meta.env.DEV) console.error('Prepared query failed:', error);
throw error;
@@ -472,8 +584,8 @@ export const testArrayParams = async (): Promise<{ success: boolean; error?: str
let testNodeId: string | null = null;
for (const tableName of NODE_TABLES) {
try {
const nodeResult = await conn.query(`MATCH (n:${tableName}) RETURN n.id AS id LIMIT 1`);
const nodeRows = await getQueryRows(nodeResult);
const nodeResult = await conn.query(`MATCH (n:${escapeTableName(tableName)}) RETURN n.id AS id LIMIT 1`);
const nodeRows = await nodeResult.getAllRows();
const nodeRow = nodeRows[0];
if (nodeRow) {
testNodeId = nodeRow.id ?? nodeRow[0];
@@ -506,24 +618,39 @@ export const testArrayParams = async (): Promise<{ success: boolean; error?: str
await stmt.close();
// Verify it was stored
const verifyResult = await conn.query(
`MATCH (e:${EMBEDDING_TABLE_NAME} {nodeId: '${testNodeId}'}) RETURN e.embedding AS emb`
// Verify it was stored (using prepared statement to avoid injection)
const verifyStmt = await conn.prepare(
`MATCH (e:${EMBEDDING_TABLE_NAME} {nodeId: $nodeId}) RETURN e.embedding AS emb`
);
const verifyRows = await getQueryRows(verifyResult);
const verifyRow = verifyRows[0];
const storedEmb = verifyRow?.emb ?? verifyRow?.[0];
if (storedEmb && Array.isArray(storedEmb) && storedEmb.length === 384) {
if (import.meta.env.DEV) {
console.log('✅ Array params WORK! Stored embedding length:', storedEmb.length);
try {
if (!verifyStmt.isSuccess()) {
const errMsg = await verifyStmt.getErrorMessage();
return { success: false, error: `Verify prepare failed: ${errMsg}` };
}
return { success: true };
} else {
return {
success: false,
error: `Embedding not stored correctly. Got: ${typeof storedEmb}, length: ${storedEmb?.length}`
};
const verifyResult = await conn.execute(verifyStmt, { nodeId: testNodeId });
const verifyRows = await verifyResult.getAllRows();
const verifyRow = verifyRows[0];
const storedEmb = verifyRow?.emb ?? verifyRow?.[0];
// Clean up test embedding
try {
const cleanupStmt = await conn.prepare(`MATCH (e:${EMBEDDING_TABLE_NAME} {nodeId: $nodeId}) DELETE e`);
try { await conn.execute(cleanupStmt, { nodeId: testNodeId }); } finally { await cleanupStmt.close(); }
} catch {}
if (storedEmb && Array.isArray(storedEmb) && storedEmb.length === 384) {
if (import.meta.env.DEV) {
console.log('✅ Array params WORK! Stored embedding length:', storedEmb.length);
}
return { success: true };
} else {
return {
success: false,
error: `Embedding not stored correctly. Got: ${typeof storedEmb}, length: ${storedEmb?.length}`
};
}
} finally {
await verifyStmt.close();
}
} catch (error) {
const errorMsg = error instanceof Error ? error.message : String(error);
+1 -1
View File
@@ -26,7 +26,7 @@ export type NodeTableName = typeof NODE_TABLES[number];
export const REL_TABLE_NAME = 'CodeRelation';
// Valid relation types
export const REL_TYPES = ['CONTAINS', 'DEFINES', 'IMPORTS', 'CALLS', 'EXTENDS', 'IMPLEMENTS', 'MEMBER_OF', 'STEP_IN_PROCESS'] as const;
export const REL_TYPES = ['CONTAINS', 'DEFINES', 'IMPORTS', 'CALLS', 'EXTENDS', 'IMPLEMENTS', 'MEMBER_OF', 'STEP_IN_PROCESS', 'HAS_METHOD', 'HAS_PROPERTY', 'OVERRIDES', 'ACCESSES', 'INHERITS', 'USES', 'DECORATES'] as const;
export type RelType = typeof REL_TYPES[number];
// ============================================================================
+39 -12
View File
@@ -22,12 +22,14 @@ import type {
OllamaConfig,
OpenRouterConfig,
MiniMaxConfig,
GLMConfig,
AgentStreamChunk,
} from './types';
import {
type CodebaseContext,
buildDynamicSystemPrompt,
} from './context-builder';
import { DEFAULT_OLLAMA_BASE_URL, DEFAULT_OPENROUTER_BASE_URL } from '../../config/ui-constants';
/**
* System prompt for the Graph RAG agent
@@ -184,7 +186,7 @@ export const createChatModel = (config: ProviderConfig): BaseChatModel => {
case 'ollama': {
const ollamaConfig = config as OllamaConfig;
return new ChatOllama({
baseUrl: ollamaConfig.baseUrl ?? 'http://localhost:11434',
baseUrl: ollamaConfig.baseUrl ?? DEFAULT_OLLAMA_BASE_URL,
model: ollamaConfig.model,
temperature: ollamaConfig.temperature ?? 0.1,
streaming: true,
@@ -203,7 +205,6 @@ export const createChatModel = (config: ProviderConfig): BaseChatModel => {
if (import.meta.env.DEV) {
console.log('🌐 OpenRouter config:', {
hasApiKey: !!openRouterConfig.apiKey,
apiKeyLength: openRouterConfig.apiKey?.length || 0,
model: openRouterConfig.model,
baseUrl: openRouterConfig.baseUrl,
});
@@ -221,7 +222,7 @@ export const createChatModel = (config: ProviderConfig): BaseChatModel => {
maxTokens: openRouterConfig.maxTokens,
configuration: {
apiKey: openRouterConfig.apiKey, // Ensure client receives it
baseURL: openRouterConfig.baseUrl ?? 'https://openrouter.ai/api/v1',
baseURL: openRouterConfig.baseUrl ?? DEFAULT_OPENROUTER_BASE_URL,
},
streaming: true,
});
@@ -246,6 +247,26 @@ export const createChatModel = (config: ProviderConfig): BaseChatModel => {
});
}
case 'glm': {
const glmConfig = config as GLMConfig;
if (!glmConfig.apiKey || glmConfig.apiKey.trim() === '') {
throw new Error('GLM API key is required but was not provided');
}
return new ChatOpenAI({
apiKey: glmConfig.apiKey,
modelName: glmConfig.model,
temperature: glmConfig.temperature ?? 0.1,
maxTokens: glmConfig.maxTokens,
configuration: {
apiKey: glmConfig.apiKey,
baseURL: glmConfig.baseUrl ?? 'https://api.z.ai/api/coding/paas/v4',
},
streaming: true,
});
}
default:
throw new Error(`Unsupported provider: ${(config as any).provider}`);
}
@@ -355,8 +376,8 @@ export async function* streamAgentResponse(
const yieldedToolCalls = new Set<string>();
const yieldedToolResults = new Set<string>();
let lastProcessedMsgCount = formattedMessages.length;
// Track if all tools are done (for distinguishing reasoning vs final content)
let allToolsDone = true;
// Track pending tool calls (for distinguishing reasoning vs final content)
let pendingToolCalls = 0;
// Track if we've seen any tool calls in this response turn.
// Anything before the first tool call should be treated as "reasoning/narration"
// so the UI can show the Cursor-like loop: plan → tool → update → tool → answer.
@@ -420,7 +441,7 @@ export async function* streamAgentResponse(
const isReasoning =
!hasSeenToolCallThisTurn ||
toolCalls.length > 0 ||
!allToolsDone;
pendingToolCalls > 0;
yield {
type: isReasoning ? 'reasoning' : 'content',
[isReasoning ? 'reasoning' : 'content']: content,
@@ -430,17 +451,23 @@ export async function* streamAgentResponse(
// Track tool calls from message chunks
if (toolCalls.length > 0) {
hasSeenToolCallThisTurn = true;
allToolsDone = false;
pendingToolCalls += toolCalls.length;
for (const tc of toolCalls) {
const toolId = tc.id || `tool-${Date.now()}-${Math.random().toString(36).slice(2)}`;
if (!yieldedToolCalls.has(toolId)) {
yieldedToolCalls.add(toolId);
let parsedArgs: Record<string, any>;
try {
parsedArgs = tc.function?.arguments ? JSON.parse(tc.function.arguments) : {};
} catch {
parsedArgs = {};
}
yield {
type: 'tool_call',
toolCall: {
id: toolId,
name: tc.name || tc.function?.name || 'unknown',
args: tc.args || (tc.function?.arguments ? JSON.parse(tc.function.arguments) : {}),
args: tc.args || parsedArgs,
status: 'running',
},
};
@@ -465,8 +492,8 @@ export async function* streamAgentResponse(
status: 'completed',
},
};
// After tool result, next AI content could be reasoning or final
allToolsDone = true;
// After tool result, decrement pending count
pendingToolCalls = Math.max(0, pendingToolCalls - 1);
}
}
}
@@ -486,7 +513,7 @@ export async function* streamAgentResponse(
for (const tc of toolCalls) {
const toolId = tc.id || `tool-${Date.now()}`;
if (!yieldedToolCalls.has(toolId)) {
allToolsDone = false;
pendingToolCalls++;
yieldedToolCalls.add(toolId);
yield {
type: 'tool_call',
@@ -517,7 +544,7 @@ export async function* streamAgentResponse(
status: 'completed',
},
};
allToolsDone = true;
pendingToolCalls = Math.max(0, pendingToolCalls - 1);
}
}
}
+164 -113
View File
@@ -16,56 +16,92 @@ import {
OllamaConfig,
OpenRouterConfig,
MiniMaxConfig,
GLMConfig,
ProviderConfig,
} from './types';
import { DEFAULT_OPENROUTER_BASE_URL, DEFAULT_OLLAMA_BASE_URL } from '../../config/ui-constants';
const STORAGE_KEY = 'gitnexus-llm-settings';
const mergeWithDefaults = (parsed?: Partial<LLMSettings> | null): LLMSettings => ({
...DEFAULT_LLM_SETTINGS,
...parsed,
openai: {
...DEFAULT_LLM_SETTINGS.openai,
...parsed?.openai,
},
azureOpenAI: {
...DEFAULT_LLM_SETTINGS.azureOpenAI,
...parsed?.azureOpenAI,
},
gemini: {
...DEFAULT_LLM_SETTINGS.gemini,
...parsed?.gemini,
},
anthropic: {
...DEFAULT_LLM_SETTINGS.anthropic,
...parsed?.anthropic,
},
ollama: {
...DEFAULT_LLM_SETTINGS.ollama,
...parsed?.ollama,
},
openrouter: {
...DEFAULT_LLM_SETTINGS.openrouter,
...parsed?.openrouter,
},
minimax: {
...DEFAULT_LLM_SETTINGS.minimax,
...parsed?.minimax,
},
glm: {
...DEFAULT_LLM_SETTINGS.glm,
...parsed?.glm,
},
});
const readSettings = (storage: Storage): Partial<LLMSettings> | null => {
const raw = storage.getItem(STORAGE_KEY);
if (!raw) return null;
try {
return JSON.parse(raw) as Partial<LLMSettings>;
} catch (error) {
console.warn('Failed to parse LLM settings:', error);
return null;
}
};
const writeSettings = (storage: Storage, settings: LLMSettings): void => {
storage.setItem(STORAGE_KEY, JSON.stringify(settings));
};
/**
* Load settings from localStorage
* Load settings from sessionStorage (migrates legacy localStorage once).
*/
export const loadSettings = (): LLMSettings => {
try {
const stored = localStorage.getItem(STORAGE_KEY);
if (!stored) {
return DEFAULT_LLM_SETTINGS;
const sessionData = typeof sessionStorage !== 'undefined' ? readSettings(sessionStorage) : null;
if (sessionData) {
return mergeWithDefaults(sessionData);
}
const parsed = JSON.parse(stored) as Partial<LLMSettings>;
// Merge with defaults to handle new fields
return {
...DEFAULT_LLM_SETTINGS,
...parsed,
openai: {
...DEFAULT_LLM_SETTINGS.openai,
...parsed.openai,
},
azureOpenAI: {
...DEFAULT_LLM_SETTINGS.azureOpenAI,
...parsed.azureOpenAI,
},
gemini: {
...DEFAULT_LLM_SETTINGS.gemini,
...parsed.gemini,
},
anthropic: {
...DEFAULT_LLM_SETTINGS.anthropic,
...parsed.anthropic,
},
ollama: {
...DEFAULT_LLM_SETTINGS.ollama,
...parsed.ollama,
},
openrouter: {
...DEFAULT_LLM_SETTINGS.openrouter,
...parsed.openrouter,
},
minimax: {
...DEFAULT_LLM_SETTINGS.minimax,
...parsed.minimax,
},
};
const legacyData = typeof localStorage !== 'undefined' ? readSettings(localStorage) : null;
if (legacyData) {
const merged = mergeWithDefaults(legacyData);
try {
if (typeof sessionStorage !== 'undefined') {
writeSettings(sessionStorage, merged);
}
if (typeof localStorage !== 'undefined') {
localStorage.removeItem(STORAGE_KEY);
}
} catch (error) {
console.warn('Failed to migrate legacy LLM settings to sessionStorage:', error);
}
return merged;
}
return DEFAULT_LLM_SETTINGS;
} catch (error) {
console.warn('Failed to load LLM settings:', error);
return DEFAULT_LLM_SETTINGS;
@@ -73,11 +109,13 @@ export const loadSettings = (): LLMSettings => {
};
/**
* Save settings to localStorage
* Save settings to sessionStorage
*/
export const saveSettings = (settings: LLMSettings): void => {
try {
localStorage.setItem(STORAGE_KEY, JSON.stringify(settings));
if (typeof sessionStorage !== 'undefined') {
writeSettings(sessionStorage, settings);
}
} catch (error) {
console.error('Failed to save LLM settings:', error);
}
@@ -94,7 +132,9 @@ export const updateProviderSettings = <T extends LLMProvider>(
T extends 'gemini' ? Partial<Omit<GeminiConfig, 'provider'>> :
T extends 'anthropic' ? Partial<Omit<AnthropicConfig, 'provider'>> :
T extends 'ollama' ? Partial<Omit<OllamaConfig, 'provider'>> :
T extends 'openrouter' ? Partial<Omit<OpenRouterConfig, 'provider'>> :
T extends 'minimax' ? Partial<Omit<MiniMaxConfig, 'provider'>> :
T extends 'glm' ? Partial<Omit<GLMConfig, 'provider'>> :
never
>
): LLMSettings => {
@@ -179,6 +219,17 @@ export const updateProviderSettings = <T extends LLMProvider>(
saveSettings(updated);
return updated;
}
case 'glm': {
const updated: LLMSettings = {
...current,
glm: {
...(current.glm ?? {}),
...(updates as Partial<Omit<GLMConfig, 'provider'>>),
},
};
saveSettings(updated);
return updated;
}
default: {
// Should be unreachable due to T extends LLMProvider, but keep a safe fallback
const updated: LLMSettings = { ...current };
@@ -204,77 +255,64 @@ export const setActiveProvider = (provider: LLMProvider): LLMSettings => {
/**
* Get the current provider configuration
*/
type ProviderBuilder = (settings: LLMSettings) => ProviderConfig | null;
const providerBuilders: Record<LLMProvider, ProviderBuilder> = {
openai: (settings) => {
if (!settings.openai?.apiKey) return null;
return { provider: 'openai', ...settings.openai } as OpenAIConfig;
},
'azure-openai': (settings) => {
if (!settings.azureOpenAI?.apiKey || !settings.azureOpenAI?.endpoint) return null;
return { provider: 'azure-openai', ...settings.azureOpenAI } as AzureOpenAIConfig;
},
gemini: (settings) => {
if (!settings.gemini?.apiKey) return null;
return { provider: 'gemini', ...settings.gemini } as GeminiConfig;
},
anthropic: (settings) => {
if (!settings.anthropic?.apiKey) return null;
return { provider: 'anthropic', ...settings.anthropic } as AnthropicConfig;
},
ollama: (settings) => {
return {
provider: 'ollama',
...settings.ollama,
baseUrl: settings.ollama?.baseUrl ?? DEFAULT_OLLAMA_BASE_URL,
} as OllamaConfig;
},
openrouter: (settings) => {
if (!settings.openrouter?.apiKey || settings.openrouter.apiKey.trim() === '') return null;
return {
provider: 'openrouter',
apiKey: settings.openrouter.apiKey,
model: settings.openrouter.model || '',
baseUrl: settings.openrouter.baseUrl || DEFAULT_OPENROUTER_BASE_URL,
temperature: settings.openrouter.temperature,
maxTokens: settings.openrouter.maxTokens,
} as OpenRouterConfig;
},
minimax: (settings) => {
if (!settings.minimax?.apiKey) return null;
return { provider: 'minimax', ...settings.minimax } as MiniMaxConfig;
},
glm: (settings) => {
if (!settings.glm?.apiKey) return null;
return {
provider: 'glm',
apiKey: settings.glm.apiKey,
model: settings.glm.model || 'GLM-5',
baseUrl: settings.glm.baseUrl || 'https://api.z.ai/api/coding/paas/v4',
temperature: settings.glm.temperature,
maxTokens: settings.glm.maxTokens,
} as GLMConfig;
},
};
export const getActiveProviderConfig = (): ProviderConfig | null => {
const settings = loadSettings();
switch (settings.activeProvider) {
case 'openai':
if (!settings.openai?.apiKey) {
return null;
}
return {
provider: 'openai',
...settings.openai,
} as OpenAIConfig;
case 'azure-openai':
if (!settings.azureOpenAI?.apiKey || !settings.azureOpenAI?.endpoint) {
return null;
}
return {
provider: 'azure-openai',
...settings.azureOpenAI,
} as AzureOpenAIConfig;
case 'gemini':
if (!settings.gemini?.apiKey) {
return null;
}
return {
provider: 'gemini',
...settings.gemini,
} as GeminiConfig;
case 'anthropic':
if (!settings.anthropic?.apiKey) {
return null;
}
return {
provider: 'anthropic',
...settings.anthropic,
} as AnthropicConfig;
case 'ollama':
return {
provider: 'ollama',
...settings.ollama,
} as OllamaConfig;
case 'openrouter':
if (!settings.openrouter?.apiKey || settings.openrouter.apiKey.trim() === '') {
return null;
}
return {
provider: 'openrouter',
apiKey: settings.openrouter.apiKey,
model: settings.openrouter.model || '',
baseUrl: settings.openrouter.baseUrl || 'https://openrouter.ai/api/v1',
temperature: settings.openrouter.temperature,
maxTokens: settings.openrouter.maxTokens,
} as OpenRouterConfig;
case 'minimax':
if (!settings.minimax?.apiKey) {
return null;
}
return {
provider: 'minimax',
...settings.minimax,
} as MiniMaxConfig;
default:
return null;
}
const builder = providerBuilders[settings.activeProvider];
return builder ? builder(settings) : null;
};
/**
@@ -288,7 +326,16 @@ export const isProviderConfigured = (): boolean => {
* Clear all settings (reset to defaults)
*/
export const clearSettings = (): void => {
localStorage.removeItem(STORAGE_KEY);
try {
if (typeof sessionStorage !== 'undefined') {
sessionStorage.removeItem(STORAGE_KEY);
}
if (typeof localStorage !== 'undefined') {
localStorage.removeItem(STORAGE_KEY);
}
} catch (error) {
console.warn('Failed to clear LLM settings:', error);
}
};
/**
@@ -310,6 +357,8 @@ export const getProviderDisplayName = (provider: LLMProvider): string => {
return 'OpenRouter';
case 'minimax':
return 'MiniMax';
case 'glm':
return 'GLM (Z.AI)';
default:
return provider;
}
@@ -333,6 +382,8 @@ export const getAvailableModels = (provider: LLMProvider): string[] => {
return ['llama3.2', 'llama3.1', 'mistral', 'codellama', 'deepseek-coder'];
case 'minimax':
return ['MiniMax-M2.5', 'MiniMax-M2.5-highspeed'];
case 'glm':
return ['GLM-5', 'GLM-5-Turbo', 'GLM-4.7', 'GLM-4.5'];
default:
return [];
}
@@ -343,7 +394,7 @@ export const getAvailableModels = (provider: LLMProvider): string[] => {
*/
export const fetchOpenRouterModels = async (): Promise<Array<{ id: string; name: string }>> => {
try {
const response = await fetch('https://openrouter.ai/api/v1/models');
const response = await fetch(`${DEFAULT_OPENROUTER_BASE_URL}/models`);
if (!response.ok) throw new Error('Failed to fetch models');
const data = await response.json();
return data.data.map((model: any) => ({
+22 -5
View File
@@ -15,6 +15,13 @@ import { tool } from '@langchain/core/tools';
import { z } from 'zod';
// Note: GRAPH_SCHEMA_DESCRIPTION from './types' is available if needed for additional context
import { WebGPUNotAvailableError, embedText, embeddingToArray, initEmbedder, isEmbedderReady } from '../embeddings/embedder';
import { NODE_TABLES, REL_TYPES } from '../lbug/schema';
const validLabel = (label: string): boolean =>
(NODE_TABLES as readonly string[]).includes(label);
const validRelType = (t: string): boolean =>
(REL_TYPES as readonly string[]).includes(t);
/**
* Tool factory - creates tools bound to the LadybugDB query functions
@@ -96,11 +103,12 @@ export const createGraphRAGTools = (
if (nodeId) {
try {
const nodeLabel = nodeId.split(':')[0];
if (!validLabel(nodeLabel)) throw new Error('invalid label');
const connectionsQuery = `
MATCH (n:${nodeLabel} {id: '${nodeId.replace(/'/g, "''")}'})
OPTIONAL MATCH (n)-[r1:CodeRelation]->(dst)
OPTIONAL MATCH (src)-[r2:CodeRelation]->(n)
RETURN
RETURN
collect(DISTINCT {name: dst.name, type: r1.type, confidence: r1.confidence}) AS outgoing,
collect(DISTINCT {name: src.name, type: r2.type, confidence: r2.confidence}) AS incoming
LIMIT 1
@@ -136,6 +144,7 @@ export const createGraphRAGTools = (
if (nodeId) {
try {
const nodeLabel = nodeId.split(':')[0];
if (!validLabel(nodeLabel)) throw new Error('invalid label');
const clusterQuery = `
MATCH (n:${nodeLabel} {id: '${nodeId.replace(/'/g, "''")}'})
MATCH (n)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
@@ -158,6 +167,7 @@ export const createGraphRAGTools = (
if (nodeId) {
try {
const nodeLabel = nodeId.split(':')[0];
if (!validLabel(nodeLabel)) throw new Error('invalid label');
const processQuery = `
MATCH (n:${nodeLabel} {id: '${nodeId.replace(/'/g, "''")}'})
MATCH (n)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
@@ -783,7 +793,11 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`,
const name = getRowValue(symbolRow, 1, 'name');
const filePath = getRowValue(symbolRow, 2, 'filePath');
const nodeType = getRowValue(symbolRow, 3, 'nodeType');
if (!validLabel(nodeType)) {
return `Unknown node type "${nodeType}" for symbol "${target}".`;
}
const clusterQuery = `
MATCH (n:${nodeType} {id: '${String(nodeId).replace(/'/g, "''")}'})
MATCH (n)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
@@ -898,10 +912,13 @@ MATCH (n:Function {id: emb.nodeId}) RETURN n`,
// Default to usage-based relation types (exclude CONTAINS, DEFINES for impact analysis)
const defaultRelTypes = ['CALLS', 'IMPORTS', 'EXTENDS', 'IMPLEMENTS'];
const activeRelTypes = relationTypes && relationTypes.length > 0
? relationTypes
const activeRelTypes = relationTypes && relationTypes.length > 0
? relationTypes.filter(t => validRelType(t))
: defaultRelTypes;
const relTypeFilter = activeRelTypes.map(t => `'${t}'`).join(', ');
if (activeRelTypes.length === 0) {
return `No valid relation types provided. Valid types: ${(REL_TYPES as readonly string[]).join(', ')}`;
}
const relTypeFilter = activeRelTypes.map(t => `'${t.replace(/'/g, "''")}'`).join(', ');
const directionLabel = direction === 'upstream'
? 'Files that DEPEND ON this (breakage risk)'
+23 -5
View File
@@ -2,13 +2,14 @@
* LLM Provider Types
*
* Type definitions for multi-provider LLM support.
* Supports Azure OpenAI and Google Gemini (with extensibility for others).
* Supports OpenAI, Azure OpenAI, Gemini, Anthropic, Ollama, OpenRouter, MiniMax, and GLM5.
*/
/**
* Supported LLM providers
*/
export type LLMProvider = 'openai' | 'azure-openai' | 'gemini' | 'anthropic' | 'ollama' | 'openrouter' | 'minimax';
import { DEFAULT_OLLAMA_BASE_URL, DEFAULT_OPENROUTER_BASE_URL } from '../../config/ui-constants';
export type LLMProvider = 'openai' | 'azure-openai' | 'gemini' | 'anthropic' | 'ollama' | 'openrouter' | 'minimax' | 'glm';
/**
* Base configuration shared by all providers
@@ -87,10 +88,20 @@ export interface MiniMaxConfig extends BaseProviderConfig {
model: string; // e.g., 'MiniMax-M2.5', 'MiniMax-M2.5-highspeed'
}
/**
* GLM (Z.AI) configuration — OpenAI-compatible API
*/
export interface GLMConfig extends BaseProviderConfig {
provider: 'glm';
apiKey: string;
model: string; // e.g., 'GLM-4.7', 'GLM-4.5', 'GLM-4.5-Air', 'GLM-5'
baseUrl?: string; // defaults to https://api.z.ai/api/coding/paas/v4
}
/**
* Union type for all provider configurations
*/
export type ProviderConfig = OpenAIConfig | AzureOpenAIConfig | GeminiConfig | AnthropicConfig | OllamaConfig | OpenRouterConfig | MiniMaxConfig;
export type ProviderConfig = OpenAIConfig | AzureOpenAIConfig | GeminiConfig | AnthropicConfig | OllamaConfig | OpenRouterConfig | MiniMaxConfig | GLMConfig;
/**
* Stored settings (what goes to localStorage)
@@ -108,6 +119,7 @@ export interface LLMSettings {
ollama?: Partial<Omit<OllamaConfig, 'provider'>>;
openrouter?: Partial<Omit<OpenRouterConfig, 'provider'>>;
minimax?: Partial<Omit<MiniMaxConfig, 'provider'>>;
glm?: Partial<Omit<GLMConfig, 'provider'>>;
// Intelligent Clustering Settings
intelligentClustering: boolean;
@@ -148,14 +160,14 @@ export const DEFAULT_LLM_SETTINGS: LLMSettings = {
temperature: 0.1,
},
ollama: {
baseUrl: 'http://localhost:11434',
baseUrl: DEFAULT_OLLAMA_BASE_URL,
model: 'llama3.2',
temperature: 0.1,
},
openrouter: {
apiKey: '',
model: '',
baseUrl: 'https://openrouter.ai/api/v1',
baseUrl: DEFAULT_OPENROUTER_BASE_URL,
temperature: 0.1,
},
minimax: {
@@ -163,6 +175,12 @@ export const DEFAULT_LLM_SETTINGS: LLMSettings = {
model: 'MiniMax-M2.5',
temperature: 0.1,
},
glm: {
apiKey: '',
model: 'GLM-5',
baseUrl: 'https://api.z.ai/api/coding/paas/v4',
temperature: 0.1,
},
};
/**
@@ -0,0 +1,75 @@
import { createContext, useContext, useCallback, useMemo, useState, ReactNode } from 'react';
import type { KnowledgeGraph, GraphNode, NodeLabel } from '../../core/graph/types';
import { DEFAULT_VISIBLE_LABELS, DEFAULT_VISIBLE_EDGES, type EdgeType } from '../../lib/constants';
interface GraphStateContextValue {
graph: KnowledgeGraph | null;
setGraph: (graph: KnowledgeGraph | null) => void;
fileContents: Map<string, string>;
setFileContents: (contents: Map<string, string>) => void;
selectedNode: GraphNode | null;
setSelectedNode: (node: GraphNode | null) => void;
visibleLabels: NodeLabel[];
toggleLabelVisibility: (label: NodeLabel) => void;
visibleEdgeTypes: EdgeType[];
toggleEdgeVisibility: (edgeType: EdgeType) => void;
depthFilter: number | null;
setDepthFilter: (depth: number | null) => void;
highlightedNodeIds: Set<string>;
setHighlightedNodeIds: (ids: Set<string>) => void;
}
const GraphStateContext = createContext<GraphStateContextValue | null>(null);
export const GraphStateProvider = ({ children }: { children: ReactNode }) => {
const [graph, setGraph] = useState<KnowledgeGraph | null>(null);
const [fileContents, setFileContents] = useState<Map<string, string>>(new Map());
const [selectedNode, setSelectedNode] = useState<GraphNode | null>(null);
const [visibleLabels, setVisibleLabels] = useState<NodeLabel[]>(DEFAULT_VISIBLE_LABELS);
const [visibleEdgeTypes, setVisibleEdgeTypes] = useState<EdgeType[]>(DEFAULT_VISIBLE_EDGES);
const [depthFilter, setDepthFilter] = useState<number | null>(null);
const [highlightedNodeIds, setHighlightedNodeIds] = useState<Set<string>>(new Set());
const toggleLabelVisibility = useCallback((label: NodeLabel) => {
setVisibleLabels(prev =>
prev.includes(label) ? prev.filter(l => l !== label) : [...prev, label]
);
}, []);
const toggleEdgeVisibility = useCallback((edgeType: EdgeType) => {
setVisibleEdgeTypes(prev =>
prev.includes(edgeType) ? prev.filter(e => e !== edgeType) : [...prev, edgeType]
);
}, []);
const value = useMemo<GraphStateContextValue>(() => ({
graph,
setGraph,
fileContents,
setFileContents,
selectedNode,
setSelectedNode,
visibleLabels,
toggleLabelVisibility,
visibleEdgeTypes,
toggleEdgeVisibility,
depthFilter,
setDepthFilter,
highlightedNodeIds,
setHighlightedNodeIds,
}), [graph, fileContents, selectedNode, visibleLabels, visibleEdgeTypes, depthFilter, highlightedNodeIds]);
return (
<GraphStateContext.Provider value={value}>
{children}
</GraphStateContext.Provider>
);
};
export const useGraphState = (): GraphStateContextValue => {
const ctx = useContext(GraphStateContext);
if (!ctx) {
throw new Error('useGraphState must be used within a GraphStateProvider');
}
return ctx;
};
+98 -114
View File
@@ -1,18 +1,21 @@
import { createContext, useContext, useState, useCallback, useRef, useEffect, ReactNode } from 'react';
import { createContext, useContext, useState, useCallback, useRef, useEffect, useMemo, ReactNode } from 'react';
import * as Comlink from 'comlink';
import { KnowledgeGraph, GraphNode, GraphRelationship, NodeLabel } from '../core/graph/types';
import { PipelineProgress, PipelineResult, deserializePipelineResult } from '../types/pipeline';
import { createKnowledgeGraph } from '../core/graph/graph';
import { DEFAULT_VISIBLE_LABELS } from '../lib/constants';
import type { IngestionWorkerApi } from '../workers/ingestion.worker';
import type { FileEntry } from '../services/zip';
import type { EmbeddingProgress, SemanticSearchResult } from '../core/embeddings/types';
import type { LLMSettings, ProviderConfig, AgentStreamChunk, ChatMessage, ToolCallInfo, MessageStep } from '../core/llm/types';
import { loadSettings, getActiveProviderConfig, saveSettings } from '../core/llm/settings-service';
import type { AgentMessage } from '../core/llm/agent';
import { DEFAULT_VISIBLE_EDGES, type EdgeType } from '../lib/constants';
import { type EdgeType } from '../lib/constants';
import type { RepoSummary, ConnectToServerResult } from '../services/server-connection';
import { fetchRepos, connectToServer } from '../services/server-connection';
import { ERROR_RESET_DELAY_MS } from '../config/ui-constants';
import { normalizePath, resolveFilePath as resolvePathFromContents } from '../lib/path-resolution';
import { FILE_REF_REGEX, NODE_REF_REGEX } from '../lib/grounding-patterns';
import { GraphStateProvider, useGraphState } from './app-state/graph';
export type ViewMode = 'onboarding' | 'loading' | 'exploring';
export type RightPanelTab = 'code' | 'chat';
@@ -74,6 +77,8 @@ interface AppState {
setRightPanelTab: (tab: RightPanelTab) => void;
openCodePanel: () => void;
openChatPanel: () => void;
helpDialogBoxOpen: boolean;
setHelpDialogBoxOpen: (open: boolean) => void;
// Filters
visibleLabels: NodeLabel[];
@@ -134,6 +139,7 @@ interface AppState {
// Embedding methods
startEmbeddings: (forceDevice?: 'webgpu' | 'wasm') => Promise<void>;
startEmbeddingsWithFallback: () => void;
semanticSearch: (query: string, k?: number) => Promise<SemanticSearchResult[]>;
semanticSearchWithContext: (query: string, k?: number, hops?: number) => Promise<any[]>;
isEmbeddingReady: boolean;
@@ -145,9 +151,7 @@ interface AppState {
llmSettings: LLMSettings;
updateLLMSettings: (updates: Partial<LLMSettings>) => void;
isSettingsPanelOpen: boolean;
isHelpDialogBoxOpen: boolean;
setSettingsPanelOpen: (open: boolean) => void;
setHelpDialogBoxOpen: (open: boolean) => void;
isAgentReady: boolean;
isAgentInitializing: boolean;
agentError: string | null;
@@ -177,20 +181,37 @@ interface AppState {
const AppStateContext = createContext<AppState | null>(null);
export const AppStateProvider = ({ children }: { children: ReactNode }) => {
export const AppStateProvider = ({ children }: { children: ReactNode }) => (
<GraphStateProvider>
<AppStateProviderInner>{children}</AppStateProviderInner>
</GraphStateProvider>
);
const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
// View state
const [viewMode, setViewMode] = useState<ViewMode>('onboarding');
// Graph data
const [graph, setGraph] = useState<KnowledgeGraph | null>(null);
const [fileContents, setFileContents] = useState<Map<string, string>>(new Map());
// Selection
const [selectedNode, setSelectedNode] = useState<GraphNode | null>(null);
const {
graph,
setGraph,
fileContents,
setFileContents,
selectedNode,
setSelectedNode,
visibleLabels,
toggleLabelVisibility,
visibleEdgeTypes,
toggleEdgeVisibility,
depthFilter,
setDepthFilter,
highlightedNodeIds,
setHighlightedNodeIds,
} = useGraphState();
// Right Panel
const [isRightPanelOpen, setRightPanelOpen] = useState(false);
const [rightPanelTab, setRightPanelTab] = useState<RightPanelTab>('code');
const [helpDialogBoxOpen, setHelpDialogBoxOpen] = useState(false);
const openCodePanel = useCallback(() => {
// Legacy API: used by graph/tree selection.
@@ -204,15 +225,7 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
setRightPanelTab('chat');
}, []);
// Filters
const [visibleLabels, setVisibleLabels] = useState<NodeLabel[]>(DEFAULT_VISIBLE_LABELS);
const [visibleEdgeTypes, setVisibleEdgeTypes] = useState<EdgeType[]>(DEFAULT_VISIBLE_EDGES);
// Depth filter
const [depthFilter, setDepthFilter] = useState<number | null>(null);
// Query state
const [highlightedNodeIds, setHighlightedNodeIds] = useState<Set<string>>(new Set());
const [queryResult, setQueryResult] = useState<QueryResult | null>(null);
// AI highlights (separate from user/query highlights)
@@ -298,7 +311,6 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
// LLM/Agent state
const [llmSettings, setLLMSettings] = useState<LLMSettings>(loadSettings);
const [isSettingsPanelOpen, setSettingsPanelOpen] = useState(false);
const [isHelpDialogBoxOpen, setHelpDialogBoxOpen] = useState(false);
const [isAgentReady, setIsAgentReady] = useState(false);
const [isAgentInitializing, setIsAgentInitializing] = useState(false);
const [agentError, setAgentError] = useState<string | null>(null);
@@ -313,54 +325,24 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
const [isCodePanelOpen, setCodePanelOpen] = useState(false);
const [codeReferenceFocus, setCodeReferenceFocus] = useState<CodeReferenceFocus | null>(null);
const normalizePath = useCallback((p: string) => {
return p.replace(/\\/g, '/').replace(/^\.?\//, '');
}, []);
const resolveFilePath = useCallback((requestedPath: string): string | null => {
const req = normalizePath(requestedPath).toLowerCase();
if (!req) return null;
return resolvePathFromContents(fileContents, requestedPath);
}, [fileContents]);
// Exact match first
for (const key of fileContents.keys()) {
if (normalizePath(key).toLowerCase() === req) return key;
}
// Ends-with match (best for partial paths like "src/foo.ts")
let best: { path: string; score: number } | null = null;
for (const key of fileContents.keys()) {
const norm = normalizePath(key).toLowerCase();
if (norm.endsWith(req)) {
const score = 1000 - norm.length; // shorter is better
if (!best || score > best.score) best = { path: key, score };
const fileNodeByPath = useMemo(() => {
if (!graph) return new Map<string, string>();
const map = new Map<string, string>();
for (const n of graph.nodes) {
if (n.label === 'File') {
map.set(normalizePath(n.properties.filePath), n.id);
}
}
if (best) return best.path;
// Segment match fallback
const segs = req.split('/').filter(Boolean);
for (const key of fileContents.keys()) {
const normSegs = normalizePath(key).toLowerCase().split('/').filter(Boolean);
let idx = 0;
for (const s of segs) {
const found = normSegs.findIndex((x, i) => i >= idx && x.includes(s));
if (found === -1) { idx = -1; break; }
idx = found + 1;
}
if (idx !== -1) return key;
}
return null;
}, [fileContents, normalizePath]);
return map;
}, [graph]);
const findFileNodeId = useCallback((filePath: string): string | undefined => {
if (!graph) return undefined;
const target = normalizePath(filePath);
const fileNode = graph.nodes.find(
(n) => n.label === 'File' && normalizePath(n.properties.filePath) === target
);
return fileNode?.id;
}, [graph, normalizePath]);
return fileNodeByPath.get(normalizePath(filePath));
}, [fileNodeByPath]);
// Code References methods
const addCodeReference = useCallback((ref: Omit<CodeReference, 'id'>) => {
@@ -419,7 +401,7 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
}
return kept;
});
}, [queryResult, selectedNode]);
}, [selectedNode]);
// Auto-add a code reference when the user selects a node in the graph/tree
useEffect(() => {
@@ -546,6 +528,25 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
}
}, []);
const startEmbeddingsWithFallback = useCallback(() => {
// Skip auto-start in automated/headless runs to avoid WebGPU errors and long downloads.
const isPlaywright =
(typeof navigator !== 'undefined' && navigator.webdriver) ||
(typeof import.meta !== 'undefined' && typeof import.meta.env !== 'undefined' && import.meta.env.VITE_PLAYWRIGHT_TEST) ||
(typeof process !== 'undefined' && process.env.PLAYWRIGHT_TEST);
if (isPlaywright) {
setEmbeddingStatus('idle');
return;
}
startEmbeddings().catch((err) => {
if (err?.name === 'WebGPUNotAvailableError' || err?.message?.includes('WebGPU')) {
startEmbeddings('wasm').catch(console.warn);
} else {
console.warn('Embeddings auto-start failed:', err);
}
});
}, [startEmbeddings]);
const semanticSearch = useCallback(async (
query: string,
k: number = 10
@@ -709,6 +710,15 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
}
});
};
let pendingUpdate = false;
const scheduleMessageUpdate = () => {
if (pendingUpdate) return;
pendingUpdate = true;
requestAnimationFrame(() => {
pendingUpdate = false;
updateMessage();
});
};
try {
const onChunk = Comlink.proxy((chunk: AgentStreamChunk) => {
@@ -731,7 +741,7 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
content: chunk.reasoning,
});
}
updateMessage();
scheduleMessageUpdate();
}
break;
@@ -754,7 +764,7 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
content: chunk.content,
});
}
updateMessage();
scheduleMessageUpdate();
// Parse inline grounding references and add them to the Code References panel.
// Supports: [[file.ts:10-25]] (file refs) and [[Class:View]] (node refs)
@@ -765,7 +775,7 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
// Pattern 1: File refs - [[path/file.ext]] or [[path/file.ext:line]] or [[path/file.ext:line-line]]
// Line numbers are optional
const fileRefRegex = /\[\[([a-zA-Z0-9_\-./\\]+\.[a-zA-Z0-9]+)(?::(\d+)(?:[-–](\d+))?)?\]\]/g;
const fileRefRegex = new RegExp(FILE_REF_REGEX.source, FILE_REF_REGEX.flags);
let fileMatch: RegExpExecArray | null;
while ((fileMatch = fileRefRegex.exec(fullText)) !== null) {
const rawPath = fileMatch[1].trim();
@@ -791,7 +801,7 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
}
// Pattern 2: Node refs - [[Type:Name]] or [[graph:Type:Name]]
const nodeRefRegex = /\[\[(?:graph:)?(Class|Function|Method|Interface|File|Folder|Variable|Enum|Type|CodeElement):([^\]]+)\]\]/g;
const nodeRefRegex = new RegExp(NODE_REF_REGEX.source, NODE_REF_REGEX.flags);
let nodeMatch: RegExpExecArray | null;
while ((nodeMatch = nodeRefRegex.exec(fullText)) !== null) {
const nodeType = nodeMatch[1];
@@ -832,7 +842,7 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
toolCall: tc,
});
setCurrentToolCalls(prev => [...prev, tc]);
updateMessage();
scheduleMessageUpdate();
}
break;
@@ -891,7 +901,7 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
return prev;
});
updateMessage();
scheduleMessageUpdate();
// Parse highlight marker from tool results
if (tc.result) {
@@ -900,15 +910,15 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
const rawIds = highlightMatch[1].split(',').map((id: string) => id.trim()).filter(Boolean);
if (rawIds.length > 0 && graph) {
const matchedIds = new Set<string>();
const graphNodeIds = graph.nodes.map(n => n.id);
const graphNodeIdSet = new Set(graph.nodes.map(n => n.id));
for (const rawId of rawIds) {
if (graphNodeIds.includes(rawId)) {
if (graphNodeIdSet.has(rawId)) {
matchedIds.add(rawId);
} else {
const found = graphNodeIds.find(gid =>
gid.endsWith(rawId) || gid.endsWith(':' + rawId)
);
const found = graph.nodes.find(n =>
n.id.endsWith(rawId) || n.id.endsWith(':' + rawId)
)?.id;
if (found) {
matchedIds.add(found);
}
@@ -929,15 +939,15 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
const rawIds = impactMatch[1].split(',').map((id: string) => id.trim()).filter(Boolean);
if (rawIds.length > 0 && graph) {
const matchedIds = new Set<string>();
const graphNodeIds = graph.nodes.map(n => n.id);
const graphNodeIdSet = new Set(graph.nodes.map(n => n.id));
for (const rawId of rawIds) {
if (graphNodeIds.includes(rawId)) {
if (graphNodeIdSet.has(rawId)) {
matchedIds.add(rawId);
} else {
const found = graphNodeIds.find(gid =>
gid.endsWith(rawId) || gid.endsWith(':' + rawId)
);
const found = graph.nodes.find(n =>
n.id.endsWith(rawId) || n.id.endsWith(':' + rawId)
)?.id;
if (found) {
matchedIds.add(found);
}
@@ -961,7 +971,7 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
case 'done':
// Finalize the assistant message - just call updateMessage one more time
updateMessage();
scheduleMessageUpdate();
break;
}
});
@@ -997,7 +1007,6 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
setProgress({ phase: 'extracting', percent: 0, message: 'Switching repository...', detail: `Loading ${repoName}` });
setViewMode('loading');
setIsAgentReady(false);
// Clear stale graph state from previous repo (highlights, selections, blast radius)
@@ -1027,7 +1036,7 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
// Reuse the same handleServerConnect logic inline
const repoPath = result.repoInfo.repoPath;
const pName = result.repoInfo.name || repoPath.split('/').pop() || 'server-project';
const pName = repoName || result.repoInfo.name || repoPath.split('/').pop() || 'server-project';
setProjectName(pName);
const graph = createKnowledgeGraph();
@@ -1046,13 +1055,7 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
await initializeAgent(pName);
}
setViewMode('exploring');
startEmbeddings().catch((err) => {
if (err?.name === 'WebGPUNotAvailableError' || err?.message?.includes('WebGPU')) {
startEmbeddings('wasm').catch(console.warn);
} else {
console.warn('Embeddings auto-start failed:', err);
}
});
startEmbeddingsWithFallback();
setProgress(null);
} catch (err) {
console.warn('Failed to load graph into LadybugDB:', err);
@@ -1071,9 +1074,9 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
});
setIsAgentReady(false);
await apiRef.current?.disposeAgent();
setTimeout(() => { setViewMode('exploring'); setProgress(null); }, 3000);
setTimeout(() => { setViewMode('exploring'); setProgress(null); }, ERROR_RESET_DELAY_MS);
}
}, [serverBaseUrl, setProgress, setViewMode, setProjectName, setGraph, setFileContents, loadServerGraph, initializeAgent, startEmbeddings, setHighlightedNodeIds, clearAIToolHighlights, clearAICitationHighlights, clearBlastRadius, setSelectedNode, setQueryResult, setCodeReferences, setCodePanelOpen, setCodeReferenceFocus]);
}, [serverBaseUrl, setProgress, setViewMode, setProjectName, setGraph, setFileContents, loadServerGraph, initializeAgent, startEmbeddingsWithFallback, setHighlightedNodeIds, clearAIToolHighlights, clearAICitationHighlights, clearBlastRadius, setSelectedNode, setQueryResult, setCodeReferences, setCodePanelOpen, setCodeReferenceFocus]);
const removeCodeReference = useCallback((id: string) => {
setCodeReferences(prev => {
@@ -1107,26 +1110,6 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
setCodeReferenceFocus(null);
}, []);
const toggleLabelVisibility = useCallback((label: NodeLabel) => {
setVisibleLabels(prev => {
if (prev.includes(label)) {
return prev.filter(l => l !== label);
} else {
return [...prev, label];
}
});
}, []);
const toggleEdgeVisibility = useCallback((edgeType: EdgeType) => {
setVisibleEdgeTypes(prev => {
if (prev.includes(edgeType)) {
return prev.filter(t => t !== edgeType);
} else {
return [...prev, edgeType];
}
});
}, []);
const value: AppState = {
viewMode,
setViewMode,
@@ -1142,6 +1125,8 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
setRightPanelTab,
openCodePanel,
openChatPanel,
helpDialogBoxOpen,
setHelpDialogBoxOpen,
visibleLabels,
toggleLabelVisibility,
visibleEdgeTypes,
@@ -1184,6 +1169,7 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
embeddingStatus,
embeddingProgress,
startEmbeddings,
startEmbeddingsWithFallback,
semanticSearch,
semanticSearchWithContext,
isEmbeddingReady: embeddingStatus === 'ready',
@@ -1194,8 +1180,6 @@ export const AppStateProvider = ({ children }: { children: ReactNode }) => {
updateLLMSettings,
isSettingsPanelOpen,
setSettingsPanelOpen,
isHelpDialogBoxOpen,
setHelpDialogBoxOpen,
isAgentReady,
isAgentInitializing,
agentError,
+7 -7
View File
@@ -6,16 +6,13 @@ import {
getBackendUrl,
type BackendRepo,
} from '../services/backend';
import { BACKEND_URL_DEBOUNCE_MS, DEFAULT_BACKEND_URL } from '../config/ui-constants';
// ── localStorage keys ────────────────────────────────────────────────────────
const LS_URL_KEY = 'gitnexus-backend-url';
const LS_REPO_KEY = 'gitnexus-backend-repo';
const DEFAULT_URL = 'http://localhost:4747';
// ── Debounce delay ───────────────────────────────────────────────────────────
const DEBOUNCE_MS = 500;
const DEFAULT_URL = DEFAULT_BACKEND_URL;
// ── Public interface ─────────────────────────────────────────────────────────
@@ -90,7 +87,10 @@ export function useBackend(): UseBackendResult {
// Re-check: still the latest probe?
if (id !== probeIdRef.current) return false;
setRepos(repoList);
} catch {
} catch (err) {
if (import.meta.env.DEV) {
console.warn('Failed to fetch repos:', err);
}
if (id === probeIdRef.current) {
setRepos([]);
}
@@ -133,7 +133,7 @@ export function useBackend(): UseBackendResult {
debounceRef.current = setTimeout(() => {
debounceRef.current = null;
void probe();
}, DEBOUNCE_MS);
}, BACKEND_URL_DEBOUNCE_MS);
},
[probe],
);
+34
View File
@@ -19,6 +19,23 @@ export const NODE_COLORS: Record<NodeLabel, string> = {
CodeElement: '#64748b', // Slate - muted
Community: '#818cf8', // Indigo light - cluster indicator
Process: '#f43f5e', // Rose - execution flow indicator
Section: '#60a5fa', // Blue light - structural section
Struct: '#f59e0b', // Amber - like Class
Trait: '#ec4899', // Pink - like Interface
Impl: '#14b8a6', // Teal - like Method
TypeAlias: '#a78bfa', // Violet light - like Type
Const: '#64748b', // Slate - like Variable
Static: '#64748b', // Slate - like Variable
Namespace: '#7c3aed', // Violet - like Module
Union: '#f97316', // Orange - like Enum
Typedef: '#a78bfa', // Violet light - like Type
Macro: '#eab308', // Yellow - like Decorator
Property: '#64748b', // Slate - like Variable
Record: '#f59e0b', // Amber - like Class
Delegate: '#14b8a6', // Teal - like Method
Annotation: '#eab308', // Yellow - like Decorator
Constructor: '#10b981', // Emerald - like Function
Template: '#a78bfa', // Violet light - like Type
};
// Node sizes by type - clear visual hierarchy with dramatic size differences
@@ -41,6 +58,23 @@ export const NODE_SIZES: Record<NodeLabel, number> = {
CodeElement: 2, // Generic small
Community: 0, // Hidden by default - metadata node
Process: 0, // Hidden by default - metadata node
Section: 8, // Structural section - similar to Folder
Struct: 8, // Like Class
Trait: 7, // Like Interface
Impl: 3, // Like Method
TypeAlias: 3, // Like Type
Const: 2, // Like Variable
Static: 2, // Like Variable
Namespace: 13, // Like Module
Union: 5, // Like Enum
Typedef: 3, // Like Type
Macro: 2, // Like Decorator
Property: 2, // Like Variable
Record: 8, // Like Class
Delegate: 3, // Like Method
Annotation: 2, // Like Decorator
Constructor: 4, // Like Function
Template: 3, // Like Type
};
// Community color palette for cluster-based coloring
@@ -0,0 +1,7 @@
// Shared regex patterns for grounding references in chat/markdown.
// Pattern 1: File refs - [[path/file.ext]] or [[path/file.ext:line]] or [[path/file.ext:line-line]]
// Line numbers are optional.
export const FILE_REF_REGEX = /\[\[([a-zA-Z0-9_\-./\\]+\.[a-zA-Z0-9]+)(?::(\d+)(?:[-–](\d+))?)?\]\]/g;
// Pattern 2: Node refs - [[Type:Name]] or [[graph:Type:Name]]
export const NODE_REF_REGEX = /\[\[(?:graph:)?(Class|Function|Method|Interface|File|Folder|Variable|Enum|Type|CodeElement):([^\]]+)\]\]/g;
+73
View File
@@ -0,0 +1,73 @@
/**
* Centralized icon re-exports from lucide-react.
*
* All components import icons from this module (@/lib/lucide-icons) rather
* than directly from lucide-react. This provides a single place to manage
* which icons are used and allows future optimization (e.g., tree-shaking
* configuration, icon subset bundling) without touching every component.
*/
export {
AlertCircle,
AlertTriangle,
ArrowRight,
Brain,
Box,
Braces,
Check,
ChevronDown,
ChevronRight,
ChevronUp,
Code,
Copy,
Eye,
EyeOff,
FileArchive,
FileCode,
Filter,
FlaskConical,
Focus,
Folder,
FolderOpen,
GitBranch,
Github,
Globe,
Hash,
Heart,
HelpCircle,
Home,
Key,
Layers,
Lightbulb,
LightbulbOff,
Loader2,
Maximize2,
MousePointerClick,
PanelLeft,
PanelLeftClose,
PanelRightClose,
Pause,
Play,
RefreshCw,
Rocket,
RotateCcw,
Search,
Send,
Server,
Settings,
SkipForward,
Snail,
Sparkles,
Square,
Star,
Table,
Target,
Terminal,
Trash2,
Upload,
User,
Variable,
X,
Zap,
ZoomIn,
ZoomOut,
} from 'lucide-react';
@@ -26,6 +26,7 @@ export interface ProcessData {
steps: ProcessStep[];
edges?: ProcessEdge[]; // CALLS edges between steps for branching
clusters?: string[];
rawMermaid?: string; // AI-generated mermaid code (sanitized before rendering)
}
/**
+45
View File
@@ -0,0 +1,45 @@
// Utilities for normalizing and resolving file paths referenced in chat and code panels.
export const normalizePath = (p: string): string => {
return p.replace(/\\/g, '/').replace(/^\.?\//, '');
};
/**
* Resolve a user-supplied path (which may be partial) to an exact file path in the repo.
* Follows the same heuristics previously embedded in useAppState:
* 1) exact match, 2) ends-with match (prefers shorter paths), 3) segment containment.
*/
export const resolveFilePath = (fileContents: Map<string, string>, requestedPath: string): string | null => {
const req = normalizePath(requestedPath).toLowerCase();
if (!req) return null;
// Exact match first
for (const key of fileContents.keys()) {
if (normalizePath(key).toLowerCase() === req) return key;
}
// Ends-with match (best for partial paths like "src/foo.ts")
let best: { path: string; score: number } | null = null;
for (const key of fileContents.keys()) {
const norm = normalizePath(key).toLowerCase();
if (norm.endsWith(req)) {
const score = 1000 - norm.length; // shorter is better
if (!best || score > best.score) best = { path: key, score };
}
}
if (best) return best.path;
// Segment match fallback
const segs = req.split('/').filter(Boolean);
for (const key of fileContents.keys()) {
const normSegs = normalizePath(key).toLowerCase().split('/').filter(Boolean);
let idx = 0;
for (const s of segs) {
const found = normSegs.findIndex((x, i) => i >= idx && x.includes(s));
if (found === -1) { idx = -1; break; }
idx = found + 1;
}
if (idx !== -1) return key;
}
return null;
};
+2 -3
View File
@@ -12,9 +12,8 @@ declare module '@ladybugdb/wasm-core' {
close(): Promise<void>;
}
export interface QueryResult {
getAll?(): Promise<any[]>;
getAllRows?(): Promise<any[]>;
getAllObjects?(): Promise<any[]>;
getAll(): Promise<any[]>;
getAllRows(): Promise<any[]>;
hasNext(): Promise<boolean>;
getNext(): Promise<any>;
}
+47 -36
View File
@@ -1,7 +1,5 @@
import * as Comlink from 'comlink';
import { runIngestionPipeline, runPipelineFromFiles } from '../core/ingestion/pipeline';
import { createKnowledgeGraph } from '../core/graph/graph';
import type { GraphNode, GraphRelationship } from '../core/graph/types';
import { PipelineProgress, SerializablePipelineResult, serializePipelineResult } from '../types/pipeline';
import { FileEntry } from '../services/zip';
import {
@@ -14,15 +12,17 @@ import { isEmbedderReady, disposeEmbedder } from '../core/embeddings/embedder';
import type { EmbeddingProgress, SemanticSearchResult } from '../core/embeddings/types';
import type { ProviderConfig, AgentStreamChunk } from '../core/llm/types';
import { createGraphRAGAgent, streamAgentResponse, type AgentMessage, createChatModel } from '../core/llm/agent';
import { createKnowledgeGraph } from '../core/graph/graph';
import type { GraphNode, GraphRelationship } from '../core/graph/types';
import { SystemMessage } from '@langchain/core/messages';
import { enrichClustersBatch, ClusterMemberInfo, ClusterEnrichment } from '../core/ingestion/cluster-enricher';
import { CommunityNode } from '../core/ingestion/community-processor';
import { PipelineResult } from '../types/pipeline';
import { buildCodebaseContext, type CodebaseContext } from '../core/llm/context-builder';
import {
buildBM25Index,
searchBM25,
isBM25Ready,
import {
buildBM25Index,
searchBM25,
isBM25Ready,
getBM25Stats,
mergeWithRRF,
type HybridSearchResult,
@@ -176,7 +176,7 @@ const createHttpHybridSearch = (backendUrl: string, repo: string) => {
endLine: s.endLine,
content: s.content ?? '',
sources: ['bm25', 'semantic'],
score: 1 - (i * 0.02),
score: Math.max(0, 1 - (i * 0.02)),
}));
const defs: any[] = (data.definitions ?? []).map((d: any, i: number) => ({
@@ -186,7 +186,7 @@ const createHttpHybridSearch = (backendUrl: string, repo: string) => {
filePath: d.filePath,
content: '',
sources: ['bm25'],
score: 0.5 - (i * 0.02),
score: Math.max(0, 0.5 - (i * 0.02)),
}));
return [...symbols, ...defs].slice(0, k);
@@ -644,6 +644,12 @@ const workerApi = {
/**
* Initialize the Graph RAG agent in backend mode (HTTP-backed tools).
* Uses HTTP wrappers instead of local LadybugDB for all tool queries.
*
* NOTE: Currently not called by any UI flow. The server-connect path
* downloads the full graph and uses local WASM queries via initializeAgent.
* This method is retained for future large-repo mode where downloading
* the entire graph to the browser would be impractical.
*
* @param config - Provider configuration for the LLM
* @param backendUrl - Base URL of the gitnexus serve backend
* @param repoName - Repository name on the backend
@@ -785,8 +791,10 @@ const workerApi = {
throw new Error('No graph loaded. Please ingest a repository first.');
}
enrichmentCancelled = false;
const { graph } = currentGraphResult;
// Filter for community nodes
const communityNodes = graph.nodes
.filter(n => n.label === 'Community')
@@ -808,15 +816,22 @@ const workerApi = {
// Initialize map
communityNodes.forEach(c => memberMap.set(c.id, []));
// Build a Map for O(1) node lookups instead of O(N) find per relationship
const nodeById = new Map(graph.nodes.map(n => [n.id, n]));
// Find all MEMBER_OF edges
graph.relationships.forEach(rel => {
for (const rel of graph.relationships) {
if (enrichmentCancelled) {
console.log('Enrichment cancelled, stopping');
break;
}
if (rel.type === 'MEMBER_OF') {
const communityId = rel.targetId;
const memberId = rel.sourceId; // MEMBER_OF goes Member -> Community
if (memberMap.has(communityId)) {
// Find member node details
const memberNode = graph.nodes.find(n => n.id === memberId);
const memberNode = nodeById.get(memberId);
if (memberNode) {
memberMap.get(communityId)?.push({
name: memberNode.properties.name,
@@ -826,7 +841,7 @@ const workerApi = {
}
}
}
});
}
// Create LLM client adapter for LangChain model
const chatModel = createChatModel(providerConfig);
@@ -864,32 +879,28 @@ const workerApi = {
}
});
// Update LadybugDB with new data
// Update LadybugDB with new data using prepared statements
try {
const lbug = await getLbugAdapter();
onProgress(enrichments.size, enrichments.size); // Done
// Update one by one via Cypher (simplest for now)
for (const [id, enrichment] of enrichments.entries()) {
// Escape strings for Cypher - replace backslash first, then quotes
const escapeCypher = (str: string) => str.replace(/\\/g, '\\\\').replace(/"/g, '\\"');
const keywordsStr = JSON.stringify(enrichment.keywords);
const descStr = escapeCypher(enrichment.description);
const nameStr = escapeCypher(enrichment.name);
const escapedId = escapeCypher(id);
const query = `
MATCH (c:Community {id: "${escapedId}"})
SET c.label = "${nameStr}",
c.keywords = ${keywordsStr},
c.description = "${descStr}",
c.enrichedBy = "llm"
`;
await lbug.executeQuery(query);
}
const paramsList = Array.from(enrichments.entries()).map(([id, enrichment]) => ({
id,
label: enrichment.name,
keywords: enrichment.keywords,
description: enrichment.description,
}));
const updateQuery = `
MATCH (c:Community {id: $id})
SET c.label = $label,
c.keywords = $keywords,
c.description = $description,
c.enrichedBy = "llm"
`;
await lbug.executeWithReusedStatement(updateQuery, paramsList);
} catch (err) {
console.error('Failed to update LadybugDB with enrichment:', err);
+66
View File
@@ -0,0 +1,66 @@
/**
* Shared test data factories for graph structures.
* No test code — pure data exports.
*/
import type { GraphNode, GraphRelationship } from '../../src/core/graph/types';
export function createFileNode(name: string, filePath?: string): GraphNode {
return {
id: `File:${filePath ?? name}`,
label: 'File',
properties: { name, filePath: filePath ?? name },
};
}
export function createFunctionNode(name: string, filePath: string, line = 1): GraphNode {
return {
id: `Function:${filePath}:${name}:${line}`,
label: 'Function',
properties: { name, filePath, startLine: line, endLine: line + 10 },
};
}
export function createClassNode(name: string, filePath: string): GraphNode {
return {
id: `Class:${filePath}:${name}`,
label: 'Class',
properties: { name, filePath },
};
}
export function createProcessNode(id: string, label: string, type: 'cross_community' | 'intra_community' = 'cross_community'): GraphNode {
return {
id,
label: 'Process',
properties: {
name: label,
heuristicLabel: label,
processType: type,
stepCount: 3,
communities: ['cluster-a', 'cluster-b'],
} as any,
};
}
export function createCallsRelationship(sourceId: string, targetId: string): GraphRelationship {
return {
id: `${sourceId}_CALLS_${targetId}`,
sourceId,
targetId,
type: 'CALLS',
confidence: 0.9,
reason: 'same-file',
};
}
export function createContainsRelationship(sourceId: string, targetId: string): GraphRelationship {
return {
id: `${sourceId}_CONTAINS_${targetId}`,
sourceId,
targetId,
type: 'CONTAINS',
confidence: 1.0,
reason: '',
};
}
+4 -2
View File
@@ -1,6 +1,8 @@
import { beforeEach } from 'vitest';
import '@testing-library/jest-dom/vitest';
// Reset storage between tests
beforeEach(() => {
sessionStorage.clear();
localStorage.clear();
sessionStorage.removeItem('gitnexus-llm-settings');
localStorage.removeItem('gitnexus-llm-settings'); // legacy key (migration)
});
+81
View File
@@ -0,0 +1,81 @@
import { describe, expect, it } from 'vitest';
import {
NODE_COLORS,
NODE_SIZES,
COMMUNITY_COLORS,
getCommunityColor,
DEFAULT_VISIBLE_LABELS,
FILTERABLE_LABELS,
ALL_EDGE_TYPES,
DEFAULT_VISIBLE_EDGES,
EDGE_INFO,
} from '../../src/lib/constants';
describe('NODE_COLORS', () => {
it('has a color for every node label used in NODE_SIZES', () => {
for (const label of Object.keys(NODE_SIZES)) {
expect(NODE_COLORS).toHaveProperty(label);
expect(NODE_COLORS[label as keyof typeof NODE_COLORS]).toMatch(/^#[0-9a-f]{6}$/i);
}
});
});
describe('NODE_SIZES', () => {
it('gives Project the largest size', () => {
const maxLabel = Object.entries(NODE_SIZES).reduce((a, b) => a[1] > b[1] ? a : b);
expect(maxLabel[0]).toBe('Project');
});
it('gives structural nodes larger sizes than code nodes', () => {
expect(NODE_SIZES.Folder).toBeGreaterThan(NODE_SIZES.Function);
expect(NODE_SIZES.File).toBeGreaterThan(NODE_SIZES.Variable);
});
});
describe('getCommunityColor', () => {
it('returns valid hex colors', () => {
for (let i = 0; i < 20; i++) {
expect(getCommunityColor(i)).toMatch(/^#[0-9a-f]{6}$/i);
}
});
it('wraps around the palette', () => {
const paletteSize = COMMUNITY_COLORS.length;
expect(getCommunityColor(0)).toBe(getCommunityColor(paletteSize));
expect(getCommunityColor(1)).toBe(getCommunityColor(paletteSize + 1));
});
});
describe('DEFAULT_VISIBLE_LABELS', () => {
it('includes common structural and code labels', () => {
expect(DEFAULT_VISIBLE_LABELS).toContain('File');
expect(DEFAULT_VISIBLE_LABELS).toContain('Function');
expect(DEFAULT_VISIBLE_LABELS).toContain('Class');
});
it('excludes noisy labels by default', () => {
expect(DEFAULT_VISIBLE_LABELS).not.toContain('Variable');
expect(DEFAULT_VISIBLE_LABELS).not.toContain('Import');
});
});
describe('edge types', () => {
it('ALL_EDGE_TYPES contains all EDGE_INFO keys', () => {
const edgeInfoKeys = Object.keys(EDGE_INFO).sort();
const allEdgeTypes = [...ALL_EDGE_TYPES].sort();
expect(edgeInfoKeys).toEqual(allEdgeTypes);
});
it('DEFAULT_VISIBLE_EDGES is a subset of ALL_EDGE_TYPES', () => {
for (const type of DEFAULT_VISIBLE_EDGES) {
expect(ALL_EDGE_TYPES).toContain(type);
}
});
it('EDGE_INFO entries have color and label', () => {
for (const info of Object.values(EDGE_INFO)) {
expect(info.color).toMatch(/^#[0-9a-f]{6}$/i);
expect(info.label.length).toBeGreaterThan(0);
}
});
});
@@ -0,0 +1,201 @@
import { describe, expect, it } from 'vitest';
import { generateAllCSVs } from '../../src/core/lbug/csv-generator';
import { createKnowledgeGraph } from '../../src/core/graph/graph';
import { NODE_TABLES } from '../../src/core/lbug/schema';
describe('generateAllCSVs', () => {
it('generates CSV for all NODE_TABLES present in the graph', () => {
const graph = createKnowledgeGraph();
// Add one node per multi-language table type
const testTables = ['Struct', 'Enum', 'Trait', 'Impl', 'Macro', 'TypeAlias'] as const;
for (const label of testTables) {
graph.addNode({
id: `${label}:test.rs:MyItem`,
label,
properties: { name: 'MyItem', filePath: 'test.rs', startLine: 1, endLine: 10 },
});
}
const csvData = generateAllCSVs(graph, new Map());
for (const label of testTables) {
const csv = csvData.nodes.get(label);
expect(csv, `CSV for ${label} should be generated`).toBeDefined();
expect(csv!.split('\n').length, `CSV for ${label} should have header + 1 row`).toBeGreaterThanOrEqual(2);
}
});
it('multi-language table CSVs have 6 columns (no isExported)', () => {
const graph = createKnowledgeGraph();
graph.addNode({
id: 'Struct:lib.rs:Point',
label: 'Struct',
properties: { name: 'Point', filePath: 'lib.rs', startLine: 5, endLine: 15 },
});
const csvData = generateAllCSVs(graph, new Map());
const csv = csvData.nodes.get('Struct')!;
const header = csv.split('\n')[0];
const columns = header.split(',').length;
expect(columns).toBe(6); // id, name, filePath, startLine, endLine, content
});
it('community keywords with commas are properly CSV-escaped', () => {
const graph = createKnowledgeGraph();
graph.addNode({
id: 'comm_0',
label: 'Community',
properties: {
name: 'TestCluster',
heuristicLabel: 'test',
keywords: ['auth, login', 'user management'],
description: 'test community',
cohesion: 0.8,
symbolCount: 5,
},
});
const csvData = generateAllCSVs(graph, new Map());
const csv = csvData.nodes.get('Community')!;
const rows = csv.split('\n');
expect(rows.length).toBeGreaterThanOrEqual(2);
// The keywords field should be quoted (RFC 4180) since it contains commas
const dataRow = rows[1];
expect(dataRow).toContain('auth');
expect(dataRow).toContain('login');
});
it('generates File CSV with content from fileContents map', () => {
const graph = createKnowledgeGraph();
graph.addNode({
id: 'File:src/index.ts',
label: 'File',
properties: { name: 'index.ts', filePath: 'src/index.ts' },
});
const fileContents = new Map([['src/index.ts', 'console.log("hello")']]);
const csvData = generateAllCSVs(graph, fileContents);
const csv = csvData.nodes.get('File')!;
expect(csv).toContain('hello');
});
it('returns a relCSV string', () => {
const graph = createKnowledgeGraph();
graph.addNode({
id: 'Function:a.ts:foo',
label: 'Function',
properties: { name: 'foo', filePath: 'a.ts', startLine: 1, endLine: 5 },
});
graph.addNode({
id: 'Function:a.ts:bar',
label: 'Function',
properties: { name: 'bar', filePath: 'a.ts', startLine: 10, endLine: 15 },
});
graph.addRelationship({
sourceId: 'Function:a.ts:foo',
targetId: 'Function:a.ts:bar',
type: 'CALLS',
properties: {},
});
const csvData = generateAllCSVs(graph, new Map());
expect(csvData.relCSV).toContain('CALLS');
expect(csvData.relCSV).toContain('Function:a.ts:foo');
});
it('handles all NODE_TABLES without crashing', () => {
const graph = createKnowledgeGraph();
for (const table of NODE_TABLES) {
if (table === 'File' || table === 'Folder' || table === 'Community' || table === 'Process') continue;
graph.addNode({
id: `${table}:test:item`,
label: table,
properties: { name: 'item', filePath: 'test', startLine: 1, endLine: 2 },
});
}
expect(() => generateAllCSVs(graph, new Map())).not.toThrow();
const csvData = generateAllCSVs(graph, new Map());
expect(csvData.nodes.size).toBeGreaterThan(0);
});
// ── Negative tests ──────────────────────────────────────────────
it('empty graph produces no node CSVs with data rows', () => {
const graph = createKnowledgeGraph();
const csvData = generateAllCSVs(graph, new Map());
// Nodes map may have header-only entries or be empty
for (const [, csv] of csvData.nodes.entries()) {
const rows = csv.split('\n').filter(r => r.trim());
expect(rows.length).toBeLessThanOrEqual(1); // header only, no data
}
});
it('empty graph produces relCSV with only a header', () => {
const graph = createKnowledgeGraph();
const csvData = generateAllCSVs(graph, new Map());
const lines = csvData.relCSV.split('\n').filter(r => r.trim());
expect(lines.length).toBeLessThanOrEqual(1);
});
it('node with double quotes in name is properly escaped', () => {
const graph = createKnowledgeGraph();
graph.addNode({
id: 'Function:a.ts:say"hello"',
label: 'Function',
properties: { name: 'say"hello"', filePath: 'a.ts', startLine: 1, endLine: 5 },
});
const csvData = generateAllCSVs(graph, new Map());
const csv = csvData.nodes.get('Function')!;
// RFC 4180: double quotes inside fields are doubled
expect(csv).toContain('""');
});
it('file node without matching fileContents gets empty content', () => {
const graph = createKnowledgeGraph();
graph.addNode({
id: 'File:missing.ts',
label: 'File',
properties: { name: 'missing.ts', filePath: 'missing.ts' },
});
const csvData = generateAllCSVs(graph, new Map()); // no file contents
const csv = csvData.nodes.get('File')!;
const rows = csv.split('\n');
expect(rows.length).toBeGreaterThanOrEqual(2);
// Content field should be empty (not undefined or crash)
});
it('community with empty keywords array produces valid CSV', () => {
const graph = createKnowledgeGraph();
graph.addNode({
id: 'comm_1',
label: 'Community',
properties: {
name: 'EmptyCluster',
heuristicLabel: 'empty',
keywords: [],
description: '',
cohesion: 0,
symbolCount: 0,
},
});
expect(() => generateAllCSVs(graph, new Map())).not.toThrow();
const csvData = generateAllCSVs(graph, new Map());
expect(csvData.nodes.get('Community')).toBeDefined();
});
it('node labels not in NODE_TABLES are silently skipped', () => {
const graph = createKnowledgeGraph();
graph.addNode({
id: 'FakeLabel:test:item',
label: 'FakeLabel' as any,
properties: { name: 'item', filePath: 'test' },
});
// Should not crash — unknown labels are just not in any CSV
expect(() => generateAllCSVs(graph, new Map())).not.toThrow();
});
});
+68
View File
@@ -0,0 +1,68 @@
import { describe, expect, it } from 'vitest';
import { createKnowledgeGraph } from '../../src/core/graph/graph';
import { createFileNode, createFunctionNode, createCallsRelationship, createContainsRelationship } from '../fixtures/graph';
describe('createKnowledgeGraph', () => {
it('starts empty', () => {
const graph = createKnowledgeGraph();
expect(graph.nodeCount).toBe(0);
expect(graph.relationshipCount).toBe(0);
expect(graph.nodes).toEqual([]);
expect(graph.relationships).toEqual([]);
});
it('adds nodes', () => {
const graph = createKnowledgeGraph();
const node = createFileNode('index.ts', 'src/index.ts');
graph.addNode(node);
expect(graph.nodeCount).toBe(1);
expect(graph.nodes[0].id).toBe('File:src/index.ts');
});
it('deduplicates nodes by id', () => {
const graph = createKnowledgeGraph();
const node = createFileNode('index.ts', 'src/index.ts');
const duplicateNode = createFileNode('index.ts', 'src/index.ts');
graph.addNode(node);
graph.addNode(duplicateNode);
expect(graph.nodeCount).toBe(1);
});
it('adds relationships', () => {
const graph = createKnowledgeGraph();
const rel = createCallsRelationship('fn:a', 'fn:b');
graph.addRelationship(rel);
expect(graph.relationshipCount).toBe(1);
expect(graph.relationships[0].type).toBe('CALLS');
});
it('deduplicates relationships by id', () => {
const graph = createKnowledgeGraph();
const rel = createCallsRelationship('fn:a', 'fn:b');
const duplicateRel = createCallsRelationship('fn:a', 'fn:b');
graph.addRelationship(rel);
graph.addRelationship(duplicateRel);
expect(graph.relationshipCount).toBe(1);
});
it('builds a multi-node graph', () => {
const graph = createKnowledgeGraph();
const file = createFileNode('app.ts', 'src/app.ts');
const fn1 = createFunctionNode('main', 'src/app.ts', 1);
const fn2 = createFunctionNode('helper', 'src/app.ts', 20);
graph.addNode(file);
graph.addNode(fn1);
graph.addNode(fn2);
graph.addRelationship(createContainsRelationship(file.id, fn1.id));
graph.addRelationship(createContainsRelationship(file.id, fn2.id));
graph.addRelationship(createCallsRelationship(fn1.id, fn2.id));
expect(graph.nodeCount).toBe(3);
expect(graph.relationshipCount).toBe(3);
});
});
@@ -0,0 +1,107 @@
import { describe, expect, it } from 'vitest';
import { generateProcessMermaid, generateSimpleMermaid } from '../../src/lib/mermaid-generator';
import type { ProcessData } from '../../src/lib/mermaid-generator';
describe('generateProcessMermaid', () => {
it('returns placeholder for empty steps', () => {
const process: ProcessData = {
id: 'p1',
label: 'Empty',
processType: 'intra_community',
steps: [],
};
expect(generateProcessMermaid(process)).toContain('No steps found');
});
it('generates a linear chain without edges', () => {
const process: ProcessData = {
id: 'p1',
label: 'GET -> Handler',
processType: 'intra_community',
steps: [
{ id: 'fn:a', name: 'handleGet', filePath: 'src/routes.ts', stepNumber: 1 },
{ id: 'fn:b', name: 'validate', filePath: 'src/validate.ts', stepNumber: 2 },
{ id: 'fn:c', name: 'respond', filePath: 'src/respond.ts', stepNumber: 3 },
],
};
const result = generateProcessMermaid(process);
expect(result).toContain('graph TD');
expect(result).toContain('handleGet');
expect(result).toContain('validate');
expect(result).toContain('respond');
// Linear chain: a -> b -> c
expect(result).toContain('-->');
});
it('uses CALLS edges when provided', () => {
const process: ProcessData = {
id: 'p1',
label: 'Branching',
processType: 'intra_community',
steps: [
{ id: 'fn:a', name: 'entry', filePath: 'src/a.ts', stepNumber: 1 },
{ id: 'fn:b', name: 'branchA', filePath: 'src/b.ts', stepNumber: 2 },
{ id: 'fn:c', name: 'branchB', filePath: 'src/c.ts', stepNumber: 3 },
],
edges: [
{ from: 'fn:a', to: 'fn:b', type: 'CALLS' },
{ from: 'fn:a', to: 'fn:c', type: 'CALLS' },
],
};
const result = generateProcessMermaid(process);
// Both edges should appear
expect(result).toContain('fn_a --> fn_b');
expect(result).toContain('fn_a --> fn_c');
});
it('applies entry and terminal classes', () => {
const process: ProcessData = {
id: 'p1',
label: 'Flow',
processType: 'intra_community',
steps: [
{ id: 'fn:start', name: 'start', filePath: 'src/a.ts', stepNumber: 1 },
{ id: 'fn:end', name: 'end', filePath: 'src/b.ts', stepNumber: 2 },
],
};
const result = generateProcessMermaid(process);
expect(result).toContain(':::entry');
expect(result).toContain(':::terminal');
});
it('uses subgraphs for cross-community processes with clusters', () => {
const process: ProcessData = {
id: 'p1',
label: 'Cross',
processType: 'cross_community',
steps: [
{ id: 'fn:a', name: 'a', filePath: 'src/a.ts', stepNumber: 1, cluster: 'Auth' },
{ id: 'fn:b', name: 'b', filePath: 'src/b.ts', stepNumber: 2, cluster: 'DB' },
],
};
const result = generateProcessMermaid(process);
expect(result).toContain('subgraph');
expect(result).toContain('Auth');
expect(result).toContain('DB');
});
});
describe('generateSimpleMermaid', () => {
it('generates a preview with entry and terminal', () => {
const result = generateSimpleMermaid('POST -> ShouldRedact', 5);
expect(result).toContain('graph LR');
expect(result).toContain('POST');
expect(result).toContain('ShouldRedact');
expect(result).toContain('3 steps');
});
it('handles labels without arrow', () => {
const result = generateSimpleMermaid('SingleNode', 2);
expect(result).toContain('graph LR');
expect(result).toContain('SingleNode');
});
});
@@ -0,0 +1,31 @@
import { describe, expect, it } from 'vitest';
import { normalizePath, resolveFilePath } from '../../src/lib/path-resolution';
describe('path-resolution utilities', () => {
const contents = new Map<string, string>([
['src/components/Header.tsx', ''],
['src/core/utils/index.ts', ''],
['README.md', ''],
['src/lib/path-resolution.ts', ''],
]);
it('normalizes leading ./ and backslashes', () => {
expect(normalizePath('./src\\components\\Header.tsx')).toBe('src/components/Header.tsx');
});
it('prefers exact matches', () => {
expect(resolveFilePath(contents, 'src/components/Header.tsx')).toBe('src/components/Header.tsx');
});
it('resolves ends-with partials', () => {
expect(resolveFilePath(contents, 'core/utils/index.ts')).toBe('src/core/utils/index.ts');
});
it('falls back to segment matching', () => {
expect(resolveFilePath(contents, 'lib/path')).toBe('src/lib/path-resolution.ts');
});
it('returns null for empty requests', () => {
expect(resolveFilePath(contents, '')).toBeNull();
});
});
@@ -0,0 +1,74 @@
import { describe, expect, it } from 'vitest';
// ==========================================================================
// PR4 Performance Optimizations — verify behavior preserved after changes
// Tests the pure functions underlying the O(1) lookup optimizations.
// ==========================================================================
describe('nodeById Map — O(1) lookup correctness', () => {
// Positive: Map.get returns correct node
it('Map provides O(1) lookup by ID', () => {
const nodes = [
{ id: 'Function:a.ts:foo', label: 'Function', name: 'foo' },
{ id: 'Class:b.ts:Bar', label: 'Class', name: 'Bar' },
{ id: 'File:c.ts', label: 'File', name: 'c.ts' },
];
const nodeById = new Map(nodes.map(n => [n.id, n]));
expect(nodeById.get('Function:a.ts:foo')?.name).toBe('foo');
expect(nodeById.get('Class:b.ts:Bar')?.name).toBe('Bar');
expect(nodeById.get('File:c.ts')?.label).toBe('File');
});
// Positive: handles duplicate IDs (last wins)
it('last node wins on duplicate IDs', () => {
const nodes = [
{ id: 'File:a.ts', label: 'File', name: 'first' },
{ id: 'File:a.ts', label: 'File', name: 'second' },
];
const nodeById = new Map(nodes.map(n => [n.id, n]));
expect(nodeById.get('File:a.ts')?.name).toBe('second');
expect(nodeById.size).toBe(1);
});
// Negative: missing ID returns undefined
it('returns undefined for non-existent ID', () => {
const nodeById = new Map([['File:a.ts', { id: 'File:a.ts' }]]);
expect(nodeById.get('NonExistent:x')).toBeUndefined();
});
// Negative: empty map
it('empty Map returns undefined for any key', () => {
const nodeById = new Map<string, any>();
expect(nodeById.get('anything')).toBeUndefined();
});
});
describe('Set.has — O(1) highlight matching', () => {
// Positive: exact match
it('Set.has returns true for present IDs', () => {
const idSet = new Set(['Function:a.ts:foo', 'Class:b.ts:Bar']);
expect(idSet.has('Function:a.ts:foo')).toBe(true);
expect(idSet.has('Class:b.ts:Bar')).toBe(true);
});
// Negative: missing ID
it('Set.has returns false for absent IDs', () => {
const idSet = new Set(['Function:a.ts:foo']);
expect(idSet.has('Function:a.ts:bar')).toBe(false);
expect(idSet.has('')).toBe(false);
});
// Positive: works with graph node IDs containing special chars
it('handles IDs with colons, dots, and slashes', () => {
const idSet = new Set(['Function:src/utils/path-resolver.ts:resolveFile']);
expect(idSet.has('Function:src/utils/path-resolver.ts:resolveFile')).toBe(true);
});
// Negative: case sensitive
it('is case-sensitive', () => {
const idSet = new Set(['Function:a.ts:Foo']);
expect(idSet.has('Function:a.ts:foo')).toBe(false);
});
});
@@ -0,0 +1,225 @@
import { describe, expect, it } from 'vitest';
import { NODE_TABLES, REL_TYPES } from '../../src/core/lbug/schema';
// ---------------------------------------------------------------------------
// Recreate the security guards locally so we can test the exact logic used in
// production without exporting private helpers.
//
// Source locations:
// validLabel / validRelType -- gitnexus-web/src/core/llm/tools.ts
// isSafeId -- gitnexus-web/src/components/ProcessesPanel.tsx
// readOnly guard (regex) -- gitnexus-web/src/core/lbug/lbug-adapter.ts
// ---------------------------------------------------------------------------
const validLabel = (label: string): boolean =>
(NODE_TABLES as readonly string[]).includes(label);
const validRelType = (t: string): boolean =>
(REL_TYPES as readonly string[]).includes(t);
const isSafeId = (id: string): boolean =>
/^[a-zA-Z0-9_:.\-/@]+$/.test(id);
const isWriteQuery = (cypher: string): boolean => {
const stripped = cypher.replace(/'[^']*'|"[^"]*"/g, '').toUpperCase();
return /\b(CREATE|DELETE|SET|MERGE|REMOVE|DROP|DETACH)\b/.test(stripped);
};
// ===========================================================================
// validLabel
// ===========================================================================
describe('validLabel – NODE_TABLES membership', () => {
it.each([
'Function', 'Class', 'File', 'Process', 'Community',
])('accepts known label "%s"', (label) => {
expect(validLabel(label)).toBe(true);
});
it.each([
'Struct', 'Enum', 'Trait', 'Impl', 'Macro', 'Typedef',
'Union', 'Namespace', 'TypeAlias', 'Const', 'Static',
'Property', 'Record', 'Delegate', 'Annotation',
'Constructor', 'Template', 'Module',
])('accepts multi-language label "%s"', (label) => {
expect(validLabel(label)).toBe(true);
});
it.each([
['empty string', ''],
['SQL keyword', 'DROP'],
['random word', 'foo'],
['Cypher injection', '})-[:R]->(x)'],
['label with semicolon', 'Function;DELETE'],
['lowercase (case matters)', 'function'],
['lowercase class', 'class'],
['whitespace padded', ' File '],
['numeric', '123'],
])('rejects invalid label: %s', (_desc, label) => {
expect(validLabel(label)).toBe(false);
});
it('NODE_TABLES contains all expected core labels', () => {
const core = ['File', 'Folder', 'Function', 'Class', 'Interface', 'Method', 'CodeElement', 'Community', 'Process'];
for (const label of core) {
expect((NODE_TABLES as readonly string[]).includes(label)).toBe(true);
}
});
});
// ===========================================================================
// validRelType
// ===========================================================================
describe('validRelType – REL_TYPES membership', () => {
it.each(
[...REL_TYPES]
)('accepts known relation type "%s"', (relType) => {
expect(validRelType(relType)).toBe(true);
});
it.each([
['empty string', ''],
['SQL keyword', 'DROP'],
['injection attempt', 'CALLS;DELETE'],
['lowercase', 'calls'],
['nonexistent type', 'FRIEND_OF'],
['padded', ' CALLS '],
])('rejects invalid relation type: %s', (_desc, relType) => {
expect(validRelType(relType)).toBe(false);
});
it('REL_TYPES has at least the base types', () => {
// Guard against accidental removal of relation types
expect(REL_TYPES.length).toBeGreaterThanOrEqual(8);
});
});
// ===========================================================================
// isSafeId
// ===========================================================================
describe('isSafeId – identifier allowlist regex', () => {
it.each([
['namespaced id', 'Function:myFunc'],
['underscore id', 'proc_5'],
['class id', 'Class:MyClass'],
['dotted name', 'Module:path.to.thing'],
['with hyphen', 'File:my-file.ts'],
['community id', 'comm_5'],
['file path id', 'File:src/index.ts'],
['nested path id', 'Function:src/utils/helpers.ts:doStuff'],
['scoped npm package', 'Module:@scope/pkg'],
['angular-style id', 'Module:@angular/core'],
])('accepts valid ID: %s', (_desc, id) => {
expect(isSafeId(id)).toBe(true);
});
it.each([
['with spaces', 'Process:my process'],
])('rejects ID with unsafe chars: %s', (_desc, id) => {
expect(isSafeId(id)).toBe(false);
});
it('rejects empty string', () => {
expect(isSafeId('')).toBe(false);
});
it.each([
['SQL injection', "'; DROP TABLE"],
['command substitution', '$(command)'],
['XSS attempt', '<script>'],
['JSON injection', '{id: "x"}'],
])('rejects injection attempt: %s', (_desc, id) => {
expect(isSafeId(id)).toBe(false);
});
it.each([
['open paren', '('],
['close paren', ')'],
['open bracket', '['],
['close bracket', ']'],
['open brace', '{'],
['close brace', '}'],
['backtick', '`'],
['double quote', '"'],
['single quote', "'"],
])('rejects Cypher metacharacter: %s', (_desc, ch) => {
expect(isSafeId(ch)).toBe(false);
});
it.each([
['embedded paren', 'func(x)'],
['embedded bracket', 'arr[0]'],
['embedded brace', '{key}'],
['embedded backtick', 'id`inject'],
])('rejects id containing metacharacter: %s', (_desc, id) => {
expect(isSafeId(id)).toBe(false);
});
});
// ===========================================================================
// readOnly guard – write-operation regex
// ===========================================================================
describe('readOnly guard – write-operation detection', () => {
describe('allows read-only queries (should NOT match)', () => {
it.each([
['simple match', 'MATCH (n) RETURN n'],
['filtered match', 'MATCH (n:Function) WHERE n.name = "test" RETURN n'],
['with relationship', 'MATCH (a)-[r:CodeRelation]->(b) RETURN a, r, b'],
['with count', 'MATCH (n) RETURN count(n)'],
['with ordering', 'MATCH (n) RETURN n ORDER BY n.name LIMIT 10'],
['call procedure', 'CALL db.schema.nodeTypeProperties()'],
])('%s', (_desc, cypher) => {
expect(isWriteQuery(cypher)).toBe(false);
});
});
describe('blocks write operations (should match)', () => {
it.each([
['DELETE node', 'MATCH (n) DELETE n'],
['CREATE node', 'CREATE (n:Test)'],
['SET property', 'MATCH (n) SET n.x = 1'],
['MERGE node', 'MERGE (n:Test {id: "1"})'],
['REMOVE property', 'MATCH (n) REMOVE n.x'],
['DETACH DELETE', 'MATCH (n) DETACH DELETE n'],
['DROP (DDL)', 'DROP TABLE x'],
])('%s', (_desc, cypher) => {
expect(isWriteQuery(cypher)).toBe(true);
});
});
describe('handles tricky cases', () => {
it('detects write keyword even when embedded in longer query', () => {
const cypher = 'MATCH (n:Function) WHERE n.name = "handler" DELETE n';
expect(isWriteQuery(cypher)).toBe(true);
});
it('detects mixed-case write keywords via toUpperCase()', () => {
expect(isWriteQuery('match (n) delete n')).toBe(true);
expect(isWriteQuery('Match (n) Set n.x = 1')).toBe(true);
});
// Keywords inside quoted strings are stripped before checking,
// so they don't trigger false positives.
it('allows "delete" inside a quoted string value', () => {
expect(isWriteQuery('MATCH (n) WHERE n.name CONTAINS "delete" RETURN n')).toBe(false);
});
it('allows "CREATE" inside single-quoted string', () => {
expect(isWriteQuery("MATCH (n) WHERE n.name = 'CREATE_USER' RETURN n")).toBe(false);
});
it('still blocks DELETE outside quotes', () => {
expect(isWriteQuery('MATCH (n) WHERE n.name = "foo" DELETE n')).toBe(true);
});
// Verify the word-boundary prevents false positives on substrings that
// are NOT Cypher write keywords.
it('does not match partial keywords like "CREATED" or "SETTING"', () => {
expect(isWriteQuery('MATCH (n) WHERE n.status = "CREATED" RETURN n')).toBe(false);
expect(isWriteQuery('MATCH (n) WHERE n.label = "SETTING" RETURN n')).toBe(false);
});
it('does not flag the word "create" inside a property name like "createdAt"', () => {
expect(isWriteQuery('MATCH (n) RETURN n.createdAt')).toBe(false);
});
});
});
@@ -0,0 +1,76 @@
import { describe, expect, it } from 'vitest';
import { normalizeServerUrl, extractFileContents } from '../../src/services/server-connection';
import type { GraphNode } from '../../src/core/graph/types';
describe('normalizeServerUrl', () => {
it('adds http:// to localhost', () => {
expect(normalizeServerUrl('localhost:4747')).toBe('http://localhost:4747/api');
});
it('adds http:// to 127.0.0.1', () => {
expect(normalizeServerUrl('127.0.0.1:4747')).toBe('http://127.0.0.1:4747/api');
});
it('adds https:// to non-local hosts', () => {
expect(normalizeServerUrl('example.com')).toBe('https://example.com/api');
});
it('strips trailing slashes', () => {
expect(normalizeServerUrl('http://localhost:4747/')).toBe('http://localhost:4747/api');
expect(normalizeServerUrl('http://localhost:4747///')).toBe('http://localhost:4747/api');
});
it('does not double-append /api', () => {
expect(normalizeServerUrl('http://localhost:4747/api')).toBe('http://localhost:4747/api');
});
it('trims whitespace', () => {
expect(normalizeServerUrl(' localhost:4747 ')).toBe('http://localhost:4747/api');
});
it('preserves existing https://', () => {
expect(normalizeServerUrl('https://gitnexus.example.com')).toBe('https://gitnexus.example.com/api');
});
});
describe('extractFileContents', () => {
it('extracts content from File nodes', () => {
const nodes: GraphNode[] = [
{
id: 'File:src/index.ts',
label: 'File',
properties: { name: 'index.ts', filePath: 'src/index.ts', content: 'console.log("hello")' } as any,
},
];
const result = extractFileContents(nodes);
expect(result['src/index.ts']).toBe('console.log("hello")');
});
it('ignores non-File nodes', () => {
const nodes: GraphNode[] = [
{
id: 'Function:main',
label: 'Function',
properties: { name: 'main', filePath: 'src/index.ts', content: 'fn body' } as any,
},
];
const result = extractFileContents(nodes);
expect(Object.keys(result)).toHaveLength(0);
});
it('ignores File nodes without content', () => {
const nodes: GraphNode[] = [
{
id: 'File:src/empty.ts',
label: 'File',
properties: { name: 'empty.ts', filePath: 'src/empty.ts' },
},
];
const result = extractFileContents(nodes);
expect(Object.keys(result)).toHaveLength(0);
});
it('returns empty object for empty input', () => {
expect(extractFileContents([])).toEqual({});
});
});
@@ -0,0 +1,145 @@
import { describe, expect, it } from 'vitest';
import {
loadSettings,
saveSettings,
setActiveProvider,
getActiveProviderConfig,
isProviderConfigured,
clearSettings,
getProviderDisplayName,
getAvailableModels,
} from '../../src/core/llm/settings-service';
describe('loadSettings', () => {
it('returns defaults when nothing is stored', () => {
const settings = loadSettings();
expect(settings.activeProvider).toBeDefined();
expect(settings.openai).toBeDefined();
expect(settings.ollama).toBeDefined();
});
it('merges stored values with defaults', () => {
sessionStorage.setItem('gitnexus-llm-settings', JSON.stringify({
activeProvider: 'ollama',
ollama: { model: 'qwen3-coder:30b' },
}));
const settings = loadSettings();
expect(settings.activeProvider).toBe('ollama');
expect(settings.ollama.model).toBe('qwen3-coder:30b');
// Should still have other provider defaults
expect(settings.openai).toBeDefined();
});
it('returns defaults on corrupted JSON', () => {
sessionStorage.setItem('gitnexus-llm-settings', 'not-json{{{');
const settings = loadSettings();
expect(settings.activeProvider).toBeDefined();
});
it('migrates legacy localStorage to sessionStorage', () => {
localStorage.setItem('gitnexus-llm-settings', JSON.stringify({
activeProvider: 'ollama',
ollama: { model: 'migrated-model' },
}));
const settings = loadSettings();
expect(settings.ollama.model).toBe('migrated-model');
expect(sessionStorage.getItem('gitnexus-llm-settings')).not.toBeNull();
expect(localStorage.getItem('gitnexus-llm-settings')).toBeNull();
});
});
describe('saveSettings / clearSettings', () => {
it('persists settings to sessionStorage', () => {
const settings = loadSettings();
settings.activeProvider = 'anthropic';
saveSettings(settings);
expect(loadSettings().activeProvider).toBe('anthropic');
});
it('clearSettings removes settings from both storages', () => {
saveSettings({ ...loadSettings(), activeProvider: 'anthropic' });
expect(sessionStorage.getItem('gitnexus-llm-settings')).not.toBeNull();
clearSettings();
expect(sessionStorage.getItem('gitnexus-llm-settings')).toBeNull();
expect(localStorage.getItem('gitnexus-llm-settings')).toBeNull();
});
});
describe('setActiveProvider', () => {
it('changes the active provider and persists', () => {
setActiveProvider('gemini');
expect(loadSettings().activeProvider).toBe('gemini');
});
});
describe('getActiveProviderConfig', () => {
it('returns null for unconfigured providers requiring API keys', () => {
setActiveProvider('openai');
expect(getActiveProviderConfig()).toBeNull();
});
it('returns config for ollama without API key', () => {
setActiveProvider('ollama');
const config = getActiveProviderConfig();
expect(config).not.toBeNull();
expect(config!.provider).toBe('ollama');
});
it('returns config for openai when API key is set', () => {
const settings = loadSettings();
settings.activeProvider = 'openai';
settings.openai = { ...settings.openai, apiKey: 'sk-test-123' };
saveSettings(settings);
const config = getActiveProviderConfig();
expect(config).not.toBeNull();
expect(config!.provider).toBe('openai');
});
it('returns null for openrouter with empty API key', () => {
const settings = loadSettings();
settings.activeProvider = 'openrouter';
settings.openrouter = { ...settings.openrouter, apiKey: ' ' };
saveSettings(settings);
expect(getActiveProviderConfig()).toBeNull();
});
});
describe('isProviderConfigured', () => {
it('returns false when provider requires API key and none is set', () => {
// Manually build a clean openai config with no API key
saveSettings({ ...loadSettings(), activeProvider: 'openai', openai: { apiKey: '', model: 'gpt-4o', temperature: 0.1 } });
expect(isProviderConfigured()).toBe(false);
});
it('returns true for ollama (no key required)', () => {
setActiveProvider('ollama');
expect(isProviderConfigured()).toBe(true);
});
});
describe('getProviderDisplayName', () => {
it('returns human-readable names', () => {
expect(getProviderDisplayName('openai')).toBe('OpenAI');
expect(getProviderDisplayName('azure-openai')).toBe('Azure OpenAI');
expect(getProviderDisplayName('gemini')).toBe('Google Gemini');
expect(getProviderDisplayName('anthropic')).toBe('Anthropic');
expect(getProviderDisplayName('ollama')).toBe('Ollama (Local)');
expect(getProviderDisplayName('openrouter')).toBe('OpenRouter');
});
});
describe('getAvailableModels', () => {
it('returns models for known providers', () => {
expect(getAvailableModels('openai').length).toBeGreaterThan(0);
expect(getAvailableModels('ollama').length).toBeGreaterThan(0);
expect(getAvailableModels('anthropic')).toContain('claude-sonnet-4-20250514');
});
it('returns empty array for unknown provider', () => {
expect(getAvailableModels('unknown' as any)).toEqual([]);
});
});
+17
View File
@@ -0,0 +1,17 @@
import { describe, expect, it } from 'vitest';
import { generateId } from '../../src/lib/utils';
describe('generateId', () => {
it('creates label:name format', () => {
expect(generateId('File', 'index.ts')).toBe('File:index.ts');
expect(generateId('Function', 'main')).toBe('Function:main');
});
it('handles empty strings', () => {
expect(generateId('', '')).toBe(':');
});
it('preserves special characters in name', () => {
expect(generateId('File', 'src/components/App.tsx')).toBe('File:src/components/App.tsx');
});
});
+30 -7
View File
@@ -1,17 +1,40 @@
import { defineConfig } from 'vitest/config';
import react from '@vitejs/plugin-react';
import path from 'path';
export default defineConfig({
test: {
environment: 'jsdom',
globals: true,
setupFiles: ['test/setup.ts'],
include: ['test/**/*.test.ts'],
exclude: ['**/node_modules/**', '**/dist/**'],
},
plugins: [react()],
resolve: {
alias: {
'@': path.resolve(__dirname, './src'),
'@anthropic-ai/sdk/lib/transform-json-schema': path.resolve(__dirname, 'node_modules/@anthropic-ai/sdk/lib/transform-json-schema.mjs'),
'mermaid': path.resolve(__dirname, 'node_modules/mermaid/dist/mermaid.esm.min.mjs'),
},
},
test: {
globals: true,
environment: 'jsdom',
setupFiles: ['./test/setup.ts'],
include: ['test/**/*.test.{ts,tsx}'],
testTimeout: 15000,
coverage: {
provider: 'v8',
include: ['src/**/*.{ts,tsx}'],
exclude: [
'src/workers/**', // Web workers (require worker env)
'src/core/lbug/**', // WASM (requires SharedArrayBuffer)
'src/core/tree-sitter/**', // WASM (requires tree-sitter binaries)
'src/core/embeddings/**', // WASM (requires ML model)
'src/main.tsx', // Entry point
'src/vite-env.d.ts', // Type declarations
],
thresholds: {
statements: 10,
branches: 10,
functions: 10,
lines: 10,
},
},
},
});
+1 -1
View File
@@ -1,4 +1,4 @@
FROM node:22-bookworm
FROM node:20-bookworm
WORKDIR /app
RUN apt-get update && apt-get install -y python3 make g++ && rm -rf /var/lib/apt/lists/*
COPY . .
+4
View File
@@ -53,7 +53,11 @@ If you prefer to configure manually instead of using `gitnexus setup`:
### Claude Code (full support — MCP + skills + hooks)
```bash
# macOS / Linux
claude mcp add gitnexus -- npx -y gitnexus@latest mcp
# Windows
claude mcp add gitnexus -- cmd /c npx -y gitnexus@latest mcp
```
### Codex (full support — MCP + skills)
+14 -14
View File
@@ -1,12 +1,12 @@
{
"name": "gitnexus",
"version": "1.4.7",
"version": "1.4.8",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "gitnexus",
"version": "1.4.7",
"version": "1.4.8",
"hasInstallScript": true,
"license": "PolyForm-Noncommercial-1.0.0",
"dependencies": {
@@ -26,7 +26,7 @@
"mnemonist": "^0.39.0",
"onnxruntime-node": "^1.24.0",
"pandemonium": "^2.4.0",
"tree-sitter": "^0.21.0",
"tree-sitter": "0.22.4",
"tree-sitter-c": "^0.21.0",
"tree-sitter-c-sharp": "^0.21.0",
"tree-sitter-cpp": "^0.22.0",
@@ -56,11 +56,11 @@
"vitest": "^4.0.18"
},
"engines": {
"node": ">=18.0.0"
"node": ">=20.0.0"
},
"optionalDependencies": {
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-swift": "^0.6.0"
"tree-sitter-swift": "0.7.1"
}
},
"node_modules/@babel/helper-string-parser": {
@@ -4958,14 +4958,14 @@
}
},
"node_modules/tree-sitter": {
"version": "0.21.1",
"resolved": "https://registry.npmjs.org/tree-sitter/-/tree-sitter-0.21.1.tgz",
"integrity": "sha512-7dxoA6kYvtgWw80265MyqJlkRl4yawIjO7S5MigytjELkX43fV2WsAXzsNfO7sBpPPCF5Gp0+XzHk0DwLCq3xQ==",
"version": "0.22.4",
"resolved": "https://registry.npmjs.org/tree-sitter/-/tree-sitter-0.22.4.tgz",
"integrity": "sha512-usbHZP9/oxNsUY65MQUsduGRqDHQOou1cagUSwjhoSYAmSahjQDAVsh9s+SlZkn8X8+O1FULRGwHu7AFP3kjzg==",
"hasInstallScript": true,
"license": "MIT",
"dependencies": {
"node-addon-api": "^8.0.0",
"node-gyp-build": "^4.8.0"
"node-addon-api": "^8.3.0",
"node-gyp-build": "^4.8.4"
}
},
"node_modules/tree-sitter-c": {
@@ -5284,9 +5284,9 @@
"license": "MIT"
},
"node_modules/tree-sitter-swift": {
"version": "0.6.0",
"resolved": "https://registry.npmjs.org/tree-sitter-swift/-/tree-sitter-swift-0.6.0.tgz",
"integrity": "sha512-9vOJZes4/UFjBr4COHtp6ZHVuZYwfChSQbpneXQog04dAstfx5px3ybVX2cN+ylvLqsvVpmXLpidxxgF2rDQ7A==",
"version": "0.7.1",
"resolved": "https://registry.npmjs.org/tree-sitter-swift/-/tree-sitter-swift-0.7.1.tgz",
"integrity": "sha512-pneKVTuGamaBsqqqfB9BvNQjktzh/0IVPR54jLB5Fq/JTDQwYHd0Wo6pVyZ5jAYpbztzq+rJ/rpL9ruxTmSoKw==",
"hasInstallScript": true,
"license": "MIT",
"optional": true,
@@ -5297,7 +5297,7 @@
"which": "2.0.2"
},
"peerDependencies": {
"tree-sitter": "^0.21.1"
"tree-sitter": "^0.22.1"
},
"peerDependenciesMeta": {
"tree_sitter": {
+5 -4
View File
@@ -66,7 +66,7 @@
"mnemonist": "^0.39.0",
"onnxruntime-node": "^1.24.0",
"pandemonium": "^2.4.0",
"tree-sitter": "^0.21.0",
"tree-sitter": "0.22.4",
"tree-sitter-c": "^0.21.0",
"tree-sitter-c-sharp": "^0.21.0",
"tree-sitter-cpp": "^0.22.0",
@@ -82,7 +82,7 @@
},
"optionalDependencies": {
"tree-sitter-kotlin": "^0.3.8",
"tree-sitter-swift": "^0.6.0"
"tree-sitter-swift": "0.7.1"
},
"devDependencies": {
"@types/cli-progress": "^3.11.6",
@@ -99,9 +99,10 @@
"overrides": {
"@huggingface/transformers": {
"onnxruntime-node": "$onnxruntime-node"
}
},
"tree-sitter": "0.22.4"
},
"engines": {
"node": ">=18.0.0"
"node": ">=20.0.0"
}
}
+4 -2
View File
@@ -41,8 +41,10 @@ try {
let needsRebuild = false;
if (content.includes('"actions"')) {
// Strip Python-style comments (#) before JSON parsing
const cleaned = content.replace(/#[^\n]*/g, '');
// Strip Python-style comments (#) and trailing commas before JSON parsing
const cleaned = content
.replace(/#[^\n]*/g, '') // Remove # comments
.replace(/,(\s*[\]}])/g, '$1'); // Remove trailing commas before ] or }
const gyp = JSON.parse(cleaned);
if (gyp.targets && gyp.targets[0] && gyp.targets[0].actions) {
+11 -10
View File
@@ -111,20 +111,21 @@ async function setupCursor(result: SetupResult): Promise<void> {
async function setupClaudeCode(result: SetupResult): Promise<void> {
const claudeDir = path.join(os.homedir(), '.claude');
const hasClaude = await dirExists(claudeDir);
if (!hasClaude) {
if (!(await dirExists(claudeDir))) {
result.skipped.push('Claude Code (not installed)');
return;
}
// Claude Code uses a JSON settings file at ~/.claude.json or claude mcp add
console.log('');
console.log(' Claude Code detected. Run this command to add GitNexus MCP:');
console.log('');
console.log(' claude mcp add gitnexus -- npx -y gitnexus mcp');
console.log('');
result.configured.push('Claude Code (MCP manual step printed)');
// Claude Code stores MCP config in ~/.claude.json
const mcpPath = path.join(os.homedir(), '.claude.json');
try {
const existing = await readJsonFile(mcpPath);
const updated = mergeMcpConfig(existing);
await writeJsonFile(mcpPath, updated);
result.configured.push('Claude Code');
} catch (err: any) {
result.errors.push(`Claude Code: ${err.message}`);
}
}
/**
+3 -4
View File
@@ -9,14 +9,13 @@
* ----------------------------------|------------------------------------------|---------------------------
* tree-sitter-queries.ts | Query string + LANGUAGE_QUERIES entry | (required)
* export-detection.ts | ExportChecker function + table entry | (required)
* import-resolution.ts | Resolver in importResolvers | resolveStandard(...)
* import-resolution.ts | namedBindingExtractors entry | undefined
* call-routing.ts | callRouters entry | noRouting
* import-resolvers/<lang>.ts | Exported resolve<Lang>Import function | resolveStandard(...)
* call-routing.ts | CallRouter function (or noRouting) | noRouting
* entry-point-scoring.ts | ENTRY_POINT_PATTERNS entry | []
* framework-detection.ts | AST_FRAMEWORK_PATTERNS entry | []
* type-extractors/<lang>.ts | New file + index.ts import | (required)
* resolvers/<lang>.ts | Resolver file (if non-standard) | (only if resolveStandard insufficient)
* named-binding-extraction.ts | Extractor (if named imports) | (only if language has named imports)
* named-bindings/<lang>.ts | Extractor (if named imports) | (only if language has named imports)
*
* 4. Also check these files for language-specific if-checks (no compile-time guard):
* - mro-processor.ts (MRO strategy selection)
+14 -1
View File
@@ -33,7 +33,9 @@ export type NodeLabel =
| 'Annotation'
| 'Constructor'
| 'Template'
| 'Section';
| 'Section'
| 'Route' // API route endpoint (e.g., /api/grants)
| 'Tool'; // MCP tool definition
import { SupportedLanguages } from '../../config/supported-languages.js';
@@ -69,6 +71,12 @@ export type NodeProperties = {
// Section-specific (markdown heading level, 1-6)
level?: number,
returnType?: string,
// Response shape (top-level keys from NextResponse.json({...}) / res.json({...}))
responseKeys?: string[],
// Error response shape (top-level keys from .json() calls with status >= 400)
errorKeys?: string[],
// Middleware wrapper chain (outermost first): ['withRateLimit', 'withCSRF', 'withAuth']
middleware?: string[],
}
export type RelationshipType =
@@ -87,6 +95,11 @@ export type RelationshipType =
| 'ACCESSES'
| 'MEMBER_OF'
| 'STEP_IN_PROCESS'
| 'HANDLES_ROUTE' // Function/File → Route (handler serves this endpoint)
| 'FETCHES' // Function/File → Route (consumer calls this endpoint)
| 'HANDLES_TOOL' // Function/File → Tool (handler implements this tool)
| 'ENTRY_POINT_OF' // Route/Tool → Process (this endpoint starts this execution flow)
| 'WRAPS' // Function → Function (middleware wrapper chain) — Reserved: future middleware graph traversal (not yet emitted)
export interface GraphNode {
id: string,
+257 -30
View File
@@ -5,33 +5,30 @@ import Parser from 'tree-sitter';
import type { ResolutionContext } from './resolution-context.js';
import { TIER_CONFIDENCE, type ResolutionTier } from './resolution-context.js';
import { isLanguageAvailable, loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { getProvider } from './languages/index.js';
import { generateId } from '../../lib/utils.js';
import { getLanguageFromFilename } from './utils/language-detection.js';
import { isVerboseIngestionEnabled } from './utils/verbose.js';
import { yieldToEventLoop } from './utils/event-loop.js';
import { FUNCTION_NODE_TYPES, extractFunctionName, findEnclosingClassId } from './utils/ast-helpers.js';
import { isBuiltInOrNoise } from './utils/noise-filter.js';
import {
getLanguageFromFilename,
isVerboseIngestionEnabled,
yieldToEventLoop,
FUNCTION_NODE_TYPES,
extractFunctionName,
isBuiltInOrNoise,
countCallArguments,
inferCallForm,
extractReceiverName,
extractReceiverNode,
findEnclosingClassId,
CALL_EXPRESSION_TYPES,
extractMixedChain,
type MixedChainStep,
} from './utils.js';
} from './utils/call-analysis.js';
import { buildTypeEnv, isSubclassOf } from './type-env.js';
import type { ConstructorBinding } from './type-env.js';
import { getTreeSitterBufferSize } from './constants.js';
import type { ExtractedCall, ExtractedAssignment, ExtractedHeritage, ExtractedRoute, FileConstructorBindings } from './workers/parse-worker.js';
import { callRouters } from './call-routing.js';
import type { ExtractedCall, ExtractedAssignment, ExtractedHeritage, ExtractedRoute, ExtractedFetchCall, FileConstructorBindings } from './workers/parse-worker.js';
import { normalizeFetchURL, routeMatches } from './route-extractors/nextjs.js';
import { extractReturnTypeName, stripNullable } from './type-extractors/shared.js';
import { typeConfigs } from './type-extractors/index.js';
import type { LiteralTypeInferrer } from './type-extractors/types.js';
import type { SyntaxNode } from './utils.js';
import type { SyntaxNode } from './utils/ast-helpers.js';
/** Per-file resolved type bindings for exported symbols.
* Populated during call processing, consumed by Phase 14 re-resolution pass. */
@@ -84,12 +81,12 @@ export function buildImportedRawReturnTypes(
/** Collect resolved type bindings for exported file-scope symbols.
* Uses graph node isExported flag — does NOT require isExported on SymbolDefinition. */
function collectExportedBindings(
typeEnv: { readonly env: ReadonlyMap<string, ReadonlyMap<string, string>> },
typeEnv: { fileScope(): ReadonlyMap<string, string> },
filePath: string,
symbolTable: { lookupExact(filePath: string, name: string): string | undefined },
graph: { getNode(id: string): { properties?: { isExported?: boolean } } | undefined },
): Map<string, string> | null {
const fileScope = typeEnv.env.get('');
const fileScope = typeEnv.fileScope();
if (!fileScope || fileScope.size === 0) return null;
const exported = new Map<string, string>();
@@ -189,7 +186,8 @@ const TYPE_PRESERVING_METHODS = new Set([
const findEnclosingFunction = (
node: SyntaxNode,
filePath: string,
ctx: ResolutionContext
ctx: ResolutionContext,
provider: import('./language-provider.js').LanguageProvider,
): string | null => {
let current = node.parent;
@@ -203,7 +201,13 @@ const findEnclosingFunction = (
return resolved.candidates[0].nodeId;
}
return generateId(label, `${filePath}:${funcName}`);
// Apply labelOverride so label matches the definition phase (single source of truth).
let finalLabel = label;
if (provider.labelOverride) {
const override = provider.labelOverride(current, label);
if (override !== null) finalLabel = override;
}
return generateId(finalLabel, `${filePath}:${funcName}`);
}
}
current = current.parent;
@@ -312,7 +316,8 @@ export const processCalls = async (
continue;
}
const queryStr = LANGUAGE_QUERIES[language];
const provider = getProvider(language);
const queryStr = provider.treeSitterQueries;
if (!queryStr) continue;
await loadLanguage(language, file.path);
@@ -338,8 +343,6 @@ export const processCalls = async (
continue;
}
const lang = getLanguageFromFilename(file.path);
// Pre-pass: extract heritage from query matches to build parentMap for buildTypeEnv.
// Heritage-processor runs in PARALLEL, so graph edges don't exist when buildTypeEnv runs.
const fileParentMap = new Map<string, string[]>();
@@ -373,14 +376,14 @@ export const processCalls = async (
const importedBindings = importedBindingsMap?.get(file.path);
const importedReturnTypes = importedReturnTypesMap?.get(file.path);
const importedRawReturnTypes = importedRawReturnTypesMap?.get(file.path);
const typeEnv = lang ? buildTypeEnv(tree, lang, { symbolTable: ctx.symbols, parentMap, importedBindings, importedReturnTypes, importedRawReturnTypes }) : null;
const typeEnv = buildTypeEnv(tree, language, { symbolTable: ctx.symbols, parentMap, importedBindings, importedReturnTypes, importedRawReturnTypes });
if (typeEnv && exportedTypeMap) {
const fileExports = collectExportedBindings(typeEnv, file.path, ctx.symbols, graph);
if (fileExports) exportedTypeMap.set(file.path, fileExports);
}
const callRouter = callRouters[language];
const callRouter = provider.callRouter;
const verifiedReceivers = typeEnv && typeEnv.constructorBindings.length > 0
const verifiedReceivers = typeEnv.constructorBindings.length > 0
? verifyConstructorBindings(typeEnv.constructorBindings, file.path, ctx)
: new Map<string, string>();
const receiverIndex = buildReceiverTypeIndex(verifiedReceivers);
@@ -402,7 +405,7 @@ export const processCalls = async (
}
// Fall back to verified constructor bindings (mirrors CALLS resolution tier 2)
if (!receiverTypeName && receiverText && receiverIndex.size > 0) {
const enclosing = findEnclosingFunction(captureMap['assignment'], file.path, ctx);
const enclosing = findEnclosingFunction(captureMap['assignment'], file.path, ctx, provider);
const funcName = enclosing ? extractFuncNameFromSourceId(enclosing) : '';
receiverTypeName = lookupReceiverType(receiverIndex, funcName, receiverText);
}
@@ -416,7 +419,7 @@ export const processCalls = async (
}
}
if (receiverTypeName) {
const enclosing = findEnclosingFunction(captureMap['assignment'], file.path, ctx);
const enclosing = findEnclosingFunction(captureMap['assignment'], file.path, ctx, provider);
const srcId = enclosing || generateId('File', file.path);
// Defer resolution: Ruby attr_accessor properties are registered during
// this same loop, so cross-file lookups fail if the declaring file hasn't
@@ -435,7 +438,7 @@ export const processCalls = async (
const calledName = nameNode.text;
const routed = callRouter(calledName, captureMap['call']);
const routed = callRouter?.(calledName, captureMap['call']);
if (routed) {
switch (routed.kind) {
case 'skip':
@@ -536,7 +539,7 @@ export const processCalls = async (
}
// Fall back to verified constructor bindings for return type inference
if (!receiverTypeName && receiverName && receiverIndex.size > 0) {
const enclosingFunc = findEnclosingFunction(callNode, file.path, ctx);
const enclosingFunc = findEnclosingFunction(callNode, file.path, ctx, provider);
const funcName = enclosingFunc ? extractFuncNameFromSourceId(enclosingFunc) : '';
receiverTypeName = lookupReceiverType(receiverIndex, funcName, receiverName);
}
@@ -552,7 +555,7 @@ export const processCalls = async (
}
}
// Hoist sourceId so it's available for ACCESSES edge emission during chain walk.
const enclosingFuncId = findEnclosingFunction(callNode, file.path, ctx);
const enclosingFuncId = findEnclosingFunction(callNode, file.path, ctx, provider);
const sourceId = enclosingFuncId || generateId('File', file.path);
// Fall back to mixed chain resolution when the receiver is a complex expression
@@ -590,7 +593,7 @@ export const processCalls = async (
// Build overload hints for languages with inferLiteralType (Java/Kotlin/C#/C++).
// Only used when multiple candidates survive arity filtering — ~1-3% of calls.
const langConfig = lang ? typeConfigs[lang as keyof typeof typeConfigs] : undefined;
const langConfig = provider.typeConfig;
const hints: OverloadHints | undefined = langConfig?.inferLiteralType
? { callNode, inferLiteralType: langConfig.inferLiteralType }
: undefined;
@@ -830,6 +833,18 @@ const resolveCallTarget = (
let filteredCandidates = filterCallableCandidates(tiered.candidates, call.argCount, call.callForm);
// Swift/Kotlin: constructor calls look like free function calls (no `new` keyword).
// If free-form filtering found no callable candidates but the symbol resolves to a
// Class/Struct, retry with constructor form so CONSTRUCTOR_TARGET_TYPES applies.
if (filteredCandidates.length === 0 && call.callForm === 'free') {
const hasTypeTarget = tiered.candidates.some(c =>
c.type === 'Class' || c.type === 'Struct' || c.type === 'Enum',
);
if (hasTypeTarget) {
filteredCandidates = filterCallableCandidates(tiered.candidates, call.argCount, 'constructor');
}
}
// Module-qualified constructor pattern: e.g. Python `import models; models.User()`.
// The attribute access gives callForm='member', but the callee may be a Class — a valid
// constructor target. Re-try with constructor-form filtering so that `module.ClassName()`
@@ -904,7 +919,20 @@ const resolveCallTarget = (
if (disambiguated) return toResolveResult(disambiguated, tiered.tier);
}
if (filteredCandidates.length !== 1) return null;
if (filteredCandidates.length !== 1) {
// Deduplicate: Swift extensions create multiple Class nodes with the same name.
// When all candidates share the same type and differ only by file (extension vs
// primary definition), they represent the same symbol. Prefer the primary
// definition (shortest file path: Product.swift over ProductExtension.swift).
if (filteredCandidates.length > 1) {
const allSameType = filteredCandidates.every(c => c.type === filteredCandidates[0].type);
if (allSameType && (filteredCandidates[0].type === 'Class' || filteredCandidates[0].type === 'Struct')) {
const sorted = [...filteredCandidates].sort((a, b) => a.filePath.length - b.filePath.length);
return toResolveResult(sorted[0], tiered.tier);
}
}
return null;
}
return toResolveResult(filteredCandidates[0], tiered.tier);
};
@@ -1358,3 +1386,202 @@ export const processRoutesFromExtracted = async (
onProgress?.(extractedRoutes.length, extractedRoutes.length);
};
/**
* Extract property access keys from a consumer file's source code near fetch calls.
*
* Looks for three patterns after a fetch/response variable assignment:
* 1. Destructuring: `const { data, pagination } = await res.json()`
* 2. Property access: `response.data`, `result.items`
* 3. Optional chaining: `data?.key1?.key2`
*
* Returns deduplicated top-level property names accessed on the response.
*
* NOTE: This scans the entire file content, not just code near a specific fetch call.
* If a file has multiple fetch calls to different routes, all accessed keys are
* attributed to each fetch. This is an acceptable tradeoff for regex-based extraction.
*/
/** Common method names on response/data objects that are NOT property accesses */
const RESPONSE_METHOD_BLOCKLIST = new Set([
'json', 'text', 'blob', 'arrayBuffer', 'formData', 'ok', 'status', 'headers',
'then', 'catch', 'finally', 'clone',
'map', 'filter', 'forEach', 'reduce', 'find', 'some', 'every',
'length', 'toString', 'valueOf',
'push', 'pop', 'shift', 'unshift', 'splice', 'slice', 'concat', 'join',
'sort', 'reverse', 'includes', 'indexOf', 'keys', 'values', 'entries',
]);
export const extractConsumerAccessedKeys = (content: string): string[] => {
const keys = new Set<string>();
// Pattern 1: Destructuring from .json() — const { key1, key2 } = await res.json()
// Also matches: const { key1, key2 } = await (await fetch(...)).json()
const destructurePattern = /(?:const|let|var)\s+\{([^}]+)\}\s*=\s*(?:await\s+)?(?:\w+\.json\s*\(\)|(?:await\s+)?(?:fetch|axios|got)\s*\([^)]*\)(?:\.then\s*\([^)]*\))?(?:\.json\s*\(\))?)/g;
let match;
while ((match = destructurePattern.exec(content)) !== null) {
const destructuredBody = match[1];
// Extract identifiers from destructuring, handling renamed bindings (key: alias)
const keyPattern = /(\w+)\s*(?::\s*\w+)?/g;
let keyMatch;
while ((keyMatch = keyPattern.exec(destructuredBody)) !== null) {
keys.add(keyMatch[1]);
}
}
// Pattern 2: Destructuring from a data/result/response/json variable
// e.g., const { items, total } = data; or const { error } = result;
const dataVarDestructure = /(?:const|let|var)\s+\{([^}]+)\}\s*=\s*(?:data|result|response|json|body|res)\b/g;
while ((match = dataVarDestructure.exec(content)) !== null) {
const destructuredBody = match[1];
const keyPattern = /(\w+)\s*(?::\s*\w+)?/g;
let keyMatch;
while ((keyMatch = keyPattern.exec(destructuredBody)) !== null) {
keys.add(keyMatch[1]);
}
}
// Pattern 3: Property access on common response variable names
// Matches: data.key, response.key, result.key, json.key, body.key
// Also matches optional chaining: data?.key
const propAccessPattern = /\b(?:data|response|result|json|body|res)\s*(?:\?\.|\.)(\w+)/g;
while ((match = propAccessPattern.exec(content)) !== null) {
const key = match[1];
// Skip common method calls that aren't property accesses
if (!RESPONSE_METHOD_BLOCKLIST.has(key)) {
keys.add(key);
}
}
return [...keys];
};
/**
* Create FETCHES edges from extracted fetch() calls to matching Route nodes.
* When consumerContents is provided, extracts property access patterns from
* consumer files and encodes them in the edge reason field.
*/
export const processNextjsFetchRoutes = (
graph: KnowledgeGraph,
fetchCalls: ExtractedFetchCall[],
routeRegistry: Map<string, string>, // routeURL → handlerFilePath
consumerContents?: Map<string, string>, // filePath → file content
) => {
// Pre-count how many routes each consumer file matches (for confidence attribution)
const routeCountByFile = new Map<string, number>();
for (const call of fetchCalls) {
const normalized = normalizeFetchURL(call.fetchURL);
if (!normalized) continue;
for (const [routeURL] of routeRegistry) {
if (routeMatches(normalized, routeURL)) {
routeCountByFile.set(call.filePath, (routeCountByFile.get(call.filePath) ?? 0) + 1);
break;
}
}
}
for (const call of fetchCalls) {
const normalized = normalizeFetchURL(call.fetchURL);
if (!normalized) continue;
for (const [routeURL] of routeRegistry) {
if (routeMatches(normalized, routeURL)) {
const sourceId = generateId('File', call.filePath);
const routeNodeId = generateId('Route', routeURL);
// Extract consumer accessed keys if file content is available
let reason = 'fetch-url-match';
if (consumerContents) {
const content = consumerContents.get(call.filePath);
if (content) {
const accessedKeys = extractConsumerAccessedKeys(content);
if (accessedKeys.length > 0) {
reason = `fetch-url-match|keys:${accessedKeys.join(',')}`;
}
}
}
// Encode multi-fetch count so downstream can set confidence
const fetchCount = routeCountByFile.get(call.filePath) ?? 1;
if (fetchCount > 1) {
reason = `${reason}|fetches:${fetchCount}`;
}
graph.addRelationship({
id: generateId('FETCHES', `${sourceId}->${routeNodeId}`),
sourceId,
targetId: routeNodeId,
type: 'FETCHES',
confidence: 0.9,
reason,
});
break;
}
}
}
};
/**
* Extract fetch() calls from source files (sequential path).
* Workers handle this via tree-sitter captures in parse-worker; this function
* provides the same extraction for the sequential fallback path.
*/
export const extractFetchCallsFromFiles = async (
files: { path: string; content: string }[],
astCache: ASTCache,
): Promise<ExtractedFetchCall[]> => {
const parser = await loadParser();
const result: ExtractedFetchCall[] = [];
for (const file of files) {
const language = getLanguageFromFilename(file.path);
if (!language) continue;
if (!isLanguageAvailable(language)) continue;
const provider = getProvider(language);
const queryStr = provider.treeSitterQueries;
if (!queryStr) continue;
await loadLanguage(language, file.path);
let tree = astCache.get(file.path);
if (!tree) {
try {
tree = parser.parse(file.content, undefined, { bufferSize: getTreeSitterBufferSize(file.content.length) });
} catch { continue; }
astCache.set(file.path, tree);
}
let matches;
try {
const lang = parser.getLanguage();
const query = new Parser.Query(lang, queryStr);
matches = query.matches(tree.rootNode);
} catch { continue; }
for (const match of matches) {
const captureMap: Record<string, any> = {};
match.captures.forEach(c => captureMap[c.name] = c.node);
if (captureMap['route.fetch']) {
const urlNode = captureMap['route.url'] ?? captureMap['route.template_url'];
if (urlNode) {
result.push({
filePath: file.path,
fetchURL: urlNode.text,
lineNumber: captureMap['route.fetch'].startPosition.row,
});
}
} else if (captureMap['http_client'] && captureMap['http_client.url']) {
const method = captureMap['http_client.method']?.text;
const url = captureMap['http_client.url'].text;
const HTTP_CLIENT_ONLY = new Set(['head', 'options', 'request', 'ajax']);
if (method && HTTP_CLIENT_ONLY.has(method) && url.startsWith('/')) {
result.push({ filePath: file.path, fetchURL: url, lineNumber: captureMap['http_client'].startPosition.row });
}
}
}
}
return result;
};
+3 -23
View File
@@ -11,7 +11,7 @@
* Keep both copies in sync until a shared package is introduced.
*/
import { SupportedLanguages } from '../../config/supported-languages.js';
import type { SyntaxNode } from './utils/ast-helpers.js';
// ── Call routing dispatch table ─────────────────────────────────────────────
@@ -26,29 +26,9 @@ export type CallRoutingResult = RubyCallRouting | null;
*/
export type CallRouter = (
calledName: string,
callNode: any,
callNode: SyntaxNode,
) => CallRoutingResult;
/** No-op router: returns null for every call (passthrough to normal processing) */
const noRouting: CallRouter = () => null;
/** Per-language call routing. noRouting = no special routing (normal call processing) */
export const callRouters = {
[SupportedLanguages.JavaScript]: noRouting,
[SupportedLanguages.TypeScript]: noRouting,
[SupportedLanguages.Python]: noRouting,
[SupportedLanguages.Java]: noRouting,
[SupportedLanguages.Kotlin]: noRouting,
[SupportedLanguages.Go]: noRouting,
[SupportedLanguages.Rust]: noRouting,
[SupportedLanguages.CSharp]: noRouting,
[SupportedLanguages.PHP]: noRouting,
[SupportedLanguages.Swift]: noRouting,
[SupportedLanguages.CPlusPlus]: noRouting,
[SupportedLanguages.C]: noRouting,
[SupportedLanguages.Ruby]: routeRubyCall,
} satisfies Record<SupportedLanguages, CallRouter>;
// ── Result types ────────────────────────────────────────────────────────────
export type RubyCallRouting =
@@ -91,7 +71,7 @@ const MAX_PARENT_DEPTH = 50;
* @param callNode - The tree-sitter `call` AST node
* @returns A discriminated union describing the call's semantic role
*/
export function routeRubyCall(calledName: string, callNode: any): RubyCallRouting {
export function routeRubyCall(calledName: string, callNode: SyntaxNode): RubyCallRouting {
// ── require / require_relative → import ─────────────────────────────────
if (calledName === 'require' || calledName === 'require_relative') {
const argList = callNode.childForFieldName?.('arguments');
@@ -0,0 +1,627 @@
/**
* COBOL Processor
*
* Standalone regex-based processor for COBOL and JCL files.
* Follows the markdown-processor.ts pattern: takes (graph, files, allPathSet),
* does its own extraction, and writes directly to the graph.
*
* Pipeline:
* 1. Separate programs from copybooks
* 2. Build copybook map (name -> content)
* 3. For each program: expand COPY statements, then run regex extraction
* 4. Map CobolRegexResults to graph nodes and relationships
* 5. Optionally process JCL files for job-step cross-references
*/
import path from 'node:path';
import { generateId } from '../../lib/utils.js';
import type { KnowledgeGraph, GraphNode } from '../graph/types.js';
import {
preprocessCobolSource,
extractCobolSymbolsWithRegex,
type CobolRegexResults,
} from './cobol/cobol-preprocessor.js';
import { expandCopies } from './cobol/cobol-copy-expander.js';
import { processJclFiles } from './cobol/jcl-processor.js';
// ---------------------------------------------------------------------------
// File detection
// ---------------------------------------------------------------------------
const COBOL_EXTENSIONS = new Set([
'.cob', '.cbl', '.cobol', '.cpy', '.copybook',
]);
const JCL_EXTENSIONS = new Set(['.jcl', '.job', '.proc']);
const COPYBOOK_EXTENSIONS = new Set(['.cpy', '.copybook']);
interface CobolFile {
path: string;
content: string;
}
export interface CobolProcessResult {
programs: number;
paragraphs: number;
sections: number;
dataItems: number;
calls: number;
copies: number;
execSqlBlocks: number;
execCicsBlocks: number;
entryPoints: number;
moves: number;
fileDeclarations: number;
jclJobs: number;
jclSteps: number;
}
/** Returns true if the file is a COBOL or copybook file. */
export function isCobolFile(filePath: string): boolean {
return COBOL_EXTENSIONS.has(path.extname(filePath).toLowerCase());
}
/** Returns true if the file is a JCL file. */
export function isJclFile(filePath: string): boolean {
return JCL_EXTENSIONS.has(path.extname(filePath).toLowerCase());
}
/** Returns true if the file is a COBOL copybook. */
function isCopybook(filePath: string): boolean {
return COPYBOOK_EXTENSIONS.has(path.extname(filePath).toLowerCase());
}
// ---------------------------------------------------------------------------
// Main processor
// ---------------------------------------------------------------------------
/**
* Process COBOL and JCL files into the knowledge graph.
*
* @param graph - The in-memory knowledge graph
* @param files - Array of { path, content } for COBOL/JCL files
* @param allPathSet - Set of all file paths in the repository
* @returns Summary of what was extracted
*/
export const processCobol = (
graph: KnowledgeGraph,
files: CobolFile[],
allPathSet: Set<string>,
): CobolProcessResult => {
const result: CobolProcessResult = {
programs: 0,
paragraphs: 0,
sections: 0,
dataItems: 0,
calls: 0,
copies: 0,
execSqlBlocks: 0,
execCicsBlocks: 0,
entryPoints: 0,
moves: 0,
fileDeclarations: 0,
jclJobs: 0,
jclSteps: 0,
};
// ── 1. Separate programs, copybooks, and JCL ───────────────────────
const programs: CobolFile[] = [];
const copybooks: CobolFile[] = [];
const jclFiles: CobolFile[] = [];
for (const file of files) {
const ext = path.extname(file.path).toLowerCase();
if (JCL_EXTENSIONS.has(ext)) {
jclFiles.push(file);
} else if (isCopybook(file.path)) {
copybooks.push(file);
} else if (COBOL_EXTENSIONS.has(ext)) {
programs.push(file);
}
}
// ── 2. Build copybook map (uppercase name -> content) ──────────────
const copybookMap = new Map<string, { content: string; path: string }>();
for (const cb of copybooks) {
const name = path.basename(cb.path, path.extname(cb.path)).toUpperCase();
copybookMap.set(name, { content: cb.content, path: cb.path });
}
// Resolve and read callbacks for expandCopies
const resolveCopy = (name: string): string | null => {
const entry = copybookMap.get(name.toUpperCase());
return entry ? entry.path : null;
};
const readCopy = (copyPath: string): string | null => {
// Find by path match
for (const [, entry] of copybookMap) {
if (entry.path === copyPath) return entry.content;
}
return null;
};
// Track module names for cross-program CALL resolution
const moduleNodeIds = new Map<string, string>(); // uppercase program name -> node id
// ── 3. Process each COBOL program ──────────────────────────────────
for (const file of programs) {
const fileNodeId = generateId('File', file.path);
// Skip if file node doesn't exist (structure-processor creates it)
if (!graph.getNode(fileNodeId)) continue;
// Preprocess: clean patch markers
const cleaned = preprocessCobolSource(file.content);
// Expand COPY statements
const { expandedContent, copyResolutions } = expandCopies(
cleaned, file.path, resolveCopy, readCopy,
);
// Extract symbols from expanded source
const extracted = extractCobolSymbolsWithRegex(expandedContent, file.path);
// Map to graph
mapToGraph(graph, extracted, file, copyResolutions, moduleNodeIds);
// Accumulate stats
result.programs += extracted.programName ? 1 : 0;
result.paragraphs += extracted.paragraphs.length;
result.sections += extracted.sections.length;
result.dataItems += extracted.dataItems.length;
result.calls += extracted.calls.length;
result.copies += extracted.copies.length;
result.execSqlBlocks += extracted.execSqlBlocks.length;
result.execCicsBlocks += extracted.execCicsBlocks.length;
result.entryPoints += extracted.entryPoints.length;
result.moves += extracted.moves.length;
result.fileDeclarations += extracted.fileDeclarations.length;
}
// ── 4. Second pass: resolve cross-program CALL targets ─────────────
// During mapToGraph, early programs create unresolved CALL edges
// (target = <unresolved>:PROGNAME) because later programs haven't
// been registered in moduleNodeIds yet. Now that ALL programs are
// processed, re-scan unresolved CALLS edges and patch them.
// This covers both `cobol-call-unresolved` and CICS LINK/XCTL edges
// whose targets contain `<unresolved>:`.
graph.forEachRelationship(rel => {
if (rel.type !== 'CALLS') return;
const match = rel.targetId.match(/<unresolved>:(.+)/);
if (!match) return;
const resolvedId = moduleNodeIds.get(match[1]);
if (!resolvedId) return;
if (rel.reason?.startsWith('cobol-call-unresolved')) {
// Replace unresolved CALL with resolved edge
graph.addRelationship({
id: rel.id + ':resolved',
type: 'CALLS',
sourceId: rel.sourceId,
targetId: resolvedId,
confidence: 0.95,
reason: 'cobol-call',
});
} else if (rel.reason === 'cics-link-unresolved' || rel.reason === 'cics-xctl-unresolved') {
// Replace unresolved CICS LINK/XCTL with resolved edge
graph.addRelationship({
id: rel.id + ':resolved',
type: 'CALLS',
sourceId: rel.sourceId,
targetId: resolvedId,
confidence: 0.95,
reason: rel.reason.replace('-unresolved', ''),
});
}
});
// ── 5. Process JCL files ───────────────────────────────────────────
if (jclFiles.length > 0) {
const jclPaths = jclFiles.map(f => f.path);
const jclContents = new Map<string, string>();
for (const f of jclFiles) {
jclContents.set(f.path, f.content);
}
const jclResult = processJclFiles(graph, jclPaths, jclContents);
result.jclJobs += jclResult.jobCount;
result.jclSteps += jclResult.stepCount;
}
return result;
};
// ---------------------------------------------------------------------------
// Graph mapping
// ---------------------------------------------------------------------------
/** Resolve a data item name to its Property node id, if it exists and is not FILLER. */
function findDataItemNode(
name: string,
dataItems: CobolRegexResults['dataItems'],
filePath: string,
): string | undefined {
const item = dataItems.find(d => d.name.toUpperCase() === name.toUpperCase());
if (!item || item.name === 'FILLER') return undefined;
return generateId('Property', `${filePath}:${item.name}`);
}
function mapToGraph(
graph: KnowledgeGraph,
extracted: CobolRegexResults,
file: CobolFile,
copyResolutions: Array<{ copyTarget: string; resolvedPath: string | null; line: number }>,
moduleNodeIds: Map<string, string>,
): void {
const { path: filePath, content } = file;
const lines = content.split('\n');
const fileNodeId = generateId('File', filePath);
// ── PROGRAM-ID -> Module node ────────────────────────────────────
let moduleId: string | undefined;
if (extracted.programName) {
moduleId = generateId('Module', `${filePath}:${extracted.programName}`);
graph.addNode({
id: moduleId,
label: 'Module',
properties: {
name: extracted.programName,
filePath,
startLine: 1,
endLine: lines.length,
language: 'cobol' as any,
isExported: true,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${fileNodeId}->${moduleId}`),
type: 'CONTAINS',
sourceId: fileNodeId,
targetId: moduleId,
confidence: 1.0,
reason: 'cobol-program-id',
});
moduleNodeIds.set(extracted.programName.toUpperCase(), moduleId);
}
const parentId = moduleId ?? fileNodeId;
// ── SECTIONs -> Namespace nodes ──────────────────────────────────
const sectionNodeIds = new Map<string, string>();
for (let i = 0; i < extracted.sections.length; i++) {
const sec = extracted.sections[i];
const nextLine = i + 1 < extracted.sections.length
? extracted.sections[i + 1].line - 1
: lines.length;
const secId = generateId('Namespace', `${filePath}:${sec.name}`);
graph.addNode({
id: secId,
label: 'Namespace',
properties: {
name: sec.name,
filePath,
startLine: sec.line,
endLine: nextLine,
language: 'cobol' as any,
isExported: true,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${parentId}->${secId}`),
type: 'CONTAINS',
sourceId: parentId,
targetId: secId,
confidence: 1.0,
reason: 'cobol-section',
});
sectionNodeIds.set(sec.name.toUpperCase(), secId);
}
// ── PARAGRAPHs -> Function nodes ─────────────────────────────────
const paraNodeIds = new Map<string, string>();
for (let i = 0; i < extracted.paragraphs.length; i++) {
const para = extracted.paragraphs[i];
const nextLine = i + 1 < extracted.paragraphs.length
? extracted.paragraphs[i + 1].line - 1
: lines.length;
const paraId = generateId('Function', `${filePath}:${para.name}`);
graph.addNode({
id: paraId,
label: 'Function',
properties: {
name: para.name,
filePath,
startLine: para.line,
endLine: nextLine,
language: 'cobol' as any,
isExported: true,
},
});
// Parent: find the containing section, or fall back to module/file
const containerId = findContainingSection(para.line, extracted.sections, sectionNodeIds) ?? parentId;
graph.addRelationship({
id: generateId('CONTAINS', `${containerId}->${paraId}`),
type: 'CONTAINS',
sourceId: containerId,
targetId: paraId,
confidence: 1.0,
reason: 'cobol-paragraph',
});
paraNodeIds.set(para.name.toUpperCase(), paraId);
}
// ── Data items -> Property nodes ─────────────────────────────────
for (const item of extracted.dataItems) {
if (item.name === 'FILLER') continue; // Skip anonymous fillers
const propId = generateId('Property', `${filePath}:${item.name}`);
graph.addNode({
id: propId,
label: 'Property',
properties: {
name: item.name,
filePath,
startLine: item.line,
endLine: item.line,
language: 'cobol' as any,
description: `level:${item.level} section:${item.section}${item.pic ? ` pic:${item.pic}` : ''}`,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${parentId}->${propId}`),
type: 'CONTAINS',
sourceId: parentId,
targetId: propId,
confidence: 1.0,
reason: 'cobol-data-item',
});
}
// ── PERFORM -> CALLS relationship (intra-file) ──────────────────
for (const perf of extracted.performs) {
const targetId = paraNodeIds.get(perf.target.toUpperCase())
?? sectionNodeIds.get(perf.target.toUpperCase());
if (!targetId) continue;
// Source: the paragraph containing the PERFORM, or the module
const sourceId = perf.caller
? (paraNodeIds.get(perf.caller.toUpperCase()) ?? parentId)
: parentId;
graph.addRelationship({
id: generateId('CALLS', `${sourceId}->perform->${targetId}:L${perf.line}`),
type: 'CALLS',
sourceId,
targetId,
confidence: 1.0,
reason: 'cobol-perform',
});
// PERFORM THRU -> expanded CALLS edge to thru target
if (perf.thruTarget) {
const thruTargetId = paraNodeIds.get(perf.thruTarget.toUpperCase())
?? sectionNodeIds.get(perf.thruTarget.toUpperCase());
if (thruTargetId && thruTargetId !== targetId) {
graph.addRelationship({
id: generateId('CALLS', `${sourceId}->perform-thru->${thruTargetId}:L${perf.line}`),
type: 'CALLS',
sourceId,
targetId: thruTargetId,
confidence: 1.0,
reason: 'cobol-perform-thru',
});
}
}
}
// ── CALL -> CALLS relationship (cross-program) ──────────────────
for (const call of extracted.calls) {
const targetModuleId = moduleNodeIds.get(call.target.toUpperCase());
// Create edge even if target not yet known — use a synthetic target id
const targetId = targetModuleId
?? generateId('Module', `<unresolved>:${call.target.toUpperCase()}`);
graph.addRelationship({
id: generateId('CALLS', `${parentId}->call->${call.target}:L${call.line}`),
type: 'CALLS',
sourceId: parentId,
targetId,
confidence: targetModuleId ? 0.95 : 0.5,
reason: targetModuleId ? 'cobol-call' : 'cobol-call-unresolved',
});
}
// ── COPY -> IMPORTS relationship ─────────────────────────────────
for (const res of copyResolutions) {
if (!res.resolvedPath) continue;
const targetFileId = generateId('File', res.resolvedPath);
graph.addRelationship({
id: generateId('IMPORTS', `${fileNodeId}->${targetFileId}:${res.copyTarget}`),
type: 'IMPORTS',
sourceId: fileNodeId,
targetId: targetFileId,
confidence: 1.0,
reason: 'cobol-copy',
});
}
// ── EXEC SQL blocks -> CodeElement nodes + ACCESSES edges ──────
for (const sql of extracted.execSqlBlocks) {
const sqlId = generateId('CodeElement', `${filePath}:exec-sql:L${sql.line}`);
graph.addNode({
id: sqlId,
label: 'CodeElement',
properties: {
name: `EXEC SQL ${sql.operation}`,
filePath,
startLine: sql.line,
endLine: sql.line,
language: 'cobol' as any,
description: `tables:[${sql.tables.join(',')}] cursors:[${sql.cursors.join(',')}]`,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${parentId}->${sqlId}`),
type: 'CONTAINS',
sourceId: parentId,
targetId: sqlId,
confidence: 1.0,
reason: 'cobol-exec-sql',
});
// ACCESSES edges to tables
for (const table of sql.tables) {
const tableId = generateId('Record', `<db>:${table}`);
graph.addRelationship({
id: generateId('ACCESSES', `${sqlId}->${tableId}:${sql.operation}`),
type: 'ACCESSES',
sourceId: sqlId,
targetId: tableId,
confidence: 0.9,
reason: `sql-${sql.operation.toLowerCase()}`,
});
}
}
// ── EXEC CICS blocks -> CodeElement nodes + CALLS edges ────────
for (const cics of extracted.execCicsBlocks) {
const cicsId = generateId('CodeElement', `${filePath}:exec-cics:L${cics.line}`);
graph.addNode({
id: cicsId,
label: 'CodeElement',
properties: {
name: `EXEC CICS ${cics.command}`,
filePath,
startLine: cics.line,
endLine: cics.line,
language: 'cobol' as any,
description: cics.mapName ? `map:${cics.mapName}` : cics.programName ? `program:${cics.programName}` : undefined,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${parentId}->${cicsId}`),
type: 'CONTAINS',
sourceId: parentId,
targetId: cicsId,
confidence: 1.0,
reason: 'cobol-exec-cics',
});
// LINK/XCTL -> cross-program CALLS
if (cics.programName && (cics.command === 'LINK' || cics.command === 'XCTL')) {
const cicsTargetModuleId = moduleNodeIds.get(cics.programName.toUpperCase());
const targetId = cicsTargetModuleId
?? generateId('Module', `<unresolved>:${cics.programName.toUpperCase()}`);
const cicsReason = `cics-${cics.command.toLowerCase()}`;
graph.addRelationship({
id: generateId('CALLS', `${parentId}->cics-${cics.command.toLowerCase()}->${cics.programName}:L${cics.line}`),
type: 'CALLS',
sourceId: parentId,
targetId,
confidence: cicsTargetModuleId ? 0.95 : 0.5,
reason: cicsTargetModuleId ? cicsReason : `${cicsReason}-unresolved`,
});
}
}
// ── ENTRY points -> Constructor nodes ──────────────────────────
for (const entry of extracted.entryPoints) {
const entryId = generateId('Constructor', `${filePath}:${entry.name}`);
graph.addNode({
id: entryId,
label: 'Constructor',
properties: {
name: entry.name,
filePath,
startLine: entry.line,
endLine: entry.line,
language: 'cobol' as any,
isExported: true,
description: entry.parameters.length > 0 ? `using:${entry.parameters.join(',')}` : undefined,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${parentId}->${entryId}`),
type: 'CONTAINS',
sourceId: parentId,
targetId: entryId,
confidence: 1.0,
reason: 'cobol-entry-point',
});
// Register in moduleNodeIds for cross-program resolution
moduleNodeIds.set(entry.name.toUpperCase(), entryId);
}
// ── MOVE data flow -> ACCESSES edges (read/write) ──────────────
for (const move of extracted.moves) {
const fromPropId = findDataItemNode(move.from, extracted.dataItems, filePath);
const toPropId = findDataItemNode(move.to, extracted.dataItems, filePath);
const callerId = move.caller
? (paraNodeIds.get(move.caller.toUpperCase()) ?? parentId)
: parentId;
if (fromPropId) {
graph.addRelationship({
id: generateId('ACCESSES', `${callerId}->read->${move.from}:L${move.line}`),
type: 'ACCESSES',
sourceId: callerId,
targetId: fromPropId,
confidence: 0.9,
reason: move.corresponding ? 'cobol-move-corresponding-read' : 'cobol-move-read',
});
}
if (toPropId) {
graph.addRelationship({
id: generateId('ACCESSES', `${callerId}->write->${move.to}:L${move.line}`),
type: 'ACCESSES',
sourceId: callerId,
targetId: toPropId,
confidence: 0.9,
reason: move.corresponding ? 'cobol-move-corresponding-write' : 'cobol-move-write',
});
}
}
// ── File declarations -> Record nodes ──────────────────────────
for (const fd of extracted.fileDeclarations) {
const fdId = generateId('Record', `${filePath}:${fd.selectName}`);
graph.addNode({
id: fdId,
label: 'Record',
properties: {
name: fd.selectName,
filePath,
startLine: fd.line,
endLine: fd.line,
language: 'cobol' as any,
description: `assign:${fd.assignTo}${fd.organization ? ` org:${fd.organization}` : ''}${fd.access ? ` access:${fd.access}` : ''}`,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${parentId}->${fdId}`),
type: 'CONTAINS',
sourceId: parentId,
targetId: fdId,
confidence: 1.0,
reason: 'cobol-file-declaration',
});
}
}
// ---------------------------------------------------------------------------
// Helpers
// ---------------------------------------------------------------------------
/** Find the section that contains a given line number. */
function findContainingSection(
line: number,
sections: Array<{ name: string; line: number }>,
sectionNodeIds: Map<string, string>,
): string | undefined {
// Sections are in order; find the last section whose start line <= the target line
let best: string | undefined;
for (const sec of sections) {
if (sec.line <= line) {
best = sectionNodeIds.get(sec.name.toUpperCase());
} else {
break;
}
}
return best;
}
@@ -0,0 +1,446 @@
/**
* COBOL COPY statement expansion engine.
*
* Expands COPY statements by inlining copybook content, applying REPLACING
* transformations (LEADING, TRAILING, EXACT), and handling nested copies
* with cycle detection.
*
* This is a preprocessing step that runs BEFORE extractCobolSymbolsWithRegex.
* The caller should run preprocessCobolSource first to clean patch markers.
*
* Supported syntax:
* COPY CPSESP.
* COPY "WORKGRID.CPY".
* COPY CPSESP REPLACING LEADING "ESP-" BY "LK-ESP-"
* LEADING "KPSESPL" BY "LK-KPSESPL".
* COPY ANAZI REPLACING "ANAZI-KEY" BY "LK-KEY".
*/
// ---------------------------------------------------------------------------
// Public interfaces
// ---------------------------------------------------------------------------
export interface CopyReplacing {
type: 'LEADING' | 'TRAILING' | 'EXACT';
from: string;
to: string;
}
export interface CopyResolution {
copyTarget: string;
resolvedPath: string | null;
line: number;
replacing: CopyReplacing[];
}
export interface CopyExpansionResult {
expandedContent: string;
copyResolutions: CopyResolution[];
expansionDepth: number;
}
// ---------------------------------------------------------------------------
// Constants
// ---------------------------------------------------------------------------
export const DEFAULT_MAX_DEPTH = 10;
/** COBOL identifier pattern: starts with letter, contains letters, digits, hyphens. */
const RE_COBOL_IDENTIFIER = /\b([A-Z][A-Z0-9-]*)\b/gi;
// ---------------------------------------------------------------------------
// Private helpers
// ---------------------------------------------------------------------------
/**
* Strip inline comments (Italian-style `|` comments).
* Only strips if `|` appears in the code area (col 7+).
*/
function stripInlineComment(line: string): string {
const idx = line.indexOf('|');
return idx >= 0 ? line.substring(0, idx) : line;
}
/**
* Check if a line is a COBOL comment (indicator in col 7 is `*` or `/`).
*/
function isCommentLine(line: string): boolean {
return line.length >= 7 && (line[6] === '*' || line[6] === '/');
}
/**
* Check if a line is a continuation line (indicator in col 7 is `-`).
*/
function isContinuationLine(line: string): boolean {
return line.length >= 7 && line[6] === '-';
}
/**
* Merge continuation lines into their predecessors.
* Returns an array of logical lines with their original starting line numbers.
*/
function mergeLogicalLines(
rawLines: string[],
): Array<{ text: string; lineNum: number }> {
const logical: Array<{ text: string; lineNum: number }> = [];
for (let i = 0; i < rawLines.length; i++) {
const raw = rawLines[i];
// Skip comment lines
if (isCommentLine(raw)) {
logical.push({ text: '', lineNum: i });
continue;
}
// Continuation: merge into previous logical line
if (isContinuationLine(raw)) {
if (logical.length > 0) {
const prev = logical[logical.length - 1];
const continuation = raw.length > 7 ? raw.substring(7).trimStart() : '';
prev.text += continuation;
}
// Push empty placeholder to preserve line count
logical.push({ text: '', lineNum: i });
continue;
}
// Normal line: strip inline comments
const cleaned = stripInlineComment(raw);
logical.push({ text: cleaned, lineNum: i });
}
return logical;
}
// ---------------------------------------------------------------------------
// COPY statement parsing
// ---------------------------------------------------------------------------
interface ParsedCopyStatement {
startLine: number;
endLine: number;
target: string;
replacing: CopyReplacing[];
}
/**
* Parse REPLACING clause text into structured replacements.
*
* Input examples:
* LEADING "ESP-" BY "LK-ESP-" LEADING "KPSESPL" BY "LK-KPSESPL"
* "ANAZI-KEY" BY "LK-KEY"
* TRAILING "-IN" BY "-OUT"
*/
function parseReplacingClause(text: string): CopyReplacing[] {
const replacings: CopyReplacing[] = [];
if (!text || text.trim().length === 0) return replacings;
// Tokenize: split on whitespace, preserving quoted strings
const tokens: string[] = [];
const tokenRe = /"([^"]*)"|(\S+)/g;
let tm: RegExpExecArray | null;
while ((tm = tokenRe.exec(text)) !== null) {
// Store the matched content; for quoted strings, keep the inner value
// but mark them so we can distinguish. We'll store all as plain strings
// and track which were quoted separately.
tokens.push(tm[1] !== undefined ? tm[1] : tm[2]);
}
// Parse token stream: [LEADING|TRAILING]? <from> BY <to>
let i = 0;
while (i < tokens.length) {
let type: CopyReplacing['type'] = 'EXACT';
const upper = tokens[i].toUpperCase();
// Check for type modifier
if (upper === 'LEADING') {
type = 'LEADING';
i++;
} else if (upper === 'TRAILING') {
type = 'TRAILING';
i++;
}
if (i >= tokens.length) break;
const from = tokens[i];
i++;
// Expect BY keyword
if (i >= tokens.length) break;
if (tokens[i].toUpperCase() !== 'BY') {
// Malformed — skip this token and try to resync
continue;
}
i++; // skip BY
if (i >= tokens.length) break;
const to = tokens[i];
i++;
replacings.push({ type, from, to });
}
return replacings;
}
/**
* Scan logical lines for COPY statements.
* COPY statements can span multiple lines and terminate with a period.
*/
function parseCopyStatements(
logicalLines: Array<{ text: string; lineNum: number }>,
): ParsedCopyStatement[] {
const results: ParsedCopyStatement[] = [];
let accumulator: string | null = null;
let startLine = 0;
let endLine = 0;
for (let i = 0; i < logicalLines.length; i++) {
const { text, lineNum } = logicalLines[i];
if (text.length === 0) continue;
// Check for COPY keyword start (not inside a string context)
const copyStart = text.match(/\bCOPY\b/i);
if (accumulator === null) {
if (!copyStart) continue;
// Start accumulating from the COPY keyword onwards
const copyIdx = copyStart.index!;
accumulator = text.substring(copyIdx);
startLine = lineNum;
endLine = lineNum;
} else {
// Continue accumulating
accumulator += ' ' + text.trim();
endLine = lineNum;
}
// Check if statement terminates (period at end of accumulated text)
if (accumulator !== null && /\.\s*$/.test(accumulator)) {
const parsed = parseSingleCopyStatement(accumulator, startLine, endLine);
if (parsed) {
results.push(parsed);
}
accumulator = null;
}
}
// If there's an unterminated COPY (missing period), try to parse what we have
if (accumulator !== null) {
const parsed = parseSingleCopyStatement(accumulator, startLine, endLine);
if (parsed) {
results.push(parsed);
}
}
return results;
}
/**
* Parse a single complete COPY statement string.
*
* Formats:
* COPY target.
* COPY "target".
* COPY target REPLACING ... .
*/
function parseSingleCopyStatement(
stmt: string,
startLine: number,
endLine: number,
): ParsedCopyStatement | null {
// Strip terminating period
const text = stmt.replace(/\.\s*$/, '').trim();
// Extract target: COPY <target> or COPY "<target>" or COPY '<target>'
const targetMatch = text.match(/^COPY\s+(?:"([^"]+)"|'([^']+)'|([A-Z][A-Z0-9-]*))/i);
if (!targetMatch) return null;
const target = targetMatch[1] || targetMatch[2] || targetMatch[3];
// Extract REPLACING clause if present
let replacing: CopyReplacing[] = [];
const replacingIdx = text.search(/\bREPLACING\b/i);
if (replacingIdx >= 0) {
const replacingText = text.substring(replacingIdx + 'REPLACING'.length);
replacing = parseReplacingClause(replacingText);
}
return { startLine, endLine, target, replacing };
}
// ---------------------------------------------------------------------------
// REPLACING application
// ---------------------------------------------------------------------------
/**
* Apply REPLACING transformations to copybook content.
*
* LEADING: replace prefix in COBOL identifiers.
* TRAILING: replace suffix in COBOL identifiers.
* EXACT: replace exact token matches.
*/
function applyReplacing(content: string, replacings: CopyReplacing[]): string {
if (replacings.length === 0) return content;
return content.replace(RE_COBOL_IDENTIFIER, (match) => {
for (const r of replacings) {
const upper = match.toUpperCase();
const from = r.from.toUpperCase();
const to = r.to.toUpperCase();
switch (r.type) {
case 'LEADING':
if (upper.startsWith(from)) {
return to + match.substring(from.length);
}
break;
case 'TRAILING':
if (upper.endsWith(from)) {
return match.substring(0, match.length - from.length) + to;
}
break;
case 'EXACT':
if (upper === from) {
return to;
}
break;
}
}
return match;
});
}
// ---------------------------------------------------------------------------
// Main expansion engine
// ---------------------------------------------------------------------------
/**
* Expand COBOL COPY statements by inlining copybook content.
*
* @param content - Source COBOL content (after preprocessCobolSource)
* @param filePath - Path of the source file (for diagnostics)
* @param resolveFile - Maps a COPY target name to a filesystem path, or null if not found
* @param readFile - Reads file content by path, or null if unreadable
* @param maxDepth - Maximum nesting depth for recursive expansion (default: 10)
* @returns Expanded content, resolution metadata, and maximum depth reached
*/
export function expandCopies(
content: string,
filePath: string,
resolveFile: (name: string) => string | null,
readFile: (path: string) => string | null,
maxDepth: number = DEFAULT_MAX_DEPTH,
/** Optional shared set to deduplicate circular-COPY warnings across multiple calls. */
warnedCircular: Set<string> = new Set<string>(),
): CopyExpansionResult {
const allResolutions: CopyResolution[] = [];
let maxDepthReached = 0;
const expanded = expandRecursive(content, filePath, 0, new Set<string>());
return {
expandedContent: expanded,
copyResolutions: allResolutions,
expansionDepth: maxDepthReached,
};
/**
* Recursively expand COPY statements in content.
*
* @param src - Source content to expand
* @param srcPath - Path of the file being expanded (for cycle detection logging)
* @param depth - Current recursion depth
* @param visited - Set of already-visited copybook paths (cycle detection)
*/
function expandRecursive(
src: string,
srcPath: string,
depth: number,
visited: Set<string>,
): string {
if (depth > maxDepthReached) {
maxDepthReached = depth;
}
const rawLines = src.split('\n');
const logicalLines = mergeLogicalLines(rawLines);
const copyStatements = parseCopyStatements(logicalLines);
// No COPY statements — return as-is
if (copyStatements.length === 0) return src;
// Process COPY statements in reverse order so line numbers stay valid
// as we splice content
const outputLines = [...rawLines];
for (let ci = copyStatements.length - 1; ci >= 0; ci--) {
const cs = copyStatements[ci];
// Resolve the copybook path
const resolvedPath = resolveFile(cs.target);
// Record resolution metadata
allResolutions.push({
copyTarget: cs.target,
resolvedPath,
line: cs.startLine,
replacing: cs.replacing,
});
// Cannot resolve — keep original lines
if (resolvedPath === null) {
continue;
}
// Cycle detection
if (visited.has(resolvedPath)) {
if (!warnedCircular.has(resolvedPath)) {
warnedCircular.add(resolvedPath);
console.warn(
`[cobol-copy-expander] Circular COPY detected: ${cs.target} (${resolvedPath}) ` +
`includes itself. Skipping expansion.`,
);
}
continue;
}
// Max depth exceeded — keep unexpanded
if (depth >= maxDepth) {
console.warn(
`[cobol-copy-expander] Max expansion depth (${maxDepth}) reached for ` +
`COPY ${cs.target} in ${srcPath}. Skipping expansion.`,
);
continue;
}
// Read the copybook content
const copybookContent = readFile(resolvedPath);
if (copybookContent === null) {
continue;
}
// Apply REPLACING transformations
const replaced = applyReplacing(copybookContent, cs.replacing);
// Recurse into the copybook for nested COPYs
const nestedVisited = new Set(visited);
nestedVisited.add(resolvedPath);
const expandedCopybook = expandRecursive(
replaced,
resolvedPath,
depth + 1,
nestedVisited,
);
// Splice: replace the COPY statement lines with expanded content
const expansionLines = expandedCopybook.split('\n');
const removeCount = cs.endLine - cs.startLine + 1;
outputLines.splice(cs.startLine, removeCount, ...expansionLines);
}
return outputLines.join('\n');
}
}
@@ -0,0 +1,911 @@
/**
* COBOL source pre-processing and regex-based symbol extraction.
*
* DESIGN DECISION — Why regex instead of a full parser (ANTLR4, tree-sitter):
*
* 1. Performance: Regex processes ~1ms/file vs 50-200ms/file for ANTLR4/tree-sitter.
* On EPAGHE (14k COBOL files), this is ~14 seconds vs 12-47 minutes.
*
* 2. Reliability: tree-sitter-cobol@0.0.1's external scanner hangs indefinitely
* on ~5% of production files (no timeout possible). ANTLR4's proleap-cobol-parser
* is a Java project — using it from Node.js requires Java subprocesses or
* extracting .g4 grammars and generating JS/TS targets (significant effort).
*
* 3. Dialect compatibility: GnuCOBOL with Italian comments, patch markers in
* cols 1-6 (mzADD, estero, etc.), and vendor extensions. Formal grammars
* target COBOL-85 and would need dialect modifications.
*
* 4. Industry precedent: ctags, GitHub code navigation, and Sourcegraph all use
* regex-based extraction for code indexing. Full parsing is only needed for
* compilation or semantic analysis, not symbol extraction.
*
* 5. Determinism: Every regex pattern is tested with canonical COBOL input
* (see test/unit/cobol-preprocessor.test.ts). Same input always produces
* same output — no grammar ambiguity or parser state issues.
*
* This module provides:
* 1. preprocessCobolSource() — cleans patch markers (kept for potential future use)
* 2. extractCobolSymbolsWithRegex() — single-pass state machine COBOL extraction
*/
// ---------------------------------------------------------------------------
// Public interfaces
// ---------------------------------------------------------------------------
export interface CobolRegexResults {
programName: string | null;
paragraphs: Array<{ name: string; line: number }>;
sections: Array<{ name: string; line: number }>;
performs: Array<{ caller: string | null; target: string; thruTarget?: string; line: number }>;
calls: Array<{ target: string; line: number }>;
copies: Array<{ target: string; line: number }>;
dataItems: Array<{
name: string;
level: number;
line: number;
pic?: string;
usage?: string;
occurs?: number;
redefines?: string;
values?: string[];
section: 'working-storage' | 'linkage' | 'file' | 'local-storage' | 'unknown';
}>;
fileDeclarations: Array<{
selectName: string;
assignTo: string;
organization?: string;
access?: string;
recordKey?: string;
fileStatus?: string;
line: number;
}>;
fdEntries: Array<{
fdName: string;
recordName?: string;
line: number;
}>;
programMetadata: {
author?: string;
dateWritten?: string;
};
// Phase 2: EXEC blocks
execSqlBlocks: Array<{
line: number;
tables: string[];
cursors: string[];
hostVariables: string[];
operation: 'SELECT' | 'INSERT' | 'UPDATE' | 'DELETE' | 'DECLARE' | 'OPEN' | 'CLOSE' | 'FETCH' | 'OTHER';
}>;
execCicsBlocks: Array<{
line: number;
command: string;
mapName?: string;
programName?: string;
transId?: string;
}>;
// Phase 3: Linkage + Data Flow
procedureUsing: string[];
entryPoints: Array<{
name: string;
parameters: string[];
line: number;
}>;
moves: Array<{
from: string;
to: string;
line: number;
caller: string | null;
corresponding: boolean;
}>;
}
// ---------------------------------------------------------------------------
// Preserved exactly: preprocessCobolSource
// ---------------------------------------------------------------------------
/**
* Normalize COBOL source for regex-based extraction.
*
* The COBOL fixed-format sequence number area (columns 1-6) is semantically
* irrelevant to parsing — compilers and tools always ignore it. This
* function replaces ANY non-space content in columns 1-6 with spaces so that
* position-sensitive regexes (paragraph/section detection, data-item anchors,
* etc.) work identically whether the file carries:
* • numeric sequence numbers (000100 … 999999)
* • alphabetic patch markers (mzADD, estero, #patch, …)
* • the COBOL default of all spaces
*
* Preserves exact line count for position mapping.
*/
export function preprocessCobolSource(content: string): string {
const lines = content.split('\n');
for (let i = 0; i < lines.length; i++) {
const line = lines[i];
if (line.length < 7) continue;
const seq = line.substring(0, 6);
// Replace any non-space character in the sequence area with spaces.
// This covers numeric sequence numbers (000100), alphabetic patch markers
// (mzADD, estero), '#'-prefixed markers, and mixed sequences.
if (/\S/.test(seq)) {
lines[i] = ' ' + line.substring(6);
}
}
return lines.join('\n');
}
// ---------------------------------------------------------------------------
// Preserved exactly: EXCLUDED_PARA_NAMES
// ---------------------------------------------------------------------------
const EXCLUDED_PARA_NAMES = new Set([
'DECLARATIVES', 'END', 'PROCEDURE', 'IDENTIFICATION',
'ENVIRONMENT', 'DATA', 'WORKING-STORAGE', 'LINKAGE',
'FILE', 'LOCAL-STORAGE', 'COMMUNICATION', 'REPORT',
'SCREEN', 'INPUT-OUTPUT', 'CONFIGURATION',
]);
// ---------------------------------------------------------------------------
// State machine types
// ---------------------------------------------------------------------------
type Division = 'identification' | 'environment' | 'data' | 'procedure' | null;
type DataSection = 'working-storage' | 'linkage' | 'file' | 'local-storage' | 'unknown';
type EnvironmentSection = 'input-output' | 'configuration' | null;
// ---------------------------------------------------------------------------
// Regex constants (compiled once, reused across calls)
// ---------------------------------------------------------------------------
const RE_DIVISION = /\b(IDENTIFICATION|ENVIRONMENT|DATA|PROCEDURE)\s+DIVISION\b/i;
const RE_SECTION = /\b(WORKING-STORAGE|LINKAGE|FILE|LOCAL-STORAGE|INPUT-OUTPUT|CONFIGURATION)\s+SECTION\b/i;
// IDENTIFICATION DIVISION
const RE_PROGRAM_ID = /\bPROGRAM-ID\.\s*([A-Z][A-Z0-9-]*)/i;
const RE_AUTHOR = /^\s+AUTHOR\.\s*(.+)/i;
const RE_DATE_WRITTEN = /^\s+DATE-WRITTEN\.\s*(.+)/i;
// ENVIRONMENT DIVISION — SELECT
const RE_SELECT_START = /\bSELECT\s+([A-Z][A-Z0-9-]+)/i;
// DATA DIVISION
const RE_FD = /^\s+FD\s+([A-Z][A-Z0-9-]+)/i;
const RE_DATA_ITEM = /^\s+(\d{1,2})\s+([A-Z][A-Z0-9-]+)\s*(.*)/i;
const RE_ANONYMOUS_REDEFINES = /^\s+(\d{1,2})\s+REDEFINES\s+([A-Z][A-Z0-9-]+)/i;
const RE_88_LEVEL = /^\s+88\s+([A-Z][A-Z0-9-]+)\s+VALUES?\s+(?:ARE\s+)?(.+)/i;
// PROCEDURE DIVISION
const RE_PROC_SECTION = /^ ([A-Z][A-Z0-9-]+)\s+SECTION\.\s*$/;
const RE_PROC_PARAGRAPH = /^ ([A-Z][A-Z0-9-]+)\.\s*$/;
const RE_PERFORM = /\bPERFORM\s+([A-Z][A-Z0-9-]+)(?:\s+THRU\s+([A-Z][A-Z0-9-]+))?/i;
// ALL DIVISIONS
// Both double-quoted ("PROG") and single-quoted ('PROG') targets are valid COBOL.
// Use separate alternation groups so quotes must match (prevents "PROG' false-matches).
const RE_CALL = /\bCALL\s+(?:"([^"]+)"|'([^']+)')/i;
const RE_COPY_UNQUOTED = /\bCOPY\s+([A-Z][A-Z0-9-]+)(?:\s|\.)/i;
const RE_COPY_QUOTED = /\bCOPY\s+(?:"([^"]+)"|'([^']+)')(?:\s|\.)/i;
// EXEC blocks
const RE_EXEC_SQL_START = /\bEXEC\s+SQL\b/i;
const RE_EXEC_CICS_START = /\bEXEC\s+CICS\b/i;
const RE_END_EXEC = /\bEND-EXEC\b/i;
// PROCEDURE DIVISION USING
const RE_PROC_USING = /\bPROCEDURE\s+DIVISION\s+USING\s+([\s\S]*?)(?:\.|$)/i;
// ENTRY point
const RE_ENTRY = /\bENTRY\s+"([^"]+)"(?:\s+USING\s+([\s\S]*?))?(?:\.|$)/i;
// MOVE statement
const RE_MOVE = /\bMOVE\s+(CORRESPONDING\s+)?([A-Z][A-Z0-9-]+)\s+TO\s+([A-Z][A-Z0-9-]+)/i;
const MOVE_SKIP = new Set([
'SPACES', 'ZEROS', 'ZEROES', 'LOW-VALUES', 'LOW-VALUE',
'HIGH-VALUES', 'HIGH-VALUE', 'QUOTES', 'QUOTE', 'ALL',
]);
// PERFORM: keywords that may follow PERFORM but are NOT paragraph/section names.
// Inline PERFORM loops (UNTIL, VARYING) and inline test clauses (WITH TEST,
// FOREVER) must not be stored as perform-target false positives.
const PERFORM_KEYWORD_SKIP = new Set([
'UNTIL', 'VARYING', 'WITH', 'TEST', 'FOREVER',
]);
// ---------------------------------------------------------------------------
// Private helper: strip Italian inline comments (| and everything after)
// ---------------------------------------------------------------------------
function stripInlineComment(line: string): string {
const idx = line.indexOf('|');
return idx >= 0 ? line.substring(0, idx) : line;
}
// ---------------------------------------------------------------------------
// Private helper: parse data item trailing clauses (PIC, USAGE, etc.)
// ---------------------------------------------------------------------------
function parseDataItemClauses(rest: string): {
pic?: string;
usage?: string;
redefines?: string;
occurs?: number;
} {
const result: { pic?: string; usage?: string; redefines?: string; occurs?: number } = {};
// Strip trailing period for easier parsing
const text = rest.replace(/\.\s*$/, '');
// PIC / PICTURE [IS] <picture-string>
const picMatch = text.match(/\bPIC(?:TURE)?\s+(?:IS\s+)?(\S+)/i);
if (picMatch) {
result.pic = picMatch[1];
}
// USAGE [IS] <usage-type> — including non-standard COMP-6, COMP-X etc.
const usageMatch = text.match(/\bUSAGE\s+(?:IS\s+)?(COMP(?:UTATIONAL)?(?:-[0-9X])?|BINARY|PACKED-DECIMAL|DISPLAY|INDEX|POINTER|NATIONAL)\b/i);
if (usageMatch) {
result.usage = usageMatch[1].toUpperCase();
} else {
// Standalone COMP variants without USAGE keyword
const compMatch = text.match(/\b(COMP(?:UTATIONAL)?(?:-[0-9X])?|BINARY|PACKED-DECIMAL)\b/i);
if (compMatch) {
result.usage = compMatch[1].toUpperCase();
}
}
// REDEFINES <name>
const redefMatch = text.match(/\bREDEFINES\s+([A-Z][A-Z0-9-]+)/i);
if (redefMatch) {
result.redefines = redefMatch[1];
}
// OCCURS <n> [TIMES]
const occursMatch = text.match(/\bOCCURS\s+(\d+)/i);
if (occursMatch) {
result.occurs = parseInt(occursMatch[1], 10);
}
return result;
}
// ---------------------------------------------------------------------------
// Private helper: parse 88-level condition values
// ---------------------------------------------------------------------------
function parseConditionValues(valuesStr: string): string[] {
// Strip trailing period
const text = valuesStr.replace(/\.\s*$/, '').trim();
const values: string[] = [];
// Match quoted strings: "O" "Y" "I"
const quotedRe = /"([^"]*)"/g;
let qm: RegExpExecArray | null;
let hasQuoted = false;
while ((qm = quotedRe.exec(text)) !== null) {
values.push(qm[1]);
hasQuoted = true;
}
if (hasQuoted) return values;
// No quotes — split on whitespace, filtering out THRU/THROUGH keywords
// Handle: 11 12 16 17 21 or 1 THRU 5
const tokens = text.split(/\s+/);
for (const token of tokens) {
const upper = token.toUpperCase();
if (upper === 'THRU' || upper === 'THROUGH') {
// Keep THRU ranges as combined value: prev THRU next is already captured
// by having both sides in the array
continue;
}
if (token.length > 0) {
values.push(token);
}
}
return values;
}
// ---------------------------------------------------------------------------
// Private helper: parse accumulated multi-line SELECT statement
// ---------------------------------------------------------------------------
interface FileDeclaration {
selectName: string;
assignTo: string;
organization?: string;
access?: string;
recordKey?: string;
fileStatus?: string;
line: number;
}
function parseSelectStatement(stmt: string, startLine: number): FileDeclaration | null {
// Normalize whitespace
const text = stmt.replace(/\s+/g, ' ').trim();
const nameMatch = text.match(/^SELECT\s+([A-Z][A-Z0-9-]+)/i);
if (!nameMatch) return null;
const result: FileDeclaration = {
selectName: nameMatch[1],
assignTo: '',
line: startLine,
};
const assignMatch = text.match(/\bASSIGN\s+(?:TO\s+)?("([^"]+)"|([A-Z][A-Z0-9-]*))/i);
if (assignMatch) {
result.assignTo = assignMatch[2] || assignMatch[3] || '';
}
const orgMatch = text.match(/\bORGANIZATION\s+(?:IS\s+)?(SEQUENTIAL|INDEXED|RELATIVE|LINE\s+SEQUENTIAL)/i);
if (orgMatch) {
result.organization = orgMatch[1].toUpperCase();
}
const accessMatch = text.match(/\bACCESS\s+(?:MODE\s+)?(?:IS\s+)?(SEQUENTIAL|RANDOM|DYNAMIC)/i);
if (accessMatch) {
result.access = accessMatch[1].toUpperCase();
}
const keyMatch = text.match(/\bRECORD\s+KEY\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)/i);
if (keyMatch) {
result.recordKey = keyMatch[1];
}
// FILE STATUS IS / STATUS IS
const statusMatch = text.match(/\b(?:FILE\s+)?STATUS\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)/i);
if (statusMatch) {
result.fileStatus = statusMatch[1];
}
return result;
}
// ---------------------------------------------------------------------------
// Private helper: parse EXEC SQL block
// ---------------------------------------------------------------------------
type SqlOperation = 'SELECT' | 'INSERT' | 'UPDATE' | 'DELETE' | 'DECLARE' | 'OPEN' | 'CLOSE' | 'FETCH' | 'OTHER';
function parseExecSqlBlock(block: string, line: number): CobolRegexResults['execSqlBlocks'][number] {
// Strip EXEC SQL ... END-EXEC wrapper
const body = block
.replace(/\bEXEC\s+SQL\b/i, '')
.replace(/\bEND-EXEC\b/i, '')
.replace(/\s+/g, ' ')
.trim();
// Determine operation from first SQL keyword
const firstWord = body.split(/\s+/)[0]?.toUpperCase() || '';
const OP_MAP: Record<string, SqlOperation> = {
SELECT: 'SELECT', INSERT: 'INSERT', UPDATE: 'UPDATE', DELETE: 'DELETE',
DECLARE: 'DECLARE', OPEN: 'OPEN', CLOSE: 'CLOSE', FETCH: 'FETCH',
};
const operation: SqlOperation = OP_MAP[firstWord] || 'OTHER';
// Extract table names from FROM, INTO (INSERT), UPDATE, DELETE FROM, JOIN
const tables: string[] = [];
const tablePatterns = [
/\bFROM\s+([A-Z][A-Z0-9_]+)/gi,
/\bINTO\s+([A-Z][A-Z0-9_]+)/gi,
/\bUPDATE\s+([A-Z][A-Z0-9_]+)/gi,
/\bJOIN\s+([A-Z][A-Z0-9_]+)/gi,
];
for (const re of tablePatterns) {
let m: RegExpExecArray | null;
while ((m = re.exec(body)) !== null) {
const name = m[1].toUpperCase();
// Skip host variables and SQL keywords
if (!name.startsWith(':') && !tables.includes(name)) {
tables.push(name);
}
}
}
// Extract cursor names from DECLARE ... CURSOR
const cursors: string[] = [];
const cursorRe = /\bDECLARE\s+([A-Z][A-Z0-9_-]+)\s+CURSOR\b/gi;
let cm: RegExpExecArray | null;
while ((cm = cursorRe.exec(body)) !== null) {
cursors.push(cm[1]);
}
// Extract host variables: :VARIABLE-NAME (strip the colon)
const hostVariables: string[] = [];
const hostRe = /:([A-Z][A-Z0-9-]+)/gi;
let hm: RegExpExecArray | null;
while ((hm = hostRe.exec(body)) !== null) {
const name = hm[1];
if (!hostVariables.includes(name)) {
hostVariables.push(name);
}
}
return { line, tables, cursors, hostVariables, operation };
}
// ---------------------------------------------------------------------------
// Private helper: parse EXEC CICS block
// ---------------------------------------------------------------------------
function parseExecCicsBlock(block: string, line: number): CobolRegexResults['execCicsBlocks'][number] {
// Strip EXEC CICS ... END-EXEC wrapper
const body = block
.replace(/\bEXEC\s+CICS\b/i, '')
.replace(/\bEND-EXEC\b/i, '')
.replace(/\s+/g, ' ')
.trim();
// Command: first keyword(s) — handle two-word commands like SEND MAP, RECEIVE MAP
const twoWordCommands = ['SEND MAP', 'RECEIVE MAP', 'SEND TEXT', 'SEND CONTROL', 'READ NEXT', 'READ PREV'];
let command = '';
const upperBody = body.toUpperCase();
for (const twoWord of twoWordCommands) {
if (upperBody.startsWith(twoWord)) {
command = twoWord;
break;
}
}
if (!command) {
command = body.split(/\s+/)[0]?.toUpperCase() || '';
}
const result: CobolRegexResults['execCicsBlocks'][number] = { line, command };
// MAP name: MAP('name') or MAP("name")
const mapMatch = body.match(/\bMAP\s*\(\s*['"]([^'"]+)['"]\s*\)/i);
if (mapMatch) result.mapName = mapMatch[1];
// PROGRAM name: PROGRAM('name') or PROGRAM("name")
const progMatch = body.match(/\bPROGRAM\s*\(\s*['"]([^'"]+)['"]\s*\)/i);
if (progMatch) result.programName = progMatch[1];
// TRANSID: TRANSID('name') or TRANSID("name")
const transMatch = body.match(/\bTRANSID\s*\(\s*['"]([^'"]+)['"]\s*\)/i);
if (transMatch) result.transId = transMatch[1];
return result;
}
// ---------------------------------------------------------------------------
// Main extraction: single-pass state machine
// ---------------------------------------------------------------------------
/**
* Extract COBOL symbols using a single-pass state machine.
* Extracts program name, paragraphs, sections, CALL, PERFORM, COPY,
* data items, file declarations, FD entries, and program metadata.
*/
export function extractCobolSymbolsWithRegex(
content: string,
_filePath: string,
): CobolRegexResults {
const rawLines = content.split('\n');
const result: CobolRegexResults = {
programName: null,
paragraphs: [],
sections: [],
performs: [],
calls: [],
copies: [],
dataItems: [],
fileDeclarations: [],
fdEntries: [],
programMetadata: {},
execSqlBlocks: [],
execCicsBlocks: [],
procedureUsing: [],
entryPoints: [],
moves: [],
};
// --- State ---
let currentDivision: Division = null;
let currentDataSection: DataSection = 'unknown';
let currentEnvSection: EnvironmentSection = null;
let currentParagraph: string | null = null;
// SELECT accumulator (multi-line)
let selectAccum: string | null = null;
let selectStartLine = 0;
// EXEC block accumulator (multi-line EXEC SQL / EXEC CICS)
let execAccum: { type: 'sql' | 'cics'; lines: string; startLine: number } | null = null;
// FD tracking: after seeing FD, the next 01-level data item is its record
let pendingFdName: string | null = null;
let pendingFdLine = 0;
// Continuation line buffer
let pendingLine: string | null = null;
let pendingLineNumber = 0;
// --- Process each raw line ---
for (let i = 0; i < rawLines.length; i++) {
const raw = rawLines[i];
// Skip lines too short to have indicator area
if (raw.length < 7) {
// If there's a pending continuation, flush it
if (pendingLine !== null) {
processLogicalLine(pendingLine, pendingLineNumber);
pendingLine = null;
}
continue;
}
const indicator = raw[6];
// Comment line: indicator is '*' or '/'
if (indicator === '*' || indicator === '/') {
continue;
}
// Continuation line: indicator is '-'
if (indicator === '-') {
if (pendingLine !== null) {
// Append continuation (area B content, trimmed leading spaces)
const continuation = raw.substring(7).trimStart();
pendingLine += continuation;
}
continue;
}
// Normal line — flush any pending continuation first
if (pendingLine !== null) {
processLogicalLine(pendingLine, pendingLineNumber);
pendingLine = null;
}
// Strip inline Italian comments, then use area A+B (from col 7 onwards,
// but keep full line for indentation-sensitive paragraph/section detection)
const cleaned = stripInlineComment(raw);
// Buffer as new pending logical line
pendingLine = cleaned;
pendingLineNumber = i;
}
// Flush final pending line
if (pendingLine !== null) {
processLogicalLine(pendingLine, pendingLineNumber);
}
// Flush any pending SELECT
flushSelect();
// If we saw an FD but never found its record, emit it without a record name
if (pendingFdName !== null) {
result.fdEntries.push({ fdName: pendingFdName, line: pendingFdLine });
pendingFdName = null;
}
return result;
// =========================================================================
// Inner function: process one logical line (after continuation merging)
// =========================================================================
function processLogicalLine(line: string, lineNum: number): void {
// --- EXEC block accumulation (spans any division) ---
if (execAccum !== null) {
execAccum.lines += ' ' + line;
if (RE_END_EXEC.test(line)) {
if (execAccum.type === 'sql') {
result.execSqlBlocks.push(parseExecSqlBlock(execAccum.lines, execAccum.startLine));
} else {
result.execCicsBlocks.push(parseExecCicsBlock(execAccum.lines, execAccum.startLine));
}
execAccum = null;
}
return; // While accumulating, skip normal processing
}
// Check for EXEC SQL / EXEC CICS start
if (RE_EXEC_SQL_START.test(line)) {
execAccum = { type: 'sql', lines: line, startLine: lineNum };
// If END-EXEC is on the same line, finalize immediately
if (RE_END_EXEC.test(line)) {
result.execSqlBlocks.push(parseExecSqlBlock(execAccum.lines, execAccum.startLine));
execAccum = null;
}
return;
}
if (RE_EXEC_CICS_START.test(line)) {
execAccum = { type: 'cics', lines: line, startLine: lineNum };
if (RE_END_EXEC.test(line)) {
result.execCicsBlocks.push(parseExecCicsBlock(execAccum.lines, execAccum.startLine));
execAccum = null;
}
return;
}
// --- Division transitions ---
const divMatch = line.match(RE_DIVISION);
if (divMatch) {
// Flush SELECT if transitioning out of environment
flushSelect();
const divName = divMatch[1].toUpperCase();
switch (divName) {
case 'IDENTIFICATION': currentDivision = 'identification'; break;
case 'ENVIRONMENT': currentDivision = 'environment'; currentEnvSection = null; break;
case 'DATA': currentDivision = 'data'; currentDataSection = 'unknown'; break;
case 'PROCEDURE': {
currentDivision = 'procedure';
currentParagraph = null;
const procUsingMatch = line.match(RE_PROC_USING);
if (procUsingMatch) {
result.procedureUsing = procUsingMatch[1].trim().split(/\s+/).filter(s => s.length > 0);
}
break;
}
}
return;
}
// --- Section transitions ---
const secMatch = line.match(RE_SECTION);
if (secMatch) {
flushSelect();
const secName = secMatch[1].toUpperCase();
switch (secName) {
case 'WORKING-STORAGE': currentDivision = 'data'; currentDataSection = 'working-storage'; break;
case 'LINKAGE': currentDivision = 'data'; currentDataSection = 'linkage'; break;
case 'FILE': currentDivision = 'data'; currentDataSection = 'file'; break;
case 'LOCAL-STORAGE': currentDivision = 'data'; currentDataSection = 'local-storage'; break;
case 'INPUT-OUTPUT': currentDivision = 'environment'; currentEnvSection = 'input-output'; break;
case 'CONFIGURATION': currentDivision = 'environment'; currentEnvSection = 'configuration'; break;
}
return;
}
// --- COPY (all divisions) ---
const copyQMatch = line.match(RE_COPY_QUOTED);
if (copyQMatch) {
result.copies.push({ target: copyQMatch[1] ?? copyQMatch[2], line: lineNum });
} else {
const copyUMatch = line.match(RE_COPY_UNQUOTED);
if (copyUMatch) {
result.copies.push({ target: copyUMatch[1], line: lineNum });
}
}
// --- CALL (all divisions, typically procedure) ---
const callMatch = line.match(RE_CALL);
if (callMatch) {
result.calls.push({ target: callMatch[1] ?? callMatch[2], line: lineNum });
}
// --- Division-specific extraction ---
switch (currentDivision) {
case 'identification':
extractIdentification(line, lineNum);
break;
case 'environment':
extractEnvironment(line, lineNum);
break;
case 'data':
extractData(line, lineNum);
break;
case 'procedure':
extractProcedure(line, lineNum);
break;
}
}
// =========================================================================
// IDENTIFICATION DIVISION extraction
// =========================================================================
function extractIdentification(line: string, _lineNum: number): void {
if (result.programName === null) {
const m = line.match(RE_PROGRAM_ID);
if (m) {
result.programName = m[1];
return;
}
}
const authorMatch = line.match(RE_AUTHOR);
if (authorMatch) {
result.programMetadata.author = authorMatch[1].replace(/\.\s*$/, '').trim();
return;
}
const dateMatch = line.match(RE_DATE_WRITTEN);
if (dateMatch) {
result.programMetadata.dateWritten = dateMatch[1].replace(/\.\s*$/, '').trim();
}
}
// =========================================================================
// ENVIRONMENT DIVISION extraction
// =========================================================================
function extractEnvironment(line: string, lineNum: number): void {
if (currentEnvSection !== 'input-output') return;
// Check for new SELECT statement
const selMatch = line.match(RE_SELECT_START);
if (selMatch) {
// Flush any previous SELECT
flushSelect();
selectAccum = line.trim();
selectStartLine = lineNum;
} else if (selectAccum !== null) {
// Accumulate continuation of current SELECT
selectAccum += ' ' + line.trim();
}
// Check if current SELECT is terminated (ends with period)
if (selectAccum !== null && /\.\s*$/.test(selectAccum)) {
flushSelect();
}
}
function flushSelect(): void {
if (selectAccum === null) return;
const decl = parseSelectStatement(selectAccum, selectStartLine);
if (decl) {
result.fileDeclarations.push(decl);
}
selectAccum = null;
}
// =========================================================================
// DATA DIVISION extraction
// =========================================================================
function extractData(line: string, lineNum: number): void {
// FD entry
const fdMatch = line.match(RE_FD);
if (fdMatch) {
// Flush any previous FD without a record
if (pendingFdName !== null) {
result.fdEntries.push({ fdName: pendingFdName, line: pendingFdLine });
}
pendingFdName = fdMatch[1];
pendingFdLine = lineNum;
return;
}
// 88-level condition names
const lv88Match = line.match(RE_88_LEVEL);
if (lv88Match) {
const name = lv88Match[1];
const values = parseConditionValues(lv88Match[2]);
result.dataItems.push({
name,
level: 88,
line: lineNum,
values,
section: currentDataSection,
});
return;
}
// Anonymous REDEFINES (no name, e.g. "01 REDEFINES WK-PERIVAL.")
const anonRedefMatch = line.match(RE_ANONYMOUS_REDEFINES);
if (anonRedefMatch) {
// Check it's truly anonymous: the second capture is not a valid data name
// followed by more clauses — it's the REDEFINES target directly after level
const level = parseInt(anonRedefMatch[1], 10);
// Only skip if this is genuinely "NN REDEFINES target" with no name between
// We detect this by checking the full data item regex does NOT match
// (because RE_DATA_ITEM expects a name before any clauses)
const dataMatch = line.match(RE_DATA_ITEM);
if (!dataMatch || dataMatch[2].toUpperCase() === 'REDEFINES') {
// Truly anonymous — skip, no node
return;
}
}
// Standard data items: level 01-49, 66, 77
const dataMatch = line.match(RE_DATA_ITEM);
if (dataMatch) {
const level = parseInt(dataMatch[1], 10);
const name = dataMatch[2];
const rest = dataMatch[3] || '';
// Skip FILLER
if (name.toUpperCase() === 'FILLER') return;
// Valid levels: 01-49, 66, 77
if ((level >= 1 && level <= 49) || level === 66 || level === 77) {
const clauses = parseDataItemClauses(rest);
const item: CobolRegexResults['dataItems'][number] = {
name,
level,
line: lineNum,
section: currentDataSection,
};
if (clauses.pic) item.pic = clauses.pic;
if (clauses.usage) item.usage = clauses.usage;
if (clauses.occurs !== undefined) item.occurs = clauses.occurs;
if (clauses.redefines) item.redefines = clauses.redefines;
result.dataItems.push(item);
// If there's a pending FD and this is a 01-level, it's the FD's record
if (pendingFdName !== null && level === 1) {
result.fdEntries.push({
fdName: pendingFdName,
recordName: name,
line: pendingFdLine,
});
pendingFdName = null;
}
}
}
}
// =========================================================================
// PROCEDURE DIVISION extraction
// =========================================================================
function extractProcedure(line: string, lineNum: number): void {
// Section header
const secMatch = line.match(RE_PROC_SECTION);
if (secMatch) {
const name = secMatch[1];
if (!EXCLUDED_PARA_NAMES.has(name) && !name.includes('DIVISION')) {
result.sections.push({ name, line: lineNum });
currentParagraph = name;
}
return;
}
// Paragraph header
const paraMatch = line.match(RE_PROC_PARAGRAPH);
if (paraMatch) {
const name = paraMatch[1];
if (!EXCLUDED_PARA_NAMES.has(name) && !name.includes('DIVISION') && !name.includes('SECTION')) {
result.paragraphs.push({ name, line: lineNum });
currentParagraph = name;
}
return;
}
// PERFORM
const perfMatch = line.match(RE_PERFORM);
if (perfMatch) {
const target = perfMatch[1];
// Skip COBOL inline-perform keywords that are not paragraph names
if (!PERFORM_KEYWORD_SKIP.has(target.toUpperCase())) {
result.performs.push({
caller: currentParagraph,
target,
thruTarget: perfMatch[2] || undefined,
line: lineNum,
});
}
}
// ENTRY point
const entryMatch = line.match(RE_ENTRY);
if (entryMatch) {
result.entryPoints.push({
name: entryMatch[1],
parameters: entryMatch[2] ? entryMatch[2].trim().split(/\s+/).filter(s => s.length > 0) : [],
line: lineNum,
});
}
// MOVE statement (skip literals and figurative constants)
const moveMatch = line.match(RE_MOVE);
if (moveMatch) {
const from = moveMatch[2].toUpperCase();
if (!MOVE_SKIP.has(from)) {
result.moves.push({
from: moveMatch[2],
to: moveMatch[3],
line: lineNum,
caller: currentParagraph,
corresponding: !!moveMatch[1],
});
}
}
}
}
@@ -0,0 +1,266 @@
/**
* JCL Parser — Regex single-pass extraction.
*
* Extracts JCL constructs from mainframe job streams:
* - JOB statements (job name, CLASS, MSGCLASS)
* - EXEC statements (step -> program or proc)
* - DD statements (dataset references, DISP)
* - PROC definitions (in-stream and catalogued)
* - INCLUDE MEMBER= directives
* - SET symbolic parameters
* - IF/ELSE/ENDIF conditional execution
* - JCLLIB ORDER= search paths
*
* Pattern follows cobol-preprocessor.ts — regex-only, no tree-sitter.
*/
export interface JclParseResults {
jobs: Array<{ name: string; line: number; class?: string; msgclass?: string }>;
steps: Array<{ name: string; jobName: string; program?: string; proc?: string; line: number }>;
ddStatements: Array<{ ddName: string; stepName: string; dataset?: string; disp?: string; line: number }>;
procs: Array<{ name: string; line: number; isInStream: boolean }>;
includes: Array<{ member: string; line: number }>;
sets: Array<{ variable: string; value: string; line: number }>;
jcllib: Array<{ order: string[]; line: number }>;
conditionals: Array<{ type: 'IF' | 'ELSE' | 'ENDIF'; condition?: string; line: number }>;
}
// ── JCL statement patterns ─────────────────────────────────────────────
// JCL continuation: line ends with a non-blank in col 72, next line starts with //
// We handle continuations by joining lines before matching.
/** Match //jobname JOB ... */
const JOB_RE = /^\/\/(\w{1,8})\s+JOB\s+(.*)/i;
/** Match //stepname EXEC PGM=program or //stepname EXEC procname */
const EXEC_RE = /^\/\/(\w{1,8})\s+EXEC\s+(.*)/i;
/** Match //ddname DD ... */
const DD_RE = /^\/\/(\w{1,8})\s+DD\s+(.*)/i;
/** Match // JCLLIB ORDER=(lib1,lib2,...) */
const JCLLIB_RE = /^\/\/\s+JCLLIB\s+ORDER=\(([^)]+)\)/i;
/** Match // IF condition THEN */
const IF_RE = /^\/\/\s+IF\s+(.+)\s+THEN/i;
/** Match // ELSE */
const ELSE_RE = /^\/\/\s+ELSE\b/i;
/** Match // ENDIF */
const ENDIF_RE = /^\/\/\s+ENDIF\b/i;
/** Match // INCLUDE MEMBER=name */
const INCLUDE_RE = /^\/\/\s+INCLUDE\s+MEMBER=(\w+)/i;
/** Match // SET var=value */
const SET_RE = /^\/\/\s+SET\s+(\w+)=(.+)/i;
/** Match // PROC or //name PROC */
const PROC_RE = /^\/\/(\w*)\s+PROC\b/i;
/** Match // PEND */
const PEND_RE = /^\/\/\s+PEND\b/i;
// ── Parameter extractors ───────────────────────────────────────────────
function extractParam(params: string, key: string): string | undefined {
// Match KEY=VALUE or KEY='VALUE' in JCL parameter string
const re = new RegExp(`${key}=(?:'([^']*)'|(\\S+?))(?:[,\\s]|$)`, 'i');
const m = params.match(re);
return m ? (m[1] ?? m[2]) : undefined;
}
function extractPgm(params: string): string | undefined {
return extractParam(params, 'PGM');
}
function extractProc(params: string): string | undefined {
// If no PGM= keyword, the first positional parameter is the proc name
if (/PGM=/i.test(params)) return undefined;
const cleaned = params.replace(/,.*/, '').trim();
// Proc name is the first token (no = sign)
if (cleaned && !cleaned.includes('=')) {
return cleaned.replace(/[,\s].*/s, '').toUpperCase();
}
return undefined;
}
function extractDsn(params: string): string | undefined {
return extractParam(params, 'DSN') ?? extractParam(params, 'DSNAME');
}
function extractDisp(params: string): string | undefined {
const m = params.match(/DISP=\(?\s*([^),\s]+)/i);
return m ? m[1] : undefined;
}
/**
* Parse a JCL file and extract all constructs.
*
* @param content - Raw JCL file content
* @param filePath - Path for diagnostics (not used in extraction)
* @returns Parsed JCL results
*/
export function parseJcl(content: string, filePath: string): JclParseResults {
const results: JclParseResults = {
jobs: [],
steps: [],
ddStatements: [],
procs: [],
includes: [],
sets: [],
jcllib: [],
conditionals: [],
};
const rawLines = content.split('\n');
// Join continuation lines: a line ending with non-blank in col 71 (0-indexed)
// followed by a line starting with // is a continuation.
const lines: Array<{ text: string; lineNum: number }> = [];
let i = 0;
while (i < rawLines.length) {
let line = rawLines[i];
const lineNum = i + 1;
// JCL continuation: if line is exactly 72+ chars and col 72 is non-blank
// and the next line starts with //, join them.
while (
i + 1 < rawLines.length &&
line.length >= 72 &&
line[71] !== ' ' &&
rawLines[i + 1].startsWith('//')
) {
i++;
// Continuation text starts after // and leading spaces
const contText = rawLines[i].substring(2).replace(/^\s+/, ' ');
// Remove the continuation marker (col 72+) from current line
line = line.substring(0, 71).trimEnd() + contText;
}
lines.push({ text: line, lineNum });
i++;
}
let currentJobName = '';
let currentStepName = '';
let inInStreamProc = false;
let inStreamProcName = '';
for (const { text, lineNum } of lines) {
// Skip JCL comments (starting with //* )
if (text.startsWith('//*')) continue;
// Skip non-JCL lines (don't start with //)
if (!text.startsWith('//')) continue;
// PROC definition (in-stream)
const procMatch = text.match(PROC_RE);
if (procMatch) {
const procName = procMatch[1] || inStreamProcName;
if (procName) {
results.procs.push({ name: procName.toUpperCase(), line: lineNum, isInStream: true });
}
inInStreamProc = true;
inStreamProcName = procName?.toUpperCase() || '';
continue;
}
// PEND (end of in-stream proc)
if (PEND_RE.test(text)) {
inInStreamProc = false;
inStreamProcName = '';
continue;
}
// JCLLIB ORDER=
const jcllibMatch = text.match(JCLLIB_RE);
if (jcllibMatch) {
const libs = jcllibMatch[1].split(',').map(s => s.trim().replace(/'/g, ''));
results.jcllib.push({ order: libs, line: lineNum });
continue;
}
// IF/ELSE/ENDIF
const ifMatch = text.match(IF_RE);
if (ifMatch) {
results.conditionals.push({ type: 'IF', condition: ifMatch[1].trim(), line: lineNum });
continue;
}
if (ELSE_RE.test(text)) {
results.conditionals.push({ type: 'ELSE', line: lineNum });
continue;
}
if (ENDIF_RE.test(text)) {
results.conditionals.push({ type: 'ENDIF', line: lineNum });
continue;
}
// INCLUDE MEMBER=
const includeMatch = text.match(INCLUDE_RE);
if (includeMatch) {
results.includes.push({ member: includeMatch[1].toUpperCase(), line: lineNum });
continue;
}
// SET var=value
const setMatch = text.match(SET_RE);
if (setMatch) {
results.sets.push({
variable: setMatch[1].toUpperCase(),
value: setMatch[2].trim().replace(/,\s*$/, ''),
line: lineNum,
});
continue;
}
// JOB statement
const jobMatch = text.match(JOB_RE);
if (jobMatch) {
currentJobName = jobMatch[1].toUpperCase();
const params = jobMatch[2];
results.jobs.push({
name: currentJobName,
line: lineNum,
class: extractParam(params, 'CLASS'),
msgclass: extractParam(params, 'MSGCLASS'),
});
continue;
}
// EXEC statement
const execMatch = text.match(EXEC_RE);
if (execMatch) {
currentStepName = execMatch[1].toUpperCase();
const params = execMatch[2];
const pgm = extractPgm(params);
const proc = pgm ? undefined : extractProc(params);
results.steps.push({
name: currentStepName,
jobName: currentJobName,
program: pgm?.toUpperCase(),
proc: proc?.toUpperCase(),
line: lineNum,
});
continue;
}
// DD statement
const ddMatch = text.match(DD_RE);
if (ddMatch) {
const ddName = ddMatch[1].toUpperCase();
const params = ddMatch[2];
results.ddStatements.push({
ddName,
stepName: currentStepName,
dataset: extractDsn(params)?.toUpperCase(),
disp: extractDisp(params)?.toUpperCase(),
line: lineNum,
});
continue;
}
}
return results;
}
@@ -0,0 +1,264 @@
/**
* JCL Processor — Converts JCL parse results into graph nodes and edges.
*
* Maps JCL entities to existing graph types (no new tables):
* - Job -> CodeElement (description: "jcl-job class:A msgclass:X")
* - Step -> CodeElement (description: "jcl-step pgm:PROGRAMNAME")
* - Dataset -> CodeElement (description: "jcl-dataset disp:SHR")
* - PROC -> Module
*
* Edges:
* - Job CONTAINS Step
* - Step CALLS Module (when PGM= matches an indexed program)
* - Step references Dataset (CALLS edge with reason "jcl-dd")
* - Job/Step IMPORTS PROC
*
* Pattern follows detectCrossProgamContracts() in pipeline.ts.
*/
import { parseJcl, type JclParseResults } from './jcl-parser.js';
import type { KnowledgeGraph } from '../../graph/types.js';
import { generateId } from '../../../lib/utils.js';
export interface JclProcessResult {
jobCount: number;
stepCount: number;
datasetCount: number;
programLinks: number;
}
/**
* Process JCL files and integrate into the knowledge graph.
*
* @param graph - The in-memory knowledge graph
* @param jclPaths - File paths of JCL files
* @param jclContents - Map of path -> file content
* @returns Summary of what was added
*/
export function processJclFiles(
graph: KnowledgeGraph,
jclPaths: string[],
jclContents: Map<string, string>,
): JclProcessResult {
let jobCount = 0;
let stepCount = 0;
let datasetCount = 0;
let programLinks = 0;
// Collect all Module names for step -> program linking
const moduleNames = new Map<string, string>(); // uppercase name -> node id
graph.forEachNode(node => {
if (node.label === 'Module') {
moduleNames.set(node.properties.name?.toUpperCase(), node.id);
}
});
for (const filePath of jclPaths) {
const content = jclContents.get(filePath);
if (!content) continue;
const parsed = parseJcl(content, filePath);
const result = integrateJclResults(graph, parsed, filePath, moduleNames);
jobCount += result.jobCount;
stepCount += result.stepCount;
datasetCount += result.datasetCount;
programLinks += result.programLinks;
}
return { jobCount, stepCount, datasetCount, programLinks };
}
function integrateJclResults(
graph: KnowledgeGraph,
parsed: JclParseResults,
filePath: string,
moduleNames: Map<string, string>,
): JclProcessResult {
let jobCount = 0;
let stepCount = 0;
let datasetCount = 0;
let programLinks = 0;
// Track step node IDs for DD -> step linking
const stepNodeIds = new Map<string, string>(); // stepName -> nodeId
// 1. Create Job nodes
for (const job of parsed.jobs) {
const jobId = generateId('CodeElement', `${filePath}:job:${job.name}`);
const classPart = job.class ? ` class:${job.class}` : '';
const msgPart = job.msgclass ? ` msgclass:${job.msgclass}` : '';
graph.addNode({
id: jobId,
label: 'CodeElement',
properties: {
name: job.name,
filePath,
startLine: job.line,
endLine: job.line,
description: `jcl-job${classPart}${msgPart}`,
},
});
// Link File -> Job (CONTAINS)
const fileId = generateId('File', filePath);
graph.addRelationship({
id: `${fileId}_contains_${jobId}`,
type: 'CONTAINS',
sourceId: fileId,
targetId: jobId,
confidence: 1.0,
reason: 'jcl-job',
});
jobCount++;
}
// 2. Create Step nodes and link to programs
for (const step of parsed.steps) {
const stepId = generateId('CodeElement', `${filePath}:step:${step.jobName}:${step.name}`);
const pgmPart = step.program ? ` pgm:${step.program}` : '';
const procPart = step.proc ? ` proc:${step.proc}` : '';
graph.addNode({
id: stepId,
label: 'CodeElement',
properties: {
name: step.name,
filePath,
startLine: step.line,
endLine: step.line,
description: `jcl-step${pgmPart}${procPart}`,
},
});
stepNodeIds.set(step.name, stepId);
// Link Job -> Step (CONTAINS)
if (step.jobName) {
const jobId = generateId('CodeElement', `${filePath}:job:${step.jobName}`);
graph.addRelationship({
id: `${jobId}_contains_${stepId}`,
type: 'CONTAINS',
sourceId: jobId,
targetId: stepId,
confidence: 1.0,
reason: 'jcl-step',
});
}
// Link Step -> Module (CALLS) when PGM= matches an indexed program
if (step.program) {
const moduleId = moduleNames.get(step.program.toUpperCase());
if (moduleId) {
graph.addRelationship({
id: `${stepId}_calls_${moduleId}`,
type: 'CALLS',
sourceId: stepId,
targetId: moduleId,
confidence: 0.95,
reason: 'jcl-exec-pgm',
});
programLinks++;
}
}
// Link Step -> PROC (CALLS) — PROC as Module
if (step.proc) {
const procModuleId = moduleNames.get(step.proc.toUpperCase());
if (procModuleId) {
graph.addRelationship({
id: `${stepId}_calls_proc_${procModuleId}`,
type: 'CALLS',
sourceId: stepId,
targetId: procModuleId,
confidence: 0.9,
reason: 'jcl-exec-proc',
});
}
}
stepCount++;
}
// 3. Create Dataset nodes from DD statements
const seenDatasets = new Set<string>();
for (const dd of parsed.ddStatements) {
if (!dd.dataset) continue;
// Create dataset node (deduplicated per file)
const datasetKey = `${filePath}:dataset:${dd.dataset}`;
const datasetId = generateId('CodeElement', datasetKey);
if (!seenDatasets.has(dd.dataset)) {
const dispPart = dd.disp ? ` disp:${dd.disp}` : '';
graph.addNode({
id: datasetId,
label: 'CodeElement',
properties: {
name: dd.dataset,
filePath,
startLine: dd.line,
endLine: dd.line,
description: `jcl-dataset${dispPart}`,
},
});
seenDatasets.add(dd.dataset);
datasetCount++;
}
// Link Step -> Dataset (CALLS with reason jcl-dd)
const stepId = stepNodeIds.get(dd.stepName);
if (stepId) {
graph.addRelationship({
id: `${stepId}_dd_${dd.ddName}_${datasetId}`,
type: 'CALLS',
sourceId: stepId,
targetId: datasetId,
confidence: 0.85,
reason: `jcl-dd:${dd.ddName}`,
});
}
}
// 4. Create PROC nodes (in-stream procs as Module)
for (const proc of parsed.procs) {
if (!proc.isInStream) continue;
const procId = generateId('Module', `${filePath}:proc:${proc.name}`);
graph.addNode({
id: procId,
label: 'Module',
properties: {
name: proc.name,
filePath,
startLine: proc.line,
endLine: proc.line,
description: 'jcl-proc-instream',
},
});
// Register for step linking
moduleNames.set(proc.name.toUpperCase(), procId);
}
// 5. INCLUDE directives -> IMPORTS edges
for (const inc of parsed.includes) {
const moduleId = moduleNames.get(inc.member.toUpperCase());
if (moduleId) {
const fileId = generateId('File', filePath);
graph.addRelationship({
id: `${fileId}_includes_${moduleId}`,
type: 'IMPORTS',
sourceId: fileId,
targetId: moduleId,
confidence: 0.9,
reason: 'jcl-include',
});
}
}
return { jobCount, stepCount, datasetCount, programLinks };
}
@@ -40,7 +40,7 @@ const UNIVERSAL_ENTRY_POINT_PATTERNS: RegExp[] = [
/^emit[A-Z]/, // emitEvent
];
const ENTRY_POINT_PATTERNS = {
export const ENTRY_POINT_PATTERNS = {
// JavaScript/TypeScript
[SupportedLanguages.JavaScript]: [
/^use[A-Z]/, // React hooks (useEffect, etc.)
@@ -216,9 +216,9 @@ const ENTRY_POINT_PATTERNS = {
/** Pre-computed merged patterns (universal + language-specific) to avoid per-call array allocation. */
const MERGED_ENTRY_POINT_PATTERNS = Object.fromEntries(
(Object.keys(ENTRY_POINT_PATTERNS) as SupportedLanguages[]).map(lang => [
Object.values(SupportedLanguages).map(lang => [
lang,
[...UNIVERSAL_ENTRY_POINT_PATTERNS, ...ENTRY_POINT_PATTERNS[lang]],
[...UNIVERSAL_ENTRY_POINT_PATTERNS, ...(ENTRY_POINT_PATTERNS[lang] ?? [])],
])
) as Record<SupportedLanguages, RegExp[]>;
+27 -51
View File
@@ -7,18 +7,17 @@
* Shared between parse-worker.ts (worker pool) and parsing-processor.ts (sequential fallback).
*/
import { findSiblingChild, SyntaxNode } from './utils.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
import { findSiblingChild, type SyntaxNode } from './utils/ast-helpers.js';
/** Handler type: given a node and symbol name, return true if the symbol is exported/public. */
type ExportChecker = (node: SyntaxNode, name: string) => boolean;
export type ExportChecker = (node: SyntaxNode, name: string) => boolean;
// ============================================================================
// Per-language export checkers
// ============================================================================
/** JS/TS: walk ancestors looking for export_statement or export_specifier. */
const tsExportChecker: ExportChecker = (node, _name) => {
export const tsExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
const type = current.type;
@@ -37,10 +36,10 @@ const tsExportChecker: ExportChecker = (node, _name) => {
};
/** Python: public if no leading underscore (convention). */
const pythonExportChecker: ExportChecker = (_node, name) => !name.startsWith('_');
export const pythonExportChecker: ExportChecker = (_node, name) => !name.startsWith('_');
/** Java: check for 'public' modifier — modifiers are siblings of the name node, not parents. */
const javaExportChecker: ExportChecker = (node, _name) => {
export const javaExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.parent) {
@@ -76,7 +75,7 @@ const CSHARP_DECL_TYPES = new Set([
* C#: modifier nodes are SIBLINGS of the name node inside the declaration.
* Walk up to the declaration node, then scan its direct children.
*/
const csharpExportChecker: ExportChecker = (node, _name) => {
export const csharpExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (CSHARP_DECL_TYPES.has(current.type)) {
@@ -92,7 +91,7 @@ const csharpExportChecker: ExportChecker = (node, _name) => {
};
/** Go: uppercase first letter = exported. */
const goExportChecker: ExportChecker = (_node, name) => {
export const goExportChecker: ExportChecker = (_node, name) => {
if (name.length === 0) return false;
const first = name[0];
return first === first.toUpperCase() && first !== first.toLowerCase();
@@ -110,7 +109,7 @@ const RUST_DECL_TYPES = new Set([
* (function_item, struct_item, etc.), not a parent. Walk up to the declaration node,
* then scan its direct children.
*/
const rustExportChecker: ExportChecker = (node, _name) => {
export const rustExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (RUST_DECL_TYPES.has(current.type)) {
@@ -129,7 +128,7 @@ const rustExportChecker: ExportChecker = (node, _name) => {
* Kotlin: default visibility is public (unlike Java).
* visibility_modifier is inside modifiers, a sibling of the name node within the declaration.
*/
const kotlinExportChecker: ExportChecker = (node, _name) => {
export const kotlinExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.parent) {
@@ -152,7 +151,7 @@ const kotlinExportChecker: ExportChecker = (node, _name) => {
* marked 'static' are file-scoped (not exported). C++ anonymous namespaces
* (namespace { ... }) also give internal linkage.
*/
const cCppExportChecker: ExportChecker = (node, _name) => {
export const cCppExportChecker: ExportChecker = (node, _name) => {
let cur: SyntaxNode | null = node;
while (cur) {
if (cur.type === 'function_definition' || cur.type === 'declaration') {
@@ -174,7 +173,7 @@ const cCppExportChecker: ExportChecker = (node, _name) => {
};
/** PHP: check for visibility modifier or top-level scope. */
const phpExportChecker: ExportChecker = (node, _name) => {
export const phpExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.type === 'class_declaration' ||
@@ -192,52 +191,29 @@ const phpExportChecker: ExportChecker = (node, _name) => {
return true;
};
/** Swift: check for 'public' or 'open' access modifiers. */
const swiftExportChecker: ExportChecker = (node, _name) => {
/**
* Swift: treat symbols as exported unless explicitly marked private/fileprivate.
*
* Swift's default access level is `internal`, which means visible to all files
* in the same module/target. Since GitNexus indexes at the target level,
* `internal` symbols should be treated as exported (cross-file visible).
* Only `private` and `fileprivate` symbols are truly file-scoped.
*/
export const swiftExportChecker: ExportChecker = (node, _name) => {
let current: SyntaxNode | null = node;
while (current) {
if (current.type === 'modifiers' || current.type === 'visibility_modifier') {
const text = current.text || '';
if (text.includes('public') || text.includes('open')) return true;
// Exclude private(set)/fileprivate(set) — only the setter is restricted,
// the symbol itself is still readable cross-file.
if (/\b(private|fileprivate)\b(?!\s*\()/.test(text)) return false;
}
current = current.parent;
}
return false;
// Default (internal), public, and open are all cross-file visible
return true;
};
// ============================================================================
// Exhaustive dispatch table — satisfies enforces all SupportedLanguages are covered
// ============================================================================
/** Ruby: all top-level definitions are public (no export syntax). */
export const rubyExportChecker: ExportChecker = (_node, _name) => true;
const exportCheckers = {
[SupportedLanguages.JavaScript]: tsExportChecker,
[SupportedLanguages.TypeScript]: tsExportChecker,
[SupportedLanguages.Python]: pythonExportChecker,
[SupportedLanguages.Java]: javaExportChecker,
[SupportedLanguages.CSharp]: csharpExportChecker,
[SupportedLanguages.Go]: goExportChecker,
[SupportedLanguages.Rust]: rustExportChecker,
[SupportedLanguages.Kotlin]: kotlinExportChecker,
[SupportedLanguages.C]: cCppExportChecker,
[SupportedLanguages.CPlusPlus]: cCppExportChecker,
[SupportedLanguages.PHP]: phpExportChecker,
[SupportedLanguages.Swift]: swiftExportChecker,
[SupportedLanguages.Ruby]: (_node, _name) => true,
} satisfies Record<SupportedLanguages, ExportChecker>;
// ============================================================================
// Public API
// ============================================================================
/**
* Check if a tree-sitter node is exported/public in its language.
* @param node - The tree-sitter AST node
* @param name - The symbol name
* @param language - The programming language
* @returns true if the symbol is exported/public
*/
export const isNodeExported = (node: SyntaxNode, name: string, language: SupportedLanguages): boolean => {
const checker = exportCheckers[language];
if (!checker) return false;
return checker(node, name);
};
@@ -471,7 +471,7 @@ interface AstFrameworkPatternConfig {
patterns: string[];
}
const AST_FRAMEWORK_PATTERNS_BY_LANGUAGE = {
export const AST_FRAMEWORK_PATTERNS_BY_LANGUAGE = {
[SupportedLanguages.JavaScript]: [
{ framework: 'nestjs', entryPointMultiplier: 3.2, reason: 'nestjs-decorator', patterns: FRAMEWORK_AST_PATTERNS.nestjs },
],
@@ -18,24 +18,23 @@ import { KnowledgeGraph } from '../graph/types.js';
import { ASTCache } from './ast-cache.js';
import Parser from 'tree-sitter';
import { isLanguageAvailable, loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { generateId } from '../../lib/utils.js';
import { getLanguageFromFilename, isVerboseIngestionEnabled, yieldToEventLoop } from './utils.js';
import { getLanguageFromFilename } from './utils/language-detection.js';
import { isVerboseIngestionEnabled } from './utils/verbose.js';
import { yieldToEventLoop } from './utils/event-loop.js';
import { SupportedLanguages } from '../../config/supported-languages.js';
import { getProvider } from './languages/index.js';
import { getTreeSitterBufferSize } from './constants.js';
import type { ExtractedHeritage } from './workers/parse-worker.js';
import type { ResolutionContext } from './resolution-context.js';
import { TIER_CONFIDENCE } from './resolution-context.js';
/** C#/Java convention: interfaces start with I followed by an uppercase letter */
const INTERFACE_NAME_RE = /^I[A-Z]/;
/**
* Determine whether a heritage.extends capture is actually an IMPLEMENTS relationship.
* Uses the symbol table first (authoritative — Tier 1); falls back to a language-gated
* heuristic for external symbols not present in the graph:
* - C# / Java: `I[A-Z]` naming convention
* - Swift: default IMPLEMENTS (protocol conformance is the norm)
* Uses the symbol table first (authoritative — Tier 1); falls back to provider-defined
* heuristics for external symbols not present in the graph:
* - interfaceNamePattern: matched against parent name (e.g., /^I[A-Z]/ for C#/Java)
* - heritageDefaultEdge: 'IMPLEMENTS' causes all unresolved parents to map to IMPLEMENTS
* - All others: default EXTENDS
*/
const resolveExtendsType = (
@@ -51,13 +50,12 @@ const resolveExtendsType = (
? { type: 'IMPLEMENTS', idPrefix: 'Interface' }
: { type: 'EXTENDS', idPrefix: 'Class' };
}
// Unresolved symbol — fall back to language-specific heuristic
if (language === SupportedLanguages.CSharp || language === SupportedLanguages.Java) {
if (INTERFACE_NAME_RE.test(parentName)) {
return { type: 'IMPLEMENTS', idPrefix: 'Interface' };
}
} else if (language === SupportedLanguages.Swift) {
// Protocol conformance is far more common than class inheritance in Swift
// Unresolved symbol — fall back to provider-defined heuristics
const provider = getProvider(language);
if (provider.interfaceNamePattern?.test(parentName)) {
return { type: 'IMPLEMENTS', idPrefix: 'Interface' };
}
if (provider.heritageDefaultEdge === 'IMPLEMENTS') {
return { type: 'IMPLEMENTS', idPrefix: 'Interface' };
}
return { type: 'EXTENDS', idPrefix: 'Class' };
@@ -117,7 +115,8 @@ export const processHeritage = async (
continue;
}
const queryStr = LANGUAGE_QUERIES[language];
const provider = getProvider(language);
const queryStr = provider.treeSitterQueries;
if (!queryStr) continue;
// 2. Load the language
+74 -32
View File
@@ -2,27 +2,22 @@ import { KnowledgeGraph } from '../graph/types.js';
import { ASTCache } from './ast-cache.js';
import Parser from 'tree-sitter';
import { isLanguageAvailable, loadParser, loadLanguage } from '../tree-sitter/parser-loader.js';
import { LANGUAGE_QUERIES } from './tree-sitter-queries.js';
import { getProvider, getProviderForFile, providersWithImplicitWiring } from './languages/index.js';
import type { LanguageProvider } from './language-provider.js';
import { generateId } from '../../lib/utils.js';
import { getLanguageFromFilename, isVerboseIngestionEnabled, yieldToEventLoop } from './utils.js';
import { getLanguageFromFilename } from './utils/language-detection.js';
import { isVerboseIngestionEnabled } from './utils/verbose.js';
import { yieldToEventLoop } from './utils/event-loop.js';
import type { ExtractedImport } from './workers/parse-worker.js';
import { getTreeSitterBufferSize } from './constants.js';
import { loadImportConfigs } from './language-config.js';
import { buildSuffixIndex } from './resolvers/index.js';
import { callRouters } from './call-routing.js';
import { buildSuffixIndex } from './import-resolvers/utils.js';
import type { ResolutionContext, ModuleAliasMap } from './resolution-context.js';
import type { SuffixIndex } from './resolvers/index.js';
import { importResolvers, namedBindingExtractors, preprocessImportPath } from './import-resolution.js';
import type { ImportResult, ResolveCtx, NamedBinding } from './import-resolution.js';
import type { SuffixIndex } from './import-resolvers/utils.js';
import type { ImportResult, ResolveCtx, ImportResolutionContext } from './import-resolvers/types.js';
import type { NamedBinding } from './named-bindings/types.js';
import type { SyntaxNode } from './utils/ast-helpers.js';
// Re-export resolver types for consumers
export type {
SuffixIndex,
TsconfigPaths,
GoModuleConfig,
CSharpProjectConfig,
ComposerConfig
} from './resolvers/index.js';
const isDev = process.env.NODE_ENV === 'development';
@@ -30,6 +25,32 @@ const isDev = process.env.NODE_ENV === 'development';
// Stores all files that a given file imports from
export type ImportMap = Map<string, Set<string>>;
/** Group files by provider (only those with implicit import wiring), then call each wirer
* with its own language's files. O(n) over files, O(1) per provider lookup. */
function wireImplicitImports(
files: string[],
importMap: Map<string, Set<string>>,
addImportEdge: (src: string, target: string) => void,
projectConfig: unknown,
): void {
if (providersWithImplicitWiring.length === 0) return;
const grouped = new Map<LanguageProvider, string[]>();
for (const file of files) {
const provider = getProviderForFile(file);
if (!provider?.implicitImportWirer) continue;
let list = grouped.get(provider);
if (!list) { list = []; grouped.set(provider, list); }
list.push(file);
}
for (const [provider, langFiles] of grouped) {
if (langFiles.length > 1) {
provider.implicitImportWirer(langFiles, importMap, addImportEdge, projectConfig);
}
}
}
// Type: Map<FilePath, Set<PackageDirSuffix>>
// Stores Go package directory suffixes imported by a file (e.g., "/internal/auth/").
// Avoids expanding every Go package import into N individual ImportMap edges.
@@ -56,14 +77,7 @@ export function isFileInPackageDir(filePath: string, dirSuffix: string): boolean
return !afterDir.includes('/');
}
/** Pre-built lookup structures for import resolution. Build once, reuse across chunks. */
export interface ImportResolutionContext {
allFilePaths: Set<string>;
allFileList: string[];
normalizedFileList: string[];
index: SuffixIndex;
resolveCache: Map<string, string | null>;
}
// ImportResolutionContext is defined in ./import-resolvers/types.ts — re-exported here for consumers.
export function buildImportResolutionContext(allPaths: string[]): ImportResolutionContext {
const allFileList = allPaths;
@@ -74,7 +88,30 @@ export function buildImportResolutionContext(allPaths: string[]): ImportResoluti
}
// Config loaders extracted to ./language-config.ts (Phase 2 refactor)
// Resolver dispatch tables are in ./import-resolution.ts — imported above
// Resolver types are in ./import-resolvers/types.ts; named binding types in ./named-bindings/types.ts
// ============================================================================
// Import path preprocessing
// ============================================================================
/**
* Clean and preprocess a raw import source text into a resolved import path.
* Strips quotes/angle brackets (universal) and applies provider-specific
* transformations (currently only Kotlin wildcard import detection).
*/
export function preprocessImportPath(
sourceText: string,
importNode: SyntaxNode,
provider: LanguageProvider,
): string | null {
const cleaned = sourceText.replace(/['"<>]/g, '');
// Defense-in-depth: reject null bytes and control characters (matches Ruby call-routing pattern)
if (!cleaned || cleaned.length > 2048 || /[\x00-\x1f]/.test(cleaned)) return null;
if (provider.importPathPreprocessor) {
return provider.importPathPreprocessor(cleaned, importNode);
}
return cleaned;
}
/** Create IMPORTS edge helpers that share a resolved-count tracker. */
function createImportEdgeHelpers(graph: KnowledgeGraph, importMap: ImportMap) {
@@ -242,7 +279,8 @@ export const processImports = async (
continue;
}
const queryStr = LANGUAGE_QUERIES[language];
const provider = getProvider(language);
const queryStr = provider.treeSitterQueries;
if (!queryStr) continue;
// 2. ALWAYS load the language before querying (parser is stateful)
@@ -298,12 +336,12 @@ export const processImports = async (
return;
}
const rawImportPath = preprocessImportPath(sourceNode.text, captureMap['import'], language);
const rawImportPath = preprocessImportPath(sourceNode.text, captureMap['import'], provider);
if (!rawImportPath) return;
totalImportsFound++;
const result = importResolvers[language](rawImportPath, file.path, resolveCtx);
const extractor = namedBindingExtractors[language];
const result = provider.importResolver(rawImportPath, file.path, resolveCtx);
const extractor = provider.namedBindingExtractor;
const bindings = namedImportMap && extractor ? extractor(captureMap['import']) : undefined;
applyImportResult(result, file.path, importMap, packageMap, addImportEdge, addImportGraphEdge, bindings, namedImportMap, moduleAliasMap);
}
@@ -312,11 +350,10 @@ export const processImports = async (
if (captureMap['call']) {
const callNameNode = captureMap['call.name'];
if (callNameNode) {
const callRouter = callRouters[language];
const routed = callRouter(callNameNode.text, captureMap['call']);
const routed = provider.callRouter?.(callNameNode.text, captureMap['call']);
if (routed && routed.kind === 'import') {
totalImportsFound++;
const result = importResolvers[language](routed.importPath, file.path, resolveCtx);
const result = provider.importResolver(routed.importPath, file.path, resolveCtx);
applyImportResult(result, file.path, importMap, packageMap, addImportEdge, addImportGraphEdge);
}
}
@@ -326,6 +363,8 @@ export const processImports = async (
// Tree is now owned by the LRU cache — no manual delete needed
}
wireImplicitImports(allFileList, importMap, addImportEdge, configs);
if (skippedByLang && skippedByLang.size > 0) {
for (const [lang, count] of skippedByLang.entries()) {
console.warn(
@@ -389,13 +428,16 @@ export const processImportsFromExtracted = async (
for (const imp of fileImports) {
totalImportsFound++;
const result = importResolvers[imp.language](imp.rawImportPath, filePath, resolveCtx);
const provider = getProvider(imp.language);
const result = provider.importResolver(imp.rawImportPath, filePath, resolveCtx);
applyImportResult(result, filePath, importMap, packageMap, addImportEdge, addImportGraphEdge, imp.namedBindings, namedImportMap, moduleAliasMap);
}
}
onProgress?.(totalFiles, totalFiles);
wireImplicitImports(files.map(f => f.path), importMap, addImportEdge, configs);
if (isDev) {
console.log(`📊 Import processing (fast path): ${getResolvedCount()}/${totalImportsFound} imports resolved to graph edges`);
}
@@ -1,385 +0,0 @@
/**
* Import Resolution Dispatch
*
* Per-language dispatch table for import resolution and named binding extraction.
* Replaces the 120-line if-chain in resolveLanguageImport() and the 7-branch
* dispatch in extractNamedBindings() with a single table lookup each.
*
* Follows the existing ExportChecker / CallRouter pattern:
* - Function aliases (not interfaces) to avoid megamorphic inline-cache issues
* - `satisfies Record<SupportedLanguages, ...>` for compile-time exhaustiveness
* - Const dispatch table — configs are accessed via ctx.configs at call time
*/
import { SupportedLanguages } from '../../config/supported-languages.js';
import type { SyntaxNode } from './utils.js';
import {
KOTLIN_EXTENSIONS,
appendKotlinWildcard,
resolveJvmWildcard,
resolveJvmMemberImport,
resolveGoPackageDir,
resolveGoPackage,
resolveCSharpImport as resolveCSharpImportHelper,
resolveCSharpNamespaceDir,
resolvePhpImport as resolvePhpImportHelper,
resolveRustImport as resolveRustImportHelper,
resolveRubyImport as resolveRubyImportHelper,
resolvePythonImport as resolvePythonImportHelper,
resolveImportPath,
} from './resolvers/index.js';
import type {
SuffixIndex,
TsconfigPaths,
GoModuleConfig,
CSharpProjectConfig,
ComposerConfig,
} from './resolvers/index.js';
import type { SwiftPackageConfig } from './language-config.js';
import {
extractTsNamedBindings,
extractPythonNamedBindings,
extractKotlinNamedBindings,
extractRustNamedBindings,
extractPhpNamedBindings,
extractCsharpNamedBindings,
extractJavaNamedBindings,
} from './named-binding-extraction.js';
import type { ImportResolutionContext } from './import-processor.js';
// ============================================================================
// Types
// ============================================================================
/**
* Result of resolving an import via language-specific dispatch.
* - 'files': resolved to one or more files -> add to ImportMap
* - 'package': resolved to a directory -> add graph edges + store dirSuffix in PackageMap
* - null: no resolution (external dependency, etc.)
*/
export type ImportResult =
| { kind: 'files'; files: string[] }
| { kind: 'package'; files: string[]; dirSuffix: string }
| null;
/** Bundled language-specific configs loaded once per ingestion run. */
export interface ImportConfigs {
tsconfigPaths: TsconfigPaths | null;
goModule: GoModuleConfig | null;
composerConfig: ComposerConfig | null;
swiftPackageConfig: SwiftPackageConfig | null;
csharpConfigs: CSharpProjectConfig[];
}
/** Full context for import resolution: file lookups + language configs. */
export interface ResolveCtx extends ImportResolutionContext {
configs: ImportConfigs;
}
/** Per-language import resolver -- function alias matching ExportChecker/CallRouter pattern. */
export type ImportResolverFn = (
rawImportPath: string,
filePath: string,
resolveCtx: ResolveCtx,
) => ImportResult;
/** A single named import binding: local name in the importing file and exported name from the source.
* When `isModuleAlias` is true, the binding represents a Python `import X as Y` module alias
* and is routed to moduleAliasMap instead of namedImportMap during import processing. */
export interface NamedBinding { local: string; exported: string; isModuleAlias?: boolean }
/** Per-language named binding extractor -- optional (returns undefined if language has no named imports). */
type NamedBindingExtractorFn = (importNode: SyntaxNode) => NamedBinding[] | undefined;
// ============================================================================
// Import path preprocessing
// ============================================================================
/**
* Clean and preprocess a raw import source text into a resolved import path.
* Strips quotes/angle brackets (universal) and applies language-specific
* transformations (currently only Kotlin wildcard import detection).
*/
export function preprocessImportPath(
sourceText: string,
importNode: SyntaxNode,
language: SupportedLanguages,
): string | null {
const cleaned = sourceText.replace(/['"<>]/g, '');
// Defense-in-depth: reject null bytes and control characters (matches Ruby call-routing pattern)
if (!cleaned || cleaned.length > 2048 || /[\x00-\x1f]/.test(cleaned)) return null;
if (language === SupportedLanguages.Kotlin) {
return appendKotlinWildcard(cleaned, importNode);
}
return cleaned;
}
// ============================================================================
// Per-language resolver functions
// ============================================================================
/**
* Standard single-file resolution (TS/JS/C/C++ and fallback for other languages).
* Handles relative imports, tsconfig path aliases, and suffix matching.
*/
function resolveStandard(
rawImportPath: string,
filePath: string,
ctx: ResolveCtx,
language: SupportedLanguages,
): ImportResult {
const resolvedPath = resolveImportPath(
filePath,
rawImportPath,
ctx.allFilePaths,
ctx.allFileList,
ctx.normalizedFileList,
ctx.resolveCache,
language,
ctx.configs.tsconfigPaths,
ctx.index,
);
return resolvedPath ? { kind: 'files', files: [resolvedPath] } : null;
}
/** Java: JVM wildcard -> member import -> standard fallthrough */
function resolveJavaImport(
rawImportPath: string,
filePath: string,
ctx: ResolveCtx,
): ImportResult {
if (rawImportPath.endsWith('.*')) {
const matchedFiles = resolveJvmWildcard(rawImportPath, ctx.normalizedFileList, ctx.allFileList, ['.java'], ctx.index);
if (matchedFiles.length > 0) return { kind: 'files', files: matchedFiles };
} else {
const memberResolved = resolveJvmMemberImport(rawImportPath, ctx.normalizedFileList, ctx.allFileList, ['.java'], ctx.index);
if (memberResolved) return { kind: 'files', files: [memberResolved] };
}
return resolveStandard(rawImportPath, filePath, ctx, SupportedLanguages.Java);
}
/**
* Kotlin: JVM wildcard/member with Java-interop fallback -> top-level function imports -> standard.
* Kotlin can import from .kt/.kts files OR from .java files (Java interop).
*/
function resolveKotlinImport(
rawImportPath: string,
filePath: string,
ctx: ResolveCtx,
): ImportResult {
if (rawImportPath.endsWith('.*')) {
const matchedFiles = resolveJvmWildcard(rawImportPath, ctx.normalizedFileList, ctx.allFileList, KOTLIN_EXTENSIONS, ctx.index);
if (matchedFiles.length === 0) {
const javaMatches = resolveJvmWildcard(rawImportPath, ctx.normalizedFileList, ctx.allFileList, ['.java'], ctx.index);
if (javaMatches.length > 0) return { kind: 'files', files: javaMatches };
}
if (matchedFiles.length > 0) return { kind: 'files', files: matchedFiles };
} else {
let memberResolved = resolveJvmMemberImport(rawImportPath, ctx.normalizedFileList, ctx.allFileList, KOTLIN_EXTENSIONS, ctx.index);
if (!memberResolved) {
memberResolved = resolveJvmMemberImport(rawImportPath, ctx.normalizedFileList, ctx.allFileList, ['.java'], ctx.index);
}
if (memberResolved) return { kind: 'files', files: [memberResolved] };
// Kotlin: top-level function imports (e.g. import models.getUser) have only 2 segments,
// which resolveJvmMemberImport skips (requires >=3). Fall back to package-directory scan
// for lowercase last segments (function/property imports). Uppercase last segments
// (class imports like models.User) fall through to standard suffix resolution.
const segments = rawImportPath.split('.');
const lastSeg = segments[segments.length - 1];
if (segments.length >= 2 && lastSeg[0] && lastSeg[0] === lastSeg[0].toLowerCase()) {
const pkgWildcard = segments.slice(0, -1).join('.') + '.*';
let dirFiles = resolveJvmWildcard(pkgWildcard, ctx.normalizedFileList, ctx.allFileList, KOTLIN_EXTENSIONS, ctx.index);
if (dirFiles.length === 0) {
dirFiles = resolveJvmWildcard(pkgWildcard, ctx.normalizedFileList, ctx.allFileList, ['.java'], ctx.index);
}
if (dirFiles.length > 0) return { kind: 'files', files: dirFiles };
}
}
return resolveStandard(rawImportPath, filePath, ctx, SupportedLanguages.Kotlin);
}
/** Go: package-level imports via go.mod module path. */
function resolveGoImport(
rawImportPath: string,
filePath: string,
ctx: ResolveCtx,
): ImportResult {
const goModule = ctx.configs.goModule;
if (goModule && rawImportPath.startsWith(goModule.modulePath)) {
const pkgSuffix = resolveGoPackageDir(rawImportPath, goModule);
if (pkgSuffix) {
const pkgFiles = resolveGoPackage(rawImportPath, goModule, ctx.normalizedFileList, ctx.allFileList);
if (pkgFiles.length > 0) {
return { kind: 'package', files: pkgFiles, dirSuffix: pkgSuffix };
}
}
// Fall through if no files found (package might be external)
}
return resolveStandard(rawImportPath, filePath, ctx, SupportedLanguages.Go);
}
/** C#: namespace-based resolution via .csproj configs, with suffix-match fallback. */
function resolveCSharpImportDispatch(
rawImportPath: string,
filePath: string,
ctx: ResolveCtx,
): ImportResult {
const csharpConfigs = ctx.configs.csharpConfigs;
if (csharpConfigs.length > 0) {
const resolvedFiles = resolveCSharpImportHelper(rawImportPath, csharpConfigs, ctx.normalizedFileList, ctx.allFileList, ctx.index);
if (resolvedFiles.length > 1) {
const dirSuffix = resolveCSharpNamespaceDir(rawImportPath, csharpConfigs);
if (dirSuffix) {
return { kind: 'package', files: resolvedFiles, dirSuffix };
}
}
if (resolvedFiles.length > 0) return { kind: 'files', files: resolvedFiles };
}
return resolveStandard(rawImportPath, filePath, ctx, SupportedLanguages.CSharp);
}
/** PHP: namespace-based resolution via composer.json PSR-4. */
function resolvePhpImportDispatch(
rawImportPath: string,
_filePath: string,
ctx: ResolveCtx,
): ImportResult {
const resolved = resolvePhpImportHelper(rawImportPath, ctx.configs.composerConfig, ctx.allFilePaths, ctx.normalizedFileList, ctx.allFileList, ctx.index);
return resolved ? { kind: 'files', files: [resolved] } : null;
}
/** Swift: module imports via Package.swift target map. */
function resolveSwiftImportDispatch(
rawImportPath: string,
_filePath: string,
ctx: ResolveCtx,
): ImportResult {
const swiftPackageConfig = ctx.configs.swiftPackageConfig;
if (swiftPackageConfig) {
const targetDir = swiftPackageConfig.targets.get(rawImportPath);
if (targetDir) {
const dirPrefix = targetDir + '/';
const files: string[] = [];
for (let i = 0; i < ctx.normalizedFileList.length; i++) {
if (ctx.normalizedFileList[i].startsWith(dirPrefix) && ctx.normalizedFileList[i].endsWith('.swift')) {
files.push(ctx.allFileList[i]);
}
}
if (files.length > 0) return { kind: 'files', files };
}
}
return null; // External framework (Foundation, UIKit, etc.)
}
/**
* Python: relative imports (PEP 328) + proximity-based bare imports.
* Falls through to standard suffix resolution when proximity finds no match.
*/
function resolvePythonImportDispatch(
rawImportPath: string,
filePath: string,
ctx: ResolveCtx,
): ImportResult {
const resolved = resolvePythonImportHelper(filePath, rawImportPath, ctx.allFilePaths);
if (resolved) return { kind: 'files', files: [resolved] };
if (rawImportPath.startsWith('.')) return null; // relative but unresolved -- don't suffix-match
return resolveStandard(rawImportPath, filePath, ctx, SupportedLanguages.Python);
}
/** Ruby: require / require_relative. */
function resolveRubyImportDispatch(
rawImportPath: string,
_filePath: string,
ctx: ResolveCtx,
): ImportResult {
const resolved = resolveRubyImportHelper(rawImportPath, ctx.normalizedFileList, ctx.allFileList, ctx.index);
return resolved ? { kind: 'files', files: [resolved] } : null;
}
/** Rust: expand grouped imports: use {crate::a, crate::b} and use crate::models::{User, Repo}. */
function resolveRustImportDispatch(
rawImportPath: string,
filePath: string,
ctx: ResolveCtx,
): ImportResult {
// Top-level grouped: use {crate::a, crate::b}
if (rawImportPath.startsWith('{') && rawImportPath.endsWith('}')) {
const inner = rawImportPath.slice(1, -1);
const parts = inner.split(',').map(p => p.trim()).filter(Boolean);
const resolved: string[] = [];
for (const part of parts) {
const r = resolveRustImportHelper(filePath, part, ctx.allFilePaths);
if (r) resolved.push(r);
}
return resolved.length > 0 ? { kind: 'files', files: resolved } : null;
}
// Scoped grouped: use crate::models::{User, Repo}
const braceIdx = rawImportPath.indexOf('::{');
if (braceIdx !== -1 && rawImportPath.endsWith('}')) {
const pathPrefix = rawImportPath.substring(0, braceIdx);
const braceContent = rawImportPath.substring(braceIdx + 3, rawImportPath.length - 1);
const items = braceContent.split(',').map(s => s.trim()).filter(Boolean);
const resolved: string[] = [];
for (const item of items) {
// Handle `use crate::models::{User, Repo as R}` — strip alias for resolution
const itemName = item.includes(' as ') ? item.split(' as ')[0].trim() : item;
const r = resolveRustImportHelper(filePath, `${pathPrefix}::${itemName}`, ctx.allFilePaths);
if (r) resolved.push(r);
}
if (resolved.length > 0) return { kind: 'files', files: resolved };
// Fallback: resolve the prefix path itself (e.g. crate::models -> models.rs)
const prefixResult = resolveRustImportHelper(filePath, pathPrefix, ctx.allFilePaths);
if (prefixResult) return { kind: 'files', files: [prefixResult] };
}
return resolveStandard(rawImportPath, filePath, ctx, SupportedLanguages.Rust);
}
// ============================================================================
// Dispatch tables
// ============================================================================
/**
* Per-language import resolver dispatch table.
* Configs are accessed via ctx.configs at call time — no factory closure needed.
* Each resolver encapsulates the full resolution flow for its language, including
* fallthrough to standard resolution where appropriate.
*/
export const importResolvers = {
[SupportedLanguages.JavaScript]: (raw, fp, ctx) => resolveStandard(raw, fp, ctx, SupportedLanguages.JavaScript),
[SupportedLanguages.TypeScript]: (raw, fp, ctx) => resolveStandard(raw, fp, ctx, SupportedLanguages.TypeScript),
[SupportedLanguages.Python]: (raw, fp, ctx) => resolvePythonImportDispatch(raw, fp, ctx),
[SupportedLanguages.Java]: (raw, fp, ctx) => resolveJavaImport(raw, fp, ctx),
[SupportedLanguages.C]: (raw, fp, ctx) => resolveStandard(raw, fp, ctx, SupportedLanguages.C),
[SupportedLanguages.CPlusPlus]: (raw, fp, ctx) => resolveStandard(raw, fp, ctx, SupportedLanguages.CPlusPlus),
[SupportedLanguages.CSharp]: (raw, fp, ctx) => resolveCSharpImportDispatch(raw, fp, ctx),
[SupportedLanguages.Go]: (raw, fp, ctx) => resolveGoImport(raw, fp, ctx),
[SupportedLanguages.Ruby]: (raw, fp, ctx) => resolveRubyImportDispatch(raw, fp, ctx),
[SupportedLanguages.Rust]: (raw, fp, ctx) => resolveRustImportDispatch(raw, fp, ctx),
[SupportedLanguages.PHP]: (raw, fp, ctx) => resolvePhpImportDispatch(raw, fp, ctx),
[SupportedLanguages.Kotlin]: (raw, fp, ctx) => resolveKotlinImport(raw, fp, ctx),
[SupportedLanguages.Swift]: (raw, fp, ctx) => resolveSwiftImportDispatch(raw, fp, ctx),
} satisfies Record<SupportedLanguages, ImportResolverFn>;
/**
* Per-language named binding extractor dispatch table.
* Languages with whole-module import semantics (Go, Ruby, C/C++, Swift) return undefined --
* their bindings are synthesized post-parse by synthesizeWildcardImportBindings() in pipeline.ts.
*/
export const namedBindingExtractors = {
[SupportedLanguages.JavaScript]: extractTsNamedBindings,
[SupportedLanguages.TypeScript]: extractTsNamedBindings,
[SupportedLanguages.Python]: extractPythonNamedBindings,
[SupportedLanguages.Java]: extractJavaNamedBindings,
[SupportedLanguages.C]: undefined,
[SupportedLanguages.CPlusPlus]: undefined,
[SupportedLanguages.CSharp]: extractCsharpNamedBindings,
[SupportedLanguages.Go]: undefined,
[SupportedLanguages.Ruby]: undefined,
[SupportedLanguages.Rust]: extractRustNamedBindings,
[SupportedLanguages.PHP]: extractPhpNamedBindings,
[SupportedLanguages.Kotlin]: extractKotlinNamedBindings,
[SupportedLanguages.Swift]: undefined,
} satisfies Record<SupportedLanguages, NamedBindingExtractorFn | undefined>;
@@ -5,20 +5,16 @@
import type { SuffixIndex } from './utils.js';
import { suffixResolve } from './utils.js';
/** C# project config parsed from .csproj files */
export interface CSharpProjectConfig {
/** Root namespace from <RootNamespace> or assembly name (default: project directory name) */
rootNamespace: string;
/** Directory containing the .csproj file */
projectDir: string;
}
import { SupportedLanguages } from '../../../config/supported-languages.js';
import type { ImportResult, ResolveCtx } from './types.js';
import { resolveStandard } from './standard.js';
import type { CSharpProjectConfig } from '../language-config.js';
/**
* Resolve a C# using-directive import path to matching .cs files.
* Resolve a C# using-directive import path to matching .cs files (low-level helper).
* Tries single-file match first, then directory match for namespace imports.
*/
export function resolveCSharpImport(
export function resolveCSharpImportInternal(
importPath: string,
csharpConfigs: CSharpProjectConfig[],
normalizedFileList: string[],
@@ -126,3 +122,23 @@ export function resolveCSharpNamespaceDir(
return null;
}
/** C#: namespace-based resolution via .csproj configs, with suffix-match fallback. */
export function resolveCSharpImport(
rawImportPath: string,
filePath: string,
ctx: ResolveCtx,
): ImportResult {
const csharpConfigs = ctx.configs.csharpConfigs;
if (csharpConfigs.length > 0) {
const resolvedFiles = resolveCSharpImportInternal(rawImportPath, csharpConfigs, ctx.normalizedFileList, ctx.allFileList, ctx.index);
if (resolvedFiles.length > 1) {
const dirSuffix = resolveCSharpNamespaceDir(rawImportPath, csharpConfigs);
if (dirSuffix) {
return { kind: 'package', files: resolvedFiles, dirSuffix };
}
}
if (resolvedFiles.length > 0) return { kind: 'files', files: resolvedFiles };
}
return resolveStandard(rawImportPath, filePath, ctx, SupportedLanguages.CSharp);
}
@@ -3,11 +3,10 @@
* Handles Go module path-based package imports.
*/
/** Go module config parsed from go.mod */
export interface GoModuleConfig {
/** Module path (e.g., "github.com/user/repo") */
modulePath: string;
}
import { SupportedLanguages } from '../../../config/supported-languages.js';
import type { ImportResult, ResolveCtx } from './types.js';
import { resolveStandard } from './standard.js';
import type { GoModuleConfig } from '../language-config.js';
/**
* Extract the package directory suffix from a Go import path.
@@ -56,3 +55,23 @@ export function resolveGoPackage(
return matches;
}
/** Go: package-level imports via go.mod module path. */
export function resolveGoImport(
rawImportPath: string,
filePath: string,
ctx: ResolveCtx,
): ImportResult {
const goModule = ctx.configs.goModule;
if (goModule && rawImportPath.startsWith(goModule.modulePath)) {
const pkgSuffix = resolveGoPackageDir(rawImportPath, goModule);
if (pkgSuffix) {
const pkgFiles = resolveGoPackage(rawImportPath, goModule, ctx.normalizedFileList, ctx.allFileList);
if (pkgFiles.length > 0) {
return { kind: 'package', files: pkgFiles, dirSuffix: pkgSuffix };
}
}
// Fall through if no files found (package might be external)
}
return resolveStandard(rawImportPath, filePath, ctx, SupportedLanguages.Go);
}
@@ -4,7 +4,10 @@
*/
import type { SuffixIndex } from './utils.js';
import type { SyntaxNode } from '../utils.js';
import type { SyntaxNode } from '../utils/ast-helpers.js';
import { SupportedLanguages } from '../../../config/supported-languages.js';
import type { ImportResult, ResolveCtx } from './types.js';
import { resolveStandard } from './standard.js';
/** Kotlin file extensions for JVM resolver reuse */
export const KOTLIN_EXTENSIONS: readonly string[] = ['.kt', '.kts'];
@@ -118,3 +121,60 @@ export function resolveJvmMemberImport(
return null;
}
/** Java: JVM wildcard -> member import -> standard fallthrough */
export function resolveJavaImport(
rawImportPath: string,
filePath: string,
ctx: ResolveCtx,
): ImportResult {
if (rawImportPath.endsWith('.*')) {
const matchedFiles = resolveJvmWildcard(rawImportPath, ctx.normalizedFileList, ctx.allFileList, ['.java'], ctx.index);
if (matchedFiles.length > 0) return { kind: 'files', files: matchedFiles };
} else {
const memberResolved = resolveJvmMemberImport(rawImportPath, ctx.normalizedFileList, ctx.allFileList, ['.java'], ctx.index);
if (memberResolved) return { kind: 'files', files: [memberResolved] };
}
return resolveStandard(rawImportPath, filePath, ctx, SupportedLanguages.Java);
}
/**
* Kotlin: JVM wildcard/member with Java-interop fallback -> top-level function imports -> standard.
* Kotlin can import from .kt/.kts files OR from .java files (Java interop).
*/
export function resolveKotlinImport(
rawImportPath: string,
filePath: string,
ctx: ResolveCtx,
): ImportResult {
if (rawImportPath.endsWith('.*')) {
const matchedFiles = resolveJvmWildcard(rawImportPath, ctx.normalizedFileList, ctx.allFileList, KOTLIN_EXTENSIONS, ctx.index);
if (matchedFiles.length === 0) {
const javaMatches = resolveJvmWildcard(rawImportPath, ctx.normalizedFileList, ctx.allFileList, ['.java'], ctx.index);
if (javaMatches.length > 0) return { kind: 'files', files: javaMatches };
}
if (matchedFiles.length > 0) return { kind: 'files', files: matchedFiles };
} else {
let memberResolved = resolveJvmMemberImport(rawImportPath, ctx.normalizedFileList, ctx.allFileList, KOTLIN_EXTENSIONS, ctx.index);
if (!memberResolved) {
memberResolved = resolveJvmMemberImport(rawImportPath, ctx.normalizedFileList, ctx.allFileList, ['.java'], ctx.index);
}
if (memberResolved) return { kind: 'files', files: [memberResolved] };
// Kotlin: top-level function imports (e.g. import models.getUser) have only 2 segments,
// which resolveJvmMemberImport skips (requires >=3). Fall back to package-directory scan
// for lowercase last segments (function/property imports). Uppercase last segments
// (class imports like models.User) fall through to standard suffix resolution.
const segments = rawImportPath.split('.');
const lastSeg = segments[segments.length - 1];
if (segments.length >= 2 && lastSeg[0] && lastSeg[0] === lastSeg[0].toLowerCase()) {
const pkgWildcard = segments.slice(0, -1).join('.') + '.*';
let dirFiles = resolveJvmWildcard(pkgWildcard, ctx.normalizedFileList, ctx.allFileList, KOTLIN_EXTENSIONS, ctx.index);
if (dirFiles.length === 0) {
dirFiles = resolveJvmWildcard(pkgWildcard, ctx.normalizedFileList, ctx.allFileList, ['.java'], ctx.index);
}
if (dirFiles.length > 0) return { kind: 'files', files: dirFiles };
}
}
return resolveStandard(rawImportPath, filePath, ctx, SupportedLanguages.Kotlin);
}
@@ -5,15 +5,8 @@
import type { SuffixIndex } from './utils.js';
import { suffixResolve } from './utils.js';
/** PHP Composer PSR-4 autoload config */
export interface ComposerConfig {
/** Map of namespace prefix -> directory (e.g., "App\\" -> "app/") */
psr4: Map<string, string>;
/** PSR-4 entries sorted by namespace length descending (longest match wins).
* Cached once at config load time to avoid re-sorting on every import. */
psr4Sorted?: readonly [string, string][];
}
import type { ImportResult, ResolveCtx } from './types.js';
import type { ComposerConfig } from '../language-config.js';
/** Get or compute the sorted PSR-4 entries (cached after first call). */
function getSortedPsr4(config: ComposerConfig): readonly [string, string][] {
@@ -25,7 +18,7 @@ function getSortedPsr4(config: ComposerConfig): readonly [string, string][] {
}
/**
* Resolve a PHP use-statement import path using PSR-4 mappings.
* Resolve a PHP use-statement import path using PSR-4 mappings (low-level helper).
* e.g. "App\Http\Controllers\UserController" -> "app/Http/Controllers/UserController.php"
*
* For function/constant imports (use function App\Models\getUser), the last
@@ -39,7 +32,7 @@ function getSortedPsr4(config: ComposerConfig): readonly [string, string][] {
* a known limitation — PHP function imports cannot be resolved to a specific file
* without parsing all candidate files.
*/
export function resolvePhpImport(
export function resolvePhpImportInternal(
importPath: string,
composerConfig: ComposerConfig | null,
allFiles: Set<string>,
@@ -96,3 +89,13 @@ export function resolvePhpImport(
const pathParts = normalized.split('/').filter(Boolean);
return suffixResolve(pathParts, normalizedFileList, allFileList, index);
}
/** PHP: namespace-based resolution via composer.json PSR-4. */
export function resolvePhpImport(
rawImportPath: string,
_filePath: string,
ctx: ResolveCtx,
): ImportResult {
const resolved = resolvePhpImportInternal(rawImportPath, ctx.configs.composerConfig, ctx.allFilePaths, ctx.normalizedFileList, ctx.allFileList, ctx.index);
return resolved ? { kind: 'files', files: [resolved] } : null;
}
@@ -4,9 +4,12 @@
*/
import { tryResolveWithExtensions } from './utils.js';
import { SupportedLanguages } from '../../../config/supported-languages.js';
import type { ImportResult, ResolveCtx } from './types.js';
import { resolveStandard } from './standard.js';
/**
* Resolve a Python import to a file path.
* Resolve a Python import to a file path (low-level helper).
*
* 1. Relative (PEP 328): `.module`, `..module` — 1 dot = current package, each extra dot goes up one level.
* 2. Proximity bare import: static heuristic — checks the importer's own directory first.
@@ -19,7 +22,7 @@ import { tryResolveWithExtensions } from './utils.js';
*
* Returns null to let the caller fall through to suffixResolve.
*/
export function resolvePythonImport(
export function resolvePythonImportInternal(
currentFile: string,
importPath: string,
allFiles: Set<string>,
@@ -57,3 +60,18 @@ export function resolvePythonImport(
return null;
}
/**
* Python: relative imports (PEP 328) + proximity-based bare imports.
* Falls through to standard suffix resolution when proximity finds no match.
*/
export function resolvePythonImport(
rawImportPath: string,
filePath: string,
ctx: ResolveCtx,
): ImportResult {
const resolved = resolvePythonImportInternal(filePath, rawImportPath, ctx.allFilePaths);
if (resolved) return { kind: 'files', files: [resolved] };
if (rawImportPath.startsWith('.')) return null; // relative but unresolved -- don't suffix-match
return resolveStandard(rawImportPath, filePath, ctx, SupportedLanguages.Python);
}
@@ -5,14 +5,15 @@
import type { SuffixIndex } from './utils.js';
import { suffixResolve } from './utils.js';
import type { ImportResult, ResolveCtx } from './types.js';
/**
* Resolve a Ruby require/require_relative path to a matching .rb file.
* Resolve a Ruby require/require_relative path to a matching .rb file (low-level helper).
*
* require_relative paths are pre-normalized to './' prefix by the caller.
* require paths use suffix matching (gem-style paths like 'json', 'net/http').
*/
export function resolveRubyImport(
export function resolveRubyImportInternal(
importPath: string,
normalizedFileList: string[],
allFileList: string[],
@@ -21,3 +22,13 @@ export function resolveRubyImport(
const pathParts = importPath.replace(/^\.\//, '').split('/').filter(Boolean);
return suffixResolve(pathParts, normalizedFileList, allFileList, index);
}
/** Ruby: require / require_relative. */
export function resolveRubyImport(
rawImportPath: string,
_filePath: string,
ctx: ResolveCtx,
): ImportResult {
const resolved = resolveRubyImportInternal(rawImportPath, ctx.normalizedFileList, ctx.allFileList, ctx.index);
return resolved ? { kind: 'files', files: [resolved] } : null;
}

Some files were not shown because too many files have changed in this diff Show More