Compare commits

...
Author SHA1 Message Date
copilot-swe-agent[bot]andmagyargergo 1d7782e4e1 fix(cobol): single-quote CALL/COPY, sequence number stripping, PERFORM keyword filtering
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/afed41bd-98b4-48e8-a9e5-ebf0b97deaa8
2026-03-24 17:10:16 +00:00
copilot-swe-agent[bot]andmagyargergo 7ce1371ad8 Initial plan: review COBOL processor completeness
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/afed41bd-98b4-48e8-a9e5-ebf0b97deaa8
2026-03-24 16:59:25 +00:00
copilot-swe-agent[bot] ae437dc30e Initial plan 2026-03-24 16:48:12 +00:00
Gergo Magyar 832789f288 test(cobol): exhaustive 57-test suite with strict exact assertions
Complete rewrite of COBOL integration tests using ground-truth approach:
dump the full graph, then assert EVERY node and EVERY edge.

57 tests across 9 sections:
- Node completeness: Module(3), Function(13), Namespace(2), Property(21),
  Record(1), CodeElement(8), Constructor(1) — exact sorted arrays
- Edge completeness: 22 tests covering every type+reason combination
  with exact source→target pairs
- Cross-program resolution: 6 tests verifying CALL, CICS LINK/XCTL, JCL
- COPY expansion: copybook data items in RPTGEN
- Section hierarchy: exact paragraph membership per section
- Data item ownership: exact per-module breakdown
- MOVE data flow: exact read/write pairs
- JCL integration: job/step/dataset containment
- Grand totals: CALLS(22), CONTAINS(48), IMPORTS(1), ACCESSES(7)

Fixture enhancements:
- CUSTUPDT.cbl: added INIT-SECTION + PROCESSING-SECTION, PERFORM THRU
- AUDITLOG.cbl: added ENTRY "AUDITLOG-BATCH"
- RPTGEN.cbl: added EXEC CICS XCTL

Zero fuzzy assertions — every expect uses toBe(N) or toEqual([...sorted]).
2026-03-24 16:10:29 +00:00
Gergo Magyar 41b0d8dfad test(cobol): add 26 integration tests with exact assertions + fix CICS resolution bug
Integration tests (test/integration/resolvers/cobol.test.ts):
- 26 tests covering full COBOL system extraction
- ALL assertions use exact toBe(N) — zero fuzzy assertions
- Fixtures: CUSTUPDT.cbl, AUDITLOG.cbl, CUSTDAT.cpy, RPTGEN.cbl, RUNJOBS.jcl

Bug fix (cobol-processor.ts):
- CICS LINK/XCTL cross-program resolution was broken — edges were
  created with "resolved" reason but pointing to <unresolved> targets
- Fix: use cics-link-unresolved / cics-xctl-unresolved suffix pattern
  matching the existing cobol-call-unresolved pattern
- Second-pass resolver now patches both CALL and CICS unresolved edges

All 3915 tests pass, 0 failures.
2026-03-24 15:41:00 +00:00
Gergo Magyar 9760f966cb feat(cobol): enrich graph with EXEC SQL/CICS, ENTRY points, MOVE data flow, PERFORM THRU
Maps the remaining 60% of CobolRegexResults to the graph:
- EXEC SQL blocks → CodeElement nodes + ACCESSES edges to DB tables
- EXEC CICS LINK/XCTL → CodeElement nodes + cross-program CALLS edges
- ENTRY points → Constructor nodes (registered for cross-program resolution)
- MOVE statements → ACCESSES edges (read/write data flow tracking)
- PERFORM THRU → expanded CALLS edges for range targets
- File declarations → Record nodes with assignment metadata
- Cross-program CALL 2nd pass: resolves unresolved targets after all programs processed
2026-03-24 15:20:26 +00:00
Gergo Magyar 88c89c42e6 docs: document custom processor pattern in pipeline.ts
Add comment block at the custom processor integration point
documenting the pattern for future non-tree-sitter language additions.
2026-03-24 14:59:09 +00:00
Gergo Magyar 4af677e637 feat: add COBOL language support with regex extraction pipeline
Standalone COBOL processor following the markdown-processor.ts pattern:
- No LanguageProvider modification — COBOL uses regex, not tree-sitter
- No SupportedLanguages enum change — standalone processor pattern

New files:
- cobol-processor.ts — orchestrator (processCobol, isCobolFile, isJclFile)
- cobol/cobol-preprocessor.ts — regex state machine extraction (~888 LOC)
- cobol/cobol-copy-expander.ts — COPY statement expansion with circular detection
- cobol/jcl-parser.ts — JCL job/step/DD extraction
- cobol/jcl-processor.ts — JCL graph node creation

Extraction produces:
- Module nodes (PROGRAM-ID)
- Function nodes (paragraphs)
- Namespace nodes (sections)
- Property nodes (data items)
- CALLS edges (PERFORM intra-file, CALL cross-program)
- IMPORTS edges (COPY statements)
- CONTAINS edges (section → paragraph hierarchy)

Pipeline integration: single processCobol() call in Phase 2.6

54 new tests (33 COBOL + 21 JCL), all 3889 tests pass.
2026-03-24 14:39:25 +00:00
22 changed files with 5738 additions and 10 deletions
+100
View File
@@ -0,0 +1,100 @@
# COBOL Code Indexing
GitNexus indexes COBOL codebases using a **regex-only extraction** strategy, bypassing tree-sitter entirely. This document explains why, how the pipeline works, and links to detailed sub-documents.
## Why Regex-Only?
The tree-sitter-cobol grammar (v0.0.1) has three critical limitations that make it unusable for production indexing:
| Issue | Impact | Severity |
|-------|--------|----------|
| External scanner hangs on ~5% of files | No timeout mechanism exists for the C scanner; the process blocks indefinitely | **Blocking** |
| Only ~15% of paragraph headers detected | Most procedure-division paragraphs are invisible to the grammar | High |
| Patch markers in cols 1-6 cause parse errors | Enterprise COBOL uses non-standard sequence area content (e.g., `mzADD`, `estero`, `#FIX`) | High |
Because the external scanner hang cannot be interrupted (there is no `setTimeoutMicros` equivalent for tree-sitter), using tree-sitter-cobol would hang the indexing pipeline on a non-trivial fraction of real-world files.
The regex-only approach provides:
- **Speed**: ~1ms per file average extraction time
- **Reliability**: zero hangs, zero crashes across 13,000+ files
- **Coverage**: captures all critical symbols -- program name, paragraphs, sections, CALL, PERFORM, COPY, data items (01-77, 88-level), file declarations, FD entries, EXEC SQL/CICS blocks, ENTRY points, and MOVE statements
## Architecture
```mermaid
flowchart TD
A[Repository Scan] --> B{File Detection}
B -->|Extension match| C[COBOL file]
B -->|GITNEXUS_COBOL_DIRS match| C
B -->|No match| Z[Skip]
C --> D{Copybook?}
D -->|Yes| E[Add to Copybook Map]
D -->|No| F[Source Program]
E --> G[COPY Expansion Engine]
F --> G
G -->|Inline copybook content| H[Expanded Source]
H --> I[Patch Marker Cleanup]
I --> J[Regex State Machine]
J --> K[Extracted Symbols]
K --> L[Graph Model Builder]
L --> M[Knowledge Graph]
subgraph "Per-Chunk Processing"
G
H
I
J
K
L
end
subgraph "Post-Processing"
M --> N[Community Detection]
M --> O[Process Detection]
M --> P[Contract Detection]
end
style J fill:#e8f5e9,stroke:#2e7d32
style G fill:#e3f2fd,stroke:#1565c0
```
## COBOL vs Tree-Sitter Languages
| Feature | COBOL (Regex) | Tree-Sitter Languages |
|---------|--------------|----------------------|
| Parser | Single-pass regex state machine | tree-sitter grammar + queries |
| Speed | ~1ms/file | ~5ms/file |
| AST available | No | Yes |
| COPY expansion | Yes (pre-processing step) | N/A |
| Deep indexing | Data items, SQL, CICS, FD, ENTRY | Type annotations, generics, etc. |
| Call extraction | PERFORM (intra-file) + CALL (cross-program) | AST-based call site detection |
| Import extraction | COPY statements | `import`/`require`/`use`/`#include` |
| Coverage | All critical symbols | Language-dependent query coverage |
| Failure mode | Never hangs | External scanner can hang (COBOL only) |
## Sub-Documents
| Document | Description |
|----------|-------------|
| [File Detection](./file-detection.md) | Extension mapping, `GITNEXUS_COBOL_DIRS`, copybook classification |
| [COPY Expansion](./copy-expansion.md) | Copybook inlining, REPLACING transformations, cycle detection |
| [Regex Extraction](./regex-extraction.md) | State machine, regex patterns, line processing |
| [Deep Indexing](./deep-indexing.md) | Data items, EXEC SQL/CICS, file declarations, FD, ENTRY, MOVE |
| [Graph Model](./graph-model.md) | COBOL-specific node types, edge types, full annotated example |
| [Performance](./performance.md) | Benchmarks, worker pool tuning, caps, troubleshooting |
## Key Source Files
| File | Purpose |
|------|---------|
| `gitnexus/src/core/ingestion/cobol-preprocessor.ts` | Patch marker cleanup + regex extraction engine |
| `gitnexus/src/core/ingestion/cobol-copy-expander.ts` | COPY statement expansion with REPLACING |
| `gitnexus/src/core/ingestion/utils.ts` | `getLanguageFromPath`, `getLanguageFromFilename` |
| `gitnexus/src/core/ingestion/pipeline.ts` | `isCobolCopybook`, `expandCobolCopies`, `detectCrossProgamContracts` |
| `gitnexus/src/core/ingestion/workers/parse-worker.ts` | `processCobolRegexOnly` -- graph model builder |
| `gitnexus/src/core/ingestion/workers/worker-pool.ts` | Configurable sub-batch size for COBOL |
+145
View File
@@ -0,0 +1,145 @@
# COBOL COPY Expansion
The COPY statement is COBOL's include mechanism -- analogous to `#include` in C or `import` in modern languages. GitNexus expands COPY statements **before** regex extraction so that symbols defined inside copybooks (data items, paragraphs, etc.) are visible in the program's extracted graph.
## Supported Syntax
### Basic COPY
```cobol
COPY CPSESP.
COPY "WORKGRID.CPY".
```
Inlines the content of the named copybook, replacing the COPY line(s).
### COPY with REPLACING
```cobol
COPY CPSESP REPLACING "ANAZI-KEY" BY "LK-KEY".
COPY CPSESP REPLACING LEADING "ESP-" BY "LK-ESP-"
LEADING "KPSESPL" BY "LK-KPSESPL".
COPY LINKAGE REPLACING TRAILING "-IN" BY "-OUT".
```
Three REPLACING types are supported:
| Type | Syntax | Behavior | Example |
| ------------ | ------------------------------------ | --------------------------------------- | -------------------------------- |
| **EXACT** | `REPLACING "OLD" BY "NEW"` | Replace exact identifier matches | `ANAZI-KEY` becomes `LK-KEY` |
| **LEADING** | `REPLACING LEADING "PFX-" BY "NEW-"` | Replace prefix on all COBOL identifiers | `ESP-NAME` becomes `LK-ESP-NAME` |
| **TRAILING** | `REPLACING TRAILING "-IN" BY "-OUT"` | Replace suffix on all COBOL identifiers | `DATA-IN` becomes `DATA-OUT` |
Multiple REPLACING clauses can appear in a single COPY statement. They are applied in order to each COBOL identifier in the copybook content.
### Multi-Line COPY
COPY statements can span multiple lines (standard COBOL continuation rules apply):
```cobol
COPY CPSESP REPLACING
- LEADING "ESP-" BY "LK-ESP-"
- LEADING "KPSESPL" BY "LK-KPSESPL".
```
Continuation lines (indicator `-` in column 7) are merged before COPY statement scanning.
## Expansion Flow
```mermaid
sequenceDiagram
participant Pipeline
participant Expander as COPY Expander
participant Resolver
participant Reader
Pipeline->>Pipeline: Identify all COBOL files
Pipeline->>Pipeline: Classify copybooks vs programs
Pipeline->>Reader: Read all copybook content upfront
Reader-->>Pipeline: Copybook content map (name -> content)
loop For each source file in chunk
Pipeline->>Expander: expandCopies(content, filePath, resolveFile, readFile)
Expander->>Expander: Merge continuation lines
Expander->>Expander: Detect COPY statements via regex
loop For each COPY statement (reverse order)
Expander->>Resolver: resolveFile(copyTarget)
Resolver-->>Expander: Copybook key or null
alt Resolved successfully
Expander->>Reader: readFile(resolvedKey)
Reader-->>Expander: Copybook content
Expander->>Expander: Apply REPLACING transformations
Expander->>Expander: Recurse for nested COPYs (depth + 1)
Expander->>Expander: Splice expanded content into output
else Not resolved
Expander->>Expander: Keep original COPY line
end
end
Expander-->>Pipeline: Expanded content + resolution metadata
Pipeline->>Pipeline: Replace file content with expanded content
end
```
## Cycle Detection
Circular COPY references (e.g., copybook A includes copybook B which includes copybook A) are detected and handled:
1. Each expansion chain maintains a `visited` set of resolved copybook paths
2. If a copybook path is already in the visited set, the expansion is skipped
3. A `warnedCircular` set (shared across all files in a chunk) deduplicates warning messages
Known circular copybooks in PROJECT-NAME: `ANAZI`, `ANDIP`, `QDIPE` (self-referential includes).
## Max Depth
Nested COPY expansion is limited to **10 levels** (`DEFAULT_MAX_DEPTH`). If a COPY chain exceeds this depth, a warning is logged and the remaining COPY statements are left unexpanded.
## REPLACING Application Detail
The REPLACING engine works by scanning all COBOL identifiers (matching `\b[A-Z][A-Z0-9-]*\b`) in the copybook content and applying each replacement rule:
```
Original copybook content:
05 ESP-NAME PIC X(30).
05 ESP-CODE PIC X(10).
05 KPSESPL-FLAG PIC X(01).
After REPLACING LEADING "ESP-" BY "LK-ESP-" LEADING "KPSESPL" BY "LK-KPSESPL":
05 LK-ESP-NAME PIC X(30).
05 LK-ESP-CODE PIC X(10).
05 LK-KPSESPL-FLAG PIC X(01).
```
For LEADING replacements, the engine checks if each identifier starts with the `from` prefix (case-insensitive) and replaces only the prefix portion, preserving the rest of the identifier.
For TRAILING replacements, the same logic applies to suffixes.
For EXACT replacements, only identifiers that match the `from` value exactly (case-insensitive) are replaced.
## Copybook Resolution
The resolver tries multiple strategies to match a COPY target name to a copybook file:
1. **Exact match**: `COPY CPSESP` resolves to copybook named `CPSESP`
2. **Strip extension**: `COPY WORKGRID.CPY` strips `.CPY` and resolves to `WORKGRID`
3. **Add extension**: `COPY CPSESP` tries `CPSESP.CPY` and `CPSESP.COPY`
If no match is found, the COPY statement is left in place (unexpanded) and a resolution record with `resolvedPath: null` is created.
## Pipeline Integration
The expansion runs **per chunk**, after file content is read but before dispatch to worker threads:
1. All copybook files are read upfront (they are typically small, collectively under 100MB)
2. Per chunk, the copybook map is merged with chunk content (in case a chunk contains copybooks)
3. Only programs (not copybooks themselves) undergo expansion
4. The expanded content replaces the original content in-place before worker dispatch
## Source Files
- `gitnexus/src/core/ingestion/cobol-copy-expander.ts` -- `expandCopies()`, `parseReplacingClause()`, `applyReplacing()`
- `gitnexus/src/core/ingestion/pipeline.ts` -- `expandCobolCopies()`, copybook map construction, chunk integration
+265
View File
@@ -0,0 +1,265 @@
# COBOL Deep Indexing
Beyond basic symbol extraction (program name, paragraphs, CALL, PERFORM, COPY), GitNexus performs deep indexing of COBOL-specific constructs: data items, EXEC SQL/CICS blocks, file declarations, FD entries, ENTRY points, and MOVE statements.
## Data Items
### Level Numbers
| Level Range | Meaning | Graph Node Type |
|-------------|---------|-----------------|
| 01 | Record (group item) | `Record` |
| 02-49 | Elementary/group items | `Property` |
| 66 | RENAMES | `Property` |
| 77 | Independent item | `Property` |
| 88 | Condition name | `Const` |
FILLER items are skipped (no useful name for the graph).
### Clauses Parsed
The `parseDataItemClauses()` function extracts these clauses from the trailing text of a data item declaration:
| Clause | Pattern | Example |
|--------|---------|---------|
| `PIC` / `PICTURE` | `\bPIC(?:TURE)?\s+(?:IS\s+)?(\S+)` | `PIC X(30)`, `PICTURE IS 9(5)V99` |
| `USAGE` | `\bUSAGE\s+(?:IS\s+)?(COMP\|BINARY\|...)` | `USAGE IS COMP-3`, `BINARY` |
| `REDEFINES` | `\bREDEFINES\s+([A-Z][A-Z0-9-]+)` | `REDEFINES WK-DATE-NUM` |
| `OCCURS` | `\bOCCURS\s+(\d+)` | `OCCURS 12 TIMES` |
Standalone COMP variants (without the `USAGE` keyword) are also detected: `COMP`, `COMP-1` through `COMP-6`, `COMP-X`, `BINARY`, `PACKED-DECIMAL`.
### Data Hierarchy
Data items form a hierarchical structure based on level numbers. The extractor uses a **stack algorithm**:
```
Processing order:
01 WK-RECORD -> push {01, WK-RECORD} -> parent: Module
05 WK-NAME -> push {05, WK-NAME} -> parent: WK-RECORD (01 < 05)
10 WK-FIRST -> push {10, WK-FIRST} -> parent: WK-NAME (05 < 10)
10 WK-LAST -> pop WK-FIRST, push -> parent: WK-NAME (05 < 10)
05 WK-CODE -> pop WK-LAST, WK-NAME -> parent: WK-RECORD (01 < 05)
88 WK-ACTIVE -> (88 handled separately) -> parent: WK-CODE
```
The stack maintains items where each entry's level is strictly less than the next. When a new item arrives with a level <= the top of stack, items are popped until the stack top has a smaller level. A `CONTAINS` edge is created from the stack top to the new item.
For 88-level condition names, the parent is the immediately preceding non-88 data item (found by scanning backwards).
### Annotated Example
```cobol
01 WK-EMPLOYEE.
05 WK-EMP-ID PIC 9(6).
05 WK-EMP-NAME PIC X(30).
05 WK-EMP-STATUS PIC X(01).
88 WK-ACTIVE VALUE "A".
88 WK-INACTIVE VALUE "I".
05 WK-SALARY PIC 9(7)V99 COMP-3.
05 WK-DEPT PIC X(04) OCCURS 3 TIMES.
```
Produces:
- `Record` node: `WK-EMPLOYEE` (level 01, section: working-storage)
- `Property` nodes: `WK-EMP-ID`, `WK-EMP-NAME`, `WK-EMP-STATUS`, `WK-SALARY`, `WK-DEPT`
- `Const` nodes: `WK-ACTIVE` (values: `A`), `WK-INACTIVE` (values: `I`)
- `CONTAINS` edges: `WK-EMPLOYEE -> WK-EMP-ID`, `WK-EMPLOYEE -> WK-EMP-NAME`, etc.
- `CONTAINS` edges: `WK-EMP-STATUS -> WK-ACTIVE`, `WK-EMP-STATUS -> WK-INACTIVE`
### Data Item Cap
A maximum of **500 data items per file** (`MAX_DATA_ITEMS_PER_FILE`) are processed. Some COBOL programs (especially after COPY expansion) can have 10,000+ data items, which would cause graph bloat and push the V8 relationship Map past its 16.7M entry limit across thousands of files.
The cap applies after extraction: the first 500 items in source order are kept. Since 01-level records appear first, critical top-level structure is preserved.
## EXEC SQL
EXEC SQL blocks are accumulated across lines between `EXEC SQL` and `END-EXEC`, then parsed as a unit.
### Operation Classification
The first SQL keyword determines the operation:
| First Keyword | Operation |
|---------------|-----------|
| `SELECT` | SELECT |
| `INSERT` | INSERT |
| `UPDATE` | UPDATE |
| `DELETE` | DELETE |
| `DECLARE` | DECLARE |
| `OPEN` | OPEN |
| `CLOSE` | CLOSE |
| `FETCH` | FETCH |
| *(anything else)* | OTHER |
### Table Extraction
Tables are extracted from SQL clauses:
| Clause Pattern | Example |
|----------------|---------|
| `FROM <table>` | `SELECT * FROM EMPLOYEES` |
| `INTO <table>` | `INSERT INTO EMPLOYEES` |
| `UPDATE <table>` | `UPDATE EMPLOYEES SET ...` |
| `JOIN <table>` | `LEFT JOIN DEPARTMENTS ON ...` |
### Cursor Detection
```cobol
EXEC SQL
DECLARE C-EMPLOYEES CURSOR FOR
SELECT EMP-ID, EMP-NAME FROM EMPLOYEES
WHERE DEPT = :WK-DEPT
END-EXEC
```
Extracts: cursor `C-EMPLOYEES`, table `EMPLOYEES`, host variable `WK-DEPT`.
### Host Variables
Host variables are COBOL variables referenced in SQL with a `:` prefix. The colon is stripped:
```sql
WHERE EMP-ID = :WK-EMP-ID AND DEPT = :WK-DEPT
```
Extracts: `WK-EMP-ID`, `WK-DEPT`.
### Graph Output
- `CodeElement` node per table, with description `sql-table op:{OP}`
- `CodeElement` node per cursor, with description `sql-cursor`
- `ACCESSES` edge from Module to each CodeElement
- Deduplication: if the same table appears in multiple SQL blocks, only one node is created
## EXEC CICS
EXEC CICS blocks are accumulated and parsed similarly to SQL blocks.
### Command Detection
Two-word commands are detected first (matched against the block start):
```
SEND MAP, RECEIVE MAP, SEND TEXT, SEND CONTROL, READ NEXT, READ PREV
```
If no two-word command matches, the first word is used (e.g., `LINK`, `XCTL`, `RETURN`, `READ`, `WRITE`).
### Extraction
| Element | Pattern | Example |
|---------|---------|---------|
| MAP name | `MAP('name')` or `MAP("name")` | `EXEC CICS SEND MAP('EMPMENU')` |
| PROGRAM name | `PROGRAM('name')` or `PROGRAM("name")` | `EXEC CICS LINK PROGRAM('BGTABUP')` |
| TRANSID | `TRANSID('name')` or `TRANSID("name")` | `EXEC CICS START TRANSID('EMP1')` |
### Graph Output
- MAP: `CodeElement` node with description `cics-map cmd:{CMD}` + `ACCESSES` edge from Module
- PROGRAM: `CALLS` edge (cross-program call via CICS LINK/XCTL)
- TRANSID: `CodeElement` node with description `cics-transid cmd:{CMD}` + `ACCESSES` edge from Module
### Annotated Example
```cobol
EXEC CICS
SEND MAP('EMPMENU')
MAPSET('EMPSET')
FROM(WK-MAP-DATA)
ERASE
END-EXEC
```
Produces:
- `CodeElement` node: `EMPMENU` (description: `cics-map cmd:SEND MAP`)
- `ACCESSES` edge: Module -> `EMPMENU`
## File Declarations
SELECT statements in the INPUT-OUTPUT SECTION are accumulated across multiple lines (until a period terminator) and parsed for:
| Clause | Pattern | Example |
|--------|---------|---------|
| SELECT | `SELECT <name>` | `SELECT MASTER-FILE` |
| ASSIGN | `ASSIGN TO <file>` | `ASSIGN TO "MASTER.DAT"` |
| ORGANIZATION | `ORGANIZATION IS <type>` | `ORGANIZATION IS INDEXED` |
| ACCESS | `ACCESS MODE IS <mode>` | `ACCESS MODE IS DYNAMIC` |
| RECORD KEY | `RECORD KEY IS <field>` | `RECORD KEY IS WK-EMP-ID` |
| FILE STATUS | `FILE STATUS IS <field>` | `FILE STATUS IS WK-FILE-STATUS` |
### Graph Output
- `CodeElement` node with description containing all parsed clauses (e.g., `select org:INDEXED access:DYNAMIC key:WK-EMP-ID status:WK-FILE-STATUS assign:MASTER.DAT`)
- `RECORD_KEY_OF` edge: from Property node to CodeElement (confidence 0.8)
- `FILE_STATUS_OF` edge: from Property node to CodeElement (confidence 0.8)
## FD Entries
FD (File Description) entries associate a file name with its record layout:
```cobol
FD MASTER-FILE.
01 MASTER-RECORD.
05 MR-EMP-ID PIC 9(6).
05 MR-EMP-NAME PIC X(30).
```
The extractor tracks `pendingFdName` state: when an `FD` line is seen, the next 01-level data item becomes its record.
### Graph Output
- `CodeElement` node with description `fd record:{recordName}`
- `CONTAINS` edge: FD CodeElement -> Record node
- `CONTAINS` edge: SELECT CodeElement -> FD CodeElement (linking file declaration to file description)
## ENTRY Points
The `ENTRY` statement defines additional entry points into a COBOL program (in addition to the main program entry):
```cobol
ENTRY "SUBPROG" USING WK-PARAM-1 WK-PARAM-2.
```
### Graph Output
- `Constructor` node with description `entry params:{param1},{param2}` (or just `entry` if no parameters)
- `CONTAINS` edge: Module -> Constructor
- Symbol table entry (so the entry point is discoverable by name)
## PROCEDURE DIVISION USING
```cobol
PROCEDURE DIVISION USING WK-INPUT-REC WK-OUTPUT-REC.
```
The USING clause identifies parameters received by the program from its caller.
### Graph Output
- `RECEIVES` edge: Module -> Property (for each parameter name, confidence 0.8)
## MOVE Statements
MOVE statements are extracted but currently only stored in the regex results (not emitted as graph edges):
```cobol
MOVE WK-NAME TO OUT-NAME.
MOVE CORRESPONDING WK-INPUT TO WK-OUTPUT.
```
### Extraction Details
- Source and target identifiers are captured
- `CORRESPONDING` keyword is tracked (bulk field-by-field move)
- Figurative constants (SPACES, ZEROS, LOW-VALUES, HIGH-VALUES, QUOTES, ALL) are skipped
- The enclosing paragraph (`caller`) is tracked for context
DATA_FLOW edges from MOVE statements are reserved for a future release.
## Source Files
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- All extraction logic, clause parsers, EXEC block parsers
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `processCobolRegexOnly()`, graph node/edge emission
- `gitnexus/src/core/ingestion/parsing-processor.ts` -- Sequential fallback with same `MAX_DATA_ITEMS_PER_FILE` cap
+126
View File
@@ -0,0 +1,126 @@
# COBOL File Detection
GitNexus detects COBOL files through two mechanisms: extension-based mapping and directory-based override for extensionless files. This document covers both, plus the copybook/program classification logic.
## Extension Mapping
### Program Extensions
| Extension | Type |
|-----------|------|
| `.cbl` | COBOL program |
| `.cob` | COBOL program |
| `.cobol` | COBOL program |
### Copybook Extensions
| Extension | Type | Notes |
|-----------|------|-------|
| `.cpy` | Copybook | Standard |
| `.copy` | Copybook | Standard |
| `.gnm` / `.GNM` | Copybook | Enterprise (GnuCOBOL naming) |
| `.fd` / `.FD` | Copybook | File Description fragment |
| `.wrk` / `.WRK` | Copybook | Working-Storage fragment |
| `.sel` / `.SEL` | Copybook | SELECT clause fragment |
| `.open` / `.OPEN` | Copybook | File OPEN fragment |
| `.close` / `.CLOSE` | Copybook | File CLOSE fragment |
| `.ini` / `.INI` | Copybook | Initialization fragment |
| `.def` / `.DEF` | Copybook | Definition fragment |
All extension matching is case-sensitive in `getLanguageFromFilename` (the extensions above are matched as written, including uppercase variants like `.GNM`).
## Extensionless File Detection: `GITNEXUS_COBOL_DIRS`
Many enterprise COBOL repositories use extensionless files -- the filename alone identifies the program (e.g., `s/BGTABFL` is the source for program `BGTABFL`). GitNexus handles this via the `GITNEXUS_COBOL_DIRS` environment variable.
### Configuration
Set `GITNEXUS_COBOL_DIRS` to a comma-separated list of directory names:
```bash
# Files in s/, c/, and wfproc/ directories (at any depth) are treated as COBOL
export GITNEXUS_COBOL_DIRS=s,c,wfproc
```
The matching is **case-insensitive** and checks all path segments:
- `/repo/s/BGTABFL` -- matches segment `s` -- COBOL
- `/repo/src/c/CPSESP` -- matches segment `c` -- COBOL
- `/repo/wfproc/WF001` -- matches segment `wfproc` -- COBOL
- `/repo/docs/README` -- no matching segment -- skipped
### Decision Tree
```mermaid
flowchart TD
A[getLanguageFromPath] --> B[getLanguageFromFilename]
B --> C{Known extension?}
C -->|Yes .cbl/.cob/.cobol/.cpy/...| D[Return COBOL]
C -->|Yes .ts/.py/.java/...| E[Return other language]
C -->|No match| F{Has extension?}
F -->|"Has dot in basename"| G[Return null]
F -->|"No dot = extensionless"| H{GITNEXUS_COBOL_DIRS set?}
H -->|No| G
H -->|Yes| I{Any path segment<br/>matches a configured dir?}
I -->|Yes| D
I -->|No| G
style D fill:#e8f5e9,stroke:#2e7d32
style G fill:#ffebee,stroke:#c62828
```
### Implementation Detail
The `GITNEXUS_COBOL_DIRS` value is parsed once (on first call) and cached in a `Set<string>`:
```typescript
// From gitnexus/src/core/ingestion/utils.ts
const getCobolDirs = (): Set<string> => {
if (_cobolDirs) return _cobolDirs;
const raw = process.env.GITNEXUS_COBOL_DIRS;
_cobolDirs = raw
? new Set(raw.split(',').map(d => d.trim().toLowerCase()))
: new Set();
return _cobolDirs;
};
```
The path segment check splits the full path on `/` and tests each segment against the cached set.
## Copybook vs Program Classification
After a file is identified as COBOL, it must be classified as either a **program** (to be parsed for symbols) or a **copybook** (to be loaded into the copybook map for COPY expansion).
### Classification Rules
A COBOL file is classified as a **copybook** if ANY of these conditions is true:
1. It has a recognized copybook extension (`.cpy`, `.copy`, `.gnm`, `.fd`, `.wrk`, `.sel`, `.open`, `.close`, `.ini`, `.def`)
2. It is an extensionless file whose path contains a directory segment matching one of: `c`, `copy`, `copybooks`, `copylib`, `cpy`
A file is classified as a **program** if:
1. It has a program extension (`.cbl`, `.cob`, `.cobol`), OR
2. It is extensionless and does NOT match any copybook directory pattern
### Copybook Name Resolution
Copybook names are derived from the filename:
- Strip the extension (if any)
- Convert to uppercase
Examples:
- `c/CPSESP` -- name: `CPSESP`
- `copy/workgrid.cpy` -- name: `WORKGRID`
- `c/ANAZI.GNM` -- name: `ANAZI`
This name is used to resolve `COPY CPSESP.` statements during expansion.
## Source Files
- `gitnexus/src/core/ingestion/utils.ts` -- `getLanguageFromPath()`, `getLanguageFromFilename()`, `getCobolDirs()`
- `gitnexus/src/core/ingestion/pipeline.ts` -- `isCobolCopybook()`, `getCopybookName()`, `COPYBOOK_EXTENSIONS`, `COBOL_PROGRAM_EXTENSIONS`
+193
View File
@@ -0,0 +1,193 @@
# COBOL Graph Model
This document describes the graph nodes and edges that GitNexus creates for COBOL codebases. The COBOL graph model is richer than most tree-sitter languages because it captures domain-specific constructs: file declarations, FD entries, data hierarchies, SQL tables, CICS maps, and cross-program contracts.
## Entity-Relationship Diagram
```mermaid
erDiagram
File ||--o{ Module : DEFINES
File ||--o{ Function : DEFINES
File ||--o{ Namespace : DEFINES
File ||--o{ Record : DEFINES
File ||--o{ Property : DEFINES
File ||--o{ Const : DEFINES
File ||--o{ CodeElement : DEFINES
File ||--o{ Constructor : DEFINES
File }o--o{ File : IMPORTS
Module ||--o{ Record : CONTAINS
Module ||--o{ Constructor : CONTAINS
Module }o--o{ CodeElement : ACCESSES
Module }o--o{ Module : CALLS
Module }o--o{ Module : CONTRACTS
Module }o--o{ Property : RECEIVES
Record ||--o{ Property : CONTAINS
Record ||--o{ Const : CONTAINS
Record }o--o{ Record : REDEFINES
Property ||--o{ Property : CONTAINS
Property ||--o{ Const : CONTAINS
Property }o--o{ Property : REDEFINES
Property }o--o{ CodeElement : RECORD_KEY_OF
Property }o--o{ CodeElement : FILE_STATUS_OF
CodeElement ||--o{ CodeElement : CONTAINS
CodeElement ||--o{ Record : CONTAINS
Function }o--o{ Function : CALLS
```
## Node Types
| Node Type | COBOL Concept | Created From | Example |
|-----------|--------------|--------------|---------|
| `Module` | PROGRAM-ID | `PROGRAM-ID. BGTABFL` | Name: `BGTABFL`, description may include author and date |
| `Function` | Paragraph | `PROCESS-RECORD.` at column 8 | Name: `PROCESS-RECORD` |
| `Namespace` | Procedure section | `MAIN-LOGIC SECTION.` at column 8 | Name: `MAIN-LOGIC` |
| `Record` | 01-level data item | `01 WK-EMPLOYEE.` | Description: `level:01 section:working-storage` |
| `Property` | 02-49/66/77 data item | `05 WK-NAME PIC X(30).` | Description: `level:05 pic:X(30) section:working-storage` |
| `Const` | 88-level condition | `88 WK-ACTIVE VALUE "A".` | Description: `level:88 values:A` |
| `CodeElement` | SELECT, FD, SQL table, CICS map, cursor, transid | Various | Description varies by subtype |
| `Constructor` | ENTRY point | `ENTRY "SUBPROG" USING WK-DATA` | Description: `entry params:WK-DATA` |
### CodeElement Subtypes
CodeElement is used for multiple COBOL constructs, distinguished by their description prefix:
| Subtype | ID Pattern | Description Format | Example |
|---------|-----------|-------------------|---------|
| File SELECT | `CodeElement:{path}:SELECT:{name}` | `select org:INDEXED access:DYNAMIC ...` | `SELECT MASTER-FILE` |
| FD entry | `CodeElement:{path}:FD:{name}` | `fd record:{recordName}` | `FD MASTER-FILE` |
| SQL table | `CodeElement:{path}:sql-table:{name}` | `sql-table op:SELECT` | Table `EMPLOYEES` |
| SQL cursor | `CodeElement:{path}:sql-cursor:{name}` | `sql-cursor` | Cursor `C-EMPLOYEES` |
| CICS map | `CodeElement:{path}:cics-map:{name}` | `cics-map cmd:SEND MAP` | Map `EMPMENU` |
| CICS transid | `CodeElement:{path}:cics-transid:{name}` | `cics-transid cmd:START` | Transid `EMP1` |
## Edge Types
| Edge Type | Source | Target | Created By | Confidence | Example |
|-----------|--------|--------|-----------|------------|---------|
| `DEFINES` | File | any node | File defines its symbols | 1.0 | File -> Module `BGTABFL` |
| `CALLS` | Function | Function | `PERFORM X [THRU Y]` | (via call-processor) | `PROCESS-RECORD` -> `CALC-TAX` |
| `CALLS` | Module | Module | `CALL "BGTABUP"` | (via call-processor) | `BGTABFL` -> `BGTABUP` |
| `CALLS` | Module | Module | `EXEC CICS LINK PROGRAM('X')` | (via call-processor) | `BGTABFL` -> `BGTABUP` |
| `IMPORTS` | File | File | `COPY copybook` | (via import-processor) | Source file -> Copybook file |
| `CONTAINS` | Module | Record | Data hierarchy root | 1.0 | `BGTABFL` -> `WK-EMPLOYEE` |
| `CONTAINS` | Record | Property | Data hierarchy | 1.0 | `WK-EMPLOYEE` -> `WK-NAME` |
| `CONTAINS` | Property | Property | Nested data items | 1.0 | `WK-ADDRESS` -> `WK-CITY` |
| `CONTAINS` | Record/Property | Const | 88-level parent | 1.0 | `WK-STATUS` -> `WK-ACTIVE` |
| `CONTAINS` | CodeElement (FD) | Record | FD record link | 1.0 | `FD:MASTER-FILE` -> `MASTER-RECORD` |
| `CONTAINS` | CodeElement (SELECT) | CodeElement (FD) | SELECT-FD link | 0.9 | `SELECT:MASTER-FILE` -> `FD:MASTER-FILE` |
| `CONTAINS` | Module | Constructor | ENTRY in module | 1.0 | `BGTABFL` -> `SUBPROG` |
| `REDEFINES` | Record | Record | `01 X REDEFINES Y` | 1.0 | `WK-DATE-NUM` -> `WK-DATE-ALPHA` |
| `REDEFINES` | Property | Property | `05 X REDEFINES Y` | 1.0 | `WK-CODE-NUM` -> `WK-CODE-ALPHA` |
| `RECORD_KEY_OF` | Property | CodeElement (SELECT) | `RECORD KEY IS field` | 0.8 | `WK-EMP-ID` -> `SELECT:MASTER-FILE` |
| `FILE_STATUS_OF` | Property | CodeElement (SELECT) | `FILE STATUS IS field` | 0.8 | `WK-FS` -> `SELECT:MASTER-FILE` |
| `ACCESSES` | Module | CodeElement | EXEC SQL/CICS | 0.9 | `BGTABFL` -> `sql-table:EMPLOYEES` |
| `RECEIVES` | Module | Property | `PROCEDURE USING` | 0.8 | `BGTABFL` -> `WK-INPUT-REC` |
| `CONTRACTS` | Module | Module | Shared copybook detection | 0.9 | `BGTABFL` -> `BGTABUP` (via `CPSESP`) |
## Full Annotated Example
Given this COBOL program:
```cobol
IDENTIFICATION DIVISION.
PROGRAM-ID. EMPMAINT.
AUTHOR. Development Team.
ENVIRONMENT DIVISION.
INPUT-OUTPUT SECTION.
FILE-CONTROL.
SELECT EMP-FILE
ASSIGN TO "EMPLOYEE.DAT"
ORGANIZATION IS INDEXED
ACCESS MODE IS DYNAMIC
RECORD KEY IS EMP-ID
FILE STATUS IS WS-FILE-STATUS.
DATA DIVISION.
FILE SECTION.
FD EMP-FILE.
01 EMP-RECORD.
05 EMP-ID PIC 9(6).
05 EMP-NAME PIC X(30).
WORKING-STORAGE SECTION.
01 WS-FLAGS.
05 WS-FILE-STATUS PIC X(02).
05 WS-EOF-FLAG PIC X(01).
88 WS-EOF VALUE "Y".
LINKAGE SECTION.
01 LK-SEARCH-KEY PIC 9(6).
PROCEDURE DIVISION USING LK-SEARCH-KEY.
MAIN-LOGIC SECTION.
MAIN-START.
PERFORM OPEN-FILE
PERFORM PROCESS-RECORDS
PERFORM CLOSE-FILE
STOP RUN.
OPEN-FILE.
OPEN I-O EMP-FILE.
PROCESS-RECORDS.
MOVE LK-SEARCH-KEY TO EMP-ID
EXEC SQL
SELECT EMP_SALARY INTO :WS-SALARY
FROM EMPLOYEES
WHERE EMP_ID = :EMP-ID
END-EXEC
CALL "EMPREPORT".
CLOSE-FILE.
CLOSE EMP-FILE.
```
The graph produced contains:
**Nodes:**
- `Module`: EMPMAINT (description: `author:Development Team`)
- `Namespace`: MAIN-LOGIC
- `Function`: MAIN-START, OPEN-FILE, PROCESS-RECORDS, CLOSE-FILE
- `Record`: EMP-RECORD, WS-FLAGS, LK-SEARCH-KEY
- `Property`: EMP-ID, EMP-NAME, WS-FILE-STATUS, WS-EOF-FLAG
- `Const`: WS-EOF (values: Y)
- `CodeElement`: SELECT:EMP-FILE, FD:EMP-FILE, sql-table:EMPLOYEES
- (COPY imports, if any, would produce File IMPORTS edges)
**Edges:**
- `DEFINES`: File -> all nodes
- `CONTAINS`: EMPMAINT -> EMP-RECORD, EMPMAINT -> WS-FLAGS, EMPMAINT -> LK-SEARCH-KEY
- `CONTAINS`: EMP-RECORD -> EMP-ID, EMP-RECORD -> EMP-NAME
- `CONTAINS`: WS-FLAGS -> WS-FILE-STATUS, WS-FLAGS -> WS-EOF-FLAG
- `CONTAINS`: WS-EOF-FLAG -> WS-EOF
- `CONTAINS`: FD:EMP-FILE -> EMP-RECORD
- `CONTAINS`: SELECT:EMP-FILE -> FD:EMP-FILE
- `CALLS`: MAIN-START -> OPEN-FILE, MAIN-START -> PROCESS-RECORDS, MAIN-START -> CLOSE-FILE
- `CALLS`: EMPMAINT -> EMPREPORT (external CALL)
- `ACCESSES`: EMPMAINT -> sql-table:EMPLOYEES
- `RECEIVES`: EMPMAINT -> LK-SEARCH-KEY (PROCEDURE USING)
- `RECORD_KEY_OF`: EMP-ID -> SELECT:EMP-FILE
- `FILE_STATUS_OF`: WS-FILE-STATUS -> SELECT:EMP-FILE
## How COBOL Differs from Tree-Sitter Languages
| Aspect | COBOL | Tree-Sitter Languages |
|--------|-------|----------------------|
| Node variety | 8 types (Module, Function, Namespace, Record, Property, Const, CodeElement, Constructor) | Typically 4-6 (Function, Class, Method, Interface, Module, Const) |
| Domain edges | RECORD_KEY_OF, FILE_STATUS_OF, ACCESSES, RECEIVES, CONTRACTS, REDEFINES | Primarily CALLS, IMPORTS, EXTENDS, IMPLEMENTS |
| Data hierarchy | Deep CONTAINS chains (01 -> 05 -> 10 -> 88) | Flat class members |
| Cross-program calls | CALL "name" + CICS LINK PROGRAM | Import-based resolution |
| Contract detection | Shared COPY copybook between caller/callee | Not applicable |
| Metadata | AUTHOR, DATE-WRITTEN on Module | JSDoc/docstring (not indexed) |
## Source Files
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `processCobolRegexOnly()`, node/edge emission logic
- `gitnexus/src/core/ingestion/pipeline.ts` -- `detectCrossProgamContracts()` for CONTRACTS edges
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- `CobolRegexResults` interface (all extracted data)
+232
View File
@@ -0,0 +1,232 @@
# COBOL Performance and Tuning
This document covers real-world benchmarks, worker pool configuration, memory management, known limitations, and troubleshooting for COBOL indexing.
## PROJECT-NAME Benchmark
The PROJECT-NAME project is a large Italian payroll system written in COBOL. It serves as the primary benchmark for COBOL indexing performance.
### Input
| Metric | Value |
| --------------------------- | ---------------------------------------------------------------------------- |
| Paths scanned | 14,217 |
| Parseable files | 13,129 |
| Total source size | 224 MB |
| Chunks | 12 (at 20 MB budget) |
| Copybooks loaded | 2,976 |
| Copybooks used in expansion | 2,955 |
| Key directories | `s/` (7773 programs), `c/` (3036 copybooks), `wfproc/` (1973 workflow files) |
### Output
| Metric | Value |
| ---------------------- | ------ |
| Graph nodes | 2.79M |
| Graph edges | 5.67M |
| Clusters (communities) | 16,679 |
| Execution flows | 300 |
### Timing
| Phase | Duration |
| ------------------------------- | ----------------- |
| Total | ~251s |
| KuzuDB write | 132s |
| Full-text search indexing | 6.7s |
| Regex extraction (avg per file) | ~1ms |
| COPY expansion + deep indexing | Remainder (~112s) |
### Indexing Command
```bash
cd /path/to/PROJECT-NAME
GITNEXUS_COBOL_DIRS=s,c,wfproc GITNEXUS_VERBOSE=1 node --max-old-space-size=8192 \
/path/to/gitnexus/dist/cli/index.js analyze --force
```
## Worker Pool Tuning
### Sub-Batch Size
The worker pool splits each worker's chunk into sub-batches to bound peak memory per `postMessage` serialization. COBOL repos use a smaller sub-batch size than the default:
| Parameter | Default | COBOL Mode |
| --------------------- | ----------- | ------------------- |
| Sub-batch size | 1,500 files | 200 files |
| Per sub-batch timeout | 120s | 120s (configurable) |
**Why 200?** COBOL regex extraction + preprocessing takes ~1ms per file on average, but with COPY expansion and deep indexing the effective time is ~150ms per file. At sub-batch size 1500, that would be ~225s per sub-batch, exceeding the 120s timeout.
COBOL mode is activated automatically when `GITNEXUS_COBOL_DIRS` is set:
```typescript
// From pipeline.ts
const cobolSubBatch = process.env.GITNEXUS_COBOL_DIRS ? 200 : undefined;
workerPool = createWorkerPool(workerUrl, undefined, cobolSubBatch);
```
### Worker Count
Workers default to `min(8, cpus - 1)`. For COBOL repos, this is usually sufficient since regex extraction is CPU-bound but fast. The bottleneck is typically KuzuDB write, not extraction.
### Timeout Configuration
| Environment Variable | Default | Purpose |
| ------------------------------------ | --------------- | --------------------------------------------------- |
| `GITNEXUS_WORKER_TIMEOUT_MS` | 120,000 (2 min) | Per sub-batch processing timeout |
| `GITNEXUS_WORKER_STARTUP_TIMEOUT_MS` | 60,000 (1 min) | Worker initialization timeout (tree-sitter loading) |
For COBOL-only repos, worker startup is faster because tree-sitter native modules are loaded lazily (skipped entirely if only COBOL files are present).
## Data Item Cap
### Configuration
```typescript
const MAX_DATA_ITEMS_PER_FILE = 500;
```
This constant appears in both `parse-worker.ts` (worker path) and `parsing-processor.ts` (sequential fallback).
### Rationale
Some COBOL programs, especially after COPY expansion, can have 10,000+ data items. At that scale:
- The in-memory relationship Map (for CONTAINS, REDEFINES, etc.) approaches the V8 16.7M entry limit across thousands of files
- KuzuDB write time increases linearly with edge count
- Most deep-nested items (level 20+) are rarely queried individually
### Impact
The cap truncates data items beyond the 500th in source order. Since 01-level Records appear first in COBOL source, the cap preserves:
- All 01-level record definitions
- The most important 02-49 level items (those closest to the record root)
- 88-level conditions associated with early items
To increase the cap for specific needs, modify the `MAX_DATA_ITEMS_PER_FILE` constant in both files.
## Memory Management
### COPY Expansion Memory
All copybook content is loaded upfront into a Map before chunk processing begins. For PROJECT-NAME:
- 2,976 copybooks, typically under 100MB total
- The Map is shared (read-only) across chunk iterations
- Per-chunk, the copybook map is merged with chunk file content (in case a chunk contains copybooks not in the pre-loaded set)
- After all chunks are processed, the copybook map is freed (`cobolCopybookContents = undefined`)
### Chunk Budget
Source files are grouped into chunks of max 20MB (`CHUNK_BYTE_BUDGET`). Each chunk's lifecycle:
1. Read file content into memory
2. Expand COPY statements (mutates content in-place)
3. Dispatch to workers for extraction
4. Workers return serialized results
5. Merge results into graph
6. Chunk content goes out of scope (GC reclaims)
This ensures only ~20MB of source + ~200-400MB of working memory (ASTs, extracted records, serialization) is active at any time.
### Shared Warning Deduplication
The `warnedCircular` set (used by the COPY expansion engine) is shared across all files in a chunk. This prevents the same circular copybook warning (e.g., `ANAZI includes itself`) from being logged thousands of times.
## Known Limitations
| Limitation | Impact | Workaround |
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------- |
| tree-sitter-cobol hangs on ~5% of files | Cannot use tree-sitter for COBOL | Regex-only extraction (current approach) |
| Data item cap (500/file) | May miss deeply nested items in large programs | Increase `MAX_DATA_ITEMS_PER_FILE` in source |
| Circular copybooks (ANAZI, ANDIP, QDIPE) | Self-referential includes cannot be expanded | Detected and skipped with warning |
| wfproc/ files may not be pure COBOL | Workflow files may produce extraction noise | Exclude `wfproc` from `GITNEXUS_COBOL_DIRS` if problematic |
| No MOVE DATA_FLOW edges yet | Data flow between variables not in graph | Reserved for future release |
| Continuation line handling | Some complex multi-line continuations (especially in string literals spanning 3+ lines) may not merge correctly | Known edge case; affects <0.1% of lines |
| Single-line EXEC blocks | `EXEC SQL SELECT ... END-EXEC` on one line is handled, but pathological nesting is not | Extremely rare in practice |
| Extension case sensitivity | `.GNM` and `.gnm` are matched differently | Use the exact case from the codebase |
## Troubleshooting
### "COPY expansion failed"
```
[pipeline] COPY expansion failed for s/BGTABFL: Cannot read properties of null
```
**Cause:** A copybook referenced by a COPY statement cannot be found.
**Fix:**
1. Verify `GITNEXUS_COBOL_DIRS` includes the directory containing copybooks (typically `c`)
2. Check that copybook filenames match the COPY target (case-insensitive, after stripping extensions)
3. Ensure copybook files are not in `.gitignore`
### Worker sub-batch timeout
```
Worker 3 sub-batch timed out after 120s (chunk: 200 items)
```
**Cause:** A sub-batch took longer than the timeout. Typically happens when one file is extremely large (50,000+ lines after COPY expansion).
**Fix:** Increase the timeout:
```bash
GITNEXUS_WORKER_TIMEOUT_MS=300000 gitnexus analyze
```
### Memory errors (heap out of memory)
```
FATAL ERROR: CALL_AND_RETRY_LAST Allocation failed - JavaScript heap out of memory
```
**Fix:** Increase Node.js heap size:
```bash
node --max-old-space-size=16384 /path/to/gitnexus/dist/cli/index.js analyze
```
For very large repos (>500MB source), consider `--max-old-space-size=32768`.
### Concurrent analyze corruption
**Rule:** Only ONE `gitnexus analyze` process should run at a time per repository. Concurrent writes to KuzuDB corrupt the database.
If corruption occurs:
```bash
# Remove the KuzuDB directory and re-index
rm -rf .gitnexus/kuzu
gitnexus analyze --force
```
### Slow KuzuDB write phase
The KuzuDB write phase (132s for PROJECT-NAME) is the bottleneck for large COBOL repos. This is proportional to the number of nodes and edges being written. Reducing `MAX_DATA_ITEMS_PER_FILE` or excluding non-essential directories from `GITNEXUS_COBOL_DIRS` can help.
### Verbose output
Enable verbose logging to see per-phase timing and statistics:
```bash
GITNEXUS_VERBOSE=1 gitnexus analyze
```
This outputs:
- Scan statistics (paths, parseable files, chunk count)
- Worker pool configuration (worker count, sub-batch size)
- COPY expansion statistics (copybooks loaded, files expanded)
- Community and process detection results
- Contract detection results
## Source Files
- `gitnexus/src/core/ingestion/workers/worker-pool.ts` -- `DEFAULT_SUB_BATCH_SIZE`, `SUB_BATCH_TIMEOUT_MS`, `WORKER_STARTUP_TIMEOUT_MS`
- `gitnexus/src/core/ingestion/pipeline.ts` -- `CHUNK_BYTE_BUDGET`, COBOL sub-batch configuration, chunk lifecycle
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `MAX_DATA_ITEMS_PER_FILE`, `processCobolRegexOnly()`
- `gitnexus/src/core/ingestion/parsing-processor.ts` -- Sequential fallback `MAX_DATA_ITEMS_PER_FILE`
@@ -0,0 +1,186 @@
# COBOL Regex Extraction
The `extractCobolSymbolsWithRegex()` function in `cobol-preprocessor.ts` performs single-pass, state-machine-driven extraction of all COBOL symbols. This document describes the state machine, line processing flow, and every regex pattern used.
## State Machine: Division Tracking
The extractor tracks which COBOL division is currently being processed. Division transitions are detected by the `RE_DIVISION` pattern.
```mermaid
stateDiagram-v2
[*] --> null : Start of file
null --> identification : IDENTIFICATION DIVISION
identification --> environment : ENVIRONMENT DIVISION
environment --> data : DATA DIVISION
data --> procedure : PROCEDURE DIVISION
note right of identification
Extracts: PROGRAM-ID, AUTHOR, DATE-WRITTEN
end note
note right of environment
Extracts: SELECT ... ASSIGN ... (file declarations)
end note
note right of data
Extracts: FD entries, data items (01-77, 88), COPY
end note
note right of procedure
Extracts: paragraphs, sections, PERFORM, CALL,
ENTRY, MOVE, EXEC SQL/CICS
end note
```
## State Machine: Data Section Tracking
Within the DATA DIVISION, a secondary state machine tracks the current section to tag data items with their origin.
```mermaid
stateDiagram-v2
[*] --> unknown : DATA DIVISION entered
unknown --> working_storage : WORKING-STORAGE SECTION
unknown --> linkage : LINKAGE SECTION
unknown --> file : FILE SECTION
unknown --> local_storage : LOCAL-STORAGE SECTION
working_storage --> linkage : LINKAGE SECTION
working_storage --> file : FILE SECTION
linkage --> working_storage : WORKING-STORAGE SECTION
file --> working_storage : WORKING-STORAGE SECTION
file --> linkage : LINKAGE SECTION
local_storage --> working_storage : WORKING-STORAGE SECTION
```
Within the ENVIRONMENT DIVISION, the `currentEnvSection` tracks whether we are in `INPUT-OUTPUT` or `CONFIGURATION` section. SELECT statement accumulation only occurs in `INPUT-OUTPUT`.
## Line Processing Flow
Each raw source line goes through this pipeline:
```
Raw line
|
v
Length < 7? ---------> Skip (flush pending if any)
|
v
Indicator col 7
|
+-- '*' or '/' -----> Comment: skip entirely
|
+-- '-' ------------> Continuation: append to pending line
|
+-- other ----------> Normal: flush pending, strip inline comments (|),
buffer as new pending logical line
```
After all lines are processed, the final pending line is flushed, along with any accumulated SELECT statement.
### Inline Comment Stripping
Enterprise COBOL (particularly Italian dialect) uses the pipe character `|` as an inline comment marker. Everything from `|` to end of line is stripped before processing.
### Patch Marker Handling
The `preprocessCobolSource()` function (run before extraction in the worker) replaces non-standard content in columns 1-6. Standard COBOL expects spaces or digit sequence numbers in this area. If any letter or `#` character is found, the entire sequence area is replaced with 6 spaces:
```
Before: mzADD MOVE WK-AMT TO WK-TOTAL
After: MOVE WK-AMT TO WK-TOTAL
```
This preserves exact line count for position mapping.
## Regex Pattern Reference
All patterns are compiled once as module-level constants and reused across calls.
### Division and Section Detection
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_DIVISION` | `\b(IDENTIFICATION\|ENVIRONMENT\|DATA\|PROCEDURE)\s+DIVISION\b` | Division boundary | `PROCEDURE DIVISION` |
| `RE_SECTION` | `\b(WORKING-STORAGE\|LINKAGE\|FILE\|LOCAL-STORAGE\|INPUT-OUTPUT\|CONFIGURATION)\s+SECTION\b` | Section boundary | `WORKING-STORAGE SECTION` |
### IDENTIFICATION DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_PROGRAM_ID` | `\bPROGRAM-ID\.\s*([A-Z][A-Z0-9-]*)` | Program name | `PROGRAM-ID. BGTABFL` |
| `RE_AUTHOR` | `^\s+AUTHOR\.\s*(.+)` | Author metadata | `AUTHOR. D. Smith` |
| `RE_DATE_WRITTEN` | `^\s+DATE-WRITTEN\.\s*(.+)` | Date metadata | `DATE-WRITTEN. 2024-01-15` |
### ENVIRONMENT DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_SELECT_START` | `\bSELECT\s+([A-Z][A-Z0-9-]+)` | File SELECT start | `SELECT MASTER-FILE` |
SELECT statements are accumulated across multiple lines until a period terminator is found, then parsed for ASSIGN, ORGANIZATION, ACCESS, RECORD KEY, and FILE STATUS clauses.
### DATA DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_FD` | `^\s+FD\s+([A-Z][A-Z0-9-]+)` | File description | `FD MASTER-FILE` |
| `RE_DATA_ITEM` | `^\s+(\d{1,2})\s+([A-Z][A-Z0-9-]+)\s*(.*)` | Data item (01-77) | `05 WK-NAME PIC X(30)` |
| `RE_ANONYMOUS_REDEFINES` | `^\s+(\d{1,2})\s+REDEFINES\s+([A-Z][A-Z0-9-]+)` | Anonymous REDEFINES | `01 REDEFINES WK-REC` |
| `RE_88_LEVEL` | `^\s+88\s+([A-Z][A-Z0-9-]+)\s+VALUES?\s+(?:ARE\s+)?(.+)` | Condition name | `88 WK-ACTIVE VALUE "Y"` |
The trailing clauses of `RE_DATA_ITEM` are parsed by `parseDataItemClauses()` for PIC, USAGE, OCCURS, and REDEFINES.
### PROCEDURE DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_PROC_SECTION` | `^ ([A-Z][A-Z0-9-]+)\s+SECTION\.\s*$` | Procedure section header | ` MAIN-LOGIC SECTION.` |
| `RE_PROC_PARAGRAPH` | `^ ([A-Z][A-Z0-9-]+)\.\s*$` | Paragraph header | ` PROCESS-RECORD.` |
| `RE_PERFORM` | `\bPERFORM\s+([A-Z][A-Z0-9-]+)(?:\s+THRU\s+([A-Z][A-Z0-9-]+))?` | PERFORM call | `PERFORM CALC-TAX THRU CALC-TAX-EXIT` |
| `RE_PROC_USING` | `\bPROCEDURE\s+DIVISION\s+USING\s+([\s\S]*?)(?:\.\|$)` | USING parameters | `PROCEDURE DIVISION USING WK-PARAM` |
| `RE_ENTRY` | `\bENTRY\s+"([^"]+)"(?:\s+USING\s+([\s\S]*?))?(?:\.\|$)` | ENTRY point | `ENTRY "SUBPROG" USING WK-DATA` |
| `RE_MOVE` | `\bMOVE\s+(CORRESPONDING\s+)?([A-Z][A-Z0-9-]+)\s+TO\s+([A-Z][A-Z0-9-]+)` | MOVE statement | `MOVE WK-NAME TO OUT-NAME` |
Note: `RE_PROC_SECTION` and `RE_PROC_PARAGRAPH` require exactly 7 spaces of leading indentation (COBOL area A starting at column 8). This is the standard COBOL paragraph indentation.
### All-Division Patterns
These patterns are checked regardless of current division:
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_CALL` | `\bCALL\s+"([^"]+)"` | External program call | `CALL "BGTABUP"` |
| `RE_COPY_UNQUOTED` | `\bCOPY\s+([A-Z][A-Z0-9-]+)(?:\s\|\.)` | COPY (unquoted) | `COPY CPSESP.` |
| `RE_COPY_QUOTED` | `\bCOPY\s+"([^"]+)"(?:\s\|\.)` | COPY (quoted) | `COPY "WORKGRID.CPY".` |
### EXEC Block Patterns
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_EXEC_SQL_START` | `\bEXEC\s+SQL\b` | Start of EXEC SQL block | `EXEC SQL` |
| `RE_EXEC_CICS_START` | `\bEXEC\s+CICS\b` | Start of EXEC CICS block | `EXEC CICS` |
| `RE_END_EXEC` | `\bEND-EXEC\b` | End of EXEC block | `END-EXEC` |
EXEC blocks accumulate all lines between `EXEC SQL/CICS` and `END-EXEC`, then delegate to `parseExecSqlBlock()` or `parseExecCicsBlock()` for detailed extraction.
## Excluded Paragraph Names
The following names are excluded from paragraph detection to avoid false positives from division/section headers:
```
DECLARATIVES, END, PROCEDURE, IDENTIFICATION,
ENVIRONMENT, DATA, WORKING-STORAGE, LINKAGE,
FILE, LOCAL-STORAGE, COMMUNICATION, REPORT,
SCREEN, INPUT-OUTPUT, CONFIGURATION
```
Additionally, paragraph candidates containing `DIVISION` or `SECTION` as substrings are excluded.
## MOVE Skip List (Figurative Constants)
MOVE statements where the source is a figurative constant are skipped:
```
SPACES, ZEROS, ZEROES, LOW-VALUES, LOW-VALUE,
HIGH-VALUES, HIGH-VALUE, QUOTES, QUOTE, ALL
```
## Source Files
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- `preprocessCobolSource()`, `extractCobolSymbolsWithRegex()`, all regex constants
+3 -10
View File
@@ -1,12 +1,12 @@
{
"name": "gitnexus",
"version": "1.4.7",
"version": "1.4.8",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "gitnexus",
"version": "1.4.7",
"version": "1.4.8",
"hasInstallScript": true,
"license": "PolyForm-Noncommercial-1.0.0",
"dependencies": {
@@ -56,7 +56,7 @@
"vitest": "^4.0.18"
},
"engines": {
"node": ">=18.0.0"
"node": ">=20.0.0"
},
"optionalDependencies": {
"tree-sitter-kotlin": "^0.3.8",
@@ -3130,7 +3130,6 @@
"resolved": "https://registry.npmjs.org/express/-/express-4.22.1.tgz",
"integrity": "sha512-F2X8g9P1X7uCPZMA3MVf9wcTqlyNp7IhH5qPCI0izhaOIYXaW9L535tGA3qmjRzpH+bZczqq7hVKxTR4NWnu+g==",
"license": "MIT",
"peer": true,
"dependencies": {
"accepts": "~1.3.8",
"array-flatten": "1.1.1",
@@ -4185,7 +4184,6 @@
"integrity": "sha512-5gTmgEY/sqK6gFXLIsQNH19lWb4ebPDLA4SdLP7dsWkIXHWlG66oPuVvXSGFPppYZz8ZDZq0dYYrbHfBCVUb1Q==",
"dev": true,
"license": "MIT",
"peer": true,
"engines": {
"node": ">=12"
},
@@ -4965,7 +4963,6 @@
"integrity": "sha512-usbHZP9/oxNsUY65MQUsduGRqDHQOou1cagUSwjhoSYAmSahjQDAVsh9s+SlZkn8X8+O1FULRGwHu7AFP3kjzg==",
"hasInstallScript": true,
"license": "MIT",
"peer": true,
"dependencies": {
"node-addon-api": "^8.3.0",
"node-gyp-build": "^4.8.4"
@@ -5368,7 +5365,6 @@
"integrity": "sha512-5C1sg4USs1lfG0GFb2RLXsdpXqBSEhAaA/0kPL01wxzpMqLILNxIxIOKiILz+cdg/pLnOUxFYOR5yhHU666wbw==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"esbuild": "~0.27.0",
"get-tsconfig": "^4.7.5"
@@ -5489,7 +5485,6 @@
"integrity": "sha512-w+N7Hifpc3gRjZ63vYBXA56dvvRlNWRczTdmCBBa+CotUzAPf5b7YMdMR/8CQoeYE5LX3W4wj6RYTgonm1b9DA==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"esbuild": "^0.27.0",
"fdir": "^6.5.0",
@@ -5565,7 +5560,6 @@
"integrity": "sha512-hOQuK7h0FGKgBAas7v0mSAsnvrIgAvWmRFjmzpJ7SwFHH3g1k2u37JtYwOwmEKhK6ZO3v9ggDBBm0La1LCK4uQ==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"@vitest/expect": "4.0.18",
"@vitest/mocker": "4.0.18",
@@ -5849,7 +5843,6 @@
"resolved": "https://registry.npmjs.org/zod/-/zod-4.3.6.tgz",
"integrity": "sha512-rftlrkhHZOcjDwkGlnUtZZkvaPHCsDATp4pGpuOOMDaTdDDXF91wuVDJoWoPsKX/3YPQ5fHuF3STjcYyKr+Qhg==",
"license": "MIT",
"peer": true,
"funding": {
"url": "https://github.com/sponsors/colinhacks"
}
@@ -0,0 +1,627 @@
/**
* COBOL Processor
*
* Standalone regex-based processor for COBOL and JCL files.
* Follows the markdown-processor.ts pattern: takes (graph, files, allPathSet),
* does its own extraction, and writes directly to the graph.
*
* Pipeline:
* 1. Separate programs from copybooks
* 2. Build copybook map (name -> content)
* 3. For each program: expand COPY statements, then run regex extraction
* 4. Map CobolRegexResults to graph nodes and relationships
* 5. Optionally process JCL files for job-step cross-references
*/
import path from 'node:path';
import { generateId } from '../../lib/utils.js';
import type { KnowledgeGraph, GraphNode } from '../graph/types.js';
import {
preprocessCobolSource,
extractCobolSymbolsWithRegex,
type CobolRegexResults,
} from './cobol/cobol-preprocessor.js';
import { expandCopies } from './cobol/cobol-copy-expander.js';
import { processJclFiles } from './cobol/jcl-processor.js';
// ---------------------------------------------------------------------------
// File detection
// ---------------------------------------------------------------------------
const COBOL_EXTENSIONS = new Set([
'.cob', '.cbl', '.cobol', '.cpy', '.copybook',
]);
const JCL_EXTENSIONS = new Set(['.jcl', '.job', '.proc']);
const COPYBOOK_EXTENSIONS = new Set(['.cpy', '.copybook']);
interface CobolFile {
path: string;
content: string;
}
export interface CobolProcessResult {
programs: number;
paragraphs: number;
sections: number;
dataItems: number;
calls: number;
copies: number;
execSqlBlocks: number;
execCicsBlocks: number;
entryPoints: number;
moves: number;
fileDeclarations: number;
jclJobs: number;
jclSteps: number;
}
/** Returns true if the file is a COBOL or copybook file. */
export function isCobolFile(filePath: string): boolean {
return COBOL_EXTENSIONS.has(path.extname(filePath).toLowerCase());
}
/** Returns true if the file is a JCL file. */
export function isJclFile(filePath: string): boolean {
return JCL_EXTENSIONS.has(path.extname(filePath).toLowerCase());
}
/** Returns true if the file is a COBOL copybook. */
function isCopybook(filePath: string): boolean {
return COPYBOOK_EXTENSIONS.has(path.extname(filePath).toLowerCase());
}
// ---------------------------------------------------------------------------
// Main processor
// ---------------------------------------------------------------------------
/**
* Process COBOL and JCL files into the knowledge graph.
*
* @param graph - The in-memory knowledge graph
* @param files - Array of { path, content } for COBOL/JCL files
* @param allPathSet - Set of all file paths in the repository
* @returns Summary of what was extracted
*/
export const processCobol = (
graph: KnowledgeGraph,
files: CobolFile[],
allPathSet: Set<string>,
): CobolProcessResult => {
const result: CobolProcessResult = {
programs: 0,
paragraphs: 0,
sections: 0,
dataItems: 0,
calls: 0,
copies: 0,
execSqlBlocks: 0,
execCicsBlocks: 0,
entryPoints: 0,
moves: 0,
fileDeclarations: 0,
jclJobs: 0,
jclSteps: 0,
};
// ── 1. Separate programs, copybooks, and JCL ───────────────────────
const programs: CobolFile[] = [];
const copybooks: CobolFile[] = [];
const jclFiles: CobolFile[] = [];
for (const file of files) {
const ext = path.extname(file.path).toLowerCase();
if (JCL_EXTENSIONS.has(ext)) {
jclFiles.push(file);
} else if (isCopybook(file.path)) {
copybooks.push(file);
} else if (COBOL_EXTENSIONS.has(ext)) {
programs.push(file);
}
}
// ── 2. Build copybook map (uppercase name -> content) ──────────────
const copybookMap = new Map<string, { content: string; path: string }>();
for (const cb of copybooks) {
const name = path.basename(cb.path, path.extname(cb.path)).toUpperCase();
copybookMap.set(name, { content: cb.content, path: cb.path });
}
// Resolve and read callbacks for expandCopies
const resolveCopy = (name: string): string | null => {
const entry = copybookMap.get(name.toUpperCase());
return entry ? entry.path : null;
};
const readCopy = (copyPath: string): string | null => {
// Find by path match
for (const [, entry] of copybookMap) {
if (entry.path === copyPath) return entry.content;
}
return null;
};
// Track module names for cross-program CALL resolution
const moduleNodeIds = new Map<string, string>(); // uppercase program name -> node id
// ── 3. Process each COBOL program ──────────────────────────────────
for (const file of programs) {
const fileNodeId = generateId('File', file.path);
// Skip if file node doesn't exist (structure-processor creates it)
if (!graph.getNode(fileNodeId)) continue;
// Preprocess: clean patch markers
const cleaned = preprocessCobolSource(file.content);
// Expand COPY statements
const { expandedContent, copyResolutions } = expandCopies(
cleaned, file.path, resolveCopy, readCopy,
);
// Extract symbols from expanded source
const extracted = extractCobolSymbolsWithRegex(expandedContent, file.path);
// Map to graph
mapToGraph(graph, extracted, file, copyResolutions, moduleNodeIds);
// Accumulate stats
result.programs += extracted.programName ? 1 : 0;
result.paragraphs += extracted.paragraphs.length;
result.sections += extracted.sections.length;
result.dataItems += extracted.dataItems.length;
result.calls += extracted.calls.length;
result.copies += extracted.copies.length;
result.execSqlBlocks += extracted.execSqlBlocks.length;
result.execCicsBlocks += extracted.execCicsBlocks.length;
result.entryPoints += extracted.entryPoints.length;
result.moves += extracted.moves.length;
result.fileDeclarations += extracted.fileDeclarations.length;
}
// ── 4. Second pass: resolve cross-program CALL targets ─────────────
// During mapToGraph, early programs create unresolved CALL edges
// (target = <unresolved>:PROGNAME) because later programs haven't
// been registered in moduleNodeIds yet. Now that ALL programs are
// processed, re-scan unresolved CALLS edges and patch them.
// This covers both `cobol-call-unresolved` and CICS LINK/XCTL edges
// whose targets contain `<unresolved>:`.
graph.forEachRelationship(rel => {
if (rel.type !== 'CALLS') return;
const match = rel.targetId.match(/<unresolved>:(.+)/);
if (!match) return;
const resolvedId = moduleNodeIds.get(match[1]);
if (!resolvedId) return;
if (rel.reason?.startsWith('cobol-call-unresolved')) {
// Replace unresolved CALL with resolved edge
graph.addRelationship({
id: rel.id + ':resolved',
type: 'CALLS',
sourceId: rel.sourceId,
targetId: resolvedId,
confidence: 0.95,
reason: 'cobol-call',
});
} else if (rel.reason === 'cics-link-unresolved' || rel.reason === 'cics-xctl-unresolved') {
// Replace unresolved CICS LINK/XCTL with resolved edge
graph.addRelationship({
id: rel.id + ':resolved',
type: 'CALLS',
sourceId: rel.sourceId,
targetId: resolvedId,
confidence: 0.95,
reason: rel.reason.replace('-unresolved', ''),
});
}
});
// ── 5. Process JCL files ───────────────────────────────────────────
if (jclFiles.length > 0) {
const jclPaths = jclFiles.map(f => f.path);
const jclContents = new Map<string, string>();
for (const f of jclFiles) {
jclContents.set(f.path, f.content);
}
const jclResult = processJclFiles(graph, jclPaths, jclContents);
result.jclJobs += jclResult.jobCount;
result.jclSteps += jclResult.stepCount;
}
return result;
};
// ---------------------------------------------------------------------------
// Graph mapping
// ---------------------------------------------------------------------------
/** Resolve a data item name to its Property node id, if it exists and is not FILLER. */
function findDataItemNode(
name: string,
dataItems: CobolRegexResults['dataItems'],
filePath: string,
): string | undefined {
const item = dataItems.find(d => d.name.toUpperCase() === name.toUpperCase());
if (!item || item.name === 'FILLER') return undefined;
return generateId('Property', `${filePath}:${item.name}`);
}
function mapToGraph(
graph: KnowledgeGraph,
extracted: CobolRegexResults,
file: CobolFile,
copyResolutions: Array<{ copyTarget: string; resolvedPath: string | null; line: number }>,
moduleNodeIds: Map<string, string>,
): void {
const { path: filePath, content } = file;
const lines = content.split('\n');
const fileNodeId = generateId('File', filePath);
// ── PROGRAM-ID -> Module node ────────────────────────────────────
let moduleId: string | undefined;
if (extracted.programName) {
moduleId = generateId('Module', `${filePath}:${extracted.programName}`);
graph.addNode({
id: moduleId,
label: 'Module',
properties: {
name: extracted.programName,
filePath,
startLine: 1,
endLine: lines.length,
language: 'cobol' as any,
isExported: true,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${fileNodeId}->${moduleId}`),
type: 'CONTAINS',
sourceId: fileNodeId,
targetId: moduleId,
confidence: 1.0,
reason: 'cobol-program-id',
});
moduleNodeIds.set(extracted.programName.toUpperCase(), moduleId);
}
const parentId = moduleId ?? fileNodeId;
// ── SECTIONs -> Namespace nodes ──────────────────────────────────
const sectionNodeIds = new Map<string, string>();
for (let i = 0; i < extracted.sections.length; i++) {
const sec = extracted.sections[i];
const nextLine = i + 1 < extracted.sections.length
? extracted.sections[i + 1].line - 1
: lines.length;
const secId = generateId('Namespace', `${filePath}:${sec.name}`);
graph.addNode({
id: secId,
label: 'Namespace',
properties: {
name: sec.name,
filePath,
startLine: sec.line,
endLine: nextLine,
language: 'cobol' as any,
isExported: true,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${parentId}->${secId}`),
type: 'CONTAINS',
sourceId: parentId,
targetId: secId,
confidence: 1.0,
reason: 'cobol-section',
});
sectionNodeIds.set(sec.name.toUpperCase(), secId);
}
// ── PARAGRAPHs -> Function nodes ─────────────────────────────────
const paraNodeIds = new Map<string, string>();
for (let i = 0; i < extracted.paragraphs.length; i++) {
const para = extracted.paragraphs[i];
const nextLine = i + 1 < extracted.paragraphs.length
? extracted.paragraphs[i + 1].line - 1
: lines.length;
const paraId = generateId('Function', `${filePath}:${para.name}`);
graph.addNode({
id: paraId,
label: 'Function',
properties: {
name: para.name,
filePath,
startLine: para.line,
endLine: nextLine,
language: 'cobol' as any,
isExported: true,
},
});
// Parent: find the containing section, or fall back to module/file
const containerId = findContainingSection(para.line, extracted.sections, sectionNodeIds) ?? parentId;
graph.addRelationship({
id: generateId('CONTAINS', `${containerId}->${paraId}`),
type: 'CONTAINS',
sourceId: containerId,
targetId: paraId,
confidence: 1.0,
reason: 'cobol-paragraph',
});
paraNodeIds.set(para.name.toUpperCase(), paraId);
}
// ── Data items -> Property nodes ─────────────────────────────────
for (const item of extracted.dataItems) {
if (item.name === 'FILLER') continue; // Skip anonymous fillers
const propId = generateId('Property', `${filePath}:${item.name}`);
graph.addNode({
id: propId,
label: 'Property',
properties: {
name: item.name,
filePath,
startLine: item.line,
endLine: item.line,
language: 'cobol' as any,
description: `level:${item.level} section:${item.section}${item.pic ? ` pic:${item.pic}` : ''}`,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${parentId}->${propId}`),
type: 'CONTAINS',
sourceId: parentId,
targetId: propId,
confidence: 1.0,
reason: 'cobol-data-item',
});
}
// ── PERFORM -> CALLS relationship (intra-file) ──────────────────
for (const perf of extracted.performs) {
const targetId = paraNodeIds.get(perf.target.toUpperCase())
?? sectionNodeIds.get(perf.target.toUpperCase());
if (!targetId) continue;
// Source: the paragraph containing the PERFORM, or the module
const sourceId = perf.caller
? (paraNodeIds.get(perf.caller.toUpperCase()) ?? parentId)
: parentId;
graph.addRelationship({
id: generateId('CALLS', `${sourceId}->perform->${targetId}:L${perf.line}`),
type: 'CALLS',
sourceId,
targetId,
confidence: 1.0,
reason: 'cobol-perform',
});
// PERFORM THRU -> expanded CALLS edge to thru target
if (perf.thruTarget) {
const thruTargetId = paraNodeIds.get(perf.thruTarget.toUpperCase())
?? sectionNodeIds.get(perf.thruTarget.toUpperCase());
if (thruTargetId && thruTargetId !== targetId) {
graph.addRelationship({
id: generateId('CALLS', `${sourceId}->perform-thru->${thruTargetId}:L${perf.line}`),
type: 'CALLS',
sourceId,
targetId: thruTargetId,
confidence: 1.0,
reason: 'cobol-perform-thru',
});
}
}
}
// ── CALL -> CALLS relationship (cross-program) ──────────────────
for (const call of extracted.calls) {
const targetModuleId = moduleNodeIds.get(call.target.toUpperCase());
// Create edge even if target not yet known — use a synthetic target id
const targetId = targetModuleId
?? generateId('Module', `<unresolved>:${call.target.toUpperCase()}`);
graph.addRelationship({
id: generateId('CALLS', `${parentId}->call->${call.target}:L${call.line}`),
type: 'CALLS',
sourceId: parentId,
targetId,
confidence: targetModuleId ? 0.95 : 0.5,
reason: targetModuleId ? 'cobol-call' : 'cobol-call-unresolved',
});
}
// ── COPY -> IMPORTS relationship ─────────────────────────────────
for (const res of copyResolutions) {
if (!res.resolvedPath) continue;
const targetFileId = generateId('File', res.resolvedPath);
graph.addRelationship({
id: generateId('IMPORTS', `${fileNodeId}->${targetFileId}:${res.copyTarget}`),
type: 'IMPORTS',
sourceId: fileNodeId,
targetId: targetFileId,
confidence: 1.0,
reason: 'cobol-copy',
});
}
// ── EXEC SQL blocks -> CodeElement nodes + ACCESSES edges ──────
for (const sql of extracted.execSqlBlocks) {
const sqlId = generateId('CodeElement', `${filePath}:exec-sql:L${sql.line}`);
graph.addNode({
id: sqlId,
label: 'CodeElement',
properties: {
name: `EXEC SQL ${sql.operation}`,
filePath,
startLine: sql.line,
endLine: sql.line,
language: 'cobol' as any,
description: `tables:[${sql.tables.join(',')}] cursors:[${sql.cursors.join(',')}]`,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${parentId}->${sqlId}`),
type: 'CONTAINS',
sourceId: parentId,
targetId: sqlId,
confidence: 1.0,
reason: 'cobol-exec-sql',
});
// ACCESSES edges to tables
for (const table of sql.tables) {
const tableId = generateId('Record', `<db>:${table}`);
graph.addRelationship({
id: generateId('ACCESSES', `${sqlId}->${tableId}:${sql.operation}`),
type: 'ACCESSES',
sourceId: sqlId,
targetId: tableId,
confidence: 0.9,
reason: `sql-${sql.operation.toLowerCase()}`,
});
}
}
// ── EXEC CICS blocks -> CodeElement nodes + CALLS edges ────────
for (const cics of extracted.execCicsBlocks) {
const cicsId = generateId('CodeElement', `${filePath}:exec-cics:L${cics.line}`);
graph.addNode({
id: cicsId,
label: 'CodeElement',
properties: {
name: `EXEC CICS ${cics.command}`,
filePath,
startLine: cics.line,
endLine: cics.line,
language: 'cobol' as any,
description: cics.mapName ? `map:${cics.mapName}` : cics.programName ? `program:${cics.programName}` : undefined,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${parentId}->${cicsId}`),
type: 'CONTAINS',
sourceId: parentId,
targetId: cicsId,
confidence: 1.0,
reason: 'cobol-exec-cics',
});
// LINK/XCTL -> cross-program CALLS
if (cics.programName && (cics.command === 'LINK' || cics.command === 'XCTL')) {
const cicsTargetModuleId = moduleNodeIds.get(cics.programName.toUpperCase());
const targetId = cicsTargetModuleId
?? generateId('Module', `<unresolved>:${cics.programName.toUpperCase()}`);
const cicsReason = `cics-${cics.command.toLowerCase()}`;
graph.addRelationship({
id: generateId('CALLS', `${parentId}->cics-${cics.command.toLowerCase()}->${cics.programName}:L${cics.line}`),
type: 'CALLS',
sourceId: parentId,
targetId,
confidence: cicsTargetModuleId ? 0.95 : 0.5,
reason: cicsTargetModuleId ? cicsReason : `${cicsReason}-unresolved`,
});
}
}
// ── ENTRY points -> Constructor nodes ──────────────────────────
for (const entry of extracted.entryPoints) {
const entryId = generateId('Constructor', `${filePath}:${entry.name}`);
graph.addNode({
id: entryId,
label: 'Constructor',
properties: {
name: entry.name,
filePath,
startLine: entry.line,
endLine: entry.line,
language: 'cobol' as any,
isExported: true,
description: entry.parameters.length > 0 ? `using:${entry.parameters.join(',')}` : undefined,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${parentId}->${entryId}`),
type: 'CONTAINS',
sourceId: parentId,
targetId: entryId,
confidence: 1.0,
reason: 'cobol-entry-point',
});
// Register in moduleNodeIds for cross-program resolution
moduleNodeIds.set(entry.name.toUpperCase(), entryId);
}
// ── MOVE data flow -> ACCESSES edges (read/write) ──────────────
for (const move of extracted.moves) {
const fromPropId = findDataItemNode(move.from, extracted.dataItems, filePath);
const toPropId = findDataItemNode(move.to, extracted.dataItems, filePath);
const callerId = move.caller
? (paraNodeIds.get(move.caller.toUpperCase()) ?? parentId)
: parentId;
if (fromPropId) {
graph.addRelationship({
id: generateId('ACCESSES', `${callerId}->read->${move.from}:L${move.line}`),
type: 'ACCESSES',
sourceId: callerId,
targetId: fromPropId,
confidence: 0.9,
reason: move.corresponding ? 'cobol-move-corresponding-read' : 'cobol-move-read',
});
}
if (toPropId) {
graph.addRelationship({
id: generateId('ACCESSES', `${callerId}->write->${move.to}:L${move.line}`),
type: 'ACCESSES',
sourceId: callerId,
targetId: toPropId,
confidence: 0.9,
reason: move.corresponding ? 'cobol-move-corresponding-write' : 'cobol-move-write',
});
}
}
// ── File declarations -> Record nodes ──────────────────────────
for (const fd of extracted.fileDeclarations) {
const fdId = generateId('Record', `${filePath}:${fd.selectName}`);
graph.addNode({
id: fdId,
label: 'Record',
properties: {
name: fd.selectName,
filePath,
startLine: fd.line,
endLine: fd.line,
language: 'cobol' as any,
description: `assign:${fd.assignTo}${fd.organization ? ` org:${fd.organization}` : ''}${fd.access ? ` access:${fd.access}` : ''}`,
},
});
graph.addRelationship({
id: generateId('CONTAINS', `${parentId}->${fdId}`),
type: 'CONTAINS',
sourceId: parentId,
targetId: fdId,
confidence: 1.0,
reason: 'cobol-file-declaration',
});
}
}
// ---------------------------------------------------------------------------
// Helpers
// ---------------------------------------------------------------------------
/** Find the section that contains a given line number. */
function findContainingSection(
line: number,
sections: Array<{ name: string; line: number }>,
sectionNodeIds: Map<string, string>,
): string | undefined {
// Sections are in order; find the last section whose start line <= the target line
let best: string | undefined;
for (const sec of sections) {
if (sec.line <= line) {
best = sectionNodeIds.get(sec.name.toUpperCase());
} else {
break;
}
}
return best;
}
@@ -0,0 +1,446 @@
/**
* COBOL COPY statement expansion engine.
*
* Expands COPY statements by inlining copybook content, applying REPLACING
* transformations (LEADING, TRAILING, EXACT), and handling nested copies
* with cycle detection.
*
* This is a preprocessing step that runs BEFORE extractCobolSymbolsWithRegex.
* The caller should run preprocessCobolSource first to clean patch markers.
*
* Supported syntax:
* COPY CPSESP.
* COPY "WORKGRID.CPY".
* COPY CPSESP REPLACING LEADING "ESP-" BY "LK-ESP-"
* LEADING "KPSESPL" BY "LK-KPSESPL".
* COPY ANAZI REPLACING "ANAZI-KEY" BY "LK-KEY".
*/
// ---------------------------------------------------------------------------
// Public interfaces
// ---------------------------------------------------------------------------
export interface CopyReplacing {
type: 'LEADING' | 'TRAILING' | 'EXACT';
from: string;
to: string;
}
export interface CopyResolution {
copyTarget: string;
resolvedPath: string | null;
line: number;
replacing: CopyReplacing[];
}
export interface CopyExpansionResult {
expandedContent: string;
copyResolutions: CopyResolution[];
expansionDepth: number;
}
// ---------------------------------------------------------------------------
// Constants
// ---------------------------------------------------------------------------
export const DEFAULT_MAX_DEPTH = 10;
/** COBOL identifier pattern: starts with letter, contains letters, digits, hyphens. */
const RE_COBOL_IDENTIFIER = /\b([A-Z][A-Z0-9-]*)\b/gi;
// ---------------------------------------------------------------------------
// Private helpers
// ---------------------------------------------------------------------------
/**
* Strip inline comments (Italian-style `|` comments).
* Only strips if `|` appears in the code area (col 7+).
*/
function stripInlineComment(line: string): string {
const idx = line.indexOf('|');
return idx >= 0 ? line.substring(0, idx) : line;
}
/**
* Check if a line is a COBOL comment (indicator in col 7 is `*` or `/`).
*/
function isCommentLine(line: string): boolean {
return line.length >= 7 && (line[6] === '*' || line[6] === '/');
}
/**
* Check if a line is a continuation line (indicator in col 7 is `-`).
*/
function isContinuationLine(line: string): boolean {
return line.length >= 7 && line[6] === '-';
}
/**
* Merge continuation lines into their predecessors.
* Returns an array of logical lines with their original starting line numbers.
*/
function mergeLogicalLines(
rawLines: string[],
): Array<{ text: string; lineNum: number }> {
const logical: Array<{ text: string; lineNum: number }> = [];
for (let i = 0; i < rawLines.length; i++) {
const raw = rawLines[i];
// Skip comment lines
if (isCommentLine(raw)) {
logical.push({ text: '', lineNum: i });
continue;
}
// Continuation: merge into previous logical line
if (isContinuationLine(raw)) {
if (logical.length > 0) {
const prev = logical[logical.length - 1];
const continuation = raw.length > 7 ? raw.substring(7).trimStart() : '';
prev.text += continuation;
}
// Push empty placeholder to preserve line count
logical.push({ text: '', lineNum: i });
continue;
}
// Normal line: strip inline comments
const cleaned = stripInlineComment(raw);
logical.push({ text: cleaned, lineNum: i });
}
return logical;
}
// ---------------------------------------------------------------------------
// COPY statement parsing
// ---------------------------------------------------------------------------
interface ParsedCopyStatement {
startLine: number;
endLine: number;
target: string;
replacing: CopyReplacing[];
}
/**
* Parse REPLACING clause text into structured replacements.
*
* Input examples:
* LEADING "ESP-" BY "LK-ESP-" LEADING "KPSESPL" BY "LK-KPSESPL"
* "ANAZI-KEY" BY "LK-KEY"
* TRAILING "-IN" BY "-OUT"
*/
function parseReplacingClause(text: string): CopyReplacing[] {
const replacings: CopyReplacing[] = [];
if (!text || text.trim().length === 0) return replacings;
// Tokenize: split on whitespace, preserving quoted strings
const tokens: string[] = [];
const tokenRe = /"([^"]*)"|(\S+)/g;
let tm: RegExpExecArray | null;
while ((tm = tokenRe.exec(text)) !== null) {
// Store the matched content; for quoted strings, keep the inner value
// but mark them so we can distinguish. We'll store all as plain strings
// and track which were quoted separately.
tokens.push(tm[1] !== undefined ? tm[1] : tm[2]);
}
// Parse token stream: [LEADING|TRAILING]? <from> BY <to>
let i = 0;
while (i < tokens.length) {
let type: CopyReplacing['type'] = 'EXACT';
const upper = tokens[i].toUpperCase();
// Check for type modifier
if (upper === 'LEADING') {
type = 'LEADING';
i++;
} else if (upper === 'TRAILING') {
type = 'TRAILING';
i++;
}
if (i >= tokens.length) break;
const from = tokens[i];
i++;
// Expect BY keyword
if (i >= tokens.length) break;
if (tokens[i].toUpperCase() !== 'BY') {
// Malformed — skip this token and try to resync
continue;
}
i++; // skip BY
if (i >= tokens.length) break;
const to = tokens[i];
i++;
replacings.push({ type, from, to });
}
return replacings;
}
/**
* Scan logical lines for COPY statements.
* COPY statements can span multiple lines and terminate with a period.
*/
function parseCopyStatements(
logicalLines: Array<{ text: string; lineNum: number }>,
): ParsedCopyStatement[] {
const results: ParsedCopyStatement[] = [];
let accumulator: string | null = null;
let startLine = 0;
let endLine = 0;
for (let i = 0; i < logicalLines.length; i++) {
const { text, lineNum } = logicalLines[i];
if (text.length === 0) continue;
// Check for COPY keyword start (not inside a string context)
const copyStart = text.match(/\bCOPY\b/i);
if (accumulator === null) {
if (!copyStart) continue;
// Start accumulating from the COPY keyword onwards
const copyIdx = copyStart.index!;
accumulator = text.substring(copyIdx);
startLine = lineNum;
endLine = lineNum;
} else {
// Continue accumulating
accumulator += ' ' + text.trim();
endLine = lineNum;
}
// Check if statement terminates (period at end of accumulated text)
if (accumulator !== null && /\.\s*$/.test(accumulator)) {
const parsed = parseSingleCopyStatement(accumulator, startLine, endLine);
if (parsed) {
results.push(parsed);
}
accumulator = null;
}
}
// If there's an unterminated COPY (missing period), try to parse what we have
if (accumulator !== null) {
const parsed = parseSingleCopyStatement(accumulator, startLine, endLine);
if (parsed) {
results.push(parsed);
}
}
return results;
}
/**
* Parse a single complete COPY statement string.
*
* Formats:
* COPY target.
* COPY "target".
* COPY target REPLACING ... .
*/
function parseSingleCopyStatement(
stmt: string,
startLine: number,
endLine: number,
): ParsedCopyStatement | null {
// Strip terminating period
const text = stmt.replace(/\.\s*$/, '').trim();
// Extract target: COPY <target> or COPY "<target>" or COPY '<target>'
const targetMatch = text.match(/^COPY\s+(?:"([^"]+)"|'([^']+)'|([A-Z][A-Z0-9-]*))/i);
if (!targetMatch) return null;
const target = targetMatch[1] || targetMatch[2] || targetMatch[3];
// Extract REPLACING clause if present
let replacing: CopyReplacing[] = [];
const replacingIdx = text.search(/\bREPLACING\b/i);
if (replacingIdx >= 0) {
const replacingText = text.substring(replacingIdx + 'REPLACING'.length);
replacing = parseReplacingClause(replacingText);
}
return { startLine, endLine, target, replacing };
}
// ---------------------------------------------------------------------------
// REPLACING application
// ---------------------------------------------------------------------------
/**
* Apply REPLACING transformations to copybook content.
*
* LEADING: replace prefix in COBOL identifiers.
* TRAILING: replace suffix in COBOL identifiers.
* EXACT: replace exact token matches.
*/
function applyReplacing(content: string, replacings: CopyReplacing[]): string {
if (replacings.length === 0) return content;
return content.replace(RE_COBOL_IDENTIFIER, (match) => {
for (const r of replacings) {
const upper = match.toUpperCase();
const from = r.from.toUpperCase();
const to = r.to.toUpperCase();
switch (r.type) {
case 'LEADING':
if (upper.startsWith(from)) {
return to + match.substring(from.length);
}
break;
case 'TRAILING':
if (upper.endsWith(from)) {
return match.substring(0, match.length - from.length) + to;
}
break;
case 'EXACT':
if (upper === from) {
return to;
}
break;
}
}
return match;
});
}
// ---------------------------------------------------------------------------
// Main expansion engine
// ---------------------------------------------------------------------------
/**
* Expand COBOL COPY statements by inlining copybook content.
*
* @param content - Source COBOL content (after preprocessCobolSource)
* @param filePath - Path of the source file (for diagnostics)
* @param resolveFile - Maps a COPY target name to a filesystem path, or null if not found
* @param readFile - Reads file content by path, or null if unreadable
* @param maxDepth - Maximum nesting depth for recursive expansion (default: 10)
* @returns Expanded content, resolution metadata, and maximum depth reached
*/
export function expandCopies(
content: string,
filePath: string,
resolveFile: (name: string) => string | null,
readFile: (path: string) => string | null,
maxDepth: number = DEFAULT_MAX_DEPTH,
/** Optional shared set to deduplicate circular-COPY warnings across multiple calls. */
warnedCircular: Set<string> = new Set<string>(),
): CopyExpansionResult {
const allResolutions: CopyResolution[] = [];
let maxDepthReached = 0;
const expanded = expandRecursive(content, filePath, 0, new Set<string>());
return {
expandedContent: expanded,
copyResolutions: allResolutions,
expansionDepth: maxDepthReached,
};
/**
* Recursively expand COPY statements in content.
*
* @param src - Source content to expand
* @param srcPath - Path of the file being expanded (for cycle detection logging)
* @param depth - Current recursion depth
* @param visited - Set of already-visited copybook paths (cycle detection)
*/
function expandRecursive(
src: string,
srcPath: string,
depth: number,
visited: Set<string>,
): string {
if (depth > maxDepthReached) {
maxDepthReached = depth;
}
const rawLines = src.split('\n');
const logicalLines = mergeLogicalLines(rawLines);
const copyStatements = parseCopyStatements(logicalLines);
// No COPY statements — return as-is
if (copyStatements.length === 0) return src;
// Process COPY statements in reverse order so line numbers stay valid
// as we splice content
const outputLines = [...rawLines];
for (let ci = copyStatements.length - 1; ci >= 0; ci--) {
const cs = copyStatements[ci];
// Resolve the copybook path
const resolvedPath = resolveFile(cs.target);
// Record resolution metadata
allResolutions.push({
copyTarget: cs.target,
resolvedPath,
line: cs.startLine,
replacing: cs.replacing,
});
// Cannot resolve — keep original lines
if (resolvedPath === null) {
continue;
}
// Cycle detection
if (visited.has(resolvedPath)) {
if (!warnedCircular.has(resolvedPath)) {
warnedCircular.add(resolvedPath);
console.warn(
`[cobol-copy-expander] Circular COPY detected: ${cs.target} (${resolvedPath}) ` +
`includes itself. Skipping expansion.`,
);
}
continue;
}
// Max depth exceeded — keep unexpanded
if (depth >= maxDepth) {
console.warn(
`[cobol-copy-expander] Max expansion depth (${maxDepth}) reached for ` +
`COPY ${cs.target} in ${srcPath}. Skipping expansion.`,
);
continue;
}
// Read the copybook content
const copybookContent = readFile(resolvedPath);
if (copybookContent === null) {
continue;
}
// Apply REPLACING transformations
const replaced = applyReplacing(copybookContent, cs.replacing);
// Recurse into the copybook for nested COPYs
const nestedVisited = new Set(visited);
nestedVisited.add(resolvedPath);
const expandedCopybook = expandRecursive(
replaced,
resolvedPath,
depth + 1,
nestedVisited,
);
// Splice: replace the COPY statement lines with expanded content
const expansionLines = expandedCopybook.split('\n');
const removeCount = cs.endLine - cs.startLine + 1;
outputLines.splice(cs.startLine, removeCount, ...expansionLines);
}
return outputLines.join('\n');
}
}
@@ -0,0 +1,911 @@
/**
* COBOL source pre-processing and regex-based symbol extraction.
*
* DESIGN DECISION — Why regex instead of a full parser (ANTLR4, tree-sitter):
*
* 1. Performance: Regex processes ~1ms/file vs 50-200ms/file for ANTLR4/tree-sitter.
* On EPAGHE (14k COBOL files), this is ~14 seconds vs 12-47 minutes.
*
* 2. Reliability: tree-sitter-cobol@0.0.1's external scanner hangs indefinitely
* on ~5% of production files (no timeout possible). ANTLR4's proleap-cobol-parser
* is a Java project — using it from Node.js requires Java subprocesses or
* extracting .g4 grammars and generating JS/TS targets (significant effort).
*
* 3. Dialect compatibility: GnuCOBOL with Italian comments, patch markers in
* cols 1-6 (mzADD, estero, etc.), and vendor extensions. Formal grammars
* target COBOL-85 and would need dialect modifications.
*
* 4. Industry precedent: ctags, GitHub code navigation, and Sourcegraph all use
* regex-based extraction for code indexing. Full parsing is only needed for
* compilation or semantic analysis, not symbol extraction.
*
* 5. Determinism: Every regex pattern is tested with canonical COBOL input
* (see test/unit/cobol-preprocessor.test.ts). Same input always produces
* same output — no grammar ambiguity or parser state issues.
*
* This module provides:
* 1. preprocessCobolSource() — cleans patch markers (kept for potential future use)
* 2. extractCobolSymbolsWithRegex() — single-pass state machine COBOL extraction
*/
// ---------------------------------------------------------------------------
// Public interfaces
// ---------------------------------------------------------------------------
export interface CobolRegexResults {
programName: string | null;
paragraphs: Array<{ name: string; line: number }>;
sections: Array<{ name: string; line: number }>;
performs: Array<{ caller: string | null; target: string; thruTarget?: string; line: number }>;
calls: Array<{ target: string; line: number }>;
copies: Array<{ target: string; line: number }>;
dataItems: Array<{
name: string;
level: number;
line: number;
pic?: string;
usage?: string;
occurs?: number;
redefines?: string;
values?: string[];
section: 'working-storage' | 'linkage' | 'file' | 'local-storage' | 'unknown';
}>;
fileDeclarations: Array<{
selectName: string;
assignTo: string;
organization?: string;
access?: string;
recordKey?: string;
fileStatus?: string;
line: number;
}>;
fdEntries: Array<{
fdName: string;
recordName?: string;
line: number;
}>;
programMetadata: {
author?: string;
dateWritten?: string;
};
// Phase 2: EXEC blocks
execSqlBlocks: Array<{
line: number;
tables: string[];
cursors: string[];
hostVariables: string[];
operation: 'SELECT' | 'INSERT' | 'UPDATE' | 'DELETE' | 'DECLARE' | 'OPEN' | 'CLOSE' | 'FETCH' | 'OTHER';
}>;
execCicsBlocks: Array<{
line: number;
command: string;
mapName?: string;
programName?: string;
transId?: string;
}>;
// Phase 3: Linkage + Data Flow
procedureUsing: string[];
entryPoints: Array<{
name: string;
parameters: string[];
line: number;
}>;
moves: Array<{
from: string;
to: string;
line: number;
caller: string | null;
corresponding: boolean;
}>;
}
// ---------------------------------------------------------------------------
// Preserved exactly: preprocessCobolSource
// ---------------------------------------------------------------------------
/**
* Normalize COBOL source for regex-based extraction.
*
* The COBOL fixed-format sequence number area (columns 1-6) is semantically
* irrelevant to parsing — compilers and tools always ignore it. This
* function replaces ANY non-space content in columns 1-6 with spaces so that
* position-sensitive regexes (paragraph/section detection, data-item anchors,
* etc.) work identically whether the file carries:
* • numeric sequence numbers (000100 … 999999)
* • alphabetic patch markers (mzADD, estero, #patch, …)
* • the COBOL default of all spaces
*
* Preserves exact line count for position mapping.
*/
export function preprocessCobolSource(content: string): string {
const lines = content.split('\n');
for (let i = 0; i < lines.length; i++) {
const line = lines[i];
if (line.length < 7) continue;
const seq = line.substring(0, 6);
// Replace any non-space character in the sequence area with spaces.
// This covers numeric sequence numbers (000100), alphabetic patch markers
// (mzADD, estero), '#'-prefixed markers, and mixed sequences.
if (/\S/.test(seq)) {
lines[i] = ' ' + line.substring(6);
}
}
return lines.join('\n');
}
// ---------------------------------------------------------------------------
// Preserved exactly: EXCLUDED_PARA_NAMES
// ---------------------------------------------------------------------------
const EXCLUDED_PARA_NAMES = new Set([
'DECLARATIVES', 'END', 'PROCEDURE', 'IDENTIFICATION',
'ENVIRONMENT', 'DATA', 'WORKING-STORAGE', 'LINKAGE',
'FILE', 'LOCAL-STORAGE', 'COMMUNICATION', 'REPORT',
'SCREEN', 'INPUT-OUTPUT', 'CONFIGURATION',
]);
// ---------------------------------------------------------------------------
// State machine types
// ---------------------------------------------------------------------------
type Division = 'identification' | 'environment' | 'data' | 'procedure' | null;
type DataSection = 'working-storage' | 'linkage' | 'file' | 'local-storage' | 'unknown';
type EnvironmentSection = 'input-output' | 'configuration' | null;
// ---------------------------------------------------------------------------
// Regex constants (compiled once, reused across calls)
// ---------------------------------------------------------------------------
const RE_DIVISION = /\b(IDENTIFICATION|ENVIRONMENT|DATA|PROCEDURE)\s+DIVISION\b/i;
const RE_SECTION = /\b(WORKING-STORAGE|LINKAGE|FILE|LOCAL-STORAGE|INPUT-OUTPUT|CONFIGURATION)\s+SECTION\b/i;
// IDENTIFICATION DIVISION
const RE_PROGRAM_ID = /\bPROGRAM-ID\.\s*([A-Z][A-Z0-9-]*)/i;
const RE_AUTHOR = /^\s+AUTHOR\.\s*(.+)/i;
const RE_DATE_WRITTEN = /^\s+DATE-WRITTEN\.\s*(.+)/i;
// ENVIRONMENT DIVISION — SELECT
const RE_SELECT_START = /\bSELECT\s+([A-Z][A-Z0-9-]+)/i;
// DATA DIVISION
const RE_FD = /^\s+FD\s+([A-Z][A-Z0-9-]+)/i;
const RE_DATA_ITEM = /^\s+(\d{1,2})\s+([A-Z][A-Z0-9-]+)\s*(.*)/i;
const RE_ANONYMOUS_REDEFINES = /^\s+(\d{1,2})\s+REDEFINES\s+([A-Z][A-Z0-9-]+)/i;
const RE_88_LEVEL = /^\s+88\s+([A-Z][A-Z0-9-]+)\s+VALUES?\s+(?:ARE\s+)?(.+)/i;
// PROCEDURE DIVISION
const RE_PROC_SECTION = /^ ([A-Z][A-Z0-9-]+)\s+SECTION\.\s*$/;
const RE_PROC_PARAGRAPH = /^ ([A-Z][A-Z0-9-]+)\.\s*$/;
const RE_PERFORM = /\bPERFORM\s+([A-Z][A-Z0-9-]+)(?:\s+THRU\s+([A-Z][A-Z0-9-]+))?/i;
// ALL DIVISIONS
// Both double-quoted ("PROG") and single-quoted ('PROG') targets are valid COBOL.
// Use separate alternation groups so quotes must match (prevents "PROG' false-matches).
const RE_CALL = /\bCALL\s+(?:"([^"]+)"|'([^']+)')/i;
const RE_COPY_UNQUOTED = /\bCOPY\s+([A-Z][A-Z0-9-]+)(?:\s|\.)/i;
const RE_COPY_QUOTED = /\bCOPY\s+(?:"([^"]+)"|'([^']+)')(?:\s|\.)/i;
// EXEC blocks
const RE_EXEC_SQL_START = /\bEXEC\s+SQL\b/i;
const RE_EXEC_CICS_START = /\bEXEC\s+CICS\b/i;
const RE_END_EXEC = /\bEND-EXEC\b/i;
// PROCEDURE DIVISION USING
const RE_PROC_USING = /\bPROCEDURE\s+DIVISION\s+USING\s+([\s\S]*?)(?:\.|$)/i;
// ENTRY point
const RE_ENTRY = /\bENTRY\s+"([^"]+)"(?:\s+USING\s+([\s\S]*?))?(?:\.|$)/i;
// MOVE statement
const RE_MOVE = /\bMOVE\s+(CORRESPONDING\s+)?([A-Z][A-Z0-9-]+)\s+TO\s+([A-Z][A-Z0-9-]+)/i;
const MOVE_SKIP = new Set([
'SPACES', 'ZEROS', 'ZEROES', 'LOW-VALUES', 'LOW-VALUE',
'HIGH-VALUES', 'HIGH-VALUE', 'QUOTES', 'QUOTE', 'ALL',
]);
// PERFORM: keywords that may follow PERFORM but are NOT paragraph/section names.
// Inline PERFORM loops (UNTIL, VARYING) and inline test clauses (WITH TEST,
// FOREVER) must not be stored as perform-target false positives.
const PERFORM_KEYWORD_SKIP = new Set([
'UNTIL', 'VARYING', 'WITH', 'TEST', 'FOREVER',
]);
// ---------------------------------------------------------------------------
// Private helper: strip Italian inline comments (| and everything after)
// ---------------------------------------------------------------------------
function stripInlineComment(line: string): string {
const idx = line.indexOf('|');
return idx >= 0 ? line.substring(0, idx) : line;
}
// ---------------------------------------------------------------------------
// Private helper: parse data item trailing clauses (PIC, USAGE, etc.)
// ---------------------------------------------------------------------------
function parseDataItemClauses(rest: string): {
pic?: string;
usage?: string;
redefines?: string;
occurs?: number;
} {
const result: { pic?: string; usage?: string; redefines?: string; occurs?: number } = {};
// Strip trailing period for easier parsing
const text = rest.replace(/\.\s*$/, '');
// PIC / PICTURE [IS] <picture-string>
const picMatch = text.match(/\bPIC(?:TURE)?\s+(?:IS\s+)?(\S+)/i);
if (picMatch) {
result.pic = picMatch[1];
}
// USAGE [IS] <usage-type> — including non-standard COMP-6, COMP-X etc.
const usageMatch = text.match(/\bUSAGE\s+(?:IS\s+)?(COMP(?:UTATIONAL)?(?:-[0-9X])?|BINARY|PACKED-DECIMAL|DISPLAY|INDEX|POINTER|NATIONAL)\b/i);
if (usageMatch) {
result.usage = usageMatch[1].toUpperCase();
} else {
// Standalone COMP variants without USAGE keyword
const compMatch = text.match(/\b(COMP(?:UTATIONAL)?(?:-[0-9X])?|BINARY|PACKED-DECIMAL)\b/i);
if (compMatch) {
result.usage = compMatch[1].toUpperCase();
}
}
// REDEFINES <name>
const redefMatch = text.match(/\bREDEFINES\s+([A-Z][A-Z0-9-]+)/i);
if (redefMatch) {
result.redefines = redefMatch[1];
}
// OCCURS <n> [TIMES]
const occursMatch = text.match(/\bOCCURS\s+(\d+)/i);
if (occursMatch) {
result.occurs = parseInt(occursMatch[1], 10);
}
return result;
}
// ---------------------------------------------------------------------------
// Private helper: parse 88-level condition values
// ---------------------------------------------------------------------------
function parseConditionValues(valuesStr: string): string[] {
// Strip trailing period
const text = valuesStr.replace(/\.\s*$/, '').trim();
const values: string[] = [];
// Match quoted strings: "O" "Y" "I"
const quotedRe = /"([^"]*)"/g;
let qm: RegExpExecArray | null;
let hasQuoted = false;
while ((qm = quotedRe.exec(text)) !== null) {
values.push(qm[1]);
hasQuoted = true;
}
if (hasQuoted) return values;
// No quotes — split on whitespace, filtering out THRU/THROUGH keywords
// Handle: 11 12 16 17 21 or 1 THRU 5
const tokens = text.split(/\s+/);
for (const token of tokens) {
const upper = token.toUpperCase();
if (upper === 'THRU' || upper === 'THROUGH') {
// Keep THRU ranges as combined value: prev THRU next is already captured
// by having both sides in the array
continue;
}
if (token.length > 0) {
values.push(token);
}
}
return values;
}
// ---------------------------------------------------------------------------
// Private helper: parse accumulated multi-line SELECT statement
// ---------------------------------------------------------------------------
interface FileDeclaration {
selectName: string;
assignTo: string;
organization?: string;
access?: string;
recordKey?: string;
fileStatus?: string;
line: number;
}
function parseSelectStatement(stmt: string, startLine: number): FileDeclaration | null {
// Normalize whitespace
const text = stmt.replace(/\s+/g, ' ').trim();
const nameMatch = text.match(/^SELECT\s+([A-Z][A-Z0-9-]+)/i);
if (!nameMatch) return null;
const result: FileDeclaration = {
selectName: nameMatch[1],
assignTo: '',
line: startLine,
};
const assignMatch = text.match(/\bASSIGN\s+(?:TO\s+)?("([^"]+)"|([A-Z][A-Z0-9-]*))/i);
if (assignMatch) {
result.assignTo = assignMatch[2] || assignMatch[3] || '';
}
const orgMatch = text.match(/\bORGANIZATION\s+(?:IS\s+)?(SEQUENTIAL|INDEXED|RELATIVE|LINE\s+SEQUENTIAL)/i);
if (orgMatch) {
result.organization = orgMatch[1].toUpperCase();
}
const accessMatch = text.match(/\bACCESS\s+(?:MODE\s+)?(?:IS\s+)?(SEQUENTIAL|RANDOM|DYNAMIC)/i);
if (accessMatch) {
result.access = accessMatch[1].toUpperCase();
}
const keyMatch = text.match(/\bRECORD\s+KEY\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)/i);
if (keyMatch) {
result.recordKey = keyMatch[1];
}
// FILE STATUS IS / STATUS IS
const statusMatch = text.match(/\b(?:FILE\s+)?STATUS\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)/i);
if (statusMatch) {
result.fileStatus = statusMatch[1];
}
return result;
}
// ---------------------------------------------------------------------------
// Private helper: parse EXEC SQL block
// ---------------------------------------------------------------------------
type SqlOperation = 'SELECT' | 'INSERT' | 'UPDATE' | 'DELETE' | 'DECLARE' | 'OPEN' | 'CLOSE' | 'FETCH' | 'OTHER';
function parseExecSqlBlock(block: string, line: number): CobolRegexResults['execSqlBlocks'][number] {
// Strip EXEC SQL ... END-EXEC wrapper
const body = block
.replace(/\bEXEC\s+SQL\b/i, '')
.replace(/\bEND-EXEC\b/i, '')
.replace(/\s+/g, ' ')
.trim();
// Determine operation from first SQL keyword
const firstWord = body.split(/\s+/)[0]?.toUpperCase() || '';
const OP_MAP: Record<string, SqlOperation> = {
SELECT: 'SELECT', INSERT: 'INSERT', UPDATE: 'UPDATE', DELETE: 'DELETE',
DECLARE: 'DECLARE', OPEN: 'OPEN', CLOSE: 'CLOSE', FETCH: 'FETCH',
};
const operation: SqlOperation = OP_MAP[firstWord] || 'OTHER';
// Extract table names from FROM, INTO (INSERT), UPDATE, DELETE FROM, JOIN
const tables: string[] = [];
const tablePatterns = [
/\bFROM\s+([A-Z][A-Z0-9_]+)/gi,
/\bINTO\s+([A-Z][A-Z0-9_]+)/gi,
/\bUPDATE\s+([A-Z][A-Z0-9_]+)/gi,
/\bJOIN\s+([A-Z][A-Z0-9_]+)/gi,
];
for (const re of tablePatterns) {
let m: RegExpExecArray | null;
while ((m = re.exec(body)) !== null) {
const name = m[1].toUpperCase();
// Skip host variables and SQL keywords
if (!name.startsWith(':') && !tables.includes(name)) {
tables.push(name);
}
}
}
// Extract cursor names from DECLARE ... CURSOR
const cursors: string[] = [];
const cursorRe = /\bDECLARE\s+([A-Z][A-Z0-9_-]+)\s+CURSOR\b/gi;
let cm: RegExpExecArray | null;
while ((cm = cursorRe.exec(body)) !== null) {
cursors.push(cm[1]);
}
// Extract host variables: :VARIABLE-NAME (strip the colon)
const hostVariables: string[] = [];
const hostRe = /:([A-Z][A-Z0-9-]+)/gi;
let hm: RegExpExecArray | null;
while ((hm = hostRe.exec(body)) !== null) {
const name = hm[1];
if (!hostVariables.includes(name)) {
hostVariables.push(name);
}
}
return { line, tables, cursors, hostVariables, operation };
}
// ---------------------------------------------------------------------------
// Private helper: parse EXEC CICS block
// ---------------------------------------------------------------------------
function parseExecCicsBlock(block: string, line: number): CobolRegexResults['execCicsBlocks'][number] {
// Strip EXEC CICS ... END-EXEC wrapper
const body = block
.replace(/\bEXEC\s+CICS\b/i, '')
.replace(/\bEND-EXEC\b/i, '')
.replace(/\s+/g, ' ')
.trim();
// Command: first keyword(s) — handle two-word commands like SEND MAP, RECEIVE MAP
const twoWordCommands = ['SEND MAP', 'RECEIVE MAP', 'SEND TEXT', 'SEND CONTROL', 'READ NEXT', 'READ PREV'];
let command = '';
const upperBody = body.toUpperCase();
for (const twoWord of twoWordCommands) {
if (upperBody.startsWith(twoWord)) {
command = twoWord;
break;
}
}
if (!command) {
command = body.split(/\s+/)[0]?.toUpperCase() || '';
}
const result: CobolRegexResults['execCicsBlocks'][number] = { line, command };
// MAP name: MAP('name') or MAP("name")
const mapMatch = body.match(/\bMAP\s*\(\s*['"]([^'"]+)['"]\s*\)/i);
if (mapMatch) result.mapName = mapMatch[1];
// PROGRAM name: PROGRAM('name') or PROGRAM("name")
const progMatch = body.match(/\bPROGRAM\s*\(\s*['"]([^'"]+)['"]\s*\)/i);
if (progMatch) result.programName = progMatch[1];
// TRANSID: TRANSID('name') or TRANSID("name")
const transMatch = body.match(/\bTRANSID\s*\(\s*['"]([^'"]+)['"]\s*\)/i);
if (transMatch) result.transId = transMatch[1];
return result;
}
// ---------------------------------------------------------------------------
// Main extraction: single-pass state machine
// ---------------------------------------------------------------------------
/**
* Extract COBOL symbols using a single-pass state machine.
* Extracts program name, paragraphs, sections, CALL, PERFORM, COPY,
* data items, file declarations, FD entries, and program metadata.
*/
export function extractCobolSymbolsWithRegex(
content: string,
_filePath: string,
): CobolRegexResults {
const rawLines = content.split('\n');
const result: CobolRegexResults = {
programName: null,
paragraphs: [],
sections: [],
performs: [],
calls: [],
copies: [],
dataItems: [],
fileDeclarations: [],
fdEntries: [],
programMetadata: {},
execSqlBlocks: [],
execCicsBlocks: [],
procedureUsing: [],
entryPoints: [],
moves: [],
};
// --- State ---
let currentDivision: Division = null;
let currentDataSection: DataSection = 'unknown';
let currentEnvSection: EnvironmentSection = null;
let currentParagraph: string | null = null;
// SELECT accumulator (multi-line)
let selectAccum: string | null = null;
let selectStartLine = 0;
// EXEC block accumulator (multi-line EXEC SQL / EXEC CICS)
let execAccum: { type: 'sql' | 'cics'; lines: string; startLine: number } | null = null;
// FD tracking: after seeing FD, the next 01-level data item is its record
let pendingFdName: string | null = null;
let pendingFdLine = 0;
// Continuation line buffer
let pendingLine: string | null = null;
let pendingLineNumber = 0;
// --- Process each raw line ---
for (let i = 0; i < rawLines.length; i++) {
const raw = rawLines[i];
// Skip lines too short to have indicator area
if (raw.length < 7) {
// If there's a pending continuation, flush it
if (pendingLine !== null) {
processLogicalLine(pendingLine, pendingLineNumber);
pendingLine = null;
}
continue;
}
const indicator = raw[6];
// Comment line: indicator is '*' or '/'
if (indicator === '*' || indicator === '/') {
continue;
}
// Continuation line: indicator is '-'
if (indicator === '-') {
if (pendingLine !== null) {
// Append continuation (area B content, trimmed leading spaces)
const continuation = raw.substring(7).trimStart();
pendingLine += continuation;
}
continue;
}
// Normal line — flush any pending continuation first
if (pendingLine !== null) {
processLogicalLine(pendingLine, pendingLineNumber);
pendingLine = null;
}
// Strip inline Italian comments, then use area A+B (from col 7 onwards,
// but keep full line for indentation-sensitive paragraph/section detection)
const cleaned = stripInlineComment(raw);
// Buffer as new pending logical line
pendingLine = cleaned;
pendingLineNumber = i;
}
// Flush final pending line
if (pendingLine !== null) {
processLogicalLine(pendingLine, pendingLineNumber);
}
// Flush any pending SELECT
flushSelect();
// If we saw an FD but never found its record, emit it without a record name
if (pendingFdName !== null) {
result.fdEntries.push({ fdName: pendingFdName, line: pendingFdLine });
pendingFdName = null;
}
return result;
// =========================================================================
// Inner function: process one logical line (after continuation merging)
// =========================================================================
function processLogicalLine(line: string, lineNum: number): void {
// --- EXEC block accumulation (spans any division) ---
if (execAccum !== null) {
execAccum.lines += ' ' + line;
if (RE_END_EXEC.test(line)) {
if (execAccum.type === 'sql') {
result.execSqlBlocks.push(parseExecSqlBlock(execAccum.lines, execAccum.startLine));
} else {
result.execCicsBlocks.push(parseExecCicsBlock(execAccum.lines, execAccum.startLine));
}
execAccum = null;
}
return; // While accumulating, skip normal processing
}
// Check for EXEC SQL / EXEC CICS start
if (RE_EXEC_SQL_START.test(line)) {
execAccum = { type: 'sql', lines: line, startLine: lineNum };
// If END-EXEC is on the same line, finalize immediately
if (RE_END_EXEC.test(line)) {
result.execSqlBlocks.push(parseExecSqlBlock(execAccum.lines, execAccum.startLine));
execAccum = null;
}
return;
}
if (RE_EXEC_CICS_START.test(line)) {
execAccum = { type: 'cics', lines: line, startLine: lineNum };
if (RE_END_EXEC.test(line)) {
result.execCicsBlocks.push(parseExecCicsBlock(execAccum.lines, execAccum.startLine));
execAccum = null;
}
return;
}
// --- Division transitions ---
const divMatch = line.match(RE_DIVISION);
if (divMatch) {
// Flush SELECT if transitioning out of environment
flushSelect();
const divName = divMatch[1].toUpperCase();
switch (divName) {
case 'IDENTIFICATION': currentDivision = 'identification'; break;
case 'ENVIRONMENT': currentDivision = 'environment'; currentEnvSection = null; break;
case 'DATA': currentDivision = 'data'; currentDataSection = 'unknown'; break;
case 'PROCEDURE': {
currentDivision = 'procedure';
currentParagraph = null;
const procUsingMatch = line.match(RE_PROC_USING);
if (procUsingMatch) {
result.procedureUsing = procUsingMatch[1].trim().split(/\s+/).filter(s => s.length > 0);
}
break;
}
}
return;
}
// --- Section transitions ---
const secMatch = line.match(RE_SECTION);
if (secMatch) {
flushSelect();
const secName = secMatch[1].toUpperCase();
switch (secName) {
case 'WORKING-STORAGE': currentDivision = 'data'; currentDataSection = 'working-storage'; break;
case 'LINKAGE': currentDivision = 'data'; currentDataSection = 'linkage'; break;
case 'FILE': currentDivision = 'data'; currentDataSection = 'file'; break;
case 'LOCAL-STORAGE': currentDivision = 'data'; currentDataSection = 'local-storage'; break;
case 'INPUT-OUTPUT': currentDivision = 'environment'; currentEnvSection = 'input-output'; break;
case 'CONFIGURATION': currentDivision = 'environment'; currentEnvSection = 'configuration'; break;
}
return;
}
// --- COPY (all divisions) ---
const copyQMatch = line.match(RE_COPY_QUOTED);
if (copyQMatch) {
result.copies.push({ target: copyQMatch[1] ?? copyQMatch[2], line: lineNum });
} else {
const copyUMatch = line.match(RE_COPY_UNQUOTED);
if (copyUMatch) {
result.copies.push({ target: copyUMatch[1], line: lineNum });
}
}
// --- CALL (all divisions, typically procedure) ---
const callMatch = line.match(RE_CALL);
if (callMatch) {
result.calls.push({ target: callMatch[1] ?? callMatch[2], line: lineNum });
}
// --- Division-specific extraction ---
switch (currentDivision) {
case 'identification':
extractIdentification(line, lineNum);
break;
case 'environment':
extractEnvironment(line, lineNum);
break;
case 'data':
extractData(line, lineNum);
break;
case 'procedure':
extractProcedure(line, lineNum);
break;
}
}
// =========================================================================
// IDENTIFICATION DIVISION extraction
// =========================================================================
function extractIdentification(line: string, _lineNum: number): void {
if (result.programName === null) {
const m = line.match(RE_PROGRAM_ID);
if (m) {
result.programName = m[1];
return;
}
}
const authorMatch = line.match(RE_AUTHOR);
if (authorMatch) {
result.programMetadata.author = authorMatch[1].replace(/\.\s*$/, '').trim();
return;
}
const dateMatch = line.match(RE_DATE_WRITTEN);
if (dateMatch) {
result.programMetadata.dateWritten = dateMatch[1].replace(/\.\s*$/, '').trim();
}
}
// =========================================================================
// ENVIRONMENT DIVISION extraction
// =========================================================================
function extractEnvironment(line: string, lineNum: number): void {
if (currentEnvSection !== 'input-output') return;
// Check for new SELECT statement
const selMatch = line.match(RE_SELECT_START);
if (selMatch) {
// Flush any previous SELECT
flushSelect();
selectAccum = line.trim();
selectStartLine = lineNum;
} else if (selectAccum !== null) {
// Accumulate continuation of current SELECT
selectAccum += ' ' + line.trim();
}
// Check if current SELECT is terminated (ends with period)
if (selectAccum !== null && /\.\s*$/.test(selectAccum)) {
flushSelect();
}
}
function flushSelect(): void {
if (selectAccum === null) return;
const decl = parseSelectStatement(selectAccum, selectStartLine);
if (decl) {
result.fileDeclarations.push(decl);
}
selectAccum = null;
}
// =========================================================================
// DATA DIVISION extraction
// =========================================================================
function extractData(line: string, lineNum: number): void {
// FD entry
const fdMatch = line.match(RE_FD);
if (fdMatch) {
// Flush any previous FD without a record
if (pendingFdName !== null) {
result.fdEntries.push({ fdName: pendingFdName, line: pendingFdLine });
}
pendingFdName = fdMatch[1];
pendingFdLine = lineNum;
return;
}
// 88-level condition names
const lv88Match = line.match(RE_88_LEVEL);
if (lv88Match) {
const name = lv88Match[1];
const values = parseConditionValues(lv88Match[2]);
result.dataItems.push({
name,
level: 88,
line: lineNum,
values,
section: currentDataSection,
});
return;
}
// Anonymous REDEFINES (no name, e.g. "01 REDEFINES WK-PERIVAL.")
const anonRedefMatch = line.match(RE_ANONYMOUS_REDEFINES);
if (anonRedefMatch) {
// Check it's truly anonymous: the second capture is not a valid data name
// followed by more clauses — it's the REDEFINES target directly after level
const level = parseInt(anonRedefMatch[1], 10);
// Only skip if this is genuinely "NN REDEFINES target" with no name between
// We detect this by checking the full data item regex does NOT match
// (because RE_DATA_ITEM expects a name before any clauses)
const dataMatch = line.match(RE_DATA_ITEM);
if (!dataMatch || dataMatch[2].toUpperCase() === 'REDEFINES') {
// Truly anonymous — skip, no node
return;
}
}
// Standard data items: level 01-49, 66, 77
const dataMatch = line.match(RE_DATA_ITEM);
if (dataMatch) {
const level = parseInt(dataMatch[1], 10);
const name = dataMatch[2];
const rest = dataMatch[3] || '';
// Skip FILLER
if (name.toUpperCase() === 'FILLER') return;
// Valid levels: 01-49, 66, 77
if ((level >= 1 && level <= 49) || level === 66 || level === 77) {
const clauses = parseDataItemClauses(rest);
const item: CobolRegexResults['dataItems'][number] = {
name,
level,
line: lineNum,
section: currentDataSection,
};
if (clauses.pic) item.pic = clauses.pic;
if (clauses.usage) item.usage = clauses.usage;
if (clauses.occurs !== undefined) item.occurs = clauses.occurs;
if (clauses.redefines) item.redefines = clauses.redefines;
result.dataItems.push(item);
// If there's a pending FD and this is a 01-level, it's the FD's record
if (pendingFdName !== null && level === 1) {
result.fdEntries.push({
fdName: pendingFdName,
recordName: name,
line: pendingFdLine,
});
pendingFdName = null;
}
}
}
}
// =========================================================================
// PROCEDURE DIVISION extraction
// =========================================================================
function extractProcedure(line: string, lineNum: number): void {
// Section header
const secMatch = line.match(RE_PROC_SECTION);
if (secMatch) {
const name = secMatch[1];
if (!EXCLUDED_PARA_NAMES.has(name) && !name.includes('DIVISION')) {
result.sections.push({ name, line: lineNum });
currentParagraph = name;
}
return;
}
// Paragraph header
const paraMatch = line.match(RE_PROC_PARAGRAPH);
if (paraMatch) {
const name = paraMatch[1];
if (!EXCLUDED_PARA_NAMES.has(name) && !name.includes('DIVISION') && !name.includes('SECTION')) {
result.paragraphs.push({ name, line: lineNum });
currentParagraph = name;
}
return;
}
// PERFORM
const perfMatch = line.match(RE_PERFORM);
if (perfMatch) {
const target = perfMatch[1];
// Skip COBOL inline-perform keywords that are not paragraph names
if (!PERFORM_KEYWORD_SKIP.has(target.toUpperCase())) {
result.performs.push({
caller: currentParagraph,
target,
thruTarget: perfMatch[2] || undefined,
line: lineNum,
});
}
}
// ENTRY point
const entryMatch = line.match(RE_ENTRY);
if (entryMatch) {
result.entryPoints.push({
name: entryMatch[1],
parameters: entryMatch[2] ? entryMatch[2].trim().split(/\s+/).filter(s => s.length > 0) : [],
line: lineNum,
});
}
// MOVE statement (skip literals and figurative constants)
const moveMatch = line.match(RE_MOVE);
if (moveMatch) {
const from = moveMatch[2].toUpperCase();
if (!MOVE_SKIP.has(from)) {
result.moves.push({
from: moveMatch[2],
to: moveMatch[3],
line: lineNum,
caller: currentParagraph,
corresponding: !!moveMatch[1],
});
}
}
}
}
@@ -0,0 +1,266 @@
/**
* JCL Parser — Regex single-pass extraction.
*
* Extracts JCL constructs from mainframe job streams:
* - JOB statements (job name, CLASS, MSGCLASS)
* - EXEC statements (step -> program or proc)
* - DD statements (dataset references, DISP)
* - PROC definitions (in-stream and catalogued)
* - INCLUDE MEMBER= directives
* - SET symbolic parameters
* - IF/ELSE/ENDIF conditional execution
* - JCLLIB ORDER= search paths
*
* Pattern follows cobol-preprocessor.ts — regex-only, no tree-sitter.
*/
export interface JclParseResults {
jobs: Array<{ name: string; line: number; class?: string; msgclass?: string }>;
steps: Array<{ name: string; jobName: string; program?: string; proc?: string; line: number }>;
ddStatements: Array<{ ddName: string; stepName: string; dataset?: string; disp?: string; line: number }>;
procs: Array<{ name: string; line: number; isInStream: boolean }>;
includes: Array<{ member: string; line: number }>;
sets: Array<{ variable: string; value: string; line: number }>;
jcllib: Array<{ order: string[]; line: number }>;
conditionals: Array<{ type: 'IF' | 'ELSE' | 'ENDIF'; condition?: string; line: number }>;
}
// ── JCL statement patterns ─────────────────────────────────────────────
// JCL continuation: line ends with a non-blank in col 72, next line starts with //
// We handle continuations by joining lines before matching.
/** Match //jobname JOB ... */
const JOB_RE = /^\/\/(\w{1,8})\s+JOB\s+(.*)/i;
/** Match //stepname EXEC PGM=program or //stepname EXEC procname */
const EXEC_RE = /^\/\/(\w{1,8})\s+EXEC\s+(.*)/i;
/** Match //ddname DD ... */
const DD_RE = /^\/\/(\w{1,8})\s+DD\s+(.*)/i;
/** Match // JCLLIB ORDER=(lib1,lib2,...) */
const JCLLIB_RE = /^\/\/\s+JCLLIB\s+ORDER=\(([^)]+)\)/i;
/** Match // IF condition THEN */
const IF_RE = /^\/\/\s+IF\s+(.+)\s+THEN/i;
/** Match // ELSE */
const ELSE_RE = /^\/\/\s+ELSE\b/i;
/** Match // ENDIF */
const ENDIF_RE = /^\/\/\s+ENDIF\b/i;
/** Match // INCLUDE MEMBER=name */
const INCLUDE_RE = /^\/\/\s+INCLUDE\s+MEMBER=(\w+)/i;
/** Match // SET var=value */
const SET_RE = /^\/\/\s+SET\s+(\w+)=(.+)/i;
/** Match // PROC or //name PROC */
const PROC_RE = /^\/\/(\w*)\s+PROC\b/i;
/** Match // PEND */
const PEND_RE = /^\/\/\s+PEND\b/i;
// ── Parameter extractors ───────────────────────────────────────────────
function extractParam(params: string, key: string): string | undefined {
// Match KEY=VALUE or KEY='VALUE' in JCL parameter string
const re = new RegExp(`${key}=(?:'([^']*)'|(\\S+?))(?:[,\\s]|$)`, 'i');
const m = params.match(re);
return m ? (m[1] ?? m[2]) : undefined;
}
function extractPgm(params: string): string | undefined {
return extractParam(params, 'PGM');
}
function extractProc(params: string): string | undefined {
// If no PGM= keyword, the first positional parameter is the proc name
if (/PGM=/i.test(params)) return undefined;
const cleaned = params.replace(/,.*/, '').trim();
// Proc name is the first token (no = sign)
if (cleaned && !cleaned.includes('=')) {
return cleaned.replace(/[,\s].*/s, '').toUpperCase();
}
return undefined;
}
function extractDsn(params: string): string | undefined {
return extractParam(params, 'DSN') ?? extractParam(params, 'DSNAME');
}
function extractDisp(params: string): string | undefined {
const m = params.match(/DISP=\(?\s*([^),\s]+)/i);
return m ? m[1] : undefined;
}
/**
* Parse a JCL file and extract all constructs.
*
* @param content - Raw JCL file content
* @param filePath - Path for diagnostics (not used in extraction)
* @returns Parsed JCL results
*/
export function parseJcl(content: string, filePath: string): JclParseResults {
const results: JclParseResults = {
jobs: [],
steps: [],
ddStatements: [],
procs: [],
includes: [],
sets: [],
jcllib: [],
conditionals: [],
};
const rawLines = content.split('\n');
// Join continuation lines: a line ending with non-blank in col 71 (0-indexed)
// followed by a line starting with // is a continuation.
const lines: Array<{ text: string; lineNum: number }> = [];
let i = 0;
while (i < rawLines.length) {
let line = rawLines[i];
const lineNum = i + 1;
// JCL continuation: if line is exactly 72+ chars and col 72 is non-blank
// and the next line starts with //, join them.
while (
i + 1 < rawLines.length &&
line.length >= 72 &&
line[71] !== ' ' &&
rawLines[i + 1].startsWith('//')
) {
i++;
// Continuation text starts after // and leading spaces
const contText = rawLines[i].substring(2).replace(/^\s+/, ' ');
// Remove the continuation marker (col 72+) from current line
line = line.substring(0, 71).trimEnd() + contText;
}
lines.push({ text: line, lineNum });
i++;
}
let currentJobName = '';
let currentStepName = '';
let inInStreamProc = false;
let inStreamProcName = '';
for (const { text, lineNum } of lines) {
// Skip JCL comments (starting with //* )
if (text.startsWith('//*')) continue;
// Skip non-JCL lines (don't start with //)
if (!text.startsWith('//')) continue;
// PROC definition (in-stream)
const procMatch = text.match(PROC_RE);
if (procMatch) {
const procName = procMatch[1] || inStreamProcName;
if (procName) {
results.procs.push({ name: procName.toUpperCase(), line: lineNum, isInStream: true });
}
inInStreamProc = true;
inStreamProcName = procName?.toUpperCase() || '';
continue;
}
// PEND (end of in-stream proc)
if (PEND_RE.test(text)) {
inInStreamProc = false;
inStreamProcName = '';
continue;
}
// JCLLIB ORDER=
const jcllibMatch = text.match(JCLLIB_RE);
if (jcllibMatch) {
const libs = jcllibMatch[1].split(',').map(s => s.trim().replace(/'/g, ''));
results.jcllib.push({ order: libs, line: lineNum });
continue;
}
// IF/ELSE/ENDIF
const ifMatch = text.match(IF_RE);
if (ifMatch) {
results.conditionals.push({ type: 'IF', condition: ifMatch[1].trim(), line: lineNum });
continue;
}
if (ELSE_RE.test(text)) {
results.conditionals.push({ type: 'ELSE', line: lineNum });
continue;
}
if (ENDIF_RE.test(text)) {
results.conditionals.push({ type: 'ENDIF', line: lineNum });
continue;
}
// INCLUDE MEMBER=
const includeMatch = text.match(INCLUDE_RE);
if (includeMatch) {
results.includes.push({ member: includeMatch[1].toUpperCase(), line: lineNum });
continue;
}
// SET var=value
const setMatch = text.match(SET_RE);
if (setMatch) {
results.sets.push({
variable: setMatch[1].toUpperCase(),
value: setMatch[2].trim().replace(/,\s*$/, ''),
line: lineNum,
});
continue;
}
// JOB statement
const jobMatch = text.match(JOB_RE);
if (jobMatch) {
currentJobName = jobMatch[1].toUpperCase();
const params = jobMatch[2];
results.jobs.push({
name: currentJobName,
line: lineNum,
class: extractParam(params, 'CLASS'),
msgclass: extractParam(params, 'MSGCLASS'),
});
continue;
}
// EXEC statement
const execMatch = text.match(EXEC_RE);
if (execMatch) {
currentStepName = execMatch[1].toUpperCase();
const params = execMatch[2];
const pgm = extractPgm(params);
const proc = pgm ? undefined : extractProc(params);
results.steps.push({
name: currentStepName,
jobName: currentJobName,
program: pgm?.toUpperCase(),
proc: proc?.toUpperCase(),
line: lineNum,
});
continue;
}
// DD statement
const ddMatch = text.match(DD_RE);
if (ddMatch) {
const ddName = ddMatch[1].toUpperCase();
const params = ddMatch[2];
results.ddStatements.push({
ddName,
stepName: currentStepName,
dataset: extractDsn(params)?.toUpperCase(),
disp: extractDisp(params)?.toUpperCase(),
line: lineNum,
});
continue;
}
}
return results;
}
@@ -0,0 +1,264 @@
/**
* JCL Processor — Converts JCL parse results into graph nodes and edges.
*
* Maps JCL entities to existing graph types (no new tables):
* - Job -> CodeElement (description: "jcl-job class:A msgclass:X")
* - Step -> CodeElement (description: "jcl-step pgm:PROGRAMNAME")
* - Dataset -> CodeElement (description: "jcl-dataset disp:SHR")
* - PROC -> Module
*
* Edges:
* - Job CONTAINS Step
* - Step CALLS Module (when PGM= matches an indexed program)
* - Step references Dataset (CALLS edge with reason "jcl-dd")
* - Job/Step IMPORTS PROC
*
* Pattern follows detectCrossProgamContracts() in pipeline.ts.
*/
import { parseJcl, type JclParseResults } from './jcl-parser.js';
import type { KnowledgeGraph } from '../../graph/types.js';
import { generateId } from '../../../lib/utils.js';
export interface JclProcessResult {
jobCount: number;
stepCount: number;
datasetCount: number;
programLinks: number;
}
/**
* Process JCL files and integrate into the knowledge graph.
*
* @param graph - The in-memory knowledge graph
* @param jclPaths - File paths of JCL files
* @param jclContents - Map of path -> file content
* @returns Summary of what was added
*/
export function processJclFiles(
graph: KnowledgeGraph,
jclPaths: string[],
jclContents: Map<string, string>,
): JclProcessResult {
let jobCount = 0;
let stepCount = 0;
let datasetCount = 0;
let programLinks = 0;
// Collect all Module names for step -> program linking
const moduleNames = new Map<string, string>(); // uppercase name -> node id
graph.forEachNode(node => {
if (node.label === 'Module') {
moduleNames.set(node.properties.name?.toUpperCase(), node.id);
}
});
for (const filePath of jclPaths) {
const content = jclContents.get(filePath);
if (!content) continue;
const parsed = parseJcl(content, filePath);
const result = integrateJclResults(graph, parsed, filePath, moduleNames);
jobCount += result.jobCount;
stepCount += result.stepCount;
datasetCount += result.datasetCount;
programLinks += result.programLinks;
}
return { jobCount, stepCount, datasetCount, programLinks };
}
function integrateJclResults(
graph: KnowledgeGraph,
parsed: JclParseResults,
filePath: string,
moduleNames: Map<string, string>,
): JclProcessResult {
let jobCount = 0;
let stepCount = 0;
let datasetCount = 0;
let programLinks = 0;
// Track step node IDs for DD -> step linking
const stepNodeIds = new Map<string, string>(); // stepName -> nodeId
// 1. Create Job nodes
for (const job of parsed.jobs) {
const jobId = generateId('CodeElement', `${filePath}:job:${job.name}`);
const classPart = job.class ? ` class:${job.class}` : '';
const msgPart = job.msgclass ? ` msgclass:${job.msgclass}` : '';
graph.addNode({
id: jobId,
label: 'CodeElement',
properties: {
name: job.name,
filePath,
startLine: job.line,
endLine: job.line,
description: `jcl-job${classPart}${msgPart}`,
},
});
// Link File -> Job (CONTAINS)
const fileId = generateId('File', filePath);
graph.addRelationship({
id: `${fileId}_contains_${jobId}`,
type: 'CONTAINS',
sourceId: fileId,
targetId: jobId,
confidence: 1.0,
reason: 'jcl-job',
});
jobCount++;
}
// 2. Create Step nodes and link to programs
for (const step of parsed.steps) {
const stepId = generateId('CodeElement', `${filePath}:step:${step.jobName}:${step.name}`);
const pgmPart = step.program ? ` pgm:${step.program}` : '';
const procPart = step.proc ? ` proc:${step.proc}` : '';
graph.addNode({
id: stepId,
label: 'CodeElement',
properties: {
name: step.name,
filePath,
startLine: step.line,
endLine: step.line,
description: `jcl-step${pgmPart}${procPart}`,
},
});
stepNodeIds.set(step.name, stepId);
// Link Job -> Step (CONTAINS)
if (step.jobName) {
const jobId = generateId('CodeElement', `${filePath}:job:${step.jobName}`);
graph.addRelationship({
id: `${jobId}_contains_${stepId}`,
type: 'CONTAINS',
sourceId: jobId,
targetId: stepId,
confidence: 1.0,
reason: 'jcl-step',
});
}
// Link Step -> Module (CALLS) when PGM= matches an indexed program
if (step.program) {
const moduleId = moduleNames.get(step.program.toUpperCase());
if (moduleId) {
graph.addRelationship({
id: `${stepId}_calls_${moduleId}`,
type: 'CALLS',
sourceId: stepId,
targetId: moduleId,
confidence: 0.95,
reason: 'jcl-exec-pgm',
});
programLinks++;
}
}
// Link Step -> PROC (CALLS) — PROC as Module
if (step.proc) {
const procModuleId = moduleNames.get(step.proc.toUpperCase());
if (procModuleId) {
graph.addRelationship({
id: `${stepId}_calls_proc_${procModuleId}`,
type: 'CALLS',
sourceId: stepId,
targetId: procModuleId,
confidence: 0.9,
reason: 'jcl-exec-proc',
});
}
}
stepCount++;
}
// 3. Create Dataset nodes from DD statements
const seenDatasets = new Set<string>();
for (const dd of parsed.ddStatements) {
if (!dd.dataset) continue;
// Create dataset node (deduplicated per file)
const datasetKey = `${filePath}:dataset:${dd.dataset}`;
const datasetId = generateId('CodeElement', datasetKey);
if (!seenDatasets.has(dd.dataset)) {
const dispPart = dd.disp ? ` disp:${dd.disp}` : '';
graph.addNode({
id: datasetId,
label: 'CodeElement',
properties: {
name: dd.dataset,
filePath,
startLine: dd.line,
endLine: dd.line,
description: `jcl-dataset${dispPart}`,
},
});
seenDatasets.add(dd.dataset);
datasetCount++;
}
// Link Step -> Dataset (CALLS with reason jcl-dd)
const stepId = stepNodeIds.get(dd.stepName);
if (stepId) {
graph.addRelationship({
id: `${stepId}_dd_${dd.ddName}_${datasetId}`,
type: 'CALLS',
sourceId: stepId,
targetId: datasetId,
confidence: 0.85,
reason: `jcl-dd:${dd.ddName}`,
});
}
}
// 4. Create PROC nodes (in-stream procs as Module)
for (const proc of parsed.procs) {
if (!proc.isInStream) continue;
const procId = generateId('Module', `${filePath}:proc:${proc.name}`);
graph.addNode({
id: procId,
label: 'Module',
properties: {
name: proc.name,
filePath,
startLine: proc.line,
endLine: proc.line,
description: 'jcl-proc-instream',
},
});
// Register for step linking
moduleNames.set(proc.name.toUpperCase(), procId);
}
// 5. INCLUDE directives -> IMPORTS edges
for (const inc of parsed.includes) {
const moduleId = moduleNames.get(inc.member.toUpperCase());
if (moduleId) {
const fileId = generateId('File', filePath);
graph.addRelationship({
id: `${fileId}_includes_${moduleId}`,
type: 'IMPORTS',
sourceId: fileId,
targetId: moduleId,
confidence: 0.9,
reason: 'jcl-include',
});
}
}
return { jobCount, stepCount, datasetCount, programLinks };
}
+29
View File
@@ -1,6 +1,7 @@
import { createKnowledgeGraph } from '../graph/graph.js';
import { processStructure } from './structure-processor.js';
import { processMarkdown } from './markdown-processor.js';
import { processCobol, isCobolFile, isJclFile } from './cobol-processor.js';
import { processParsing } from './parsing-processor.js';
import {
processImports,
@@ -458,6 +459,14 @@ async function runScanAndStructure(
stats: { filesProcessed: totalFiles, totalFiles, nodesCreated: graph.nodeCount },
});
// ── Custom (non-tree-sitter) processors ─────────────────────────────
// Each custom processor follows the pattern in markdown-processor.ts:
// 1. Export a process function: (graph, files, allPathSet) => result
// 2. Export a file detection function: (path) => boolean
// 3. Filter files by extension, write nodes/edges directly to graph
// To add a new language: create a new processor file, import it here,
// and add a filter-read-call-log block following the pattern below.
// ── Phase 2.5: Markdown processing (headings + cross-links) ────────
const mdScanned = scannedFiles.filter(f => f.path.endsWith('.md') || f.path.endsWith('.mdx'));
if (mdScanned.length > 0) {
@@ -472,6 +481,26 @@ async function runScanAndStructure(
}
}
// ── Phase 2.6: COBOL processing (regex extraction, no tree-sitter) ──
const cobolScanned = scannedFiles.filter(f => isCobolFile(f.path) || isJclFile(f.path));
if (cobolScanned.length > 0) {
const cobolContents = await readFileContents(repoPath, cobolScanned.map(f => f.path));
const cobolFiles = cobolScanned
.filter(f => cobolContents.has(f.path))
.map(f => ({ path: f.path, content: cobolContents.get(f.path)! }));
const allPathSet = new Set(allPaths);
const cobolResult = processCobol(graph, cobolFiles, allPathSet);
if (isDev) {
console.log(` COBOL: ${cobolResult.programs} programs, ${cobolResult.paragraphs} paragraphs, ${cobolResult.sections} sections from ${cobolFiles.length} files`);
if (cobolResult.execSqlBlocks > 0 || cobolResult.execCicsBlocks > 0 || cobolResult.entryPoints > 0) {
console.log(` COBOL enriched: ${cobolResult.execSqlBlocks} SQL blocks, ${cobolResult.execCicsBlocks} CICS blocks, ${cobolResult.entryPoints} entry points, ${cobolResult.moves} moves, ${cobolResult.fileDeclarations} file declarations`);
}
if (cobolResult.jclJobs > 0) {
console.log(` JCL: ${cobolResult.jclJobs} jobs, ${cobolResult.jclSteps} steps`);
}
}
}
return { scannedFiles, allPaths, totalFiles };
}
@@ -0,0 +1,25 @@
IDENTIFICATION DIVISION.
PROGRAM-ID. AUDITLOG.
DATA DIVISION.
WORKING-STORAGE SECTION.
01 WS-LOG-MESSAGE PIC X(80).
01 WS-TIMESTAMP PIC X(26).
LINKAGE SECTION.
01 LS-CUST-ID PIC 9(8).
01 LS-AMOUNT PIC 9(7)V99.
PROCEDURE DIVISION USING LS-CUST-ID LS-AMOUNT.
MAIN-PARAGRAPH.
PERFORM WRITE-LOG
GOBACK.
WRITE-LOG.
STRING 'Customer ' LS-CUST-ID ' amount ' LS-AMOUNT
DELIMITED BY SIZE INTO WS-LOG-MESSAGE
DISPLAY WS-LOG-MESSAGE.
ENTRY "AUDITLOG-BATCH" USING LS-CUST-ID.
DISPLAY 'Batch audit for ' LS-CUST-ID
GOBACK.
@@ -0,0 +1,6 @@
01 WS-CUSTOMER-DATA.
05 WS-CUST-CODE PIC X(10).
05 WS-CUST-TYPE PIC X(3).
88 PREMIUM-CUSTOMER VALUE 'PRM'.
88 REGULAR-CUSTOMER VALUE 'REG'.
05 WS-CUST-ADDR PIC X(50).
@@ -0,0 +1,60 @@
IDENTIFICATION DIVISION.
PROGRAM-ID. CUSTUPDT.
AUTHOR. TEST.
ENVIRONMENT DIVISION.
INPUT-OUTPUT SECTION.
FILE-CONTROL.
SELECT CUSTOMER-FILE ASSIGN TO 'CUSTFILE'
ORGANIZATION IS INDEXED
ACCESS IS DYNAMIC
RECORD KEY IS CUST-ID
FILE STATUS IS WS-FILE-STATUS.
DATA DIVISION.
FILE SECTION.
FD CUSTOMER-FILE.
01 CUSTOMER-RECORD.
05 CUST-ID PIC 9(8).
05 CUST-NAME PIC X(30).
05 CUST-BALANCE PIC 9(7)V99.
WORKING-STORAGE SECTION.
01 WS-FILE-STATUS PIC XX.
01 WS-CUSTOMER-NAME PIC X(30).
01 WS-AMOUNT PIC 9(7)V99.
01 WS-EOF PIC 9 VALUE 0.
88 END-OF-FILE VALUE 1.
PROCEDURE DIVISION.
INIT-SECTION SECTION.
MAIN-PARAGRAPH.
PERFORM INIT-PARAGRAPH
PERFORM PROCESS-PARAGRAPH
PERFORM CLEANUP-PARAGRAPH
STOP RUN.
INIT-PARAGRAPH.
OPEN I-O CUSTOMER-FILE
MOVE SPACES TO WS-CUSTOMER-NAME.
PROCESSING-SECTION SECTION.
PROCESS-PARAGRAPH.
PERFORM READ-CUSTOMER THRU WRITE-CUSTOMER
CALL "AUDITLOG" USING CUST-ID WS-AMOUNT.
READ-CUSTOMER.
READ CUSTOMER-FILE
NOT AT END
MOVE CUST-NAME TO WS-CUSTOMER-NAME
END-READ.
UPDATE-BALANCE.
ADD WS-AMOUNT TO CUST-BALANCE
MOVE WS-AMOUNT TO CUST-BALANCE.
WRITE-CUSTOMER.
REWRITE CUSTOMER-RECORD.
CLEANUP-PARAGRAPH.
CLOSE CUSTOMER-FILE.
@@ -0,0 +1,40 @@
IDENTIFICATION DIVISION.
PROGRAM-ID. RPTGEN.
DATA DIVISION.
WORKING-STORAGE SECTION.
COPY CUSTDAT.
01 WS-REPORT-LINE PIC X(132).
01 WS-SQL-CODE PIC S9(9) COMP.
PROCEDURE DIVISION.
MAIN-PARAGRAPH.
PERFORM FETCH-DATA
PERFORM FORMAT-REPORT
PERFORM SEND-SCREEN
CALL "CUSTUPDT"
STOP RUN.
FETCH-DATA.
EXEC SQL
SELECT CUST_NAME, CUST_BALANCE
FROM CUSTOMER
WHERE CUST_ID = :WS-CUST-CODE
END-EXEC.
FORMAT-REPORT.
MOVE WS-CUST-CODE TO WS-REPORT-LINE
PERFORM MAIN-PARAGRAPH THRU FORMAT-REPORT.
SEND-SCREEN.
EXEC CICS
SEND MAP('CUSTRPT') MAPSET('CUSTSET')
END-EXEC.
EXEC CICS
LINK PROGRAM('AUDITLOG')
END-EXEC.
EXEC CICS
XCTL PROGRAM('CUSTUPDT')
END-EXEC.
@@ -0,0 +1,5 @@
//CUSTJOB JOB (ACCT),'CUSTOMER UPDATE',CLASS=A,MSGCLASS=X
//STEP1 EXEC PGM=CUSTUPDT
//CUSTFILE DD DSN=PROD.CUSTOMER.MASTER,DISP=SHR
//STEP2 EXEC PGM=RPTGEN
//SYSOUT DD SYSOUT=*
@@ -0,0 +1,760 @@
/**
* COBOL: Exhaustive strict integration test.
*
* Every single node and edge produced by the COBOL/JCL pipeline is asserted
* with exact counts and exact sorted lists. No fuzzy assertions.
*
* Ground truth captured from the cobol-app fixture:
* CUSTUPDT.cbl — 3 programs, 2 sections, 13 paragraphs, 21 data items,
* AUDITLOG.cbl 1 file declaration, 1 COPY, 1 EXEC SQL, 3 EXEC CICS,
* RPTGEN.cbl 1 ENTRY point, 3 MOVE pairs, 2 JCL jobs, 2 JCL steps,
* CUSTDAT.cpy 1 JCL dataset, cross-program CALL/LINK/XCTL resolution.
* RUNJOBS.jcl
*/
import { describe, it, expect, beforeAll } from 'vitest';
import path from 'path';
import {
FIXTURES, getRelationships, getNodesByLabel, edgeSet,
runPipelineFromRepo, type PipelineResult,
} from './helpers.js';
describe('COBOL full system extraction', () => {
let result: PipelineResult;
beforeAll(async () => {
result = await runPipelineFromRepo(
path.join(FIXTURES, 'cobol-app'),
() => {},
{ skipGraphPhases: true }, // COBOL is regex-based, not in SupportedLanguages enum
);
}, 60000);
// =====================================================================
// NODE COMPLETENESS -- assert exact count and exact sorted list per label
// =====================================================================
describe('node completeness', () => {
it('produces exactly 3 Module nodes', () => {
const modules = getNodesByLabel(result, 'Module');
expect(modules.length).toBe(3);
expect(modules).toEqual(['AUDITLOG', 'CUSTUPDT', 'RPTGEN']);
});
it('produces exactly 13 Function nodes (paragraphs across all programs)', () => {
const funcs = getNodesByLabel(result, 'Function');
expect(funcs.length).toBe(13);
// getNodesByLabel returns sorted names; MAIN-PARAGRAPH appears 3 times
// (once per program: CUSTUPDT, RPTGEN, AUDITLOG — separate graph nodes
// with different filePaths but same name, all returned by getNodesByLabel)
expect(funcs).toEqual([
'CLEANUP-PARAGRAPH', // CUSTUPDT
'FETCH-DATA', // RPTGEN
'FORMAT-REPORT', // RPTGEN
'INIT-PARAGRAPH', // CUSTUPDT
'MAIN-PARAGRAPH', // AUDITLOG
'MAIN-PARAGRAPH', // CUSTUPDT
'MAIN-PARAGRAPH', // RPTGEN
'PROCESS-PARAGRAPH', // CUSTUPDT
'READ-CUSTOMER', // CUSTUPDT
'SEND-SCREEN', // RPTGEN
'UPDATE-BALANCE', // CUSTUPDT
'WRITE-CUSTOMER', // CUSTUPDT
'WRITE-LOG', // AUDITLOG
]);
});
it('produces exactly 2 Namespace nodes (PROCEDURE DIVISION sections)', () => {
const ns = getNodesByLabel(result, 'Namespace');
expect(ns.length).toBe(2);
expect(ns).toEqual(['INIT-SECTION', 'PROCESSING-SECTION']);
});
it('produces exactly 21 Property nodes (data items + 88-levels)', () => {
const props = getNodesByLabel(result, 'Property');
expect(props.length).toBe(21);
expect(props).toEqual([
'CUST-BALANCE',
'CUST-ID',
'CUST-NAME',
'CUSTOMER-RECORD',
'END-OF-FILE',
'LS-AMOUNT',
'LS-CUST-ID',
'PREMIUM-CUSTOMER',
'REGULAR-CUSTOMER',
'WS-AMOUNT',
'WS-CUST-ADDR',
'WS-CUST-CODE',
'WS-CUST-TYPE',
'WS-CUSTOMER-DATA',
'WS-CUSTOMER-NAME',
'WS-EOF',
'WS-FILE-STATUS',
'WS-LOG-MESSAGE',
'WS-REPORT-LINE',
'WS-SQL-CODE',
'WS-TIMESTAMP',
]);
});
it('produces exactly 1 Record node (file declaration)', () => {
const records = getNodesByLabel(result, 'Record');
expect(records.length).toBe(1);
expect(records).toEqual(['CUSTOMER-FILE']);
});
it('produces exactly 8 CodeElement nodes (EXEC blocks + JCL entities)', () => {
const ce = getNodesByLabel(result, 'CodeElement');
expect(ce.length).toBe(8);
expect(ce).toEqual([
'CUSTJOB',
'EXEC CICS LINK',
'EXEC CICS SEND MAP',
'EXEC CICS XCTL',
'EXEC SQL SELECT',
'PROD.CUSTOMER.MASTER',
'STEP1',
'STEP2',
]);
});
it('produces exactly 1 Constructor node (ENTRY point)', () => {
const constructors = getNodesByLabel(result, 'Constructor');
expect(constructors.length).toBe(1);
expect(constructors).toEqual(['AUDITLOG-BATCH']);
});
});
// =====================================================================
// EDGE COMPLETENESS -- assert exact count and exact pairs per type+reason
// =====================================================================
describe('edge completeness', () => {
// -- ACCESSES edges -------------------------------------------------
it('produces exactly 3 ACCESSES edges with reason cobol-move-read', () => {
const edges = getRelationships(result, 'ACCESSES')
.filter(e => e.rel.reason === 'cobol-move-read');
expect(edges.length).toBe(3);
expect(edgeSet(edges)).toEqual([
'FORMAT-REPORT \u2192 WS-CUST-CODE',
'READ-CUSTOMER \u2192 CUST-NAME',
'UPDATE-BALANCE \u2192 WS-AMOUNT',
]);
});
it('produces exactly 3 ACCESSES edges with reason cobol-move-write', () => {
const edges = getRelationships(result, 'ACCESSES')
.filter(e => e.rel.reason === 'cobol-move-write');
expect(edges.length).toBe(3);
expect(edgeSet(edges)).toEqual([
'FORMAT-REPORT \u2192 WS-REPORT-LINE',
'READ-CUSTOMER \u2192 WS-CUSTOMER-NAME',
'UPDATE-BALANCE \u2192 CUST-BALANCE',
]);
});
it('produces exactly 1 ACCESSES edge with reason sql-select (synthetic target)', () => {
// The sql-select edge targets a synthetic Record node (<db>:CUSTOMER) that
// is not materialized in the graph. We verify by filtering on reason only,
// since getRelationships resolves sourceId/targetId to node names when nodes exist.
const allAccesses = getRelationships(result, 'ACCESSES');
const sqlAccesses = allAccesses.filter(e => e.rel.reason === 'sql-select');
expect(sqlAccesses.length).toBe(1);
expect(sqlAccesses[0].source).toBe('EXEC SQL SELECT');
});
it('produces exactly 7 total ACCESSES edges', () => {
const edges = getRelationships(result, 'ACCESSES');
expect(edges.length).toBe(7);
});
// -- CALLS edges: cobol-perform -------------------------------------
it('produces exactly 9 CALLS edges with reason cobol-perform', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.rel.reason === 'cobol-perform');
expect(edges.length).toBe(9);
expect(edgeSet(edges)).toEqual([
'FORMAT-REPORT \u2192 MAIN-PARAGRAPH',
'MAIN-PARAGRAPH \u2192 CLEANUP-PARAGRAPH',
'MAIN-PARAGRAPH \u2192 FETCH-DATA',
'MAIN-PARAGRAPH \u2192 FORMAT-REPORT',
'MAIN-PARAGRAPH \u2192 INIT-PARAGRAPH',
'MAIN-PARAGRAPH \u2192 PROCESS-PARAGRAPH',
'MAIN-PARAGRAPH \u2192 SEND-SCREEN',
'MAIN-PARAGRAPH \u2192 WRITE-LOG',
'PROCESS-PARAGRAPH \u2192 READ-CUSTOMER',
]);
});
// -- CALLS edges: cobol-perform-thru --------------------------------
it('produces exactly 2 CALLS edges with reason cobol-perform-thru', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.rel.reason === 'cobol-perform-thru');
expect(edges.length).toBe(2);
expect(edgeSet(edges)).toEqual([
'FORMAT-REPORT \u2192 FORMAT-REPORT',
'PROCESS-PARAGRAPH \u2192 WRITE-CUSTOMER',
]);
});
// -- CALLS edges: cobol-call ----------------------------------------
it('produces exactly 2 CALLS edges with reason cobol-call', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.rel.reason === 'cobol-call');
expect(edges.length).toBe(2);
expect(edgeSet(edges)).toEqual([
'CUSTUPDT \u2192 AUDITLOG',
'RPTGEN \u2192 CUSTUPDT',
]);
});
// -- CALLS edges: cics-link -----------------------------------------
it('produces exactly 1 CALLS edge with reason cics-link', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.rel.reason === 'cics-link');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual([
'RPTGEN \u2192 AUDITLOG',
]);
});
// -- CALLS edges: cics-xctl -----------------------------------------
it('produces exactly 1 CALLS edge with reason cics-xctl', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.rel.reason === 'cics-xctl');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual([
'RPTGEN \u2192 CUSTUPDT',
]);
});
// -- CALLS edges: unresolved (retained from first pass) ---------------
it('produces exactly 2 CALLS edges with reason cobol-call-unresolved', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.rel.reason === 'cobol-call-unresolved');
expect(edges.length).toBe(2);
// CUSTUPDT -> AUDITLOG and RPTGEN -> CUSTUPDT were initially unresolved
// because the target module had not yet been processed at the time.
// The second pass adds resolved edges but does NOT remove these.
expect(edges.map(e => e.source).sort()).toEqual(['CUSTUPDT', 'RPTGEN']);
});
it('produces exactly 1 CALLS edge with reason cics-link-unresolved', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.rel.reason === 'cics-link-unresolved');
expect(edges.length).toBe(1);
expect(edges[0].source).toBe('RPTGEN');
});
it('produces exactly 1 CALLS edge with reason cics-xctl-unresolved', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.rel.reason === 'cics-xctl-unresolved');
expect(edges.length).toBe(1);
expect(edges[0].source).toBe('RPTGEN');
});
// -- CALLS edges: jcl-exec-pgm --------------------------------------
it('produces exactly 2 CALLS edges with reason jcl-exec-pgm', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.rel.reason === 'jcl-exec-pgm');
expect(edges.length).toBe(2);
expect(edgeSet(edges)).toEqual([
'STEP1 \u2192 CUSTUPDT',
'STEP2 \u2192 RPTGEN',
]);
});
// -- CALLS edges: jcl-dd:CUSTFILE -----------------------------------
it('produces exactly 1 CALLS edge with reason jcl-dd:CUSTFILE', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.rel.reason === 'jcl-dd:CUSTFILE');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual([
'STEP1 \u2192 PROD.CUSTOMER.MASTER',
]);
});
// -- CONTAINS edges: cobol-program-id -------------------------------
it('produces exactly 3 CONTAINS edges with reason cobol-program-id', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.rel.reason === 'cobol-program-id');
expect(edges.length).toBe(3);
expect(edgeSet(edges)).toEqual([
'AUDITLOG.cbl \u2192 AUDITLOG',
'CUSTUPDT.cbl \u2192 CUSTUPDT',
'RPTGEN.cbl \u2192 RPTGEN',
]);
});
// -- CONTAINS edges: cobol-section ----------------------------------
it('produces exactly 2 CONTAINS edges with reason cobol-section', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.rel.reason === 'cobol-section');
expect(edges.length).toBe(2);
expect(edgeSet(edges)).toEqual([
'CUSTUPDT \u2192 INIT-SECTION',
'CUSTUPDT \u2192 PROCESSING-SECTION',
]);
});
// -- CONTAINS edges: cobol-paragraph --------------------------------
it('produces exactly 13 CONTAINS edges with reason cobol-paragraph', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.rel.reason === 'cobol-paragraph');
expect(edges.length).toBe(13);
expect(edgeSet(edges)).toEqual([
'AUDITLOG \u2192 MAIN-PARAGRAPH',
'AUDITLOG \u2192 WRITE-LOG',
'INIT-SECTION \u2192 INIT-PARAGRAPH',
'INIT-SECTION \u2192 MAIN-PARAGRAPH',
'PROCESSING-SECTION \u2192 CLEANUP-PARAGRAPH',
'PROCESSING-SECTION \u2192 PROCESS-PARAGRAPH',
'PROCESSING-SECTION \u2192 READ-CUSTOMER',
'PROCESSING-SECTION \u2192 UPDATE-BALANCE',
'PROCESSING-SECTION \u2192 WRITE-CUSTOMER',
'RPTGEN \u2192 FETCH-DATA',
'RPTGEN \u2192 FORMAT-REPORT',
'RPTGEN \u2192 MAIN-PARAGRAPH',
'RPTGEN \u2192 SEND-SCREEN',
]);
});
// -- CONTAINS edges: cobol-data-item --------------------------------
it('produces exactly 21 CONTAINS edges with reason cobol-data-item', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.rel.reason === 'cobol-data-item');
expect(edges.length).toBe(21);
expect(edgeSet(edges)).toEqual([
'AUDITLOG \u2192 LS-AMOUNT',
'AUDITLOG \u2192 LS-CUST-ID',
'AUDITLOG \u2192 WS-LOG-MESSAGE',
'AUDITLOG \u2192 WS-TIMESTAMP',
'CUSTUPDT \u2192 CUST-BALANCE',
'CUSTUPDT \u2192 CUST-ID',
'CUSTUPDT \u2192 CUST-NAME',
'CUSTUPDT \u2192 CUSTOMER-RECORD',
'CUSTUPDT \u2192 END-OF-FILE',
'CUSTUPDT \u2192 WS-AMOUNT',
'CUSTUPDT \u2192 WS-CUSTOMER-NAME',
'CUSTUPDT \u2192 WS-EOF',
'CUSTUPDT \u2192 WS-FILE-STATUS',
'RPTGEN \u2192 PREMIUM-CUSTOMER',
'RPTGEN \u2192 REGULAR-CUSTOMER',
'RPTGEN \u2192 WS-CUST-ADDR',
'RPTGEN \u2192 WS-CUST-CODE',
'RPTGEN \u2192 WS-CUST-TYPE',
'RPTGEN \u2192 WS-CUSTOMER-DATA',
'RPTGEN \u2192 WS-REPORT-LINE',
'RPTGEN \u2192 WS-SQL-CODE',
]);
});
// -- CONTAINS edges: cobol-exec-sql ---------------------------------
it('produces exactly 1 CONTAINS edge with reason cobol-exec-sql', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.rel.reason === 'cobol-exec-sql');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual([
'RPTGEN \u2192 EXEC SQL SELECT',
]);
});
// -- CONTAINS edges: cobol-exec-cics --------------------------------
it('produces exactly 3 CONTAINS edges with reason cobol-exec-cics', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.rel.reason === 'cobol-exec-cics');
expect(edges.length).toBe(3);
expect(edgeSet(edges)).toEqual([
'RPTGEN \u2192 EXEC CICS LINK',
'RPTGEN \u2192 EXEC CICS SEND MAP',
'RPTGEN \u2192 EXEC CICS XCTL',
]);
});
// -- CONTAINS edges: cobol-entry-point ------------------------------
it('produces exactly 1 CONTAINS edge with reason cobol-entry-point', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.rel.reason === 'cobol-entry-point');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual([
'AUDITLOG \u2192 AUDITLOG-BATCH',
]);
});
// -- CONTAINS edges: cobol-file-declaration -------------------------
it('produces exactly 1 CONTAINS edge with reason cobol-file-declaration', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.rel.reason === 'cobol-file-declaration');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual([
'CUSTUPDT \u2192 CUSTOMER-FILE',
]);
});
// -- CONTAINS edges: jcl-job ----------------------------------------
it('produces exactly 1 CONTAINS edge with reason jcl-job', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.rel.reason === 'jcl-job');
expect(edges.length).toBe(1);
expect(edgeSet(edges)).toEqual([
'RUNJOBS.jcl \u2192 CUSTJOB',
]);
});
// -- CONTAINS edges: jcl-step ---------------------------------------
it('produces exactly 2 CONTAINS edges with reason jcl-step', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.rel.reason === 'jcl-step');
expect(edges.length).toBe(2);
expect(edgeSet(edges)).toEqual([
'CUSTJOB \u2192 STEP1',
'CUSTJOB \u2192 STEP2',
]);
});
// -- IMPORTS edges: cobol-copy --------------------------------------
it('produces exactly 1 IMPORTS edge with reason cobol-copy', () => {
const edges = getRelationships(result, 'IMPORTS')
.filter(e => e.rel.reason === 'cobol-copy');
expect(edges.length).toBe(1);
expect(edges[0].sourceFilePath).toMatch(/RPTGEN\.cbl$/);
expect(edges[0].targetFilePath).toMatch(/CUSTDAT\.cpy$/);
});
});
// =====================================================================
// CROSS-PROGRAM RESOLUTION -- verify specific resolved edges
// =====================================================================
describe('cross-program resolution', () => {
it('CUSTUPDT CALL "AUDITLOG" resolves to Module node', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.source === 'CUSTUPDT' && e.target === 'AUDITLOG' && e.rel.reason === 'cobol-call');
expect(edges.length).toBe(1);
expect(edges[0].sourceLabel).toBe('Module');
expect(edges[0].targetLabel).toBe('Module');
});
it('RPTGEN CALL "CUSTUPDT" resolves to Module node', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.source === 'RPTGEN' && e.target === 'CUSTUPDT' && e.rel.reason === 'cobol-call');
expect(edges.length).toBe(1);
expect(edges[0].sourceLabel).toBe('Module');
expect(edges[0].targetLabel).toBe('Module');
});
it('RPTGEN CICS LINK AUDITLOG resolves to Module node', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.source === 'RPTGEN' && e.target === 'AUDITLOG' && e.rel.reason === 'cics-link');
expect(edges.length).toBe(1);
expect(edges[0].sourceLabel).toBe('Module');
expect(edges[0].targetLabel).toBe('Module');
});
it('RPTGEN CICS XCTL CUSTUPDT resolves to Module node', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.source === 'RPTGEN' && e.target === 'CUSTUPDT' && e.rel.reason === 'cics-xctl');
expect(edges.length).toBe(1);
expect(edges[0].sourceLabel).toBe('Module');
expect(edges[0].targetLabel).toBe('Module');
});
it('JCL STEP1 links to CUSTUPDT Module via jcl-exec-pgm', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.source === 'STEP1' && e.target === 'CUSTUPDT' && e.rel.reason === 'jcl-exec-pgm');
expect(edges.length).toBe(1);
expect(edges[0].sourceLabel).toBe('CodeElement');
expect(edges[0].targetLabel).toBe('Module');
});
it('JCL STEP2 links to RPTGEN Module via jcl-exec-pgm', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.source === 'STEP2' && e.target === 'RPTGEN' && e.rel.reason === 'jcl-exec-pgm');
expect(edges.length).toBe(1);
expect(edges[0].sourceLabel).toBe('CodeElement');
expect(edges[0].targetLabel).toBe('Module');
});
});
// =====================================================================
// COPY EXPANSION -- verify copybook data items appear in host program
// =====================================================================
describe('COPY expansion', () => {
it('RPTGEN IMPORTS CUSTDAT copybook', () => {
const imports = getRelationships(result, 'IMPORTS')
.filter(e => e.rel.reason === 'cobol-copy');
expect(imports.length).toBe(1);
expect(imports[0].sourceFilePath).toMatch(/RPTGEN\.cbl$/);
expect(imports[0].targetFilePath).toMatch(/CUSTDAT\.cpy$/);
});
it('copybook data items appear as Property nodes owned by RPTGEN', () => {
const contains = getRelationships(result, 'CONTAINS')
.filter(e => e.source === 'RPTGEN' && e.rel.reason === 'cobol-data-item');
const targets = contains.map(e => e.target).sort();
expect(targets).toEqual([
'PREMIUM-CUSTOMER',
'REGULAR-CUSTOMER',
'WS-CUST-ADDR',
'WS-CUST-CODE',
'WS-CUST-TYPE',
'WS-CUSTOMER-DATA',
'WS-REPORT-LINE',
'WS-SQL-CODE',
]);
});
});
// =====================================================================
// SECTION-TO-PARAGRAPH HIERARCHY -- exact structure
// =====================================================================
describe('section-to-paragraph hierarchy', () => {
it('INIT-SECTION contains exactly 2 paragraphs', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.source === 'INIT-SECTION' && e.rel.reason === 'cobol-paragraph');
expect(edges.length).toBe(2);
expect(edges.map(e => e.target).sort()).toEqual([
'INIT-PARAGRAPH',
'MAIN-PARAGRAPH',
]);
});
it('PROCESSING-SECTION contains exactly 5 paragraphs', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.source === 'PROCESSING-SECTION' && e.rel.reason === 'cobol-paragraph');
expect(edges.length).toBe(5);
expect(edges.map(e => e.target).sort()).toEqual([
'CLEANUP-PARAGRAPH',
'PROCESS-PARAGRAPH',
'READ-CUSTOMER',
'UPDATE-BALANCE',
'WRITE-CUSTOMER',
]);
});
it('RPTGEN (no sections) contains exactly 4 paragraphs directly', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.source === 'RPTGEN' && e.rel.reason === 'cobol-paragraph');
expect(edges.length).toBe(4);
expect(edges.map(e => e.target).sort()).toEqual([
'FETCH-DATA',
'FORMAT-REPORT',
'MAIN-PARAGRAPH',
'SEND-SCREEN',
]);
});
it('AUDITLOG (no sections) contains exactly 2 paragraphs directly', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.source === 'AUDITLOG' && e.rel.reason === 'cobol-paragraph');
expect(edges.length).toBe(2);
expect(edges.map(e => e.target).sort()).toEqual([
'MAIN-PARAGRAPH',
'WRITE-LOG',
]);
});
});
// =====================================================================
// DATA ITEM OWNERSHIP -- exact per-module breakdown
// =====================================================================
describe('data item ownership', () => {
it('CUSTUPDT owns exactly 9 data items', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.source === 'CUSTUPDT' && e.rel.reason === 'cobol-data-item');
expect(edges.length).toBe(9);
expect(edges.map(e => e.target).sort()).toEqual([
'CUST-BALANCE',
'CUST-ID',
'CUST-NAME',
'CUSTOMER-RECORD',
'END-OF-FILE',
'WS-AMOUNT',
'WS-CUSTOMER-NAME',
'WS-EOF',
'WS-FILE-STATUS',
]);
});
it('AUDITLOG owns exactly 4 data items', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.source === 'AUDITLOG' && e.rel.reason === 'cobol-data-item');
expect(edges.length).toBe(4);
expect(edges.map(e => e.target).sort()).toEqual([
'LS-AMOUNT',
'LS-CUST-ID',
'WS-LOG-MESSAGE',
'WS-TIMESTAMP',
]);
});
it('RPTGEN owns exactly 8 data items (including expanded copybook)', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.source === 'RPTGEN' && e.rel.reason === 'cobol-data-item');
expect(edges.length).toBe(8);
expect(edges.map(e => e.target).sort()).toEqual([
'PREMIUM-CUSTOMER',
'REGULAR-CUSTOMER',
'WS-CUST-ADDR',
'WS-CUST-CODE',
'WS-CUST-TYPE',
'WS-CUSTOMER-DATA',
'WS-REPORT-LINE',
'WS-SQL-CODE',
]);
});
});
// =====================================================================
// MOVE DATA FLOW -- exact source->target pairs
// =====================================================================
describe('MOVE data flow', () => {
it('READ-CUSTOMER reads CUST-NAME and writes WS-CUSTOMER-NAME', () => {
const accesses = getRelationships(result, 'ACCESSES');
const reads = accesses.filter(e =>
e.source === 'READ-CUSTOMER' && e.rel.reason === 'cobol-move-read',
);
expect(reads.length).toBe(1);
expect(reads[0].target).toBe('CUST-NAME');
const writes = accesses.filter(e =>
e.source === 'READ-CUSTOMER' && e.rel.reason === 'cobol-move-write',
);
expect(writes.length).toBe(1);
expect(writes[0].target).toBe('WS-CUSTOMER-NAME');
});
it('UPDATE-BALANCE reads WS-AMOUNT and writes CUST-BALANCE', () => {
const accesses = getRelationships(result, 'ACCESSES');
const reads = accesses.filter(e =>
e.source === 'UPDATE-BALANCE' && e.rel.reason === 'cobol-move-read',
);
expect(reads.length).toBe(1);
expect(reads[0].target).toBe('WS-AMOUNT');
const writes = accesses.filter(e =>
e.source === 'UPDATE-BALANCE' && e.rel.reason === 'cobol-move-write',
);
expect(writes.length).toBe(1);
expect(writes[0].target).toBe('CUST-BALANCE');
});
it('FORMAT-REPORT reads WS-CUST-CODE and writes WS-REPORT-LINE', () => {
const accesses = getRelationships(result, 'ACCESSES');
const reads = accesses.filter(e =>
e.source === 'FORMAT-REPORT' && e.rel.reason === 'cobol-move-read',
);
expect(reads.length).toBe(1);
expect(reads[0].target).toBe('WS-CUST-CODE');
const writes = accesses.filter(e =>
e.source === 'FORMAT-REPORT' && e.rel.reason === 'cobol-move-write',
);
expect(writes.length).toBe(1);
expect(writes[0].target).toBe('WS-REPORT-LINE');
});
});
// =====================================================================
// JCL INTEGRATION -- exact structure
// =====================================================================
describe('JCL integration', () => {
it('CUSTJOB job is contained by RUNJOBS.jcl file', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.rel.reason === 'jcl-job');
expect(edges.length).toBe(1);
expect(edges[0].source).toBe('RUNJOBS.jcl');
expect(edges[0].target).toBe('CUSTJOB');
expect(edges[0].sourceLabel).toBe('File');
expect(edges[0].targetLabel).toBe('CodeElement');
});
it('CUSTJOB contains exactly 2 steps', () => {
const edges = getRelationships(result, 'CONTAINS')
.filter(e => e.source === 'CUSTJOB' && e.rel.reason === 'jcl-step');
expect(edges.length).toBe(2);
expect(edges.map(e => e.target).sort()).toEqual(['STEP1', 'STEP2']);
});
it('STEP1 references PROD.CUSTOMER.MASTER dataset via jcl-dd:CUSTFILE', () => {
const edges = getRelationships(result, 'CALLS')
.filter(e => e.rel.reason === 'jcl-dd:CUSTFILE');
expect(edges.length).toBe(1);
expect(edges[0].source).toBe('STEP1');
expect(edges[0].target).toBe('PROD.CUSTOMER.MASTER');
expect(edges[0].sourceLabel).toBe('CodeElement');
expect(edges[0].targetLabel).toBe('CodeElement');
});
});
// =====================================================================
// GRAND TOTALS -- ensure no unexpected edges leak in
// =====================================================================
describe('grand totals', () => {
it('produces exactly 22 total CALLS edges (18 resolved + 4 unresolved)', () => {
// Resolved edges:
// 9 cobol-perform + 2 cobol-perform-thru + 2 cobol-call +
// 1 cics-link + 1 cics-xctl + 2 jcl-exec-pgm + 1 jcl-dd:CUSTFILE = 18
// Unresolved edges (retained from first pass before cross-program resolution):
// 2 cobol-call-unresolved + 1 cics-link-unresolved + 1 cics-xctl-unresolved = 4
// Grand total: 22
const edges = getRelationships(result, 'CALLS');
expect(edges.length).toBe(22);
});
it('produces exactly 48 total CONTAINS edges', () => {
// 3 cobol-program-id + 2 cobol-section + 13 cobol-paragraph +
// 21 cobol-data-item + 1 cobol-exec-sql + 3 cobol-exec-cics +
// 1 cobol-entry-point + 1 cobol-file-declaration +
// 1 jcl-job + 2 jcl-step = 48
const edges = getRelationships(result, 'CONTAINS');
expect(edges.length).toBe(48);
});
it('produces exactly 1 total IMPORTS edge', () => {
const edges = getRelationships(result, 'IMPORTS');
expect(edges.length).toBe(1);
});
it('produces exactly 7 total ACCESSES edges', () => {
// 3 cobol-move-read + 3 cobol-move-write + 1 sql-select = 7
const edges = getRelationships(result, 'ACCESSES');
expect(edges.length).toBe(7);
});
});
});
@@ -0,0 +1,711 @@
import { describe, it, expect } from 'vitest';
import {
preprocessCobolSource,
extractCobolSymbolsWithRegex,
} from '../../src/core/ingestion/cobol/cobol-preprocessor.js';
import type { CobolRegexResults } from '../../src/core/ingestion/cobol/cobol-preprocessor.js';
// ---------------------------------------------------------------------------
// Helper: build COBOL source from an array of lines.
//
// The parser processes full raw lines including columns 1-6 (sequence area).
// Regexes anchored with ^\s+ (data items, FD, AUTHOR, etc.) require the line
// to start with whitespace, so test lines use spaces in cols 1-6 instead of
// numeric sequence numbers unless specifically testing sequence-number behavior.
//
// Column layout:
// 1-6: sequence/patch area (spaces or digits)
// 7: indicator (* comment, - continuation, / page break, space normal)
// 8-11: Area A (divisions, sections, paragraphs start here = 7 leading spaces)
// 12+: Area B (statements = 11+ leading spaces)
// ---------------------------------------------------------------------------
function cobol(...lines: string[]): string {
return lines.join('\n');
}
// ---------------------------------------------------------------------------
// preprocessCobolSource
// ---------------------------------------------------------------------------
describe('preprocessCobolSource', () => {
it('replaces alphabetic patch markers in cols 1-6 with spaces', () => {
const input = cobol(
'mzADD IDENTIFICATION DIVISION.',
'estero PROGRAM-ID. TEST1.',
);
const output = preprocessCobolSource(input);
const lines = output.split('\n');
expect(lines[0].substring(0, 6)).toBe(' ');
expect(lines[0].substring(6)).toBe(' IDENTIFICATION DIVISION.');
expect(lines[1].substring(0, 6)).toBe(' ');
});
it('strips numeric sequence numbers from cols 1-6', () => {
const input = cobol(
'000100 IDENTIFICATION DIVISION.',
'000200 PROGRAM-ID. TEST1.',
);
const output = preprocessCobolSource(input);
const lines = output.split('\n');
expect(lines[0]).toBe(' IDENTIFICATION DIVISION.');
expect(lines[1]).toBe(' PROGRAM-ID. TEST1.');
});
it('preserves lines shorter than 7 characters', () => {
const input = cobol('SHORT', ' ', '000100 IDENTIFICATION DIVISION.');
const output = preprocessCobolSource(input);
const lines = output.split('\n');
expect(lines[0]).toBe('SHORT');
expect(lines[1]).toBe(' ');
});
it('preserves exact line count (no lines added/removed)', () => {
const input = cobol(
'mzADD IDENTIFICATION DIVISION.',
'000200 PROGRAM-ID. TEST1.',
'patch# DATA DIVISION.',
'',
'000500 PROCEDURE DIVISION.',
);
const output = preprocessCobolSource(input);
expect(output.split('\n').length).toBe(input.split('\n').length);
});
});
// ---------------------------------------------------------------------------
// extractCobolSymbolsWithRegex
// ---------------------------------------------------------------------------
describe('extractCobolSymbolsWithRegex', () => {
// -------------------------------------------------------------------------
// PROGRAM-ID
// -------------------------------------------------------------------------
describe('PROGRAM-ID', () => {
it('extracts PROGRAM-ID from IDENTIFICATION DIVISION', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.programName).toBe('TESTPROG');
});
it('returns null programName for content without PROGRAM-ID', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' AUTHOR. SOMEONE.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.programName).toBeNull();
});
});
// -------------------------------------------------------------------------
// Paragraphs & Sections
// -------------------------------------------------------------------------
describe('Paragraphs & Sections', () => {
it('extracts paragraphs in PROCEDURE DIVISION (7 leading spaces)', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' DISPLAY "HELLO".',
' SUB-PARA.',
' DISPLAY "WORLD".',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.paragraphs).toHaveLength(2);
expect(r.paragraphs[0].name).toBe('MAIN-PARA');
expect(r.paragraphs[1].name).toBe('SUB-PARA');
});
it('extracts sections in PROCEDURE DIVISION', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' INIT-SECTION SECTION.',
' INIT-PARA.',
' DISPLAY "INIT".',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.sections).toHaveLength(1);
expect(r.sections[0].name).toBe('INIT-SECTION');
expect(r.paragraphs).toHaveLength(1);
expect(r.paragraphs[0].name).toBe('INIT-PARA');
});
it('excludes reserved names (DECLARATIVES, END, PROCEDURE, etc.)', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' DECLARATIVES.',
' END.',
' REAL-PARA.',
' DISPLAY "OK".',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.paragraphs.map(p => p.name)).toEqual(['REAL-PARA']);
});
it('does NOT treat IDENTIFICATION/ENVIRONMENT/DATA/WORKING-STORAGE as paragraphs', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' ENVIRONMENT DIVISION.',
' DATA DIVISION.',
' WORKING-STORAGE SECTION.',
' PROCEDURE DIVISION.',
' REAL-PARA.',
' DISPLAY "OK".',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
const names = r.paragraphs.map(p => p.name);
expect(names).not.toContain('IDENTIFICATION');
expect(names).not.toContain('ENVIRONMENT');
expect(names).not.toContain('DATA');
expect(names).not.toContain('WORKING-STORAGE');
expect(names).toContain('REAL-PARA');
});
});
// -------------------------------------------------------------------------
// CALL / PERFORM / COPY
// -------------------------------------------------------------------------
describe('CALL / PERFORM / COPY', () => {
it('extracts CALL "PROGRAM" statements', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' CALL "SUBPROG".',
' CALL "ANOTHER".',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.calls).toHaveLength(2);
expect(r.calls[0].target).toBe('SUBPROG');
expect(r.calls[1].target).toBe('ANOTHER');
});
it("extracts CALL 'PROGRAM' statements (single-quoted target)", () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
" CALL 'SUBPROG'.",
" CALL 'ANOTHER'.",
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.calls).toHaveLength(2);
expect(r.calls[0].target).toBe('SUBPROG');
expect(r.calls[1].target).toBe('ANOTHER');
});
it('extracts PERFORM paragraph-name with caller context', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' PERFORM SUB-PARA.',
' SUB-PARA.',
' DISPLAY "HELLO".',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.performs).toHaveLength(1);
expect(r.performs[0].target).toBe('SUB-PARA');
expect(r.performs[0].caller).toBe('MAIN-PARA');
});
it('extracts PERFORM ... THRU ... statements', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' PERFORM STEP-A THRU STEP-Z.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.performs).toHaveLength(1);
expect(r.performs[0].target).toBe('STEP-A');
expect(r.performs[0].thruTarget).toBe('STEP-Z');
});
it('extracts COPY copybook (unquoted)', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' DATA DIVISION.',
' WORKING-STORAGE SECTION.',
' COPY WSCOPY.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.copies).toHaveLength(1);
expect(r.copies[0].target).toBe('WSCOPY');
});
it('extracts COPY "copybook" (quoted)', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' DATA DIVISION.',
' WORKING-STORAGE SECTION.',
' COPY "MY-COPY".',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.copies).toHaveLength(1);
expect(r.copies[0].target).toBe('MY-COPY');
});
it("extracts COPY 'copybook' (single-quoted)", () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' DATA DIVISION.',
' WORKING-STORAGE SECTION.',
" COPY 'MY-COPY'.",
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.copies).toHaveLength(1);
expect(r.copies[0].target).toBe('MY-COPY');
});
it('does NOT store PERFORM UNTIL as a perform target', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' PERFORM UNTIL WS-EOF = 1.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.performs.map(p => p.target)).not.toContain('UNTIL');
});
it('does NOT store PERFORM VARYING as a perform target', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' PERFORM VARYING I FROM 1 BY 1 UNTIL I > 10.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.performs.map(p => p.target)).not.toContain('VARYING');
});
it('does NOT store PERFORM WITH TEST as a perform target', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' PERFORM WITH TEST AFTER UNTIL WS-DONE.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.performs.map(p => p.target)).not.toContain('WITH');
});
});
// -------------------------------------------------------------------------
// Data Division
// -------------------------------------------------------------------------
describe('Data Division', () => {
it('extracts data items with level, name, PIC, USAGE', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' DATA DIVISION.',
' WORKING-STORAGE SECTION.',
' 01 WS-RECORD.',
' 05 WS-NAME PIC X(30).',
' 05 WS-AMOUNT PIC 9(7)V99 USAGE COMP-3.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.dataItems.length).toBeGreaterThanOrEqual(3);
const wsName = r.dataItems.find(d => d.name === 'WS-NAME');
expect(wsName).toBeDefined();
expect(wsName!.level).toBe(5);
expect(wsName!.pic).toMatch(/^X\(30\)/);
const wsAmount = r.dataItems.find(d => d.name === 'WS-AMOUNT');
expect(wsAmount).toBeDefined();
expect(wsAmount!.usage).toBe('COMP-3');
});
it('extracts 88-level condition names with values', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' DATA DIVISION.',
' WORKING-STORAGE SECTION.',
' 01 WS-STATUS PIC X.',
' 88 WS-ACTIVE VALUE "A".',
' 88 WS-INACTIVE VALUE "I".',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
const active = r.dataItems.find(d => d.name === 'WS-ACTIVE');
expect(active).toBeDefined();
expect(active!.level).toBe(88);
expect(active!.values).toEqual(['A']);
const inactive = r.dataItems.find(d => d.name === 'WS-INACTIVE');
expect(inactive).toBeDefined();
expect(inactive!.values).toEqual(['I']);
});
it('extracts FD entries with record name linkage', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' DATA DIVISION.',
' FILE SECTION.',
' FD EMPLOYEE-FILE.',
' 01 EMPLOYEE-RECORD.',
' 05 EMP-ID PIC 9(5).',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.fdEntries).toHaveLength(1);
expect(r.fdEntries[0].fdName).toBe('EMPLOYEE-FILE');
expect(r.fdEntries[0].recordName).toBe('EMPLOYEE-RECORD');
});
it('skips FILLER items', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' DATA DIVISION.',
' WORKING-STORAGE SECTION.',
' 01 WS-REC.',
' 05 FILLER PIC X(10).',
' 05 WS-DATA PIC X(20).',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
const fillerItems = r.dataItems.filter(d => d.name === 'FILLER');
expect(fillerItems).toHaveLength(0);
expect(r.dataItems.find(d => d.name === 'WS-DATA')).toBeDefined();
});
it('correctly assigns data section (working-storage, linkage, file)', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' DATA DIVISION.',
' FILE SECTION.',
' FD MY-FILE.',
' 01 FILE-REC PIC X(80).',
' WORKING-STORAGE SECTION.',
' 01 WS-VAR PIC X(10).',
' LINKAGE SECTION.',
' 01 LK-VAR PIC X(10).',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
const fileRec = r.dataItems.find(d => d.name === 'FILE-REC');
expect(fileRec).toBeDefined();
expect(fileRec!.section).toBe('file');
const wsVar = r.dataItems.find(d => d.name === 'WS-VAR');
expect(wsVar).toBeDefined();
expect(wsVar!.section).toBe('working-storage');
const lkVar = r.dataItems.find(d => d.name === 'LK-VAR');
expect(lkVar).toBeDefined();
expect(lkVar!.section).toBe('linkage');
});
});
// -------------------------------------------------------------------------
// Environment Division
// -------------------------------------------------------------------------
describe('Environment Division', () => {
it('extracts SELECT ... ASSIGN TO with organization, access, record key', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' ENVIRONMENT DIVISION.',
' INPUT-OUTPUT SECTION.',
' FILE-CONTROL.',
' SELECT EMPLOYEE-FILE',
' ASSIGN TO "EMPFILE"',
' ORGANIZATION IS INDEXED',
' ACCESS MODE IS DYNAMIC',
' RECORD KEY IS EMP-ID.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.fileDeclarations).toHaveLength(1);
const fd = r.fileDeclarations[0];
expect(fd.selectName).toBe('EMPLOYEE-FILE');
expect(fd.assignTo).toBe('EMPFILE');
expect(fd.organization).toBe('INDEXED');
expect(fd.access).toBe('DYNAMIC');
expect(fd.recordKey).toBe('EMP-ID');
});
});
// -------------------------------------------------------------------------
// Sequence Numbers (fixed-format COBOL with 000100 cols 1-6)
// -------------------------------------------------------------------------
describe('Sequence Numbers', () => {
it('extracts paragraphs from fixed-format COBOL with numeric sequence numbers', () => {
// preprocessCobolSource strips cols 1-6, so the extractor sees clean lines
const src = preprocessCobolSource(cobol(
'000010 IDENTIFICATION DIVISION.',
'000020 PROGRAM-ID. SEQPROG.',
'000030 PROCEDURE DIVISION.',
'000040 MAIN-PARA.',
'000050 PERFORM SUB-PARA.',
'000060 SUB-PARA.',
'000070 DISPLAY "HI".',
));
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.programName).toBe('SEQPROG');
expect(r.paragraphs.map(p => p.name)).toEqual(['MAIN-PARA', 'SUB-PARA']);
expect(r.performs).toHaveLength(1);
expect(r.performs[0].target).toBe('SUB-PARA');
});
it('extracts data items from fixed-format COBOL with numeric sequence numbers', () => {
const src = preprocessCobolSource(cobol(
'000010 IDENTIFICATION DIVISION.',
'000020 PROGRAM-ID. SEQPROG.',
'000030 DATA DIVISION.',
'000040 WORKING-STORAGE SECTION.',
'000050 01 WS-AMOUNT PIC 9(7)V99.',
'000060 01 WS-NAME PIC X(30).',
));
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.dataItems.map(d => d.name)).toContain('WS-AMOUNT');
expect(r.dataItems.map(d => d.name)).toContain('WS-NAME');
});
});
// -------------------------------------------------------------------------
// State Machine
// -------------------------------------------------------------------------
describe('State Machine', () => {
it('correctly transitions between divisions', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' ENVIRONMENT DIVISION.',
' DATA DIVISION.',
' WORKING-STORAGE SECTION.',
' 01 WS-VAR PIC X(10).',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' DISPLAY WS-VAR.',
' STOP RUN.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.programName).toBe('TESTPROG');
expect(r.dataItems.find(d => d.name === 'WS-VAR')).toBeDefined();
expect(r.paragraphs).toHaveLength(1);
expect(r.paragraphs[0].name).toBe('MAIN-PARA');
});
it('handles continuation lines (indicator "-" in column 7)', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' CALL "VERY-LONG-PR',
' - "OGRAM-NAME".',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
// Continuation merges lines; at minimum verify no crash and paragraph found
expect(r.paragraphs).toHaveLength(1);
expect(r.paragraphs[0].name).toBe('MAIN-PARA');
});
it('skips comment lines (indicator "*" or "/" in column 7)', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' * THIS IS A COMMENT',
' / THIS IS A PAGE BREAK COMMENT',
' CALL "REALPROG".',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.calls).toHaveLength(1);
expect(r.calls[0].target).toBe('REALPROG');
});
});
// -------------------------------------------------------------------------
// EXEC Blocks
// -------------------------------------------------------------------------
describe('EXEC Blocks', () => {
it('extracts EXEC SQL blocks with tables and host variables', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' EXEC SQL',
' SELECT EMP-NAME, EMP-SALARY',
' FROM EMPLOYEE',
' WHERE EMP-ID = :WS-EMP-ID',
' END-EXEC.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.execSqlBlocks).toHaveLength(1);
const sql = r.execSqlBlocks[0];
expect(sql.operation).toBe('SELECT');
expect(sql.tables).toContain('EMPLOYEE');
expect(sql.hostVariables).toContain('WS-EMP-ID');
});
it('extracts EXEC CICS blocks with command and MAP/PROGRAM/TRANSID', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
" EXEC CICS SEND MAP('EMPMAP')",
" PROGRAM('EMPPROG')",
" TRANSID('EMPT')",
' END-EXEC.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.execCicsBlocks).toHaveLength(1);
const cics = r.execCicsBlocks[0];
expect(cics.command).toBe('SEND MAP');
expect(cics.mapName).toBe('EMPMAP');
expect(cics.programName).toBe('EMPPROG');
expect(cics.transId).toBe('EMPT');
});
it('handles single-line EXEC SQL ... END-EXEC', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' EXEC SQL DELETE FROM ORDERS WHERE ORD-ID = :WS-ORD END-EXEC.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.execSqlBlocks).toHaveLength(1);
expect(r.execSqlBlocks[0].operation).toBe('DELETE');
expect(r.execSqlBlocks[0].tables).toContain('ORDERS');
});
it('handles multi-line EXEC SQL ... END-EXEC', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' EXEC SQL',
' INSERT INTO AUDIT_LOG',
' VALUES (:WS-TIMESTAMP, :WS-USER)',
' END-EXEC.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.execSqlBlocks).toHaveLength(1);
const sql = r.execSqlBlocks[0];
expect(sql.operation).toBe('INSERT');
expect(sql.tables).toContain('AUDIT_LOG');
expect(sql.hostVariables).toContain('WS-TIMESTAMP');
expect(sql.hostVariables).toContain('WS-USER');
});
});
// -------------------------------------------------------------------------
// Linkage & Data Flow
// -------------------------------------------------------------------------
describe('Linkage & Data Flow', () => {
it('extracts PROCEDURE DIVISION USING parameters', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' DATA DIVISION.',
' LINKAGE SECTION.',
' 01 LK-PARAM1 PIC X(10).',
' 01 LK-PARAM2 PIC 9(5).',
' PROCEDURE DIVISION USING LK-PARAM1 LK-PARAM2.',
' MAIN-PARA.',
' DISPLAY LK-PARAM1.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.procedureUsing).toEqual(['LK-PARAM1', 'LK-PARAM2']);
});
it('extracts ENTRY points with USING', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' ENTRY "ALTENTRY" USING WS-PARAM1 WS-PARAM2.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.entryPoints).toHaveLength(1);
expect(r.entryPoints[0].name).toBe('ALTENTRY');
expect(r.entryPoints[0].parameters).toEqual(['WS-PARAM1', 'WS-PARAM2']);
});
it('extracts MOVE statements (skipping figurative constants)', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' PROCEDURE DIVISION.',
' MAIN-PARA.',
' MOVE WS-SOURCE TO WS-TARGET.',
' MOVE SPACES TO WS-BLANK.',
' MOVE ZEROS TO WS-ZERO.',
' MOVE CORRESPONDING WS-REC1 TO WS-REC2.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
const moveTargets = r.moves.map(m => ({ from: m.from, to: m.to, corr: m.corresponding }));
expect(moveTargets).toContainEqual({ from: 'WS-SOURCE', to: 'WS-TARGET', corr: false });
expect(moveTargets).toContainEqual({ from: 'WS-REC1', to: 'WS-REC2', corr: true });
expect(r.moves.find(m => m.from === 'SPACES')).toBeUndefined();
expect(r.moves.find(m => m.from === 'ZEROS')).toBeUndefined();
});
});
// -------------------------------------------------------------------------
// Edge Cases
// -------------------------------------------------------------------------
describe('Edge Cases', () => {
it('empty program returns empty results', () => {
const r = extractCobolSymbolsWithRegex('', 'empty.cbl');
expect(r.programName).toBeNull();
expect(r.paragraphs).toHaveLength(0);
expect(r.sections).toHaveLength(0);
expect(r.performs).toHaveLength(0);
expect(r.calls).toHaveLength(0);
expect(r.copies).toHaveLength(0);
expect(r.dataItems).toHaveLength(0);
expect(r.fileDeclarations).toHaveLength(0);
expect(r.fdEntries).toHaveLength(0);
expect(r.execSqlBlocks).toHaveLength(0);
expect(r.execCicsBlocks).toHaveLength(0);
expect(r.procedureUsing).toHaveLength(0);
expect(r.entryPoints).toHaveLength(0);
expect(r.moves).toHaveLength(0);
});
it('extracts AUTHOR and DATE-WRITTEN from program metadata', () => {
const src = cobol(
' IDENTIFICATION DIVISION.',
' PROGRAM-ID. TESTPROG.',
' AUTHOR. JOHN DOE.',
' DATE-WRITTEN. 2025-01-15.',
);
const r = extractCobolSymbolsWithRegex(src, 'test.cbl');
expect(r.programMetadata.author).toBe('JOHN DOE');
expect(r.programMetadata.dateWritten).toBe('2025-01-15');
});
});
});
+338
View File
@@ -0,0 +1,338 @@
import { describe, it, expect } from 'vitest';
import { parseJcl } from '../../src/core/ingestion/cobol/jcl-parser.js';
import type { JclParseResults } from '../../src/core/ingestion/cobol/jcl-parser.js';
describe('parseJcl', () => {
// ── JOB statements ──────────────────────────────────────────────────
describe('JOB statements', () => {
it('extracts job name', () => {
const jcl = `//MYJOB JOB (ACCT),'MY JOB'`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.jobs).toHaveLength(1);
expect(r.jobs[0].name).toBe('MYJOB');
expect(r.jobs[0].line).toBe(1);
});
it('extracts CLASS and MSGCLASS parameters', () => {
const jcl = `//PAYJOB JOB (ACCT),'PAYROLL',CLASS=A,MSGCLASS=X`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.jobs).toHaveLength(1);
expect(r.jobs[0].name).toBe('PAYJOB');
expect(r.jobs[0].class).toBe('A');
expect(r.jobs[0].msgclass).toBe('X');
});
it('handles job with no CLASS or MSGCLASS', () => {
const jcl = `//BAREJOB JOB (ACCT),'BARE'`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.jobs).toHaveLength(1);
expect(r.jobs[0].name).toBe('BAREJOB');
expect(r.jobs[0].class).toBeUndefined();
expect(r.jobs[0].msgclass).toBeUndefined();
});
});
// ── EXEC statements ─────────────────────────────────────────────────
describe('EXEC statements', () => {
it('extracts step with PGM=program', () => {
const jcl = [
'//MYJOB JOB (ACCT)',
'//STEP1 EXEC PGM=IEFBR14',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.steps).toHaveLength(1);
expect(r.steps[0].name).toBe('STEP1');
expect(r.steps[0].program).toBe('IEFBR14');
expect(r.steps[0].proc).toBeUndefined();
});
it('extracts step with proc name (no PGM= keyword)', () => {
const jcl = [
'//MYJOB JOB (ACCT)',
'//STEP1 EXEC MYPROC',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.steps).toHaveLength(1);
expect(r.steps[0].name).toBe('STEP1');
expect(r.steps[0].program).toBeUndefined();
expect(r.steps[0].proc).toBe('MYPROC');
});
it('associates step with current job', () => {
const jcl = [
'//JOB1 JOB (ACCT)',
'//STEPA EXEC PGM=PROG1',
'//JOB2 JOB (ACCT)',
'//STEPB EXEC PGM=PROG2',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.steps).toHaveLength(2);
expect(r.steps[0].jobName).toBe('JOB1');
expect(r.steps[1].jobName).toBe('JOB2');
});
});
// ── DD statements ───────────────────────────────────────────────────
describe('DD statements', () => {
it('extracts DD name and dataset (DSN=)', () => {
const jcl = [
'//MYJOB JOB (ACCT)',
'//STEP1 EXEC PGM=IEFBR14',
'//INPUT DD DSN=MY.DATA.SET,DISP=SHR',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.ddStatements).toHaveLength(1);
expect(r.ddStatements[0].ddName).toBe('INPUT');
expect(r.ddStatements[0].dataset).toBe('MY.DATA.SET');
});
it('extracts DISP parameter', () => {
const jcl = [
'//MYJOB JOB (ACCT)',
'//STEP1 EXEC PGM=IEFBR14',
'//OUTPUT DD DSN=MY.OUT,DISP=(NEW,CATLG,DELETE)',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.ddStatements).toHaveLength(1);
expect(r.ddStatements[0].disp).toBe('NEW');
});
it('associates DD with current step', () => {
const jcl = [
'//MYJOB JOB (ACCT)',
'//STEP1 EXEC PGM=PROG1',
'//DD1 DD DSN=DS1,DISP=SHR',
'//STEP2 EXEC PGM=PROG2',
'//DD2 DD DSN=DS2,DISP=SHR',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.ddStatements).toHaveLength(2);
expect(r.ddStatements[0].stepName).toBe('STEP1');
expect(r.ddStatements[1].stepName).toBe('STEP2');
});
});
// ── PROC definitions ────────────────────────────────────────────────
describe('PROC definitions', () => {
it('extracts in-stream PROC with name', () => {
const jcl = [
'//MYPROC PROC',
'//STEP1 EXEC PGM=IEFBR14',
'// PEND',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.procs).toHaveLength(1);
expect(r.procs[0].name).toBe('MYPROC');
expect(r.procs[0].isInStream).toBe(true);
});
it('handles PROC/PEND pairs', () => {
const jcl = [
'//PROC1 PROC',
'//S1 EXEC PGM=PROG1',
'// PEND',
'//PROC2 PROC',
'//S2 EXEC PGM=PROG2',
'// PEND',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.procs).toHaveLength(2);
expect(r.procs[0].name).toBe('PROC1');
expect(r.procs[1].name).toBe('PROC2');
});
});
// ── INCLUDE / SET ───────────────────────────────────────────────────
describe('INCLUDE and SET', () => {
it('extracts INCLUDE MEMBER=name', () => {
const jcl = `// INCLUDE MEMBER=MYINCL`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.includes).toHaveLength(1);
expect(r.includes[0].member).toBe('MYINCL');
expect(r.includes[0].line).toBe(1);
});
it('extracts SET variable=value', () => {
const jcl = `// SET ENV=PROD`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.sets).toHaveLength(1);
expect(r.sets[0].variable).toBe('ENV');
expect(r.sets[0].value).toBe('PROD');
});
});
// ── Conditionals ────────────────────────────────────────────────────
describe('Conditionals', () => {
it('extracts IF condition THEN', () => {
const jcl = `// IF STEP1.RC = 0 THEN`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.conditionals).toHaveLength(1);
expect(r.conditionals[0].type).toBe('IF');
expect(r.conditionals[0].condition).toBe('STEP1.RC = 0');
});
it('extracts ELSE and ENDIF', () => {
const jcl = [
'// IF STEP1.RC = 0 THEN',
'//GOOD EXEC PGM=GOODPGM',
'// ELSE',
'//BAD EXEC PGM=BADPGM',
'// ENDIF',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.conditionals).toHaveLength(3);
expect(r.conditionals[0].type).toBe('IF');
expect(r.conditionals[1].type).toBe('ELSE');
expect(r.conditionals[1].condition).toBeUndefined();
expect(r.conditionals[2].type).toBe('ENDIF');
expect(r.conditionals[2].condition).toBeUndefined();
});
});
// ── JCLLIB ──────────────────────────────────────────────────────────
describe('JCLLIB', () => {
it('extracts JCLLIB ORDER=(lib1,lib2)', () => {
const jcl = `// JCLLIB ORDER=(SYS1.PROCLIB,USER.PROCLIB)`;
const r = parseJcl(jcl, 'test.jcl');
expect(r.jcllib).toHaveLength(1);
expect(r.jcllib[0].order).toEqual(['SYS1.PROCLIB', 'USER.PROCLIB']);
expect(r.jcllib[0].line).toBe(1);
});
});
// ── Continuation lines ──────────────────────────────────────────────
describe('Continuation lines', () => {
it('joins continuation lines (col 72 non-blank + next line starts with //)', () => {
// Build a DD line that is exactly 72 chars with non-blank at col 72 (index 71).
// The continuation line provides the DISP parameter.
// "//DD1 DD DSN=MY.VERY.LONG.DATASET.NAME.THAT.KEEPS.GOING," is 60 chars.
// Pad to 71 then add non-blank at col 72.
const base = '//DD1 DD DSN=MY.VERY.LONG.DATASET.NAME.THAT.KEEPS.GOING,';
const padding = ' '.repeat(71 - base.length);
const line1 = base + padding + 'X'; // col 72 is 'X' (non-blank) -> continuation
const line2 = '// DISP=SHR';
const jcl = [
'//MYJOB JOB (ACCT)',
'//STEP1 EXEC PGM=IEFBR14',
line1,
line2,
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
// The continuation should join the DD line so both DSN and DISP are parsed
expect(r.ddStatements).toHaveLength(1);
expect(r.ddStatements[0].ddName).toBe('DD1');
expect(r.ddStatements[0].dataset).toBe('MY.VERY.LONG.DATASET.NAME.THAT.KEEPS.GOING');
expect(r.ddStatements[0].disp).toBe('SHR');
});
});
// ── Edge cases ──────────────────────────────────────────────────────
describe('Edge cases', () => {
it('skips JCL comments (//*)', () => {
const jcl = [
'//* This is a comment',
'//MYJOB JOB (ACCT)',
'//* Another comment',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.jobs).toHaveLength(1);
expect(r.jobs[0].name).toBe('MYJOB');
});
it('skips non-JCL lines', () => {
const jcl = [
'This is not a JCL line',
'//MYJOB JOB (ACCT)',
' Some data',
'//STEP1 EXEC PGM=IEFBR14',
].join('\n');
const r = parseJcl(jcl, 'test.jcl');
expect(r.jobs).toHaveLength(1);
expect(r.steps).toHaveLength(1);
});
it('empty input returns empty results', () => {
const r = parseJcl('', 'test.jcl');
expect(r.jobs).toEqual([]);
expect(r.steps).toEqual([]);
expect(r.ddStatements).toEqual([]);
expect(r.procs).toEqual([]);
expect(r.includes).toEqual([]);
expect(r.sets).toEqual([]);
expect(r.jcllib).toEqual([]);
expect(r.conditionals).toEqual([]);
});
it('complete JCL job with multiple steps and DDs', () => {
const jcl = [
'//* Complete payroll job',
'//PAYJOB JOB (ACCT123),\'PAYROLL RUN\',CLASS=A,MSGCLASS=X',
'// JCLLIB ORDER=(PAY.PROCLIB,SYS1.PROCLIB)',
'// SET ENV=PROD',
'// INCLUDE MEMBER=STDPARMS',
'//*',
'// IF 1 = 1 THEN',
'//STEP01 EXEC PGM=PAYEXT',
'//INPUT DD DSN=PAY.MASTER,DISP=SHR',
'//OUTPUT DD DSN=PAY.EXTRACT,DISP=(NEW,CATLG,DELETE)',
'//SYSPRINT DD SYSOUT=*',
'//*',
'//STEP02 EXEC PAYCALC',
'//INFILE DD DSN=PAY.EXTRACT,DISP=SHR',
'// ELSE',
'//STEP03 EXEC PGM=IEFBR14',
'// ENDIF',
].join('\n');
const r = parseJcl(jcl, 'payroll.jcl');
// Jobs
expect(r.jobs).toHaveLength(1);
expect(r.jobs[0]).toEqual({
name: 'PAYJOB',
line: 2,
class: 'A',
msgclass: 'X',
});
// JCLLIB
expect(r.jcllib).toHaveLength(1);
expect(r.jcllib[0].order).toEqual(['PAY.PROCLIB', 'SYS1.PROCLIB']);
// SET
expect(r.sets).toHaveLength(1);
expect(r.sets[0]).toEqual({ variable: 'ENV', value: 'PROD', line: 4 });
// INCLUDE
expect(r.includes).toHaveLength(1);
expect(r.includes[0].member).toBe('STDPARMS');
// Conditionals
expect(r.conditionals).toHaveLength(3);
expect(r.conditionals[0].type).toBe('IF');
expect(r.conditionals[1].type).toBe('ELSE');
expect(r.conditionals[2].type).toBe('ENDIF');
// Steps
expect(r.steps).toHaveLength(3);
expect(r.steps[0]).toMatchObject({ name: 'STEP01', program: 'PAYEXT', jobName: 'PAYJOB' });
expect(r.steps[1]).toMatchObject({ name: 'STEP02', proc: 'PAYCALC', jobName: 'PAYJOB' });
expect(r.steps[2]).toMatchObject({ name: 'STEP03', program: 'IEFBR14', jobName: 'PAYJOB' });
// DD statements
expect(r.ddStatements).toHaveLength(4);
expect(r.ddStatements[0]).toMatchObject({ ddName: 'INPUT', stepName: 'STEP01', dataset: 'PAY.MASTER', disp: 'SHR' });
expect(r.ddStatements[1]).toMatchObject({ ddName: 'OUTPUT', stepName: 'STEP01', disp: 'NEW' });
expect(r.ddStatements[2]).toMatchObject({ ddName: 'SYSPRINT', stepName: 'STEP01' });
expect(r.ddStatements[3]).toMatchObject({ ddName: 'INFILE', stepName: 'STEP02', dataset: 'PAY.EXTRACT' });
});
});
});