Compare commits

...
Author SHA1 Message Date
gitnexus-release-bot[bot] a0ee408f4b release: v1.6.6-rc.47 2026-05-23 08:10:23 +00:00
dependabot[bot] fb94dba484 chore(deps)(deps-dev): bump tsx from 4.22.0 to 4.22.3 in /gitnexus (#1789) 2026-05-23 08:49:15 +01:00
Hugo Gu 7fc797e2ce feat: Support DeepSeek V4 API (#1594) 2026-05-23 08:05:13 +01:00
Gergő Magyar 51e667808a feat(lang-kotlin): flip Kotlin to MIGRATED_LANGUAGES + close #1756 / #1757 (refs #1746) (#1782) 2026-05-23 07:24:35 +01:00
dependabot[bot] 84ac88a741 chore(deps)(deps): bump qs from 6.14.2 to 6.15.2 in /gitnexus (#1791) 2026-05-23 06:39:05 +01:00
ChamHerry fc6007e70b feat(i18n): make web and CLI language-aware (#1748) 2026-05-23 06:14:24 +01:00
87b91c821e fix(lbug): add WAL checkpoint-threshold control (#1772)
* Initial plan

* fix(analyze): add WAL auto-checkpoint CLI control and default-off behavior

* test(analyze): share lbug auto-checkpoint parsing and align validation

* fix(analyze): always enable lbug auto-checkpoint and expose threshold control

* refactor(lbug): inline always-on auto-checkpoint constructor arg

* fix(analyze): guide checkpoint-threshold on Ladybug WAL checkpoint IO failures

* test(analyze): cover checkpoint IO guidance and add integration guard

* fix(analyze): tighten checkpoint IO detection and remove test hook

* fix(analyze): remove checkpoint test hook and tighten error matching

* fix(analyze): rename to wal-checkpoint-threshold, raise default, add manual checkpoint driver with retry

Address review feedback on PR #1772:

- Rename CLI flag, env var, AnalyzeOptions field, recovery-hint tag, and
  parser/constants from lbug-* to engine-neutral wal-* (matches the existing
  WAL_RECOVERY_SUGGESTION / isWalCorruptionError convention).
- Raise default threshold from -1 (Ladybug stock ~16 MiB) to 64 MiB so users
  on the default config no longer hit the original rename/remove race.
- Align both READMEs to publish 67108864 (64 MiB) instead of 65536 (which
  would have made the crash more frequent).
- Add wal-checkpoint-driver.ts: a periodic manual CHECKPOINT driver wrapped
  in a 3-attempt jittered retry (50/200/500 ms), driven from runFullAnalysis.
  Opt-out via GITNEXUS_WAL_MANUAL_CHECKPOINT=0. Moves the race window into a
  JS-controllable retry surface while keeping native auto-checkpoint on.
- Move LBUG_CHECKPOINT_RENAME_RE / REMOVE_RE plus the predicate (renamed to
  isLbugCheckpointIoError) into lbug-config.ts alongside isWalCorruptionError.
  Predicate is now exported. Add a permissive fallback matcher and pin the
  matched Ladybug version in comments.
- Warn instead of silently defaulting when GITNEXUS_WAL_CHECKPOINT_THRESHOLD
  is set to a non-empty unparseable value (closes the CLI-vs-env asymmetry).
- Add a typed RecoveryHint string-literal union in cli-message.ts so future
  hint tags can't drift.
- Add a real integration test under test/integration/ that triggers a
  Ladybug checkpoint IO failure via a pre-existing directory at the rename
  target (portable across platforms; no test-only injection hook).
- Add small-disk / CI caveat (32 MiB secondary suggestion) to the recovery
  hint and README env-var rows.
- Document CLI/env precedence in the analyze --help block.
- Help placeholder: <value> -> <bytes>.
- Rename analyze-lbug-auto-checkpoint.test.ts to use the new wal-* token.

* chore(lbug): remove dead jitteredDelay helper and apply prettier

- Drop unused `jitteredDelay` function flagged by CodeQL in PR #1772; the
  retry loop already inlines the same calculation with the injectable
  `randomImpl` so the helper was dead. Move the non-cryptographic-by-design
  comment next to the actual jitter site.
- Apply `prettier --write` to wal-checkpoint-driver.ts and the new
  integration test to absorb the PR autofix bot's formatting findings.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Test <test@example.com>
2026-05-22 14:46:49 +01:00
azizur100389andGergő Magyar 952ada70c5 feat(cpp): Resolve overloaded operator calls (#1754)
* feat(cpp): resolve overloaded operator calls

* fix(cpp): tighten overloaded operator resolution

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-22 13:31:06 +01:00
Gergő MagyarandTest 060fe75715 docs(lang-kotlin): refresh scope-resolver JSDoc after #1758-#1763 landed (#1781)
The scope-resolver header comment claimed forced-mode passed 154/175
(88%) and listed smart casts, cross-file iterables, method chains,
overload selection, virtual dispatch, and interface defaults as
"remaining gaps". All six landed in PRs #1774-#1779. Forced mode now
passes 175/175 (verified post-merge against `main`).

Update the header to:
- state the current forced-mode result accurately,
- enumerate the closed sub-issues so future readers can trace each
  capability back to its PR,
- and explicitly name the remaining flip blockers (#1755, #1756,
  #1757) so the next maintainer to look at this file knows exactly
  what's required before adding `Kotlin` to `MIGRATED_LANGUAGES`.

Docs-only — no behavioral changes.

Refs #1746.

Co-authored-by: Test <test@example.com>
2026-05-22 12:44:44 +01:00
d15f8bef54 feat(ingestion): log deferred resolution progress when verbose (#1741) (#1773)
* feat(ingestion): log deferred resolution progress when verbose

Add [deferred-profile] timing logs for post-chunk import, heritage, heritage-map, and legacy call resolution. Enabled on GITNEXUS_VERBOSE / analyze -v (and optionally GITNEXUS_PROFILE_DEFERRED) to diagnose analyze stalls on large repos (issue #1741).

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(ingestion): address PR #1773 production-readiness review

Move deferred call progress logs after the registry-primary skip so sites= counts match files actually resolved. Only time buildHeritageMap when heritage records exist; otherwise log an explicit skip. Add wiring tests that assert [deferred-profile] emission from buildHeritageMap and processCallsFromExtracted. Snapshot GITNEXUS_PROFILE_DEFERRED env vars in analyze CLI isolation.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ingestion): address PR #1773 code-review findings

P0
- Replace forbidden toBeGreaterThanOrEqual/toBeLessThan in
  profileElapsedMs test with exact-arithmetic vi.spyOn(hrtime.bigint)
  asserting .toBe(2.5) and .toBe(0). DoD §2.7 compliance.

P2
- Use Number() (not parseInt) when parsing
  GITNEXUS_PROFILE_DEFERRED_SLOW_MS so scientific notation like '1e9'
  doesn't silently parse to 1 and turn the slow-file log into a per-file
  log storm.
- Introduce startTimer(enabled): bigint | null and endTimer(start,
  format) helpers in deferred-resolution-profile.ts; refactor 6+
  timing blocks in parse-impl.ts and call-processor.ts to use them.
  Removes the 0n sentinel that conflated 'disabled' with 'zero
  elapsed time' and let TS narrow correctly.
- Split the call-processor file counter: filesProcessed (all iterated)
  vs resolvedFiles (post registry-primary skip). Key the every-N
  progress log and the start-of-phase log on resolvedFiles so mixed
  Python+JVM repos where the skipped language sorts first still emit
  'calls 1/1 file=...' on the first non-skipped file. Adds a wiring
  test for the mixed-language ordering case.

P3
- Restore the original isDev '🔗 E1: Seeded ...' logger.info line so
  log scrapers keyed on the emoji marker still match; emit the
  [deferred-profile] variant only when deferredProfile && !isDev.
- Move tFile = startTimer(profileCalls) below the registry-primary
  skip so skipped files don't trigger an hrtime.bigint() call.
- Document GITNEXUS_PROFILE_DEFERRED and
  GITNEXUS_PROFILE_DEFERRED_SLOW_MS in the README env-var table.

* refactor(ingestion): extract parseTruthyEnv to shared utils (U5)

Three narrow-form env-var truthy checkers (verbose.ts, registry-primary-flag.ts,
deferred-resolution-profile.ts) each had their own `'1' | 'true' | 'yes'` parser
with subtle divergences (trim or no trim, set vs disjunction). Consolidate on a
single `parseTruthyEnv(raw)` helper in utils/env.ts — the module already serves
as the centralization point for shared ingestion env constants.

logger.ts's broader `isTruthyEnv` (negative-list, pino-debug convention) stays
untouched — different intent, different semantics.

New table-driven test at test/unit/env.test.ts covers case variants,
whitespace, and rejection of falsy / unknown tokens.

* refactor(ingestion): named constants for deferred-profile log gates (U6)

Replace magic literals 10 / 100 / 3_000 / 5_000 in
deferred-resolution-profile.ts with module-private named constants
LOG_EVERY_N_VERBOSE, LOG_EVERY_N_PROFILE, DEFAULT_SLOW_MS_VERBOSE,
DEFAULT_SLOW_MS. Not exported — internal tuning knobs. Pure refactor;
existing tests assert the exact values and still pass unchanged.

* fix(ingestion): pre-pass denominator for deferred call progress (U1, A1)

The live per-file denominator in processCallsFromExtracted previously
read `totalFiles - skippedRegistryPrimaryFiles` at log time. On mixed
Python+JVM repos where the skipped language interleaves with the
resolved one, the denominator drifts upward as the loop iterates —
files iterated before later skips have been seen carry an inflated
denominator. The live ratio only self-corrects after the final file
has been classified.

Fix: one-pass pre-count over byFile.keys() before the work loop
computes resolvedTotal once. The denominator is then stable from the
first emission onward. The pre-pass runs only on the enabled path
(profileCalls=true) so the disabled path keeps zero extra work.

Adds a wiring test exercising the alternating [ts, py, ts, py, ...]
order that triggered the drift, asserting every emitted line uses
`/4` and no other denominator slips through.

* fix(ingestion): E1 enrichment log emits on both dev and profile flags (U2, A2)

The post-chunk E1 enrichment log used `if (isDev) {...} else if
(deferredProfile) {...}` which is mutually exclusive. On combined runs
(NODE_ENV=development + GITNEXUS_PROFILE_DEFERRED=1) the [deferred-
profile] line was silently swallowed — operators grepping that prefix
saw a gap between wildcard-synth and heritage timings, while the
inline comment promised dual emission.

Fix: two independent `if` statements so both branches fire when both
flags are set. The original emoji-prefixed `🔗 E1: Seeded` line keeps
its phrasing for any dev-mode log scrapers that depend on the marker.

Pinning test (parse-impl-e1-emission-shape.test.ts) reads the source
and asserts (a) both branches exist as standalone `if` statements and
(b) the closing `}` of the isDev branch is followed by `if`, not
`else if`. Source-shape pins are the right test scope for a purely
structural change — the regression we are guarding against is exactly
how a future reader greps for it.

* feat(ingestion): unresolved-side counters in heritage-map profile (U7)

The existing maxNameCartesian / ambiguousHeritageRecords counters in
buildHeritageMap only observed records where BOTH the child and parent
name lookups resolved. On JVM monorepos the actual pathological case is
one side empty (typically an unresolved external supertype with many
same-named children, or vice versa) — those records were silently
dropped from the metric.

Add `unresolvedChildLookups` and `unresolvedParentLookups` in a
separate `if (profileHeritage)` block placed immediately after the two
`lookupClassByName` calls (so it observes the unresolved cases the
length-guarded ambiguity block below cannot see). Both counters reuse
the existing childDefs / parentDefs values — no additional lookups.

Done-summary log extended to include the two new counters. Wiring test
covers both directions (unresolved parent, unresolved child) plus the
existing "both resolved" baseline now asserts the new counters report
zero for that case.

* fix(ingestion): endTimer formatter exception safety (U3)

Wrap the format callback in endTimer in a try/catch so a throwing
formatter (custom toString, JSON.stringify on a circular object,
future heavier serializers) cannot abort the deferred resolution
band. Observability code must never escalate to a load-bearing
failure mode.

On catch we emit a single `[deferred-profile] formatter error: …`
line via logDeferredProfile and return; the caller's stage continues
as if profiling had no-op'd for this timer. DoD §2.8 is satisfied —
the failure is surfaced, not silently swallowed.

Tests cover the four cases: happy path emits the formatted line, null
start no-ops without invoking the formatter, throwing formatter is
caught and surfaces one error line, non-Error throws are coerced via
String() in the message.

* fix(ingestion): defensive wrap + dropped-line counter for logDeferredProfile (U4)

Wrap logger.info inside logDeferredProfile in a try/catch so a throwing
underlying logger cannot abort the deferred resolution band. Pino with
sync:false (the current SonicBoom destination) does not throw
synchronously for `info(string)` calls, but first-use construction
paths (pino-pretty resolve, level validation) and any future transport
reconfiguration could. The wrap is belt-and-suspenders coverage; the
counter makes silent failures visible.

A module-private droppedLogLines counter accumulates dropped lines.
Two helpers — getDeferredProfileDroppedCount() and
resetDeferredProfileDroppedCount() — expose the counter. The handler
deliberately does NOT call the failing logger; that would risk an
infinite loop if the failure is steady-state.

processCallsFromExtracted resets the counter at entry (so each analyze
run gets a fresh count rather than accumulating across the process
lifetime — relevant for the MCP server, eval harness, integration
tests), and surfaces the count in the done-summary as `note: N profile
log lines dropped (logger errors)` when greater than zero. DoD §2.8
(no silent diagnostic catches) is satisfied.

Tests cover the helper API (zero at entry, idempotent reset) and the
happy path; the catch arm is pinned via source-shape assertion since
the logger Proxy can't be vi.spyOn'd directly (lazy `get` trap, no
own-property to wrap).

* docs(readme): clarify GITNEXUS_PROFILE_DEFERRED_SLOW_MS coercion (U8)

The env-var row mentioned integer / scientific notation only, but the
underlying parser (`Number(raw)` since the U2 fix in PR #1773) also
accepts decimals like `.5` and hex like `0x10`. Document the actual
acceptance set plus the non-finite / non-positive fallback so operators
setting unusual values know what to expect.

---------

Co-authored-by: Test <test@example.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-22 12:37:30 +01:00
Gergő MagyarandTest f72d9a99c6 fix(lang-kotlin): virtual dispatch via constructor type override (#1762) (#1778)
Closes #1762. `val animal: Animal = Dog(); animal.speak()` resolved
to `Animal.speak` (or no edge to Dog at all) under
`REGISTRY_PRIMARY_KOTLIN=1` because the Kotlin scope query emits BOTH
an annotation type-binding (`animal -> Animal`) and a constructor-
inferred type-binding (`animal -> Dog`). The generic scope-extractor
ranks annotation sources higher than constructor-inferred sources (see
`typeBindingStrength` in scope-extractor.ts), so the annotation
always won and `animal.speak()` dispatched against the static type.

Kotlin's virtual dispatch semantics expect the dynamic type — the
overriding `Dog.speak` should win when the RHS is a constructor call,
because that's what runs at runtime.

Fix: in `emitKotlinScopeCaptures`, suppress the `@type-binding.
annotation` capture when the underlying `property_declaration` has a
`call_expression` value sibling. The constructor-inferred capture
remains, becomes the sole binding for the variable, and receiver-bound
resolution dispatches against the constructed class (and walks its
MRO).

This is intentionally Kotlin-specific — flipping precedence globally
would change behavior for other languages whose static-type
annotations are still the right binding when present. Kotlin is the
language where the constructor RHS is the dispatch target by design.

Verification (REGISTRY_PRIMARY_KOTLIN=1):
- Forced-mode: 21 -> 20 failing of 175 (1 fewer; test 1715 in
  `test/integration/resolvers/kotlin.test.ts` now green).
- Default-mode Kotlin: 175/175 unchanged.
- Full resolver suite: 2216/2216 unchanged.
- Remaining 20 failures are tracked by sibling sub-issues
  (#1758, #1759, #1760, #1761, #1763).

Does NOT add Kotlin to MIGRATED_LANGUAGES per parent #1746 flip criteria.

Closes #1762. Refs #1746.

Co-authored-by: Test <test@example.com>
2026-05-22 10:40:59 +01:00
Gergő MagyarandTest 64efc202f6 fix(lang-kotlin): method-chain fixpoint receiver types (#1760) (#1776)
Closes #1760. Multi-step intra-file chains like

    val user = getUser()
    val addr = user.address
    val city = addr.getCity()
    city.save()

produced no `CALLS` edge for `city.save()` because the Kotlin extractor
only inferred property types for `simple_identifier` values (`val x = y`)
and call expressions with simple-identifier callees (`val x = fn()`).
Navigation expressions (`val addr = user.address`) and call expressions
with navigation-expression callees (`val city = addr.getCity()`)
returned null, leaving `addr` and `city` unbound — the chain broke
two hops before `city.save()`.

Implementation:
- `collectKotlinClassMembers(rootNode)` indexes per-file class fields
  (primary-constructor `val`/`var` params + body property declarations)
  and method return types. Per-file scope matches the existing
  extractor design.
- `inferKotlinPropertyType` gains two new cases:
  1. `navigation_expression` value — receiver type via `localTypes`,
     field type via `classMembers.fields`.
  2. `call_expression` with `navigation_expression` callee — receiver
     type via `localTypes`, method return type via `classMembers.methods`.
  Both return null when any link is unknown (safe / over-conservative).

Verification (REGISTRY_PRIMARY_KOTLIN=1):
- Forced-mode: 21 -> 20 failing of 175 (1 fewer; test 1491 in
  `test/integration/resolvers/kotlin.test.ts` now green).
- Default-mode Kotlin: 175/175 unchanged.
- Full resolver suite: 2216/2216 unchanged.
- Remaining 20 failures are tracked by sibling sub-issues
  (#1758, #1759, #1761, #1762, #1763).

Does NOT add Kotlin to MIGRATED_LANGUAGES per parent #1746 flip criteria.

Closes #1760. Refs #1746.

Co-authored-by: Test <test@example.com>
2026-05-22 10:40:46 +01:00
Gergő MagyarandTest 67cc4c6d94 fix(lang-kotlin): cross-file iterable return propagation (#1759) (#1775)
Two related bugs surfaced in REGISTRY_PRIMARY_KOTLIN=1 forced mode:

1. `import models.getRepo` silently resolved to `models/User.kt` (the
   first `.kt` file inside `models/` by iteration order) when no file
   was named after the symbol. `findKotlinFile` returned a single
   directory child as a fallback, so the importer's module-scope mirror
   only ever picked up the first arbitrary candidate — `getUser → User`
   landed but `getRepo → Repo` never did, and downstream `repo.save()`
   resolution fell through to no edge.

2. `for (x in importedCallable())` produced no for-loop type binding
   when the callee's return type lived in another file, because
   `inferKotlinIterableElementType`'s call-expression arm consulted
   only the local file's `returnTypes` map.

Fix:
- Split `findKotlinFile` into `findKotlinExactOrSuffix` (exact / suffix
  match only) and `findKotlinDirectoryChild` (legacy single-child
  fallback). Add `findKotlinPackageFiles` returning every `.kt`/`.kts`
  file inside a package directory. The resolver now fans out the
  stripped path through `findKotlinExactOrSuffix → findKotlinPackageFiles`,
  returning a `readonly string[]` candidate set. The finalize pass
  walks each candidate and picks the one whose `localDefs` actually
  export the imported name — exactly the multi-target contract
  `FinalizeHooks.resolveImportTarget` already supports.
- `inferKotlinIterableElementType` for `call_expression` now falls
  back to the callee's identifier text when the local return-type map
  has no entry. `propagateImportedReturnTypes` chain-follows
  `loopvar → callee → ElementType` once the imported `callee → Element`
  mirror lands at module scope (which now works thanks to fix #1).

Verification (REGISTRY_PRIMARY_KOTLIN=1):
- Forced-mode: 21 -> 18 failing of 175 (3 fewer; tests 487, 1242, 1251
  in test/integration/resolvers/kotlin.test.ts now green).
- Default-mode Kotlin: 175/175 unchanged.
- Full resolver suite: 2216/2216 unchanged (incl. `kotlin-calls`
  `util.OneArg.writeAudit` regression check at line 176).
- Remaining 18 failures are tracked by sibling sub-issues
  (#1758, #1760, #1761, #1762, #1763).

Does NOT add Kotlin to MIGRATED_LANGUAGES per parent #1746 flip criteria.

Closes #1759. Refs #1746.

Co-authored-by: Test <test@example.com>
2026-05-22 10:40:32 +01:00
Gergő MagyarandTest a3e7dfa8a6 fix(lang-kotlin): interface default method dispatch via implements-split MRO (#1763) (#1779)
Closes #1763. `user.validate()` on `class User : Validator` resolved
to no edge under REGISTRY_PRIMARY_KOTLIN=1 when validate() was a
default method declared on the Validator interface:

    class User(val name: String) : Validator
    interface Validator { fun validate(): Boolean = true }
    fun run() { val user = User("alice"); user.validate() }

The generic `buildMro` walks EXTENDS edges only. Kotlin classes
implement interfaces via IMPLEMENTS edges (per the parsing-processor),
so the implementor's MRO never picked up the interface's default
methods — `findOwnedMember(User, validate)` returned undefined and
no fallback walked to Validator.

Fix: replace `defaultLinearize` with a Kotlin-specific MRO builder
modeled after PHP's `buildPhpMro` (trait composition):
1. Run the generic `buildMro` (EXTENDS-only).
2. Collect direct IMPLEMENTS edges as class -> interface[] map.
3. For each class, walk its EXTENDS-MRO ancestors AND its own
   IMPLEMENTS edges to seed interface candidates, then BFS-close to
   pick up transitive interface inheritance (interface A : B).
4. Append the interface closure to the class's MRO (after the EXTENDS
   chain — Kotlin requires explicit override on conflict, so this
   ordering is a safe approximation for method lookup).
5. Classes with no EXTENDS but with IMPLEMENTS edges (the #1763
   fixture shape) get their MRO seeded directly from their interfaces.

Verification (REGISTRY_PRIMARY_KOTLIN=1):
- Forced-mode: 21 -> 20 failing of 175 (1 fewer; test 2062 in
  `test/integration/resolvers/kotlin.test.ts` now green).
- Default-mode Kotlin: 175/175 unchanged.
- Full resolver suite: 2216/2216 unchanged.
- Remaining 20 failures are tracked by sibling sub-issues
  (#1758, #1759, #1760, #1761, #1762).

Does NOT add Kotlin to MIGRATED_LANGUAGES per parent #1746 flip criteria.

Closes #1763. Refs #1746.

Co-authored-by: Test <test@example.com>
2026-05-22 10:24:02 +01:00
Gergő MagyarandTest ccf0b8b73c fix(lang-kotlin): smart-cast type refinement for when/is and if/is (#1758) (#1774)
Adds tree-sitter @scope.block captures for Kotlin when-arm bodies and
if-then bodies, plus a synthesizer that emits narrowed type-bindings
anchored on those bodies. The receiver-bound calls pass then resolves
`obj.member()` inside `is T` arms against `T` without leaking the
narrowing to sibling arms, `else` branches, or the enclosing function.

Implementation:
- query.ts: @scope.block on `(when_entry (when_condition (type_test))
  (control_structure_body))` and `(if_expression (check_expression)
  (control_structure_body))`.
- captures.ts: synthesizeKotlinSmartCastBindings walks `when_expression`
  and `if_expression` nodes; emits `@type-binding.annotation` with a
  `@type-binding.narrowed` marker so kotlinBindingScopeFor in
  simple-hooks.ts overrides the scope-extractor's auto-hoist (which
  would otherwise promote unbraced-arm bindings to the function scope
  because the body anchor coincides with the Block scope's range).
- simple-hooks.ts: kotlinBindingScopeFor checks the marker and pins the
  binding to the innermost (Block) scope.

Verification (REGISTRY_PRIMARY_KOTLIN=1):
- Forced-mode: 21 -> 9 failing of 175 (12 fewer; all 12 when/is tests
  now green: lines 957, 966, 975, 1096, 1107, 1118, 1131, 1142, 1153,
  1164, 1182, 1195 in test/integration/resolvers/kotlin.test.ts).
- Default-mode: 175/175 unchanged.
- Full resolver suite: 2216/2216 unchanged.
- Remaining 9 failures are tracked by sibling sub-issues (#1759-#1763).

Does NOT add Kotlin to MIGRATED_LANGUAGES per parent #1746 flip criteria.

Closes #1758. Refs #1746.

Co-authored-by: Test <test@example.com>
2026-05-22 09:44:19 +01:00
Gergő MagyarandTest 9ad48c173e fix(lang-kotlin): overload target-id selection by parameter types (#1761) (#1777)
Same-arity Kotlin class-method overloads collapsed onto whichever node
was registered first. `lookup("alice")` resolved to `lookup(Int)` —
not because the picker chose wrong, but because `resolveDefGraphId`
fell through to the simple-name fallback after its parameter-typed key
lookup missed.

Root cause: `populateKotlinOwners` (which calls
`populateClassOwnedMembers`) assigned `ownerId` and qualified names to
class-owned function defs but left `def.type === 'Function'`. The
graph parsing-processor, in contrast, emits a `Method` node label for
class members. `resolveDefGraphId`'s parameter-typed key lookup is
gated on `def.type === 'Method'` (graph-bridge/ids.ts:108-116), so it
was skipped for every Kotlin class method. With the type-keyed lookup
skipped, the resolver fell through to `simpleKey`, which is
first-wins by registration order — and the Int overload always
registered first in these fixtures.

Fix: `populateKotlinOwners` now upgrades `def.type` from `Function`
to `Method` after `populateClassOwnedMembers` assigns `ownerId`. This
aligns the scope-resolution model with the graph's node labels so
parameter-typed key registration and lookup operate in the same
keyspace.

Picker logic in `pickImplicitThisOverload` / `narrowOverloadCandidates`
was already correct — verified by trace: it narrowed `lookup("alice")`
to the `[String]` def. Only the graph-id lookup was broken.

Verification (REGISTRY_PRIMARY_KOTLIN=1):
- Forced-mode: 21 -> 18 failing of 175 (3 fewer; tests 1620, 1659,
  1692 in `test/integration/resolvers/kotlin.test.ts` now green).
- Default-mode Kotlin: 175/175 unchanged.
- Full resolver suite: 2216/2216 unchanged.
- Remaining 18 failures are tracked by sibling sub-issues
  (#1758, #1759, #1760, #1762, #1763).

Does NOT add Kotlin to MIGRATED_LANGUAGES per parent #1746 flip criteria.

Closes #1761. Refs #1746.

Co-authored-by: Test <test@example.com>
2026-05-22 09:22:53 +01:00
954b184248 fix(cli): apply --no-stats to keep-marker stats line (#1706) (#1765)
* fix(cli): apply --no-stats to keep-marker stats line (#1706)

The keep-marker branch of upsertGitNexusSection rebuilt the index-summary
line on every analyze and always re-injected the volatile counts,
ignoring --no-stats. For teams that commit a trimmed AGENTS.md/CLAUDE.md
with a gitnexus:keep marker, that produced recurring no-value merge
conflicts — exactly what --no-stats exists to prevent.

Thread noStats into upsertGitNexusSection. Under --no-stats the keep-path
stats line becomes "Indexed as **<name>**" with no (N symbols, ...)
parenthetical; the project name still refreshes so renames propagate.
The statsPattern parenthetical is now optional so a count-free line left
by a prior --no-stats run still matches.

* test(cli): cover count-return and AGENTS.md parity for --no-stats keep path

Addresses review findings F1 and F2 on PR #1765:

- F1: add a test that counts RETURN when --no-stats is dropped after a
  prior count-free run — guards against the flag becoming sticky.
- F2: extend the noStats+keep "drops the volatile counts" test to assert
  AGENTS.md alongside CLAUDE.md, so a future asymmetry between the two
  upsertGitNexusSection call sites is caught.

---------

Co-authored-by: Emmanuel Alawode <platforms@chowbea.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-22 07:03:40 +01:00
Copilot be3833d9c9 chore(security): upgrade @vercel/node in gitnexus-web and remediate transitive advisories (#1705) 2026-05-22 06:26:09 +01:00
dependabot[bot] dc96bb048a chore(deps)(deps-dev): bump tsx from 4.21.1 to 4.22.0 in /gitnexus (#1768) 2026-05-22 05:27:51 +01:00
Abhigyan Patwari 5a0f5e81db ci(web): use npm ci for deterministic Vercel installs (#1764) 2026-05-22 05:08:42 +01:00
dependabot[bot] 8c1983a8bf chore(deps)(deps-dev): bump @types/node in /gitnexus (#1767) 2026-05-22 04:41:57 +01:00
231ad71d40 fix(mcp): disambiguate duplicate-name repo resolution for worktrees (#1753)
* fix(mcp): disambiguate duplicate-name repo resolution for worktrees

When multiple indexed repos share the same registry name (main checkout plus linked worktrees), MCP tools no longer silently pick the first sibling. Resolution prefers the repo matching process.cwd()'s git root, throws RegistryAmbiguousTargetError when still ambiguous, and uses canonical path matching aligned with the CLI registry.

Fixes #1658. Complements worktree detect_changes fixes in #1654/#1691.

* fix(mcp): refresh registry on duplicate-name ambiguity before failing

resolveRepo now retries resolveRepoFromCache after RegistryAmbiguousTargetError so stale in-memory siblings clear when the registry changes. Adds detect_changes callTool ambiguity test, registry-refresh regression test, pickRepoHandleForCwd MCP cwd doc, and temp-dir cleanup in #1658 fixtures.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(mcp): PR #1753 review follow-ups + collision-id case bug

Address Findings 3-6 from the production-readiness review on PR #1753,
plus a latent bug surfaced while writing the F5 regression test:

- F3: drop the no-op `try { ... } catch (err) { throw err; }` wrapper
  around the miss-path retry in `resolveRepo`; the catch only re-threw.
- F4: rewrite the misleading "child/repo" example on the relative-path
  tier — `child/repo` would be classified as path-like and never reach
  this branch. Comment now describes bare, separator-free names
  resolved against `process.cwd()`.
- F5: add regression test for the stable hashed-id tier so a duplicate
  sibling can be reached by its `<name>-<hash>` id. Writing this test
  exposed that `repoId()` produced a mixed-case base64url suffix while
  `resolveRepoFromCache` lowercased the param before the Map lookup, so
  collision ids with any uppercase byte in the hash were unreachable.
  Fix: lowercase the hash in `repoId` so it survives `paramLower`.
- F6: add regression test asserting two repos sharing a name prefix
  (`project-a`, `project-b`) cause `resolveRepo("project")` to reject
  as not-found rather than silently returning the first partial match.

* refactor(mcp): tighten PR #1753 follow-up tests + pin hash length

Address three P2 maintainability findings from the ce-code-review pass
on commit aa7f2050:

- Export `REPO_ID_HASH_LENGTH` from local-backend.ts and use it in both
  `repoId()` and the hashed-id test. Closes the silent-drift hole where
  the test's inline formula could fall out of sync with the source
  without any signal.
- Extract `makeSharedPrefixFixture(nameA, nameB)` next to
  `makeDuplicateNameFixture`. Centralises the temp-dir + `.gitnexus`
  scaffolding + `duplicateFixtureDirs.push()` cleanup contract so
  future callers can't drop the cleanup step.
- Reorder the hashed-id test's comment block so the intentional-coupling
  rationale leads, before the description of the formula being mirrored.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* chore: re-run CI

---------

Co-authored-by: Test <test@example.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-21 19:21:25 +01:00
df2ed009ce fix(group): detect httpx AsyncClient alias imports (#1687)
* fix(group): detect httpx AsyncClient alias imports

* fix(group): anchor httpx dotted imports and skip shadowed aliases

Addresses Findings 1-3 of the production-readiness review on PR #1687.

- F1: the `(dotted_name (identifier) @module)` capture matches every
  segment of a dotted module path, so `import package.httpx as hx` and
  `from package.httpx import AsyncClient` would falsely populate the
  alias sets. Anchor the check on `moduleNode.parent?.text === 'httpx'`
  so the full dotted_name must equal `httpx`.

- F2: `moduleAliases` and `asyncClientAliases` were file-global and
  unaware of Python scope. A function-local rebind like
  `AsyncClient = lambda: MockClient()` left the alias entry intact and
  any subsequent `client = AsyncClient(); client.get(...)` emitted a
  false-positive consumer contract. Walk every
  `(assignment left: (identifier) @name)` whose name matches an alias,
  record the enclosing function/class scope as poisoned, and skip
  direct- and module-attribute matches when the call site is inside
  that scope chain.

- F3: extend the existing fixture with dotted-package look-alikes and
  three local-shadow cases (`shadow_direct_alias`, `shadow_module_alias`,
  `shadow_direct_context`) and assert the would-be FP contractIds are
  not emitted.

- F6: refresh the module-level docstring to mention the supported
  import-alias forms and the shadow-exclusion behavior.

* refactor(group): tighten httpx alias shadow detection and broaden tests

Follow-up addressing the residual review findings on PR #1687.

- Replace inline scope-key construction in isAliasShadowed with a
  getScopeKey call so the two helpers cannot drift apart (M1).
- Collapse the double tree traversal in collectHttpxAsyncClients: build
  one combined alias set and pass it to a single
  collectAliasShadowScopes call (perf, P2).
- Add a `shadowScopeKey` helper that returns the scope a rebind actually
  shadows under Python LEGB rules: function scope for in-function
  rebinds, 'module' for top-level rebinds, and `null` for class-body
  rebinds (class attributes do not shadow bare-name lookups in methods).
  Removes the previous blanket `scopeKey === 'module'` skip and now
  correctly poisons module-level rebinds (correctness #1).
- Extend `ALIAS_SHADOW_PATTERNS` to cover tuple, list, and pattern_list
  destructuring targets (correctness #2).
- Rename `ALIAS_REBIND_PATTERNS` to `ALIAS_SHADOW_PATTERNS` and update
  the block comment to say "shadowed" rather than "poisoned" (M4).
- Collapse `callScopeKeys` to a single-line return; the dead Set wrap
  was misleading future readers (M2).

Tests:
- New negative fixtures for 3-segment dotted import
  (`import a.b.c.httpx as deep_evil`), relative import
  (`from .httpx import AsyncClient as rel_evil_async`), tuple
  destructuring rebind, and an isolated file exercising the module-level
  rebind path (T1, correctness #2, expanded F2).
- New positive fixture confirming that a class-body assignment of
  `AsyncClient` does NOT poison the surrounding methods.
- Add a positive control assertion for `module_direct_client` so the
  dotted-package negative assertions cannot pass vacuously (T3).

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Test <test@example.com>
2026-05-21 18:24:27 +01:00
luyua9andGergő Magyar dd3527327d feat(ingestion): Link object literal methods to exported bindings (#1718)
* fix: link object literal methods to exported bindings

* fix(ingestion): bridge object-literal value receivers in scope-resolution (PR #1718 review)

Addresses adversarial production-readiness review on PR #1718 / issue #1358:
- F1 (caller resolution) — setting `ownerId` on object-literal method symbols
  alone is not sufficient; the scope-resolution receiver-bound resolver only
  consults class-like or type-annotated bindings, so lowercase value receivers
  (`export const fooService = {...}; fooService.getUser(...)`) never reach the
  owner-indexed lookup. Adds a Case 5 value-receiver bridge in
  receiver-bound-calls.ts that resolves the receiver name as a Const/Variable
  binding, translates its def to the canonical graph node id, and emits the
  CALLS edge via the owner-indexed method registry.
- F2 (boundary guard) — rewrites findObjectLiteralBindingInfo as an explicit
  two-phase AST walk: Phase A tracks object-literal depth (returns null for
  nested literals and pre-declarator function/class boundaries — IIFE
  patterns); Phase B walks the declarator's ancestors and rejects function,
  class, and block-statement containers (if / for / while / try / catch /
  switch / etc.) before reaching program/export_statement. Prevents false
  HAS_METHOD edges for locally-scoped or block-scoped object literals.
- F4 — drops the dead `ownerName` field from ObjectLiteralBindingInfo.

Constraint: TS/JS are scope-resolution migrated per RFC #909; the legacy
Call-Resolution DAG (call-processor.ts) is intentionally left untouched.

Tests:
- test/integration/ast-helpers-object-literal-binding.test.ts (13 cases) —
  pins helper semantics: happy paths, function/arrow/class-ctor boundaries,
  nested literals, block scope (if / for-of / try), IIFE, assignment
  expressions without declarator.
- test/integration/object-literal-owner-resolution.test.ts (9 cases) —
  drives the full pipeline against an on-disk fixture: sequential CALLS edge
  emission (issue #1358 proof), worker-mode parity, negative local binding,
  and nested-literal attribution boundary.

Full sweep: 2958/2958 integration + 6056/6056 unit tests pass.

* refactor(ingestion): address code-review findings on object-literal owner resolution

Multi-agent code review on the prior commit surfaced 7 actionable findings,
all walked through and applied here. None change observable behavior for
issue #1358's fix; all harden correctness, predicate stability, and test
signal.

- #1 (P1 / 3-reviewer corroboration): Case 5 in receiver-bound-calls.ts no
  longer hand-builds graph.addRelationship + a dedup key. New
  tryEmitEdgeWithExplicitTargetId in edges.ts takes a pre-resolved target
  id (the canonical Method nodeId from the parser) and reuses every
  invariant of tryEmitEdge: dedup-key format, collapse-flag honoring,
  caller-id resolution, rel-id shape, mapReferenceKindToEdgeType for
  read/write ACCESSES. This also lands the adversarial reviewer's "F2"
  follow-up (hardcoded type: 'CALLS' for non-call sites) for free.

- #2 (P2 cross-reviewer): findValueBindingInScope's predicate inverted
  from denylist ("not class-like and not callable") to explicit allowlist
  matching reconcileOwnership's registration set:
  Const | Variable | Property | Static. Extracted as isOwnableValueLabel
  so future NodeLabel additions require an explicit opt-in.

- #6 (P2): walkScopeChain<T>() extracted; both findClassBindingInScope
  and findValueBindingInScope now route through it. Local scope.bindings
  are exhausted BEFORE lookupBindingsAt (imported/augmented) at every
  scope level — preserves JavaScript lexical scoping where a local const
  shadows an imported binding of the same name. Behavior was already
  correct in findClassBindingInScope but was implicit; now it is the
  walker's explicit, documented contract.

- #7 (P2): scope-walker duplication closed. findClassBindingInScope and
  findValueBindingInScope reduce to thin wrappers over walkScopeChain
  with their respective predicate. findClassBindingInScope keeps its
  qualifiedNames + dotted-name fallback tail.

- #3 (P2): parse-worker.ts hoists `const ownerId = enclosingClassId ??
  objectLiteralOwnerInfo?.ownerId` once before the symbol push, dropping
  the duplicated coalesce + `as string` cast. Matches the cast-free
  pattern at parsing-processor.ts:793. HAS_METHOD emit site reuses the
  same hoisted local.

- #4 (P2): object-literal-owner-resolution.test.ts Test A's CALLS-edge
  assertion no longer matches by name alone. .toEqual now pins the
  canonical target id (Method:src/service.ts:getUser#1 via generateId),
  confidence (0.85), and reason ('import-resolved'). A regression that
  emits the edge at confidence=0, with the wrong reason, or against a
  phantom Method node now fails the test.

- #5 (P2): worker-parity test adds a CI tripwire — when CI=1 and
  dist/parse-worker.js is missing, throw at module top with a clear
  message. Locally, skipIf(!hasDistWorker) keeps the fast-iteration
  experience; CI cannot pass with U3 (worker-path ownerId) unverified.

Verification: tsc --noEmit clean. Targeted regression sweep on
ast-helpers-object-literal-binding (13), object-literal-owner-resolution
(9), has-method (60), cross-file-binding (40) — 122/122 pass. Full unit
sweep: 6056/6056. Integration suite: 1 pre-existing Windows-flake in
worker-pool.test.ts (passes 28/28 in isolation) unrelated to this diff.

* refactor(scope-resolution): align Const label emission with legacy DAG (PR #1718 review F1)

Eliminates the architectural fragility surfaced by PR #1718's adversarial review
Finding 1. Previously, normalizeNodeLabel('const') returned 'Variable' while
the legacy DAG parse phase emits 'Const' graph nodes (via @definition.const
capture for lexical_declaration). PR #1718's Case 5 value-receiver bridge
resolved correctly only because resolveDefGraphId happened to fall back to
simpleKey after the qualified-key miss — accidental correctness.

After this change, scope-resolution defs for `const x = ...` declarations
report def.type === 'Const', matching the graph node label. resolveDefGraphId's
qualified-key path now hits on the first try; the simple-key fallback is no
longer load-bearing for value receivers and can be tightened in future without
silently breaking Case 5.

Audit completeness verification:
- Grep `\bVariable\b` across src/core/ingestion/scope-resolution/ surfaced two
  consumer sites that already accept both labels: reconcile-ownership.ts:101+168
  (`def.type === 'Variable' || def.type === 'Const' || ...`) and
  walkers.ts:207 isOwnableValueLabel (`Const | Variable | Property | Static`).
  No language hook in src/core/ingestion/languages/ branches on
  `def.type === 'Variable'` for what's actually a const declaration.
- Sentinel stress test (the full unit + integration suite run with the
  renamed label in place): 6137/6137 unit tests pass; 2967/2967 integration
  tests pass. One pre-existing Windows-only flake on worker-pool.test.ts when
  run alongside the full integration suite (passes 28/28 in isolation,
  unrelated to scope-extractor — same flake observed before this diff).

The variable mapping (`'variable' → 'Variable'`) is preserved for `var`
declarations, matching the legacy DAG's `@definition.variable` capture for
variable_declaration. The split now mirrors the parse-phase capture
distinction exactly.

Per plan docs/plans/2026-05-21-002-feat-pr1718-followups-class-instance-and-label-normalization-plan.md
U4 + U5. T1 (class-instance singleton resolution from issue #1358's second
sub-case) is deferred to a standalone pre-plan investigation, not shipped
here.

* test(ingestion): add regression coverage for issue #1358 singleton sub-cases

Closes the remaining sub-cases of issue #1358 surfaced by PR #1718's
adversarial review (Finding 4, NOTED): the class-instance singleton
(`export const fooService = new FooService();`) and the factory-pattern
singleton (`export const fooService = makeFooService();`).

Pre-plan investigation (per docs/plans/2026-05-21-002 § "Pre-Plan
Investigation Task (T1)") confirmed Outcome A for both patterns — they
already resolve end-to-end through scope-resolution's
`@type-binding.constructor` capture (languages/typescript/query.ts:489-511)
+ `propagateImportedReturnTypes` chain-follow
(scope-resolution/passes/imported-return-types.ts:114) + receiver-bound
Case 4 simple typeBinding lookup (receiver-bound-calls.ts:625). The
mechanism was wired correctly before this session; the regression-net
wasn't.

This test pins the behavior:
- Pattern 1: `caller → FooService.getUser` CALLS edge with
  confidence 0.85 and reason 'import-resolved'
- Pattern 2: same edge shape via factory chain-follow (the
  `@type-binding.alias` capture for `const u = find()` style)

Both assertions use exact `.toEqual([{...}])` shape pinning so a future
regression that targets a phantom Method node, emits at lower confidence,
or drops the cross-file import-resolved reason fails loudly.

Verification: 5/5 pass, 127/127 in targeted regression sweep including
object-literal-owner-resolution.test.ts, ast-helpers-object-literal-
binding.test.ts, has-method.test.ts, and cross-file-binding.test.ts.

No production code change. The class methods get a class-qualified node id
(`Method:src/service.ts:FooService.getUser#1`) distinguishing them from
same-name methods on other classes — distinct from the bare-name node id
shape PR #1718's object-literal case uses.

* test(resolvers): add class-instance + factory-pattern singleton coverage for TS/JS (issue #1358)

Closes the remaining sub-cases of issue #1358 surfaced by PR #1718's
adversarial review (Finding 4). PR #1718 fixed object-literal-shorthand
singletons (`export const fooService = { getUser() {} }`); this commit adds
parallel coverage for the two other singleton shapes that resolve through
the existing scope-resolution chain:

  // Pattern 1 — class-instance singleton
  export class FooService { getUser(id) { ... } }
  export const fooService = new FooService();

  // Pattern 2 — factory-pattern singleton
  export class FooService { getUser(id) { ... } }
  export function makeFooService() { return new FooService(); }
  export const fooService = makeFooService();

Pre-plan investigation (per local plan docs/plans/2026-05-21-002 § "Pre-Plan
Investigation Task (T1)") confirmed Outcome A — both patterns already
resolve end-to-end through:
  - `@type-binding.constructor` capture (languages/{typescript,javascript}/
    query.ts) seeds `fooService → FooService` at parse time
  - `propagateImportedReturnTypes` (scope-resolution/passes/
    imported-return-types.ts:114) mirrors the typeBinding cross-file
  - Receiver-bound Case 4 simple typeBinding lookup
    (scope-resolution/passes/receiver-bound-calls.ts:625) MRO-walks
    FooService and emits the CALLS edge to getUser

Tests added per language × pattern (5 each, 10 total):
- node existence (Class, Method, Function, Const, plus Function for the
  factory pattern's `makeFooService`)
- HAS_METHOD edge from class to method (class-instance variant)
- CALLS edge from caller to `getUser` with `targetFilePath: 'src/service.{ts,js}'`,
  `reason: 'import-resolved'`, `confidence: 0.85` — exact `.toEqual([{...}])`
  shape pinning so a regression that emits at lower confidence or drops the
  cross-file reason fails loudly

Fixtures placed under the existing `test/fixtures/lang-resolution/` convention.
Tests appended to `test/integration/resolvers/{typescript,javascript}.test.ts`,
matching the in-file pattern of every other resolver scenario.

Also supersedes and removes the standalone
`test/integration/class-instance-and-factory-singleton-resolution.test.ts`
introduced earlier in this PR session (`0df91b77`) — the proper home for
language-resolver scenarios is the per-language resolver test file alongside
similar fixtures (`javascript-self-this-resolution`, `javascript-cross-file`,
`typescript-tsconfig-paths`, etc.). One canonical location for the scenario,
not two.

Verification: 10/10 new singleton tests pass; 297/297 full TS+JS resolver
suite pass (no regression in any existing resolver test).

* test(resolvers): gate TS/JS singleton tests behind scope-resolution parity (CI run 26223603426)

The class-instance and factory-pattern singleton CALLS-edge resolution
tests added in c8e573bc rely on scope-resolution-only mechanisms
(`@type-binding.constructor` capture + `propagateImportedReturnTypes`
mirror + receiver-bound Case 4). The `scope-parity / typescript parity`
and `scope-parity / javascript parity` CI jobs run with
`REGISTRY_PRIMARY_TYPESCRIPT=0` / `REGISTRY_PRIMARY_JAVASCRIPT=0` and
exercise the legacy DAG path, which has no cross-file constructor-derived
typeBinding propagation. Verified by job 77202610819 (TS parity) and
77202610869 (JS parity) failing with:

  × resolves caller.fooService.getUser() to FooService.getUser via constructor-inferred typeBinding
  × resolves caller.fooService.getUser() through the factory chain to FooService.getUser

Note: my local Windows shell-prefix env-var invocation did not propagate
the flag into vitest workers correctly (the cpp parity gate's 47-skipped
behavior masked the issue when I ran an ad-hoc comparison), so the
empirical "both modes pass" finding I posted earlier was wrong. CI is the
source of truth.

Changes:
- test/integration/resolvers/helpers.ts: add `typescript` and `javascript`
  entries to `LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES` for the 2 CALLS-edge
  resolution tests in each language. Node-existence and HAS_METHOD
  assertions are NOT excluded — those pass under legacy DAG (parser-level
  emission is intact).
- test/integration/resolvers/typescript.test.ts: drop the `it` import from
  vitest; replace with `const it = createResolverParityIt('typescript');`
  shadow (matches the c/cpp/csharp/go pattern at the top of those files).
- test/integration/resolvers/javascript.test.ts: same shadow with
  `createResolverParityIt('javascript')`.

Verification:
- Default mode (registry-primary): 297/297 TS+JS resolver tests pass.
- Legacy DAG mode: the 4 listed singleton CALLS-edge tests will skip; all
  other singleton assertions (node existence + HAS_METHOD edge) continue
  to run and pass under both modes.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-21 17:18:27 +01:00
Gergő MagyarandCursor d3de5fa5d5 fix(install): materialize vendored grammars to fix Windows EPERM (#1728) (#1729)
* fix(install): materialize vendored grammars to fix Windows EPERM (#1728)

Stop using file: optionalDependencies for tree-sitter-dart/proto/swift,
which made npm symlink vendor paths on install and fail on Windows without
symlink privileges. Copy vendor trees into node_modules at postinstall
instead; keep native builds and #836 vendor hygiene.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(install): atomic materialize swap + fail-soft tests (#1728, #836)

Hardens PR #1729 against two issues the original implementation could
still hit:

1. Torn-state on rmSync→cpSync. The previous loop deleted the
   destination before copying. If cpSync threw — the exact Windows EPERM
   scenario this PR targets — a previously-working grammar was silently
   wiped. Now we copy to {dest}.materialize-tmp first and renameSync into
   place, so an interrupted copy leaves the prior materialization intact.

2. Fail-soft try/catch had no test coverage. Adds two POSIX-only tests
   (chmod 0o555 to deterministically force cpSync to throw) that verify
   (a) a single grammar failure does not abort the other two, and (b) an
   existing materialization survives a partial-copy failure. Skipped on
   Windows where chmod doesn't enforce write restriction; runs on Linux
   CI.

Other test improvements locking in the install-hygiene invariants:

- All three vendored grammars (dart/proto/swift) checked, not just dart.
- GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1 short-circuit is exercised.
- Vendor cleanliness (#836): no node_modules/build under vendor/.
- Idempotent re-runs (clean overwrite verified via sentinel file).
- Missing-vendor warn+continue path now has explicit coverage.
- Vendored package manifests asserted to carry no install script or
  runtime dependencies.
- package.json optionalDependencies asserted free of vendored grammars.
- package-lock.json assertion tightened from `if (entry !== undefined)
  { expect(entry.link).not.toBe(true); }` (vacuous when entry is absent,
  i.e. the expected post-fix state) to `expect(...).toBeUndefined()`.

Verified locally:
- npx tsc --noEmit: clean
- vitest test/unit/materialize-vendor-grammars.test.ts: 8 pass + 2
  POSIX-only skipped on Windows
- npm pack tarball: no vendor/*/node_modules or vendor/*/build entries
- Isolated global install (clean + upgrade + SKIP env) into temp prefix:
  succeeds; gitnexus --version → 1.6.5; vendor stays clean post-install.

* fix(install): address review feedback — Swift parity, atomicity, CI smoke

Resolves all findings from the automated production-readiness review on
verify/issue-1728-symlink.

Swift warning parity (review #2):
  Add tree-sitter-swift to OPTIONAL_GRAMMARS in src/cli/optional-grammars.ts
  alongside Dart and Proto. Before this commit, Swift was materialized at
  postinstall and probed by build-tree-sitter-swift.cjs but the runtime
  warnMissingOptionalGrammars() never warned when it failed to load —
  users got silent Swift degradation from the optional-grammars surface
  (parser-loader's separate unavailableNote only fires on demand). Now
  the warning path matches the materialize path.

README env-var table (review #1):
  Update the GITNEXUS_SKIP_OPTIONAL_GRAMMARS row at README.md line 248 to
  list all three vendored grammars (dart, proto, swift). The quick note
  earlier in the README already mentioned all three; only the table row
  was stale.

Atomicity hardening (review #3):
  materialize-vendor-grammars.cjs now copies to {dest}.materialize-tmp,
  renames the existing dest to {dest}.materialize-bak (if present), then
  renames the partial into dest, then removes the backup. If the
  partial→dest rename fails (e.g. Windows AV scanner racing the swap),
  the catch block restores from backup so the previously-materialized
  grammar is preserved. Closes the narrow torn-state window where the
  prior implementation could leave dest deleted after rmSync succeeded
  but renameSync failed.

Swift probe docs (review #4):
  build-tree-sitter-swift.cjs script header rewritten to describe what
  the script actually does — probe node-gyp-build at install time so
  missing-prebuild failures surface as install-time warnings instead of
  first-parse runtime errors. The script does not "activate" anything;
  the runtime require() in parser-loader does the actual load. Console
  warning text updated to match ("prebuild probe" not "activation").

Windows packaged-install smoke test (review #5):
  New CI job `packaged-install-smoke` in .github/workflows/ci-tests.yml
  matrices on windows-latest and ubuntu-latest. Runs npm pack, installs
  the produced tarball globally into RUNNER_TEMP, then asserts:
    * no vendor/*/node_modules or vendor/*/build (#836 invariant)
    * tree-sitter-{dart,proto,swift} in node_modules are real
      directories, not junctions/symlinks (#1728 invariant)
    * gitnexus --version runs against the installed CLI
  Closes the coverage gap where the existing windows-latest job only
  ran `npm ci` in the source checkout — exercising postinstall but not
  the tarball reify step that historically tripped EPERM.

Verified locally:
  npx tsc --noEmit: clean
  vitest test/unit/materialize-vendor-grammars.test.ts test/unit/cli-commands.test.ts:
    18 pass + 2 POSIX-only skipped on Windows
  prettier + eslint on all changed files: clean

* fix(ci): disable credential persistence on packaged-install-smoke checkout

GitHub Advanced Security (zizmor artipacked) flagged the new
packaged-install-smoke job's actions/checkout step as a potential
credential-persistence risk. The job runs `npm pack` + global install
and never pushes back, so the GITHUB_TOKEN that checkout would persist
in .git/config provides no value and only widens the leak surface (any
future artifact-upload step in this job would carry the token).

Disable persistence explicitly via `persist-credentials: false` on this
job's checkout. Scoped to the new job — pre-existing checkouts above
are left unchanged.

* fix(ci): use find instead of ls for tarball lookup (SC2012)

actionlint shellcheck SC2012 flagged `TARBALL=$(ls gitnexus-*.tgz | head -n1)`.
Switch to `find . -maxdepth 1 -name 'gitnexus-*.tgz' -print -quit` which
handles non-alphanumeric filenames safely. Also add an explicit
empty-result check so the failure mode is a clear error message instead
of a silent `npm install -g ""` later.

* fix(tests): sabotage vendor src (not partial path) in POSIX fail-soft tests

The fail-soft tests in materialize-vendor-grammars.test.ts pre-chmod'd
the destination's .materialize-tmp partial directory to 0o555 to force
cpSync to throw. After the atomicity rewrite (`fix(install): atomic
materialize swap + fail-soft tests`), the materialize script now starts
each grammar's loop with `fs.rmSync(partial, { force: true })`, which
deletes the chmod'd sabotage before cpSync runs — so cpSync succeeds and
the partial is then renamed into dest, leaving the test's `finally`
block with no path to chmod back (ENOENT) and the assertion that proto
remained unmaterialized failing because it materialized cleanly.

Fix: sabotage the *vendor source* directory (which the script reads from
but never modifies) by chmod'ing it to 0o000. cpSync then fails on
readdir, the catch block fires per-grammar, dart and swift still
materialize from their unaffected sources, and the existing-dest
preservation test verifies that a sabotaged second-run leaves the prior
materialization (and its sentinel file) intact.

Tests now pass locally (8 pass + 2 POSIX-only skipped on Windows) and
should pass on macOS/Ubuntu CI where the sabotage runs.

* fix(tests): restrict fail-soft tests to Linux (macOS Node cpSync abort)

Node 22 on macOS aborts the process with `libc++abi: terminating due
to uncaught exception filesystem_error` when fs.cpSync hits a source
directory it can't read — the abort happens at the C++ filesystem layer
and bypasses Node's JS try/catch entirely (nodejs/node#51399). My
chmod-0o000-the-source sabotage strategy triggers this SIGABRT on
macOS CI before the production script's `try { cpSync } catch` ever
runs, so the test sees a child-process crash instead of the fail-soft
warning it's verifying.

The production script's fail-soft is correct on Linux (where EACCES
surfaces as a normal JS exception) and effectively untestable on macOS
via permission sabotage. Real installs don't hit this — npm always
ships vendor/ with readable permissions — so the macOS gap is a test
artifact, not a behavior gap.

Restrict the two chmod-based tests to Linux only by replacing
`skipOnWin` with `linuxOnly`. Linux CI continues to verify both the
one-grammar-fails-others-succeed and existing-materialization-preserved
invariants. macOS and Windows runs skip these two scenarios; the other
8 tests still run on every platform.

* fix(tests): remove materialize unit tests, rely on CI smoke job

The materialize-vendor-grammars.test.ts file has been a recurring source
of platform-specific CI noise:

  - Windows: chmod doesn't enforce read/write restrictions the way POSIX
    does, so the fail-soft tests had to be skipped there.
  - macOS Node 22: cpSync against an unreadable source aborts the process
    with a libc++ filesystem_error (nodejs/node#51399) that bypasses JS
    try/catch entirely — making the chmod-based fail-soft tests
    unrunnable on macOS too.
  - The "vendor-cleanliness" and "idempotency" tests on Windows
    intermittently flake due to fs.cpSync timing on the GitHub runner.

The invariants these tests verified are now covered by stronger,
more realistic surfaces:

  - packaged-install-smoke (ci-tests.yml): runs `npm pack` then
    `npm install -g ./gitnexus-*.tgz` on windows-latest and
    ubuntu-latest, then asserts no vendor/*/node_modules,
    no vendor/*/build (#836), no junctions/symlinks on the
    materialized grammar directories (#1728), and a working
    `gitnexus --version`. This is the actual end-user install path.

  - cli-commands.test.ts (kept, unmodified): asserts package.json
    declares no `file:` optionalDependencies for vendored grammars,
    the Swift vendor manifest carries no install script or
    dependencies, and the postinstall chain runs
    materialize-vendor-grammars.cjs + build-tree-sitter-swift.cjs.
    These are static manifest checks — deterministic, fast, no
    flake risk.

Removing the dynamic script-execution tests trades unit-level coverage
for end-to-end smoke coverage that actually exercises the
`file:` → cpSync change against a real npm install lifecycle, on
the platform the fix targets (windows-latest).

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-21 16:47:22 +01:00
4c06d64a3b chore(deps)(deps): bump zod from 4.3.6 to 4.4.3 in /gitnexus-web (#1736)
Bumps [zod](https://github.com/colinhacks/zod) from 4.3.6 to 4.4.3.
- [Release notes](https://github.com/colinhacks/zod/releases)
- [Commits](https://github.com/colinhacks/zod/compare/v4.3.6...v4.4.3)

---
updated-dependencies:
- dependency-name: zod
  dependency-version: 4.4.3
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-21 16:25:35 +01:00
ChamHerryandwangxc 2a3d14057a fix(analyze): prevent cache-hit native workers from aborting (#1751)
* fix(analyze): prevent cache-hit native workers from aborting

Delay parse worker startup until a cache miss requires it, fall back to sequential parsing when initial worker readiness fails, and preserve analyzer diagnostics/progress when heap respawn captures child output.

Constraint: Node 25 and tree-sitter/N-API worker initialization can abort before ready, while warm-cache analysis should not start workers at all.

Rejected: Treating status-134/SIGABRT as heap OOM unconditionally | native worker aborts require distinct recovery guidance and stderr/stdout evidence.

Rejected: cli-progress noTTYOutput for respawn progress | it appends newline frames instead of preserving one-line redraw UX.

Confidence: high

Scope-risk: moderate

Directive: Keep parse-worker creation behind confirmed cache misses and preserve TTY-style progress when respawn pipes stderr for crash classification.

Tested: GitNexus impact analysis for ensureHeap, runChunkedParseAndResolve, createWorkerPool, WorkerPool, walkRepositoryPaths; GitNexus detect_changes scoped to staged worktree; targeted vitest for analyze respawn, parse lazy cache, filesystem walker, worker pool; npx tsc --noEmit; npm run build; NODE_OPTIONS='--max-old-space-size=8192' npm test.

Not-tested: Windows terminal rendering and published npm package install path.

* ci(docker): tolerate slower arm64 TypeScript builds

Docker PR builds run gitnexus prepare under QEMU for linux/arm64, where the fixed 120s TypeScript timeout can kill otherwise healthy builds. Increase the default timeout and allow GITNEXUS_BUILD_TIMEOUT_MS to tune slower environments without changing the build steps.

Constraint: PR #1751 Docker Build & Push gitnexus failed with spawnSync /bin/sh ETIMEDOUT while running node_modules/.bin/tsc in scripts/build.js.\nRejected: Rerunning CI only | the failure was the build script's deterministic timeout boundary under arm64 emulation, not a code assertion.\nConfidence: high\nScope-risk: narrow\nDirective: Keep build timeout changes in scripts/build.js configurable; do not hide real compiler failures, only allow slower successful compiles to finish.\nTested: GitNexus impact for gitnexus/scripts/build.js reported LOW; gitnexus detect_changes reported 1 changed file, 0 affected processes, low risk; git diff --check; gitnexus npm run build.\nNot-tested: GitHub Docker arm64 build rerun before pushing; local Docker multi-platform build under QEMU.

* fix(analyze): truncate respawn progress safely

Preserve complete ANSI escape sequences and grapheme boundaries when the respawn progress terminal shim truncates wrapped output, so the shim does not emit dangling escape bytes or split surrogate pairs while keeping raw writes untouched.

Constraint: Claude review on PR #1751 flagged `s.slice(0, width)` in createAnsiPipeTerminal.write() as a latent terminal-corruption risk.
Rejected: Adding a display-width dependency | a local helper is sufficient for this narrow respawn terminal shim and avoids new dependency churn.
Rejected: Changing silent status-134 classification | current tests already document the output-less 134 fallback as heap guidance.
Confidence: high
Scope-risk: narrow
Directive: Keep respawn terminal writes ANSI-aware and preserve rawWrite bypass semantics for callers that intentionally write control sequences.
Tested: GitNexus impact for createAnsiPipeTerminal reported LOW; GitNexus detect_changes reported 2 changed files, 3 affected processes, medium risk; targeted vitest for analyze respawn progress and heap respawn; gitnexus npx tsc --noEmit; prettier check for changed files; eslint for changed files.
Not-tested: Full npm test suite; manual terminal rendering on Windows.

---------

Co-authored-by: wangxc <wangxc_a_bj@si-tech.com.cn>
2026-05-21 16:17:02 +01:00
a9fef2c68d fix(lbug): keep serve stable when sidecars are missing (#1747)
* fix(lbug): keep serve stable when sidecars are missing

Shared missing-shadow WAL recovery prevents repeated read-only open warnings when LadybugDB sidecars are absent, while the Express preflight fix keeps `gitnexus serve` compatible with Express 5 route parsing.

Constraint: LadybugDB read-only replay can require a `.shadow` sidecar that may be absent after interrupted writes or checkpoint edge cases.
Rejected: keep reactive WARN-only quarantine in each adapter | it leaves repeated user-visible warnings and duplicate recovery behavior.
Confidence: high
Scope-risk: broad
Directive: Do not silently delete large orphan WALs; only quarantine tiny orphan WALs before open and keep large WALs for explicit recovery.
Tested: cd gitnexus && npx vitest run test/unit/sidecar-recovery.test.ts test/unit/lbug-adapter-wal-schema.test.ts test/unit/pool-wal-recovery.test.ts test/unit/web-ui-serving.test.ts && npx tsc --noEmit
Not-tested: full npm test in this split branch; full unit suite passed on the source branch before PR split.

Co-authored-by: OmX <omx@oh-my-codex.dev>

* fix(lbug): pool-caller ENOENT guard, symmetric size gate, permission-aware errors (PR #1747 review)

Addresses the production-readiness review of PR #1747 (Findings 1, 2, 3 of 6).
Findings 4, 5, 6 are deferred to follow-ups per the plan.

1. ENOENT-tolerance scoped to pool-adapter callers only
   - `quarantineWalForMissingShadow` stays strict in `sidecar-recovery.ts`.
     The direct adapter calls it inside `acquireInitLock` (cross-process
     file lock) — ENOENT there means the file vanished under lock and
     remains a real bug to surface.
   - New `tryQuarantineForMissingShadow` local helper in `pool-adapter.ts`
     returns a discriminated union { kind: 'quarantined', path } |
     { kind: 'peer-handled' }. Catches ENOENT, re-verifies via
     statIfExists, and converts to 'peer-handled' only when WAL really
     is gone. Defensive: if ENOENT but WAL still present, throws as
     classified error rather than silently returning success.

2. Symmetric WAL-size gate on both recovery paths
   - `refuseLargeWalQuarantine` applied in both
     `reopenReadOnlyAfterMissingShadow` and
     `reopenWritableAfterMissingShadow`. Closes the read-only data-loss
     vector (large orphan WAL silently discarded would never be replayed
     by a later writable open).

3. Permission-aware error classifier
   - New `renameFailureMessage` and `isPermissionRenameError` in
     `sidecar-recovery.ts`. EACCES / EPERM / EBUSY now surface a
     permission-specific message pointing at ACLs, AV exclusions, and
     file-locks. Other codes (ENOSPC, EROFS, EIO, ENOENT) fall through
     to `shadowSidecarRecoveryMessage`.
   - Used at both pool-adapter and direct-adapter caller catches around
     `quarantineWalForMissingShadow`.
   - `doInitLbug`'s pass-through classifier extended to include the new
     permission message. The lock-retry substring match tightened so
     "file-lock error" in the permission message is not mistaken for a
     LadybugDB lock-retry trigger.

Tests
   - sidecar-recovery.test.ts: 7 new tests for `renameFailureMessage` and
     `isPermissionRenameError`.
   - pool-wal-recovery.test.ts: 6 new tests covering ENOENT race,
     EACCES/EPERM/EBUSY classification, ENOSPC fallthrough, and the
     defensive "WAL still present after ENOENT" branch.
   - lbug-adapter-wal-schema.test.ts: 5 new tests covering the symmetric
     size gate on both recovery paths, including the boundary at exactly
     TINY_ORPHAN_WAL_BYTES (4096) and the off-by-one at 4097.

Deferred (tracked as follow-up work)
   - Brittle LadybugDB error-string matching (Finding 4).
   - PNA header end-to-end coverage gap (Finding 5).
   - warnedKeys module-global persistence (Finding 6).
   - Cross-process init lock for pool-adapter.

* fix(lbug): dedup shadow-replay predicate + counter-based warn anti-spam (PR #1747 review, Findings 4 & 6)

Smallest viable response to the two remaining non-blocking findings from the
production-readiness review of PR #1747. An earlier-revision plan proposed
regex widening + a near-miss detector + per-dbPath warn scoping; an
adversarial doc-review found those defended against hypothetical strings
LadybugDB does not produce, added observability theater with no recovery
behavior change, and did not actually fix the long-running gitnexus serve
case for hot dbPaths (where finalizeLbugSidecarsAfterClose rarely fires).
Scope shrunk to dedup + counter-based — strictly behavior-changing and
fully testable.

Finding 4 — dedup + version-coupling markers
   - `isReadOnlyShadowReplayError` was inlined in both `lbug-adapter.ts:451`
     and `pool-adapter.ts:317`. Centralized as an export from
     `sidecar-recovery.ts`. The two local copies are removed; both adapters
     now import from the shared module.
   - Both LadybugDB-coupled predicates (`isMissingShadowSidecarError` and
     `isReadOnlyShadowReplayError`) gain a `// LADYBUGDB-CONTRACT:` marker
     comment citing `@ladybugdb/core ^0.16.1`. When bumping LadybugDB,
     `git grep "LADYBUGDB-CONTRACT"` enumerates every version-coupled spot.
   - Strict matcher unchanged — when LadybugDB actually changes the error
     format, the failure mode stays loud (raw native error propagates) and
     the markers make every affected predicate trivially greppable.

Finding 6 — counter-based warn anti-spam
   - `warnedKeys: Set<string>` → `warnedKeyCounts: Map<string, number>`.
     `warnOnce` keeps its signature `(logger, key, message)` and keying
     convention unchanged — the swap is internal.
   - `WARN_MILESTONES = [1, 10, 100, 1000, 10000]`. Logarithmic spacing
     gives O(log N) warns for a condition that fires N times. Past the
     first occurrence the warn message is suffixed with "(Nth occurrence
     of this condition)" so persistence is visible in the log line itself.
   - Solves the long-running serve case: a hot dbPath hitting the same
     condition 100 times now fires 3 warns (occurrences 1, 10, 100)
     instead of 1 warn + 99 silent debug lines.

Tests (10 new in sidecar-recovery.test.ts, all green)
   - Centralized isReadOnlyShadowReplayError: positive match, false-positive
     guard, structural assertion that the duplicate regex is gone from both
     adapter files, LADYBUGDB-CONTRACT marker count.
   - Counter-based warnOnce: milestone-at-10 with suffix, milestone-at-100,
     key isolation across dbPaths, reset zeroes the counter, first-occurrence
     message does NOT carry the suffix.

Deferred (tracked separately)
   - Finding 5 — PNA header end-to-end coverage gap (CORS boundary is sound).
   - LadybugDB structured error codes (if/when the library exposes them).
   - Per-call milestone configurability — re-open if tuning is needed.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* ci: trigger CI rebuild

---------

Co-authored-by: wangxc <wangxc_a_bj@si-tech.com.cn>
Co-authored-by: OmX <omx@oh-my-codex.dev>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-21 12:35:43 +01:00
dependabot[bot] 8d71847791 chore(deps)(deps): bump @tailwindcss/vite in /gitnexus-web (#1734)
Bumps [@tailwindcss/vite](https://github.com/tailwindlabs/tailwindcss/tree/HEAD/packages/@tailwindcss-vite) from 4.2.4 to 4.3.0.
- [Release notes](https://github.com/tailwindlabs/tailwindcss/releases)
- [Changelog](https://github.com/tailwindlabs/tailwindcss/blob/main/CHANGELOG.md)
- [Commits](https://github.com/tailwindlabs/tailwindcss/commits/v4.3.0/packages/@tailwindcss-vite)

---
updated-dependencies:
- dependency-name: "@tailwindcss/vite"
  dependency-version: 4.3.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:07:41 +01:00
dependabot[bot] 061e123d72 chore(deps): bump actions/dependency-review-action from 4.9.0 to 5.0.0 (#1739)
Bumps [actions/dependency-review-action](https://github.com/actions/dependency-review-action) from 4.9.0 to 5.0.0.
- [Release notes](https://github.com/actions/dependency-review-action/releases)
- [Commits](https://github.com/actions/dependency-review-action/compare/2031cfc080254a8a887f58cffee85186f0e49e48...a1d282b36b6f3519aa1f3fc636f609c47dddb294)

---
updated-dependencies:
- dependency-name: actions/dependency-review-action
  dependency-version: 5.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:07:22 +01:00
dependabot[bot] a4954368ad chore(deps): bump release-drafter/release-drafter from 7.2.1 to 7.3.0 (#1740)
Bumps [release-drafter/release-drafter](https://github.com/release-drafter/release-drafter) from 7.2.1 to 7.3.0.
- [Release notes](https://github.com/release-drafter/release-drafter/releases)
- [Commits](https://github.com/release-drafter/release-drafter/compare/563bf132657a13ded0b01fcb723c5a58cdd824e2...c2e2804cc59f45f57076a99af580d0fedb697927)

---
updated-dependencies:
- dependency-name: release-drafter/release-drafter
  dependency-version: 7.3.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:07:07 +01:00
dependabot[bot] be1071143a chore(deps)(deps): bump react-syntax-highlighter in /gitnexus-web (#1731)
Bumps [react-syntax-highlighter](https://github.com/react-syntax-highlighter/react-syntax-highlighter) from 16.1.0 to 16.1.1.
- [Release notes](https://github.com/react-syntax-highlighter/react-syntax-highlighter/releases)
- [Changelog](https://github.com/react-syntax-highlighter/react-syntax-highlighter/blob/master/CHANGELOG.MD)
- [Commits](https://github.com/react-syntax-highlighter/react-syntax-highlighter/compare/v16.1.0...v16.1.1)

---
updated-dependencies:
- dependency-name: react-syntax-highlighter
  dependency-version: 16.1.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:06:41 +01:00
dependabot[bot] 3d8aa7f435 chore(deps): bump github/codeql-action from 4.35.3 to 4.35.4 (#1738)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.35.3 to 4.35.4.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/e46ed2cbd01164d986452f91f178727624ae40d7...68bde559dea0fdcac2102bfdf6230c5f70eb485e)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.35.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:06:15 +01:00
dependabot[bot] 4606e24f25 chore(deps)(deps): bump dompurify from 3.4.2 to 3.4.3 in /gitnexus-web (#1735)
Bumps [dompurify](https://github.com/cure53/DOMPurify) from 3.4.2 to 3.4.3.
- [Release notes](https://github.com/cure53/DOMPurify/releases)
- [Commits](https://github.com/cure53/DOMPurify/compare/3.4.2...3.4.3)

---
updated-dependencies:
- dependency-name: dompurify
  dependency-version: 3.4.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-21 12:06:02 +01:00
74653a8ffc feat(web): Support GitLab repository urls. (#1565)
Add GitLab URL input mode alongside existing GitHub and local modes:
- GitLab URL validation for gitlab.com and self-hosted instances
- GitLab icon component (custom SVG, matching existing GitHub icon pattern)
- Mode tab UI with GitLab option
- Backend API integration for GitLab HTTPS URLs

No token configuration included — public repositories supported only.

Close: #378


AI-model: kimi-for-coding/k2p6

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-21 10:52:01 +01:00
Gergő MagyarandCursor 8db51184ab fix(server): restore gitnexus serve startup under Express 5 (#1749)
* fix(server): restore gitnexus serve startup under Express 5

Express 5 rejects app.options('*'), which broke CI e2e when the backend
failed to start. Move PNA middleware before cors so preflight responses
include Access-Control-Allow-Private-Network, and add regression tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(server): address PR review — prettier, ephemeral port, cleanup

- Format integration and rate-limit test files for CI quality/format
- Use OS-assigned port instead of random 47xxx range
- Remove per-test GITNEXUS_HOME temp dir in afterEach
- Use regex for PNA-before-cors structural guard (indent-agnostic)

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-21 10:18:09 +01:00
1b5c6e5b6a feat(ingestion): add Kotlin scope resolver (#1727)
* feat(ingestion): add Kotlin scope resolver

* fix(ingestion): tighten Kotlin scope captures

---------

Co-authored-by: Shining <xuenning@qiyi.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-21 08:52:23 +01:00
CopilotandGergő Magyar c34c36036f fix(workers): resilient + zero-copy ingestion worker pool — prevent analyze hangs on TS-root-scale loads (#1693)
* Initial plan

* fix: skip worker-timeout files in sequential fallback and optimize TS capture node lookup

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0e53743e-0600-4690-bd0d-198894daef58

* refactor: clarify TS capture helpers after validation feedback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/0e53743e-0600-4690-bd0d-198894daef58

* fix(workers): exclude in-flight file on worker error/exit, not just singleton timeout

WorkerPoolDispatchError previously surfaced the stalled path only for the
singleton-timeout final-fail branch. Worker `error` and `exit` events (and
the msg-channel `error` reply) fell back to plain `Error`, so the sequential
fallback re-attempted every file in the active job — re-hanging on the same
pathological file when the worker crashed mid-parse.

Lift the in-flight-file inference into `inFlightExcludePath(job, lastProgress)`
and wire it into the three remaining in-pool failure sites. `lastProgress` is
already in `runWorker` scope, so `items[lastProgress]` (the next file the
worker was about to acknowledge) is the best single guess at the culprit;
earlier files are still re-tried sequentially. Returns `[]` when no path is
determinable (`lastProgress >= items.length`, or path missing/non-string) so
sequential retries the whole job.

Replacement-worker startup failures stay plain `Error` (no job context); the
result-before-flush protocol bug stays plain `Error` (code fault, not file).

Tests cover the three new exclusion paths plus a negative test confirming
non-WorkerPoolDispatchError throws fall through to full sequential retry.

* fix(review): apply autofix feedback

- Use cause-neutral "worker-excluded" label in skip messages and tests now
  that worker error/exit paths share the same exclusion contract as
  singleton-timeout (correctness + maintainability reviewers).
- Add JSDoc to findSelfOrAncestorOfType{s} explaining the parent-walk
  short-circuit vs root-DFS fallback (maintainability reviewer).

* feat(workers): resilient + scalable worker pool

Restructures `createWorkerPool` so a single bad file no longer kills the
pool for the rest of an analyze run. Five interlocking layers:

1. **Auto-respawn on error/exit** — worker death triggers `replaceWorker`
   on the same slot, bounded by `maxRespawnsPerSlot` (default 3). The slot
   is dropped from rotation when the budget is exhausted; other slots
   keep running.

2. **Circuit breaker** — replaces the permanent `poolBroken=true` with a
   consecutive-failure counter. The pool only trips after
   `consecutiveFailureThreshold` deaths (default `max(3, poolSize)`) with
   no successful job in between. A successful job resets the counter so
   transient bursts of bad files don't escalate.

3. **Session-scoped file quarantine** — paths identified as the in-flight
   file at the moment of a worker death are added to a `Set<string>` on
   the pool. `dispatch()` filters quarantined items up front (they never
   reach a worker again this pool lifetime). Exposed via the new
   `WorkerPool.getQuarantinedPaths()` so callers can log/route them.
   `processParsing` surfaces the per-chunk quarantine summary alongside
   the existing fallback-exclusion log.

4. **Authoritative in-flight tracking** — `parse-worker.ts` emits
   `{type:'starting-file', path}` before each file. The pool tracks this
   per slot and uses it for crash attribution, falling back to the
   `items[lastProgress]` heuristic only when no starting-file has been
   observed (very-early crash, older worker build). Closes the
   reorder/race concerns raised by reviewers C1 and R3 in the earlier
   review run.

5. **Per-job cumulative timeout budget** — each `WorkerJob` tracks the
   total wall time spent across attempts/splits/retries. When the budget
   is exhausted (default 5x `subBatchIdleTimeoutMs`), the pool surfaces
   the in-flight path instead of letting exponential backoff balloon
   into multi-hour stalls.

Cross-layer wiring: a new `wakeIdleSlots` helper kicks any non-busy live
slot when items are requeued (after a death or split-retry), so a dropped
slot doesn't strand work in the queue. `recoverAndResume` consolidates
the per-job teardown shared by the three in-pool death sites (`error`,
`exit`, msg-channel `error`).

New env knobs: `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT`,
`GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS`,
`GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD`.
New `WorkerPoolOptions.workerFactory` injection point for unit tests.

Tests: 12 new unit tests using a FakeWorker mock cover quarantine
seeding, slot-respawn, slot-drop after budget, breaker trip + reset,
and quarantine filtering. Plus option-resolution tests for the three
new env vars. All 19 worker-pool/-fallback/-options tests pass; full
unit suite 6040 passed / 30 skipped / 0 failed.

* fix(workers): apply code-review fixes (12 findings)

Walks through every finding from ce-code-review run
20260519-094648-3549cf5e. All 12 picked Apply.

Critical:
- F1 — Layer 5 cumulative-timeout exhaustion no longer silently drops
  the rest of the job. `requeueRemainder` is now invoked before
  `handleWorkerDeath` in both Layer 5 and singleton-final-fail give-up
  paths so non-quarantined items get re-tried by another worker.
- F2 — idle-timer recovery overhaul. `!shouldContinue` branch no
  longer calls `replaceWorker` (double-spawn race with the
  `handleWorkerDeath` inside `requeueAfterTimeout`). `shouldContinue`
  branch now enforces `maxRespawnsPerSlot` before respawning, closing
  the budget-bypass for the timeout-retry path. Also fixes premature
  `maybeDone` by simplifying the bookkeeping.
- F3 — `requeueRemainder` no longer pre-charges `cumulativeTimeoutMs`
  by `job.timeoutMs`. The death itself consumed no budget, so the
  next `requeueAfterTimeout` was double-billing the first attempt.
- F4 — `WorkerPool.getQuarantinedPaths` is now optional on the
  interface, matching the defensive `?.()` call site and the existing
  mocks. Removes the contract-vs-callsite contradiction.
- F5 — per-job unattributed-death tracking. When a worker dies with
  no exclusion attribution, `requeueRemainder` tracks death count per
  `startIndex`. First time: re-queue intact. Second time: quarantine
  items[0] as best guess, or drop the job entirely when items lack
  paths. Bounds the death loop the original design admitted to.
- F6 — per-slot consecutive-failure counter. Replaces the pool-wide
  scalar so a chronically-failing slot trips the breaker on its own
  streak instead of being masked by another slot's successes.

Smaller:
- F7 — exhaustiveness `never` check on `WorkerOutgoingMessage` union.
- F8 — recursive `runWorker` on fully-quarantined jobs converted to
  a while-loop.
- F9 — `tripBreaker` calls `reject(err)` BEFORE awaiting
  `worker.terminate()`. A stuck terminate no longer blocks the caller.
- F10 — `parsing-processor.ts` quarantine log de-duplicates per pool
  instance via a `WeakMap`. Only newly-quarantined paths are logged
  in each chunk; the per-chunk count still surfaces via progress.
- F11 — extract `firstPath` local in `requeueAfterTimeout`; eliminates
  double `itemPath` call and the `unknown as string` cast.

Tests (F12, 6 new):
- crash-error event path (errorHandler).
- F5 drop-branch coverage via items without `.path`.
- Common-case unattributable crash falling back to items[0] heuristic.
- `replaceWorker` startup failure (workerFactory emits 'exit' before
  'online').
- All-slots-dropped breaker trip.
- `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` env override.

Residual gap (deferred): no unit test exercises the Layer 5
cumulative-budget runtime path — requires fake-timer interleaving
with FakeWorker that's too brittle for this iteration. Tracked.

Unit suite: 257 files / 6056 passed / 30 skipped / 0 failed.

* test(workers): integration tests for resilience layers + fix requeue-after-timeout flow

Adds 6 new real-worker integration tests covering the PR #1693
resilience layers + fixes 3 follow-on bugs surfaced while writing them.

New integration coverage (real worker threads + temp fixture scripts):

- `respawns the slot after worker process.exit and finishes the work on
  the replacement` — exercises Layer 1 auto-respawn + Layer 3 quarantine
  through real IPC.
- `attributes exactly via authoritative starting-file message on worker
  crash` — Layer 4 end-to-end: starting-file message → exact quarantine
  attribution (not the items[0] heuristic).
- `quarantine filters subsequent dispatches without sending to a worker`
  — second dispatch's sub-batch payload audited via filesystem; the
  quarantined path is never sent across the message channel.
- `drops a slot after maxRespawnsPerSlot and continues on the survivor`
  — 2-slot pool, slot dies twice past budget, survivor finishes
  re-queued remainder.
- `trips the circuit breaker on cascading per-slot consecutive failures`
  — single-slot pool, dies on every job, breaker trips after
  consecutiveFailureThreshold with WorkerPoolDispatchError carrying
  the cumulative quarantine.
- `survives a worker error event (uncaught throw) the same as a
  process.exit` — validates recoverAndResume on the errorHandler path
  via a real worker `throw` (not just process.exit).

Bug fixes uncovered while writing these tests:

1. **Stack-overflow recursion in runWorker's no-worker branch** —
   `if (!worker) { ...; wakeIdleSlots(); maybeDone(); }` recursed
   indefinitely when multiple slots were mid-respawn simultaneously
   (wakeIdleSlots → runWorker → no worker → wakeIdleSlots → …).
   Removed the wakeIdleSlots call: the slot's own respawn IIFE owns
   runWorker post-respawn, and other slots will pick up work via
   finishJob's runWorker.

2. **requeueAfterTimeout dispatched work before respawn completed** —
   the F2 fix had `requeueAfterTimeout` `void`-discarding
   `handleWorkerDeath`, so the `!shouldContinue` IIFE had no way to
   know when the respawn finished. New design: `requeueAfterTimeout`
   returns a `TimeoutDecision` discriminated union; the IIFE owns
   the death-and-respawn-and-dispatch orchestration in an async
   closure so it can `await handleWorkerDeath` and then call
   `runWorker` deterministically.

3. **Stalled-singleton + protocol-error + replacement-startup-crash
   tests** had stale contracts predating the resilience refactor. The
   stalled-singleton no longer rejects (it quarantines + resolves
   `[]`); the protocol-error rejection message now mentions
   "circuit breaker tripped"; the replacement-startup-crash test
   documents the known `waitForWorkerOnline` race (online fires
   before the worker's main script runs, so a top-level throw looks
   like a successful spawn) — the test asserts the file is
   quarantined via the second-idle-timeout give-up path.

Full suite: 334 files / 8982 passed / 43 skipped / 0 failed (second
run; first run had a Vitest-reported flake from an uncaught worker
exception bleeding into the test report — repeated runs are clean).

* perf(workers): raise pool cap to cores-1 + defer per-chunk extraction to keep workers busy

User reported 4-5% CPU utilization on a multi-core machine during
ingestion. Two structural reasons:

1. **Pool cap.** `createWorkerPool` resolved size as
   `Math.min(8, max(1, os.cpus().length - 1))` — a 16-core box got 8
   workers (50% theoretical max). U1 lifts the default to
   `min(16, max(1, cores - 1))`, exposes `GITNEXUS_WORKER_POOL_SIZE`
   env override, and adds `--workers <N>` CLI flag (`0` disables the
   pool for sequential fallback).

2. **Per-chunk extraction serialized the loop.** Per chunk:
   dispatch → await workers → main-thread `processImportsFromExtracted`
   + `processHeritageFromExtracted` + `processRoutesFromExtracted`
   + `synthesizeWildcardImportBindings` + `seedCrossFileReceiverTypes`
   → next chunk dispatch. Workers sat idle through every extraction
   block. U2 (revised from the plan's pipelined-chunks design) defers
   these passes to a single end-of-loop batch. Chunk loop becomes
   parse + merge + accumulate. Resolution sees strictly-more-info
   (full repo graph) so cross-chunk import/heritage targets resolve at
   least as well as before. Memory cost: `deferredWorkerImports`
   accumulates across chunks; bounded by total file count, acceptable.

Plan deviation note: the plan called for an in-flight chunk pipeline
(N concurrent dispatches with bounded memory). That design needed
either a `processParsing` API refactor or duplicating its catch-block
fallback in `parse-impl`. The deferred-extraction approach delivers
the same "workers stay busy" outcome with much smaller surface area
and zero changes to `processParsing`. The `GITNEXUS_PARSE_CHUNK_CONCURRENCY`
env var documented in U2 of the plan is therefore not implemented in
this commit; if memory growth from `deferredWorkerImports` becomes
a problem at very-large-repo scale, a bounded sliding-window variant
can land as a follow-up.

Tests:
- New `test/unit/analyze-worker-pool-size.test.ts` covers --workers
  validation (5 invalid inputs rejected with exit code 1 + clear
  error; valid integers set the env var; `--workers 0` routes to
  sequential).
- Extended `worker-pool-resilience.test.ts` with `resolveAutoPoolSize`
  scenarios: env override, env=0, env above cap, invalid env fallback,
  auto-formula match, integer return type.
- Full unit suite: 6097 / 6127 passed / 30 skipped / 0 failed.
- Full integration suite (second run): 77 / 78 passed / 1 skipped /
  0 failed. First run had a known cosmetic flake from an uncaught
  worker exception bleeding into the test reporter.

Resilience contract from PR #1693 preserved: per-slot respawn budget,
circuit breaker, quarantine, authoritative in-flight tracking,
cumulative timeout budget — all unchanged.

New env vars surfaced in --help: GITNEXUS_WORKER_POOL_SIZE,
GITNEXUS_PARSE_CHUNK_CONCURRENCY (reserved for future bounded
pipelining).

* docs(readme): document --workers CLI flag

* feat(workers): add getStats() and per-chunk throughput logging

* test(workers): cleanup leaked temp-dirs and drop duplicate option-resolution block

- Add afterEach to worker-pool-resilience.test.ts cleaning up the per-test temp
  directory created by beforeEach (~25 stale dirs per CI run previously).
- Delete the duplicated describe('worker pool option resolution', ...) block.
  Verified the first block (lines 490-532) is a strict superset (includes the
  GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS env test the second block omitted),
  so deletion loses no test coverage.

Addresses PR #1693 review findings L2 (temp-dir leak) and L3 (duplicate block).

* feat(cli): thread --workers via PipelineOptions + snapshot/restore CLI env

Resolves PR #1693 review B2 (env-var leak in long-running hosts):

- --workers is now threaded through AnalyzeOptions -> runFullAnalysis
  -> PipelineOptions.workerPoolSize -> createWorkerPool's explicit
  poolSize arg, bypassing the GITNEXUS_WORKER_POOL_SIZE env channel.
  The env var remains as a back-compat fallback inside resolveAutoPoolSize
  for operators who set it directly.
- analyzeCommand and wikiCommand snapshot the GITNEXUS_* env vars they
  mutate at function entry and restore them in finally. Inner *Impl
  extraction keeps the diff surgical (no body re-indent). process.exit(0)
  on the CLI success path still terminates the process; restoration
  matters for programmatic callers (tests, long-running hosts) reaching
  early-return paths or the alreadyUpToDate fast path.
- Tests updated to assert the new behavior:
    analyze-worker-pool-size.test.ts: workerPoolSize flows through
      runFullAnalysis options; env is not mutated; back-to-back calls
      see their own values, not the previous call's leak.
    analyze-worker-timeout.test.ts: env IS set during the runFullAnalysis
      call (captured via mockImplementation) and restored after, proving
      the timeout reaches downstream while the leak fix holds.
- Also addresses L4: afterEach NODE_OPTIONS restore so back-to-back test
  runs don't accumulate --max-old-space-size=8192 tokens.

Addresses PR #1693 review B2 (blocker) and L4 (test polish).

* feat(workers): harden worker lifecycle (messageerror + availableParallelism + ready handshake)

Resolves PR #1693 review H1, H2, M4:

H1 - messageerror handler at every dispatch site
  V8 deserialization failure on postMessage previously left the message
  silently lost; the pool would wait out the idle timeout (default 30s)
  instead of treating it as worker death. The dispatch loop now wires
  worker.once('messageerror', ...) alongside error/exit and routes through
  recoverAndResume so the existing per-slot respawn budget, in-flight
  file attribution, and circuit-breaker layers fire as designed.

H2 - resolveAutoPoolSize uses os.availableParallelism()
  Mirrors the pattern at capabilities.ts:85 (defaultEmbeddingThreads).
  os.cpus().length returns the host CPU count, which over-sizes the pool
  on cgroup-limited containers, taskset-restricted runtimes, and CI
  runners with explicit CPU quotas. Falls back to os.cpus().length on
  Node < 18.14.

M4 - worker-side ready handshake replaces online-trust
  parse-worker.ts now emits {type: 'ready'} after all top-of-script
  initialization completes, BEFORE the message handler is attached. The
  pool's renamed waitForWorkerReady listens for this message under a
  bounded WORKER_READY_TIMEOUT_MS (5s) budget instead of trusting Node's
  online event - which fires when the worker thread starts, BEFORE the
  script body runs, letting init crashes slip past pool startup. ready
  is added to WorkerOutgoingMessage with an exhaustiveness-checked
  no-op branch in the dispatch handler (defensive: the message is
  consumed by waitForWorkerReady before dispatch handlers attach).
  messageerror is wired into waitForWorkerReady the same way.

Test scaffolding:
  - FakeWorker emits {type: 'ready'} in addition to 'online' so
    replacement workers in unit tests don't hit the 5s budget.
  - Integration test ad-hoc worker scripts go through a writeReadyWorker
    helper that prepends the ready handshake. Tests intending to script
    "crash BEFORE ready" can bypass the helper.

61/61 worker-pool unit tests pass; 28/28 integration tests pass.

* feat(parse-impl): monotonic progress + verbose-gated throughput log + seed-before-build

Resolves PR #1693 review M2, M3, L1, L5 in a single parse-impl.ts pass:

M2 - Monotonic progress through deferred phase (no more "stuck at 82%")
  Previously the deferred resolution stages (imports, heritage, routes,
  calls) all emitted percent: 82 — the UI looked frozen for the duration
  of the deferred work, which on large repos is several seconds to minutes
  and visually identical to the hang PR #1693 set out to fix.
  Redistributed:
    parse phase:  20-70 (was 20-82)
    imports:      70-75
    heritage:     75-80
    routes:       80-85
    calls:        85-95
  Each deferred stage now advances through its own band via the existing
  per-batch progress callback. Skipped stages (zero deferred input) leave
  their band as a no-op jump - the next stage still starts at its own
  band, preserving strict monotonicity. The "no parseable files" early
  return now jumps to 95 (was 82), and the duplicate "Parsing N files..."
  announcement is suppressed when totalParseable === 0 to avoid a
  non-monotonic 95 -> 20 regression that pre-existed (uncovered by the
  new monotonic test).

M3 - Throughput log gated on `--verbose`, not just NODE_ENV=development
  The per-chunk files/s log was gated on `isDev`, so operators running
  `gitnexus analyze --verbose` in a production install never saw it.
  Now fires when (isDev || isVerboseIngestionEnabled()) — matches the
  documented promise that `--verbose` shows tuning observability.

L1 - Typo rename: `chunkChunkStartMs` -> `chunkStartMs`

L5 - `buildExportedTypeMapFromGraph` runs BEFORE `seedCrossFileReceiverTypes`
  Previously the seeding branch was reached with `exportedTypeMap.size === 0`
  in the worker path (the map was only built far below, AFTER the seeding
  branch), so the seed dead-coded itself silently and call resolution
  never got the cross-file receiver-type enrichment. Now the map is
  populated from the in-progress graph before the seed call; the
  post-parse builder remains as a defensive sequential-path fallback,
  guarded by `size === 0` so we don't pay the cost twice on the worker
  path. Net win: cross-file CALLS edges that previously had no receiver
  type now get enriched.

New test: parse-impl-progress-monotonic.test.ts
  Asserts the emitted percent stream is strictly non-decreasing across
  the parse + deferred phases, and that the deferred band (>=70) is
  actually reached. Also pins the "no parseable files" path to exactly
  [95] so the 95 -> 20 regression we just fixed can't re-emerge.

* feat(parse-impl): bounded chunk concurrency via file-pre-fetch pipeline

Resolves PR #1693 review B1 (GITNEXUS_PARSE_CHUNK_CONCURRENCY documented
in --help but unimplemented).

The chunk loop now pre-fetches chunk file contents up to
`parseChunkConcurrency` chunks ahead of the worker-dispatch cursor so
disk I/O overlaps with worker compute. Worker dispatch itself stays
serial because WorkerPool.dispatch is not reentrant — concurrent calls
would race on the shared per-slot busy/in-flight state, regressing the
hang/resilience work this PR is built on. The pre-fetch path is the
honest interpretation of "concurrent in-flight parse chunks" that the
help text advertises: I/O overlap, not parallel worker dispatch.

Concurrency value resolution:
  1. PipelineOptions.parseChunkConcurrency (threaded from CLI)
  2. GITNEXUS_PARSE_CHUNK_CONCURRENCY env var
  3. Default 2 (matches the help text)

F4 (wildcard-synthesis ordering) is preserved: deferred-state
aggregation runs in chunkIdx order because the for-loop iterates
sequentially after awaiting each chunk's pre-fetched contents.
Cross-chunk processors (processImportsFromExtracted,
synthesizeWildcardImportBindings, etc.) still run only after all
chunks complete — they see deterministic input regardless of
file-read completion order.

Concurrency=1 produces behavior identical to the pure-serial loop;
that's the regression baseline.

New test: parse-impl-chunk-concurrency.test.ts
  - Asserts graph output is identical (nodeCount + relationshipCount)
    between parseChunkConcurrency=1 and =2 — the critical correctness
    invariant. Exact .toBe(N) comparisons per DoD §2.7 (the second run's
    counts must equal the first run's exactly).
  - Pins specific fixture symbols (foo/bar/Baz) under both
    parseChunkConcurrency=1 and the env-fallback (3) path.
  - Env-fallback test confirms GITNEXUS_PARSE_CHUNK_CONCURRENCY is
    honored when the option is undefined.

* test(workers): pin cumulative-timeout exhaustion behavior

Resolves PR #1693 review M6: the existing resilience suite asserts only
the *default value* of maxCumulativeTimeoutMs (5x subBatchIdleTimeoutMs),
not that dispatch actually aborts the offending job when the cumulative
wall-clock budget is exhausted. Without this test, a future refactor
could remove the exhaustion branch in requeueAfterTimeout and the suite
would stay green while the pool sat in retry loops for an hour on a
real production stall.

Scenario:
  subBatchIdleTimeoutMs    = 100ms
  timeoutBackoffFactor     = 10
  maxCumulativeTimeoutMs   = 300ms

Single file, HangingWorker that never responds. First attempt times
out at 100ms (cumulative=100). The next backoff (1000ms, cumulative
1100ms) exceeds the 300ms cap, so requeueAfterTimeout returns
give-up on the first timeout retry and the file goes to the session
quarantine. Asserts:
  - pool.getQuarantinedPaths() includes 'src/stuck.ts' after dispatch
  - if dispatch rejected, the error is a WorkerPoolDispatchError
    (the typed surface that routes to sequential fallback)

Uses a local minimal HangingWorker double rather than the full
action-scripted FakeWorker from worker-pool-resilience.test.ts —
the inverse pattern (always hang) doesn't need the scripted-action
machinery and keeps the test file focused on the one behavior.

* docs(readme): add environment-variables reference table

Resolves PR #1693 review L6: operator-facing env vars were either
mentioned inline (GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS) or only
documented via `gitnexus --help`, with no single place to look up
the full set. The new "Environment variables" subsection under the
Quick Start CLI block lists every operator-facing knob with default,
effect, and tuning guidance, matching the names in cli/index.ts
addHelpText post-U2 / U1.

Covers:
  GITNEXUS_WORKER_POOL_SIZE           (--workers)
  GITNEXUS_PARSE_CHUNK_CONCURRENCY    (newly real per U1)
  GITNEXUS_VERBOSE                    (--verbose)
  GITNEXUS_MAX_FILE_SIZE              (--max-file-size)
  GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS (--worker-timeout × 1000)
  GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES
  GITNEXUS_CHUNK_BYTE_BUDGET
  GITNEXUS_NO_GITIGNORE
  GITNEXUS_SKIP_OPTIONAL_GRAMMARS

CLI flag vs env-var precedence is stated explicitly (CLI > env > default)
so operators running long-lived hosts (MCP server, eval-server) know
which channel wins.

* test(workers): pin quarantine path round-trip and non-normalization contract

Resolves PR #1693 review M5 (Windows quarantine path-normalization
coverage). worker-pool.ts quarantines paths via a Set<string> keyed by
exact string equality. The existing suite never asserted this contract,
which lets a future "helpfully normalizing" refactor on one side of the
pipeline (caller, worker, or pool) silently break quarantine filtering
on Windows.

This file pins the contract from both directions:

1. Round-trip: a path the caller dispatches with backslashes
   (src\bad.ts) flows through starting-file -> death -> quarantine ->
   next-dispatch filter verbatim. The replacement worker never sees the
   re-dispatched bad path because the pool's pre-dispatch filter
   short-circuits it.

2. Non-normalization: quarantining src\poison.ts does NOT filter
   src/poison.ts. Whoever changes that contract has to update this test
   alongside (the load-bearing assertion catches accidental
   path.normalize() calls in the quarantine path).

Runs on every platform — the path strings are test-injected, so the
test exercises the same code path regardless of the host's path.sep.
Used a self-contained FakeWorker that emits {type:'ready'} for U3's
waitForWorkerReady handshake, so the test doesn't depend on the larger
worker-pool-resilience.test.ts harness.

* test(typescript): pin capture-anchor rewrite invariants (B5 regression)

Resolves PR #1693 review B5: the captures.ts ancestor-walk rewrite
(findSelfOrAncestorOfType[s] + pickFirstNode replacing the prior
findNodeAtRange-from-root path) was semantically equivalent to its
predecessor per Lane 4 of the production-readiness review, but the
existing typescript-captures.test.ts didn't pin the specific sharp
edges where an over-aggressive walk would silently break captures.
This file does.

Each test exercises a capture class whose anchor type is one the
rewrite explicitly handles:

  - member call obj.foo() -> @reference.call.member (call_expression
    anchor walks to self)
  - dynamic import import("./helper") -> raw @import.dynamic gets
    decomposed by splitImportStatement into @import.statement with
    @import.kind=dynamic + @import.source stripped of quotes
  - JSX <Foo /> in .tsx -> @reference.call.free emitted (TSX query
    pattern, query.ts:899-905) but @declaration.parameter-count is
    NOT synthesized because findSelfOrAncestorOfType('call_expression')
    returns null on a jsx_self_closing_element anchor. Pre-rewrite the
    range lookup also returned null. Pinning this contract catches
    accidental "walk JSX -> outer call" refactors.
  - constructor `new Foo(1,2)` -> @reference.call.constructor (new_expression
    anchor walks to self)
  - named/namespace import + re-export -> @import.statement (one each)
  - class method override -> @declaration.method per class, no collapse
  - member read obj.foo (no call) -> @reference.read.member

All assertions use exact .toBe(N) per DoD §2.7.

* test(parse-impl): pin multi-chunk graph equivalence under deferred extraction

Resolves PR #1693 review B4: the deferred-extraction reorder (moving
processImportsFromExtracted / Heritage / Routes / Wildcard /
ReceiverTypes from per-chunk to end-of-loop) was proven observably
equivalent by Lane 4 of the production-readiness review. Until now,
the existing suite never asserted cross-chunk graph equivalence,
which lets a future refactor that accidentally tightens the per-chunk
vs end-of-loop coupling silently break cross-chunk resolution.

This test forces multi-chunk parsing on a small fixture by setting
GITNEXUS_CHUNK_BYTE_BUDGET=64 BEFORE the parse-impl module loads
(the budget is captured at module load via vi.resetModules — a future
move to function-scope env reads is U14 in Phase 2). Then runs the
same fixture under a 10MB budget (single chunk) and asserts the two
graphs are byte-identical: same nodeCount, same relationshipCount,
exact .toBe(N) per DoD §2.7.

Fixture: 3-file class hierarchy with cross-file inheritance — Animal
(a.ts) -> Dog extends Animal (b.ts) -> makeDog returns Dog (c.ts).
Forces the resolver to chain imports + heritage across chunks. A
second test pins specific symbol names (Animal, Dog, makeDog, speak,
bark) in the multi-chunk graph so a regression in chunk-boundary
resolution surfaces as a missing-symbol failure with a specific
diagnostic instead of a bare count mismatch.

* test(parse-impl): wall-clock integration pinning multi-chunk pipeline (B3)

Resolves PR #1693 review B3 — the final P0/P1 merge blocker. With this
test, all five doc-review blockers (B1-B5) are pinned by regression
coverage.

The PR's headline claim is "analyze no longer hangs on TS-root-shaped
loads". The existing suite pins each resilience layer (worker-pool-
resilience.test.ts), the deferred-extraction equivalence (U7), and
the chunk-concurrency contract (U1). What was missing: a single
end-to-end run that exercises the full chunked parse-and-resolve
path on a multi-chunk fixture, BOUNDED by a wall-clock budget so a
regression that re-introduces the hang fails this test loudly via
timeout rather than slipping past as a count drift.

Implementation:
  - 17-file synthetic fixture: 15 small modules (one function each),
    one "realistic dense" complex.ts (30 functions + class + interface),
    and an index.ts re-exporting them. Forces cross-chunk import
    chains.
  - GITNEXUS_CHUNK_BYTE_BUDGET=64 via vi.resetModules forces multi-chunk
    parsing on the small fixture.
  - Promise.race with 30s timeout: a hang fails as
    "exceeded WALL_CLOCK_BUDGET_MS — likely the hang B3 was meant to
    prevent", not as a bounds-only inequality (DoD §2.7 distinction —
    hang-detector via exception, not regression-mask via inequality).
  - Exact .toBe(true) assertions on specific expected symbols
    (fn0..fn14, Service, Config, configure, describe, complex0/15/29)
    so a silent mid-chunk crash that exits 0 without producing graph
    data also fails this test, not just the hang case.

Scope: runs the sequential-fallback path (skipWorkers: true) because
the full real-worker scenario requires a built dist/parse-worker.js
and ~60s wall-clock per run — appropriate for a CI-integration job,
not vitest. The load-bearing invariants pinned here catch the bulk
of B3's concern; the dist-worker swap is a Phase 2 follow-up
documented in the file header.

* refactor(parse-impl): move chunk-byte-budget env read to function scope

Resolves PR #1693 review F7 / U14: pre-U14, `CHUNK_BYTE_BUDGET` was a
module-load IIFE constant that captured `GITNEXUS_CHUNK_BYTE_BUDGET`
once and froze the value for the module's lifetime. That defeated
per-call option threading (a future
`PipelineOptions.chunkByteBudget` was silently no-op'd because the
function body read the frozen module-level constant) AND forced tests
to use `vi.resetModules` to vary chunk layout. The U7
deferred-extraction test and the U6 multi-chunk integration test
both used the workaround.

After this change:

  - `DEFAULT_CHUNK_BYTE_BUDGET = 2 * 1024 * 1024` stays as a
    module-level constant — purely a default, no env access.
  - `resolveChunkByteBudget(options)` runs per call: option wins,
    then env, then default. Same options-first/env-fallback/default
    pattern as resolveAutoPoolSize and the U1 parseChunkConcurrency
    resolver — keeps the ingestion code's configuration model uniform.
  - `PipelineOptions.chunkByteBudget?` added with documentation that
    threading through options lets long-running hosts (eval-server,
    MCP daemon) size per-call without leaking process.env state
    across analyze invocations.

New test (parse-impl-env-reads.test.ts) pins all four behaviors:
  1. option-first: option present + env present -> option wins
  2. env-fallback: option absent + env present -> env wins
  3. default-fallback: both absent -> 2 MB default
  4. per-call: two back-to-back runs in the same vitest worker with
     different chunkByteBudget option values observe their OWN values,
     proving the module-load freeze is gone (no vi.resetModules in
     this test — that's the invariant being verified).

All four assertions use exact `.toBe(N)` per DoD §2.7. The chunk
count is observed by parsing the `Parsing chunk X/Y` progress message
stream — a stable proxy that doesn't require exposing internal
parse-impl counter state.

Note: U7 and U6 tests still use `vi.resetModules` because they were
written before this change. A follow-up cleanup could simplify those
tests (drop the resetModules dance, pass chunkByteBudget via options),
but they pass as-is so this commit doesn't touch them.

* feat(workers): per-slot generation counter for late-event protection (U12)

Adds a monotonic per-slot generation counter to createWorkerPool's
state. Each successful worker replacement (replaceWorker) bumps the
slot's counter exactly once — atomically with the workers[slotIndex]
swap, so observers (getStats) see the new (worker, generation) pair
consistently. Handler closures in the dispatch loop capture the
slot's generation at attach time and short-circuit when they fire
on a stale generation.

In the current implementation, cleanup() synchronously removes
listeners on a Worker instance the moment a death is observed, so
no listener naturally fires on a stale generation — the guard is a
defensive layer protecting against any future refactor that loosens
cleanup() ordering or re-attaches handlers across the swap. The
load-bearing observable is the slotGenerations[] array exposed via
WorkerPoolStats so operators (and tests) can confirm a slot was
actually replaced and not just the same worker recycled.

Implementation:
  - const slotGenerations: number[] = new Array(size).fill(0) in
    createWorkerPool's per-pool state, alongside respawnCount and
    consecutiveFailuresPerSlot.
  - replaceWorker: slotGenerations[workerIndex]++ AFTER the
    workers[workerIndex] = replacement swap (only on the success
    branch — drop-slot paths leave the counter unchanged).
  - runWorker dispatch loop: const slotGen = slotGenerations[workerIndex]
    captured before handler attachment; every handler (handler /
    errorHandler / exitHandler / messageErrorHandler) starts with
    `if (slotGenerations[workerIndex] !== slotGen) return`.
  - WorkerPoolStats gains `readonly slotGenerations: readonly number[]`.
  - getStats() returns slotGenerations.slice() so callers can't mutate
    pool state by writing to the returned array.

Two existing toEqual snapshots in worker-pool-resilience.test.ts
extended with the new slotGenerations field (both expect all-zeros —
neither test scenario triggers a respawn).

New test file (worker-pool-slot-generation.test.ts, 4 tests):
  1. Fresh pool: every slot at generation 0.
  2. Successful crash + respawn: generation bumps to 1 exactly once.
  3. Crash that drops the slot (maxRespawnsPerSlot:0): generation
     stays at 0 because no successful respawn happened. The dispatch
     rejection on breaker trip is the expected outcome here; the
     load-bearing assertion is the post-rejection stats.
  4. Multi-slot independence: one slot crashing bumps only that
     slot's generation, not the other. Order-independent via sort()
     because the round-robin assignment isn't pinned by contract.

All assertions exact .toEqual / .toBe per DoD §2.7.

* docs(bench): add parse-throughput benchmark scaffold (R13)

Resolves PR #1693 review R13 (benchmark artifact requirement).

Creates `gitnexus/bench/parse-throughput.md` documenting:

- Synthetic fixture spec (same shape as the U6 integration test, so
  CI smoke baseline and ad-hoc benchmark exercise the same paths).
- What to measure (wall-clock, peak heap, chunk count, getStats
  snapshot) and the hardware-shape metadata to record alongside.
- Harness recipe — vitest + env-var overrides to exercise sequential
  fallback vs worker-pool paths.
- Latest-measurement table with placeholder rows for the three paths
  (sequential, workers+concurrency, workers single-threaded) and an
  explicit "Status: scaffold — fill in before merging" callout. The
  U6 test's observed ~6 s wall-clock is captured as a smoke-baseline.
- Operator-tuning quick reference cross-linked to the README env-var
  section (U11) so the doc is actionable without re-reading the PR.
- "What this benchmark does NOT measure" section explicitly scoping
  the artifact's limits (synthetic ≠ real-repo, throughput-only ≠
  resilience-tested, Phase 3 IPC repack row reserved for U16-U17).

Mitigates the doc-review SG5 "static doc drift" concern via:
  1. Explicit "regenerate this file before merging" callout at the top.
  2. Self-contained methodology so anyone can re-run the numbers.
  3. Cross-links to the U6 integration test that already bounds the
     wall-clock as part of the CI suite — so "is it still completing?"
     is regression-tested even if the numbers in this doc drift.

The standalone harness script (`bench/scripts/parse-throughput.ts`)
remains a stretch goal per the original plan. The U6 vitest with
verbose ingestion logs covers the primary observability gap until
the standalone harness lands.

* perf(parse-impl): free deferred-extraction arrays after consumption (U15 lightweight M1)

PR #1693 review M1 noted that the deferred-extraction accumulator
arrays (`deferredWorkerImports`, `deferredWorkerCalls`,
`deferredWorkerHeritage`, `deferredConstructorBindings`,
`deferredAssignments`) were retained until function return, making
peak accumulator memory O(repo) instead of O(in-flight stage).

This commit implements the LIGHTWEIGHT version: free each array
immediately after its last consumer drains/reads it, dropping peak
accumulator memory progressively through the deferred-extraction
stages. The structural per-chunk streaming variant (the original
U15 framing) is deliberately deferred — the doc-review's adversarial
reviewer (A4) flagged it as defending unmeasured memory pressure,
and the simpler array-clearing captures the bulk of the benefit
without committing to a scheduling-strategy decision (microtask vs
parallel extractor task vs worker-side) that profile data should
inform.

Clears added:

  1. After `processImportsFromExtracted` (the sole consumer of
     `deferredWorkerImports`): clear the imports array before
     the heavier heritage/calls stages run.
  2. After `buildHeritageMap` (the LAST consumer of the raw
     `deferredWorkerHeritage` records — processCallsFromExtracted
     reads from the derived `fullWorkerHeritageMap` instead):
     clear the heritage array before the call-resolution stage.
  3. After `processAssignmentsFromExtracted` (the joint last
     consumer with processCallsFromExtracted for the calls/
     bindings/assignments triple): clear all three before
     downstream graph-build / scope-resolution uses its own
     working memory.

Arrays returned in the function result object (allFetchCalls,
allExtractedRoutes, allDecoratorRoutes, allToolDefs, allORMQueries,
allParsedFiles) intentionally stay live — downstream consumers
need them.

Graph-output equivalence is preserved (U7 multi-chunk equivalence
test passes — the clears happen AFTER each array's last consumer
has copied data into the graph or derived structures).

* feat(workers): introduce protocol.ts wire-format module (U16, IPC scaffold)

Defines the binary frame for worker-thread IPC as an isolated, fully-tested
module. Production wiring is deferred to U17 — shipping the wire-format
contract first de-risks the migration by establishing a single source of
truth for the byte layout. Resolves the scaffold half of PR #1693 review
R12.

Wire layout (per message, single buffer):

  +---------+-----------+---------------------+
  | tag     | length    | payload bytes …     |
  | 1 byte  | 4 bytes   |                     |
  +---------+-----------+---------------------+

  tag    : MessageTag enum value (0x01 DispatchJob ... 0x08 Ready)
  length : little-endian uint32 byte count for the payload region
  payload: UTF-8 JSON-encoded value, possibly "null"

Why JSON for the body (rather than per-shape binary encoders): the
doc-review adversarial reviewer (A2) flagged that a true per-shape
binary encoder for the result message — which carries nested
heterogeneous extracted-call / import / heritage / route arrays —
would be 500-1500 LOC and a substantial maintenance burden. The
honest perf win the IPC repack targets is moving file CONTENTS via
ArrayBuffer transferList (zero-copy ownership transfer for the
largest single piece of state in any message). That win is captured
by U17 layering transferList over the bulk file-content payload while
keeping this module's framing for the surrounding metadata. If U18
benchmark data shows the JSON body is itself a bottleneck after U17
lands, a follow-up unit can swap to per-shape binary encoding behind
the same encodeMessage / decodeMessage surface without changing the
frame.

API:
  - MessageTag (const object): stable byte tags 0x01..0x08
  - PROTOCOL_HEADER_BYTES = 5
  - ProtocolDecodeError extends Error: distinct class so U17's
    pool-side handler can route protocol violations through the
    existing messageerror recovery layer (U3 H1) distinctly from
    other failure classes
  - encodeMessage(tag, payload): Buffer
  - decodeMessage(buf): { tag, payload }
  - Uses Buffer#subarray instead of the deprecated Buffer#slice

Tests (18, all exact-equality per DoD §2.7):
  - byte layout (tag at offset 0, length LE uint32 at offset 1)
  - empty/null payload encodes to 5-byte header + 4-byte "null" body
  - round-trip for every MessageTag with representative payloads
  - non-ASCII path string (UTF-8 byte-length boundary)
  - 9 MB payload (well past the existing 8 MB sub-batch budget)
  - decode errors surface as ProtocolDecodeError, not generic Error:
      * buffer < header size
      * tag outside valid range
      * declared length exceeds buffer
      * payload bytes are not valid JSON
  - error class name is preserved through prototype chain so callers
    can `err instanceof ProtocolDecodeError` reliably

* refactor(workers): extract quarantine into its own module (U13 partial)

Honest partial U13: extract the quarantine resilience layer (Layer 3
of the 5-layer model) into a dedicated module with a small explicit
interface. The full 5-module split that the original plan named was
flagged by doc-review A10 as abstraction-without-multi-consumer-demand
("Each has exactly one consumer: worker-pool.ts. None of these layers
is imported elsewhere in the codebase pre-extraction, and the plan
doesn't identify any future consumer.") This commit ships the smallest
self-contained layer as a named module to validate the factory +
interface pattern with minimal risk. The remaining four layers
(respawn-budget, cumulative-timeout, circuit-breaker, slot-attribution)
stay inline until a real second consumer emerges (e.g., a non-parse
worker pool that reuses the same resilience layers).

Module shape (`workers/quarantine.ts`, ~30 LOC):

  interface Quarantine {
    add(path: string): void;
    has(path: string): boolean;
    snapshot(): string[];   // defensive copy
    readonly size: number;  // getter, reflects state at access time
  }
  function createQuarantine(): Quarantine

Replaces in `worker-pool.ts`:
  - `const quarantined: Set<string> = new Set()` -> `createQuarantine()`
  - `quarantined.has(p)`            -> `quarantine.has(p)` (2 sites)
  - `quarantined.add(p)`            -> `quarantine.add(p)` (2 sites)
  - `quarantined.size`              -> `quarantine.size` (2 sites)
  - `Array.from(quarantined)`       -> `quarantine.snapshot()` (6 sites)

Public worker-pool.ts API is unchanged — `getQuarantinedPaths()` still
returns the same defensive `string[]` copy. The behavioral contract is
preserved: paths are quarantined as opaque strings (the U9 / M5
non-normalization contract still holds — see the new dedicated test).

Tests:
  - 8 isolated unit tests for the quarantine module — pins the
    interface contract (empty start, add/has/size, dedup on repeated
    add, no separator normalization, snapshot defensive copy + freshness,
    size-getter live behavior).
  - All 86 existing worker-pool tests pass unchanged — they exercise
    the quarantine through the pool and act as the regression net for
    behavior preservation.

Why not the full 5-module extraction in this commit: doc-review A10's
concern is real — a single-consumer abstraction adds module-boundary
overhead (5 sets of imports, 5 dedicated test files, 5 interfaces to
keep in sync with worker-pool) without any structural benefit until a
second consumer materializes. Extracting one validates the pattern;
the remaining four can be moved on demand.

* feat(workers): wire protocol.ts encoded IPC into parse-worker + pool (U17)

Production worker IPC now uses the U16 binary wire format (1-byte tag +
4-byte LE length + UTF-8 JSON body) end-to-end. The pool encodes every
outgoing `sub-batch` / `flush` dispatch via `encodeMessage`; the worker
decodes incoming frames via `decodeMessage` and encodes its `ready`,
`starting-file`, `progress`, `sub-batch-done`, `result`, `warning`, and
`error` outputs the same way.

The load-bearing correctness fix is making `decodeMessage` accept
`Uint8Array` rather than only `Buffer`: Node's `worker_threads`
`postMessage` structured-clones the payload, which strips the `Buffer`
prototype on the receive side. A frame sent as `Buffer` arrives as a
plain `Uint8Array`, and `Buffer.isBuffer(raw)` returns false — so the
first attempt at U17 (gating decode on `Buffer.isBuffer`) silently
treated every incoming frame as POJO and the worker never responded.
The fix adopts the underlying memory zero-copy via
`Buffer.from(view.buffer, view.byteOffset, view.byteLength)` and uses
`raw instanceof Uint8Array` at every call site (parse-worker decode,
pool dispatch handler, pool ready-handshake handler, FakeWorker test
mocks, and the integration-test worker preamble).

The pool stays tolerant of POJO incoming so unit-test FakeWorkers
don't need rewriting — only the new outgoing encoded dispatches require
the test scaffolding to decode on receive, which the test FakeWorkers
and the integration test's inline `parentPort.on` wrapper now do.

The slot-drop integration test was rewritten from a shared-counter-file
race (which pre-U17 timing happened to land on the assertion-friendly
counter==2 endpoint, but post-U17 protocol decoding latency shifted to
counter==1 and produced 3 quarantines instead of 2) to a deterministic
path-based crash trigger: slot 0 crashes on a.ts, respawns, crashes on
the requeued b.ts, slot is dropped after budget exhausted; slot 1
handles [c.ts, d.ts] normally. Outcome no longer depends on inter-worker
file-write ordering.

Protocol coverage adds two regression tests pinning the Uint8Array
decode path: structured-clone-stripped frames decode identically to
their Buffer originals, and Uint8Array views with non-zero byteOffset
into a wider ArrayBuffer also decode correctly (catches `Buffer.from(uint8)`
copying semantics if a future refactor loses the zero-copy adoption).

All 94 worker-pool tests (9 files, unit + integration) pass; the full
unit suite (6128 tests across 268 files) passes unchanged.

* perf(workers): zero-copy file content transfer via transferList (U19)

Pool dispatch now hoists `{path, content: string}[]` file contents OUT
of the U17 JSON envelope into separately-allocated `Uint8Array`s whose
ArrayBuffers are passed to `worker.postMessage`'s `transferList` for
zero-copy ownership transfer. The envelope itself carries only
lightweight metadata (`{path, byteLength}` per file) and is structure-
cloned the same as before.

What this saves vs U17 baseline:

- **JSON.stringify of file contents on main thread** drops to zero —
  the envelope is now O(paths + sizes), not O(total bytes). For a 200-
  file sub-batch of 10 KB TS files, that's ~2 MB of escape processing
  per dispatch that disappears. JSON.stringify's per-character branch
  on quotes/backslashes/control chars is roughly 2x slower than
  UTF-8 transcode in TextEncoder, so the replacement is a CPU win
  even though it adds a single TextEncoder.encode per file.
- **Structured-clone memcpy of file contents** drops to zero — the
  contents' backing ArrayBuffers are ownership-transferred, not copied
  into the worker's heap. The envelope's struct-clone cost is now
  proportional to metadata size only.
- **JSON.parse on worker thread** likewise no longer scales with
  content size. Worker decodes each `Uint8Array` to string via
  `TextDecoder` lazily at the parse boundary — runs on the worker
  thread, parallel with continued main-thread work, vs U17's
  sequential JSON.parse blocking the worker before processBatch can
  start.

Pipelining: TextEncoder.encode (main) and TextDecoder.decode (worker)
can both run while the OTHER side is doing useful work. Under U17,
struct-clone was a synchronous main-thread blocker.

The ArrayBuffer ownership contract is load-bearing:

- File-content `Uint8Array`s are allocated via `TextEncoder.encode`,
  NOT `Buffer.from(str, 'utf8')`. TextEncoder produces a dedicated
  ArrayBuffer per call; `Buffer.from(str)` carves from Node's shared
  `Buffer.poolSize` slab for small strings, so transferring one
  pool-backed Buffer's ArrayBuffer would detach every other Buffer
  that shares that slab — silent data corruption.
- The envelope itself is NOT transferred. It MAY be pool-backed by
  `encodeMessage`, and at ~30-80 bytes/file the struct-clone cost is
  negligible. Not transferring avoids the same detach-collateral risk
  the contents path is careful to dodge.

Detection is strict: every input element must have both `path: string`
and `content: string`. A single non-conforming element disqualifies
the whole batch from the transfer path and falls back to the legacy
single-Uint8Array `encodeMessage` envelope. Safer than partial
transfer (which would split a sub-batch into mixed-shape messages
the worker can't reassemble).

`parse-worker.ts` `decodeIncomingMessage` recognizes the hybrid
`{envelope, contents}` shape, decodes the envelope, zips metadata
positionally with the contents array, decodes UTF-8 → string per file,
and hands the reassembled `ParseWorkerInput[]` to the existing
`processBatch`. Identical downstream behavior to U17 — the IPC
optimization is invisible above this line.

Test scaffolding (3 FakeWorkers + 1 integration-test preamble) gain a
`decodeDispatchedMessage` helper that tolerates BOTH shapes (legacy
single-frame Uint8Array AND the new hybrid envelope+contents) so the
in-process unit mocks keep their existing action-scripting API and the
9 ad-hoc integration test workers keep their `msg.type === 'sub-batch'`
handlers unchanged.

`buildDispatchMessage` is now exported from worker-pool.ts so its
contract can be tested in isolation. A new
`test/unit/worker-pool-transferlist.test.ts` pins:
  - hybrid shape produced for parse-worker inputs
  - transferList carries one ArrayBuffer per file in input order
  - envelope decodes to metadata only (no `content` field)
  - content bytes round-trip byte-for-byte through UTF-8 (ASCII,
    multi-byte, surrogate-pair emoji)
  - each content's ArrayBuffer is independently allocated (no pool
    sharing) — the load-bearing transfer-safety invariant
  - non-parse shapes, empty arrays, and mixed-conformance arrays all
    fall back to the legacy single-frame path

All 271 test files (6166 unit + integration tests) pass.

* fix(workers,tests,docs): apply ce-code-review findings (16 items)

Walks the full set of findings from a multi-agent code review (11
reviewers, 1 maintainability dispatch lost to tool-permission denial)
of the PR #1693 branch. All 16 actionable findings — 4 P1, 4 P2,
8 P3 — applied in a single pass against a consistent tree. Tests
pass (269/269 unit files, 29/29 integration).

P1 — bounds-only / disguised-bounds assertions across 4 test files
(per user-memory DoD §2.7):
  - worker-pool.test.ts: 5 sites — `nodes.length > 0` dropped (redundant
    after `.toContain('validateInput')`); `files.length >= 4` pinned to
    `.toBe(7)` (mini-repo/src has exactly 7 .ts files); `results.length
    > 0` pinned to `.toHaveLength(1)` (default sub-batch absorbs all 7);
    `result.fileCount >= 0` pinned to `.toBe(1)` (empty file is still
    "processed"); `warnRecords.length > 0` replaced with content-
    predicate `/respawn|dropping|replacement|did not report ready/`
    (catches silenced warnings); `fallbackExcludePaths.length > 0`
    pinned to exact `['one.ts', 'two.ts']` (deterministic given the
    single-slot pool + 2 items + per-item starting-file).
  - parse-impl-fallback.test.ts: 3 sites — `astCacheClearCalls >= 1`
    pinned to exact 4 (per-chunk × 2 + finally × 2); the two error-path
    delta checks pinned to exact +2 and +3 (verified empirically).
  - parse-impl-progress-monotonic.test.ts: `percents.length > 0` →
    `.not.toEqual([])`; per-element `Math.max(prev, cur)` tautology
    replaced with direct `if (cur < prev) throw`; final-percent
    `Math.min(last, 95)` tautology pinned to exact `.toBe(70)` (3-file
    skipWorkers fixture's deferred band lands at the band start).
  - parse-impl-large-fixture.test.ts: `Math.min(elapsedMs, BUDGET)`
    tautology removed; Promise.race rejection is the load-bearing
    wall-clock check.

P1 — terminate() lacks `.catch` mask:
  - worker-pool.ts terminate() now matches the `.catch(() => undefined)`
    pattern used at every other internal terminate site. Prevents a
    hung/OOM worker's terminate rejection from masking the original
    pipeline error when called from parse-impl.ts's finally block, and
    guarantees `workers.length = 0` / `activeSlots.clear()` always run.

P1 — hybrid envelope length-mismatch + null-payload silent data loss:
  - parse-worker.ts decodeIncomingMessage: explicit non-null-and-typed
    check before `.type` access (decodeMessage permits null payloads
    per encodeMessage contract); explicit length-equality assertion
    between `decoded.files` and `contents` before zipping. Without
    these, `TextDecoder.decode(undefined)` silently returns "" and
    produces empty-content graph nodes — a contract violation that
    used to be undetectable. Both throws route through the outer
    try/catch → worker `error` reply → pool's recoverAndResume.

P1 — unsafe casts at the IPC boundary:
  - buildDispatchMessage now uses a properly-typed `isParseWorkerItemArray`
    type guard. The narrowed branch accesses `item.path` and
    `item.content` as statically-typed strings — a future rename of
    `ParseWorkerInput.content` would fail to compile inside the branch
    instead of silently mismatching at runtime. The remaining
    decodeMessage payload casts are bounded by the F3/F6 runtime
    guards.

P2 — idle-timeout retry bypasses circuit breaker:
  - worker-pool.ts timeout-retry IIFE now increments
    `consecutiveFailuresPerSlot[workerIndex]` alongside `respawnCount`.
    A slot that consistently times out (vs crashes) now trips the
    per-slot breaker, instead of consuming its full respawn budget
    over potentially tens of minutes without the breaker firing.

P2 — null/non-object worker message crashes pool handler:
  - Dispatch handler in worker-pool.ts now guards `null /
    non-object / no string type discriminant` before `msg.type` access
    and routes through recoverAndResume on violation. Previously a
    legitimate `null` payload would throw TypeError out of the
    EventEmitter listener → uncaughtException on main, crashing the
    analyze.

P2 — workerPoolSize === 0 creates unusable pool:
  - parse-impl.ts now treats `workerPoolSize === 0` as `skipWorkers`
    at the gate. Matches the PipelineOptions docstring contract ("0
    disables the pool entirely — equivalent to skipWorkers"); avoids
    constructing a pool that rejects every dispatch and logs
    "Worker pool parsing stopped" per chunk.

P2 — encodeMessage 2-buffer allocation per frame:
  - protocol.ts encodeMessage coalesced to a single
    `Buffer.allocUnsafe + writeUInt8 + writeUInt32LE + buf.write
    (string, offset, 'utf8')`. Drops the intermediate
    `Buffer.from(JSON.stringify(...), 'utf8')` allocation + memcpy.
    Length pre-check via `Buffer.byteLength(string, 'utf8')` surfaces
    the uint32 cap before any allocation.

P3 — slotGenerations made optional on WorkerPoolStats so external
  implementations of getStats() that predate U12 don't compile-break;
  in-repo callers already use optional chaining.

P3 — buildDispatchMessage marked `@internal` so it isn't surfaced as
  public API by typedoc / api-extractor (it's a test-only export).

P3 — verboseThroughputLog hoisted above the chunk loop (env vars can't
  change mid-run; one O(env-read) per analyze, not per chunk).

P3 — corrected the messageerror routing comment in worker-pool.ts
  dispatch handler. `ProtocolDecodeError` is caught by the surrounding
  try/catch — distinct from `messageerror`, which fires for V8
  structured-clone failures before the message body would reach the
  handler.

P3 — initial pool spawn now uses a `Promise.allSettled` ready-handshake
  gate symmetric with `replaceWorker`. Dispatch awaits this gate before
  selecting slots, so an init-crashing initial worker is dropped from
  `activeSlots` and a downstream OOM/missing-native-binding failure
  surfaces in seconds (bounded by WORKER_READY_TIMEOUT_MS) rather than
  waiting for the first idle timeout (30s default).

P3 — `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT`,
  `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS`,
  `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` added to:
    - CLI `--help` text in src/cli/index.ts
    - Root README env-var table
    - gitnexus/README troubleshooting section (new "Worker pool
      resilience tuning" subsection)

P3 — CLI `catch (e: any)` / `catch (err: any)` in analyze.ts replaced
  with `catch (err: unknown)` + narrowed access; matches modern TS
  best practice and the codebase pattern at other catch sites.

P3 — `WorkerPoolStats.terminated: boolean` field added (optional, for
  backward compatibility). `terminate()` sets it true; `getStats()`
  surfaces it. Distinguishes graceful shutdown from a circuit-breaker
  trip in observability surfaces.

Coverage / advisory items not addressed in this commit (kept in the
report only):
  - maintainability reviewer failed (Read/Bash denied) — god-module
    audit on worker-pool.ts (~1400 LOC) carried as residual risk
  - quarantine case-sensitivity contract unpinned (adversarial #8)
  - WORKER_READY_TIMEOUT_MS env-configurability (adversarial #2)
  - chunk-byte-budget × parseChunkConcurrency memory multiplier doc
    (adversarial #5)
  - MCP discoverability gaps for env vars / verbose (agent-native W1/W2)
  - bench/parse-throughput.md scaffold-with-TBD-rows (PS RR-003)

* fix(parsing): sequential gap-fill for worker-quarantined chunk files (U20.U1)

When the worker pool's Layer 3 quarantine filters one or more files
out of a chunk's dispatch, the worker results returned to
processParsing are silently narrower than the input chunk. Without
this reparse, the graph for this run would be missing every quarantined
file's symbols/imports/calls/heritage with no failure signal.

After the existing per-chunk quarantine log emits in
processParsing's worker-path try-block, run processParsingSequential
on JUST the quarantined-in-chunk files. The sequential path writes
directly to the graph, so symbols for those files land alongside
worker output for the surviving files.

Mirrors the WorkerPoolDispatchError catch-block's processParsingSequential
call shape — same signature, same args, same scopeTreeCache wiring.
Emits a structured warn naming `reparsedPaths` so operators can
observe the sequential fall-through.

This fixes the in-run side of the corruption Codex's adversarial
review of PR #1693 flagged. The cross-run side (chunk-cache
poisoning) is closed by U20.U2 in a follow-up commit.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* fix(parse-impl): suppress chunk-cache write when any chunk file was quarantined (U20.U2)

The chunk hash at parse-impl.ts:424-428 is computed from every file
in the chunk. The worker pool's Layer 3 quarantine
(worker-pool.ts createQuarantine) filters quarantined files out of
dispatch, so `rawResults` reflects only the surviving files. Before
this commit, the write at line 500-507 stored that partial result
under the full-coverage chunk hash — and on the next analyze with
unchanged content, the cache HIT branch (line 439-464) silently
replayed the incomplete result. Symbols from the quarantined file
were missing from the graph for as long as the cache survived.

Codex's adversarial review of PR #1693 flagged this as a silent-
corruption class because there's no failure signal: no warn log
during the replay, no graph-equivalence check, no exit code change.
The corruption only surfaces if an operator notices a missing symbol
in `gitnexus_query` output.

Guard the write with `chunkFiles.some(f => quarantineSet.has(f.path))`.
When any chunk file is in the worker pool's cumulative quarantine
snapshot, skip the `parseCache.entries.set` call. Emits a verbose-
only info log so operators investigating "why aren't my chunks
caching" have a diagnostic trail.

Skipping the write means the next analyze gets a cache miss for this
chunk and re-dispatches it. Quarantine is session-scoped (a fresh
createWorkerPool starts with an empty quarantine), so the new pool
gives the quarantined file another chance. If quarantine fires again,
U20.U1's sequential gap-fill still produces a complete graph for that
run; the cache stays empty for the chunk until a fully-clean
dispatch lands.

The cache-hit replay branch at parse-impl.ts:439-464 is unchanged.
Its contract strengthens: "cache entries are complete" becomes true
post-fix, but the replay code doesn't need to know that.

Closes the cross-run side of the Codex finding. U20.U3 adds the
regression test.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* test(parse-impl): integration regression for quarantine + chunk-cache (U20.U3)

Pins the U20 fix end-to-end via REAL `worker_threads` + `createWorkerPool`.
Mirrors the writeReadyWorker pattern from `test/integration/worker-pool.test.ts`
— inline READY_PREAMBLE + custom test worker script that:

  1. Decodes the U17/U19 IPC protocol (Buffer frame OR hybrid envelope/
     contents shape) the same way the production parse-worker does.
  2. Emits a `{type:'ready'}` handshake so the pool's
     `waitForWorkerReady` resolves promptly.
  3. On a sub-batch containing `poison.ts`, emits starting-file +
     `process.exit(134)`. The pool attributes the death to `poison.ts`
     via the in-flight signal and adds it to the session-scoped
     quarantine.
  4. On a sub-batch without poison, synthesizes a minimal valid
     `ParseWorkerResult` with one `Function` node per file (no
     tree-sitter dep in the test worker — the synthesized nodes give
     `mergeChunkResults` deterministic content for the graph).

Assertions exercise both fix layers:

  - U1 (sequential gap-fill in processParsing): the graph contains a
    `Function` node named `poison` AFTER the run. The custom worker
    never emits anything for `poison.ts`, so the only path for that
    symbol to reach the graph is `processParsing`'s sequential
    reparse of the quarantined-in-chunk file using the real
    tree-sitter parser against the actual source.
  - U2 (cache-write suppression in runChunkedParseAndResolve):
    `parseCache.entries` does NOT contain the chunk hash after the
    run; `parseCache.usedKeys` DOES contain it (chunk processed,
    cache write specifically skipped).
  - Cross-run: a second pass over the same fixture with the same
    parseCache and a fresh worker pool re-dispatches the chunk
    (cache empty), the worker crashes again, sequential gap-fill
    runs again, and the cache stays empty. Pins the round-trip
    contract.

Adds `workerUrlForTest?: URL` to PipelineOptions — same `@internal`
test-only injection precedent as `workerThresholdsForTest` (already
in PipelineOptions for thresholds). When set, parse-impl uses the
provided URL instead of the src/ → dist/ resolution dance. Production
call sites never set this field; the only consumer today is this
integration test.

Why integration over unit:
  - The fix lives at the boundary between parsing-processor.ts and
    parse-impl.ts under a real WorkerPool. Unit-mocking the
    worker-pool module bypasses the structured-clone boundary, the
    dispatch lifecycle, and the actual quarantine flow — it verifies
    the test setup rather than the contract. The real worker thread
    executing through the U17/U19 IPC protocol IS the load-bearing
    surface.
  - User-explicit preference (saved as
    feedback_integration_over_vimock.md memory). For worker-pool /
    parse-impl / IPC-touching code: write integration tests under
    test/integration/ using writeReadyWorker patterns; avoid
    vi.mock on worker-pool.js.

Test wall-clock: under 2s; both `it` blocks together complete in
~1.8s under the existing CI conditions.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* refactor(parsing): remove sequential-parser fallback (U20 design pivot)

The worker pool's resilience layers — respawn budget, circuit breaker,
quarantine, slot-attribution, cumulative timeout — are now the SOLE
contract for handling worker failures. Two sequential-reparse paths
are removed from processParsing:

1. **U20.U1 sequential gap-fill for quarantined chunk files** (just
   added in commit 7dd489e9, now reverted). The pre-emptive rescue
   would re-run processParsingSequential on the file that ALREADY
   killed a worker — which for the most common quarantine cause
   (tree-sitter native SIGSEGV on a pathological file) re-triggers
   the same native crash on the main thread, killing the entire
   analyze. The "rescue" turned silent missing-symbols into a louder
   analyze-wide crash. Drop the rescue; accept the per-run gap.

2. **Pre-existing WorkerPoolDispatchError catch-block sequential
   fallback** (in production since PR #1693's resilience layer
   landed). Same risk class — when the pool exhausts its respawn
   budget / trips the circuit breaker, the failing files are
   precisely the ones likely to crash a sequential parser too. The
   "graceful degradation" hid pool failures behind degraded-but-
   completing analyze runs, making operational issues harder to
   surface and diagnose. Drop the catch-block; WorkerPoolDispatchError
   propagates to the analyze entry point where the user sees a clear
   hard signal.

What stays:
- The `skipWorkers: true` / small-repo path that uses
  `processParsingSequential` as the EXPLICIT primary path (not a
  fallback). Caller-driven opt-out and tiny-repo perf optimization
  are different intents.
- U2's chunk-cache write suppression in parse-impl.ts (commit
  7c9c9556). When quarantine fires, the chunk stays uncached so the
  next analyze with a fresh pool retries the file cleanly. That's
  the cross-run correctness Codex's adversarial review actually
  asked for.
- The per-chunk quarantine warn log (parsing-processor.ts) — operators
  see which files were skipped, both immediately and across runs.

What changed:
- `processParsing` worker-path try-block: unwrapped. The
  `processParsingWithWorkers` call is now direct (no try/catch
  wrapping); errors propagate to the chunk-loop caller.
- `parsing-worker-fallback.test.ts` rewritten: the previous 5 tests
  asserted graceful sequential-fallback behavior. Replaced with 3
  tests pinning the new contract — raw Error propagates, WorkerPool-
  DispatchError propagates with fallbackExcludePaths intact, normal
  quarantine signal does NOT throw and surfaces via progress detail.
- `parse-impl-quarantine-cache-skip.test.ts` (U20 integration test)
  updated: poison.ts is NOT in the post-run graph; surviving files
  are; chunk-cache stays empty; second pass re-dispatches and leaves
  cache empty.
- Plan doc updated to mark R1 as dropped and explain the U20 pivot
  in the Summary.

User decision: explicit directive ("let's remove the sequential
fallback entirely we must rely on entirely that the parallel process
is resilient enough to work itself through the code base"). The pool's
resilience layers are designed for this — respawn budget, circuit
breaker, quarantine, slot-generation, cumulative-timeout cap — and
adding a layer below them was redundant insurance with real downside.

Tests: 269/269 unit files (6135 tests) green. 31/31 worker-pool +
parse-impl integration tests green. The 2 reported "errors" in the
integration run are the pre-existing intentional-process.exit unhandled-
exception leaks from test workers — unchanged by U20.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* fix(workers,tests,docs): address ce-ultrareview findings F1/F2/F3/F4

Multi-lane review run on the PR #1693 branch surfaced four addressable
items beyond the blocking three.

F1 (minor, CodeQL): unused `findMatch` helper in
test/unit/scope-resolution/typescript/typescript-captures-anchor.test.ts:28
removed. `countMatchesTsx` flagged by the same CodeQL pass is a false
positive — it's called at line 88 by the JSX-anchor regression tests
so the rewrite case actually fires under TSX, not just TS.

F2 (medium, docs): bench/parse-throughput.md retitled as
"(scaffold)" with an explicit "no measurement data has been collected
yet" note above the table. The self-contradictory "Regenerate this
file before merging any PR that touches the ingestion pipeline"
instruction is dropped — the file ships intentionally without
numbers; the load-bearing perf-regression protection lives in
test/integration/parse-impl-large-fixture.test.ts (U6, 30s
Promise.race wall-clock budget). The Latest measurement section now
preserves the ~6s sequential observation as a smoke reference, not as
a regression target.

F3 (low, API hygiene): `WorkerPoolDispatchError.fallbackExcludePaths`
renamed to `quarantinedPaths`. The "fallback" terminology was
load-bearing under the pre-U20 design when `processParsing`'s
sequential-fallback catch-block consumed it to filter the fallback
file list. After commit be1f65c removed that catch-block, no
production code reads the field — but it stays populated by the pool
because the snapshot is genuinely useful operator diagnostics when
the breaker trips. The rename clarifies the field's actual semantics
(here are the files the pool quarantined before it tripped) without
changing wire behavior. Definition + the lone surviving in-pool
comment reference + both test assertions updated.

F4 (low → real fix, reliability): timeout-retry IIFE in
worker-pool.ts now consults `consecutiveFailureThreshold` and trips
the circuit breaker when the per-slot consecutive-failure count
crosses it. Closes a gap left by ce-code-review's REL-02 patch — that
fix added the `consecutiveFailuresPerSlot[workerIndex]++` increment
in the timeout-retry path but did NOT add the corresponding
threshold-check + tripBreaker call. Result: chronic pure-timeout
deaths accumulated counts that never tripped the breaker until the
slot also hit `respawnCount > maxRespawnsPerSlot`. Now timeouts and
crashes are structurally treated the same way by the breaker, which
is what the REL-02 increment was meant to enable. Test coverage:
worker-pool-resilience.test.ts already exercises the breaker via the
shared handleWorkerDeath path; this new branch traces the same
trip semantics with a different entry point, so the breaker-tripped
state is observable via the same `getStats().poolBroken` and
`WorkerPoolDispatchError.quarantinedPaths` surface.

Out of scope here (caller actions or future PRs):
  - F5 (info): cumulative-quarantine cache check is safe in practice
    because chunks are alphabetically deterministic; no action.
  - F6 (low): exit-code-0 quarantine exemption — pre-existing P2
    residual, bounded by quarantine + respawn budget; deferred.
  - F7 (info): dispatch non-reentrancy contract documented but not
    enforced; no production caller violates it; deferred.
  - PR title `[WIP]` removal — happens on GitHub side.

Tests: 274/274 test files (6185 passing, 30 skipped). The single
"error" in the integration runner is the pre-existing intentional-
process.exit unhandled-exception leak from the deliberate startup-
crash test worker, unchanged by these fixes.

* fix(workers): swap protocol body from JSON to V8 serialize/deserialize

CI scope-parity tests on Ubuntu surfaced silent data loss in the
worker IPC: `Phase 'scopeResolution' failed: scope.typeBindings is not
iterable` (Python, Go) and `importerModule.typeBindings.has is not a
function` (Python). Plus three #1066 large-file regression tests
(Python / C# / TypeScript) failed because call relationships weren't
resolving from the worker output.

**Root cause:** U17 introduced `JSON.stringify`/`JSON.parse` as the
protocol body codec. JSON has no representation for `Map`, `Set`,
`Date`, `RegExp`, `BigInt`, `TypedArray`, `undefined` values, or
circular refs — `JSON.stringify(someMap)` returns `"{}"`. Production
scope-resolution code keys data structures on Maps throughout
(`ParsedFile.scopes[*].typeBindings: ReadonlyMap<string, TypeRef>`,
plus `bindings`, `bySourceScope`, `byTargetDef`, the finalize-algorithm
edge indexes, etc.). The JSON round-trip silently turned every Map
into an empty object, manifesting downstream as iteration / `.has`
calls failing on the decoded payload.

**Fix:** replace the JSON body with `node:v8`'s `serialize` /
`deserialize`. That's the same structured-clone algorithm Node's
`worker.postMessage` uses natively — bit-for-bit compatible with the
pre-U17 implicit-clone path. Full type fidelity for Map, Set, Date,
RegExp, BigInt, TypedArray, undefined values, and circular refs. No
external dependency.

A previous iteration of this fix attempted to bolt a Map/Set
replacer+reviver onto the JSON path. Rejected in favor of V8
serialization because:
  - the JSON tag-marker approach requires per-type registration
    (Map, Set; then Date, RegExp, BigInt would each need their own
    sentinels); V8 handles them all uniformly
  - keys to JSON-encode would still need handling for nested types
    (and the marker approach doesn't survive nested Maps-in-Maps
    cleanly without recursive replacer logic)
  - V8 is faster than JSON for object-heavy payloads anyway (binary
    format, no string escaping pass)
  - the user-explicit ask was "a much more generic solution that will
    work for everything" — V8 serialization IS the generic solution

Trade-offs documented in the module header:
  - body bytes are opaque (binary, not human-readable) — debugging
    requires `v8.deserialize` ad-hoc; protocol.test.ts exercises every
    supported MessageTag including the new type-fidelity cases as a
    regression net.
  - format is tied to the running Node major. Pool always spawns
    workers on the same Node instance the main thread runs, so this is
    moot in production. Would matter if frames ever persisted to disk
    (nothing does today).

Protocol test file rewritten:
  - drops the JSON-specific byte-layout assertions (e.g. `body must
    equal "null" string`) — replaced with V8-derived expected lengths
  - adds a "structured-clone type fidelity" describe block that pins
    Map, nested Map, Set, Date, RegExp, BigInt, TypedArray, undefined
    values, and circular-ref round-trips. These are the load-bearing
    regression tests preventing a future "optimize" PR from quietly
    swapping V8 back to JSON.
  - the bad-body decode-error test now uses arbitrary non-V8 bytes
    instead of `{not-json}` — same intent.

Integration test READY_PREAMBLEs (worker-pool.test.ts and
parse-impl-quarantine-cache-skip.test.ts) update their inline
decoders to use `v8.deserialize` matching the production codec.
Both files have a standalone CJS worker preamble that can't import
dist/protocol.js by relative path, so the V8 dependency is required
via `node:v8` directly.

Tests: 271/271 unit files (6163 tests + 30 skipped). 28/28
worker-pool integration. 3/3 parse-impl integration. 791/791
scope-parity tests (the four CI-failing files: python.test.ts,
go.test.ts, typescript.test.ts, csharp.test.ts) all green again.

References plan: docs/plans/2026-05-20-002-fix-chunk-cache-corruption-on-worker-quarantine-plan.md

* refactor(workers): drop protocol.ts; use native postMessage + transferList

The protocol.ts framing layer was redundant — Node's `worker.postMessage`
already runs V8 structured-clone internally, the same algorithm that
backed `v8.serialize`. Wrapping V8.serialize → Buffer →
postMessage(struct-clone-Buffer) was a double-walk: one full
structured-clone pass to produce the Buffer, then another pass when
postMessage cloned that Buffer across threads. This commit cuts the
wrapper layer; workers and pool exchange POJO directly via
`worker.postMessage(value, transferList)`, with file-content
`ArrayBuffer`s in `transferList` for zero-copy ownership transfer.

What changes:

- **Deleted** `src/core/ingestion/workers/protocol.ts` (~180 LOC) +
  `test/unit/workers/protocol.test.ts` (~250 LOC). The MessageTag
  enum / ProtocolDecodeError / encodeMessage / decodeMessage surface
  is gone. Tag-based routing is replaced by the `msg.type`
  discriminant that every receive site already checks. Protocol-decode
  errors map to Node's `messageerror` event (V8 deserialization
  failures during postMessage), which the pool already wires to
  `recoverAndResume`.
- **`worker-pool.ts`**: `decodeIncomingWorkerMessage` removed; handlers
  receive POJO directly. `buildDispatchMessage` now returns
  `{message: {type:'sub-batch', files: [{path, content: Uint8Array}]},
  transferList: ArrayBuffer[]}`. The Uint8Array-per-content allocation
  via `TextEncoder.encode` is preserved (it's the load-bearing
  transfer-safety contract that keeps content out of Node's shared
  `Buffer.poolSize` slab). Flush dispatch is now plain
  `worker.postMessage({type:'flush'})`.
- **`parse-worker.ts`**: `decodeIncomingMessage` removed. The message
  handler receives POJO directly; the only conversion is
  `Uint8Array → string` for sub-batch file contents at the
  `decodeSubBatchFiles` boundary, before handing to `processBatch`.
  Outgoing messages are emitted as POJO via plain
  `parentPort.postMessage({type:'starting-file', ...})` etc. The
  `sharedHybridDecoder` is now `sharedContentDecoder` (same intent,
  clearer name for the simpler shape).
- **Test scaffolding**: FakeWorkers in `worker-pool-resilience`,
  `worker-pool-windows-quarantine`, and `worker-pool-slot-generation`
  drop their `decodeMessage` import + `decodeDispatchedMessage` helper.
  The helpers stay (still convert `files[i].content` Uint8Array →
  string for test-action introspection) but no longer touch any
  protocol framing — just shape-check for sub-batch.
- **Integration READY_PREAMBLEs** (worker-pool.test.ts and
  parse-impl-quarantine-cache-skip.test.ts): drop the inline
  v8.deserialize + envelope-unzip logic; the preamble is now just
  the ready handshake + a `parentPort.on` wrapper that converts
  `files[i].content` Uint8Array → string for the ad-hoc test worker
  scripts.
- **`worker-pool-transferlist.test.ts`**: contract tests updated for
  the new buildDispatchMessage shape — no `envelope` field anymore;
  `message.files[i].content` is Uint8Array; transferList holds each
  content.buffer in input order. Pool-slab independence still pinned.

What stays the same:

- Zero-copy file-content transfer via transferList — every file's
  ArrayBuffer is ownership-transferred to the worker (no copy).
- Full structured-clone type fidelity — Map / Set / Date / RegExp /
  BigInt / TypedArray / undefined / circular refs all preserved by
  Node's native postMessage. The V8 fix from commit 06f6957e is
  inherent in this path; there's no JSON layer to lose them.
- TextEncoder-per-content allocation — keeps content buffers out of
  the shared `Buffer.poolSize` slab so transferring one cannot detach
  another.
- The pool's resilience layers (respawn, breaker, quarantine,
  starting-file attribution, cumulative timeout, ready handshake,
  slot-generation guard) — unchanged.
- U20 chunk-cache write suppression on quarantine — unchanged.

Net: ~430 LOC removed (protocol.ts + tests + inline decoders + helpers),
~120 LOC simplified in worker-pool.ts and parse-worker.ts. One less
serialization pass per message on the hot path.

Tests: 270/270 unit files (6133 + 30 skipped). 822/822 integration
tests including the four CI-failing scope-parity files (Python, Go,
TypeScript, C#) — the V8-fidelity contract holds via native
postMessage with no explicit serializer. The single "error" reported
in worker-pool.test.ts is the pre-existing intentional
process.exit unhandled-exception artifact from the deliberate
startup-crash test, unchanged by this commit.

* refactor(parse-worker): drop legacy single-message dispatch mode

The `parentPort.on('message', ...)` handler had an `Array.isArray(msg)`
branch left over from a pre-sub-batch dispatch shape — the pool used
to send the items array directly, before the worker pool added
sub-batching and the `{type:'sub-batch', files: ...}` envelope.

No production caller has dispatched that shape since the sub-batching
refactor landed; verified by grepping the repo for `postMessage([`
patterns (zero matches). The `ParseWorkerInput[]` arm in the
`WorkerIncomingMessage` discriminated union also blocked
exhaustiveness narrowing — flagged by the kieran-typescript code
review (RR-01) as "if a future unit removes the legacy array path,
this arm should be dropped." Dropping it now.

What changes:
  - Remove the `Array.isArray(msg)` branch from the message handler.
  - Drop `ParseWorkerInput[]` from the `WorkerIncomingMessage` union;
    it's now a clean `{type:'sub-batch'} | {type:'flush'}` discriminated
    union, so the dispatch switch is exhaustive over `msg.type`.

Tests: 71/71 worker-pool unit + integration tests green (resilience,
slot-generation, windows-quarantine, transferlist, parsing-worker-
fallback, worker-pool integration, parse-impl-quarantine-cache-skip).

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 20:39:35 +01:00
azizur100389andGergő Magyar aa8f4d6efe fix(group): Union HTTP graph and source contracts (#1709)
* Union HTTP graph and source contracts

* test(group): Document HTTP source union follow-ups

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 17:44:07 +01:00
Shane Thurston Wijaya 4d2ed0e525 fix(eval-server): localhost now doesn't normalize into IPv4 instead lets OS decide which to bind (#1722)
* fix(eval-server): localhost now doesn't normalize into IPv4 instead lets OS decide which to bind

* fix(eval-server): EADDRNOTAVAIL now treats as potential IPv6

* test(eval-server): new integration test for --host localhost

* docs(eval-server): updated eval/README.md based on latest update

* fix(eval-server): clarify EADDRNOTAVAIL diagnostic, guard server.address(), and soften localhost docs
2026-05-20 16:14:13 +01:00
df1882d36b fix(ingestion): surface skipped large-file paths by default (#1659) (#1661)
* fix(ingestion): surface skipped large-file paths by default (#1659)

The 512 KB skip threshold in filesystem-walker is necessary, but the
existing warning only said "Skipped N large files" with no paths unless
GITNEXUS_VERBOSE=1 was set. In a repo with one or two oversized first-
party source files (e.g. a 17K-line cron handler), every IMPORTS/CALLS
edge from that file silently disappeared and the surface looked like a
Python resolver bug. Issue #1659 was filed against the resolver for
exactly that reason, but the resolver was fine; the file was being
dropped before parse.

Changes:
  * Always print up to 5 skipped paths after the count line.
  * If more than 5 were skipped, append "...and N more" with a hint to
    set GITNEXUS_VERBOSE=1 for the full list.
  * When running at the default threshold, emit a one-line hint about
    GITNEXUS_MAX_FILE_SIZE=<KB> so operators know how to widen it.
  * Cover the new behavior with three additional tests in the existing
    filesystem-walker integration suite, plus a new describe block for
    the >5 preview-cap case.

Verified end-to-end on a 680-file Python repo that hit #1659: before
the patch, "Skipped 3 large files (>512KB, ...)" was the only signal
and impact upstream of a function called from cron.py returned 1 of 5
real callers; after the patch the cron file is listed by name with the
hint, and running with GITNEXUS_MAX_FILE_SIZE=1024 brings the missing
callers back (impactedCount 1 -> 9).

* fix(ingestion): address #1661 adversarial review follow-ups (F1/F2/F3)

Three non-blocking nits flagged by the adversarial review on #1661:

F1 (output stability) — skippedLargePaths was populated by concurrent
fs.stat callbacks in batches of 32, so push order within a batch was
completion-order rather than input-order. The default preview's "first
5" could vary across runs on the same repo. Fix: sort the array before
slicing. New test asserts the verbose output is in sorted order.

F2 (boundary coverage) — the preview-cap describe block created 8
large files, so the SKIPPED_PREVIEW_CAP = 5 comparison was never
exercised at the exact <= boundary. A future off-by-one (<= → <) would
not fail the suite. Fix: add two tests, one with exactly 5 files (all
listed, no truncation) and one with exactly 6 files (5 listed plus
"...and 1 more").

F3 (hint accuracy) — isDefault compared effective bytes, so an
operator who explicitly set GITNEXUS_MAX_FILE_SIZE=512 (the same KB as
the default) would still see the "Set GITNEXUS_MAX_FILE_SIZE=<KB>..."
hint. Fix: gate the hint on whether the env var is unset, not on the
resulting byte value. New test pins the explicit-default-value case.

All 34 filesystem-walker tests pass (was 30; +4 new). Prettier clean,
typecheck clean for the changed files.

---------

Co-authored-by: scotjelinski <58397194+scotjelinski@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 14:37:58 +01:00
f350ae278a feat: Add analyze --repair-fts, enforce FTS verification, and harden repair safeguards (#1720)
* Initial plan

* feat(analyze): add --repair-fts and verify FTS index rebuilds

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/dccb3673-af86-43aa-aede-2e1449399775

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor(fts): tighten repair/verify messaging and option naming

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/dccb3673-af86-43aa-aede-2e1449399775

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: highlight analyze --repair-fts vs --force in READMEs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/61edc967-debc-419f-9f51-aebf2ef08d22

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(analyze): guard repair mode against missing graph store

* fix(cli): reject --repair-fts with --force

* test(analyze): document repair-store fixture intent

* test(analyze): tidy repair failure fixtures and constants

* test(analyze): clarify mock constants in repair tests

* test(analyze): rename simulated missing-index constant

* test(analyze): clarify mocked graph shape in full-verify test

* refactor(analyze): finalize flag validation and test clarity

* test(skip-git): avoid hard failing when FTS extension is unavailable

* test(skip-git): log visible FTS-unavailable test skips

* test(skip-git): tighten FTS-unavailable error detection

* test(skip-git): simplify FTS-unavailable message checks

* test(skip-git): avoid HOME pointing at parent repo in fixture env

* fix(analyze): address Claude follow-up findings for repair guardrails

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* fix(repair-fts): clarify invalid graph-store preflight errors

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* test(analyze): strengthen assertions for conflict and missing-store errors

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* fix(repair-fts): make invalid graph-store type errors explicit

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

* fix(repair-fts): improve graph-store type diagnostics

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2f7243d3-ba16-4d83-86e5-17e6c58a3b0d

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 13:37:04 +01:00
azizur100389andGergő Magyar dae70a26ea feat(cpp): Add pointer nullptr ellipsis conversion ranks (#1708)
* Add C++ pointer null ellipsis ranks

* test(cpp): Strengthen pointer overload assertions

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 12:06:51 +01:00
CopilotandGergő Magyar b4a2a4b91e fix(ingestion): Prioritize same-module Java type resolution for duplicate FQNs across modules (#1712)
* Initial plan

* Fix Java same-name type resolution with same-module priority

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/df0843e3-e244-4e0f-a94a-311df3899bd0

* Refine Java ambiguity fallback safety check

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/df0843e3-e244-4e0f-a94a-311df3899bd0

* Remove Java-specific fallback from shared scope walkers

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e882906a-2c96-411e-94a5-123a345421a9

* Harden Java module key and ambiguous owner fallback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e882906a-2c96-411e-94a5-123a345421a9

* Add negative assertions for duplicate-FQN module edges

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e882906a-2c96-411e-94a5-123a345421a9

* Make Java same-module ordering path-agnostic

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Refine generic Java path-affinity ordering safeguards

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Polish Java path-affinity ordering clarity

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Simplify Java path-affinity ordering logic

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b317a759-f6bc-4590-bd2a-628f0ee9c477

* Revert legacy DAG Java ambiguity ordering changes

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/94e50cf2-9733-4e69-a0eb-9fd38cbdb589

* Skip duplicate-FQN Java assertions in legacy parity mode

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/1b560efa-1b3b-4697-b590-c6ef447f431e

* Tighten duplicate-FQN Java CALLS edge cardinality assertions

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/67c18f93-5e56-4b15-8404-cdf1be9b4485

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 08:00:03 +01:00
Nilotpal KashyapandGergő Magyar d7e1815aa3 fix(detect-changes): guard resolveWorktreeCwd against overriding a separately-indexed worktree (#1691)
* fix(detect-changes): guard resolveWorktreeCwd against overriding a separately-indexed worktree

When the repo registry entry points to a linked worktree (both main
checkout and worktree indexed separately), resolveWorktreeCwd was
incorrectly replacing the correct worktree repoPath with the server's
main-checkout launch directory. Both share the same canonical root so
the existing same-repo check passed, causing git diff to run from the
wrong directory and return 0 changes (issue #1659).

Fix: early-exit guard — if tryRealpath(repoPath) differs from
tryRealpath(getCanonicalRepoRoot(repoPath)), repoPath is itself a
linked worktree and is returned unchanged. Auto-detection only fires
when repoPath equals the canonical main-checkout root.

Also normalises the launchCanonical comparison in the auto-detect path
to use tryRealpath for cross-platform consistency.

Regression test: 'returns worktreeDir unchanged when repoPath IS a
linked worktree and launchCwd is the main checkout'.

* test(detect-changes): add worktreeA→worktreeB case and assumption comment

Cover the missing case from the production-readiness review:
repoPath = wt-A (indexed), launchCwd = wt-B (server on a different
linked worktree). The guard fires on repoPath being a worktree
regardless of launchCwd, so wt-A is returned unchanged.

Also add an inline comment documenting the assumption that repoPath
is a git root or linked-worktree root (not an arbitrary subdirectory),
as noted in Finding 2 of the review.

* refactor(detect-changes): validate repoPath is a git root before canonical comparison

Instead of relying on a comment asserting repoPath is always a git
root, call getGitRoot(repoPath) first. Only if the result matches
repoPath itself do we call getCanonicalRepoRoot and apply the guard.

This eliminates the over-classification risk for subdirectory repoPath
values and makes the assumption explicit in code. repoCanonical is
shared across both the guard and the auto-detect block.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-20 06:46:16 +01:00
dependabot[bot] 92ad0f5491 chore(deps): bump idna in /eval in the uv group across 1 directory (#1713) 2026-05-20 05:38:26 +01:00
LocallyInsaneDBandGergő Magyar 803f0bed5f fix(lbug): probe-then-load FTS extension on Windows (#1690) (#1692)
* fix(lbug): probe-then-load FTS extension on Windows (#1690)

The Windows skip-on-process.platform==='win32' guard in pool-adapter.ts
hard-skipped loadFTSExtension() for every Windows host, even when the
FTS extension binary was already present locally at
~/.lbdb/extension/<version>/win_amd64/fts/libfts.lbug_extension.

That left BM25 silently degraded on Windows hosts that had a working
extension on disk, with no error path — `gitnexus doctor` still reported
FTS as available, but query returned 0 BM25 hits.

This patch adds hasLocalWinFtsExtension() which probes
~/.lbdb/extension/*/win_amd64/fts/ before the Windows skip. When a binary
is on disk we call loadFTSExtension(..., { policy: 'load-only' }); the
crashing install path documented in #1199 / #1217 is never exercised at
query time, and LadybugDB's version-specific resolution combined with
the ExtensionManager's tryLoad try/catch handles stale or zero-byte
sibling version dirs cleanly (no dlopen attempted on a stale binary).
When no binary is on disk at all, we fall back to the upstream skip so
install-time SIGSEGV continues to be avoided.

Verified on Windows 10 + Node 22.19.0 + gitnexus 1.6.5 +
@ladybugdb/core 0.16.1 with the FTS extension cached at 0.16.0:

  * BM25 timing goes from 0 → ~250-326ms on previously-zero queries
  * gitnexus context / impact / cypher unaffected
  * Adversarial-mixed-state run (real 0.16.0 binary + zero-byte stubs at
    0.15.0, 0.16.1, 0.17.0): exits 0, no SIGSEGV, FTS resolves to the
    real 0.16.0 binary, BM25 returns real hits
  * Stub-only state at the resolution path (0.16.0, zero-byte): exits 0,
    emits "FTS extension unavailable; load-only policy: extension not
    pre-installed", FTS marked unavailable cleanly via markUnavailable
    in extension-loader.ts — no silent greenlight

Closes #1690

* test(lbug): cover hasLocalWinFtsExtension probe + format pool-adapter

- Export hasLocalWinFtsExtension and add lbug-pool-win-fts-probe.test.ts
  with 7 cases against a real tmpdir + os.homedir spy:
    * missing ~/.lbdb/extension dir -> false
    * extension root present but no version dirs -> false
    * one version dir with binary present -> true
    * zero-byte stub at probe path -> true (LOAD failure handled downstream)
    * multi-version with binary only in a non-first dir -> true
    * multi-version with no binary anywhere (Nix/Bazel/MDM tree) -> false
    * fs.readdir throws (EACCES) -> false

  The Windows conditional in doInitLbug / initLbugWithDb is intentionally
  not unit-isolated: it reduces to `probe ? load : true` over a fully
  constructed lbug.Database + Connection pool, which the
  test/integration/lbug-pool*.test.ts suites already exercise on the
  windows-latest CI matrix.

- Apply prettier format to the fs.stat() call in pool-adapter.ts,
  resolving the quality/format CI failure surfaced by gitnexus/autofix.

Addresses DoD §2.7 test-coverage blocker raised in the production-
readiness review on #1692, and the dir-exists-no-file regression case
raised on #1690.

Refs #1690.

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-19 12:09:21 +01:00
55f8d442f6 fix(mcp): setup fallback on Windows when global gitnexus resolves to a non-spawnable shim (#1694)
* Initial plan

* fix: avoid invalid Windows MCP shim paths

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a052306e-483a-42d0-b65a-2646906457c7

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: cover .ps1 windows mcp fallback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/5aaed570-2a0b-4ed9-a0ac-ca099ce5675e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: assert windows fallback for cursor and codex

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/5aaed570-2a0b-4ed9-a0ac-ca099ce5675e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-19 08:17:26 +01:00
18167400c4 chore(deps)(deps): bump express and @types/express in /gitnexus (#872)
* chore(deps)(deps): bump express and @types/express in /gitnexus

Bumps [express](https://github.com/expressjs/express) and [@types/express](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/express). These dependencies needed to be updated together.

Updates `express` from 4.22.1 to 5.2.1
- [Release notes](https://github.com/expressjs/express/releases)
- [Changelog](https://github.com/expressjs/express/blob/master/History.md)
- [Commits](https://github.com/expressjs/express/compare/v4.22.1...v5.2.1)

Updates `@types/express` from 4.17.25 to 5.0.6
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/express)

---
updated-dependencies:
- dependency-name: "@types/express"
  dependency-version: 5.0.6
  dependency-type: direct:development
  update-type: version-update:semver-major
- dependency-name: express
  dependency-version: 5.2.1
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

* fix(server): normalize jobId param for Express 5 SSE routes

Express 5 types req.params values as string | string[]. mountSSEProgress uses a dynamic route path so TypeScript cannot narrow jobId; assert it once with assertString and reuse in the SSE progress callback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(ci): retrigger CI

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-19 07:43:07 +01:00
dad1ca7ab5 chore(deps)(deps): bump zod from 3.25.76 to 4.3.6 in /gitnexus-web (#1464)
Bumps [zod](https://github.com/colinhacks/zod) from 3.25.76 to 4.3.6.
- [Release notes](https://github.com/colinhacks/zod/releases)
- [Commits](https://github.com/colinhacks/zod/compare/v3.25.76...v4.3.6)

---
updated-dependencies:
- dependency-name: zod
  dependency-version: 4.3.6
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:56:04 +01:00
c746f30c90 chore(deps)(deps): bump langsmith (#1552)
Bumps the npm_and_yarn group with 1 update in the /gitnexus-web directory: [langsmith](https://github.com/langchain-ai/langsmith-sdk).


Updates `langsmith` from 0.5.23 to 0.6.3
- [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases)
- [Commits](https://github.com/langchain-ai/langsmith-sdk/commits)

---
updated-dependencies:
- dependency-name: langsmith
  dependency-version: 0.6.3
  dependency-type: indirect
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:55:44 +01:00
15a667ae5e chore(deps)(deps-dev): bump @vitest/coverage-v8 in /gitnexus (#1604)
Bumps [@vitest/coverage-v8](https://github.com/vitest-dev/vitest/tree/HEAD/packages/coverage-v8) from 4.1.5 to 4.1.6.
- [Release notes](https://github.com/vitest-dev/vitest/releases)
- [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.6/packages/coverage-v8)

---
updated-dependencies:
- dependency-name: "@vitest/coverage-v8"
  dependency-version: 4.1.6
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:55:24 +01:00
6210d80f1e chore(deps)(deps-dev): bump tsx from 4.21.0 to 4.21.1 in /gitnexus (#1698)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.21.0 to 4.21.1.
- [Release notes](https://github.com/privatenumber/tsx/releases)
- [Changelog](https://github.com/privatenumber/tsx/blob/master/release.config.cjs)
- [Commits](https://github.com/privatenumber/tsx/compare/v4.21.0...v4.21.1)

---
updated-dependencies:
- dependency-name: tsx
  dependency-version: 4.21.1
  dependency-type: direct:development
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:55:04 +01:00
637cfca39c chore(deps)(deps): bump express-rate-limit in /gitnexus (#1697)
Bumps [express-rate-limit](https://github.com/express-rate-limit/express-rate-limit) from 8.5.1 to 8.5.2.
- [Release notes](https://github.com/express-rate-limit/express-rate-limit/releases)
- [Commits](https://github.com/express-rate-limit/express-rate-limit/compare/v8.5.1...v8.5.2)

---
updated-dependencies:
- dependency-name: express-rate-limit
  dependency-version: 8.5.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Abhigyan Patwari <126312502+abhigyanpatwari@users.noreply.github.com>
2026-05-19 06:54:33 +01:00
dependabot[bot]andGergő Magyar 73543a4714 chore(deps)(deps-dev): bump @types/node in /gitnexus (#1696)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 25.6.2 to 25.7.0.
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node)

---
updated-dependencies:
- dependency-name: "@types/node"
  dependency-version: 25.7.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-19 06:54:14 +01:00
DuduPhudu b37974fdac feat(javascript): migrate JavaScript to scope-based resolution (RFC #909 Ring 3, issue #928) (#1640) 2026-05-19 06:23:13 +01:00
dependabot[bot] ade2069633 chore(deps)(deps): bump brace-expansion from 5.0.5 to 5.0.6 in /gitnexus (#1689) 2026-05-19 05:35:27 +01:00
azizur100389 5f0c0eba0e feat(cpp): Expand type_traits constraint registry (#1648) 2026-05-18 21:10:18 +01:00
Gergő Magyar 2632bcccc0 fix(api): open lbug read-only for /api/graph, /api/search, /api/grep (#1686) 2026-05-18 19:57:35 +01:00
Gergő Magyar c9199b654f fix(test): retry Windows temp cleanup in cli-e2e teardown (#1688) 2026-05-18 18:17:54 +01:00
Shane Thurston Wijaya 33f18ceaa2 feat(eval-server): added --host for user configured host IP instead of system hardcoded IP (127.0.0.1) (#1667)
* feat(eval-server): added --host for user configured host IP instead of system hardcoded IP (127.0.0.1)

* fix(eval-server): localhost value in --host now returns 127.0.0.1 instead of the raw input to fix wrong address, handled error for ipv6 disabled containers

* feat(eval-server): add --host flag with validation and error handling

  Co-Authored-By: Val Vladescu <val.vladescu@thirdbridge.com>

* fix(eval-server): bracketed IPv6 addresses to remove ambiguity

* docs(eval-server): document --host flag, READY signal format, and parser migration note

* fix(eval-server): use actual bound port in READY signal; strengthen --host e2e tests

  Co-Authored-By: Val Vladescu <val.vladescu@thirdbridge.com>

* feat(eval): wire eval-server --host through gitnexus_docker.py

* docs(eval): added guidance for docker user

* docs(eval): revise the imprecise documentation

* fix(e2e): updated original stdout for new format
2026-05-18 16:00:42 +01:00
c30833fad3 perf(scope-resolution): use owner-keyed lookup for Step 2 member resolution (#1657)
* perf(scope-resolution): use owner-keyed lookup for Step 2 member resolution (#1656)

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(scope-resolution): index Const/Static in FieldRegistry for Step 2 lookup

Extend FieldRegistry to hold multiple defs per (owner, name), reconcile Const and Static into the owner-keyed index, and wire lookupAllByOwner through the production hook so Step 2 does not drop field kinds the registry never indexed. Pass explicitReceiver on read/write reference sites and document undefined-vs-empty hook semantics for defs fallback.

Co-authored-by: Cursor <cursoragent@cursor.com>

* perf(scope-resolution): centralize O(1) owned-member hook and guard hot path

Extract lookupOwnedMembersByOwner for the production Step 2 hook so merges stay O(1) per registry with no defs.byId scan. Add a perf-contract unit test that throws if byId.values runs when the hook is wired. Reuse a frozen empty sentinel on double miss to avoid per-probe allocations.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: drop unused buildFieldRegistry import

* chore(scope-resolution): apply ce-code-review safe_auto fixes

- Drop unreachable return + unused values() capture in perf-contract trap (Finding #7)
- Type lookupOwnedMembersByOwner ownerDefId as DefId (Finding #9)
- Add Static-kind Step 2 lookup test mirroring the Const case (Finding #11)

* docs(field-registry): document lookupFieldByOwner first-wins semantics

Audit of all 6 production callers (call-processor.ts:2279, walkers.ts:535,
receiver-bound-calls.ts:380+730, type-env.ts:627+631) confirms none depends
on last-wins precedence — all treat the return as a generic 'field with
this name owned by this class'. Clarify the JSDoc to surface the semantic
change introduced when FieldRegistry moved from last-wins to append-order
storage (ce-code-review finding #2).

* test(scope-resolution): extend Step 2 perf contract to implicit-self, MRO, field paths

Adds three sibling tests under the Step 2 perf contract describe block, each
asserting defs.byId.values() does NOT execute when ownedMembersByOwner is wired:

- implicit-self receiver via typeBindings.self (no explicitReceiver branch)
- 2-level MRO chain (Child extends Parent, save resolves on Parent at depth 1)
- FieldRegistry read via Step 2 (property lookup, separate registry path)

Pins the perf invariant on every distinct entry into walkReceiverTypeBinding
so a regression bypassing the hook on any sub-path now fails CI immediately
(ce-code-review finding #8).

* test(resolve-references): cover arity-overload filtering via resolveReferenceSites

Pins the orchestration-layer wiring of providers.arityCompatibility:
hook returns [save(arity 1), save(arity 2)], referenceSite.arity = 1,
arityCompatibility verdicts 'compatible'/'incompatible' by parameterCount,
exactly one reference emitted with toDef = the arity-1 overload.

registries.test.ts already covered arity at the buildMethodRegistry level;
this adds the missing entry-point check that resolveReferenceSites threads
providers correctly through to lookupCore.Step5 (ce-code-review finding #10).

* test(resolve-references): add hook-on vs hook-off parity test

Runs resolveReferenceSites twice on the same fixture (Parent.save method
hit + Child.name field hit, Child extends Parent MRO chain) — once with
ownedMembersByOwner wired to a synthetic registry, once with the hook
absent so collectOwnedMembers takes the defs.byId fallback. Asserts:

- stats are identical (sitesProcessed / referencesEmitted / unresolved)
- referenceIndex.bySourceScope entries have equal length
- toDef sets are equal
- each per-site reference (including evidence and depth) is .toEqual

Locks the semantic-parity claim in code while both paths still exist.
Will be removed alongside the fallback in finding #1 (ce-code-review #3).

* test(typescript): probe Step 2 MRO walk against ambient (declare class) base

Adds typescript-ambient-base-class fixture with an export declare class
AmbientBase + Derived extends AmbientBase and a call site d.ambientMethod().
Integration assertions:

- Both classes are detected
- EXTENDS edge Derived → AmbientBase emitted
- CALLS edge to ambient.ts:ambientMethod resolved via MRO walk

Probes the ce-code-review #6 concern that ambient-only owners (whose
bodies are never parsed) might be silently skipped by Step 2 after the
owner-keyed lookup change. Result: the call resolves correctly — the
method signature inside the declare class body still flows through
reconcileOwnership into model.methods, so the hook returns the right
ancestor hits. Residual risk is empirically closed.

* feat(scope-resolution): route nested types via owner-keyed TypeRegistry

Closes the Step 2 contract footgun where 'hook returns [] = authoritative
miss' silently dropped any owned def whose NodeLabel was outside the
method/field if-chain in reconcileOwnership.

- TypeRegistry: add nestedByOwner Map + lookupAllByOwner(owner, simple)
  + registerByOwner(owner, simple, def). Mirrors MethodRegistry/
  FieldRegistry shape; cleared with the rest on cascade clear.
- reconcileOwnership: route class-like NodeLabels (Class/Interface/Enum/
  Struct/Union/Trait/TypeAlias/Typedef/Record/Delegate/Annotation/
  Template/Namespace) via types.registerByOwner. New nestedTypesRegistered
  stat. Idempotent skip via nodeId match.
- validateOwnershipParity: extend the I9 invariant check to nested types.
- lookupOwnedMembersByOwner: merge methods + fields + nested-type hits;
  short-circuit when any one source contributes the full result.

Unblocks future receiver-MRO registries that need to resolve 'Outer.Inner'
through the receiver's type-binding chain (ce-code-review finding #5a).

* refactor(scope-resolution): make ownedMembersByOwner required; delete byId fallback

Per ce-code-review finding #1, the optional-hook design encoded a silent
O(|defs|) perf cliff into the type system: any RegistryContext built
without the hook regressed Step 2 to scanning every def per probe with
no warning. Production wires the hook unconditionally; the fallback was
exercised only by tests.

- RegistryContext.ownedMembersByOwner: required, returns readonly
  SymbolDefinition[] (no | undefined). Implementations MUST return [] on
  authoritative miss.
- collectOwnedMembers in lookup-core.ts collapses to a one-line forward
  to the hook; the defs.byId.values() scan and simpleNameOf helper are
  deleted (simpleNameOf had no other consumers).
- ResolveReferencesInput.ownedMembersByOwner: required to match.
- Tests: drop three fallback-path tests (registries Const fallback,
  resolveReferenceSites no-hook fallback, resolveReferenceSites Const-
  undefined fallback) and the hook-vs-fallback parity test added by
  finding #3. makeCtx in registries.test.ts now defaults to a real
  owner-keyed scan over the test fixture defs so tests that don't care
  about the hook keep working.

* perf(free-call-fallback): cache global callables by simple name once per pass

pickUniqueGlobalCallable scanned scopes.defs.byId.values() on every
free-call fallback site. After PR #1656 fixed Step 2, this scan became
the dominant remaining O(|defs|) hot path on large repos (ce-code-review
finding #4).

- buildGlobalCallableIndex builds a Map<simpleName, SymbolDefinition[]>
  over scopes.defs once at the top of emitFreeCallFallback. Same filter
  the per-site scan applied: Function / Method / Constructor, keyed by
  the last .-segment of qualifiedName.
- pickUniqueGlobalCallable consumes the prebuilt index via O(1) Map.get
  instead of iterating every def. Per-site complexity drops from
  O(|defs|) to O(|defs with this simple name|).
- Cost: O(|defs|) once per pass instead of O(|defs| * |free-call sites|).

Subsequent narrowing (arity, conversion-rank) and the model-side fallback
(model.symbols.lookupCallableByName + model.methods.lookupMethodByName)
are unchanged.

* chore(autofix): apply prettier + eslint fixes via /autofix command

* ci: trigger build

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Gergo Magyar <abhigyan1.patwari@gmail.com>
2026-05-18 13:14:27 +01:00
Copilot 7d500390b9 fix: Use Ladybug native read-only enforcement and prepared statement execution for Cypher query paths (#1655) 2026-05-18 06:54:24 +01:00
Nilotpal Kashyap bdc0439a10 feat(detect-changes): support git worktrees (#1654) 2026-05-17 20:54:41 +01:00
Shane Thurston Wijaya 105efd0f7c feat(wiki): added --lang <lang> flags to gitnexus wiki for multilanguage wiki generation support (#1613) 2026-05-17 19:54:02 +01:00
Copilot 493827222d fix(ingestion): Raise analyze auto-heap to 16GB and tighten cross-platform OOM guidance for UE5-scale repositories (#1652) 2026-05-17 16:28:07 +01:00
Copilot ed50a6729f fix(wiki): Remove the hidden 60s default timeout, validate gitnexus wiki timeout/retry flags, and surface timeout errors (#1651) 2026-05-17 12:03:54 +01:00
Nilotpal Kashyap dfbe68ad24 fix(lbug): issue #1647, detect WAL corruption in schema init and surface recovery (#1650) 2026-05-17 10:46:45 +01:00
azizur100389andGergő Magyar 2376912ca7 feat(ingestion): Add C++ parameter type class sidecar (#1642)
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-16 21:44:26 +01:00
Zander Raycraft a4dfebd073 feat(cpp): sfinae filter (#1623)
* feat(cpp): SFINAE-aware overload filter — drops candidates whose enable_if_t / requires constraints fail (#1579)

* fix(cpp):  SFINAE follow-ups for is_integral_v/is_arithmetic_v bool and char support, an unqualified F1 test fixture, and parameter-lookup gap documentation (#1579) -> claude feedback

* revert: reverting all changes to .md files
2026-05-16 20:23:13 +01:00
Gergő Magyar 42d4fcaf6f chore: release v1.6.5 (#1645) 2026-05-16 17:11:25 +01:00
a26ac55fb0 fix(lbug): Recover gitnexus analyze from orphan LadybugDB sidecars when main DB file is missing (#1622)
* Initial plan

* fix: recover from orphan lbug sidecars on init

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6e8ea6e8-f9ab-46ff-9c1b-4d2c73a6452c

* test: strengthen orphan sidecar recovery coverage

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6e8ea6e8-f9ab-46ff-9c1b-4d2c73a6452c

* fix(lbug): only clean orphan sidecars when DB is missing

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a34217b5-0e98-4949-bae1-2a50933f291e

* test(lbug): cover no-cleanup path when db file exists

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a34217b5-0e98-4949-bae1-2a50933f291e

* test(lbug): use errno-shaped ENOENT mocks for sidecar recovery

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a34217b5-0e98-4949-bae1-2a50933f291e

* test(lbug): cover partial sidecar and unlink-failure recovery cases

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a34217b5-0e98-4949-bae1-2a50933f291e

* refactor(lbug): tighten ENOENT detection and test naming

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a34217b5-0e98-4949-bae1-2a50933f291e

* test(lbug): normalize errno mock helpers across sidecar tests

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a34217b5-0e98-4949-bae1-2a50933f291e

* docs(lbug): annotate orphan `.wal.checkpoint` cleanup provenance

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7edf5156-43e0-412d-87a4-bf4b2934deac

* test(lbug): clarify unlink-failure path test intent

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7edf5156-43e0-412d-87a4-bf4b2934deac

* fix(lbug): handle orphan-sidecar cleanup error paths explicitly

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/93294b2c-57f6-459c-8eb2-86e3b8920fb0

* refactor(lbug): extract errno and error-summary helpers

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/93294b2c-57f6-459c-8eb2-86e3b8920fb0

* test(lbug): expand non-ENOENT lstat coverage and remove magic number

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/93294b2c-57f6-459c-8eb2-86e3b8920fb0

* test(lbug): add native integration test for orphan sidecar recovery

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2dd28264-4604-430a-a249-af52afd29245

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(lbug): annotate best-effort catch in integration test cleanup

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2dd28264-4604-430a-a249-af52afd29245

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(lbug): add cross-process init lock for orphan sidecar cleanup with integration tests

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e4cbcfec-a252-449d-8d65-2f3570a253f8

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor(lbug): use INIT_LOCK_STALE_MS in stale lock detection and address review feedback

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e4cbcfec-a252-449d-8d65-2f3570a253f8

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* style(lbug): fix Prettier line-length violation in acquireInitLock fs.open call

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a140b567-0e9b-4ec9-a158-9fe6b8685ec2

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(lbug): ensure parent directory exists before creating init lock file

acquireInitLock tried to create `${dbPath}.init.lock` using O_CREAT | O_EXCL,
but on a fresh repo the parent directory (`.gitnexus/`) doesn't exist yet —
the mkdir call was inside the locked section. This caused ENOENT failures
on all platforms (Windows, macOS, Ubuntu) during `gitnexus analyze`.

Move mkdir to before the lock file creation attempt.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6883dc3c-36eb-4907-bcd8-61d23e2c641a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(lbug): verify acquireInitLock succeeds when parent directory does not exist

Adds an integration test proving the fix from the previous commit:
acquireInitLock now creates the parent directory before attempting
to create the lock file, preventing ENOENT on fresh repos.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6883dc3c-36eb-4907-bcd8-61d23e2c641a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-16 11:45:32 +01:00
azizur100389andGergő Magyar 467c14caa2 feat(cpp): standard-conversion-sequence ranking for overload resolution (#1606)
* feat(cpp): add standard-conversion-sequence ranking to overload resolution (#1578)

Introduce `ConversionRankFn` abstraction and `cppConversionRank` implementation
to disambiguate C++ overloaded calls by argument-to-parameter conversion cost.
Exact type match (rank 0) beats standard arithmetic conversion (rank 2), which
beats non-viable mismatch (Infinity). Thread the rank function through
`narrowOverloadCandidates`, `pickImplicitThisOverload`, `pickOverload`, and
`pickUniqueGlobalCallable` via the `ScopeResolver.conversionRankFn` contract.
Add `findAllCallableBindingsInScope` scope walker for collecting all overloads
at the first binding scope. Guard against false ambiguity suppression when
candidates span different files (local-shadows-import preservation).

* fix: address Claude review findings on conversion-rank PR

Finding 1 (HIGH): add tests that exercise the conversion ranker.
  - p('a') with p(int)/p(double): char→int promotion (rank 1) beats
    char→double conversion (rank 2), forcing step 4b in
    narrowOverloadCandidates. Exact-type filter misses both overloads.
  - h(42, 2.5) with h(int,int)/h(double,double): multi-arg tied total
    score forces the ranker, both candidates score 2 → suppressed.

Finding 2 (HIGH): unify multi-candidate suppression across all paths.
  - Non-ADL free-call: suppress when narrowed.length > 1 (same-file
    guard), mirroring ADL merged-candidate behavior.
  - ADL ordinary-only: same pattern.
  - pickOverload: return OVERLOAD_AMBIGUOUS when candidates.length > 1
    after normalized-ambiguity check.
  - Case 0.5 (this receiver): set ambiguous=true when narrowed > 1.

Finding 3+4 (MEDIUM): implement rank-1 integral promotions.
  - char→int and bool→int now return rank 1 (ISO C++ [conv.prom]).
  - Updated comment to remove misleading ISO table header; document
    only the post-normalization ranking that is actually implemented.
  - Updated ConversionRankFn JSDoc in overload-narrowing.ts.

218/218 C++ tests pass (registry-primary). Legacy: 186+32.

* fix: implement pairwise dominance comparison for overload ranking

Replace the summed per-slot conversion cost with ISO C++-aligned
pairwise dominance comparison ([over.ics.rank]). F1 is better than
F2 only when F1 is not worse for every argument and strictly better
for at least one. Non-dominated candidates are returned; if multiple
remain they are genuinely ambiguous.

This fixes false CALLS edges for asymmetric multi-arg overloads:
h('a', 2.5) against h(int,int) / h(double,double) — the old summed
cost picked h(double,double) (cost 2 < 3), but ISO C++ considers
the call ambiguous because h(int,int) is better at arg 0 via char
promotion. The pairwise check correctly finds neither dominates.

Add h('a', 2.5) test case asserting zero CALLS edges alongside
the existing h(42, 2.5) symmetric-tie test.

218/218 C++ tests pass (registry-primary). Legacy: 186+32.

* docs: update step 4b JSDoc to reflect pairwise dominance

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-16 11:15:21 +01:00
fa06c5610b fix: resolve cross-file type propagation stall on large repos (#1626)
* Initial plan

* fix: add time-based deadline to cross-file type propagation to prevent stalling on large repos

Adds a 2-minute wall-clock time limit (DEFAULT_CROSS_FILE_ELAPSED_MS) to
runCrossFileBindingPropagation. When exceeded, the phase gracefully stops
and logs a warning. Users can override via GITNEXUS_CROSS_FILE_TIMEOUT_MS
env var. This prevents the analyze command from stalling for hours on very
large repositories where per-file re-resolution is expensive.

Fixes the reported issue where gitnexus analyze stalls at "Cross-file type
propagation" for several hours on repos with 15000+ files.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b8341947-557c-4111-a3a8-991ba455ab01

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: root cause - cache tree-sitter queries across files, add live progress reporting

Root cause: cross-file propagation called processCalls() with 1 file at a time,
causing Parser.Query to be recompiled from the query string for every single file
(O(N) compilations vs O(1) for the whole phase). Additionally, progress was only
reported once at the start, making the phase appear completely frozen.

Fixes:
- Add optional `compiledQueryCache` parameter to `processCalls` so callers that
  invoke it with single-file batches can share compiled query objects across calls.
  The cross-file phase now compiles each language's query string exactly once and
  reuses it for all files of that language (e.g. 1 TypeScript compile for 595+ files).
- Pre-count candidate files and emit onProgress every 25 files showing
  "Cross-file type propagation (N/M files)..." so the UI shows real movement
  instead of a frozen bar.
- Keep the wall-clock deadline (GITNEXUS_CROSS_FILE_TIMEOUT_MS) as a safety
  net for pathological inputs.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/f5028cc8-4bc9-4309-8ffb-798fe2bd7a0a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: address code review - use SupportedLanguages key type, rename queryCache to compiledQueryCache

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/f5028cc8-4bc9-4309-8ffb-798fe2bd7a0a

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(cross-file): remove wall-clock timeout from type propagation

The query compilation cache and live progress reporting address the
original stall; the 2-minute deadline could truncate cross-file work on
large repos. MAX_CROSS_FILE_REPROCESS (2000) remains as the only cap.

* test(cross-file): verify compiledQueryCache is shared across all processCalls invocations

Finding 1: O(N) query recompilation was fixed by sharing a compiledQueryCache Map
across all processCalls invocations in runCrossFileBindingPropagation. This test
verifies the fix is correctly wired: the same Map instance is passed as the
12th argument to every call, proving queries are compiled once per language,
not once per file.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3ab768d9-3993-4882-9d8f-17f7fcbd086e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cross-file): verify live progress events are emitted with N/M format

Finding 2: frozen progress display was fixed by emitting onProgress every 25 files
with "Cross-file type propagation (N/M files)..." messages instead of calling it
once at phase start. This test verifies the fix with 50 candidate files: expects
onProgress called 3 times (1 initial + at 25 + at 50) with correct N/M counters.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3ab768d9-3993-4882-9d8f-17f7fcbd086e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(cross-file): skip registry-primary language files before readFileContents

Finding 3 (from comment 4466231612): cross-file-impl was calling processCalls
for every candidate file even when that file's language is registry-primary
(TypeScript, C++, Python, Go, C#, PHP, C — since AGENTS.md v1.7.0). processCalls
would immediately skip those files via its own isRegistryPrimary guard, but
cross-file-impl still paid the full cost: readFileContents I/O, buildImportedReturnTypes,
buildImportedRawReturnTypes, and Map allocation — all discarded.

Fix: check isRegistryPrimary(lang) in both the totalCandidates pre-count loop
and the levelCandidates builder, before any file I/O or map building. This
eliminates 595+ no-op processCalls invocations on large TypeScript repos.

Test: mocks isRegistryPrimary to always return true and verifies that
processCalls is never invoked and result is 0. The mock also defaults to false
in beforeEach so existing tests using .ts files are unaffected.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3ab768d9-3993-4882-9d8f-17f7fcbd086e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* refactor(test): address code review - simplify mock factory, name the arg index constant

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/3ab768d9-3993-4882-9d8f-17f7fcbd086e

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-16 10:02:40 +01:00
Gergő Magyar f28185d67e fix(ci): bump publish job to Node 24 for npm OIDC support (#1628)
PR #1627's npm install -g npm@latest step crashed mid-install with MODULE_NOT_FOUND: promise-retry — a known fragility when npm self-upgrades. Node 22's bundled npm is 10.9.x (no OIDC). Fix: bump publish job's node-version to 24, which ships with npm 11.x natively. Package consumers unaffected (this Node version is only used during publish; engines.node is >=22.0.0; ci-tests.yml continues testing on Node 22).
2026-05-16 09:18:08 +01:00
Gergő Magyar f69c382bcb fix(ci): engage npm Trusted Publishing OIDC properly (#1627)
First live-fire RC publish after #1610 failed at npm publish with E404. The if: failure() cleanup correctly auto-deleted the partial v-tag and rc-marker, but OIDC never engaged. Root cause: two coordinated upstream bugs.

1. actions/setup-node@v6 with registry-url: writes _authToken into the runner .npmrc AND exports NODE_AUTH_TOKEN from its token: input (defaulting to github.token). npm publish sends GITHUB_TOKEN as the bearer and the registry returns 404. OIDC never tried because npm thinks it already has a credential. See actions/setup-node#1440.

2. The Node 22 runner ships with npm 10.9.x. npm Trusted Publishing OIDC support requires npm >= 11.5.1.

Fix: omit registry-url: from the setup-node step (per the consensus workaround in community discussion #176761), and add npm install -g npm@latest before publish. --provenance flag is NOT added; npm auto-attaches provenance under Trusted Publishing.

Sources:
- https://github.com/actions/setup-node/issues/1440
- https://github.com/orgs/community/discussions/176761
- https://docs.npmjs.com/trusted-publishers/
2026-05-16 08:36:49 +01:00
Gergő Magyar 83fbd4be26 refactor(ci): unify release pipeline under publish.yml (#1610)
Collapse release-candidate.yml into publish.yml so there is exactly one workflow that publishes gitnexus to npm, creates GitHub Releases, and triggers Docker builds — for both release candidates and stable releases. Closes #1609 architecturally.

A first-stage `route` job classifies push-to-main / push-tag / workflow_dispatch into `rc` / `stable` modes and fails closed on malformed shapes. RC path runs rc-guard → ci.yml → publish (mint GitHub App token → checkout with persist-credentials:false → resolve next rc version → atomic v-tag + rc/<SHA> marker push → vtag integrity gate → npm publish via OIDC → GitHub prerelease → if: failure() cleanup) → docker.yml. Stable path verifies package.json matches the tag and publishes to `latest` via OIDC (no docker).

Hardening:

  • Self-trigger prevention via negative-glob `tags: ['v*', '!v*-rc.*']` — the bug class behind #1609 cannot recur.
  • Two distinct actions/checkout steps per mode (no conditional `token:` expression footgun).
  • Workflow-level `permissions: {}` deny-all + per-job grants; `id-token: write` only where OIDC is used.
  • npm Trusted Publishing replaces NPM_TOKEN (delete the secret after the first successful publish).
  • GitHub App installation token (actions/create-github-app-token@v3.2.0) replaces the long-lived RELEASE_PUSH_TOKEN PAT (delete after first successful RC).
  • vtag integrity gate fails closed on empty / mode-mismatched output (prevents Release named `main` from a github.ref fallback).
  • Annotation-injection sanitization on every logged ref.
  • Explicit `secrets:` passthrough on docker.yml (DOCKERHUB_USERNAME, DOCKERHUB_TOKEN); ci.yml no longer inherits anything.
  • `if: failure()` cleanup auto-deletes v-tag + rc-marker on partial failure (eliminates the external-consumer phantom-version ingestion window).
  • ACTIONS_STEP_DEBUG window closed via `set +x` wrap on the inline auth-header compute.
  • Curated retry-loud error handling on `gh api` bot-user-id lookup and `npx semver`.

Pre-merge validation:

  • 10-reviewer multi-agent code-review pass; 14 findings fixed inline (commit 820cefae), 6 deferred to follow-ups.
  • End-to-end dry-run rehearsal via workflow_dispatch (run 25919563064) validated route classification, rc-guard, App token mint, RC checkout, version resolver, vtag synthetic-regex check, and faithful tarball pack at the bumped version.
  • All zizmor findings on the unification commits closed.
  • Branch-protection required checks all green.

Post-merge actions:

  • After the first successful RC, delete the `NPM_TOKEN` and `RELEASE_PUSH_TOKEN` secrets — they are no longer used.
  • The first real RC after merge is the live-fire test for steps dry-run could not exercise (atomic tag push, real npm OIDC handshake, GitHub Release creation, docker.yml under explicit secrets passthrough). The if: failure() cleanup step handles the partial-failure recovery automatically; the Rollback Runbook in CONTRIBUTING.md covers the rare cases auto-cleanup can't reach.
2026-05-16 07:46:56 +01:00
263ca353a6 fix: shard parse cache persistence on large repos (#1580)
* fix: shard parse cache persistence on large repos

* fix(parse-cache): validate shard keys, docs, and sharded-cache tests

- Reject non-sha256-hex keys from index.json before path.join (path traversal).

- saveParseCache: skip invalid keys defensively; try/catch per-shard JSON.stringify.

- Clarify save comment (tmp dir + rename vs atomic).

- Tests: hex keys throughout, traversal keys, multi-shard, version-mismatch+legacy, second save, legacy removal.

- AGENTS.md / GUARDRAILS.md: document .gitnexus/parse-cache/ vs legacy parse-cache.json.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

---------

Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-16 07:19:06 +01:00
azizur100389 8500f18e5f fix(cpp): detect same-name ambiguity across inline namespace children (#1564) (#1600) 2026-05-15 19:11:09 +01:00
Copilotandmagyargergo aed370b931 feat: C++ ADL V2: merge ordinary and ADL free-call candidates before overload selection (#1599)
* Initial plan

* Merge C++ ADL and ordinary free-call candidate sets

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/1aea3511-3471-4ec2-9819-0fb27ac40b89

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* Address review feedback on merged ADL ambiguity suppression

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/1aea3511-3471-4ec2-9819-0fb27ac40b89

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: apply prettier to C++ ADL resolver fallback files

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9b9c1494-bc69-4db5-a89d-69eb816bab82

* docs: update ADL ambiguity comments to merged narrowing flow

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9b9c1494-bc69-4db5-a89d-69eb816bab82

* fix: suppress global fallback when merged ADL narrowing yields zero candidates

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9b9c1494-bc69-4db5-a89d-69eb816bab82

* docs: clarify free-call fallback comment for ADL merged path

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/9b9c1494-bc69-4db5-a89d-69eb816bab82

* feat: ADL Gap 2 — enum-typed arguments contribute enclosing namespace

ISO C++ [basic.lookup.argdep] §2: "If T is an enumeration type, its
associated namespace is the namespace in which it is defined."

- Add Enum to findCppClassDefBySimpleName type filter
- Map Enum defs to enclosing namespace in populateCppAssociatedNamespaces
- Add test fixture cpp-adl-enum-arg with color::Channel enum

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/ca8987b6-365e-4034-af56-ca3f9b439902

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat: ADL Gap 6 — inline namespace expansion in associated set

ISO C++ inline namespaces are transparent for ADL: if a namespace is
in the associated set, candidates declared in its inline-namespace
children are also reachable.

- Expand pickCppAdlCandidates to scan inline-namespace children of
  associated namespaces (via isCppInlineNamespaceScope predicate)
- Add test fixture cpp-adl-inline-ns-expansion: Event in outer audit,
  record in inline v1, other::record(int) forces arity disambiguation

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/ca8987b6-365e-4034-af56-ca3f9b439902

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat: ADL Gap 1 — hidden friend functions visible via ADL

ISO C++ [basic.lookup.argdep] §2: friend functions declared inside a
class body are visible via ADL when the class is an associated class.

- Exempt friend_declaration from cppLabelOverride's class-body function
  suppression (c-cpp.ts) so friend function defs are captured
- Scan Function scopes that are direct children of associated Class
  scopes in pickCppAdlCandidates (adl.ts) to find hidden friends
- Add test fixture cpp-adl-hidden-friend: `friend void process(Foo&)`
  declared inside lib::Foo, resolved via ADL from app::run()

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/ca8987b6-365e-4034-af56-ca3f9b439902

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* feat: ADL Gap 3 — non-function ordinary lookup suppresses ADL

ISO C++ [basic.lookup.unqual] §7: if ordinary unqualified lookup finds
a name that is not a function or function template, ADL is not performed.

- Add hasNonCallableBindingInScope walker in walkers.ts
- In free-call-fallback, check for non-callable binding before invoking
  ADL; when found, bypass resolveAdlCandidates entirely
- Add test fixture cpp-adl-non-function-blocks: variable `int record`
  shadows the function name, blocking ADL from finding audit::record

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/ca8987b6-365e-4034-af56-ca3f9b439902

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: use nearest-scope semantics for ADL non-callable blocker check

Finding 1: `hasNonCallableBindingInScope` walked the entire scope chain,
which could incorrectly suppress ADL when an inner scope had a callable
and an outer scope had a non-callable for the same name. Per ISO C++
`[basic.lookup.unqual]` §7, ADL is blocked only when ordinary lookup
itself finds a non-function — if ordinary lookup stops at an inner scope
where only callables exist, ADL should still fire.

Replace the separate `hasNonCallableBindingInScope` + `findAllCallable
BindingsInScope` calls with a combined `findCallableBindingsAndAdlBlocker`
walker that stops at the first scope with ANY binding for the name and
returns both `{ callables, nonCallableFound }`. One pass, one stop.

Fixture: cpp-adl-inner-callable-outer-noncallable — inner scope has
callable `swap(int,int)`, outer scope has `int swap = 0`. ADL fires and
resolves to `data::swap(Pair&,Pair&)` via argTypes narrowing.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a2f97daf-17fd-4891-8b10-a81e44d32808

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix: block-scope function declaration suppresses ADL

Finding 2: ISO C++ [basic.lookup.argdep] lists three ADL blockers:
1. class member declaration (handled by pickImplicitThisOverload)
2. block-scope function declaration NOT a using-declaration (NEW)
3. non-function/non-template declaration (handled by nonCallableFound)

Extend `findCallableBindingsAndAdlBlocker` to return `blockScopeDeclFound`
when a callable is found at a Function or Block scope — indicating a local
forward declaration that should suppress ADL per standard.

`free-call-fallback.ts` now checks both `nonCallableFound` and
`blockScopeDeclFound` to determine ADL suppression.

Fixture: cpp-adl-block-scope-decl-blocks — `void record(int);` declared
inside function body prevents ADL from discovering audit::record.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a2f97daf-17fd-4891-8b10-a81e44d32808

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* docs: update stale ADL_AMBIGUOUS comment in unqualified-ref-collision fixture

Finding 3: The `ADL_AMBIGUOUS` sentinel was removed by this PR (replaced
by `isOverloadAmbiguousAfterNormalization` in merged-narrowing). Update
the fixture comment to reference the current mechanism.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a2f97daf-17fd-4891-8b10-a81e44d32808

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test: add legacy-parity expected failures for ADL blocker tests

The new ADL nearest-scope blocker and block-scope function declaration
tests rely on scope-resolution-only mechanisms not present in the legacy
DAG path. Register them in LEGACY_RESOLVER_PARITY_EXPECTED_FAILURES.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a2f97daf-17fd-4891-8b10-a81e44d32808

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore: revert unrelated prettier-plugin-tailwindcss devDep addition

The `prettier-plugin-tailwindcss` dependency was accidentally added while
running local prettier; it is not needed for the C++ ADL changes.

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/a2f97daf-17fd-4891-8b10-a81e44d32808

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
2026-05-15 15:41:14 +01:00
99b8c7b03b feat: C++ ADL V2: free-function reference args contribute enclosing namespace (#1598)
* Initial plan

* cpp ADL V2: free-function reference args contribute enclosing namespace

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/24805583-c0c4-4ef8-978f-b874bd917947

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* merge: resolve conflicts with origin/main and fix overloaded fixture app.cpp

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/136aeffe-45da-47e2-95dd-e3883e85fad7

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(finding-1): replace ISO C++ [basic.lookup.argdep] misstatement with GitNexus-approximation label

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fa37c0dd-65f9-4dd4-9811-617227a37073

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(finding-2): verify Function/Method exists in namespace before contributing via qualified_identifier arg

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fa37c0dd-65f9-4dd4-9811-617227a37073

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(finding-3): function parameters in parameter_list no longer misclassified as free-function refs

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fa37c0dd-65f9-4dd4-9811-617227a37073

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* doc(finding-4): document typedef/using-aliased function-pointer limitation in lookupAdlIdentifierType

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fa37c0dd-65f9-4dd4-9811-617227a37073

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(finding-5): add negative fixtures for local-fp shadowing free-func and unqualified namespace collision

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/fa37c0dd-65f9-4dd4-9811-617227a37073

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(legacy-parity): skip two new negative-fixture tests from legacy DAG parity run

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/64ffbaf1-f442-4a2b-8542-4afa500d9182

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-15 10:55:44 +01:00
813acd7ec5 feat: C++ ADL V2: include base-class associated namespaces via MRO (#1597)
* Initial plan

* fix(cpp): include base-class namespaces in ADL candidate selection

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7df6f692-1af9-43e6-82de-099ed43a60cb

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): clarify ADL base-namespace test names

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/7df6f692-1af9-43e6-82de-099ed43a60cb

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): remove stale legacy parity expected-failure entry

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/1a55a5e8-ae91-44bc-9b21-9324cdfea3de

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): assert base-namespace ADL tests are not parity skips

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b3726b70-e797-4f37-955d-7d61fd28d338

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(cpp): avoid MRO amplification on ambiguous class-name ADL lookup

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b3726b70-e797-4f37-955d-7d61fd28d338

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): strengthen ADL base-namespace target identity assertions

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b3726b70-e797-4f37-955d-7d61fd28d338

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): add ADL negative cases for anonymous and unresolved bases

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/b3726b70-e797-4f37-955d-7d61fd28d338

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): fix anonymous-base parity expectation and formatting

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/6a2e3cf9-beea-435c-8494-6a7a00af0f1e

* fix(cpp): propagate unnamed-namespace members through #include in registry-primary resolver

Anonymous-namespace contents in a header (e.g. `namespace { void f(); }`)
are reachable by unqualified lookup in any TU that #includes the header
per ISO C++ [basic.namespace.anon]/1 (the unnamed namespace behaves as
if a `using namespace unique;` is inserted into the enclosing scope, with
per-TU `unique`). The registry-primary path was filtering these defs out
of `expandCppWildcardNames` via both the structural Namespace-owner check
and the `isFileLocal` mark, so `hidden_probe(d)` from a TU including the
header resolved to nothing while the legacy DAG returned the correct edge.

Track anonymous-`namespace_definition` source ranges at capture time,
resolve them to ScopeIds in `populateOwners` (parallels inline-namespace
handling), and exempt those scopes from the two wildcard-expansion filters
plus the `populateCppNonGloballyVisible` structural set. `markFileLocal`
is preserved so the global free-call fallback still blocks cross-TU leaks
for files that do NOT #include the declaring file (cpp-anon-ns-cross-file
guard still passes).

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-05-15 07:54:23 +01:00
Copilot 7fbf302018 feat: C++ ADL V2: include template-specialization associated namespaces (with nested template args) (#1596) 2026-05-15 05:04:42 +01:00
dependabot[bot] 5ee1122330 chore(deps)(deps-dev): bump vitest from 4.1.5 to 4.1.6 in /gitnexus (#1605) 2026-05-14 22:28:31 +01:00
Copilot cdac8a691a feat: C++ ADL V2: include class-typed reference args (incl. rvalue refs) in associated-namespace lookup (#1595) 2026-05-14 20:25:12 +01:00
b00ba2ab47 feat(cpp): resolve template-body this-> + using ns::name calls in scope resolver (#1590)
* Initial plan

* fix(cpp): resolve this-> and using-name calls in template bodies

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/d9d91945-f19c-4fd2-9b52-b0ebc9aa34b6

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* fix(cpp): treat duplicate using-name hits as ambiguous

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/d9d91945-f19c-4fd2-9b52-b0ebc9aa34b6

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* chore(autofix): apply prettier + eslint fixes via /autofix command

* fix(cpp): gate this-receiver path and harden overload semantics

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/030a1842-c698-460d-ae2a-95037e6def73

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): add positive this-> overload case and document field shadowing

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/030a1842-c698-460d-ae2a-95037e6def73

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

* test(cpp): skip new template-this assertions in legacy parity lane

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/27002f6e-6331-41e3-8175-9d9e4691927c

Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: magyargergo <11230420+magyargergo@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-14 18:18:36 +01:00
CopilotandGergő Magyar c2193318b5 feat(cpp): Enable C++ ADL for class pointer arguments and exclude function pointers (#1592)
* Initial plan

* fix: unwrap cpp adl pointer argument types

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2e9c8549-e062-410c-9ce3-66ba0a181590

* chore: tighten cpp adl function-pointer guard

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2e9c8549-e062-410c-9ce3-66ba0a181590

* docs: clarify cpp adl implementation comments

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/2e9c8549-e062-410c-9ce3-66ba0a181590

* fix: avoid aborting cpp adl declaration scan

Agent-Logs-Url: https://github.com/abhigyanpatwari/GitNexus/sessions/e54f1d4b-9aac-407c-9b5e-b5f3ea0534ea

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-05-14 17:56:32 +01:00
438 changed files with 38532 additions and 7427 deletions
@@ -17,11 +17,11 @@ npx gitnexus analyze
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates CLAUDE.md / AGENTS.md context files.
| Flag | Effect |
| ------------------- | ------------------------------------------------------------------------------------------------------- |
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
| Flag | Effect |
| -------------- | ---------------------------------------------------------------- |
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
+102
View File
@@ -75,3 +75,105 @@ jobs:
build: 'true'
- run: npx vitest run
working-directory: gitnexus
# End-to-end smoke test for the #1728 packaging fix: pack the published
# tarball, install it globally into a temp prefix, and assert no junction
# creation (the EPERM root cause) plus working CLI plus vendor cleanliness
# (#836). Runs on windows-latest because that is the platform the fix
# targets; the in-repo `npm ci` job above only exercises the dev-tree path
# and skips the tarball reify step where the historical EPERM occurred.
packaged-install-smoke:
name: packaged install smoke (${{ matrix.os }})
strategy:
fail-fast: false
matrix:
os: [windows-latest, ubuntu-latest]
runs-on: ${{ matrix.os }}
timeout-minutes: 15
steps:
# persist-credentials: false — this job runs npm pack + npm install -g
# from a tarball and never pushes back; the token in .git/config would
# be at risk of leaking through any future artifact-upload step
# (zizmor artipacked audit). Disable upfront.
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- uses: ./.github/actions/setup-gitnexus
with:
build: 'true'
- name: Pack gitnexus tarball
shell: bash
run: npm pack
working-directory: gitnexus
- name: Install gitnexus tarball into isolated prefix
shell: bash
run: |
set -euo pipefail
PREFIX="$RUNNER_TEMP/gitnexus-smoke"
mkdir -p "$PREFIX"
TARBALL=$(find . -maxdepth 1 -name 'gitnexus-*.tgz' -print -quit)
if [ -z "$TARBALL" ]; then
echo "ERROR: no gitnexus-*.tgz tarball found in $(pwd)" >&2
exit 1
fi
echo "Installing $TARBALL into $PREFIX"
npm install -g --prefix "$PREFIX" "./$TARBALL" --no-audit --no-fund
echo "PREFIX=$PREFIX" >> "$GITHUB_ENV"
working-directory: gitnexus
- name: Assert no junctions or vendor build artifacts
shell: bash
run: |
set -euo pipefail
# Locate the installed gitnexus package across npm prefix layouts
# (lib/node_modules on POSIX, node_modules on Windows).
for candidate in "$PREFIX/lib/node_modules/gitnexus" "$PREFIX/node_modules/gitnexus"; do
if [ -d "$candidate" ]; then
INSTALLED="$candidate"
break
fi
done
if [ -z "${INSTALLED:-}" ]; then
echo "ERROR: installed gitnexus package not found under $PREFIX" >&2
ls -la "$PREFIX" || true
exit 1
fi
echo "Installed package at: $INSTALLED"
# #836 invariant: no node_modules/ or build/ under any vendor/*.
BAD=$(find "$INSTALLED/vendor" \( -name node_modules -o -name build \) -print 2>/dev/null || true)
if [ -n "$BAD" ]; then
echo "ERROR: vendor tree contains forbidden build artifacts (#836):" >&2
echo "$BAD" >&2
exit 1
fi
# #1728 invariant: materialized grammar dirs are real directories,
# not junctions/symlinks (which is what the EPERM regression created).
for name in tree-sitter-dart tree-sitter-proto tree-sitter-swift; do
entry="$INSTALLED/node_modules/$name"
if [ ! -e "$entry" ]; then
echo "WARN: $name not materialized (toolchain/prebuild may be unavailable on $RUNNER_OS)"
continue
fi
if [ -L "$entry" ]; then
echo "ERROR: $entry is a symlink/junction — #1728 regression" >&2
exit 1
fi
if [ ! -d "$entry" ]; then
echo "ERROR: $entry is not a directory" >&2
exit 1
fi
done
- name: Assert gitnexus --version works
shell: bash
run: |
set -euo pipefail
if [ "$RUNNER_OS" = "Windows" ]; then
"$PREFIX/gitnexus.cmd" --version
else
"$PREFIX/bin/gitnexus" --version
fi
+8 -8
View File
@@ -11,14 +11,14 @@ permissions:
# Concurrency convention: see CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
# Hardcoded `CI-` prefix (not `${{ github.workflow }}`) because this workflow is
# invoked as a reusable workflow from publish.yml and release-candidate.yml. In
# called-workflow context `github.workflow` evaluation is ambiguous across GitHub
# Actions versions, and a prefix that could resolve to the caller's name would
# share a concurrency group with the caller → deadlock. A literal prefix is
# immune. Direct `pull_request` invocations use `CI-<ref>`; invocations from a
# reusable-workflow caller fall into a per-run-unique group that never serializes
# with the caller. `push` to main is handled by release-candidate.yml, which
# calls this workflow once before publishing.
# invoked as a reusable workflow from publish.yml. In called-workflow context
# `github.workflow` evaluation is ambiguous across GitHub Actions versions, and a
# prefix that could resolve to the caller's name would share a concurrency group
# with the caller → deadlock. A literal prefix is immune. Direct `pull_request`
# invocations use `CI-<ref>`; invocations from a reusable-workflow caller fall
# into a per-run-unique group that never serializes with the caller. `push` to
# main is handled by publish.yml (RC mode), which calls this workflow once
# before publishing.
concurrency:
group: ${{ github.event_name == 'pull_request' && format('CI-{0}', github.ref) || format('CI-nested-{0}', github.run_id) }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
+2 -2
View File
@@ -48,7 +48,7 @@ jobs:
persist-credentials: false
- name: Initialize CodeQL
uses: github/codeql-action/init@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/init@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
languages: ${{ matrix.language }}
queries: security-and-quality
@@ -69,6 +69,6 @@ jobs:
- '**/test/fixtures/**'
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/analyze@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
category: '/language:${{ matrix.language }}'
+1 -1
View File
@@ -33,7 +33,7 @@ jobs:
persist-credentials: false
- name: Dependency Review
uses: actions/dependency-review-action@2031cfc080254a8a887f58cffee85186f0e49e48 # v4.9.0
uses: actions/dependency-review-action@a1d282b36b6f3519aa1f3fc636f609c47dddb294 # v5.0.0
with:
fail-on-severity: high
comment-summary-in-pr: on-failure
+10 -1
View File
@@ -25,6 +25,15 @@ on:
a gitnexus/package.json whose version matches the tag.
required: true
type: string
# Explicit secret contract — callers pass these by name. Replaces the
# blanket `secrets: inherit` pattern (zizmor `secrets-inherit` audit).
# GHCR auth uses the implicit GITHUB_TOKEN; only Docker Hub credentials
# need to be passed through.
secrets:
DOCKERHUB_USERNAME:
required: true
DOCKERHUB_TOKEN:
required: true
permissions:
contents: read
@@ -73,7 +82,7 @@ jobs:
steps:
# Only the workflow_call path requires a non-empty `inputs.tag` — callers
# (e.g. release-candidate.yml) must pass the RC tag explicitly. On direct
# (publish.yml in RC mode) must pass the RC tag explicitly. On direct
# tag pushes the tag comes from `github.ref`, so `inputs.tag` is always
# empty and validating it here would break every real release (#1064).
# The downstream "Verify tag matches gitnexus/package.json version" step
+1 -1
View File
@@ -108,7 +108,7 @@ jobs:
# Pinned to v7.2.0. Verify SHA via:
# gh api repos/release-drafter/release-drafter/git/refs/tags/v7.2.0
# v7 removed `disable-releaser`; use `dry-run: true` to only autolabel.
- uses: release-drafter/release-drafter@563bf132657a13ded0b01fcb723c5a58cdd824e2 # v7.2.1
- uses: release-drafter/release-drafter@c2e2804cc59f45f57076a99af580d0fedb697927 # v7.3.0
with:
config-name: release-drafter.yml
dry-run: true
+830 -34
View File
@@ -1,62 +1,421 @@
name: Publish to npm
name: Publish
# ─────────────────────────────────────────────────────────────────────────────
# Sole publisher for the `gitnexus` npm package, GitHub Releases, and Docker
# images. Replaces the former two-workflow design — see issue #1609 for the
# double-publish race this unification closes.
#
# Two release modes, both routed through this file:
# • Release candidate (rc) — triggered by push to `main` or workflow_dispatch.
# The RC path computes the next rc version, applies it in-CI, pushes a
# detached release commit with v<X.Y.Z>-rc.<N> + rc/<SHA> marker
# atomically, then publishes to npm with --tag rc and creates a GitHub
# prerelease. RC-only docker.yml invocation follows.
# • Stable — triggered by push of a v<X.Y.Z> tag (no -rc.*
# suffix). Verifies package.json matches the tag, publishes to npm with
# --tag latest, creates a stable GitHub Release. No docker (RC-only).
#
# ⚠️ SELF-TRIGGER INVARIANT — DO NOT WEAKEN ⚠️
# The `tags:` filter below uses a negative glob `'!v*-rc.*'` to prevent the
# workflow from re-triggering itself when the RC path pushes its own v-tag.
# Without this exclusion, every RC publish double-fires (the bug fixed by
# #1609). If a NEW prerelease channel is introduced (e.g. `-beta.N`,
# `-alpha.N`, `-next.N`), the negative-glob list MUST be extended in
# lock-step or self-trigger returns. The same invariant applies to the
# `Classify` step further below — its accepted-tag regex must align with
# the trigger filter's exclusion list.
# ─────────────────────────────────────────────────────────────────────────────
on:
push:
branches: [main]
paths-ignore:
- '**.md'
- 'docs/**'
- 'LICENSE'
tags:
# Negative-globbed exclusion of RC tags this workflow itself produces
# (see the SELF-TRIGGER INVARIANT in the header comment).
- 'v*'
# No workflow-level permissions — scoped per job below.
- '!v*-rc.*'
workflow_dispatch:
inputs:
bump:
description: >-
Cycle policy. 'auto' (default) continues the active rc cycle on
this branch if there is one, otherwise bumps patch from latest.
Choose 'patch' / 'minor' / 'major' to explicitly start or reset
an rc cycle.
required: false
default: 'auto'
type: choice
options:
- auto
- patch
- minor
- major
force:
description: 'Publish even when HEAD already has an rc marker'
required: false
default: 'false'
type: choice
options:
- 'false'
- 'true'
# Workflow-level deny-all; each job declares the minimum it needs.
permissions: {}
# Concurrency convention: see CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
# Tag refs are unique per release, so distinct tags run in parallel. Re-pushes of the
# same tag serialize. cancel-in-progress: false — never cancel a publish mid-flight.
# Distinct refs (refs/heads/main, refs/tags/v*) run in parallel. The
# release-PR-skip in rc-guard is the load-bearing invariant that prevents
# an RC main-push and a stable tag-push colliding on the same release commit.
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: false
jobs:
# ── Phase 1: classify the triggering event into a release mode ─────────────
route:
name: Classify release event
runs-on: ubuntu-latest
timeout-minutes: 2
permissions:
contents: read
outputs:
mode: ${{ steps.classify.outputs.mode }}
head_sha: ${{ steps.classify.outputs.head_sha }}
bump_input: ${{ inputs.bump }}
force_input: ${{ inputs.force }}
steps:
- name: Classify
id: classify
shell: bash
env:
EVENT_NAME: ${{ github.event_name }}
GH_REF: ${{ github.ref }}
GH_REF_NAME: ${{ github.ref_name }}
run: |
set -euo pipefail
HEAD_SHA="${GITHUB_SHA}"
echo "head_sha=${HEAD_SHA}" >> "$GITHUB_OUTPUT"
# Sanitize before logging (annotation-injection defense in depth).
REF_SAFE="${GH_REF//::/__}"
REF_NAME_SAFE="${GH_REF_NAME//::/__}"
echo "event=${EVENT_NAME} ref=${REF_SAFE} ref_name=${REF_NAME_SAFE}"
MODE=""
case "${EVENT_NAME}" in
workflow_dispatch)
# Manual dispatch is only valid on main — that's the only ref
# where a real publish makes sense.
if [ "${GH_REF}" = "refs/heads/main" ]; then
MODE="rc"
else
echo "::error::workflow_dispatch is only permitted on refs/heads/main (got ${REF_SAFE})."
exit 1
fi
;;
push)
case "${GH_REF}" in
refs/heads/main)
MODE="rc"
;;
refs/tags/v*)
# The trigger filter already excluded v*-rc.* tags. Anything
# reaching here is either a stable semver or a malformed v*.
TAG="${GH_REF#refs/tags/}"
if [[ "${TAG}" =~ ^v[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
MODE="stable"
else
echo "::error::malformed v* tag rejected: ${REF_NAME_SAFE}"
echo "::error::stable tags must match ^v[0-9]+\\.[0-9]+\\.[0-9]+\$"
exit 1
fi
;;
*)
echo "::error::unexpected push ref ${REF_SAFE} reached publish workflow."
exit 1
;;
esac
;;
*)
echo "::error::unsupported event ${EVENT_NAME}."
exit 1
;;
esac
echo "mode=${MODE}" >> "$GITHUB_OUTPUT"
echo "Classified as mode=${MODE}"
# ── Phase 2 (RC only): dedup marker + release-PR skip ──────────────────────
rc-guard:
name: RC guard (marker + release-PR skip)
needs: route
if: needs.route.outputs.mode == 'rc'
runs-on: ubuntu-latest
timeout-minutes: 5
permissions:
contents: read
pull-requests: read
outputs:
should_run: ${{ steps.decide.outputs.should_run }}
head_sha: ${{ steps.decide.outputs.head_sha }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
fetch-depth: 0
fetch-tags: true
# rc-guard reads only — no git pushes from this job. Skip the
# default extraheader credential persistence (artipacked audit).
persist-credentials: false
- name: Decide
id: decide
shell: bash
env:
FORCE: ${{ inputs.force }}
BUMP_INPUT: ${{ inputs.bump }}
EVENT_NAME: ${{ github.event_name }}
GH_TOKEN: ${{ github.token }}
REPO: ${{ github.repository }}
run: |
set -euo pipefail
HEAD_SHA=$(git rev-parse HEAD)
echo "head_sha=$HEAD_SHA" >> "$GITHUB_OUTPUT"
if [ "$FORCE" = "true" ]; then
echo "Force flag set — running regardless of marker tag."
echo "should_run=true" >> "$GITHUB_OUTPUT"
exit 0
fi
# Explicit cycle reset on dispatch bypasses dedup.
if [ "$EVENT_NAME" = "workflow_dispatch" ] \
&& [ -n "${BUMP_INPUT:-}" ] \
&& [ "${BUMP_INPUT:-auto}" != "auto" ]; then
echo "Explicit bump=$BUMP_INPUT — bypassing marker dedup."
echo "should_run=true" >> "$GITHUB_OUTPUT"
exit 0
fi
# ── Skip when the merge commit corresponds to a release ───────────
# This skip is load-bearing: it prevents an RC build firing on the
# release-PR commit from racing the imminent stable-tag push on the
# same SHA. Two complementary checks:
# 1. HEAD subject matches `chore: release vX.Y.Z` (the canonical
# release-PR title). Anchored to require the bare title or the
# squash-merge `(#NNNN)` suffix exactly. Case-insensitive so
# `Chore: Release v1.2.3` (IDE auto-capitalization) still
# matches — prior commit-author conventions left the door open.
# 2. Squash-merged PR carries the `release` label.
# Either match suppresses the rc build — stable releases publish on
# the v-tag instead.
HEAD_SUBJECT="$(git log -1 --pretty=%s HEAD)"
# Sanitize GitHub-Actions annotation prefixes before logging — even
# though %s strips newlines, a crafted subject containing `::error::`
# could forge log annotations.
HEAD_SUBJECT_SAFE="${HEAD_SUBJECT//::/__}"
RELEASE_SUBJECT_RE='^chore:[[:space:]]*release[[:space:]]+v[0-9]+\.[0-9]+\.[0-9]+([[:space:]]+\(#[0-9]+\))?$'
shopt -s nocasematch
if [[ "$HEAD_SUBJECT" =~ $RELEASE_SUBJECT_RE ]]; then
shopt -u nocasematch
echo "HEAD commit subject matches a release commit — skipping rc."
echo " subject (sanitised): $HEAD_SUBJECT_SAFE"
echo "should_run=false" >> "$GITHUB_OUTPUT"
exit 0
fi
shopt -u nocasematch
# Squash-merge commits include `(#NNNN)` at the end of the subject.
if [[ "$HEAD_SUBJECT" =~ \(#([0-9]+)\)[[:space:]]*$ ]]; then
PR_NUM="${BASH_REMATCH[1]}"
echo "Detected squash-merge of PR #$PR_NUM — checking labels."
if LABELS_JSON="$(gh pr view "$PR_NUM" --repo "$REPO" --json labels 2>/dev/null)"; then
if printf '%s' "$LABELS_JSON" | jq -e '.labels[] | select(.name == "release")' >/dev/null; then
echo "PR #$PR_NUM has the 'release' label — skipping rc."
echo "should_run=false" >> "$GITHUB_OUTPUT"
exit 0
fi
echo "PR #$PR_NUM has no 'release' label — proceeding."
else
# Lookup failure is not fatal — fall through to dedup check.
echo "::warning::Could not read labels for PR #${PR_NUM} — falling through."
fi
fi
# Dedup: is there already an rc/<HEAD_SHA> marker pointing at HEAD?
MARKER="rc/${HEAD_SHA}"
if git rev-parse "refs/tags/$MARKER" >/dev/null 2>&1; then
echo "HEAD already has marker $MARKER — skipping."
echo "should_run=false" >> "$GITHUB_OUTPUT"
else
echo "No marker on HEAD — proceeding."
echo "should_run=true" >> "$GITHUB_OUTPUT"
fi
# ── Phase 3: reusable CI gate ──────────────────────────────────────────────
# Runs for both rc (when guard says go) and stable. No `secrets:` passed —
# ci.yml and its entire reusable-workflow chain (ci-quality, ci-tests,
# ci-e2e, ci-scope-parity, ci-report) reference zero `secrets.*` values;
# passing any would be unused surface. GITHUB_TOKEN is implicit.
ci:
needs: [route, rc-guard]
if: ${{ always() && (needs.route.outputs.mode == 'stable' || needs.rc-guard.outputs.should_run == 'true') }}
uses: ./.github/workflows/ci.yml
permissions:
contents: read
actions: read
# No pull-requests:write — `ci.yml`'s save-pr-meta job is gated on
# `github.event_name == 'pull_request'`, so it never runs during a
# tag-triggered publish. Least-privilege for release-critical paths.
# ── Phase 4: publish to npm + push refs (RC path) ──────────────────────────
# INVARIANT: `timeout-minutes` MUST stay below the App-token TTL (~60 min
# for actions/create-github-app-token installation tokens). The atomic
# tag-push step relies on the token minted at job start; if the job ever
# runs longer than the TTL, the push fails with an opaque 401. If you
# need to raise the timeout, re-mint the token immediately before the
# `Create and push rc tags` step instead.
publish:
needs: ci
name: Publish to npm
needs: [route, rc-guard, ci]
if: ${{ always() && needs.ci.result == 'success' && (needs.route.outputs.mode == 'stable' || needs.rc-guard.outputs.should_run == 'true') }}
runs-on: ubuntu-latest
timeout-minutes: 15
timeout-minutes: 20
permissions:
# contents: write — RC path needs it for `git push --atomic` (v-tag +
# marker). Stable path runs in the same job and inherits the grant; it
# never invokes `git push`, so the elevated scope is unused there.
# id-token: write — npm provenance attestation.
contents: write
id-token: write
outputs:
# Two distinct step IDs feed this output; exactly one fires per run.
vtag: ${{ steps.rc-tags.outputs.vtag || steps.stable-vtag.outputs.vtag }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
# ── Mint short-lived GitHub App token (RC only) ──────────────────────
# Industry direction (2025-2026): GitHub Apps with
# `actions/create-github-app-token` over long-lived PATs for
# workflow-touching tag pushes. Same fine-grained permission surface,
# ~1h expiry, not tied to a user seat, organizationally auditable.
# Replaces a prior fine-grained PAT.
#
# Required secrets (set in repo Settings → Secrets and variables → Actions):
# secrets.RELEASE_APP_ID — the App's numeric ID
# secrets.RELEASE_APP_PRIVATE_KEY — the App's PEM private key
# (The App ID is technically not sensitive — it's visible on the App's
# settings page — but storing it as a secret is harmless and avoids
# mixing storage classes for the same App.)
# The App must be installed on this repository with:
# - Contents: write (push the v-tag and rc marker)
# - Workflows: write (because the v-tag's tree may touch
# .github/workflows/**, which the default
# GITHUB_TOKEN cannot author)
# - Metadata: read (required for the `gh api /users/<slug>[bot]`
# bot-identity lookup in the tag-push step)
- name: Mint GitHub App token (RC)
if: needs.route.outputs.mode == 'rc'
id: app-token
uses: actions/create-github-app-token@bcd2ba49218906704ab6c1aa796996da409d3eb1 # v3.2.0
with:
# `client-id` is the renamed input that supersedes the deprecated
# `app-id` in v3.x. The action accepts the App's numeric ID or
# its Client ID under this name. We pass the numeric App ID,
# which the action resolves correctly.
client-id: ${{ secrets.RELEASE_APP_ID }}
private-key: ${{ secrets.RELEASE_APP_PRIVATE_KEY }}
# ── Separate checkout steps per mode ─────────────────────────────────
# Conditional `token:` expressions are footguns: empty string passed to
# actions/checkout fails opaquely, and `|| github.token` silently
# degrades a missing token to GITHUB_TOKEN, masking auth failures until
# the eventual `git push`. Two distinct steps make the auth contract
# explicit and fail loudly at checkout when the App token mint failed
# on the RC path.
- name: Checkout (RC)
if: needs.route.outputs.mode == 'rc'
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
fetch-depth: 0
fetch-tags: true
# Short-lived GitHub App installation token. Required because the
# v-tag push lands at a SHA whose tree may touch
# `.github/workflows/**`, which the default GITHUB_TOKEN cannot
# author.
token: ${{ steps.app-token.outputs.token }}
# Do not persist the token in .git/config (artipacked audit). The
# RC tag push uses an inline `http.extraheader` at push time only;
# the credential never lands on disk. See the
# `Create and push rc tags` step below.
persist-credentials: false
- name: Checkout (stable)
if: needs.route.outputs.mode == 'stable'
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
# No `token:` — actions/checkout uses GITHUB_TOKEN by default. Stable
# path performs no git pushes; the default scope is sufficient.
with:
# No git pushes from the stable path either. Skip credential
# persistence (artipacked audit).
persist-credentials: false
- name: Working-tree sanity
# Defense in depth (mirrors the vtag integrity gate, but on the input side):
# if a route-mode regression skipped both checkout `if:` gates, all
# downstream steps would run on a bare runner and produce confusing
# ENOENT errors. Fail loudly and early here instead.
shell: bash
run: |
if [ ! -f gitnexus/package.json ]; then
echo "::error::no working tree at gitnexus/package.json — route classification likely failed silently."
exit 1
fi
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 22
registry-url: https://registry.npmjs.org
# Hermetic install for the published artifact — no cache carry-over
# from non-tag contexts. setup-node v5+ caches by default when a
# packageManager field is present in package.json, so the explicit
# opt-out is required to clear the zizmor cache-poisoning audit.
# ~30s slower per release; runs rarely.
# Node 24 ships with npm >= 11.5.x, which is the minimum that
# supports npm Trusted Publishing OIDC. Node 22 ships with npm
# 10.9.x (no OIDC) and `npm install -g npm@latest` to self-upgrade
# is fragile — it can crash the in-flight reify with
# `MODULE_NOT_FOUND` on `promise-retry` etc. Bumping the Node
# version is the clean fix; the package's `engines` field is
# `>=22.0.0` so consumer-side compatibility is unaffected (this
# Node version is only used during publish, not by package users).
node-version: 24
# `registry-url:` is intentionally OMITTED. Under npm Trusted
# Publishing, OIDC only engages when no credential is configured.
# Setting `registry-url:` would make setup-node write
# `//registry.npmjs.org/:_authToken=${NODE_AUTH_TOKEN}` into the
# runner's .npmrc AND export NODE_AUTH_TOKEN from its `token:`
# input (default github.token). `npm publish` would then attempt
# GITHUB_TOKEN as the npm token, get rejected with 404, and OIDC
# would never be tried. See actions/setup-node#1440 and the GitHub
# Community discussion #176761 for the upstream bug and consensus
# workaround.
#
# Hermetic install for published artifacts — opt out of the v5+
# default packageManager-based caching (clears the zizmor
# cache-poisoning audit). ~30s slower per release; runs rarely.
package-manager-cache: false
- name: Build gitnexus-shared
run: npm install && npm run build
working-directory: gitnexus-shared
- run: npm ci
- name: Install gitnexus dependencies
run: npm ci
working-directory: gitnexus
- name: Verify version consistency
# ── Stable-only: verify the tag and package.json agree ───────────────
- name: Verify version consistency (stable)
if: needs.route.outputs.mode == 'stable'
shell: bash
working-directory: gitnexus
run: |
set -euo pipefail
TAG_VERSION="${GITHUB_REF#refs/tags/v}"
if ! [[ "$TAG_VERSION" =~ ^[0-9]+\.[0-9]+\.[0-9]+(-[a-zA-Z0-9.]+)?$ ]]; then
echo "::error::Tag does not follow semver: v$TAG_VERSION"
# Stable mode REJECTS prerelease suffixes — those are filtered at
# trigger by the negative-glob filter, but defend at the bash layer too.
if ! [[ "$TAG_VERSION" =~ ^[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
echo "::error::Stable tag must be ^v[0-9]+.[0-9]+.[0-9]+$ — got v$TAG_VERSION"
exit 1
fi
PKG_VERSION=$(node -p "require('./package.json').version")
@@ -65,24 +424,376 @@ jobs:
exit 1
fi
echo "Version verified: $PKG_VERSION"
working-directory: gitnexus
- name: Build
# ── RC-only: compute the next rc version against the live registry ──
- name: Resolve rc version (rc)
id: rc-version
if: needs.route.outputs.mode == 'rc'
shell: bash
working-directory: gitnexus
env:
BUMP_INPUT: ${{ inputs.bump }}
EVENT_NAME: ${{ github.event_name }}
PKG_NAME: gitnexus
run: |
set -euo pipefail
# 1. Current published `latest` — the floor for any new rc base.
# Only E404 ("never published") falls back to package.json; any
# other error (network, auth, malformed response) fails fast
# (retry-loud policy: never silently substitute on transient errors).
NPM_STDERR_LATEST="$(mktemp)"
if CURRENT_LATEST="$(npm view "$PKG_NAME" version 2>"$NPM_STDERR_LATEST")"; then
:
else
if grep -qiE 'E404|not found' "$NPM_STDERR_LATEST"; then
CURRENT_LATEST="$(node -p "require('./package.json').version")"
echo "Package not on registry (E404) — seeding from package.json: $CURRENT_LATEST"
else
echo "::error::npm registry unreachable for 'view version':" >&2
cat "$NPM_STDERR_LATEST" >&2
rm -f "$NPM_STDERR_LATEST"
exit 1
fi
fi
rm -f "$NPM_STDERR_LATEST"
CURRENT_LATEST_CLEAN="${CURRENT_LATEST%%-*}"
# 2. Full version list — needed for the counter and active-cycle
# inference. Same E404-only fallback.
NPM_STDERR_VERSIONS="$(mktemp)"
if VERSIONS_JSON="$(npm view "$PKG_NAME" versions --json 2>"$NPM_STDERR_VERSIONS")"; then
:
else
if grep -qiE 'E404|not found' "$NPM_STDERR_VERSIONS"; then
VERSIONS_JSON='[]'
echo "No published versions for $PKG_NAME yet (E404)."
else
echo "::error::npm registry unreachable for 'view versions':" >&2
cat "$NPM_STDERR_VERSIONS" >&2
rm -f "$NPM_STDERR_VERSIONS"
exit 1
fi
fi
rm -f "$NPM_STDERR_VERSIONS"
# 3. Base selection.
# - workflow_dispatch + bump != auto → explicit cycle reset.
# - Otherwise (push, or dispatch with bump=auto) → continue the
# highest active rc base > latest if any; else patch from latest.
# Curated wrapper around `npx semver` — bare npx errors are noisy
# and don't distinguish registry-unreachable from invalid-bump-spec.
semver_bump() {
local kind="$1" current="$2" stderr_file out
stderr_file="$(mktemp)"
if out="$(npx --yes -p semver@7 semver -i "$kind" "$current" 2>"$stderr_file")"; then
rm -f "$stderr_file"
printf '%s' "$out"
return 0
fi
echo "::error::semver bump failed (kind=${kind}, current=${current}):" >&2
cat "$stderr_file" >&2
rm -f "$stderr_file"
return 1
}
if [ "$EVENT_NAME" = "workflow_dispatch" ] \
&& [ -n "${BUMP_INPUT:-}" ] \
&& [ "${BUMP_INPUT:-auto}" != "auto" ]; then
BASE="$(semver_bump "$BUMP_INPUT" "$CURRENT_LATEST_CLEAN")"
echo "Explicit bump=$BUMP_INPUT → BASE=$BASE"
else
cat > /tmp/active_base.mjs <<'NODESCRIPT'
const latest = process.env.LATEST;
let v;
try { v = JSON.parse(process.env.VERSIONS_JSON); } catch { v = []; }
if (!Array.isArray(v)) v = [v];
const parse = s => s.split(".").map(n => parseInt(n, 10));
const gt = (a, b) => {
const [A, B] = [parse(a), parse(b)];
for (let i = 0; i < 3; i++) if (A[i] !== B[i]) return A[i] > B[i];
return false;
};
const bases = new Set();
for (const s of v) {
const m = /^(\d+\.\d+\.\d+)-rc\.\d+$/.exec(s);
if (m && gt(m[1], latest)) bases.add(m[1]);
}
if (!bases.size) { process.stdout.write(""); process.exit(0); }
const sorted = [...bases].sort((a, b) => gt(a, b) ? 1 : -1);
process.stdout.write(sorted[sorted.length - 1]);
NODESCRIPT
ACTIVE_BASE="$(LATEST="$CURRENT_LATEST_CLEAN" VERSIONS_JSON="$VERSIONS_JSON" node /tmp/active_base.mjs)"
if [ -n "$ACTIVE_BASE" ]; then
BASE="$ACTIVE_BASE"
echo "Continuing active rc cycle → BASE=$BASE"
else
BASE="$(semver_bump patch "$CURRENT_LATEST_CLEAN")"
echo "No active rc cycle → patch bump from latest → BASE=$BASE"
fi
fi
# 4. Counter: 1 + max existing N for `${BASE}-rc.*`, else 1.
cat > /tmp/next_rc.mjs <<'NODESCRIPT'
const base = process.env.BASE;
const prefix = base + "-rc.";
let v;
try { v = JSON.parse(process.env.VERSIONS_JSON); } catch { v = []; }
if (!Array.isArray(v)) v = [v];
const ns = v
.filter(s => typeof s === "string" && s.startsWith(prefix))
.map(s => parseInt(s.slice(prefix.length), 10))
.filter(n => Number.isInteger(n) && n >= 0);
process.stdout.write(String(ns.length ? Math.max(...ns) + 1 : 1));
NODESCRIPT
NEXT_N="$(BASE="$BASE" VERSIONS_JSON="$VERSIONS_JSON" node /tmp/next_rc.mjs)"
RC_VERSION="${BASE}-rc.${NEXT_N}"
echo "Computed rc: $RC_VERSION"
# 5. Defensive: if the exact version already exists on the registry
# (race with another run), abort before re-publishing.
NPM_STDERR_EXISTS="$(mktemp)"
if npm view "$PKG_NAME@$RC_VERSION" version 2>"$NPM_STDERR_EXISTS" >/dev/null; then
rm -f "$NPM_STDERR_EXISTS"
echo "::error::Version $RC_VERSION already exists on npm — aborting."
exit 1
else
if grep -qiE 'E404|not found' "$NPM_STDERR_EXISTS"; then
rm -f "$NPM_STDERR_EXISTS"
# Version doesn't exist — safe to proceed.
else
echo "::error::npm registry unreachable for existence check:" >&2
cat "$NPM_STDERR_EXISTS" >&2
rm -f "$NPM_STDERR_EXISTS"
exit 1
fi
fi
{
echo "base=$BASE"
echo "rc_n=$NEXT_N"
echo "rc_version=$RC_VERSION"
} >> "$GITHUB_OUTPUT"
- name: Apply rc version in-CI
if: needs.route.outputs.mode == 'rc'
shell: bash
working-directory: gitnexus
run: |
set -euo pipefail
npm version "${{ steps.rc-version.outputs.rc_version }}" \
--no-git-tag-version --allow-same-version
- name: Build gitnexus
run: npm run build
working-directory: gitnexus
- name: Dry-run publish
run: npm publish --dry-run
working-directory: gitnexus
- name: Publish to npm
run: npm publish --provenance --access public
# Cheap verification that the tarball assembles before the real publish.
shell: bash
working-directory: gitnexus
env:
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
NPM_TAG: ${{ needs.route.outputs.mode == 'rc' && 'rc' || 'latest' }}
run: npm publish --dry-run --tag "$NPM_TAG"
- name: Extract release notes from CHANGELOG
# ── Acquire the "rc lock" BEFORE publishing (idempotency anchor) ─────
# We create two refs and push atomically:
# v<RC_VERSION> → annotated tag on a detached release commit whose
# tree contains the rewritten package.json, so the
# tag's source matches the npm tarball.
# rc/<HEAD_SHA> → lightweight tag on HEAD; the guard's dedup key.
# Push fails → nothing published. Push succeeds, npm fails → marker
# blocks retries until manual cleanup (see Rollback Runbook in plan).
- name: Create and push rc tags
id: rc-tags
if: needs.route.outputs.mode == 'rc'
shell: bash
working-directory: gitnexus
env:
RC_VERSION: ${{ steps.rc-version.outputs.rc_version }}
HEAD_SHA: ${{ needs.rc-guard.outputs.head_sha }}
# Short-lived GitHub App token. Auth is supplied inline at push
# time via `http.extraheader` (per GitHub's documented
# x-access-token Basic pattern). It is NOT persisted in
# .git/config (artipacked audit) — checkout above ran with
# `persist-credentials: false`.
PUSH_TOKEN: ${{ steps.app-token.outputs.token }}
# App's slug from create-github-app-token (e.g. `gitnexus-release-bot`).
# Used to attribute the release commit to the App identity rather
# than the generic github-actions[bot]. The bot's numeric user-id
# is resolved at runtime via the GitHub API (the action does not
# expose it directly as of v3.2.0).
APP_SLUG: ${{ steps.app-token.outputs.app-slug }}
GH_TOKEN: ${{ steps.app-token.outputs.token }}
run: |
set -euo pipefail
VTAG="v${RC_VERSION}"
MARKER="rc/${HEAD_SHA}"
# Resolve the App's bot user-id and construct the noreply email
# in the GitHub-canonical `<id>+<slug>[bot]@users.noreply.github.com`
# shape. `[bot]` is part of the actual login on GitHub.
#
# The lookup is wrapped in a bounded retry because the first RC
# after App installation may hit propagation delay (404), and
# transient api.github.com 5xx during heavy org activity is a real
# failure class. Without retry, every transient blip aborts the
# entire release after CI has already succeeded.
BOT_LOGIN="${APP_SLUG}[bot]"
BOT_USER_ID=""
api_stderr="$(mktemp)"
for attempt in 1 2 3; do
if BOT_USER_ID="$(gh api "/users/${BOT_LOGIN}" --jq .id 2>"$api_stderr")" \
&& [[ "${BOT_USER_ID}" =~ ^[0-9]+$ ]]; then
break
fi
BOT_USER_ID=""
if [ "$attempt" -lt 3 ]; then
echo "::warning::bot user-id lookup attempt ${attempt} failed; retrying in $((attempt * 5))s"
sleep $((attempt * 5))
fi
done
if ! [[ "${BOT_USER_ID}" =~ ^[0-9]+$ ]]; then
echo "::error::Could not resolve bot user-id for ${BOT_LOGIN} after 3 attempts."
echo "::error::gh api stderr:"
cat "$api_stderr" >&2 || true
echo "::error::Common causes: (a) newly-installed App — user record still propagating to /users/ (wait ~5min, redispatch with force=true); (b) App lacks Metadata: read permission; (c) transient api.github.com 5xx (redispatch)."
rm -f "$api_stderr"
exit 1
fi
rm -f "$api_stderr"
git config user.name "${BOT_LOGIN}"
git config user.email "${BOT_USER_ID}+${BOT_LOGIN}@users.noreply.github.com"
# Detached release commit with the version bump — main stays
# pristine, but the v-tag's tree matches the published package
# exactly (release-integrity).
git add package.json package-lock.json 2>/dev/null || git add package.json
git commit -m "release: ${VTAG}" --allow-empty
RELEASE_SHA="$(git rev-parse HEAD)"
echo "Detached release commit: $RELEASE_SHA"
git tag -a "$VTAG" "$RELEASE_SHA" -m "$VTAG"
git tag "$MARKER" "$HEAD_SHA"
# Inline auth header. The base64-encoded form is masked as well
# as the raw token, because GitHub's secret-masker only masks the
# raw value — any subsequent `set -x` / GIT_TRACE line would
# otherwise expose the encoded credential.
#
# `set +x` wraps the compute+mask pair so that if an operator
# enables ACTIONS_STEP_DEBUG=true for triage (which turns on
# `set -x` globally), the assignment is NOT traced for the one
# line between compute and mask-registration. Without this wrap,
# debug mode would log `+ auth_header='Authorization: Basic <encoded>'`
# exposing a still-valid (~1h) App token.
{ set +x; } 2>/dev/null
auth_header="Authorization: Basic $(printf 'x-access-token:%s' "${PUSH_TOKEN}" | base64 -w0)"
echo "::add-mask::${auth_header}"
# Re-enable tracing only when explicitly requested via step-debug.
if [ "${ACTIONS_STEP_DEBUG:-false}" = "true" ]; then set -x; fi
# Atomic push of both refs. If either would clobber an existing
# remote ref, the push fails and we stop before npm publish.
git -c http.extraheader="${auth_header}" \
push --atomic origin "refs/tags/$VTAG" "refs/tags/$MARKER"
{
echo "vtag=$VTAG"
echo "marker=$MARKER"
echo "release_sha=$RELEASE_SHA"
} >> "$GITHUB_OUTPUT"
- name: Set vtag (stable)
id: stable-vtag
if: needs.route.outputs.mode == 'stable'
shell: bash
# github.ref_name flows in via env to avoid templating into the
# shell source (template-injection audit). Even though refs are
# constrained by git naming rules, the env-passthrough pattern
# makes injection structurally impossible.
env:
REF_NAME: ${{ github.ref_name }}
run: |
echo "vtag=${REF_NAME}" >> "$GITHUB_OUTPUT"
# ── vtag integrity gate ──────────────────────────────────────────────
# Fail closed before any artifact-producing step (npm publish, Release,
# Docker) runs against an empty or mode-mismatched vtag. Prevents the
# silent "Release named main" / "Docker tagged from ref fallback"
# failure modes that the previous draft was vulnerable to.
- name: vtag integrity gate
id: vtag-gate
shell: bash
env:
MODE: ${{ needs.route.outputs.mode }}
VTAG: ${{ steps.rc-tags.outputs.vtag || steps.stable-vtag.outputs.vtag }}
run: |
set -euo pipefail
if [ -z "$VTAG" ]; then
echo "::error::vtag is empty — refusing to create GitHub Release or trigger Docker."
exit 1
fi
case "$MODE" in
rc)
if ! [[ "$VTAG" =~ ^v[0-9]+\.[0-9]+\.[0-9]+-rc\.[0-9]+$ ]]; then
echo "::error::vtag '${VTAG}' does not match rc shape ^v[0-9]+.[0-9]+.[0-9]+-rc.[0-9]+$"
exit 1
fi
;;
stable)
if ! [[ "$VTAG" =~ ^v[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
echo "::error::vtag '${VTAG}' does not match stable shape ^v[0-9]+.[0-9]+.[0-9]+$"
exit 1
fi
;;
*)
echo "::error::unknown mode '${MODE}' at vtag integrity gate."
exit 1
;;
esac
echo "vtag verified: ${VTAG} (mode=${MODE})"
echo "vtag=${VTAG}" >> "$GITHUB_OUTPUT"
# npm Trusted Publishing (GA'd 2025-07-31). OIDC authentication only
# engages when no npm credential is configured anywhere — the absence
# is the signal. Two upstream behaviors had to be neutralized for
# this to work:
#
# 1. setup-node's `registry-url:` is omitted (see the setup-node
# step above). With it, setup-node writes
# `//registry.npmjs.org/:_authToken=${NODE_AUTH_TOKEN}` into
# .npmrc and exports NODE_AUTH_TOKEN from `token:` (defaulting
# to github.token). npm publish then sends GITHUB_TOKEN as the
# bearer credential and the registry returns 404. OIDC is never
# tried because npm thinks it already has a credential.
# 2. The runner's bundled npm (10.9.x on Node 22) has no OIDC
# support; the upgrade step above pins it to >= 11.5.1.
#
# Provenance is auto-attached by the registry on trusted-publisher
# publishes — no --provenance flag needed.
#
# Prerequisite: register the package as a trusted publisher at
# https://www.npmjs.com/package/gitnexus/access (Publishing access →
# Trusted Publishers → GitHub Actions):
# Owner: abhigyanpatwari
# Repository: GitNexus
# Workflow: publish.yml
# Environment: (none)
- name: Publish to npm
shell: bash
working-directory: gitnexus
env:
NPM_TAG: ${{ needs.route.outputs.mode == 'rc' && 'rc' || 'latest' }}
run: npm publish --access public --tag "$NPM_TAG"
# ── Stable-only: pull CHANGELOG body if present ──────────────────────
- name: Extract release notes from CHANGELOG (stable)
id: changelog
if: needs.route.outputs.mode == 'stable'
shell: bash
run: |
VERSION="${GITHUB_REF#refs/tags/v}"
@@ -98,5 +809,90 @@ jobs:
- name: Create GitHub Release
uses: softprops/action-gh-release@b4309332981a82ec1c5618f44dd2e27cc8bfbfda # v2
with:
body_path: ${{ steps.changelog.outputs.fallback == 'false' && '/tmp/release-notes.md' || '' }}
generate_release_notes: ${{ steps.changelog.outputs.fallback == 'true' }}
tag_name: ${{ steps.vtag-gate.outputs.vtag }}
name: >-
${{ needs.route.outputs.mode == 'rc'
&& format('Release Candidate {0}', steps.vtag-gate.outputs.vtag)
|| steps.vtag-gate.outputs.vtag }}
prerelease: ${{ needs.route.outputs.mode == 'rc' }}
make_latest: ${{ needs.route.outputs.mode == 'stable' && 'true' || 'false' }}
# Stable: prefer CHANGELOG body, fall back to auto-generated.
# RC: always auto-generated + the prerelease body block below.
body_path: >-
${{ needs.route.outputs.mode == 'stable' && steps.changelog.outputs.fallback == 'false'
&& '/tmp/release-notes.md' || '' }}
generate_release_notes: >-
${{ needs.route.outputs.mode == 'rc'
|| steps.changelog.outputs.fallback == 'true' }}
body: >-
${{ needs.route.outputs.mode == 'rc' && format(
'Automated release candidate build from `main`.{0}{0}**npm:** `npm install gitnexus@rc`{0}**Version:** `{1}`{0}**Target base:** `{2}` (rc #{3}){0}**Source commit (main):** {4}{0}**Release commit (versioned tree):** {5}{0}{0}Release candidates are pre-stable builds intended for early testing. Stable releases remain on the `latest` dist-tag.',
'\n',
steps.rc-version.outputs.rc_version,
steps.rc-version.outputs.base,
steps.rc-version.outputs.rc_n,
needs.rc-guard.outputs.head_sha,
steps.rc-tags.outputs.release_sha
) || '' }}
# ── RC partial-failure cleanup ───────────────────────────────────────
# If anything after the atomic tag-push step failed (npm publish
# blew up, GitHub Release call timed out, etc.), the v-tag and
# rc/<SHA> marker are already on origin. External consumers
# (Renovate, Dependabot, Releases RSS) can ingest a phantom tag for
# a version that was never published to npm. This step deletes them
# automatically so the operator's recovery is just "redispatch with
# force=true on the next commit", not a manual ref cleanup.
#
# Scoped strictly to RC + real (non-dry-run) + the rc-tags step
# actually produced a vtag (otherwise nothing to clean up). The
# App token is still valid (~1h TTL, job timeout 20min).
- name: Cleanup pushed tags on partial failure
if: ${{ failure() && needs.route.outputs.mode == 'rc' && steps.rc-tags.outputs.vtag != '' }}
shell: bash
working-directory: gitnexus
env:
VTAG: ${{ steps.rc-tags.outputs.vtag }}
MARKER: ${{ steps.rc-tags.outputs.marker }}
PUSH_TOKEN: ${{ steps.app-token.outputs.token }}
run: |
set -uo pipefail
echo "::warning::Publish step failed after tag push. Cleaning up remote refs to prevent phantom-version ingestion by downstream consumers."
{ set +x; } 2>/dev/null
auth_header="Authorization: Basic $(printf 'x-access-token:%s' "${PUSH_TOKEN}" | base64 -w0)"
echo "::add-mask::${auth_header}"
if [ "${ACTIONS_STEP_DEBUG:-false}" = "true" ]; then set -x; fi
# Delete v-tag and marker. Each delete is best-effort — if one
# is already absent (atomic push partially rejected, or earlier
# cleanup ran), the other still gets attempted.
for ref in "refs/tags/${VTAG}" "refs/tags/${MARKER}"; do
if git -c http.extraheader="${auth_header}" push origin --delete "${ref}" 2>&1; then
echo "deleted origin ${ref}"
else
echo "::warning::could not delete origin ${ref} — may already be absent or protected. Manual cleanup may be required."
fi
done
echo "::notice::Cleanup complete. To retry the release, redispatch the workflow with force=true on the same SHA, or push a new commit to main."
# ── Phase 5 (RC only): Docker images ───────────────────────────────────────
# R6: Docker remains RC-only. Stable Docker builds are explicitly deferred.
# Secrets are passed explicitly (not via `secrets: inherit`) so the
# callee's secret surface is auditable from the caller's source.
docker:
name: Build & Push RC Docker images
needs: [route, publish]
if: ${{ needs.route.outputs.mode == 'rc' && needs.publish.outputs.vtag != '' }}
uses: ./.github/workflows/docker.yml
secrets:
DOCKERHUB_USERNAME: ${{ secrets.DOCKERHUB_USERNAME }}
DOCKERHUB_TOKEN: ${{ secrets.DOCKERHUB_TOKEN }}
permissions:
contents: read
packages: write
id-token: write
attestations: write
with:
tag: ${{ needs.publish.outputs.vtag }}
-459
View File
@@ -1,459 +0,0 @@
name: Release Candidate
on:
# Publish a release-candidate build whenever a merge/commit lands on main.
# Docs/README-only changes are filtered out so prose updates don't
# cut a release.
push:
branches: [main]
paths-ignore:
- '**.md'
- 'docs/**'
- 'LICENSE'
workflow_dispatch:
inputs:
bump:
description: >-
Cycle policy. 'auto' (default) continues the active rc cycle on
this branch if there is one, otherwise bumps patch from latest.
Choose 'patch' / 'minor' / 'major' to explicitly start or reset
an rc cycle.
required: false
default: 'auto'
type: choice
options:
- auto
- patch
- minor
- major
force:
description: 'Publish even when HEAD already has an rc marker'
required: false
default: 'false'
type: choice
options:
- 'false'
- 'true'
# No workflow-level permissions — scoped per job below.
permissions: {}
# Concurrency convention: see CONTRIBUTING.md → "GitHub Actions — Concurrency Convention".
# Serialize all runs on the same ref (push + workflow_dispatch) to prevent two publishes
# racing on the rc counter. cancel-in-progress: false — the earlier merge publishes first.
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: false
jobs:
# ── Skip when HEAD already has an rc marker (retry / duplicate dispatch) ──
# The marker is a lightweight tag `rc/<HEAD_SHA>` pushed *before* `npm
# publish`, so a failed publish leaves the marker in place and the guard
# refuses to re-publish. Recovery path after a partial failure:
# git push --delete origin rc/<HEAD_SHA> v<RC_VERSION>
# then redispatch with force=true.
guard:
name: Check if release candidate should run
runs-on: ubuntu-latest
timeout-minutes: 5
permissions:
contents: read
pull-requests: read # read PR labels on the merge commit
outputs:
should_run: ${{ steps.decide.outputs.should_run }}
head_sha: ${{ steps.decide.outputs.head_sha }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
fetch-depth: 0
fetch-tags: true
- name: Decide
id: decide
shell: bash
env:
FORCE: ${{ inputs.force }}
BUMP_INPUT: ${{ inputs.bump }}
EVENT_NAME: ${{ github.event_name }}
GH_TOKEN: ${{ github.token }}
REPO: ${{ github.repository }}
run: |
set -euo pipefail
HEAD_SHA=$(git rev-parse HEAD)
echo "head_sha=$HEAD_SHA" >> "$GITHUB_OUTPUT"
if [ "$FORCE" = "true" ]; then
echo "Force flag set — running regardless of marker tag."
echo "should_run=true" >> "$GITHUB_OUTPUT"
exit 0
fi
# An explicit cycle reset on dispatch (bump != auto) also bypasses
# the dedup guard — the maintainer is deliberately asking for a
# new rc from the same commit.
if [ "$EVENT_NAME" = "workflow_dispatch" ] \
&& [ -n "${BUMP_INPUT:-}" ] \
&& [ "${BUMP_INPUT:-auto}" != "auto" ]; then
echo "Explicit bump=$BUMP_INPUT — bypassing marker dedup."
echo "should_run=true" >> "$GITHUB_OUTPUT"
exit 0
fi
# ── Skip when the merge commit corresponds to a release ─────────
# Two complementary checks (belt-and-suspenders):
# 1. The HEAD commit subject matches `chore: release vX.Y.Z`
# (the canonical release-PR title in this repo). Anchored
# at both ends to require the bare title or the squash-merge
# `(#NNNN)` suffix exactly — rejects noisy variants like
# `chore: release v1.0.0 (something unrelated)`.
# 2. The squash-merged PR carries the `release` label.
# Either match suppresses the rc build — stable releases publish
# via publish.yml on the v-tag, so the rc cycle should pause for
# them rather than racing the npm publish.
HEAD_SUBJECT="$(git log -1 --pretty=%s HEAD)"
# Sanitise GitHub-Actions annotation prefixes before logging the
# raw subject — defence-in-depth so a hypothetical commit subject
# containing `::error::` or `::set-output::` cannot forge log
# annotations even though %s strips newlines.
HEAD_SUBJECT_SAFE="${HEAD_SUBJECT//::/__}"
RELEASE_SUBJECT_RE='^chore:[[:space:]]*release[[:space:]]+v[0-9]+\.[0-9]+\.[0-9]+([[:space:]]+\(#[0-9]+\))?$'
if [[ "$HEAD_SUBJECT" =~ $RELEASE_SUBJECT_RE ]]; then
echo "HEAD commit subject matches a release commit — skipping rc."
echo " subject (sanitised): $HEAD_SUBJECT_SAFE"
echo "should_run=false" >> "$GITHUB_OUTPUT"
exit 0
fi
# Squash-merge commits include `(#NNNN)` at the end of the subject.
if [[ "$HEAD_SUBJECT" =~ \(#([0-9]+)\)[[:space:]]*$ ]]; then
PR_NUM="${BASH_REMATCH[1]}"
echo "Detected squash-merge of PR #$PR_NUM — checking labels."
if LABELS_JSON="$(gh pr view "$PR_NUM" --repo "$REPO" --json labels 2>/dev/null)"; then
if printf '%s' "$LABELS_JSON" | jq -e '.labels[] | select(.name == "release")' >/dev/null; then
echo "PR #$PR_NUM has the 'release' label — skipping rc."
echo "should_run=false" >> "$GITHUB_OUTPUT"
exit 0
fi
echo "PR #$PR_NUM has no 'release' label — proceeding."
else
# Lookup failure is not fatal — fall through to the dedup check
# so a transient GH API hiccup doesn't silently suppress rc builds.
echo "::warning::Could not read labels for PR #${PR_NUM} — falling through."
fi
fi
# Dedup: is there already an rc/<HEAD_SHA> marker pointing at HEAD?
MARKER="rc/${HEAD_SHA}"
if git rev-parse "refs/tags/$MARKER" >/dev/null 2>&1; then
echo "HEAD already has marker $MARKER — skipping."
echo "should_run=false" >> "$GITHUB_OUTPUT"
else
echo "No marker on HEAD — proceeding."
echo "should_run=true" >> "$GITHUB_OUTPUT"
fi
# ── Reuse the stable CI workflow ─────────────────────────────────────
ci:
needs: guard
if: needs.guard.outputs.should_run == 'true'
uses: ./.github/workflows/ci.yml
permissions:
contents: read
secrets: inherit
# ── Publish the rc build to npm + create GitHub prerelease ───────────
publish:
name: Publish release candidate to npm
needs: [guard, ci]
if: needs.guard.outputs.should_run == 'true'
runs-on: ubuntu-latest
timeout-minutes: 20
permissions:
# The default GITHUB_TOKEN cannot be granted `workflows: write`, so
# tag pushes that reach a commit which modified `.github/workflows/**`
# are rejected with: "refusing to allow a GitHub App to create or
# update workflow ... without `workflows` permission". We pass a
# fine-grained PAT (RELEASE_PUSH_TOKEN, scoped to this repo with
# Contents: write + Workflows: write) to `actions/checkout` so that
# the subsequent `git push --atomic` of the v-tag and rc marker
# carries the PAT's identity. Job-level GITHUB_TOKEN keeps its
# scoped permissions for everything else (npm provenance, etc.).
contents: write # push rc tag + marker (via PAT)
id-token: write # npm provenance
outputs:
vtag: ${{ steps.reltag.outputs.vtag }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
fetch-depth: 0
fetch-tags: true
# Use the PAT so `origin` is preauthed for `git push`. Without
# this the default GITHUB_TOKEN is wired into the remote, and a
# workflows-touching tag push is rejected — see the permissions
# block above.
token: ${{ secrets.RELEASE_PUSH_TOKEN }}
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: 22
registry-url: https://registry.npmjs.org
# Hermetic install — release-candidate produces shipped artifacts.
# setup-node v5+ caches by default when a packageManager field is
# present in package.json; explicit opt-out is required to clear
# the zizmor cache-poisoning audit. See cache-poisoning audit.
package-manager-cache: false
- name: Build gitnexus-shared
run: npm install && npm run build
working-directory: gitnexus-shared
- name: Install gitnexus dependencies
run: npm ci
working-directory: gitnexus
- name: Resolve rc version
id: version
shell: bash
working-directory: gitnexus
env:
BUMP_INPUT: ${{ inputs.bump }}
EVENT_NAME: ${{ github.event_name }}
PKG_NAME: gitnexus
run: |
set -euo pipefail
# 1. Current published `latest` — the floor for any new rc base.
# Only E404 ("never published") falls back to package.json; any
# other error (network, auth, malformed response) fails fast.
NPM_STDERR_LATEST="$(mktemp)"
if CURRENT_LATEST="$(npm view "$PKG_NAME" version 2>"$NPM_STDERR_LATEST")"; then
:
else
if grep -q 'E404' "$NPM_STDERR_LATEST"; then
CURRENT_LATEST="$(node -p "require('./package.json').version")"
echo "Package not on registry (E404) — seeding from package.json: $CURRENT_LATEST"
else
echo "::error::npm registry unreachable for 'view version':" >&2
cat "$NPM_STDERR_LATEST" >&2
rm -f "$NPM_STDERR_LATEST"
exit 1
fi
fi
rm -f "$NPM_STDERR_LATEST"
CURRENT_LATEST_CLEAN="${CURRENT_LATEST%%-*}"
# 2. Full version list — needed for the counter and for active-cycle
# inference. Same E404-only fallback.
NPM_STDERR_VERSIONS="$(mktemp)"
if VERSIONS_JSON="$(npm view "$PKG_NAME" versions --json 2>"$NPM_STDERR_VERSIONS")"; then
:
else
if grep -q 'E404' "$NPM_STDERR_VERSIONS"; then
VERSIONS_JSON='[]'
echo "No published versions for $PKG_NAME yet (E404)."
else
echo "::error::npm registry unreachable for 'view versions':" >&2
cat "$NPM_STDERR_VERSIONS" >&2
rm -f "$NPM_STDERR_VERSIONS"
exit 1
fi
fi
rm -f "$NPM_STDERR_VERSIONS"
# 3. Base selection.
# - workflow_dispatch + bump ∈ {patch,minor,major} → explicit cycle
# reset from latest.
# - Everything else (push, or dispatch with bump=auto) → continue
# the highest active rc base > latest if one exists; else
# default to patch from latest.
if [ "$EVENT_NAME" = "workflow_dispatch" ] \
&& [ -n "${BUMP_INPUT:-}" ] \
&& [ "${BUMP_INPUT:-auto}" != "auto" ]; then
BASE="$(npx --yes -p semver@7 semver -i "$BUMP_INPUT" "$CURRENT_LATEST_CLEAN")"
echo "Explicit bump=$BUMP_INPUT → BASE=$BASE"
else
cat > /tmp/active_base.mjs <<'NODESCRIPT'
const latest = process.env.LATEST;
let v;
try { v = JSON.parse(process.env.VERSIONS_JSON); } catch { v = []; }
if (!Array.isArray(v)) v = [v];
const parse = s => s.split(".").map(n => parseInt(n, 10));
const gt = (a, b) => {
const [A, B] = [parse(a), parse(b)];
for (let i = 0; i < 3; i++) if (A[i] !== B[i]) return A[i] > B[i];
return false;
};
const bases = new Set();
for (const s of v) {
const m = /^(\d+\.\d+\.\d+)-rc\.\d+$/.exec(s);
if (m && gt(m[1], latest)) bases.add(m[1]);
}
if (!bases.size) { process.stdout.write(""); process.exit(0); }
const sorted = [...bases].sort((a, b) => gt(a, b) ? 1 : -1);
process.stdout.write(sorted[sorted.length - 1]);
NODESCRIPT
ACTIVE_BASE="$(LATEST="$CURRENT_LATEST_CLEAN" VERSIONS_JSON="$VERSIONS_JSON" node /tmp/active_base.mjs)"
if [ -n "$ACTIVE_BASE" ]; then
BASE="$ACTIVE_BASE"
echo "Continuing active rc cycle → BASE=$BASE"
else
BASE="$(npx --yes -p semver@7 semver -i patch "$CURRENT_LATEST_CLEAN")"
echo "No active rc cycle → patch bump from latest → BASE=$BASE"
fi
fi
# 4. Counter: 1 + max existing N for `${BASE}-rc.*`, else 1.
cat > /tmp/next_rc.mjs <<'NODESCRIPT'
const base = process.env.BASE;
const prefix = base + "-rc.";
let v;
try { v = JSON.parse(process.env.VERSIONS_JSON); } catch { v = []; }
if (!Array.isArray(v)) v = [v];
const ns = v
.filter(s => typeof s === "string" && s.startsWith(prefix))
.map(s => parseInt(s.slice(prefix.length), 10))
.filter(n => Number.isInteger(n) && n >= 0);
process.stdout.write(String(ns.length ? Math.max(...ns) + 1 : 1));
NODESCRIPT
NEXT_N="$(BASE="$BASE" VERSIONS_JSON="$VERSIONS_JSON" node /tmp/next_rc.mjs)"
RC_VERSION="${BASE}-rc.${NEXT_N}"
echo "Computed rc: $RC_VERSION"
# 5. Defensive: if the exact version already exists on the registry
# (e.g., race with another run), abort before re-publishing.
# Same E404-only pattern used above — a transient network
# failure must fail loudly, not pretend the version is missing.
NPM_STDERR_EXISTS="$(mktemp)"
if npm view "$PKG_NAME@$RC_VERSION" version 2>"$NPM_STDERR_EXISTS" >/dev/null; then
rm -f "$NPM_STDERR_EXISTS"
echo "::error::Version $RC_VERSION already exists on npm — aborting."
exit 1
else
if grep -qiE 'E404|not found' "$NPM_STDERR_EXISTS"; then
rm -f "$NPM_STDERR_EXISTS"
# Version doesn't exist — safe to proceed.
else
echo "::error::npm registry unreachable for existence check:" >&2
cat "$NPM_STDERR_EXISTS" >&2
rm -f "$NPM_STDERR_EXISTS"
exit 1
fi
fi
{
echo "base=$BASE"
echo "rc_n=$NEXT_N"
echo "rc_version=$RC_VERSION"
} >> "$GITHUB_OUTPUT"
- name: Apply rc version in-CI
shell: bash
working-directory: gitnexus
run: |
set -euo pipefail
npm version "${{ steps.version.outputs.rc_version }}" \
--no-git-tag-version --allow-same-version
- name: Build gitnexus
run: npm run build
working-directory: gitnexus
- name: Dry-run publish
run: npm publish --dry-run --tag rc
working-directory: gitnexus
# ── Acquire the "rc lock" BEFORE publishing (fixes idempotency) ─────
# We create two tags and push them atomically:
# v<RC_VERSION> → annotated tag on a detached release commit
# whose tree contains the rewritten package.json
# (so the tag's source matches the npm tarball)
# rc/<HEAD_SHA> → lightweight tag on HEAD; the guard's dedup key
# If this push fails, nothing is published — safe.
# If this push succeeds but npm publish fails, the marker stays on
# the remote and blocks retries until an operator manually cleans up.
- name: Create and push rc tags
id: reltag
shell: bash
working-directory: gitnexus
env:
RC_VERSION: ${{ steps.version.outputs.rc_version }}
HEAD_SHA: ${{ needs.guard.outputs.head_sha }}
run: |
set -euo pipefail
VTAG="v${RC_VERSION}"
MARKER="rc/${HEAD_SHA}"
git config user.name 'github-actions[bot]'
git config user.email '41898282+github-actions[bot]@users.noreply.github.com'
# Detached release commit with the version bump — keeps `main`
# pristine but gives the v-tag a tree that matches the published
# package contents exactly (fixes release-integrity gap).
git add package.json package-lock.json 2>/dev/null || git add package.json
git commit -m "release: ${VTAG}" --allow-empty
RELEASE_SHA="$(git rev-parse HEAD)"
echo "Detached release commit: $RELEASE_SHA"
# Annotated release tag on the release commit.
git tag -a "$VTAG" "$RELEASE_SHA" -m "$VTAG"
# Lightweight marker on the user-visible HEAD for the guard.
git tag "$MARKER" "$HEAD_SHA"
# Atomic push of both refs. If either would clobber an existing
# remote ref, the push fails and we stop before npm publish.
git push --atomic origin "refs/tags/$VTAG" "refs/tags/$MARKER"
{
echo "vtag=$VTAG"
echo "marker=$MARKER"
echo "release_sha=$RELEASE_SHA"
} >> "$GITHUB_OUTPUT"
- name: Publish to npm (rc dist-tag)
run: npm publish --provenance --access public --tag rc
working-directory: gitnexus
env:
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Create GitHub prerelease
uses: softprops/action-gh-release@b4309332981a82ec1c5618f44dd2e27cc8bfbfda # v2
with:
tag_name: ${{ steps.reltag.outputs.vtag }}
name: Release Candidate ${{ steps.reltag.outputs.vtag }}
prerelease: true
make_latest: 'false'
generate_release_notes: true
body: |
Automated release candidate build from `main`.
**npm:** `npm install gitnexus@rc`
**Version:** `${{ steps.version.outputs.rc_version }}`
**Target base:** `${{ steps.version.outputs.base }}` (rc #${{ steps.version.outputs.rc_n }})
**Source commit (main):** ${{ needs.guard.outputs.head_sha }}
**Release commit (versioned tree):** ${{ steps.reltag.outputs.release_sha }}
Release candidates are pre-stable builds intended for early testing.
Stable releases remain on the `latest` dist-tag.
# ── Build & push RC Docker images ────────────────────────────────────
# Calls docker.yml as a reusable workflow so that the build, signing, and
# attestation logic stays in one place. The publish job exposes `vtag`
# (e.g. `v1.2.3-rc.1`) as an output so we can pass it as the tag input.
# RC images are signed with Cosign keyless signing; the OIDC identity
# will be `docker.yml@refs/heads/main` (the caller's ref) rather than a
# tag ref — see README.md § Docker for the correct verify command for RCs.
docker:
name: Build & Push RC Docker images
needs: [guard, publish]
if: needs.guard.outputs.should_run == 'true' && needs.publish.outputs.vtag != ''
uses: ./.github/workflows/docker.yml
# Reusable workflows do not receive caller secrets unless inherited; without
# this, DOCKERHUB_* / GITHUB_TOKEN are empty in docker.yml → "Username and
# password required" on Docker Hub login (see same pattern on `ci:` above).
secrets: inherit
permissions:
contents: read
packages: write
id-token: write
attestations: write
with:
tag: ${{ needs.publish.outputs.vtag }}
+1 -1
View File
@@ -53,6 +53,6 @@ jobs:
retention-days: 5
- name: Upload to Security tab
uses: github/codeql-action/upload-sarif@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/upload-sarif@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
sarif_file: results.sarif
+1 -1
View File
@@ -76,7 +76,7 @@ jobs:
exit-code: '0'
- name: Upload to Security tab
uses: github/codeql-action/upload-sarif@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/upload-sarif@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
sarif_file: trivy-${{ matrix.image.name }}.sarif
category: trivy-${{ matrix.image.name }}
+1 -1
View File
@@ -76,7 +76,7 @@ jobs:
continue-on-error: true
- name: Upload SARIF
uses: github/codeql-action/upload-sarif@e46ed2cbd01164d986452f91f178727624ae40d7 # v4.35.3
uses: github/codeql-action/upload-sarif@68bde559dea0fdcac2102bfdf6230c5f70eb485e # v4.35.4
with:
sarif_file: zizmor.sarif
category: zizmor
+5 -4
View File
@@ -37,7 +37,8 @@ rules:
- pr-labeler.yml
# Note: cache-poisoning is NOT exempted. The two prior findings in
# publish.yml and release-candidate.yml were fixed structurally by
# dropping `cache: npm` from those workflows (matches the pattern used
# by PyO3/maturin for the same audit). See the commit that added this
# file for the rationale.
# publish.yml and the former release-candidate.yml were fixed structurally
# by dropping `cache: npm` from those workflows (matches the pattern used
# by PyO3/maturin for the same audit). After the publish-workflow
# unification (issue #1609), only publish.yml remains; the same
# cache-poisoning hardening applies there.
+2 -2
View File
@@ -68,8 +68,8 @@ gitnexus-web/test-results/
eval/.coverage
eval/.hypothesis/
# Design docs (local only)
docs/plans/
# Local docs
docs/
gitnexus/test/fixtures/mini-repo/*.md
gitnexus/test/fixtures/mini-repo/.claude
+42 -108
View File
@@ -48,6 +48,7 @@ Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.
| Date | Version | Change |
|------|---------|--------|
| 2026-05-22 | 1.8.0 | Kotlin added to `MIGRATED_LANGUAGES` (registry-primary call resolution by default). Closes #1756 (companion-vs-instance dispatch) and #1757 (lambda scopes); refs #1746. RFC §6.4 corpus criterion waived (corpus-mode wiring is #927-scope); fixture criterion met. |
| 2026-04-23 | 1.7.0 | TypeScript added to `MIGRATED_LANGUAGES` (registry-primary call resolution by default). |
| 2026-04-20 | 1.6.0 | Added scope-resolution pipeline pointer (RFC #909 Ring 3); Python migrated to registry-primary. |
| 2026-04-19 | 1.5.0 | Cross-repo impact (#794): `impact`/`query`/`context` accept `repo: "@<group>"` + `service`. Removed `group_query`/`group_contracts`/`group_status` MCP tools; added `gitnexus://group/{name}/contracts` and `gitnexus://group/{name}/status` resources. |
@@ -62,131 +63,64 @@ Commands and gotchas live under **Repo reference** below and in **[CONTRIBUTING.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
Indexed as **GitNexus** (4325 symbols, 10556 relationships, 300 execution flows). Use MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any tool warns the index is stale, run `npx gitnexus analyze` first.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** `gitnexus_impact({target: "symbolName", direction: "upstream"})` — report blast radius to the user.
- **MUST run `gitnexus_detect_changes()` before committing** — verify only expected symbols and flows are affected.
- **MUST warn the user** if impact returns HIGH or CRITICAL risk.
- Explore unfamiliar code with `gitnexus_query({query: "concept"})` (process-grouped, ranked) instead of grepping.
- Full context on a symbol: `gitnexus_context({name: "symbolName"})`.
## When Debugging
1. `gitnexus_query({query: "<error or symptom>"})` — find related execution flows
2. `gitnexus_context({name: "<suspect function>"})` — callers, callees, process participation
3. `READ gitnexus://repo/GitNexus/process/{processName}` — trace flow step by step
4. Regressions: `gitnexus_detect_changes({scope: "compare", base_ref: "main"})`
## When Refactoring
- **Rename:** `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` first. Graph edits are safe; text_search edits need manual review.
- **Extract/Split:** `gitnexus_context` (incoming/outgoing refs) then `gitnexus_impact` (upstream callers) before moving code.
- **After any refactor:** `gitnexus_detect_changes({scope: "all"})` to verify scope.
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Never Do
- Edit a symbol without running `gitnexus_impact` first.
- Ignore HIGH/CRITICAL risk warnings.
- Rename with find-and-replace — use `gitnexus_rename`.
- Commit without `gitnexus_detect_changes()`.
- Add language-specific behavior to shared ingestion code (`gitnexus/src/core/ingestion/`) — use a `LanguageProvider` hook. Seeing `provider.mroStrategy === 'xxx'` or an import from `languages/xxx.ts` in shared code means stop and add a hook.
## Tools Quick Reference
| Tool | When to use | Example |
|------|-------------|---------|
| `list_repos` | Discover indexed repos | `gitnexus_list_repos({})` |
| `query` | Find code by concept | `gitnexus_query({query: "auth validation"})` |
| `context` | 360-degree view of one symbol | `gitnexus_context({name: "validateUser"})` |
| `impact` | Blast radius before editing | `gitnexus_impact({target: "X", direction: "upstream"})` |
| `detect_changes` | Pre-commit scope check | `gitnexus_detect_changes({scope: "staged"})` |
| `rename` | Safe multi-file rename | `gitnexus_rename({symbol_name: "old", new_name: "new", dry_run: true})` |
| `cypher` | Custom graph queries | `gitnexus_cypher({query: "MATCH ..."})` |
| `api_impact` | Pre-change API route impact | `gitnexus_api_impact({route: "/api/users", method: "GET"})` |
| `route_map` | Route → handler → consumer map | `gitnexus_route_map({})` |
| `tool_map` | MCP/RPC tool definitions | `gitnexus_tool_map({})` |
| `shape_check` | Response shape vs consumer access | `gitnexus_shape_check({route: "/api/users"})` |
| `group_list` | List repo groups | `gitnexus_group_list({})` |
| `group_sync` | Rebuild group Contract Registry | `gitnexus_group_sync({name: "myGroup"})` |
| `query` (group mode) | Cross-repo search in a group (RRF-merged) | `gitnexus_query({repo: "@myGroup", query: "auth"})` |
| `context` (group mode) | 360° view across all member repos | `gitnexus_context({repo: "@myGroup", name: "validateUser"})` |
| `impact` (group mode) | Cross-repo blast radius via Contract Bridge | `gitnexus_impact({repo: "@myGroup", target: "X", direction: "upstream"})` |
> Group mode: pass `repo: "@<groupName>"` to fan out across all member repos, or `repo: "@<groupName>/<memberPath>"` to target a single member (path keys from `group.yaml`). Optional `service: "<monorepo/path>"` filters by service root. Group-level state (contracts, staleness) lives in the resources table below — there are **no** `group_query` / `group_context` / `group_impact` / `group_contracts` / `group_status` MCP tools.
>
> For a full walkthrough of setting up a group across multiple repos that communicate over gRPC, see [docs/guides/microservices-grpc.md](docs/guides/microservices-grpc.md).
## Impact Risk Levels
| Depth | Meaning | Action |
|-------|---------|--------|
| d=1 | WILL BREAK — direct callers/importers | MUST update |
| d=2 | LIKELY AFFECTED — indirect deps | Should test |
| d=3 | MAY NEED TESTING — transitive | Test if critical path |
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, index freshness |
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
| `gitnexus://group/{name}/contracts` | Group Contract Registry (provider/consumer rows + cross-links) |
| `gitnexus://group/{name}/status` | Per-member index + Contract Registry staleness report |
## Self-Check Before Finishing
## CLI
1. `gitnexus_impact` was run for all modified symbols
2. No HIGH/CRITICAL warnings were ignored
3. `gitnexus_detect_changes()` confirms expected scope
4. All d=1 dependents were updated
## Keeping the Index Fresh
```bash
npx gitnexus analyze # incremental by default; preserves embeddings
npx gitnexus analyze --force # full rebuild from scratch (opt out of incremental)
npx gitnexus analyze --embeddings # also generate embeddings for new/changed nodes
npx gitnexus analyze --drop-embeddings # explicit opt-in to wipe existing embeddings
```
`analyze` runs **incrementally by default**. The pipeline still parses every file every run (cross-file resolution requires it), but tree-sitter parsing is **served from a content-addressed cache** at `.gitnexus/parse-cache.json` for chunks whose file contents haven't changed since the last run. Only changed-file rows (and their importers) are rewritten in LadybugDB; unchanged-file rows are preserved. Output is byte-equivalent to a full rebuild. Pass `--force` to wipe and re-index from scratch (e.g., to recover from a corrupt index, or after upgrading GitNexus).
The parse cache key is **content-addressed and version-tagged**: it survives `--force` runs, and is automatically invalidated by a `gitnexus` package upgrade (so a new tree-sitter grammar doesn't silently replay stale parse output). Safe to delete `.gitnexus/parse-cache.json` at any time — it'll be rebuilt on the next analyze.
Check `.gitnexus/meta.json` `stats.embeddings` (0 = none). A plain `analyze` no longer drops existing vectors — pass `--drop-embeddings` to wipe.
> Claude Code: PostToolUse hook detects a stale index after `git commit` and `git merge` and prompts the agent to run `analyze`. The hook does not invoke `analyze` itself.
## CLI Skills
| Task | Skill file |
|------|-----------|
| Architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Debugging / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Refactoring | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools/resources/schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| CLI commands (index, status, clean, wiki) | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
## Hook env knobs
The Claude Code hook (`gitnexus/hooks/claude/gitnexus-hook.cjs` and the mirrored plugin copy under `gitnexus-claude-plugin/hooks/`) honours these env vars. Defaults work for normal installations; set them only to override resolution. All path overrides ignore values that do not exist on disk and fall through to the standard resolution chain.
| Env var | Type | Default | Purpose |
|---------|------|---------|---------|
| `GITNEXUS_HOOK_CLI_PATH` | path | resolved via package layout / `require.resolve` | Override path to the `gitnexus` CLI entry the hook spawns for `augment`. |
| `GITNEXUS_HOOK_LSOF_PATH` | path | `lsof` on `PATH` (with `/usr/bin/lsof`, `/usr/sbin/lsof`, `/sbin/lsof` fallbacks) | Override POSIX `lsof` location for the DB-lock probe. |
| `GITNEXUS_HOOK_PS_PATH` | path | `ps` on `PATH` (with `/bin/ps`, `/usr/bin/ps` fallbacks) | Override POSIX `ps` location. |
| `GITNEXUS_HOOK_POWERSHELL_PATH` | path | `%SystemRoot%\System32\WindowsPowerShell\v1.0\powershell.exe` (then `SysWOW64`, then `powershell.exe` on `PATH`) | Override Windows PowerShell location used by the Restart-Manager probe. |
| `GITNEXUS_HOOK_LINUX_PROC_BUDGET_MS` | integer ms | `1200` | Max wall-clock for the Linux `/proc` fd scan before bailing out to the `lsof` fallback. |
| `GITNEXUS_HOOK_RM_TARGET` | path | derived | Restart-Manager target file (the LadybugDB path under `.gitnexus/`). Set internally by the hook; rarely overridden manually. |
| `GITNEXUS_DEBUG` | boolean (`1`/`true`) | unset | Verbose stderr from the hook: prints discarded augment-stderr prefixes and one-shot `.ps1` load-failure warnings. |
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (239 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Extractors area (135 symbols) | `.claude/skills/generated/extractors/SKILL.md` |
| Work in the Components area (112 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Lbug area (96 symbols) | `.claude/skills/generated/lbug/SKILL.md` |
| Work in the Group area (94 symbols) | `.claude/skills/generated/group/SKILL.md` |
| Work in the Cli area (92 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Configs area (92 symbols) | `.claude/skills/generated/configs/SKILL.md` |
| Work in the Type-extractors area (90 symbols) | `.claude/skills/generated/type-extractors/SKILL.md` |
| Work in the Hooks area (88 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Unit area (80 symbols) | `.claude/skills/generated/unit/SKILL.md` |
| Work in the Cpp area (73 symbols) | `.claude/skills/generated/cpp/SKILL.md` |
| Work in the Scope-resolution area (72 symbols) | `.claude/skills/generated/scope-resolution/SKILL.md` |
| Work in the Server area (66 symbols) | `.claude/skills/generated/server/SKILL.md` |
| Work in the Local area (61 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Wiki area (60 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Workers area (57 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Embeddings area (56 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Typescript area (53 symbols) | `.claude/skills/generated/typescript/SKILL.md` |
| Work in the Storage area (51 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Php area (48 symbols) | `.claude/skills/generated/php/SKILL.md` |
<!-- gitnexus:end -->
+64
View File
@@ -52,3 +52,67 @@ If always-on instructions grow, load deep conventions via conditional reads (e.g
## GitNexus rules
See the `<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in **[AGENTS.md](AGENTS.md)** for the canonical MCP tools, impact analysis rules, and index instructions.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/GitNexus/context` | Codebase overview, check index freshness |
| `gitnexus://repo/GitNexus/clusters` | All functional areas |
| `gitnexus://repo/GitNexus/processes` | All execution flows |
| `gitnexus://repo/GitNexus/process/{name}` | Step-by-step execution trace |
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
| Work in the Ingestion area (239 symbols) | `.claude/skills/generated/ingestion/SKILL.md` |
| Work in the Extractors area (135 symbols) | `.claude/skills/generated/extractors/SKILL.md` |
| Work in the Components area (112 symbols) | `.claude/skills/generated/components/SKILL.md` |
| Work in the Lbug area (96 symbols) | `.claude/skills/generated/lbug/SKILL.md` |
| Work in the Group area (94 symbols) | `.claude/skills/generated/group/SKILL.md` |
| Work in the Cli area (92 symbols) | `.claude/skills/generated/cli/SKILL.md` |
| Work in the Configs area (92 symbols) | `.claude/skills/generated/configs/SKILL.md` |
| Work in the Type-extractors area (90 symbols) | `.claude/skills/generated/type-extractors/SKILL.md` |
| Work in the Hooks area (88 symbols) | `.claude/skills/generated/hooks/SKILL.md` |
| Work in the Unit area (80 symbols) | `.claude/skills/generated/unit/SKILL.md` |
| Work in the Cpp area (73 symbols) | `.claude/skills/generated/cpp/SKILL.md` |
| Work in the Scope-resolution area (72 symbols) | `.claude/skills/generated/scope-resolution/SKILL.md` |
| Work in the Server area (66 symbols) | `.claude/skills/generated/server/SKILL.md` |
| Work in the Local area (61 symbols) | `.claude/skills/generated/local/SKILL.md` |
| Work in the Wiki area (60 symbols) | `.claude/skills/generated/wiki/SKILL.md` |
| Work in the Workers area (57 symbols) | `.claude/skills/generated/workers/SKILL.md` |
| Work in the Embeddings area (56 symbols) | `.claude/skills/generated/embeddings/SKILL.md` |
| Work in the Typescript area (53 symbols) | `.claude/skills/generated/typescript/SKILL.md` |
| Work in the Storage area (51 symbols) | `.claude/skills/generated/storage/SKILL.md` |
| Work in the Php area (48 symbols) | `.claude/skills/generated/php/SKILL.md` |
<!-- gitnexus:end -->
+57 -27
View File
@@ -144,16 +144,18 @@ If you use coding agents, follow project context files (e.g. `AGENTS.md`, `CLAUD
## Releases
Two publish workflows ship `gitnexus` to npm:
One workflow ships `gitnexus` to npm — `.github/workflows/publish.yml`. It
routes between two modes based on the triggering event:
- **Stable** (`.github/workflows/publish.yml`) — triggered by pushing any `v*`
tag. Publishes to the `latest` dist-tag with a changelog-backed GitHub
release. Maintainers are expected to tag from `main` as a convention; the
workflow itself does not enforce branch reachability.
- **Release Candidate** (`.github/workflows/release-candidate.yml`) — runs on
every push to `main` (typically a merged PR) plus manual dispatch. Docs-only
changes are skipped via `paths-ignore`. Publishes to the `rc` dist-tag with
version `X.Y.Z-rc.N` and a GitHub prerelease, where:
- **Stable mode** — triggered by pushing any `v<X.Y.Z>` tag (no `-rc.*`
suffix; RC tags are excluded at trigger via a negative glob). Publishes to
the `latest` dist-tag with a changelog-backed GitHub release. Maintainers
are expected to tag from `main` as a convention; the workflow itself does
not enforce branch reachability. No Docker build (RC-only).
- **Release-candidate mode** — runs on every push to `main` (typically a
merged PR) plus manual `workflow_dispatch`. Docs-only changes are skipped
via `paths-ignore`. Publishes to the `rc` dist-tag with version
`X.Y.Z-rc.N` and a GitHub prerelease, where:
- `X.Y.Z` is selected automatically. On push (and on dispatch with
`bump: auto`, the default) the workflow **continues the active rc cycle**:
if the registry already has `X.Y.Z-rc.*` versions with `X.Y.Z` > current
@@ -170,36 +172,64 @@ Two publish workflows ship `gitnexus` to npm:
caller's ref — see README.md § Docker for the verify command).
Idempotency: the workflow pushes an `rc/<HEAD_SHA>` marker tag and a
`v<RC>` release tag **atomically, before** calling `npm publish`. The guard
refuses to re-run once the marker exists, so a post-publish failure will
not mint a duplicate rc for the same commit. The `v<RC>` tag points at a
detached release commit whose `package.json` matches the npm tarball
exactly (traceable releases). Recovery after a partial failure:
`v<RC>` release tag **atomically, before** calling `npm publish`. The
RC guard refuses to re-run once the marker exists, so a post-publish
failure will not mint a duplicate rc for the same commit. The `v<RC>`
tag points at a detached release commit whose `package.json` matches
the npm tarball exactly (traceable releases). The RC tag is excluded
from this workflow's `push: tags:` filter, so it does **not** re-trigger
publishing — preventing the double-publish failure mode tracked in #1609.
Recovery after a partial failure: the workflow's `if: failure()` cleanup
step in the `publish` job auto-deletes the v-tag and marker on most
post-publish failures, so the typical retry is just:
```bash
gh workflow run publish.yml --ref main -f force=true
# or push a new commit to main, which will cut a fresh RC
```
If auto-cleanup didn't run (e.g. the cleanup step itself failed, or the
failure happened in the route/rc-guard phase before the marker was
pushed), manual cleanup is:
```bash
git push --delete origin rc/<HEAD_SHA> v<RC>
# then redispatch the workflow with force: true
# then redispatch with force: true
```
**Release-PR-skip subject pattern.** The rc-guard job recognizes a
squash-merged release commit by matching the commit subject against
`^chore: release vX.Y.Z` (optionally followed by ` (#NNNN)` for the
squash-merge PR-number suffix). Match is case-insensitive — `Chore: Release v1.2.3`
works too. PRs that should suppress the RC build must either use this
subject shape, or carry the `release` label so the label-based fallback
fires. Other release-style subjects (`chore(release): v1.2.3`,
`release: v1.2.3`) will NOT trigger the skip — please name the release
PR exactly `chore: release vX.Y.Z` to keep the dedup deterministic.
**Docker-only partial failure:** if `publish` succeeds (npm tarball + tags
are live) but the `docker` job subsequently fails (e.g. GHCR flakiness),
the npm RC is already published and the `rc/<HEAD_SHA>` marker is in place.
Re-running `release-candidate.yml` with `force: true` will abort at the
"Version already exists on npm" guard. To recover without cutting a new RC:
Recovery without cutting a new RC:
```bash
# 1. Manually trigger only the docker workflow, passing the existing RC tag:
gh workflow run docker.yml --ref main -f tag=v<RC_VERSION>
# (requires a workflow_dispatch trigger on docker.yml — see note below)
# Re-run only the failed docker job from the original workflow run:
gh run rerun <run-id> --failed
```
Because `docker.yml` intentionally has no `workflow_dispatch` (images are
tag-driven by design), the practical recovery options are:
- Wait for the next commit on `main`, which will cut a new RC that includes
the Docker build.
- Manually run `docker build` + `docker push` locally and sign with Cosign
against the same digest.
- Delete `rc/<HEAD_SHA>` and `v<RC>` tags, then redispatch with `force: true` to re-run the full RC pipeline (cuts a new RC number).
Find the run ID via `gh run list --workflow=publish.yml --branch main`.
`docker.yml` intentionally has no `workflow_dispatch` trigger (images are
tag-driven by design), so the gh-run-rerun path is the supported recovery.
**GitHub Release transient failure** (npm publish succeeded, Release step
failed): the npm artifact is live but no GitHub Release page exists.
Recover by either re-running the failed job (`gh run rerun <run-id> --failed`),
or creating the Release manually:
```bash
gh release create v<RC> --prerelease --generate-notes # RC
gh release create v<X.Y.Z> --notes-file gitnexus/CHANGELOG.md # stable
```
The rc workflow never moves `latest`. To verify after a change, inspect dist-tags:
+1 -1
View File
@@ -36,7 +36,7 @@ Format: **Trigger → Instruction → Reason**. Append new Signs when the same m
### Index seems corrupt or "incremental" is misbehaving
- **Trigger:** `analyze` produces unexpected results, or `meta.json.incrementalInProgress` is set, or the index is in a half-state after a crash.
- **Do:** `npx gitnexus analyze --force` to rebuild from scratch. The dirty-flag check forces this automatically when a previous incremental run didn't complete cleanly, but `--force` is the manual escape hatch. Safe to delete `.gitnexus/parse-cache.json` at any time — content-addressed, will be regenerated.
- **Do:** `npx gitnexus analyze --force` to rebuild from scratch. The dirty-flag check forces this automatically when a previous incremental run didn't complete cleanly, but `--force` is the manual escape hatch. Safe to delete the `.gitnexus/parse-cache/` directory (and any legacy `.gitnexus/parse-cache.json`) at any time — content-addressed, will be regenerated.
- **Why:** Incremental writeback is selective DB row replacement; if the on-disk state is inconsistent for any reason, a full rebuild is the cheapest path back to a known-good index.
### Embeddings vanished after analyze
+106 -82
View File
@@ -1,4 +1,5 @@
# GitNexus
**⚠️ Important Notice:** GitNexus has NO official cryptocurrency, token, or coin. Any token/coin using the GitNexus name on Pump.fun or any other platform is **not affiliated with, endorsed by, or created by** this project or its maintainers. Do not purchase any cryptocurrency claiming association with GitNexus.
<div align="center">
@@ -30,14 +31,9 @@
Indexes any codebase into a knowledge graph — every dependency, call chain, cluster, and execution flow — then exposes it through smart tools so AI agents never miss code.
https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
> *Like DeepWiki, but deeper.* DeepWiki helps you *understand* code. GitNexus lets you *analyze* it — because a knowledge graph tracks every relationship, not just descriptions.
> _Like DeepWiki, but deeper._ DeepWiki helps you _understand_ code. GitNexus lets you _analyze_ it — because a knowledge graph tracks every relationship, not just descriptions.
**TL;DR:** The **Web UI** is a quick way to chat with any repo. The **CLI + MCP** is how you make your AI agent actually reliable — it gives Cursor, Claude Code, Codex, and friends a deep architectural view of your codebase so they stop missing dependencies, breaking call chains, and shipping blind edits. Even smaller models get full architectural clarity, making it compete with Goliath models.
@@ -47,18 +43,17 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
[![Star History Chart](https://api.star-history.com/svg?repos=abhigyanpatwari/GitNexus&type=date&legend=top-left)](https://www.star-history.com/#abhigyanpatwari/GitNexus&type=date&legend=top-left)
## Two Ways to Use GitNexus
| | **CLI + MCP** | **Web UI** |
| ----------------- | -------------------------------------------------------------- | ------------------------------------------------------------ |
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
| **For** | Daily development with Cursor, Claude Code, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Privacy** | Everything local, no network | Everything in-browser, no server |
| | **CLI + MCP** | **Web UI** |
| ----------- | --------------------------------------------------------------------- | -------------------------------------------------------------------- |
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
| **For** | Daily development with Cursor, Claude Code, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Privacy** | Everything local, no network | Everything in-browser, no server |
> **Bridge mode:** `gitnexus serve` connects the two — the web UI auto-detects the local server and can browse all your CLI-indexed repos without re-uploading or re-indexing.
@@ -69,6 +64,7 @@ https://github.com/user-attachments/assets/172685ba-8e54-4ea7-9ad1-e31a3398da72
GitNexus is available as an **enterprise offering** - either as a fully managed **SaaS** or a **self-hosted** deployment. Also available for **commercial use** of the OSS version with proper licensing.
Enterprise includes:
- **PR Review** - automated blast radius analysis on pull requests
- **Auto-updating Code Wiki** - always up-to-date documentation (Code Wiki is also available in OSS)
- **Auto-reindexing** - knowledge graph stays fresh automatically
@@ -77,6 +73,7 @@ Enterprise includes:
- **Priority feature/language support** - request new languages or features
**Upcoming:**
- Auto regression forensics
- End-to-end test generation
@@ -109,7 +106,7 @@ That's it. This indexes the codebase, installs agent skills, registers Claude Co
To configure MCP for your editor, run `npx gitnexus setup` once — or set it up manually below.
> **Faster install (no C++ toolchain needed):** set `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` before `npm install -g gitnexus` to skip the native `tree-sitter-dart` and `tree-sitter-proto` builds. Dart/Proto files won't be parsed, but install completes in seconds without `python3`/`make`/`g++`. Strict `=1` only — any other value falls through to the rebuild.
> **Faster install (no C++ toolchain needed):** set `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` before `npm install -g gitnexus` to skip vendored grammar materialize/build (`tree-sitter-dart`, `tree-sitter-proto`, `tree-sitter-swift`). Dart/Proto/Swift files won't be parsed, but install completes in seconds without `python3`/`make`/`g++`. Strict `=1` only — any other value falls through to the rebuild.
### MCP Setup
@@ -117,13 +114,13 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
### Editor Support
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
| --------------------- | --- | ------ | -------------------- | -------------- |
| **Claude Code** | Yes | Yes | Yes (PreToolUse + PostToolUse) | **Full** |
| **Cursor** | Yes | Yes | Yes (postToolUse, [manual install](gitnexus-cursor-integration/README.md#hook-install)) | **Full** |
| **Codex** | Yes | Yes | — | MCP + Skills |
| **Windsurf** | Yes | — | — | MCP |
| **OpenCode** | Yes | Yes | — | MCP + Skills |
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
| --------------- | --- | ------ | --------------------------------------------------------------------------------------- | ------------ |
| **Claude Code** | Yes | Yes | Yes (PreToolUse + PostToolUse) | **Full** |
| **Cursor** | Yes | Yes | Yes (postToolUse, [manual install](gitnexus-cursor-integration/README.md#hook-install)) | **Full** |
| **Codex** | Yes | Yes | — | MCP + Skills |
| **Windsurf** | Yes | — | — | MCP |
| **OpenCode** | Yes | Yes | — | MCP + Skills |
> **Claude Code** gets the deepest integration: MCP tools + agent skills + PreToolUse hooks that enrich searches with graph context + PostToolUse hooks that detect a stale index after commits and prompt the agent to reindex.
@@ -131,10 +128,10 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
Built by the community — not officially maintained, but worth checking out.
| Project | Author | Description |
|---------|--------|-------------|
| [pi-gitnexus](https://github.com/tintinweb/pi-gitnexus) | [@tintinweb](https://github.com/tintinweb) | GitNexus plugin for [pi](https://pi.dev) — `pi install npm:pi-gitnexus` |
| [gitnexus-stable-ops](https://github.com/ShunsukeHayashi/gitnexus-stable-ops) | [@ShunsukeHayashi](https://github.com/ShunsukeHayashi) | Stable ops & deployment workflows (Miyabi ecosystem) |
| Project | Author | Description |
| ----------------------------------------------------------------------------- | ------------------------------------------------------ | ----------------------------------------------------------------------- |
| [pi-gitnexus](https://github.com/tintinweb/pi-gitnexus) | [@tintinweb](https://github.com/tintinweb) | GitNexus plugin for [pi](https://pi.dev) — `pi install npm:pi-gitnexus` |
| [gitnexus-stable-ops](https://github.com/ShunsukeHayashi/gitnexus-stable-ops) | [@ShunsukeHayashi](https://github.com/ShunsukeHayashi) | Stable ops & deployment workflows (Miyabi ecosystem) |
> Have a project built on GitNexus? Open a PR to add it here!
@@ -197,7 +194,8 @@ args = ["-y", "gitnexus@latest", "mcp"]
```bash
gitnexus setup # Configure MCP for your editors (one-time)
gitnexus analyze [path] # Index a repository (or update stale index)
gitnexus analyze --force # Force full re-index
gitnexus analyze --repair-fts # Fast path: rebuild/verify only FTS indexes on existing index data
gitnexus analyze --force # Full rebuild: re-parse + graph rebuild + FTS rebuild
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
@@ -205,6 +203,8 @@ gitnexus analyze --skip-git # Index folders that are not Git repositories
gitnexus analyze --embeddings # Enable embedding generation (slower, better search)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus analyze --worker-timeout 60 # Increase worker idle timeout for slow parses
gitnexus analyze --wal-checkpoint-threshold 67108864 # 64 MiB. Control LadybugDB WAL auto-checkpoint threshold (default: 67108864 = 64 MiB; -1 keeps Ladybug stock ~16 MiB)
gitnexus analyze --workers <n> # Parse worker pool size (default: cores-1, capped at 16; 0 = sequential)
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
gitnexus serve # Start local HTTP server (multi-repo) for web UI connection
gitnexus list # List all indexed repositories
@@ -229,6 +229,28 @@ gitnexus group status <name> # Check staleness of repos in a group
If `analyze` reports a worker parse timeout on a large or unusual repository, it keeps running and falls back safely. To give slow worker jobs more time, use `gitnexus analyze --worker-timeout 60` or set `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=60000`. For very large files, `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` controls the worker job byte budget.
#### Environment variables
Most `analyze` knobs are also CLI flags (`--workers`, `--worker-timeout`, `--max-file-size`, `--verbose`). Use the env-var form when you'd otherwise repeat the same flag every run, or when invoking GitNexus from a long-running host (MCP server, eval-server, CI shell) that already manages its own environment. CLI flags take precedence over env vars; env vars take precedence over built-in defaults.
| Variable | Default | Effect | Tune when… |
| -------------------------------------- | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
| `GITNEXUS_WORKER_POOL_SIZE` | `cores - 1`, capped at 16 | Parse worker pool size. `0` disables the pool (sequential fallback). Equivalent to `--workers <n>`. | Constrained containers (cgroup CPU limits), CI runners with explicit quotas, or debugging a worker-only crash via `0`. |
| `GITNEXUS_PARSE_CHUNK_CONCURRENCY` | `2` | Number of chunks whose file contents may be read into memory in parallel while the pool dispatches the current chunk. Worker dispatch itself stays serial. | Repos large enough to chunk (multi-MB total source) where disk I/O is a measurable fraction of analyze wall-clock. |
| `GITNEXUS_VERBOSE` | unset | When `1`, enables verbose ingestion logs (skipped-file warnings, per-chunk throughput, parse-cache stats). Equivalent to `--verbose`. | Debugging an analyze that "completed" but seems to have missed files; tuning `--workers` / chunk concurrency against observable throughput. |
| `GITNEXUS_PROFILE_DEFERRED` | unset | When `1`, emits `[deferred-profile]` timing/progress logs for the post-chunk deferred resolution band (imports → heritage → buildHeritageMap → legacy call resolution). Implied by `GITNEXUS_VERBOSE`. | Diagnosing analyze stalls in "Resolving calls (all chunks)" on large Java/Kotlin repos (issue #1741) without the full verbose ingestion noise. |
| `GITNEXUS_PROFILE_DEFERRED_SLOW_MS` | `3000` (verbose) / `5000` | Per-file threshold in ms above which `processCallsFromExtracted` emits a `slow file …` log line. Parsed via `Number()`: accepts integers (`5000`), scientific notation (`2.5e3`), decimals (`.5`), and hex (`0x10`). Non-finite or non-positive values fall back to the default. | Hunting a few outlier files dominating the deferred call-resolution stage; lower to surface more, raise to focus only on the worst. |
| `GITNEXUS_MAX_FILE_SIZE` | `512` (KB) | Walker skip threshold in KB. Hard cap is `32768` (tree-sitter buffer ceiling). Equivalent to `--max-file-size <kb>`. | Indexing repos with intentionally-large source files (generated parsers, vendored bundles) that should still be parsed. |
| `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS` | `30000` | Worker idle timeout in milliseconds before retry/fallback. Equivalent to `--worker-timeout <seconds>` × 1000. | Slow-parsing files (large minified JS, deeply-nested TS types) that legitimately need more than 30s. |
| `GITNEXUS_WAL_CHECKPOINT_THRESHOLD` | `67108864` (64 MiB) | LadybugDB WAL auto-checkpoint threshold in bytes. Equivalent to `--wal-checkpoint-threshold <bytes>`. `-1` keeps LadybugDB's stock threshold (~16 MiB). Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. | You need a larger or smaller WAL auto-checkpoint threshold for your analyze workload. |
| `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` | `8388608` (8 MB) | Per-job byte budget the pool will send to a worker in one `postMessage`. | Very large individual files; mostly diagnostic — bumping past 8 MB risks structured-clone memory pressure. |
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per worker slot before the slot is dropped from the active rotation. Bounds respawn loops on a chronically-crashing slot. | Hosts where a flaky worker should retry more (raise) or fail-fast (lower) before the slot is dropped. |
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Combined with `timeoutBackoffFactor`, prevents exponentially-growing retries from stalling for hours. | Slow files that legitimately need long total retry windows; lower to fail-fast on stalls. |
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD`| `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, every subsequent dispatch rejects until a fresh pool is created. | Hosts where a SIGSEGV-prone native grammar should trip the breaker sooner; CI runners that should fail loudly. |
| `GITNEXUS_CHUNK_BYTE_BUDGET` | `2097152` (2 MB) | Chunk boundary used for cache-key composition and dispatch. Smaller = finer-grained cache hits but more dispatch overhead. | Tuning incremental-analyze cache behavior on monorepos. |
| `GITNEXUS_NO_GITIGNORE` | unset | When set, skips `.gitignore` parsing. `.gitnexusignore` is still honored. | Indexing a repo whose `.gitignore` excludes files you actually want indexed (e.g., generated code committed for cross-repo lookup). |
| `GITNEXUS_SKIP_OPTIONAL_GRAMMARS` | unset | When `=1` strictly, skips vendored grammar materialize/build for `tree-sitter-dart`, `tree-sitter-proto`, and `tree-sitter-swift` at install time. | Installing on a host without a C++ toolchain or where Swift prebuilds don't match; you're willing to skip Dart/Proto/Swift parsing. |
#### Publishing to understand-quickly (opt-in)
[`looptech-ai/understand-quickly`](https://github.com/looptech-ai/understand-quickly) is a public registry of code-knowledge graphs that lists `gitnexus@1` as a first-class format. After registering your repo once (`npx @understand-quickly/cli add` or the [wizard](https://looptech-ai.github.io/understand-quickly/add.html)), `gitnexus publish` fires a single `repository_dispatch` event so the registry resyncs your entry on demand instead of waiting for the nightly job.
@@ -239,27 +261,27 @@ It is opt-in and a no-op without `UNDERSTAND_QUICKLY_TOKEN` — a fine-grained G
**16 tools** exposed via MCP (11 per-repo + 5 group):
| Tool | What It Does | `repo` Param |
| ------------------ | ----------------------------------------------------------------- | -------------- |
| `list_repos` | Discover all indexed repositories | — |
| `query` | Process-grouped hybrid search (BM25 + semantic + RRF) | Optional |
| `context` | 360-degree symbol view — categorized refs, process participation | Optional |
| `impact` | Blast radius analysis with depth grouping and confidence | Optional |
| `detect_changes` | Git-diff impact — maps changed lines to affected processes | Optional |
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
| `cypher` | Raw Cypher graph queries | Optional |
| `group_list` | List configured repository groups | — |
| `group_sync` | Extract contracts and match across repos/services | — |
| `group_contracts`| Inspect extracted contracts and cross-links | — |
| `group_query` | Search execution flows across all repos in a group | — |
| `group_status` | Check staleness of repos in a group | — |
| Tool | What It Does | `repo` Param |
| ----------------- | ---------------------------------------------------------------- | ------------ |
| `list_repos` | Discover all indexed repositories | — |
| `query` | Process-grouped hybrid search (BM25 + semantic + RRF) | Optional |
| `context` | 360-degree symbol view — categorized refs, process participation | Optional |
| `impact` | Blast radius analysis with depth grouping and confidence | Optional |
| `detect_changes` | Git-diff impact — maps changed lines to affected processes | Optional |
| `rename` | Multi-file coordinated rename with graph + text search | Optional |
| `cypher` | Raw Cypher graph queries | Optional |
| `group_list` | List configured repository groups | — |
| `group_sync` | Extract contracts and match across repos/services | — |
| `group_contracts` | Inspect extracted contracts and cross-links | — |
| `group_query` | Search execution flows across all repos in a group | — |
| `group_status` | Check staleness of repos in a group | — |
> When only one repo is indexed, the `repo` parameter is optional. With multiple repos, specify which one: `query({query: "auth", repo: "my-app"})`.
**Resources** for instant context:
| Resource | Purpose |
| ----------------------------------------- | ---------------------------------------------------- |
| Resource | Purpose |
| --------------------------------------- | ---------------------------------------------------- |
| `gitnexus://repos` | List all indexed repositories (read this first) |
| `gitnexus://repo/{name}/context` | Codebase stats, staleness check, and available tools |
| `gitnexus://repo/{name}/clusters` | All functional clusters with cohesion scores |
@@ -270,9 +292,9 @@ It is opt-in and a no-op without `UNDERSTAND_QUICKLY_TOKEN` — a fine-grained G
**2 MCP prompts** for guided workflows:
| Prompt | What It Does |
| ----------------- | ------------------------------------------------------------------------- |
| `detect_impact` | Pre-commit change analysis — scope, affected processes, risk level |
| Prompt | What It Does |
| --------------- | ------------------------------------------------------------------------- |
| `detect_impact` | Pre-commit change analysis — scope, affected processes, risk level |
| `generate_map` | Architecture documentation from the knowledge graph with mermaid diagrams |
**4 agent skills** installed to `.claude/skills/` automatically:
@@ -359,10 +381,10 @@ npx gitnexus@latest serve
The official Docker setup ships **two signed images** orchestrated by `docker-compose.yaml`. Each image is published to both **GitHub Container Registry** (GHCR) and **Docker Hub** — same build, same digest, same Cosign signature — so pick whichever registry you prefer:
| Purpose | GHCR (default in `docker-compose.yaml`) | Docker Hub mirror |
| ---------------------------------------------------------------------- | --------------------------------------------- | ------------------------------------------- |
| CLI / `gitnexus serve` backend (HTTP API on port `4747`, MCP, indexer) | `ghcr.io/abhigyanpatwari/gitnexus:latest` | `akonlabs/gitnexus:latest` |
| Static web UI (port `4173`) | `ghcr.io/abhigyanpatwari/gitnexus-web:latest` | `akonlabs/gitnexus-web:latest` |
| Purpose | GHCR (default in `docker-compose.yaml`) | Docker Hub mirror |
| ---------------------------------------------------------------------- | --------------------------------------------- | ------------------------------ |
| CLI / `gitnexus serve` backend (HTTP API on port `4747`, MCP, indexer) | `ghcr.io/abhigyanpatwari/gitnexus:latest` | `akonlabs/gitnexus:latest` |
| Static web UI (port `4173`) | `ghcr.io/abhigyanpatwari/gitnexus-web:latest` | `akonlabs/gitnexus-web:latest` |
> **Heads-up — image rename.** Earlier releases published the web UI under
> `ghcr.io/abhigyanpatwari/gitnexus`. Starting with the introduction of the
@@ -429,7 +451,7 @@ The Docker images are version-locked to the npm package:
Both registries receive the same digest from a single build step, so you can
pull from either and the signature verifies identically.
- Release-candidate images (e.g. `:1.7.0-rc.1`) are published alongside each
RC npm release. They are built by `release-candidate.yml` calling `docker.yml`
RC npm release. They are built by `publish.yml` calling `docker.yml`
as a reusable workflow after the RC tag is created and pushed.
- `:latest` is auto-promoted only from non-prerelease tags by the Docker
metadata action, so it always points at a real, npm-published version.
@@ -462,7 +484,7 @@ registries because both sets of tags were signed at the same digest in one
workflow run.
**Release candidates** — signed from `refs/heads/main` (the caller's ref when
`release-candidate.yml` invokes `docker.yml` as a reusable workflow):
`publish.yml` invokes `docker.yml` as a reusable workflow):
```bash
cosign verify ghcr.io/abhigyanpatwari/gitnexus:1.7.0-rc.1 \
@@ -578,22 +600,22 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
### Supported Languages
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
|----------|---------|----------------|---------|----------|-----------------|---------------------|--------|------------|-------------|
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Language | Imports | Named Bindings | Exports | Heritage | Type Annotations | Constructor Inference | Config | Frameworks | Entry Points |
| ---------- | ------- | -------------- | ------- | -------- | ---------------- | --------------------- | ------ | ---------- | ------------ |
| TypeScript | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| JavaScript | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ |
| Python | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Java | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Kotlin | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| C# | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Go | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Rust | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| PHP | ✓ | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ |
| Ruby | ✓ | — | ✓ | ✓ | — | ✓ | — | ✓ | ✓ |
| Swift | — | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
@@ -725,9 +747,11 @@ gitnexus wiki --force
# Increase the timeout or retries for large codebase or slow LLM providers
gitnexus wiki --timeout <seconds> # Per-attempt LLM request timeout in seconds (default: 60)
gitnexus wiki --timeout <seconds> # LLM request timeout in seconds (default: disabled)
gitnexus wiki --retries <n> # Max LLM retry attempts per request (default: 3)
# Change the language generation for wiki
gitnexus wiki --lang <lang> # Output language for generated documentation (e.g. english, chinese, spanish, japanese)
```
The wiki generator reads the indexed graph structure, groups files into modules via LLM, generates per-module documentation pages, and creates an overview page — all with cross-references to the knowledge graph.
@@ -736,16 +760,16 @@ The wiki generator reads the indexed graph structure, groups files into modules
## Tech Stack
| Layer | CLI | Web |
| ------------------------- | ------------------------------------- | --------------------------------------- |
| Layer | CLI | Web |
| ------------------- | ------------------------------------- | --------------------------------------- |
| **Runtime** | Node.js (native) | Browser (WASM) |
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
| **Database** | LadybugDB native | LadybugDB WASM |
| **Database** | LadybugDB native | LadybugDB WASM |
| **Embeddings** | HuggingFace transformers.js (GPU/CPU) | transformers.js (WebGPU/WASM) |
| **Search** | BM25 + semantic + RRF | BM25 + semantic + RRF |
| **Agent Interface** | MCP (stdio) | LangChain ReAct agent |
| **Visualization** | — | Sigma.js + Graphology (WebGL) |
| **Frontend** | — | React 18, TypeScript, Vite, Tailwind v4 |
| **Visualization** | — | Sigma.js + Graphology (WebGL) |
| **Frontend** | — | React 18, TypeScript, Vite, Tailwind v4 |
| **Clustering** | Graphology | Graphology |
| **Concurrency** | Worker threads + async | Web Workers + Comlink |
@@ -761,12 +785,12 @@ The wiki generator reads the indexed graph structure, groups files into modules
### Recently Completed
- [X] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
- [X] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
- [X] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
- [X] Multi-Repo MCP, Zero-Config Setup, 14 Language Support
- [X] Community Detection, Process Detection, Confidence Scoring
- [X] Hybrid Search, Vector Index
- [x] Constructor-Inferred Type Resolution, `self`/`this` Receiver Mapping
- [x] Wiki Generation, Multi-File Rename, Git-Diff Impact Analysis
- [x] Process-Grouped Search, 360-Degree Context, Claude Code Hooks
- [x] Multi-Repo MCP, Zero-Config Setup, 14 Language Support
- [x] Community Detection, Process Detection, Confidence Scoring
- [x] Hybrid Search, Vector Index
---
-100
View File
@@ -1,100 +0,0 @@
# COBOL Code Indexing
GitNexus indexes COBOL codebases using a **regex-only extraction** strategy, bypassing tree-sitter entirely. This document explains why, how the pipeline works, and links to detailed sub-documents.
## Why Regex-Only?
The tree-sitter-cobol grammar (v0.0.1) has three critical limitations that make it unusable for production indexing:
| Issue | Impact | Severity |
|-------|--------|----------|
| External scanner hangs on ~5% of files | No timeout mechanism exists for the C scanner; the process blocks indefinitely | **Blocking** |
| Only ~15% of paragraph headers detected | Most procedure-division paragraphs are invisible to the grammar | High |
| Patch markers in cols 1-6 cause parse errors | Enterprise COBOL uses non-standard sequence area content (e.g., `mzADD`, `estero`, `#FIX`) | High |
Because the external scanner hang cannot be interrupted (there is no `setTimeoutMicros` equivalent for tree-sitter), using tree-sitter-cobol would hang the indexing pipeline on a non-trivial fraction of real-world files.
The regex-only approach provides:
- **Speed**: ~1ms per file average extraction time
- **Reliability**: zero hangs, zero crashes across 13,000+ files
- **Coverage**: captures all critical symbols -- program name, paragraphs, sections, CALL, PERFORM, COPY, data items (01-77, 88-level), file declarations, FD entries, EXEC SQL/CICS blocks, ENTRY points, and MOVE statements
## Architecture
```mermaid
flowchart TD
A[Repository Scan] --> B{File Detection}
B -->|Extension match| C[COBOL file]
B -->|GITNEXUS_COBOL_DIRS match| C
B -->|No match| Z[Skip]
C --> D{Copybook?}
D -->|Yes| E[Add to Copybook Map]
D -->|No| F[Source Program]
E --> G[COPY Expansion Engine]
F --> G
G -->|Inline copybook content| H[Expanded Source]
H --> I[Patch Marker Cleanup]
I --> J[Regex State Machine]
J --> K[Extracted Symbols]
K --> L[Graph Model Builder]
L --> M[Knowledge Graph]
subgraph "Per-Chunk Processing"
G
H
I
J
K
L
end
subgraph "Post-Processing"
M --> N[Community Detection]
M --> O[Process Detection]
M --> P[Contract Detection]
end
style J fill:#e8f5e9,stroke:#2e7d32
style G fill:#e3f2fd,stroke:#1565c0
```
## COBOL vs Tree-Sitter Languages
| Feature | COBOL (Regex) | Tree-Sitter Languages |
|---------|--------------|----------------------|
| Parser | Single-pass regex state machine | tree-sitter grammar + queries |
| Speed | ~1ms/file | ~5ms/file |
| AST available | No | Yes |
| COPY expansion | Yes (pre-processing step) | N/A |
| Deep indexing | Data items, SQL, CICS, FD, ENTRY | Type annotations, generics, etc. |
| Call extraction | PERFORM (intra-file) + CALL (cross-program) | AST-based call site detection |
| Import extraction | COPY statements | `import`/`require`/`use`/`#include` |
| Coverage | All critical symbols | Language-dependent query coverage |
| Failure mode | Never hangs | External scanner can hang (COBOL only) |
## Sub-Documents
| Document | Description |
|----------|-------------|
| [File Detection](./file-detection.md) | Extension mapping, `GITNEXUS_COBOL_DIRS`, copybook classification |
| [COPY Expansion](./copy-expansion.md) | Copybook inlining, REPLACING transformations, cycle detection |
| [Regex Extraction](./regex-extraction.md) | State machine, regex patterns, line processing |
| [Deep Indexing](./deep-indexing.md) | Data items, EXEC SQL/CICS, file declarations, FD, ENTRY, MOVE |
| [Graph Model](./graph-model.md) | COBOL-specific node types, edge types, full annotated example |
| [Performance](./performance.md) | Benchmarks, worker pool tuning, caps, troubleshooting |
## Key Source Files
| File | Purpose |
|------|---------|
| `gitnexus/src/core/ingestion/cobol-preprocessor.ts` | Patch marker cleanup + regex extraction engine |
| `gitnexus/src/core/ingestion/cobol-copy-expander.ts` | COPY statement expansion with REPLACING |
| `gitnexus/src/core/ingestion/utils.ts` | `getLanguageFromPath`, `getLanguageFromFilename` |
| `gitnexus/src/core/ingestion/pipeline.ts` | `isCobolCopybook`, `expandCobolCopies`, `detectCrossProgamContracts` |
| `gitnexus/src/core/ingestion/workers/parse-worker.ts` | `processCobolRegexOnly` -- graph model builder |
| `gitnexus/src/core/ingestion/workers/worker-pool.ts` | Configurable sub-batch size for COBOL |
-157
View File
@@ -1,157 +0,0 @@
# COBOL COPY Expansion
The COPY statement is COBOL's include mechanism -- analogous to `#include` in C or `import` in modern languages. GitNexus expands COPY statements **before** regex extraction so that symbols defined inside copybooks (data items, paragraphs, etc.) are visible in the program's extracted graph.
## Supported Syntax
### Basic COPY
```cobol
COPY CPSESP.
COPY "WORKGRID.CPY".
```
Inlines the content of the named copybook, replacing the COPY line(s).
### COPY with REPLACING
```cobol
COPY CPSESP REPLACING "ANAZI-KEY" BY "LK-KEY".
COPY CPSESP REPLACING LEADING "ESP-" BY "LK-ESP-"
LEADING "KPSESPL" BY "LK-KPSESPL".
COPY LINKAGE REPLACING TRAILING "-IN" BY "-OUT".
```
Three REPLACING types are supported:
| Type | Syntax | Behavior | Example |
| ------------ | ------------------------------------ | --------------------------------------- | -------------------------------- |
| **EXACT** | `REPLACING "OLD" BY "NEW"` | Replace exact identifier matches | `ANAZI-KEY` becomes `LK-KEY` |
| **LEADING** | `REPLACING LEADING "PFX-" BY "NEW-"` | Replace prefix on all COBOL identifiers | `ESP-NAME` becomes `LK-ESP-NAME` |
| **TRAILING** | `REPLACING TRAILING "-IN" BY "-OUT"` | Replace suffix on all COBOL identifiers | `DATA-IN` becomes `DATA-OUT` |
Multiple REPLACING clauses can appear in a single COPY statement. They are applied in order to each COBOL identifier in the copybook content.
### Multi-Line COPY
COPY statements can span multiple lines (standard COBOL continuation rules apply):
```cobol
COPY CPSESP REPLACING
- LEADING "ESP-" BY "LK-ESP-"
- LEADING "KPSESPL" BY "LK-KPSESPL".
```
Continuation lines (indicator `-` in column 7) are merged before COPY statement scanning.
## Expansion Flow
```mermaid
sequenceDiagram
participant Pipeline
participant Expander as COPY Expander
participant Resolver
participant Reader
Pipeline->>Pipeline: Identify all COBOL files
Pipeline->>Pipeline: Classify copybooks vs programs
Pipeline->>Reader: Read all copybook content upfront
Reader-->>Pipeline: Copybook content map (name -> content)
loop For each source file in chunk
Pipeline->>Expander: expandCopies(content, filePath, resolveFile, readFile)
Expander->>Expander: Merge continuation lines
Expander->>Expander: Detect COPY statements via regex
loop For each COPY statement (reverse order)
Expander->>Resolver: resolveFile(copyTarget)
Resolver-->>Expander: Copybook key or null
alt Resolved successfully
Expander->>Reader: readFile(resolvedKey)
Reader-->>Expander: Copybook content
Expander->>Expander: Apply REPLACING transformations
Expander->>Expander: Recurse for nested COPYs (depth + 1)
Expander->>Expander: Splice expanded content into output
else Not resolved
Expander->>Expander: Keep original COPY line
end
end
Expander-->>Pipeline: Expanded content + resolution metadata
Pipeline->>Pipeline: Replace file content with expanded content
end
```
The return type `CopyExpansionResult` contains `expandedContent` and `copyResolutions`. The `expansionDepth` field has been removed from the return type (it was unused by callers).
COPY statement line numbers in `CopyResolution` are 1-based (consistent with the preprocessor's line numbering). The splice operation that replaces COPY lines with expanded content adjusts for 0-based array indexing internally.
## Cycle Detection
Circular COPY references (e.g., copybook A includes copybook B which includes copybook A) are detected and handled:
1. Each expansion chain maintains a `visited` set of resolved copybook paths
2. If a copybook path is already in the visited set, the expansion is skipped
3. A `warnedCircular` set (internal to `expandCopies()`, not a parameter) deduplicates warning messages within a single file expansion
Known circular copybooks in PROJECT-NAME: `ANAZI`, `ANDIP`, `QDIPE` (self-referential includes).
## Max Depth
Nested COPY expansion is limited to **10 levels** (`DEFAULT_MAX_DEPTH`). If a COPY chain exceeds this depth, a warning is logged and the remaining COPY statements are left unexpanded.
## Max Total Expansions
A breadth amplification guard caps the total number of COPY expansions across all branches within a single file to **500** (`MAX_TOTAL_EXPANSIONS`). This prevents exponential blowup from diamond-shaped COPY graphs where N copybooks each include N other copybooks. Once the limit is reached, further COPY statements in that file are left unexpanded and a single warning is logged.
## REPLACING Application Detail
The REPLACING engine works by scanning all COBOL identifiers (matching `\b[A-Z][A-Z0-9-]*\b`) in the copybook content and applying each replacement rule:
```
Original copybook content:
05 ESP-NAME PIC X(30).
05 ESP-CODE PIC X(10).
05 KPSESPL-FLAG PIC X(01).
After REPLACING LEADING "ESP-" BY "LK-ESP-" LEADING "KPSESPL" BY "LK-KPSESPL":
05 LK-ESP-NAME PIC X(30).
05 LK-ESP-CODE PIC X(10).
05 LK-KPSESPL-FLAG PIC X(01).
```
For LEADING replacements, the engine checks if each identifier starts with the `from` prefix (case-insensitive) and replaces only the prefix portion, preserving the rest of the identifier.
For TRAILING replacements, the same logic applies to suffixes.
For EXACT replacements, only identifiers that match the `from` value exactly (case-insensitive) are replaced.
## Copybook Resolution
The resolver tries multiple strategies to match a COPY target name to a copybook file:
1. **Exact match**: `COPY CPSESP` resolves to copybook named `CPSESP`
2. **Strip extension**: `COPY WORKGRID.CPY` strips `.CPY` and resolves to `WORKGRID`
3. **Add extension**: `COPY CPSESP` tries `CPSESP.CPY` and `CPSESP.COPY`
If no match is found, the COPY statement is left in place (unexpanded) and a resolution record with `resolvedPath: null` is created.
## Pipeline Integration
The expansion runs **per chunk**, after file content is read but before dispatch to worker threads:
1. All copybook files are read upfront (they are typically small, collectively under 100MB)
2. Per chunk, the copybook map is merged with chunk content (in case a chunk contains copybooks)
3. Only programs (not copybooks themselves) undergo expansion
4. The expanded content replaces the original content in-place before worker dispatch
## Inline Comment Handling
The copy expander's `stripInlineComment()` helper is quote-aware: pipe characters (`|`) inside single- or double-quoted strings are preserved. This matches the same quote-aware logic used by the preprocessor.
## Source Files
- `gitnexus/src/core/ingestion/cobol-copy-expander.ts` -- `expandCopies()`, `parseReplacingClause()`, `applyReplacing()`
- `gitnexus/src/core/ingestion/pipeline.ts` -- `expandCobolCopies()`, copybook map construction, chunk integration
-312
View File
@@ -1,312 +0,0 @@
# COBOL Deep Indexing
Beyond basic symbol extraction (program name, paragraphs, CALL, PERFORM, COPY), GitNexus performs deep indexing of COBOL-specific constructs: data items, EXEC SQL/CICS blocks, file declarations, FD entries, ENTRY points, and MOVE statements.
## Data Items
### Level Numbers
| Level Range | Meaning | Graph Node Type |
|-------------|---------|-----------------|
| 01 | Record (group item) | `Record` |
| 02-49 | Elementary/group items | `Property` |
| 66 | RENAMES | `Property` |
| 77 | Independent item | `Property` |
| 88 | Condition name | `Const` |
FILLER items are skipped (no useful name for the graph).
### Clauses Parsed
The `parseDataItemClauses()` function extracts these clauses from the trailing text of a data item declaration:
| Clause | Pattern | Example |
|--------|---------|---------|
| `PIC` / `PICTURE` | `\bPIC(?:TURE)?\s+(?:IS\s+)?(\S+)` | `PIC X(30)`, `PICTURE IS 9(5)V99` |
| `USAGE` | `\bUSAGE\s+(?:IS\s+)?(COMP\|BINARY\|...)` | `USAGE IS COMP-3`, `BINARY` |
| `REDEFINES` | `\bREDEFINES\s+([A-Z][A-Z0-9-]+)` | `REDEFINES WK-DATE-NUM` |
| `OCCURS` | `\bOCCURS\s+(\d+)` | `OCCURS 12 TIMES` |
Standalone COMP variants (without the `USAGE` keyword) are also detected: `COMP`, `COMP-1` through `COMP-6`, `COMP-X`, `BINARY`, `PACKED-DECIMAL`.
### Data Hierarchy
Data items form a hierarchical structure based on level numbers. The extractor uses a **stack algorithm**:
```
Processing order:
01 WK-RECORD -> push {01, WK-RECORD} -> parent: Module
05 WK-NAME -> push {05, WK-NAME} -> parent: WK-RECORD (01 < 05)
10 WK-FIRST -> push {10, WK-FIRST} -> parent: WK-NAME (05 < 10)
10 WK-LAST -> pop WK-FIRST, push -> parent: WK-NAME (05 < 10)
05 WK-CODE -> pop WK-LAST, WK-NAME -> parent: WK-RECORD (01 < 05)
88 WK-ACTIVE -> (88 handled separately) -> parent: WK-CODE
```
The stack maintains items where each entry's level is strictly less than the next. When a new item arrives with a level <= the top of stack, items are popped until the stack top has a smaller level. A `CONTAINS` edge is created from the stack top to the new item.
For 88-level condition names, the parent is the immediately preceding non-88 data item (found by scanning backwards).
### Annotated Example
```cobol
01 WK-EMPLOYEE.
05 WK-EMP-ID PIC 9(6).
05 WK-EMP-NAME PIC X(30).
05 WK-EMP-STATUS PIC X(01).
88 WK-ACTIVE VALUE "A".
88 WK-INACTIVE VALUE "I".
05 WK-SALARY PIC 9(7)V99 COMP-3.
05 WK-DEPT PIC X(04) OCCURS 3 TIMES.
```
Produces:
- `Record` node: `WK-EMPLOYEE` (level 01, section: working-storage)
- `Property` nodes: `WK-EMP-ID`, `WK-EMP-NAME`, `WK-EMP-STATUS`, `WK-SALARY`, `WK-DEPT`
- `Const` nodes: `WK-ACTIVE` (values: `A`), `WK-INACTIVE` (values: `I`)
- `CONTAINS` edges: `WK-EMPLOYEE -> WK-EMP-ID`, `WK-EMPLOYEE -> WK-EMP-NAME`, etc.
- `CONTAINS` edges: `WK-EMP-STATUS -> WK-ACTIVE`, `WK-EMP-STATUS -> WK-INACTIVE`
### Data Item Cap
A maximum of **500 data items per file** (`MAX_DATA_ITEMS_PER_FILE`) are processed. Some COBOL programs (especially after COPY expansion) can have 10,000+ data items, which would cause graph bloat and push the V8 relationship Map past its 16.7M entry limit across thousands of files.
The cap applies after extraction: the first 500 items in source order are kept. Since 01-level records appear first, critical top-level structure is preserved.
## EXEC SQL
EXEC SQL blocks are accumulated across lines between `EXEC SQL` and `END-EXEC`, then parsed as a unit.
### Operation Classification
The first SQL keyword determines the operation:
| First Keyword | Operation |
|---------------|-----------|
| `SELECT` | SELECT |
| `INSERT` | INSERT |
| `UPDATE` | UPDATE |
| `DELETE` | DELETE |
| `DECLARE` | DECLARE |
| `OPEN` | OPEN |
| `CLOSE` | CLOSE |
| `FETCH` | FETCH |
| *(anything else)* | OTHER |
### Table Extraction
Tables are extracted from SQL clauses:
| Clause Pattern | Example |
|----------------|---------|
| `FROM <table>` | `SELECT * FROM EMPLOYEES` |
| `INSERT INTO <table>` | `INSERT INTO EMPLOYEES` |
| `UPDATE <table>` | `UPDATE EMPLOYEES SET ...` |
| `JOIN <table>` | `LEFT JOIN DEPARTMENTS ON ...` |
Note: The `INTO` pattern is restricted to `INSERT INTO` to avoid false positives from `FETCH ... INTO :host-var` and `SELECT ... INTO :host-var` statements, where `INTO` introduces host variables rather than table names.
### Cursor Detection
```cobol
EXEC SQL
DECLARE C-EMPLOYEES CURSOR FOR
SELECT EMP-ID, EMP-NAME FROM EMPLOYEES
WHERE DEPT = :WK-DEPT
END-EXEC
```
Extracts: cursor `C-EMPLOYEES`, table `EMPLOYEES`, host variable `WK-DEPT`.
### Host Variables
Host variables are COBOL variables referenced in SQL with a `:` prefix. The colon is stripped:
```sql
WHERE EMP-ID = :WK-EMP-ID AND DEPT = :WK-DEPT
```
Extracts: `WK-EMP-ID`, `WK-DEPT`.
### Graph Output
- `CodeElement` node per table, with description `sql-table op:{OP}`
- `CodeElement` node per cursor, with description `sql-cursor`
- `ACCESSES` edge from Module to each CodeElement
- Deduplication: if the same table appears in multiple SQL blocks, only one node is created
## EXEC CICS
EXEC CICS blocks are accumulated and parsed similarly to SQL blocks.
### Command Detection
Two-word commands are detected first (matched against the block start):
```
SEND MAP, RECEIVE MAP, SEND TEXT, SEND CONTROL, READ NEXT, READ PREV
```
If no two-word command matches, the first word is used (e.g., `LINK`, `XCTL`, `RETURN`, `READ`, `WRITE`).
### Extraction
| Element | Pattern | Example |
|---------|---------|---------|
| MAP name | `MAP('name')` or `MAP("name")` | `EXEC CICS SEND MAP('EMPMENU')` |
| PROGRAM name | `PROGRAM('name')` or `PROGRAM("name")` | `EXEC CICS LINK PROGRAM('BGTABUP')` |
| TRANSID | `TRANSID('name')` or `TRANSID("name")` | `EXEC CICS START TRANSID('EMP1')` |
### Graph Output
- MAP: `CodeElement` node with description `cics-map cmd:{CMD}` + `ACCESSES` edge from Module
- PROGRAM: `CALLS` edge (cross-program call via CICS LINK/XCTL)
- TRANSID: `CodeElement` node with description `cics-transid cmd:{CMD}` + `ACCESSES` edge from Module
### Annotated Example
```cobol
EXEC CICS
SEND MAP('EMPMENU')
MAPSET('EMPSET')
FROM(WK-MAP-DATA)
ERASE
END-EXEC
```
Produces:
- `CodeElement` node: `EMPMENU` (description: `cics-map cmd:SEND MAP`)
- `ACCESSES` edge: Module -> `EMPMENU`
## File Declarations
SELECT statements in the INPUT-OUTPUT SECTION are accumulated across multiple lines (until a period terminator) and parsed for:
| Clause | Pattern | Example |
|--------|---------|---------|
| SELECT | `SELECT <name>` | `SELECT MASTER-FILE` |
| ASSIGN | `ASSIGN TO <file>` | `ASSIGN TO "MASTER.DAT"` |
| ORGANIZATION | `ORGANIZATION IS <type>` | `ORGANIZATION IS INDEXED` |
| ACCESS | `ACCESS MODE IS <mode>` | `ACCESS MODE IS DYNAMIC` |
| RECORD KEY | `RECORD KEY IS <field>` | `RECORD KEY IS WK-EMP-ID` |
| FILE STATUS | `FILE STATUS IS <field>` | `FILE STATUS IS WK-FILE-STATUS` |
### Graph Output
- `CodeElement` node with description containing all parsed clauses (e.g., `select org:INDEXED access:DYNAMIC key:WK-EMP-ID status:WK-FILE-STATUS assign:MASTER.DAT`)
- `RECORD_KEY_OF` edge: from Property node to CodeElement (confidence 0.8)
- `FILE_STATUS_OF` edge: from Property node to CodeElement (confidence 0.8)
## FD Entries
FD (File Description) entries associate a file name with its record layout:
```cobol
FD MASTER-FILE.
01 MASTER-RECORD.
05 MR-EMP-ID PIC 9(6).
05 MR-EMP-NAME PIC X(30).
```
The extractor tracks `pendingFdName` state: when an `FD` line is seen, the next 01-level data item becomes its record.
### Graph Output
- `CodeElement` node with description `fd record:{recordName}`
- `CONTAINS` edge: FD CodeElement -> Record node
- `CONTAINS` edge: SELECT CodeElement -> FD CodeElement (linking file declaration to file description)
## ENTRY Points
The `ENTRY` statement defines additional entry points into a COBOL program (in addition to the main program entry):
```cobol
ENTRY "SUBPROG" USING WK-PARAM-1 WK-PARAM-2.
```
### Graph Output
- `Constructor` node with description `entry params:{param1},{param2}` (or just `entry` if no parameters)
- `CONTAINS` edge: Module -> Constructor
- Symbol table entry (so the entry point is discoverable by name)
## PROCEDURE DIVISION USING
```cobol
PROCEDURE DIVISION USING WK-INPUT-REC WK-OUTPUT-REC.
```
The USING clause identifies parameters received by the program from its caller.
### Graph Output
- `RECEIVES` edge: Module -> Property (for each parameter name, confidence 0.8)
## MOVE Statements
MOVE statements produce `ACCESSES` edges in the graph:
```cobol
MOVE WK-NAME TO OUT-NAME.
MOVE CORRESPONDING WK-INPUT TO WK-OUTPUT.
MOVE CORR WK-IN TO WK-OUT.
```
### Extraction Details
- Source and target identifiers are captured
- `CORRESPONDING` and its abbreviation `CORR` are both recognized (bulk field-by-field move)
- Figurative constants (SPACES, ZEROS, LOW-VALUES, HIGH-VALUES, QUOTES, ALL) are skipped
- The enclosing paragraph (`caller`) is tracked for context
### MOVE CORRESPONDING / CORR Edge Reasons
MOVE CORRESPONDING (and CORR) produces distinct edge reasons to differentiate from simple MOVE:
| Edge | Reason (simple MOVE) | Reason (CORRESPONDING/CORR) |
|------|---------------------|-----------------------------|
| Read (source) | `cobol-move-read` | `cobol-move-corresponding-read` |
| Write (target) | `cobol-move-write` | `cobol-move-corresponding-write` |
This distinction allows queries to find bulk field-by-field moves separately from simple variable assignments.
## GO TO DEPENDING ON
The `GO TO` statement with multiple targets and a `DEPENDING ON` clause is a computed branch:
```cobol
GO TO PARA-1 PARA-2 PARA-3
DEPENDING ON WK-SELECTOR.
```
All target paragraph names are extracted and emitted as separate `gotos` entries. Each target produces a `CALLS` edge in the graph (same semantics as PERFORM). The `DEPENDING ON` variable is not currently tracked as a data-flow dependency.
## SORT INPUT/OUTPUT PROCEDURE
SORT and MERGE statements can specify procedural entry points instead of file-based I/O:
```cobol
SORT SORT-FILE ON ASCENDING KEY SORT-KEY
INPUT PROCEDURE IS PREPARE-INPUT
OUTPUT PROCEDURE IS FORMAT-OUTPUT.
```
`INPUT PROCEDURE IS` and `OUTPUT PROCEDURE IS` targets are extracted as control-flow targets (same as PERFORM). They produce `performs` entries and corresponding `CALLS` edges in the graph.
## Fixed-Format Literal Continuation
In fixed-format COBOL, string literals can span multiple lines using the continuation indicator (`-` in column 7). When a continuation line starts with a quote character, the extractor joins it with the predecessor by removing the trailing quote from the previous line and the opening quote from the continuation:
```
Line N: MOVE "THIS IS A LONG STRI
Line N+1 (cont): - "NG VALUE" TO WK-FIELD.
Merged: MOVE "THIS IS A LONG STRING VALUE" TO WK-FIELD.
```
The trailing `"` on line N and the opening `"` on line N+1 are both removed, producing a seamless literal. If no matching quote is found on the predecessor line, the continuation is appended as-is.
## Source Files
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- All extraction logic, clause parsers, EXEC block parsers
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `processCobolRegexOnly()`, graph node/edge emission
- `gitnexus/src/core/ingestion/parsing-processor.ts` -- Sequential fallback with same `MAX_DATA_ITEMS_PER_FILE` cap
-126
View File
@@ -1,126 +0,0 @@
# COBOL File Detection
GitNexus detects COBOL files through two mechanisms: extension-based mapping and directory-based override for extensionless files. This document covers both, plus the copybook/program classification logic.
## Extension Mapping
### Program Extensions
| Extension | Type |
|-----------|------|
| `.cbl` | COBOL program |
| `.cob` | COBOL program |
| `.cobol` | COBOL program |
### Copybook Extensions
| Extension | Type | Notes |
|-----------|------|-------|
| `.cpy` | Copybook | Standard |
| `.copy` | Copybook | Standard |
| `.gnm` / `.GNM` | Copybook | Enterprise (GnuCOBOL naming) |
| `.fd` / `.FD` | Copybook | File Description fragment |
| `.wrk` / `.WRK` | Copybook | Working-Storage fragment |
| `.sel` / `.SEL` | Copybook | SELECT clause fragment |
| `.open` / `.OPEN` | Copybook | File OPEN fragment |
| `.close` / `.CLOSE` | Copybook | File CLOSE fragment |
| `.ini` / `.INI` | Copybook | Initialization fragment |
| `.def` / `.DEF` | Copybook | Definition fragment |
All extension matching is case-sensitive in `getLanguageFromFilename` (the extensions above are matched as written, including uppercase variants like `.GNM`).
## Extensionless File Detection: `GITNEXUS_COBOL_DIRS`
Many enterprise COBOL repositories use extensionless files -- the filename alone identifies the program (e.g., `s/BGTABFL` is the source for program `BGTABFL`). GitNexus handles this via the `GITNEXUS_COBOL_DIRS` environment variable.
### Configuration
Set `GITNEXUS_COBOL_DIRS` to a comma-separated list of directory names:
```bash
# Files in s/, c/, and wfproc/ directories (at any depth) are treated as COBOL
export GITNEXUS_COBOL_DIRS=s,c,wfproc
```
The matching is **case-insensitive** and checks all path segments:
- `/repo/s/BGTABFL` -- matches segment `s` -- COBOL
- `/repo/src/c/CPSESP` -- matches segment `c` -- COBOL
- `/repo/wfproc/WF001` -- matches segment `wfproc` -- COBOL
- `/repo/docs/README` -- no matching segment -- skipped
### Decision Tree
```mermaid
flowchart TD
A[getLanguageFromPath] --> B[getLanguageFromFilename]
B --> C{Known extension?}
C -->|Yes .cbl/.cob/.cobol/.cpy/...| D[Return COBOL]
C -->|Yes .ts/.py/.java/...| E[Return other language]
C -->|No match| F{Has extension?}
F -->|"Has dot in basename"| G[Return null]
F -->|"No dot = extensionless"| H{GITNEXUS_COBOL_DIRS set?}
H -->|No| G
H -->|Yes| I{Any path segment<br/>matches a configured dir?}
I -->|Yes| D
I -->|No| G
style D fill:#e8f5e9,stroke:#2e7d32
style G fill:#ffebee,stroke:#c62828
```
### Implementation Detail
The `GITNEXUS_COBOL_DIRS` value is parsed once (on first call) and cached in a `Set<string>`:
```typescript
// From gitnexus/src/core/ingestion/utils.ts
const getCobolDirs = (): Set<string> => {
if (_cobolDirs) return _cobolDirs;
const raw = process.env.GITNEXUS_COBOL_DIRS;
_cobolDirs = raw
? new Set(raw.split(',').map(d => d.trim().toLowerCase()))
: new Set();
return _cobolDirs;
};
```
The path segment check splits the full path on `/` and tests each segment against the cached set.
## Copybook vs Program Classification
After a file is identified as COBOL, it must be classified as either a **program** (to be parsed for symbols) or a **copybook** (to be loaded into the copybook map for COPY expansion).
### Classification Rules
A COBOL file is classified as a **copybook** if ANY of these conditions is true:
1. It has a recognized copybook extension (`.cpy`, `.copy`, `.gnm`, `.fd`, `.wrk`, `.sel`, `.open`, `.close`, `.ini`, `.def`)
2. It is an extensionless file whose path contains a directory segment matching one of: `c`, `copy`, `copybooks`, `copylib`, `cpy`
A file is classified as a **program** if:
1. It has a program extension (`.cbl`, `.cob`, `.cobol`), OR
2. It is extensionless and does NOT match any copybook directory pattern
### Copybook Name Resolution
Copybook names are derived from the filename:
- Strip the extension (if any)
- Convert to uppercase
Examples:
- `c/CPSESP` -- name: `CPSESP`
- `copy/workgrid.cpy` -- name: `WORKGRID`
- `c/ANAZI.GNM` -- name: `ANAZI`
This name is used to resolve `COPY CPSESP.` statements during expansion.
## Source Files
- `gitnexus/src/core/ingestion/utils.ts` -- `getLanguageFromPath()`, `getLanguageFromFilename()`, `getCobolDirs()`
- `gitnexus/src/core/ingestion/pipeline.ts` -- `isCobolCopybook()`, `getCopybookName()`, `COPYBOOK_EXTENSIONS`, `COBOL_PROGRAM_EXTENSIONS`
-193
View File
@@ -1,193 +0,0 @@
# COBOL Graph Model
This document describes the graph nodes and edges that GitNexus creates for COBOL codebases. The COBOL graph model is richer than most tree-sitter languages because it captures domain-specific constructs: file declarations, FD entries, data hierarchies, SQL tables, CICS maps, and cross-program contracts.
## Entity-Relationship Diagram
```mermaid
erDiagram
File ||--o{ Module : DEFINES
File ||--o{ Function : DEFINES
File ||--o{ Namespace : DEFINES
File ||--o{ Record : DEFINES
File ||--o{ Property : DEFINES
File ||--o{ Const : DEFINES
File ||--o{ CodeElement : DEFINES
File ||--o{ Constructor : DEFINES
File }o--o{ File : IMPORTS
Module ||--o{ Record : CONTAINS
Module ||--o{ Constructor : CONTAINS
Module }o--o{ CodeElement : ACCESSES
Module }o--o{ Module : CALLS
Module }o--o{ Module : CONTRACTS
Module }o--o{ Property : RECEIVES
Record ||--o{ Property : CONTAINS
Record ||--o{ Const : CONTAINS
Record }o--o{ Record : REDEFINES
Property ||--o{ Property : CONTAINS
Property ||--o{ Const : CONTAINS
Property }o--o{ Property : REDEFINES
Property }o--o{ CodeElement : RECORD_KEY_OF
Property }o--o{ CodeElement : FILE_STATUS_OF
CodeElement ||--o{ CodeElement : CONTAINS
CodeElement ||--o{ Record : CONTAINS
Function }o--o{ Function : CALLS
```
## Node Types
| Node Type | COBOL Concept | Created From | Example |
|-----------|--------------|--------------|---------|
| `Module` | PROGRAM-ID | `PROGRAM-ID. BGTABFL` | Name: `BGTABFL`, description may include author and date |
| `Function` | Paragraph | `PROCESS-RECORD.` at column 8 | Name: `PROCESS-RECORD` |
| `Namespace` | Procedure section | `MAIN-LOGIC SECTION.` at column 8 | Name: `MAIN-LOGIC` |
| `Record` | 01-level data item | `01 WK-EMPLOYEE.` | Description: `level:01 section:working-storage` |
| `Property` | 02-49/66/77 data item | `05 WK-NAME PIC X(30).` | Description: `level:05 pic:X(30) section:working-storage` |
| `Const` | 88-level condition | `88 WK-ACTIVE VALUE "A".` | Description: `level:88 values:A` |
| `CodeElement` | SELECT, FD, SQL table, CICS map, cursor, transid | Various | Description varies by subtype |
| `Constructor` | ENTRY point | `ENTRY "SUBPROG" USING WK-DATA` | Description: `entry params:WK-DATA` |
### CodeElement Subtypes
CodeElement is used for multiple COBOL constructs, distinguished by their description prefix:
| Subtype | ID Pattern | Description Format | Example |
|---------|-----------|-------------------|---------|
| File SELECT | `CodeElement:{path}:SELECT:{name}` | `select org:INDEXED access:DYNAMIC ...` | `SELECT MASTER-FILE` |
| FD entry | `CodeElement:{path}:FD:{name}` | `fd record:{recordName}` | `FD MASTER-FILE` |
| SQL table | `CodeElement:{path}:sql-table:{name}` | `sql-table op:SELECT` | Table `EMPLOYEES` |
| SQL cursor | `CodeElement:{path}:sql-cursor:{name}` | `sql-cursor` | Cursor `C-EMPLOYEES` |
| CICS map | `CodeElement:{path}:cics-map:{name}` | `cics-map cmd:SEND MAP` | Map `EMPMENU` |
| CICS transid | `CodeElement:{path}:cics-transid:{name}` | `cics-transid cmd:START` | Transid `EMP1` |
## Edge Types
| Edge Type | Source | Target | Created By | Confidence | Example |
|-----------|--------|--------|-----------|------------|---------|
| `DEFINES` | File | any node | File defines its symbols | 1.0 | File -> Module `BGTABFL` |
| `CALLS` | Function | Function | `PERFORM X [THRU Y]` | (via call-processor) | `PROCESS-RECORD` -> `CALC-TAX` |
| `CALLS` | Module | Module | `CALL "BGTABUP"` | (via call-processor) | `BGTABFL` -> `BGTABUP` |
| `CALLS` | Module | Module | `EXEC CICS LINK PROGRAM('X')` | (via call-processor) | `BGTABFL` -> `BGTABUP` |
| `IMPORTS` | File | File | `COPY copybook` | (via import-processor) | Source file -> Copybook file |
| `CONTAINS` | Module | Record | Data hierarchy root | 1.0 | `BGTABFL` -> `WK-EMPLOYEE` |
| `CONTAINS` | Record | Property | Data hierarchy | 1.0 | `WK-EMPLOYEE` -> `WK-NAME` |
| `CONTAINS` | Property | Property | Nested data items | 1.0 | `WK-ADDRESS` -> `WK-CITY` |
| `CONTAINS` | Record/Property | Const | 88-level parent | 1.0 | `WK-STATUS` -> `WK-ACTIVE` |
| `CONTAINS` | CodeElement (FD) | Record | FD record link | 1.0 | `FD:MASTER-FILE` -> `MASTER-RECORD` |
| `CONTAINS` | CodeElement (SELECT) | CodeElement (FD) | SELECT-FD link | 0.9 | `SELECT:MASTER-FILE` -> `FD:MASTER-FILE` |
| `CONTAINS` | Module | Constructor | ENTRY in module | 1.0 | `BGTABFL` -> `SUBPROG` |
| `REDEFINES` | Record | Record | `01 X REDEFINES Y` | 1.0 | `WK-DATE-NUM` -> `WK-DATE-ALPHA` |
| `REDEFINES` | Property | Property | `05 X REDEFINES Y` | 1.0 | `WK-CODE-NUM` -> `WK-CODE-ALPHA` |
| `RECORD_KEY_OF` | Property | CodeElement (SELECT) | `RECORD KEY IS field` | 0.8 | `WK-EMP-ID` -> `SELECT:MASTER-FILE` |
| `FILE_STATUS_OF` | Property | CodeElement (SELECT) | `FILE STATUS IS field` | 0.8 | `WK-FS` -> `SELECT:MASTER-FILE` |
| `ACCESSES` | Module | CodeElement | EXEC SQL/CICS | 0.9 | `BGTABFL` -> `sql-table:EMPLOYEES` |
| `RECEIVES` | Module | Property | `PROCEDURE USING` | 0.8 | `BGTABFL` -> `WK-INPUT-REC` |
| `CONTRACTS` | Module | Module | Shared copybook detection | 0.9 | `BGTABFL` -> `BGTABUP` (via `CPSESP`) |
## Full Annotated Example
Given this COBOL program:
```cobol
IDENTIFICATION DIVISION.
PROGRAM-ID. EMPMAINT.
AUTHOR. Development Team.
ENVIRONMENT DIVISION.
INPUT-OUTPUT SECTION.
FILE-CONTROL.
SELECT EMP-FILE
ASSIGN TO "EMPLOYEE.DAT"
ORGANIZATION IS INDEXED
ACCESS MODE IS DYNAMIC
RECORD KEY IS EMP-ID
FILE STATUS IS WS-FILE-STATUS.
DATA DIVISION.
FILE SECTION.
FD EMP-FILE.
01 EMP-RECORD.
05 EMP-ID PIC 9(6).
05 EMP-NAME PIC X(30).
WORKING-STORAGE SECTION.
01 WS-FLAGS.
05 WS-FILE-STATUS PIC X(02).
05 WS-EOF-FLAG PIC X(01).
88 WS-EOF VALUE "Y".
LINKAGE SECTION.
01 LK-SEARCH-KEY PIC 9(6).
PROCEDURE DIVISION USING LK-SEARCH-KEY.
MAIN-LOGIC SECTION.
MAIN-START.
PERFORM OPEN-FILE
PERFORM PROCESS-RECORDS
PERFORM CLOSE-FILE
STOP RUN.
OPEN-FILE.
OPEN I-O EMP-FILE.
PROCESS-RECORDS.
MOVE LK-SEARCH-KEY TO EMP-ID
EXEC SQL
SELECT EMP_SALARY INTO :WS-SALARY
FROM EMPLOYEES
WHERE EMP_ID = :EMP-ID
END-EXEC
CALL "EMPREPORT".
CLOSE-FILE.
CLOSE EMP-FILE.
```
The graph produced contains:
**Nodes:**
- `Module`: EMPMAINT (description: `author:Development Team`)
- `Namespace`: MAIN-LOGIC
- `Function`: MAIN-START, OPEN-FILE, PROCESS-RECORDS, CLOSE-FILE
- `Record`: EMP-RECORD, WS-FLAGS, LK-SEARCH-KEY
- `Property`: EMP-ID, EMP-NAME, WS-FILE-STATUS, WS-EOF-FLAG
- `Const`: WS-EOF (values: Y)
- `CodeElement`: SELECT:EMP-FILE, FD:EMP-FILE, sql-table:EMPLOYEES
- (COPY imports, if any, would produce File IMPORTS edges)
**Edges:**
- `DEFINES`: File -> all nodes
- `CONTAINS`: EMPMAINT -> EMP-RECORD, EMPMAINT -> WS-FLAGS, EMPMAINT -> LK-SEARCH-KEY
- `CONTAINS`: EMP-RECORD -> EMP-ID, EMP-RECORD -> EMP-NAME
- `CONTAINS`: WS-FLAGS -> WS-FILE-STATUS, WS-FLAGS -> WS-EOF-FLAG
- `CONTAINS`: WS-EOF-FLAG -> WS-EOF
- `CONTAINS`: FD:EMP-FILE -> EMP-RECORD
- `CONTAINS`: SELECT:EMP-FILE -> FD:EMP-FILE
- `CALLS`: MAIN-START -> OPEN-FILE, MAIN-START -> PROCESS-RECORDS, MAIN-START -> CLOSE-FILE
- `CALLS`: EMPMAINT -> EMPREPORT (external CALL)
- `ACCESSES`: EMPMAINT -> sql-table:EMPLOYEES
- `RECEIVES`: EMPMAINT -> LK-SEARCH-KEY (PROCEDURE USING)
- `RECORD_KEY_OF`: EMP-ID -> SELECT:EMP-FILE
- `FILE_STATUS_OF`: WS-FILE-STATUS -> SELECT:EMP-FILE
## How COBOL Differs from Tree-Sitter Languages
| Aspect | COBOL | Tree-Sitter Languages |
|--------|-------|----------------------|
| Node variety | 8 types (Module, Function, Namespace, Record, Property, Const, CodeElement, Constructor) | Typically 4-6 (Function, Class, Method, Interface, Module, Const) |
| Domain edges | RECORD_KEY_OF, FILE_STATUS_OF, ACCESSES, RECEIVES, CONTRACTS, REDEFINES | Primarily CALLS, IMPORTS, EXTENDS, IMPLEMENTS |
| Data hierarchy | Deep CONTAINS chains (01 -> 05 -> 10 -> 88) | Flat class members |
| Cross-program calls | CALL "name" + CICS LINK PROGRAM | Import-based resolution |
| Contract detection | Shared COPY copybook between caller/callee | Not applicable |
| Metadata | AUTHOR, DATE-WRITTEN on Module | JSDoc/docstring (not indexed) |
## Source Files
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `processCobolRegexOnly()`, node/edge emission logic
- `gitnexus/src/core/ingestion/pipeline.ts` -- `detectCrossProgamContracts()` for CONTRACTS edges
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- `CobolRegexResults` interface (all extracted data)
-261
View File
@@ -1,261 +0,0 @@
# COBOL Performance and Tuning
This document covers real-world benchmarks, worker pool configuration, memory management, known limitations, and troubleshooting for COBOL indexing.
## PROJECT-NAME Benchmark
The PROJECT-NAME project is a large Italian payroll system written in COBOL. It serves as the primary benchmark for COBOL indexing performance.
### Input
| Metric | Value |
| --------------------------- | ---------------------------------------------------------------------------- |
| Paths scanned | 14,217 |
| Parseable files | 13,129 |
| Total source size | 224 MB |
| Chunks | 12 (at 20 MB budget) |
| Copybooks loaded | 2,976 |
| Copybooks used in expansion | 2,955 |
| Key directories | `s/` (7773 programs), `c/` (3036 copybooks), `wfproc/` (1973 workflow files) |
### Output
| Metric | Value |
| ---------------------- | ------ |
| Graph nodes | 2.79M |
| Graph edges | 5.67M |
| Clusters (communities) | 16,679 |
| Execution flows | 300 |
### Timing
| Phase | Duration |
| ------------------------------- | ----------------- |
| Total | ~251s |
| KuzuDB write | 132s |
| Full-text search indexing | 6.7s |
| Regex extraction (avg per file) | ~1ms |
| COPY expansion + deep indexing | Remainder (~112s) |
### Indexing Command
```bash
cd /path/to/PROJECT-NAME
GITNEXUS_COBOL_DIRS=s,c,wfproc GITNEXUS_VERBOSE=1 node --max-old-space-size=8192 \
/path/to/gitnexus/dist/cli/index.js analyze --force
```
## Open-Source Benchmarks
### CardDemo (AWS)
| Metric | Value |
| ------ | ----- |
| Graph nodes | 12,323 |
| Graph edges | 8,893 |
| Total time | 7.4s |
### ACAS
| Metric | Value |
| ------ | ----- |
| Graph nodes | 14,016 |
| Graph edges | 15,452 |
| Total time | 9.3s |
### Micro-Benchmark (Single-File Extraction)
| Metric | Value |
| ------ | ----- |
| Per-iteration | 0.65ms |
| Throughput | ~382K lines/sec |
## Worker Pool Tuning
### Sub-Batch Size
The worker pool splits each worker's chunk into sub-batches to bound peak memory per `postMessage` serialization. COBOL repos use a smaller sub-batch size than the default:
| Parameter | Default | COBOL Mode |
| --------------------- | ----------- | ------------------- |
| Sub-batch size | 1,500 files | 200 files |
| Per sub-batch timeout | 120s | 120s (configurable) |
**Why 200?** COBOL regex extraction + preprocessing takes ~1ms per file on average, but with COPY expansion and deep indexing the effective time is ~150ms per file. At sub-batch size 1500, that would be ~225s per sub-batch, exceeding the 120s timeout.
COBOL mode is activated automatically when `GITNEXUS_COBOL_DIRS` is set:
```typescript
// From pipeline.ts
const cobolSubBatch = process.env.GITNEXUS_COBOL_DIRS ? 200 : undefined;
workerPool = createWorkerPool(workerUrl, undefined, cobolSubBatch);
```
### Worker Count
Workers default to `min(8, cpus - 1)`. For COBOL repos, this is usually sufficient since regex extraction is CPU-bound but fast. The bottleneck is typically KuzuDB write, not extraction.
### Timeout Configuration
| Environment Variable | Default | Purpose |
| ------------------------------------ | --------------- | --------------------------------------------------- |
| `GITNEXUS_WORKER_TIMEOUT_MS` | 120,000 (2 min) | Per sub-batch processing timeout |
| `GITNEXUS_WORKER_STARTUP_TIMEOUT_MS` | 60,000 (1 min) | Worker initialization timeout (tree-sitter loading) |
For COBOL-only repos, worker startup is faster because tree-sitter native modules are loaded lazily (skipped entirely if only COBOL files are present).
## Data Item Cap
### Configuration
```typescript
const MAX_DATA_ITEMS_PER_FILE = 500;
```
This constant appears in both `parse-worker.ts` (worker path) and `parsing-processor.ts` (sequential fallback).
### Rationale
Some COBOL programs, especially after COPY expansion, can have 10,000+ data items. At that scale:
- The in-memory relationship Map (for CONTAINS, REDEFINES, etc.) approaches the V8 16.7M entry limit across thousands of files
- KuzuDB write time increases linearly with edge count
- Most deep-nested items (level 20+) are rarely queried individually
### Impact
The cap truncates data items beyond the 500th in source order. Since 01-level Records appear first in COBOL source, the cap preserves:
- All 01-level record definitions
- The most important 02-49 level items (those closest to the record root)
- 88-level conditions associated with early items
To increase the cap for specific needs, modify the `MAX_DATA_ITEMS_PER_FILE` constant in both files.
## Memory Management
### COPY Expansion Breadth Guard
A per-file `MAX_TOTAL_EXPANSIONS = 500` limit prevents exponential blowup from diamond-shaped COPY graphs (e.g., N copybooks each containing N COPY statements). Once the limit is reached, further COPY statements in that file are left unexpanded. See [copy-expansion.md](copy-expansion.md) for details.
### COPY Expansion Memory
All copybook content is loaded upfront into a Map before chunk processing begins. For PROJECT-NAME:
- 2,976 copybooks, typically under 100MB total
- The Map is shared (read-only) across chunk iterations
- Per-chunk, the copybook map is merged with chunk file content (in case a chunk contains copybooks not in the pre-loaded set)
- After all chunks are processed, the copybook map is freed (`cobolCopybookContents = undefined`)
### Chunk Budget
Source files are grouped into chunks of max 20MB (`CHUNK_BYTE_BUDGET`). Each chunk's lifecycle:
1. Read file content into memory
2. Expand COPY statements (mutates content in-place)
3. Dispatch to workers for extraction
4. Workers return serialized results
5. Merge results into graph
6. Chunk content goes out of scope (GC reclaims)
This ensures only ~20MB of source + ~200-400MB of working memory (ASTs, extracted records, serialization) is active at any time.
### Shared Warning Deduplication
The `warnedCircular` set (used by the COPY expansion engine) is shared across all files in a chunk. This prevents the same circular copybook warning (e.g., `ANAZI includes itself`) from being logged thousands of times.
## Known Limitations
| Limitation | Impact | Workaround |
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------- |
| tree-sitter-cobol hangs on ~5% of files | Cannot use tree-sitter for COBOL | Regex-only extraction (current approach) |
| Data item cap (500/file) | May miss deeply nested items in large programs | Increase `MAX_DATA_ITEMS_PER_FILE` in source |
| Circular copybooks (ANAZI, ANDIP, QDIPE) | Self-referential includes cannot be expanded | Detected and skipped with warning |
| wfproc/ files may not be pure COBOL | Workflow files may produce extraction noise | Exclude `wfproc` from `GITNEXUS_COBOL_DIRS` if problematic |
| No MOVE DATA_FLOW edges yet | Data flow between variables not in graph | Reserved for future release |
| Continuation line handling | Some complex multi-line continuations (especially in string literals spanning 3+ lines) may not merge correctly | Known edge case; affects <0.1% of lines |
| Single-line EXEC blocks | `EXEC SQL SELECT ... END-EXEC` on one line is handled, but pathological nesting is not | Extremely rare in practice |
| Extension case sensitivity | `.GNM` and `.gnm` are matched differently | Use the exact case from the codebase |
## Troubleshooting
### "COPY expansion failed"
```
[pipeline] COPY expansion failed for s/BGTABFL: Cannot read properties of null
```
**Cause:** A copybook referenced by a COPY statement cannot be found.
**Fix:**
1. Verify `GITNEXUS_COBOL_DIRS` includes the directory containing copybooks (typically `c`)
2. Check that copybook filenames match the COPY target (case-insensitive, after stripping extensions)
3. Ensure copybook files are not in `.gitignore`
### Worker sub-batch timeout
```
Worker 3 sub-batch timed out after 120s (chunk: 200 items)
```
**Cause:** A sub-batch took longer than the timeout. Typically happens when one file is extremely large (50,000+ lines after COPY expansion).
**Fix:** Increase the timeout:
```bash
GITNEXUS_WORKER_TIMEOUT_MS=300000 gitnexus analyze
```
### Memory errors (heap out of memory)
```
FATAL ERROR: CALL_AND_RETRY_LAST Allocation failed - JavaScript heap out of memory
```
**Fix:** Increase Node.js heap size:
```bash
node --max-old-space-size=16384 /path/to/gitnexus/dist/cli/index.js analyze
```
For very large repos (>500MB source), consider `--max-old-space-size=32768`.
### Concurrent analyze corruption
**Rule:** Only ONE `gitnexus analyze` process should run at a time per repository. Concurrent writes to KuzuDB corrupt the database.
If corruption occurs:
```bash
# Remove the KuzuDB directory and re-index
rm -rf .gitnexus/kuzu
gitnexus analyze --force
```
### Slow KuzuDB write phase
The KuzuDB write phase (132s for PROJECT-NAME) is the bottleneck for large COBOL repos. This is proportional to the number of nodes and edges being written. Reducing `MAX_DATA_ITEMS_PER_FILE` or excluding non-essential directories from `GITNEXUS_COBOL_DIRS` can help.
### Verbose output
Enable verbose logging to see per-phase timing and statistics:
```bash
GITNEXUS_VERBOSE=1 gitnexus analyze
```
This outputs:
- Scan statistics (paths, parseable files, chunk count)
- Worker pool configuration (worker count, sub-batch size)
- COPY expansion statistics (copybooks loaded, files expanded)
- Community and process detection results
- Contract detection results
## Source Files
- `gitnexus/src/core/ingestion/workers/worker-pool.ts` -- `DEFAULT_SUB_BATCH_SIZE`, `SUB_BATCH_TIMEOUT_MS`, `WORKER_STARTUP_TIMEOUT_MS`
- `gitnexus/src/core/ingestion/pipeline.ts` -- `CHUNK_BYTE_BUDGET`, COBOL sub-batch configuration, chunk lifecycle
- `gitnexus/src/core/ingestion/workers/parse-worker.ts` -- `MAX_DATA_ITEMS_PER_FILE`, `processCobolRegexOnly()`
- `gitnexus/src/core/ingestion/parsing-processor.ts` -- Sequential fallback `MAX_DATA_ITEMS_PER_FILE`
@@ -1,206 +0,0 @@
# COBOL Regex Extraction
The `extractCobolSymbolsWithRegex()` function in `cobol-preprocessor.ts` performs single-pass, state-machine-driven extraction of all COBOL symbols. This document describes the state machine, line processing flow, and every regex pattern used.
## State Machine: Division Tracking
The extractor tracks which COBOL division is currently being processed. Division transitions are detected by the `RE_DIVISION` pattern.
```mermaid
stateDiagram-v2
[*] --> null : Start of file
null --> identification : IDENTIFICATION DIVISION
identification --> environment : ENVIRONMENT DIVISION
environment --> data : DATA DIVISION
data --> procedure : PROCEDURE DIVISION
note right of identification
Extracts: PROGRAM-ID, AUTHOR, DATE-WRITTEN
end note
note right of environment
Extracts: SELECT ... ASSIGN ... (file declarations)
end note
note right of data
Extracts: FD entries, data items (01-77, 88), COPY
end note
note right of procedure
Extracts: paragraphs, sections, PERFORM, CALL,
ENTRY, MOVE, EXEC SQL/CICS
end note
```
## State Machine: Data Section Tracking
Within the DATA DIVISION, a secondary state machine tracks the current section to tag data items with their origin.
```mermaid
stateDiagram-v2
[*] --> unknown : DATA DIVISION entered
unknown --> working_storage : WORKING-STORAGE SECTION
unknown --> linkage : LINKAGE SECTION
unknown --> file : FILE SECTION
unknown --> local_storage : LOCAL-STORAGE SECTION
working_storage --> linkage : LINKAGE SECTION
working_storage --> file : FILE SECTION
linkage --> working_storage : WORKING-STORAGE SECTION
file --> working_storage : WORKING-STORAGE SECTION
file --> linkage : LINKAGE SECTION
local_storage --> working_storage : WORKING-STORAGE SECTION
```
Within the ENVIRONMENT DIVISION, the `currentEnvSection` tracks whether we are in `INPUT-OUTPUT` or `CONFIGURATION` section. SELECT statement accumulation only occurs in `INPUT-OUTPUT`.
## Line Processing Flow
Each raw source line goes through this pipeline:
```
Raw line
|
v
Length < 7? ---------> Skip (flush pending if any)
|
v
Indicator col 7
|
+-- '*' or '/' -----> Comment: skip entirely
|
+-- '-' ------------> Continuation: append to pending line
|
+-- other ----------> Normal: flush pending, strip inline comments (|),
buffer as new pending logical line
```
After all lines are processed, the final pending line is flushed, along with any accumulated SELECT statement, SORT/MERGE accumulator, and any open EXEC block (truncated file without `END-EXEC`).
### Inline Comment Stripping
Enterprise COBOL (particularly Italian dialect) uses the pipe character `|` as an inline comment marker. The `stripInlineComment()` helper is **quote-aware**: it tracks whether the scan position is inside a single- or double-quoted string and only treats `|` as a comment marker when outside quotes. Pipe characters inside string literals are preserved.
Free-format `*>` inline comment stripping uses the same quote-aware approach: the scanner walks character by character, toggling quote state, and only recognizes `*>` as a comment marker when not inside a quoted string.
### Patch Marker Handling
The `preprocessCobolSource()` function (run before extraction in the worker) replaces non-standard content in columns 1-6. Standard COBOL expects spaces or digit sequence numbers in this area. If any letter or `#` character is found, the entire sequence area is replaced with 6 spaces:
```
Before: mzADD MOVE WK-AMT TO WK-TOTAL
After: MOVE WK-AMT TO WK-TOTAL
```
This preserves exact line count for position mapping.
## Regex Pattern Reference
All patterns are compiled once as module-level constants and reused across calls.
### Division and Section Detection
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_DIVISION` | `\b(IDENTIFICATION\|ENVIRONMENT\|DATA\|PROCEDURE)\s+DIVISION\b` | Division boundary | `PROCEDURE DIVISION` |
| `RE_SECTION` | `\b(WORKING-STORAGE\|LINKAGE\|FILE\|LOCAL-STORAGE\|INPUT-OUTPUT\|CONFIGURATION)\s+SECTION\b` | Section boundary | `WORKING-STORAGE SECTION` |
### IDENTIFICATION DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_PROGRAM_ID` | `\bPROGRAM-ID\.\s*([A-Z][A-Z0-9-]*)` | Program name | `PROGRAM-ID. BGTABFL` |
| `RE_AUTHOR` | `^\s+AUTHOR\.\s*(.+)` | Author metadata | `AUTHOR. D. Smith` |
| `RE_DATE_WRITTEN` | `^\s+DATE-WRITTEN\.\s*(.+)` | Date metadata | `DATE-WRITTEN. 2024-01-15` |
### ENVIRONMENT DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_SELECT_START` | `\bSELECT\s+(?:OPTIONAL\s+)?([A-Z][A-Z0-9-]+)` | File SELECT start (with optional `SELECT OPTIONAL` support) | `SELECT MASTER-FILE`, `SELECT OPTIONAL TRANS-FILE` |
SELECT statements are accumulated across multiple lines until a period terminator is found, then parsed for ASSIGN, ORGANIZATION, ACCESS, RECORD KEY, and FILE STATUS clauses.
### DATA DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_FD` | `^\s+FD\s+([A-Z][A-Z0-9-]+)` | File description | `FD MASTER-FILE` |
| `RE_DATA_ITEM` | `^\s+(\d{1,2})\s+([A-Z][A-Z0-9-]+)\s*(.*)` | Data item (01-77) | `05 WK-NAME PIC X(30)` |
| `RE_ANONYMOUS_REDEFINES` | `^\s+(\d{1,2})\s+REDEFINES\s+([A-Z][A-Z0-9-]+)` | Anonymous REDEFINES | `01 REDEFINES WK-REC` |
| `RE_88_LEVEL` | `^\s+88\s+([A-Z][A-Z0-9-]+)\s+VALUES?\s+(?:ARE\s+)?(.+)` | Condition name | `88 WK-ACTIVE VALUE "Y"` |
The trailing clauses of `RE_DATA_ITEM` are parsed by `parseDataItemClauses()` for PIC, USAGE, OCCURS, and REDEFINES.
### PROCEDURE DIVISION
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_PROC_SECTION` | `^ ([A-Z][A-Z0-9-]+)\s+SECTION\.\s*$` | Procedure section header | ` MAIN-LOGIC SECTION.` |
| `RE_PROC_PARAGRAPH` | `^ ([A-Z][A-Z0-9-]+)\.\s*$` | Paragraph header | ` PROCESS-RECORD.` |
| `RE_PERFORM` | `\bPERFORM\s+([A-Z][A-Z0-9-]+)(?:\s+THRU\s+([A-Z][A-Z0-9-]+))?` | PERFORM call | `PERFORM CALC-TAX THRU CALC-TAX-EXIT` |
| `RE_PROC_USING` | `\bPROCEDURE\s+DIVISION\s+USING\s+([\s\S]*?)(?:\.\|$)` | USING parameters | `PROCEDURE DIVISION USING WK-PARAM` |
| `RE_ENTRY` | `\bENTRY\s+"([^"]+)"(?:\s+USING\s+([\s\S]*?))?(?:\.\|$)` | ENTRY point | `ENTRY "SUBPROG" USING WK-DATA` |
| `RE_MOVE` | `\bMOVE\s+((?:CORRESPONDING\|CORR)\s+)?([A-Z][A-Z0-9-]+)\s+TO\s+(.+)` | MOVE statement (supports CORR abbreviation and multi-target) | `MOVE WK-NAME TO OUT-NAME`, `MOVE CORR WK-IN TO WK-OUT` |
The USING parameter list (`RE_PROC_USING`) is split on `\bRETURNING\b` before tokenization -- any RETURNING clause and everything after it is excluded from the parameter list (`.split(/\bRETURNING\b/i)[0]`).
Note: `RE_PROC_SECTION` and `RE_PROC_PARAGRAPH` require exactly 7 spaces of leading indentation (COBOL area A starting at column 8). This is the standard COBOL paragraph indentation.
### All-Division Patterns
These patterns are checked regardless of current division:
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_CALL` | `\bCALL\s+"([^"]+)"` | External program call | `CALL "BGTABUP"` |
| `RE_COPY_UNQUOTED` | `\bCOPY\s+([A-Z][A-Z0-9-]+)(?:\s\|\.)` | COPY (unquoted) | `COPY CPSESP.` |
| `RE_COPY_QUOTED` | `\bCOPY\s+"([^"]+)"(?:\s\|\.)` | COPY (quoted) | `COPY "WORKGRID.CPY".` |
### SORT/MERGE Support
| Constant | Purpose |
|----------|---------|
| `SORT_CLAUSE_NOISE` | Set of SORT/MERGE clause keywords filtered from USING/GIVING file lists: `ON`, `ASCENDING`, `DESCENDING`, `KEY`, `WITH`, `DUPLICATES`, `IN`, `ORDER`, `COLLATING`, `SEQUENCE`, `IS`, `THROUGH`, `THRU`, `INPUT`, `OUTPUT`, `PROCEDURE` |
SORT and MERGE statements are accumulated across multiple lines (like SELECT) until a period terminator is found, then parsed for USING/GIVING file lists and INPUT/OUTPUT PROCEDURE targets. The `flushSort()` helper encapsulates the flush-and-parse logic, mirroring the existing `flushSelect()` pattern. Both helpers are called at EOF to handle truncated files.
### GO TO Multi-Target
`RE_GOTO` captures all paragraph names in a `GO TO` statement, including the multi-target form `GO TO p1 p2 p3 DEPENDING ON x`. The captured group contains all target names (space-separated), which are split into individual targets. Each target produces a separate `gotos` entry.
### PROGRAM-ID Detection
PROGRAM-ID is detected regardless of the current division state. This handles sibling programs that appear after `END PROGRAM` and omit the `IDENTIFICATION DIVISION` header -- the extractor will still capture the PROGRAM-ID and push a new program boundary.
### EXEC Block Patterns
| Constant | Pattern | Purpose | Example Match |
|----------|---------|---------|---------------|
| `RE_EXEC_SQL_START` | `\bEXEC\s+SQL\b` | Start of EXEC SQL block | `EXEC SQL` |
| `RE_EXEC_CICS_START` | `\bEXEC\s+CICS\b` | Start of EXEC CICS block | `EXEC CICS` |
| `RE_END_EXEC` | `\bEND-EXEC\b` | End of EXEC block | `END-EXEC` |
EXEC blocks accumulate all lines between `EXEC SQL/CICS` and `END-EXEC`, then delegate to `parseExecSqlBlock()` or `parseExecCicsBlock()` for detailed extraction.
## Excluded Paragraph Names
The following names are excluded from paragraph detection to avoid false positives from division/section headers:
```
DECLARATIVES, END, PROCEDURE, IDENTIFICATION,
ENVIRONMENT, DATA, WORKING-STORAGE, LINKAGE,
FILE, LOCAL-STORAGE, COMMUNICATION, REPORT,
SCREEN, INPUT-OUTPUT, CONFIGURATION
```
Additionally, paragraph candidates containing `DIVISION` or `SECTION` as substrings are excluded.
## MOVE Skip List (Figurative Constants)
MOVE statements where the source is a figurative constant are skipped:
```
SPACES, ZEROS, ZEROES, LOW-VALUES, LOW-VALUE,
HIGH-VALUES, HIGH-VALUE, QUOTES, QUOTE, ALL
```
## Source Files
- `gitnexus/src/core/ingestion/cobol-preprocessor.ts` -- `preprocessCobolSource()`, `extractCobolSymbolsWithRegex()`, all regex constants
-300
View File
@@ -1,300 +0,0 @@
# Using GitNexus across gRPC microservices
## When to use this guide
This guide is for teams whose product lives in **several separate Git repositories** — one per service — and whose services talk to each other over **gRPC** (possibly alongside HTTP and message topics). GitNexus indexes each repo independently, then a _group_ stitches the per-repo indexes into a single cross-repo view that the `impact`, `query`, and `context` tools can traverse. If your services live in one monorepo, much of this still applies — set each service as a member of a group and use the `service` prefix to scope queries — but the walkthrough assumes the harder multi-repo case.
## Mental model
- Each repository has its own `.gitnexus/` index (a LadybugDB graph of symbols, relationships, processes). `gitnexus analyze` in each repo produces that index completely independently.
- A **group** is a higher-level construct stored at `~/.gitnexus/groups/<group>/` that references the per-repo indexes by their registry name.
- Sync-time extractors walk each member repo and emit **contracts** — provider or consumer records keyed by a canonical `contractId` (`grpc::auth.AuthService/Login`, `http::GET::/orders`, etc.).
- The sync step matches providers and consumers that share a `contractId` and writes **cross-links** to `<groupDir>/contracts.json`. Those cross-links are what lets `impact({repo: "@<group>", target: "X"})` hop from one repo into another.
- Contracts come from three places: automatic contract extractors (`grpc-extractor`, `http-route-extractor`, `topic-extractor`), a manifest escape hatch (`config.links` in `group.yaml`), and — for same-name symbol matches where no contract is declared — the exact-match matching cascade in [`matching.ts`](../../gitnexus/src/core/group/matching.ts).
- Each repo stays editable and re-indexable on its own. Re-run `gitnexus analyze` in a repo when it changes, then `gitnexus group sync <group>` to refresh `contracts.json`. `gitnexus group status` reports which members are stale.
## Prerequisites
- GitNexus installed and runnable as `gitnexus` or `npx gitnexus` (see the root [README.md](../../README.md)).
- Each service repository checked out locally. No requirement that they share a parent directory — the group references them by registry name.
- Write access to `~/.gitnexus/` (the default gitnexus home; see `getDefaultGitnexusDir` in [`storage.ts`](../../gitnexus/src/core/group/storage.ts)).
## Step-by-step walkthrough
The example uses three services — a TypeScript API gateway, a Go orders service, and a Python inventory service — with gRPC between them. The gateway is an `orders` consumer; the orders service is both an `orders` provider and an `inventory` consumer; the inventory service is an `inventory` provider.
### 1. Index each repository
Run `analyze` from inside each service repo (or pass the path). The CLI surface lives in [`gitnexus/src/cli/analyze.ts`](../../gitnexus/src/cli/analyze.ts) and is wired in [`gitnexus/src/cli/index.ts`](../../gitnexus/src/cli/index.ts).
```bash
cd ~/code/gateway && npx gitnexus analyze
cd ~/code/orders && npx gitnexus analyze
cd ~/code/inventory && npx gitnexus analyze
```
Useful flags:
- `--force` — reindex even if up to date.
- `--embeddings` — generate embedding vectors (needed only if you want semantic search; the exact-match cross-repo cascade does **not** need them).
- `--name <alias>` — register the repo under a specific alias when two repos share a basename (e.g. two `api/` folders).
- `--skip-git` — index a checkout that isn't a git repo.
Each run writes a `.gitnexus/` folder in the repo and registers the repo in `~/.gitnexus/registry.json`. Confirm with `npx gitnexus list`.
### 2. Author `group.yaml`
Create the group directory and edit the config. Either use the CLI scaffolder or write the file directly — both produce the same shape consumed by [`config-parser.ts`](../../gitnexus/src/core/group/config-parser.ts).
```bash
npx gitnexus group create payments-platform
# or manually:
mkdir -p ~/.gitnexus/groups/payments-platform
$EDITOR ~/.gitnexus/groups/payments-platform/group.yaml
```
Minimal working `group.yaml`:
```yaml
version: 1
name: payments-platform
description: Gateway + orders + inventory (gRPC)
repos:
gateway: gateway
orders: orders
inventory: inventory
# Only add explicit links when the automatic extractors miss something —
# see "When automatic extraction isn't enough" below.
links: []
packages: {}
detect:
http: true
grpc: true
topics: true
shared_libs: true
embedding_fallback: false
matching:
bm25_threshold: 0.7
embedding_threshold: 0.65
max_candidates_per_step: 3
# Exclude noisy paths from cross-link matching (contracts are still extracted)
exclude_links_paths: [/ping, /health, /healthcheck]
exclude_links_param_only_paths: true
```
Field notes (schema in [`types.ts`](../../gitnexus/src/core/group/types.ts)):
- `version` — must be `1`. The parser rejects anything else.
- `name` — required; used for the group directory name and all CLI / MCP calls.
- `repos` — a mapping from **group path** (a logical name you choose; can be a hierarchy like `backend/orders`) to **registry name** (the name shown by `npx gitnexus list`). Both sides appear throughout the tooling: contract rows use the group path; `@<group>/<groupPath>` routes tools to a single member.
- `links` — optional manifest escape hatch, one entry per explicit cross-repo contract. Validated by the parser: `from` and `to` must be known repo paths, `type` must be one of `http | grpc | topic | lib | custom`, and `role` must be `provider | consumer`.
- `detect` — toggles per extractor family. Defaults (set in `config-parser.ts`) turn `http`, `grpc`, `topics`, and `shared_libs` on; disable the ones you don't use to speed up sync.
- `matching` — thresholds for the matching cascade. The exact match is always run; other strategies depend on indexer state. Two optional fields reduce false-positive cross-links in large groups:
- `exclude_links_paths` — list of HTTP paths to exclude from cross-link matching (default `[]`). Contracts at these paths are still extracted and visible in the registry, but they don't produce cross-repo links. Useful for health-check endpoints (`/ping`, `/health`) that every service exposes. Trailing slashes are normalized.
- `exclude_links_param_only_paths` — when `true`, exclude routes where every segment is `{param}` (e.g. `/{param}`, `/{param}/{param}`) from cross-link matching (default `false`). Mixed routes like `/users/{param}` are not affected.
### 3. Sync the group
```bash
npx gitnexus group sync payments-platform --verbose
```
What this does (see [`sync.ts`](../../gitnexus/src/core/group/sync.ts)):
1. Opens each member's per-repo LadybugDB.
2. Runs the HTTP, gRPC, and topic extractors against the source files.
3. Applies manifest `links` through [`manifest-extractor.ts`](../../gitnexus/src/core/group/extractors/manifest-extractor.ts).
4. Runs the exact-match cascade, joining providers and consumers that share a normalized `contractId`.
5. Writes `contracts.json` in the group directory.
Flags:
- `--exact-only` — stop after the exact cascade; skip BM25 and embedding fallback.
- `--skip-embeddings` — run exact plus BM25 but not embedding-based matching.
- `--allow-stale` — don't warn if a member's index is stale.
- `--json` — machine-readable output.
The same operation is available over MCP as `group_sync({ name: "payments-platform" })` — see [`tools.ts`](../../gitnexus/src/mcp/tools.ts).
### 4. Inspect the registry
Use `gitnexus group contracts` for the CLI view or read the `gitnexus://group/<name>/contracts` MCP resource for the same data.
```bash
npx gitnexus group contracts payments-platform --type grpc --json
```
A shortened response:
```json
{
"contracts": [
{
"contractId": "grpc::orders.OrderService/PlaceOrder",
"type": "grpc",
"role": "provider",
"repo": "orders",
"symbolRef": { "filePath": "internal/grpc/order_server.go", "name": "RegisterOrderServiceServer" },
"confidence": 0.8,
"meta": { "service": "OrderService", "method": "PlaceOrder", "source": "go_register" }
},
{
"contractId": "grpc::orders.OrderService/PlaceOrder",
"type": "grpc",
"role": "consumer",
"repo": "gateway",
"symbolRef": { "filePath": "src/clients/orders.ts", "name": "OrderServiceClient" },
"confidence": 0.75,
"meta": { "service": "OrderService", "source": "ts_generated_client" }
}
],
"crossLinks": [
{
"from": { "repo": "gateway", "symbolUid": "…", "symbolRef": { "filePath": "src/clients/orders.ts", "name": "OrderServiceClient" } },
"to": { "repo": "orders", "symbolUid": "…", "symbolRef": { "filePath": "internal/grpc/order_server.go", "name": "RegisterOrderServiceServer" } },
"type": "grpc",
"contractId": "grpc::orders.OrderService/PlaceOrder",
"matchType": "exact",
"confidence": 1.0
}
]
}
```
Staleness of the underlying indexes shows up in `npx gitnexus group status payments-platform` or the `gitnexus://group/<name>/status` resource.
### 5. Run cross-repo impact with `@<group>` routing
From any shell (you do **not** have to `cd` into a member repo), the normal `impact` / `query` / `context` tools accept `repo: "@<group>"` to fan out across all members, or `repo: "@<group>/<memberPath>"` to target one member. Routing is implemented in [`resolve-at-member.ts`](../../gitnexus/src/core/group/resolve-at-member.ts) and described in [`tools.ts`](../../gitnexus/src/mcp/tools.ts).
Example MCP calls:
```json
{"tool": "impact", "arguments": {
"repo": "@payments-platform/orders",
"target": "PlaceOrder",
"direction": "upstream",
"crossDepth": 2
}}
```
```json
{"tool": "query", "arguments": {
"repo": "@payments-platform",
"query": "retry logic around PlaceOrder"
}}
```
The CLI equivalents still exist for scripting:
```bash
npx gitnexus group impact payments-platform \
--repo orders --target PlaceOrder --direction upstream --cross-depth 2
```
Phase 1 walks within the anchor member; Phase 2 hops across the Contract Bridge wherever a cross-link endpoint matches an impacted symbol. See [`cross-impact.ts`](../../gitnexus/src/core/group/cross-impact.ts) for the bridge query.
## How gRPC extraction works
`GrpcExtractor` ([`grpc-extractor.ts`](../../gitnexus/src/core/group/extractors/grpc-extractor.ts)) runs two passes per member repo:
1. **Proto map.** Every `**/*.proto` file is parsed to enumerate `service Foo { rpc Bar(...) }` blocks and (transitively) resolve the package name. Each RPC method becomes a provider contract with `contractId = grpc::<package>.<Service>/<Method>` and `confidence = 0.85`. Parsing uses the vendored `tree-sitter-proto` grammar when available and falls back to a length-preserving manual parser (`extractServiceBlocks`) otherwise, so `.proto` extraction works on platforms where the grammar fails to build.
2. **Source scan.** Every source file whose extension matches [`GRPC_SCAN_GLOB`](../../gitnexus/src/core/group/extractors/grpc-patterns/index.ts) is parsed by its language plugin:
| Language | Provider signal | Consumer signal |
|----------|-----------------|-----------------|
| Go ([`go.ts`](../../gitnexus/src/core/group/extractors/grpc-patterns/go.ts)) | `pb.RegisterXxxServer(...)`, `pb.UnimplementedXxxServer` embedded in struct | `pb.NewXxxClient(conn)` |
| Java ([`java.ts`](../../gitnexus/src/core/group/extractors/grpc-patterns/java.ts)) | `extends XxxServiceGrpc.XxxServiceImplBase` (with or without `@GrpcService`) | `XxxServiceGrpc.newBlockingStub(...)`, `newStub(...)` |
| Python ([`python.ts`](../../gitnexus/src/core/group/extractors/grpc-patterns/python.ts)) | `add_XxxServicer_to_server(...)` (bare or `_pb2_grpc.` attribute form) | `XxxStub(channel)` (ignores `Mock`/`Test`/`Fake`/`Stub`) |
| Node / TS ([`node.ts`](../../gitnexus/src/core/group/extractors/grpc-patterns/node.ts)) | NestJS `@GrpcMethod('Service','Method')` | `@GrpcClient` field typed `XxxServiceClient`, `client.getService<X>('Service')`, `new XxxServiceClient(...)`, `new foo.bar.XxxService(...)` in files that call `loadPackageDefinition` |
For each source-scan detection the extractor looks up the short service name in the proto map and picks:
- `grpc::<package>.<Service>/<Method>` when a method is named and the service resolves against the proto map,
- `grpc::<package>.<Service>/*` (wildcard) when only the service is known, or
- `grpc::<ServiceName>/*` when no `.proto` is available at all.
Provider detections land at confidence 0.8 (with proto) or 0.65 (without); consumers at 0.75 or 0.55. NestJS `@GrpcMethod` is fixed at 0.8 because the decorator is self-describing.
### Matching
`matching.ts` lowercases the package/service segment before comparing contract ids, so bindings that capitalize names differently (`auth.AuthService` vs `auth.authservice`) still match. Method names are compared case-sensitively because gRPC's wire path is case-sensitive. Service-only wildcards (`grpc::pkg.Svc/*`) match any method on the same service during cross-linking.
### Known limitations
- **Ambiguous proto resolution.** If a short service name exists in more than one `.proto` file and the source-scan hit can't be narrowed down by shared directory segments (`resolveProtoConflict` refuses to guess), the extractor skips contract emission and logs a warning.
- **Proto packages must be resolvable locally.** Transitive imports that point outside the repo produce an empty package segment, which means the contract id collapses to `grpc::<Service>/<Method>`. Cross-repo matches still work as long as both sides agree on the empty package.
- **Rewrite rules are not implemented.** If the provider repo writes `grpc::orders.OrderService/PlaceOrder` and the consumer repo writes `grpc::orderspb.OrderService/PlaceOrder`, they won't cross-link automatically. Use `config.links` to declare the correspondence (see below).
- **One sync = one snapshot.** Contracts are extracted against the indexed snapshot of each repo. Re-index first, then re-sync; the `status` command and resource surface staleness.
## When automatic extraction isn't enough
The escape hatch is the `links` list in `group.yaml`, handled by [`ManifestExtractor`](../../gitnexus/src/core/group/extractors/manifest-extractor.ts). Each entry is a **one-directional** provider/consumer declaration:
```yaml
version: 1
name: payments-platform
repos:
gateway: gateway
orders: orders
inventory: inventory
links:
# Explicit gRPC method: use when naming mismatches stop the
# automatic matcher from cross-linking.
- from: gateway
to: orders
type: grpc
contract: OrderService/PlaceOrder
role: consumer
# Service-level link when you don't want to enumerate methods.
- from: orders
to: inventory
type: grpc
contract: InventoryService
role: consumer
# Works for HTTP too — use `METHOD::/path` form for the exact
# handler, or just `/path` for a method-agnostic wildcard.
- from: gateway
to: orders
type: http
contract: POST::/orders
role: consumer
```
What the manifest extractor does (see [`manifest-extractor.ts`](../../gitnexus/src/core/group/extractors/manifest-extractor.ts)):
1. Builds a canonical `contractId` with `buildContractId` — the same canonicalization used by the automatic extractors, so manifest links cross-match automatic contracts on the other side.
2. Tries to resolve each side to a real graph symbol (the `Route` node for HTTP, a `Function|Method` / `Class|Interface` for gRPC, a `Package|Module` for `lib`).
3. If resolution fails, falls back to a deterministic synthetic uid (`manifest::<repo>::<contractId>`) so both sides still line up in cross-impact — name-only links still work when the symbol isn't in the graph.
4. Emits both a provider and a consumer `StoredContract` (confidence `1.0`, `source: "manifest"`) and a `CrossLink` with `matchType: "manifest"`.
Use `links` for exactly the cases the extractor can't infer: different package names across repos (see #701), hand-rolled transports, cases where the provider repo isn't checked out locally but you still want a record, or any contract whose provider and consumer simply don't share a surface the extractors know how to pattern-match.
History: the manifest extractor used to be silently skipped by the sync pipeline; that was fixed in [#827](https://github.com/abhigyanpatwari/GitNexus/pull/827) (tracking issue #826). If you ever see `config.links` with zero cross-links in `contracts.json`, make sure you're on a build that includes that fix, then re-run `group sync`.
## Troubleshooting
1. **`contracts.json` is empty after a sync.** Either no member repo contained a recognizable gRPC pattern, or the extractors are disabled in `detect`. Confirm `detect.grpc: true` and re-run with `--verbose`.
2. **A known provider/consumer pair doesn't cross-link.** Most common cause: the package segment differs. Check the raw contract ids with `gitnexus group contracts <name> --unmatched` — if you see two same-method contracts with different package prefixes, add a manifest `links:` entry to bridge them (no automatic rewrite rules yet).
3. **`matchType: "manifest"` is missing entirely.** The extractor needs `config.links` to be non-empty and the sync pipeline to actually call it — verify you're on a post-#827 build. Empty contract rows for manifest links usually mean `resolveSymbol` couldn't find a graph match; the synthetic uid still lets cross-impact work, it just won't carry a file path.
4. **Ambiguous proto warnings.** Look for `[grpc-extractor] Ambiguous proto resolution` in the sync logs; that means a service name exists in multiple `.proto` files under the same repo and the path-distance heuristic couldn't pick a winner. Resolve by renaming the service or declaring the intended pairing in `config.links`.
5. **Cross-impact says "stale".** Both sides need a fresh per-repo index _and_ a fresh group sync. Order matters: `gitnexus analyze` in each changed repo, then `gitnexus group sync <name>`. Use `gitnexus group status <name>` to see which side is behind.
## Related docs and references
- [AGENTS.md](../../AGENTS.md) — authoritative list of MCP tools and resources, including group-mode routing and the `gitnexus://group/…` resources.
- [ARCHITECTURE.md](../../ARCHITECTURE.md) — overall data flow and the call-resolution DAG that the per-repo indexer uses.
- [`gitnexus/src/core/group/`](../../gitnexus/src/core/group/) — `service.ts`, `sync.ts`, `config-parser.ts`, `matching.ts`.
- [`gitnexus/src/core/group/extractors/grpc-extractor.ts`](../../gitnexus/src/core/group/extractors/grpc-extractor.ts) and [`grpc-patterns/`](../../gitnexus/src/core/group/extractors/grpc-patterns/) — gRPC detection.
- [`gitnexus/src/core/group/extractors/manifest-extractor.ts`](../../gitnexus/src/core/group/extractors/manifest-extractor.ts) — the `config.links` escape hatch.
- [`gitnexus/src/mcp/tools.ts`](../../gitnexus/src/mcp/tools.ts) — MCP tool schemas (`group_list`, `group_sync`, plus `@<group>` routing on `impact` / `query` / `context`).
- [`gitnexus/src/cli/group.ts`](../../gitnexus/src/cli/group.ts) — CLI command definitions and flags.
- Upstream issues: [#701](https://github.com/abhigyanpatwari/GitNexus/issues/701), [#826](https://github.com/abhigyanpatwari/GitNexus/issues/826), [#906](https://github.com/abhigyanpatwari/GitNexus/issues/906).
-185
View File
@@ -1,185 +0,0 @@
# Using GitNexus across Apache Thrift microservices
## When to use this guide
Use this guide when several repositories communicate through Apache Thrift and you want GitNexus to trace impact across provider and consumer boundaries. The walkthrough assumes each service is indexed on its own, then joined through a GitNexus group.
This is not a framework integration guide. GitNexus reads portable Thrift IDL and common Java generated-code shapes. Framework-specific wiring, service discovery, deployment metadata, and private annotations belong outside the open-source core.
## Mental model
- `.thrift` files define the canonical service contract. A method in an IDL service becomes a stable contract id in the form `thrift::<namespace>.<Service>/<Method>`.
- Service wildcard ids in the form `thrift::<namespace>.<Service>/*` are supported as manifest and matching fallback forms when a service-level link is needed.
- Java generated-code usage points GitNexus toward implementation and call sites. Providers commonly implement generated `Service.Iface`; consumers commonly hold or construct generated service interfaces or clients.
- Group sync matches provider and consumer contracts with the same id, then cross-repo impact can hop through those links.
- Framework-specific wiring should be modeled by extractor plugins, manifest links, or downstream integrations rather than hard-coded into core Thrift support.
## Fictional IDL
```thrift
namespace java billing.v1
struct PlaceOrderRequest {
1: string orderId
2: double amount
}
struct PlaceOrderResponse {
1: bool accepted
}
struct GetOrderRequest {
1: string orderId
}
struct GetOrderResponse {
1: string orderId
2: string status
}
service OrderService {
PlaceOrderResponse PlaceOrder(1: PlaceOrderRequest request)
GetOrderResponse GetOrder(1: GetOrderRequest request)
}
```
The service methods above produce canonical ids:
- `thrift::billing.v1.OrderService/PlaceOrder`
- `thrift::billing.v1.OrderService/GetOrder`
- `thrift::billing.v1.OrderService/*` as a service-level manifest or matching fallback form
## Java provider example
Generated Java code usually exposes an `Iface` interface for the service. A provider implementation can be detected when it implements that generated interface.
```java
package example.billing;
import billing.v1.GetOrderRequest;
import billing.v1.GetOrderResponse;
import billing.v1.OrderService;
import billing.v1.PlaceOrderRequest;
import billing.v1.PlaceOrderResponse;
public final class OrderServiceHandler implements OrderService.Iface {
@Override
public PlaceOrderResponse PlaceOrder(PlaceOrderRequest request) {
return new PlaceOrderResponse(true);
}
@Override
public GetOrderResponse GetOrder(GetOrderRequest request) {
return new GetOrderResponse(request.getOrderId(), "CREATED");
}
}
```
With the IDL available, GitNexus can connect the implementation to `thrift::billing.v1.OrderService/PlaceOrder` and `thrift::billing.v1.OrderService/GetOrder`.
## Java consumer examples
Consumers are strongest when Java usage can be tied back to the IDL namespace and service.
```java
package example.checkout;
import billing.v1.OrderService;
import billing.v1.PlaceOrderRequest;
public final class CheckoutWorkflow {
private final OrderService.Iface orders;
public CheckoutWorkflow(OrderService.Iface orders) {
this.orders = orders;
}
public void submit(String orderId) throws Exception {
orders.PlaceOrder(new PlaceOrderRequest(orderId, 42.0));
}
}
```
Some generated-code styles use the generated service type directly while keeping enough IDL context through imports and method calls.
```java
package example.reporting;
import billing.v1.GetOrderRequest;
import billing.v1.OrderService;
public final class OrderLookup {
private final OrderService.Client client;
public OrderLookup(OrderService.Client client) {
this.client = client;
}
public String status(String orderId) throws Exception {
return client.GetOrder(new GetOrderRequest(orderId)).getStatus();
}
}
```
When IDL context is missing, GitNexus may still emit a weaker consumer signal for generated `Iface` or `Client` shapes, but confidence is lower.
## Group configuration
New group configs enable Thrift contract detection by default. Keep `detect.thrift: true`
when a group should scan for Thrift contracts, or set it to `false` to skip Thrift
extraction for that group.
```yaml
version: 1
name: billing-platform
description: Fictional services connected by Apache Thrift
repos:
checkout: checkout-service
billing: billing-service
links: []
detect:
http: true
grpc: false
thrift: true
topics: false
shared_libs: true
```
To disable Thrift extraction explicitly:
```yaml
detect:
thrift: false
```
After indexing each member repository, run group sync to extract contracts and write cross-repo links:
```bash
npx gitnexus group sync billing-platform
```
## Manifest escape hatch
Use manifest links when automatic extraction cannot see a provider or consumer, or when generated code is wrapped behind an abstraction. Write the contract without the `thrift::` prefix; GitNexus canonicalizes it to the full Thrift contract id.
```yaml
links:
- from: checkout
to: billing
type: thrift
contract: billing.v1.OrderService/PlaceOrder
role: consumer
```
GitNexus canonicalizes that manifest entry to `thrift::billing.v1.OrderService/PlaceOrder` and uses it to connect the two repositories.
## Known limitations
- Java detection currently targets v1 generated-code patterns.
- Maven and POM dependency coordinates are not used for inference.
- Framework-specific annotations and service discovery metadata are ignored by open-source Thrift extraction.
- Ambiguous same-name services are skipped instead of guessed.
- Java consumers without IDL context are lower confidence and limited to generated `Iface` and `Client` shapes.
@@ -1,326 +0,0 @@
---
title: "feat: Complete COBOL language feature coverage for maximum knowledge graph value"
type: feat
status: active
date: 2026-03-26
origin: Feature audit from v3-integration-architect agent (session 8642401e)
---
## Enhancement Summary
**Deepened on:** 2026-03-26
**Research agents used:** COBOL expert (Phase 1+2), graph value analyst, codebase explorer
**Sections enhanced:** Phase 1 (5 features), Phase 2 (4 features), graph value ranking
### Key Improvements from Research
1. **CALL USING** is the #1 highest-value edge type (9.2/10) — fixes ~40% of missing caller references
2. **EXEC DLI** requires dual-interface support (EXEC DLI + CBLTDLI CALL) for full IMS coverage
3. **DECLARATIVES** is lowest-risk Phase 2 item — existing section/paragraph detection already captures structure
4. **SET TO TRUE** accounts for 80-90% of all SET statements — prioritize this form
5. **INSPECT** needs multi-line accumulator (like SORT) — can span 5+ continuation lines
6. **Graph value ranking**: cobol-call-using (9.2) > cobol-error-handler (9.0) > dli-gu (8.2) > cobol-string (6.2)
### New Edge Cases Discovered
- CALL USING supports mixed modes: `USING BY REFERENCE WS-A BY CONTENT WS-B BY VALUE WS-C`
- CALL USING `ADDRESS OF` and `OMITTED` must be filtered from parameter lists
- EXEC DLI can have multiple SEGMENT levels in hierarchical retrieval (use matchAll)
- DECLARATIVES can have multiple USE sections (one per file + catch-all for INPUT/OUTPUT/I-O/EXTEND)
- INSPECT TALLYING can have multiple counters in a single statement
- STRING/UNSTRING can span multiple lines (need accumulator pattern)
---
# Complete COBOL Language Feature Coverage
## Overview
Implement the remaining 25 unhandled COBOL language features and fix 10 partial features to achieve ~95% coverage (up from 71.9%). The goal is to build the richest possible knowledge graph from COBOL codebases, enabling a future `modernize` MCP command (out of scope for this plan) that would use the graph to assist with COBOL-to-modern-language migration.
## Problem Statement
The COBOL processor currently handles 54 of 89 applicable language features (71.9%). The 25 unhandled features represent real data loss in the knowledge graph:
- **Cross-program data flow** is invisible (CALL ... USING parameters not extracted)
- **IMS/DB programs** produce empty graphs (EXEC DLI not recognized)
- **String transformation logic** is invisible (STRING/UNSTRING/INSPECT not tracked)
- **SQL copybook dependencies** are missing (EXEC SQL INCLUDE not mapped)
- **Error handling flows** are lost (DECLARATIVES/USE AFTER not captured)
## Proposed Solution
Implement features in 4 phases, ordered by graph value density (edges created per LOC of implementation). Each phase is independently shippable and testable.
## Technical Approach
### Phase 1: High-Value Data Flow Edges (~150 LOC, ~8 new edge types)
The highest-ROI features: they create new ACCESSES and IMPORTS edges that directly improve impact analysis.
**Critical research finding**: Multi-line statement accumulation is the dominant challenge. CALL USING, STRING/UNSTRING, and multi-line data item clauses all span multiple lines in production COBOL. The free-format path processes each line independently — these features need statement accumulators (like SORT/SELECT) or the free-format path needs multi-line awareness. Estimated LOC increased from 110 to 150 to account for accumulator infrastructure.
#### 1.1 EXEC SQL INCLUDE -> IMPORTS edges
- **File:** `cobol-preprocessor.ts` (parseExecSqlBlock)
- **What:** Detect `INCLUDE` as the operation, extract member name, emit as a `copies[]` entry
- **Graph:** IMPORTS edge from File to included copybook/SQLCA with reason `sql-include`
- **Tests:** Unit test for `EXEC SQL INCLUDE SQLCA END-EXEC` and `EXEC SQL INCLUDE CUSTCOPY END-EXEC`
**Research insights (EXEC SQL INCLUDE):**
- DB2 member names can contain underscores: `EXEC SQL INCLUDE CUST_TBL_DCL END-EXEC` — regex must use `[A-Z][A-Z0-9_-]+`
- Quoted literal form: `EXEC SQL INCLUDE 'DBRMLIB.MEMBER' END-EXEC` (z/OS PDS qualified name)
- SQLCA/SQLDA are DB2 builtins — won't resolve to repo files. Emit unresolved IMPORTS edge (still valuable)
- No REPLACING support on EXEC SQL INCLUDE (unlike COPY)
- Add `INCLUDE` to `OP_MAP` in `parseExecSqlBlock`; extract member via `RE_SQL_INCLUDE = /^INCLUDE\s+(?:'([^']+)'|"([^"]+)"|([A-Z][A-Z0-9_-]+))/i`
#### 1.2 CALL ... USING parameter extraction -> ACCESSES edges (Graph value: 9.2/10)
- **File:** `cobol-preprocessor.ts` (processLogicalLine CALL section)
- **What:** After capturing CALL target, scan for USING clause. Extract parameter names (reuse USING_KEYWORDS filter). Store as `calls[].parameters: string[]`
- **Interface:** Add `parameters?: string[]` to calls array type in CobolRegexResults
- **File:** `cobol-processor.ts` (CALL edge block)
- **Graph:** For each USING parameter, create ACCESSES edge from caller to data item Property node with reason `cobol-call-using`
- **Tests:** `CALL 'AUDITLOG' USING CUST-ID WS-AMOUNT` -> 2 ACCESSES edges
**Research insights (CALL USING forms):**
- Mixed modes: `CALL 'PGM' USING BY REFERENCE WS-A BY CONTENT WS-B BY VALUE WS-C`
- Pointer passing: `CALL 'PGM' USING ADDRESS OF WS-A`
- Placeholder: `CALL 'PGM' USING OMITTED WS-B`
- Filter keywords: add `ADDRESS`, `OMITTED`, `LENGTH` to USING_KEYWORDS (already has BY/VALUE/REFERENCE/CONTENT)
- **Impact tool enhancement:** CALL-USING edges enable BFS traversal through parameter data flow — single most impactful edge type for COBOL impact analysis
#### 1.3 STRING/UNSTRING data flow -> ACCESSES edges
- **File:** `cobol-preprocessor.ts` (new section in extractProcedure)
- **What:** Accumulate multi-line STRING/UNSTRING until period or END-STRING/END-UNSTRING. Extract sources and INTO targets.
- **Interface:** Add `strings: Array<{ sources: string[]; target: string; type: 'string' | 'unstring'; line: number; caller: string | null }>` to CobolRegexResults
- **Graph:** read-ACCESSES on sources, write-ACCESSES on INTO target with reason `cobol-string-read` / `cobol-string-write`
- **Tests:** 2 unit tests + integration test assertions
**Research insights (STRING/UNSTRING):**
- **Needs statement accumulator** — STRING/UNSTRING always span multiple lines in production
- Terminate accumulation at: period, END-STRING/END-UNSTRING, or start of next COBOL verb
- STRING sources: identifiers before each `DELIMITED BY`. Filter: STRING, DELIMITED, BY, SIZE, ALL, INTO, WITH, POINTER, ON, OVERFLOW, NOT, END-STRING
- UNSTRING: source is first identifier after UNSTRING; INTO targets are identifiers after INTO. Filter: DELIMITER, IN, COUNT, TALLYING, OR
- WITH POINTER field is both read AND written (starting position updated)
- TALLYING IN / COUNT IN fields are write targets
- Literal sources (`'text'`) must be filtered — quote-aware tokenization needed
- **Edge case**: STRING terminated by next verb, not period — existing fixture has `STRING ... DISPLAY` without period between them
#### 1.4 OCCURS DEPENDING ON -> ACCESSES edge
- **File:** `cobol-preprocessor.ts` (parseDataItemClauses)
- **What:** Extend OCCURS regex to capture DEPENDING ON field, KEY fields, and INDEXED BY names
- **Interface:** Add `dependingOn?: string`, `occursMax?: number`, `occursKeys?: Array<{direction: string; fields: string[]}>`, `indexedBy?: string[]` to data items
- **Graph:** ACCESSES edge from table item to controlling field with reason `cobol-depends-on`
- **Tests:** `05 WS-TABLE OCCURS 100 DEPENDING ON WS-COUNT` -> edge
**Research insights (OCCURS):**
- IBM allows `OCCURS 0 TO n DEPENDING ON` (zero minimum) and `OCCURS UNBOUNDED DEPENDING ON` (V6.4)
- Subscripted controlling fields: `DEPENDING ON WS-COUNT(WS-IDX)` — strip subscripts before storing
- **Pre-existing gap**: Multi-line data item clauses without continuation indicator are NOT captured. `05 WS-TABLE\n OCCURS 100\n DEPENDING ON WS-COUNT.` — the current RE_DATA_ITEM only gets the first line, `rest` is empty. Fixing properly requires a data item accumulator (like SELECT). **Defer full fix to Phase 3; implement same-line capture now.**
- KEY IS fields: `ASCENDING KEY IS WS-KEY-1 WS-KEY-2` — capture for SEARCH ALL resolution
- INDEXED BY: `INDEXED BY IDX-1 IDX-2` — capture for SET/SEARCH context
#### 1.5 VALUE clause for standard data items
- **File:** `cobol-preprocessor.ts` (parseDataItemClauses)
- **What:** Extract VALUE using a pragmatic function that handles quoted strings, numerics, figurative constants, hex/national literals
- **Interface:** Already exists as `values?: string[]` on data items (currently only populated for 88-level)
- **Graph:** Stored in Property node description (no new edges)
- **Tests:** `01 WS-STATUS PIC X VALUE 'A'` -> values: ['A']
**Research insights (VALUE forms):**
- Hex literals: `VALUE X'F1F2F3F4'`, National: `VALUE N'text'`, DBCS: `VALUE G'text'`
- Figurative constants: SPACES, ZEROS, ZEROES, LOW-VALUES, HIGH-VALUES, QUOTES, NULL, NULLS
- ALL literal: `VALUE ALL '*'`
- Numeric with sign/decimal: `VALUE -123.45`, `VALUE +1`
- `VALUE IS` optional — both `VALUE 'A'` and `VALUE IS 'A'` valid
- **Decimal vs period ambiguity**: `VALUE 100.` — is `.` decimal or terminator? `parseDataItemClauses` already strips trailing period, so this is handled
- IBM V6.4: floating-point `VALUE 1.0E5` — extend numeric regex if needed
- Implementation: use a pragmatic `extractValue(rest)` function, not a single complex regex
### Phase 2: EXEC DLI + DECLARATIVES (~90 LOC, ~4 new edge types)
IMS/DB support and error handling flows.
#### 2.1 EXEC DLI (IMS/DB) -> ACCESSES edges (Graph value: 8.2/10)
- **File:** `cobol-preprocessor.ts` (processLogicalLine — add RE_EXEC_DLI_START check alongside SQL/CICS)
- **What:** Accumulate EXEC DLI blocks like EXEC SQL. Parse DLI verbs (GU, GN, GNP, GHU, GHN, GHNP, ISRT, DLET, REPL, CHKP, SCHD, TERM). Extract segment name, PCB number, INTO/FROM areas, WHERE fields, PSB name.
- **Interface:** Add `execDliBlocks: Array<{ line: number; verb: string; pcbNumber?: number; segmentName?: string; intoField?: string; fromField?: string; whereField?: string; psbName?: string }>` to CobolRegexResults
- **Graph:** CodeElement node + ACCESSES edge to `<ims>:<segmentName>` Record node with reason `dli-{verb}`; ACCESSES edges to INTO/FROM data areas; PSB ACCESSES for SCHD
- **Tests:** `EXEC DLI GU USING PCB(1) SEGMENT(CUSTOMER) INTO(WS-CUST) END-EXEC`
**Research insights (dual IMS interface):**
- **EXEC DLI**: Embedded command interface for CICS-DL/I programs only
- **CBLTDLI CALL**: Batch interface via `CALL 'CBLTDLI' USING function-code PCB io-area SSA1..SSA15`
- CBLTDLI is already captured as a CALL to 'CBLTDLI' — enrich with USING parameter semantics later
- Multiple SEGMENT levels in hierarchical retrieval — use `matchAll` on segment regex
- DLI verbs: GU (most common), GN, GNP, GHU, GHN, GHNP, ISRT, REPL, DLET, CHKP, SCHD, TERM, ROLL, ROLB
- **Edge case**: DLET/REPL have no SEGMENT clause (operate on current position)
- **Recommended order**: Implement AFTER DECLARATIVES and SET (lower risk, higher frequency)
#### 2.2 DECLARATIVES / USE AFTER STANDARD EXCEPTION (Graph value: 9.0/10)
- **File:** `cobol-preprocessor.ts` (processLogicalLine — detect DECLARATIVES keyword, track USE AFTER blocks)
- **What:** When `DECLARATIVES.` is encountered, switch to declaratives mode. Extract USE statements binding sections to files/modes.
- **Interface:** Add `declaratives: Array<{ sectionName: string; useType: 'error' | 'debug' | 'label' | 'reporting'; target: string; line: number }>` to CobolRegexResults
- **Graph:** ACCESSES edge from declarative Namespace to file Record with reason `cobol-declarative-error-handler`
- **Tests:** Unit test with DECLARATIVES section, integration test for error flow
**Research insights (DECLARATIVES syntax):**
- `USE AFTER STANDARD {EXCEPTION|ERROR} ON {file-name|INPUT|OUTPUT|I-O|EXTEND}`
- EXCEPTION and ERROR are synonymous; STANDARD is optional in IBM dialects
- Multiple USE sections allowed (one per file + catch-all for I/O modes)
- `END DECLARATIVES.` must NOT reset PROCEDURE DIVISION state
- `DECLARATIVES` is already in EXCLUDED_PARA_NAMES — no false paragraph risk
- Existing section/paragraph detection already captures structural elements — just need USE binding
- **Lowest risk Phase 2 item** — implement first
#### 2.3 SET statement -> ACCESSES edges
- **File:** `cobol-preprocessor.ts` (extractProcedure — new RE_SET regex)
- **Interface:** Add `sets: Array<{ targets: string[]; form: 'to-true'|'to-value'|'up-by'|'down-by'|'address-of'|'to-null'|'to-entry'; value?: string; entryTarget?: string; entryIsLiteral?: boolean; line: number; caller: string | null }>` to CobolRegexResults
- **Graph:** ACCESSES write edge with reason `cobol-set-condition` (TO TRUE), `cobol-set-index` (TO/UP/DOWN), `cobol-set-address` (ADDRESS OF). SET ENTRY with literal -> CALLS edge.
- **Tests:** `SET WS-EOF TO TRUE`, `SET IDX-1 TO 5`, `SET IDX-1 UP BY 1`
**Research insights (SET forms by frequency):**
- `SET condition TO TRUE` — 80-90% of all SET usage. Multiple targets: `SET COND-A COND-B TO TRUE`
- `SET index TO/UP BY/DOWN BY` — ~8%. Multiple indices: `SET IDX-1 IDX-2 UP BY 1`
- `SET pointer TO ADDRESS OF data-item` / `SET ADDRESS OF data-item TO pointer` — ~2%
- `SET proc-ptr TO ENTRY "PROGNAME"` — rare but creates CALLS edge (like dynamic CALL)
- Filter OF/IN qualifiers: `SET COND-A OF WS-RECORD TO TRUE` (strip OF WS-RECORD)
- **Prioritize**: SET TO TRUE alone covers 80-90% — implement this form first
#### 2.4 INSPECT -> ACCESSES edges
- **File:** `cobol-preprocessor.ts` (extractProcedure — new `inspectAccum` accumulator like SORT)
- **What:** Accumulate multi-line INSPECT until period. Extract inspected field + tally counters.
- **Interface:** Add `inspects: Array<{ inspectedField: string; counters: string[]; form: 'tallying'|'replacing'|'converting'|'tallying-replacing'; line: number; caller: string | null }>` to CobolRegexResults
- **Graph:** ACCESSES read on inspected field always; write if REPLACING/CONVERTING. Write edges for tally counters. Reason: `cobol-inspect-read`/`cobol-inspect-write`/`cobol-inspect-tally`
- **Tests:** `INSPECT WS-FIELD TALLYING WS-COUNT FOR ALL 'A'` -> read on WS-FIELD, write on WS-COUNT
**Research insights (INSPECT forms by frequency):**
- REPLACING (~60%): `INSPECT WS-STR REPLACING ALL 'A' BY 'B'`
- TALLYING (~25%): `INSPECT WS-STR TALLYING WS-CNT FOR ALL 'A'` — multiple counters possible
- CONVERTING (~10%): `INSPECT WS-STR CONVERTING 'abc' TO 'ABC'`
- Combined (~5%): TALLYING + REPLACING in single statement
- **Needs multi-line accumulator** — INSPECT frequently spans 3-5 lines in production
- Extract tally counters with `([A-Z][A-Z0-9-]+)\s+FOR\b` matchAll pattern
- Filter figurative constants (SPACES, ZEROS) using existing MOVE_SKIP set
### Phase 3: Completeness Fixes (~60 LOC)
Fix the 10 partial features and small gaps.
#### 3.1 CALL ... RETURNING extraction
- Extend RE_CALL processing to capture RETURNING target after the USING clause
- Store as `calls[].returning?: string`
- Graph: ACCESSES write edge with reason `cobol-call-returning`
#### 3.2 SELECT OPTIONAL flag preservation
- Store `isOptional: boolean` in FileDeclaration interface
- Include in Record node description
#### 3.3 ALTERNATE RECORD KEY extraction
- Add regex in parseSelectStatement: `/\bALTERNATE\s+RECORD\s+KEY\s+(?:IS\s+)?([A-Z][A-Z0-9-]+)/i`
- Store as `alternateKeys?: string[]`
#### 3.4 COMMON attribute on nested programs
- Extend RE_PROGRAM_ID: `/\bPROGRAM-ID\.\s*([A-Z][A-Z0-9-]+)(?:\s+IS\s+COMMON)?/i`
- Store `isCommon: boolean` on Module node
- Affects cross-program CALL resolution scope
#### 3.5 IS EXTERNAL / IS GLOBAL as first-class properties
- Change from usage string hack to proper boolean fields on data items
- Add `isExternal?: boolean`, `isGlobal?: boolean` to data item interface
#### 3.6 AUTHOR / DATE-WRITTEN mapped to Module node
- Already extracted as programMetadata — map to Module node properties
- `graph.addNode({ ..., properties: { ..., author, dateWritten } })`
#### 3.7 REPLACE statement
- Track REPLACE / REPLACE OFF state in preprocessor
- Apply text substitutions during preprocessing (before regex extraction)
- Complex: requires careful scoping rules
### Phase 4: Niche Features (~30 LOC)
Low-priority but nice for completeness.
#### 4.1 INITIALIZE statement -> write ACCESSES
- `/\bINITIALIZE\s+([A-Z][A-Z0-9-]+)/i`
- ACCESSES write edge with reason `cobol-initialize`
#### 4.2 Remaining IDENTIFICATION DIVISION paragraphs
- DATE-COMPILED, INSTALLATION, SECURITY, REMARKS
- Map to Module node description properties
#### 4.3 EXEC SQL INCLUDE -> IMPORTS edge (expansion)
- For EXEC SQL INCLUDE inside EXEC blocks that reference copybooks containing SQL
- Create IMPORTS edge similar to COPY
## Acceptance Criteria
### Functional Requirements
- [ ] Phase 1: All 5 features implemented with unit + integration tests
- [ ] Phase 2: All 4 features implemented with unit + integration tests
- [ ] Phase 3: All 7 partial features fixed
- [ ] Phase 4: At least 2 of 3 niche features implemented
- [ ] All existing 145 tests continue to pass
- [ ] TypeScript compiles cleanly
### Non-Functional Requirements
- [ ] No performance regression: CardDemo benchmark stays under 8s
- [ ] No file exceeds 1500 LOC (preprocessor currently 1326)
- [ ] ACAS benchmark shows increased node/edge counts (more data extracted)
- [ ] CardDemo benchmark shows increased edge counts (CALL USING, STRING, etc.)
### Quality Gates
- [ ] Each phase has its own commit
- [ ] Integration test assertions updated with exact counts per phase
- [ ] Benchmark run after each phase to track graph growth
## Dependencies & Risks
### Dependencies
- None. All changes are additive to existing COBOL processor code.
- No LanguageProvider changes needed.
- No graph schema changes needed (all new constructs map to existing node labels + edge types).
### Risks
- **preprocessor.ts size**: Currently 1326 LOC. Phase 1+2 adds ~200 LOC -> 1526 LOC. May need to extract helpers into a separate `cobol-data-flow.ts` module if it exceeds 1500.
- **REPLACE statement** (Phase 3.7) is the most complex feature — requires tracking text substitution state across logical lines. Consider deferring to a separate PR if it takes >100 LOC.
- **EXEC DLI** (Phase 2.1) is only testable against IMS codebases. Need fixture data or synthetic test cases.
## Graph Value Ranking by MCP Tool Impact
Research agent analyzed all 5 MCP tools (query, context, impact, detect_changes, rename) against planned edge types:
| Edge Type | QUERY | CONTEXT | IMPACT | DETECT | RENAME | **Overall** |
|-----------|-------|---------|--------|--------|--------|-------------|
| `cobol-call-using` | 4/5 | 5/5 | 5/5 | 4/5 | 4/5 | **9.2/10** |
| `cobol-error-handler` | 5/5 | 4/5 | 5/5 | 5/5 | 2/5 | **9.0/10** |
| `dli-*` (IMS verbs) | 4/5 | 4/5 | 5/5 | 4/5 | 2/5 | **8.2/10** |
| `cobol-string-*` | 4/5 | 3/5 | 3/5 | 3/5 | 2/5 | **6.2/10** |
**Key finding**: `cobol-call-using` alone would fix ~40% of missing caller references in COBOL graphs.
## Future Considerations
This plan provides the graph data foundation for a future `modernize` MCP command (out of scope) that would:
- Use CALL USING edges to map data contracts between programs
- Use STRING/UNSTRING edges to identify data transformation logic
- Use EXEC SQL/DLI edges to map database access patterns
- Use DECLARATIVES to understand error handling architecture
- Use the complete knowledge graph to generate migration plans
**MCP tool enhancements needed** (after this plan ships):
- Add `cobol-call-using`, `cobol-error-handler`, `dli-*` to IMPACT tool's default `relationTypes` for COBOL repos
- Add confidence floors for new edge types in `IMPACT_RELATION_CONFIDENCE`
- Register new edge types in `VALID_RELATION_TYPES` set (`local-backend.ts:52`)
## Sources & References
### Internal References
- Feature audit: session 8642401e (COBOL expert agent, 123 features audited)
- Prior plans: `docs/plans/2026-03-25-feat-cobol-100-percent-feature-coverage-plan.md`
- Architecture: `docs/code-indexing/cobol/` (7 documentation files)
### External References
- COBOL features reference: mainframestechhelp.com/tutorials/cobol/features.htm
- COBOL-85 standard: ISO/IEC 1989:1985
- IBM Enterprise COBOL reference
@@ -1,725 +0,0 @@
# PR #626 HIGH-Priority Fixes Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Fix 4 HIGH-priority issues from PR #626 code review before merge.
**Architecture:** Minimal targeted fixes — each task is independent. TDD: tests first, then implementation. No refactoring beyond what's needed.
**Tech Stack:** TypeScript, Vitest, Node.js fs/path APIs
**Spec:** `docs/superpowers/specs/2026-04-02-pr626-high-fixes-design.md`
**Paths:** All file paths are relative to the monorepo root (`GitNexus/`). Git commands run from the root. The `gitnexus/` prefix is a package subdirectory, not a separate repo.
---
### Task 1: Path Traversal — Validate Group Name
**Files:**
- Modify: `gitnexus/src/core/group/storage.ts:17-19` (getGroupDir) and `:63-68` (createGroupDir)
- Test: `gitnexus/test/unit/group/storage.test.ts`
- [ ] **Step 1: Write failing tests for validateGroupName**
In `gitnexus/test/unit/group/storage.test.ts`, add `createGroupDir` and `validateGroupName` to the existing import from `'../../../src/core/group/storage.js'` (line 6-11). Then add these describe blocks at the end of the outer `describe('Group storage', ...)`:
```typescript
describe('validateGroupName', () => {
it('test_validateGroupName_traversal_path_throws', () => {
expect(() => validateGroupName('../../evil')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_slash_in_name_throws', () => {
expect(() => validateGroupName('foo/bar')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_empty_string_throws', () => {
expect(() => validateGroupName('')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_starts_with_dash_throws', () => {
expect(() => validateGroupName('-leading-dash')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_starts_with_underscore_throws', () => {
expect(() => validateGroupName('_leading')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_dots_throws', () => {
expect(() => validateGroupName('com.example')).toThrow(/Invalid group name/);
});
it('test_validateGroupName_valid_alphanumeric_passes', () => {
expect(() => validateGroupName('my-group_01')).not.toThrow();
});
it('test_validateGroupName_single_char_passes', () => {
expect(() => validateGroupName('A')).not.toThrow();
});
it('test_validateGroupName_all_digits_passes', () => {
expect(() => validateGroupName('123')).not.toThrow();
});
});
describe('getGroupDir rejects invalid names', () => {
it('test_getGroupDir_traversal_throws', () => {
expect(() => getGroupDir(tmpDir, '../../etc')).toThrow(/Invalid group name/);
});
it('test_getGroupDir_valid_name_returns_path', () => {
const dir = getGroupDir(tmpDir, 'company');
expect(dir).toBe(path.join(tmpDir, 'groups', 'company'));
});
});
describe('createGroupDir rejects invalid names', () => {
it('test_createGroupDir_traversal_throws', async () => {
await expect(createGroupDir(tmpDir, '../evil')).rejects.toThrow(/Invalid group name/);
});
});
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `cd gitnexus && npx vitest run test/unit/group/storage.test.ts`
Expected: FAIL — `validateGroupName` is not exported, `getGroupDir` does not throw.
- [ ] **Step 3: Implement validateGroupName and wire into getGroupDir and createGroupDir**
In `gitnexus/src/core/group/storage.ts`, add the validation function before `getGroupDir` and call it:
```typescript
const GROUP_NAME_RE = /^[a-zA-Z0-9][a-zA-Z0-9_-]*$/;
export function validateGroupName(name: string): void {
if (!GROUP_NAME_RE.test(name)) {
throw new Error(
`Invalid group name "${name}". Names must start with a letter or digit and contain only [a-zA-Z0-9_-].`,
);
}
}
export function getGroupDir(gitnexusDir: string, groupName: string): string {
validateGroupName(groupName);
return path.join(gitnexusDir, 'groups', groupName);
}
```
`createGroupDir` already calls `getGroupDir` at line 68, so it inherits validation automatically. No change needed in `createGroupDir`.
- [ ] **Step 4: Run tests to verify they pass**
Run: `cd gitnexus && npx vitest run test/unit/group/storage.test.ts`
Expected: ALL PASS
- [ ] **Step 5: Commit**
```bash
cd gitnexus && git add src/core/group/storage.ts test/unit/group/storage.test.ts
git commit -m "fix(group): validate group name to prevent path traversal
Add validateGroupName() with regex [a-zA-Z0-9][a-zA-Z0-9_-]*.
Called in getGroupDir (defense in depth) which covers all CLI entry
points: create, add, remove, status, sync.
Addresses PR #626 review item 1 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
### Task 2: Directory Exclusions in Service Boundary Detector
**Files:**
- Modify: `gitnexus/src/core/group/service-boundary-detector.ts:24-51` (add constant), `:78` (walkForBoundaries), `:130` (hasSourceFilesInSubdirs)
- Test: `gitnexus/test/unit/group/service-boundary-detector.test.ts`
- [ ] **Step 1: Write failing tests for excluded directories**
Add this describe block inside the existing `detectServiceBoundaries` describe in `gitnexus/test/unit/group/service-boundary-detector.test.ts`:
```typescript
it('test_detect_skips_vendor_directory', async () => {
writeFile('services/auth/package.json', '{}');
writeFile('services/auth/src/index.ts', '');
// vendor should be skipped — its contents should not create a boundary
writeFile('vendor/some-dep/package.json', '{}');
writeFile('vendor/some-dep/src/lib.go', '');
const boundaries = await detectServiceBoundaries(tmpDir);
const paths = boundaries.map((b) => b.servicePath);
expect(paths).toContain('services/auth');
expect(paths).not.toContain('vendor/some-dep');
});
it('test_detect_skips_target_directory', async () => {
writeFile('services/api/go.mod', 'module api');
writeFile('services/api/main.go', '');
writeFile('target/classes/Main.java', '');
writeFile('target/pom.xml', '<project/>');
const boundaries = await detectServiceBoundaries(tmpDir);
const paths = boundaries.map((b) => b.servicePath);
expect(paths).toContain('services/api');
expect(paths).not.toContain('target');
});
it('test_detect_skips_pycache_directory', async () => {
writeFile('services/ml/pyproject.toml', '[project]');
writeFile('services/ml/model.py', '');
// __pycache__ with a marker + source files — would be detected as
// a boundary if not excluded, since it has package.json + .py file
writeFile('__pycache__/package.json', '{}');
writeFile('__pycache__/cached.py', '');
const boundaries = await detectServiceBoundaries(tmpDir);
const paths = boundaries.map((b) => b.servicePath);
expect(paths).toContain('services/ml');
expect(paths.every((p) => !p.includes('__pycache__'))).toBe(true);
});
it('test_detect_skips_dotfile_directories_regression', async () => {
writeFile('services/api/package.json', '{}');
writeFile('services/api/src/index.ts', '');
writeFile('.hidden/package.json', '{}');
writeFile('.hidden/src/index.ts', '');
const boundaries = await detectServiceBoundaries(tmpDir);
const paths = boundaries.map((b) => b.servicePath);
expect(paths).toContain('services/api');
expect(paths).not.toContain('.hidden');
});
it('test_detect_does_not_skip_regular_source_directories', async () => {
writeFile('services/api/package.json', '{}');
writeFile('services/api/src/index.ts', '');
const boundaries = await detectServiceBoundaries(tmpDir);
expect(boundaries).toHaveLength(1);
expect(boundaries[0].serviceName).toBe('api');
});
```
- [ ] **Step 2: Run tests to verify `vendor` and `target` tests fail**
Run: `cd gitnexus && npx vitest run test/unit/group/service-boundary-detector.test.ts`
Expected: `test_detect_skips_vendor_directory` and `test_detect_skips_target_directory` FAIL (vendor/target not excluded). Other new tests may pass since dotfile exclusion already exists.
- [ ] **Step 3: Add EXCLUDED_DIRS constant and update both walking functions**
In `gitnexus/src/core/group/service-boundary-detector.ts`:
After `SOURCE_EXTENSIONS` (after line 51), add:
```typescript
const EXCLUDED_DIRS = new Set([
'node_modules',
'vendor',
'target',
'build',
'dist',
'__pycache__',
'.venv',
'venv',
'.tox',
'.mypy_cache',
'.gradle',
'.mvn',
'out',
'bin',
]);
```
In `walkForBoundaries`, replace line 78:
```typescript
if (entry.name.startsWith('.') || entry.name === 'node_modules') continue;
```
with:
```typescript
if (entry.name.startsWith('.') || EXCLUDED_DIRS.has(entry.name)) continue;
```
In `hasSourceFilesInSubdirs`, replace line 130:
```typescript
if (entry.isDirectory() && !entry.name.startsWith('.') && entry.name !== 'node_modules') {
```
with:
```typescript
if (entry.isDirectory() && !entry.name.startsWith('.') && !EXCLUDED_DIRS.has(entry.name)) {
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `cd gitnexus && npx vitest run test/unit/group/service-boundary-detector.test.ts`
Expected: ALL PASS
- [ ] **Step 5: Commit**
```bash
cd gitnexus && git add src/core/group/service-boundary-detector.ts test/unit/group/service-boundary-detector.test.ts
git commit -m "fix(group): add directory exclusions to service boundary detector
Add EXCLUDED_DIRS set: vendor, target, build, dist, __pycache__,
.venv, venv, .tox, .mypy_cache, .gradle, .mvn, out, bin.
Applied in walkForBoundaries and hasSourceFilesInSubdirs.
Replaces inline node_modules check.
Addresses PR #626 review item 3 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
### Task 3: Remove Double-Close of LadybugDB Pools
**Files:**
- Modify: `gitnexus/src/cli/group.ts:160` (remove import), `:187-189` (remove finally block body)
- Test: `gitnexus/test/unit/group/sync.test.ts` (add pool cleanup test)
- Test: `gitnexus/test/integration/group/group-cli.test.ts` (verify no blanket close in source)
- [ ] **Step 1: Write unit tests for per-id pool cleanup in sync.ts**
Add to `gitnexus/test/unit/group/sync.test.ts`, inside the existing `describe('syncGroup', ...)`:
```typescript
it('test_syncGroup_closes_only_opened_pools', async () => {
const config = makeConfig({
'app/backend': 'backend-repo',
'app/frontend': 'frontend-repo',
});
const closedIds: string[] = [];
// Mock initLbug/closeLbug via per-repo override that tracks pool lifecycle
const { vi } = await import('vitest');
const poolAdapter = await import('../../../src/core/lbug/pool-adapter.js');
const initSpy = vi.spyOn(poolAdapter, 'initLbug').mockResolvedValue(undefined);
const closeSpy = vi.spyOn(poolAdapter, 'closeLbug').mockImplementation(async (id?: string) => {
if (id) closedIds.push(id);
});
try {
await syncGroup(config, {
resolveRepoHandle: async (_name, groupPath) => ({
id: groupPath.replace(/\//g, '-'),
path: groupPath,
repoPath: '/tmp/' + groupPath,
storagePath: '/tmp/' + groupPath + '/.gitnexus',
}),
skipWrite: true,
}).catch(() => {});
// Regardless of extraction errors, closeLbug should be called per id
// closeLbug should only receive specific pool ids, never undefined/empty
for (const id of closedIds) {
expect(id).toBeTruthy();
expect(typeof id).toBe('string');
}
// No blanket close (no-arg call)
const blanketCalls = closeSpy.mock.calls.filter((args) => args.length === 0 || !args[0]);
expect(blanketCalls).toHaveLength(0);
} finally {
initSpy.mockRestore();
closeSpy.mockRestore();
}
});
```
- [ ] **Step 2: Run sync unit test to verify it passes (sync.ts already does per-id cleanup)**
Run: `cd gitnexus && npx vitest run test/unit/group/sync.test.ts`
Expected: PASS — sync.ts already cleans up correctly. This test locks the behavior.
- [ ] **Step 3: Write test verifying CLI source has no blanket closeLbug()**
Add to `gitnexus/test/integration/group/group-cli.test.ts`:
```typescript
it('test_sync_command_source_does_not_call_blanket_closeLbug', () => {
const cliGroupPath = path.join(repoRoot, 'src', 'cli', 'group.ts');
const source = fs.readFileSync(cliGroupPath, 'utf-8');
// closeLbug() without arguments (blanket close) must not appear.
// closeLbug(id) with argument is fine (that's in sync.ts, not here).
// Match closeLbug() but not closeLbug(someArg)
const blanketClosePattern = /closeLbug\s*\(\s*\)/;
expect(source).not.toMatch(blanketClosePattern);
});
```
- [ ] **Step 4: Run test to verify it fails**
Run: `cd gitnexus && npx vitest run test/integration/group/group-cli.test.ts`
Expected: FAIL — `closeLbug()` (no args) exists at line 188.
- [ ] **Step 5: Remove blanket closeLbug() from cli/group.ts**
In `gitnexus/src/cli/group.ts`:
Remove the `closeLbug` import at line 160:
```typescript
const { closeLbug } = await import('../core/lbug/pool-adapter.js');
```
Replace the try/finally wrapper (lines 162-189):
```typescript
try {
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
console.log(`Syncing group "${name}" (${Object.keys(config.repos).length} repos)...\n`);
const result = await syncGroup(config, {
groupDir,
allowStale: Boolean(opts.allowStale),
verbose: Boolean(opts.verbose),
skipEmbeddings: Boolean(opts.skipEmbeddings),
exactOnly: Boolean(opts.exactOnly),
});
if (opts.json) {
console.log(JSON.stringify(result, null, 2));
} else {
console.log(`\nMatching cascade:`);
const exactLinks = result.crossLinks.filter((l) => l.matchType === 'exact');
console.log(` exact: ${exactLinks.length} cross-links (confidence 1.0)`);
console.log(` unmatched: ${result.unmatched.length} contracts`);
console.log(
`\nWrote contracts.json (${result.contracts.length} contracts, ${result.crossLinks.length} cross-links)`,
);
}
} finally {
await closeLbug().catch(() => {});
}
```
Becomes (remove try/finally entirely, since sync.ts handles its own cleanup):
```typescript
const groupDir = getGroupDir(getDefaultGitnexusDir(), name);
const config = await loadGroupConfig(groupDir);
console.log(`Syncing group "${name}" (${Object.keys(config.repos).length} repos)...\n`);
const result = await syncGroup(config, {
groupDir,
allowStale: Boolean(opts.allowStale),
verbose: Boolean(opts.verbose),
skipEmbeddings: Boolean(opts.skipEmbeddings),
exactOnly: Boolean(opts.exactOnly),
});
if (opts.json) {
console.log(JSON.stringify(result, null, 2));
} else {
console.log(`\nMatching cascade:`);
const exactLinks = result.crossLinks.filter((l) => l.matchType === 'exact');
console.log(` exact: ${exactLinks.length} cross-links (confidence 1.0)`);
console.log(` unmatched: ${result.unmatched.length} contracts`);
console.log(
`\nWrote contracts.json (${result.contracts.length} contracts, ${result.crossLinks.length} cross-links)`,
);
}
```
- [ ] **Step 6: Run tests to verify they pass**
Run: `cd gitnexus && npx vitest run test/integration/group/group-cli.test.ts test/unit/group/sync.test.ts`
Expected: ALL PASS
- [ ] **Step 7: Commit**
```bash
cd gitnexus && git add src/cli/group.ts test/integration/group/group-cli.test.ts test/unit/group/sync.test.ts
git commit -m "fix(group): remove blanket closeLbug() from CLI sync command
sync.ts already closes pools per-id in its finally block.
The blanket closeLbug() in cli/group.ts tears down ALL active pools
including unrelated ones in MCP server context.
Addresses PR #626 review item 4 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
### Task 4: gRPC Proto Regex — Brace-Depth Counter
**Files:**
- Modify: `gitnexus/src/core/group/extractors/grpc-extractor.ts:101-130` (parseProtoFile)
- Test: `gitnexus/test/unit/group/grpc-extractor.test.ts`
- [ ] **Step 1: Write failing tests for nested braces in proto services**
Add this describe block inside the existing `proto file parsing` describe in `gitnexus/test/unit/group/grpc-extractor.test.ts`:
```typescript
it('test_extract_proto_with_google_api_http_nested_braces', async () => {
writeFile(
'api/gateway.proto',
`syntax = "proto3";
package gateway.v1;
import "google/api/annotations.proto";
service GatewayService {
rpc GetUser (GetUserRequest) returns (UserResponse) {
option (google.api.http) = {
get: "/v1/users/{user_id}"
};
}
rpc CreateUser (CreateUserRequest) returns (UserResponse) {
option (google.api.http) = {
post: "/v1/users"
body: "*"
};
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter(
(c) => c.role === 'provider' && c.symbolRef.filePath === 'api/gateway.proto',
);
expect(providers).toHaveLength(2);
const ids = providers.map((c) => c.contractId).sort();
expect(ids).toEqual([
'grpc::gateway.v1.GatewayService/CreateUser',
'grpc::gateway.v1.GatewayService/GetUser',
]);
});
it('test_extract_proto_with_multiple_services', async () => {
writeFile(
'api/multi.proto',
`syntax = "proto3";
package multi;
service ServiceA {
rpc MethodA (Req) returns (Res);
}
service ServiceB {
rpc MethodB1 (Req) returns (Res);
rpc MethodB2 (Req) returns (Res);
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter(
(c) => c.role === 'provider' && c.symbolRef.filePath === 'api/multi.proto',
);
expect(providers).toHaveLength(3);
const ids = providers.map((c) => c.contractId).sort();
expect(ids).toEqual([
'grpc::multi.ServiceA/MethodA',
'grpc::multi.ServiceB/MethodB1',
'grpc::multi.ServiceB/MethodB2',
]);
});
it('test_extract_proto_with_nested_option_blocks_in_rpc', async () => {
writeFile(
'api/nested.proto',
`syntax = "proto3";
package nested;
service DeepService {
rpc DeepMethod (Req) returns (Res) {
option (google.api.http) = {
post: "/v1/deep"
body: "*"
additional_bindings {
get: "/v1/deep/{id}"
}
};
}
}`,
);
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter(
(c) => c.role === 'provider' && c.symbolRef.filePath === 'api/nested.proto',
);
expect(providers).toHaveLength(1);
expect(providers[0].contractId).toBe('grpc::nested.DeepService/DeepMethod');
});
it('test_extract_proto_malformed_unclosed_brace_skips_service', async () => {
writeFile(
'api/broken.proto',
`syntax = "proto3";
package broken;
service IncompleteService {
rpc SomeMethod (Req) returns (Res);
// Missing closing brace — EOF before depth returns to 0
`,
);
// Should not throw; incomplete service is silently skipped
const contracts = await extractor.extract(null, tmpDir, makeRepo(tmpDir));
const providers = contracts.filter(
(c) => c.role === 'provider' && c.symbolRef.filePath === 'api/broken.proto',
);
// The old regex would find partial match; the new parser should skip it
expect(providers).toHaveLength(0);
});
```
- [ ] **Step 2: Run tests to verify the nested brace test fails**
Run: `cd gitnexus && npx vitest run test/unit/group/grpc-extractor.test.ts`
Expected: `test_extract_proto_with_google_api_http_nested_braces` FAIL — regex stops at first `}` inside the `option` block.
- [ ] **Step 3: Replace serviceRe regex with extractServiceBlocks function**
In `gitnexus/src/core/group/extractors/grpc-extractor.ts`, replace the `parseProtoFile` method (lines 101-130):
```typescript
private parseProtoFile(content: string, filePath: string): ExtractedContract[] {
const out: ExtractedContract[] = [];
const pkgMatch = content.match(/^package\s+([\w.]+)\s*;/m);
const pkg = pkgMatch ? pkgMatch[1] : '';
for (const { name: serviceName, body } of extractServiceBlocks(content)) {
const rpcRe = /rpc\s+(\w+)\s*\(/g;
let rpcMatch: RegExpExecArray | null;
while ((rpcMatch = rpcRe.exec(body)) !== null) {
const methodName = rpcMatch[1];
const cid = contractId(pkg, serviceName, methodName);
out.push(
makeContract(cid, 'provider', filePath, `${serviceName}.${methodName}`, 0.85, {
package: pkg,
service: serviceName,
method: methodName,
source: 'proto',
}),
);
}
}
return out;
}
```
Add this function before the class (e.g. after `serviceOnlyContractId`, around line 26):
```typescript
function extractServiceBlocks(content: string): Array<{ name: string; body: string }> {
const results: Array<{ name: string; body: string }> = [];
const headerRe = /service\s+(\w+)\s*\{/g;
let headerMatch: RegExpExecArray | null;
while ((headerMatch = headerRe.exec(content)) !== null) {
const serviceName = headerMatch[1];
const bodyStart = headerMatch.index + headerMatch[0].length;
let depth = 1;
let pos = bodyStart;
while (pos < content.length && depth > 0) {
const ch = content[pos];
if (ch === '{') depth++;
else if (ch === '}') depth--;
pos++;
}
// If EOF before depth returns to 0, skip incomplete service
if (depth !== 0) continue;
// body is between opening { (consumed by regex) and closing } (pos is one past it)
const body = content.slice(bodyStart, pos - 1);
results.push({ name: serviceName, body });
}
return results;
}
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `cd gitnexus && npx vitest run test/unit/group/grpc-extractor.test.ts`
Expected: ALL PASS (including existing regression tests)
- [ ] **Step 5: Commit**
```bash
cd gitnexus && git add src/core/group/extractors/grpc-extractor.ts test/unit/group/grpc-extractor.test.ts
git commit -m "fix(group): replace gRPC proto regex with brace-depth counter
The serviceRe regex used [^}]* which stopped at the first '}'.
Proto services with google.api.http annotations contain nested {}
blocks, causing methods to be missed.
New extractServiceBlocks() uses a brace-depth counter (init depth=1
after opening {, scan char-by-char). Malformed protos with unclosed
braces are silently skipped.
Addresses PR #626 review item 2 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
---
### Task 5: Run Full Test Suite
- [ ] **Step 1: Run all group-related tests**
Run: `cd gitnexus && npx vitest run test/unit/group/ test/integration/group/`
Expected: ALL PASS
- [ ] **Step 2: Run full test suite to catch regressions**
Run: `cd gitnexus && npx vitest run`
Expected: ALL PASS, 0 failures
- [ ] **Step 3: Run typecheck**
Run: `cd gitnexus && npx tsc --noEmit`
Expected: No errors
---
### Task 6: CLI Integration Smoke Test
- [ ] **Step 1: Add CLI smoke test for path traversal**
Add to `gitnexus/test/integration/group/group-cli.test.ts` inside the existing `group CLI` describe:
```typescript
it('test_create_with_invalid_name_fails', () => {
const result = runGroup(['create', '../../evil']);
expect(result.status).not.toBe(0);
expect(result.stderr).toContain('Invalid group name');
});
```
- [ ] **Step 2: Run test**
Run: `cd gitnexus && npx vitest run test/integration/group/group-cli.test.ts`
Expected: ALL PASS
- [ ] **Step 3: Commit**
```bash
cd gitnexus && git add test/integration/group/group-cli.test.ts
git commit -m "test(group): add CLI smoke test for path traversal rejection
Verifies that 'group create ../../evil' fails with Invalid group name.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>"
```
@@ -1,175 +0,0 @@
# PR #626 HIGH-Priority Fixes Design
**Date:** 2026-04-02
**PR:** abhigyanpatwari/GitNexus#626 — Intra-repo service communication tracking
**Scope:** 4 HIGH-priority issues identified by abhigyanpatwari and xkonjin
**Approach:** Minimal targeted fixes (option A) — no refactoring, no scope creep
---
## Fix 1: Path Traversal via Group Name
**File:** `gitnexus/src/core/group/storage.ts`
**Risk:** A group name like `../../etc` creates directories outside the intended path.
### Solution
Add `validateGroupName(name: string): void` that enforces `/^[a-zA-Z0-9][a-zA-Z0-9_-]*$/`.
- Call in `createGroupDir` (primary entry point)
- Call in `getGroupDir` (defense in depth)
- Throw descriptive error on invalid names
**Legacy:** Groups already on disk with names outside this pattern are not auto-renamed; only new `create` / resolved paths are validated.
### Why regex over path.resolve + startsWith
- abhigyanpatwari explicitly requested `[a-zA-Z0-9_-]`
- Stricter: disallows spaces, dots, Unicode edge cases
- Simpler to reason about
### Tests
- `../../evil` throws
- `foo/bar` throws
- Empty string throws
- `my-group_01` passes
- `A` (single char) passes
- CLI smoke: one integration test that hits `getGroupDir` / `createGroupDir` (e.g. `group create` or `group add`) with an invalid name proves wiring for every subcommand that resolves a group through storage
### CLI/API entry points accepting groupName
All paths flow through `getGroupDir` (which validates), so coverage is implicit. For reference:
| Command | Entry | Calls |
|---------|-------|-------|
| `group create` | `cli/group.ts` action | `createGroupDir` -> `getGroupDir` |
| `group add` | `cli/group.ts` action | `getGroupDir` |
| `group remove` | `cli/group.ts` action | `getGroupDir` |
| `group list` | `cli/group.ts` action | reads `groups/` dir directly — no traversal risk (reads, not writes) |
| `group status` | `cli/group.ts` action | `getGroupDir` |
| `group sync` | `cli/group.ts` action | `getGroupDir` |
**`listGroups`:** Reads directory names from disk without validation. Not a write path, so no traversal risk. May surface manually-created directories with non-conforming names — accepted as-is, not in scope.
---
## Fix 2: gRPC Proto Regex -> Brace-Depth Counter
**File:** `gitnexus/src/core/group/extractors/grpc-extractor.ts`
**Risk:** `serviceRe = /service\s+(\w+)\s*\{([^}]*)}/gs` stops at first `}`. Proto services with `google.api.http` annotations inside RPCs contain nested `{ }` blocks.
### Solution
Replace `serviceRe` regex with `extractServiceBlocks(content: string): Array<{ name: string; body: string }>`:
1. Use regex only to find `service <Name> {` start positions (regex consumes the opening `{`)
2. Initialise depth to 1 immediately after the opening `{`
3. Scan forward char by char: `{` -> depth++, `}` -> depth--; collect into body
4. Stop when depth reaches 0 (the matching closing `}`)
5. Return name + body pairs
Inner `rpcRe` regex remains unchanged — it operates on the already-extracted body.
**Malformed input:** If EOF is reached before `depth` returns to 0, skip the incomplete service (do not add to results). Lock this in the test.
**Scope limitation (v1):** Brace-depth only — no lexer for string literals or comments containing `{`/`}`. Sufficient for `google.api.http` annotations. Known false positive: braces inside `//` comments or quoted strings within proto options. Accepted for v1; a proper proto lexer is out of scope.
### Tests
- Proto with single service, no nesting (regression)
- Proto with `google.api.http` nested braces inside RPC options
- Proto with multiple services
- Proto with nested `option` blocks inside RPC (e.g. `google.api.http`)
- Malformed proto with unclosed brace (graceful handling)
---
## Fix 3: Directory Exclusions in Service Boundary Detector
**File:** `gitnexus/src/core/group/service-boundary-detector.ts`
**Risk:** Walks entire repo tree, only skipping dotfiles and `node_modules`. Extremely slow on repos with `vendor/`, `target/`, `__pycache__/`, `.venv/`.
### Solution
Create `EXCLUDED_DIRS` as a `Set<string>` (alongside existing `SERVICE_MARKERS`, `SOURCE_EXTENSIONS`), for example:
```text
node_modules, vendor, target, build, dist,
__pycache__, .venv, venv, .tox, .mypy_cache,
.gradle, .mvn, out, bin
```
(Implement as `new Set([...])` — the list above is the membership, not a string literal.)
Apply in both:
- `walkForBoundaries` (line 77-78) — replace current inline `=== 'node_modules'` check with `EXCLUDED_DIRS.has(entry.name)`
- `hasSourceFilesInSubdirs` (line 130) — replace `entry.name !== 'node_modules'` with `!EXCLUDED_DIRS.has(entry.name)`
Note: remove the old `=== 'node_modules'` literal from both locations — it is covered by `EXCLUDED_DIRS`.
Dotfile exclusion (`.` prefix) remains as a separate check since it's a pattern, not a name.
Exclusions apply only to `isDirectory()` entries — file names are never checked against `EXCLUDED_DIRS`.
**Tradeoff:** Rare layouts that keep source under names like `out/` or `bin/` will be skipped; accepted for performance on typical monorepos.
**Case sensitivity:** `Set.has` is case-sensitive (matches current `=== 'node_modules'` behavior). Windows case-insensitive FS not handled — accepted as-is, consistent with existing code.
### Tests
- Directory named `vendor/` is skipped
- Directory named `target/` is skipped
- Directory named `__pycache__/` is skipped
- Regular source directories are NOT skipped
- Dotfile directories still skipped (regression)
---
## Fix 4: Double-Close of LadybugDB Pools
**Files:**
- `gitnexus/src/core/group/sync.ts` (lines 155-157) — per-id cleanup (KEEP)
- `gitnexus/src/cli/group.ts` (line 188) — blanket `closeLbug()` (REMOVE)
**Risk:** In MCP server context, `closeLbug()` without arguments tears down ALL active pools, including ones from unrelated operations.
### Solution
Remove the `closeLbug()` call (no arguments) from `cli/group.ts` finally block. The per-id cleanup in `sync.ts` is sufficient:
```typescript
// sync.ts — KEEP: cleans up only pools opened by this sync
finally {
for (const id of [...new Set(openPoolIds)]) {
await closeLbug(id).catch(() => {});
}
}
```
```typescript
// cli/group.ts — REMOVE: blanket close that kills all pools
finally {
await closeLbug().catch(() => {}); // DELETE THIS
}
```
Remove the `closeLbug` import from `cli/group.ts` — after removing the `finally` call it has no remaining usages.
### Tests (unit level — mock pool adapter)
- `syncGroup` closes only the pools it opened (mock `closeLbug`, assert called with specific ids)
- Two-pool scenario: sync opens pools A and B, both closed in finally; pool C (opened elsewhere) not touched
- CLI `sync` command does not call blanket `closeLbug()` (verify no zero-arg call in source — static check or grep-based test)
---
## Out of Scope
- JSON -> LadybugDB migration (tracked in #606)
- MEDIUM/LOW issues (items 5-10 from review summary)
- Test gap coverage beyond what's needed for these 4 fixes
- Any refactoring or architectural changes
## Execution Order
Fixes are independent — can be implemented in parallel or any order.
Recommended order for review clarity: 1 -> 3 -> 4 -> 2 (simplest to most complex).
+57 -2
View File
@@ -162,8 +162,8 @@ Each mode has a `system_{mode}.jinja` + `instance_{mode}.jinja` pair. The agent
```
Agent → bash command → /usr/local/bin/gitnexus-query
→ curl localhost:4848/tool/query (fast path: eval-server, ~100ms)
→ npx gitnexus query (fallback: cold CLI, ~5-10s)
→ curl http://127.0.0.1:4848/tool/query (fast path: eval-server, ~100ms)
→ npx gitnexus query (fallback: cold CLI, ~5-10s)
```
Each tool script in `/usr/local/bin/` is standalone — no sourcing, no env inheritance needed. This is critical because mini-swe-agent runs every command via `subprocess.run` in a fresh subshell.
@@ -176,6 +176,61 @@ The eval-server is a lightweight HTTP daemon that:
- Includes next-step hints to guide tool chaining (query → context → impact → fix)
- Auto-shuts down after idle timeout
**CLI flags:**
| Flag | Default | Purpose |
|------|---------|---------|
| `--port <port>` | `4848` | Port to listen on |
| `--host <host>` | `127.0.0.1` | Bind address — use `0.0.0.0` for cross-container access |
| `--idle-timeout <seconds>` | `0` (disabled) | Auto-shutdown after N seconds of inactivity |
**READY signal:**
When the server is ready, it writes to stdout:
```
# IPv4
GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848
# IPv6 (bracketed to avoid colon ambiguity)
GITNEXUS_EVAL_SERVER_READY:[::1]:4848
```
Parse the port as the last colon-segment (`split(':').pop()`) — not `split(':')[1]`, which breaks for IPv6 and for non-loopback IPv4 hosts added in this release.
### Custom port and host
`run_eval.py` does not expose `--port` or `--host` as CLI flags. Configure them in your mode YAML under the `environment:` key:
```yaml
# configs/modes/native_augment.yaml (or whichever mode you're running)
environment:
eval_server_port: 4849 # change if 4848 is already in use on the host
eval_server_host: "0.0.0.0" # bind all interfaces — needed for cross-container setups
```
Defaults are `port: 4848` and `host: 127.0.0.1` (loopback only). Use `0.0.0.0` only when the agent container needs to reach the eval-server from a separate network namespace. The health probe and tool scripts connect via the configured bind host (defaulting to `127.0.0.1`), which is reachable for both loopback and all-interface binds.
`"localhost"` is also a valid `eval_server_host` value. The OS resolves it at bind time — typically `127.0.0.1` on dual-stack or IPv4-only systems, and `::1` on IPv6-only systems. The exact result depends on your `/etc/hosts` and `gai.conf`. The READY signal will reflect the actual bound address (e.g. `GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848` or `GITNEXUS_EVAL_SERVER_READY:[::1]:4848`), not the literal string `localhost`. Use this when you want the server to bind to whichever loopback address the OS prefers rather than forcing IPv4.
**Running eval-server directly in Docker / Docker Compose:**
```bash
# Bind to all interfaces so sibling containers can reach it
gitnexus eval-server --host 0.0.0.0 --port 4848
# Then probe from a sibling container via its service hostname
curl http://eval-container:4848/health
```
If you need a non-default port (e.g. to avoid conflicts), pass `--port <port>` alongside `--host`. The READY signal will reflect both:
```
GITNEXUS_EVAL_SERVER_READY:0.0.0.0:5000
```
Parse the port as the last colon-segment (`split(':').pop()`) — safe for both IPv4 and bracketed IPv6 forms.
### Index caching
SWE-bench repos repeat (Django has 200+ instances at different commits). The harness caches GitNexus indexes per `(repo, commit)` hash in `~/.gitnexus-eval-cache/` to avoid redundant re-indexing.
+18 -5
View File
@@ -39,6 +39,7 @@ logger = logging.getLogger("gitnexus_docker")
DEFAULT_CACHE_DIR = Path.home() / ".gitnexus-eval-cache"
EVAL_SERVER_PORT = 4848
EVAL_SERVER_HOST = "127.0.0.1"
class GitNexusDockerEnvironment(DockerEnvironment):
@@ -62,6 +63,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
skip_embeddings: bool = True,
gitnexus_timeout: int = 120,
eval_server_port: int = EVAL_SERVER_PORT,
eval_server_host: str = EVAL_SERVER_HOST,
**kwargs,
):
super().__init__(**kwargs)
@@ -70,6 +72,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
self.skip_embeddings = skip_embeddings
self.gitnexus_timeout = gitnexus_timeout
self.eval_server_port = eval_server_port
self.eval_server_host = eval_server_host
self.index_time: float = 0.0
self._gitnexus_ready = False
@@ -165,22 +168,29 @@ class GitNexusDockerEnvironment(DockerEnvironment):
def _start_eval_server(self):
"""Start the GitNexus eval-server daemon in the background."""
logger.info(f"Starting eval-server on port {self.eval_server_port}...")
logger.info(
f"Starting eval-server on {self.eval_server_host}:{self.eval_server_port}..."
)
self.execute({
"command": (
f"nohup npx gitnexus eval-server --port {self.eval_server_port} "
f"--host {self.eval_server_host} "
f"--idle-timeout 600 "
f"> /tmp/gitnexus-eval-server.log 2>&1 &"
),
"timeout": 5,
})
# Use 127.0.0.1 for the health probe — reachable whether server binds
# loopback or all interfaces (0.0.0.0), avoiding DNS resolution issues.
health_host = "127.0.0.1"
# Wait for the server to be ready (up to ~15s for KuzuDB init)
for i in range(EVAL_SERVER_HEALTH_RETRIES):
time.sleep(EVAL_SERVER_HEALTH_INTERVAL_SECONDS)
health = self.execute({
"command": f"curl -sf http://127.0.0.1:{self.eval_server_port}/health 2>/dev/null || echo 'NOT_READY'",
"command": f"curl -sf http://{health_host}:{self.eval_server_port}/health 2>/dev/null || echo 'NOT_READY'",
"timeout": EVAL_SERVER_HEALTH_TIMEOUT_SECONDS,
})
output = health.get("output", "").strip()
@@ -201,7 +211,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
)
@staticmethod
def _render_tool_script(spec: ToolScriptSpec, port: str) -> str:
def _render_tool_script(spec: ToolScriptSpec, port: str, host: str = EVAL_SERVER_HOST) -> str:
"""
Render a standalone bash script for a GitNexus tool.
@@ -212,6 +222,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
if spec.endpoint:
lines.append(f'PORT="${{GITNEXUS_EVAL_PORT:-{port}}}"')
lines.append(f'HOST="${{GITNEXUS_EVAL_HOST:-{host}}}"')
if spec.header:
lines.append(spec.header.strip())
@@ -221,7 +232,7 @@ class GitNexusDockerEnvironment(DockerEnvironment):
if spec.endpoint:
lines.append(
f'result=$(curl -sf -X POST "http://127.0.0.1:${{PORT}}{spec.endpoint}" '
f'result=$(curl -sf -X POST "http://${{HOST}}:${{PORT}}{spec.endpoint}" '
'-H "Content-Type: application/json" -d "$payload" 2>/dev/null)'
)
lines.append('if [ $? -eq 0 ] && [ -n "$result" ]; then echo "$result"; exit 0; fi')
@@ -244,9 +255,10 @@ class GitNexusDockerEnvironment(DockerEnvironment):
Uses heredocs with quoted delimiter to avoid all quoting/escaping issues.
"""
port = str(self.eval_server_port)
host = self.eval_server_host
for spec in TOOL_SPECS.values():
script_content = self._render_tool_script(spec, port).strip()
script_content = self._render_tool_script(spec, port, host).strip()
# Use heredoc with quoted delimiter — prevents all variable expansion and quoting issues
self.execute({
"command": (
@@ -387,5 +399,6 @@ class GitNexusDockerEnvironment(DockerEnvironment):
"index_time_seconds": round(self.index_time, 2),
"skip_embeddings": self.skip_embeddings,
"eval_server_port": self.eval_server_port,
"eval_server_host": self.eval_server_host,
}
return base
Generated
+3 -3
View File
@@ -760,11 +760,11 @@ wheels = [
[[package]]
name = "idna"
version = "3.11"
version = "3.15"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/6f/6d/0703ccc57f3a7233505399edb88de3cbd678da106337b9fcde432b65ed60/idna-3.11.tar.gz", hash = "sha256:795dafcc9c04ed0c1fb032c2aa73654d8e8c5023a7df64a53f39190ada629902", size = 194582, upload-time = "2025-10-12T14:55:20.501Z" }
sdist = { url = "https://files.pythonhosted.org/packages/82/77/7b3966d0b9d1d31a36ddf1746926a11dface89a83409bf1483f0237aa758/idna-3.15.tar.gz", hash = "sha256:ca962446ea538f7092a95e057da437618e886f4d349216d2b1e294abfdb65fdc", size = 199245, upload-time = "2026-05-12T22:45:57.011Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/0e/61/66938bbb5fc52dbdf84594873d5b51fb1f7c7794e9c0f5bd885f30bc507b/idna-3.11-py3-none-any.whl", hash = "sha256:771a87f49d9defaf64091e6e6fe9c18d4833f140bd19464795bc32d966ca37ea", size = 71008, upload-time = "2025-10-12T14:55:18.883Z" },
{ url = "https://files.pythonhosted.org/packages/d2/23/408243171aa9aaba178d3e2559159c24c1171a641aa83b67bdd3394ead8e/idna-3.15-py3-none-any.whl", hash = "sha256:048adeaf8c2d788c40fee287673ccaa74c24ffd8dcf09ffa555a2fbb59f10ac8", size = 72340, upload-time = "2026-05-12T22:45:55.733Z" },
]
[[package]]
@@ -56,15 +56,15 @@ Generates repository documentation from the knowledge graph using an LLM. Requir
| Flag | Effect |
|------|--------|
| `--force` | Force full regeneration |
| `--force` | Force full regeneration, also required to re-gerenate an existing wiki in a different language |
| `--model <model>` | LLM model (default: minimax/minimax-m2.5) |
| `--base-url <url>` | LLM API base URL |
| `--api-key <key>` | LLM API key |
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
| `--gist` | Publish wiki as a public GitHub Gist |
| `--timeout <seconds>` | Per-attempt LLM request timeout in seconds (default: 60) |
| `--timeout <seconds>` | LLM request timeout in seconds (default: disabled) |
| `--retries <n>` | Max LLM retry attempts per request (default: 3) |
| `--lang <lang>` | Output language for generated documentation (e.g. english, chinese, spanish, japanese)|
### list — Show all indexed repos
```bash
+3 -1
View File
@@ -26,7 +26,7 @@ export type { PipelinePhase, PipelineProgress } from './pipeline.js';
// ─── Scope-based resolution — RFC #909 (Ring 1 #910) ────────────────────────
// Data model (RFC §2)
export type { SymbolDefinition } from './scope-resolution/symbol-definition.js';
export type { ParameterTypeClass, SymbolDefinition } from './scope-resolution/symbol-definition.js';
export type {
ScopeId,
DefId,
@@ -127,8 +127,10 @@ export { CLASS_KINDS, METHOD_KINDS, FIELD_KINDS } from './scope-resolution/regis
export type {
RegistryContext,
RegistryProviders,
OwnedMembersByOwnerLookup,
OwnerScopedContributor,
ArityVerdict,
ConstraintContext,
} from './scope-resolution/registries/context.js';
// Scope tree spine + position lookup (RFC §2.2 + §3.1; Ring 2 SHARED #912)
@@ -21,6 +21,7 @@
* (defined in `./types.ts`).
*/
import type { ParameterTypeClass } from './symbol-definition.js';
import type { Range, ScopeId } from './types.js';
/**
@@ -79,4 +80,11 @@ export interface ReferenceSite {
* (C#: `42` → `'int'`, `"alice"` → `'string'`).
*/
readonly argumentTypes?: readonly string[];
/**
* Optional per-argument type-shape sidecar for languages that need
* cv/ref/pointer distinctions during constraint filtering. This is
* intentionally separate from `argumentTypes`, which stays normalized
* for existing overload narrowing and conversion-rank logic.
*/
readonly argumentTypeClasses?: readonly ParameterTypeClass[];
}
@@ -13,7 +13,7 @@
*/
import type { NodeLabel } from '../../graph/types.js';
import type { SymbolDefinition } from '../symbol-definition.js';
import type { ParameterTypeClass, SymbolDefinition } from '../symbol-definition.js';
import type { Callsite, DefId } from '../types.js';
import type { DefIndex } from '../def-index.js';
import type { QualifiedNameIndex } from '../qualified-name-index.js';
@@ -30,10 +30,50 @@ export interface RegistryProviders {
* when absent, every candidate receives `'unknown'` (neutral signal).
*/
arityCompatibility?(callsite: Callsite, def: SymbolDefinition): ArityVerdict;
/**
* Language-specific constraint compatibility between a callsite and a
* candidate `def`. Mirrors `arityCompatibility` and shares its three-valued
* verdict shape; the third value `'unknown'` MUST keep the candidate
* (monotonicity: adding a predicate can only narrow correctly, never
* produce a wrong edge). Consulted by `narrowOverloadCandidates` after
* arity + type filters when a candidate carries `templateConstraints`.
*
* Optional; when absent the constraint filter is a pass-through. Languages
* with no constrained-overload semantics leave this undefined.
*/
constraintCompatibility?(
callsite: Callsite,
def: SymbolDefinition,
ctx: ConstraintContext,
): ArityVerdict;
}
export type ArityVerdict = 'compatible' | 'unknown' | 'incompatible';
/**
* Context threaded into `constraintCompatibility`. Kept minimal in the
* Tier-A scope (only `argumentTypes`, riding here until a separate
* `Callsite`-widening refactor moves them onto the call site directly).
* Future Tier-B graph-aware predicates (`is_base_of_v`, etc.) will widen
* this interface with `lookupTypeByName` and similar helpers.
*/
export interface ConstraintContext {
/**
* Per-slot argument types at the call site, normalized per the language
* adapter. Empty string means unknown. Same convention as
* `narrowOverloadCandidates`' `argTypes` parameter.
*/
readonly argumentTypes?: readonly string[];
/**
* Optional shape-preserving sidecar aligned with `argumentTypes`.
* Unknown or unsupported slots should be omitted by producers or
* marked with `indirection: 'unknown'`; consumers must preserve the
* monotonic fallback and return 'unknown' instead of guessing.
*/
readonly argumentTypeClasses?: readonly ParameterTypeClass[];
}
// ─── Owner-scoped contributor (concrete shape for `RegistryContributor`) ────
/**
@@ -60,6 +100,19 @@ export interface OwnerScopedContributor {
byName(name: string): readonly SymbolDefinition[];
}
/**
* Required owner-keyed lookup hook for Step 2 receiver/MRO member walks.
* Production callers wire this to the SemanticModel's authoritative
* method/field/nested-type registries so each `(ownerDefId, memberName)`
* probe is O(1). Implementations MUST return `[]` on an indexed miss —
* Step 2 treats `[]` as authoritative and does not consult `defs` for a
* fallback scan.
*/
export type OwnedMembersByOwnerLookup = (
ownerDefId: DefId,
memberName: string,
) => readonly SymbolDefinition[];
// ─── Top-level context threaded through every lookup ───────────────────────
export interface RegistryContext {
@@ -67,6 +120,7 @@ export interface RegistryContext {
readonly defs: DefIndex;
readonly qualifiedNames: QualifiedNameIndex;
readonly moduleScopes: ModuleScopeIndex;
readonly ownedMembersByOwner: OwnedMembersByOwnerLookup;
/**
* Method-dispatch index; required for method/field registries that
* honor `useReceiverTypeBinding`. Omit for class-only lookups.
@@ -27,8 +27,10 @@
* is true, resolve the receiver's type at `startScope` (from
* `scope.typeBindings`), then walk the MRO via
* `MethodDispatchIndex.mroFor(ownerDefId)`. Membership per owner comes
* through `RegistryContext.methodDispatch` + owner lookups into
* `scope.ownedDefs`; each hit records a raw signal with the owner's
* through an optional `RegistryContext.ownedMembersByOwner` hook when
* supplied (`undefined` → fall back to `defs.byId`; `[]` → indexed
* miss), otherwise via the compatibility fallback scan over
* `defs.byId`; each hit records a raw signal with the owner's
* MRO depth.
*
* **Step 3 — Owner-scoped contributor.** When
@@ -263,13 +265,14 @@ function walkReceiverTypeBinding(
// Walk the owner itself at depth 0, then its MRO chain.
const walk: DefId[] = [ownerDefId, ...ctx.methodDispatch.mroFor(ownerDefId)];
for (let mroDepth = 0; mroDepth < walk.length; mroDepth++) {
const currentOwnerId = walk[mroDepth]!;
let mroDepth = 0;
for (const currentOwnerId of walk) {
const members = collectOwnedMembers(currentOwnerId, name, ctx);
for (const def of members) {
if (!acceptedKinds.has(def.type)) continue;
recordTypeBindingHit(perCandidate, def, mroDepth, ownerDefId);
}
mroDepth++;
}
}
@@ -333,23 +336,7 @@ function collectOwnedMembers(
memberName: string,
ctx: RegistryContext,
): readonly SymbolDefinition[] {
// An owner's members are defs whose `ownerId === ownerDefId` and whose
// simple name matches `memberName`. We iterate `defs.byId` — O(D) per
// call today. A future by-owner index would make this O(K); tracked as
// a follow-up optimization before Ring 3 flips go production.
const out: SymbolDefinition[] = [];
for (const def of ctx.defs.byId.values()) {
if (def.ownerId !== ownerDefId) continue;
if (simpleNameOf(def) !== memberName) continue;
out.push(def);
}
return out;
}
function simpleNameOf(def: SymbolDefinition): string | undefined {
if (def.qualifiedName === undefined || def.qualifiedName.length === 0) return undefined;
const dot = def.qualifiedName.lastIndexOf('.');
return dot === -1 ? def.qualifiedName : def.qualifiedName.slice(dot + 1);
return ctx.ownedMembersByOwner(ownerDefId, memberName);
}
function recordTypeBindingHit(
@@ -11,6 +11,17 @@
import type { NodeLabel } from '../graph/types.js';
export interface ParameterTypeClass {
/** Normalized base type, matching the coarse `parameterTypes` vocabulary when known. */
base: string;
/** Top-level cv signal preserved from the original C++ parameter spelling. */
cv: 'none' | 'const' | 'volatile' | 'const volatile' | 'unknown';
/** Coarse value/reference/pointer shape. */
indirection: 'value' | 'lvalue-ref' | 'rvalue-ref' | 'pointer' | 'unknown';
/** Number of pointer markers when indirection is `pointer`; otherwise 0. */
pointerDepth: number;
}
export interface SymbolDefinition {
nodeId: string;
filePath: string;
@@ -26,12 +37,22 @@ export interface SymbolDefinition {
/** Per-parameter type names for overload disambiguation (e.g. ['int', 'String']).
* Populated when parameter types are resolvable from AST (any typed language). */
parameterTypes?: string[];
/** Additive per-parameter type shape sidecar for languages that need cv/ref/pointer distinctions.
* Does not participate in graph node identity unless a resolver explicitly opts in. */
parameterTypeClasses?: ParameterTypeClass[];
/** Raw return type text extracted from AST (e.g. 'User', 'Promise<User>') */
returnType?: string;
/** Declared type for non-callable symbols — fields/properties (e.g. 'Address', 'List<User>') */
declaredType?: string;
/** Generic/template specialization arguments for class-like symbols (e.g. ['User'], ['T*']). */
templateArguments?: string[];
/** Per-language constraint payload for template / generic overloads
* (e.g. C++ `enable_if_t<P, T>` predicate trees, C++20 `requires` clauses).
* Opaque to shared code — the producing language adapter owns the shape
* and is the only consumer. Read via the optional
* `ScopeResolver.constraintCompatibility` hook during overload narrowing.
* Absent for symbols that have no constraints (the common case). */
templateConstraints?: unknown;
/** Links Method/Constructor/Property to owning Class/Struct/Trait nodeId */
ownerId?: string;
}
@@ -0,0 +1,71 @@
import { test, expect, type Page } from '@playwright/test';
const BACKEND_URL = 'http://localhost:4747';
const REPO_NAME = 'mock-repo';
async function mockBackend(page: Page) {
const repo = {
name: REPO_NAME,
path: '/tmp/mock-repo',
repoPath: '/tmp/mock-repo',
indexedAt: new Date().toISOString(),
stats: { files: 1, nodes: 0, edges: 0, processes: 0 },
};
await page.route(
(url) => url.origin === BACKEND_URL && url.pathname === '/api/repos',
(route) => route.fulfill({ json: [repo] }),
);
await page.route(
(url) => url.origin === BACKEND_URL && url.pathname === '/api/repo',
(route) => route.fulfill({ json: repo }),
);
await page.route(
(url) => url.origin === BACKEND_URL && url.pathname === '/api/graph',
(route) => route.fulfill({ json: { nodes: [], relationships: [] } }),
);
await page.route(
(url) => url.origin === BACKEND_URL && url.pathname === '/api/heartbeat',
(route) =>
route.fulfill({
status: 200,
headers: { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-cache' },
body: ':ok\n\n',
}),
);
}
async function enterExploringView(page: Page) {
await page.goto('/');
await page.locator('[data-testid="landing-repo-card"]').first().click();
await expect(page.getByTestId('language-switcher')).toBeVisible({ timeout: 20_000 });
}
test.describe('language switching', () => {
test('switches Header language, updates document metadata, and persists after reload', async ({
page,
}) => {
await mockBackend(page);
await page.goto('/');
await page.evaluate(() => window.localStorage.clear());
await enterExploringView(page);
await page.getByTestId('language-switcher').selectOption('zh-CN');
await expect(page.locator('html')).toHaveAttribute('lang', 'zh-CN');
await expect(page.getByText('觉得不错就点星')).toBeVisible();
await expect
.poll(() => page.evaluate(() => window.localStorage.getItem('gitnexus.lng')))
.toBe('zh-CN');
await page.reload();
await expect(page.getByTestId('language-switcher')).toHaveValue('zh-CN', { timeout: 20_000 });
await expect(page.locator('html')).toHaveAttribute('lang', 'zh-CN');
await expect(page.getByText('觉得不错就点星')).toBeVisible();
await page.getByTestId('language-switcher').selectOption('en');
await expect(page.locator('html')).toHaveAttribute('lang', 'en');
});
});
+546 -403
View File
File diff suppressed because it is too large Load Diff
+22 -6
View File
@@ -18,7 +18,6 @@
"test:e2e:report": "playwright show-report"
},
"dependencies": {
"gitnexus-shared": "file:../gitnexus-shared",
"@langchain/anthropic": "^1.3.29",
"@langchain/core": "^1.1.44",
"@langchain/google-genai": "^2.1.30",
@@ -26,16 +25,19 @@
"@langchain/ollama": "^1.2.6",
"@langchain/openai": "^1.4.5",
"@sigma/edge-curve": "^3.1.0",
"@tailwindcss/vite": "^4.2.4",
"@tailwindcss/vite": "^4.3.0",
"axios": "^1.16.0",
"d3": "^7.9.0",
"dompurify": "^3.4.2",
"dompurify": "^3.4.3",
"gitnexus-shared": "file:../gitnexus-shared",
"graphology": "^0.26.0",
"graphology-indices": "^0.17.0",
"graphology-layout-force": "^0.2.4",
"graphology-layout-forceatlas2": "^0.10.1",
"graphology-layout-noverlap": "^0.4.2",
"graphology-utils": "^2.3.0",
"i18next": "^26.2.0",
"i18next-browser-languagedetector": "^8.2.1",
"langchain": "^1.3.5",
"lru-cache": "^11.2.4",
"lucide-react": "^1.14.0",
@@ -44,14 +46,15 @@
"pandemonium": "^2.4.0",
"react": "^19.2.5",
"react-dom": "^19.2.6",
"react-i18next": "^17.0.8",
"react-markdown": "^10.1.0",
"react-syntax-highlighter": "^16.1.0",
"react-syntax-highlighter": "^16.1.1",
"react-zoom-pan-pinch": "^4.0.3",
"remark-gfm": "^4.0.1",
"sigma": "^3.0.2",
"tailwindcss": "^4.2.4",
"uuid": "^14.0.0",
"zod": "^3.25.76"
"zod": "^4.4.3"
},
"devDependencies": {
"@babel/types": "^7.29.0",
@@ -64,7 +67,7 @@
"@types/react": "^19.2.14",
"@types/react-dom": "^19.2.3",
"@types/react-syntax-highlighter": "^15.5.13",
"@vercel/node": "^5.5.16",
"@vercel/node": "^5.8.2",
"@vitejs/plugin-react": "^5.1.4",
"@vitest/coverage-v8": "^4.1.5",
"jsdom": "^29.1.1",
@@ -73,5 +76,18 @@
"vite": "^8.0.11",
"vitest": "^4.1.5",
"wait-on": "^9.0.5"
},
"overrides": {
"@vercel/static-config": {
"ajv": "8.18.0"
},
"@vercel/node": {
"path-to-regexp": "6.3.0",
"undici": "6.24.0"
},
"@vercel/python-analysis": {
"minimatch": "10.2.3",
"smol-toml": "1.6.1"
}
}
}
+19 -12
View File
@@ -21,8 +21,11 @@ import {
type BackendRepo,
} from './services/backend-client';
import { ERROR_RESET_DELAY_MS } from './config/ui-constants';
import { formatBackendError } from './i18n/error-messages';
import { useTranslation } from 'react-i18next';
const AppContent = () => {
const { t } = useTranslation(['common', 'errors']);
const {
viewMode,
setViewMode,
@@ -54,7 +57,6 @@ const AppContent = () => {
async (result: ConnectResult): Promise<void> => {
// Use the canonical repo name from the server response so all subsequent
// backend calls (queries, search, grep, readFile) scope to this repo.
const repoName = result.repoInfo.name;
const repoPath = result.repoInfo.repoPath ?? result.repoInfo.path;
// Normalize both Windows (\) and Unix (/) path separators before splitting
const projectName =
@@ -104,6 +106,11 @@ const AppContent = () => {
// Auto-connect when ?server or ?project query param is present (bookmarkable shortcut)
const autoConnectRan = useRef(false);
const tRef = useRef(t);
useEffect(() => {
tRef.current = t;
}, [t]);
useEffect(() => {
if (autoConnectRan.current) return;
const params = new URLSearchParams(window.location.search);
@@ -116,8 +123,8 @@ const AppContent = () => {
setProgress({
phase: 'extracting',
percent: 0,
message: 'Connecting to server...',
detail: 'Validating server',
message: tRef.current('common:progress.connecting'),
detail: tRef.current('common:progress.validatingServer'),
});
setViewMode('loading');
@@ -132,8 +139,8 @@ const AppContent = () => {
setProgress({
phase: 'extracting',
percent: 5,
message: 'Connecting to server...',
detail: 'Validating server',
message: tRef.current('common:progress.connecting'),
detail: tRef.current('common:progress.validatingServer'),
});
} else if (phase === 'downloading') {
const pct = total ? Math.round((downloaded / total) * 90) + 5 : 50;
@@ -141,15 +148,15 @@ const AppContent = () => {
setProgress({
phase: 'extracting',
percent: pct,
message: 'Downloading graph...',
detail: `${mb} MB downloaded`,
message: tRef.current('common:progress.downloadingGraph'),
detail: tRef.current('common:progress.downloadedMb', { mb }),
});
} else if (phase === 'extracting') {
setProgress({
phase: 'extracting',
percent: 97,
message: 'Processing...',
detail: 'Extracting file contents',
message: tRef.current('common:progress.processing'),
detail: tRef.current('common:progress.extractingFileContents'),
});
}
},
@@ -173,8 +180,8 @@ const AppContent = () => {
setProgress({
phase: 'error',
percent: 0,
message: 'Failed to connect to server',
detail: err instanceof Error ? err.message : 'Unknown error',
message: tRef.current('errors:connectFailed'),
detail: formatBackendError(err, tRef.current),
});
setTimeout(() => {
setViewMode('onboarding');
@@ -299,7 +306,7 @@ const AppContent = () => {
{serverDisconnected && (
<div className="fixed bottom-12 left-1/2 z-50 -translate-x-1/2 rounded-lg border border-yellow-500/30 bg-yellow-900/80 px-4 py-2 text-sm text-yellow-200 shadow-lg backdrop-blur">
Server connection lost — reconnecting&hellip;
{t('errors:backend.reconnecting')}
</div>
)}
@@ -17,6 +17,7 @@
import { Sparkles, Github } from '@/lib/lucide-icons';
import { RepoAnalyzer } from './RepoAnalyzer';
import { useTranslation } from 'react-i18next';
interface AnalyzeOnboardingProps {
/** Called when analysis finishes and the repo is ready to load. */
@@ -24,6 +25,8 @@ interface AnalyzeOnboardingProps {
}
export const AnalyzeOnboarding = ({ onComplete }: AnalyzeOnboardingProps) => {
const { t } = useTranslation('onboarding');
return (
<div className="relative animate-fade-in overflow-hidden rounded-3xl border border-border-default bg-surface p-7">
{/* Ambient glows — mirrors OnboardingGuide aesthetic */}
@@ -47,11 +50,10 @@ export const AnalyzeOnboarding = ({ onComplete }: AnalyzeOnboardingProps) => {
</div>
<h2 className="text-lg leading-snug font-semibold text-text-primary">
Analyze your first repository
{t('analyzeFirst.title')}
</h2>
<p className="mx-auto mt-1.5 max-w-xs text-sm leading-relaxed text-text-secondary">
Paste a GitHub URL and GitNexus will clone it, parse the code, and build a live
knowledge graph — right in your browser.
{t('analyzeFirst.description')}
</p>
</div>
</div>
@@ -63,7 +65,7 @@ export const AnalyzeOnboarding = ({ onComplete }: AnalyzeOnboardingProps) => {
{/* Footer hint */}
<p className="mt-5 text-center text-[11px] leading-relaxed text-text-muted">
Public repos only &middot; Cloned locally by the server &middot; No data leaves your machine
{t('analyzeFirst.footer')}
</p>
</div>
);
@@ -1,33 +1,16 @@
import { useState, useEffect } from 'react';
import { X } from '@/lib/lucide-icons';
import type { JobProgress as AnalyzeJobProgress } from '../services/backend-client';
import { useTranslation } from 'react-i18next';
import { translateAnalyzePhase } from '../i18n/progress';
interface AnalyzeProgressProps {
progress: AnalyzeJobProgress;
onCancel: () => void;
}
const PHASE_LABELS: Record<string, string> = {
queued: 'Queued',
cloning: 'Cloning repository',
pulling: 'Pulling latest',
extracting: 'Scanning files',
structure: 'Building structure',
parsing: 'Parsing code',
imports: 'Resolving imports',
calls: 'Tracing calls',
heritage: 'Extracting inheritance',
communities: 'Detecting communities',
processes: 'Detecting processes',
complete: 'Pipeline complete',
lbug: 'Loading into database',
fts: 'Creating search indexes',
embeddings: 'Generating embeddings',
done: 'Done',
retrying: 'Retrying after crash',
};
export const AnalyzeProgress = ({ progress, onCancel }: AnalyzeProgressProps) => {
const { t } = useTranslation('common');
const [startTime] = useState(() => Date.now());
const [elapsed, setElapsed] = useState(0);
@@ -38,11 +21,11 @@ export const AnalyzeProgress = ({ progress, onCancel }: AnalyzeProgressProps) =>
const formatElapsed = (ms: number) => {
const s = Math.floor(ms / 1000);
if (s < 60) return `${s}s`;
return `${Math.floor(s / 60)}m ${s % 60}s`;
if (s < 60) return t('units.elapsedSeconds', { seconds: s });
return t('units.elapsedMinutesSeconds', { minutes: Math.floor(s / 60), seconds: s % 60 });
};
const label = PHASE_LABELS[progress.phase] || progress.message || progress.phase;
const label = translateAnalyzePhase(progress.phase, progress.message, t);
const pct = Math.max(0, Math.min(100, progress.percent));
return (
@@ -69,7 +52,7 @@ export const AnalyzeProgress = ({ progress, onCancel }: AnalyzeProgressProps) =>
className="flex items-center gap-1.5 rounded-lg bg-red-500/10 px-3 py-1.5 text-xs text-red-400 transition-all duration-200 hover:bg-red-500/20"
>
<X className="h-3.5 w-3.5" />
Cancel
{t('actions.cancel')}
</button>
</div>
</div>
@@ -17,6 +17,7 @@ import { useAppState } from '../hooks/useAppState';
import { type GraphNode, getSyntaxLanguageFromFilename } from 'gitnexus-shared';
import { NODE_COLORS } from '../lib/constants';
import { readFile, type ReadFileResult } from '../services/backend-client';
import { useTranslation } from 'react-i18next';
const getSyntaxLanguage = (filePath: string | undefined): string => {
if (!filePath) return 'text';
@@ -46,6 +47,7 @@ export interface CodeReferencesPanelProps {
}
export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) => {
const { t } = useTranslation(['common', 'graph']);
const {
graph,
selectedNode,
@@ -294,14 +296,14 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
<button
onClick={() => setIsCollapsed(false)}
className="rounded p-2 text-text-secondary transition-colors hover:bg-cyan-500/10 hover:text-cyan-400"
title="Expand Code Panel"
title={t('graph:codePanel.expand')}
>
<PanelLeft className="h-5 w-5" />
</button>
<div className="my-1 h-px w-6 bg-border-subtle" />
{showSelectedViewer && (
<div className="rotate-90 text-[9px] font-medium tracking-wide whitespace-nowrap text-amber-400">
SELECTED
{t('graph:codePanel.selected')}
</div>
)}
{showCitations && (
@@ -325,20 +327,22 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
<div
onMouseDown={startResize}
className="absolute top-0 right-0 h-full w-2 cursor-col-resize bg-transparent transition-colors hover:bg-cyan-500/25"
title="Drag to resize"
title={t('graph:codePanel.dragResize')}
/>
{/* Header */}
<div className="flex items-center justify-between border-b border-border-subtle bg-gradient-to-r from-elevated/60 to-surface/60 px-3 py-2.5">
<div className="flex items-center gap-2">
<Code className="h-4 w-4 text-cyan-400" />
<span className="text-sm font-semibold text-text-primary">Code Inspector</span>
<span className="text-sm font-semibold text-text-primary">
{t('graph:codePanel.title')}
</span>
</div>
<div className="flex items-center gap-1.5">
{showCitations && (
<button
onClick={() => clearCodeReferences()}
className="rounded p-1.5 text-text-muted transition-colors hover:bg-red-500/10 hover:text-red-400"
title="Clear AI citations"
title={t('graph:codePanel.clearCitations')}
>
<Trash2 className="h-4 w-4" />
</button>
@@ -346,7 +350,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
<button
onClick={() => setIsCollapsed(true)}
className="rounded p-1.5 text-text-muted transition-colors hover:bg-hover hover:text-text-primary"
title="Collapse Panel"
title={t('common:actions.collapse')}
>
<PanelLeftClose className="h-4 w-4" />
</button>
@@ -361,7 +365,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
<div className="flex items-center gap-1.5 rounded-md border border-amber-500/25 bg-amber-500/15 px-2 py-0.5">
<MousePointerClick className="h-3 w-3 text-amber-400" />
<span className="text-[10px] font-semibold tracking-wide text-amber-300 uppercase">
Selected
{t('graph:codePanel.selected')}
</span>
</div>
<FileCode className="ml-1 h-3.5 w-3.5 text-amber-400/70" />
@@ -372,7 +376,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
<button
onClick={() => setSelectedNode(null)}
className="rounded p-1 text-text-muted transition-colors hover:bg-amber-500/10 hover:text-amber-400"
title="Clear selection"
title={t('graph:codePanel.clearSelection')}
>
<X className="h-4 w-4" />
</button>
@@ -381,7 +385,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
{isLoadingFile ? (
<div className="flex items-center justify-center gap-2 py-8 text-text-muted">
<Loader2 className="h-4 w-4 animate-spin" />
<span className="text-sm">Loading source...</span>
<span className="text-sm">{t('graph:codePanel.loadingSource')}</span>
</div>
) : selectedFileContent ? (
<SyntaxHighlighter
@@ -420,12 +424,9 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
) : (
<div className="px-3 py-3 text-sm text-text-muted">
{selectedIsFile ? (
<>
Code not available in memory for{' '}
<span className="font-mono">{selectedFilePath}</span>
</>
<>{t('graph:codePanel.codeNotAvailable', { path: selectedFilePath })}</>
) : (
<>Select a file node to preview its contents.</>
<>{t('graph:codePanel.selectFile')}</>
)}
</div>
)}
@@ -446,11 +447,11 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
<div className="flex items-center gap-1.5 rounded-md border border-cyan-500/25 bg-cyan-500/15 px-2 py-0.5">
<Sparkles className="h-3 w-3 text-cyan-400" />
<span className="text-[10px] font-semibold tracking-wide text-cyan-300 uppercase">
AI Citations
{t('graph:codePanel.aiCitations')}
</span>
</div>
<span className="ml-1 text-xs text-text-muted">
{aiReferences.length} reference{aiReferences.length !== 1 ? 's' : ''}
{t('graph:codePanel.references', { count: aiReferences.length })}
</span>
</div>
<div className="scrollbar-thin min-h-0 flex-1 space-y-3 overflow-y-auto p-3">
@@ -483,9 +484,9 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
<span
className="mt-0.5 flex-shrink-0 rounded px-2 py-0.5 text-[10px] font-semibold tracking-wide uppercase"
style={{ backgroundColor: nodeColor, color: '#06060a' }}
title={ref.label ?? 'Code'}
title={ref.label ?? t('graph:codePanel.code')}
>
{ref.label ?? 'Code'}
{ref.label ?? t('graph:codePanel.code')}
</span>
<div className="min-w-0 flex-1">
<div className="truncate text-xs font-medium text-text-primary">
@@ -501,7 +502,10 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
</span>
)}
{totalLines > 0 && (
<span className="text-text-muted"> • {totalLines} lines</span>
<span className="text-text-muted">
{' '}
• {t('graph:codePanel.lines', { count: totalLines })}
</span>
)}
</div>
</div>
@@ -518,7 +522,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
onFocusNode(nodeId);
}}
className="rounded p-1.5 text-text-muted transition-colors hover:bg-hover hover:text-text-primary"
title="Focus in graph"
title={t('common:actions.focusInGraph')}
>
<Target className="h-4 w-4" />
</button>
@@ -526,7 +530,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
<button
onClick={() => removeCodeReference(ref.id)}
className="rounded p-1.5 text-text-muted transition-colors hover:bg-hover hover:text-text-primary"
title="Remove"
title={t('common:actions.remove')}
>
<X className="h-4 w-4" />
</button>
@@ -572,8 +576,7 @@ export const CodeReferencesPanel = ({ onFocusNode }: CodeReferencesPanelProps) =
</SyntaxHighlighter>
) : (
<div className="px-3 py-3 text-sm text-text-muted">
Code not available in memory for{' '}
<span className="font-mono">{ref.filePath}</span>
{t('graph:codePanel.codeNotAvailable', { path: ref.filePath })}
</div>
)}
</div>
+22 -12
View File
@@ -10,6 +10,8 @@ import { useBackend } from '../hooks/useBackend';
import { OnboardingGuide } from './OnboardingGuide';
import { AnalyzeOnboarding } from './AnalyzeOnboarding';
import { RepoLanding } from './RepoLanding';
import { useTranslation } from 'react-i18next';
import { formatBackendError } from '../i18n/error-messages';
interface DropZoneProps {
onServerConnect?: (result: ConnectResult, serverUrl?: string) => void | Promise<void>;
@@ -60,6 +62,8 @@ function Crossfade({ activeKey, children }: { activeKey: string; children: React
// ── Phase cards ─────────────────────────────────────────────────────────────
function SuccessCard() {
const { t } = useTranslation('onboarding');
return (
<div
className="relative overflow-hidden rounded-3xl border border-emerald-500/20 bg-surface p-7"
@@ -76,10 +80,10 @@ function SuccessCard() {
</div>
<h2 className="mb-2 text-center text-lg font-semibold text-emerald-400">
Server Connected
{t('success.title')}
</h2>
<p className="text-center text-sm leading-relaxed text-text-secondary">
Preparing your code knowledge graph...
{t('success.description')}
</p>
{/* Subtle progress hint */}
@@ -100,6 +104,8 @@ function SuccessCard() {
}
function LoadingCard({ message }: { message: string }) {
const { t } = useTranslation(['common', 'onboarding']);
return (
<div
className="relative overflow-hidden rounded-3xl border border-accent/20 bg-surface p-7"
@@ -116,10 +122,10 @@ function LoadingCard({ message }: { message: string }) {
</div>
<h2 className="mb-2 text-center text-lg font-semibold text-text-primary">
{message || 'Connecting...'}
{message || t('common:progress.connectingShort')}
</h2>
<p className="text-center text-sm leading-relaxed text-text-secondary">
This may take a moment for large repositories
{t('onboarding:loading.largeRepoHint')}
</p>
{/* Decorative sparkle */}
@@ -134,6 +140,7 @@ function LoadingCard({ message }: { message: string }) {
// ── DropZone ─────────────────────────────────────────────────────────────────
export const DropZone = ({ onServerConnect }: DropZoneProps) => {
const { t } = useTranslation(['common', 'errors']);
const [error, setError] = useState<string | null>(null);
// Backend polling for server detection
@@ -163,7 +170,7 @@ export const DropZone = ({ onServerConnect }: DropZoneProps) => {
// appropriate screen (landing with repo cards, or analyze for zero repos).
const handleAutoConnect = async () => {
setPhase('loading');
setLoadingMessage('Connecting...');
setLoadingMessage(t('common:progress.connectingShort'));
setError(null);
try {
@@ -179,8 +186,7 @@ export const DropZone = ({ onServerConnect }: DropZoneProps) => {
setPhase('landing');
} catch (err) {
if ((err as Error).name === 'AbortError') return;
const message = err instanceof Error ? err.message : 'Failed to connect';
setError(message);
setError(formatBackendError(err, t));
setPhase('onboarding');
}
};
@@ -193,7 +199,7 @@ export const DropZone = ({ onServerConnect }: DropZoneProps) => {
const connectToRepo = (repoName: string) => {
autoConnectRan.current = true;
setPhase('loading');
setLoadingMessage('Loading graph...');
setLoadingMessage(t('common:progress.loadingGraph'));
setError(null);
(async () => {
@@ -204,13 +210,17 @@ export const DropZone = ({ onServerConnect }: DropZoneProps) => {
detectedBackendUrl,
(p, downloaded, total) => {
if (p === 'validating') {
setLoadingMessage('Validating server...');
setLoadingMessage(t('common:progress.validatingServerEllipsis'));
} else if (p === 'downloading') {
const pct = total ? Math.round((downloaded / total) * 100) : null;
const mb = (downloaded / (1024 * 1024)).toFixed(1);
setLoadingMessage(pct ? `Downloading graph... ${pct}%` : `Downloading... ${mb} MB`);
setLoadingMessage(
pct
? t('common:progress.downloadingWithPercent', { percent: pct })
: t('common:progress.downloadingMb', { mb }),
);
} else if (p === 'extracting') {
setLoadingMessage('Processing graph...');
setLoadingMessage(t('common:progress.processingGraph'));
}
},
abortController.signal,
@@ -221,7 +231,7 @@ export const DropZone = ({ onServerConnect }: DropZoneProps) => {
}
} catch (err) {
if ((err as Error).name === 'AbortError') return;
setError(err instanceof Error ? err.message : 'Failed to load graph');
setError(formatBackendError(err, t));
setPhase(detectedRepos.length > 0 ? 'landing' : 'analyze');
} finally {
abortControllerRef.current = null;
@@ -2,12 +2,14 @@ import { Brain, Loader2, Check, AlertCircle, Zap } from '@/lib/lucide-icons';
import { useAppState } from '../hooks/useAppState';
import { useState } from 'react';
import { WebGPUFallbackDialog } from './WebGPUFallbackDialog';
import { useTranslation } from 'react-i18next';
/**
* Embedding status indicator and trigger button
* Shows in header when graph is loaded
*/
export const EmbeddingStatus = () => {
const { t } = useTranslation('graph');
const { embeddingStatus, embeddingProgress, startEmbeddings, graph, viewMode, serverBaseUrl } =
useAppState();
@@ -63,10 +65,10 @@ export const EmbeddingStatus = () => {
<button
onClick={() => handleStartEmbeddings()}
className="group flex items-center gap-2 rounded-lg border border-border-subtle bg-surface px-3 py-1.5 text-sm text-text-secondary transition-all hover:border-accent/50 hover:bg-hover hover:text-text-primary"
title="Generate embeddings for semantic search"
title={t('embedding.generateTitle')}
>
<Brain className="h-4 w-4 text-node-interface transition-colors group-hover:text-accent" />
<span className="hidden sm:inline">Enable Semantic Search</span>
<span className="hidden sm:inline">{t('embedding.enable')}</span>
<Zap className="h-3 w-3 text-text-muted" />
</button>
</div>
@@ -83,7 +85,7 @@ export const EmbeddingStatus = () => {
<div className="flex items-center gap-2.5 rounded-lg border border-accent/30 bg-surface px-3 py-1.5 text-sm">
<Loader2 className="h-4 w-4 animate-spin text-accent" />
<div className="flex flex-col gap-0.5">
<span className="text-xs text-text-secondary">Loading AI model...</span>
<span className="text-xs text-text-secondary">{t('embedding.loadingModel')}</span>
<div className="h-1 w-24 overflow-hidden rounded-full bg-elevated">
<div
className="h-full rounded-full bg-gradient-to-r from-accent to-node-interface transition-all duration-300"
@@ -108,7 +110,7 @@ export const EmbeddingStatus = () => {
<Loader2 className="h-4 w-4 animate-spin text-node-function" />
<div className="flex flex-col gap-0.5">
<span className="text-xs text-text-secondary">
Embedding {processed}/{total} nodes
{t('embedding.embeddingNodes', { processed, total })}
</span>
<div className="h-1 w-24 overflow-hidden rounded-full bg-elevated">
<div
@@ -126,7 +128,7 @@ export const EmbeddingStatus = () => {
return (
<div className="flex items-center gap-2 rounded-lg border border-node-interface/30 bg-surface px-3 py-1.5 text-sm text-text-secondary">
<Loader2 className="h-4 w-4 animate-spin text-node-interface" />
<span className="text-xs">Creating vector index...</span>
<span className="text-xs">{t('embedding.creatingIndex')}</span>
</div>
);
}
@@ -136,10 +138,10 @@ export const EmbeddingStatus = () => {
return (
<div
className="flex items-center gap-2 rounded-lg border border-node-function/30 bg-node-function/10 px-3 py-1.5 text-sm text-node-function"
title="Semantic search is ready! Use natural language in the AI chat."
title={t('embedding.readyTitle')}
>
<Check className="h-4 w-4" />
<span className="text-xs font-medium">Semantic Ready</span>
<span className="text-xs font-medium">{t('embedding.ready')}</span>
</div>
);
}
@@ -151,10 +153,10 @@ export const EmbeddingStatus = () => {
<button
onClick={() => handleStartEmbeddings()}
className="flex items-center gap-2 rounded-lg border border-red-500/30 bg-red-500/10 px-3 py-1.5 text-sm text-red-400 transition-colors hover:bg-red-500/20"
title="Embedding failed. Click to retry."
title={t('embedding.errorTitle')}
>
<AlertCircle className="h-4 w-4" />
<span className="text-xs">Failed - Retry</span>
<span className="text-xs">{t('embedding.failedRetry')}</span>
</button>
{fallbackDialog}
</>
+29 -29
View File
@@ -19,6 +19,7 @@ import {
Type,
} from '@/lib/lucide-icons';
import { useAppState } from '../hooks/useAppState';
import { useTranslation } from 'react-i18next';
import { FILTERABLE_LABELS, NODE_COLORS, ALL_EDGE_TYPES, EDGE_INFO } from '../lib/constants';
import type { GraphNode, NodeLabel } from 'gitnexus-shared';
@@ -211,6 +212,7 @@ interface FileTreePanelProps {
}
export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
const { t } = useTranslation(['common', 'graph']);
const {
graph,
visibleLabels,
@@ -303,7 +305,7 @@ export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
<button
onClick={() => setIsCollapsed(false)}
className="rounded p-2 text-text-secondary transition-colors hover:bg-hover hover:text-text-primary"
title="Expand Panel"
title={t('graph:fileTree.expandPanel')}
>
<PanelLeft className="h-5 w-5" />
</button>
@@ -314,7 +316,7 @@ export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
setActiveTab('files');
}}
className={`rounded p-2 transition-colors ${activeTab === 'files' ? 'bg-accent/10 text-accent' : 'text-text-secondary hover:bg-hover hover:text-text-primary'}`}
title="File Explorer"
title={t('graph:fileTree.fileExplorer')}
>
<Folder className="h-5 w-5" />
</button>
@@ -324,7 +326,7 @@ export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
setActiveTab('filters');
}}
className={`rounded p-2 transition-colors ${activeTab === 'filters' ? 'bg-accent/10 text-accent' : 'text-text-secondary hover:bg-hover hover:text-text-primary'}`}
title="Filters"
title={t('graph:fileTree.filters')}
>
<Filter className="h-5 w-5" />
</button>
@@ -345,7 +347,7 @@ export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
: 'text-text-secondary hover:bg-hover hover:text-text-primary'
}`}
>
Explorer
{t('graph:fileTree.explorer')}
</button>
<button
onClick={() => setActiveTab('filters')}
@@ -355,13 +357,13 @@ export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
: 'text-text-secondary hover:bg-hover hover:text-text-primary'
}`}
>
Filters
{t('graph:fileTree.filters')}
</button>
</div>
<button
onClick={() => setIsCollapsed(true)}
className="rounded p-1 text-text-muted transition-colors hover:bg-hover hover:text-text-primary"
title="Collapse Panel"
title={t('graph:fileTree.collapsePanel')}
>
<PanelLeftClose className="h-4 w-4" />
</button>
@@ -375,7 +377,7 @@ export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
<Search className="absolute top-1/2 left-2.5 h-3.5 w-3.5 -translate-y-1/2 text-text-muted" />
<input
type="text"
placeholder="Search files..."
placeholder={t('graph:fileTree.searchFiles')}
value={searchQuery}
onChange={(e) => setSearchQuery(e.target.value)}
className="w-full rounded border border-border-subtle bg-elevated py-1.5 pr-3 pl-8 text-xs text-text-primary placeholder:text-text-muted focus:border-accent focus:outline-none"
@@ -386,7 +388,9 @@ export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
{/* File tree */}
<div className="scrollbar-thin flex-1 overflow-y-auto py-2">
{fileTree.length === 0 ? (
<div className="px-3 py-4 text-center text-xs text-text-muted">No files loaded</div>
<div className="px-3 py-4 text-center text-xs text-text-muted">
{t('graph:fileTree.noFilesLoaded')}
</div>
) : (
fileTree.map((node) => (
<TreeItem
@@ -409,11 +413,9 @@ export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
<div className="scrollbar-thin flex-1 overflow-y-auto p-3">
<div className="mb-3">
<h3 className="mb-2 text-xs font-medium tracking-wide text-text-secondary uppercase">
Node Types
{t('graph:fileTree.nodeTypes')}
</h3>
<p className="mb-3 text-[11px] text-text-muted">
Toggle visibility of node types in the graph
</p>
<p className="mb-3 text-[11px] text-text-muted">{t('graph:fileTree.nodeTypesDesc')}</p>
</div>
<div className="flex flex-col gap-1">
@@ -449,11 +451,9 @@ export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
{/* Edge Type Toggles */}
<div className="mt-6 border-t border-border-subtle pt-4">
<h3 className="mb-2 text-xs font-medium tracking-wide text-text-secondary uppercase">
Edge Types
{t('graph:fileTree.edgeTypes')}
</h3>
<p className="mb-3 text-[11px] text-text-muted">
Toggle visibility of relationship types
</p>
<p className="mb-3 text-[11px] text-text-muted">{t('graph:fileTree.edgeTypesDesc')}</p>
<div className="flex flex-col gap-1">
{ALL_EDGE_TYPES.map((edgeType) => {
@@ -488,19 +488,17 @@ export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
<div className="mt-6 border-t border-border-subtle pt-4">
<h3 className="mb-2 text-xs font-medium tracking-wide text-text-secondary uppercase">
<Target className="mr-1.5 inline h-3 w-3" />
Focus Depth
{t('graph:fileTree.focusDepth')}
</h3>
<p className="mb-3 text-[11px] text-text-muted">
Show nodes within N hops of selection
</p>
<p className="mb-3 text-[11px] text-text-muted">{t('graph:fileTree.focusDepthDesc')}</p>
<div className="flex flex-wrap gap-1.5">
{[
{ value: null, label: 'All' },
{ value: 1, label: '1 hop' },
{ value: 2, label: '2 hops' },
{ value: 3, label: '3 hops' },
{ value: 5, label: '5 hops' },
{ value: null, label: t('graph:fileTree.all') },
{ value: 1, label: t('graph:fileTree.hops', { count: 1 }) },
{ value: 2, label: t('graph:fileTree.hops', { count: 2 }) },
{ value: 3, label: t('graph:fileTree.hops', { count: 3 }) },
{ value: 5, label: t('graph:fileTree.hops', { count: 5 }) },
].map(({ value, label }) => (
<button
key={label}
@@ -517,14 +515,16 @@ export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
</div>
{depthFilter !== null && !selectedNode && (
<p className="mt-2 text-[10px] text-amber-400">Select a node to apply depth filter</p>
<p className="mt-2 text-[10px] text-amber-400">
{t('graph:fileTree.selectNodeDepth')}
</p>
)}
</div>
{/* Legend */}
<div className="mt-6 border-t border-border-subtle pt-4">
<h3 className="mb-3 text-xs font-medium tracking-wide text-text-secondary uppercase">
Color Legend
{t('graph:fileTree.colorLegend')}
</h3>
<div className="grid grid-cols-2 gap-2">
{(
@@ -558,8 +558,8 @@ export const FileTreePanel = ({ onFocusNode }: FileTreePanelProps) => {
{graph && (
<div className="border-t border-border-subtle bg-elevated/50 px-3 py-2">
<div className="flex items-center justify-between text-[10px] text-text-muted">
<span>{graph.nodes.length} nodes</span>
<span>{graph.relationships.length} edges</span>
<span>{t('common:counts.nodes', { count: graph.nodes.length })}</span>
<span>{t('common:counts.edges', { count: graph.relationships.length })}</span>
</div>
</div>
)}
+15 -9
View File
@@ -21,12 +21,14 @@ import {
import type { GraphNode } from 'gitnexus-shared';
import { QueryFAB } from './QueryFAB';
import Graph from 'graphology';
import { useTranslation } from 'react-i18next';
export interface GraphCanvasHandle {
focusNode: (nodeId: string) => void;
}
export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
const { t } = useTranslation('graph');
const {
graph,
setSelectedNode,
@@ -268,7 +270,7 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
onClick={handleClearSelection}
className="ml-2 rounded px-2 py-0.5 text-xs text-text-secondary transition-colors hover:bg-white/10 hover:text-text-primary"
>
Clear
{t('canvas.clear')}
</button>
</div>
)}
@@ -278,21 +280,21 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
<button
onClick={zoomIn}
className="flex h-9 w-9 items-center justify-center rounded-md border border-border-subtle bg-elevated text-text-secondary transition-colors hover:bg-hover hover:text-text-primary"
title="Zoom In"
title={t('canvas.zoomIn')}
>
<ZoomIn className="h-4 w-4" />
</button>
<button
onClick={zoomOut}
className="flex h-9 w-9 items-center justify-center rounded-md border border-border-subtle bg-elevated text-text-secondary transition-colors hover:bg-hover hover:text-text-primary"
title="Zoom Out"
title={t('canvas.zoomOut')}
>
<ZoomOut className="h-4 w-4" />
</button>
<button
onClick={resetZoom}
className="flex h-9 w-9 items-center justify-center rounded-md border border-border-subtle bg-elevated text-text-secondary transition-colors hover:bg-hover hover:text-text-primary"
title="Fit to Screen"
title={t('canvas.fit')}
>
<Maximize2 className="h-4 w-4" />
</button>
@@ -305,7 +307,7 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
<button
onClick={handleFocusSelected}
className="flex h-9 w-9 items-center justify-center rounded-md border border-accent/30 bg-accent/20 text-accent transition-colors hover:bg-accent/30"
title="Focus on Selected Node"
title={t('canvas.focusSelected')}
>
<Focus className="h-4 w-4" />
</button>
@@ -316,7 +318,7 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
<button
onClick={handleClearSelection}
className="flex h-9 w-9 items-center justify-center rounded-md border border-border-subtle bg-elevated text-text-secondary transition-colors hover:bg-hover hover:text-text-primary"
title="Clear Selection"
title={t('canvas.clearSelection')}
>
<RotateCcw className="h-4 w-4" />
</button>
@@ -333,7 +335,7 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
? 'animate-pulse border-accent bg-accent text-white shadow-glow'
: 'border-border-subtle bg-elevated text-text-secondary hover:bg-hover hover:text-text-primary'
} `}
title={isLayoutRunning ? 'Stop Layout' : 'Run Layout Again'}
title={isLayoutRunning ? t('canvas.stopLayout') : t('canvas.runLayout')}
>
{isLayoutRunning ? <Pause className="h-4 w-4" /> : <Play className="h-4 w-4" />}
</button>
@@ -343,7 +345,9 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
{isLayoutRunning && (
<div className="absolute bottom-4 left-1/2 z-10 flex -translate-x-1/2 animate-fade-in items-center gap-2 rounded-full border border-emerald-500/30 bg-emerald-500/20 px-3 py-1.5 backdrop-blur-sm">
<div className="h-2 w-2 animate-ping rounded-full bg-emerald-400" />
<span className="text-xs font-medium text-emerald-400">Layout optimizing...</span>
<span className="text-xs font-medium text-emerald-400">
{t('canvas.layoutOptimizing')}
</span>
</div>
)}
@@ -359,7 +363,9 @@ export const GraphCanvas = forwardRef<GraphCanvasHandle>((_, ref) => {
? 'flex h-10 w-10 items-center justify-center rounded-lg border border-cyan-400/40 bg-cyan-500/15 text-cyan-200 transition-colors hover:border-cyan-300/60 hover:bg-cyan-500/20'
: 'flex h-10 w-10 items-center justify-center rounded-lg border border-border-subtle bg-elevated text-text-muted transition-colors hover:bg-hover hover:text-text-primary'
}
title={isAIHighlightsEnabled ? 'Turn off all highlights' : 'Turn on AI highlights'}
title={
isAIHighlightsEnabled ? t('canvas.turnOffHighlights') : t('canvas.turnOnHighlights')
}
data-testid="ai-highlights-toggle"
>
{isAIHighlightsEnabled ? (
+25 -16
View File
@@ -21,9 +21,12 @@ import {
type JobProgress,
} from '../services/backend-client';
import { useState, useMemo, useRef, useEffect } from 'react';
import { useTranslation } from 'react-i18next';
import { GraphNode } from 'gitnexus-shared';
import { EmbeddingStatus } from './EmbeddingStatus';
import { RepoAnalyzer } from './RepoAnalyzer';
import { LanguageSwitcher } from './LanguageSwitcher';
import { translateProgressMessage } from '../i18n/progress';
// Color mapping for node types in search results
const NODE_TYPE_COLORS: Record<string, string> = {
@@ -55,6 +58,7 @@ export const Header = ({
onAnalyzeComplete,
onReposChanged,
}: HeaderProps) => {
const { t } = useTranslation(['common', 'header']);
const {
projectName,
graph,
@@ -208,7 +212,7 @@ export const Header = ({
{availableRepos.length > 0 && (
<div>
<div className="px-3 pt-2.5 pb-1.5 text-[10px] font-medium tracking-wider text-text-muted uppercase">
Repositories
{t('header:repositories')}
</div>
{availableRepos.map((repo) => (
<div
@@ -232,7 +236,7 @@ export const Header = ({
</span>
{repo.name === projectName && (
<span className="shrink-0 font-mono text-[10px] text-accent">
active
{t('header:active')}
</span>
)}
</button>
@@ -245,7 +249,7 @@ export const Header = ({
setReanalyzeProgress({
phase: 'queued',
percent: 0,
message: 'Starting...',
message: t('common:progress.starting'),
});
try {
const { jobId } = await startAnalyze({
@@ -282,8 +286,8 @@ export const Header = ({
}`}
title={
reanalyzing === repo.name
? 'Re-analyzing...'
: `Re-analyze ${repo.name}`
? t('header:reanalyzing')
: t('header:reanalyzeRepo', { repoName: repo.name })
}
>
<RefreshCw
@@ -317,7 +321,7 @@ export const Header = ({
}
}}
className="cursor-pointer rounded p-1 text-text-muted/0 transition-all group-hover:text-text-muted hover:!text-red-400"
title={`Delete ${repo.name}`}
title={t('header:deleteRepo', { repoName: repo.name })}
>
<Trash2 className="h-3.5 w-3.5" />
</button>
@@ -332,7 +336,10 @@ export const Header = ({
<div className="mb-1.5 flex items-center gap-2">
<Loader2 className="h-3 w-3 shrink-0 animate-spin text-accent" />
<span className="truncate text-xs text-text-secondary">
Re-analyzing {reanalyzing}: {reanalyzeProgress.message}
{t('header:reanalyzingRepo', {
repoName: reanalyzing,
message: translateProgressMessage(reanalyzeProgress.message, t),
})}
</span>
</div>
<div className="h-1 overflow-hidden rounded-full bg-elevated">
@@ -359,7 +366,7 @@ export const Header = ({
>
<Sparkles className="h-3.5 w-3.5 shrink-0 text-accent" />
<span className="text-sm text-text-secondary">
Analyze a new repository...
{t('header:analyzeNew')}
</span>
</button>
</div>
@@ -378,7 +385,7 @@ export const Header = ({
<input
ref={inputRef}
type="text"
placeholder="Search nodes..."
placeholder={t('header:searchNodes')}
value={searchQuery}
onChange={(e) => {
setSearchQuery(e.target.value);
@@ -399,7 +406,7 @@ export const Header = ({
<div className="absolute top-full right-0 left-0 z-50 mt-1 overflow-hidden rounded-xl border border-border-subtle bg-surface shadow-xl">
{searchResults.length === 0 ? (
<div className="px-4 py-3 text-sm text-text-muted">
No nodes found for &ldquo;{searchQuery}&rdquo;
{t('header:noNodesFound', { query: searchQuery })}
</div>
) : (
<div className="max-h-80 overflow-y-auto">
@@ -441,7 +448,7 @@ export const Header = ({
className="group flex items-center gap-2 rounded-lg bg-gradient-to-r from-purple-600 to-pink-600 px-3.5 py-2 text-sm font-medium text-white shadow-lg transition-all duration-200 hover:-translate-y-0.5 hover:from-purple-500 hover:to-pink-500 hover:shadow-xl"
>
<Github className="h-4 w-4" />
<span className="hidden sm:inline">Star if cool</span>
<span className="hidden sm:inline">{t('header:starIfCool')}</span>
<Star className="h-3.5 w-3.5 transition-all group-hover:fill-yellow-300 group-hover:text-yellow-300" />
<span className="hidden sm:inline">✨</span>
</a>
@@ -449,24 +456,26 @@ export const Header = ({
{/* Stats */}
{graph && (
<div className="mr-2 flex items-center gap-4 text-xs text-text-muted">
<span>{nodeCount} nodes</span>
<span>{edgeCount} edges</span>
<span>{t('common:counts.nodes', { count: nodeCount })}</span>
<span>{t('common:counts.edges', { count: edgeCount })}</span>
</div>
)}
{/* Embedding Status */}
<EmbeddingStatus />
<LanguageSwitcher />
{/* Icon buttons */}
<button
onClick={() => setSettingsPanelOpen(true)}
className="flex h-9 w-9 cursor-pointer items-center justify-center rounded-md text-text-secondary transition-colors hover:bg-hover hover:text-text-primary"
title="AI Settings"
title={t('header:aiSettings')}
>
<Settings className="h-4.5 w-4.5" />
</button>
<button
title="Help"
title={t('header:help')}
onClick={() => setHelpDialogBoxOpen(true)}
className="flex h-9 w-9 cursor-pointer items-center justify-center rounded-md text-text-secondary transition-colors hover:bg-hover hover:text-text-primary"
>
@@ -483,7 +492,7 @@ export const Header = ({
} `}
>
<Sparkles className="h-4 w-4" />
<span>Nexus AI</span>
<span>{t('common:app.nexusAI')}</span>
</button>
</div>
</header>
+99 -102
View File
@@ -1,5 +1,6 @@
import React, { useState } from 'react';
import { X, GitBranch, Search, Filter, Zap, Keyboard, BarChart2, HelpCircle } from 'lucide-react';
import { useTranslation } from 'react-i18next';
interface HelpPanelProps {
isOpen: boolean;
@@ -12,34 +13,33 @@ type TabId = 'overview' | 'graph' | 'search' | 'ai' | 'shortcuts' | 'status';
interface Tab {
id: TabId;
label: string;
icon: React.ReactNode;
}
const tabs: Tab[] = [
{ id: 'overview', label: 'Overview', icon: <HelpCircle className="h-4 w-4" /> },
{ id: 'graph', label: 'Graph & nodes', icon: <GitBranch className="h-4 w-4" /> },
{ id: 'search', label: 'Search & filter', icon: <Search className="h-4 w-4" /> },
{ id: 'ai', label: 'Nexus AI', icon: <Zap className="h-4 w-4" /> },
{ id: 'shortcuts', label: 'Shortcuts', icon: <Keyboard className="h-4 w-4" /> },
{ id: 'status', label: 'Status bar', icon: <BarChart2 className="h-4 w-4" /> },
{ id: 'overview', icon: <HelpCircle className="h-4 w-4" /> },
{ id: 'graph', icon: <GitBranch className="h-4 w-4" /> },
{ id: 'search', icon: <Search className="h-4 w-4" /> },
{ id: 'ai', icon: <Zap className="h-4 w-4" /> },
{ id: 'shortcuts', icon: <Keyboard className="h-4 w-4" /> },
{ id: 'status', icon: <BarChart2 className="h-4 w-4" /> },
];
const shortcuts = [
{ label: 'Search nodes', mac: '⌘ K', win: 'Ctrl K' },
{ label: 'Deselect / close', mac: 'Esc', win: 'Esc' },
{ labelKey: 'shortcuts.searchNodes', mac: '⌘ K', win: 'Ctrl K' },
{ labelKey: 'shortcuts.deselectClose', mac: 'Esc', win: 'Esc' },
];
const nodeColors = [
{ color: '#10b981', label: 'Function', desc: 'Function declarations' },
{ color: '#3b82f6', label: 'File', desc: 'Source files' },
{ color: '#f59e0b', label: 'Class', desc: 'Class declarations' },
{ color: '#14b8a6', label: 'Method', desc: 'Class methods' },
{ color: '#ec4899', label: 'Interface', desc: 'TypeScript interfaces' },
{ color: '#6366f1', label: 'Folder', desc: 'Directory nodes' },
{ color: '#10b981', labelKey: 'nodeTypes.function', descKey: 'nodeTypes.functionDesc' },
{ color: '#3b82f6', labelKey: 'nodeTypes.file', descKey: 'nodeTypes.fileDesc' },
{ color: '#f59e0b', labelKey: 'nodeTypes.class', descKey: 'nodeTypes.classDesc' },
{ color: '#14b8a6', labelKey: 'nodeTypes.method', descKey: 'nodeTypes.methodDesc' },
{ color: '#ec4899', labelKey: 'nodeTypes.interface', descKey: 'nodeTypes.interfaceDesc' },
{ color: '#6366f1', labelKey: 'nodeTypes.folder', descKey: 'nodeTypes.folderDesc' },
];
const getStatusItems = (nodeCount: number, edgeCount: number) => [
const getStatusItems = (t: (key: string) => string, nodeCount: number, edgeCount: number) => [
{
badge: (
<span
@@ -53,8 +53,8 @@ const getStatusItems = (nodeCount: number, edgeCount: number) => [
}}
/>
),
title: 'Ready',
desc: 'Graph is fully loaded and interactive',
title: t('status.ready'),
desc: t('status.readyDesc'),
},
{
badge: (
@@ -62,8 +62,8 @@ const getStatusItems = (nodeCount: number, edgeCount: number) => [
{nodeCount}
</span>
),
title: 'Nodes count',
desc: 'Total files and symbols in the graph',
title: t('status.nodesCount'),
desc: t('status.nodesCountDesc'),
},
{
badge: (
@@ -71,8 +71,8 @@ const getStatusItems = (nodeCount: number, edgeCount: number) => [
{edgeCount}
</span>
),
title: 'Edges count',
desc: 'Import / dependency connections',
title: t('status.edgesCount'),
desc: t('status.edgesCountDesc'),
},
{
badge: (
@@ -85,11 +85,11 @@ const getStatusItems = (nodeCount: number, edgeCount: number) => [
whiteSpace: 'nowrap',
}}
>
Semantic Ready
{t('status.semanticReadyBadge')}
</span>
),
title: 'AI index status',
desc: 'Repo is fully indexed for AI queries',
title: t('status.aiIndexStatus'),
desc: t('status.aiIndexStatusDesc'),
},
// { badge: <span style={{ fontSize: 11, fontWeight: 500, color: '#9ca3af', flexShrink: 0 }}>typescript</span>, title: 'Language', desc: 'Primary language detected in the repo' },
];
@@ -119,6 +119,8 @@ function TabContent({
nodeCount: number;
edgeCount: number;
}) {
const { t } = useTranslation('help');
if (active === 'overview')
return (
<div style={{ display: 'flex', flexDirection: 'column', gap: 10 }}>
@@ -131,7 +133,7 @@ function TabContent({
letterSpacing: '0.08em',
}}
>
Getting started
{t('overview.gettingStarted')}
</p>
<div
@@ -143,11 +145,10 @@ function TabContent({
}}
>
<p style={{ fontSize: 13, fontWeight: 500, color: '#e2e2e8', margin: '0 0 4px' }}>
What is GitNexus?
{t('overview.whatIsTitle')}
</p>
<p style={{ fontSize: 12, color: '#9ca3af', margin: 0, lineHeight: 1.6 }}>
An interactive graph explorer for your codebase. Every file, function, and import
becomes a node you can explore, query, and navigate visually.
{t('overview.whatIsDescription')}
</p>
</div>
@@ -160,11 +161,10 @@ function TabContent({
}}
>
<p style={{ fontSize: 13, fontWeight: 500, color: '#e2e2e8', margin: '0 0 4px' }}>
Your current repo
{t('overview.currentRepoTitle')}
</p>
<p style={{ fontSize: 12, color: '#9ca3af', margin: 0, lineHeight: 1.6 }}>
Loaded: <span style={{ color: '#a78bfa', fontFamily: 'monospace' }}></span> {nodeCount}{' '}
nodes · {edgeCount} edges
{t('overview.loadedCounts', { nodeCount, edgeCount })}
</p>
</div>
@@ -177,15 +177,16 @@ function TabContent({
}}
>
<p style={{ fontSize: 13, fontWeight: 500, color: '#e2e2e8', margin: '0 0 4px' }}>
Three ways to explore
{t('overview.threeWaysTitle')}
</p>
<p style={{ fontSize: 12, color: '#9ca3af', margin: 0, lineHeight: 1.6 }}>
<strong style={{ color: '#e2e2e8', fontWeight: 500 }}>1.</strong> Click nodes to inspect
<strong style={{ color: '#e2e2e8', fontWeight: 500 }}>1.</strong>{' '}
{t('overview.wayInspect')}
<br />
<strong style={{ color: '#e2e2e8', fontWeight: 500 }}>2.</strong> Search by name or type
<strong style={{ color: '#e2e2e8', fontWeight: 500 }}>2.</strong>{' '}
{t('overview.waySearch')}
<br />
<strong style={{ color: '#e2e2e8', fontWeight: 500 }}>3.</strong> Ask Nexus AI a natural
language question
<strong style={{ color: '#e2e2e8', fontWeight: 500 }}>3.</strong> {t('overview.wayAsk')}
</p>
</div>
@@ -198,11 +199,11 @@ function TabContent({
}}
>
<p style={{ fontSize: 13, fontWeight: 500, color: '#e2e2e8', margin: '0 0 4px' }}>
Navigation
{t('overview.navigationTitle')}
</p>
<p style={{ fontSize: 12, color: '#9ca3af', margin: 0, lineHeight: 1.6 }}>
· Scroll to zoom <br />
· Click and drag to pan <br />· Double-click a node to focus its subgraph
· {t('overview.navZoom')} <br />· {t('overview.navPan')} <br />·{' '}
{t('overview.navFocus')}
</p>
</div>
</div>
@@ -220,44 +221,44 @@ function TabContent({
letterSpacing: '0.08em',
}}
>
Node color legend
{t('graph.nodeColorLegend')}
</p>
{nodeColors.map(({ color, label, desc }) => (
<div key={label} style={{ display: 'flex', gap: 10, alignItems: 'flex-start' }}>
<span
style={{
width: 12,
height: 12,
borderRadius: '50%',
background: color,
flexShrink: 0,
marginTop: 2,
}}
/>
<div>
<p style={{ fontSize: 12, fontWeight: 500, color: '#e2e2e8', margin: '0 0 2px' }}>
{label} nodes
</p>
<p style={{ fontSize: 12, color: '#9ca3af', margin: 0 }}>{desc}</p>
{nodeColors.map(({ color, labelKey, descKey }) => {
const label = t(labelKey);
return (
<div key={labelKey} style={{ display: 'flex', gap: 10, alignItems: 'flex-start' }}>
<span
style={{
width: 12,
height: 12,
borderRadius: '50%',
background: color,
flexShrink: 0,
marginTop: 2,
}}
/>
<div>
<p style={{ fontSize: 12, fontWeight: 500, color: '#e2e2e8', margin: '0 0 2px' }}>
{t('graph.nodeLabel', { label })}
</p>
<p style={{ fontSize: 12, color: '#9ca3af', margin: 0 }}>{t(descKey)}</p>
</div>
</div>
</div>
))}
);
})}
<div style={{ borderTop: '0.5px solid rgba(255,255,255,0.08)', margin: '4px 0' }} />
<p style={{ fontSize: 12, color: '#9ca3af', margin: 0, lineHeight: 1.6 }}>
Node <strong style={{ color: '#e2e2e8', fontWeight: 500 }}>size</strong> reflects
connection count — larger nodes are depended on by more files. Edges point from importer →
imported.
{t('graph.sizeDescription')}
</p>
<div
style={{ background: 'rgba(255,255,255,0.04)', borderRadius: 10, padding: '10px 14px' }}
>
<p style={{ fontSize: 12, color: '#9ca3af', margin: 0, lineHeight: 1.6 }}>
Click any node to open its detail panel — showing imports, exports, and reverse
dependencies.
{t('graph.detailDescription')}
</p>
</div>
</div>
@@ -275,7 +276,7 @@ function TabContent({
letterSpacing: '0.08em',
}}
>
Search & filter
{t('search.title')}
</p>
<div
@@ -284,12 +285,11 @@ function TabContent({
<div style={{ display: 'flex', alignItems: 'center', gap: 8, marginBottom: 6 }}>
<kbd style={kbdStyle}>⌘K</kbd>/<kbd style={kbdStyle}>Ctrl K</kbd>
<p style={{ fontSize: 12, fontWeight: 500, color: '#e2e2e8', margin: 0 }}>
Search nodes
{t('search.searchNodes')}
</p>
</div>
<p style={{ fontSize: 12, color: '#9ca3af', margin: 0, lineHeight: 1.6 }}>
Search by filename, function name, or import path. Matching nodes are highlighted live
in the graph.
{t('search.searchDescription')}
</p>
</div>
@@ -299,12 +299,11 @@ function TabContent({
<div style={{ display: 'flex', alignItems: 'center', gap: 8, marginBottom: 6 }}>
<Filter style={{ width: 14, height: 14, color: '#a78bfa', flexShrink: 0 }} />
<p style={{ fontSize: 12, fontWeight: 500, color: '#e2e2e8', margin: 0 }}>
Filter panel
{t('search.filterPanel')}
</p>
</div>
<p style={{ fontSize: 12, color: '#9ca3af', margin: 0, lineHeight: 1.6 }}>
Use the filter icon in the left sidebar to isolate specific node types, hide leaf nodes,
or focus on a depth range from a selected root.
{t('search.filterDescription')}
</p>
</div>
@@ -312,13 +311,13 @@ function TabContent({
style={{ background: 'rgba(255,255,255,0.04)', borderRadius: 10, padding: '12px 14px' }}
>
<p style={{ fontSize: 12, fontWeight: 500, color: '#e2e2e8', margin: '0 0 6px' }}>
Search syntax
{t('search.syntax')}
</p>
{[
{ query: 'auth', hint: 'match by name fragment' },
{ query: './utils/', hint: 'match by path prefix' },
{ query: 'type:config', hint: 'filter by node type' },
].map(({ query, hint }) => (
{ query: 'auth', hintKey: 'search.hints.nameFragment' },
{ query: './utils/', hintKey: 'search.hints.pathPrefix' },
{ query: 'type:config', hintKey: 'search.hints.nodeType' },
].map(({ query, hintKey }) => (
<div
key={query}
style={{ display: 'flex', alignItems: 'baseline', gap: 8, marginBottom: 4 }}
@@ -336,7 +335,7 @@ function TabContent({
>
{query}
</code>
<span style={{ fontSize: 12, color: '#6b7280' }}>{hint}</span>
<span style={{ fontSize: 12, color: '#6b7280' }}>{t(hintKey)}</span>
</div>
))}
</div>
@@ -355,7 +354,7 @@ function TabContent({
letterSpacing: '0.08em',
}}
>
Nexus AI
{t('ai.title')}
</p>
<div
@@ -367,20 +366,19 @@ function TabContent({
}}
>
<p style={{ fontSize: 12, fontWeight: 500, color: '#a78bfa', margin: '0 0 4px' }}>
✓ Semantic Ready
{t('ai.semanticReady')}
</p>
<p style={{ fontSize: 12, color: '#9ca3af', margin: 0, lineHeight: 1.6 }}>
Your repo is indexed and ready for semantic queries. Nexus AI understands code structure
and relationships, not just file names.
{t('ai.description')}
</p>
</div>
<p style={{ fontSize: 12, color: '#9ca3af', margin: '4px 0 2px' }}>Try asking:</p>
<p style={{ fontSize: 12, color: '#9ca3af', margin: '4px 0 2px' }}>{t('tryAsking')}</p>
{[
'"Which files depend on the auth module?"',
'"Find circular dependencies in this repo"',
'"What are the most connected components?"',
'"Show me all files that import useEffect"',
t('ai.questions.dependencies'),
t('ai.questions.circular'),
t('ai.questions.connected'),
t('ai.questions.imports'),
].map((q) => (
<div
key={q}
@@ -400,8 +398,7 @@ function TabContent({
<div style={{ borderTop: '0.5px solid rgba(255,255,255,0.08)', margin: '4px 0' }} />
<p style={{ fontSize: 12, color: '#6b7280', margin: 0, lineHeight: 1.6 }}>
Open the prompt via the <span style={{ color: '#e2e2e8' }}>Nexus AI</span> button
(top-right).
{t('ai.openPrompt')}
</p>
</div>
);
@@ -428,7 +425,7 @@ function TabContent({
letterSpacing: '0.08em',
}}
>
Action
{t('shortcuts.columns.action')}
</span>
<span
style={{
@@ -454,9 +451,9 @@ function TabContent({
</span>
</div>
{shortcuts.map(({ label, mac, win }, i) => (
{shortcuts.map(({ labelKey, mac, win }, i) => (
<div
key={label}
key={labelKey}
style={{
display: 'grid',
gridTemplateColumns: '1fr 80px 88px',
@@ -467,7 +464,7 @@ function TabContent({
i < shortcuts.length - 1 ? '0.5px solid rgba(255,255,255,0.05)' : 'none',
}}
>
<span style={{ fontSize: 12, color: '#9ca3af' }}>{label}</span>
<span style={{ fontSize: 12, color: '#9ca3af' }}>{t(labelKey)}</span>
<span style={{ display: 'flex', justifyContent: 'center' }}>
<kbd style={kbdStyle}>{mac}</kbd>
</span>
@@ -491,9 +488,9 @@ function TabContent({
letterSpacing: '0.08em',
}}
>
Status bar explained
{t('status.explained')}
</p>
{getStatusItems(nodeCount, edgeCount).map(({ badge, title, desc }) => (
{getStatusItems(t, nodeCount, edgeCount).map(({ badge, title, desc }) => (
<div
key={title}
style={{
@@ -521,7 +518,9 @@ function TabContent({
}
export const HelpPanel = ({ isOpen, onClose, nodeCount, edgeCount }: HelpPanelProps) => {
const { t } = useTranslation('help');
const [active, setActive] = useState<TabId>('overview');
const localizedTabs = tabs.map((tab) => ({ ...tab, label: t(`tabs.${tab.id}`) }));
if (!isOpen) return null;
@@ -592,9 +591,9 @@ export const HelpPanel = ({ isOpen, onClose, nodeCount, edgeCount }: HelpPanelPr
</div>
<div>
<h2 style={{ fontSize: 16, fontWeight: 600, color: '#e2e2e8', margin: 0 }}>
Help & Reference
{t('title')}
</h2>
<p style={{ fontSize: 12, color: '#6b7280', margin: 0 }}>GitNexus — graph explorer</p>
<p style={{ fontSize: 12, color: '#6b7280', margin: 0 }}>{t('footer')}</p>
</div>
</div>
<button
@@ -632,7 +631,7 @@ export const HelpPanel = ({ isOpen, onClose, nodeCount, edgeCount }: HelpPanelPr
gap: 2,
}}
>
{tabs.map(({ id, label, icon }) => {
{localizedTabs.map(({ id, label, icon }) => {
const isActive = active === id;
return (
<button
@@ -699,16 +698,14 @@ export const HelpPanel = ({ isOpen, onClose, nodeCount, edgeCount }: HelpPanelPr
background: 'rgba(255,255,255,0.01)',
}}
>
<span style={{ fontSize: 11, color: '#4b5563' }}>
GitNexus — open source codebase graph explorer
</span>
<span style={{ fontSize: 11, color: '#4b5563' }}>{t('footerLong')}</span>
<a
href="https://github.com/abhigyanpatwari/GitNexus"
target="_blank"
rel="noopener noreferrer"
style={{ fontSize: 11, color: '#a78bfa', textDecoration: 'none' }}
>
Docs & GitHub ↗
{t('docsGithub')}
</a>
</div>
</div>
@@ -0,0 +1,43 @@
import { Globe } from '@/lib/lucide-icons';
import { useTranslation } from 'react-i18next';
import { SUPPORTED_LANGUAGES, type SupportedLanguage } from '../i18n/languages';
export const LanguageSwitcher = () => {
const { t, i18n } = useTranslation('header');
const currentLanguage = i18n.resolvedLanguage || i18n.language;
const currentLanguageMetadata =
SUPPORTED_LANGUAGES.find(
(language) => language.code.toLowerCase() === currentLanguage.toLowerCase(),
) ?? SUPPORTED_LANGUAGES[0];
const handleChange = (language: SupportedLanguage) => {
void i18n.changeLanguage(language);
};
return (
<label
className="flex h-9 items-center gap-1.5 rounded-md border border-border-subtle bg-surface px-2 text-text-secondary transition-colors hover:border-border-default hover:bg-hover hover:text-text-primary"
title={t('selectLanguage')}
>
<Globe className="h-4 w-4" aria-hidden="true" />
<span className="sr-only">{t('language')}</span>
<select
data-testid="language-switcher"
value={currentLanguageMetadata.code}
aria-label={t('selectLanguage')}
onChange={(event) => handleChange(event.target.value as SupportedLanguage)}
className="cursor-pointer border-none bg-transparent text-xs font-medium outline-none"
>
{SUPPORTED_LANGUAGES.map((language) => (
<option
key={language.code}
value={language.code}
className="bg-surface text-text-primary"
>
{language.nativeName}
</option>
))}
</select>
</label>
);
};
+13 -4
View File
@@ -1,10 +1,16 @@
import type { PipelineProgress } from 'gitnexus-shared';
import { useTranslation } from 'react-i18next';
import { translateProgressMessage } from '../i18n/progress';
interface LoadingOverlayProps {
progress: PipelineProgress;
}
export const LoadingOverlay = ({ progress }: LoadingOverlayProps) => {
const { t } = useTranslation(['common', 'graph']);
const message = translateProgressMessage(progress.message, t);
const detail = translateProgressMessage(progress.detail, t);
return (
<div className="fixed inset-0 z-50 flex flex-col items-center justify-center bg-void">
{/* Background gradient effects */}
@@ -32,11 +38,11 @@ export const LoadingOverlay = ({ progress }: LoadingOverlayProps) => {
{/* Status text */}
<div className="text-center">
<p className="mb-1 font-mono text-sm text-text-secondary">
{progress.message}
{message}
<span className="animate-pulse">|</span>
</p>
{progress.detail && (
<p className="max-w-md truncate font-mono text-xs text-text-muted">{progress.detail}</p>
<p className="max-w-md truncate font-mono text-xs text-text-muted">{detail}</p>
)}
</div>
@@ -46,12 +52,15 @@ export const LoadingOverlay = ({ progress }: LoadingOverlayProps) => {
<div className="flex items-center gap-2">
<span className="h-2 w-2 rounded-full bg-node-file" />
<span>
{progress.stats.filesProcessed} / {progress.stats.totalFiles} files
{t('graph:loading.filesProgress', {
processed: progress.stats.filesProcessed,
total: progress.stats.totalFiles,
})}
</span>
</div>
<div className="flex items-center gap-2">
<span className="h-2 w-2 rounded-full bg-node-function" />
<span>{progress.stats.nodesCreated} nodes</span>
<span>{t('common:counts.nodes', { count: progress.stats.nodesCreated })}</span>
</div>
</div>
)}
@@ -5,6 +5,7 @@ import { Prism as SyntaxHighlighter } from 'react-syntax-highlighter';
import { vscDarkPlus } from 'react-syntax-highlighter/dist/esm/styles/prism';
import { MermaidDiagram } from './MermaidDiagram';
import { ToolCallCard } from './ToolCallCard';
import { useTranslation } from 'react-i18next';
import { Copy, Check } from '@/lib/lucide-icons';
// Custom syntax theme
@@ -38,6 +39,7 @@ export const MarkdownRenderer: React.FC<MarkdownRendererProps> = ({
toolCalls,
showCopyButton = false,
}) => {
const { t } = useTranslation('common');
const [copied, setCopied] = useState(false);
const copyTimerRef = useRef<ReturnType<typeof setTimeout>>(undefined);
@@ -125,7 +127,9 @@ export const MarkdownRenderer: React.FC<MarkdownRendererProps> = ({
href={hrefStr}
onClick={(e) => handleLinkClick(e, hrefStr)}
className={`${baseParams} ${colorParams}`}
title={isNodeRef ? `View ${inner} in Code panel` : `Open in Code panel • ${inner}`}
title={t(isNodeRef ? 'chat.viewNodeInCodePanel' : 'chat.openInCodePanel', {
inner,
})}
{...props}
>
<span className="text-inherit">{children}</span>
@@ -182,7 +186,7 @@ export const MarkdownRenderer: React.FC<MarkdownRendererProps> = ({
},
pre: ({ children }: any) => <>{children}</>,
}),
[handleLinkClick],
[handleLinkClick, t],
);
return (
@@ -205,14 +209,14 @@ export const MarkdownRenderer: React.FC<MarkdownRendererProps> = ({
<button
onClick={handleCopy}
className="flex items-center gap-1.5 rounded border border-transparent px-2 py-1 text-xs text-text-muted transition-all hover:border-border-subtle hover:bg-surface hover:text-text-primary"
title="Copy to clipboard"
title={t('actions.copy')}
>
{copied ? (
<Check className="h-3.5 w-3.5 text-emerald-400" />
) : (
<Copy className="h-3.5 w-3.5" />
)}
<span>{copied ? 'Copied' : 'Copy'}</span>
<span>{copied ? t('actions.copied') : t('actions.copy')}</span>
</button>
</div>
)}
+10 -6
View File
@@ -1,4 +1,5 @@
import { Suspense, useEffect, useRef, useState, lazy } from 'react';
import { useTranslation } from 'react-i18next';
import mermaid from 'mermaid';
import DOMPurify from 'dompurify';
import { AlertTriangle, Maximize2 } from '@/lib/lucide-icons';
@@ -55,6 +56,7 @@ interface MermaidDiagramProps {
}
export const MermaidDiagram = ({ code }: MermaidDiagramProps) => {
const { t } = useTranslation(['graph']);
const containerRef = useRef<HTMLDivElement>(null);
const [error, setError] = useState<string | null>(null);
const [showModal, setShowModal] = useState(false);
@@ -98,7 +100,7 @@ export const MermaidDiagram = ({ code }: MermaidDiagramProps) => {
const processData: any = showModal
? {
id: 'ai-generated',
label: 'AI Generated Diagram',
label: t('graph:diagram.aiGenerated'),
processType: 'intra_community',
steps: [], // Empty - we'll render raw mermaid
edges: [],
@@ -112,12 +114,12 @@ export const MermaidDiagram = ({ code }: MermaidDiagramProps) => {
<div className="my-3 rounded-lg border border-rose-500/30 bg-rose-500/10 p-4">
<div className="mb-2 flex items-center gap-2 text-sm text-rose-300">
<AlertTriangle className="h-4 w-4" />
<span className="font-medium">Diagram Error</span>
<span className="font-medium">{t('graph:diagram.error')}</span>
</div>
<pre className="font-mono text-xs whitespace-pre-wrap text-rose-200/70">{error}</pre>
<details className="mt-2">
<summary className="cursor-pointer text-xs text-text-muted hover:text-text-secondary">
Show source
{t('graph:diagram.showSource')}
</summary>
<pre className="mt-2 overflow-x-auto rounded bg-surface p-2 text-xs text-text-muted">
{code}
@@ -134,12 +136,12 @@ export const MermaidDiagram = ({ code }: MermaidDiagramProps) => {
{/* Header */}
<div className="flex items-center justify-between border-b border-border-subtle bg-surface/60 px-3 py-2">
<span className="text-[10px] font-medium tracking-wider text-text-muted uppercase">
Diagram
{t('graph:diagram.label')}
</span>
<button
onClick={() => setShowModal(true)}
className="rounded p-1 text-text-muted transition-colors hover:bg-hover hover:text-text-primary"
title="Expand"
title={t('graph:diagram.expandTitle')}
>
<Maximize2 className="h-3.5 w-3.5" />
</button>
@@ -161,7 +163,9 @@ export const MermaidDiagram = ({ code }: MermaidDiagramProps) => {
{/* Use ProcessFlowModal for expansion */}
{showModal && processData && (
<Suspense fallback={<div className="p-4 text-sm text-text-muted">Loading diagram…</div>}>
<Suspense
fallback={<div className="p-4 text-sm text-text-muted">{t('graph:diagram.loading')}</div>}
>
<ProcessFlowModal process={processData} onClose={() => setShowModal(false)} />
</Suspense>
)}
+23 -21
View File
@@ -1,6 +1,7 @@
import { useState, useRef, useEffect } from 'react';
import { Check, Copy, Terminal, Server, Zap, Sparkles } from '@/lib/lucide-icons';
import { REQUIRED_NODE_VERSION } from '../config/ui-constants';
import { useTranslation } from 'react-i18next';
// ── Design constants ─────────────────────────────────────────────────────────
@@ -9,6 +10,7 @@ const isDev = import.meta.env.DEV;
// ── Copy-to-clipboard button ─────────────────────────────────────────────────
function CopyButton({ text }: { text: string }) {
const { t } = useTranslation('onboarding');
const [copied, setCopied] = useState(false);
const timerRef = useRef<ReturnType<typeof setTimeout> | null>(null);
@@ -32,7 +34,7 @@ function CopyButton({ text }: { text: string }) {
return (
<button
onClick={handleCopy}
aria-label={copied ? 'Copied!' : 'Copy to clipboard'}
aria-label={copied ? t('guide.copiedAria') : t('guide.copyAria')}
className={`shrink-0 cursor-pointer rounded-md px-2 py-1 transition-all duration-200 focus-visible:ring-2 focus-visible:ring-accent/40 focus-visible:outline-none ${
copied
? 'bg-emerald-400/10 text-emerald-400'
@@ -128,6 +130,7 @@ function StepRow({
description?: string;
children?: React.ReactNode;
}) {
const { t } = useTranslation('onboarding');
const isVisible = state !== 'waiting';
return (
@@ -151,7 +154,7 @@ function StepRow({
</span>
{state === 'done' && (
<span className="animate-fade-in font-mono text-[10px] tracking-wider text-emerald-400/60 uppercase">
done
{t('guide.done')}
</span>
)}
</div>
@@ -168,6 +171,8 @@ function StepRow({
// ── Polling status bar ────────────────────────────────────────────────────────
function PollingBar() {
const { t } = useTranslation('onboarding');
return (
<div
className="flex animate-fade-in items-center gap-3 rounded-xl border border-accent/15 bg-accent/5 px-4 py-3"
@@ -183,12 +188,12 @@ function PollingBar() {
<div className="min-w-0 flex-1">
<p className="text-xs font-medium text-text-secondary">
Listening for server
{t('guide.listeningForServer')}
<span className="ml-0.5 inline-flex text-text-muted">
<span className="animate-pulse">...</span>
</span>
</p>
<p className="mt-0.5 text-[11px] text-text-muted">Will auto-connect when detected</p>
<p className="mt-0.5 text-[11px] text-text-muted">{t('guide.willAutoConnect')}</p>
</div>
</div>
);
@@ -201,8 +206,9 @@ interface OnboardingGuideProps {
}
export const OnboardingGuide = ({ isPolling }: OnboardingGuideProps) => {
const { t } = useTranslation('onboarding');
const primary = isDev ? 'npm run --prefix gitnexus serve' : 'npx gitnexus@latest serve';
const termLabel = isDev ? 'Start backend' : 'Terminal';
const termLabel = isDev ? t('guide.startBackend') : t('guide.terminal');
// Step states: step 1 = copy command, step 2 = run/wait, step 3 = auto-connect
// Once polling starts the user has presumably run the command — mark step 1 done.
@@ -226,12 +232,10 @@ export const OnboardingGuide = ({ isPolling }: OnboardingGuideProps) => {
</span>
</div>
<h2 className="text-lg leading-snug font-semibold text-text-primary">
Start your local server
{t('guide.startServer')}
</h2>
<p className="mx-auto mt-1 max-w-xs text-sm leading-relaxed text-text-secondary">
{isDev
? 'Fire up the Express backend in a separate terminal to unlock the full graph.'
: 'One command is all it takes. The browser connects automatically.'}
{isDev ? t('guide.devDescription') : t('guide.prodDescription')}
</p>
</div>
</div>
@@ -248,8 +252,8 @@ export const OnboardingGuide = ({ isPolling }: OnboardingGuideProps) => {
<StepRow
state={step1State}
number={1}
title="Copy the command"
description={isPolling ? undefined : 'Click the icon in the terminal to copy.'}
title={t('guide.copyCommand')}
description={isPolling ? undefined : t('guide.copyCommandDescription')}
>
<TerminalWindow command={primary} label={termLabel} isActive={step1State === 'active'} />
@@ -259,13 +263,13 @@ export const OnboardingGuide = ({ isPolling }: OnboardingGuideProps) => {
<div className="my-3 flex items-center gap-3">
<div className="h-px flex-1 bg-border-subtle" />
<span className="text-[11px] tracking-widest text-text-muted uppercase">
or install globally
{t('guide.orInstallGlobally')}
</span>
<div className="h-px flex-1 bg-border-subtle" />
</div>
<TerminalWindow
command="npm install -g gitnexus && gitnexus serve"
label="Global install"
label={t('guide.globalInstall')}
isActive={false}
/>
</>
@@ -276,10 +280,8 @@ export const OnboardingGuide = ({ isPolling }: OnboardingGuideProps) => {
<StepRow
state={step2State}
number={2}
title={isPolling ? 'Waiting for server to start' : 'Paste and run in your terminal'}
description={
isPolling ? undefined : 'Open a terminal at the project root, paste, and hit Enter.'
}
title={isPolling ? t('guide.waitingForServer') : t('guide.pasteAndRun')}
description={isPolling ? undefined : t('guide.pasteAndRunDescription')}
>
{isPolling && <PollingBar />}
</StepRow>
@@ -288,8 +290,8 @@ export const OnboardingGuide = ({ isPolling }: OnboardingGuideProps) => {
<StepRow
state={step3State}
number={3}
title="Auto-connects and opens the graph"
description="No refresh needed — the page detects the server automatically."
title={t('guide.autoConnects')}
description={t('guide.autoConnectsDescription')}
/>
</div>
@@ -297,7 +299,7 @@ export const OnboardingGuide = ({ isPolling }: OnboardingGuideProps) => {
<div className="mt-6 flex items-center justify-center gap-1.5 border-t border-border-subtle pt-5 text-xs text-text-muted">
<Server className="h-3 w-3 shrink-0" />
<span>
Requires{' '}
{t('guide.requires')}{' '}
<a
href="https://nodejs.org"
target="_blank"
@@ -309,7 +311,7 @@ export const OnboardingGuide = ({ isPolling }: OnboardingGuideProps) => {
</span>
<span className="mx-1 text-border-default">·</span>
<Terminal className="h-3 w-3 shrink-0" />
<span>Port 4747</span>
<span>{t('guide.port')}</span>
</div>
</div>
);
@@ -5,6 +5,7 @@
*/
import { useEffect, useRef, useCallback, useState } from 'react';
import { useTranslation } from 'react-i18next';
import { Copy, Focus, ZoomIn, ZoomOut } from 'lucide-react';
import mermaid from 'mermaid';
import DOMPurify from 'dompurify';
@@ -59,6 +60,7 @@ export const ProcessFlowModal = ({
onFocusInGraph,
isFullScreen = false,
}: ProcessFlowModalProps) => {
const { t } = useTranslation(['graph', 'common']);
const containerRef = useRef<HTMLDivElement>(null);
const diagramRef = useRef<HTMLDivElement>(null);
const scrollContainerRef = useRef<HTMLDivElement>(null);
@@ -171,13 +173,13 @@ export const ProcessFlowModal = ({
diagramRef.current!.innerHTML = `
<div class="text-center p-8">
<div class="text-red-400 text-sm font-medium mb-2">
${isSizeError ? '📊 Diagram Too Large' : '⚠️ Render Error'}
${isSizeError ? t('graph:processFlow.diagramTooLarge') : t('graph:processFlow.renderError')}
</div>
<div class="text-slate-400 text-xs max-w-md">
${
isSizeError
? `This diagram has ${process.steps?.length || 0} steps and is too complex to render. Try viewing individual processes instead of "All Processes".`
: `Unable to render diagram. Steps: ${process.steps?.length || 0}`
? t('graph:processFlow.tooComplex', { count: process.steps?.length || 0 })
: t('graph:processFlow.unableToRender', { count: process.steps?.length || 0 })
}
</div>
</div>
@@ -186,7 +188,7 @@ export const ProcessFlowModal = ({
};
renderDiagram();
}, [process]);
}, [process, t]);
// Close on escape
useEffect(() => {
@@ -242,7 +244,9 @@ export const ProcessFlowModal = ({
{/* Header */}
<div className="relative z-10 border-b border-white/10 px-6 py-5">
<h2 className="text-lg font-semibold text-white">Process: {process.label}</h2>
<h2 className="text-lg font-semibold text-white">
{t('graph:processFlow.title', { label: process.label })}
</h2>
</div>
{/* Diagram */}
@@ -271,7 +275,7 @@ export const ProcessFlowModal = ({
<button
onClick={handleZoomOut}
className="rounded-md p-2 text-slate-300 transition-all hover:bg-white/10 hover:text-white"
title="Zoom out (-)"
title={t('graph:processFlow.zoomOutTitle')}
>
<ZoomOut className="h-4 w-4" />
</button>
@@ -281,7 +285,7 @@ export const ProcessFlowModal = ({
<button
onClick={handleZoomIn}
className="rounded-md p-2 text-slate-300 transition-all hover:bg-white/10 hover:text-white"
title="Zoom in (+)"
title={t('graph:processFlow.zoomInTitle')}
>
<ZoomIn className="h-4 w-4" />
</button>
@@ -289,9 +293,9 @@ export const ProcessFlowModal = ({
<button
onClick={resetView}
className="flex items-center gap-2 rounded-lg border border-white/10 bg-white/5 px-4 py-2.5 text-sm font-medium text-slate-300 transition-all hover:bg-white/10 hover:text-white"
title="Reset zoom and pan"
title={t('graph:processFlow.resetTitle')}
>
Reset View
{t('graph:processFlow.resetView')}
</button>
{onFocusInGraph && (
<button
@@ -299,7 +303,7 @@ export const ProcessFlowModal = ({
className="flex items-center gap-2 rounded-lg bg-cyan-400 px-5 py-2.5 text-sm font-medium text-slate-900 shadow-lg shadow-cyan-500/20 transition-all hover:bg-cyan-300"
>
<Focus className="h-4 w-4" />
Toggle Focus
{t('graph:processFlow.toggleFocus')}
</button>
)}
<button
@@ -307,13 +311,13 @@ export const ProcessFlowModal = ({
className="flex items-center gap-2 rounded-lg bg-purple-600 px-5 py-2.5 text-sm font-medium text-white shadow-lg shadow-purple-500/20 transition-all hover:bg-purple-500"
>
<Copy className="h-4 w-4" />
Copy Mermaid
{t('graph:processFlow.copyMermaid')}
</button>
<button
onClick={onClose}
className="rounded-lg border border-white/10 bg-white/5 px-5 py-2.5 text-sm font-medium text-slate-300 transition-all hover:bg-white/10 hover:text-white"
>
Close
{t('common:actions.close')}
</button>
</div>
</div>
+34 -20
View File
@@ -6,6 +6,7 @@
*/
import { useState, useMemo, useCallback, useEffect } from 'react';
import { useTranslation } from 'react-i18next';
import {
GitBranch,
Search,
@@ -26,6 +27,7 @@ import type { ProcessData, ProcessStep } from '../lib/mermaid-generator';
const isSafeId = (id: string): boolean => /^[a-zA-Z0-9_:.\-/@]+$/.test(id);
export const ProcessesPanel = () => {
const { t } = useTranslation(['graph']);
const { graph, runQuery, setHighlightedNodeIds, highlightedNodeIds } = useAppState();
const [searchQuery, setSearchQuery] = useState('');
const [selectedProcess, setSelectedProcess] = useState<ProcessData | null>(null);
@@ -120,7 +122,7 @@ export const ProcessesPanel = () => {
if (!allStepsMap.has(stepId)) {
allStepsMap.set(stepId, {
id: stepId,
name: row.name || row[1] || 'Unknown',
name: row.name || row[1] || t('graph:processes.unknownStep'),
filePath: row.filePath || row[2],
stepNumber: row.stepNumber || row.step || row[3] || 0,
});
@@ -158,7 +160,7 @@ export const ProcessesPanel = () => {
const combinedProcessData: ProcessData = {
id: 'combined-all',
label: `All Processes (${allProcessIds.length} combined)`,
label: t('graph:processes.allProcessesLabel', { count: allProcessIds.length }),
processType: 'cross_community', // Treat as cross-community for styling
steps: allSteps,
edges: allEdges,
@@ -171,7 +173,7 @@ export const ProcessesPanel = () => {
} finally {
setLoadingProcess(null);
}
}, [processes, runQuery]);
}, [processes, runQuery, t]);
// Load process steps and open modal
const handleViewProcess = useCallback(
@@ -191,7 +193,7 @@ export const ProcessesPanel = () => {
const steps: ProcessStep[] = stepsResult.map((row: any) => ({
id: row.id || row[0],
name: row.name || row[1] || 'Unknown',
name: row.name || row[1] || t('graph:processes.unknownStep'),
filePath: row.filePath || row[2],
stepNumber: row.stepNumber || row.step || row[3] || 0,
}));
@@ -244,7 +246,7 @@ export const ProcessesPanel = () => {
setLoadingProcess(null);
}
},
[runQuery, graph],
[runQuery, graph, t],
);
// Cache for process steps (so we don't re-query when toggling focus)
@@ -327,10 +329,11 @@ export const ProcessesPanel = () => {
<div className="mb-4 flex h-14 w-14 items-center justify-center rounded-xl bg-surface">
<GitBranch className="h-7 w-7 text-text-muted" />
</div>
<h3 className="mb-2 text-base font-medium text-text-primary">No Processes Detected</h3>
<h3 className="mb-2 text-base font-medium text-text-primary">
{t('graph:processes.emptyTitle')}
</h3>
<p className="max-w-xs text-sm text-text-secondary">
Processes are execution flows traced from entry points. Load a codebase to see detected
processes.
{t('graph:processes.emptyDescription')}
</p>
</div>
);
@@ -347,7 +350,7 @@ export const ProcessesPanel = () => {
type="text"
value={searchQuery}
onChange={(e) => setSearchQuery(e.target.value)}
placeholder="Filter processes..."
placeholder={t('graph:processes.filterPlaceholder')}
className="flex-1 border-none bg-transparent text-sm text-text-primary outline-none placeholder:text-text-muted"
/>
</div>
@@ -356,7 +359,7 @@ export const ProcessesPanel = () => {
className="flex items-center gap-2 text-xs text-text-muted"
data-testid="process-list-loaded"
>
<span>{totalCount} processes detected</span>
<span>{t('graph:processes.detected', { count: totalCount })}</span>
</div>
</div>
@@ -374,9 +377,11 @@ export const ProcessesPanel = () => {
</div>
<div className="flex-1">
<h4 className="text-sm font-medium text-text-primary group-hover:text-cyan-200">
Full Process Map
{t('graph:processes.fullMap')}
</h4>
<p className="text-xs text-text-muted">View combined map of {totalCount} processes</p>
<p className="text-xs text-text-muted">
{t('graph:processes.viewCombined', { count: totalCount })}
</p>
</div>
{loadingProcess === 'all' ? (
<span className="mr-1 animate-spin">
@@ -401,7 +406,9 @@ export const ProcessesPanel = () => {
<ChevronRight className="h-4 w-4 text-text-muted" />
)}
<Zap className="h-4 w-4 text-amber-400" />
<span className="text-sm font-medium text-text-primary">Cross-Community</span>
<span className="text-sm font-medium text-text-primary">
{t('graph:processes.crossCommunity')}
</span>
<span className="ml-auto rounded-full bg-surface px-2 py-0.5 text-xs text-text-muted">
{filteredProcesses.cross.length}
</span>
@@ -438,7 +445,9 @@ export const ProcessesPanel = () => {
<ChevronRight className="h-4 w-4 text-text-muted" />
)}
<Home className="h-4 w-4 text-emerald-400" />
<span className="text-sm font-medium text-text-primary">Intra-Community</span>
<span className="text-sm font-medium text-text-primary">
{t('graph:processes.intraCommunity')}
</span>
<span className="ml-auto rounded-full bg-surface px-2 py-0.5 text-xs text-text-muted">
{filteredProcesses.intra.length}
</span>
@@ -492,6 +501,7 @@ const ProcessItem = ({
onView,
onToggleFocus,
}: ProcessItemProps) => {
const { t } = useTranslation(['graph']);
// Determine row styling - focused gets special highlight
const rowClass = isFocused
? 'bg-amber-950/40 border border-amber-500/50 ring-1 ring-amber-400/30'
@@ -508,11 +518,11 @@ const ProcessItem = ({
<div className="min-w-0 flex-1">
<div className="truncate text-sm text-text-primary">{process.label}</div>
<div className="flex items-center gap-2 text-xs text-text-muted">
<span>{process.stepCount} steps</span>
<span>{t('graph:processes.steps', { count: process.stepCount })}</span>
{process.clusters.length > 0 && (
<>
<span>•</span>
<span>{process.clusters.length} clusters</span>
<span>{t('graph:processes.clusters', { count: process.clusters.length })}</span>
</>
)}
</div>
@@ -525,7 +535,11 @@ const ProcessItem = ({
? 'animate-pulse border border-amber-400/40 bg-amber-500/20 text-amber-400 opacity-100 hover:bg-amber-500/30 hover:text-amber-300'
: 'border border-white/10 bg-white/5 text-text-muted opacity-0 group-hover:opacity-100 hover:border-cyan-400/40 hover:bg-cyan-500/20 hover:text-cyan-400'
}`}
title={isFocused ? 'Click to remove highlight from graph' : 'Click to highlight in graph'}
title={
isFocused
? t('graph:processes.removeHighlightTitle')
: t('graph:processes.highlightTitle')
}
data-testid="process-highlight-button"
>
<Lightbulb className="h-4 w-4" />
@@ -541,16 +555,16 @@ const ProcessItem = ({
}`}
>
{isLoading ? (
<span className="animate-pulse">Loading...</span>
<span className="animate-pulse">{t('graph:processes.loading')}</span>
) : isSelected ? (
<>
<Eye className="h-3.5 w-3.5" />
Viewing
{t('graph:processes.viewing')}
</>
) : (
<>
<Eye className="h-3.5 w-3.5" />
View
{t('graph:processes.view')}
</>
)}
</button>
+32 -20
View File
@@ -10,31 +10,33 @@ import {
Table,
} from '@/lib/lucide-icons';
import { useAppState } from '../hooks/useAppState';
import { useTranslation } from 'react-i18next';
const EXAMPLE_QUERIES = [
{
label: 'All Functions',
labelKey: 'functions',
query: `MATCH (n:Function) RETURN n.id AS id, n.name AS name, n.filePath AS path LIMIT 50`,
},
{
label: 'All Classes',
labelKey: 'classes',
query: `MATCH (n:Class) RETURN n.id AS id, n.name AS name, n.filePath AS path LIMIT 50`,
},
{
label: 'All Interfaces',
labelKey: 'interfaces',
query: `MATCH (n:Interface) RETURN n.id AS id, n.name AS name, n.filePath AS path LIMIT 50`,
},
{
label: 'Function Calls',
labelKey: 'calls',
query: `MATCH (a:File)-[r:CodeRelation {type: 'CALLS'}]->(b:Function) RETURN a.id AS id, a.name AS caller, b.name AS callee LIMIT 50`,
},
{
label: 'Import Dependencies',
labelKey: 'imports',
query: `MATCH (a:File)-[r:CodeRelation {type: 'IMPORTS'}]->(b:File) RETURN a.id AS id, a.name AS from, b.name AS imports LIMIT 50`,
},
];
export const QueryFAB = () => {
const { t } = useTranslation(['common', 'graph']);
const {
setHighlightedNodeIds,
setQueryResult,
@@ -86,13 +88,13 @@ export const QueryFAB = () => {
if (!query.trim() || isRunning) return;
if (!graph) {
setError('No project loaded. Load a project first.');
setError(t('graph:queryFab.noProject'));
return;
}
const ready = await isDatabaseReady();
if (!ready) {
setError('Database not ready. Please wait for loading to complete.');
setError(t('graph:queryFab.dbNotReady'));
return;
}
@@ -147,13 +149,22 @@ export const QueryFAB = () => {
setQueryResult({ rows, nodeIds, executionTime });
setHighlightedNodeIds(new Set(nodeIds));
} catch (err) {
setError(err instanceof Error ? err.message : 'Query execution failed');
setError(err instanceof Error ? err.message : t('graph:queryFab.executionFailed'));
setQueryResult(null);
setHighlightedNodeIds(new Set());
} finally {
setIsRunning(false);
}
}, [query, isRunning, graph, isDatabaseReady, runQuery, setHighlightedNodeIds, setQueryResult]);
}, [
query,
isRunning,
graph,
isDatabaseReady,
runQuery,
setHighlightedNodeIds,
setQueryResult,
t,
]);
const handleKeyDown = (e: React.KeyboardEvent) => {
if (e.key === 'Enter' && (e.ctrlKey || e.metaKey)) {
@@ -189,7 +200,7 @@ export const QueryFAB = () => {
className="group absolute bottom-4 left-4 z-20 flex items-center gap-2 rounded-xl bg-gradient-to-r from-cyan-500 to-teal-500 px-4 py-2.5 text-sm font-medium text-white shadow-[0_0_20px_rgba(6,182,212,0.4)] transition-all duration-200 hover:-translate-y-0.5 hover:shadow-[0_0_30px_rgba(6,182,212,0.6)]"
>
<Terminal className="h-4 w-4" />
<span>Query</span>
<span>{t('graph:queryFab.query')}</span>
{queryResult && queryResult.nodeIds.length > 0 && (
<span className="ml-1 rounded-md bg-white/20 px-1.5 py-0.5 text-xs font-semibold">
{queryResult.nodeIds.length}
@@ -209,7 +220,7 @@ export const QueryFAB = () => {
<div className="flex h-7 w-7 items-center justify-center rounded-lg bg-gradient-to-br from-cyan-500 to-teal-500">
<Terminal className="h-4 w-4 text-white" />
</div>
<span className="text-sm font-medium">Cypher Query</span>
<span className="text-sm font-medium">{t('graph:queryFab.cypherQuery')}</span>
</div>
<button
onClick={handleClose}
@@ -239,7 +250,7 @@ export const QueryFAB = () => {
className="flex items-center gap-1.5 rounded-md px-3 py-1.5 text-xs text-text-secondary transition-colors hover:bg-hover hover:text-text-primary"
>
<Sparkles className="h-3.5 w-3.5" />
<span>Examples</span>
<span>{t('graph:queryFab.examples')}</span>
<ChevronDown
className={`h-3.5 w-3.5 transition-transform ${showExamples ? 'rotate-180' : ''}`}
/>
@@ -249,11 +260,11 @@ export const QueryFAB = () => {
<div className="absolute bottom-full left-0 mb-2 w-64 animate-fade-in rounded-lg border border-border-subtle bg-surface py-1 shadow-xl">
{EXAMPLE_QUERIES.map((example) => (
<button
key={example.label}
key={example.labelKey}
onClick={() => handleSelectExample(example.query)}
className="w-full px-3 py-2 text-left text-sm text-text-secondary transition-colors hover:bg-hover hover:text-text-primary"
>
{example.label}
{t(`graph:queryFab.exampleLabels.${example.labelKey}`)}
</button>
))}
</div>
@@ -266,7 +277,7 @@ export const QueryFAB = () => {
onClick={handleClear}
className="rounded-md px-3 py-1.5 text-xs text-text-secondary transition-colors hover:bg-hover hover:text-text-primary"
>
Clear
{t('graph:queryFab.clear')}
</button>
)}
<button
@@ -279,7 +290,7 @@ export const QueryFAB = () => {
) : (
<Play className="h-3.5 w-3.5" />
)}
<span>Run</span>
<span>{t('graph:queryFab.run')}</span>
<kbd className="ml-1 rounded bg-white/20 px-1 py-0.5 text-[10px]">⌘↵</kbd>
</button>
</div>
@@ -297,12 +308,13 @@ export const QueryFAB = () => {
<div className="flex items-center justify-between bg-cyan-500/5 px-4 py-2.5">
<div className="flex items-center gap-3 text-xs">
<span className="text-text-secondary">
<span className="font-semibold text-cyan-400">{queryResult.rows.length}</span> rows
<span className="font-semibold text-cyan-400">{queryResult.rows.length}</span>{' '}
{t('graph:queryFab.rows')}
</span>
{queryResult.nodeIds.length > 0 && (
<span className="text-text-secondary">
<span className="font-semibold text-cyan-400">{queryResult.nodeIds.length}</span>{' '}
highlighted
{t('graph:queryFab.highlighted')}
</span>
)}
<span className="text-text-muted">{queryResult.executionTime.toFixed(1)}ms</span>
@@ -313,7 +325,7 @@ export const QueryFAB = () => {
onClick={clearQueryHighlights}
className="text-xs text-text-muted transition-colors hover:text-text-primary"
>
Clear
{t('graph:queryFab.clear')}
</button>
)}
<button
@@ -362,7 +374,7 @@ export const QueryFAB = () => {
</table>
{queryResult.rows.length > 50 && (
<div className="border-t border-border-subtle bg-surface px-3 py-2 text-xs text-text-muted">
Showing 50 of {queryResult.rows.length} rows
{t('graph:queryFab.showingRows', { count: queryResult.rows.length })}
</div>
)}
</div>
+126 -23
View File
@@ -9,6 +9,7 @@
import { useState, useRef, useEffect, useId } from 'react';
import {
Github,
Gitlab,
FolderOpen,
Loader2,
Check,
@@ -23,23 +24,35 @@ import {
type JobProgress,
} from '../services/backend-client';
import { AnalyzeProgress } from './AnalyzeProgress';
import { useTranslation } from 'react-i18next';
// ── Helpers ──────────────────────────────────────────────────────────────────
type InputMode = 'github' | 'local';
type InputMode = 'github' | 'gitlab' | 'local';
const GITHUB_RE = /^https?:\/\/(www\.)?github\.com\/[^/\s]+\/[^/\s]+/i;
const GITLAB_RE = /^https?:\/\/[^/\s]+\/[^/\s]+\/[^/\s]+(\/.*)?$/i;
const IS_WINDOWS = navigator.userAgent.toLowerCase().includes('win');
function isValidGithubUrl(value: string): boolean {
return GITHUB_RE.test(value.trim());
}
function isValidGitlabUrl(value: string): boolean {
return GITLAB_RE.test(value.trim());
}
// ── Mode tabs ────────────────────────────────────────────────────────────────
function ModeTabs({ mode, onChange }: { mode: InputMode; onChange: (m: InputMode) => void }) {
const { t } = useTranslation('onboarding');
return (
<div className="flex gap-1 rounded-lg bg-elevated p-1" role="tablist" aria-label="Input type">
<div
className="flex gap-1 rounded-lg bg-elevated p-1"
role="tablist"
aria-label={t('repoAnalyzer.inputType')}
>
<button
role="tab"
aria-selected={mode === 'github'}
@@ -51,7 +64,20 @@ function ModeTabs({ mode, onChange }: { mode: InputMode; onChange: (m: InputMode
} `}
>
<Github className="h-3 w-3" />
GitHub URL
{t('repoAnalyzer.githubUrl')}
</button>
<button
role="tab"
aria-selected={mode === 'gitlab'}
onClick={() => onChange('gitlab')}
className={`flex flex-1 cursor-pointer items-center justify-center gap-1.5 rounded-md px-3 py-1.5 text-xs font-medium transition-all duration-150 ${
mode === 'gitlab'
? 'bg-accent text-white shadow-sm'
: 'text-text-muted hover:text-text-secondary'
} `}
>
<Gitlab className="h-3 w-3" />
{t('repoAnalyzer.gitlabUrl')}
</button>
<button
role="tab"
@@ -64,7 +90,7 @@ function ModeTabs({ mode, onChange }: { mode: InputMode; onChange: (m: InputMode
} `}
>
<FolderOpen className="h-3 w-3" />
Local Folder
{t('repoAnalyzer.localFolder')}
</button>
</div>
);
@@ -83,6 +109,7 @@ function AnalyzeButton({
onClick: () => void;
variant: 'onboarding' | 'sheet';
}) {
const { t } = useTranslation('onboarding');
const sizeClass =
variant === 'onboarding' ? 'w-full px-5 py-3.5 text-sm' : 'w-full px-4 py-3 text-sm';
return (
@@ -96,7 +123,7 @@ function AnalyzeButton({
} `}
>
{isLoading ? <Loader2 className="h-4 w-4 animate-spin" /> : <Sparkles className="h-4 w-4" />}
<span>{isLoading ? 'Starting analysis...' : 'Analyze Repository'}</span>
<span>{isLoading ? t('repoAnalyzer.starting') : t('repoAnalyzer.analyzeRepository')}</span>
{canSubmit && !isLoading && <ArrowRight className="h-3.5 w-3.5" />}
</button>
);
@@ -105,6 +132,8 @@ function AnalyzeButton({
// ── Done state ───────────────────────────────────────────────────────────────
function DoneState({ repoName }: { repoName: string }) {
const { t } = useTranslation('onboarding');
return (
<div
className="flex animate-fade-in flex-col items-center gap-3 py-4"
@@ -115,10 +144,10 @@ function DoneState({ repoName }: { repoName: string }) {
<Check className="h-6 w-6 text-emerald-400" />
</div>
<div className="text-center">
<p className="text-sm font-medium text-emerald-400">Analysis complete</p>
<p className="text-sm font-medium text-emerald-400">{t('repoAnalyzer.complete')}</p>
<p className="mt-0.5 font-mono text-xs text-text-muted">{repoName}</p>
</div>
<p className="text-xs text-text-secondary">Loading graph...</p>
<p className="text-xs text-text-secondary">{t('repoAnalyzer.loadingGraph')}</p>
</div>
);
}
@@ -134,17 +163,19 @@ export interface RepoAnalyzerProps {
}
export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProps) => {
const { t } = useTranslation(['common', 'errors', 'onboarding']);
const inputId = useId();
const folderInputRef = useRef<HTMLInputElement>(null);
const [mode, setMode] = useState<InputMode>('github');
const [githubUrl, setGithubUrl] = useState('');
const [gitlabUrl, setGitlabUrl] = useState('');
const [localPath, setLocalPath] = useState('');
const [phase, setPhase] = useState<InternalPhase>('input');
const [validationError, setValidationError] = useState<string | null>(null);
const [progress, setProgress] = useState<JobProgress>({
phase: 'queued',
percent: 0,
message: 'Queued',
message: t('common:analyzePhases.queued'),
});
const [completedRepoName, setCompletedRepoName] = useState('');
@@ -162,6 +193,7 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
const handleModeChange = (m: InputMode) => {
setMode(m);
setGithubUrl('');
setGitlabUrl('');
setLocalPath('');
setValidationError(null);
};
@@ -175,15 +207,21 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
const canSubmit =
mode === 'github'
? isValidGithubUrl(githubUrl) && (phase === 'input' || phase === 'error')
: localPath.trim().length > 1 && (phase === 'input' || phase === 'error');
: mode === 'gitlab'
? isValidGitlabUrl(gitlabUrl) && (phase === 'input' || phase === 'error')
: localPath.trim().length > 1 && (phase === 'input' || phase === 'error');
const handleAnalyze = async () => {
if (mode === 'github' && !isValidGithubUrl(githubUrl)) {
setValidationError('Please enter a valid GitHub repository URL.');
setValidationError(t('errors:invalidGithubUrl'));
return;
}
if (mode === 'gitlab' && !isValidGitlabUrl(gitlabUrl)) {
setValidationError('Please enter a valid GitLab repository URL.');
return;
}
if (mode === 'local' && localPath.trim().length < 2) {
setValidationError('Please enter a folder path.');
setValidationError(t('errors:missingFolderPath'));
return;
}
@@ -191,18 +229,30 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
setPhase('starting');
try {
const request = mode === 'github' ? { url: githubUrl.trim() } : { path: localPath.trim() };
const request =
mode === 'github'
? { url: githubUrl.trim() }
: mode === 'gitlab'
? { url: gitlabUrl.trim() }
: { path: localPath.trim() };
const { jobId } = await startAnalyze(request);
jobIdRef.current = jobId;
setPhase('analyzing');
const nameSource = mode === 'github' ? githubUrl.trim() : localPath.trim();
const nameSource =
mode === 'github'
? githubUrl.trim()
: mode === 'gitlab'
? gitlabUrl.trim()
: localPath.trim();
const controller = streamAnalyzeProgress(
jobId,
(p) => setProgress(p),
(data) => {
const name =
data.repoName ?? nameSource.split(/[/\\]/).filter(Boolean).at(-1) ?? 'repository';
data.repoName ??
nameSource.split(/[/\\]/).filter(Boolean).at(-1) ??
t('onboarding:repoAnalyzer.defaultRepoName');
setCompletedRepoName(name);
setPhase('done');
sseControllerRef.current = null;
@@ -212,13 +262,13 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
}, 1200);
},
(errMsg) => {
setValidationError(errMsg || 'Analysis failed. Check server logs.');
setValidationError(errMsg || t('errors:analysisFailed'));
setPhase('error');
},
);
sseControllerRef.current = controller;
} catch (err) {
setValidationError(err instanceof Error ? err.message : 'Failed to start analysis');
setValidationError(err instanceof Error ? err.message : t('errors:startAnalysisFailed'));
setPhase('error');
}
};
@@ -233,7 +283,7 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
jobIdRef.current = null;
}
setPhase('input');
setProgress({ phase: 'queued', percent: 0, message: 'Queued' });
setProgress({ phase: 'queued', percent: 0, message: t('common:analyzePhases.queued') });
};
const isLoading = phase === 'starting';
@@ -252,7 +302,7 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
htmlFor={inputId}
className="block text-xs font-medium tracking-wider text-text-secondary uppercase"
>
GitHub Repository URL
{t('onboarding:repoAnalyzer.githubRepositoryUrl')}
</label>
<div
className={`flex items-center gap-3 rounded-xl border bg-void px-4 py-3.5 transition-all duration-200 ${
@@ -297,6 +347,59 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
</div>
)}
{/* GitLab URL input */}
{showInput && mode === 'gitlab' && (
<div className="space-y-2">
<label
htmlFor={inputId}
className="block text-xs font-medium tracking-wider text-text-secondary uppercase"
>
{t('onboarding:repoAnalyzer.gitlabRepositoryUrl')}
</label>
<div
className={`flex items-center gap-3 rounded-xl border bg-void px-4 py-3.5 transition-all duration-200 ${
validationError && phase === 'error'
? 'border-red-500/50'
: isValidGitlabUrl(gitlabUrl)
? 'border-accent/50 shadow-[0_0_0_3px_rgba(124,58,237,0.08)]'
: 'border-border-default focus-within:border-accent/40'
} `}
>
<Gitlab className="h-4 w-4 shrink-0 text-text-muted" />
<input
id={inputId}
type="url"
value={gitlabUrl}
onChange={(e) => {
setGitlabUrl(e.target.value);
if (validationError) setValidationError(null);
}}
onKeyDown={(e) => {
if (e.key === 'Enter' && canSubmit && !isLoading) {
e.preventDefault();
handleAnalyze();
}
}}
disabled={isLoading}
placeholder="https://gitlab.com/owner/repo"
autoComplete="url"
spellCheck={false}
className="flex-1 border-none bg-transparent font-mono text-sm text-text-primary outline-none placeholder:text-text-muted disabled:opacity-50"
/>
{gitlabUrl.length > 10 && (
<div className="shrink-0">
{isValidGitlabUrl(gitlabUrl) ? (
<Check className="h-3.5 w-3.5 text-emerald-400" />
) : (
<AlertCircle className="h-3.5 w-3.5 text-text-muted" />
)}
</div>
)}
</div>
<p className="text-xs text-text-muted">{t('onboarding:repoAnalyzer.gitlabSupported')}</p>
</div>
)}
{/* Local folder input */}
{showInput && mode === 'local' && (
<div className="space-y-2">
@@ -304,7 +407,7 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
htmlFor={`${inputId}-local`}
className="block text-xs font-medium tracking-wider text-text-secondary uppercase"
>
Local Folder Path
{t('onboarding:repoAnalyzer.localFolderPath')}
</label>
<div
className={`flex items-center gap-3 rounded-xl border bg-void px-4 py-3.5 transition-all duration-200 ${
@@ -367,7 +470,7 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
className="flex w-full cursor-pointer items-center justify-center gap-2 rounded-lg border border-border-subtle bg-elevated px-3 py-2 text-xs font-medium text-text-secondary transition-all duration-150 hover:bg-hover hover:text-text-primary disabled:opacity-50"
>
<FolderOpen className="h-3.5 w-3.5" />
Browse for folder
{t('onboarding:repoAnalyzer.browseForFolder')}
</button>
</div>
)}
@@ -410,14 +513,14 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
}}
className="flex-1 cursor-pointer rounded-xl border border-border-subtle bg-elevated px-4 py-2.5 text-sm text-text-secondary transition-all duration-200 hover:bg-hover hover:text-text-primary"
>
Try again
{t('common:actions.tryAgain')}
</button>
{onCancel && (
<button
onClick={onCancel}
className="cursor-pointer px-4 py-2.5 text-sm text-text-muted transition-colors hover:text-text-secondary"
>
Dismiss
{t('common:actions.dismiss')}
</button>
)}
</div>
@@ -429,7 +532,7 @@ export const RepoAnalyzer = ({ variant, onComplete, onCancel }: RepoAnalyzerProp
onClick={onCancel}
className="w-full cursor-pointer py-1 text-xs text-text-muted transition-colors hover:text-text-secondary"
>
Hide (analysis continues in background)
{t('onboarding:repoAnalyzer.hideBackground')}
</button>
)}
</div>
+19 -14
View File
@@ -15,26 +15,29 @@
import { Sparkles, ArrowRight, GitBranch, FileCode, Layers } from '@/lib/lucide-icons';
import { RepoAnalyzer } from './RepoAnalyzer';
import type { BackendRepo } from '../services/backend-client';
import type { TFunction } from 'i18next';
import { useTranslation } from 'react-i18next';
// ── Helpers ──────────────────────────────────────────────────────────────────
function formatRelativeTime(dateStr: string): string {
function formatRelativeTime(dateStr: string, t: TFunction): string {
const date = new Date(dateStr);
const now = new Date();
const diffMs = now.getTime() - date.getTime();
const diffMins = Math.floor(diffMs / 60_000);
if (diffMins < 1) return 'just now';
if (diffMins < 60) return `${diffMins}m ago`;
if (diffMins < 1) return t('onboarding:landing.time.justNow');
if (diffMins < 60) return t('onboarding:landing.time.minutesAgo', { count: diffMins });
const diffHours = Math.floor(diffMins / 60);
if (diffHours < 24) return `${diffHours}h ago`;
if (diffHours < 24) return t('onboarding:landing.time.hoursAgo', { count: diffHours });
const diffDays = Math.floor(diffHours / 24);
if (diffDays < 30) return `${diffDays}d ago`;
if (diffDays < 30) return t('onboarding:landing.time.daysAgo', { count: diffDays });
return date.toLocaleDateString();
}
// ── Repo card ────────────────────────────────────────────────────────────────
function RepoCard({ repo, onClick }: { repo: BackendRepo; onClick: () => void }) {
const { t } = useTranslation(['common', 'onboarding']);
const stats = repo.stats;
return (
@@ -53,7 +56,7 @@ function RepoCard({ repo, onClick }: { repo: BackendRepo; onClick: () => void })
</div>
{repo.indexedAt && (
<p className="mt-1 pl-6 text-xs text-text-muted">
Indexed {formatRelativeTime(repo.indexedAt)}
{t('onboarding:landing.indexed', { time: formatRelativeTime(repo.indexedAt, t) })}
</p>
)}
</div>
@@ -64,17 +67,18 @@ function RepoCard({ repo, onClick }: { repo: BackendRepo; onClick: () => void })
<div className="mt-3 flex flex-wrap gap-2 pl-6">
{stats.files != null && (
<span className="inline-flex items-center gap-1 rounded-md bg-void px-2 py-0.5 text-[11px] text-text-muted">
<FileCode className="h-3 w-3" /> {stats.files.toLocaleString()} files
<FileCode className="h-3 w-3" /> {t('common:counts.files', { count: stats.files })}
</span>
)}
{stats.nodes != null && (
<span className="inline-flex items-center gap-1 rounded-md bg-void px-2 py-0.5 text-[11px] text-text-muted">
<Layers className="h-3 w-3" /> {stats.nodes.toLocaleString()} symbols
<Layers className="h-3 w-3" /> {t('common:counts.symbols', { count: stats.nodes })}
</span>
)}
{stats.processes != null && stats.processes > 0 && (
<span className="inline-flex items-center gap-1 rounded-md bg-void px-2 py-0.5 text-[11px] text-text-muted">
<Sparkles className="h-3 w-3" /> {stats.processes} flows
<Sparkles className="h-3 w-3" />{' '}
{t('common:counts.flows', { count: stats.processes })}
</span>
)}
</div>
@@ -92,6 +96,8 @@ interface RepoLandingProps {
}
export const RepoLanding = ({ repos, onSelectRepo, onAnalyzeComplete }: RepoLandingProps) => {
const { t } = useTranslation('onboarding');
return (
<div className="relative animate-fade-in overflow-hidden rounded-3xl border border-border-default bg-surface p-7">
{/* Ambient glows — mirrors OnboardingGuide aesthetic */}
@@ -109,10 +115,10 @@ export const RepoLanding = ({ repos, onSelectRepo, onAnalyzeComplete }: RepoLand
</div>
<h2 className="text-lg leading-snug font-semibold text-text-primary">
Choose a repository
{t('landing.chooseRepository')}
</h2>
<p className="mx-auto mt-1.5 max-w-xs text-sm leading-relaxed text-text-secondary">
Select an indexed repository to explore, or analyze a new one.
{t('landing.description')}
</p>
</div>
</div>
@@ -128,7 +134,7 @@ export const RepoLanding = ({ repos, onSelectRepo, onAnalyzeComplete }: RepoLand
<div className="mb-5 flex items-center gap-3">
<div className="h-px flex-1 bg-border-subtle" />
<span className="text-[11px] tracking-widest text-text-muted uppercase">
or analyze new
{t('landing.orAnalyzeNew')}
</span>
<div className="h-px flex-1 bg-border-subtle" />
</div>
@@ -140,8 +146,7 @@ export const RepoLanding = ({ repos, onSelectRepo, onAnalyzeComplete }: RepoLand
{/* Footer hint */}
<p className="mt-5 text-center text-[11px] leading-relaxed text-text-muted">
Public &amp; private repos &middot; Cloned locally by the server &middot; No data leaves
your machine
{t('landing.footer')}
</p>
</div>
);
+24 -23
View File
@@ -16,7 +16,9 @@ import { ToolCallCard } from './ToolCallCard';
import { isProviderConfigured } from '../core/llm/settings-service';
import { MarkdownRenderer } from './MarkdownRenderer';
import { ProcessesPanel } from './ProcessesPanel';
import { useTranslation } from 'react-i18next';
export const RightPanel = () => {
const { t } = useTranslation(['chat', 'common']);
const {
isRightPanelOpen,
setRightPanelOpen,
@@ -202,10 +204,10 @@ export const RightPanel = () => {
};
const chatSuggestions = [
'Explain the project architecture',
'What does this project do?',
'Show me the most important files',
'Find all API handlers',
t('chat:suggestions.architecture'),
t('chat:suggestions.whatDoes'),
t('chat:suggestions.importantFiles'),
t('chat:suggestions.apiHandlers'),
];
if (!isRightPanelOpen) return null;
@@ -225,7 +227,7 @@ export const RightPanel = () => {
}`}
>
<Sparkles className="h-3.5 w-3.5" />
<span>Nexus AI</span>
<span>{t('chat:tabs.chat')}</span>
</button>
{/* Processes Tab */}
@@ -238,9 +240,9 @@ export const RightPanel = () => {
}`}
>
<GitBranch className="h-3.5 w-3.5" />
<span>Processes</span>
<span>{t('chat:tabs.processes')}</span>
<span className="rounded-full bg-gradient-to-r from-violet-500 to-fuchsia-500 px-1.5 py-0.5 text-[10px] font-semibold text-white">
NEW
{t('chat:newBadge')}
</span>
</button>
</div>
@@ -249,7 +251,7 @@ export const RightPanel = () => {
<button
onClick={() => setRightPanelOpen(false)}
className="rounded p-1.5 text-text-muted transition-colors hover:bg-hover hover:text-text-primary"
title="Close Panel"
title={t('chat:actions.closePanel')}
>
<PanelRightClose className="h-4 w-4" />
</button>
@@ -270,12 +272,12 @@ export const RightPanel = () => {
<div className="ml-auto flex items-center gap-2">
{!isAgentReady && (
<span className="rounded-full border border-amber-500/30 bg-amber-500/15 px-2 py-1 text-[11px] text-amber-300">
Configure AI
{t('chat:badges.configureAI')}
</span>
)}
{isAgentInitializing && (
<span className="flex items-center gap-1 rounded-full border border-border-subtle bg-surface px-2 py-1 text-[11px] text-text-muted">
<Loader2 className="h-3 w-3 animate-spin" /> Connecting
<Loader2 className="h-3 w-3 animate-spin" /> {t('chat:badges.connecting')}
</span>
)}
</div>
@@ -296,10 +298,9 @@ export const RightPanel = () => {
<div className="mb-4 flex h-14 w-14 items-center justify-center rounded-xl bg-gradient-to-br from-accent to-node-interface text-2xl shadow-glow">
🧠
</div>
<h3 className="mb-2 text-base font-medium">Ask me anything</h3>
<h3 className="mb-2 text-base font-medium">{t('chat:empty.title')}</h3>
<p className="mb-5 text-sm leading-relaxed text-text-secondary">
I can help you understand the architecture, find functions, or explain
connections.
{t('chat:empty.description')}
</p>
<div className="flex flex-wrap justify-center gap-2">
{chatSuggestions.map((suggestion) => (
@@ -323,7 +324,7 @@ export const RightPanel = () => {
<div className="mb-2 flex items-center gap-2">
<User className="h-4 w-4 text-text-muted" />
<span className="text-xs font-medium tracking-wide text-text-muted uppercase">
You
{t('chat:roles.you')}
</span>
</div>
<div className="pl-6 text-sm text-text-primary">{message.content}</div>
@@ -336,7 +337,7 @@ export const RightPanel = () => {
<div className="mb-3 flex items-center gap-2">
<Sparkles className="h-4 w-4 text-accent" />
<span className="text-xs font-medium tracking-wide text-text-muted uppercase">
Nexus AI
{t('chat:roles.assistant')}
</span>
{isChatLoading && message === chatMessages[chatMessages.length - 1] && (
<Loader2 className="h-3 w-3 animate-spin text-accent" />
@@ -394,7 +395,7 @@ export const RightPanel = () => {
{/* Scroll to bottom */}
<button
aria-label="Scroll to bottom"
aria-label={t('chat:actions.scrollBottom')}
onClick={() => scrollToBottom()}
className={`absolute bottom-20 left-1/2 z-10 -translate-x-1/2 rounded-full border border-border-subtle bg-elevated px-3 py-1.5 text-xs text-text-secondary shadow-lg transition-all duration-200 hover:border-accent hover:text-accent ${
!isAtBottom && chatMessages.length > 0
@@ -403,7 +404,7 @@ export const RightPanel = () => {
}`}
>
<ArrowDown className="mr-1 inline h-3.5 w-3.5" />
Scroll to bottom
{t('chat:actions.scrollBottom')}
</button>
{/* Input */}
@@ -414,7 +415,7 @@ export const RightPanel = () => {
value={chatInput}
onChange={(e) => setChatInput(e.target.value)}
onKeyDown={handleKeyDown}
placeholder="Ask about the codebase..."
placeholder={t('chat:input.placeholder')}
rows={1}
className="scrollbar-thin min-h-[36px] flex-1 resize-none border-none bg-transparent text-sm text-text-primary outline-none placeholder:text-text-muted"
style={{ height: '36px', overflowY: 'hidden' }}
@@ -422,15 +423,15 @@ export const RightPanel = () => {
<button
onClick={clearChat}
className="px-2 py-1 text-xs text-text-muted transition-colors hover:text-text-primary"
title="Clear chat"
title={t('chat:actions.clearChat')}
>
Clear
{t('common:actions.clear')}
</button>
{isChatLoading ? (
<button
onClick={stopChatResponse}
className="flex h-9 w-9 items-center justify-center rounded-md bg-red-500/80 text-white transition-all hover:bg-red-500"
title="Stop response"
title={t('chat:actions.stopResponse')}
>
<Square className="h-3.5 w-3.5 fill-current" />
</button>
@@ -449,8 +450,8 @@ export const RightPanel = () => {
<AlertTriangle className="h-3.5 w-3.5" />
<span>
{isProviderConfigured()
? 'Initializing AI agent...'
: 'Configure an LLM provider to enable chat.'}
? t('chat:input.initializing')
: t('chat:input.configureProvider')}
</span>
</div>
)}
+139 -81
View File
@@ -23,6 +23,7 @@ import {
import type { LLMSettings, LLMProvider } from '../core/llm/types';
import { DEFAULT_OLLAMA_BASE_URL } from '../config/ui-constants';
import { ProviderConfigCard } from './settings/ProviderConfigCard';
import { useTranslation } from 'react-i18next';
interface SettingsPanelProps {
isOpen: boolean;
@@ -51,6 +52,7 @@ const OpenRouterModelCombobox = ({
isLoading,
onLoadModels,
}: OpenRouterModelComboboxProps) => {
const { t } = useTranslation('settings');
const [isOpen, setIsOpen] = useState(false);
const [searchTerm, setSearchTerm] = useState('');
const inputRef = useRef<HTMLInputElement>(null);
@@ -142,7 +144,7 @@ const OpenRouterModelCombobox = ({
value={searchTerm}
onChange={handleInputChange}
onKeyDown={handleKeyDown}
placeholder="Search or type model ID..."
placeholder={t('searchModelPlaceholder')}
className="flex-1 bg-transparent font-mono text-sm text-text-primary outline-none placeholder:text-text-muted"
onClick={(e) => e.stopPropagation()}
/>
@@ -150,7 +152,7 @@ const OpenRouterModelCombobox = ({
<span
className={`flex-1 truncate font-mono text-sm ${value ? 'text-text-primary' : 'text-text-muted'}`}
>
{displayValue || 'Select or type a model...'}
{displayValue || t('selectModelPlaceholder')}
</span>
)}
<div className="flex items-center gap-1">
@@ -167,20 +169,20 @@ const OpenRouterModelCombobox = ({
{isLoading ? (
<div className="flex items-center justify-center gap-2 px-4 py-6 text-center text-sm text-text-muted">
<Loader2 className="h-4 w-4 animate-spin" />
Loading models...
{t('loadingModels')}
</div>
) : filteredModels.length === 0 ? (
<div className="px-4 py-4 text-center">
{models.length === 0 ? (
<div className="text-sm text-text-muted">
<Search className="mx-auto mb-2 h-5 w-5 opacity-50" />
<p>Type a model ID or press Enter</p>
<p className="mt-1 text-xs">e.g. openai/gpt-4o</p>
<p>{t('customModelHint')}</p>
<p className="mt-1 text-xs">{t('customModelExample')}</p>
</div>
) : (
<div className="text-sm text-text-muted">
<p>No models match "{searchTerm}"</p>
<p className="mt-1 text-xs">Press Enter to use as custom ID</p>
<p>{t('noModelsMatch', { searchTerm })}</p>
<p className="mt-1 text-xs">{t('pressEnterCustom')}</p>
</div>
)}
</div>
@@ -198,7 +200,7 @@ const OpenRouterModelCombobox = ({
))}
{filteredModels.length > 50 && (
<div className="border-t border-border-subtle px-4 py-2 text-center text-xs text-text-muted">
+{filteredModels.length - 50} more • Refine your search
{t('moreModels', { count: filteredModels.length - 50 })}
</div>
)}
</div>
@@ -248,6 +250,7 @@ export const SettingsPanel = ({
isBackendConnected,
onBackendUrlChange,
}: SettingsPanelProps) => {
const { t } = useTranslation(['common', 'settings']);
const [settings, setSettings] = useState<LLMSettings>(loadSettings);
const [showApiKey, setShowApiKey] = useState<Record<string, boolean>>({});
const [saveStatus, setSaveStatus] = useState<'idle' | 'saved' | 'error'>('idle');
@@ -338,6 +341,7 @@ export const SettingsPanel = ({
'openrouter',
'minimax',
'glm',
'deepseek',
];
return (
@@ -354,8 +358,8 @@ export const SettingsPanel = ({
<Brain className="h-5 w-5 text-accent" />
</div>
<div>
<h2 className="text-lg font-semibold text-text-primary">AI Settings</h2>
<p className="text-xs text-text-muted">Configure your LLM provider</p>
<h2 className="text-lg font-semibold text-text-primary">{t('settings:title')}</h2>
<p className="text-xs text-text-muted">{t('settings:subtitle')}</p>
</div>
</div>
<button
@@ -371,16 +375,18 @@ export const SettingsPanel = ({
{/* Local Server */}
{backendUrl !== undefined && onBackendUrlChange && (
<div className="space-y-3">
<label className="block text-sm font-medium text-text-secondary">Local Server</label>
<label className="block text-sm font-medium text-text-secondary">
{t('settings:localServer')}
</label>
<div className="space-y-2">
<div className="mb-2 flex items-center gap-2">
<Server className="h-4 w-4 text-text-muted" />
<span className="text-sm text-text-secondary">Backend URL</span>
<span className="text-sm text-text-secondary">{t('settings:backendUrl')}</span>
<span
className={`h-2 w-2 rounded-full ${isBackendConnected ? 'bg-green-400' : 'bg-red-400'}`}
/>
<span className="text-xs text-text-muted">
{isBackendConnected ? 'Connected' : 'Not connected'}
{isBackendConnected ? t('settings:connected') : t('settings:notConnected')}
</span>
</div>
<input
@@ -390,17 +396,16 @@ export const SettingsPanel = ({
placeholder="http://localhost:4747"
className="w-full rounded-xl border border-border-subtle bg-elevated px-4 py-3 font-mono text-sm text-text-primary transition-all outline-none placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20"
/>
<p className="text-xs text-text-muted">
Run <code className="rounded bg-elevated px-1 py-0.5">gitnexus serve</code> to
start the local server
</p>
<p className="text-xs text-text-muted">{t('settings:runServeHint')}</p>
</div>
</div>
)}
{/* Provider Selection */}
<div className="space-y-3">
<label className="block text-sm font-medium text-text-secondary">Provider</label>
<label className="block text-sm font-medium text-text-secondary">
{t('settings:provider')}
</label>
<div className="grid grid-cols-1 gap-3 sm:grid-cols-3">
{providers.map((provider) => (
<button
@@ -429,7 +434,9 @@ export const SettingsPanel = ({
? '⚡'
: provider === 'glm'
? '🔮'
: '☁️'}
: provider === 'deepseek'
? '🐋'
: '☁️'}
</div>
<span className="font-medium">{getProviderDisplayName(provider)}</span>
</button>
@@ -438,7 +445,7 @@ export const SettingsPanel = ({
</div>
<div className="rounded-xl border border-amber-500/30 bg-amber-500/10 p-3 text-xs text-amber-200">
API keys are stored in session storage and will be cleared when you close this tab.
{t('settings:apiKeySession')}
</div>
{/* OpenAI Settings */}
@@ -447,10 +454,10 @@ export const SettingsPanel = ({
title="OpenAI"
apiKey={{
value: settings.openai?.apiKey ?? '',
placeholder: 'Enter your OpenAI API key',
helperText: 'Get your API key from',
placeholder: t('settings:providers.openai.apiKeyPlaceholder'),
helperText: t('settings:providers.openai.helperText'),
helperLink: 'https://platform.openai.com/api-keys',
helperLinkLabel: 'OpenAI Platform',
helperLinkLabel: t('settings:providers.openai.helperLinkLabel'),
isVisible: !!showApiKey['openai'],
onChange: (value) =>
setSettings((prev) => ({
@@ -461,7 +468,7 @@ export const SettingsPanel = ({
}}
model={{
value: settings.openai?.model ?? 'gpt-5.2-chat',
placeholder: 'e.g., gpt-4o, gpt-4-turbo, gpt-3.5-turbo',
placeholder: t('settings:providers.openai.modelPlaceholder'),
onChange: (value) =>
setSettings((prev) => ({
...prev,
@@ -472,7 +479,8 @@ export const SettingsPanel = ({
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
<Server className="h-4 w-4" />
Base URL <span className="font-normal text-text-muted">(optional)</span>
{t('settings:baseUrl')}{' '}
<span className="font-normal text-text-muted">({t('settings:optional')})</span>
</label>
<input
type="url"
@@ -483,12 +491,11 @@ export const SettingsPanel = ({
openai: { ...prev.openai!, baseUrl: e.target.value },
}))
}
placeholder="https://api.openai.com/v1 (default)"
placeholder={t('settings:providers.openai.baseUrlPlaceholder')}
className="w-full rounded-xl border border-border-subtle bg-elevated px-4 py-3 text-text-primary transition-all outline-none placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20"
/>
<p className="text-xs text-text-muted">
Leave empty to use the default OpenAI API. Set a custom URL for proxies or
compatible APIs.
{t('settings:providers.openai.baseUrlHint')}
</p>
</div>
</ProviderConfigCard>
@@ -500,10 +507,10 @@ export const SettingsPanel = ({
title="Google Gemini"
apiKey={{
value: settings.gemini?.apiKey ?? '',
placeholder: 'Enter your Google AI API key',
helperText: 'Get your API key from',
placeholder: t('settings:providers.gemini.apiKeyPlaceholder'),
helperText: t('settings:providers.gemini.helperText'),
helperLink: 'https://aistudio.google.com/app/apikey',
helperLinkLabel: 'Google AI Studio',
helperLinkLabel: t('settings:providers.gemini.helperLinkLabel'),
isVisible: !!showApiKey['gemini'],
onChange: (value) =>
setSettings((prev) => ({
@@ -514,7 +521,7 @@ export const SettingsPanel = ({
}}
model={{
value: settings.gemini?.model ?? 'gemini-2.0-flash',
placeholder: 'e.g., gemini-2.0-flash, gemini-1.5-pro',
placeholder: t('settings:providers.gemini.modelPlaceholder'),
onChange: (value) =>
setSettings((prev) => ({
...prev,
@@ -530,10 +537,10 @@ export const SettingsPanel = ({
title="Anthropic"
apiKey={{
value: settings.anthropic?.apiKey ?? '',
placeholder: 'Enter your Anthropic API key',
helperText: 'Get your API key from',
placeholder: t('settings:providers.anthropic.apiKeyPlaceholder'),
helperText: t('settings:providers.anthropic.helperText'),
helperLink: 'https://console.anthropic.com/settings/keys',
helperLinkLabel: 'Anthropic Console',
helperLinkLabel: t('settings:providers.anthropic.helperLinkLabel'),
isVisible: !!showApiKey['anthropic'],
onChange: (value) =>
setSettings((prev) => ({
@@ -544,7 +551,7 @@ export const SettingsPanel = ({
}}
model={{
value: settings.anthropic?.model ?? 'claude-sonnet-4-20250514',
placeholder: 'e.g., claude-sonnet-4-20250514, claude-3-opus',
placeholder: t('settings:providers.anthropic.modelPlaceholder'),
onChange: (value) =>
setSettings((prev) => ({
...prev,
@@ -560,7 +567,7 @@ export const SettingsPanel = ({
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
<Key className="h-4 w-4" />
API Key
{t('settings:apiKey')}
</label>
<div className="relative">
<input
@@ -572,7 +579,7 @@ export const SettingsPanel = ({
azureOpenAI: { ...prev.azureOpenAI!, apiKey: e.target.value },
}))
}
placeholder="Enter your Azure OpenAI API key"
placeholder={t('settings:providers.azure.apiKeyPlaceholder')}
className="w-full rounded-xl border border-border-subtle bg-elevated px-4 py-3 pr-12 text-text-primary transition-all outline-none placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20"
/>
<button
@@ -592,7 +599,7 @@ export const SettingsPanel = ({
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
<Server className="h-4 w-4" />
Endpoint
{t('settings:endpoint')}
</label>
<input
type="url"
@@ -609,7 +616,9 @@ export const SettingsPanel = ({
</div>
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">Deployment Name</label>
<label className="text-sm font-medium text-text-secondary">
{t('settings:deploymentName')}
</label>
<input
type="text"
value={settings.azureOpenAI?.deploymentName ?? ''}
@@ -619,14 +628,16 @@ export const SettingsPanel = ({
azureOpenAI: { ...prev.azureOpenAI!, deploymentName: e.target.value },
}))
}
placeholder="e.g., gpt-4o-deployment"
placeholder={t('settings:providers.azure.deploymentNamePlaceholder')}
className="w-full rounded-xl border border-border-subtle bg-elevated px-4 py-3 text-text-primary transition-all outline-none placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20"
/>
</div>
<div className="grid grid-cols-2 gap-4">
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">Model</label>
<label className="text-sm font-medium text-text-secondary">
{t('settings:model')}
</label>
<input
type="text"
value={settings.azureOpenAI?.model ?? 'gpt-4o'}
@@ -642,7 +653,9 @@ export const SettingsPanel = ({
</div>
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">API Version</label>
<label className="text-sm font-medium text-text-secondary">
{t('settings:apiVersion')}
</label>
<input
type="text"
value={settings.azureOpenAI?.apiVersion ?? '2024-08-01-preview'}
@@ -659,14 +672,14 @@ export const SettingsPanel = ({
</div>
<p className="text-xs text-text-muted">
Configure your Azure OpenAI service in the{' '}
{t('settings:azureHint')}{' '}
<a
href="https://portal.azure.com/#view/Microsoft_Azure_ProjectOxford/CognitiveServicesHub/~/OpenAI"
target="_blank"
rel="noopener noreferrer"
className="text-accent hover:underline"
>
Azure Portal
{t('settings:azurePortal')}
</a>
</p>
</div>
@@ -678,7 +691,8 @@ export const SettingsPanel = ({
{/* How to run Ollama */}
<div className="rounded-xl border border-amber-500/30 bg-amber-500/10 p-3">
<p className="text-xs leading-relaxed text-amber-300">
<span className="font-medium">📋 Quick Start:</span> Install Ollama from{' '}
<span className="font-medium">{t('settings:providers.ollama.quickStart')}</span>{' '}
{t('settings:providers.ollama.installFrom')}{' '}
<a
href="https://ollama.ai"
target="_blank"
@@ -687,7 +701,7 @@ export const SettingsPanel = ({
>
ollama.ai
</a>
, then run:
{t('settings:providers.ollama.thenRun')}
</p>
<code className="mt-2 block rounded-lg bg-black/30 px-3 py-2 font-mono text-sm text-amber-200">
ollama serve
@@ -697,7 +711,7 @@ export const SettingsPanel = ({
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
<Server className="h-4 w-4" />
Base URL
{t('settings:baseUrl')}
</label>
<div className="flex gap-2">
<input
@@ -719,18 +733,21 @@ export const SettingsPanel = ({
}
disabled={isCheckingOllama}
className="rounded-xl border border-border-subtle bg-elevated px-3 py-3 text-text-secondary transition-colors hover:border-accent/50 hover:text-text-primary disabled:opacity-50"
title="Check connection"
title={t('settings:checkConnection')}
>
<RefreshCw className={`h-4 w-4 ${isCheckingOllama ? 'animate-spin' : ''}`} />
</button>
</div>
<p className="text-xs text-text-muted">
Default port is <code className="rounded bg-elevated px-1 py-0.5">11434</code>.
{t('settings:defaultPort')}{' '}
<code className="rounded bg-elevated px-1 py-0.5">11434</code>.
</p>
</div>
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">Model</label>
<label className="text-sm font-medium text-text-secondary">
{t('settings:model')}
</label>
{ollamaError && !isCheckingOllama && (
<div className="rounded-lg border border-red-500/30 bg-red-500/10 p-2">
@@ -750,11 +767,11 @@ export const SettingsPanel = ({
ollama: { ...prev.ollama!, model: e.target.value },
}))
}
placeholder="e.g., llama3.2, mistral, codellama"
placeholder={t('settings:providers.ollama.modelPlaceholder')}
className="w-full rounded-xl border border-border-subtle bg-elevated px-4 py-3 font-mono text-sm text-text-primary transition-all outline-none placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20"
/>
<p className="text-xs text-text-muted">
Pull a model with{' '}
{t('settings:pullModel')}{' '}
<code className="rounded bg-elevated px-1 py-0.5">ollama pull llama3.2</code>
</p>
</div>
@@ -767,10 +784,10 @@ export const SettingsPanel = ({
title="OpenRouter"
apiKey={{
value: settings.openrouter?.apiKey ?? '',
placeholder: 'Enter your OpenRouter API key',
helperText: 'Get your API key from',
placeholder: t('settings:providers.openrouter.apiKeyPlaceholder'),
helperText: t('settings:providers.openrouter.helperText'),
helperLink: 'https://openrouter.ai/keys',
helperLinkLabel: 'OpenRouter Keys',
helperLinkLabel: t('settings:providers.openrouter.helperLinkLabel'),
isVisible: !!showApiKey['openrouter'],
onChange: (value) =>
setSettings((prev) => ({
@@ -781,7 +798,9 @@ export const SettingsPanel = ({
}}
>
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">Model</label>
<label className="text-sm font-medium text-text-secondary">
{t('settings:model')}
</label>
<OpenRouterModelCombobox
value={settings.openrouter?.model ?? ''}
onChange={(model) =>
@@ -795,14 +814,14 @@ export const SettingsPanel = ({
onLoadModels={loadOpenRouterModels}
/>
<p className="text-xs text-text-muted">
Browse all models at{' '}
{t('settings:browseModels')}{' '}
<a
href="https://openrouter.ai/models"
target="_blank"
rel="noopener noreferrer"
className="text-accent hover:underline"
>
OpenRouter Models
{t('settings:openRouterModels')}
</a>
</p>
</div>
@@ -815,10 +834,10 @@ export const SettingsPanel = ({
title="MiniMax"
apiKey={{
value: settings.minimax?.apiKey ?? '',
placeholder: 'Enter your MiniMax API key',
helperText: 'Get your API key from',
placeholder: t('settings:providers.minimax.apiKeyPlaceholder'),
helperText: t('settings:providers.minimax.helperText'),
helperLink: 'https://platform.minimax.io',
helperLinkLabel: 'MiniMax Platform',
helperLinkLabel: t('settings:providers.minimax.helperLinkLabel'),
isVisible: !!showApiKey['minimax'],
onChange: (value) =>
setSettings((prev) => ({
@@ -829,24 +848,61 @@ export const SettingsPanel = ({
}}
model={{
value: settings.minimax?.model ?? 'MiniMax-M2.5',
placeholder: 'e.g., MiniMax-M2.5, MiniMax-M2.5-highspeed',
placeholder: t('settings:providers.minimax.modelPlaceholder'),
onChange: (value) =>
setSettings((prev) => ({
...prev,
minimax: { ...prev.minimax!, model: value },
})),
helperText: 'Available: MiniMax-M2.5 (default), MiniMax-M2.5-highspeed (faster)',
helperText: t('settings:providers.minimax.helperModel'),
}}
/>
)}
{/* DeepSeek Settings */}
{settings.activeProvider === 'deepseek' && (
<ProviderConfigCard
title="DeepSeek"
apiKey={{
value: settings.deepseek?.apiKey ?? '',
placeholder: 'Enter your DeepSeek API key',
helperText: 'Get your API key from',
helperLink: 'https://platform.deepseek.com/api_keys',
helperLinkLabel: 'DeepSeek Platform',
isVisible: !!showApiKey['deepseek'],
onChange: (value) =>
setSettings((prev) => ({
...prev,
deepseek: { ...prev.deepseek!, apiKey: value },
})),
onToggleVisibility: () => toggleApiKeyVisibility('deepseek'),
}}
model={{
value: settings.deepseek?.model ?? 'deepseek-v4-flash',
placeholder: 'e.g., deepseek-v4-flash, deepseek-v4-pro, deepseek-chat',
onChange: (value) =>
setSettings((prev) => ({
...prev,
deepseek: { ...prev.deepseek!, model: value },
})),
helperText:
'deepseek-v4-flash (default), deepseek-v4-pro, deepseek-chat (V3), deepseek-reasoner (R1)',
}}
>
<p className="text-xs text-text-muted">
Compatible via OpenAI API format. The deepseek-reasoner model uses thinking mode and
requires round-tripping reasoning content.
</p>
</ProviderConfigCard>
)}
{/* GLM Settings */}
{settings.activeProvider === 'glm' && (
<div className="animate-fade-in space-y-4">
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
<Key className="h-4 w-4" />
API Key
{t('settings:apiKey')}
</label>
<div className="relative">
<input
@@ -858,7 +914,7 @@ export const SettingsPanel = ({
glm: { ...prev.glm!, apiKey: e.target.value },
}))
}
placeholder="Enter your Z.AI API key"
placeholder={t('settings:providers.glm.apiKeyPlaceholder')}
className="w-full rounded-xl border border-border-subtle bg-elevated px-4 py-3 pr-12 text-text-primary transition-all outline-none placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20"
/>
<button
@@ -874,20 +930,22 @@ export const SettingsPanel = ({
</button>
</div>
<p className="text-xs text-text-muted">
Get your API key from{' '}
{t('settings:providers.openai.helperText')}{' '}
<a
href="https://docs.z.ai"
target="_blank"
rel="noopener noreferrer"
className="text-accent hover:underline"
>
Z.AI Platform
{t('settings:zaiPlatform')}
</a>
</p>
</div>
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">Model</label>
<label className="text-sm font-medium text-text-secondary">
{t('settings:model')}
</label>
<select
value={settings.glm?.model ?? 'GLM-5'}
onChange={(e) =>
@@ -907,7 +965,9 @@ export const SettingsPanel = ({
</div>
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">Base URL</label>
<label className="text-sm font-medium text-text-secondary">
{t('settings:baseUrl')}
</label>
<input
type="text"
value={settings.glm?.baseUrl ?? 'https://api.z.ai/api/coding/paas/v4'}
@@ -920,9 +980,7 @@ export const SettingsPanel = ({
placeholder="https://api.z.ai/api/coding/paas/v4"
className="w-full rounded-xl border border-border-subtle bg-elevated px-4 py-3 font-mono text-sm text-text-primary transition-all outline-none placeholder:text-text-muted focus:border-accent focus:ring-2 focus:ring-accent/20"
/>
<p className="text-xs text-text-muted">
Coding API (default). Use https://api.z.ai/api/paas/v4 for the general API.
</p>
<p className="text-xs text-text-muted">{t('settings:glmCodingApi')}</p>
</div>
</div>
)}
@@ -934,10 +992,10 @@ export const SettingsPanel = ({
🔒
</div>
<div className="text-xs leading-relaxed text-text-muted">
<span className="font-medium text-text-secondary">Privacy:</span> Your API keys are
stored only in your browser's session storage and are cleared when the tab closes.
They're sent directly to the LLM provider when you chat. Your code never leaves your
machine.
<span className="font-medium text-text-secondary">
{t('settings:privacyLabel')}
</span>{' '}
{t('settings:privacyFull')}
</div>
</div>
</div>
@@ -949,13 +1007,13 @@ export const SettingsPanel = ({
{saveStatus === 'saved' && (
<span className="flex animate-fade-in items-center gap-1.5 text-green-400">
<Check className="h-4 w-4" />
Settings saved
{t('settings:settingsSaved')}
</span>
)}
{saveStatus === 'error' && (
<span className="flex animate-fade-in items-center gap-1.5 text-red-400">
<AlertCircle className="h-4 w-4" />
Failed to save
{t('settings:failedToSave')}
</span>
)}
</div>
@@ -964,13 +1022,13 @@ export const SettingsPanel = ({
onClick={onClose}
className="px-4 py-2 text-sm text-text-secondary transition-colors hover:text-text-primary"
>
Cancel
{t('common:actions.cancel')}
</button>
<button
onClick={handleSave}
className="rounded-lg bg-accent px-5 py-2 text-sm font-medium text-white transition-colors hover:bg-accent-dim"
>
Save Settings
{t('settings:saveSettings')}
</button>
</div>
</div>
+9 -6
View File
@@ -1,9 +1,12 @@
import { useMemo } from 'react';
import { Heart } from '@/lib/lucide-icons';
import { useAppState } from '../hooks/useAppState';
import { useTranslation } from 'react-i18next';
import { translateProgressMessage } from '../i18n/progress';
export const StatusBar = () => {
const { graph, progress } = useAppState();
const { t } = useTranslation(['common', 'graph']);
const nodeCount = graph?.nodes.length ?? 0;
const edgeCount = graph?.relationships.length ?? 0;
@@ -37,12 +40,12 @@ export const StatusBar = () => {
style={{ width: `${progress.percent}%` }}
/>
</div>
<span>{progress.message}</span>
<span>{translateProgressMessage(progress.message, t)}</span>
</>
) : (
<div className="flex items-center gap-1.5" data-testid="status-ready">
<span className="h-1.5 w-1.5 rounded-full bg-node-function" />
<span>Ready</span>
<span>{t('common:progress.ready')}</span>
</div>
)}
</div>
@@ -56,10 +59,10 @@ export const StatusBar = () => {
>
<Heart className="h-3.5 w-3.5 animate-pulse fill-pink-500/40 text-pink-500 transition-all duration-200 group-hover:scale-110 group-hover:fill-pink-500" />
<span className="text-[11px] font-medium text-pink-400 transition-colors group-hover:text-pink-300">
Sponsor
{t('graph:statusBar.sponsor')}
</span>
<span className="hidden text-[10px] text-pink-300/50 italic transition-colors group-hover:text-pink-300/80 md:inline">
need to buy some API credits to run SWE-bench 😅
{t('graph:statusBar.sponsorHint')}
</span>
</a>
@@ -67,9 +70,9 @@ export const StatusBar = () => {
<div className="flex items-center gap-3" data-testid="graph-stats">
{graph && (
<>
<span>{nodeCount} nodes</span>
<span>{t('common:counts.nodes', { count: nodeCount })}</span>
<span className="text-border-default">•</span>
<span>{edgeCount} edges</span>
<span>{t('common:counts.edges', { count: edgeCount })}</span>
{primaryLanguage && (
<>
<span className="text-border-default">•</span>
+20 -17
View File
@@ -15,6 +15,8 @@ import {
AlertCircle,
} from '@/lib/lucide-icons';
import type { ToolCallInfo } from '../core/llm/types';
import type { TFunction } from 'i18next';
import { useTranslation } from 'react-i18next';
interface ToolCallCardProps {
toolCall: ToolCallInfo;
@@ -25,7 +27,7 @@ interface ToolCallCardProps {
/**
* Format tool arguments for display
*/
const formatArgs = (args: Record<string, unknown>): string => {
const formatArgs = (args: Record<string, unknown>, t: TFunction): string => {
if (!args || Object.keys(args).length === 0) {
return '';
}
@@ -34,7 +36,7 @@ const formatArgs = (args: Record<string, unknown>): string => {
if ('cypher' in args && typeof args.cypher === 'string') {
let result = '';
if ('query' in args && typeof args.query === 'string') {
result += `Search: "${args.query}"\n\n`;
result += t('graph:toolCall.searchPrefix', { query: args.query }) + '\n\n';
}
result += args.cypher;
return result;
@@ -88,24 +90,25 @@ const getStatusDisplay = (status: ToolCallInfo['status']) => {
/**
* Get a friendly display name for the tool
*/
const getToolDisplayName = (name: string): string => {
const getToolDisplayName = (name: string, t: TFunction): string => {
const names: Record<string, string> = {
// Current 7-tool architecture
search: '🔍 Search Code',
cypher: '🔗 Cypher Query',
grep: '🔎 Pattern Search',
read: '📄 Read File',
overview: '🗺️ Codebase Overview',
explore: '🔬 Deep Dive',
impact: '💥 Impact Analysis',
search: t('graph:toolCall.tools.search'),
cypher: t('graph:toolCall.tools.cypher'),
grep: t('graph:toolCall.tools.grep'),
read: t('graph:toolCall.tools.read'),
overview: t('graph:toolCall.tools.overview'),
explore: t('graph:toolCall.tools.explore'),
impact: t('graph:toolCall.tools.impact'),
};
return names[name] || name;
};
export const ToolCallCard = ({ toolCall, defaultExpanded = false }: ToolCallCardProps) => {
const { t } = useTranslation(['common', 'graph']);
const [isExpanded, setIsExpanded] = useState(defaultExpanded);
const status = getStatusDisplay(toolCall.status);
const formattedArgs = formatArgs(toolCall.args);
const formattedArgs = formatArgs(toolCall.args, t);
return (
<div
@@ -131,13 +134,13 @@ export const ToolCallCard = ({ toolCall, defaultExpanded = false }: ToolCallCard
{/* Tool name */}
<span className="flex-1 text-sm font-medium text-text-primary">
{getToolDisplayName(toolCall.name)}
{getToolDisplayName(toolCall.name, t)}
</span>
{/* Status indicator */}
<span className={`flex items-center gap-1 text-xs ${status.color}`}>
{status.icon}
<span className="capitalize">{toolCall.status}</span>
<span className="capitalize">{t(`graph:toolCall.status.${toolCall.status}`)}</span>
</span>
</div>
@@ -148,7 +151,7 @@ export const ToolCallCard = ({ toolCall, defaultExpanded = false }: ToolCallCard
{formattedArgs && (
<div className="border-b border-border-subtle/50 px-3 py-2">
<div className="mb-1.5 text-[10px] tracking-wider text-text-muted uppercase">
{toolCall.name === 'cypher' ? 'Query' : 'Input'}
{toolCall.name === 'cypher' ? t('graph:toolCall.query') : t('graph:toolCall.input')}
</div>
<pre className="overflow-x-auto rounded bg-surface/50 p-2 font-mono text-xs whitespace-pre-wrap text-text-secondary">
{formattedArgs}
@@ -160,12 +163,12 @@ export const ToolCallCard = ({ toolCall, defaultExpanded = false }: ToolCallCard
{toolCall.result && (
<div className="px-3 py-2">
<div className="mb-1.5 text-[10px] tracking-wider text-text-muted uppercase">
Result
{t('graph:toolCall.result')}
</div>
<div className="max-h-[400px] overflow-y-auto rounded bg-surface/50">
<pre className="p-2 font-mono text-xs whitespace-pre-wrap text-text-secondary">
{toolCall.result.length > 3000
? toolCall.result.slice(0, 3000) + '\n\n... (truncated)'
? toolCall.result.slice(0, 3000) + '\n\n' + t('common:progress.truncated')
: toolCall.result}
</pre>
</div>
@@ -176,7 +179,7 @@ export const ToolCallCard = ({ toolCall, defaultExpanded = false }: ToolCallCard
{toolCall.status === 'running' && !toolCall.result && (
<div className="flex items-center gap-2 px-3 py-3 text-xs text-text-muted">
<Loader2 className="h-3 w-3 animate-spin" />
<span>Executing...</span>
<span>{t('common:progress.executing')}</span>
</div>
)}
</div>
@@ -1,5 +1,6 @@
import { useState, useEffect } from 'react';
import { X, Snail, Rocket, SkipForward } from '@/lib/lucide-icons';
import { useTranslation } from 'react-i18next';
interface WebGPUFallbackDialogProps {
isOpen: boolean;
@@ -20,6 +21,7 @@ export const WebGPUFallbackDialog = ({
onSkip,
nodeCount,
}: WebGPUFallbackDialogProps) => {
const { t } = useTranslation('graph');
const [isAnimating, setIsAnimating] = useState(true);
const [isVisible, setIsVisible] = useState(false);
@@ -69,10 +71,10 @@ export const WebGPUFallbackDialog = ({
🤔
</div>
<div>
<h2 className="text-lg font-semibold text-text-primary">WebGPU said "nope"</h2>
<p className="mt-0.5 text-sm text-text-muted">
Your browser doesn't support GPU acceleration
</p>
<h2 className="text-lg font-semibold text-text-primary">
{t('embedding.fallback.title')}
</h2>
<p className="mt-0.5 text-sm text-text-muted">{t('embedding.fallback.subtitle')}</p>
</div>
</div>
</div>
@@ -80,24 +82,31 @@ export const WebGPUFallbackDialog = ({
{/* Content */}
<div className="space-y-4 px-6 py-5">
<p className="text-sm leading-relaxed text-text-secondary">
Couldn't create embeddings with WebGPU, so semantic search (Graph RAG) won't be as
smart. The graph still works fine though!
{t('embedding.fallback.description')}
</p>
<div className="rounded-lg border border-border-subtle bg-elevated/50 p-4">
<p className="text-sm text-text-secondary">
<span className="font-medium text-text-primary">Your options:</span>
<span className="font-medium text-text-primary">
{t('embedding.fallback.options')}
</span>
</p>
<ul className="mt-2 space-y-1.5 text-sm text-text-muted">
<li className="flex items-start gap-2">
<Snail className="mt-0.5 h-4 w-4 flex-shrink-0 text-amber-400" />
<span>
<strong className="text-text-secondary">Use CPU</strong> — Works but{' '}
{isSmallCodebase ? 'a bit' : 'way'} slower
<strong className="text-text-secondary">{t('embedding.fallback.useCpu')}</strong>{' '}
—{' '}
{isSmallCodebase
? t('embedding.fallback.useCpuDescriptionSmall')
: t('embedding.fallback.useCpuDescriptionLarge')}
{nodeCount > 0 && (
<span className="text-text-muted">
{' '}
(~{estimatedMinutes} min for {nodeCount} nodes)
{t('embedding.fallback.estimated', {
minutes: estimatedMinutes,
count: nodeCount,
})}
</span>
)}
</span>
@@ -105,8 +114,8 @@ export const WebGPUFallbackDialog = ({
<li className="flex items-start gap-2">
<SkipForward className="mt-0.5 h-4 w-4 flex-shrink-0 text-blue-400" />
<span>
<strong className="text-text-secondary">Skip it</strong> — Graph works, just no AI
semantic search
<strong className="text-text-secondary">{t('embedding.fallback.skipIt')}</strong>{' '}
— {t('embedding.fallback.skipDescription')}
</span>
</li>
</ul>
@@ -115,11 +124,11 @@ export const WebGPUFallbackDialog = ({
{isSmallCodebase && (
<p className="flex items-center gap-1.5 rounded-lg bg-node-function/10 px-3 py-2 text-xs text-node-function">
<Rocket className="h-3.5 w-3.5" />
Small codebase detected! CPU should be fine.
{t('embedding.fallback.smallCodebase')}
</p>
)}
<p className="text-xs text-text-muted">💡 Tip: Try Chrome or Edge for WebGPU support</p>
<p className="text-xs text-text-muted">{t('embedding.fallback.tip')}</p>
</div>
{/* Actions */}
@@ -129,7 +138,7 @@ export const WebGPUFallbackDialog = ({
className="flex flex-1 items-center justify-center gap-2 rounded-lg border border-border-subtle bg-surface px-4 py-2.5 text-sm font-medium text-text-secondary transition-all hover:bg-hover hover:text-text-primary"
>
<SkipForward className="h-4 w-4" />
Skip Embeddings
{t('embedding.fallback.skipEmbeddings')}
</button>
<button
onClick={onUseCPU}
@@ -140,7 +149,9 @@ export const WebGPUFallbackDialog = ({
}`}
>
<Snail className="h-4 w-4" />
Use CPU {isSmallCodebase ? '(Recommended)' : '(Slow)'}
{isSmallCodebase
? t('embedding.fallback.useCpuRecommended')
: t('embedding.fallback.useCpuSlow')}
</button>
</div>
</div>
@@ -1,5 +1,6 @@
import { ReactNode } from 'react';
import { Eye, EyeOff, Key } from '@/lib/lucide-icons';
import { useTranslation } from 'react-i18next';
type ApiKeyField = {
value: string;
@@ -35,6 +36,8 @@ export const ProviderConfigCard = ({
model,
children,
}: ProviderConfigCardProps) => {
const { t } = useTranslation('settings');
return (
<div className="animate-fade-in space-y-4">
<div className="flex items-center justify-between">
@@ -48,7 +51,7 @@ export const ProviderConfigCard = ({
<div className="space-y-2">
<label className="flex items-center gap-2 text-sm font-medium text-text-secondary">
<Key className="h-4 w-4" />
API Key
{t('apiKey')}
</label>
<div className="relative">
<input
@@ -76,7 +79,7 @@ export const ProviderConfigCard = ({
rel="noopener noreferrer"
className="text-accent hover:underline"
>
{apiKey.helperLinkLabel ?? 'Learn more'}
{apiKey.helperLinkLabel ?? t('learnMore')}
</a>
) : null}
</p>
@@ -87,7 +90,7 @@ export const ProviderConfigCard = ({
{model && (
<div className="space-y-2">
<label className="text-sm font-medium text-text-secondary">
{model.label ?? 'Model'}
{model.label ?? t('model')}
</label>
<input
type="text"
+109 -13
View File
@@ -6,7 +6,13 @@
*/
import { createReactAgent } from '@langchain/langgraph/prebuilt';
import { SystemMessage } from '@langchain/core/messages';
import {
SystemMessage,
HumanMessage,
AIMessage,
ToolMessage,
type BaseMessage,
} from '@langchain/core/messages';
import { ChatOpenAI, AzureChatOpenAI } from '@langchain/openai';
import { ChatGoogleGenerativeAI } from '@langchain/google-genai';
import { ChatAnthropic } from '@langchain/anthropic';
@@ -23,10 +29,17 @@ import type {
OpenRouterConfig,
MiniMaxConfig,
GLMConfig,
DeepSeekConfig,
AgentStreamChunk,
AgentHistoryMessage,
} from './types';
import { type CodebaseContext, buildDynamicSystemPrompt } from './context-builder';
import { DEFAULT_OLLAMA_BASE_URL, DEFAULT_OPENROUTER_BASE_URL } from '../../config/ui-constants';
import {
DeepSeekChatOpenAI,
normalizeMessageContent,
normalizeToolCalls,
} from './deepseek-chat-model';
/**
* System prompt for the Graph RAG agent
@@ -124,6 +137,7 @@ When generating diagrams:
BAD: A[User's Data] --> B(Process & Save)
GOOD: A["User Data"] --> B["Process and Save"]
`;
export const createChatModel = (config: ProviderConfig): BaseChatModel => {
switch (config.provider) {
case 'openai': {
@@ -264,6 +278,26 @@ export const createChatModel = (config: ProviderConfig): BaseChatModel => {
});
}
case 'deepseek': {
const deepseekConfig = config as DeepSeekConfig;
if (!deepseekConfig.apiKey || deepseekConfig.apiKey.trim() === '') {
throw new Error('DeepSeek API key is required but was not provided');
}
return new DeepSeekChatOpenAI({
apiKey: deepseekConfig.apiKey,
modelName: deepseekConfig.model,
temperature: deepseekConfig.temperature ?? 0.1,
maxTokens: deepseekConfig.maxTokens,
configuration: {
apiKey: deepseekConfig.apiKey,
baseURL: 'https://api.deepseek.com',
},
streaming: true,
});
}
default:
throw new Error(`Unsupported provider: ${(config as any).provider}`);
}
@@ -324,11 +358,65 @@ export const createGraphRAGAgent = (
/**
* Message type for agent conversation
*/
export interface AgentMessage {
role: 'user' | 'assistant';
content: string;
export type AgentMessage = { role: 'user'; content: string } | AgentHistoryMessage;
export interface AgentRuntimeOptions {
/** Capture assistant/tool messages for providers that require exact transcript replay. */
captureHistory?: boolean;
}
export const buildLangChainMessages = (messages: AgentMessage[]): BaseMessage[] =>
messages.map((message) => {
if (message.role === 'user') {
return new HumanMessage(message.content);
}
if (message.role === 'tool') {
return new ToolMessage({
content: message.content,
tool_call_id: message.toolCallId,
...(message.name ? { name: message.name } : {}),
});
}
return new AIMessage({
content: message.content,
...(typeof message.reasoningContent === 'string'
? { additional_kwargs: { reasoning_content: message.reasoningContent } }
: {}),
...(message.toolCalls?.length ? { tool_calls: message.toolCalls } : {}),
} as any);
});
export const serializeAgentHistoryMessages = (
messages: unknown[],
startIndex = 0,
): AgentHistoryMessage[] => {
const serialized: AgentHistoryMessage[] = [];
for (const rawMessage of messages.slice(startIndex)) {
const msg: any = rawMessage;
const msgType = msg?._getType?.() || msg?.type || msg?.constructor?.name || 'unknown';
if (msgType === 'ai' || msgType === 'AIMessage') {
const reasoningContent = (msg.additional_kwargs || msg.kwargs)?.reasoning_content;
const toolCalls = normalizeToolCalls(msg.tool_calls);
serialized.push({
role: 'assistant',
content: normalizeMessageContent(msg.content),
...(toolCalls?.length && typeof reasoningContent === 'string' ? { reasoningContent } : {}),
...(toolCalls?.length ? { toolCalls } : {}),
});
continue;
}
if (msgType === 'tool' || msgType === 'ToolMessage') {
serialized.push({
role: 'tool',
content: normalizeMessageContent(msg.content),
toolCallId: String(msg.tool_call_id ?? ''),
...(typeof msg.name === 'string' ? { name: msg.name } : {}),
});
}
}
return serialized;
};
/**
* Stream a response from the agent
* Uses BOTH streamModes for best of both worlds:
@@ -340,12 +428,10 @@ export interface AgentMessage {
export async function* streamAgentResponse(
agent: ReturnType<typeof createReactAgent>,
messages: AgentMessage[],
options: AgentRuntimeOptions = {},
): AsyncGenerator<AgentStreamChunk> {
try {
const formattedMessages = messages.map((m) => ({
role: m.role,
content: m.content,
}));
const formattedMessages = buildLangChainMessages(messages);
// Use BOTH modes: 'values' for structure, 'messages' for token streaming
const stream = await agent.stream({ messages: formattedMessages }, {
@@ -364,6 +450,9 @@ export async function* streamAgentResponse(
// Anything before the first tool call should be treated as "reasoning/narration"
// so the UI can show the Cursor-like loop: plan → tool → update → tool → answer.
let hasSeenToolCallThisTurn = false;
// Track the last set of messages so we can persist the raw assistant/tool
// transcript for the next user turn.
let lastStepMessages: any[] | null = null;
for await (const event of stream) {
// Events come as [streamMode, data] tuples when using multiple modes
@@ -482,6 +571,9 @@ export async function* streamAgentResponse(
// Handle 'values' mode - state snapshots for structure
if (mode === 'values' && data?.messages) {
const stepMessages = data.messages || [];
if (options.captureHistory) {
lastStepMessages = stepMessages;
}
// Process new messages for tool calls/results we might have missed
for (let i = lastProcessedMsgCount; i < stepMessages.length; i++) {
@@ -539,7 +631,14 @@ export async function* streamAgentResponse(
if (import.meta.env.DEV) {
console.log('✅ Stream completed normally, yielding done');
}
yield { type: 'done' };
yield {
type: 'done',
historyMessages:
options.captureHistory && lastStepMessages
? serializeAgentHistoryMessages(lastStepMessages, formattedMessages.length)
: undefined,
};
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
// DEBUG: Stream error
@@ -561,10 +660,7 @@ export const invokeAgent = async (
agent: ReturnType<typeof createReactAgent>,
messages: AgentMessage[],
): Promise<string> => {
const formattedMessages = messages.map((m) => ({
role: m.role,
content: m.content,
}));
const formattedMessages = buildLangChainMessages(messages);
const result = await agent.invoke({ messages: formattedMessages });
@@ -0,0 +1,257 @@
import {
ChatOpenAI,
ChatOpenAICompletions,
type ChatOpenAICallOptions,
type ChatOpenAICompletionsCallOptions,
type ChatOpenAIFields,
} from '@langchain/openai';
import type { BaseMessage } from '@langchain/core/messages';
import type { BaseLanguageModelInput } from '@langchain/core/language_models/base';
import type { AIMessageChunk } from '@langchain/core/messages';
import type { Runnable } from '@langchain/core/runnables';
import type { CallbackManagerForLLMRun } from '@langchain/core/callbacks/manager';
import type { ChatGenerationChunk, ChatResult } from '@langchain/core/outputs';
import type { AgentToolCall } from './types';
/**
* DeepSeek's thinking-mode chat API requires assistant `reasoning_content`
* from prior turns to be replayed verbatim on the next request. LangChain
* preserves the inbound value on `AIMessage.additional_kwargs`, but its
* OpenAI-compatible outbound converter currently drops that provider-specific
* field. This completions subclass keeps the behavior scoped to DeepSeek by
* replacing only the serialized request messages immediately before the
* DeepSeek API call.
*/
export class DeepSeekChatOpenAICompletions<
CallOptions extends ChatOpenAICompletionsCallOptions = ChatOpenAICompletionsCallOptions,
> extends ChatOpenAICompletions<CallOptions> {
private activeMessages: BaseMessage[] | null = null;
private setActiveMessages(messages: BaseMessage[]): void {
if (this.activeMessages !== null) {
throw new Error('DeepSeekChatOpenAICompletions does not support overlapping requests');
}
this.activeMessages = messages;
}
override async _generate(
messages: BaseMessage[],
options: this['ParsedCallOptions'],
runManager?: CallbackManagerForLLMRun,
): Promise<ChatResult> {
this.setActiveMessages(messages);
try {
return await super._generate(messages, options, runManager);
} finally {
this.activeMessages = null;
}
}
override async *_streamResponseChunks(
messages: BaseMessage[],
options: this['ParsedCallOptions'],
runManager?: CallbackManagerForLLMRun,
): AsyncGenerator<ChatGenerationChunk> {
this.setActiveMessages(messages);
try {
yield* super._streamResponseChunks(messages, options, runManager);
} finally {
this.activeMessages = null;
}
}
override async completionWithRetry(request: any, requestOptions?: any): Promise<any> {
const messages = this.activeMessages
? buildDeepSeekRequestMessages(this.activeMessages)
: request.messages;
return super.completionWithRetry({ ...request, messages }, requestOptions);
}
}
/**
* OpenAI-compatible DeepSeek chat model with a DeepSeek-specific completions
* serializer. Keeping this as a subclass avoids provider checks in the shared
* agent streaming path and ensures LangChain `withConfig()` clones used by tool
* binding retain the same request serialization behavior.
*/
export class DeepSeekChatOpenAI<
CallOptions extends ChatOpenAICallOptions = ChatOpenAICallOptions,
> extends ChatOpenAI<CallOptions> {
private readonly deepSeekFields: ChatOpenAIFields;
constructor(fields: ChatOpenAIFields) {
const deepSeekFields = {
...fields,
completions: new DeepSeekChatOpenAICompletions(fields),
} as ChatOpenAIFields;
super(deepSeekFields);
this.deepSeekFields = deepSeekFields;
}
override withConfig(
config: Partial<CallOptions>,
): Runnable<BaseLanguageModelInput, AIMessageChunk, CallOptions> {
// Mirror ChatOpenAI.withConfig() for this LangChain version, but keep the
// DeepSeek subclass. Calling super.withConfig() would drop our custom
// completions serializer by returning a plain ChatOpenAI instance.
const newModel = new DeepSeekChatOpenAI<CallOptions>(this.deepSeekFields);
newModel.defaultOptions = {
...this.defaultOptions,
...config,
} as typeof this.defaultOptions;
return newModel;
}
}
export const normalizeMessageContent = (content: unknown): string => {
if (typeof content === 'string') return content;
if (Array.isArray(content)) {
return content
.filter((block: any) => block?.type === 'text' || typeof block === 'string')
.map((block: any) => (typeof block === 'string' ? block : block.text || ''))
.join('');
}
if (content == null) return '';
return String(content);
};
const normalizeToolCallArgs = (toolCall: any): Record<string, unknown> => {
if (toolCall?.args && typeof toolCall.args === 'object') {
return toolCall.args as Record<string, unknown>;
}
try {
return toolCall?.function?.arguments ? JSON.parse(toolCall.function.arguments) : {};
} catch {
return {};
}
};
export const normalizeToolCalls = (toolCalls: unknown): AgentToolCall[] | undefined => {
if (!Array.isArray(toolCalls) || toolCalls.length === 0) return undefined;
return toolCalls.map((toolCall: any) => ({
id: typeof toolCall?.id === 'string' ? toolCall.id : undefined,
name: toolCall?.name || toolCall?.function?.name || 'unknown',
args: normalizeToolCallArgs(toolCall),
type: typeof toolCall?.type === 'string' ? toolCall.type : 'tool_call',
}));
};
const stringifyToolArguments = (args: unknown): string => {
if (typeof args === 'string') return args;
try {
return JSON.stringify(args ?? {});
} catch {
return '{}';
}
};
const normalizeOpenAIContent = (content: unknown): string | Array<Record<string, unknown>> => {
if (typeof content === 'string') return content;
if (!Array.isArray(content)) return normalizeMessageContent(content);
const blocks = content.flatMap((block: any) => {
if (typeof block === 'string') {
return [{ type: 'text', text: block }];
}
if (block?.type === 'text' && typeof block.text === 'string') {
return [{ type: 'text', text: block.text }];
}
return [];
});
if (blocks.length === 0) return '';
if (blocks.length === 1) return blocks[0].text as string;
return blocks;
};
const getOpenAIRole = (message: any): string => {
const messageType =
message?._getType?.() || message?.type || message?.constructor?.name || 'unknown';
if ((message.additional_kwargs || {}).__openai_role__ === 'developer') {
return 'developer';
}
switch (messageType) {
case 'human':
case 'HumanMessage':
return 'user';
case 'ai':
case 'AIMessage':
return 'assistant';
case 'system':
case 'SystemMessage':
return 'system';
case 'tool':
case 'ToolMessage':
return 'tool';
case 'function':
case 'FunctionMessage':
return 'function';
default:
return typeof message.role === 'string' ? message.role : 'user';
}
};
export const buildDeepSeekRequestMessages = (
messages: Array<BaseMessage | Record<string, unknown>>,
): Array<Record<string, unknown>> =>
messages.map((message: any) => {
const role = getOpenAIRole(message);
const additionalKwargs =
message.additional_kwargs && typeof message.additional_kwargs === 'object'
? message.additional_kwargs
: {};
const requestMessage: Record<string, unknown> = {
role,
content: normalizeOpenAIContent(message.content),
};
if (typeof message.name === 'string' && message.name.length > 0) {
requestMessage.name = message.name;
}
if (role === 'assistant') {
const toolCalls = Array.isArray(message.tool_calls)
? message.tool_calls
: Array.isArray(additionalKwargs.tool_calls)
? additionalKwargs.tool_calls
: undefined;
if (toolCalls?.length) {
requestMessage.tool_calls = toolCalls.map((toolCall: any) => {
if (toolCall?.function) {
return {
id: toolCall.id,
type: toolCall.type ?? 'function',
function: {
name: toolCall.function.name,
arguments: stringifyToolArguments(toolCall.function.arguments),
},
};
}
return {
id: toolCall?.id,
type: 'function',
function: {
name: toolCall?.name ?? 'unknown',
arguments: stringifyToolArguments(toolCall?.args),
},
};
});
}
if (additionalKwargs.function_call != null) {
requestMessage.function_call = additionalKwargs.function_call;
}
if (toolCalls?.length && typeof additionalKwargs.reasoning_content === 'string') {
requestMessage.reasoning_content = additionalKwargs.reasoning_content;
}
return requestMessage;
}
if (role === 'tool' && typeof message.tool_call_id === 'string') {
requestMessage.tool_call_id = message.tool_call_id;
}
if (role === 'function' && typeof message.name === 'string') {
requestMessage.name = message.name;
}
return requestMessage;
});
+45 -1
View File
@@ -17,6 +17,7 @@ import {
OpenRouterConfig,
MiniMaxConfig,
GLMConfig,
DeepSeekConfig,
ProviderConfig,
} from './types';
import { DEFAULT_OPENROUTER_BASE_URL, DEFAULT_OLLAMA_BASE_URL } from '../../config/ui-constants';
@@ -59,6 +60,10 @@ const mergeWithDefaults = (parsed?: Partial<LLMSettings> | null): LLMSettings =>
...DEFAULT_LLM_SETTINGS.glm,
...parsed?.glm,
},
deepseek: {
...DEFAULT_LLM_SETTINGS.deepseek,
...parsed?.deepseek,
},
});
const readSettings = (storage: Storage): Partial<LLMSettings> | null => {
@@ -144,7 +149,9 @@ export const updateProviderSettings = <T extends LLMProvider>(
? Partial<Omit<MiniMaxConfig, 'provider'>>
: T extends 'glm'
? Partial<Omit<GLMConfig, 'provider'>>
: never
: T extends 'deepseek'
? Partial<Omit<DeepSeekConfig, 'provider'>>
: never
>,
): LLMSettings => {
const current = loadSettings();
@@ -239,6 +246,17 @@ export const updateProviderSettings = <T extends LLMProvider>(
saveSettings(updated);
return updated;
}
case 'deepseek': {
const updated: LLMSettings = {
...current,
deepseek: {
...(current.deepseek ?? {}),
...(updates as Partial<Omit<DeepSeekConfig, 'provider'>>),
},
};
saveSettings(updated);
return updated;
}
default: {
// Should be unreachable due to T extends LLMProvider, but keep a safe fallback
const updated: LLMSettings = { ...current };
@@ -316,6 +334,10 @@ const providerBuilders: Record<LLMProvider, ProviderBuilder> = {
maxTokens: settings.glm.maxTokens,
} as GLMConfig;
},
deepseek: (settings) => {
if (!settings.deepseek?.apiKey) return null;
return { provider: 'deepseek', ...settings.deepseek } as DeepSeekConfig;
},
};
export const getActiveProviderConfig = (): ProviderConfig | null => {
@@ -347,6 +369,24 @@ export const clearSettings = (): void => {
}
};
interface ProviderCapabilities {
/** Provider requires hidden assistant/tool transcript replay across turns. */
preserveAssistantTranscript: boolean;
}
const DEFAULT_PROVIDER_CAPABILITIES: ProviderCapabilities = {
preserveAssistantTranscript: false,
};
const PROVIDER_CAPABILITIES: Partial<Record<LLMProvider, ProviderCapabilities>> = {
deepseek: { preserveAssistantTranscript: true },
};
export const getProviderCapabilities = (provider: LLMProvider): ProviderCapabilities => ({
...DEFAULT_PROVIDER_CAPABILITIES,
...PROVIDER_CAPABILITIES[provider],
});
/**
* Get display name for a provider
*/
@@ -368,6 +408,8 @@ export const getProviderDisplayName = (provider: LLMProvider): string => {
return 'MiniMax';
case 'glm':
return 'GLM (Z.AI)';
case 'deepseek':
return 'DeepSeek';
default:
return provider;
}
@@ -398,6 +440,8 @@ export const getAvailableModels = (provider: LLMProvider): string[] => {
return ['MiniMax-M2.5', 'MiniMax-M2.5-highspeed'];
case 'glm':
return ['GLM-5', 'GLM-5-Turbo', 'GLM-4.7', 'GLM-4.5'];
case 'deepseek':
return ['deepseek-v4-flash', 'deepseek-v4-pro', 'deepseek-chat', 'deepseek-reasoner'];
default:
return [];
}
+52 -3
View File
@@ -2,7 +2,7 @@
* LLM Provider Types
*
* Type definitions for multi-provider LLM support.
* Supports OpenAI, Azure OpenAI, Gemini, Anthropic, Ollama, OpenRouter, MiniMax, and GLM5.
* Supports OpenAI, Azure OpenAI, Gemini, Anthropic, Ollama, OpenRouter, MiniMax, GLM, and DeepSeek.
*/
/**
@@ -17,7 +17,8 @@ export type LLMProvider =
| 'ollama'
| 'openrouter'
| 'minimax'
| 'glm';
| 'glm'
| 'deepseek';
/**
* Base configuration shared by all providers
@@ -106,6 +107,15 @@ export interface GLMConfig extends BaseProviderConfig {
baseUrl?: string; // defaults to https://api.z.ai/api/coding/paas/v4
}
/**
* DeepSeek configuration — OpenAI-compatible API
*/
export interface DeepSeekConfig extends BaseProviderConfig {
provider: 'deepseek';
apiKey: string;
model: string; // e.g., 'deepseek-v4-flash', 'deepseek-v4-pro'
}
/**
* Union type for all provider configurations
*/
@@ -117,7 +127,8 @@ export type ProviderConfig =
| OllamaConfig
| OpenRouterConfig
| MiniMaxConfig
| GLMConfig;
| GLMConfig
| DeepSeekConfig;
/**
* Stored settings (what goes to localStorage)
@@ -136,6 +147,7 @@ export interface LLMSettings {
openrouter?: Partial<Omit<OpenRouterConfig, 'provider'>>;
minimax?: Partial<Omit<MiniMaxConfig, 'provider'>>;
glm?: Partial<Omit<GLMConfig, 'provider'>>;
deepseek?: Partial<Omit<DeepSeekConfig, 'provider'>>;
// Intelligent Clustering Settings
intelligentClustering: boolean;
@@ -197,6 +209,11 @@ export const DEFAULT_LLM_SETTINGS: LLMSettings = {
baseUrl: 'https://api.z.ai/api/coding/paas/v4',
temperature: 0.1,
},
deepseek: {
apiKey: '',
model: 'deepseek-v4-flash',
temperature: 0.1,
},
};
/**
@@ -219,6 +236,8 @@ export interface ChatMessage {
id: string;
role: 'user' | 'assistant' | 'tool';
content: string;
/** Hidden raw transcript for reconstructing future agent turns */
historyMessages?: AgentHistoryMessage[];
/** @deprecated Use steps instead for proper ordering */
toolCalls?: ToolCallInfo[];
/** Ordered steps: reasoning, tool calls, and final content interleaved */
@@ -238,6 +257,34 @@ export interface ToolCallInfo {
status: 'pending' | 'running' | 'completed' | 'error';
}
/**
* Minimal tool-call payload needed to reconstruct prior assistant turns.
*/
export interface AgentToolCall {
id?: string;
name: string;
args: Record<string, unknown>;
type: 'tool_call';
}
/**
* Hidden per-turn transcript we keep so providers like DeepSeek can replay
* the original assistant/tool exchange on later user turns.
*/
export type AgentHistoryMessage =
| {
role: 'assistant';
content: string;
reasoningContent?: string;
toolCalls?: AgentToolCall[];
}
| {
role: 'tool';
content: string;
toolCallId: string;
name?: string;
};
/**
* Streaming chunk from agent
* Now supports step-based streaming where each step is a distinct message
@@ -248,6 +295,8 @@ export interface AgentStreamChunk {
reasoning?: string;
/** Final answer content (streamed token by token) */
content?: string;
/** Hidden raw transcript for reconstructing future agent turns */
historyMessages?: AgentHistoryMessage[];
/** Tool call information */
toolCall?: ToolCallInfo;
/** Error message */
+56 -18
View File
@@ -18,7 +18,12 @@ import type {
ToolCallInfo,
MessageStep,
} from '../core/llm/types';
import { loadSettings, getActiveProviderConfig, saveSettings } from '../core/llm/settings-service';
import {
loadSettings,
getActiveProviderConfig,
getProviderCapabilities,
saveSettings,
} from '../core/llm/settings-service';
import type { AgentMessage } from '../core/llm/agent';
import { type EdgeType } from '../lib/constants';
import {
@@ -35,10 +40,18 @@ import {
type JobProgress,
} from '../services/backend-client';
import { ERROR_RESET_DELAY_MS } from '../config/ui-constants';
import i18n from '../i18n';
import { normalizePath } from '../lib/path-resolution';
import { FILE_REF_REGEX, NODE_REF_REGEX } from '../lib/grounding-patterns';
import { GraphStateProvider, useGraphState } from './app-state/graph';
export const AUTO_START_EMBEDDINGS_STORAGE_KEY = 'gitnexus.autoStartEmbeddings';
export const shouldAutoStartEmbeddings = (): boolean => {
if (typeof window === 'undefined' || !window.localStorage) return false;
return window.localStorage.getItem(AUTO_START_EMBEDDINGS_STORAGE_KEY) === 'true';
};
export type ViewMode = 'onboarding' | 'loading' | 'exploring';
export type RightPanelTab = 'code' | 'chat';
export type EmbeddingStatus = 'idle' | 'loading' | 'embedding' | 'indexing' | 'ready' | 'error';
@@ -529,6 +542,10 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
setEmbeddingStatus('idle');
return;
}
if (!shouldAutoStartEmbeddings()) {
setEmbeddingStatus('idle');
return;
}
startEmbeddings().catch((err) => {
console.warn('Embeddings auto-start failed:', err);
});
@@ -623,6 +640,8 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
const sendChatMessage = useCallback(
async (message: string): Promise<void> => {
if (isChatLoading) return;
// Refresh Code panel for the new question: keep user-pinned refs, clear old AI citations
clearAICodeReferences();
// Also clear previous tool-driven AI highlights (highlight_in_graph)
@@ -649,7 +668,7 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
const assistantMessage: ChatMessage = {
id: `assistant-${Date.now()}`,
role: 'assistant',
content: 'Wait a moment, vector index is being created.',
content: i18n.t('common:chat.waitForVectorIndex'),
timestamp: Date.now(),
};
setChatMessages((prev) => [...prev, assistantMessage]);
@@ -662,11 +681,23 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
setIsChatLoading(true);
setCurrentToolCalls([]);
const providerCapabilities = getProviderCapabilities(llmSettings.activeProvider);
// Prepare message history for agent (convert our format to AgentMessage format)
const history: AgentMessage[] = [...chatMessages, userMessage].map((m) => ({
role: m.role === 'tool' ? 'assistant' : m.role,
content: m.content,
}));
const history: AgentMessage[] = [...chatMessages, userMessage].flatMap<AgentMessage>((m) => {
if (m.role === 'user') {
return [{ role: 'user', content: m.content }];
}
if (m.role === 'tool') {
return m.toolCallId
? [{ role: 'tool', content: m.content, toolCallId: m.toolCallId }]
: [];
}
if (providerCapabilities.preserveAssistantTranscript && m.historyMessages?.length) {
return m.historyMessages;
}
return [{ role: 'assistant', content: m.content }];
});
// Create placeholder for assistant response
const assistantMessageId = `assistant-${Date.now()}`;
@@ -675,6 +706,7 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
// Keep toolCalls for backwards compat and currentToolCalls state
const toolCallsForMessage: ToolCallInfo[] = [];
let stepCounter = 0;
let assistantHistoryMessages: ChatMessage['historyMessages'];
// Helper to update the message with current steps
const updateMessage = () => {
@@ -691,6 +723,7 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
id: assistantMessageId,
role: 'assistant' as const,
content,
historyMessages: assistantHistoryMessages,
steps: [...stepsForMessage],
toolCalls: [...toolCallsForMessage],
timestamp: existing?.timestamp ?? Date.now(),
@@ -973,6 +1006,9 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
break;
case 'done':
assistantHistoryMessages = providerCapabilities.preserveAssistantTranscript
? chunk.historyMessages
: undefined;
// Finalize the assistant message - just call updateMessage one more time
scheduleMessageUpdate();
break;
@@ -984,10 +1020,11 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
const agent = agentRef.current;
if (!agent) throw new Error('Agent not initialized');
const { streamAgentResponse } = await import('../core/llm/agent');
for await (const chunk of streamAgentResponse(agent, history)) {
for await (const chunk of streamAgentResponse(agent, history, {
captureHistory: providerCapabilities.preserveAssistantTranscript,
})) {
onChunk(chunk);
}
onChunk({ type: 'done' });
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
setAgentError(message);
@@ -1007,6 +1044,7 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
clearAIToolHighlights,
graph,
embeddingStatus,
isChatLoading,
],
);
@@ -1032,8 +1070,8 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
setProgress({
phase: 'extracting',
percent: 0,
message: 'Switching repository...',
detail: `Loading ${repoName}`,
message: i18n.t('common:progress.switchingRepository'),
detail: i18n.t('common:progress.loadingRepository', { repo: repoName }),
});
setViewMode('loading');
setIsAgentReady(false);
@@ -1061,8 +1099,8 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
setProgress({
phase: 'extracting',
percent: 5,
message: 'Switching repository...',
detail: 'Validating',
message: i18n.t('common:progress.switchingRepository'),
detail: i18n.t('common:progress.validating'),
});
} else if (phase === 'downloading') {
const pct = total ? Math.round((downloaded / total) * 90) + 5 : 50;
@@ -1070,15 +1108,15 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
setProgress({
phase: 'extracting',
percent: pct,
message: 'Downloading graph...',
detail: `${mb} MB downloaded`,
message: i18n.t('common:progress.downloadingGraph'),
detail: i18n.t('common:progress.downloadedMb', { mb }),
});
} else if (phase === 'extracting') {
setProgress({
phase: 'extracting',
percent: 97,
message: 'Processing...',
detail: 'Extracting file contents',
message: i18n.t('common:progress.processing'),
detail: i18n.t('common:progress.extractingFileContents'),
});
}
},
@@ -1110,8 +1148,8 @@ const AppStateProviderInner = ({ children }: { children: ReactNode }) => {
setProgress({
phase: 'error',
percent: 0,
message: 'Failed to switch repository',
detail: err instanceof Error ? err.message : 'Unknown error',
message: i18n.t('common:progress.failedSwitchRepository'),
detail: err instanceof Error ? err.message : i18n.t('common:progress.unknownError'),
});
setIsAgentReady(false);
agentRef.current = null;
+27
View File
@@ -0,0 +1,27 @@
import type { TFunction } from 'i18next';
import { BackendError } from '../services/backend-client';
export function formatBackendError(error: unknown, t: TFunction): string {
if (error instanceof BackendError) {
const seconds = error.retryAfterMs ? Math.ceil(error.retryAfterMs / 1000) : undefined;
const fallback = error.message || t('errors:unknown');
switch (error.code) {
case 'network':
return t('errors:backend.network', { defaultValue: fallback });
case 'timeout':
return t('errors:backend.timeout', { defaultValue: fallback });
case 'rate_limited':
return t('errors:backend.rateLimited', { seconds, defaultValue: fallback });
case 'not_found':
return t('errors:backend.notFound', { defaultValue: fallback });
case 'client':
return t('errors:backend.client', { message: error.message, defaultValue: fallback });
case 'server':
return t('errors:backend.server', { message: error.message, defaultValue: fallback });
default:
return fallback;
}
}
return error instanceof Error ? error.message : t('errors:unknown');
}
+70
View File
@@ -0,0 +1,70 @@
import i18n from 'i18next';
import LanguageDetector from 'i18next-browser-languagedetector';
import { initReactI18next } from 'react-i18next';
import {
DEFAULT_LANGUAGE,
SUPPORTED_LANGUAGE_CODES,
getLanguageMetadata,
normalizeSupportedLanguage,
} from './languages';
import { namespaceList, resources } from './resources';
const DEFAULT_NAMESPACE = 'common';
export const LANGUAGE_STORAGE_KEY = 'gitnexus.lng';
function syncDocumentLanguage(language: string | undefined): void {
if (typeof document === 'undefined') return;
const metadata = getLanguageMetadata(language);
document.documentElement.lang = metadata.code;
document.documentElement.dir = metadata.dir;
}
function convertDetectedLanguage(language: string): string {
return normalizeSupportedLanguage(language) ?? DEFAULT_LANGUAGE;
}
function persistSupportedLanguage(language: string | undefined): void {
const normalized = normalizeSupportedLanguage(language);
if (!normalized || typeof window === 'undefined') return;
try {
window.localStorage.setItem(LANGUAGE_STORAGE_KEY, normalized);
} catch {
// localStorage may be unavailable in restricted browser contexts.
}
}
export const i18nReady = i18n
.use(LanguageDetector)
.use(initReactI18next)
.init({
resources,
fallbackLng: DEFAULT_LANGUAGE,
supportedLngs: SUPPORTED_LANGUAGE_CODES,
load: 'currentOnly',
ns: namespaceList,
defaultNS: DEFAULT_NAMESPACE,
fallbackNS: false,
returnEmptyString: false,
interpolation: { escapeValue: false },
react: { useSuspense: false },
detection: {
order: ['querystring', 'localStorage', 'navigator', 'htmlTag'],
lookupQuerystring: 'lng',
lookupLocalStorage: LANGUAGE_STORAGE_KEY,
caches: [],
convertDetectedLanguage,
},
})
.then(() => {
const language = i18n.resolvedLanguage || i18n.language;
syncDocumentLanguage(language);
persistSupportedLanguage(language);
});
i18n.on('languageChanged', (language) => {
const resolvedLanguage = i18n.resolvedLanguage || language;
syncDocumentLanguage(resolvedLanguage);
persistSupportedLanguage(resolvedLanguage);
});
export default i18n;
+42
View File
@@ -0,0 +1,42 @@
export type SupportedLanguage = 'en' | 'zh-CN';
export interface LanguageMetadata {
code: SupportedLanguage;
nativeName: string;
englishName: string;
dir: 'ltr' | 'rtl';
}
export const DEFAULT_LANGUAGE: SupportedLanguage = 'en';
export const SUPPORTED_LANGUAGES: LanguageMetadata[] = [
{ code: 'en', nativeName: 'English', englishName: 'English', dir: 'ltr' },
{ code: 'zh-CN', nativeName: '简体中文', englishName: 'Simplified Chinese', dir: 'ltr' },
];
export const SUPPORTED_LANGUAGE_CODES = SUPPORTED_LANGUAGES.map((language) => language.code);
export function normalizeSupportedLanguage(
code: string | undefined | null,
): SupportedLanguage | null {
const normalized = code?.trim().split('.')[0]?.replace(/_/g, '-').toLowerCase();
if (!normalized) return null;
if (normalized === 'en' || normalized.startsWith('en-')) return 'en';
if (
normalized === 'zh' ||
normalized === 'zh-cn' ||
normalized.startsWith('zh-cn-') ||
normalized === 'zh-hans' ||
normalized.startsWith('zh-hans-')
) {
return 'zh-CN';
}
return null;
}
export function getLanguageMetadata(code: string | undefined): LanguageMetadata {
const normalized = normalizeSupportedLanguage(code);
return (
SUPPORTED_LANGUAGES.find((language) => language.code === normalized) ?? SUPPORTED_LANGUAGES[0]
);
}
+31
View File
@@ -0,0 +1,31 @@
import type { TFunction } from 'i18next';
export function translateAnalyzePhase(
phase: string,
message: string | undefined,
t: TFunction,
): string {
const key = `common:analyzePhases.${phase}`;
const translated = t(key, { defaultValue: '' });
return translated || message || phase;
}
export function translateProgressMessage(message: string | undefined, t: TFunction): string {
if (!message) return '';
const key = PROGRESS_MESSAGE_KEYS[message];
return key ? t(key) : message;
}
const PROGRESS_MESSAGE_KEYS: Record<string, string> = {
'Connecting...': 'common:progress.connectingShort',
'Connecting to server...': 'common:progress.connecting',
'Validating server': 'common:progress.validatingServer',
'Validating server...': 'common:progress.validatingServerEllipsis',
'Downloading graph...': 'common:progress.downloadingGraph',
'Extracting file contents': 'common:progress.extractingFileContents',
'Processing...': 'common:progress.processing',
'Processing graph...': 'common:progress.processingGraph',
'Loading graph...': 'common:progress.loadingGraph',
Queued: 'common:analyzePhases.queued',
'Starting...': 'common:progress.starting',
};
+20
View File
@@ -0,0 +1,20 @@
import type { Resource } from 'i18next';
const localeModules = import.meta.glob('../locales/*/*.json', {
eager: true,
import: 'default',
}) as Record<string, Record<string, unknown>>;
export const resources: Resource = {};
export const namespaces = new Set<string>();
for (const [path, translations] of Object.entries(localeModules)) {
const match = path.match(/\.\.\/locales\/([^/]+)\/([^/.]+)\.json$/);
if (!match) continue;
const [, language, namespace] = match;
resources[language] ??= {};
resources[language][namespace] = translations;
namespaces.add(namespace);
}
export const namespaceList = Array.from(namespaces).sort();
+61
View File
@@ -123,6 +123,67 @@ export {
* defaults to `currentColor`, so Tailwind `text-*` utilities work the same as
* with any other icon in this module.
*/
/**
* GitLab tanuki mark — SVG path data from simple-icons (CC0-1.0).
*
* GitLab's logo (the tanuki/fox-head) is a registered trademark of GitLab Inc.
* We use it here only to indicate GitLab source-repo integration.
*
* API-compatible with `lucide-react` icons (`LucideProps`).
*/
export const Gitlab = forwardRef<SVGSVGElement, LucideProps>(function Gitlab(
{
size = 24,
color = 'currentColor',
className,
strokeWidth: _strokeWidth,
absoluteStrokeWidth: _absoluteStrokeWidth,
...rest
},
ref,
) {
const numericSize = typeof size === 'string' ? Number.parseFloat(size) : size;
const useSmallVariant = Number.isFinite(numericSize) && (numericSize as number) <= 16;
if (useSmallVariant) {
return (
<svg
ref={ref}
xmlns="http://www.w3.org/2000/svg"
width={size}
height={size}
viewBox="0 0 16 16"
fill={color}
className={className}
{...rest}
>
<path d="M8 15.282l1.855-5.717H6.145L8 15.282z" />
<path d="M8 15.282L6.145 9.565H2.333L8 15.282z" />
<path d="M2.333 9.565l-.944-2.942c-.09-.267.067-.553.333-.553h3.153L2.333 9.565z" />
<path d="M4.875 6.07L6.145 9.565H2.333l2.542-3.495z" />
<path d="M13.667 9.565l.944-2.942c.09-.267-.067-.553-.333-.553h-3.153l2.542 3.495z" />
<path d="M11.125 6.07L9.855 9.565h3.812l-2.542-3.495z" />
<path d="M8 15.282l1.855-5.717H6.145L8 15.282z" />
</svg>
);
}
return (
<svg
ref={ref}
xmlns="http://www.w3.org/2000/svg"
width={size}
height={size}
viewBox="0 0 24 24"
fill={color}
className={className}
{...rest}
>
<path d="m23.6004 9.5927-.0337-.0862L20.3.9814a.851.851 0 0 0-.3362-.405.8748.8748 0 0 0-.9997.0539.8748.8748 0 0 0-.29.4399l-2.2055 6.748H7.5375l-2.2057-6.748a.8573.8573 0 0 0-.29-.4412.8748.8748 0 0 0-.9997-.0537.8585.8585 0 0 0-.3362.4049L.4332 9.5015l-.0325.0862a6.0657 6.0657 0 0 0 2.0119 7.0105l.0113.0087.03.0213 4.976 3.7264 2.462 1.8633 1.4995 1.1321a1.0085 1.0085 0 0 0 1.2197 0l1.4995-1.1321 2.4619-1.8633 5.006-3.7489.0125-.01a6.0682 6.0682 0 0 0 2.0094-7.003z" />
</svg>
);
});
export const Github = forwardRef<SVGSVGElement, LucideProps>(function Github(
{
size = 24,
+36
View File
@@ -0,0 +1,36 @@
{
"tabs": {
"chat": "Nexus AI",
"processes": "Processes"
},
"suggestions": {
"architecture": "Explain the project architecture",
"whatDoes": "What does this project do?",
"importantFiles": "Show me the most important files",
"apiHandlers": "Find all API handlers"
},
"empty": {
"title": "Ask me anything",
"description": "I can help you understand the architecture, find functions, or explain connections."
},
"input": {
"placeholder": "Ask about the codebase...",
"initializing": "Initializing AI agent...",
"configureProvider": "Configure an LLM provider to enable chat."
},
"actions": {
"closePanel": "Close Panel",
"scrollBottom": "Scroll to bottom",
"clearChat": "Clear chat",
"stopResponse": "Stop response"
},
"badges": {
"configureAI": "Configure AI",
"connecting": "Connecting"
},
"roles": {
"you": "You",
"assistant": "Nexus AI"
},
"newBadge": "NEW"
}
+85
View File
@@ -0,0 +1,85 @@
{
"app": {
"name": "GitNexus",
"nexusAI": "Nexus AI"
},
"actions": {
"cancel": "Cancel",
"dismiss": "Dismiss",
"tryAgain": "Try again",
"hide": "Hide",
"retry": "Retry",
"copy": "Copy",
"copied": "Copied",
"close": "Close",
"run": "Run",
"clear": "Clear",
"remove": "Remove",
"focusInGraph": "Focus in graph",
"expand": "Expand",
"collapse": "Collapse"
},
"chat": {
"viewNodeInCodePanel": "View {{inner}} in Code panel",
"openInCodePanel": "Open in Code panel • {{inner}}",
"waitForVectorIndex": "Wait a moment, vector index is being created."
},
"counts": {
"files_one": "{{count}} file",
"files_other": "{{count}} files",
"nodes_one": "{{count}} node",
"nodes_other": "{{count}} nodes",
"edges_one": "{{count}} edge",
"edges_other": "{{count}} edges",
"symbols_one": "{{count}} symbol",
"symbols_other": "{{count}} symbols",
"flows_one": "{{count}} flow",
"flows_other": "{{count}} flows"
},
"progress": {
"connecting": "Connecting to server...",
"connectingShort": "Connecting...",
"validatingServer": "Validating server",
"validatingServerEllipsis": "Validating server...",
"downloadingGraph": "Downloading graph...",
"downloadedMb": "{{mb}} MB downloaded",
"downloadingWithPercent": "Downloading graph... {{percent}}%",
"downloadingMb": "Downloading... {{mb}} MB",
"processing": "Processing...",
"processingGraph": "Processing graph...",
"extractingFileContents": "Extracting file contents",
"loadingGraph": "Loading graph...",
"starting": "Starting...",
"executing": "Executing...",
"truncated": "... (truncated)",
"ready": "Ready",
"switchingRepository": "Switching repository...",
"loadingRepository": "Loading {{repo}}",
"validating": "Validating",
"failedSwitchRepository": "Failed to switch repository",
"unknownError": "Unknown error"
},
"analyzePhases": {
"queued": "Queued",
"cloning": "Cloning repository",
"pulling": "Pulling latest",
"extracting": "Scanning files",
"structure": "Building structure",
"parsing": "Parsing code",
"imports": "Resolving imports",
"calls": "Tracing calls",
"heritage": "Extracting inheritance",
"communities": "Detecting communities",
"processes": "Detecting processes",
"complete": "Pipeline complete",
"lbug": "Loading into database",
"fts": "Creating search indexes",
"embeddings": "Generating embeddings",
"done": "Done",
"retrying": "Retrying after crash"
},
"units": {
"elapsedSeconds": "{{seconds}}s",
"elapsedMinutesSeconds": "{{minutes}}m {{seconds}}s"
}
}
+19
View File
@@ -0,0 +1,19 @@
{
"unknown": "Unknown error",
"connectFailed": "Failed to connect to server",
"loadGraphFailed": "Failed to load graph",
"failedToConnect": "Failed to connect",
"analysisFailed": "Analysis failed. Check server logs.",
"startAnalysisFailed": "Failed to start analysis",
"invalidGithubUrl": "Please enter a valid GitHub repository URL.",
"missingFolderPath": "Please enter a folder path.",
"backend": {
"reconnecting": "Server connection lost. Reconnecting…",
"network": "Unable to reach the GitNexus server. Make sure `gitnexus serve` is running.",
"timeout": "The server took too long to respond. Try again in a moment.",
"rateLimited": "Too many requests. Try again in {{seconds}}s.",
"notFound": "The requested repository or resource was not found.",
"client": "Request failed: {{message}}",
"server": "Server error: {{message}}"
}
}
+174
View File
@@ -0,0 +1,174 @@
{
"statusBar": {
"sponsor": "Sponsor",
"sponsorHint": "need to buy some API credits to run SWE-bench 😅"
},
"loading": {
"filesProgress": "{{processed}} / {{total}} files"
},
"toolCall": {
"status": {
"running": "running",
"completed": "completed",
"error": "error"
},
"tools": {
"search": "🔍 Search Code",
"cypher": "🔗 Cypher Query",
"grep": "🔎 Pattern Search",
"read": "📄 Read File",
"overview": "🗺️ Codebase Overview",
"explore": "🔬 Deep Dive",
"impact": "💥 Impact Analysis"
},
"query": "Query",
"input": "Input",
"result": "Result",
"searchPrefix": "Search: \"{{query}}\""
},
"embedding": {
"generateTitle": "Generate embeddings for semantic search",
"enable": "Enable Semantic Search",
"loadingModel": "Loading AI model...",
"embeddingNodes": "Embedding {{processed}}/{{total}} nodes",
"creatingIndex": "Creating vector index...",
"readyTitle": "Semantic search is ready! Use natural language in the AI chat.",
"ready": "Semantic Ready",
"errorTitle": "Embedding failed. Click to retry.",
"failedRetry": "Failed - Retry",
"fallback": {
"title": "WebGPU said \"nope\"",
"subtitle": "Your browser doesn't support GPU acceleration",
"description": "Couldn't create embeddings with WebGPU, so semantic search (Graph RAG) won't be as smart. The graph still works fine though!",
"options": "Your options:",
"useCpu": "Use CPU",
"useCpuDescriptionSmall": "Works but a bit slower",
"useCpuDescriptionLarge": "Works but way slower",
"estimated": "(~{{minutes}} min for {{count}} nodes)",
"skipIt": "Skip it",
"skipDescription": "Graph works, just no AI semantic search",
"smallCodebase": "Small codebase detected! CPU should be fine.",
"tip": "💡 Tip: Try Chrome or Edge for WebGPU support",
"skipEmbeddings": "Skip Embeddings",
"useCpuRecommended": "Use CPU (Recommended)",
"useCpuSlow": "Use CPU (Slow)"
}
},
"queryFab": {
"query": "Query",
"cypherQuery": "Cypher Query",
"examples": "Examples",
"run": "Run",
"noProject": "No project loaded. Load a project first.",
"dbNotReady": "Database not ready. Please wait for loading to complete.",
"executionFailed": "Query execution failed",
"exampleLabels": {
"functions": "All Functions",
"classes": "All Classes",
"interfaces": "All Interfaces",
"calls": "Function Calls",
"imports": "Import Dependencies"
},
"clear": "Clear",
"rows": "rows",
"highlighted": "highlighted",
"showingRows": "Showing 50 of {{count}} rows"
},
"fileTree": {
"expandPanel": "Expand Panel",
"fileExplorer": "File Explorer",
"filters": "Filters",
"collapsePanel": "Collapse Panel",
"searchFiles": "Search files...",
"noFilesLoaded": "No files loaded",
"all": "All",
"selectNodeDepth": "Select a node to apply depth filter",
"explorer": "Explorer",
"nodeTypes": "Node Types",
"nodeTypesDesc": "Toggle visibility of node types in the graph",
"edgeTypes": "Edge Types",
"edgeTypesDesc": "Toggle visibility of relationship types",
"focusDepth": "Focus Depth",
"focusDepthDesc": "Show nodes within N hops of selection",
"hops_one": "{{count}} hop",
"hops_other": "{{count}} hops",
"colorLegend": "Color Legend"
},
"codePanel": {
"expand": "Expand Code Panel",
"dragResize": "Drag to resize",
"title": "Code Inspector",
"clearCitations": "Clear AI citations",
"clearSelection": "Clear selection",
"loadingSource": "Loading source...",
"selectFile": "Select a file node to preview its contents.",
"code": "Code",
"selected": "Selected",
"aiCitations": "AI Citations",
"references_one": "{{count}} reference",
"references_other": "{{count}} references",
"lines_one": "{{count}} line",
"lines_other": "{{count}} lines",
"codeNotAvailable": "Code not available in memory for {{path}}"
},
"canvas": {
"zoomIn": "Zoom In",
"zoomOut": "Zoom Out",
"fit": "Fit to Screen",
"focusSelected": "Focus on Selected Node",
"clearSelection": "Clear Selection",
"clear": "Clear",
"stopLayout": "Stop Layout",
"runLayout": "Run Layout Again",
"layoutOptimizing": "Layout optimizing...",
"turnOffHighlights": "Turn off all highlights",
"turnOnHighlights": "Turn on AI highlights"
},
"processes": {
"unknownStep": "Unknown",
"allProcessesLabel_one": "All Processes ({{count}} combined)",
"allProcessesLabel_other": "All Processes ({{count}} combined)",
"emptyTitle": "No Processes Detected",
"emptyDescription": "Processes are execution flows traced from entry points. Load a codebase to see detected processes.",
"filterPlaceholder": "Filter processes...",
"detected_one": "{{count}} process detected",
"detected_other": "{{count}} processes detected",
"fullMap": "Full Process Map",
"viewCombined_one": "View combined map of {{count}} process",
"viewCombined_other": "View combined map of {{count}} processes",
"crossCommunity": "Cross-Community",
"intraCommunity": "Intra-Community",
"steps_one": "{{count}} step",
"steps_other": "{{count}} steps",
"clusters_one": "{{count}} cluster",
"clusters_other": "{{count}} clusters",
"highlightTitle": "Click to highlight in graph",
"removeHighlightTitle": "Click to remove highlight from graph",
"loading": "Loading...",
"viewing": "Viewing",
"view": "View"
},
"processFlow": {
"title": "Process: {{label}}",
"diagramTooLarge": "📊 Diagram Too Large",
"renderError": "⚠️ Render Error",
"tooComplex_one": "This diagram has {{count}} step and is too complex to render. Try viewing individual processes instead of \"All Processes\".",
"tooComplex_other": "This diagram has {{count}} steps and is too complex to render. Try viewing individual processes instead of \"All Processes\".",
"unableToRender_one": "Unable to render diagram. Steps: {{count}}",
"unableToRender_other": "Unable to render diagram. Steps: {{count}}",
"zoomOutTitle": "Zoom out (-)",
"zoomInTitle": "Zoom in (+)",
"resetTitle": "Reset zoom and pan",
"resetView": "Reset View",
"toggleFocus": "Toggle Focus",
"copyMermaid": "Copy Mermaid"
},
"diagram": {
"aiGenerated": "AI Generated Diagram",
"error": "Diagram Error",
"showSource": "Show source",
"label": "Diagram",
"expandTitle": "Expand",
"loading": "Loading diagram…"
}
}
+16
View File
@@ -0,0 +1,16 @@
{
"repositories": "Repositories",
"active": "active",
"reanalyzing": "Re-analyzing...",
"reanalyzeRepo": "Re-analyze {{repoName}}",
"deleteRepo": "Delete {{repoName}}",
"reanalyzingRepo": "Re-analyzing {{repoName}}: {{message}}",
"analyzeNew": "Analyze a new repository...",
"searchNodes": "Search nodes...",
"noNodesFound": "No nodes found for \"{{query}}\"",
"starIfCool": "Star if cool",
"aiSettings": "AI Settings",
"help": "Help",
"language": "Language",
"selectLanguage": "Select language"
}
+96
View File
@@ -0,0 +1,96 @@
{
"tabs": {
"overview": "Overview",
"ai": "Nexus AI",
"shortcuts": "Shortcuts",
"status": "Status bar",
"graph": "Graph & nodes",
"search": "Search & filter"
},
"shortcuts": {
"searchNodes": "Search nodes",
"deselectClose": "Deselect / close",
"columns": {
"action": "Action",
"mac": "Mac",
"windows": "Windows"
}
},
"nodeTypes": {
"function": "Function",
"functionDesc": "Function declarations",
"file": "File",
"fileDesc": "Source files",
"class": "Class",
"classDesc": "Class declarations",
"method": "Method",
"methodDesc": "Class methods",
"interface": "Interface",
"interfaceDesc": "TypeScript interfaces",
"folder": "Folder",
"folderDesc": "Directory nodes"
},
"status": {
"ready": "Ready",
"readyDesc": "Graph is fully loaded and interactive",
"nodesCount": "Nodes count",
"nodesCountDesc": "Total files and symbols in the graph",
"edgesCount": "Edges count",
"edgesCountDesc": "Import / dependency connections",
"aiIndexStatus": "AI index status",
"aiIndexStatusDesc": "Repo is fully indexed for AI queries",
"semanticReadyBadge": "Semantic Ready",
"explained": "Status bar explained"
},
"tryAsking": "Try asking:",
"footer": "GitNexus — graph explorer",
"title": "Help & Reference",
"footerLong": "GitNexus — open source codebase graph explorer",
"docsGithub": "Docs & GitHub ↗",
"overview": {
"gettingStarted": "Getting started",
"whatIsTitle": "What is GitNexus?",
"whatIsDescription": "An interactive graph explorer for your codebase. Every file, function, and import becomes a node you can explore, query, and navigate visually.",
"currentRepoTitle": "Your current repo",
"loadedCounts": "Loaded: {{nodeCount}} nodes · {{edgeCount}} edges",
"threeWaysTitle": "Three ways to explore",
"wayInspect": "Click nodes to inspect",
"waySearch": "Search by name or type",
"wayAsk": "Ask Nexus AI a natural language question",
"navigationTitle": "Navigation",
"navZoom": "Scroll to zoom",
"navPan": "Click and drag to pan",
"navFocus": "Double-click a node to focus its subgraph"
},
"graph": {
"nodeColorLegend": "Node color legend",
"nodeLabel": "{{label}} nodes",
"sizeDescription": "Node size reflects connection count — larger nodes are depended on by more files. Edges point from importer → imported.",
"detailDescription": "Click any node to open its detail panel — showing imports, exports, and reverse dependencies."
},
"search": {
"title": "Search & filter",
"searchNodes": "Search nodes",
"searchDescription": "Search by filename, function name, or import path. Matching nodes are highlighted live in the graph.",
"filterPanel": "Filter panel",
"filterDescription": "Use the filter icon in the left sidebar to isolate specific node types, hide leaf nodes, or focus on a depth range from a selected root.",
"syntax": "Search syntax",
"hints": {
"nameFragment": "match by name fragment",
"pathPrefix": "match by path prefix",
"nodeType": "filter by node type"
}
},
"ai": {
"title": "Nexus AI",
"semanticReady": "✓ Semantic Ready",
"description": "Your repo is indexed and ready for semantic queries. Nexus AI understands code structure and relationships, not just file names.",
"questions": {
"dependencies": "\"Which files depend on the auth module?\"",
"circular": "\"Find circular dependencies in this repo\"",
"connected": "\"What are the most connected components?\"",
"imports": "\"Show me all files that import useEffect\""
},
"openPrompt": "Open the prompt via the Nexus AI button (top-right)."
}
}
@@ -0,0 +1,67 @@
{
"success": {
"title": "Server Connected",
"description": "Preparing your code knowledge graph..."
},
"loading": {
"largeRepoHint": "This may take a moment for large repositories"
},
"guide": {
"copyAria": "Copy to clipboard",
"copiedAria": "Copied!",
"startServer": "Start your local server",
"devDescription": "Fire up the Express backend in a separate terminal to unlock the full graph.",
"prodDescription": "One command is all it takes. The browser connects automatically.",
"copyCommand": "Copy the command",
"copyCommandDescription": "Click the icon in the terminal to copy.",
"done": "done",
"orInstallGlobally": "or install globally",
"globalInstall": "Global install",
"startBackend": "Start backend",
"terminal": "Terminal",
"waitingForServer": "Waiting for server to start",
"pasteAndRun": "Paste and run in your terminal",
"pasteAndRunDescription": "Open a terminal at the project root, paste, and hit Enter.",
"listeningForServer": "Listening for server",
"willAutoConnect": "Will auto-connect when detected",
"autoConnects": "Auto-connects and opens the graph",
"autoConnectsDescription": "No refresh needed — the page detects the server automatically.",
"requires": "Requires",
"port": "Port 4747"
},
"analyzeFirst": {
"title": "Analyze your first repository",
"description": "Paste a GitHub URL and GitNexus will clone it, parse the code, and build a live knowledge graph — right in your browser.",
"footer": "Public repos only · Cloned locally by the server · No data leaves your machine"
},
"landing": {
"chooseRepository": "Choose a repository",
"description": "Select an indexed repository to explore, or analyze a new one.",
"indexed": "Indexed {{time}}",
"orAnalyzeNew": "or analyze new",
"footer": "Public & private repos · Cloned locally by the server · No data leaves your machine",
"time": {
"justNow": "just now",
"minutesAgo": "{{count}}m ago",
"hoursAgo": "{{count}}h ago",
"daysAgo": "{{count}}d ago"
}
},
"repoAnalyzer": {
"inputType": "Input type",
"githubUrl": "GitHub URL",
"gitlabUrl": "GitLab URL",
"localFolder": "Local Folder",
"starting": "Starting analysis...",
"analyzeRepository": "Analyze Repository",
"complete": "Analysis complete",
"loadingGraph": "Loading graph...",
"defaultRepoName": "repository",
"githubRepositoryUrl": "GitHub Repository URL",
"gitlabRepositoryUrl": "GitLab Repository URL",
"gitlabSupported": "Supports GitLab.com and self-hosted GitLab instances.",
"localFolderPath": "Local Folder Path",
"browseForFolder": "Browse for folder",
"hideBackground": "Hide (analysis continues in background)"
}
}
+93
View File
@@ -0,0 +1,93 @@
{
"title": "AI Settings",
"subtitle": "Configure your LLM provider",
"localServer": "Local Server",
"backendUrl": "Backend URL",
"connected": "Connected",
"notConnected": "Not connected",
"runServeHint": "Run `gitnexus serve` to connect the web UI to a local backend.",
"provider": "Provider",
"apiKey": "API Key",
"learnMore": "Learn more",
"model": "Model",
"searchModelPlaceholder": "Search or type model ID...",
"selectModelPlaceholder": "Select or type a model...",
"customModelHint": "Type a model ID or press Enter",
"customModelExample": "e.g. openai/gpt-4o",
"pressEnterCustom": "Press Enter to use as custom ID",
"baseUrl": "Base URL",
"optional": "optional",
"deploymentName": "Deployment Name",
"apiVersion": "API Version",
"checkConnection": "Check connection",
"privacyLabel": "Privacy:",
"privacyText": "Your API keys are stored locally in this browser.",
"providers": {
"openai": {
"description": "Use OpenAI models for chat and code reasoning.",
"apiKeyPlaceholder": "Enter your OpenAI API key",
"helperText": "Get your API key from",
"helperLinkLabel": "OpenAI Platform",
"modelPlaceholder": "e.g., gpt-4o, gpt-4-turbo, gpt-3.5-turbo",
"baseUrlPlaceholder": "https://api.openai.com/v1 (default)",
"baseUrlHint": "Leave empty to use the default OpenAI API. Set a custom URL for proxies or compatible APIs."
},
"gemini": {
"description": "Use Google Gemini models.",
"apiKeyPlaceholder": "Enter your Google AI API key",
"helperText": "Get your API key from",
"helperLinkLabel": "Google AI Studio",
"modelPlaceholder": "e.g., gemini-2.0-flash, gemini-1.5-pro"
},
"anthropic": {
"description": "Use Anthropic Claude models.",
"apiKeyPlaceholder": "Enter your Anthropic API key",
"helperText": "Get your API key from",
"helperLinkLabel": "Anthropic Console",
"modelPlaceholder": "e.g., claude-sonnet-4-20250514, claude-3-opus"
},
"azure": {
"apiKeyPlaceholder": "Enter your Azure OpenAI API key",
"deploymentNamePlaceholder": "e.g., gpt-4o-deployment"
},
"ollama": {
"quickStart": "📋 Quick Start:",
"installFrom": "Install Ollama from",
"thenRun": ", then run:",
"modelPlaceholder": "e.g., llama3.2, mistral, codellama"
},
"openrouter": {
"apiKeyPlaceholder": "Enter your OpenRouter API key",
"helperText": "Get your API key from",
"helperLinkLabel": "OpenRouter Keys"
},
"minimax": {
"apiKeyPlaceholder": "Enter your MiniMax API key",
"helperText": "Get your API key from",
"helperLinkLabel": "MiniMax Platform",
"modelPlaceholder": "e.g., MiniMax-M2.5, MiniMax-M2.5-highspeed",
"helperModel": "Available: MiniMax-M2.5 (default), MiniMax-M2.5-highspeed (faster)"
},
"glm": {
"apiKeyPlaceholder": "Enter your Z.AI API key"
}
},
"loadingModels": "Loading models...",
"noModelsMatch": "No models match \"{{searchTerm}}\"",
"moreModels": "+{{count}} more • Refine your search",
"apiKeySession": "API keys are stored in session storage and will be cleared when you close this tab.",
"startLocalServer": "start the local server",
"settingsSaved": "Settings saved",
"failedToSave": "Failed to save",
"saveSettings": "Save Settings",
"endpoint": "Endpoint",
"azurePortal": "Azure Portal",
"azureHint": "Configure your Azure OpenAI service in the",
"defaultPort": "Default port is",
"pullModel": "Pull a model with",
"browseModels": "Browse all models at",
"openRouterModels": "OpenRouter Models",
"zaiPlatform": "Z.AI Platform",
"glmCodingApi": "Coding API (default). Use https://api.z.ai/api/paas/v4 for the general API.",
"privacyFull": "Your API keys are stored only in your browser's session storage and are cleared when the tab closes. They're sent directly to the LLM provider when you chat. Your code never leaves your machine."
}
+36
View File
@@ -0,0 +1,36 @@
{
"tabs": {
"chat": "Nexus AI",
"processes": "流程"
},
"suggestions": {
"architecture": "解释项目架构",
"whatDoes": "这个项目是做什么的?",
"importantFiles": "显示最重要的文件",
"apiHandlers": "查找所有 API 处理器"
},
"empty": {
"title": "可以问我任何问题",
"description": "我可以帮你理解架构、查找函数或解释连接关系。"
},
"input": {
"placeholder": "询问这个代码库...",
"initializing": "正在初始化 AI Agent...",
"configureProvider": "配置 LLM 提供商以启用聊天。"
},
"actions": {
"closePanel": "关闭面板",
"scrollBottom": "滚动到底部",
"clearChat": "清空聊天",
"stopResponse": "停止响应"
},
"badges": {
"configureAI": "配置 AI",
"connecting": "连接中"
},
"roles": {
"you": "你",
"assistant": "Nexus AI"
},
"newBadge": "新"
}
@@ -0,0 +1,85 @@
{
"app": {
"name": "GitNexus",
"nexusAI": "Nexus AI"
},
"actions": {
"cancel": "取消",
"dismiss": "关闭",
"tryAgain": "重试",
"hide": "隐藏",
"retry": "重试",
"copy": "复制",
"copied": "已复制",
"close": "关闭",
"run": "运行",
"clear": "清空",
"remove": "移除",
"focusInGraph": "在图中聚焦",
"expand": "展开",
"collapse": "折叠"
},
"chat": {
"viewNodeInCodePanel": "在代码面板中查看 {{inner}}",
"openInCodePanel": "在代码面板中打开 • {{inner}}",
"waitForVectorIndex": "请稍候,正在创建向量索引。"
},
"counts": {
"files_one": "{{count}} 个文件",
"files_other": "{{count}} 个文件",
"nodes_one": "{{count}} 个节点",
"nodes_other": "{{count}} 个节点",
"edges_one": "{{count}} 条边",
"edges_other": "{{count}} 条边",
"symbols_one": "{{count}} 个符号",
"symbols_other": "{{count}} 个符号",
"flows_one": "{{count}} 条流程",
"flows_other": "{{count}} 条流程"
},
"progress": {
"connecting": "正在连接服务器...",
"connectingShort": "正在连接...",
"validatingServer": "正在验证服务器",
"validatingServerEllipsis": "正在验证服务器...",
"downloadingGraph": "正在下载图数据...",
"downloadedMb": "已下载 {{mb}} MB",
"downloadingWithPercent": "正在下载图数据... {{percent}}%",
"downloadingMb": "正在下载... {{mb}} MB",
"processing": "正在处理...",
"processingGraph": "正在处理图数据...",
"extractingFileContents": "正在提取文件内容",
"loadingGraph": "正在加载图数据...",
"starting": "正在启动...",
"executing": "正在执行...",
"truncated": "...(已截断)",
"ready": "就绪",
"switchingRepository": "正在切换仓库...",
"loadingRepository": "正在加载 {{repo}}",
"validating": "正在验证",
"failedSwitchRepository": "切换仓库失败",
"unknownError": "未知错误"
},
"analyzePhases": {
"queued": "已排队",
"cloning": "正在克隆仓库",
"pulling": "正在拉取最新代码",
"extracting": "正在扫描文件",
"structure": "正在构建结构",
"parsing": "正在解析代码",
"imports": "正在解析导入",
"calls": "正在追踪调用",
"heritage": "正在提取继承关系",
"communities": "正在检测社区",
"processes": "正在检测流程",
"complete": "流水线完成",
"lbug": "正在加载数据库",
"fts": "正在创建搜索索引",
"embeddings": "正在生成嵌入向量",
"done": "完成",
"retrying": "崩溃后正在重试"
},
"units": {
"elapsedSeconds": "{{seconds}} 秒",
"elapsedMinutesSeconds": "{{minutes}} 分 {{seconds}} 秒"
}
}
@@ -0,0 +1,19 @@
{
"unknown": "未知错误",
"connectFailed": "连接服务器失败",
"loadGraphFailed": "加载图数据失败",
"failedToConnect": "连接失败",
"analysisFailed": "分析失败,请检查服务器日志。",
"startAnalysisFailed": "启动分析失败",
"invalidGithubUrl": "请输入有效的 GitHub 仓库 URL。",
"missingFolderPath": "请输入文件夹路径。",
"backend": {
"reconnecting": "服务器连接已断开,正在重连…",
"network": "无法连接 GitNexus 服务器,请确认 `gitnexus serve` 正在运行。",
"timeout": "服务器响应超时,请稍后重试。",
"rateLimited": "请求过于频繁,请在 {{seconds}} 秒后重试。",
"notFound": "未找到请求的仓库或资源。",
"client": "请求失败:{{message}}",
"server": "服务器错误:{{message}}"
}
}
+174
View File
@@ -0,0 +1,174 @@
{
"statusBar": {
"sponsor": "赞助",
"sponsorHint": "需要买点 API 额度跑 SWE-bench 😅"
},
"loading": {
"filesProgress": "{{processed}} / {{total}} 个文件"
},
"toolCall": {
"status": {
"running": "运行中",
"completed": "已完成",
"error": "错误"
},
"tools": {
"search": "🔍 搜索代码",
"cypher": "🔗 Cypher 查询",
"grep": "🔎 模式搜索",
"read": "📄 读取文件",
"overview": "🗺️ 代码库概览",
"explore": "🔬 深入分析",
"impact": "💥 影响分析"
},
"query": "查询",
"input": "输入",
"result": "结果",
"searchPrefix": "搜索:\"{{query}}\""
},
"embedding": {
"generateTitle": "为语义搜索生成嵌入向量",
"enable": "启用语义搜索",
"loadingModel": "正在加载 AI 模型...",
"embeddingNodes": "正在嵌入 {{processed}}/{{total}} 个节点",
"creatingIndex": "正在创建向量索引...",
"readyTitle": "语义搜索已就绪!可在 AI 聊天中使用自然语言。",
"ready": "语义就绪",
"errorTitle": "嵌入失败,点击重试。",
"failedRetry": "失败 - 重试",
"fallback": {
"title": "WebGPU 拒绝了请求",
"subtitle": "你的浏览器不支持 GPU 加速",
"description": "无法用 WebGPU 创建嵌入向量,因此语义搜索(Graph RAG)不会那么智能,但图谱仍可正常使用。",
"options": "可选方案:",
"useCpu": "使用 CPU",
"useCpuDescriptionSmall": "可用,但会稍慢",
"useCpuDescriptionLarge": "可用,但会慢很多",
"estimated": "(约 {{minutes}} 分钟,{{count}} 个节点)",
"skipIt": "跳过",
"skipDescription": "图谱可用,但没有 AI 语义搜索",
"smallCodebase": "检测到小型代码库!CPU 应该没问题。",
"tip": "💡 提示:可尝试 Chrome 或 Edge 以获得 WebGPU 支持",
"skipEmbeddings": "跳过嵌入",
"useCpuRecommended": "使用 CPU(推荐)",
"useCpuSlow": "使用 CPU(较慢)"
}
},
"queryFab": {
"query": "查询",
"cypherQuery": "Cypher 查询",
"examples": "示例",
"run": "运行",
"noProject": "未加载项目,请先加载项目。",
"dbNotReady": "数据库尚未就绪,请等待加载完成。",
"executionFailed": "查询执行失败",
"exampleLabels": {
"functions": "所有函数",
"classes": "所有类",
"interfaces": "所有接口",
"calls": "函数调用",
"imports": "导入依赖"
},
"clear": "清除",
"rows": "行",
"highlighted": "已高亮",
"showingRows": "显示 {{count}} 行中的前 50 行"
},
"fileTree": {
"expandPanel": "展开面板",
"fileExplorer": "文件浏览器",
"filters": "过滤器",
"collapsePanel": "折叠面板",
"searchFiles": "搜索文件...",
"noFilesLoaded": "未加载文件",
"all": "全部",
"selectNodeDepth": "选择节点后才能应用深度过滤",
"explorer": "浏览器",
"nodeTypes": "节点类型",
"nodeTypesDesc": "切换图中节点类型的可见性",
"edgeTypes": "边类型",
"edgeTypesDesc": "切换关系类型的可见性",
"focusDepth": "聚焦深度",
"focusDepthDesc": "显示所选节点 N 跳内的节点",
"hops_one": "{{count}} 跳",
"hops_other": "{{count}} 跳",
"colorLegend": "颜色图例"
},
"codePanel": {
"expand": "展开代码面板",
"dragResize": "拖动调整大小",
"title": "代码检查器",
"clearCitations": "清除 AI 引用",
"clearSelection": "清除选择",
"loadingSource": "正在加载源码...",
"selectFile": "选择文件节点以预览其内容。",
"code": "代码",
"selected": "已选择",
"aiCitations": "AI 引用",
"references_one": "{{count}} 条引用",
"references_other": "{{count}} 条引用",
"lines_one": "{{count}} 行",
"lines_other": "{{count}} 行",
"codeNotAvailable": "内存中没有 {{path}} 的代码内容"
},
"canvas": {
"zoomIn": "放大",
"zoomOut": "缩小",
"fit": "适应屏幕",
"focusSelected": "聚焦所选节点",
"clearSelection": "清除选择",
"clear": "清除",
"stopLayout": "停止布局",
"runLayout": "重新运行布局",
"layoutOptimizing": "正在优化布局...",
"turnOffHighlights": "关闭全部高亮",
"turnOnHighlights": "开启 AI 高亮"
},
"processes": {
"unknownStep": "未知",
"allProcessesLabel_one": "全部流程(合并 {{count}} 个)",
"allProcessesLabel_other": "全部流程(合并 {{count}} 个)",
"emptyTitle": "未检测到流程",
"emptyDescription": "流程是从入口点追踪出的执行链路。加载代码库后即可查看检测到的流程。",
"filterPlaceholder": "过滤流程...",
"detected_one": "检测到 {{count}} 个流程",
"detected_other": "检测到 {{count}} 个流程",
"fullMap": "完整流程图",
"viewCombined_one": "查看 {{count}} 个流程的合并图",
"viewCombined_other": "查看 {{count}} 个流程的合并图",
"crossCommunity": "跨社区",
"intraCommunity": "社区内",
"steps_one": "{{count}} 步",
"steps_other": "{{count}} 步",
"clusters_one": "{{count}} 个聚类",
"clusters_other": "{{count}} 个聚类",
"highlightTitle": "在图谱中高亮",
"removeHighlightTitle": "移除图谱高亮",
"loading": "加载中...",
"viewing": "查看中",
"view": "查看"
},
"processFlow": {
"title": "流程:{{label}}",
"diagramTooLarge": "📊 图表过大",
"renderError": "⚠️ 渲染错误",
"tooComplex_one": "该图表包含 {{count}} 个步骤,复杂度过高无法渲染。请查看单个流程,而不是“全部流程”。",
"tooComplex_other": "该图表包含 {{count}} 个步骤,复杂度过高无法渲染。请查看单个流程,而不是“全部流程”。",
"unableToRender_one": "无法渲染图表。步骤数:{{count}}",
"unableToRender_other": "无法渲染图表。步骤数:{{count}}",
"zoomOutTitle": "缩小 (-)",
"zoomInTitle": "放大 (+)",
"resetTitle": "重置缩放和平移",
"resetView": "重置视图",
"toggleFocus": "切换聚焦",
"copyMermaid": "复制 Mermaid"
},
"diagram": {
"aiGenerated": "AI 生成图表",
"error": "图表错误",
"showSource": "显示源码",
"label": "图表",
"expandTitle": "展开",
"loading": "正在加载图表…"
}
}
@@ -0,0 +1,16 @@
{
"repositories": "仓库",
"active": "当前",
"reanalyzing": "正在重新分析...",
"reanalyzeRepo": "重新分析 {{repoName}}",
"deleteRepo": "删除 {{repoName}}",
"reanalyzingRepo": "正在重新分析 {{repoName}}:{{message}}",
"analyzeNew": "分析新仓库...",
"searchNodes": "搜索节点...",
"noNodesFound": "未找到“{{query}}”相关节点",
"starIfCool": "觉得不错就点星",
"aiSettings": "AI 设置",
"help": "帮助",
"language": "语言",
"selectLanguage": "选择语言"
}
+96
View File
@@ -0,0 +1,96 @@
{
"tabs": {
"overview": "概览",
"ai": "Nexus AI",
"shortcuts": "快捷键",
"status": "状态栏",
"graph": "图谱与节点",
"search": "搜索与过滤"
},
"shortcuts": {
"searchNodes": "搜索节点",
"deselectClose": "取消选择 / 关闭",
"columns": {
"action": "操作",
"mac": "Mac",
"windows": "Windows"
}
},
"nodeTypes": {
"function": "函数",
"functionDesc": "函数声明",
"file": "文件",
"fileDesc": "源码文件",
"class": "类",
"classDesc": "类声明",
"method": "方法",
"methodDesc": "类方法",
"interface": "接口",
"interfaceDesc": "TypeScript 接口",
"folder": "文件夹",
"folderDesc": "目录节点"
},
"status": {
"ready": "就绪",
"readyDesc": "图谱已完全加载并可交互",
"nodesCount": "节点数",
"nodesCountDesc": "图谱中的文件与符号总数",
"edgesCount": "边数",
"edgesCountDesc": "导入 / 依赖连接",
"aiIndexStatus": "AI 索引状态",
"aiIndexStatusDesc": "仓库已完成 AI 查询索引",
"semanticReadyBadge": "语义就绪",
"explained": "状态栏说明"
},
"tryAsking": "可以尝试问:",
"footer": "GitNexus — 图谱浏览器",
"title": "帮助与参考",
"footerLong": "GitNexus — 开源代码库图谱浏览器",
"docsGithub": "文档与 GitHub ↗",
"overview": {
"gettingStarted": "开始使用",
"whatIsTitle": "GitNexus 是什么?",
"whatIsDescription": "GitNexus 是代码库的交互式图谱浏览器。每个文件、函数和导入关系都会变成可探索、可查询、可视化导航的节点。",
"currentRepoTitle": "当前仓库",
"loadedCounts": "已加载:{{nodeCount}} 个节点 · {{edgeCount}} 条边",
"threeWaysTitle": "三种探索方式",
"wayInspect": "点击节点查看详情",
"waySearch": "按名称或类型搜索",
"wayAsk": "向 Nexus AI 提出自然语言问题",
"navigationTitle": "导航",
"navZoom": "滚动缩放",
"navPan": "点击并拖动进行平移",
"navFocus": "双击节点聚焦其子图"
},
"graph": {
"nodeColorLegend": "节点颜色图例",
"nodeLabel": "{{label}}节点",
"sizeDescription": "节点大小反映连接数量——越大的节点被越多文件依赖。边的方向表示从导入方 → 被导入方。",
"detailDescription": "点击任意节点可打开详情面板,查看导入、导出和反向依赖。"
},
"search": {
"title": "搜索与过滤",
"searchNodes": "搜索节点",
"searchDescription": "可按文件名、函数名或导入路径搜索,匹配的节点会在图谱中实时高亮。",
"filterPanel": "过滤面板",
"filterDescription": "使用左侧栏的过滤图标隔离特定节点类型、隐藏叶子节点,或以选中节点为根聚焦指定深度范围。",
"syntax": "搜索语法",
"hints": {
"nameFragment": "按名称片段匹配",
"pathPrefix": "按路径前缀匹配",
"nodeType": "按节点类型过滤"
}
},
"ai": {
"title": "Nexus AI",
"semanticReady": "✓ 语义索引就绪",
"description": "仓库已完成索引,可进行语义查询。Nexus AI 理解代码结构和关系,而不只是文件名。",
"questions": {
"dependencies": "“哪些文件依赖 auth 模块?”",
"circular": "“找出这个仓库中的循环依赖”",
"connected": "“哪些组件连接最密集?”",
"imports": "“显示所有导入 useEffect 的文件”"
},
"openPrompt": "点击右上角的 Nexus AI 按钮打开提问面板。"
}
}
@@ -0,0 +1,67 @@
{
"success": {
"title": "服务器已连接",
"description": "正在准备代码知识图谱..."
},
"loading": {
"largeRepoHint": "大型仓库可能需要一些时间"
},
"guide": {
"copyAria": "复制到剪贴板",
"copiedAria": "已复制!",
"startServer": "启动本地服务器",
"devDescription": "在另一个终端启动 Express 后端,即可启用完整图谱。",
"prodDescription": "只需一条命令,浏览器会自动连接。",
"copyCommand": "复制命令",
"copyCommandDescription": "点击终端右侧图标复制。",
"done": "完成",
"orInstallGlobally": "或全局安装",
"globalInstall": "全局安装",
"startBackend": "启动后端",
"terminal": "终端",
"waitingForServer": "等待服务器启动",
"pasteAndRun": "粘贴并在终端运行",
"pasteAndRunDescription": "在项目根目录打开终端,粘贴命令并回车。",
"listeningForServer": "正在监听服务器",
"willAutoConnect": "检测到后会自动连接",
"autoConnects": "自动连接并打开图谱",
"autoConnectsDescription": "无需刷新,页面会自动检测服务器。",
"requires": "需要",
"port": "端口 4747"
},
"analyzeFirst": {
"title": "分析你的第一个仓库",
"description": "粘贴 GitHub URL,GitNexus 会克隆仓库、解析代码,并直接在浏览器中构建实时知识图谱。",
"footer": "仅支持公开仓库 · 服务器本地克隆 · 数据不会离开你的机器"
},
"landing": {
"chooseRepository": "选择仓库",
"description": "选择一个已索引仓库开始探索,或分析一个新仓库。",
"indexed": "索引于 {{time}}",
"orAnalyzeNew": "或分析新仓库",
"footer": "支持公开与私有仓库 · 服务器本地克隆 · 数据不会离开你的机器",
"time": {
"justNow": "刚刚",
"minutesAgo": "{{count}} 分钟前",
"hoursAgo": "{{count}} 小时前",
"daysAgo": "{{count}} 天前"
}
},
"repoAnalyzer": {
"inputType": "输入类型",
"githubUrl": "GitHub URL",
"gitlabUrl": "GitLab URL",
"localFolder": "本地文件夹",
"starting": "正在启动分析...",
"analyzeRepository": "分析仓库",
"complete": "分析完成",
"loadingGraph": "正在加载图数据...",
"defaultRepoName": "仓库",
"githubRepositoryUrl": "GitHub 仓库 URL",
"gitlabRepositoryUrl": "GitLab 仓库 URL",
"gitlabSupported": "支持 GitLab.com 和自托管 GitLab 实例。",
"localFolderPath": "本地文件夹路径",
"browseForFolder": "浏览文件夹",
"hideBackground": "隐藏(分析继续在后台进行)"
}
}
@@ -0,0 +1,93 @@
{
"title": "AI 设置",
"subtitle": "配置你的 LLM 提供商",
"localServer": "本地服务器",
"backendUrl": "后端 URL",
"connected": "已连接",
"notConnected": "未连接",
"runServeHint": "运行 `gitnexus serve` 将 Web UI 连接到本地后端。",
"provider": "提供商",
"apiKey": "API Key",
"learnMore": "了解更多",
"model": "模型",
"searchModelPlaceholder": "搜索或输入模型 ID...",
"selectModelPlaceholder": "选择或输入模型...",
"customModelHint": "输入模型 ID 或按 Enter",
"customModelExample": "例如 openai/gpt-4o",
"pressEnterCustom": "按 Enter 使用自定义 ID",
"baseUrl": "Base URL",
"optional": "可选",
"deploymentName": "部署名称",
"apiVersion": "API 版本",
"checkConnection": "检查连接",
"privacyLabel": "隐私:",
"privacyText": "你的 API Key 仅保存在此浏览器本地。",
"providers": {
"openai": {
"description": "使用 OpenAI 模型进行聊天和代码推理。",
"apiKeyPlaceholder": "输入 OpenAI API Key",
"helperText": "从这里获取 API Key:",
"helperLinkLabel": "OpenAI Platform",
"modelPlaceholder": "例如:gpt-4o、gpt-4-turbo、gpt-3.5-turbo",
"baseUrlPlaceholder": "https://api.openai.com/v1(默认)",
"baseUrlHint": "留空则使用默认 OpenAI API。可为代理或兼容 API 设置自定义 URL。"
},
"gemini": {
"description": "使用 Google Gemini 模型。",
"apiKeyPlaceholder": "输入 Google AI API Key",
"helperText": "从这里获取 API Key:",
"helperLinkLabel": "Google AI Studio",
"modelPlaceholder": "例如:gemini-2.0-flash、gemini-1.5-pro"
},
"anthropic": {
"description": "使用 Anthropic Claude 模型。",
"apiKeyPlaceholder": "输入 Anthropic API Key",
"helperText": "从这里获取 API Key:",
"helperLinkLabel": "Anthropic Console",
"modelPlaceholder": "例如:claude-sonnet-4-20250514、claude-3-opus"
},
"azure": {
"apiKeyPlaceholder": "输入 Azure OpenAI API Key",
"deploymentNamePlaceholder": "例如:gpt-4o-deployment"
},
"ollama": {
"quickStart": "📋 快速开始:",
"installFrom": "从这里安装 Ollama:",
"thenRun": ",然后运行:",
"modelPlaceholder": "例如:llama3.2、mistral、codellama"
},
"openrouter": {
"apiKeyPlaceholder": "输入 OpenRouter API Key",
"helperText": "从这里获取 API Key:",
"helperLinkLabel": "OpenRouter Keys"
},
"minimax": {
"apiKeyPlaceholder": "输入 MiniMax API Key",
"helperText": "从这里获取 API Key:",
"helperLinkLabel": "MiniMax Platform",
"modelPlaceholder": "例如:MiniMax-M2.5、MiniMax-M2.5-highspeed",
"helperModel": "可用:MiniMax-M2.5(默认)、MiniMax-M2.5-highspeed(更快)"
},
"glm": {
"apiKeyPlaceholder": "输入 Z.AI API Key"
}
},
"loadingModels": "正在加载模型...",
"noModelsMatch": "没有匹配“{{searchTerm}}”的模型",
"moreModels": "+{{count}} 个更多模型 • 缩小搜索范围",
"apiKeySession": "API Key 保存在会话存储中,关闭此标签页后会被清除。",
"startLocalServer": "启动本地服务器",
"settingsSaved": "设置已保存",
"failedToSave": "保存失败",
"saveSettings": "保存设置",
"endpoint": "端点",
"azurePortal": "Azure Portal",
"azureHint": "在这里配置 Azure OpenAI 服务:",
"defaultPort": "默认端口为",
"pullModel": "使用以下命令拉取模型",
"browseModels": "在这里浏览全部模型:",
"openRouterModels": "OpenRouter Models",
"zaiPlatform": "Z.AI Platform",
"glmCodingApi": "Coding API(默认)。通用 API 请使用 https://api.z.ai/api/paas/v4。",
"privacyFull": "你的 API Key 只保存在此浏览器的会话存储中,关闭标签页后会清除。聊天时会直接发送给 LLM 提供商,你的代码不会离开本机。"
}
+1
View File
@@ -1,6 +1,7 @@
import React from 'react';
import ReactDOM from 'react-dom/client';
import App from './App';
import './i18n';
import './index.css';
ReactDOM.createRoot(document.getElementById('root') as HTMLElement).render(
+33
View File
@@ -1,8 +1,41 @@
import { beforeEach } from 'vitest';
import '@testing-library/jest-dom/vitest';
const I18N_LANGUAGE_STORAGE_KEY = 'gitnexus.lng';
function ensureStorage(name: 'localStorage' | 'sessionStorage') {
const current = globalThis[name];
if (
current &&
typeof current.getItem === 'function' &&
typeof current.removeItem === 'function'
) {
return;
}
const store = new Map<string, string>();
Object.defineProperty(globalThis, name, {
configurable: true,
value: {
getItem: (key: string) => store.get(key) ?? null,
setItem: (key: string, value: string) => store.set(key, String(value)),
removeItem: (key: string) => store.delete(key),
clear: () => store.clear(),
key: (index: number) => Array.from(store.keys())[index] ?? null,
get length() {
return store.size;
},
},
});
}
ensureStorage('localStorage');
ensureStorage('sessionStorage');
localStorage.removeItem(I18N_LANGUAGE_STORAGE_KEY);
// Reset storage between tests
beforeEach(() => {
sessionStorage.removeItem('gitnexus-llm-settings');
localStorage.removeItem('gitnexus-llm-settings'); // legacy key (migration)
localStorage.removeItem(I18N_LANGUAGE_STORAGE_KEY);
});
@@ -0,0 +1,396 @@
import { describe, expect, it } from 'vitest';
import {
buildLangChainMessages,
createChatModel,
serializeAgentHistoryMessages,
type AgentMessage,
} from '../../src/core/llm/agent';
import {
buildDeepSeekRequestMessages,
DeepSeekChatOpenAI,
DeepSeekChatOpenAICompletions,
} from '../../src/core/llm/deepseek-chat-model';
describe('buildLangChainMessages', () => {
it('reconstructs assistant tool-call turns for replay', () => {
const messages: AgentMessage[] = [
{ role: 'user', content: 'Check the weather' },
{
role: 'assistant',
content: 'Let me check that.',
reasoningContent: '',
toolCalls: [
{
id: 'call_weather',
name: 'get_weather',
args: { location: 'Hangzhou' },
type: 'tool_call',
},
],
},
{
role: 'tool',
content: 'Cloudy 7~13°C',
toolCallId: 'call_weather',
name: 'get_weather',
},
];
const langChainMessages = buildLangChainMessages(messages);
expect(langChainMessages).toHaveLength(3);
expect((langChainMessages[1] as any).additional_kwargs.reasoning_content).toBe('');
expect((langChainMessages[1] as any).tool_calls).toEqual([
{
id: 'call_weather',
name: 'get_weather',
args: { location: 'Hangzhou' },
type: 'tool_call',
},
]);
expect((langChainMessages[2] as any).tool_call_id).toBe('call_weather');
});
});
describe('serializeAgentHistoryMessages', () => {
it('captures assistant and tool messages from a completed turn', () => {
const serialized = serializeAgentHistoryMessages(
[
{ _getType: () => 'human', content: 'old prompt' },
{
_getType: () => 'ai',
content: 'Let me check that.',
additional_kwargs: { reasoning_content: 'Need weather tool.' },
tool_calls: [
{
id: 'call_weather',
name: 'get_weather',
args: { location: 'Hangzhou' },
type: 'tool_call',
},
],
},
{
_getType: () => 'tool',
content: 'Cloudy 7~13°C',
tool_call_id: 'call_weather',
name: 'get_weather',
},
{
_getType: () => 'ai',
content: 'Tomorrow will be cloudy.',
additional_kwargs: { reasoning_content: 'Result received.' },
},
],
1,
);
expect(serialized).toEqual([
{
role: 'assistant',
content: 'Let me check that.',
reasoningContent: 'Need weather tool.',
toolCalls: [
{
id: 'call_weather',
name: 'get_weather',
args: { location: 'Hangzhou' },
type: 'tool_call',
},
],
},
{
role: 'tool',
content: 'Cloudy 7~13°C',
toolCallId: 'call_weather',
name: 'get_weather',
},
{
role: 'assistant',
content: 'Tomorrow will be cloudy.',
},
]);
});
});
describe('buildDeepSeekRequestMessages', () => {
it('preserves reasoning_content on assistant tool-call messages', () => {
const requestMessages = buildDeepSeekRequestMessages(
buildLangChainMessages([
{ role: 'user', content: '如何支持Gitlab Repo' },
{
role: 'assistant',
content: '',
reasoningContent: 'I should inspect the repository support flow first.',
toolCalls: [
{
id: 'call_1',
name: 'search',
args: { query: 'Gitlab repo support' },
type: 'tool_call',
},
],
},
{
role: 'tool',
content: 'No matches',
toolCallId: 'call_1',
name: 'search',
},
]),
);
expect(requestMessages).toEqual([
{ role: 'user', content: '如何支持Gitlab Repo' },
{
role: 'assistant',
content: '',
reasoning_content: 'I should inspect the repository support flow first.',
tool_calls: [
{
id: 'call_1',
type: 'function',
function: {
name: 'search',
arguments: '{"query":"Gitlab repo support"}',
},
},
],
},
{
role: 'tool',
content: 'No matches',
name: 'search',
tool_call_id: 'call_1',
},
]);
});
});
it('drops reasoning_content from assistant messages without tool calls', () => {
const messages = buildLangChainMessages([
{ role: 'user', content: 'Hello' },
{
role: 'assistant',
content: 'Hi there',
reasoningContent: 'I should greet the user.',
},
]);
const requestMessages = buildDeepSeekRequestMessages(messages);
expect(requestMessages).toEqual([
{ role: 'user', content: 'Hello' },
{ role: 'assistant', content: 'Hi there' },
]);
});
it('drops reasoningContent from serialized assistant messages without tool calls', () => {
const serialized = serializeAgentHistoryMessages(
[
{
_getType: () => 'ai',
content: 'Simple answer.',
additional_kwargs: { reasoning_content: 'Thinking about it.' },
},
],
0,
);
expect(serialized).toEqual([
{
role: 'assistant',
content: 'Simple answer.',
},
]);
});
describe('createChatModel', () => {
it('keeps DeepSeek model subclasses on withConfig clones used for tool binding', () => {
const model = createChatModel({
provider: 'deepseek',
apiKey: 'test-key',
model: 'deepseek-v4-flash',
temperature: 0.1,
} as any) as any;
expect(model).toBeInstanceOf(DeepSeekChatOpenAI);
expect(model.completions).toBeInstanceOf(DeepSeekChatOpenAICompletions);
const clonedModel = model.withConfig({ tools: [] }) as any;
expect(clonedModel).toBeInstanceOf(DeepSeekChatOpenAI);
expect(clonedModel.completions).toBeInstanceOf(DeepSeekChatOpenAICompletions);
});
it('uses DeepSeek serialization on withConfig clones', async () => {
const model = createChatModel({
provider: 'deepseek',
apiKey: 'test-key',
model: 'deepseek-v4-flash',
temperature: 0.1,
} as any) as any;
const clonedModel = model.withConfig({ tools: [] }) as any;
clonedModel.completions.streaming = false;
let capturedRequest: any;
clonedModel.completions.client = {
chat: {
completions: {
create: async (request: any) => {
capturedRequest = request;
return {
choices: [
{
message: { role: 'assistant', content: 'ok' },
finish_reason: 'stop',
},
],
};
},
},
},
};
await clonedModel.completions._generate(
buildLangChainMessages([
{ role: 'user', content: 'Check the weather' },
{
role: 'assistant',
content: '',
reasoningContent: 'Need the weather tool.',
toolCalls: [
{
id: 'call_weather',
name: 'get_weather',
args: { location: 'Hangzhou' },
type: 'tool_call',
},
],
},
{
role: 'tool',
content: 'Cloudy 7~13°C',
toolCallId: 'call_weather',
name: 'get_weather',
},
]),
{ stream: false },
);
expect(capturedRequest.messages[1].reasoning_content).toBe('Need the weather tool.');
expect(capturedRequest.messages[1].tool_calls[0].function.arguments).toBe(
'{"location":"Hangzhou"}',
);
expect(capturedRequest.messages[2].tool_call_id).toBe('call_weather');
});
it('preserves reasoning_content through the streaming path used by DeepSeek tool calls', async () => {
const model = createChatModel({
provider: 'deepseek',
apiKey: 'test-key',
model: 'deepseek-v4-flash',
temperature: 0.1,
} as any) as any;
model.completions.streaming = true;
async function* mockStream() {
yield {
id: 'chatcmpl-1',
model: 'deepseek-v4-flash',
choices: [
{
index: 0,
delta: {
role: 'assistant',
reasoning_content: 'Need the weather tool.',
},
},
],
};
yield {
id: 'chatcmpl-1',
model: 'deepseek-v4-flash',
choices: [
{
index: 0,
delta: {
tool_calls: [
{
index: 0,
id: 'call_weather',
type: 'function',
function: {
name: 'get_weather',
arguments: '{"location":"Hangzhou"}',
},
},
],
},
finish_reason: 'tool_calls',
},
],
};
}
model.completions.client = {
chat: {
completions: {
create: async () => mockStream(),
},
},
};
let streamedMessage: any;
for await (const chunk of model.completions._streamResponseChunks(
buildLangChainMessages([{ role: 'user', content: 'Check the weather' }]),
{},
)) {
streamedMessage = streamedMessage ? streamedMessage.concat(chunk.message) : chunk.message;
}
expect(streamedMessage.additional_kwargs.reasoning_content).toBe('Need the weather tool.');
expect(streamedMessage.tool_calls).toEqual([
{
id: 'call_weather',
name: 'get_weather',
args: { location: 'Hangzhou' },
type: 'tool_call',
},
]);
expect(serializeAgentHistoryMessages([streamedMessage], 0)).toEqual([
{
role: 'assistant',
content: '',
reasoningContent: 'Need the weather tool.',
toolCalls: [
{
id: 'call_weather',
name: 'get_weather',
args: { location: 'Hangzhou' },
type: 'tool_call',
},
],
},
]);
});
it('rejects overlapping DeepSeek requests before reusing active messages', async () => {
const model = createChatModel({
provider: 'deepseek',
apiKey: 'test-key',
model: 'deepseek-v4-flash',
temperature: 0.1,
} as any) as any;
model.completions.activeMessages = buildLangChainMessages([{ role: 'user', content: 'busy' }]);
await expect(
model.completions._generate(
buildLangChainMessages([{ role: 'user', content: 'Check the weather' }]),
{ stream: false },
),
).rejects.toThrow('DeepSeekChatOpenAICompletions does not support overlapping requests');
});
});
@@ -0,0 +1,25 @@
import { beforeEach, describe, expect, it } from 'vitest';
import {
AUTO_START_EMBEDDINGS_STORAGE_KEY,
shouldAutoStartEmbeddings,
} from '../../src/hooks/useAppState';
describe('embedding auto-start gate', () => {
beforeEach(() => {
window.localStorage.clear();
});
it('defaults to disabled so connecting a repo remains read-only', () => {
expect(shouldAutoStartEmbeddings()).toBe(false);
});
it('allows opt-in through localStorage', () => {
window.localStorage.setItem(AUTO_START_EMBEDDINGS_STORAGE_KEY, 'true');
expect(shouldAutoStartEmbeddings()).toBe(true);
});
it('treats any non-true value as disabled', () => {
window.localStorage.setItem(AUTO_START_EMBEDDINGS_STORAGE_KEY, 'false');
expect(shouldAutoStartEmbeddings()).toBe(false);
});
});

Some files were not shown because too many files have changed in this diff Show More